Target following model training method and device, equipment and storage medium

By using the training loss of the box generation network and the heatmap generation network to train the target following model, the problem of low accuracy of the target following model in video is solved, and the clear perception and accurate following of the target subject is achieved.

CN117274296BActive Publication Date: 2025-11-07TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
CN202210678305.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2025-11-07
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

In existing technologies, target following models suffer from problems such as target ghosting, trailing, blurring, shadows, and background interference in videos, resulting in low following accuracy.

Method used

By acquiring image sample sequences of the second annotation results corresponding to the target subject, a second training loss is obtained using a box generation network, and combined with the first training loss of a heatmap generation network, the target following model is trained to improve the subject perception and following accuracy of the target.

Benefits of technology

It reduces the interference of ambiguous information such as target edges and background on the target following model, achieves clear perception and accurate following of the target, and improves the accuracy and clarity of the target following model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117274296B_ABST
    Figure CN117274296B_ABST
Patent Text Reader

Abstract

The application discloses a target following model training method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: acquiring a plurality of first image samples with target corresponding first annotation results and a plurality of second image samples with subject corresponding second annotation results; acquiring a first training loss corresponding to a heat map generation network based on the target first image sample and the remaining first image sample; acquiring a second training loss corresponding to a frame generation network based on the target second image sample and the remaining second image sample; and training a target following model for following the subject of the target based on the first training loss and the second training loss. The embodiments of the application can be applied to the scenes of artificial intelligence, intelligent transportation, and auxiliary driving, and the application can reduce the interference of blurred information such as the edge part and the background of the target, thereby improving the following accuracy and the following definiteness of the target.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of artificial intelligence, and particularly relate to a target following model training method and device, equipment and a storage medium. BACKGROUND

[0002] With the development of artificial intelligence technology, its research and application in the field of VOT (Video Object Tracking) are also increasing.

[0003] In the related art, a target following model that can be used to follow a target in a video is trained through a video frame sequence with the annotation information of the whole target. However, in actual application scenarios, the target or the background can be in a motion state, and the target can have problems such as residual image, trailing, blur, shadow, background interference, occlusion, etc. The target image often does not have clear and explicit target edges, thereby causing the target following model obtained by the related art to have low following accuracy. SUMMARY

[0004] Embodiments of the present application provide a target following model training method, device, equipment and storage medium, which can improve the following accuracy of the target following model. The technical solution can include the following content.

[0005] According to an aspect of an embodiment of the present application, a target following model training method is provided, the target following model including a heat map generation network and a frame generation network, and the method includes:

[0006] obtaining a first image sample sequence and a second image sample sequence; wherein the first image sample sequence includes a plurality of first image samples with a first annotation result corresponding to a target in time sequence, and the second image sample sequence includes a plurality of second image samples with a second annotation result corresponding to a main body of the target in time sequence;

[0007] obtaining a first training loss corresponding to the heat map generation network based on a target first image sample and a remaining first image sample in the first image sample sequence; wherein the first training loss is used to represent the difference between a first output result corresponding to the heat map generation network and the first annotation result, the target first image sample is used to indicate the following target of the target following model, and the first output result is used to represent the predicted distribution of the target;

[0008] obtain a second training loss corresponding to the bounding box generation network based on a target second image sample and a remaining second image sample in the second image sample sequence; the second training loss is used to represent a difference between a second output result corresponding to the bounding box generation network and the second annotation result, the target second image sample is used to indicate a subject of a following target of the target following model, and the second output result is used to represent a predicted distribution of the subject of the target;

[0009] train the target following model based on the first training loss and the second training loss to obtain a trained target following model; the trained target following model is used to follow the subject of the target.

[0010] According to an aspect of an embodiment of the present application, a device for training a target following model is provided, the target following model comprising a heat map generation network and a bounding box generation network, and the device comprises:

[0011] a sample sequence obtaining module, configured to obtain a first image sample sequence and a second image sample sequence; the first image sample sequence comprises a plurality of first image samples with a first annotation result corresponding to a target in time sequence, and the second image sample sequence comprises a plurality of second image samples with a second annotation result corresponding to a subject of the target in time sequence;

[0012] a first loss obtaining module, configured to obtain a first training loss corresponding to the heat map generation network based on a target first image sample and a remaining first image sample in the first image sample sequence; the first training loss is used to represent a difference between a first output result corresponding to the heat map generation network and the first annotation result, the target first image sample is used to indicate a following target of the target following model, and the first output result is used to represent a predicted distribution of the target;

[0013] a second loss obtaining module, configured to obtain a second training loss corresponding to the bounding box generation network based on a target second image sample and a remaining second image sample in the second image sample sequence; the second training loss is used to represent a difference between a second output result corresponding to the bounding box generation network and the second annotation result, the target second image sample is used to indicate a subject of a following target of the target following model, and the second output result is used to represent a predicted distribution of the subject of the target;

[0014] a following model training module, configured to train the target following model based on the first training loss and the second training loss to obtain a trained target following model; the trained target following model is used to follow the subject of the target.

[0015] According to an aspect of the embodiments of the present application, a computer device is provided, comprising a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the training method of the target following model.

[0016] According to an aspect of the embodiments of the present application, a computer readable storage medium is provided, the readable storage medium storing a computer program, the computer program being loaded and executed by a processor to implement the training method of the target following model.

[0017] According to an aspect of the embodiments of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the computer device execute the training method of the target following model.

[0018] The technical solutions provided by the embodiments of the present application have at least the following beneficial effects.

[0019] By using the plurality of second image samples corresponding to the second annotation result of the subject with the target sorted by time, the second training loss corresponding to the frame generation network is obtained, and the target following model is trained through the second training loss, which realizes training the target following model based on the clear and explicit subject of the target, reduces the interference of the fuzzy information (such as residual image, tailing, blur, occlusion, etc.) of the edge part, background, etc. of the target on the target following model, so that the target following model can accurately perceive and follow the subject of the target, thereby improving the following accuracy and following explicitness of the target. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0021] Figure 1 is a schematic diagram of the scheme implementation environment provided by an embodiment of the present application;

[0022] Figure 2 is a schematic diagram of the heat map generation network provided by an embodiment of the present application;

[0023] Figure 3 is a schematic diagram of the frame generation network provided by an embodiment of the present application;

[0024] Figure 4 is a flowchart of a training method of a target following model provided by an embodiment of the present application;

[0025] Figure 5 is a schematic diagram of a use method of a target following model provided by an embodiment of the present application;

[0026] Figure 6 is a block diagram of a training device of a target following model provided by an embodiment of the present application;

[0027] Figure 7 is a structural schematic diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0029] Artificial intelligence (AI) is to use digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use the knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.

[0030] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0031] Computer Vision (CV) is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes to identify and measure targets, and further perform image processing to make computer processing more suitable for human eye observation or image transmission to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, and map construction.

[0032] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and other disciplines. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. It is applied in various fields of artificial intelligence. Machine learning and deep learning generally include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.

[0033] The technical solutions provided by the embodiments of the present application relate to computer vision technology and machine learning technology of artificial intelligence. The computer vision technology is used to extract features of image samples, obtain a heat map corresponding to a target, and obtain a predicted bounding box corresponding to a subject of the target based on the heat map corresponding to the target. Then, the machine learning technology is used to train a target following model (such as a heat map generation network and a frame generation network in the target following model) based on the predicted bounding box corresponding to the subject of the target, so as to obtain a trained target following model.

[0034] The execution subject of each step of the method provided by the embodiments of the present application can be a computer device, which refers to an electronic device with data calculation, processing, and storage capabilities. The computer device can be a terminal such as a PC (Personal Computer), a tablet computer, a smart phone, a wearable device, a smart robot, a vehicle-mounted terminal, or a server. The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0035] The technical solutions provided in the embodiments of the present application are applicable to any scene requiring a target following task, such as a video target following scene, an intelligent traffic scene, an auxiliary driving scene, an image analysis scene, a television live broadcast or relay scene, a medical scene, and the like. The technical solutions provided in the embodiments of the present application can improve the following accuracy and explicitness of a target.

[0036] The model structure and training method of the target following model provided in the embodiments of the present application will be described in detail below.

[0037] Reference is made to Figure 1 , which shows a schematic diagram of a scheme implementation environment provided in an embodiment of the present application. The scheme implementation environment can include a model training device 10 and a model using device 20.

[0038] The model training device 10 can be an electronic device such as a PC, a computer, a tablet computer, a server, a smart robot, a vehicle-mounted terminal, or other electronic devices with strong computing capability. The model training device 10 is configured to train a target following model 30.

[0039] In the embodiments of the present application, the target following model 30 is a neural network model that can be used for target following (such as video target following), target trajectory acquisition, and the like. Exemplarily, in a target following scene, the target following model 30 can be used to follow the subject of a target object in a video. In a target trajectory acquisition scene, the target following model 30 can be used to generate a moving trajectory of a target object. The embodiments of the present application do not limit the tasks to which the target following model 30 can be applied. The target following model 30 in the embodiments of the present application can support the following tasks of the subject of one or more targets.

[0040] Optionally, the model training device 10 can train the target following model 30 in a machine learning manner, so that the target following model 30 has good following performance.

[0041] The trained target following model 30 described above can be deployed in the model using device 20 for use, to provide target following, target trajectory acquisition, and the like. The model using device 20 can be a terminal device such as a mobile phone, a computer, a smart television, a multimedia playing device, a wearable device, a medical device, a vehicle-mounted terminal, or a server, and the embodiments of the present application do not limit this.

[0042] In some embodiments, as shown in Figure 1 , the target following model 30 can include a heat map generation network 310 and a frame generation network 320.

[0043] The heat map generation network 310 is configured to process the image sample (e.g., the first image sample or the second image sample) to obtain a first output result corresponding to each image sample, which is used to represent a predicted distribution (e.g., a predicted position and a predicted shape) of the target, for example, a key pixel point corresponding to the target in the image sample. For example, the first output result can be a heat map corresponding to the target, which can be a two-dimensional array consistent with the size (length and width) of the input image sample. Each value in the heat map is used to represent the possibility of the corresponding pixel belonging to the target. According to the heat map corresponding to the target, the key pixel point corresponding to the target can be determined.

[0044] In one example, referring to Figure 2 , the heat map generation network 310 can include a feature extraction layer 311 and a cross-correlation layer 312.

[0045] The feature extraction layer 311 is configured to perform feature extraction on the image sample to obtain a feature map corresponding to the image sample. The feature extraction layer 311 is a twin network structure, and each branch is represented by f θ , and the two f θ are ResNet (Residual Network) with shared weights. Each branch of the feature extraction layer 311 can be composed of the first four stages of ResNet (e.g., ResNet-50).

[0046] For example, for the target image sample 313 and the remaining image sample 314, the feature extraction layer 311 is used to obtain a feature map 315 corresponding to the target image sample 313 and a feature map 316 corresponding to the remaining image sample 314, respectively. The target image sample 313 is used to indicate the target followed by the target following model 30. The target image sample 313 can be the first image sample in the image sample sequence, and the remaining image sample 314 can be any image sample in the image sample sequence except the target image sample 313. There is one or more same targets in the target image sample 313 and the remaining image sample 314, which can be a vehicle, an animal, a person, another target object, etc. The image sample sequence includes a plurality of image samples sorted by time. The image sample sequence can refer to a sequence of all or part of video frames corresponding to a video file or a video stream. The image sample sequence can also include multiple images with the same target, which is not limited in the embodiments of the present application.

[0047] The cross-correlation layer 312 is configured to perform deep cross-correlation (denoted by *d) on the feature map 315 corresponding to the target image sample 313 and the feature map 316 corresponding to the remaining image sample 314 to obtain a heat map 317 corresponding to the target in the remaining image sample 314.

[0048] In one feasible example, the heat map generation network 310 is a trained neural network, and the target following model 30 can be obtained by directly splicing the trained heat map generation network 310 and the box generation network 320, so that the heat map generation network 310 does not need to be trained again. Alternatively, the heat map generation network 310 can be a corresponding neural network for generating a heat map in the SiamMask algorithm, the D3S algorithm (A Discriminative Single Shot Segmentation Tracker, a target following algorithm), or the like.

[0049] The box generation network 320 is configured to process the first output result of the heat map generation network 310 to obtain a second output result corresponding to the image sample, and the second output result is used to represent a predicted distribution (such as a predicted position and a predicted shape) of the target subject, for example, the second output result can be a predicted bounding box surrounding the target subject.

[0050] In one example, referring to Figure 3 The box generation network 320 is a lightweight neural network. The box generation network 320 can include three convolutional layers, three fully connected layers, and a de-overfitting layer (i.e., a Dropout layer). For example, each convolutional layer can be a convolutional network with a 3*3 convolutional kernel, a stride of 1, and a padding of 1. For example, in order to facilitate the structural design and training of the box generation network 320, the heat map corresponding to the target can be first pooled into a feature map with a size of 8*8, the first convolutional layer can convolve the 8*8 feature map into an 8*8*64 feature map, the second convolutional layer can still convolve the 8*8*64 feature map into an 8*8*64 feature map, the third convolutional layer can convolve the 8*8*64 feature map into an 8*8*128 feature map, the first fully connected layer can obtain a 128-dimensional vector based on the 8*8*128 feature map, the de-overfitting layer can de-overfit the 128-dimensional vector to obtain a 64-dimensional vector, and the last two fully connected layers can process the 64-dimensional vector to obtain a 5-dimensional vector, which is used to represent a predicted horizontal position, a predicted vertical position, a predicted width, a predicted height, and a predicted angle of the predicted bounding box. The predicted angle of the predicted bounding box is a predicted angle between the predicted bounding box and the horizontal coordinate axis.

[0051] Designing the box generation network 320 as a lightweight neural network not only enables the prediction of the bounding box based on the heat map, but also simplifies the network structure of the box generation network 320 and improves the generation efficiency of the predicted bounding box.

[0052] The model architecture of the target following model is not limited in the embodiments of the present application, and can be adaptively adjusted according to actual use requirements, such as increasing a convolutional layer, reducing a convolutional layer, and the like. Understandably, the model capable of realizing the following function of the subject of the target should be within the protection scope of the embodiments of the present application.

[0053] Exemplarily, a first loss corresponding to the target following model 30 can be obtained by the heat map generation network 310 based on the first image sample sequence with the first annotation result of the target, and a second loss corresponding to the target following model 30 can be obtained by the heat map generation network 310 and the frame generation network 320 based on the second image sample sequence with the second annotation result of the subject of the target, and the target following model 30 can be trained based on the first loss and the second loss, so as to obtain the target following model 30 for following the subject of the target.

[0054] The model architecture of the target following model is introduced above, and the training method of the target following model will be described in detail below.

[0055] Please refer to Figure 4 which shows a flowchart of the training method of the target following model provided by an embodiment of the present application. The execution subject of each step of the method can be the model training device introduced above. The method can include the following steps (401-404).

[0056] Step 401, a first image sample sequence and a second image sample sequence are obtained; wherein the first image sample sequence includes a plurality of first image samples with first annotation results of a target in time sequence, and the second image sample sequence includes a plurality of second image samples with second annotation results of a subject of the target in time sequence.

[0057] The first image sample sequence can refer to an image sample sequence with first annotation results of a target, and the image sample sequence can correspond to all or part of a certain video file, a certain video stream, and the like. The image sample sequence can also include a plurality of images with the same target, which is not limited in the embodiments of the present application. For example, a part of a video clip is extracted from a certain video file, a video frame sequence corresponding to the part of the video clip is obtained, and the target in each video frame in the video frame sequence is annotated (such as a bounding box annotation), so as to obtain the first image sample sequence. The first annotation result is used to indicate the position and size of the target in the first image sample. For example, in the case that the first annotation result is a bounding box of the target, the bounding box can be used to indicate the real position and real size of the target in the first image sample.

[0058] The second image sample sequence can refer to an image sample sequence with a second annotation result corresponding to the subject of the target. The image sample sequence corresponding to the second image sample sequence can be the same as the image sample sequence corresponding to the first image sample sequence. The second annotation result is used to indicate the position and size of the subject of the target in the second image sample. For example, in the case of the second annotation result being a bounding box of the subject of the target, the bounding box can be used to indicate the real position and real size of the subject of the target in the second image sample.

[0059] In the embodiments of the present application, the target can refer to a vehicle, an animal, a person, other target objects, etc. The subject of the target can refer to a subject part or an identifiable part of the target. For example, the subject of the vehicle can refer to the cockpit of the vehicle, and the subject of the animal can refer to the torso of the animal.

[0060] In one example, the acquisition process of the first image sample sequence and the second image sample sequence can be as follows:

[0061] 1. From the data set with mask annotation results, select an image sample sequence with mask annotation results of the target.

[0062] Optionally, based on the mask annotation result of the target, the first annotation result (such as a rectangular bounding box) corresponding to the target can be constructed, thereby obtaining the first image sample sequence. For example, from the COCO data set, an image sample sequence with mask annotation results of the target is obtained, and then based on the image sample sequence with mask annotation results of the target, a first image sample sequence corresponding to the target is obtained.

[0063] 2. For each image sample in the image sample sequence, a real bounding box enclosing the subject of the target is constructed.

[0064] Optionally, based on the mask annotation result corresponding to the subject of the target, a real bounding box (such as a rectangular bounding box) enclosing the mask pixels corresponding to the subject of the target can be constructed.

[0065] 3. The real bounding box corresponding to each image sample is adjusted to minimize the number of subject mask pixels outside the real bounding box and minimize the number of non-subject mask pixels inside the real bounding box, thereby obtaining multiple second image samples with second annotation results corresponding to the subject of the target.

[0066] Optionally, the process can be represented as follows:

[0067]

[0068] Wherein, v represents a real bounding box corresponding to the subject of the target (such as the 5-dimensional vector described above), B represents a pixel region corresponding to v, x is a pixel point; M is a function, which returns 1 if x belongs to the mask annotation result corresponding to the subject of the target, otherwise returns 0; T is another function, which returns 1 if the input is 1, otherwise returns 0; a is a weight coefficient.

[0069] The purpose of the formula is to simultaneously perform two tasks to find the optimal real bounding box for the mask annotation result corresponding to the subject of the target: minimize the number of subject mask pixels outside the real bounding box and minimize the number of non-subject mask pixels inside the real bounding box, and adjust the weight of the two tasks using the hyperparameter a.

[0070] The task of minimizing the number of subject mask pixels outside the real bounding box can drive the real bounding box to shrink, exclude non-subject mask pixels and discard the edge part of the target, and at the same time induce the real bounding box to converge to the subject part of the target with high subject mask pixel density. The task of minimizing the number of non-subject mask pixels inside the real bounding box can constrain the real bounding box in the opposite direction to avoid excessive shrinking of the real bounding box, which can cause too many key pixels corresponding to the subject to be excluded outside the real bounding box. By adjusting the hyperparameter a, a more reasonable real bounding box corresponding to the subject of the target can be obtained.

[0071] By automatically annotating the second annotation result for the second image sample, the problems of low efficiency and non-uniform annotation standards in manual annotation can be avoided, thereby improving the annotation efficiency and accuracy of the second annotation result. Alternatively, if necessary, manually annotating the second annotation result for the second image sample is also a feasible way.

[0072] In step 402, a first training loss corresponding to the heat map generation network is obtained based on a target first image sample and remaining first image samples in the first image sample sequence; wherein the first training loss is used to represent the difference between the first output result corresponding to the heat map generation network and the first annotation result, the target first image sample is used to indicate the target followed by the target following model, and the first output result is used to represent the predicted distribution of the target.

[0073] The target first image sample can refer to the first image sample in the first image sample sequence. For example, taking the video frame sequence corresponding to the first image sample sequence as an example, the target first image sample can refer to the first video frame in the video frame sequence. Alternatively, the target in the target first image sample can be taken as the target to be followed, and the target in each remaining first image sample can be followed. The remaining first image sample can refer to any first image sample in the first image sample sequence except the target first image sample. The heat map generation network and the first output result are the same as described in the above embodiments, which will not be repeated here.

[0074] In one embodiment, taking the corresponding heat map of the first output result as an example, the first training loss can be obtained as follows:

[0075] 1. For each remaining first image sample, the first output result corresponding to the remaining first image sample is obtained based on the target first image sample and the remaining first image sample through the heat map generation network.

[0076] For example, the feature map corresponding to the target first image sample and the feature corresponding to the remaining first image sample are obtained through the heat map generation network, respectively, and then the feature map corresponding to the target first image sample and the feature corresponding to the remaining first image sample are deeply cross-correlated to obtain the heat map of the target correspondence in the remaining first image sample.

[0077] 2. The first training loss corresponding to the remaining first image sample is obtained based on the difference between the first output result corresponding to the remaining first image sample and the first annotation result.

[0078] In some examples, the mask loss can be obtained based on the difference between the mask data corresponding to the first output result and the real mask data corresponding to the first annotation result, and the confidence loss can be obtained based on the confidence of the mask data containing the target corresponding to the first output result using the Logistic loss algorithm, and the sum of the mask loss and the confidence loss is determined as the first training loss.

[0079] In other examples, the mask loss can be obtained based on the difference between the mask data corresponding to the first output result and the real mask data corresponding to the first annotation result, the box loss can be obtained based on the difference between the bounding box corresponding to the first output result and the real bounding box corresponding to the first annotation result using the Smooth L1 algorithm, and the confidence loss can be obtained based on the confidence of the bounding box being foreground or background corresponding to the first output result using the cross-entropy loss algorithm, and the sum of the mask loss, the box loss and the confidence loss is determined as the first training loss.

[0080] 3. The first training loss corresponding to the heat map generation network is obtained based on the first training loss corresponding to each remaining first image sample.

[0081] For the first image sample sequence, a sum value of respective first training losses corresponding to respective residual first image samples can be determined as a total training loss corresponding to the heat map generation network. However, in an actual training process, the network parameters of the heat map generation network are adjusted once based on the first training loss corresponding to each residual first image sample, so as to iteratively complete the training of the heat map generation network. Optionally, the first training loss corresponding to each residual first image sample can be obtained by the target following model trained by the first training loss corresponding to the previous residual first image sample.

[0082] At step 403, a second training loss corresponding to the box generation network is obtained based on a target second image sample and a residual second image sample in the second image sample sequence; the second training loss is used to represent a difference between a second output result corresponding to the box generation network and a second annotation result, and the target second image sample is used to indicate a subject of a following target of the target following model, and the second output result is used to represent a predicted distribution of the subject of the target.

[0083] The target second image sample can refer to a first image sample in the second image sample sequence. For example, taking the video frame sequence corresponding to the second image sample sequence as an example, the target second image sample can refer to a first video frame in the video frame sequence. Optionally, the subject of the target in the target second image sample can be taken as the following target, and the subject of the target in each residual first image sample can be followed. The residual second image sample can refer to any second image sample in the second image sample sequence except the target second image sample. The box generation network and the second output result are the same as described in the above embodiments, and will not be described here.

[0084] In one example, the process of obtaining the second training loss can be as follows:

[0085] 1. For each residual second image sample, a first output result corresponding to the residual second image sample is obtained by the heat map generation network based on the target second image sample and the residual second image sample.

[0086] For example, the feature map corresponding to the target second image sample and the feature corresponding to the residual second image sample are obtained by the heat map generation network respectively, and then the feature map corresponding to the target second image sample and the feature corresponding to the residual second image sample are subjected to deep cross-correlation to obtain a heat map corresponding to the target in the residual second image sample.

[0087] 2. A second output result corresponding to the residual second image sample is obtained by the box generation network based on the first output result corresponding to the residual second image sample.

[0088] For example, a series of processes such as pooling, convolution, de- overfitting, full connection calculation, etc. are performed on the heat map corresponding to the target in the remaining second image samples to obtain a predicted bounding box corresponding to the subject of the target, that is, a second output result. In the embodiments of the present application, the predicted bounding box corresponding to the subject of the target is obtained directly based on the heat map corresponding to the target. Since the heat map has clear abstract information, it is not necessary to analyze the texture, color and other features, thereby improving the running efficiency of the frame generation network, and further improving the running efficiency of the target following model. In addition, since no additional network is required for processing the texture, color and other features, the network structure of the frame generation network is simplified, and the model architecture of the target following model is further simplified.

[0089] Optionally, the predicted bounding box can be rotatable, for example, the predicted bounding box can be a rotatable rectangular box. In the embodiments of the present application, the rotatable rectangular box refers to a rectangular box surrounding the subject of the target. The angle of the rectangular box surrounding the subject of the target can be different at different time instants in a video, that is, the rectangular box can be rotatable. Since the angle of the rotatable rectangular box at different time instants can be different, the short side or the long side of the rotatable rectangular box does not need to be always parallel to the horizontal coordinate axis. Therefore, the angle of the rotatable rectangular box can be adaptively adjusted according to the change of the rotation state of the subject of the target, thereby having the ability to adaptively rotate following the change of the subject of the target. Moreover, the rotatable rectangular box can be scaled in a reasonable direction according to the corresponding angle, so as to reasonably surround the key pixel points corresponding to the subject of the target, without being limited by a fixed angle and being able to only scale the rectangular box in the horizontal or vertical direction. The introduction of too many non-key pixel points is avoided, so that the rotatable rectangular box can more accurately fit the subject of the target, thereby further improving the target following accuracy.

[0090] In one example, the minimum circumscribed rectangle of the key pixel points corresponding to the subject of the target can be determined as the rotatable rectangular box corresponding to the subject of the target.

[0091] In another example, the generation of the rotatable rectangular box is performed with the target of minimizing the number of key pixel points outside the rotatable rectangular box, minimizing the number of non-key pixel points inside the rotatable rectangular box, the proportion of the number of key pixel points in the rotatable rectangular box being less than a first key pixel point threshold, and the proportion of the number of non-key pixel points in the rotatable rectangular box being less than a second key pixel point threshold, so as to obtain the rotatable rectangular box corresponding to the subject of the target. The first key pixel point threshold and the second key pixel point threshold can be set and adjusted according to actual use requirements.

[0092] Optionally, the predicted horizontal position and the predicted vertical position can be defined as the coordinates of the center point of the rotatable rectangular frame in pixels, the predicted width and the predicted height can be defined as the width and the length of the rotatable rectangular frame in a mathematical sense, and the predicted angle can be defined as the angle between the long side of the rotatable rectangular frame and the x-coordinate axis.

[0093] 3. Based on the difference between the second output result corresponding to the remaining second image sample and the second annotation result, a second training loss corresponding to the remaining second image sample is obtained.

[0094] Optionally, the Smooth L1 algorithm can be used to obtain the second training loss corresponding to the remaining second image sample based on the difference between the second output result corresponding to the remaining second image sample and the second annotation result. Optionally, the second training loss corresponding to each remaining second image sample can be obtained by training the target following model with the second training loss corresponding to the previous remaining second image sample.

[0095] In one example, taking the predicted bounding box of the subject targeted by the second output result as an example, the process of obtaining the second training loss can be as follows:

[0096] Based on the difference between the predicted horizontal position corresponding to the predicted bounding box of the remaining second image sample and the real horizontal position corresponding to the second annotation result of the remaining second image sample, a horizontal position loss corresponding to the remaining second image sample is obtained.

[0097] The horizontal position loss can be represented as follows:

[0098] where x p is the predicted horizontal position, x r is the real horizontal position, w r is the real width.

[0099] Based on the difference between the predicted vertical position corresponding to the predicted bounding box of the remaining second image sample and the real vertical position corresponding to the second annotation result of the remaining second image sample, a vertical position loss corresponding to the remaining second image sample is obtained.

[0100] The vertical position loss can be represented as follows:

[0101] where y p is the predicted vertical position, y r is the real vertical position, h r is the real height.

[0102] The width loss of the remaining second image sample is obtained based on a difference between a predicted width corresponding to the predicted bounding box corresponding to the remaining second image sample and a real width corresponding to the second annotation result corresponding to the remaining second image sample.

[0103] The width loss can be represented as follows:

[0104] where w p is the predicted width.

[0105] The height loss of the remaining second image sample is obtained based on a difference between a predicted height corresponding to the predicted bounding box corresponding to the remaining second image sample and a real height corresponding to the second annotation result corresponding to the remaining second image sample.

[0106] The height loss can be represented as follows:

[0107] where h p is the predicted height.

[0108] The angle loss of the remaining second image sample is obtained based on a difference between a predicted angle corresponding to the predicted bounding box corresponding to the remaining second image sample and a real angle corresponding to the second annotation result corresponding to the remaining second image sample.

[0109] The angle loss can be represented as follows:

[0110] d θ = θ p - θ r ; where θ p is the predicted height and θ r is the real height.

[0111] The second training loss corresponding to the remaining second image sample is obtained by summing the horizontal position loss, the vertical position loss, the width loss, the height loss, and the angle loss.

[0112] 4. The second training loss corresponding to the bounding box generation network is obtained based on the second training loss corresponding to each of the remaining second image samples.

[0113] For the sequence of second image samples, the sum of the second training loss corresponding to each of the remaining second image samples can be determined as the total training loss corresponding to the heat map generation network. However, in the actual training process, the network parameters of the bounding box generation network are adjusted once based on the second training loss corresponding to each of the remaining second image samples, so that the training of the bounding box generation network is iteratively completed.

[0114] At step 404, the target following model is trained based on the first training loss and the second training loss, to obtain a trained target following model; wherein the trained target following model is used to follow the subject of the target.

[0115] Since the purpose of the heat map generation network (e.g., for obtaining a heat map corresponding to a target) and the purpose of the box generation network (e.g., for obtaining a predicted bounding box corresponding to the subject of a target) are not consistent, the two networks are trained separately in the embodiments of the present application, and after the network parameters of the two networks change stably, the network parameters of the two networks are fine-tuned at the same time, so that the data processing mode of the heat map generation network and the data processing mode of the box generation network are mutually compatible, which helps to output a target following model with stable and excellent performance.

[0116] In one example, the training process of the target following model can be as follows:

[0117] 1. Adjust the network parameters of the heat map generation network based on the first training loss to obtain a first-stage target following model.

[0118] The network parameters of the heat map generation network in the first-stage target following model are preliminarily adjusted.

[0119] Optionally, the network parameters of the heat map generation network can be iteratively adjusted based on the first training loss corresponding to each remaining first image sample, to obtain the first-stage target following model.

[0120] 2. Adjust the network parameters of the box generation network of the first-stage target following model based on the second training loss and with a first learning rate, to obtain a second-stage target following model.

[0121] The network parameters of the heat map generation network and the network parameters of the box generation network in the second-stage target following model are preliminarily adjusted.

[0122] During the adjustment of the network parameters of the box generation network of the first-stage target following model, the heat map generation network is frozen, so that its network parameters are not adjusted in this stage of training process, that is, only the network parameters of the box generation network are adjusted. Optionally, the network parameters of the box generation network can be iteratively adjusted based on the second training loss corresponding to each remaining second image sample, to obtain the second-stage target following model.

[0123] The first learning rate can be adaptively set and adjusted according to an empirical value, which is not limited in the embodiments of the present application.

[0124] 3. Adjust the network parameters of the second-stage target following model based on the second training loss and with a second learning rate, to obtain a trained target following model.

[0125] Optionally, the second training loss in the embodiments of the present application can be obtained by the target following model of the first stage, or can be obtained by the target following model of the second stage, or can be obtained by the untrained target following model, and the embodiments of the present application do not limit this.

[0126] The network parameters of the target following model (including the network parameters of the heat map generation network and the network parameters of the frame generation network) can be iteratively fine-tuned based on the second training loss corresponding to each remaining second image sample, so as to obtain the trained target following model.

[0127] The second learning rate is less than the first learning rate. For example, the second learning rate can be set to 0.001 times of the first learning rate.

[0128] In one example, the trained target following model is used to continuously mark the subject of the target in each input image of the input image sequence. Illustratively, the predicted bounding box of the subject of the target in each input image corresponding to the input image sequence is obtained by the trained target following model, and the subject of the target is continuously highlighted based on the predicted bounding box.

[0129] For example, taking the input image sequence as a video frame sequence as an example, the target video frame (such as the first frame video frame marked with the target to be followed) and the video frame to be processed (such as each video frame in the video frame sequence except the target video frame) are processed by the trained target following model, so that the predicted bounding box of the subject of the target in each video frame to be processed is output frame by frame, and the subject of the target in each video frame to be processed is continuously marked and displayed based on the predicted bounding box of the subject of the target in each video frame to be processed.

[0130] In another example, the trained target following model is used to obtain the moving track of the target in the input image sequence. Illustratively, the predicted bounding box of the subject of the target in each input image corresponding to the input image sequence is obtained by the trained target following model, and the moving track of the target in the input image sequence is obtained based on the position of the predicted bounding box.

[0131] For example, taking the input image sequence as a video frame sequence as an example, the target video frame (such as the first frame video frame marked with the target to be followed) and the video frame to be processed (such as each video frame in the video frame sequence except the target video frame) are processed by the trained target following model, so that the predicted bounding box of the subject of the target in each video frame to be processed is output frame by frame, and the moving track of the target in the input image sequence is obtained according to the position of the predicted bounding box of the subject of the target in each video frame to be processed.

[0132] Optionally, the execution subject in the use process of the trained target following model can be the same as the execution subject in the training process of the target following model, such as the model training device 10 described above, and the execution subject in the use process of the trained target following model can also be an additional execution subject, such as the model use device 20 described above, and the embodiments of the present application do not limit this.

[0133] To sum up, the technical scheme provided by the embodiments of the present application, by the plurality of second image samples corresponding to the second annotation result of the target subject sorted by time, the second training loss corresponding to the frame generation network is obtained, and the target following model is trained through the second training loss, which realizes training the target following model based on the clear and explicit target subject, reduces the interference of the fuzzy information (such as residual image, tailing, blur, occlusion, etc.) corresponding to the target on the target following model, so that the target following model can accurately perceive and follow the target subject, thereby improving the following accuracy and explicitness of the target.

[0134] In addition, by respectively training the heat map generation network and the frame generation network first, and then fine-tuning the network parameters of the two networks when the network parameters of the heat map generation network and the frame generation network change smoothly, the data processing mode of the heat map generation network and the data processing mode of the frame generation network are mutually compatible, thereby improving the following stability and accuracy of the target following model.

[0135] In addition, by setting the frame generation network as a lightweight neural network, not only can the target subject corresponding to the predicted bounding box be generated based on the heat map, but also the network structure of the frame generation network can be simplified and the generation efficiency of the predicted bounding box can be improved.

[0136] In one exemplary embodiment, the training method of the target following model is introduced by taking the target following model for following the target object in the video as an example, which can include the following contents.

[0137] The video sample file is obtained, as well as the first video frame sequence corresponding to the first annotation result of the target object in the video sample file, and the second video frame sequence corresponding to the second annotation result of the target subject. Wherein, the target object can be any object in the picture displayed by the video sample file, such as animals, vehicles, people, etc.

[0138] The first frame of the first video frame sequence is taken as a fixed input, and the second frame of the first video frame sequence is input into the target following model, and the heat map corresponding to the target object in the second frame is obtained through the heat map generation network of the target following model.

[0139] Based on the heat map corresponding to the target object in the second video frame and the first annotation result corresponding to the second video frame, a first training loss corresponding to the second video frame is obtained.

[0140] Based on the first training loss corresponding to the second video frame, only the network parameters of the heat map generation network are adjusted, and then the first training loss corresponding to each of the remaining video frames in the first video frame sequence is obtained frame by frame, and the network parameters of the heat map generation network are adjusted in turn until the network parameters of the heat map generation network tend to be stable, and a target following model in a first stage is obtained.

[0141] The first frame video in the second video sequence is taken as a fixed input, and the second video frame in the second video frame sequence is input to the target following model in the second stage, to obtain a new rotatable rectangular frame corresponding to the target object in the second video frame, and based on the new rotatable rectangular frame corresponding to the target object in the second video frame and the second annotation result corresponding to the second video frame, a new second training loss corresponding to the second video frame is obtained.

[0142] Based on the second training loss corresponding to the second video frame, only the network parameters of the frame generation network are adjusted, and then the second training loss corresponding to each of the remaining video frames in the second video frame sequence is obtained frame by frame, and the network parameters of the frame generation network are adjusted in turn until the network parameters of the frame generation network tend to be stable, and a target following model in a second stage is obtained.

[0143] The first frame video in the second video sequence is taken as a fixed input, and the second video frame in the second video frame sequence is input to the target following model in the second stage, to obtain a new rotatable rectangular frame corresponding to the target object in the second video frame, and based on the new rotatable rectangular frame corresponding to the target object in the second video frame and the second annotation result corresponding to the second video frame, a new second training loss corresponding to the second video frame is obtained.

[0144] Based on the new second training loss corresponding to the second video frame, the network parameters of the heat map generation network and the network parameters of the frame generation network are fine-tuned, and then the new second training loss corresponding to each of the remaining video frames in the second video frame sequence is obtained frame by frame, and the network parameters of the heat map generation network and the network parameters of the frame generation network are fine-tuned in turn until the network parameters of the heat map generation network and the network parameters of the frame generation network tend to be stable, and a trained target following model is obtained. The trained target following model can be used for following the subject of the target object in the video.

[0145] In one example, reference is made toFigure 5 The first frame of video frame marked with the target object and the sequence of video frames to be processed are input into the target following model, and the target following model marks the main body of the target object in the sequence of video frames to be processed through the rotatable rectangular frame frame by frame. The heat map generation network 501 obtains the heat map corresponding to the target object as a whole in the sequence of video frames to be processed based on the first frame of video frame and the sequence of video frames to be processed, and the frame generation network 502 generates the rotatable rectangular frame corresponding to the main body of the target object in the sequence of video frames to be processed based on the heat map corresponding to the target object as a whole in the sequence of video frames to be processed.

[0146] For example, in a television relay sprint race, the specified target athlete is intelligently followed. Obviously, in the race, the limbs of the athlete will be blurred in some frames due to rapid movement, and may be occluded or overlapped with the limbs of adjacent athletes in the picture. The technical solution provided in the embodiment of the application can perceive the main body of the target athlete in this scenario and tend to label the rotating rectangular frame on the torso of the target athlete, while ignoring the limbs, thereby providing a stable and clear following effect.

[0147] In summary, the technical solution provided in the embodiment of the application obtains the second training loss corresponding to the frame generation network through the plurality of second image samples with the second annotation results corresponding to the main body of the target in time sequence, and trains the target following model through the second training loss, thereby realizing training of the target following model based on the clear and explicit main body of the target, reducing the interference of the blurred information (such as ghosting, tailing, blurring, occlusion, etc.) corresponding to the edge part, background, etc. of the target on the target following model, so that the target following model can accurately perceive and follow the main body of the target, thereby improving the following accuracy and explicitness of the target.

[0148] The following is an apparatus embodiment of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the apparatus embodiments of the present application, please refer to the method embodiments of the present application.

[0149] Please refer to Figure 6 which shows a block diagram of a training device of a target following model provided in an embodiment of the present application. The device can be used to implement the training method of the target following model. The device 600 can include a sample sequence acquisition module 601, a first loss acquisition module 602, a second loss acquisition module 603, and a following model training module 604.

[0150] The sample sequence obtaining module 601 is configured to obtain a first image sample sequence and a second image sample sequence; the first image sample sequence includes a plurality of first image samples with a first annotation result corresponding to a target in time sequence, and the second image sample sequence includes a plurality of second image samples with a second annotation result corresponding to a subject of the target in time sequence.

[0151] The first loss obtaining module 602 is configured to obtain a first training loss corresponding to the heat map generation network based on a target first image sample and a remaining first image sample in the first image sample sequence; the first training loss is used to represent a difference between a first output result corresponding to the heat map generation network and the first annotation result, the target first image sample is used to indicate a following target of the target following model, and the first output result is used to represent a predicted distribution of the target.

[0152] The second loss obtaining module 603 is configured to obtain a second training loss corresponding to the frame generation network based on a target second image sample and a remaining second image sample in the second image sample sequence; the second training loss is used to represent a difference between a second output result corresponding to the frame generation network and the second annotation result, the target second image sample is used to indicate a subject of the following target of the target following model, and the second output result is used to represent a predicted distribution of the subject of the target.

[0153] The following model training module 604 is configured to train the target following model based on the first training loss and the second training loss to obtain a trained target following model; the trained target following model is used to follow the subject of the target.

[0154] In an example embodiment, the second loss obtaining module 603 is configured to:

[0155] For each of the remaining second image samples, the first output result corresponding to the remaining second image sample is obtained based on the target second image sample and the remaining second image sample by the heat map generation network;

[0156] The second output result corresponding to the remaining second image sample is obtained based on the first output result corresponding to the remaining second image sample by the frame generation network;

[0157] The second training loss corresponding to the remaining second image sample is obtained based on a difference between the second output result corresponding to the remaining second image sample and the second annotation result;

[0158] obtain the second training loss corresponding to the bounding box generation network based on the second training loss corresponding to each of the remaining second image samples.

[0159] In an example embodiment, the second output result corresponding to the bounding box generation network is a predicted bounding box surrounding a main body of the target, the predicted bounding box corresponds to a predicted horizontal position, a predicted vertical position, a predicted width, a predicted height and a predicted angle, the predicted angle corresponding to the predicted bounding box is a predicted angle between the predicted bounding box and a horizontal coordinate axis; the second loss obtaining module 603 is further configured to:

[0160] obtain a horizontal position loss corresponding to the remaining second image samples based on a difference between the predicted horizontal position corresponding to the predicted bounding box corresponding to the remaining second image samples and a real horizontal position corresponding to the second annotation result corresponding to the remaining second image samples;

[0161] obtain a vertical position loss corresponding to the remaining second image samples based on a difference between the predicted vertical position corresponding to the predicted bounding box corresponding to the remaining second image samples and a real vertical position corresponding to the second annotation result corresponding to the remaining second image samples;

[0162] obtain a width loss corresponding to the remaining second image samples based on a difference between the predicted width corresponding to the predicted bounding box corresponding to the remaining second image samples and a real width corresponding to the second annotation result corresponding to the remaining second image samples;

[0163] obtain a height loss corresponding to the remaining second image samples based on a difference between the predicted height corresponding to the predicted bounding box corresponding to the remaining second image samples and a real height corresponding to the second annotation result corresponding to the remaining second image samples;

[0164] obtain an angle loss corresponding to the remaining second image samples based on a difference between the predicted angle corresponding to the predicted bounding box corresponding to the remaining second image samples and a real angle corresponding to the second annotation result corresponding to the remaining second image samples;

[0165] sum the horizontal position loss, the vertical position loss, the width loss, the height loss and the angle loss to obtain the second training loss corresponding to the remaining second image samples.

[0166] In an example embodiment, the first output result corresponding to the heat map generation network is a heat map corresponding to the target, each value in the heat map is used to represent a possibility that a corresponding pixel belongs to the target; the first loss obtaining module 602 is configured to:

[0167] For each of the remaining first image samples, a first output result corresponding to the remaining first image sample is obtained by the heat map generation network based on the target first image sample and the remaining first image sample;

[0168] A first training loss corresponding to the remaining first image sample is obtained based on a difference between the first output result corresponding to the remaining first image sample and a first label result;

[0169] A first training loss corresponding to the heat map generation network is obtained based on the first training loss corresponding to each of the remaining first image samples.

[0170] In an example embodiment, the following model training module 604 is configured to:

[0171] The heat map generation network is adjusted in parameters based on the first training loss, to obtain a target following model at a first stage;

[0172] The box generation network of the target following model at the first stage is adjusted in parameters based on the second training loss at a first learning rate, to obtain a target following model at a second stage;

[0173] The target following model at the second stage is adjusted in parameters based on the second training loss at a second learning rate, to obtain the target following model at the training completion;

[0174] The second learning rate is less than the first learning rate.

[0175] In an example embodiment, the sample sequence obtaining module 601 is configured to:

[0176] From a data set with mask label results, a sequence of image samples with mask label results of the target is selected;

[0177] For each image sample in the sequence of image samples, a real bounding box enclosing a subject of the target is constructed;

[0178] The real bounding box corresponding to each image sample is adjusted to obtain a plurality of second image samples with second label results of the subject corresponding to the target, with the goal of minimizing the number of subject mask pixels outside the real bounding box and minimizing the number of non-subject mask pixels inside the real bounding box.

[0179] In an example embodiment, the apparatus 600 further comprises a target subject display module and a moving trajectory obtaining module (not shown in the figure). Figure 6

[0180] ​The target subject display module is configured to acquire a predicted bounding box of the target subject in each input image corresponding to the input image sequence by using the trained target following model, and continuously highlight the target subject based on the predicted bounding box.

[0181] The movement trajectory acquisition module is configured to acquire a predicted bounding box of the target subject in each input image corresponding to the input image sequence by using the trained target following model, and acquire a movement trajectory of the target in the input image sequence based on the position of the predicted bounding box.

[0182] In summary, the technical scheme provided by the embodiments of the present application acquires the second training loss of the frame generation network by using the plurality of second image samples corresponding to the second annotation results of the target subject sorted according to time, and trains the target following model by using the second training loss, thereby realizing training of the target following model based on clear and explicit target subjects, reducing the interference of fuzzy information (such as residual image, tailing, blur, occlusion, etc.) of the target on the target following model, and enabling the target following model to accurately perceive and follow the target subject, thereby improving the following accuracy and explicitness of the target.

[0183] It should be noted that the device provided in the above embodiments is only used as an example to divide the above functional modules in realizing the functions, and in actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is described in detail in the method embodiments, which will not be described here.

[0184] Please refer to Figure 7 which shows a structural schematic diagram of a computer device provided in an embodiment of the present application. The computer device can be any electronic device with data calculation, processing and storage functions, and the computer device can be realized as Figure 1 The model training device 10 and / or the model using device 20 in the scheme implementation environment shown. Specifically, it can include the following contents.

[0185] The computer device 700 includes a central processing unit (such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array), etc.) 701, a system memory 704, including a RAM (Random-Access Memory) 702 and a ROM (Read-Only Memory) 703, and a system bus 705 that couples the system memory 704 to the central processing unit 701. The computer device 700 also includes an input / output system (I / O system) 706 that helps transfer information between the various devices within the server, and a mass storage device 707 for storing an operating system 713, application programs 714, and other program modules 715.

[0186] In some embodiments, the input / output system 706 includes a display 708 for displaying information and an input device 709, such as a mouse, keyboard, or the like, for inputting information by a user. The display 708 and input device 709 are both connected to the central processing unit 701 through an input / output controller 710 that is connected to the system bus 705. The input / output system 706 can also include the input / output controller 710 for receiving and processing input from a number of other devices, such as a keyboard, mouse, or electronic stylus, etc. Similarly, the input / output controller 710 also provides output to a display screen, printer, or other type of output device.

[0187] The mass storage device 707 is connected to the central processing unit 701 through a mass storage controller (not shown) that is connected to the system bus 705. The mass storage device 707 and its associated computer-readable media provide non-volatile storage for the computer device 700. That is, the mass storage device 707 can include a computer-readable medium (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.

[0188] Without loss of generality, the computer readable medium can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other solid state memory technology, CD-ROM, DVD (Digital Video Disc), or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices. It should be understood by those skilled in the art that computer storage media does not limit to the above-mentioned several types. The system memory 704 and the mass storage device 707 mentioned above can be collectively referred to as memory.

[0189] According to the embodiments of the present application, the computer device 700 can also run on a remote computer connected to the network through a network such as the Internet. That is, the computer device 700 can be connected to the network 712 through the network interface unit 711 connected to the system bus 705, or can be connected to other types of network or remote computer system (not shown) using the network interface unit 711.

[0190] The memory further includes a computer program stored in the memory and configured to be executed by one or more processors to implement the training method of the target following model.

[0191] In one exemplary embodiment, a computer readable storage medium is also provided, and the storage medium stores a computer program which, when executed by a processor, implements the training method of the target following model.

[0192] Optionally, the computer readable storage medium can include ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives) or optical disc, etc. Among them, the random access memory can include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0193] In an example embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the training method of the target following model.

[0194] It should be noted that the information (including but not limited to object device information, object personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the object or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the image samples, videos, annotation results and the like involved in the present application are obtained under sufficient authorization.

[0195] It should be understood that "multiple" mentioned in the present document refers to two or more. "And / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents that the associated objects before and after it are in an "or" relationship. In addition, the step numbers described in the present document only exemplarily show a possible execution order between steps, and in some other embodiments, the above steps can also be executed in a non-numbered order, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in an order opposite to the illustration, and the embodiments of the present application do not limit this.

[0196] The above only describes example embodiments of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for a target following model, characterized in that, The target following model comprises a heat map generation network and a box generation network, and the method comprises: obtaining a first image sample sequence and a second image sample sequence; wherein the first image sample sequence comprises a plurality of first image samples with target corresponding first annotation results sorted by time, and the second image sample sequence comprises a plurality of second image samples with subject corresponding second annotation results of the target sorted by time; for a target first image sample and a remaining first image sample in the first image sample sequence, obtaining a feature map corresponding to the target first image sample and a feature corresponding to the remaining first image sample through the heat map generation network, and performing deep cross correlation on the feature map corresponding to the target first image sample and the feature map corresponding to the remaining first image sample to obtain a first output result corresponding to the remaining first image sample, the target first image sample being used to indicate a following target of the target following model, the first output result being used to represent a predicted distribution of the target, and the first output result corresponding to the remaining first image sample being a heat map corresponding to the target in the remaining first image sample; based on the first output result corresponding to the remaining first image sample in the first image sample sequence, obtaining a first training loss corresponding to the heat map generation network; wherein the first training loss is used to represent a difference between the first output result corresponding to the remaining first image sample and the first annotation result; for a target second image sample and a remaining second image sample in the second image sample sequence, obtaining a feature map corresponding to the target second image sample and a feature corresponding to the remaining second image sample through the heat map generation network, and performing deep cross correlation on the feature map corresponding to the target second image sample and the feature map corresponding to the remaining second image sample to obtain a first output result corresponding to the remaining second image sample, the target second image sample being used to indicate a subject of the following target of the target following model, and the first output result corresponding to the remaining second image sample being a heat map corresponding to the target in the remaining second image sample; obtaining a second output result corresponding to the remaining second image sample based on the first output result corresponding to the remaining second image sample through the box generation network, the second output result being used to represent a predicted distribution of the subject of the target; based on the second output result corresponding to the remaining second image sample in the second image sample sequence, obtaining a second training loss corresponding to the box generation network; wherein the second training loss is used to represent a difference between the second output result corresponding to the remaining second image sample and the second annotation result; based on the first training loss and the second training loss, training the target following model to obtain a trained target following model; wherein the trained target following model is used to follow the subject of the target.

2. The method of claim 1, wherein, The second training loss corresponding to the bounding box generation network is obtained based on the second output result corresponding to each of the remaining second image samples in the second image sample sequence. The second training loss corresponding to each of the remaining second image samples is obtained based on a difference between the second output result corresponding to the remaining second image sample and a second annotation result corresponding to the remaining second image sample. The second training loss corresponding to the bounding box generation network is obtained based on the second training loss corresponding to each of the remaining second image samples.

3. The method of claim 2, wherein, The second output result corresponding to the bounding box generation network is a predicted bounding box surrounding a main body of the target, and the predicted bounding box corresponds to a predicted horizontal position, a predicted vertical position, a predicted width, a predicted height, and a predicted angle, wherein the predicted angle of the predicted bounding box is a predicted included angle between the predicted bounding box and a horizontal coordinate axis. The second training loss corresponding to each of the remaining second image samples is obtained based on a difference between the second output result corresponding to the remaining second image sample and a second annotation result corresponding to the remaining second image sample. The horizontal position loss corresponding to the remaining second image sample is obtained based on a difference between the predicted horizontal position corresponding to the predicted bounding box corresponding to the remaining second image sample and a real horizontal position corresponding to the second annotation result corresponding to the remaining second image sample. The vertical position loss corresponding to the remaining second image sample is obtained based on a difference between the predicted vertical position corresponding to the predicted bounding box corresponding to the remaining second image sample and a real vertical position corresponding to the second annotation result corresponding to the remaining second image sample. The width loss corresponding to the remaining second image sample is obtained based on a difference between the predicted width corresponding to the predicted bounding box corresponding to the remaining second image sample and a real width corresponding to the second annotation result corresponding to the remaining second image sample. The height loss corresponding to the remaining second image sample is obtained based on a difference between the predicted height corresponding to the predicted bounding box corresponding to the remaining second image sample and a real height corresponding to the second annotation result corresponding to the remaining second image sample. The angle loss corresponding to the remaining second image sample is obtained based on a difference between the predicted angle corresponding to the predicted bounding box corresponding to the remaining second image sample and a real angle corresponding to the second annotation result corresponding to the remaining second image sample. The second training loss corresponding to the remaining second image sample is obtained by summing the horizontal position loss, the vertical position loss, the width loss, the height loss, and the angle loss.

4. The method of claim 1, wherein, Each value in the heat map is used to represent a possibility that a corresponding pixel belongs to the target. The first training loss corresponding to the heat map generation network is obtained based on the first output result corresponding to each of the remaining first image samples in the first image sample sequence. For each of the remaining first image samples in the first image sample sequence, a first training loss corresponding to the remaining first image sample is obtained based on a difference between a first output result corresponding to the remaining first image sample and a first annotation result; A first training loss corresponding to the heat map generation network is obtained based on the first training loss corresponding to each of the remaining first image samples.

5. The method of claim 1, wherein, The training of the target following model based on the first training loss and the second training loss to obtain a trained target following model comprises: Parameter adjustment of the heat map generation network based on the first training loss to obtain a first-stage target following model; Parameter adjustment of the bounding box generation network of the first-stage target following model based on the second training loss at a first learning rate to obtain a second-stage target following model; Parameter adjustment of the second-stage target following model based on the second training loss at a second learning rate to obtain the trained target following model; The second learning rate is less than the first learning rate.

6. The method of claim 1, wherein, The second image sample is obtained as follows: From a data set with mask annotation results, an image sample sequence with mask annotation results of the target is selected; For each image sample in the image sample sequence, a real bounding box enclosing the subject of the target is constructed; The real bounding box corresponding to each image sample is adjusted to obtain a plurality of second image samples with second annotation results of the subject of the target, with the goal of minimizing the number of subject mask pixels outside the real bounding box and minimizing the number of non-subject mask pixels inside the real bounding box.

7. The method of claim 1, wherein, The method further comprises: Through the trained target following model, a predicted bounding box of the subject of the target in each input image corresponding to an input image sequence is obtained, and the subject of the target is continuously highlighted based on the predicted bounding box; Or, Through the trained target following model, a predicted bounding box of the subject of the target in each input image corresponding to an input image sequence is obtained, and a moving trajectory of the target in the input image sequence is obtained based on the position of the predicted bounding box.

8. A training device for a target following model, characterized in that, The target following model comprises a heat map generation network and a bounding box generation network, and the device comprises: A sample sequence acquisition module is configured to acquire a first image sample sequence and a second image sample sequence; wherein the first image sample sequence comprises a plurality of first image samples with first annotation results corresponding to a target sorted by time, and the second image sample sequence comprises a plurality of second image samples with second annotation results corresponding to a subject of the target sorted by time; The first loss obtaining module is configured to, for a target first image sample and remaining first image samples in the first image sample sequence, obtain a feature map corresponding to the target first image sample and features corresponding to the remaining first image samples by using the heat map generation network, and perform deep cross-correlation on the feature map corresponding to the target first image sample and the feature maps corresponding to the remaining first image samples to obtain first output results corresponding to the remaining first image samples, the target first image sample being used to indicate a following target of the target following model, the first output results being used to represent a predicted distribution of the target, and the first output results corresponding to the remaining first image samples being heat maps corresponding to targets in the remaining first image samples; and obtain a first training loss corresponding to the heat map generation network based on the first output results corresponding to the remaining first image samples in the first image sample sequence, the first training loss being used to represent a difference between the first output results corresponding to the remaining first image samples and the first annotation results. The second loss obtaining module is configured to, for a target second image sample and remaining second image samples in the second image sample sequence, obtain a feature map corresponding to the target second image sample and features corresponding to the remaining second image samples by using the heat map generation network, and perform deep cross-correlation on the feature map corresponding to the target second image sample and the feature maps corresponding to the remaining second image samples to obtain first output results corresponding to the remaining second image samples, the target second image sample being used to indicate a subject of the following target of the target following model, the first output results corresponding to the remaining second image samples being heat maps corresponding to targets in the remaining second image samples; obtain second output results corresponding to the remaining second image samples based on the first output results corresponding to the remaining second image samples by using the box generation network, the second output results being used to represent a predicted distribution of the subject of the target; and obtain a second training loss corresponding to the box generation network based on the second output results corresponding to the remaining second image samples in the second image sample sequence, the second training loss being used to represent a difference between the second output results corresponding to the remaining second image samples and the second annotation results. The following model training module is configured to train the target following model based on the first training loss and the second training loss to obtain a trained target following model, the trained target following model being used to follow the subject of the target.

9. A computer device, comprising: The computer device includes a processor and a memory, and the memory stores a computer program, which is loaded and executed by the processor to implement the target following model training method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, which is loaded and executed by the processor to implement the target following model training method according to any one of claims 1 to 7.

11. A computer program product, characterised in that, The computer program product comprises computer instructions executed by a processor to implement the training method of the target following model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video real-time multi-target detection and tracking method and device based on deep learning

    CN112288770A

  • Target tracking method and device and electronic equipment

    CN112967315A

Cited By

  • Object picking optimization

    US20240058953A1