A target positioning method, system, storage medium, and terminal device
By acquiring the feature information and optical flow information of the video frame and combining the candidate position information, the precise positioning of the target object in the video is achieved, the problem of poor positioning effect in the prior art is solved, and the positioning accuracy and robustness are improved.
Patent Information
- Application Number
- CN202110495900.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-05-07
AI Technical Summary
The existing target positioning methods are difficult to accurately locate the target object in the video under complex backgrounds, small video interfaces, frequent objects entering and exiting, and shaking of the shooting device.
By obtaining the characteristic information of the reference frame image and the to-process frame image in the video to be processed, and combining the optical flow information, the candidate position information of the target object is determined, and the precise positioning of the target object is finally achieved.
The noise of the static appearance feature is offset by the moving characteristics of the image, which improves the positioning accuracy of the target object and reduces the impact of occlusion and background interference.
Smart Images

Figure CN113763420B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing based on artificial intelligence, and particularly relates to a target positioning method, system, storage medium and terminal device. Background Art
[0002] Since the development of target positioning technology in the 1960s, the main target positioning algorithms are divided into two categories. One is the positioning algorithm based on correlation filtering, and the other is the positioning algorithm based on deep learning. Among them, the positioning algorithm based on correlation filtering mainly locates the target through the cross-correlation of two pictures, and applies the Fourier transform to convert the convolution operation in the spatial domain to the frequency domain, greatly improving the operation speed; while the positioning algorithm based on deep learning mainly uses the machine learning model of artificial intelligence to extract features from pictures, and locates the target object in the pictures based on the extracted features.
[0003] However, in practical applications, due to various factors such as complex backgrounds, small video interfaces, frequent entry and exit of multiple objects in the video, occlusion between objects, and shaking of the shooting device during video shooting in some video scenarios, the positioning effect of the target object in the video by the existing target positioning methods is less than satisfactory. Summary of the Invention
[0004] Embodiments of the present invention provide a target positioning method, system, storage medium and terminal device, which more accurately realize the positioning of the target object.
[0005] An embodiment of the present invention provides a target positioning method, including:
[0006] Obtaining first feature information of a target picture block included in a reference frame image in a video to be processed, and obtaining second feature information of a frame image to be processed in the video to be processed; wherein, the target picture block is an image including a target object;
[0007] Determining first candidate position information of the target object in the frame image to be processed according to the first feature information and the second feature information;
[0008] Obtaining first optical flow information of the reference frame image, obtaining second optical flow information according to the frame image to be processed, and obtaining reference position information of the target picture block in the reference frame image;
[0009] Determining second candidate position information of the target object in the frame image to be processed according to the first optical flow information, the second optical flow information and the reference position information;
[0010] Determining the position information of the target object in the frame image to be processed according to the first candidate position information and the second candidate position information.
[0011] Another embodiment of the present invention provides a target positioning system, including:
[0012] A feature acquisition unit, configured to acquire first feature information of a target picture block included in a reference frame image in a video to be processed, and acquire second feature information of a frame image to be processed in the video to be processed; wherein, the target picture block is an image including a target object;
[0013] A first candidate unit, configured to determine first candidate position information of the target object in the frame image to be processed according to the first feature information and the second feature information;
[0014] An optical flow information unit, configured to acquire first optical flow information of the reference frame image, acquire second optical flow information according to the frame image to be processed, and acquire reference position information of the target picture block in the reference frame image;
[0015] A second candidate unit, configured to determine second candidate position information of the target object in the frame image to be processed according to the first optical flow information, the second optical flow information and the reference position information;
[0016] A position determination unit, configured to determine position information of the target object in the frame image to be processed according to the first candidate position information and the second candidate position information.
[0017] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium, which stores a plurality of computer programs, and the computer programs are suitable for being loaded and executed by a processor to perform the target positioning method as described in an embodiment of the present invention.
[0018] Another aspect of the embodiments of the present invention further provides a terminal device, including a processor and a memory;
[0019] The memory is used to store a plurality of computer programs, and the computer programs are used to be loaded and executed by the processor to perform the target positioning method as described in an embodiment of the present invention; the processor is used to implement each computer program in the plurality of computer programs.
[0020] It can be seen that in the method of this embodiment, the target positioning system locates the target object based on the static appearance features of the image, that is, determines the first candidate position information of the target object in the image to be processed according to the first feature information and the second feature information of the target picture block and the image to be processed; then combines the positioning of the target object based on the motion features of the image, that is, determines the second candidate position information of the target object in the image to be processed according to the first optical flow information and the second optical flow information of the reference frame image and the image to be processed and the reference position information of the target picture block, and further realizes the final positioning of the target object according to the first candidate position information and the second candidate position information. In this way, the motion features of the image can offset the noise in the process of locating the target object by the static appearance features of the image as much as possible, such as the occlusion of the target object and background interference, etc., making the final positioning of the target object more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 is a schematic diagram of a target positioning method provided by an embodiment of the present invention;
[0023] Figure 2 is a flowchart of a target positioning method provided by an embodiment of the present invention;
[0024] Figure 3 is a schematic diagram of matching a target picture block with an image within a candidate box in an embodiment of the present invention;
[0025] Figure 4 is a flowchart of a method for training an appearance feature model in an embodiment of the present invention;
[0026] Figure 5 is a schematic diagram of the logical structure of an appearance feature model in an embodiment of the present invention;
[0027] Figure 6 is a flowchart of a method for training a motion feature model in an embodiment of the present invention;
[0028] Figure 7 is a schematic diagram of the structure of a target positioning system in an application embodiment of the present invention;
[0029] Figure 8 is a schematic diagram of a distributed system to which the target positioning method in another application embodiment of the present invention is applied;
[0030] Figure 9 It is a schematic diagram of the block structure in another application embodiment of the present invention;
[0031] Figure 10 It is a schematic diagram of the structure of a target positioning system provided by an embodiment of the present invention;
[0032] Figure 11 It is a schematic diagram of the structure of a terminal device provided by an embodiment of the present invention. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] An embodiment of the present invention provides a target positioning method, which can be mainly applied to locate target objects in each frame of a video (especially a short video), such as Figure 1 As shown, in the embodiment of the present invention, the target positioning system can locate the target object according to the following method:
[0036] Obtain the first feature information of the target picture block included in the reference frame image of the video to be processed, and obtain the second feature information of the frame image to be processed in the video to be processed; wherein, the target picture block is an image containing a target object; according to the first feature information and the second feature information, determine the first candidate position information of the target object in the frame image to be processed; obtain the first optical flow information of the reference frame image, obtain the second optical flow information according to the frame image to be processed, and obtain the reference position information of the target picture block in the reference frame image; according to the first optical flow information, the second optical flow information and the reference position information, determine the second candidate position information of the target object in the frame image to be processed; according to the first candidate position information and the second candidate position information, determine the position information of the target object in the frame image to be processed.
[0037] The above determination of the first candidate position information can be achieved through an appearance feature model, and the second candidate position information can be achieved through a motion feature model. Both the appearance feature model and the motion feature model are machine learning models based on artificial intelligence. Among them, Artificial Intelligence (AI) is the theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making.
[0038] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0039] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0040] In this way, through the motion features of the image, the noise of the static appearance features of the image during the target object localization process can be offset as much as possible, such as the occlusion of the target object and background interference, etc., making the final localization of the target object more accurate.
[0041] An embodiment of the present invention provides a target localization method, which is mainly the method executed by the above-mentioned target localization system. The flowchart is as Figure 2 shown, including:
[0042] Step 101, obtain the first feature information of the target picture block included in the reference frame image in the video to be processed, and obtain the second feature information of the frame image to be processed in the video to be processed, where the target picture block is an image containing the target object.
[0043] It can be understood that the target localization method of the embodiment of the present invention is mainly to localize the target object in the video to be processed. A certain frame image (such as the first frame image) in the video to be processed is used as the reference frame image, and the target picture block in the reference frame image is specified. Through the method of this embodiment, the target object can be localized for each frame image (i.e., the above-mentioned frame image to be processed) other than the reference frame image in the video to be processed.
[0044] Specifically, when the target localization system obtains the first feature information of the target picture block and the second feature information of the frame image to be processed, the features can be extracted through a convolutional neural network.
[0045] Step 102, determine the first candidate position information of the target object in the frame image to be processed according to the first feature information and the second feature information.
[0046] Specifically, the target positioning system first determines multiple candidate boxes where the target object is located in the image of the frame to be processed based on the first feature information and the second feature information. Then, it matches the feature information of the images within each candidate box with the first feature information of the target picture block in the reference frame image to obtain the matching degree corresponding to each candidate box. Finally, it determines the first candidate position information of the target object in the image of the frame to be processed based on the matching degrees corresponding to each candidate box. For example, it can directly determine the position information of the candidate box with the highest matching degree as the first candidate position information of the target object in the image of the frame to be processed.
[0047] Among them, when the target positioning system matches the feature information of the images within each candidate box with the first feature information, it can perform feature matching at multiple granularities, that is, through global feature matching and local feature matching. Specifically, it matches the global feature information of the images within each candidate box with the first feature information to obtain the first sub-matching degree; it divides the images within each candidate box and the target picture block into multiple regions according to the same strategy, and respectively matches the local feature information of each region in the images within each candidate box with the local feature information of the corresponding region in the target picture block to obtain the second sub-matching degree; it determines the matching degree corresponding to each candidate box according to the first sub-matching degree and the second sub-matching degree. For example, it takes the weighted sum value of the first sub-matching degree and the second sub-matching degree corresponding to a candidate box as the matching degree corresponding to this candidate box, etc.
[0048] For example Figure 3 As shown, both the target picture block and the image within a candidate box are divided into upper and lower region images. During the calculation process, it matches the global feature 1 of the target picture block with the global feature 2 of the image within the candidate box to obtain the first sub-matching degree, matches the local feature 11 of the upper region of the target picture block with the local feature 21 of the upper region of the image within the candidate box to obtain the second sub-matching degree 1, and matches the local feature 12 of the lower region of the target picture block with the local feature 22 of the lower region of the image within the candidate box to obtain the second sub-matching degree 2. Then, it can obtain the matching degree between the candidate box image and the target picture block based on the first sub-matching and the second sub-matching degrees 1 and 2.
[0049] It should be noted that when the target positioning system executes the above steps 101 and 102, it can use a trained appearance feature model to execute. This appearance feature model is mainly used to determine the target object in the image of the frame to be processed based on the static features of the image of the frame to be processed. Among them, the appearance feature model is a machine learning model based on artificial intelligence, which can generally be trained by a certain method, and the operation logic of the trained appearance feature model is stored in the target positioning system.
[0050] Step 103: Obtain the first optical flow information of the reference frame image, obtain the second optical flow information according to the image to be processed, and obtain the reference position information of the target picture block in the reference frame image.
[0051] Here, the reference position information of the target picture block in the reference frame image may be the coordinate information of the points on the contour line of the target picture block, etc. Generally, the shape of the contour line of the target picture block is rectangular, then the reference position information is specifically the coordinate information of the four vertices of the rectangle in the reference frame image.
[0052] The optical flow of any image refers to the instantaneous velocity of the pixel motion of a spatial moving object on the observation imaging plane, and uses the change of pixels in the time domain in the image sequence and the correlation between adjacent frame images to find the corresponding relationship between the previous frame image and the current frame image, so that the motion information of the object between adjacent frame images can be calculated according to the optical flow of the image. In this embodiment, the obtained first optical flow information and second optical flow information may be dense optical flow, that is, the optical flow information for each pixel point in the image, so that the information of the image will not be lost. Among them, the second optical flow information may be the optical flow information of the image to be processed, or the optical flow information of the image obtained after preprocessing the image to be processed. Here, the preprocessing may include processing such as transforming the image to be processed into the coordinate system of the reference frame image.
[0053] Among them, since the shooting device may move during the process of shooting the video to be processed, the optical flow information of each frame image obtained in the above step 103 not only includes the motion state of the object in the image, but also includes the motion state of the shooting device. In order to eliminate the interference of the movement of the shooting device as much as possible, in this embodiment, the target positioning system will preprocess the image to be processed, that is, transform the image to be processed into the coordinate system of the reference frame image to obtain the transformed image to be processed, and then obtain the second optical flow information of the transformed image to be processed.
[0054] Specifically, when the target positioning system transforms the image to be processed into the coordinate system of the reference frame image, it can first calculate the homography matrix between the reference frame image and the image to be processed according to the information of the background region feature points in the reference frame image and the information of the background region feature points in the image to be processed, and then transform the image to be processed into the coordinate system of the reference frame image according to the calculated homography matrix.
[0055] Among them, if two cameras capture two images A and B of the same space, and there is a transformation between one image A and the other image B, and this transformation is a one-to-one correspondence, this relationship can be represented by a matrix, and this matrix is the homography matrix. In this embodiment, the homography matrix refers to the correspondence of the transformation between the reference frame image and the image to be processed. The feature points in the background area of the reference frame image refer to the pixel points in the area of the reference frame image except for the target picture block, and the feature points in the background area of the image to be processed refer to the pixel points in the area of the image to be processed except for the target object. The information of each feature point can be described by various description methods, such as Speeded Up Robust Features (SURF) feature points, etc.
[0056] Denote H r,i as the homography matrix between the reference frame image f r and the i-th frame image f i (i.e., the image to be processed as described above). The feature point set is denoted as f r b , f i b which are the information of the feature points in the background areas of the reference frame image and the i-th frame image respectively. The homography matrix can be obtained by the Random Sample Consensus (RANSAC) algorithm shown in the following formula 1-1, and the image to be processed can be transformed into the coordinate system of the reference frame image by the following formula 1-2.
[0057] H r,i = RANSAC(P(f r b , f i b )) (1-1)
[0058] H r,i × f r b = f i b (1-2)
[0059] Step 104: Determine the second candidate position information of the target object in the image to be processed according to the first optical flow information, the second optical flow information, and the reference position information.
[0060] It can be understood that when the target positioning system executes the above steps 103 and 104, a trained motion feature model can be used to execute. The motion feature model is mainly used to determine the target object in the to-be-processed frame image according to the features of the motion of the to-be-processed frame image. Among them, the motion feature model is a machine learning model based on artificial intelligence, which can generally be obtained through a certain method of training, and the operation logic of the trained motion feature model is stored in the target positioning system.
[0061] Step 105, determine the position information of the target object in the to-be-processed frame image according to the first candidate position information and the second candidate position information.
[0062] Specifically, in one case, the target positioning system can directly perform a certain calculation on the first candidate position information and the second candidate position information to obtain the position information of the target object.
[0063] In another case, the target positioning system will first update the first candidate position information to obtain the updated candidate position information, and perform a certain calculation on the updated candidate position information and the second candidate position information to obtain the final position information of the target object. Specifically:
[0064] The target positioning system will set multiple update rates for the first candidate position information; calculate the updated candidate position information under each update rate according to the update rate, the first candidate position information, and the position information of the object target in the previous frame image of the to-be-processed frame image; calculate the distance between the sub-image corresponding to each updated candidate position information and the target picture block in the to-be-processed frame image respectively, and select the updated candidate position information corresponding to the sub-image with the smallest distance; determine the position information of the target object in the to-be-processed frame image according to the selected updated candidate position information and the second candidate position information.
[0065] Among them, in order to increase the stability of determining the position information of the target object according to the appearance features of the to-be-processed frame image (i.e., the above second feature information), especially when the target object encounters short-term occlusion and sudden interference, the stability of positioning the target object can be ensured. The target positioning system will update the first candidate position information according to the position information of the target object in the previous frame image. For example, the following formulas (2-1) and (2-2) can be used to obtain the size of the updated frame where the target object is located, and then the updated candidate position information can be obtained:
[0066] τ(w,lr) = w'*(1 - lr) + w*lr (2-1)
[0067] τ(h,lr) = h'*(1 - lr) + h*lr (2-2)
[0068] Wherein, w' and h' are respectively the width and height of the bounding box where the target object is located in the previous frame image of the frame image to be processed, and w and h are respectively the width and height of the bounding box where the target object is located in the frame image to be processed; lr is the weighting ratio, that is, the update rate. In this embodiment, multiple update rates can be set, and one of the most suitable update rates can be adaptively selected from the multiple update rates to obtain the updated candidate position information.
[0069] After obtaining an updated candidate position information for each update rate, the distance between the sub-image corresponding to each updated candidate position information and the target picture block in the frame image to be processed can be calculated according to the following formula 2-3, where is the feature information of the sub-image corresponding to the updated candidate position information, is the feature information of the target picture block:
[0070]
[0071] Then, select the updated candidate position information corresponding to the sub-image with the smallest distance. The size of the bounding box where the sub-image with the smallest distance is located can be obtained according to formulas 2-4 and 2-5, and then the corresponding updated candidate position information can be obtained. Among them, the multiple update rates are respectively lr, and lr*γ:
[0072]
[0073]
[0074] It should be noted that when the target object encounters similar background interference and short-term occlusion, if only through the methods of steps 101 and 102 above, some information of the static appearance features of the frame image to be processed (i.e., the above-mentioned second feature information) will inevitably be lost, which will make the positioning of the target object in the frame image to be processed not very accurate. In this embodiment, the target positioning system not only needs to consider the appearance features of the frame image to be processed, but also needs to combine the motion features of the frame image to be processed, that is, by the methods of steps 103 and 104, the second optical flow information of the frame image to be processed is obtained, so that the finally obtained position information of the target object is more accurate, and the accuracy of positioning the target object in the frame image to be processed is improved.
[0075] It can be seen that in the method of this embodiment, the target positioning system locates the target object based on the static appearance features of the image, that is, determines the first candidate position information of the target object in the to-be-processed frame image according to the first feature information and the second feature information of the target picture block and the to-be-processed frame image; then combines the positioning of the target object based on the motion features of the image, that is, determines the second candidate position information of the target object in the to-be-processed frame image according to the first optical flow information and the second optical flow information of the reference frame image and the to-be-processed frame image and the reference position information of the target picture block, and further realizes the final positioning of the target object according to the first candidate position information and the second candidate position information. In this way, the motion features of the image can offset the noise in the target object positioning process caused by the static appearance features of the image as much as possible, such as occlusion and background interference of the target object, making the final positioning of the target object more accurate.
[0076] In a specific embodiment, the above steps 101 and 102 can be implemented by an appearance feature model based on artificial intelligence, and the appearance feature model can be trained through the following steps. The flowchart is as Figure 4 shown and includes:
[0077] Step 201, determine the initial appearance feature model, where the initial appearance feature model includes an appearance feature extraction module, a position regression module, a classification module, and a prediction module.
[0078] It can be understood that when the target positioning system determines the initial appearance feature model, it will determine the multi-layer structure included in the initial appearance feature model and the initial values of the parameters in each layer structure. Among them, the parameters in each layer structure refer to the fixed parameters used in the calculation process of each layer structure in the initial appearance feature model, which do not need to be assigned values at any time, such as parameters such as parameter scale, number of network layers, and user vector length.
[0079] Such as Figure 5As shown in the figure, the structure of the initial appearance feature model may specifically include: an appearance feature extraction module, which is used to extract the feature information of the to-be-processed frame image and the target picture block respectively, generally a siamese network; a position regression module, which is used to determine the position information of the candidate box where the target object is located in the to-be-processed frame image according to the feature information of the to-be-processed frame image and the target picture block extracted by the feature extraction module; a prediction module, which is used to select the position information of a certain candidate box as the position information of the sample object in the sample image. Specifically, the matching score between the image in each candidate box and the target picture block can be calculated. If the matching score corresponding to a certain candidate box is greater than a certain threshold, the position information of this candidate box is used as the position information of the target object; a classification module, which is used to determine whether the image in the candidate box where the target object is located determined by the position regression module belongs to the target object. Specifically, the probability information that the image in the candidate box where the target object is located belongs to the target object can be output. If this probability information is greater than the preset value, the image in the box where the target object is located belongs to the target object.
[0080] It should be noted that, as Figure 5 shown in the figure, generally in the specific implementation, the classification module in the initial appearance feature model may include two. That is, a classification module is connected after the position regression module, and this classification module also needs to classify based on the feature information extracted by the above appearance feature extraction module when classifying; another classification model is also connected after the prediction module, which is used to determine whether the image corresponding to the position information predicted by the prediction module belongs to the sample object.
[0081] Step 202: Determine the first training sample. The first training sample includes multiple first sample image groups. Each first sample image group includes a sample object picture block, at least one sample image, the position annotation information of multiple sample boxes in the sample image, and the type annotation of whether each sample box belongs to the box where the sample object is located.
[0082] Step 203: The appearance feature extraction module is used to obtain the feature information of the sample object picture block and the sample image respectively. The position regression module determines the position information of the candidate box where the sample object is located in the sample image according to the feature information of the sample object picture block and the sample image. The prediction module selects the position information of a certain candidate box as the position information of the sample object in the sample image. The classification module determines the type information of whether the image in the candidate box where the sample object is located determined by the position regression module belongs to the sample object.
[0083] Among them, the prediction module can specifically perform feature matching between the images within each candidate box and the sample object picture block, and calculate the matching scores between the images within each candidate box where the sample object is located and the sample object picture block. If the matching score corresponding to a certain candidate box is greater than a certain threshold, the position information of this candidate box is used as the position information of the sample object. Specifically, when performing feature matching between the images within each candidate box and the sample object picture block, the global features and local features between the images within the candidate box and the sample object picture block can be respectively matched.
[0084] Step 204: According to the position information obtained by the prediction module, the position annotation information in the first training sample, the type information determined by the classification module, and the type annotation in the first training sample, adjust the initial appearance feature model to obtain the final appearance feature model.
[0085] Specifically, the target positioning system first calculates a first loss function related to the appearance feature extraction module, the position regression module, and the prediction module according to the position information obtained by the prediction module and the position annotation information in the first training sample. This first loss function is used to indicate the error between the position information of the sample object obtained by the appearance feature extraction module, the position regression module, and the prediction module, and the actual position information of the sample object in each sample image in the first training sample (obtained according to the position annotation information), such as the cross-entropy loss function, etc.; according to the type information obtained by the classification module and the type annotation in the first training sample, calculate a second loss function related to the appearance feature extraction module, the position regression module, and the classification module. This second loss function is used to indicate the error between the type information obtained by the appearance feature extraction module and the position regression module and the actual type of the image within the candidate box in each sample image in the first training sample (obtained according to the type annotation); then calculate an overall loss function according to the first loss function and the second loss function, such as the overall loss function being the weighted sum value of the first loss function and the second loss function, etc.; and then adjust the parameter values of the parameters in the above initial appearance feature model according to the overall loss function.
[0086] The training process of the appearance feature model is to minimize the value of the above error as much as possible. This training process continuously optimizes the parameter values of the parameters in the initial appearance feature model determined in the above step 201 through a series of mathematical optimization means such as backpropagation for derivation and gradient descent, and makes the calculated value of the above overall loss function drop to the lowest.
[0087] Specifically, when the function value of the calculated overall loss function is relatively large, such as greater than a preset value, the parameter values need to be changed, such as reducing the weight value connected to a certain neuron, etc., so that the function value of the overall loss function calculated according to the adjusted parameter values decreases.
[0088] In a specific implementation process, as described above Figure 5 As shown, the appearance feature model of this embodiment is mainly divided into two stages. In the first stage, the position information of the candidate box where the sample object is located is obtained through the appearance feature extraction module and the position regression module. At the same time, an image within the candidate box determined by the position regression module is classified using a classification model. In the second stage, the position information of the sample object is obtained through the prediction module. At the same time, an image within the candidate box determined by the prediction module can also be classified using another classification model. In this way, the overall loss function calculated by the target positioning system can be divided into the loss functions of the first stage and the second stage. And the loss function of each stage can include two parts, namely, the part of the position regression module (or prediction module) and the part of the classification model. In this way, through the supervision of the two parts, the adjustment of the parameter values in the appearance feature model can be made more accurate.
[0089] Specifically, the loss function L 1,reg based on the position regression module in the first stage and the loss function L 2,reg based on the prediction module in the second stage can be calculated through the following formulas 3-1 to 3-5, and the overall loss function can be calculated through the following formula 3-6, where L 1,cls and L 2,cls are the loss functions based on the classification module in the two stages:
[0090]
[0091]
[0092]
[0093]
[0094]
[0095] L = γ 1 L 1,cls + γ 2 L 1,reg + γ 3 L 2,cls + γ 4 L 2,reg (3-6)
[0096] Where, A x , A y , A ω , A h are the position information obtained through the initial appearance feature model, specifically the center point coordinates and width and height of the box where the sample object is located, T x , T y , T ω , Th are respectively the center coordinates, width, and height of the sample box marked in the first training sample, and γ 1 , γ 2 , γ 3 , γ 4 is the weight value.
[0097] In addition, it should be noted that the above steps 203 to 204 are an adjustment of the parameter values of the parameters in the initial appearance feature model based on the position information obtained by the initial appearance feature model and the type information obtained by the classification module. In actual applications, the above steps 203 to 204 need to be continuously looped until the adjustment of the parameter values meets a certain stop condition.
[0098] Therefore, after the target positioning system executes the above steps 201 to 204 of the embodiment, it is also necessary to determine whether the current adjustment of the parameter values meets the preset stop condition. When it is met, the process ends; when it is not met, the initial appearance feature model after adjusting the parameter values is returned to execute the above steps 203 to 204. The preset stop conditions include, but are not limited to, any one of the following conditions: the difference between the currently adjusted parameter value and the parameter value adjusted last time is less than a threshold, that is, the adjusted parameter value reaches convergence; and the number of times of adjusting the parameter values is equal to the preset number of times, etc.
[0099] In another specific embodiment, the above steps 103 and 104 can be implemented by a motion feature model based on artificial intelligence, and the motion feature model can be trained through the following steps. The flowchart is as Figure 6 shown, including:
[0100] Step 301, determine the initial motion feature model, which includes a motion feature extraction module and a position determination module.
[0101] It can be understood that when the target positioning system determines the initial motion feature model, it will determine the multi-layer structure included in the initial motion feature model and the initial values of the parameters in each layer structure. Among them, the parameters in each layer structure refer to the fixed parameters used in the calculation process of each layer structure in the initial motion feature model, which do not need to be assigned values at any time, such as parameters such as parameter scale, number of network layers, and user vector length.
[0102] The structure of the initial motion feature model can specifically include: a motion feature extraction module, which is used to extract the features of the optical flow information of any two frames of images and the features of the reference position information of the target picture block in a certain frame of image respectively; a position determination module, which is used to determine the position information of the target object in the other frame of the above any two frames of images according to the features extracted by the motion feature extraction module.
[0103] Step 302: Determine the second training sample. The second training sample includes multiple groups of second sample images. Each group of second sample images includes the optical flow information corresponding to two sample images and the position annotation information of the sample object in the two sample images respectively.
[0104] Step 303: The motion feature extraction module extracts the features of the optical flow information of each sample image in the group of second sample images, and the features of the position annotation information of the sample object in a certain sample image. The position determination module determines the position information of the sample object in another sample image within the group of second sample images according to the features extracted by the motion feature extraction module.
[0105] Step 304: Adjust the initial motion feature model according to the position information obtained by the position determination module and the position annotation information in the second training sample to obtain the final motion feature model.
[0106] Specifically, the target positioning system first calculates the loss function related to the motion feature extraction module and the position determination module according to the position information obtained by the position determination module and the position annotation information in the second training sample. This loss function is used to indicate the difference between the position information of the sample object obtained by the motion feature extraction module and the position determination module and the actual position information of the sample object in each sample image in the second training sample (obtained according to the position annotation information). Then, the parameter values in the above initial motion feature model are adjusted according to the calculated loss function.
[0107] The training process of the motion feature model is to minimize the value of the above difference. This training process continuously optimizes the parameter values in the initial motion feature model determined in Step 301 through a series of mathematical optimization means such as backpropagation derivative and gradient descent, and makes the calculated value of the above loss function drop to the lowest. Specifically, when the function value of the calculated loss function is relatively large, such as greater than a preset value, the parameter values need to be changed, such as reducing the weight value of a neuron connection, etc., so that the function value of the loss function calculated according to the adjusted parameter values decreases.
[0108] In a specific implementation, the loss function calculated by the target positioning system based on the motion feature model adopts the (D-IOU) loss. Specifically, it can be expressed by the following formula 4:
[0109]
[0110] where ρ is the Euclidean distance, b and b gtThey are respectively the position information determined by the position determination module (specifically, the center coordinates of the box where the sample object is located, i.e., the prediction box) and the position annotation information in the second training label sample (specifically, the center coordinates of the box where the sample object is located in the second training sample, i.e., the annotation box), and c is the diagonal length of the smallest rectangle enclosing the prediction box and the annotation box.
[0111] In addition, it should be noted that the above steps 303 to 304 are an adjustment of the parameter values of the parameters in the initial motion feature model based on the position information obtained by the initial motion feature model. In actual applications, it is necessary to continuously loop through the above steps 303 to 304 until the adjustment of the parameter values meets a certain stop condition.
[0112] Therefore, after the target positioning system executes the above steps 301 to 304 of the embodiment, it is also necessary to determine whether the current adjustment of the parameter values meets the preset stop condition. When it is met, the process ends; when it is not met, for the initial motion feature model after adjusting the parameter values, return to execute the above steps 303 to 304. Among them, the preset stop conditions include but are not limited to any one of the following conditions: the difference between the currently adjusted parameter value and the parameter value adjusted last time is less than a threshold, that is, the adjusted parameter value reaches convergence; and the number of times of adjusting the parameter values is equal to the preset number, etc.
[0113] The following uses a specific application example to illustrate the target positioning method in the present invention. For example Figure 7 As shown, the target positioning system in this embodiment is a multi-cue two-stage locator, denoted as M-SPM, which may include an appearance feature model, a motion feature model, an adaptive update module, and an output module, where:
[0114] The appearance feature model is used to extract the feature information of the target picture block and the image of the frame to be processed, and determine the first candidate position information of the target object in the image of the frame to be processed according to the extracted feature information.
[0115] Specifically, the appearance feature module is specifically a siamese network, and the specific structure is as above Figure 5 As shown, in the first stage, the first feature information of the target picture block and the second feature information of the image of the frame to be processed can be respectively extracted through the appearance feature extraction module. The first feature information is convolved on the second feature information as a convolution kernel to determine the position information of multiple candidate boxes where the target object is located in the image of the frame to be processed, and at the same time obtain the probability information of whether the image in each candidate box belongs to the target object. Furthermore, the position information of the k candidate boxes with the highest probability information, that is is passed into the second stage, and k can be 48.
[0116] In the second stage, specifically, the prediction module can intercept each candidate box c through the region of interest align (RoIAlign) operation i to obtain the feature information of each candidate box from the features of the fourth and sixth layers of the image within the candidate box and match the feature information of each candidate box with the feature information of the target image patch The matching network can specifically be a two-layer convolutional fully connected layer to obtain the matching score of each candidate box, and then determine one of the candidate boxes as the box where the target object is located based on the matching score. Among them, when matching the feature information of each candidate box with the feature information of the target image patch multi-granularity feature matching can be mainly adopted, that is, global feature information matching and local feature information matching, to obtain the corresponding matching scores respectively, and fuse and aggregate the matching scores to obtain the matching score corresponding to any candidate box. In this way, the interference of the complex background in the image to be processed can be resisted, making the finally predicted first candidate position information more accurate.
[0117] The adaptive update module is used to adaptively determine the update rate, and based on the determined update rate and the position information of the target object in the previous frame image of the image to be processed determined by the appearance feature model, update the first candidate position information determined by the appearance feature model to obtain the updated candidate position information and transmit it to the output module.
[0118] Among them, in the visualization results of practical applications, it can be found that the update rate adaptively selected by the adaptive update module is basically related to the movement speed of the target object. When the movement speed is large, the position change of the target object between the front and back two frame images is large, then the adaptive update module considers a larger ratio of the position information of the target object in the current image to be processed when updating the above first candidate position information, and vice versa.
[0119] The motion feature model is used to determine the second candidate position information of the target object in the image to be processed according to the first optical flow information of the reference frame image, the reference position information of the target object in the reference frame image, and the second optical flow information of the image to be processed (or the preprocessed image to be processed). Then, in a specific embodiment, the target positioning system may further include a preprocessing module for preprocessing the image to be processed. For example, the image to be processed is transformed into the coordinate system of the reference frame image, which can eliminate the interference caused by the movement of the shooting device between the image to be processed and the reference frame image.
[0120] The first optical flow information and the second optical flow information can be calculated using the Gunnar Farneback algorithm to obtain a dense optical flow. After the first optical flow information and the reference position information enter the long short-term memory network (Long Short-Term Memory, LSTM), it includes an encoder and a decoder, and after the second optical flow information enters the decoder of the LSTM, the second candidate position information of the target object in the frame image to be processed is output.
[0121] The output module is used to determine the position information of the target object in the frame image to be processed according to the updated candidate position information obtained by the above-mentioned adaptive update module and the second candidate position information obtained by the motion feature model.
[0122] The target positioning method of this embodiment mainly includes the following two parts:
[0123] (a) Train to obtain appearance feature model and motion feature model.
[0124] On the one hand, when training the appearance feature model, the above Figure 4 The method shown is used for training, wherein, when determining the first training sample, four public data sets can be selected, including: video data sets VID and YoutubeBB, and detection data sets DET and COCO.
[0125] Specifically, a video clip can be randomly selected from the video data set, and then a frame of image can be randomly extracted from the video clip as the reference frame image where the sample object image block is located, so as to obtain the sample object image block, and the sample object image block can be preprocessed by adding enhancement methods such as blur flipping. Afterwards, when selecting the sample image, since each frame of the video data set VID is annotated, when selecting the sample image, a frame of image can be randomly extracted within the range of 100 frames before and after a reference frame image as the sample image; and the images in the video data set YoutubeBB are annotated with one frame of image per second, so when selecting the sample image, a frame of image can be extracted from multiple frames (such as 3 frames) of images with annotation information before and after the reference frame image as the sample image. Furthermore, in order to avoid the center preference caused by too much white filling on the edge due to the deep network, after selecting the sample image, random translation and other methods can be added to preprocess the sample image, that is, the sample object is randomly moved a certain distance from the center of the sample image to enhance the learning of the appearance feature model.
[0126] In the process of training the appearance feature model, a detection set can be selected to detect the trained appearance feature model. The detection set can be selected from the detection data set. Specifically, pictures of the same or different categories can be directly selected to form positive and negative picture pairs to input the trained appearance feature model.
[0127] During the process of adjusting the parameter values in the initial appearance feature model, the Stochastic Gradient Descent (SGD) optimizer can be used for adjustment, and the learning rate is 0.0001. Moreover, in order to make the network converge more stably, the parameter values in the backbone network (i.e., the above-mentioned appearance feature extraction module) can be frozen in the first 10 epochs, that is, the parameter values in the backbone network are not adjusted, and only the network in the second stage (i.e., the prediction module) and the classification branch and regression branch in the first stage (i.e., the classification module and the location regression module) are trained. The parameter values in the backbone network start to be trained in the 11th epoch.
[0128] On the other hand, when training the motion feature model, the method shown above Figure 6 can be used for training. Among them, when selecting the second training sample, the dataset VID can be used. For each video clip in the dataset VID, a small video clip of 7 consecutive frames of images is randomly selected. The first 6 frames of images are observation frames, and the 7th frame of image is the prediction frame to predict the position information of the sample object in the 7th frame of image. Moreover, when training the motion feature model, the Adaptive Moment Estimation (ADAM) optimizer can be used to adjust the parameter values in the initial motion feature model, and the learning rate can be 0.001.
[0129] (2) Locate the target object in any video.
[0130] For any video, when locating the target object in the video, a certain frame of the video (usually the first frame) can be used as the reference frame image, and an image cut is made from the reference frame image to obtain a target picture block including the target object. Then, the target picture block and other frame images in the video except the reference frame image can be input into the above-mentioned target location system. In this way, the target location system can use the other frame images as the to-be-processed frame images to obtain the position information of the target object.
[0131] Among them, in the process of locating the target object in the video, the target positioning system can first use the first frame image in the video as the reference frame image, and determine the position information of the box where the target object is located in the first frame image. For the subsequent 5 frames of images of the first frame image, the prediction results are obtained by using the appearance feature model plus the adaptive update module. Starting from the 7th frame image, the motion feature model is activated, that is, the prediction results are obtained by using the appearance feature model, the motion feature model and the adaptive update model. The observable number of frames is taken as 6, that is, the position information of the target object in any subsequent frame image is inferred through the optical flow information and feature information of the first 6 frame images. The structures of the motion feature model and the appearance feature model also extract the features of the candidate boxes in the form of ROI Align in the backbone network (that is, the above-mentioned appearance feature extraction module) and calculate the cosine distance with the target feature respectively, and take the one with the smaller cosine distance as the final prediction output.
[0132] In the specific practice process, on the one hand, the existing baseline model is used to locate the target in the video, and after adding some specific functions (such as adding multi-granularity feature matching, etc.) to the baseline model and then locating the target in the video, the evaluation indexes, namely accuracy, robustness and average expected overlap rate, are calculated respectively, as shown in Table 1 below. Among them, the accuracy rate is the average intersection over union, the robustness is the total number of frames in which the locator fails to locate, and the average expected overlap rate is the average of the intersection over union obtained with different frames as the maximum frame without re-initialization in a video. It can be seen that after adding the motion feature model to the baseline model, the performance is greatly improved, that is, the accuracy rate is improved, the robustness is reduced, and the average expected overlap rate is also improved.
[0133]
[0134] Table 1
[0135] On the other hand, after using the model of the existing locator to locate the target in the video and using the locator in the embodiment of the present invention, that is, M-SPM, to locate the target in the video, the evaluation indexes, namely success rate and precision (or normalized precision), are calculated respectively, as shown in Table 2 below. Among them, the success rate refers to the average of the proportions of successful frames in the evaluated video at each threshold from 0 to 1 interval, generally with an interval of 0.05, and the average of the proportions of successful frames at 20 thresholds is calculated; the precision, also known as the accuracy rate, refers to the proportion of frames in which the Euclidean distance between the position determined by the locator model and the position marked in the training sample is less than the specified distance threshold. Here, the threshold generally takes values from 0 to 51 with an interval of 1; the normalized precision is to normalize the calculated precision, and the normalized precision mainly takes into account that the calculation of the original precision index is sensitive to the image resolution and the size of the box, so normalization is carried out.
[0136] Among them, the models of existing locators may include: SINT, ECO, DSiam, VITAL, StructSiam, Siam-BM, DaSiamRPN, ATOM, SPM, SiamRPN++, DiMP, SiamBAN, MAML, and ROAM. When training the models of each locator, training samples are respectively selected from the datasets OTB100 and LaSOT. It can be seen that no matter which dataset is used to train the model M-SPM of the locator in the embodiments of the present invention, and after target localization is performed using M-SPM, the success rate and accuracy are greatly improved, and the effect is better when the dataset OTB100 is used to train the model M-SPM.
[0137]
[0138] Table 2
[0139] On the other hand, after target localization of a video is performed using the models of existing locators and the locator in the embodiments of the present invention, i.e., M-SPM, the evaluation indexes, namely accuracy, robustness, and average expected overlap rate, are respectively calculated as shown in Table 3 below. Among them, the models of existing locators may include: LADCF, MFT, SiamRPN, SiamDW, SPM, ATOM, SiamRPN++, SiamMask, SiamBAN, SiamR-CNN, and MAML. When training the models of each locator, training samples are respectively selected from the datasets VOT2018 and VOT2019. It can be seen that no matter which dataset is used to train the model M-SPM of the locator in the embodiments of the present invention, and after target localization is performed using M-SPM, the accuracy rate and average expected overlap rate are greatly improved, and the robustness is reduced. The effect is better when the dataset VOT2018 is used to train the model M-SPM.
[0140]
[0141]
[0142] Table 3
[0143] On the other hand, after target localization of a video is performed using the models of existing locators and the locator in the embodiments of the present invention, i.e., M-SPM, the evaluation indexes, namely success rate and accuracy, are respectively calculated as shown in Table 4 below. Among them, the models of existing locators may include: SiamRPN, SiamMask, and DROL. When training the models of each locator, training samples are selected from the self-constructed dataset. It can be seen that after target localization is performed using the model M-SPM of the locator in the embodiments of the present invention, the impulse power and accuracy are greatly improved.
[0144] Locator Success rate (↑) Accuracy (↑) SiamRPN 0.616 0.406 SiamMask 0.641 0.441 DROL 0.643 0.441 M-SPM 0.649 0.469
[0145] Table 4
[0146] The following uses another specific application example to illustrate the target positioning method in the present invention. The target positioning system in the embodiment of the present invention is mainly a distributed system 100, and this distributed system may include a client 300 and multiple nodes 200 (any form of computing device accessing the network, such as a server, a user terminal), and the client 300 and the nodes 200 are connected in the form of network communication.
[0147] Taking the distributed system as a blockchain system as an example, see Figure 8 It is an optional structural schematic diagram of the distributed system 100 provided by the embodiment of the present invention applied to the blockchain system, formed by multiple nodes 200 (any form of computing device accessing the network, such as a server, a user terminal) and a client 300. A peer-to-peer (P2P, Peer To Peer) network is formed among the nodes, and the P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In the distributed system, any machine such as a server or a terminal can join and become a node, and the node includes a hardware layer, an intermediate layer, an operating system layer, and an application layer.
[0148] See Figure 8 The functions of each node in the shown blockchain system involve the following functions:
[0149] 1) Routing, a basic function of the node, used to support communication between nodes.
[0150] In addition to the routing function, the node can also have the following functions:
[0151] 2) Application, used to be deployed in the blockchain, to implement specific services according to actual business requirements, record the data related to the implemented functions to form record data, carry a digital signature in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system. When other nodes verify the source and integrity of the record data successfully, the record data is added to the temporary block.
[0152] For example, the service implemented by the application also includes the code for implementing the target positioning function, and this target positioning function mainly includes:
[0153] Obtain the first feature information of the target picture block included in the reference frame image in the video to be processed, and obtain the second feature information of the frame image to be processed in the video to be processed; wherein, the target picture block is an image including a target object; determine the first candidate position information of the target object in the frame image to be processed according to the first feature information and the second feature information; obtain the first optical flow information of the reference frame image, obtain the second optical flow information according to the frame image to be processed, and obtain the reference position information of the target picture block in the reference frame image; determine the second candidate position information of the target object in the frame image to be processed according to the first optical flow information, the second optical flow information and the reference position information; determine the position information of the target object in the frame image to be processed according to the first candidate position information and the second candidate position information.
[0154] 3) The blockchain includes a series of blocks (Blocks) that are sequentially connected in the order of generation time. Once a new block is added to the blockchain, it will not be removed again. The block records the record data submitted by nodes in the blockchain system.
[0155] See Figure 9 This is an optional schematic diagram of the block structure provided by the embodiment of the present invention. Each block includes the hash value of the transaction record stored in this block (the hash value of this block), and the hash value of the previous block. Each block is connected through the hash value to form a blockchain. In addition, the block may also include information such as the timestamp when the block is generated. The blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains relevant information for verifying the validity of its information (anti-counterfeiting) and generating the next block.
[0156] The embodiment of the present invention also provides a target positioning system, and its structural schematic diagram is as Figure 10 shown, and specifically may include:
[0157] The feature acquisition unit 10 is used to obtain the first feature information of the target picture block included in the reference frame image in the video to be processed, and obtain the second feature information of the frame image to be processed in the video to be processed; wherein, the target picture block is an image including a target object.
[0158] The first candidate unit 11 is used to determine the first candidate position information of the target object in the frame image to be processed according to the first feature information and the second feature information obtained by the feature acquisition unit 10.
[0159] The first candidate unit 11 is specifically configured to determine multiple candidate boxes where the target object is located in the to-be-processed frame image according to the first feature information and the second feature information; match the feature information of the images within each candidate box with the first feature information of the target picture block in the reference frame image to obtain the matching degree corresponding to each candidate box; and determine the first candidate position information of the target object in the to-be-processed frame image according to the matching degree corresponding to each candidate box.
[0160] Among them, when the first candidate unit 11 matches the feature information of the images within each candidate box with the first feature information of the target picture block in the reference frame image to obtain the matching degree corresponding to each candidate box, it is specifically configured to match the global feature information of the images within each candidate box with the first feature information to obtain a first sub-matching degree; divide the images within each candidate box and the target picture block into multiple regions according to the same strategy, and respectively match the local feature information of each region in the images within each candidate box with the local feature information of the corresponding region in the target picture block to obtain a second sub-matching degree; and determine the matching degree corresponding to each candidate box according to the first sub-matching degree and the second sub-matching degree.
[0161] The optical flow information unit 12 is configured to obtain the first optical flow information of the reference frame image, obtain the second optical flow information according to the to-be-processed frame image, and obtain the reference position information of the target picture block in the reference frame image.
[0162] The second candidate unit 13 is configured to determine the second candidate position information of the target object in the to-be-processed frame image according to the first optical flow information, the second optical flow information, and the reference position information obtained by the optical flow information unit 12.
[0163] The position determination unit 14 is configured to determine the position information of the target object in the to-be-processed frame image according to the first candidate position information determined by the first candidate unit 11 and the second candidate position information determined by the second candidate unit 13.
[0164] The position determination unit 14 is specifically configured to set multiple update rates for the first candidate position information; calculate the updated candidate position information under each update rate according to the update rate, the first candidate position information, and the position information of the object target in the previous frame image of the to-be-processed frame image; calculate the distance between the sub-image corresponding to each updated candidate position information in the to-be-processed frame image and the target picture block, and select the updated candidate position information corresponding to the sub-image with the smallest distance; and determine the position information of the target object in the to-be-processed frame image according to the selected updated candidate position information and the second candidate position information.
[0165] Furthermore, the target positioning system of this embodiment may further include:
[0166] A training unit 15 is configured to determine an initial appearance feature model, where the initial appearance feature model includes an appearance feature extraction module, a position regression module, a prediction module, and a classification module; determine a first training sample, where the first training sample includes a plurality of first sample image groups, each first sample image group includes a sample object picture block, at least one sample image, position annotation information of a plurality of sample frames in the sample image, and type annotation indicating whether each sample frame belongs to the frame where the sample object is located; respectively obtain feature information of the sample object picture block and the sample image through the appearance feature extraction module, the position regression module determines position information of a candidate frame where the sample object is located in the sample image according to the feature information of the sample object picture block and the sample image, the prediction module selects the position information of a certain candidate frame as the position information of the sample object in the sample image, and the classification module determines whether the image within the candidate frame where the sample object is located determined by the position regression module belongs to the type information of the sample object; adjust the initial appearance feature model according to the position information obtained by the prediction module and the position annotation information in the first training sample, and the type information determined by the classification module and the type annotation in the first training sample, so as to obtain a final appearance feature model. In this way, the above-mentioned feature acquisition unit 10 and the first candidate unit 11 can determine the first candidate position information by using the appearance feature model trained by the training unit 15.
[0167] The training unit 15 is further configured to stop adjusting the parameter value when the number of adjustments to the parameter value is equal to a preset number, or when the difference between the currently adjusted parameter value and the parameter value adjusted last time is less than a threshold.
[0168] Further, the target positioning system of this embodiment may further include:
[0169] A preprocessing unit 16 is configured to convert the to-be-processed frame image into the coordinate system of the reference frame image to obtain a converted processed frame image; when the optical flow information unit 12 obtains second optical flow information according to the to-be-processed frame image, it is specifically configured to obtain the second optical flow information of the converted processed frame image.
[0170] Wherein, when the preprocessing unit 16 converts the to-be-processed frame image into the coordinate system of the reference frame image to obtain a converted processed frame image, it is specifically configured to calculate a homography matrix between the reference frame image and the to-be-processed frame image according to the information of the background region feature points in the reference frame image and the information of the background region feature points in the to-be-processed frame image; convert the to-be-processed frame image into the coordinate system of the reference frame image according to the homography matrix.
[0171] In this embodiment, the target positioning system can use the motion features of the image to offset, as much as possible, the interference of the static appearance features of the image in the process of target object positioning, such as the occlusion of the target object and background interference, etc., so that the final positioning of the target object is more accurate.
[0172] An embodiment of the present invention further provides a terminal device, and its structural schematic diagram is as Figure 11 shown. The terminal device may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 20 (for example, one or more processors) and a memory 21, and one or more storage media 22 for storing application programs 221 or data 222 (for example, one or more mass storage devices). Among them, the memory 21 and the storage media 22 may be transient storage or persistent storage. The program stored in the storage media 22 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the terminal device. Further, the central processing unit 20 may be configured to communicate with the storage media 22 and execute a series of instruction operations in the storage media 22 on the terminal device.
[0173] Specifically, the application program 221 stored in the storage media 22 includes an application program for target positioning, and this program may include the feature acquisition unit 10, the first candidate unit 11, the optical flow information unit 12, the second candidate unit 13, the position determination unit 14, the training unit 15, and the preprocessing unit 16 in the above target positioning system, which will not be elaborated here. Further, the central processing unit 20 may be configured to communicate with the storage media 22 and execute a series of operations corresponding to the application program for target positioning stored in the storage media 22 on the terminal device.
[0174] The terminal device may further include one or more power supplies 23, one or more wired or wireless network interfaces 24, one or more input / output interfaces 25, and / or one or more operating systems 223, such as WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0175] The steps performed by the target positioning system in the above method embodiment may be based on the Figure 11 structure of the shown terminal device.
[0176] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium, which stores multiple computer programs, and the computer programs are adapted to be loaded and executed by a processor to perform the target positioning method performed by the above target positioning system.
[0177] Another aspect of the embodiments of the present invention further provides a terminal device, including a processor and a memory;
[0178] The memory is used to store a plurality of computer programs, and the computer programs are used to be loaded and executed by the processor to perform the target positioning method executed by the above-mentioned target positioning system; the processor is used to implement each of the plurality of computer programs.
[0179] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disc, etc.
[0180] The above has introduced in detail a target positioning method, system, storage medium and terminal device provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A target localization method, characterized in that, it includes: obtaining first feature information of a target picture block included in a reference frame image in a video to be processed, and obtaining second feature information of a frame image to be processed in the video to be processed; wherein, the target picture block is an image containing a target object; determining first candidate position information of the target object in the frame image to be processed according to the first feature information and the second feature information; obtaining first optical flow information of the reference frame image, obtaining second optical flow information according to the frame image to be processed, and obtaining reference position information of the target picture block in the reference frame image; determining second candidate position information of the target object in the frame image to be processed according to the first optical flow information, the second optical flow information and the reference position information; determining position information of the target object in the frame image to be processed according to the first candidate position information and the second candidate position information.
2. The method according to claim 1, characterized in that, the determining the first candidate position information of the target object in the frame image to be processed according to the first feature information and the second feature information specifically includes: determining a plurality of candidate boxes where the target object is located in the frame image to be processed according to the first feature information and the second feature information; matching the feature information of the images within each candidate box with the first feature information of the target picture block in the reference frame image to respectively obtain the matching degrees corresponding to the candidate boxes; determining the first candidate position information of the target object in the frame image to be processed according to the matching degrees corresponding to the candidate boxes.
3. The method according to claim 2, characterized in that, the matching the feature information of the images within each candidate box with the first feature information of the target picture block in the reference frame image to respectively obtain the matching degrees corresponding to the candidate boxes specifically includes: matching the global feature information of the images within each candidate box with the first feature information to obtain a first sub - matching degree; dividing the images within each candidate box and the target picture block into multiple regions according to the same strategy, and respectively matching the local feature information of each region in the images within each candidate box with the local feature information of the corresponding region in the target picture block to obtain a second sub - matching degree; determining the matching degrees corresponding to the candidate boxes according to the first sub - matching degree and the second sub - matching degree.
4. The method according to claim 1, characterized in that, the method further includes: determining an initial appearance feature model, where the initial appearance feature model includes an appearance feature extraction module, a position regression module, a prediction module and a classification module; determining a first training sample, where the first training sample includes a plurality of first sample image groups, and each first sample image group includes a sample object picture block, at least one sample image, position annotation information of multiple sample boxes in the sample image, and type annotation indicating whether each sample box belongs to the box where the sample object is located; The feature extraction module for appearance respectively obtains the feature information of the sample object picture block and the sample image. The position regression module determines the position information of the candidate box where the sample object is located in the sample image according to the feature information of the sample object picture block and the sample image. The prediction module selects the position information of a certain candidate box as the position information of the sample object in the sample image. The classification module determines whether the image within the candidate box where the sample object is located determined by the position regression module belongs to the type information of the sample object; According to the position information obtained by the prediction module and the position annotation information in the first training sample, and the type information determined by the classification module and the type annotation in the first training sample, adjust the initial appearance feature model to obtain the final appearance feature model.
5. The method according to claim 4, characterized in that, when the number of times of adjusting the parameter value in the initial appearance feature model is equal to the preset number of times, or if the difference between the currently adjusted parameter value and the parameter value adjusted last time is less than a threshold value, then stop adjusting the parameter value.
6. The method according to any one of claims 1 to 5, characterized in that, before obtaining the second optical flow information according to the to-be-processed frame image, the method further includes: transforming the to-be-processed frame image into the coordinate system of the reference frame image to obtain the transformed processed frame image; then obtaining the second optical flow information according to the to-be-processed frame image specifically includes: obtaining the second optical flow information of the transformed processed frame image.
7. The method according to claim 6, characterized in that, transforming the to-be-processed frame image into the coordinate system of the reference frame image to obtain the transformed processed frame image specifically includes: calculating the homography matrix between the reference frame image and the to-be-processed frame image according to the information of the background region feature points in the reference frame image and the information of the background region feature points in the to-be-processed frame image; transforming the to-be-processed frame image into the coordinate system of the reference frame image according to the homography matrix.
8. The method according to any one of claims 1 to 5, characterized in that, determining the position information of the target object in the to-be-processed frame image according to the first candidate position information and the second candidate position information specifically includes: setting a plurality of update rates for the first candidate position information; respectively calculating the updated candidate position information at each update rate according to the update rate, the first candidate position information and the position information of the object target in the previous frame image of the to-be-processed frame image; respectively calculating the distance between the sub-image corresponding to each updated candidate position information in the to-be-processed frame image and the target picture block, and selecting the updated candidate position information corresponding to the sub-image with the smallest distance; determining the position information of the target object in the to-be-processed frame image according to the selected updated candidate position information and the second candidate position information.
9. A target positioning system, characterized in that, comprising: A feature acquisition unit, configured to acquire first feature information of a target picture block included in a reference frame image in a video to be processed, and acquire second feature information of a frame image to be processed in the video to be processed; wherein, the target picture block is an image including a target object; A first candidate unit, configured to determine first candidate position information of the target object in the frame image to be processed according to the first feature information and the second feature information; An optical flow information unit, configured to acquire first optical flow information of the reference frame image, acquire second optical flow information according to the frame image to be processed, and acquire reference position information of the target picture block in the reference frame image; A second candidate unit, configured to determine second candidate position information of the target object in the frame image to be processed according to the first optical flow information, the second optical flow information and the reference position information; A position determination unit, configured to determine position information of the target object in the frame image to be processed according to the first candidate position information and the second candidate position information.
10. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a plurality of computer programs, and the computer programs are adapted to be loaded and executed by a processor to perform the target positioning method according to any one of claims 1 to 4.
11. A terminal device, characterized in that, it includes a processor and a memory; the memory is used to store a plurality of computer programs, and the computer programs are used to be loaded and executed by the processor to perform the target positioning method according to any one of claims 1 to 4; the processor is used to implement each of the computer programs in the plurality of computer programs.
Citation Information
Patent Citations
Target tracking method
CN109584269A
Target object tracking method, device and equipment, and computer readable storage medium
CN112132866A