Method, apparatus, electronic device and storage medium for positioning a tracked object in a video
By combining image and language features with shared backbone network, fusion features are generated to accurately locate tracking objects in videos, solving the problem of inaccurate positioning in the prior art and realizing accurate positioning of tracking objects in dynamic videos.
Patent Information
- Application Number
- CN202210673113.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-06-14
AI Technical Summary
Existing positioning models cannot accurately locate tracking objects in videos, especially in complex and dynamic scenarios, which makes it impossible for electronic devices to effectively locate tracking objects.
The preset shared backbone network is used to combine the image features and language features of video frame images, and to aggregate the network and shared images and language backbone networks through frame-intensive features to generate fusion image features and fusion language features, thereby accurately positioning the tracking object.
Accurate positioning of tracked objects in video is achieved, the positioning inaccurate problem caused by the limitations of the positioning model in the prior art is solved, and the positioning accuracy of electronic devices in dynamic video is improved.
Smart Images

Figure CN115222768B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly to a method, apparatus, electronic device and storage medium for locating a tracking object in a video. Background Art
[0002] With the development of science and technology, image recognition technology has become more and more mature. When an electronic device locates a tracking object in an image, it usually uses referential expression understanding, that is, locates a static tracking object according to a natural language description statement. However, this method cannot locate a complex and dynamic tracking object, that is, cannot locate a tracking object in a video.
[0003] Existing methods for locating a tracking object in a video may include: the electronic device locates the tracking object based on a video-natural language referential expression understanding model of a target tracking framework, or the electronic device locates the tracking object based on a video-natural language referential expression understanding model of one-stage object detection. However, due to the corresponding limitations of the above two models, the electronic device cannot accurately locate the tracking object in the video. Summary of the Invention
[0004] The present invention provides a method, apparatus, electronic device and storage medium for locating a tracking object in a video, so as to solve the defect that in the prior art, due to the corresponding limitations of the existing location model, the electronic device cannot accurately locate the tracking object in the video based on the existing location model, and realize that the electronic device effectively and accurately locates the tracking object in the video to be processed based on a preset shared backbone network, in combination with the image features and language features of the video frame image.
[0005] The present invention provides a method for locating a tracking object in a video, including:
[0006] In the process of locating the tracking object in the current frame image of the video to be processed, obtaining the current image feature and the current language feature corresponding to the current frame image;
[0007] According to the current image feature and the current language feature, based on the preset shared backbone network, obtaining the fused image feature and the fused language feature corresponding to the current frame image;
[0008] According to the fused image feature and the fused language feature, determining the location result of the tracking object.
[0009] A method for locating a tracked object in a video provided by the present invention. Obtaining the current image features corresponding to the current frame image includes: obtaining the first image features corresponding to the key frame image in the video to be processed, where the key frame image is any one of each frame image in the video to be processed; obtaining the second image features corresponding to the adjacent frame images of the key frame image; and based on the first image features and the second image features, obtaining the current image features corresponding to the current frame image based on a preset frame dense feature aggregation network.
[0010] A method for locating a tracked object in a video provided by the present invention. Based on the first image features and the second image features, obtaining the current image features corresponding to the current frame image based on a preset frame dense feature aggregation network includes: based on the preset frame dense feature aggregation network, obtaining a normalized weight matrix according to the first image features and the second image features; and determining the current image features corresponding to the current frame image according to the first image features and the normalized weight matrix.
[0011] A method for locating a tracked object in a video provided by the present invention. Based on the current image features and the current language features, obtaining the fused image features and the fused language features corresponding to the current frame image based on a preset shared backbone network includes: obtaining visual vector features based on the current image features and the preset shared image backbone network; obtaining a first similarity matrix based on the current language features and the visual vector features and the preset shared image backbone network; obtaining a second similarity matrix based on the current language features and the visual vector features and a preset shared language backbone network; determining the fused image features corresponding to the current frame image according to the current language features and the first similarity matrix; and determining the fused language features corresponding to the current frame image according to the visual feature vector and the second similarity matrix.
[0012] A method for locating a tracked object in a video provided by the present invention. After obtaining the first similarity matrix based on the current language features and the visual vector features and the preset shared image backbone network, the method further includes: obtaining the candidate position of the tracked object in the current image features; and adding a first constraint function to the first similarity matrix according to the candidate position.
[0013] A method for locating a tracked object in a video provided by the present invention. Determining the positioning result of the tracked object according to the fused image features and the fused language features includes: determining the language expression sentence features according to the fused language features; determining a first language conditional vector and a second speech conditional vector according to the speech expression sentence features; and determining the positioning result of the tracked object according to the fused image features, the first language conditional vector, and the second speech conditional vector.
[0014] A method for locating a tracked object in a video provided by the present invention further includes: obtaining a first constraint function corresponding to the first similarity matrix and a second constraint function corresponding to the second similarity matrix; determining a positioning regression loss function corresponding to the preset shared backbone network according to the first constraint function and the second constraint function; and determining a total loss function corresponding to the preset shared backbone network according to the positioning regression loss function.
[0015] The present invention also provides a positioning device, including:
[0016] An obtaining module, configured to obtain a current image feature and a current language feature corresponding to the current frame image during the process of locating a tracked object in a video to be processed.
[0017] A determining module, configured to obtain a fused image feature and a fused language feature corresponding to the current frame image based on a preset shared backbone network according to the current image feature and the current language feature; and determine a positioning result of the tracked object according to the fused image feature and the fused language feature.
[0018] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the method for locating a tracked object in a video as described in any one of the above when executing the program.
[0019] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and the computer program implements the method for locating a tracked object in a video as described in any one of the above when executed by a processor.
[0020] The present invention also provides a computer program product, including a computer program, where the computer program implements the method for locating a tracked object in a video as described in any one of the above when executed by a processor.
[0021] The video object tracking positioning method, device, electronic device, and storage medium provided by the present invention. The method may include: during the process of positioning the tracking object in the current frame image of the video to be processed, obtaining the current image feature and the current language feature corresponding to the current frame image; then, based on the preset shared backbone network according to the current image feature and the current language feature, obtaining the relatively accurate fused image feature and the fused language feature corresponding to the current frame image; finally, accurately determining the positioning result of the tracking object according to the fused image feature and the fused language feature, so as to achieve accurate positioning of the tracking object in the video to be processed. This method is used to solve the defect that in the prior art, due to the limitations of the existing positioning model, the electronic device cannot accurately position the tracking object in the video based on the existing positioning model, and realizes that the electronic device combines the image feature and the language feature of the video frame image based on the preset shared backbone network to effectively and accurately position the tracking object in the video to be processed. Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0023] Figure 1 It is one of the flowcharts of the video object tracking positioning method provided by the present invention;
[0024] Figure 2 It is another flowchart of the video object tracking positioning method provided by the present invention;
[0025] Figure 3 It is the third flowchart of the video object tracking positioning method provided by the present invention;
[0026] Figure 4 It is the structural schematic diagram of the positioning device provided by the present invention;
[0027] Figure 5 It is the structural schematic diagram of the electronic device provided by the present invention. Detailed Embodiments
[0028] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] In the prior art, when an electronic device locates a tracking object based on a video-natural language referring expression understanding model of a target tracking framework, since the performance of the tracking framework depends on the quality of the tracking template selected by the electronic device, the electronic device usually initializes the tracking template with the tracking target area corresponding to the first frame image in the video to be processed. However, when there is no labeled data to assist the electronic device in selecting the tracking template, if the electronic device only uses an image referring expression understanding model to locate the tracking object in the first frame image, the positioning result will be inaccurate, and further, the quality of the tracking template selected by the electronic device will be poor. That is to say, the video-natural language referring expression understanding model of the electronic device based on the target tracking framework cannot accurately locate the tracking object.
[0030] When an electronic device locates a tracking object based on a video-natural language referring expression understanding model of one-stage object detection, the electronic device only uses the adjacent frame image of the key frame image of the video for image feature collaborative learning. Although the electronic device establishes a connection for the image information between video frames, since the time sequence of two adjacent video frame images is relatively close, the image feature information corresponding to the two video frame images has strong similarity, resulting in the electronic device being unable to fully establish the image feature relationship between video frames, and thus unable to accurately obtain the changes in the motion, appearance, etc. of the dynamic tracking object in the video frames. That is, the video-natural language referring expression understanding model of the electronic device based on one-stage object detection cannot accurately locate the tracking object.
[0031] It should be noted that the electronic device involved in the embodiments of the present invention may include, but is not limited to, at least one of the following: a computer terminal, a mobile terminal, a wearable device, etc.
[0032] The execution subject of the embodiments of the present invention may be a positioning device or an electronic device. Hereinafter, the embodiments of the present invention will be further described by taking the electronic device as an example.
[0033] As shown in 1, the flowchart of the method for locating a tracking object in a video provided by the present invention may include:
[0034] 101. In the process of locating the tracking object in the current frame image of the video to be processed, the current image feature and the current language feature corresponding to the current frame image are obtained.
[0035] Among them, the video to be processed generally refers to various technologies that capture, record, process, store, transmit, and reproduce a series of static images in the form of electrical signals. That is, the video to be processed may include multiple frames of images.
[0036] The current frame image refers to the frame image corresponding to the current moment in the video to be processed.
[0037] The tracking object refers to a dynamic referent that the electronic device needs to locate, and there are changes in motion and / or appearance, etc. of this referent in the video to be processed.
[0038] The current image feature refers to the pixel feature of the tracking object in the current frame image.
[0039] The current language feature refers to the language expression feature of the tracking object in the current frame image.
[0040] In some embodiments, the electronic device can first obtain a key frame image and the first image feature corresponding to the key frame image from multiple frames of images in the video to be processed, and the key frame image is any frame image in the multiple frames of images; then, the electronic device obtains the adjacent frame image of the key frame image and the second image feature corresponding to the adjacent frame image; then, the electronic device can obtain the current image feature corresponding to the current frame image according to the first image feature and the second image feature and a preset frame dense feature aggregation network.
[0041] Among them, the preset frame dense feature aggregation network is used to adaptively generate the weighted value of the corresponding position points of the first image feature and the second image feature according to the first image feature and the second image feature; then, based on the weighted value, establish the image feature connection between the key frame image and the adjacent frame image; then, perform frame dense weighted aggregation on the video frame images within the adjacent time series corresponding to the key frame image to obtain a more accurate current image feature corresponding to the current frame image.
[0042] In some embodiments, the preset frame dense feature aggregation network can effectively avoid the problem of inaccurate positioning of the referent in the front and rear frame images in the video-natural language referential expression understanding model of one-stage object detection in the prior art.
[0043] In some embodiments, the electronic device obtains the current language feature corresponding to the current frame image based on a preset shared language backbone network.
[0044] Among them, the preset shared language backbone network is used to extract the language expression feature in the current frame image, and the speech expression feature may include description statement features.
[0045] 102. Based on the current image features and current language features, and based on a preset shared backbone network, obtain the fused image features and fused language features corresponding to the current frame image.
[0046] Among them, the preset shared backbone network may include: a preset shared image backbone network and a preset shared language backbone network.
[0047] The preset shared image backbone network is used to determine the fused image features corresponding to the current frame image; the preset shared language backbone network is further used to extract the fused language features corresponding to the current frame image.
[0048] In some embodiments, the preset shared backbone network is a video referential expression understanding network based on multi-stage image-natural language cross-generation fusion. This preset shared backbone network adopts a one-stage object detection framework, which can effectively avoid the problem of selecting a tracking template in the existing video-natural language referential expression understanding model based on a target tracking framework.
[0049] After the electronic device obtains the current image features and current language features, since the current image features and current language features cannot accurately locate the tracking object, the electronic device needs to fuse the current image features and the current language features according to different feature methods to obtain the corresponding fused image features and fused language features. These fused image features and fused language features are relatively accurate, so that the electronic device can accurately locate the tracking object subsequently.
[0050] In some embodiments, the electronic device, based on the language-image generation branch in the preset shared backbone network, obtains the fused image features corresponding to the current frame image according to the current language features; the electronic device, based on the image-language generation branch in the preset shared backbone network, obtains the fused language features corresponding to the current frame image according to the current image features. Among them, the generation timing of the fused image features and the generation timing of the fused language features are not limited.
[0051] The electronic device can, in a cross-modal generation manner, supplement and improve the image information of the current frame image, and at the same time, supplement and improve the language information of the current frame image, so as to obtain relatively accurate fused image features and fused language features.
[0052] 103. Determine the positioning result of the tracking object according to the fused image features and fused language features.
[0053] The electronic device, based on the relatively accurate fused image features and fused language features, accurately performs referential expression understanding on the tracking object, so as to accurately obtain the positioning result of the tracking object.
[0054] Among them, the above-mentioned referential expression understanding refers to locating the tracking object in all frame images of the video to be processed according to the natural language description statement, and using the inter-frame information of the video to solve the problem of dynamic changes of the tracking object.
[0055] Optionally, the positioning result may include the positioning box prediction result.
[0056] In some embodiments, the electronic device may first obtain two corresponding language condition vectors under two different language conditions according to the fused language features; then, the electronic device may accurately determine the positioning box prediction result of the tracking object according to the fused image features and these two language condition vectors.
[0057] Optionally, after step 103, the method may further include: the electronic device outputs the positioning result to ensure that the user can intuitively obtain the positioning result.
[0058] In the embodiments of the present invention, in the process of positioning the tracking object in the current frame image of the video to be processed, the current image feature and the current language feature corresponding to the current frame image are obtained; then, according to the current image feature and the current language feature, based on a preset shared backbone network, the relatively accurate fused image feature and fused language feature corresponding to the current frame image can be obtained; finally, according to the fused image feature and the fused language feature, the positioning result of the tracking object is accurately determined, so as to realize the accurate positioning of the tracking object in the video to be processed. This method is used to solve the defect that in the prior art, due to the limitations of the existing positioning model, the electronic device cannot accurately position the tracking object in the video based on the existing positioning model, and realizes that the electronic device combines the image feature and the language feature of the video frame image based on the preset shared backbone network to effectively and accurately position the tracking object in the video to be processed.
[0059] As Figure 2 shown, it is a schematic flowchart of the method for positioning a tracking object in a video provided by the present invention, which may include:
[0060] 201. Obtain the first image feature corresponding to the key frame image in the video to be processed.
[0061] Among them, the key frame image is any one of each frame image in the video to be processed.
[0062] The key frame image refers to the key frame I t corresponding image in the video to be processed.
[0063] Optionally, the key frame image may be the current frame image.
[0064] 202. Obtain the second image feature corresponding to the adjacent frame image of the key frame image.
[0065] Among them, adjacent frame images refer to key frame I t and adjacent frames [I t-τ in the adjacent time sequence (t - τ, t + τ). t+τ corresponding images.
[0066] Optionally, the adjacent time sequence can be set before the electronic device leaves the factory or can be user - defined, and no specific limitation is made here.
[0067] Optionally, the adjacent frame image can be the image corresponding to the adjacent frame of the current frame in the adjacent time sequence.
[0068] In some embodiments, based on a preset shared image backbone network, the electronic device can extract image features with a preset number of scales from the key frame image and the adjacent frame image.
[0069] Optionally, the preset number can be set before the electronic device leaves the factory or can be user - defined according to a large amount of experimental data, and no specific limitation is made here.
[0070] Exemplarily, assuming that the preset number is 3, then the electronic device extracts image features with 3 scales from the key frame image and the adjacent frame image, which are 1 / 32, 1 / 16, and 1 / 8 of the image size of the video to be processed respectively.
[0071] In some embodiments, the first image feature and the second image feature are the maximum scales obtained by the electronic device through upsampling.
[0072] 203. Based on the first image feature and the second image feature, based on a preset frame - dense feature aggregation network, obtain the current image feature corresponding to the current frame image, and obtain the current language feature corresponding to the current frame image.
[0073] In some embodiments, the electronic device fuses the first image feature and the second image feature in a splicing manner to obtain the image features corresponding to each frame of the video to be processed, and uses each image feature as the input parameter of the preset frame - dense feature aggregation network.
[0074] Optionally, the electronic device obtaining the current image feature corresponding to the current frame image based on the first image feature and the second image feature and based on a preset frame - dense feature aggregation network may include: the electronic device obtains a normalized weight matrix based on the preset frame - dense feature aggregation network according to the first image feature and the second image feature; the electronic device determines the current image feature corresponding to the current frame image according to the first image feature and the normalized weight matrix.
[0075] Optionally, the electronic device may obtain a normalized weight matrix based on a preset frame-dense feature aggregation network according to the first image feature and the second image feature, which may include: the electronic device obtains a first weight matrix based on a weight formula in the preset frame-dense feature aggregation network; the electronic device obtains the normalized weight matrix based on a normalization formula.
[0076] Among them, the weight formula is: W x→t = Ψ(Ω (3) ([F x ; F t )));
[0077] W x→t represents the first weight matrix; F x represents the first image feature; F t represents the second image feature; [;] represents feature vector concatenation; Ω (3) (·) represents three convolutional layers with a rectified linear unit (ReLU) activation function; Ψ(·) represents a convolutional layer without an activation function.
[0078] Among them, the normalization formula is ∑ x∈[t-τ,t+τ] w x→t = 1, w x→t ∈W x→t ;
[0079] w x→t is any matrix in the first weight matrix W x→t .
[0080] The electronic device first concatenates the first image feature F x and the second image feature F t based on the preset frame-dense feature aggregation network, and then performs corresponding processing on the concatenation result in three convolutional layers with ReLU activation functions and one convolutional layer without an activation function to obtain the first weight matrix W x in the feature map space between the first image feature F t and the second image feature F x→t ; then, the electronic device normalizes the first weight matrix W x→t element by element along the adjacent time series (t-τ, t+τ) dimension using a softmax function to obtain the normalized weight matrix.
[0081] Optionally, the electronic device may determine a current image feature corresponding to the current frame image according to the first image feature and the normalized weight matrix, which may include: the electronic device obtains the current image feature corresponding to the current frame image according to an image feature formula.
[0082] Among them, the image feature formula is
[0083] represents the current image features; ⊙ represents element-wise multiplication.
[0084] The electronic device can obtain information on changes in the motion and / or appearance, etc. of the tracking object in each frame of the image at adjacent time series by means of an image feature formula, that is, the electronic device performs weighted aggregation on the second image feature at the spatial position of the first image feature in a manner of adaptively generating a weighted matrix, thereby assisting the first image feature F x in the feature learning of the preset frame-dense feature aggregation network.
[0085] 204. Based on the current image features and the preset shared image backbone network, obtain the visual vector features.
[0086] The electronic device can obtain the current image features corresponding to the frame images in multiple stages according to the video to be processed; then, the electronic device obtains the visual vector features corresponding to each current image feature based on the preset shared image backbone network according to each current image feature.
[0087] Optionally, the electronic device obtaining the visual vector features based on the current image features and the preset shared image backbone network may include: the electronic device obtains the position coordinate vector corresponding to the current image features; the electronic device obtains the visual vector features based on the current image features and the position coordinate vector and the preset shared image backbone network.
[0088] Among them, the position coordinate vector is represented by i represents the horizontal pixel position of the current image feature; j represents the vertical pixel position of the current image feature; w represents the width of the current frame image; h represents the height of the current frame image.
[0089] After obtaining the current image features, the electronic device can splice the current image features and the position coordinate vector; then perform feature transformation on the splicing result in the convolutional layer of the language-image generation branch in the preset shared image backbone network to obtain the visual vector features corresponding to the current image features.
[0090] 205. Based on the current language features and the visual vector features and the preset shared image backbone network, obtain the first similarity matrix.
[0091] Optionally, the electronic device obtaining the first similarity matrix based on the current language features and the visual vector features and the preset shared image backbone network may include: the electronic device obtains the first similarity matrix according to the first similarity formula in the preset shared image backbone network.
[0092] Among them, the first similarity formula is
[0093] s lv represents the first similarity matrix, represents the visual vector feature corresponding to the current frame image in the k-th stage; represents the current language feature corresponding to the current frame image in the k-th stage; f v represents the visual feature vector the visual feature element in the corresponding matrix; f l represents the current language feature the language feature element in the corresponding matrix; l represents the current language feature length.
[0094] In some embodiments, the first similarity matrix refers to the language-image similarity matrix.
[0095] The electronic device can perform element-by-element calculations on the visual feature elements and language feature elements to obtain the first similarity matrix; then, the electronic device can activate the first similarity matrix along the column vector dimension with the Softmax function. In the activated first similarity matrix, the column elements represent the similarity between each element at the element position of the current frame image and each word of the language expression, and the sum of the similarities is 1. That is, the electronic device obtains the element features corresponding to each element position in the current frame image, and how much feature information each word of the description statement provides.
[0096] Optionally, after step 205, the method may further include: the electronic device obtains the candidate position corresponding to the tracking object in the current image features; the electronic device adds a first constraint function to the first similarity matrix according to the candidate position.
[0097] Among them, the first constraint function is
[0098]
[0099] N represents the number of stages; y k(m) ∈{0,1} represents the element of the localization ground truth template matrix, where the candidate position corresponding to the localization ground truth of the tracking object is 1, and the non-candidate position is 0.
[0100] The electronic device can, according to the candidate position corresponding to the localization ground truth of the tracking object in the current frame image of the k-th stage, in the language-image generation branch, for the first similarity matrix Add a first constraint function to improve the ability of image features and language features to generate each other. Since the optimal candidate position can be the geometric center corresponding to the ground truth box of the tracked object, the subsequent electronic device can effectively constrain the generation of fused image features by the current language features using this first constraint function to improve the positioning accuracy of the tracked object.
[0101] 206. Obtain a second similarity matrix based on the current language features and visual vector features and a preset shared language backbone network.
[0102] Optionally, for the electronic device to obtain a second similarity matrix based on the current language features and visual vector features and a preset shared language backbone network, it may include: the electronic device obtains the second similarity matrix according to the second similarity formula in the preset shared language backbone network.
[0103] Among them, the second similarity formula is
[0104] s vl represents the second similarity matrix.
[0105] In some embodiments, the second similarity matrix refers to an image-language similarity matrix.
[0106] The second similarity matrix and the first similarity matrix are not in the relationship of transpose matrices. Each column in the second similarity matrix represents the similarity relationship between each language expression word and the elements in the current image features, and the sum of the similarities is 1. That is to say, the electronic device obtains the feature of each word in the description statement of the current frame image, and how much feature information each element in the current frame image needs to contribute.
[0107] Optionally, after step 206, the method may further include: the electronic device obtains the candidate position corresponding to the tracked object in the current language features; the electronic device adds a second constraint function to the second similarity matrix according to the candidate position.
[0108] Among them, the second constraint function is
[0109]
[0110] represents the template vector of the language features, with the position of the word being 1 and the position without the word being 0.
[0111] The electronic device can, according to the description statement features, add a second constraint function to the second similarity matrix s in the image-language generation branch vl to improve the ability of image features and language features to generate each other, thereby effectively constraining the subsequent electronic device to generate fused language features according to the current image features to improve the positioning accuracy of the tracked object.
[0112] 207. Determine the fused image features corresponding to the current frame image according to the current language features and the first similarity matrix.
[0113] Optionally, for the electronic device to determine the fused image features corresponding to the current frame image according to the visual feature vector and the first similarity matrix, it may include: the electronic device obtains the target image features corresponding to the current frame image according to the first formula; the electronic device determines the fused image features according to the target image features.
[0114] Among them, the first formula is
[0115] represents the target image features corresponding to the current frame image in the k-th stage.
[0116] The electronic device establishes a connection between the language-image similarity matrix and the current language features one by one, and supplements the image information of the current frame image in a cross-modal generation manner.
[0117] In some embodiments, after the electronic device obtains the target image features it first concatenates the target image features with the visual vector features Then, the electronic device performs feature transformation on the concatenated result in a convolutional layer; then, the electronic device adds the result of the feature transformation to the visual vector features element by element through a residual connection to obtain the fused image features In this way, it can effectively ensure the forward propagation and gradient backpropagation of the fused image features
[0118] 208. Determine the fused language features corresponding to the current frame image according to the visual feature vector and the second similarity matrix.
[0119] Optionally, for the electronic device to determine the fused language features corresponding to the current frame image according to the visual feature vector and the second similarity matrix, it may include: the electronic device obtains the target language features corresponding to the current frame image according to the second formula; the electronic device determines the fused language features according to the target language features.
[0120] Among them, the second formula is
[0121] represents the target language features corresponding to the current frame image in the k-th stage.
[0122] The electronic device establishes a connection between the image - language similarity matrix and the current image features one by one, and supplements the language information of the current frame image in a cross - modal generation manner.
[0123] In some embodiments, after the electronic device obtains the target language feature it first concatenates the target language feature with the current language feature Then, the electronic device learns the concatenation result in the fully - connected layer of the ReLU activation function. Next, the electronic device adds the learned result to the current language feature element - by - element through a residual connection to obtain the fused language feature corresponding to the next stage at the k - th stage
[0124] 209. Determine the language expression sentence feature according to the fused language feature.
[0125] Optionally, the electronic device determines the language expression sentence feature according to the fused language feature, which may include: the electronic device obtains the language expression sentence feature according to an aggregation formula.
[0126] Among them, the aggregation formula is:
[0127] u j =tanh(W q q j );
[0128] F w represents the language expression sentence feature; α j represents the first intermediate parameter; u j represents the second intermediate parameter; W q represents the fully - connected layer; q j represents the fused language feature corresponding to the j - th word in.
[0129] The electronic device first aggregates the fused language feature obtained from the image - language generation branch through a fully - connected layer and the Softmax activation function in an attention - weighted manner to obtain a more accurate language expression sentence feature F w .
[0130] 210. Determine the first language condition vector and the second speech condition vector according to the speech expression sentence feature.
[0131] Optionally, the electronic device determines a first language conditional vector and a second speech conditional vector according to the speech expression sentence features, which may include: the electronic device obtains the first language conditional vector according to a first language formula; the electronic device obtains the second language conditional vector according to a second language formula.
[0132] Among them, the first language formula is γ k =tanh(W γ F w +b γ );
[0133] The second language formula is β k =tanh(W β F w +b β );
[0134] γ k represents the first language conditional vector; β k represents the second language conditional vector; W γ represents the first learnable parameter matrix; W β represents the second learnable parameter matrix; b γ represents the first learnable parameter value; b β represents the second learnable parameter value.
[0135] In some embodiments, γ k refers to the scaling scale; β k refers to the translation size.
[0136] Optionally, the first learnable parameter data frame W γ , the second learnable parameter matrix W β , the first learnable parameter value b γ and the second learnable parameter value b β are pre-trained and learned by the electronic device.
[0137] Optionally, after step 210, the method may further include: the electronic device copies, splices, and adjusts the sizes of the first language conditional vector and the second language conditional vector to obtain a new first language conditional vector and a new second language conditional vector.
[0138] The size adjustment means that the electronic device adjusts the sizes of the first language conditional vector and the second language conditional vector to the same size as the size of the current frame image.
[0139] 211. Determine the positioning result of the tracking object according to the fused image features, the first language conditional vector, and the second speech conditional vector.
[0140] Optionally, the electronic device determines the positioning result of the tracking object according to the fused image feature, the first language conditional vector, and the second speech conditional vector, which may include: the electronic device obtains the target feature corresponding to the tracking object according to the target formula; the electronic device determines the positioning result of the tracking object according to the target feature.
[0141] Among them, the target formula is
[0142] represents the target feature.
[0143] After obtaining the target feature, under the guidance of the expression statement, the electronic device uses the scaling scale γ and the translation size β k to refine the fused image feature k to maximize the fusion between the current image feature and the current language feature. to refine the fused image feature to maximize the fusion between the current image feature and the current language feature.
[0144] After obtaining the target feature, the electronic device can obtain the prediction result of the positioning box corresponding to the tracking object in the image to be processed through learning of several convolutional layers.
[0145] In the embodiment of the present invention, in the process of positioning the tracking object in the current frame image of the video to be processed, the current image feature corresponding to the current frame image can be accurately determined according to the first image feature corresponding to the obtained key frame image and the second image feature corresponding to the adjacent frame image, and the current language feature corresponding to the current frame image can be obtained; then, according to the current image feature and the current language feature, based on the preset shared backbone network, the relatively accurate fused image feature and fused language feature corresponding to the current frame image can be obtained; finally, according to the fused image feature and the fused language feature, the positioning result of the tracking object can be accurately determined, so as to realize the accurate positioning of the tracking object in the video to be processed. This method is used to solve the defect that in the prior art, due to the limitations of the existing positioning model, the electronic device cannot accurately position the tracking object in the video based on the existing positioning model, and realizes that the electronic device effectively and accurately positions the tracking object in the video to be processed based on the preset shared backbone network, combining the image feature and language feature of the video frame image.
[0146] As Figure 3 shown, the flowchart of the method for positioning the tracking object in the video provided by the present invention may include:
[0147] 301. Obtain the first image feature corresponding to the key frame image in the video to be processed.
[0148] Among them, the key-frame image is any one of the images in each frame of the video to be processed.
[0149] 302. Obtain the second image feature corresponding to the adjacent frame image of the key-frame image.
[0150] 303. Based on the first image feature and the second image feature, and based on a preset frame dense feature aggregation network, obtain the current image feature corresponding to the current frame image, and obtain the current language feature corresponding to the current frame image.
[0151] 304. Based on the current image feature and based on a preset shared image backbone network, obtain the visual vector feature.
[0152] 305. Based on the current language feature and the visual vector feature, and based on a preset shared image backbone network, obtain the first similarity matrix.
[0153] 306. Based on the current language feature and the visual vector feature, and based on a preset shared language backbone network, obtain the second similarity matrix.
[0154] 307. Based on the visual feature vector and the first similarity matrix, determine the fused image feature corresponding to the current frame image.
[0155] 308. Based on the current language feature and the second similarity matrix, determine the fused language feature corresponding to the current frame image.
[0156] 309. Based on the fused language feature, determine the language expression sentence feature.
[0157] 310. Based on the speech expression sentence feature, determine the first language conditional vector and the second speech conditional vector.
[0158] 311. Based on the fused image feature, the first language conditional vector and the second speech conditional vector, determine the positioning result of the tracking object.
[0159] It should be noted that steps 301 and 311 are similar to Figure 2 the steps 201-211 shown, and will not be specifically elaborated here.
[0160] 312. Obtain the first constraint function corresponding to the first similarity matrix and the second constraint function corresponding to the second similarity matrix.
[0161] It should be noted that step 312 has been described in detail in Figure 2 the steps 205-206 shown, and will not be specifically elaborated here.
[0162] 313. Based on the first constraint function and the second constraint function, determine the positioning regression loss function corresponding to the preset shared backbone network.
[0163] Among them, the localization regression loss function is
[0164]
[0165] b ∈ {b x , b y , b w , b h} represents the prediction result of the localization box; p represents the confidence corresponding to the prediction result b of the localization box; b* represents the ground truth of the localization box; p* represents the confidence corresponding to the ground truth b* of the localization box; N b represents the number of anchors for each grid in the current image feature; L box (·) represents the mean squared error loss function for regressing the localization box; L conf (·) represents the cross-entropy loss function for regressing the confidence corresponding to the localization box.
[0166] 314. Determine the total loss function corresponding to the preset shared backbone network according to the localization regression loss function.
[0167] Among them, the total loss function is L = L det + λ(L lv + L vl );
[0168] λ represents a hyperparameter used to adjust the localization box regression loss and the image-language similarity matrix constraint loss.
[0169] The electronic device obtains the total loss function. In order to further improve the mutual generation ability between the current language feature and the current image feature in the current frame image, that is to say, the cross-modal feature generation ability can be further improved.
[0170] In an embodiment of the present invention, in the process of positioning a tracking object in the current frame image of a video to be processed, the current image feature corresponding to the current frame image and the current language feature can be accurately determined according to the first image feature corresponding to the obtained key frame image and the second image feature corresponding to the adjacent frame image; then, according to the current image feature and the current language feature, multiple constraint functions are used to constrain multiple formulas in a preset shared backbone network, and a relatively accurate fused image feature and fused language feature corresponding to the current frame image can be obtained; finally, according to the fused image feature and the fused language feature, the positioning result of the tracking object is accurately determined, so as to realize the accurate positioning of the tracking object in the video to be processed. This method is used to solve the defect that in the prior art, due to the limitations of the existing positioning model, the electronic device cannot accurately position the tracking object in the video based on the existing positioning model, and realizes that the electronic device combines the image feature and language feature of the video frame image based on the preset shared backbone network to effectively and accurately position the tracking object in the video to be processed.
[0171] The positioning device provided by the present invention will be described below. The positioning device described below can be correspondingly referred to the method for positioning a tracking object in a video described above.
[0172] As Figure 4 shown, the structural schematic diagram of the positioning device provided by the present invention may include:
[0173] An obtaining module 401, configured to obtain the current image feature and the current language feature corresponding to the current frame image in the process of positioning a tracking object in the current frame image of the video to be processed;
[0174] A determining module 402, configured to obtain a fused image feature and a fused language feature corresponding to the current frame image based on a preset shared backbone network according to the current image feature and the current language feature; and determine the positioning result of the tracking object according to the fused image feature and the fused language feature.
[0175] Optionally, the obtaining module 401 is specifically configured to obtain the first image feature corresponding to the key frame image in the video to be processed, where the key frame image is any one of each frame image in the video to be processed; obtain the second image feature corresponding to the adjacent frame image of the key frame image; and obtain the current image feature corresponding to the current frame image based on a preset frame dense feature aggregation network according to the first image feature and the second image feature.
[0176] Optionally, the determining module 402 is specifically configured to obtain a normalized weight matrix based on a preset frame dense feature aggregation network according to the first image feature and the second image feature; and determine a current image feature corresponding to the current frame image according to the first image feature and the normalized weight matrix.
[0177] Optionally, the determining module 402 is specifically configured to obtain a visual vector feature based on a preset shared image backbone network according to the current image feature; obtain a first similarity matrix based on the current language feature and the visual vector feature based on the preset shared image backbone network; obtain a second similarity matrix based on the current language feature and the visual vector feature based on a preset shared language backbone network; determine a fused image feature corresponding to the current frame image according to the current language feature and the first similarity matrix; and determine a fused language feature corresponding to the current frame image according to the visual feature vector and the second similarity matrix.
[0178] Optionally, the obtaining module 401 is further configured to obtain a candidate position corresponding to the tracking object in the current image feature.
[0179] The determining module 402 is further configured to add a first constraint function to the first similarity matrix according to the candidate position.
[0180] Optionally, the determining module 402 is specifically configured to determine a language expression sentence feature according to the fused language feature; determine a first language conditional vector and a second speech conditional vector according to the speech expression sentence feature; and determine a positioning result of the tracking object according to the fused image feature, the first language conditional vector, and the second speech conditional vector.
[0181] Optionally, the obtaining module 401 is specifically configured to obtain a first constraint function corresponding to the first similarity matrix and a second constraint function corresponding to the second similarity matrix.
[0182] The determining module 402 is specifically configured to determine a positioning regression loss function corresponding to the preset shared backbone network according to the first constraint function and the second constraint function; and determine a total loss function corresponding to the preset shared backbone network according to the positioning regression loss function.
[0183] Figure 5 An example of a schematic physical structure diagram of an electronic device is shown in Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute the method for locating a tracked object in a video. The method includes: during the process of locating the tracked object in the current frame image of the video to be processed, obtaining the current image feature and the current language feature corresponding to the current frame image; based on the current image feature and the current language feature, and based on a preset shared backbone network, obtaining the fused image feature and the fused language feature corresponding to the current frame image; according to the fused image feature and the fused language feature, determining the positioning result of the tracked object.
[0184] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0185] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for locating a tracked object in a video provided by the above-mentioned various methods. The method includes: during the process of locating the tracked object in the current frame image of the video to be processed, obtaining the current image feature and the current language feature corresponding to the current frame image; based on the current image feature and the current language feature, and based on a preset shared backbone network, obtaining the fused image feature and the fused language feature corresponding to the current frame image; according to the fused image feature and the fused language feature, determining the positioning result of the tracked object.
[0186] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for locating a tracked object in a video provided by the above-mentioned various methods. The method includes: during the process of locating the tracked object in the current frame image of the video to be processed, obtaining the current image feature and the current language feature corresponding to the current frame image; based on the preset shared backbone network according to the current image feature and the current language feature, obtaining the fused image feature and the fused language feature corresponding to the current frame image; and determining the positioning result of the tracked object according to the fused image feature and the fused language feature.
[0187] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0188] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for locating an object to be tracked in a video, characterized in that, Including: During the process of locating a tracking object in the current frame image of the video to be processed, obtaining the current image feature and the current language feature corresponding to the current frame image; According to the current image feature and the current language feature, based on a preset shared backbone network, obtaining a fused image feature and a fused language feature corresponding to the current frame image; the preset shared backbone network is a video referential expression understanding network based on multi-stage image-natural language cross-generation fusion; the preset shared backbone network adopts a one-stage object detection framework; According to the fused image feature and the fused language feature, determining the positioning result of the tracking object; The step of obtaining the fused image feature and the fused language feature corresponding to the current frame image according to the current image feature and the current language feature based on a preset shared backbone network includes: Based on the current image feature and a preset shared image backbone network, obtaining a visual vector feature; Based on the current language feature and the visual vector feature, and based on the preset shared image backbone network, obtaining a first similarity matrix; Based on the current language feature and the visual vector feature, and based on a preset shared language backbone network, obtaining a second similarity matrix; Based on the current language feature and the first similarity matrix, determining the fused image feature corresponding to the current frame image; Based on the visual vector feature and the second similarity matrix, determining the fused language feature corresponding to the current frame image; The step of determining the positioning result of the tracking object according to the fused image feature and the fused language feature includes: Based on the fused language feature, determining a language expression sentence feature; Based on the language expression sentence feature, determining a first language conditional vector and a second language conditional vector; Based on the fused image feature, the first language conditional vector, and the second language conditional vector, determining the positioning result of the tracking object.
2. The positioning method according to claim 1, characterized in that, The step of obtaining the current image feature corresponding to the current frame image includes: Obtaining a first image feature corresponding to a key frame image in the video to be processed, where the key frame image is any one of each frame image in the video to be processed; Obtaining a second image feature corresponding to an adjacent frame image of the key frame image; Based on the first image feature and the second image feature, and based on a preset frame dense feature aggregation network, obtaining the current image feature corresponding to the current frame image.
3. The positioning method according to claim 2, wherein The step of obtaining the current image feature corresponding to the current frame image based on the first image feature and the second image feature and based on a preset frame dense feature aggregation network includes: Based on a preset frame dense feature aggregation network, according to the first image feature and the second image feature, obtaining a normalized weight matrix; Based on the first image feature and the normalized weight matrix, determining the current image feature corresponding to the current frame image.
4. The positioning method according to claim 1, wherein After obtaining the first similarity matrix based on the current language feature and the visual vector feature and based on the preset shared image backbone network, the method further includes: Obtain the candidate position corresponding to the tracked object in the current image features; Add a first constraint function to the first similarity matrix according to the candidate position.
5. The positioning method according to claim 1 or 4, characterized in that The method further includes: Obtain the first constraint function corresponding to the first similarity matrix and the second constraint function corresponding to the second similarity matrix; Determine the localization regression loss function corresponding to the preset shared backbone network according to the first constraint function and the second constraint function; Determine the total loss function corresponding to the preset shared backbone network according to the localization regression loss function.
6. A positioning device, characterized in that, Includes: An acquisition module, configured to obtain the current image features and the current language features corresponding to the current frame image during the process of localizing the tracked object in the current frame image of the video to be processed; A determination module, configured to obtain the fused image features and the fused language features corresponding to the current frame image based on a preset shared backbone network according to the current image features and the current language features; determine the localization result of the tracked object according to the fused image features and the fused language features; the preset shared backbone network is a video referential expression understanding network based on multi-stage image-natural language cross-generation fusion; the preset shared backbone network adopts a one-stage object detection framework; The obtaining the fused image features and the fused language features corresponding to the current frame image based on a preset shared backbone network according to the current image features and the current language features includes: Obtain visual vector features based on the current image features and a preset shared image backbone network; Obtain a first similarity matrix based on the current language features and the visual vector features and a preset shared image backbone network; Obtain a second similarity matrix based on the current language features and the visual vector features and a preset shared language backbone network; Determine the fused image features corresponding to the current frame image according to the current language features and the first similarity matrix; Determine the fused language features corresponding to the current frame image according to the visual vector features and the second similarity matrix; The determining the localization result of the tracked object according to the fused image features and the fused language features includes: Determine the language expression sentence features according to the fused language features; Determine a first language conditional vector and a second language conditional vector according to the language expression sentence features; Determine the localization result of the tracked object according to the fused image features, the first language conditional vector, and the second language conditional vector.
7. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method for localizing a tracked object in a video according to any one of claims 1 to 5 when executing the program.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by the processor, implements the method for localizing a tracked object in a video according to any one of claims 1 to 5.