Dynamic scene positioning and tracking methods, devices and media
By generating a mask to shield dynamic objects, feature extraction and matching are performed, solving the problems of positioning failure and low tracking accuracy in dynamic scene localization and tracking, and achieving higher positioning and tracking accuracy.
Patent Information
- Application Number
- CN202311149732.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-07
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2043-09-07
AI Technical Summary
Existing dynamic scene localization and tracking methods suffer from localization failures or low tracking accuracy, especially when the actual scene changes or dynamic objects are present, resulting in a decrease in localization and tracking accuracy.
By generating cue vectors, creating masks to block dynamic objects, performing feature extraction and matching, using static features to perform feature matching with preset static 3D models, estimating camera pose, and filtering out information that does not need to participate in map matching.
It improves the accuracy of positioning and tracking, enabling more accurate positioning and tracking in dynamic scenes and enhancing matching accuracy in dynamic environments.
Smart Images

Figure CN117197237B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of positioning technology, and in particular to a dynamic scene positioning and tracking method, device and medium. Background Technology
[0002] Simultaneous Localization and Mapping (SLAM) has long been a pain point in the industry regarding localization and tracking. A relatively mature approach involves pre-creating a 3D model of the target scene and storing key map information such as keyframes, feature points, and map points for localization and tracking. When users actually experience augmented reality (AR) applications, the images captured by the user's camera are matched with the stored 3D model information to achieve localization and tracking.
[0003] However, when users experience AR applications, if the actual scene changes or there are dynamic objects such as people, animals, or vehicles, the matching degree between the actual scene and the pre-stored 3D modeling information decreases, causing positioning failures or reduced tracking accuracy. Summary of the Invention
[0004] This invention provides a dynamic scene positioning and tracking method, device, and medium to solve the problems of positioning failure or low tracking accuracy in existing dynamic scene positioning and tracking methods.
[0005] To solve the above-mentioned technical problems, the present invention is implemented as follows:
[0006] In a first aspect, embodiments of the present invention provide a dynamic scene positioning and tracking method, including:
[0007] Based on the user's target operation, at least one prompt vector is generated, wherein the target operation is used to select a dynamic object in the actual scene image;
[0008] Generate a mask based on the actual scene image and the at least one cue vector;
[0009] Based on the mask, feature extraction is performed on the actual scene image to obtain static features;
[0010] The static features are matched with the preset static 3D model to obtain the feature matching results;
[0011] Based on the feature matching results, location tracking is performed;
[0012] The mask is used to determine the characteristics of the dynamic object.
[0013] Secondly, embodiments of the present invention provide a dynamic scene positioning and tracking device, comprising:
[0014] The first generation module is used to generate at least one prompt vector based on the user's target operation, wherein the target operation is used to select a dynamic object in the actual scene image;
[0015] The second generation module is used to generate a mask based on the actual scene image and the at least one cue vector;
[0016] The extraction module is used to extract features from the actual scene image based on the mask to obtain static features;
[0017] The matching module is used to perform feature matching between the static features and the preset static 3D model to obtain the feature matching result;
[0018] The location tracking module is used to perform location tracking based on the feature matching results;
[0019] The mask is used to determine the characteristics of the dynamic object.
[0020] Thirdly, embodiments of the present invention provide an electronic device, including a processor, the processor being used for:
[0021] Based on the user's target operation, at least one prompt vector is generated, wherein the target operation is used to select a dynamic object in the actual scene image;
[0022] Generate a mask based on the actual scene image and the at least one cue vector;
[0023] Based on the mask, feature extraction is performed on the actual scene image to obtain static features;
[0024] The static features are matched with the preset static 3D model to obtain the feature matching results;
[0025] Based on the feature matching results, location tracking is performed;
[0026] The mask is used to determine the characteristics of the dynamic object.
[0027] Fourthly, embodiments of the present invention provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the dynamic scene positioning and tracking method described in the first aspect above.
[0028] Fifthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the dynamic scene positioning and tracking method described in the first aspect above.
[0029] In this embodiment of the invention, the dynamic scene positioning and tracking method can generate at least one prompt vector based on the user's target operation, and use the at least one prompt vector to generate a corresponding mask. During the feature extraction and feature matching process, features within the mask range are ignored, and feature matching is performed with a preset static 3D model to estimate the camera pose. Thus, during the positioning and tracking process, information that does not need to participate in map matching can be filtered out, thereby performing positioning and tracking more accurately and improving the accuracy of positioning and tracking. Attached Figure Description
[0030] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is a flowchart of a dynamic scene positioning and tracking method provided in an embodiment of the present invention;
[0032] Figure 2 This is a schematic diagram of the structure of a dynamic scene positioning and tracking device provided in an embodiment of the present invention;
[0033] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0035] In this embodiment of the invention, a dynamic scene positioning and tracking method, device and medium are proposed to solve the problems of positioning failure or low tracking accuracy in existing dynamic scene positioning and tracking methods.
[0036] See Figure 1 , Figure 1 This is a flowchart of a dynamic scene positioning and tracking method provided in an embodiment of the present invention, such as... Figure 1As shown, the method includes the following steps:
[0037] Step 101: Generate at least one prompt vector based on the user's target operation, wherein the target operation is used to select a dynamic object in the actual scene image.
[0038] Specifically, the target operation can be an operation input by the user to the AR device, and the prompt vector can be a vector used to indicate the object corresponding to the target operation in the actual scene image during the mask generation process. The dynamic object can be a person, a vehicle, or other dynamic targets, as well as other objects that the user does not need to participate in map matching.
[0039] Step 102: Generate a mask based on the actual scene image and the at least one cue vector.
[0040] Specifically, the mask can be used to determine the location of the dynamic object, thereby shielding the dynamic object.
[0041] Step 103: Extract features from the actual scene image based on the mask to obtain static features.
[0042] Specifically, the feature extraction process can ignore the features corresponding to the mask, and the static features can be all the features of the actual scene image after removing the features corresponding to the mask.
[0043] Step 104: Perform feature matching between the static features and the preset static 3D model to obtain the feature matching result.
[0044] Specifically, the feature matching can be a process of combining the static features with the preset static 3D model to estimate visual relative position changes and track.
[0045] Step 105: Perform localization and tracking based on the feature matching results; wherein the mask is used to determine the features of the dynamic object.
[0046] Specifically, the camera pose can be estimated based on the feature matching results, and the spatial position of a specific object at different times can be tracked.
[0047] In this embodiment of the invention, the dynamic scene positioning and tracking method can generate at least one prompt vector based on the user's target operation, and use the at least one prompt vector to generate a corresponding mask. During the feature extraction and feature matching process, features within the mask range are ignored, and feature matching is performed with a preset static 3D model to estimate the camera pose. Thus, during the positioning and tracking process, information that does not need to participate in map matching can be filtered out, thereby performing positioning and tracking more accurately and improving the accuracy of positioning and tracking.
[0048] Optionally, the target operation includes at least one of the following: tap selection, box selection, voice input, and text input, and generating at least one prompt vector based on the user's target operation includes:
[0049] Based on the user's selection, obtain the coordinates and first category of the first point selected by the user;
[0050] The coordinates of the first point are encoded to obtain the first D-dimensional vector of the first point;
[0051] Add the first D-dimensional vector and the second D-dimensional vector corresponding to the first category to obtain the prompt vector of the first point;
[0052] Alternatively, generating at least one prompt vector based on the user's target action includes:
[0053] Based on the user's selection, the coordinates of the selected box and the second category are obtained, wherein the coordinates of the box include the coordinates of the second point and the coordinates of the third point;
[0054] The coordinates of the second point and the third point are respectively encoded to obtain the third D-dimensional vector of the second point and the fourth D-dimensional vector of the third point;
[0055] The third D-dimensional vector and the fifth D-dimensional vector corresponding to the second category are added together to obtain the prompt vector of the second point;
[0056] The fourth D-dimensional vector and the fifth D-dimensional vector corresponding to the second category are added together to obtain the prompt vector of the third point;
[0057] Alternatively, generating at least one prompt vector based on the user's target action includes:
[0058] The user's voice input and / or text input are encoded to obtain a prompt vector for the voice input and / or text input.
[0059] Specifically, the point selection can be achieved by the user clicking on an actual scene image on the device screen to generate a point, and then selecting a first category for that point, which can be either a foreground or background point. The box selection can be achieved by the user clicking on the device screen, expanding a box, and then selecting a second category to which the box belongs, which can also be either foreground or background. The second and third points can be the two vertices corresponding to the diagonal of the box, for example, the top-left and bottom-right vertices of the box. The voice input can be achieved by converting speech to text and then encoding it using a free-format encoder such as Compressed Lossless Image Protocol (CLIP) to obtain a prompt vector. The text input can be achieved by the user inputting background and foreground information into the platform in the form of text or commands, which can also be encoded using an encoder to obtain a prompt vector.
[0060] For example, a user clicks on an actual scene image on the device screen, generating a point. The point's coordinates are obtained, and the user selects whether the point belongs to the foreground or background. The point's coordinate information can be a D-dimensional vector obtained through positional encoding technology. Foreground and background each correspond to a learnable D-dimensional vector. The D-dimensional vector of this point is added to the D-dimensional vector corresponding to whether the point belongs to the foreground or background, resulting in a single D-dimensional vector used as the point's prompt vector. The user clicks on the device screen and expands a box, selecting whether the box belongs to the foreground or background. The prompt encoding method for the box is similar to that for the point; the coordinates of the top-left and bottom-right corners can be taken, and a prompt vector can be constructed according to the point selection method, combining them into a vector pair, i.e., two D-dimensional vectors, as the prompt vector.
[0061] It should be noted that the platform can verify the user's selected points, boxes, voice input, or text commands to prevent accidental or random selections. The platform can also automatically detect dynamic objects in the actual target scene based on feature comparison.
[0062] In this embodiment of the invention, the dynamic scene positioning and tracking method can generate a mask of a dynamic target in real time through interactive operations with the user, such as point selection, box selection, voice input, or text input. This allows the user to participate in the mask generation process by dividing the foreground and background in the actual scene or by inputting voice and text, thereby effectively improving the accuracy of positioning and tracking.
[0063] Optionally, generating a mask based on the actual scene image and the at least one cue vector includes:
[0064] Determine the embedding vector of the actual scene image, the embedding vector being used to characterize the actual scene image;
[0065] The Embedding vector and the at least one cue vector are input into the mask decoder to obtain the mask.
[0066] Specifically, the at least one cue vector may include multiple cue vectors generated by the user through point selection, box selection, voice input, or text input. These cue vectors may be generated in the same way or in different ways. The mask decoder may use at least one cue vector to perform mask decoding on the embedding vector of the actual scene image to obtain the mask indicated by the at least one cue vector.
[0067] In this embodiment of the invention, the dynamic scene localization and tracking method can determine the embedding vector of the actual scene image, and input the embedding vector and the at least one cue vector into the mask decoder to obtain a mask, thereby improving the accuracy of mask generation and further improving the accuracy of localization and tracking.
[0068] Optionally, the mask decoder includes a cross-attention mechanism layer, a multilayer perceptron, and a multilayer transposed convolution. The step of inputting the embedding vector and the at least one cue vector into the mask decoder to obtain the mask includes:
[0069] The at least one cue vector and the Embedding vector are input into the cross-attention mechanism layer to obtain the first output result;
[0070] The first output result is input into the multilayer perceptron to obtain the second output result;
[0071] The second output and the Embedding vector are input into the cross-attention mechanism layer to obtain the third output.
[0072] The second and third output results are input into the cross-attention mechanism layer to obtain the fourth output result;
[0073] The Embedding vector is input into the multi-layer transposed convolution to obtain the fifth output result;
[0074] The fourth output result is input into the multilayer perceptron to obtain the sixth output result, which has the same channel dimension as the fifth output result.
[0075] Multiply the fifth output result and the sixth output result by a dot product to obtain the mask.
[0076] Specifically, the at least one cue vector and the embedding vector are input into the cross-attention mechanism layer, and the at least one cue vector is used as a query for interaction information to obtain a first output result; the first output result is input into the multilayer perceptron for nonlinear transformation to obtain a second output result; the second output result and the embedding vector are input into the cross-attention mechanism layer, and the embedding vector is used as a query for interaction information to obtain a third output result; the second output result and the third output result are input into the cross-attention mechanism layer, and the second output result is used as a query for interaction information to obtain a fourth output result; the embedding vector is input into the multilayer transposed convolution, and upsampling is performed, with the upsampling factor being equivalent to the downsampling factor in the embedding vector generation process, i.e., restoring the embedding vector to the original image size, and the convolution yields a fifth output result; the fourth output result is input into the multilayer perceptron to obtain a sixth output result, and the sixth output result has the same channel dimension as the fifth output result; the fifth output result and the sixth output result are multiplied by a dot to obtain a mask.
[0077] In this embodiment of the invention, the dynamic scene localization and tracking method can utilize the cross-attention mechanism layer, multilayer perceptron, and multilayer transposed convolution in the mask decoder to generate a mask based on the embedding vector and the at least one cue vector, thereby improving the efficiency and accuracy of mask generation.
[0078] Optionally, the mask decoder further includes a self-attention mechanism layer, wherein inputting the at least one cue vector and the embedding vector into the cross-attention mechanism layer to obtain a first output result includes:
[0079] The at least one cue vector is input into the self-attention mechanism layer to obtain the seventh output result;
[0080] The seventh output result and the Embedding vector are input into the cross-attention mechanism layer to obtain the first output result.
[0081] Specifically, the self-attention mechanism layer can be used for information interaction between multiple cue vectors.
[0082] In this embodiment of the invention, since the prompting methods of multiple prompt vectors are not completely the same, the self-attention mechanism layer of the mask decoder can be used to interact with multiple prompt vectors. The output result after interaction and the embedding vector are input into the cross-attention mechanism layer to obtain a more effective prompt representation, thereby further improving the accuracy of localization and tracking.
[0083] Optionally, determining the embedding vector of the actual scene image includes:
[0084] The actual scene image is input into the encoder of the preset model to obtain the embedding vector of the actual scene image;
[0085] The preset model is trained based on the training dataset;
[0086] The training dataset includes a first image set and a second image set corresponding to the first image set, wherein the second image set is obtained by randomly erasing the first image set.
[0087] Specifically, the preset model can be a combination of encoder and decoder built using the Transformer approach. The decoder can be used to encode the actual scene image to obtain an embedding vector. The first image set can be publicly available images on the Internet obtained through methods such as web scraping. The second image set can be an image set obtained by randomly erasing each image in the first image set, with a maximum of 50% of the original image region remaining. The preset model is trained based on the training dataset by: inputting the images in the second image set into the encoder of the preset model, which outputs an H*W*C shaped embedding vector to represent the image. This embedding vector is then used as input, and the corresponding original image in the first image set is used as the output target to train the encoder-decoder combination. After the preset model is trained, the encoder module is used to generate the embedding vector of the actual scene image.
[0088] In this embodiment of the invention, the dynamic scene localization and tracking method can obtain the embedding vector of the actual scene image by inputting the actual scene image into the encoder of the preset model. The preset model is trained in a self-supervised manner, which can effectively reduce training time and reduce a lot of image annotation costs.
[0089] See Figure 2 , Figure 2 This is a structural schematic diagram of a dynamic scene positioning and tracking device provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the dynamic scene positioning and tracking device 200 includes:
[0090] The first generation module 201 is used to generate at least one prompt vector based on the user's target operation, wherein the target operation is used to select a dynamic object in the actual scene image;
[0091] The second generation module 202 is used to generate a mask based on the actual scene image and the at least one cue vector;
[0092] Extraction module 203 is used to extract features from the actual scene image based on the mask to obtain static features;
[0093] Matching module 204 is used to perform feature matching between the static features and the preset static 3D model to obtain feature matching results;
[0094] The positioning and tracking module 205 is used to perform positioning and tracking based on the feature matching results;
[0095] The mask is used to determine the characteristics of the dynamic object.
[0096] Optionally, the target operation includes at least one of the following: point selection, box selection, voice input, and text input, and the first generation module 201 includes:
[0097] The selection unit is used to obtain the coordinates and first category of the first point selected by the user based on the user's selection.
[0098] The first encoding unit is used to perform position encoding on the coordinates of the first point to obtain the first D-dimensional vector of the first point;
[0099] The first calculation unit is used to add the first D-dimensional vector and the second D-dimensional vector corresponding to the first category to obtain the prompt vector of the first point;
[0100] Alternatively, generating at least one prompt vector based on the user's target action includes:
[0101] The selection unit is used to obtain the coordinates of the selected box and the second category based on the user's selection, wherein the coordinates of the box include the coordinates of the second point and the coordinates of the third point;
[0102] The second encoding unit is used to perform position encoding on the coordinates of the second point and the coordinates of the third point respectively, to obtain the third D-dimensional vector of the second point and the fourth D-dimensional vector of the third point;
[0103] The second calculation unit is used to add the third D-dimensional vector and the fifth D-dimensional vector corresponding to the second category to obtain the prompt vector of the second point;
[0104] The third calculation unit is used to add the fourth D-dimensional vector and the fifth D-dimensional vector corresponding to the second category to obtain the prompt vector of the third point;
[0105] Alternatively, generating at least one prompt vector based on the user's target action includes:
[0106] The third encoding unit is used to encode the user's voice input and / or text input to obtain the prompt vector of the voice input and / or text input.
[0107] Optionally, the second generation module 202 includes:
[0108] A determining unit is configured to determine the embedding vector of the actual scene image, wherein the embedding vector is used to characterize the actual scene image;
[0109] A generation unit is used to input the Embedding vector and the at least one cue vector into a mask decoder to obtain a mask.
[0110] Optionally, the mask decoder includes a cross-attention mechanism layer, a multilayer perceptron, and a multilayer transposed convolution, and the generation unit is specifically used for:
[0111] The at least one cue vector and the Embedding vector are input into the cross-attention mechanism layer to obtain the first output result;
[0112] The first output result is input into the multilayer perceptron to obtain the second output result;
[0113] The second output and the Embedding vector are input into the cross-attention mechanism layer to obtain the third output.
[0114] The second and third output results are input into the cross-attention mechanism layer to obtain the fourth output result;
[0115] The Embedding vector is input into the multi-layer transposed convolution to obtain the fifth output result;
[0116] The fourth output result is input into the multilayer perceptron to obtain the sixth output result, which has the same channel dimension as the fifth output result.
[0117] Multiply the fifth output result and the sixth output result by a dot product to obtain the mask.
[0118] Optionally, the mask decoder further includes a self-attention mechanism layer, wherein inputting the at least one cue vector and the embedding vector into the cross-attention mechanism layer to obtain a first output result includes:
[0119] The at least one cue vector is input into the self-attention mechanism layer to obtain the seventh output result;
[0120] The seventh output result and the Embedding vector are input into the cross-attention mechanism layer to obtain the first output result.
[0121] Optionally, the determining unit is specifically used for:
[0122] The actual scene image is input into the encoder of the preset model to obtain the embedding vector of the actual scene image;
[0123] The preset model is trained based on the training dataset;
[0124] The training dataset includes a first image set and a second image set corresponding to the first image set, wherein the second image set is obtained by randomly erasing the first image set.
[0125] For details, see Figure 3 As shown, this embodiment of the invention also provides an electronic device, including a bus 301, a transceiver 302, an antenna 303, a bus interface 304, a processor 305, and a memory 306.
[0126] Processor 305, used for:
[0127] Based on the user's target operation, at least one prompt vector is generated, wherein the target operation is used to select a dynamic object in the actual scene image;
[0128] Generate a mask based on the actual scene image and the at least one cue vector;
[0129] Based on the mask, feature extraction is performed on the actual scene image to obtain static features;
[0130] The static features are matched with the preset static 3D model to obtain the feature matching results;
[0131] Based on the feature matching results, location tracking is performed;
[0132] The mask is used to determine the characteristics of the dynamic object.
[0133] exist Figure 3In this document, a bus architecture (represented by bus 301) is used. Bus 301 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 305 and memory represented by memory 306. Bus 301 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 304 provides an interface between bus 301 and transceiver 302. Transceiver 302 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 305 is transmitted over a wireless medium via antenna 303, which further receives data and transmits it to processor 305.
[0134] Processor 305 manages bus 301 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 306 can be used to store data used by processor 305 during operation.
[0135] Optionally, the processor 305 can be a CPU, ASIC, FPGA, or CPLD.
[0136] Optionally, the target operation includes at least one of the following: point selection, box selection, voice input, and text input, and the processor 305 is specifically used for:
[0137] Based on the user's selection, obtain the coordinates and first category of the first point selected by the user;
[0138] The coordinates of the first point are encoded to obtain the first D-dimensional vector of the first point;
[0139] Add the first D-dimensional vector and the second D-dimensional vector corresponding to the first category to obtain the prompt vector of the first point;
[0140] Alternatively, generating at least one prompt vector based on the user's target action includes:
[0141] Based on the user's selection, the coordinates of the selected box and the second category are obtained, wherein the coordinates of the box include the coordinates of the second point and the coordinates of the third point;
[0142] The coordinates of the second point and the third point are respectively encoded to obtain the third D-dimensional vector of the second point and the fourth D-dimensional vector of the third point;
[0143] The third D-dimensional vector and the fifth D-dimensional vector corresponding to the second category are added together to obtain the prompt vector of the second point;
[0144] The fourth D-dimensional vector and the fifth D-dimensional vector corresponding to the second category are added together to obtain the prompt vector of the third point;
[0145] Alternatively, generating at least one prompt vector based on the user's target action includes:
[0146] The user's voice input and / or text input are encoded to obtain a prompt vector for the voice input and / or text input.
[0147] Optionally, the processor 305 is specifically used for:
[0148] Determine the embedding vector of the actual scene image, the embedding vector being used to characterize the actual scene image;
[0149] The Embedding vector and the at least one cue vector are input into the mask decoder to obtain the mask.
[0150] Optionally, the mask decoder includes a cross-attention mechanism layer, a multilayer perceptron, and a multilayer transposed convolution, and the processor 305 is specifically used for:
[0151] The at least one cue vector and the Embedding vector are input into the cross-attention mechanism layer to obtain the first output result;
[0152] The first output result is input into the multilayer perceptron to obtain the second output result;
[0153] The second output and the Embedding vector are input into the cross-attention mechanism layer to obtain the third output.
[0154] The second and third output results are input into the cross-attention mechanism layer to obtain the fourth output result;
[0155] The Embedding vector is input into the multi-layer transposed convolution to obtain the fifth output result;
[0156] The fourth output result is input into the multilayer perceptron to obtain the sixth output result, which has the same channel dimension as the fifth output result.
[0157] Multiply the fifth output result and the sixth output result by a dot product to obtain the mask.
[0158] Optionally, the mask decoder further includes a self-attention mechanism layer, and the processor 305 is specifically used for:
[0159] The at least one cue vector is input into the self-attention mechanism layer to obtain the seventh output result;
[0160] The seventh output result and the Embedding vector are input into the cross-attention mechanism layer to obtain the first output result.
[0161] Optionally, the processor 305 is specifically used for:
[0162] The actual scene image is input into the encoder of the preset model to obtain the embedding vector of the actual scene image;
[0163] The preset model is trained based on the training dataset;
[0164] The training dataset includes a first image set and a second image set corresponding to the first image set, wherein the second image set is obtained by randomly erasing the first image set.
[0165] This invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described dynamic scene positioning and tracking method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0166] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described dynamic scene positioning and tracking method embodiments, achieving the same technical effects. To avoid repetition, further details are omitted here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0167] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0168] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0169] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.
Claims
1. A dynamic scene positioning tracking method, characterized in that, The method comprises the following steps: generating at least one prompt vector according to a target operation of a user, the target operation being used for selecting a dynamic object in an actual scene image; generating a mask mask according to the actual scene image and the at least one prompt vector; performing feature extraction on the actual scene image according to the mask mask to obtain static features; performing feature matching on the static features and a preset static three-dimensional modeling to obtain a feature matching result; performing positioning tracking according to the feature matching result; wherein the mask mask is used to determine the features of the dynamic object; the generating of the mask mask according to the actual scene image and the at least one prompt vector comprises: determining an Embedding vector of the actual scene image, the Embedding vector being used to represent the actual scene image; inputting the Embedding vector and the at least one prompt vector into a mask decoder to obtain the mask mask; the mask decoder comprises a cross-attention mechanism layer, a multi-layer perceptron and a multi-layer transposed convolution, and the inputting of the Embedding vector and the at least one prompt vector into the mask decoder to obtain the mask mask comprises: inputting the at least one prompt vector and the Embedding vector into the cross-attention mechanism layer to obtain a first output result; inputting the first output result into the multi-layer perceptron to obtain a second output result; inputting the second output result and the Embedding vector into the cross-attention mechanism layer to obtain a third output result; inputting the second output result and the third output result into the cross-attention mechanism layer to obtain a fourth output result; inputting the Embedding vector into the multi-layer transposed convolution to obtain a fifth output result; inputting the fourth output result into the multi-layer perceptron to obtain a sixth output result, the channel dimension of the sixth output result being the same as that of the fifth output result; performing point multiplication on the fifth output result and the sixth output result to obtain the mask mask.
2. The method of claim 1, wherein, The target operation comprises at least one of the following: point selection, frame selection, voice input and text input, and the generating of the at least one prompt vector according to the target operation of the user comprises: obtaining the coordinates of a first point selected by the user and a first category according to the point selection of the user; performing position encoding on the coordinates of the first point to obtain a first D-dimensional vector of the first point; performing addition on the first D-dimensional vector and a second D-dimensional vector corresponding to the first category to obtain a prompt vector of the first point; alternatively, the generating of the at least one prompt vector according to the target operation of the user comprises: obtaining the coordinates of a frame selected by the user and a second category according to the frame selection of the user, the coordinates of the frame comprising the coordinates of a second point and the coordinates of a third point; performing position encoding on the coordinates of the second point and the coordinates of the third point respectively to obtain a third D-dimensional vector of the second point and a fourth D-dimensional vector of the third point; performing addition on the third D-dimensional vector and a fifth D-dimensional vector corresponding to the second category to obtain a prompt vector of the second point; adding the fourth D-dimensional vector and a fifth D-dimensional vector corresponding to the second category, to obtain a prompt vector of the third point; Alternatively, the generating at least one prompt vector according to the target operation of the user comprises: encoding the voice input and / or text input of the user to obtain a prompt vector of the voice input and / or text input.
3. The method of claim 1, wherein, The mask decoder further comprises a self-attention mechanism layer, and the at least one prompt vector and the Embedding vector are input into the cross-attention mechanism layer to obtain a first output result, which comprises: the at least one prompt vector is input into the self-attention mechanism layer to obtain a seventh output result; the seventh output result and the Embedding vector are input into the cross-attention mechanism layer to obtain a first output result.
4. The method of claim 1, wherein, The determining the Embedding vector of the actual scene image comprises: inputting the actual scene image into an encoder of a preset model to obtain the Embedding vector of the actual scene image; The preset model is obtained by training based on a training data set; The training data set comprises a first image set and a second image set corresponding to the first image set, and the second image set is obtained by randomly erasing the first image set.
5. A dynamic scene positioning tracking device, characterized by, Comprise: The first generation module is configured to generate at least one prompt vector according to a target operation of a user, and the target operation is used to select a dynamic object in an actual scene image. The second generation module is configured to generate a mask according to the actual scene image and the at least one prompt vector. The extraction module is configured to extract features of the actual scene image according to the mask to obtain static features. The matching module is configured to perform feature matching on the static features and a preset static three-dimensional modeling to obtain a feature matching result. The positioning and tracking module is configured to perform positioning and tracking according to the feature matching result. The mask is used to determine the features of the dynamic object. The second generation module comprises: A determination unit is configured to determine an Embedding vector of the actual scene image, and the Embedding vector is used to represent the actual scene image. A generation unit is configured to input the Embedding vector and the at least one prompt vector into a mask decoder to obtain a mask. The mask decoder comprises a cross-attention mechanism layer, a multi-layer perceptron and a multi-layer transposed convolution, and the generation unit is specifically configured to: input the at least one prompt vector and the Embedding vector into the cross-attention mechanism layer to obtain a first output result; input the first output result into the multi-layer perceptron to obtain a second output result; input the second output result and the Embedding vector into the cross-attention mechanism layer to obtain a third output result; input the second output result and the third output result into the cross-attention mechanism layer to obtain a fourth output result; input the Embedding vector into the multi-layer transposed convolution to obtain a fifth output result; input the fourth output result into the multi-layer perceptron to obtain a sixth output result, the sixth output result and the fifth output result having the same channel dimension; perform point multiplication on the fifth output result and the sixth output result to obtain the mask mask.
6. An electronic device, comprising: The processor is configured to: generate at least one prompt vector according to a target operation of a user, the target operation being used to select a dynamic object in an actual scene image; generate a mask mask according to the actual scene image and the at least one prompt vector; perform feature extraction on the actual scene image according to the mask mask to obtain static features; perform feature matching on the static features and a preset static three-dimensional modeling to obtain a feature matching result; perform positioning tracking according to the feature matching result; The mask mask is used to determine the features of the dynamic object. The processor is specifically configured to: determine an Embedding vector of the actual scene image, the Embedding vector being used to represent the actual scene image; input the Embedding vector and the at least one prompt vector into a mask decoder to obtain the mask mask; The mask decoder includes a cross-attention mechanism layer, a multi-layer perceptron and a multi-layer transposed convolution, and the processor is specifically configured to: input the at least one prompt vector and the Embedding vector into the cross-attention mechanism layer to obtain a first output result; input the first output result into the multi-layer perceptron to obtain a second output result; input the second output result and the Embedding vector into the cross-attention mechanism layer to obtain a third output result; input the second output result and the third output result into the cross-attention mechanism layer to obtain a fourth output result; input the Embedding vector into the multi-layer transposed convolution to obtain a fifth output result; input the fourth output result into the multi-layer perceptron to obtain a sixth output result, the sixth output result and the fifth output result having the same channel dimension; perform point multiplication on the fifth output result and the sixth output result to obtain the mask mask.
7. An electronic device, comprising: The processor, the memory and the program stored on the memory and executable on the processor, when the program is executed by the processor, implement the steps of the dynamic scene positioning tracking method in any one of claims 1 to 4. The computer program is stored on the computer readable storage medium, and when the computer program is executed by the processor, the steps of the dynamic scene positioning tracking method in any one of claims 1 to 4 are implemented.
8. A computer-readable storage medium, characterized in that,
Citation Information
Patent Citations
Instance tracking method and device
CN112561961A
Image description method and device, equipment and storage medium
CN114581543A