Article object tracking method and device, computer device, and storage medium

By using 3D object models and object tracking models in enclosed spaces, combined with an attention mechanism, the problem of low object tracking accuracy was solved, and accurate tracking was achieved in complex situations.

CN116977901BActive Publication Date: 2026-05-29INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2023-08-03
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, object tracking methods in enclosed spaces suffer from low accuracy, especially when objects are occluded, overlapped, or interfered with by nearby objects, which can easily lead to missed or false alarms.

Method used

By acquiring multiple frames of video images captured by video image acquisition devices in a confined space and a pre-constructed 3D object model of the object to be tracked, the object tracking model is used to identify the object's location and track its trajectory. By combining an attention mechanism with the 2D representation of the 3D model, the tracking accuracy of the object is improved.

Benefits of technology

It enables accurate tracking of objects under conditions of occlusion, overlap, or interference from nearby objects, improving the tracking accuracy and recognition accuracy of objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116977901B_ABST
    Figure CN116977901B_ABST
Patent Text Reader

Abstract

The application relates to an article object tracking method and device, computer equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring multiple frames of video images of an article object to be tracked, which are captured by a video image acquisition device in a closed space, and a three-dimensional article model corresponding to the article object to be tracked, which is pre-constructed; inputting the multiple frames of video images and the three-dimensional article model into an article object tracking model, acquiring the article object position of the article object to be tracked in each frame of video image through the article object tracking model; and obtaining the object track of the article object to be tracked based on the image sequence of the multiple frames of video images and the article object position of the article object to be tracked in each frame of video image. The method can introduce the three-dimensional model of the article object to realize tracking of the article object, and can further improve the tracking accuracy of the article object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer device, storage medium, and computer program product for tracking objects. Background Technology

[0002] With the development of computer technology, a method for tracking objects in enclosed spaces, such as bank vaults, has emerged. This method involves setting up video recording devices, such as cameras, in the enclosed space to track and monitor objects within the target space.

[0003] Traditional methods for tracking and monitoring objects in enclosed spaces typically rely on background models to detect object movement. An alarm is triggered when an object disappears from the monitoring range or is not returned within a specified timeframe. Furthermore, object tracking is usually based on 2D images of the object, which is susceptible to issues such as occlusion, overlap, or interference from nearby objects, leading to missed or false alarms. Therefore, current object tracking methods generally have low accuracy. Summary of the Invention

[0004] Therefore, it is necessary to provide an object tracking method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the tracking accuracy of object tracking in response to the above-mentioned technical problems.

[0005] Firstly, this application provides a method for tracking object items, the method comprising:

[0006] Acquire multiple frames of video images of the object to be tracked captured by a video image acquisition device in a confined space, as well as a three-dimensional object model corresponding to the object to be tracked that is pre-constructed;

[0007] The multi-frame video images and the three-dimensional object model are input into the object tracking model, and the position of the object to be tracked in each frame of the video image is obtained through the object tracking model.

[0008] Based on the image order of the multi-frame video images and the position of the object to be tracked in each frame of the video images, the object trajectory of the object to be tracked is obtained.

[0009] In one embodiment, the step of inputting the multi-frame video images and the 3D object model into an object tracking model, and obtaining the object object position of the tracked object in each frame of the video images through the object tracking model, includes: inputting the current video image frame into a first branch of the object tracking model, and obtaining the image frame features corresponding to the current video image frame through the first branch; the current video image frame is any one of the multi-frame video images; inputting the 3D object model into a second branch of the object tracking model to obtain the 2D image features corresponding to the 3D object model; performing an association operation on the 2D image features and the image frame features through an attention mechanism to obtain associated features; and obtaining the object object position of the tracked object in the current video image frame based on the associated features.

[0010] In one embodiment, the first branch includes: a pose estimation module; the pose estimation module is used to obtain pose parameters corresponding to the current video image frame; the step of inputting the three-dimensional object model into the second branch of the object tracking model to obtain two-dimensional image features corresponding to the three-dimensional object model includes: inputting the three-dimensional object model and the pose parameters into the second branch, using the second branch to perform direct linear transformation processing on the pose parameters and the three-dimensional object model to obtain a two-dimensional representation image corresponding to the three-dimensional object model; and performing feature extraction on the two-dimensional representation image to obtain the two-dimensional image features.

[0011] In one embodiment, the number of the tracked object objects is multiple; the number of the three-dimensional object models is multiple, each corresponding to one of the tracked object objects; each of the three-dimensional object models is associated with a three-dimensional model identifier; the step of inputting the multi-frame video images and the three-dimensional object models into the object tracking model, and obtaining the object object position of the tracked object in each frame of the video images through the object tracking model, includes: inputting the multi-frame video images and the multiple three-dimensional object models into the object tracking model, and obtaining the position of each of the three-dimensional object models in each frame of the video images through the object tracking model. The corresponding object region; based on the object region, the position of the object corresponding to each three-dimensional object model in each frame of video image is obtained, and the three-dimensional model identifier associated with each three-dimensional object model and the association relationship between the positions of each object are constructed; the object trajectory of the object to be tracked is obtained based on the image order of the multiple frames of video images and the object position of the object to be tracked in each frame of video image, including: based on the image order of the multiple frames of video images and the object position associated with each three-dimensional model identifier in each frame of video image, the object trajectory of each object to be tracked is obtained.

[0012] In one embodiment, before acquiring multiple frames of video images of the object to be tracked captured by the video image acquisition device in the enclosed space, and the three-dimensional object model corresponding to the object to be tracked pre-constructed, the method further includes: before the object to be tracked enters the enclosed space, acquiring multiple images of the object to be tracked corresponding to different shooting perspectives; inputting each of the object images into the three-dimensional modeling model, and acquiring image features corresponding to each of the object images through the three-dimensional modeling model; using the image features to perform mesh transformation and image pooling processing on a pre-set initial ellipsoidal mesh to obtain the three-dimensional object model corresponding to the object to be tracked.

[0013] In one embodiment, the step of inputting each of the object images into a 3D modeling model and obtaining image features corresponding to each of the object images through the 3D modeling model includes: inputting each of the object images into a residual network in the 3D modeling model and obtaining a first image feature corresponding to each of the object images through the residual network; inputting each of the first image features into a visual converter network in the 3D modeling model to obtain a second image feature corresponding to each of the object images; and using each of the first image features and each of the second image features as the image features corresponding to each of the object images.

[0014] In one embodiment, the step of using the image features to perform mesh transformation and image pooling on a pre-set initial ellipsoidal mesh to obtain a 3D object model corresponding to the object to be tracked includes: stitching each of the first image features to obtain stitched image features; using the stitched image features to perform mesh transformation and image pooling on the initial ellipsoidal mesh to obtain an initial 3D object model; obtaining the current second image feature corresponding to the current object image; the current object image is any one of multiple object images corresponding to different shooting perspectives; using the current second image feature to perform mesh transformation and image pooling on the initial 3D object model to obtain a new initial 3D object model, and returning to execute the step of obtaining the current second image feature corresponding to the current object image until the current object image is the last object image; and using the initial 3D object model as the 3D object model corresponding to the object to be tracked.

[0015] In one embodiment, after obtaining the object trajectory of the object to be tracked, the method further includes: performing alarm processing in the enclosed space when the object trajectory meets preset alarm conditions; wherein the preset alarm conditions include: the distance between the current position of the object to be tracked and the initial position where the object to be tracked was placed in the enclosed space is greater than a preset distance threshold; or the object trajectory indicates that the object to be tracked has left the enclosed space and is outside the shooting range of the video image acquisition device.

[0016] Secondly, this application also provides an object tracking device, the device comprising:

[0017] The object acquisition module is used to acquire multiple frames of video images of the object to be tracked captured by a video image acquisition device in a closed space, as well as a three-dimensional object model corresponding to the object to be tracked that is pre-constructed.

[0018] The object location acquisition module is used to input the multi-frame video images and the three-dimensional object model into the object tracking model, and obtain the object location of the object to be tracked in each frame of video images through the object tracking model.

[0019] The object trajectory acquisition module is used to obtain the object trajectory of the object to be tracked based on the image order of the multi-frame video images and the position of the object in each frame of the video images.

[0020] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0021] Acquire multiple frames of video images of the object to be tracked captured by a video image acquisition device in a confined space, as well as a three-dimensional object model corresponding to the object to be tracked that is pre-constructed;

[0022] The multi-frame video images and the three-dimensional object model are input into the object tracking model, and the position of the object to be tracked in each frame of the video image is obtained through the object tracking model.

[0023] Based on the image order of the multi-frame video images and the position of the object to be tracked in each frame of the video images, the object trajectory of the object to be tracked is obtained.

[0024] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0025] Acquire multiple frames of video images of the object to be tracked captured by a video image acquisition device in a confined space, as well as a three-dimensional object model corresponding to the object to be tracked that is pre-constructed;

[0026] The multi-frame video images and the three-dimensional object model are input into the object tracking model, and the position of the object to be tracked in each frame of the video image is obtained through the object tracking model.

[0027] Based on the image order of the multi-frame video images and the position of the object to be tracked in each frame of the video images, the object trajectory of the object to be tracked is obtained.

[0028] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0029] Acquire multiple frames of video images of the object to be tracked captured by a video image acquisition device in a confined space, as well as a three-dimensional object model corresponding to the object to be tracked that is pre-constructed;

[0030] The multi-frame video images and the three-dimensional object model are input into the object tracking model, and the position of the object to be tracked in each frame of the video image is obtained through the object tracking model.

[0031] Based on the image order of the multi-frame video images and the position of the object to be tracked in each frame of the video images, the object trajectory of the object to be tracked is obtained.

[0032] The aforementioned object tracking method, apparatus, computer equipment, storage medium, and computer program product acquire multiple frames of video images of the object to be tracked captured by a video image acquisition device in a confined space, as well as a pre-constructed three-dimensional object model corresponding to the object to be tracked; input the multiple frames of video images and the three-dimensional object model into an object tracking model, and obtain the position of the object to be tracked in each frame of video images through the object tracking model; based on the image order of the multiple frames of video images and the position of the object to be tracked in each frame of video images, obtain the object trajectory of the object to be tracked. This application uses a video image acquisition device to capture multiple frames of video images of an object to be tracked in a confined space, along with a pre-constructed 3D object model. These multiple frames of video images and the 3D object model can then be input into an object tracking model. The position of the object to be tracked in each frame of the video image is obtained through this model. Based on the sequence of the video images and the position of the object in each frame, the trajectory of the object can be obtained. Compared to existing technologies that track objects based on 2D images, this application improves tracking accuracy by introducing a 3D model of the object. Attached Figure Description

[0033] Figure 1 This is a flowchart illustrating an item object tracking method in one embodiment;

[0034] Figure 2 This is a schematic diagram of the process for obtaining the location of an item object to be tracked in one embodiment;

[0035] Figure 3 This is a flowchart illustrating the process of obtaining two-dimensional image features corresponding to a three-dimensional object model in one embodiment.

[0036] Figure 4 This is a flowchart illustrating the process of constructing a 3D object model in one embodiment;

[0037] Figure 5 This is a flowchart illustrating the process of obtaining a 3D object model corresponding to the object to be tracked in one embodiment;

[0038] Figure 6 This is a schematic diagram illustrating the process of creating a 3D model of an object using images from multiple perspectives in one embodiment.

[0039] Figure 7 This is a flowchart illustrating the process of using a 3D model of an object for trajectory tracking in one embodiment.

[0040] Figure 8 This is a structural block diagram of an object tracking device in one embodiment;

[0041] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0043] In one embodiment, such as Figure 1 As shown, an object tracking method is provided. This embodiment illustrates the method applied to a server, but it is understood that the method can also be applied to a terminal, or to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0044] Step S101: Acquire multiple frames of video images of the object to be tracked captured by a video image acquisition device in a closed space, as well as a three-dimensional object model corresponding to the object to be tracked that is pre-constructed.

[0045] A confined space can refer to a sealed space used to store objects, such as a bank vault or safe. Video image acquisition equipment is set up in a confined space to capture video of the objects within it for tracking purposes. This equipment can be a camera installed in the confined space for monitoring objects; there can be one or more cameras. The object to be tracked refers to the object placed in the confined space that needs to be tracked. This object can have a pre-built 3D model, for example, a 3D model of the object can be built before it enters the confined space.

[0046] In this embodiment, when the server tracks objects in a confined space, it can first use a video image acquisition device, such as a camera, set up in the confined space to capture multiple frames of video images containing the objects to be tracked in the confined space. It can also obtain a pre-built three-dimensional model of the object to be tracked, i.e., a three-dimensional object model, stored in the server.

[0047] Step S102: Input the multi-frame video images and the 3D object model into the object tracking model, and obtain the position of the object to be tracked in each frame of the video images through the object tracking model.

[0048] The object location refers to the position information of the object to be tracked in each frame of the video image. The object tracking model is a neural network model used to identify the object location of the object to be tracked, thereby realizing object tracking. This model can accurately identify the two-dimensional image representation of the object to be tracked in the video image in multiple frames by recognizing the two-dimensional image representation corresponding to the three-dimensional object model, and thus determine the location of the object to be tracked. Therefore, even if the object is occluded, overlapped, or interfered with by nearby objects, the object tracking model can still accurately identify the location of the object to be tracked based on the two-dimensional representation of the three-dimensional model.

[0049] Specifically, after the server obtains multiple frames of video images of the object to be tracked, as well as a 3D model of the object, it can input both the multiple frames of video images and the 3D model into the object tracking model. The object tracking model then obtains the position information of the object to be tracked in each frame of the video image, thus obtaining the position of the object in each frame of the video image.

[0050] Step S103: Based on the image order of multiple video frames and the position of the object to be tracked in each video frame, the object trajectory of the object to be tracked is obtained.

[0051] Image order refers to the temporal sequence of the video frames. After obtaining the position of the object to be tracked in each frame of the video image, the server can determine the displacement trajectory of the object to be tracked, i.e., the object trajectory, according to the image order. For example, if the multiple video frames are in the order of Image 1-Image 2-Image 3, and the position of the object to be tracked in each frame is: position A in Image 1, position B in Image 2, and position C in Image 3, then the object trajectory is position A-position B-position C.

[0052] In the aforementioned object tracking method, multiple frames of video images of the object to be tracked, captured by a video image acquisition device in a confined space, and a pre-constructed 3D object model corresponding to the object to be tracked are obtained. The multiple frames of video images and the 3D object model are input into an object tracking model, which obtains the position of the object in each frame of the video image. Based on the image order of the multiple frames of video images and the position of the object in each frame, the trajectory of the object to be tracked is obtained. This application captures multiple frames of video images of the object to be tracked in a confined space using a video image acquisition device, along with a pre-constructed 3D object model. These images and the 3D object model can then be input into an object tracking model, which obtains the position of the object in each frame of the video image. Based on the order of the video images and the position of the object in each frame, the trajectory of the object can be obtained. Compared to existing technologies that track objects based on 2D images, this application improves tracking accuracy by introducing a 3D model of the object.

[0053] In one embodiment, such as Figure 2 As shown, step S102 may further include:

[0054] Step S201: Input the current video image frame into the first branch of the object tracking model, and use the first branch to find the image frame features corresponding to the current video image frame; the current video image frame is any one of the multiple video images.

[0055] The first branch refers to the object tracking model branch used to extract video image frames. In this embodiment, the object tracking model can include two branches, which are used to extract features from multiple video images and the 3D object model, respectively. The current video image frame refers to any one of the multiple video images. The server can take any one of the multiple video images as the current video image frame and input the current video image frame into the first branch of the object tracking model. The first branch extracts features from the current video image frame to obtain the image frame features corresponding to the current video image frame.

[0056] Step S202: Input the 3D object model into the second branch of the object tracking model to obtain the 2D image features corresponding to the 3D object model.

[0057] Two-dimensional image features refer to the image features of the two-dimensional representation corresponding to the three-dimensional object model. The second branch of the object tracking model is a neural network branch used to extract features from the three-dimensional object model. This branch can first convert the three-dimensional object model into a two-dimensional image representation, thereby extracting the image features corresponding to the two-dimensional image representation as the two-dimensional image features corresponding to the three-dimensional object model.

[0058] Step S203: The two-dimensional image features and image frame features are correlated using an attention mechanism to obtain correlated features;

[0059] Step S204: Based on the association features, obtain the position of the object to be tracked in the current video image frame.

[0060] The attention mechanism refers to the attention mechanism, where the server can associate the two-dimensional image features obtained in step S202 with the image frame features obtained in step S201 using the attention mechanism to obtain associated features. Subsequently, the object tracking model can use these associated features to determine the position of the object in the current video image frame.

[0061] In this embodiment, the server can extract the image frame features corresponding to the current video image frame and the two-dimensional image features corresponding to the three-dimensional object model through the first branch and the second branch, respectively, and associate the above features through the attention mechanism, so that the object tracking model can identify the position of the object to be tracked in the current video image frame, thereby further improving the accuracy of object position recognition.

[0062] Furthermore, the first branch includes: a pose estimation module; the pose estimation module is used to obtain the pose parameters corresponding to the current video image frame; such as... Figure 3 As shown, step S202 may further include:

[0063] Step S301: Input the 3D object model and pose parameters into the second branch, and use the second branch to perform direct linear transformation on the pose parameters and the 3D object model to obtain the 2D representation image corresponding to the 3D object model.

[0064] The pose estimation module is mainly used to determine the pose information of the video image acquisition device that captures the current video image frame. In this embodiment, the server can input the current video image frame into the pose estimation module in the first branch, and the pose estimation module can obtain the pose parameters of the video image acquisition device that captures the current video image frame. The pose parameters can also be input into the second branch of the object tracking model, so as to perform direct linear transformation processing on the pose parameters and the three-dimensional object model to obtain the two-dimensional representation image corresponding to the three-dimensional object model.

[0065] For example, the first branch may include a PNP pose estimation module. By inputting the current video image frame into the PNP pose estimation module in the first branch, the PNP pose estimation algorithm can be used to extract the pose parameter information of the current video image frame. The pose parameter information is then input into the direct linear transformation module (DLT) in the second branch. The DLT module performs direct linear transformation on the pose parameters and the input 3D object model to obtain the 2D representation image corresponding to the 3D object model.

[0066] Step S302: Extract features from the two-dimensional representation image to obtain two-dimensional image features.

[0067] After obtaining the two-dimensional representation image corresponding to the three-dimensional object model, feature extraction can be performed on the two-dimensional representation image. For example, the denseblock module and transformer module in the second branch can be used to extract image features, thereby obtaining the corresponding two-dimensional image features.

[0068] In this embodiment, the server can first obtain the pose parameters corresponding to the current video image frame, and then use the pose parameters and the three-dimensional object model to perform a direct linear transformation, thereby obtaining the two-dimensional image features corresponding to the three-dimensional object model, which further improves the accuracy of two-dimensional image feature extraction.

[0069] In one embodiment, the number of items to be tracked is multiple; the number of 3D item models is multiple, each corresponding to a multiple item to be tracked; each 3D item model is associated with a 3D model identifier; step S102 may further include: inputting multiple frames of video images and multiple 3D item models into the item object tracking model, obtaining the item object region corresponding to each 3D item model in each frame of video images through the item object tracking model; obtaining the item object position corresponding to each 3D item model in each frame of video images based on the item object region, and constructing the 3D model identifier associated with each 3D item model, as well as the association relationship between the positions of each item object; step S103 may further include: obtaining the object trajectory of each item to be tracked based on the image order of the multiple frames of video images and the item object positions associated with each 3D model identifier in each frame of video images.

[0070] The object region represents the area where each tracked object is located in each frame of video image, while the 3D model identifier is the identification information used to identify the 3D object model corresponding to different tracked object objects. In this embodiment, there can be multiple tracked object objects, meaning multiple object objects can be tracked simultaneously, and different tracked object objects can correspond to different 3D object models. During the tracking of multiple tracked object objects, the server can input the 3D object models of the multiple tracked object objects and multiple frames of video images into the object tracking model. Then, the object tracking model can obtain the object region corresponding to each 3D object model in each frame of video image, thereby further determining the position of the object corresponding to each 3D object model.

[0071] Then, the server can construct the association between the 3D model identifiers corresponding to each 3D object model and the object positions corresponding to each 3D object model in each of the above video frames. In this way, the trajectory corresponding to each 3D model identifier can be determined according to the image order of the multiple video frames and the object positions corresponding to each 3D model identifier, and then the object trajectory of each object to be tracked can be determined.

[0072] For example, the tracked object objects can include: object 1, object 2, and object 3, which correspond to 3D object model 1, 3D object model 2, and 3D object model 3, respectively. Each 3D model can be identified by identifier 1, identifier 2, and identifier 3. Then, the server can input 3D object model 1, 3D object model 2, and 3D object model 3 into the object tracking model, and simultaneously input multiple frames of video images, namely video image A, video image B, and video image C, into the object tracking model. This allows the acquisition of each 3D object model at the position of each frame of video image, and simultaneously establishes a correspondence between each 3D model identifier and the corresponding position in each frame of video image. That is, if 3D object model 1 is located at position A1 in video image A, then an association relationship is established between identifier 1 and position A1. Similarly, if 3D object model 2 is located at position A2 in video image A, and 3D object model 3 is located at position A3 in video image A, then an association relationship is established between identifier 2 and position A2, and between identifier 3 and position A3. The above method can establish the association between each 3D model identifier and its position in each frame of video image. Then, this association can be used to determine the trajectory of each tracked object. For example, for object 1, its associated 3D model identifier is identifier 1. The server can determine the position information associated with identifier 1 in each frame of video image, which can be position A1, position B1, and position C1. Thus, based on the image order of multiple frames of video image, the trajectory of object 1 can be determined, which can be position A1-position B1-position C1. Similarly, the position information associated with identifier 2 in each frame of video image can also be obtained to obtain the trajectory of object 2.

[0073] In this embodiment, by setting corresponding 3D model identifiers for multiple 3D object models, the object positions in each frame of video image associated with the 3D model identifiers can be used to simultaneously track the object trajectories of multiple object objects to be tracked, thereby further improving the efficiency of multi-object tracking.

[0074] In one embodiment, such as Figure 4 As shown, before step S101, the following may also be included:

[0075] Step S401: Before the object to be tracked enters the enclosed space, acquire multiple images of the object to be tracked corresponding to different shooting perspectives.

[0076] The images of the object can be taken before the object to be tracked enters the enclosed space, each corresponding to a different shooting angle. For example, for a certain object, before entering the enclosed space, such as before the object enters the vault, the object can be photographed from different shooting angles to obtain multiple images of the object corresponding to different shooting angles.

[0077] Step S402: Input the images of each object into the 3D modeling model, and obtain the image features corresponding to each object image through the 3D modeling model.

[0078] The 3D modeling model can be a pre-trained neural network model used for 3D modeling. After the server obtains images of an object to be tracked from different visual angles, it can input each object image into the 3D modeling model, which will then extract the image features corresponding to each object image.

[0079] Step S403: Using image features, perform mesh transformation and image pooling on the pre-set initial ellipsoidal mesh to obtain the three-dimensional object model corresponding to the object to be tracked.

[0080] After obtaining the image features, the image features can be used to perform mesh transformation and image pooling on the pre-set initial ellipsoidal mesh. For example, the mesh transformation and image pooling of pixel2mesh can be used to gradually optimize the initial ellipsoidal mesh, and finally form the three-dimensional object model corresponding to the object to be tracked.

[0081] In this embodiment, the terminal can also obtain multiple images of the object to be tracked corresponding to different shooting perspectives before the object enters the enclosed space. This allows the terminal to use the 3D modeling model to obtain the image features corresponding to the object images, and then perform mesh transformation and image pooling on the initialized ellipsoidal mesh, thereby further improving the modeling accuracy of the 3D object model.

[0082] Further, step S402 may further include: inputting the images of each object into the residual network in the 3D modeling model, obtaining the first image features corresponding to each object image through the residual network; inputting each first image feature into the visual converter network in the 3D modeling model, obtaining the second image features corresponding to each object image; and using each first image feature and each second image feature as the image features corresponding to each object image.

[0083] In this embodiment, the image features may include a first image feature and a second image feature. The first image feature is obtained by directly extracting features from the image of the object, i.e., shallow image features. The second image feature is obtained by further extracting features from the first image feature, i.e., deep image features. The neural network feature extraction layers used for the first image feature and the second image feature are also different. The first image feature is extracted through a residual network, i.e., resblock, while the second image feature is extracted through a visual transducer network, i.e., VIT network.

[0084] Specifically, the server can input each object image into the residual network in the 3D modeling model, and the residual network can extract the first image features of each object image. Then, the server can further input each first image feature into the visual converter network in the 3D modeling model, and the visual converter network can further encode the first image features to obtain the second image features corresponding to each object image. Thus, the first image features and the second image features are used as the image features corresponding to each object image.

[0085] In this embodiment, the server can also use the residual network in the 3D modeling model to extract the first image feature corresponding to each object image, and use the visual converter network in the 3D modeling model to further encode the first image feature to obtain the second image feature corresponding to each object image. Thus, the first image feature and the second image feature are used as the image feature corresponding to each object image. This method can further improve the accuracy of feature extraction of object images.

[0086] In addition, such as Figure 5 As shown, step S403 may further include:

[0087] Step S501: Perform splicing processing on each first image feature to obtain spliced ​​image features;

[0088] Step S502: Use the features of the stitched image to perform mesh transformation and image pooling on the initialized ellipsoidal mesh to obtain the initial three-dimensional object model.

[0089] The stitched image features are image features obtained by stitching together each first image feature. For example, this can be achieved through a concatenating operation. After the server obtains the first image features corresponding to each item image, it can perform a concatenating operation on each first image feature to obtain the stitched image features. Afterward, the server can also use the stitched image features to perform mesh transformation and image pooling on the initial ellipsoidal mesh to obtain the initial 3D item model.

[0090] Step S503: Obtain the current second image features corresponding to the current object image; the current object image is any one of multiple object images corresponding to different shooting views;

[0091] Step S504: Use the current second image features to perform mesh transformation and image pooling on the initial 3D object model to obtain a new initial 3D object model, and return to execute step S503 until the current object image is the last object image.

[0092] Step S505: Use the initial 3D object model as the 3D object model corresponding to the object to be tracked.

[0093] The current item image refers to any one of multiple item image images, while the current second image feature is the second image feature corresponding to the current item image. After obtaining the initial 3D item model, the server can obtain the second image feature of the current item image according to the order of the item images. Then, it can use the second image feature to perform mesh transformation and image pooling on the initial 3D item model to update the initial 3D item model. The server can then obtain the second image feature of a new current item image according to the order of the item images again to update the initial 3D item model again, until the current item image used is the last item image. The last updated initial 3D item model is then used as the 3D item model corresponding to the item to be tracked.

[0094] For example, an object image can contain object image A, object image B, and object image C. After obtaining the second image features of each object image and the initial 3D object model obtained by using the stitched image features, the initial 3D object model can first be updated by performing mesh transformation and image pooling processing on the initial 3D object model using the second image features corresponding to object image A. Then, the initial 3D object model can be updated again using the second image features of object images B and C, thus obtaining the final 3D object model.

[0095] In this embodiment, the first image features can be stitched together first, and the stitched image features can be used to perform mesh transformation and image pooling on the initial ellipsoidal mesh to obtain an initial three-dimensional object model. Then, the second image features corresponding to each object image can be used to iteratively update the initial three-dimensional object model to generate the final three-dimensional object model. This method can further improve the accuracy of the obtained three-dimensional object model.

[0096] In one embodiment, after step S103, the method may further include: performing alarm processing in a confined space if the object trajectory meets preset alarm conditions; wherein the preset alarm conditions include: the distance between the current position of the object to be tracked and the initial position of the object to be tracked in the confined space is greater than a preset distance threshold; or the object trajectory indicates that the object to be tracked has left the confined space and is outside the shooting range of the video image acquisition device.

[0097] The preset alarm conditions can be pre-defined alarm conditions. In this embodiment, the enclosed space can be a space used to store objects, such as a vault. Therefore, it is necessary to constantly track the objects in the enclosed space to prevent their loss and to issue an alarm when an object may be lost. That is, an alarm is triggered in the enclosed space when the object's trajectory meets the preset alarm conditions. For example, the alarm condition could be that the distance between the current position of the object to be tracked and its initial position in the enclosed space is greater than a preset distance threshold. In other words, when the distance between the current position of the object to be tracked and its initial position in the enclosed space is large, it indicates that the object to be tracked has moved far, and an alarm is triggered. Alternatively, the object's trajectory could indicate that the object to be tracked has left the enclosed space and is outside the shooting range of the video image acquisition device. In other words, it indicates that the object to be tracked may leave the monitorable range of the enclosed space, and an alarm is triggered, thereby ensuring the safety of the objects.

[0098] In this embodiment, after obtaining the object trajectory of the object to be tracked, the server can also determine whether the alarm conditions are met based on the object trajectory. If the conditions are met, an alarm can be triggered. The alarm conditions include situations such as the current location of the object to be tracked being far from its initial placement location, or the object to be tracked leaving the monitoring range. This method can further ensure the safety of the object to be tracked when placed in a confined space.

[0099] In one embodiment, a method for tracking items in a bank vault is also provided. This method generates a 3D model of the items by taking multi-view photos of the items to be entered into the vault before they enter. The photos and 3D model are then input into a pre-trained multi-object tracking model to record the item's movement trajectory. The item's location can then be monitored based on the movement trajectory and the item's position. Specifically, this may include the following steps:

[0100] Step 1: Before the item enters the bank vault, take multi-view images of the item and generate a corresponding 3D model.

[0101] Among them, the images of the items can be obtained by taking pictures with camera equipment. Specifically, before an item enters the vault, the item can be photographed from multiple visual angles using camera equipment to obtain the multi-view images of the item.

[0102] Afterwards, a 3D mesh model of the item can be generated; the specific process is as follows: Figure 6 As shown, features can be extracted from multi-view object images using ResBlock, then the extracted features can be further encoded using the VIT (vision transformer) transformer, and then the initialized ellipsoidal mesh can be gradually optimized using pixel2mesh mesh transformation and image pooling, finally resulting in a 3D mesh model.

[0103] Step 2: Obtain video images of the vault and determine the trajectory of the item in the vault from the video images based on the 3D model of the item.

[0104] This involves multi-target tracking of items in the vault, assigning REIDs to items in multiple vaults, and tracking them. The specific process is as follows: Figure 7 As shown, target recognition can be performed based on the video sequence to be matched and the 3D model features of the object. The video frames and 3D model features of multiple cameras at the same time are processed each time to obtain the position at each time.

[0105] Multi-object tracking is essentially tracking based on the detection of multiple objects. A REID is assigned to each object in the first frame of the image. Multi-object re-identification is then performed on each frame. The 2D image of a specific object in each video frame has a certain similarity to its 3D model. Therefore, both the surveillance video frame and the 3D model are input into the network. Pose is obtained through PNP calculation of the surveillance video. The pose parameters and the 3D network model are then linearly transformed to obtain the corresponding 2D representation of the 3D model. After feature extraction, an attention operation is performed on the features of the target 2D graphic in the video frame to obtain the relevant features. Then, an IoU extraction is performed directly using a convolutional-transformer structure to obtain the positional information, similar to DETR.

[0106] Step 3: Track and monitor the item based on its movement trajectory and the pre-set location information.

[0107] The pre-set location information can be the pre-positioned location of items in the vault. After obtaining the movement trajectory of the items in the vault, the vault items can be monitored based on the aforementioned movement trajectory and location information. For example, alarms can be issued when the movement trajectory is too large or when the items completely leave the monitoring area.

[0108] This embodiment allows for multi-angle photography of items, 3D modeling of items, and then the use of cameras with multiple perspectives to perform multi-target, multi-view target tracking of items. By combining the global dependency extraction advantage of the transformer, the item tracking function can be optimized, thereby further improving the security of vault item storage.

[0109] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0110] Based on the same inventive concept, this application also provides an object tracking device for implementing the object tracking method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more object tracking device embodiments provided below can be found in the limitations of the object tracking method described above, and will not be repeated here.

[0111] In one embodiment, such as Figure 8 As shown, an object tracking device is provided, including: an object acquisition module 801, an object location acquisition module 802, and an object trajectory acquisition module 803, wherein:

[0112] The object acquisition module 801 is used to acquire multiple frames of video images of the object to be tracked captured by the video image acquisition device in the enclosed space, as well as the three-dimensional object model corresponding to the object to be tracked that is pre-constructed.

[0113] The object location acquisition module 802 is used to input multi-frame video images and a 3D object model into the object tracking model, and obtain the object location of the object to be tracked in each frame of video images through the object tracking model.

[0114] The object trajectory acquisition module 803 is used to obtain the object trajectory of the object to be tracked based on the image order of multiple video images and the position of the object in each video image.

[0115] In one embodiment, the object location acquisition module 802 is further configured to input the current video image frame into the first branch of the object tracking model, and obtain the image frame features corresponding to the current video image frame through the first branch; the current video image frame is any one of multiple video images; input the three-dimensional object model into the second branch of the object tracking model to obtain the two-dimensional image features corresponding to the three-dimensional object model; perform an association operation on the two-dimensional image features and the image frame features through an attention mechanism to obtain associated features; and obtain the object location of the object to be tracked in the current video image frame based on the associated features.

[0116] In one embodiment, the first branch includes: a pose estimation module; the pose estimation module is used to obtain the pose parameters corresponding to the current video image frame; the object position acquisition module 802 is further used to input the three-dimensional object model and the pose parameters into the second branch, and use the second branch to perform direct linear transformation processing on the pose parameters and the three-dimensional object model to obtain a two-dimensional representation image corresponding to the three-dimensional object model; and to perform feature extraction on the two-dimensional representation image to obtain two-dimensional image features.

[0117] In one embodiment, the number of tracked object objects is multiple; the number of 3D object models is multiple, each corresponding to a multiple tracked object object; each 3D object model is associated with a 3D model identifier; the object position acquisition module 802 is further used to input multiple frames of video images and multiple 3D object models into the object tracking model, and through the object tracking model, obtain the object object region corresponding to each 3D object model in each frame of video images; based on the object object region, obtain the object object position corresponding to each 3D object model in each frame of video images, and construct the 3D model identifier associated with each 3D object model, as well as the association relationship between the object positions; the object trajectory acquisition module 803 is further used to obtain the object trajectory of each tracked object based on the image order of multiple frames of video images and the object object positions associated with each 3D model identifier in each frame of video images.

[0118] In one embodiment, the object tracking device further includes: a 3D modeling module, used to acquire multiple images of the object to be tracked corresponding to different shooting perspectives before the object to be tracked enters a confined space; input each object image into a 3D modeling model, and obtain image features corresponding to each object image through the 3D modeling model; and use the image features to perform mesh transformation and image pooling processing on a pre-set initial ellipsoidal mesh to obtain a 3D object model corresponding to the object to be tracked.

[0119] In one embodiment, the 3D modeling module is further configured to input the images of each object into a residual network in the 3D modeling model, obtain the first image features corresponding to each object image through the residual network; input each first image feature into a visual converter network in the 3D modeling model to obtain the second image features corresponding to each object image; and use each first image feature and each second image feature as the image features corresponding to each object image.

[0120] In one embodiment, the 3D model modeling module is further configured to stitch together the first image features to obtain stitched image features; perform mesh transformation and image pooling on the initial ellipsoidal mesh using the stitched image features to obtain an initial 3D object model; obtain the current second image features corresponding to the current object image; the current object image is any one of multiple object images corresponding to different shooting views; perform mesh transformation and image pooling on the initial 3D object model using the current second image features to obtain a new initial 3D object model, and return to execute the step of obtaining the current second image features corresponding to the current object image until the current object image is the last object image; and use the initial 3D object model as the 3D object model corresponding to the object to be tracked.

[0121] In one embodiment, the object tracking device further includes: a confined space alarm module, used to perform alarm processing in a confined space when the object trajectory meets preset alarm conditions; wherein the preset alarm conditions include: the distance between the object trajectory representing the current position of the object to be tracked and the initial position of the object to be tracked in the confined space is greater than a preset distance threshold; or the object trajectory representing the object to be tracked leaving the confined space and being captured by the video image acquisition device.

[0122] Each module in the aforementioned object tracking device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0123] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores video image data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements an object tracking method.

[0124] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0125] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0126] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0127] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0128] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0129] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0130] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0131] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for tracking object items, characterized in that, The method includes: Acquire multiple frames of video images of the object to be tracked captured by a video image acquisition device in a confined space, as well as a three-dimensional object model corresponding to the object to be tracked that is pre-constructed; The multi-frame video images and the three-dimensional object model are input into the object tracking model, and the position of the object to be tracked in each frame of the video image is obtained through the object tracking model. Based on the image order of the multi-frame video images and the position of the object to be tracked in each frame of the video images, the object trajectory of the object to be tracked is obtained. The step of inputting the multi-frame video images and the 3D object model into the object tracking model, and obtaining the position of the object to be tracked in each frame of the video images through the object tracking model, includes: The current video image frame is input into the first branch of the object tracking model, and the image frame features corresponding to the current video image frame are obtained through the first branch; the current video image frame is any one of the multiple video images. The three-dimensional object model is input into the second branch of the object tracking model to obtain the two-dimensional image features corresponding to the three-dimensional object model. The two-dimensional image features and the image frame features are correlated using an attention mechanism to obtain correlated features; Based on the association features, the position of the tracked object in the current video image frame is obtained; The first branch includes: a pose estimation module; the pose estimation module is used to obtain the pose parameters corresponding to the current video image frame. The step of inputting the 3D object model into the second branch of the object tracking model to obtain the 2D image features corresponding to the 3D object model includes: The three-dimensional object model and the pose parameters are input into the second branch. The second branch is used to perform a direct linear transformation on the pose parameters and the three-dimensional object model to obtain a two-dimensional representation image corresponding to the three-dimensional object model. Feature extraction is performed on the two-dimensional representation image to obtain the two-dimensional image features.

2. The method according to claim 1, characterized in that, The number of the objects to be tracked is multiple; the number of the 3D object models is multiple, each corresponding to one of the multiple objects to be tracked; each 3D object model is associated with a 3D model identifier; the step of inputting the multiple video frames and the 3D object models into the object tracking model, and obtaining the position of the object to be tracked in each video frame through the object tracking model, includes: The multiple video frames and the multiple 3D object models are input into the object tracking model. The object tracking model is used to obtain the object object region corresponding to each 3D object model in each video frame. Based on the object region, the position of the object corresponding to each three-dimensional object model in each frame of video image is obtained, and the three-dimensional model identifier associated with each three-dimensional object model and the association relationship between the positions of each object are constructed. The process of obtaining the object trajectory of the object to be tracked based on the image order of the multi-frame video images and the position of the object in each frame of the video images includes: Based on the image order of the multi-frame video images and the object positions associated with each 3D model in each frame video image, the object trajectory of each object to be tracked is obtained.

3. The method according to claim 1, characterized in that, Before acquiring multiple frames of video images of the object to be tracked captured by the video image acquisition device in the enclosed space, and the three-dimensional object model corresponding to the object to be tracked pre-constructed, the method further includes: Before the tracked object enters the enclosed space, acquire multiple images of the tracked object corresponding to different shooting perspectives; The images of each of the aforementioned objects are input into a 3D modeling model, and the image features corresponding to each of the aforementioned object images are obtained through the 3D modeling model. Using the image features, a mesh transformation and image pooling are performed on the pre-set initial ellipsoidal mesh to obtain the three-dimensional object model corresponding to the object to be tracked.

4. The method according to claim 3, characterized in that, The step of inputting the images of each of the object objects into a 3D modeling model, and obtaining the image features corresponding to each of the object objects through the 3D modeling model, includes: The images of each of the aforementioned object objects are input into the residual network in the 3D modeling model, and the first image features corresponding to each of the aforementioned object objects are obtained through the residual network. Each of the first image features is input into the visual converter network in the 3D modeling model to obtain the second image features corresponding to each of the object images; Each of the first image features and each of the second image features are used as the image features corresponding to each of the object images.

5. The method according to claim 4, characterized in that, The process of using the image features to perform mesh transformation and image pooling on a pre-set initialized ellipsoidal mesh to obtain a 3D object model corresponding to the object to be tracked includes: The first image features are stitched together to obtain stitched image features. The initial ellipsoidal mesh is transformed and image pooling is performed using the features of the stitched image to obtain an initial three-dimensional object model; Obtain the current second image features corresponding to the current object image; the current object image is any one of multiple object images corresponding to different shooting perspectives; The initial 3D object model is subjected to mesh transformation and image pooling processing using the current second image features to obtain a new initial 3D object model. The process then returns to the step of obtaining the current second image features corresponding to the current object image until the current object image is the last object image. The initial 3D item model is used as the 3D item model corresponding to the item object to be tracked.

6. The method according to claim 1, characterized in that, After obtaining the object trajectory of the item to be tracked, the process further includes: If the object trajectory meets the preset alarm conditions, alarm processing is performed in the enclosed space. The preset alarm conditions include: The object trajectory represents the current position of the tracked item and the distance between the tracked item and its initial position in the enclosed space, which is greater than a preset distance threshold. or The object trajectory represents the shooting range of the video image acquisition device as the object to be tracked leaves the enclosed space.

7. An object tracking device, characterized in that, The device includes: The object acquisition module is used to acquire multiple frames of video images of the object to be tracked captured by a video image acquisition device in a closed space, as well as a three-dimensional object model corresponding to the object to be tracked that is pre-constructed. The object location acquisition module is used to input the multi-frame video images and the three-dimensional object model into the object tracking model, and obtain the object location of the object to be tracked in each frame of video images through the object tracking model. The object trajectory acquisition module is used to obtain the object trajectory of the object to be tracked based on the image order of the multi-frame video images and the position of the object in each frame of the video images. The object location acquisition module is further configured to input the current video image frame into the first branch of the object tracking model, and obtain the image frame features corresponding to the current video image frame through the first branch; the current video image frame is any one of the multiple video images; input the three-dimensional object model into the second branch of the object tracking model to obtain the two-dimensional image features corresponding to the three-dimensional object model; perform an association operation on the two-dimensional image features and the image frame features through an attention mechanism to obtain associated features; and obtain the object object location of the object to be tracked in the current video image frame based on the associated features. The first branch includes a pose estimation module; the pose estimation module is used to obtain the pose parameters corresponding to the current video image frame; the step of inputting the three-dimensional object model into the second branch of the object tracking model to obtain the two-dimensional image features corresponding to the three-dimensional object model includes: inputting the three-dimensional object model and the pose parameters into the second branch, using the second branch to perform direct linear transformation processing on the pose parameters and the three-dimensional object model to obtain the two-dimensional representation image corresponding to the three-dimensional object model; and performing feature extraction on the two-dimensional representation image to obtain the two-dimensional image features.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.