Cross-device object tracking method and system

Through neural network technology, the coordinated tracking of objects between multiple devices is solved, and the problem that object tracking cannot cross the camera range in the prior art is solved, achieving more efficient and accurate object tracking.

CN119963820APending Publication Date: 2025-05-09INVENTEC APPLIANCES (SHANGHAI) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510051844.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing object tracking technologies are difficult to coordinately track objects between multiple devices, especially when the object leaves the camera range, and cannot continue tracking.

Method used

Using neural network technology, continuous frames are captured through the photography module, object characteristics are detected, and object pools are compared with the peripheral devices, matching the same object and giving identifiers. The features of the target object are transmitted to the peripheral device to realize cross-device object tracking.

Benefits of technology

Cooperative object tracking between multiple devices is realized, and the object can continue to track after it leaves one camera range, improving the real-time and accuracy of object tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963820A_ABST
    Figure CN119963820A_ABST
Patent Text Reader

Abstract

The invention provides a cross-device object tracking method and system, and the method comprises the steps: shooting continuous frames through a shooting module of a device or a peripheral device, detecting at least one object in a current frame through a multi-object tracking technology, extracting object features of the object in each frame through a convolutional neural network, and carrying out the tracking of the object in each frame. And comparing the object features with the object features of the previous frame temporarily stored in the original object pool of each device, converting each object feature into an embedded vector by using an encoder of a conversion model, providing a decoder of the conversion model to compare the matched objects of the previous frame and the next frame, and matching the same objects. Next, a target object is determined from the objects, and the object characteristics of the target object are transmitted and temporarily stored to a target object pool of the device and the peripheral device, so that the device and the peripheral device can cooperatively track the target object according to the object characteristics of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an object tracking technology, and more particularly to a cross-device object tracking method and system for collaboratively tracking an object by sharing object features among multiple devices. Background Art

[0002] Current object tracking technologies are all used on fixed cameras, such as traffic cameras installed on roads, which use object recognition technology or subsequent human visual judgment to track specific people or vehicles. When the person or vehicle being tracked is in a place where there is no camera, it cannot be tracked.

[0003] Object tracking technology can also be applied to portable cameras, such as secret recorders. For example, when police and security personnel are chasing suspects, when personnel wearing portable cameras cooperate to track or capture suspicious objects, it is customary to communicate with the central control center to know the location of the suspicious object. However, if there are too many people in this group, it will be difficult for the central control center to communicate with each person individually in real time to hunt down the suspicious object. Summary of the invention

[0004] In view of the problems in the prior art, the purpose of the present invention is to provide a cross-device object tracking method and a cross-device object tracking system implemented by using a neural network, so that multiple devices can collaborate in object tracking.

[0005] According to an embodiment, in the cross-device object tracking method, the method runs in the device and at least one peripheral device, firstly shoots continuous frames through a camera module, and detects the object in the current frame using a multi-object tracking technology, then obtains the object features of the object, and then compares the object features of the previous frame temporarily stored in the original object pool of the device and the peripheral device, and the same object can be matched, that is, the same identifier can be assigned. Then, the target object is determined from the objects, and the object features of the target object are transmitted and temporarily stored in the target object pool set in the memory of the device and its peripheral device. The device and the peripheral device can collaboratively track the target object according to the object features of the target object in the target object pool.

[0006] Furthermore, in the device or peripheral device, a convolutional neural network can be used to extract object features of objects in each frame, and then an encoder of a transformation model can be used to convert each object feature into an embedding vector, so that a decoder of the transformation model can compare the objects that match the previous and next frames.

[0007] Furthermore, when the device obtains the embedding vector of the target object to perform object tracking, the embedding vector of the target object is transmitted to the peripheral device, so that the peripheral device can perform object tracking according to the embedding vector of the target object.

[0008] Furthermore, the device and one or more peripheral devices each implement an edge computing device, and therefore, each computes the convolutional neural network and the transformation model, and can perform operations only on pixels in the edge area of ​​each frame to save computing power.

[0009] Furthermore, the device and the surrounding devices may establish a group according to the correlation of the geographical locations of the devices, so that the device may transmit the object characteristics of the target object to one or more surrounding devices in the same group in a multicast manner.

[0010] In one embodiment, the device and the peripheral devices transmit the object characteristics of the target object via a central control center.

[0011] Furthermore, when the device and at least one of the surrounding devices obtain object features of an object in a current frame, the object features of an object in a previous frame in an original object pool can be compared; when it is determined from the comparison results that there are new object features, they are compared with the object features of the target object in the target object pool. When the target object is matched, it indicates that the target object appears in continuous frames taken by the device or the surrounding device, and is then shared between the device and the surrounding devices.

[0012] According to an embodiment of the cross-device object tracking system, the system may include a central control center and the device and peripheral devices, wherein the device and peripheral devices are connected to the central control center via a channel and run the cross-device object tracking method. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Other features, objectives and advantages of the present invention will become more apparent from a reading of the detailed description of non-limiting embodiments made with reference to the following accompanying drawings.

[0014] Figure 1 A schematic diagram for implementing circuit components in a device for executing a cross-device object tracking method;

[0015] Figure 2 A schematic diagram of an architecture embodiment of a cross-device object tracking system;

[0016] Figure 3 A flowchart of an embodiment of a cross-device object tracking method;

[0017] Figure 4 A flowchart of another embodiment of a cross-device object tracking method;

[0018] Figure 5 A schematic diagram of an image frame for performing edge area operations;

[0019] Figure 6 A diagram of an embodiment of an architecture for tracking objects using embedding vectors; and

[0020] Figure 7A schematic diagram of an embodiment of a neural network architecture for implementing a cross-device object tracking method.

[0021] Reference numerals:

[0022] 10 Network

[0023] 100 Central Control Center

[0024] 120 Devices

[0025] 121 Central Processing Unit

[0026] 123 Image Processing Unit

[0027] 125 Photography Module

[0028] 127 Communication module

[0029] 129 Buffer

[0030] 191 Original Object Pool

[0031] 192 Target Object Pool

[0032] 21 Device 1

[0033] 22 Installation 2

[0034] 23 Installation 3

[0035] 24 Installation 4

[0036] 200 Central Control Center

[0037] 201 Input Interface

[0038] 203 Object Selection Unit

[0039] 205 Object feature vector calculation unit

[0040] 207 Object Vector Multicast Unit

[0041] 209 Output interface

[0042] 400 Central Control Center

[0043] 50 video frames

[0044] 501 Edge Area

[0045] 503 Central Area

[0046] 505 Target

[0047] 601 First Vector

[0048] 602 Second Vector

[0049] 603 Third Vector

[0050] 61a,61b,61c Decoder

[0051] 63a,63b,63c Buffer

[0052] 631a,631b,631c Original object pool

[0053] 633a,633b,633cTarget object pool

[0054] 701 First frame

[0055] 702 Second frame

[0056] 703 Frame 3

[0057] 701' First object marking diagram

[0058] 702' Second object marking diagram

[0059] 703' Third object marking diagram

[0060] 71a,71b,71c Convolutional Neural Network

[0061] 73a,73b,73c Transformation Model Encoder

[0062] 75a,75b,75c Transformation Model Decoder

[0063] 77a,77b,77c Buffer

[0064] 771a,771b,771c Original object pool

[0065] 773a,773b,773cTarget object pool

[0066] 79a,79b,79c Matching object indicator

[0067] A,A' New object

[0068] Steps S301 to S313: Cross-device object tracking process 1

[0069] Steps S401 to S423 Cross-device object tracking process 2 DETAILED DESCRIPTION

[0070] The following is an explanation of the embodiments of the present invention through specific embodiments. Those skilled in the art can understand the advantages and effects of the present invention from the contents disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and the details in this specification can also be modified and changed in various ways based on different viewpoints and applications without departing from the concept of the present invention. In addition, the drawings of the present invention are only simple schematic illustrations and are not depicted according to actual sizes. It is stated in advance. The following embodiments will further explain the relevant technical contents of the present invention in detail, but the disclosed contents are not intended to limit the scope of protection of the present invention.

[0071] It should be understood that, although the terms "first", "second", "third", etc. may be used in this document to describe various components or signals, these components or signals should not be limited by these terms. These terms are mainly used to distinguish one component from another component, or one signal from another signal. In addition, the term "or" used in this document may include any one or more combinations of the associated listed items depending on the actual situation.

[0072] The present invention provides a cross-device object tracking method and system, wherein the cross-device object tracking method is run in a system composed of multiple devices, and the device can be a fixed or mobile electronic device with a camera function, and can be run in a cross-device object tracking system formed by a plurality of fixed or mobile cameras working in collaboration. According to an embodiment, the device running the method and one or more of its peripheral devices each implement an edge computing device, and can each compute a convolutional neural network (CNN) and a transformer model.

[0073] Figure 1 Schematic diagram of circuit components in a device for running a cross-device object tracking method, wherein device 120 is a fixed or mobile device with a photographic function, connected to a central control center 100 via a network 10, and image data or related information generated by device 120 can be transmitted to the central control center 100 in real time for processing, and an object to be tracked is determined from the image data or related information, and then the device 120 or its peripheral devices are notified to achieve the purpose of cross-device object tracking.

[0074] The main circuit of the device 120 includes a central processing unit 121 for performing data processing, which is electrically connected to an image processing unit 123. The image processing unit 123 processes the dynamic image captured by the camera module 125 and generates an image signal after image processing. The central processing unit 121 performs multi-object tracking (MOT) technology to detect objects frame by frame based on image features and obtain object features of each object. In one embodiment, the device 120 operates as an edge computing device, and the central processing unit 121 performs a convolutional neural network (CNN) to extract object features of one or more objects in each frame, and then performs a transformation model (transformer model), wherein the encoder converts each object feature into an embedded vector, so as to provide a decoder in the transformation model to compare the object vectors in the previous and next frames to match the objects.

[0075] The device 120 is provided with a buffer 129 implemented by a storage element. The buffer 129 is divided into an original object pool 191 and a target object pool 192. The original object pool 191 is used to temporarily store the object features of the object detected in the previous frame in the continuous frames, and the target object pool 192 stores the object features of the determined target object. The target object features can be obtained from the central control center 100.

[0076] The central control center 100 can obtain the object features of at least one object in each frame from the device 120, or the embedded vectors of each object converted by the encoder, and transmit them to the central control center 100 via the network 10 by the communication module 127. The central control center 100 determines the target object to be tracked so that it can be broadcast to other peripheral devices.

[0077] According to an embodiment, in a cross-device object tracking method, the object features of at least one object in the current frame are obtained through the central processing unit 121 in the device 120, and the object features of the object in the previous frame in the original object pool 191 are compared to confirm the object and determine whether a new object has entered the image. If there is a new object, an identifier is given to the new object. The target object pool 192 can store the object features of the target object for use in tracking the target object in collaboration with surrounding devices.

[0078] Figure 2 The central control center 200 is connected to multiple devices, such as device 1 21, device 2 22, device 3 23 and device 4 24. Device 1 21 captures continuous frames and detects at least one object from each frame to obtain object features, that is, the communication module thereof transmits the object features to the central control center 200 through the input interface 201.

[0079] According to an embodiment, the object tracking technology running in the central control center 200 shown in the figure is implemented by a tracking converter (trackformer) composed of a convolutional neural network and a conversion model. In another embodiment, the object tracking technology can also be implemented using the object tracking model Yolo (You Only Look Once) established by deep learning. The tracking converter includes multiple functional components such as an object selection unit 203, an object feature vector calculation unit 205, and an object vector multicast unit 207. The object selection unit 203 determines the target object to be tracked from the objects confirmed in each image frame, and uses the object feature vector calculation unit 205 to calculate the target object feature to convert it into an embedded vector of the target object, and then outputs the embedded vector of the target object through the output interface 209 through the object vector multicast unit 207 according to the communication method of the peripheral devices in the group, as shown in the figure, to the device 2 22, the device 3 23 and the device 4 24.

[0080] In this way, the device 2 22, the device 3 23 and the device 4 24 obtain the object features of the target object from the central control center 200, or the embedded vectors, which can be stored in the target object pool in each device for comparing the dynamic images taken by the camera module 125, so as to realize collaborative operation and implement the cross-device object tracking method.

[0081] Each of the devices may be a mobile device, such as a body-worn camera (BWC) or smart glasses, and multiple devices in the same field form a group, so that the central control center 200 can transmit the object features or embedded vectors of the target object to the devices in the group in a multicast manner. According to an embodiment, it can also cooperate with surrounding fixed devices so that multiple devices can collaboratively track the target object according to the object features of the target object.

[0082] like Figure 3 As shown, in a cross-device object tracking method, in a device, continuous frames are obtained in real time by shooting through a photographic module therein (step S301), and a multi-object tracking technology (MOT) detection executed in the device can be used to obtain one or more objects in the current frame (step S303), and object features of each object in each frame can be obtained by image processing technology (step S305), which can be compared with the object features of the previous frame temporarily stored in the original object pool to match the same objects in the previous and next frames, and the same identifier is assigned to the same objects in the previous and next frames (step S307), and each object is tracked according to the identifier.

[0083] Next, the device transmits the object features of one or more objects detected in each frame to the central control center, and the central control center determines the target object from the one or more objects (step S309). For example, the central control center personnel determine the target object to be tracked based on the object image features and object movement trajectory received from the device. The object features of the target object (or converted into an embedded vector) can then be transmitted to at least one peripheral device outside the device in a multicast manner (step S311), so that the device and at least one peripheral device can collaboratively track the target object based on the object features of the target object (step S313).

[0084] Figure 4 The flowchart of another embodiment of the cross-device object tracking method is shown.

[0085] According to the embodiment, the cross-device object tracking method is executed by the terminal device (including the device and the peripheral devices). However, when the computing power of the terminal device is insufficient, the central control center server can calculate the object features based on the image provided by the device. Among them, when the computing power and power of the terminal device are sufficient, the device and the peripheral device executing the cross-device object tracking method can each implement edge computing, and can respectively execute multi-object tracking technology (MOT), and the convolutional neural network and conversion model can implement the multi-object tracking technology architecture, using multiple devices in different positions to collaboratively perform object tracking.

[0086] In the device (or any peripheral device), before object tracking is performed, the image transmission mode is first entered (step S401), the decoder in the conversion model is run (step S403), and then the camera module of the device is used to capture continuous frames (step S405).

[0087] The device runs a convolutional neural network to obtain the image features of each frame and extracts the object features of one or more objects therein (step S407). The object features of each object include one or any combination of the shape, material and color of each object. Afterwards, the encoder of the conversion model running in the device converts each object feature into an embedding vector (step S409), and temporarily stores the object features of at least one object in the previous frame of the continuous frames obtained through the original object pool in each device (such as the embedding vector obtained by calculation) (step S411), so as to provide the decoder of the conversion model to compare the objects matching the previous and next frames, and determine whether it is a new object. The device can transmit the object features obtained from each frame or the embedded vector obtained by conversion to the central control center 400 (step S413), and the central control center 400 determines the target object to be tracked.

[0088] When the device or the peripheral device receives the target object vector from the central control center 400 (step S415), each device enters the object tracking mode (step S417), and places the obtained object features or embedded vectors of the target object into the target object pool of the buffer in the device (step S419). Next, in each device, the object similarity can be calculated frame by frame (step S421) according to the embedded vector of the object in each frame obtained in real time and the embedded vector of the target object temporarily stored in the target object pool, and compared with the pre-set similarity threshold to determine whether the object similarity is greater than (or equal to) this threshold (step S423). After obtaining the similarity between each object and the target object in each frame, it is determined that the object similarity is not greater than or equal to the threshold (no), indicating that the object is not the target object, and the process returns to step S405, and then performs object comparison from the next frame. If it is determined that the object similarity is greater than or equal to the threshold (yes), the target object is matched (step S425), and the purpose of cross-device collaborative tracking of objects is achieved between multiple devices, and the relevant information can be transmitted to the central control center (step S427).

[0089] It is worth mentioning that in step S423 of determining whether the object similarity is greater than or equal to the threshold, when the target object is not matched and it is confirmed that there is no target object feature in the target object pool, the process can return to step S405 to execute the subsequent object tracking program; on the other hand, if the target object is not matched, the process returns to step S421 to continue using the target object features in the target object pool to perform the next frame object similarity comparison.

[0090] For example, when multiple police officers wearing mobile devices are chasing a criminal, they can use multiple camera devices to shoot from different angles, and can be combined with fixed camera devices on site to achieve the purpose of collaborative tracking of the criminal, complementing the shortcomings of each camera device's limited field of view that makes it difficult to track the criminal.

[0091] It is worth mentioning that each mobile device can implement edge computing. However, considering the limited computing resources and power of each device, in order to save energy when performing object tracking, the computing range of each frame can be reduced, for example, only the pixels in the edge area of ​​each frame can be operated. The reason can be referred to Figure 5 A schematic diagram showing an image frame on which edge region operations are performed.

[0092] by Figure 5 Taking the displayed image frame 50 as an example, according to the purpose of detecting the target object 505, when it is necessary to confirm whether the target object enters the range of each frame, the target object will enter the image from the edge of each frame, so the central area calculation can be temporarily excluded and the edge area calculation can be focused on, such as Figure 5As shown, in order to save computing power, the image frame 50 is divided into an edge area 501 and a central area 503, and operations can be performed only on pixels in the edge area 501 of the image frame 50 to detect the target object 505. Or in another embodiment, the computing power ratio of the central area of ​​each image frame can be reduced, for example, by reducing the frame rate of the operation, reducing the number of pixels, etc.

[0093] In addition to the above-mentioned edge area calculations that can reduce computing power requirements, according to another scenario, when the target area approaches the photographic device from a distance, the target object is processed through image processing from unrecognizable object features to recognizable object features (features such as the object's appearance, material, and color, etc.). Therefore, the time interval for device calculation can be designed, and the calculation frequency is increased when the target object gradually approaches the device.

[0094] According to the above embodiment, please refer to Figure 6 A diagram showing an architecture implementation scheme of tracking objects using embedded vectors by an edge computing device or a central control center is shown, wherein a conversion model is run in a device that executes multi-object tracking technology, and a decoder in the conversion model processes object features to obtain an embedded vector, wherein an original object pool and a target object pool are implemented by a buffer, wherein the original object pool is used to store the embedded vector of the object in the previous frame, and the target object pool is used to temporarily store the embedded vector of the target object to be tracked.

[0095] According to the illustrated implementation example, when performing multi-object tracking, the device may receive an embedding vector of a designated target object to be tracked from the central control center, and may perform a normalization operation and store the embedding vector in a target object pool (633a, 633b, and 633c) of the buffer. The decoder in the device is shown as decoders 61a, 61b, and 61c frame by frame, and the buffer is schematically shown as buffers 63a, 63b, and 63c frame by frame, wherein the original object pool stores the embedding vector of the object frame by frame, schematically shown as original object pools 631a, 631b, and 631c, and the target object pool is schematically shown as target object pools 633a, 633b, and 633c frame by frame.

[0096] The device performs object tracking frame by frame, obtains the object features in each frame frame by frame in the device, and converts them into embedding vectors, which are first compared with the embedding vectors of the previous frame objects stored in the original object pool. If there are new embedding vectors that cannot be matched according to the comparison results, they will be compared with the embedding vectors of the target objects in the target object pool, and vector similarity calculations will be used. When the result is higher than the similarity threshold set by the system, it means that the target object newly entering the image is matched, and tracking can be performed and relevant information can be issued.

[0097] like Figure 6As shown, the device obtains the first vector 601 of the object in the first frame through calculation, wherein the decoder 61a of the conversion model takes out the target object embedding vector from the target object pool 633a of the buffer 63a, and calculates the similarity with the first vector 601 to confirm whether the first vector 601 includes the target object or a new object. If it is a new object, an identifier can be assigned; if it is a target object, tracking can be performed. When processing the first frame, the original object pool 631a in the buffer 63a does not store any object embedding vector. At this time, the first vector 601 can be temporarily stored in the original object pool 631b of the buffer 63b to become the comparison data of the next frame.

[0098] Next, the device obtains the second vector 602 of the second frame through calculation, and the decoder 61b obtains the embedding vector of the object from the original object pool 631b and the target object pool 633b of the buffer 63b, that is, first compares the second vector 602 with the object embedding vector of the previous frame to determine whether there is a new vector (new object), and then compares the embedding vector of the target object to confirm whether it is the target object. In order to provide the device with the embedding vector of the next frame to be processed at the next time, the second vector 602 is stored in the original object pool 631c of the buffer 63c.

[0099] Similarly, the device obtains the third vector 603 of the third frame through calculation, and the decoder 61c obtains the object embedding vector of the previous frame from the original object pool 631c of the buffer 63c, and obtains the embedding vector of the target object from the target object pool 633c, so as to compare the third vector 603 with the object embedding vector of the previous frame in the original object pool 631c to determine whether there is a new vector (new object). All new vectors are then compared with the embedding vector of the target object in the target object pool 633c to confirm whether they are the target object, so as to continue tracking the target object. By repeating Figure 6 The displayed process achieves the purpose of tracking multiple objects and target objects.

[0100] Figure 7A schematic diagram of a neural network architecture embodiment for implementing a cross-device object tracking method is provided. The cross-device object tracking method adopts a multi-object tracking (MOT) technology, wherein a tracking transformer model (trackformer) implemented by a convolutional neural network (CNN) and a transformer model (transformer) is adopted. The convolutional neural network is used to extract object features in an image, and then the object features are converted into embedding vectors by an encoder (encoder) in the transformer model. After that, the features of the camera input image and the object features of the previous frame are compared by a decoder (decoder) in the transformer model to match the same object. When there is a matching object, the identifier (ID) of the previous frame is used, and the object can be marked with the same color frame on the identified image to achieve the purpose of object tracking. The object features include the shape, material, and color of the object, etc., and the decoder is provided to compare the object features of the previous and next frames to determine whether the same object appears in the previous and next frames, thereby establishing the object trajectory.

[0101] Figure 7 The state of each component at each time point is displayed frame by frame. According to the exemplary embodiment of the diagram, when the device performs multi-object tracking, in addition to the original object pool 771a in the buffer 77a not storing any embedding vector at the beginning and therefore being empty, the original object pools 771b and 771c in the buffers 77b and 77c are subsequently used to store the embedding vector of the previous frame, and the target object pools 773a, 773b and 773c are used to store the embedding vector of the target object specified by the central control center. This diagram shows that the target object pools 773a, 773b and 773c have the same target object features to be tracked. Further, the diagram also displays the neural network architecture running in the device frame by frame, including the use of specific convolutional neural networks 71a, 71b and 71c, conversion model encoders 73a, 73b and 73c, and conversion model decoders 75a, 75b and 75c.

[0102] During operation, the device receives the embedding vector of the target object from the central control center, and can perform normalization operation (normalization) and store it in the target object pool (773a, 773b and 773c) of the buffer (77a, 77b and 77c). This illustration shows that the original object pool 771a at the beginning does not store the embedding vector, and when the target object tracking is performed frame by frame, the target object pool (773a, 773b and 773c) stores the target object features obtained from the central control center.

[0103] The device first obtains the first frame 701 at the first time, extracts the object features in the first frame 701 through the convolutional neural network 71a, converts the object features into an embedding vector through the conversion model encoder 73a, and provides the conversion model decoder 75a. If multiple objects appear, different identifiers can be assigned to facilitate distinguishing different objects. In this example, the matching object indicator 79a shows three objects appearing in the first object marking image 701' with different cross-section lines. The object embedding vector obtained from the first frame 701 will be temporarily stored in the original object pool 771b for the purpose of matching the target object in the second frame.

[0104] Next, the embedding vectors corresponding to one or more objects determined from the first frame 701 are stored in the original object pool 771b in the buffer 77b, while the target object pool 773b continues to record the embedding vectors of the target objects to be tracked.

[0105] The device obtains a second frame 702 at a second time. Similarly, the convolutional neural network 71b is used to extract object features in the second frame 702, and the object features are converted into an embedding vector through the transformation model encoder 73b. The transformation model decoder 75b is provided to compare the embedding vector in the second frame 702 with the embedding vector of the previous frame (in this case, the first frame 601) stored in the original object pool 771b to match one or more objects in the second frame 702 and determine whether there is a new object.

[0106] In the process of matching objects using the original object pool 771b in the second frame 702, if there is a new object, the conversion model decoder 75b compares the target object embedding vector in the target object pool 773b, confirms the object in the image through similarity, and determines whether there is a target object. Similarly, the matching result is indicated by the matching object indicator 79b to indicate one or more objects determined, which can be visually presented as one or more objects in the second object labeling diagram 702'.

[0107] according to Figure 7 In the example shown, initially three objects appear in the first frame 701, such as the three objects indicated by the first object mark 701', and are respectively displayed in the matching object indicator 79a, and the corresponding three object embedding vectors are stored in the original object pool 771b. Then, the object embedding vector of the second frame 702 is obtained, and the four objects that appear in the second object mark 702' are indicated by different section lines in the matching object indicator 79b. After comparing the object embedding vectors in the original object pool 771b, a new object is found, and the figure shows the new object A indicated in the second frame 702. After the embedding vector calculation, after comparing the target object embedding vector in the target object pool 673b, it can be determined that the new object A' displayed in the second object mark 702' is the target object.

[0108] Similarly, at the second time, the embedding vector obtained from the second frame 702 is stored in the original object pool 771c and provided for comparison with the next frame, and the target object pool 773c still stores the embedding vector of the target object.

[0109] Then, the device obtains the third frame 703 at the third time, extracts the object features in the third frame 703 through the convolutional neural network 71c, converts the object features into an embedding vector through the conversion model encoder 73c, and provides a conversion model decoder 75c to compare the embedding vector in the third frame 703 with the embedding vector of the previous frame stored in the original object pool 771c to match one or more objects in the third frame 703 and determine whether there is a new object; if there is a new object, the conversion model decoder 75c then compares the target object embedding vector in the target object pool 773c to confirm the object in the image through similarity and determine whether there is a target object.

[0110] This legend shows that the third frame 703 has one less object relative to the second frame 702. By comparing the embedding vector of the previous frame stored in the original object pool 771c, it is determined that the third frame 703 includes the new object A determined in the previous frame and two other objects. The matching result uses the matching object indicator 79c to indicate the three objects determined, and is correspondingly presented in a visual manner as three objects in the third object marking diagram 703', including the new object A'.

[0111] By repeating the above process, the new object A determined to be the target object can be continuously tracked, thereby achieving the purpose of tracking the target object using a neural network architecture. Furthermore, according to an embodiment of the cross-device object tracking method proposed by the present invention, each device detects the target object based on the correlation of vector information in the image between the previous and next frames, as well as the position change of the target object frame by frame. In this way, the device will provide the information of the determined target object to the central control center, which will then broadcast the information to other peripheral devices associated with the geographical location, thereby achieving the purpose of collaborative tracking of the target object by multiple devices.

[0112] Among them, the device and one or more peripheral devices establish a group based on the correlation of the geographical locations of each device, so that the device or the central control center transmits the object features of the target object to one or more peripheral devices in the same group in a multicast manner. Each peripheral device implements an edge computing device, calculates a convolutional neural network and a conversion model, and can only perform operations on pixels in the edge area of ​​each frame to achieve the purpose of cross-device object tracking proposed by the present invention.

[0113] Here is an example of multiple devices working together to track a target object. Multiple security personnel in a scene wear mobile devices that can continuously capture images of a scene. Each device operates as an edge computing device, executing multi-object tracking technology in each device to obtain the object features of objects in consecutive frames and convert them into embedding vectors for similarity calculation. Any of the devices transmits the resulting object embedding vector to the central control center, which determines the target object to be tracked, such as a suspicious person.

[0114] When executing the method of tracking an object across devices, the combination of collaborative operations can be determined according to the application scenario, such as a mobile device combined with at least one mobile device, a mobile device combined with at least one fixed device, a fixed device combined with at least one mobile device, or a combination of multiple fixed devices working in collaboration.

[0115] Since each mobile device can only capture and track objects within a limited range, the central control center transmits the embedded vector of the target object to the mobile devices worn by multiple security personnel in the same scene through multicasting. The vector is specifically stored in the target object pool of each device. It can also be transmitted to the memory of the fixed camera device in the scene. The multi-object tracking technology running in multiple devices can collaboratively track suspicious persons with different fields of view. If any device confirms that a suspicious person has been captured, an alarm message can be sent to the central control center.

[0116] In summary, according to the cross-device object tracking method and system described in the above embodiments, multiple devices are used to respectively execute multi-object tracking technology, and the object features of the target object are shared through the central control center. The target object tracking is performed by extracting object features and object similarity comparison through a neural network architecture, thereby achieving collaborative operation to realize the purpose of cross-device object tracking.

[0117] The contents disclosed above are only preferred feasible embodiments of the present invention, and are not intended to limit the scope of the present invention. Therefore, all equivalent technical changes made using the contents of the present invention's specification and drawings are included in the scope of the present invention.

Claims

1. A cross-device object tracking method, which operates in a device and at least one peripheral device in collaboration, wherein the device and one of the at least one peripheral device shoot continuous frames through a camera module, detect at least one object in a current frame using a multi-object tracking technology, and obtain the object features of the at least one object, and compare the object features of the previous frame temporarily stored in an original object pool of the device or one of the at least one peripheral device to match the same object, characterized in that: The method includes: Determine a target object from the at least one object, transmit the object characteristics of the target object and temporarily store them in a target object pool of the device and the at least one peripheral device respectively; and The device and the at least one peripheral device cooperate to track the target object according to at least one object feature of the target object temporarily stored in the respective target object pools.

2. The cross-device object tracking method according to claim 1, characterized in that: In the device or the at least one peripheral device, a convolutional neural network is used to extract the object features of the at least one object in each frame, and each object feature is converted into an embedding vector using an encoder of a conversion model, and a decoder of the conversion model is provided to compare the objects matching the previous and next frames.

3. The cross-device object tracking method according to claim 2, characterized in that: When the device or the at least one peripheral device obtains the embedding vector of the target object, object tracking is performed according to the embedding vector of the target object.

4. The cross-device object tracking method according to claim 3, characterized in that: The device and the at least one peripheral device each implement an edge computing device, each computes the convolutional neural network and the conversion model, and only performs computing on pixels in the edge area of ​​each frame.

5. The cross-device object tracking method according to claim 1, characterized in that: The device and the at least one peripheral device establish a group according to the correlation of the geographical locations of the devices, and the device transmits the object feature of the target object to the one or more peripheral devices in the same group in a multicast manner.

6. The cross-device object tracking method according to any one of claims 1 to 5, characterized in that: The object characteristics of the target object are transmitted between the device and the at least one peripheral device via a central control center.

7. The cross-device object tracking method according to claim 6, characterized in that: When the device and one of the at least one peripheral device obtain the object features of the at least one object in the current frame, the object features of the at least one object in the previous frame in the original object pool are compared; when it is determined from the comparison result that there is a new object feature, it is compared with the object features of the target object in the target object pool, and the target object is matched, indicating that the target object appears in the continuous frames taken by the device or the at least one peripheral device, and is then shared between the device and the at least one peripheral device.

8. A cross-device object tracking system, characterized in that: include: a central control centre; and A device and at least one peripheral device, wherein the device and the at least one peripheral device are connected to the central control center via a channel; The cross-device object tracking methods that are implemented include: The device and one of the at least one peripheral device shoot continuous frames through a camera device, detect at least one object in a current frame using a multi-object tracking technology, and obtain object features of the at least one object; Comparing the object features of the previous frame temporarily stored in an original object pool of the device and one of the at least one peripheral device to match the same object; Transmitting the object characteristics of the at least one object to the central control center, the central control center determines a target object, and transmits the object characteristics of the target object to the device and the at least one peripheral device, and temporarily stores them in a target object pool of each of the device and the at least one peripheral device; and The device and the at least one peripheral device cooperate to track the target object according to the object features of the target object temporarily stored in the respective target object pools.

9. The cross-device object tracking system according to claim 8, characterized in that: In the device or the at least one peripheral device, a convolutional neural network is used to extract the object features of the at least one object in each frame, and each object feature is converted into an embedding vector using an encoder of a conversion model, and a decoder of the conversion model is provided to match the objects of the previous and next frames.

10. The cross-device object tracking system according to claim 9, characterized in that: The central control center enables the device or the at least one peripheral device to obtain the embedding vector of the target object, that is, to perform object tracking according to the embedding vector of the target object.

11. The cross-device object tracking system according to any one of claims 8 to 10, characterized in that: The central control center determines the target object to be tracked, and the device and the at least one peripheral device each temporarily store the object features of the target object transmitted from the central control center through the target object pool.

12. The cross-device object tracking system according to any one of claims 11, characterized in that: When the device and one of the at least one peripheral device obtain the object features of the at least one object in the current frame, they are compared with the object features of the one or more objects in the previous frame in the original object pool; when it is determined from the comparison result that there is a new object feature, it is compared with the object feature of the target object in the target object pool, and the target object is matched, indicating that the target object appears in the continuous frames taken by the device or the one or more peripheral devices, and is then shared between the device and the at least one peripheral device.