Methods, apparatuses, and electronic devices for continuous cross-camera tracking of the same object based on different features and scenes.

By calculating the similarity of object features across multiple cameras, automatic continuous tracking across cameras is achieved, solving the problem of low efficiency of manual inspection in existing technologies and improving the accuracy and efficiency of tracking target objects over large areas.

CN114187322BActive Publication Date: 2025-12-02BEIJING BESCO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111210923.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-18
Publication Date
2025-12-02
Estimated Expiration
2041-10-18

AI Technical Summary

Technical Problem

In existing technologies, the images from multiple camera acquisition devices are displayed independently, requiring users to manually inspect and mark the movement of target objects in different areas, resulting in low efficiency and difficulty in guaranteeing accuracy.

Method used

By extracting features from target video frames and candidate video frames, calculating object feature similarity, and automatically matching the same object, continuous tracking across cameras can be achieved.

Benefits of technology

It enables automatic and continuous tracking of target objects within a large area covered by multiple cameras, improving tracking efficiency and accuracy while reducing human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114187322B_ABST
    Figure CN114187322B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, and electronic device for continuous cross-camera tracking of the same object based on different features and scenes. The method includes: extracting features from a target object determined in a target video frame to obtain target object features; extracting features from a candidate object determined in a candidate video frame to obtain candidate object features; calculating the similarity between a candidate object determined in at least one candidate video frame and the target object based on the target object features and the candidate object features; and determining that the candidate object and the target object are the same object when the similarity is greater than a first similarity threshold. This application's embodiments calculate similarity based on object features of target objects determined in video frames acquired sequentially by different acquisition devices with interactive areas, and perform object matching based on the similarity. This enables the use of multiple acquisition devices with interactive areas to track the trajectory of a target object in a large area that cannot be covered by a single acquisition device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, and electronic device for continuous tracking of the same object across cameras based on different features and scenes. Background Technology

[0002] With the rapid development of video / image acquisition technology and the decrease in the price of video acquisition equipment, organizations or individuals can now cover more areas by setting up more acquisition devices. In this case, existing technologies typically only display the captured images from multiple acquisition devices in a stitched-together manner on a large display area. However, each captured image still independently displays the image from its respective monitoring area. Therefore, users need to manually piece together the large monitoring image they need by viewing the different captured images.

[0003] In situations where multiple acquisition devices cover a large target area, when a user wants to understand the continuous trajectory of a specific target across these multiple acquisition areas, current technologies can only rely on manual inspection of multiple images or manual marking of the target's movement in different areas within those images. Such manual inspection or marking not only heavily relies on manpower, but also is extremely time-consuming, making it impossible to ensure tracking accuracy in long-duration trajectory tracking scenarios. Summary of the Invention

[0004] This application provides a method, apparatus, and electronic device for continuous cross-camera tracking of the same object based on different features and scenes, in order to solve the shortcomings of the prior art that requires manual tracking of target objects in multiple acquisition areas.

[0005] To achieve the above objectives, embodiments of this application provide a method for continuous cross-camera tracking of the same object based on different features and scenes, including:

[0006] Feature extraction is performed on the target object identified in the target video frame to obtain the target object features, wherein the target video frame is acquired by the first video source;

[0007] Feature extraction is performed on candidate objects identified in candidate video frames to obtain candidate object features, wherein the candidate video frames are acquired by a second video source, and the acquisition range of the first video source and the acquisition range of the second video source have an interactive area;

[0008] Based on the target object features and the candidate object features, calculate the similarity between the candidate object determined in at least one candidate video frame and the target object, wherein the acquisition time of the candidate video frame is later than that of the target video frame;

[0009] When the similarity is greater than the first similarity threshold, it is determined that the candidate object and the target object are the same object.

[0010] This application also provides a cross-camera continuous tracking device for the same object based on different features and scenes, including:

[0011] The first extraction module is used to extract features from a target object determined in a target video frame to obtain the features of the target object, wherein the target video frame is acquired by a first video source;

[0012] The second extraction module is used to extract features from the candidate objects identified in the candidate video frames to obtain the features of the candidate objects. The candidate video frames are acquired by a second video source, and the acquisition range of the first video source and the acquisition range of the second video source have an interactive area.

[0013] A calculation module is used to calculate the similarity between a candidate object determined in at least one of the candidate video frames and the target object based on the characteristics of the target object and the characteristics of the candidate object, wherein the acquisition time of the candidate video frame is later than that of the target video frame;

[0014] The determination module is used to determine that the candidate object and the target object are the same object when the similarity is greater than a first similarity threshold.

[0015] This application also provides an electronic device, including:

[0016] Memory, used to store programs;

[0017] A processor is configured to run the program stored in the memory, wherein the program executes the cross-camera continuous tracking method for the same object based on different features and scenes provided in the embodiments of this application.

[0018] This application also provides a computer-readable storage medium storing a computer program executable by a processor, wherein when the program is executed by the processor, it implements the cross-camera continuous tracking method for the same object based on different features and scenes provided in this application.

[0019] The present application provides a method, apparatus, and electronic device for continuous cross-camera tracking of the same object based on different features and scenes. By calculating the similarity of the target object based on the object features determined in video frames acquired by different acquisition devices with interactive areas in a sequential manner, and matching objects based on the similarity, it is possible to use multiple acquisition devices with interactive areas to track the trajectory of the target object in a large area that cannot be covered by a single acquisition device.

[0020] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0021] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0022] Figure 1 A schematic diagram illustrating an application scenario of a cross-camera continuous tracking scheme for the same object based on different features and scenes, as provided in an embodiment of this application.

[0023] Figure 2 A flowchart of an embodiment of the cross-camera continuous tracking method for the same object based on different features and scenes provided in this application;

[0024] Figure 3 A flowchart of another embodiment of the cross-camera continuous tracking method for the same object based on different features and scenes provided in this application;

[0025] Figure 4 A schematic diagram of the structure of an embodiment of a cross-camera continuous tracking device for the same object based on different features and scenes provided in this application;

[0026] Figure 5 A schematic diagram of the structure of an embodiment of the electronic device provided in this application. Detailed Implementation

[0027] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0028] Example 1

[0029] The solution provided in this application can be applied to any system that has the ability to continuously track the same object across cameras based on different features and scenes, such as a computing system with video processing capabilities, etc. Figure 1 This is a schematic diagram illustrating an application scenario of a cross-camera continuous tracking scheme for the same object based on different features and scenes, as provided in an embodiment of this application. Figure 1The scenario shown is merely one example of the principle of the technical solution of this application.

[0030] With the rapid development of video / image acquisition technology, the price of video acquisition equipment has become increasingly affordable. Organizations and individuals can now cover more areas by setting up more acquisition devices, thus achieving monitoring of a wider target area. In this scenario, current technology typically only stitches together the images from multiple acquisition devices to display them in a large display area. However, each acquisition image still independently displays the view from its respective monitoring area. Therefore, users need to manually piece together the desired large monitoring image by viewing the different acquisition images.

[0031] In particular, when using multiple acquisition devices to cover a large target area, if a user wants to understand the continuous trajectory of a specific target across these multiple acquisition areas, the existing technology, while able to identify objects within the acquired video frames, relies on manual inspection of multiple frames across different acquisition areas to determine if the objects contained within are the same object or to manually mark the target object in multiple frames to determine its movement across different areas. This manual inspection or marking is not only heavily dependent on manpower, but also extremely time-consuming, making it impossible to ensure tracking accuracy in long-duration trajectory tracking scenarios.

[0032] For example, in such Figure 1 In the scenario shown, the acquisition range 1 of acquisition device 1 and the acquisition range 2 of acquisition device 2 have a certain overlap area, i.e., an interactive area. In other words, the image in this interactive area is included in both the video frames acquired by acquisition device 1 and the video frames acquired by acquisition device 2. For example, in Figure 1 In the scenario shown, data acquisition device 1 can be installed inside a room, and data acquisition device 2 can be installed in the lobby outside the room. Therefore, Figure 1The interactive area shown can be a rectangular area centered on, for example, the sliding door of the room where acquisition device 1 is installed. This area includes a rectangular area extending a predetermined distance from the sliding door into the interior of the room where acquisition device 1 is installed, and a rectangular area extending a predetermined distance from the sliding door into the hallway where acquisition device 2 is installed. In other words, when a target object, such as a person, moves from the room through the sliding door into the hallway, it can pass through this interactive area and can appear successively in the video frames captured by acquisition device 1 and the video frames captured by acquisition device 2. Therefore, in the existing technology, when tracking object 1, as object 1 moves back and forth in room 1, the acquisition device 1 can mark object 1 in each video frame based on the characteristics of the object identified in the captured video frames, thereby achieving continuous tracking of object 1. However, when object 1 enters the interactive area and walks through the sliding door into the lobby, object 1 actually disappears from the video frames captured by the acquisition device 1. That is, the characteristics of the object extracted by the acquisition device 1 in the captured video frames are different from the characteristics of object 1. At this point, the existing technology confirms that object 1 has left the room, and the tracking process for object 1 ends. If the user of the acquisition device wants to continue tracking object 1, they need to manually search for it in the screens of other acquisition devices, for example, in the video frames captured by acquisition device 2, and then start tracking object 1 from acquisition device 2. This manual processing method is not only inefficient when using a large number of acquisition devices for joint monitoring of a large area, but also prone to missing target objects. Furthermore, because existing technologies use multiple acquisition devices for monitoring, each device independently performs acquisition and recognition operations. Therefore, tracking the movement of a target object, such as Object 1, within overlapping acquisition areas across these multiple devices relies solely on manual intervention. For example, in... Figure 1 In the scenario shown, the user of this acquisition system composed of multiple acquisition devices can track an object 1 within the acquisition range of another acquisition device (e.g., acquisition device 2) when the object 1 leaves the acquisition range of acquisition device 1. This object 1 can be designated as the target object in acquisition device 2, or the user can confirm that the object 1 has entered the acquisition range 2 of acquisition device 2 by checking the recognition result of acquisition device 2. Furthermore, the user can obtain a continuous trajectory of the object 1's movement within the acquisition ranges of acquisition devices 1 and 2 by summarizing the movement trajectories of the object 1 identified in acquisition device 1 and acquisition device 2. However, such a manual tracking method across multiple acquisition devices relies on human intervention, resulting in low efficiency and accuracy susceptible to user fatigue.

[0033] In this embodiment, after the acquisition device 1 identifies object 1, an acquisition device with an interactive area with acquisition device 1, such as acquisition device 2, compares the features of the object identified in video frames acquired after a specified time, such as the time after object 1 disappears from the acquisition range of acquisition device 1, with object 1. If they match, it is determined that object 1 has appeared within the acquisition range of acquisition device 2. For example, in Figure 1 In the scenario shown, the acquisition device 1 identifies object 1 within its acquisition range 1, and can identify object 1 and its location in each video frame, thus enabling continuous tracking of object 1's trajectory within the acquisition range 1. When object 1 moves to... Figure 1 When the interaction area between acquisition device 1 and acquisition device 2 is shown, for example, the video frame acquired at that moment can be used as the first video frame, and object 1 can be identified from it. Then, at the next moment, since object 1 has passed through the interaction area (i.e., through the sliding door between the room and the hall) and entered the hall, the features of the object identified in the second video frame acquired by acquisition device 1 at that next moment do not match the features of object 1. That is, acquisition device 1 can determine that object 1 has left the acquisition range 1. In this case, in this embodiment, acquisition device 2 can automatically compare the features of the object in the video frame acquired at that moment with the features of object 1, without requiring the user to manually select other acquisition devices. Therefore, acquisition device 2 can automatically compare the features of the identified objects one by one in its subsequent video frames to see if they match the features of object 1. And when a match is confirmed, acquisition device 2 can confirm that object 1 has appeared within its acquisition range 2. In this way, acquisition device 2 can cooperate with acquisition device 1 in the trajectory tracking processing of object 1.

[0034] In addition, in this embodiment of the application, the user may pre-specify object 1 in the video frame captured by the acquisition device 1 as the target object, and other acquisition devices that together with the acquisition device 1 constitute the acquisition system. Based on the user's specified operation on the target object, other acquisition devices can automatically extract the features of the object in all video frames after the specified operation and match and compare them with the features of the specified object 1, thereby marking or extracting all video frames containing object 1 captured at subsequent times for the user of the acquisition system to refer to or use.

[0035] Specifically, in this embodiment, when the acquisition device 1 determines that object 1 has disappeared from its acquisition range 1, a callback function can be used to pass the target object features of the determined target object to the processing thread. This processing thread can be used to compare, for example, the target object features of object 1 with the candidate object features of candidate objects determined by the acquisition device 2. For example, the features of objects acquired by multiple acquisition devices can be processed on a video frame processing server. Thus, when the server processes the video frames of acquisition device 1 and determines that object 1 has disappeared from the acquisition range 1 of acquisition device 1, the server can extract the object features of object 1 from the video frames of acquisition device 1 before the disappearance of object 1, and use these object features as parameters for the callback function of the thread processing the video frames. When this thread performs feature comparison on objects in the video frames acquired by acquisition device 2, it can call this parameter to compare with the objects in the video frames acquired by acquisition device 2 to confirm whether object 1 appears in the acquisition range 2 of acquisition device 2.

[0036] Furthermore, in, for example Figure 1In the scenario shown, when the acquisition device 1 confirms that object 1 has disappeared from its acquisition range, object 1 may leave the acquisition range 1 of acquisition device 1 through the interaction area of ​​the acquisition range of other acquisition devices. Alternatively, object 1 may have only changed its body posture, or some or all of its features may not have been acquired by acquisition device 1 due to obstruction by other objects within the acquisition range of acquisition device 1, such as furniture or other people. Therefore, in this embodiment, when the acquisition device 1 performs object recognition on video frames acquired sequentially in time and confirms, for example, that the target object of object 1 exists in an earlier video frame but disappears in a later video frame, it can first consider reducing the feature matching threshold for confirming object 1 in the video frames acquired by acquisition device 1. That is, in this embodiment, when performing object recognition on the video frames acquired by acquisition device 1, it actually first identifies each object contained in the video frame, extracts its features to perform feature matching with pre-given reference features or standard features, and confirms objects with a similarity greater than a certain threshold, such as 90%, as pre-given objects. Therefore, when object 1 only undergoes a change in body posture, such as squatting or turning around, or is obscured by other objects within the acquisition range, the video frames captured by acquisition device 1 contain fewer features of that object. Consequently, when these determined features are used as the entirety of the object's features for feature matching with pre-given reference features or previously acquired features of object 1, the matching similarity will obviously decrease, leading to an incorrect judgment that object 1 has disappeared from the acquisition range 1 of acquisition device 1. Therefore, in this embodiment, when it is determined that object 1 exists in the earlier video frame but not in the later video frame based on the feature comparison results of two temporally adjacent video frames, object 1 can be identified by lowering the feature matching threshold.

[0037] Furthermore, under the above circumstances, the tracking method of this application embodiment can also place the object determined in the later video frame into a temporary queue, and when comparing the features of the object in the subsequent video frame, in addition to comparing it with the features of a preset object, such as object 1, it can also further compare it with the features of the object in the temporary queue, thereby reducing the probability of the above-mentioned misjudgment.

[0038] Furthermore, in this embodiment, when the features of object 1 in the video frame acquired by acquisition device 1 are passed as parameters of a callback function to the thread processing the video frame acquired by acquisition device 2, a second similarity threshold can be further set. This similarity threshold can be used when the thread compares the features of the object in the video frame acquired by the acquisition device with the features of object 1 in the parameters of the callback function. For example, when the similarity between the first candidate object determined in the first candidate video frame acquired by acquisition device 2 and a preset target object such as object 1 is not greater than the first similarity threshold, the similarity between the second candidate object determined in the second candidate video frame and the target object is greater than the first similarity threshold, and the similarity between the first candidate object and the second candidate object is greater than the second similarity threshold, it is determined that the first candidate object, the second candidate object, and the target object are the same object. For example, as described above, when the similarity between object 1 and object 1 in the first candidate video frame acquired by acquisition device 2 is lower than the first similarity threshold due to changes in posture such as squatting or turning around, or due to being blocked by other objects within the acquisition range 1, according to the embodiments of this application, in the next second candidate video frame, the features of the candidate object can be compared not only with the preset features of object 1, but also with the features of the object in the first candidate video frame. If the feature similarity between the candidate object in the next second candidate video frame and object 1 is greater than the first similarity threshold, that is, the candidate object in the second candidate video frame can be considered as object 1, then the similarity between the object in the first candidate video frame and the object in the second candidate video frame can be further determined. If the similarity is greater than the second similarity threshold, then the object in the first candidate video frame can also be considered as object 1. In other words, if object 1 maintains a crouching posture when leaving the acquisition range 1 of acquisition device 1 and entering the acquisition range 2 of acquisition device 2, the feature similarity between the first candidate object identified in the first candidate video frame and object 1 is lower than the first similarity threshold. However, if object 1 then stands upright, for example, in the second candidate video frame, it is in a posture between crouching and standing upright. Therefore, the feature similarity between the second candidate object identified in the second candidate video frame and object 1 is higher than the first similarity threshold, confirming that object 1 does indeed appear in the second candidate video frame. In this case, although the feature similarity between object 1 and object 1 is low in the first candidate video frame because object 1 is crouching, the feature similarity between the crouching object 1 and the upright object 1 in the second candidate video frame is high. Therefore, this method can reduce the misjudgment that object 1 does not exist in the first candidate video frame.

[0039] Furthermore, in this embodiment, when it is detected that, for example, the target object 1 exists in the first target video frame acquired by the acquisition device 1 but does not exist in the later second target video frame, and the similarity between the first candidate object determined in the first candidate video frame acquired by the acquisition device 2 and the target object 1 is not greater than a first similarity threshold, that is, in Figure 1 In the scenario shown, object 1 exists in the first target frame but disappears in the subsequent second target video frame. According to this embodiment, object 1 may have left the acquisition range 1 of acquisition device 1. Therefore, in the first candidate video frame acquired by acquisition device 2 (which has an interaction area with acquisition device 1's acquisition range 1), a first candidate object is identified and its features are compared with those of object 1. However, the similarity is lower than the first similarity threshold, meaning object 1 is not found in the first candidate video frame acquired by acquisition device 2. In this embodiment, the target object identified in the second target video frame of acquisition device 1 can be placed in the target queue, and the first candidate object in the first candidate video frame of acquisition device 2 can be placed in the candidate queue. Next, the feature matching similarity between the second candidate object identified in the second candidate video frame of acquisition device 2 and object 1 can be calculated. If the feature matching similarity between the second candidate object in the second candidate video frame and object 1 is higher than the first similarity threshold, then object 1 is determined to appear in the second candidate video frame, and feature similarity matching calculation can continue using the features of the target objects in the target queue. Furthermore, a matching weight can be set for the target objects in the target queue, and this matching weight can be reduced as the number of times the target object fails to match at least one candidate object increases. In other words, in this embodiment, although object 1 can be placed into the target queue as a target object when it disappears in the second target video frame of the acquisition device 1, if object 1 cannot be found in the candidate video frames acquired by the acquisition device 2 for a continuous period of time, for example, if the features of the candidate objects in multiple candidate video frames have a similarity to the features of object 1 that is lower than the first similarity threshold, then the weight of object 1 can be gradually reduced, that is, the confidence of finding object 1 can be gradually reduced, thereby avoiding misjudgment. However, if object 1 is found in the candidate video frames acquired by the acquisition device 2 during the process of gradually reducing the matching weight of object 1, then the matching weight of object 1 can be immediately restored.

[0040] Therefore, the cross-camera continuous tracking scheme for the same object based on different features and scenes provided in this application calculates the similarity of the target object by determining the object features of the target object in video frames acquired by different acquisition devices with interactive areas in a sequential manner, and performs object matching based on the similarity. This enables the use of multiple acquisition devices with interactive areas to track the trajectory of the target object in a large area that cannot be covered by a single acquisition device.

[0041] The above embodiments illustrate the technical principles and exemplary application framework of the embodiments of this application. The specific technical solutions of the embodiments of this application will be further described in detail below through multiple embodiments.

[0042] Example 2

[0043] Figure 2 This is a flowchart of an embodiment of the cross-camera continuous tracking method for the same object based on different features and scenes provided in this application. The execution subject of this method can be various terminal or server devices with video or image processing capabilities, or it can be a device or chip integrated on these devices. Figure 2 As shown, the cross-camera continuous tracking method for the same object based on different features and scenes includes the following steps:

[0044] S201, extract features from the target object identified in the target video frame to obtain the target object features.

[0045] In this embodiment of the application, target video frames can be acquired from various acquisition devices, and in step S201, feature extraction is performed on the target object in the given target video frame to obtain the target object features. For example, it can be obtained by... Figure 1 The acquisition device 1 shown is used to acquire target video frames and perform object recognition processing on the target video frames to determine the object 1 contained therein as the target object. In step S201, feature extraction can be performed on the determined object 1 to obtain the target object features.

[0046] S202, extract features from the candidate objects identified in the candidate video frames to obtain the features of the candidate objects.

[0047] After obtaining the target object features in step S201, feature extraction can be performed in step S202 on the candidate objects identified from candidate video frames acquired by a second video source different from the first video source used in step S201, thereby obtaining the candidate object features. In this embodiment, for example, Figure 1The acquisition device 1 shown is used as the first video source, and the acquisition device 2, which has an interactive area with the acquisition device 1, is used as the second video source. Alternatively, a server or database storing the video can also be used as the first and second video sources, as long as the video provided is acquired by different acquisition devices with interactive areas.

[0048] S203, based on the characteristics of the target object and the characteristics of the candidate object, calculate the similarity between the candidate object and the target object determined in at least one candidate video frame.

[0049] After obtaining the target object features in step S201 and the candidate object features in step S202, the similarity between the candidate object identified in the candidate video frame and the target object features obtained in step S201 can be calculated in candidate video frames whose acquisition time is later than the target video frame. For example, in... Figure 1 In the scenario shown, in step S201, feature extraction can be performed on the target video frame identified by the acquisition device 1 to obtain the target object features of object 1. In step S202, candidate objects of candidate objects identified in the candidate video frames acquired by the acquisition device 2 can be extracted. Thus, in step S203, the similarity between the target object features and the candidate object features can be calculated to determine, for example, whether there are candidate objects whose target features can match the target object features of object 1 in the candidate video frames acquired by the acquisition device 2 later than, for example, when object 1 disappears from the acquisition range 1 of the acquisition device 1.

[0050] S204, when the similarity is greater than the first similarity threshold, it is determined that the candidate object and the target object are the same object.

[0051] For example, in step S204, it can be determined whether the target object and the candidate object are the same object based on the similarity between the target object and the candidate object calculated in step S203. For example, in step S204, when the similarity between the target object and the candidate object calculated in step S203 is greater than a first similarity threshold, such as 90%, it can be determined that, for example, the candidate object in the candidate video frame acquired by the acquisition device 2 is consistent with the object 1 that disappeared from the acquisition range of the acquisition device 1.

[0052] Therefore, the cross-camera continuous tracking method for the same object based on different features and scenes provided in this application calculates similarity based on the object features of the target object determined in video frames acquired by different acquisition devices with interactive areas in a sequential manner, and performs object matching based on similarity. This enables the use of multiple acquisition devices with interactive areas to track the trajectory of the target object in a large area that cannot be covered by a single acquisition device.

[0053] Example 3

[0054] Figure 3 This is a flowchart of another embodiment of the cross-camera continuous tracking method for the same object based on different features and scenes provided in this application. The execution subject of this method can be various terminal or server devices with video or image processing capabilities, or it can be a device or chip integrated on these devices. Figure 3 As shown, the cross-camera continuous tracking method for the same object based on different features and scenes includes the following steps:

[0055] S301, extract features from the target objects identified in multiple target video frames to obtain the features of the target objects.

[0056] In this embodiment of the application, multiple target video frames can be acquired from various acquisition devices, for example, they can be... Figure 1 The acquisition device 1 shown acquires multiple target objects in the given multiple target video frames and performs feature extraction on the target objects in step S301 to obtain the target object features. For example, it can be obtained by... Figure 1 The acquisition device 1 shown is used to acquire multiple target video frames over a period of time. For example, the acquisition device 1 can acquire multiple target video frames during the time it takes for object 1 to walk from the middle of the room to the vicinity of the sliding door in the interactive area, and perform object recognition processing on these multiple target video frames to determine that the object 1 contained therein is the target object. Thus, in step S301, feature extraction can be performed on the determined object 1 to obtain the target object features.

[0057] S302, extract features from the candidate objects identified in multiple candidate video frames to obtain the features of the candidate objects.

[0058] After obtaining the target object features in step S301, feature extraction can be performed in step S302 on the candidate objects identified from multiple candidate video frames acquired by a second video source different from the first video source used in step S301, thereby obtaining the candidate object features. In this embodiment, for example, Figure 1 The acquisition device 1 shown is used as the first video source, and the acquisition device 2, which has an interactive area with the acquisition device 1, is used as the second video source. Alternatively, a server or database storing the video can also be used as the first and second video sources, as long as the video provided is acquired by different acquisition devices with interactive areas.

[0059] Furthermore, the embodiments in this application may not be limited to... Figure 1The example shown uses one acquisition device 2 as a single second video source, but multiple acquisition devices with interactive areas with acquisition device 1 can be used as multiple second video sources. For example, in a scenario with multiple rooms sharing a hall, where the doors of two or more rooms are adjacent to each other, therefore, when, for example... Figure 1 As shown, when object 1 moves from the hall equipped with acquisition device 1 towards the vicinity of two doors of two adjacent rooms and disappears from the target video frame captured by acquisition device 1, object 1 may enter one of the two adjacent rooms. Therefore, matching weights can be set for the acquisition devices in each room, and these matching weights can be adjusted based on various auxiliary attribute information of the object. For example, the auxiliary attribute information can be one or more of the following: object 1's position information, orientation information, velocity information, and size information. For example, when object 1 is close to a certain room, the matching weight of that room can be adjusted to be higher. This way, even if the feature matching degree is low when matching the candidate object in the candidate video frame captured by the acquisition device in that room due to reasons such as object 1's change in body posture or being occluded by other objects when entering the room, the probability of successful matching can still be increased by adjusting the matching weight, thereby reducing the erroneous judgment that the candidate object in the room is not object 1 due to reasons such as changes in body posture or being occluded by other objects.

[0060] S303 uses a callback function to pass the target object characteristics of the target object determined in the first video source to the processing thread.

[0061] In this embodiment of the application, a processing thread can be used to calculate the similarity, that is, the processing thread can compare the target object features of the target object obtained in step S301 with the candidate object features of the candidate objects determined in the second video source obtained in step S302.

[0062] In the embodiments of this application, for example, in such Figure 1In the scenario shown, when acquisition device 1 determines that object 1 has disappeared from its acquisition range 1, a callback function can be used in step S303 to pass the target object features of the determined target object to the processing thread. This processing thread can be used to compare, for example, the target object features of object 1 with the candidate object features of candidate objects determined by acquisition device 2. For example, the video frame processing server can process the features of objects acquired by multiple acquisition devices. Thus, when the server processes the video frames of acquisition device 1 and determines that object 1 has disappeared from the acquisition range 1 of acquisition device 1, the server can extract the object features of object 1 from the video frames of acquisition device 1 before the moment object 1 disappeared, and use these object features as parameters for the callback function of the thread processing the video frames. When this thread performs feature comparison on objects in the video frames acquired by acquisition device 2, it can call this parameter to compare with the objects in the video frames acquired by acquisition device 2 to confirm whether object 1 appears in the acquisition range 2 of acquisition device 2.

[0063] S304, based on the characteristics of the target object and the characteristics of the candidate object, calculate the similarity between the candidate object and the target object determined in at least one candidate video frame.

[0064] After obtaining the target object features in step S301 and the candidate object features in step S302, and passing the target object features to the processing thread via a callback function in step S303, in step S304, the similarity between the candidate object determined in the candidate video frame acquired later than the target video frame and the target object obtained in step S301 can be calculated. For example, in... Figure 1 In the scenario shown, in step S301, feature extraction can be performed on the object 1, i.e. the target object, determined in multiple target video frames acquired by acquisition device 1 to obtain the target object features of object 1. In step S302, candidate objects of candidate objects determined in candidate video frames acquired by acquisition device 2 can be extracted. Thus, in step S304, a processing thread can be used to calculate the similarity between the target object and the candidate object using the target object features passed through the callback function in step S303, to determine, for example, whether there are candidate objects in the candidate video frames acquired by acquisition device 2 later than, for example, when object 1 disappears from the acquisition range 1 of acquisition device 1, where the candidate target features can match the target object features of object 1.

[0065] S305, when the similarity is greater than the first similarity threshold, it is determined that the candidate object and the target object are the same object.

[0066] For example, in step S305, it can be determined whether the target object and the candidate object are the same object based on the similarity between the target object and the candidate object calculated in step S304. For example, in step S305, when the similarity between the target object and the candidate object calculated in step S304 is greater than a first similarity threshold, such as 90%, it can be determined that, for example, the candidate object in the candidate video frame acquired by the acquisition device 2 is consistent with the object 1 that disappeared from the acquisition range of the acquisition device 1.

[0067] Furthermore, in, for example Figure 1 In the scenario shown, when the acquisition device 1 confirms that object 1 has disappeared from its acquisition range, object 1 may leave the acquisition range 1 of acquisition device 1 through the interaction area of ​​the acquisition range of other acquisition devices. Alternatively, object 1 may have only changed its body posture, or some or all of its features may not have been acquired by acquisition device 1 due to obstruction by other objects within the acquisition range of acquisition device 1, such as furniture or other people. Therefore, in this embodiment, when the acquisition device 1 performs object recognition on video frames acquired sequentially in time and confirms, for example, that the target object of object 1 exists in an earlier video frame but disappears in a later video frame, it can first consider reducing the feature matching threshold for confirming object 1 in the video frames acquired by acquisition device 1. That is, in this embodiment, when performing object recognition on the video frames acquired by acquisition device 1, it actually first identifies each object contained in the video frame, extracts its features to perform feature matching with pre-given reference features or standard features, and confirms objects with a similarity greater than a certain threshold, such as 90%, as pre-given objects. Therefore, when object 1 only undergoes a change in body posture, such as squatting or turning around, or is obscured by other objects within the acquisition range, the video frames acquired by acquisition device 1 contain fewer features of that object. Consequently, in step S305, when these determined features are used as all the features of the object and matched against pre-given reference features or previously acquired features of object 1, the matching similarity will obviously decrease, leading to an incorrect judgment that object 1 has disappeared from the acquisition range 1 of acquisition device 1. Therefore, in this embodiment, when the feature comparison results of two temporally adjacent video frames in step S305 determine that object 1 exists in the earlier video frame but not in the later video frame, object 1 can be identified by lowering the feature matching threshold.

[0068] Furthermore, in the above-mentioned case, in step S305 of this application embodiment, the object determined in the later video frame can be placed in a temporary queue, and when comparing the features of the object in the subsequent video frame, in addition to comparing it with the features of a preset object, such as object 1, it can also be compared with the features of the object in the temporary queue, thereby reducing the probability of the above-mentioned misjudgment.

[0069] Furthermore, in this embodiment, a second similarity threshold can be further set. This similarity threshold can be used when the thread compares the features of objects in the video frames captured by the acquisition device with the features of object 1 in the parameters of the callback function. For example, when it is determined in step S305 that the similarity between the first candidate object determined in the first candidate video frame captured by the acquisition device 2 and the preset target object, such as object 1, is not greater than the first similarity threshold, the similarity between the second candidate object determined in the second candidate video frame and the target object is greater than the first similarity threshold, and the similarity between the first candidate object and the second candidate object is greater than the second similarity threshold, it is determined that the first candidate object, the second candidate object, and the target object are the same object.

[0070] For example, as described above, when the similarity between object 1 and object 1 in the first candidate video frame acquired by acquisition device 2 is lower than the first similarity threshold due to changes in posture such as squatting or turning around, or due to being blocked by other objects within the acquisition range 1, according to the embodiments of this application, in step S305, the features of the candidate object in the second candidate video frame can be further compared not only with the preset features of object 1, but also with the features of the object in the first candidate video frame. If the feature similarity between the candidate object in the next second candidate video frame and object 1 is greater than the first similarity threshold, that is, the candidate object in the second candidate video frame can be considered as object 1, then the similarity between the object in the first candidate video frame and the object in the second candidate video frame can be further determined. If the similarity is greater than the second similarity threshold, then the object in the first candidate video frame can also be considered as object 1. In other words, if object 1 maintains a crouching posture when leaving the acquisition range 1 of acquisition device 1 and entering the acquisition range 2 of acquisition device 2, the feature similarity between the first candidate object identified in the first candidate video frame and object 1 is lower than the first similarity threshold. However, if object 1 then stands upright, for example, in the second candidate video frame, it is in a posture between crouching and standing upright. Therefore, the feature similarity between the second candidate object identified in the second candidate video frame and object 1 is higher than the first similarity threshold, confirming that object 1 does indeed appear in the second candidate video frame. In this case, although the feature similarity between object 1 and object 1 is low in the first candidate video frame because object 1 is crouching, the feature similarity between the crouching object 1 and the upright object 1 in the second candidate video frame is high. Therefore, this method can reduce the misjudgment that object 1 does not exist in the first candidate video frame.

[0071] Furthermore, in this embodiment of the application, step S305 may also be performed when, for example, the target object 1 exists in the first target video frame acquired by the acquisition device 1 but does not exist in the later second target video frame, and the similarity between the first candidate object determined in the first candidate video frame acquired by the acquisition device 2 and the target object 1 is not greater than a first similarity threshold, that is, in Figure 1In the scenario shown, object 1 exists in the first target frame but disappears in the subsequent second target video frame. According to this embodiment, object 1 may have left the acquisition range 1 of acquisition device 1. Therefore, in the first candidate video frame acquired by acquisition device 2 (which has an interaction area with acquisition device 1's acquisition range 1), a first candidate object is identified and its features are compared with those of object 1. However, the similarity is lower than the first similarity threshold, meaning object 1 is not found in the first candidate video frame acquired by acquisition device 2. In this embodiment, the target object identified in the second target video frame of acquisition device 1 can be placed in the target queue, and the first candidate object in the first candidate video frame of acquisition device 2 can be placed in the candidate queue. Next, the feature matching similarity between the second candidate object identified in the second candidate video frame of acquisition device 2 and object 1 can be calculated. If the feature matching similarity between the second candidate object in the second candidate video frame and object 1 is higher than the first similarity threshold, then object 1 is determined to appear in the second candidate video frame, and feature similarity matching calculation can continue using the features of the target objects in the target queue. Furthermore, a matching weight can be set for the target objects in the target queue, and this matching weight can be reduced as the number of times the target object fails to match at least one candidate object increases. In other words, in this embodiment, although object 1 can be placed into the target queue as a target object when it disappears in the second target video frame of the acquisition device 1, if object 1 cannot be found in the candidate video frames acquired by the acquisition device 2 for a continuous period of time, for example, if the features of the candidate objects in multiple candidate video frames have a similarity to the features of object 1 that is lower than the first similarity threshold, then the weight of object 1 can be gradually reduced, that is, the confidence of finding object 1 can be gradually reduced, thereby avoiding misjudgment. However, if object 1 is found in the candidate video frames acquired by the acquisition device 2 during the process of gradually reducing the matching weight of object 1, then the matching weight of object 1 can be immediately restored.

[0072] Therefore, the cross-camera continuous tracking method for the same object based on different features and scenes provided in this application calculates similarity based on the object features of the target object determined in video frames acquired by different acquisition devices with interactive areas in a sequential manner, and performs object matching based on similarity. This enables the use of multiple acquisition devices with interactive areas to track the trajectory of the target object in a large area that cannot be covered by a single acquisition device.

[0073] Example 4

[0074] Figure 4 This application provides a schematic diagram of the structure of an embodiment of a cross-camera continuous tracking device for the same object based on different features and scenes, which can be used to perform, for example... Figure 2 and Figure 3 The method steps are shown. (As shown) Figure 4 As shown, the cross-camera continuous tracking device for the same object based on different features and scenes may include: a first extraction module 41, a second extraction module 42, a calculation module 43, and a determination module 44.

[0075] The first extraction module 41 can be used to extract features from the target object determined in the target video frame to obtain the target object features.

[0076] In this embodiment of the application, target video frames can be acquired from various acquisition devices, and the first extraction module 41 can extract features of the target object in a given target video frame to obtain the features of the target object. For example, it can be obtained by... Figure 1 The acquisition device 1 shown is used to acquire target video frames and perform object recognition processing on the target video frames to determine the object 1 contained therein as the target object, so that the first extraction module 41 can perform feature extraction on the determined object 1 to obtain the target object features.

[0077] The second extraction module 42 can be used to extract features from the candidate objects identified in the candidate video frames to obtain the features of the candidate objects.

[0078] After the first extraction module 41 obtains the target object features, the second extraction module 42 can extract features from candidate video frames identified by a second video source different from the first video source used by the first extraction module 41, thereby obtaining candidate object features. In this embodiment, for example... Figure 1 The acquisition device 1 shown is used as the first video source, and the acquisition device 2, which has an interactive area with the acquisition device 1, is used as the second video source. Alternatively, a server or database storing the video can also be used as the first and second video sources, as long as the video provided is acquired by different acquisition devices with interactive areas.

[0079] The calculation module 43 can be used to calculate the similarity between the candidate object and the target object in at least one candidate video frame based on the characteristics of the target object and the characteristics of the candidate object.

[0080] After the first extraction module 41 obtains the target object features and the second extraction module 42 obtains the candidate object features, the similarity between the candidate object determined in the candidate video frame and the target object obtained by the first extraction module 41 can be calculated in candidate video frames whose acquisition time is later than the target video frame. For example, in... Figure 1In the scenario shown, the first extraction module 41 can extract features of the target object 1 determined in the target video frame acquired by the acquisition device 1 to obtain the target object features of the object 1, and the second extraction module 42 extracts the candidate objects of the candidate objects determined in the candidate video frames acquired by the acquisition device 2. Thus, the calculation module 43 can calculate the similarity between the target object features and the candidate object features to determine, for example, whether there are candidate objects whose target features can match the target object features of the object 1 in the candidate video frames acquired by the acquisition device 2 later than, for example, when the object 1 disappears from the acquisition range 1 of the acquisition device 1.

[0081] The determination module 44 can be used to determine that the candidate object and the target object are the same object when the similarity is greater than the first similarity threshold.

[0082] For example, the determining module 44 can determine whether the target object and the candidate object are the same object based on the similarity between the target object and the candidate object calculated by the calculation module 43. For example, when the similarity between the target object and the candidate object calculated by the calculation module 43 is greater than a first similarity threshold, such as 90%, the determining module 44 can determine that, for example, the candidate object in the candidate video frame acquired by the acquisition device 2 is consistent with the object 1 that disappeared from the acquisition range of the acquisition device 1.

[0083] Furthermore, in, for example Figure 1In the scenario shown, when the acquisition device 1 confirms that object 1 has disappeared from its acquisition range, object 1 may leave the acquisition range 1 of acquisition device 1 through the interaction area of ​​the acquisition range of other acquisition devices. Alternatively, object 1 may have only changed its body posture, or some or all of its features may not have been acquired by acquisition device 1 due to obstruction by other objects within the acquisition range of acquisition device 1, such as furniture or other people. Therefore, in this embodiment, when the acquisition device 1 performs object recognition on video frames acquired sequentially in time and confirms, for example, that the target object of object 1 exists in an earlier video frame but disappears in a later video frame, it can first consider reducing the feature matching threshold for confirming object 1 in the video frames acquired by acquisition device 1. That is, in this embodiment, when performing object recognition on the video frames acquired by acquisition device 1, it actually first identifies each object contained in the video frame, extracts its features to perform feature matching with pre-given reference features or standard features, and confirms objects with a similarity greater than a certain threshold, such as 90%, as pre-given objects. Therefore, when object 1 only changes its body posture, such as squatting or turning around, or is obscured by other objects within the acquisition range, the video frames acquired by acquisition device 1 contain fewer features of that object. Consequently, when the determining module 44 uses these determined features as the entirety of the object's features and performs feature matching with pre-given reference features or previously acquired features of object 1, it obviously leads to a decrease in matching similarity, resulting in an incorrect judgment that object 1 has disappeared from the acquisition range 1 of acquisition device 1. Therefore, in this embodiment, when the determining module 44 determines that object 1 exists in the earlier video frame but not in the later video frame based on the feature comparison results of two temporally adjacent video frames, object 1 can be identified by lowering the feature matching threshold.

[0084] Furthermore, in the above situation, the determination module 44 of this application embodiment can also put the object determined in the later video frame into a temporary queue, and when comparing the features of the object in the subsequent video frame, in addition to comparing it with the features of the preset object, such as object 1, it will further compare it with the features of the object in the temporary queue, thereby reducing the probability of the above misjudgment.

[0085] Furthermore, in this embodiment, a second similarity threshold can be further set. This similarity threshold can be used when the thread compares the features of objects in the video frames acquired by the acquisition device with the features of object 1 in the parameters of the callback function. For example, when the determining module 44 determines that the similarity between the first candidate object determined in the first candidate video frame acquired by the acquisition device 2 and the preset target object, such as object 1, is not greater than the first similarity threshold, the similarity between the second candidate object determined in the second candidate video frame and the target object is greater than the first similarity threshold, and the similarity between the first candidate object and the second candidate object is greater than the second similarity threshold, it is determined that the first candidate object, the second candidate object, and the target object are the same object.

[0086] For example, as described above, when the similarity between the determination module 44 and the features of object 1 in the first candidate video frame acquired by the acquisition device 2 is lower than the first similarity threshold due to changes in posture such as squatting or turning around, or due to being blocked by other objects within the acquisition range 1, the determination module 44 can further compare the features of the candidate object in the second candidate video frame not only with the preset features of object 1, but also with the features of the object in the first candidate video frame, according to the embodiments of this application. If the feature similarity between the candidate object in the next second candidate video frame and object 1 is greater than the first similarity threshold, that is, the candidate object in the second candidate video frame can be considered as object 1, then the similarity between the object in the first candidate video frame and the object in the second candidate video frame can be further determined. If the similarity is greater than the second similarity threshold, then the object in the first candidate video frame can also be considered as object 1. In other words, if object 1 maintains a crouching posture when leaving the acquisition range 1 of acquisition device 1 and entering the acquisition range 2 of acquisition device 2, the feature similarity between the first candidate object identified in the first candidate video frame and object 1 is lower than the first similarity threshold. However, if object 1 then stands upright, for example, in the second candidate video frame, it is in a posture between crouching and standing upright. Therefore, the feature similarity between the second candidate object identified in the second candidate video frame and object 1 is higher than the first similarity threshold, confirming that object 1 does indeed appear in the second candidate video frame. In this case, although the feature similarity between object 1 and object 1 is low in the first candidate video frame because object 1 is crouching, the feature similarity between the crouching object 1 and the upright object 1 in the second candidate video frame is high. Therefore, this method can reduce the misjudgment that object 1 does not exist in the first candidate video frame.

[0087] Furthermore, in this embodiment, the determining module 44 can also determine the target object 1 when it is detected that the target object 1 exists in the first target video frame acquired by the acquisition device 1 but does not exist in the later second target video frame, and the similarity between the first candidate object determined in the first candidate video frame acquired by the acquisition device 2 and the target object 1 is not greater than a first similarity threshold, that is, when Figure 1 In the scenario shown, object 1 exists in the first target frame but disappears in the subsequent second target video frame. According to this embodiment, object 1 may have left the acquisition range 1 of acquisition device 1. Therefore, in the first candidate video frame acquired by acquisition device 2 (which has an interaction area with acquisition device 1's acquisition range 1), a first candidate object is identified and its features are compared with those of object 1. However, the similarity is lower than the first similarity threshold, meaning object 1 is not found in the first candidate video frame acquired by acquisition device 2. In this embodiment, the target object identified in the second target video frame of acquisition device 1 can be placed in the target queue, and the first candidate object in the first candidate video frame of acquisition device 2 can be placed in the candidate queue. Next, the feature matching similarity between the second candidate object identified in the second candidate video frame of acquisition device 2 and object 1 can be calculated. If the feature matching similarity between the second candidate object in the second candidate video frame and object 1 is higher than the first similarity threshold, then object 1 is determined to appear in the second candidate video frame, and feature similarity matching calculation can continue using the features of the target objects in the target queue. Furthermore, a matching weight can be set for the target objects in the target queue, and this matching weight can be reduced as the number of times the target object fails to match at least one candidate object increases. In other words, in this embodiment, although object 1 can be placed into the target queue as a target object when it disappears in the second target video frame of the acquisition device 1, if object 1 cannot be found in the candidate video frames acquired by the acquisition device 2 for a continuous period of time, for example, if the features of the candidate objects in multiple candidate video frames have a similarity to the features of object 1 that is lower than the first similarity threshold, then the weight of object 1 can be gradually reduced, that is, the confidence of finding object 1 can be gradually reduced, thereby avoiding misjudgment. However, if object 1 is found in the candidate video frames acquired by the acquisition device 2 during the process of gradually reducing the matching weight of object 1, then the matching weight of object 1 can be immediately restored.

[0088] Therefore, the cross-camera continuous tracking device for the same object based on different features and scenes provided in this application calculates similarity based on the object features of the target object determined in video frames acquired by different acquisition devices with interactive areas in a sequential manner, and performs object matching based on the similarity. This enables the use of multiple acquisition devices with interactive areas to track the trajectory of the target object in a large area that cannot be covered by a single acquisition device.

[0089] Example 5

[0090] The above describes the internal functions and structure of a cross-camera continuous tracking device for the same object based on different features and scenes, which can be implemented as an electronic device. Figure 5 A schematic diagram illustrating the structure of an embodiment of the electronic device provided in this application. (See attached diagram.) Figure 5 As shown, the electronic device includes a memory 51 and a processor 52.

[0091] Memory 51 is used to store programs. In addition to the programs described above, memory 51 can also be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.

[0092] The memory 51 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0093] Processor 52 is not limited to a central processing unit (CPU), but may also be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an embedded neural network processor (NPU), or an artificial intelligence (AI) chip. Processor 52 is coupled to memory 51 and executes the program stored in memory 51. When the program runs, it executes the cross-camera continuous tracking method for the same object based on different features and scenes described in embodiments two and three above.

[0094] Furthermore, such as Figure 5 As shown, the electronic device may also include other components such as a communication component 53, a power supply component 54, an audio component 55, and a display 56. Figure 5 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 5 The components shown.

[0095] Communication component 53 is configured to facilitate wired or wireless communication between electronic devices and other devices. The electronic devices can access wireless networks based on communication standards, such as WiFi, 3G, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 53 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 53 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0096] Power supply component 54 provides power to various components of the electronic device. Power supply component 54 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.

[0097] Audio component 55 is configured to output and / or input audio signals. For example, audio component 55 includes a microphone (MIC) configured to receive external audio signals when the electronic device is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 51 or transmitted via communication component 53. In some embodiments, audio component 55 also includes a speaker for outputting audio signals.

[0098] Display 56 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation.

[0099] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for continuous tracking of the same object across cameras based on different features and scenes, comprising: Feature extraction is performed on the target object identified in the target video frame to obtain the target object features, wherein the target video frame is acquired by the first video source; Feature extraction is performed on candidate objects identified in candidate video frames to obtain candidate object features, wherein the candidate video frames are acquired by a second video source, and the acquisition range of the first video source and the acquisition range of the second video source have an interactive area; Based on the target object features and the candidate object features, calculate the similarity between the candidate object determined in at least one candidate video frame and the target object, wherein the acquisition time of the candidate video frame is later than that of the target video frame; When the similarity is greater than a first similarity threshold, it is determined that the candidate object and the target object are the same object. The method further includes: When the target object is detected to exist in the first target video frame but not in the second target video frame, and the similarity between the first candidate object determined in the first candidate video frame and the target object is not greater than the first similarity threshold, the target object is added to the target queue, and the first candidate object is added to the candidate queue. Furthermore, the matching weight of the target object is reduced as the number of times the target object fails to match at least one candidate object increases. Wherein, the second target video frame and the first target video frame are captured by the first video source, and the capture time of the second target video frame is later than that of the first target video frame; The similarity between the target object and the second candidate object determined in the second candidate video frame is calculated for matching operation. Wherein, the second target video frame is captured by the first video source, and the capture time of the second target video frame is later than that of the first target video frame; When the matching weight of the target object decreases to a preset weight threshold, the target object is removed from the target queue; When the target object successfully matches the candidate object, the initial value of the matching weight of the target object is restored; Wherein, the second video source is multiple, and the method further includes: Set matching weights for multiple second video sources; Based on the auxiliary attribute information of the target object, the matching weight of each of the second video sources is adjusted, wherein the auxiliary attribute information is one or more of position information, direction information, speed information, and size information.

2. The method for continuous cross-camera tracking of the same object based on different features and scenes according to claim 1, wherein, The first video source acquires multiple target video frames, and the method further includes: When the target object is detected to exist in the first target video frame but not in the second target video frame, the first similarity threshold for the target object is reduced. The second target video frame and the first target video frame are acquired by the first video source, and the acquisition time of the second target video frame is later than that of the first target video frame.

3. The method for continuous cross-camera tracking of the same object based on different features and scenes according to claim 1, wherein, The second video source acquires multiple candidate video frames, and the method further includes: When the similarity between the first candidate object determined in the first candidate video frame and the target object is not greater than the first similarity threshold, the similarity between the second candidate object determined in the second candidate video frame and the target object is greater than the first similarity threshold, and the similarity between the first candidate object and the second candidate object is greater than the second similarity threshold, it is determined that the first candidate object, the second candidate object, and the target object are the same object. The second candidate video frame and the first candidate video frame are acquired by the second video source, the second candidate video frame and the first candidate video frame are adjacent video frames, and the acquisition time of the second candidate video frame is later than that of the first candidate video frame.

4. The cross-camera continuous tracking method for the same object based on different features and scenes according to any one of claims 1 to 3, wherein, The method further includes: A callback function is used to pass the target object features of the target object determined in the first video source to the processing thread, wherein the processing thread is used to compare the target object features of the target object with the candidate object features of the candidate object determined in the second video source.

5. A cross-camera continuous tracking device for the same object based on different features and scenes, comprising: The first extraction module is used to extract features from a target object determined in a target video frame to obtain the features of the target object, wherein the target video frame is acquired by a first video source; The second extraction module is used to extract features from the candidate objects identified in the candidate video frames to obtain the features of the candidate objects. The candidate video frames are acquired by a second video source, and the acquisition range of the first video source and the acquisition range of the second video source have an interactive area. A calculation module is used to calculate the similarity between a candidate object determined in at least one of the candidate video frames and the target object based on the characteristics of the target object and the characteristics of the candidate object, wherein the acquisition time of the candidate video frame is later than that of the target video frame; The determination module is used to determine whether the candidate object and the target object are the same object when the similarity is greater than a first similarity threshold. The calculation module is further used for: When the target object is detected to exist in the first target video frame but not in the second target video frame, and the similarity between the first candidate object determined in the first candidate video frame and the target object is not greater than the first similarity threshold, the target object is added to the target queue, and the first candidate object is added to the candidate queue. Wherein, the second target video frame and the first target video frame are captured by the first video source, and the capture time of the second target video frame is later than that of the first target video frame; The similarity between the target object and the second candidate object determined in the second candidate video frame is calculated for matching operation. Wherein, the second target video frame is captured by the first video source, and the capture time of the second target video frame is later than that of the first target video frame; When the matching weight of the target object decreases to a preset weight threshold, the target object is removed from the target queue; When the target object successfully matches the candidate object, the initial value of the matching weight of the target object is restored; The second video source is multiple, and the calculation module is further used for Set matching weights for multiple second video sources; Based on the auxiliary attribute information of the target object, the matching weight of each of the second video sources is adjusted, wherein the auxiliary attribute information is one or more of position information, direction information, speed information, and size information.

6. An electronic device, comprising: Memory, used to store programs; A processor is configured to run the program stored in the memory to perform the cross-camera continuous tracking method for the same object based on different features and scenes as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Target tracking method and device and gun-ball linkage tracking method

    CN110414443A

  • Target detection tracking method and device, equipment and storage medium

    CN110428449A