Target tracking methods, helmet-wearing detection methods, electronic devices and storage media

By maintaining the bound target object information in the target tracking method, the problem of multiple captures caused by the occlusion of non-motorized vehicle transport target objects is solved, and continuous and efficient tracking of target objects is achieved.

CN116309697BActive Publication Date: 2026-04-03ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

During the tracking of target objects, especially when people transported by non-motorized vehicles are obscured, the target objects may not be detected for a long time, leading to the problem of multiple captures.

Method used

By ensuring that the information of bound target objects with established binding relationships is retained for the same duration as the binding relationship itself in the target tracking method, the system can guarantee that the target remains active during tracking and avoid deletion due to prolonged periods without detection. Combined with helmet-wearing detection methods, this improves the accuracy of continuous target object tracking.

Benefits of technology

It reduces the phenomenon of multiple captures during target tracking, enables continuous tracking of target objects, and improves tracking efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116309697B_ABST
    Figure CN116309697B_ABST
Patent Text Reader

Abstract

This application discloses a target tracking method, a helmet-wearing detection method, an electronic device, and a storage medium. The method includes: tracking a target moving source in a current video frame; identifying associated target objects located within a first region of the target moving source in the current video frame; determining whether the associated target object and the bound target object are the same target object based on first object information of a bound target object that has a binding relationship with the target moving source; and, in response to the association target object and the bound target object being the same target object, obtaining the object identifier of the bound target object from the first object information and using the object identifier of the bound target object as the object identifier of the associated target object. Through the above methods, this application can continuously track targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a target tracking method, a helmet-wearing detection method, an electronic device, and a storage medium. Background Technology

[0002] Currently, during the tracking of target objects transported by mobile sources, situations may arise where the target object is obscured. For example, when tracking people transported by non-motorized vehicles, the person in the back seat of the non-motorized vehicle may be obscured, making the target object undetectable for an extended period. Because the target object remains undetectable for a long time, its relevant information is deleted. However, when the target object reappears subsequently, it is assigned new information, ultimately leading to a large number of multiple captures during the tracking process. Summary of the Invention

[0003] The main technical problem addressed by this application is to provide a target tracking method, a helmet-wearing detection method, an electronic device, and a storage medium capable of continuously tracking a target.

[0004] To address the aforementioned technical problems, this application provides a target tracking method, comprising: tracking a target moving source in a current video frame; wherein the target moving source is used to carry a target object; identifying associated target objects located within a first region of the target moving source in the current video frame; determining whether the associated target object and the bound target object are the same target object based on first object information of a bound target object that has a binding relationship with the target moving source; wherein the retention time of the first object information of the bound target object follows the existence time of the binding relationship, and the binding relationship exists at least during the target tracking period, which is the period during which the target moving source can be tracked after the binding relationship is established; in response to the associated target object and the bound target object being the same target object, obtaining the object identifier of the bound target object from the first object information and using the object identifier of the bound target object as the object identifier of the associated target object; in response to the associated target object and the bound target object not being the same target object, determining whether to establish a binding relationship between the associated target object and the target moving source based on the target object to be bound to the target moving source.

[0005] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a helmet-wearing detection method, the method comprising: acquiring a current video frame; acquiring object information of at least one target object to be detected tracked in the current video frame; wherein the target object to be detected is tracked using the aforementioned target tracking method, the target movement source is a non-motorized vehicle, and the target object to be detected is a target object carried on a non-motorized vehicle; for each target object to be detected, performing helmet-wearing detection on the target object to be detected, and obtaining a helmet-wearing detection result corresponding to the target object to be detected; wherein the helmet-wearing detection result is used to indicate whether the target object to be detected is wearing a helmet.

[0006] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide an electronic device, which includes a processor and a memory, wherein the memory stores program instructions and the processor executes the program instructions to implement the above-mentioned target tracking method.

[0007] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium for storing program instructions that can be executed to implement the above-mentioned target tracking method.

[0008] In the above technical solution, the retention time of the first object information of the bound target object follows the existence time of the binding relationship. The binding relationship remains in existence at least during the target tracking period, which is the period during which the target mobile source can be tracked after the binding relationship is established. In other words, at least during the period during which the target mobile source can be tracked, the first object information of the bound target object that has established a binding relationship with the target mobile source will not be deleted because the bound target object has not been detected for a long time. This allows for continuous tracking of the bound objects of the target mobile source and reduces the problem of a large number of multiple captures during the target tracking process. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating an embodiment of the target tracking method provided in this application;

[0010] Figure 2 This is a schematic diagram of an embodiment of the target mobile source provided in this application;

[0011] Figure 3 This is a schematic diagram of the structure of an embodiment of the multi-task training model provided in this application;

[0012] Figure 4 This is a schematic diagram of an embodiment of the second region provided in this application;

[0013] Figure 5 This is a partial schematic diagram of an embodiment of the target mobile source provided in this application;

[0014] Figure 6 yes Figure 1 The flowchart of step S13 shown is a schematic diagram of one embodiment.

[0015] Figure 7 This is a flowchart illustrating another embodiment of the target tracking method provided in this application;

[0016] Figure 8 yes Figure 7 The flowchart of step S71 shown is a schematic diagram of an embodiment.

[0017] Figure 9 This is a flowchart illustrating an embodiment of the helmet wearing detection method provided in this application;

[0018] Figure 10 This is a schematic diagram of the structure of an embodiment of the helmet wearing classification model provided in this application;

[0019] Figure 11 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application;

[0020] Figure 12 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation

[0021] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0022] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0023] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0024] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the target tracking method provided in this application. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily replace it with a similar method. Figure 1 The illustrated process sequence is limited. For example... Figure 1 As shown, this embodiment includes:

[0025] Step S11: Track the target moving source in the current video frame.

[0026] In target tracking tasks, if a target cannot be detected for an extended period, it is removed from the target tracking set. When the target reappears, it is assigned a new tracking ID, which can easily lead to a large number of multiple captures in the target tracking task. The method in this embodiment is used for continuous target tracking; the targets mentioned herein include, but are not limited to, people, animals, motor vehicles, and non-motor vehicles, etc., and are not specifically limited thereto.

[0027] In this embodiment, a target moving source is tracked in the current video frame; wherein, the target moving source is used to transport the target object. The target object described herein includes, but is not limited to, people, animals, etc.

[0028] The target mobile source includes, but is not limited to, non-motorized vehicles and motorized vehicles, etc., without specific limitations here. It should be noted that, if... Figure 2 As shown, Figure 2 This is a schematic diagram of an embodiment of the target mobile source provided in this application. Taking a non-motorized vehicle as an example, the non-motorized vehicle tracked in the current video frame can be 1, 2, 3 or more, etc., and no specific limitation is made here.

[0029] In one implementation, before tracking the target moving source in the current video frame, further information is obtained.

[0030] Take the motion source information of at least one moving source in the current video frame; wherein, the specific information content included in the motion source information is not limited, such as the motion source information including the characteristics of the moving source.

[0031] Information, location information, etc. At this point, the target moving source is tracked in the current video frame, specifically:

[0032] Based on the feature information in the moving source information, the system identifies the same moving source present in the current video frame and the previous video frame, and uses this same moving source as the target moving source. In other words, it compares the current video...

[0033] By comparing the feature information of the moving source in the current video frame with that of the moving source in the previous video frame, we can determine whether the same moving source exists in the current video frame and the previous video frame.

[0034] The system uses a stable, existing moving source as the target moving source, rather than directly using the moving source detected in the current video frame. This reduces the computational load and improves the efficiency of continuous target tracking.

[0035] In one embodiment, the location information in the mobile source information is used to locate the position of the mobile source detection area, specifically including the x-axis coordinate of the upper left corner of the mobile source detection area, the upper left corner of the mobile source detection area, and the position of the upper left corner of the mobile source detection area.

[0036] The y-axis coordinate of the lower right corner, the x-axis coordinate of the lower right corner, and the y-axis coordinate of the lower right corner. Among them, such as... Figure 2 As shown, the moving source detection area is the detection box of the moving source, which is the largest bounding rectangle containing the moving source and its corresponding associated target object.

[0037] In one embodiment, the feature information of the mobile source is a 128-dimensional REID feature. In a specific embodiment, such as... Figure 3 As shown, Figure 3 This is the multi-task training model provided in this application.

[0038] A schematic diagram of one embodiment shows that the multi-task training model uses Yolov5 as the backbone network and combines it with a Feature Pyramid Network (FPN). The mobile source information includes the location information and feature information of the mobile source. The location information and the corresponding feature information can be output simultaneously in one forward inference process of the network. Compared with using two independent networks to perform location detection and feature extraction tasks separately, the method of reusing a backbone network for detection and feature extraction tasks improves the speed of obtaining the mobile source information and consumes less time. It should be noted that the location information and feature information of the mobile source share a heat map in the network output.

[0039] In one embodiment, the moving source detection area is a first area range, that is, the moving detection area is equivalent to the first area range. For example... Figure 4 As shown, Figure 4 This is a schematic diagram of an embodiment of the second region provided in this application. Because in congested situations, the target object carried by the preceding target mobile source may appear within the first region of the following target mobile source, leading to errors in the associated target object of the subsequently determined target mobile source, in other embodiments, the mobile source detection area is the area within the second region, and the first region includes the second region. In one specific embodiment, the second region is a square region with a side length equal to the width of the first region.

[0040] Step S12: Locate the associated target object within the first region of the target moving source from the current video frame.

[0041] Because the current video frame corresponds to a large video frame image, there may be multiple target objects within the video frame image; however, not all target objects need to be tracked; for example... Figure 2As shown, taking a non-motorized vehicle as the target moving source and a person as the target object as an example, the person actually needs to be continuously tracked, while pedestrians do not need to be tracked. Therefore, in this embodiment, the associated target object located within the first region of the target moving source is found from the current video frame. That is, the target object located within the first region of the target moving source is taken as the associated target object.

[0042] In one embodiment, before tracking the target moving source in the current video frame, object information of at least one target object in the current video frame is acquired. The specific content of the object information is not limited; for example, the object information may include the target object's feature information, location information, etc. Then, associated target objects located within a first region of the target moving source are identified from the current video frame. Specifically, based on the location information in the object information and the location information in the moving source information, associated target objects located within the first region of the target moving source are identified from at least one target object. In other words, by using the object information of the target object and the location information of the target moving source, it is determined whether the target object is located within the first region, and the target object located within the first region of the target moving source is identified as an associated target object.

[0043] In one embodiment, the location information in the object information is used to locate the position of the object detection region of the target object, specifically including the x-coordinate of the upper left corner, the y-coordinate of the upper left corner, the x-coordinate of the lower right corner, and the y-coordinate of the lower right corner of the object detection region. For example, Figure 2 As shown, the object detection region is the detection box of the target object, and the detection box of the target object is the largest bounding rectangle containing the target object or the largest bounding rectangle containing the head of the target object.

[0044] In one embodiment, the feature information of the target object is a 128-dimensional REID feature. In a specific embodiment, the object information includes the location information and feature information of the target object, utilizing, for example... Figure 3 The multi-task training model shown can obtain the location information and corresponding feature information of the target object simultaneously by performing one network forward inference. Compared with using two independent networks to perform location detection and feature extraction tasks separately, the method of reusing a backbone network for detection and feature extraction tasks improves the speed of obtaining the object information of the target object and consumes less time. It should be noted that the location information and feature information of the target object share the same heatmap in the network output.

[0045] In one specific implementation, the mobile source information of at least one mobile source and the object information of at least one target object are obtained by a detection model. For example, such as Figure 3 As shown, input the current video frame as follows Figure 3 The multi-task training model shown can obtain the mobile source information of at least one mobile source and the object information of at least one target object through one forward inference process.

[0046] For example, taking a non-motorized vehicle as the mobile source, a person as the target object, and both the mobile source information and the object information including location information and feature information, with the location information including the x-coordinate of the top left corner, the y-coordinate of the top left corner, the x-coordinate of the bottom right corner, and the y-coordinate of the bottom right corner; the feature information and the location information are in one-to-one correspondence. Figure 2 In other words, Figure 2 Enter to Figure 3 The multi-task training model shown acquires motion source information for two mobile sources and object information for three target objects, specifically: [[class, ul_x, ul_y, lr_x, lr_y, embedding]1, [class, ul_x, ul_y, lr_x, lr_y, embedding]2, [[class, ul_x, ul_y, lr_x, lr_y, embedding]3, [class, ul_x, ul_y, lr_x, lr_y, embedding]4 and [[class, ul_x, ul_y, lr_x, lr_y, embedding]5, where each item represents the category of the detection region, the x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, the y-coordinate of the bottom-right corner, and the 128-dimensional REID feature, respectively.

[0047] Step S13: Based on the first object information of the bound target object that has a binding relationship with the target mobile source, determine whether the associated target object and the bound target object are the same target object.

[0048] In this embodiment, based on the first object information of the bound target object that has a binding relationship with the target mobile source, it is determined whether the associated target object and the bound target object are the same target object. The retention time of the first object information of the bound target object follows the existence time of the binding relationship, and the binding relationship exists at least during the target tracking period, which is the period during which the target mobile source can be tracked after the binding relationship is established. That is, at least during the period during which the target mobile source can be tracked, the first object information of the bound target object that has a binding relationship with the target mobile source will not be deleted because the bound target object has not been detected for a long time, thus enabling continuous tracking of the bound objects of the target mobile source and reducing the problem of numerous multi-captures during the target tracking process.

[0049] For example, taking a non-motorized vehicle as the target moving source and a human head as the associated target object: Figure 2 and Figure 5As shown, Figure 5 This is a partial schematic diagram of an embodiment of the target mobile source provided in this application. Figure 5 (a) and Figure 5 (b) In the corresponding video frame, the head a of the person in the back row of non-motorized vehicle A is obscured, which will make... Figure 5 (a) and Figure 5 (b) The head a cannot be detected in the corresponding video frame; however, if the head a is a bound target object that has established a binding relationship with the non-motorized vehicle A, the first object information of the head a will not be deleted when the head a is blocked for a long time and cannot be detected, so as to achieve continuous tracking of the head a during the period when the non-motorized vehicle A can be tracked.

[0050] In one embodiment, the retention time of the first object information of the bound target object is the same as the existence time of the binding relationship. That is, as long as the binding relationship between the bound target object and the target mobile source exists, and the target mobile source is tracking the target, the first object information of the bound target object will never be deleted.

[0051] In one embodiment, the binding relationship is established when the bound target object is first detected to meet a preset binding condition in a historical video frame. The preset binding condition is that the bound target object is located within a first region of the target moving source in a preset number of historical video frames. The preset number is a positive integer, and the target moving source is tracked in every video frame between the historical video frame and the current video frame. In other words, the binding relationship with the target moving source is established after the bound target object first meets the preset binding condition to avoid errors in binding relationship establishment caused by the target being in the same frame for a short period of time.

[0052] Step S14: In response to the fact that the associated target object and the bound target object are the same target object, obtain the object identifier of the bound target object from the first object information, and use the object identifier of the bound target object as the object identifier of the associated target object.

[0053] In this embodiment, in response to the fact that the associated target object and the bound target object are the same target object, the object identifier of the bound target object is obtained from the first object information, and the object identifier of the bound target object is used as the object identifier of the associated target object. That is, after determining that the associated target object and the bound target object are the same target object, the object identifier of the bound target object is used as the object identifier of the associated target object, so that even if the bound target object disappears or appears intermittently, it can maintain a unique object identifier throughout the entire process, thereby avoiding the problem of a large number of multiple captures during the tracking process.

[0054] Step S15: In response to the fact that the associated target object and the bound target object are not the same target object, determine whether to establish a binding relationship between the associated target object and the target mobile source based on the target object to be bound to the target mobile source.

[0055] In this embodiment, in response to the fact that the associated target object and the already bound target object are not the same target object, it is determined whether to establish a binding relationship between the associated target object and the target mobile source based on the object to be bound to the target mobile source. In other words, when the associated target object and the already bound target object are not the same target object, the associated target object cannot be considered as not needing to establish a binding relationship with the target mobile source, or in other words, it cannot be considered as a target object that does not need to be tracked; further determination is needed to determine whether to establish a binding relationship between the associated target object and the target mobile source.

[0056] In the above embodiments, the retention time of the first object information of the bound target object follows the existence time of the binding relationship. The binding relationship exists at least during the target tracking period, which is the period during which the target mobile source can be tracked after the binding relationship is established. That is to say, at least during the period during which the target mobile source can be tracked, the first object information of the bound target object that has established a binding relationship with the target mobile source will not be deleted because the bound target object has not been detected for a long time, so as to continuously track the bound objects of the target mobile source and reduce the problem of a large number of multiple captures during the target tracking process.

[0057] Please see Figure 6 , Figure 6 yes Figure 1 The flowchart shown is a schematic diagram of one embodiment of step S13. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily follow the same pattern. Figure 6 The illustrated process sequence is limited. For example... Figure 6 As shown, the first object information of the bound target object includes the first object characteristics of the bound target object. In this embodiment, it includes:

[0058] Step S61: Obtain the first matching result between the first object feature and the second object feature of the associated target object.

[0059] In this embodiment, a first matching result is obtained between the first object feature and the second object feature of the associated target object. That is, a first matching result is obtained between the first object feature of each bound target object and the second object feature of the associated target object, so as to determine whether the associated target object is any of the bound target objects based on the first matching result.

[0060] Step S62: In response to the first matching result that the first object feature matches the second object feature, determine that the associated target object and the bound target object are the same target object.

[0061] In this embodiment, in response to the first matching result being a match between the first object feature and the second object feature, it is determined that the associated target object and the bound target object are the same target object. That is, when the first matching result is a match between the first object feature and the second object feature, it indicates that the associated target object matches a bound target object, and the associated target object and the matched bound target object are the same target object.

[0062] In one implementation, in response to a first matching result indicating a mismatch between the first object feature and the second object feature, it is determined that the associated target object and the bound target object are not the same target object. That is, when the first matching result indicates a mismatch between the first object feature and the second object feature, it means that the associated target object does not match any bound target object, and the associated target object and any bound target object are different target objects.

[0063] In one embodiment, in response to the association target object and the bound target object being the same target object, the first object feature of the bound target object matching the association target object is updated using the second object feature of the association target object. That is, when the association target object and the bound target object are the same target object, the first object feature of the bound target object matching the association target object is updated accordingly, specifically updated to the second object feature of the association target object.

[0064] In one implementation, in response to the association target object and the bound target object being the same target object, a region with higher region confidence and larger region size is selected from the object detection region corresponding to the association target object and the currently saved optimal object region of the bound target object, and saved as the latest optimal object region of the bound target object. That is, each bound target object has a corresponding cache area to store its corresponding optimal object region, and the corresponding optimal object region is updated during the continuous tracking of each bound target object; furthermore, the cache area only stores the optimal object region, reducing the system load. For example, if the region confidence of the currently saved optimal object region of the bound target object A is greater than the confidence of the object detection region corresponding to the matched associated target object, then the currently saved optimal object region of the bound target object A remains unchanged; if the region confidence of the currently saved optimal object region of the bound target object A is less than or greater than the confidence of the object detection region corresponding to the matched associated target object, then the object detection region corresponding to the matched associated target object is taken as the latest optimal object region of the bound target object A; if the region confidence of the currently saved optimal object region of the bound target object A is equal to the confidence of the object detection region corresponding to the matched associated target object, and the size of the object detection region corresponding to the matched associated target object is greater than the currently saved optimal object region of the bound target object A, then the object detection region corresponding to the matched associated target object is taken as the latest optimal object region of the bound target object A.

[0065] In one specific implementation, in order to improve the accuracy of subsequent detections such as helmet wearing based on the optimal object region of the bound target object, the determined optimal object region of the bound target object will be expanded so that the expanded optimal object region includes the complete bound target object.

[0066] In one implementation, such as Figure 7 As shown, Figure 7 This is a flowchart illustrating another embodiment of the target tracking method provided in this application. The target tracking method provided in this application further includes the following steps:

[0067] Step S71: In response to the fact that the associated target object and the bound target object are not the same target object, determine whether the associated target object and the bound target object are the same target object based on the second object information of the target mobile source's target object to be bound.

[0068] To prevent binding errors caused by targets appearing in the same frame for a short period of time, associated target objects that appear within the first area of ​​the target mobile source will not be directly bound to the target mobile source. Only after they meet the preset binding conditions will they be bound to the target mobile source and become bound target objects.

[0069] Therefore, in this embodiment, in response to the fact that the associated target object and the bound target object are not the same target object, the system determines whether the associated target object and the bound target object are the same object based on the second object information of the target mobile source's target object to be bound. In other words, in the association...

[0070] When the target object and the bound target object are not the same target object, the associated target object 5 cannot be regarded as a target object that does not need to establish a binding relationship with the target mobile source or that does not need to be tracked. It is necessary to further determine whether the associated target object and the target object to be bound are the same target object.

[0071] In one implementation, such as Figure 8 As shown, Figure 8 yes Figure 7 Step S71 shown in one embodiment

[0072] The flowchart illustrates that the second object information of the target object to be bound includes the third object feature of the target object to be bound (0). Based on the second object information of the target object to be bound from the target mobile source, the process is as follows:

[0073] Determining whether the associated target object and the target object to be bound are the same target object includes the following sub-steps:

[0074] Step S81: Obtain the second matching result between the second object feature and the third object feature.

[0075] In this embodiment, a second matching result is obtained between the second object feature and the third object feature. That is, the third object feature of each target object to be bound is obtained along with the associated target feature.

[0076] The second matching result between the second object features of the target object is used to determine whether the associated target object is any target object to be bound, based on the second matching result.

[0077] Step S82: In response to the second matching result that the second object feature matches the third object feature, determine that the associated target object and the target object to be bound are the same target object.

[0078] In this embodiment, in response to the second matching result being a match between the second object feature and the third object feature...

[0079] Feature matching determines that the associated target object and the target object to be bound are the same target object. In other words, when the second matching result shows that the second object feature matches the third object feature, it indicates that the associated target object matches a target object to be bound, and the associated target object and the matched target object to be bound are the same target object.

[0080] 5. In one embodiment, in response to the associated target object and the target object to be bound being the same target...

[0081] The target object is updated using the second object feature of the associated target object to the third object feature of the target object to be bound. In other words, when the associated target object and the target object to be bound are the same target object, the third object feature of the target object to be bound that matches the associated target object will be updated accordingly, specifically updated to the second object feature of the associated target object.

[0082] In one embodiment, in response to the fact that the associated target object and the target object to be bound are not the same target object, the associated target object is determined as a new target object to be bound to the target mobile source.

[0083] Step S83: In response to the second matching result that the second object feature does not match the third object feature, determine that the associated target object and the target object to be bound are not the same target object.

[0084] In one implementation, in response to a second matching result indicating a mismatch between the second object feature and the third object feature, it is determined that the associated target object and the target object to be bound are not the same target object. That is, when the second matching result indicates a mismatch between the second object feature and the third object feature, it means that the associated target object does not match any target object to be bound, and the associated target object and any target object to be bound are different target objects.

[0085] Step S72: In response to the fact that the associated target object and the target object to be bound are the same target object and the target object to be bound meets the preset binding conditions, the target object to be bound is bound to the target mobile source to establish a binding relationship, so as to become a new bound target object.

[0086] In this embodiment, in response to the fact that the associated target object and the target object to be bound are the same target object, and the target object to be bound meets the preset binding conditions, a binding relationship is established between the target object to be bound and the target mobile source, so as to make it a new bound object. That is to say, after the target object to be bound meets the preset binding conditions, a binding relationship will be established between the target object to be bound and the target mobile source, and it will be continuously tracked as a new bound target object.

[0087] In one embodiment, in response to the association target object and the target object to be bound being the same target object, the number of historical video frames in which the target object to be bound appears is counted; and in response to the number of historical video frames corresponding to the target object to be bound meeting a preset requirement, a binding relationship is established between the target object to be bound and the target mobile source. That is, when it is determined that the target object to be bound has been detected several times in history, it indicates that the target object to be bound did not appear in a short period of time, but has been continuously present within the first region of the target mobile source. Therefore, when it is determined that the number of historical video frames corresponding to the target object to be bound meets the preset requirement, it is considered that the preset binding requirement has been met, and at this time, a binding relationship is established between the target object to be bound and the target mobile source.

[0088] In one specific implementation, the target object to be bound is deleted in response to the existence time of the target object exceeding a preset time threshold. That is, for a number of target objects to be bound corresponding to the target mobile source, a queue is maintained. For each target object to be bound in the queue, if the existence time of the target object to be bound in the queue exceeds the preset time threshold, it is considered that the target object to be bound only appeared in the first area of ​​the target mobile source for a short time, and at this time the target object to be bound is deleted.

[0089] In one specific implementation, in response to the number of target objects to be bound exceeding a preset number

[0090] The threshold is used to delete the oldest pending target object (object 5) among the currently pending target objects of the target mobile source. In other words, for a number of pending target objects corresponding to a target mobile source, a threshold is maintained.

[0091] To avoid excessive queue size and maintenance issues, when a new target object is added to the queue and the total number of target objects exceeds a preset threshold, the oldest target object added to the queue is removed.

[0092] 0 Please see Figure 9 , Figure 9 This is a schematic flowchart of an embodiment of the helmet wearing detection method provided in this application. It should be noted that if substantially the same result is obtained, this embodiment is not necessarily identical. Figure 9 The illustrated process sequence is limited. For example... Figure 9 As shown, this embodiment includes:

[0093] Step S91: Obtain the current video frame.

[0094] The method in this embodiment is used for helmet wearing detection. In this embodiment, the current 5 frames of video are acquired. In one embodiment, the current frames can be obtained from local storage or cloud storage.

[0095] The previous video frame. Of course, in other implementations, the current video frame can also be acquired in real time using a video capture device, and this is not specifically limited here.

[0096] Step S92: Obtain object information of at least one target object to be detected obtained from tracking in the current video frame.

[0097] In this embodiment, object information of at least one target object to be detected is obtained from the current video frame; wherein, the target object to be detected is tracked using the target tracking method described above, the target moving source is a non-motorized vehicle, and the target object to be detected is a target object carried on a non-motorized vehicle.

[0098] Step S93: For each target object to be tested, perform helmet wearing detection on the target object to obtain the helmet wearing detection result corresponding to the target object.

[0099] In this embodiment, for each target object to be detected, a helmet wearing detection is performed on the target object to be detected to obtain the helmet wearing detection result corresponding to the target object to be detected; wherein, the helmet wearing detection result is used to indicate whether the target object to be detected is wearing a helmet.

[0100] In one implementation, such as Figure 10 As shown, Figure 10 This is a schematic diagram of an embodiment of the helmet wearing classification model provided in this application. The model utilizes a deep neural network-based helmet wearing classification model to detect helmet wearing on the target object, obtaining the helmet wearing detection result corresponding to the target object. Specifically, ResNet18 is a feature extraction network; the fc1 layer is a fully connected layer, representing a highly abstract representation of each input image after the feature extraction layer; and fc2 transforms the high-dimensional features output by fc1 into a two-dimensional one-hot encoding prediction result, outputting whether a helmet is worn.

[0101] In one specific implementation, for each target object to be detected, a helmet-wearing classification model based on a deep neural network is used to perform helmet-wearing detection on the expanded map of the optimal object region of the target object, resulting in high helmet-wearing detection accuracy. Furthermore, since only the expanded map of the optimal object region of the target object is acquired for helmet-wearing detection throughout the entire tracking sequence, a larger model with deeper network layers and greater width can be selected to improve the accuracy of helmet-wearing detection; for example, such as... Figure 10 As shown, the feature extraction network can be replaced with ResNet50, ResNet101, ResNet152, etc.

[0102] Please see Figure 11 , Figure 11This is a schematic diagram of an embodiment of the electronic device provided in this application. The electronic device 110 includes a memory 111 and a processor 112 coupled to each other. The processor 112 is used to execute program instructions stored in the memory 111 to implement the steps of any of the above-described target tracking method or helmet wearing detection method embodiments. In a specific implementation scenario, the electronic device 110 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 110 may also include mobile devices such as laptops and tablets, which are not limited here.

[0103] Specifically, processor 112 controls itself and memory 111 to implement the steps of any of the above-described target tracking method or helmet-wearing detection method embodiments. Processor 112 may also be referred to as a CPU (Central Processing Unit). Processor 112 may be an integrated circuit chip with signal processing capabilities. Processor 112 may also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor. Furthermore, processor 112 may be implemented using integrated circuit chips.

[0104] Please see Figure 12 , Figure 12 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. The computer-readable storage medium 120 of this application embodiment stores program instructions 121. When executed, these program instructions 121 implement the methods provided by any embodiment of the target tracking method or helmet wearing detection method of this application, as well as any non-conflicting combination thereof. The program instructions 121 can be formed into a program file and stored in the aforementioned computer-readable storage medium 120 in the form of a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) executes all or part of the steps of the methods of various embodiments of this application. The aforementioned computer-readable storage medium 120 includes various media capable of storing program code, such as a USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or terminal devices such as computers, servers, mobile phones, and tablets.

[0105] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

[0106] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A target tracking method, characterized in that, The method includes: A target moving source is tracked in the current video frame; wherein the target moving source is used to carry the target object; From the current video frame, identify the associated target object located within the first region of the target moving source; Based on the first object information of the bound target object that has a binding relationship with the target mobile source, it is determined whether the associated target object and the bound target object are the same target object; wherein, the retention time of the first object information of the bound target object follows the existence time of the binding relationship, the binding relationship exists at least during the target tracking period, the target tracking period is the period during which the target mobile source can be tracked after the binding relationship is established, the binding relationship is established when the bound target object is first detected to meet the preset binding conditions in the historical video frames, the preset binding conditions are that the bound target object is located within the first region of the target mobile source in a preset number of historical video frames, the preset number is a positive integer, and the target mobile source is tracked in each video frame between the historical video frames and the current video frame; In response to the fact that the associated target object and the bound target object are the same target object, the object identifier of the bound target object is obtained from the first object information, and the object identifier of the bound target object is used as the object identifier of the associated target object; In response to the fact that the associated target object and the bound target object are not the same target object, based on the target object to be bound to the target mobile source, it is determined whether to establish the binding relationship between the associated target object and the target mobile source.

2. The method according to claim 1, characterized in that, The retention time of the first object information of the bound target object is the same as the existence time of the binding relationship.

3. The method according to claim 1, characterized in that, The first object information of the bound target object includes the first object characteristics of the bound target object; determining whether the associated target object and the bound target object are the same target object based on the first object information of the bound target object that has a binding relationship with the target mobile source includes: Obtain a first matching result between the first object feature and the second object feature of the associated target object; wherein the second object feature is extracted from the current video frame; In response to the first matching result indicating that the first object feature matches the second object feature, it is determined that the associated target object and the bound target object are the same target object; and / or, In response to the first matching result indicating that the first object feature does not match the second object feature, it is determined that the associated target object and the bound target object are not the same target object; And / or, the method further includes at least one of the following steps: In response to the fact that the associated target object and the bound target object are the same target object, the first object feature of the bound target object that matches the associated target object is updated using the second object feature of the associated target object; In response to the fact that the associated target object and the bound target object are the same target object, a region with higher confidence and larger size is selected from the object detection region corresponding to the associated target object and the currently saved optimal object region of the bound target object, and saved as the latest optimal object region of the bound target object.

4. The method according to claim 1, characterized in that, In response to the fact that the associated target object and the bound target object are not the same target object, determining whether to establish the binding relationship between the associated target object and the target mobile source based on the target object to be bound to the target mobile source includes: In response to the fact that the associated target object and the bound target object are not the same target object, based on the second object information of the target mobile source to be bound, it is determined whether the associated target object and the target object to be bound are the same target object; wherein, the target object to be bound is located within the first area of ​​the target mobile source; In response to the fact that the associated target object and the target object to be bound are the same target object and the target object to be bound meets the preset binding conditions, the binding relationship is established between the target object to be bound and the target mobile source, so as to serve as the new bound target object.

5. The method according to claim 4, characterized in that, The step of establishing the binding relationship between the target object to be bound and the target mobile source in response to the fact that the associated target object and the target object to be bound are the same target object and the target object to be bound meets the preset binding conditions includes: In response to the fact that the associated target object and the target object to be bound are the same target object, the number of historical video frames in which the target object to be bound appears is counted. In response to the number of historical video frames corresponding to the target object to be bound meeting a preset number requirement, the binding relationship is established between the target object to be bound and the target mobile source; And / or, the method further includes at least one of the following steps: If the existence time of the target object to be bound exceeds a preset time threshold, the target object to be bound is deleted. In response to the number of target objects to be bound exceeding a preset threshold, the earliest existing target object to be bound among the current target objects to be bound to the target mobile source is deleted.

6. The method according to claim 4, characterized in that, The second object information of the target object to be bound includes the third object characteristics of the target object to be bound; determining whether the associated target object and the target object to be bound are the same target object based on the second object information of the target mobile source includes: Obtain the second matching result between the second object feature and the third object feature; In response to the second matching result that the second object feature matches the third object feature, it is determined that the associated target object and the target object to be bound are the same target object; In response to the second matching result indicating that the second object feature does not match the third object feature, it is determined that the associated target object and the target object to be bound are not the same target object; The method further includes at least one of the following steps: In response to the fact that the associated target object and the target object to be bound are the same target object, the third object feature of the target object to be bound that matches the associated target object is updated using the second object feature of the associated target object; In response to the fact that the associated target object and the target object to be bound are not the same target object, the associated target object is determined as a new target object to be bound to the target mobile source.

7. The method according to claim 1, characterized in that, Before tracking the target moving source in the current video frame, the method further includes: Obtain motion source information of at least one moving source and object information of at least one target object in the current video frame; The step of tracking the target moving source in the current video frame includes: Based on the feature information in the moving source information, the same moving source is determined to exist in the current video frame and the previous video frame, and the same moving source is taken as the target moving source. The step of finding the associated target object located within the first region of the target moving source from the current video frame includes: Based on the location information in the object information and the location information in the mobile source information, the associated target object located within the first area range of the target mobile source is found from the at least one target object.

8. A method for detecting helmet wearing, characterized in that, The method includes: Get the current video frame; Obtain object information of at least one target object to be detected tracked in the current video frame; wherein the target object to be detected is tracked using the target tracking method as described in any one of claims 1-7, the target moving source is a non-motorized vehicle, and the target object to be detected is a target object carried on a non-motorized vehicle; For each of the target objects to be detected, a helmet wearing detection is performed on the target object to obtain a helmet wearing detection result corresponding to the target object; wherein, the helmet wearing detection result is used to indicate whether the target object to be detected is wearing a helmet.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing program instructions, and the processor executing the program instructions to implement the target tracking method as described in any one of claims 1-7, or the helmet wearing detection method as described in claim 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program instructions that can be executed to implement the target tracking method as described in any one of claims 1-7, or the helmet-wearing detection method as described in claim 8.

Citation Information

Patent Citations

  • Target tracking method, target tracking device and computer readable medium

    CN111161320A

  • Non-motor vehicle target tracking method and device and electronic equipment

    CN114882491A