Label uniforming method based on multiple object tracking and voting and a video acquisition system
Patent Information
- Application Number
- TW114135628
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-10-08
- Filing Date
- 2025-09-17
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-09-16
AI Technical Summary
Existing deep learning-based image recognition models struggle to consistently and effectively infer the category of an object due to insufficient training data, leading to inconsistent and inaccurate labeling across different image contexts, especially when objects are viewed from different angles or under varying brightness and blur conditions.
A unified labeling method based on multi-target tracking and voting, which tracks objects across multiple frames, generates counts for different categories, and updates the inference label to the category with the highest count, effectively unifying labels across frames without requiring additional training data.
This method provides reliable, consistent, and accurate object category labeling, serving as effective training material for AI models, reducing the need for manual retraining and saving human resources, while ensuring accurate identification of object categories across varying conditions.
Smart Images

Figure TWG2TB001910783_001 
Figure TWG2TB001910783_002 
Figure TWG2TB001910783_003
Abstract
Description
[Technical Field]
[0001] A method and system for managing tags, particularly a unified tagging method and image acquisition system based on multi-target tracking and voting. [Previous Technology]
[0002] With the advancement of deep learning algorithms, the field of video recognition has also developed rapidly. A key focus of development in video recognition is how to identify and infer objects from an image segment.
[0003] An image recognition method utilizing deep learning requires training an artificial intelligence (AI) model to successfully identify objects using a large amount of training data, such as files with training labels and long-term training images. However, although there are many AI models on the market that can successfully identify objects, most AI models cannot accurately identify the category to which an object belongs. The problem is that unless an object is clearly related to a category, most training data used to train AI models to identify objects does not label the category to which an object belongs. Therefore, when training the AI model to identify objects in different image contexts using the aforementioned training data, for example, when training the AI model to identify objects from different viewpoints, with different brightness, or with different levels of blur, the AI model will not be able to consistently and effectively infer the category to which the object belongs. If we want to enable the AI model to consistently and effectively infer the category of an object, currently, the amount of training data required to train the AI model needs to be increased several times over, in order to increase the number of training labels used to label the category of an object several times over. In other words, there is currently no simpler way to enable the AI model to consistently and effectively infer the category of an object without changing or increasing the amount of training data used for object recognition.
[0004] In order to help the AI model to consistently and effectively infer the category to which the object belongs without increasing the amount of training data several times during the training phase, a new label uniforming method is needed to identify images with objects. [Summary of the Invention]
[0005] This invention provides a unified labeling method and image acquisition system based on multi-target tracking and voting.
[0006] The labeling unification method of the present invention is used to manage the inference labels of an object being tracked in different frames of an image segment, so that the inference labels in these different frames are updated and unified. In this way, without re-training any AI model, the uniform label output by the labeling unification method based on multi-object tracking and voting of the present invention can consistently label the object across the various frames of the image segment.
[0007] The label unification method based on multi-target tracking and voting of the present invention is executed by a processing unit, and the label unification method includes the following steps: receiving an image segment, wherein the image segment has multiple frames; tracking an object in the frames of the image segment using multi-target tracking (MOT); labeling the object with an inference label in each of the frames of the image segment; generating multiple counts corresponding to multiple categories; determining that the category with the highest count is the corresponding unification label; and updating the inference label of the object in each frame to the unification label.
[0008] The image acquisition system of the present invention is used to execute the tagging unification method, and the image acquisition system includes: at least one camera unit for capturing an image segment, wherein the image segment has multiple frames; a processing unit connected to the at least one camera unit; wherein the processing unit performs: receiving the image segment from the at least one camera unit; tracking an object in the frames of the image segment using multi-object tracking (MOT); tagging the object with an inference label in each of the frames of the image segment; generating multiple counts corresponding to multiple categories; determining that the category with the highest count corresponds to a unified label; and updating the inference label of the object in each frame to the unified label.
[0009] By generating multiple counts corresponding to multiple categories, determining that the category with the highest count corresponds to a unified label, and updating the inferred label of the object in each frame to the unified label, the processing unit responsible for executing the label unification method can be regarded as conducting a virtual election for the categories of the object, voting for the category with the highest count among the counts as the unified label corresponding to the object. In this way, under the premise that the input received by the present invention is limited, for example, under the premise that the input of the present invention is a limited number of frames in the image segment, and without retraining any AI model, the selected unified label will be the most probable and most reliable inference choice for the category of the object. In other words, when an AI model that has not been properly trained is used and is unable to consistently and effectively infer the category of an object from the frames of an image segment, even without investing more teaching materials and time in retraining the AI model, the processing unit executing the labeling unification method of the present invention can still easily and cost-efficiently unify the inference labels of all the frames as the unified label.
[0010] By labeling the object in the image clip with the uniform tag, the image clip can not only provide a user with reliable, consistent, and accurate information about the object, but also serve as excellent training material. This training material can be used to further train an AI model, enabling the AI model to be trained more appropriately and to more accurately distinguish the category of the object.
Implementation Method
[0019] The present invention provides a label unification method based on multi-target tracking and voting, and provides an image acquisition system responsible for executing the label unification method.
[0020] This labeling unification method is used to manage and improve the problem of inconsistent inference labels for an object in different frames of an image segment, so that the inconsistent inference labels in these different frames are updated and unified. In this way, the labeling unification method enables the object to be consistently identified in these frames of the image segment.
[0021] Please refer to Figure 1. The labeling unification method based on multi-object tracking and voting includes the following steps: Step S1: Receive an image segment, wherein the image segment has multiple frames; Step S2: Track an object in the frames of the image segment using multi-object tracking (MOT); Step S3: Label the object with an inference label in each frame of the image segment; Step S4: Generate multiple counts corresponding to multiple categories; Step S5: Determine the category with the highest count as the corresponding unified label; and Step S6: Update the inference label of the object in each frame to the unified label.
[0022] By executing steps S4 to S6, the label unification method can be viewed as a virtual election for the categories of the object, selecting the category with the highest count among the counts, that is, the category corresponding to the largest count, as the unified label for the object. In this way, given the limited input received by the present invention, for example, when the input to the present invention is a limited number of frames in the image segment, and without retraining any AI model, the selected unified label will be the most probable and most reliable inference for the category of the object. In other words, when an inadequately trained AI model is used and fails to consistently and effectively infer the category of an object from the frames of an image segment, even without investing more teaching materials and time in proper retraining, the labeling unification method of this invention can easily and cost-efficiently update the inferred labels of all frames to the unified label.
[0023] By labeling the object in the video clip with the uniform tag, the video clip not only provides a user with reliable, consistent, and accurate information about the object, but also serves as excellent training material. This training material can be used to further train an AI model, enabling the AI model to be trained more appropriately and more accurately distinguish the category of the object. The uniform tagging method of this invention can help save human resources, that is, save the human resources required to manually label the video to retrain the AI model. In other words, this invention can not only quickly generate results that can be used to consistently identify the object for the user of this invention, but also save the human resources spent on manually producing training data for the AI model to retrain on how to identify the object and its category. The object in the video clip can be any form of entity, and the entity can be an animate or an inanimate. Furthermore, the category of the object can be any information or feature of the object, such as the type or model of the object.
[0024] The image acquisition system of the present invention, which executes the unified labeling method, has a processing unit and at least one camera unit. The at least one camera unit captures an image segment, wherein the image segment has multiple frames. The processing unit is connected to the at least one camera unit, and the processing unit performs the following: receiving the image segment from the at least one camera unit; tracking an object in the frames of the image segment using multi-object tracking (MOT); labeling the object with an inference tag in each frame of the image segment; generating multiple counts corresponding to multiple categories; determining that the category with the highest count corresponds to a unified label; and updating the inference tag of the object in each frame to the unified label.
[0025] Referring to Figure 2, in one embodiment of the present invention, the tag unification method is executed by an image acquisition system 100. Specifically, the image acquisition system 100 has a processing unit 10, a memory unit 20, and a camera unit 30. The processing unit 10 is electrically connected to the memory unit 20 and the camera unit 30, and the processing unit 10 is responsible for executing the tag unification method. The camera unit 30 captures an image segment 200, and the image segment 200 has multiple frames. For example, the image segment 200 has N frames, where N is a positive integer greater than 2. These N frames include a first frame 201, a second frame 202, and a last frame 20N. An object appears in multiple frames of the image segment 200, and the object is tracked by MOT (Motion of the Object).
[0026] Referring to Figure 3, in another embodiment of the present invention, the unification method is still executed by the processing unit 10 of the image acquisition system 100, but the configuration of the image acquisition system 100 is slightly different. In this embodiment, a plurality of camera units 30 are wirelessly connected to a communication unit 40, and the communication unit 40 is electrically connected to the processing unit 10. The image segment 200 is captured by one of the camera units 30. In addition, a display unit 50 is also electrically connected to the processing unit 10. The processing unit 10 controls the display unit 50 to display the image segment 200 captured by one of the camera units 30 in real time. The processing unit 10 also controls the communication unit 40 to share the real-time broadcast of the image segment 200 to an external device 300 wirelessly connected to the communication unit 40. The external device 300 can be any electronic device capable of connecting to the Internet to view the image segment 200, such as any kind of smart portable device or computer device. Smart wearable devices can be smartphones, smart glasses, or smart VR / AR wearable devices. Computer devices can be desktop computers, tablets, or laptops.
[0027] When steps S1 to S5 are executed, the image segment 200 transmitted to the display unit 50 and the external device 300 will be marked with a high reliability tag in each frame. Although, in this case, the image segment 200 is not played on the display unit 50 and the external device 300 in real time, in reality, the image segment 200 played on the display unit 50 and the external device 300 has almost no time delay. This is because the unified tagging method of the present invention can be executed efficiently by the processing unit 10, so the image segment 200 can be marked almost instantly, so that the image segment 200 can still be generally regarded as the image of the image segment 200 played on the display unit 50 and the external device 300 in real time.
[0028] Generally, in practical applications, the category of an object is usually lacking labeling, and therefore the category typically lacks a tag that establishes an association with the object. Therefore, AI models used for object recognition can only identify the object itself, but cannot identify which category it belongs to. For example, in the field of vehicle recognition, training data for training AI models to identify vehicles on the road typically includes training content that trains the AI model to identify vehicles from multiple different perspectives and in different contexts. However, only when the vehicle's logo appears in the training content from a specific perspective will the vehicle's brand have an observable association with the logo. Therefore, most AI models can only identify the vehicle on the road, but cannot consistently identify the brand to which the vehicle belongs.
[0029] In detail, since vehicles only display their brand logos at the front and rear, images showing the front or rear of the vehicle are generally sufficient for an AI model to infer the vehicle's brand. However, because the AI model did not receive brand labels under various conditions during training, it was not trained to identify the brand under different viewing angles, brightness levels, or blur levels. Therefore, when the image analyzed by the AI model is the side view of the vehicle, it may fail to identify the brand because it cannot see the vehicle logo. This means that when the vehicle turns, such as when making a U-turn, the AI model can only identify the brand when the vehicle logo appears at the front and rear of the vehicle in the image. When the side view of the vehicle appears in the image, the AI model will begin to guess wildly about the brand with low accuracy, resulting in inconsistent and inaccurate inferences. In the field of vehicle identification technology, an embodiment of the present invention can provide a solution to the aforementioned problems. The labeling unification method of the present invention can also be applied to other technical fields to solve other kinds of technical problems. The labeling unification method of the present invention is used to manage multiple inference tags assigned to an object across multiple frames; therefore, regardless of the identification technical problem to be solved, the present invention can provide the benefit of uniformly updating the inference tags of the object in each frame to the unified tag.
[0030] Referring to Figure 4, in one embodiment, the display unit 50 is a display, and the camera unit 30 that captures the first frame 201 of the image segment 200 is a monitor. In Figure 4, the first frame 201 is presented to a first user of the present invention through the display unit 50, and in the first frame 201, the motion of a first object 210 and a second object 220 is tracked by MOT. The first object 210 and the second object 220 are each labeled with a deduction tag. A first deduction tag 211 corresponding to the first object 210 indicates that the first object 210 is deduced to be a vehicle, and a second deduction tag 221 corresponding to the second object 220 indicates that the second object 220 is deduced to be a bird.
[0031] In this embodiment, a user-defined custom voting model and an object recognition model are used to generate the first inference label 211 corresponding to the first object 210 and the second inference label 221 corresponding to the second object 220. In other words, when the processing unit 10 executes the tag unification method, the processing unit 10 also uses the custom voting model and the object recognition model stored in the memory unit 20.
[0032] The object recognition model has been pre-trained using deep learning methods to identify the object in the image clip. However, like the aforementioned AI model, the object recognition model can only identify the object and cannot infer more information about the object's category in any situation. In other words, the object recognition model can identify the first object 210 as a vehicle and the second object 220 as a bird, but it cannot consistently identify the vehicle's category, such as its model number, in any situation within the image clip. Nevertheless, the object recognition model still assists the custom voting model in identifying the object in the image clip.
[0033] Referring to Figures 5A to 5D, the processing unit 10 can use MOT to track multiple objects between different frames of the image segment, and each tracked object is tracked independently by the processing unit 10. In one example, the image segment shows a vehicle making a U-turn on a road and a bird flying across the road. The first object 210 and the second object 220 are tracked by MOT between the first frame 201, the second frame 202, a third frame 203, and the last frame 20N. The processing unit 10 can distinguish that the first object 210 is a vehicle and the second object 220 is a bird.
[0034] Because both the first object 210 and the second object 220 appear in multiple frames, MOT generates an object identification group for each of the first object 210 and the second object 220, respectively, to facilitate continuous tracking of the first object 210 and the second object 220 under different viewing angles, different brightness levels, or different degrees of blur. Because the first object 210 has its own object identification group, the first object 210 can be independently identified without being affected by the second object 220. In this embodiment, since the bird is not the object that the user wants to analyze, the user only uses the tagging unification method of the present invention to manage the unification and updating of multiple first inference tags 211 among multiple frames.
[0035] Please refer to Figure 6. This custom voting model is used when labeling the object with the inference tag and generating the counts. The user can define how to calculate and generate the counts corresponding to these categories using the custom voting model. In one embodiment, step S3 is performed by the custom voting model as follows:
[0036] Step S30: In each frame of the image segment, generate multiple confidence values corresponding to the categories, and label the object with the inference tag according to the category with the highest confidence value. The confidence values corresponding to the categories are generated based on the aforementioned object recognition model.
[0037] Further, step S4 is performed by the custom voting model as follows:
[0038] Step S40: Add the confidence values corresponding to the same category to obtain the count corresponding to this category; wherein each count is the sum of the confidence values corresponding to the same category.
[0039] Corresponding to the examples in Figures 5A to 5D, by executing step S30, the first object 210 is inferred to be the content shown in Table 1: Information about the object being inferred and its corresponding confidence value (Probability of soft tags): The first inference label 211 and its corresponding highest confidence value (Highest probability of soft tags): In the first frame 201: The vehicle belongs to: Category A: 0.7 Category B: 0.2 Category C: 0.1 Category A: 0.7 In the second frame 202: The vehicle belongs to: Category A: 0.3 Category B: 0.4 Category C: 0.3 Category B: 0.4 In the third frame 203: The vehicle belongs to: Category A: 0.6 Category B: 0.2 Category C: 0.2 Category A: 0.6 In the final frame 20N: The vehicle belongs to: Category A: 0.3 Category B: 0.3 Category C: 0.4 Category C: 0.4 Table 1
[0040] By executing step S30, the category corresponding to the highest confidence value in each frame is set as the first inference label 211 of the first object 210 in each frame. For example, in the first frame 201 shown in Figure 5A, the first object 210 is inferred to have a 100% probability of being a vehicle and a 0% probability of being a bird, and the first object 210 is inferred to have a 70% probability of belonging to category A, a 20% probability of belonging to category B, and a 10% probability of belonging to category C. Because the first object 210 is most likely to be a vehicle of category A in the first frame 201, the first inference label 211 presents "category A 0.7" in the first frame 201 to represent the most likely category of the first object 210 and its corresponding confidence value. Following the same logic, the present invention can also yield the following inferences: Ø The first object 210 has the first inference label 211 in the first frame 201, and its corresponding highest confidence value is category A 0.7. Ø The first object 210 has the first inference label 211 in the second frame 202, and its corresponding highest confidence value is category B 0.4. Ø The first object 210 has the first inference label 211 in the third frame 203, and its corresponding highest confidence value is category A 0.5. Ø The first object 210 has the first inference label 211 in the last frame 20N, and its corresponding highest confidence value is category C 0.4.
[0041] Because the vehicle's logo appears in the first frame 201 and the vehicle's logo appears in the third frame 203, the vehicle's model is more correctly (confidently) inferred to be category A. Conversely, because the vehicle's logo does not appear in the side view shown in the second frame 202 and the other side view shown in the last frame 20N, the vehicle's model is incorrectly and inconsistently inferred to be category B or category C. This inference result is not unexpected, as such inference errors have occurred in the prior art. That is to say, the monitor of this embodiment photographed the vehicle with different brightness, different blur levels, and most obviously different viewing angles between Figures 5A and 5D. When the vehicle was photographed from different viewing angles and the image segment 200 was captured, the object recognition model was challenged to determine whether the vehicle's category was correctly inferred.
[0042] In order to correct the inconsistent inference labels and unify the inference labels, the present invention votes on all possible inference results of the first inference label 211 among the frames of the image segment according to the custom voting model.
[0043] By executing step S40, the confidence values corresponding to the same category are summed to form the count. In this embodiment, among the frames of the image segment 200, all possible outcomes of the first inference label 211 are participants in the election. In this election, multiple candidates are the categories, and the number of votes obtained by each candidate is the sum of all inference probabilities corresponding to each candidate. The so-called vote count is the count. In other words, in this embodiment, the sum of all confidence values (probabilities of soft labels) among the frames is the number of votes corresponding to the categories, and the sum of these votes becomes the count, thus the count takes into account all possible outcomes of the inferences among the frames. That is, all confidence values in the frames are used to generate the counts corresponding to the categories. For example, please refer to Table 2 below: First Shadow Frame: second Shadow Frame: third Shadow Frame: at last Shadow Frame: The total number of votes is calculated by adding them together. Votes for Category A: 0.7 0.3 0.6 0.3 1.9 Votes for Category B: 0.2 0.4 0.2 0.3 1.1 Votes for category C: 0.1 0.3 0.2 0.4 1.0 Table 2
[0044] In the above example, when the election ended, category A received the most votes, so the first inference label 211 was uniformly updated to a uniform label 212 across the frames of the image segment 200, that is, each uniform label 212 in Figures 7A to 7D is displayed as category A.
[0045] In another embodiment of the present invention, when step S40 is performed, the confidence values corresponding to the same category are summed to form the count. However, only the first inference label 211 with the highest confidence value (the highest probability of the soft label) participates in the election. In other words, only the highest confidence values in each frame are summed to form the votes corresponding to those categories, and the sum of those votes becomes the counts, so the counts only take into account the most representative inference results among those frames. That is, only the highest confidence value in each frame is used to generate the counts corresponding to those categories. In this example, the data presented in Table 1 is used to generate the election situation shown in Table 3 below: First Shadow Frame: second Shadow Frame: third Shadow Frame: at last Shadow Frame: The total number of votes is calculated by adding them together. Votes for Category A: 0.7 0.6 1.3 Votes for Category B: 0.4 0.4 Votes for category C: 0.4 0.4 Table 3
[0046] In the above example, when the election ended, category A still received the most votes. Therefore, the first inference label 211 was uniformly updated to a uniform label 212 across the frames of the image segment 200. That is, in Figures 7A to 7D, each uniform label 212 is displayed as category A.
[0047] For each frame, the confidence values corresponding to different categories have only a very low probability of being equal. However, when the confidence values corresponding to different categories in a frame are equal, the processing unit 10 may attempt to regenerate the confidence values in this frame, or avoid adding the confidence values in this frame to the count.
[0048] Please refer to Figure 8. In extremely low probability cases, the counts generated by the label unification method for these categories may result in a situation where the counts are equal. To solve this problem, the label unification method further includes the following sub-steps for step S5: Step S50: Determine whether the image segment corresponds to multiple highest counts; Step S51: When it is determined that the image segment corresponds to only one highest count, set the category corresponding to the highest count as the corresponding unification label, and execute step S6; Step S52: When it is determined that the image segment corresponds to multiple highest counts, for each category, calculate the occurrence count of the inference labels in all the frames of the image segment; Step S53: Determine whether multiple highest occurrence counts have occurred; Step S54: When it is determined that multiple highest occurrence counts have occurred, generate an exception message regarding the object labels; Step S55: When it is determined that only one highest occurrence count has occurred, set the category corresponding to the highest occurrence count as the corresponding unification label, and execute step S6.
[0049] In one embodiment, the abnormal message is generated by the processing unit 10, and the abnormal message may be a text message displayed on the display unit 50 or the external device 300. The text message can indicate that the object, which is difficult to classify, has appeared unusually in the image clip.
[0050] As can be seen from the above, the present invention can perform multiple elections for the highest count and the highest occurrence frequency in order to determine which category of the uniform label 212 is most appropriate for the object.
[0051] Referring to Figures 7A to 7D, as shown in Table 2, with only a single highest count, the uniform label 212 updates the first object 210 in each frame to the corresponding category A. Each uniform label 212 only presents the result of the category to which the first object 210 belongs, without presenting the count calculated in the background. In this example, the first object 210 is a vehicle, and this vehicle belongs to category A. This result is obtained because category A, to which the vehicle belongs, was voted out from multiple possible categories, and thus category A was promoted to the uniform label 212 that unifies and updates the category to which the first object 210 belongs. By reliably, consistently, and accurately updating the uniform label 212 in the frames of the image segment 200, the user of the present invention can easily and effortlessly know the category to which an object belongs when viewing the image segment 200.
[0052] Furthermore, by consistently labeling the object with the uniform tag 212 within the image segment 200, the AI model can also use the image segment 200 as a learning material to learn how, for example, to establish a relationship between the first object 210 of the vehicle and category A without being affected by the viewing angle of the first object 210 or the limitation of not displaying a logo on the side of the vehicle. [Simplified Explanation of the Diagram]
[0011] Figure 1 is a flowchart of a labeling unification method based on multi-target tracking and voting according to the present invention.
[0012] Figure 2 is a schematic diagram of an embodiment of the image acquisition system of the present invention that performs the labeling unification method.
[0013] Figure 3 is a schematic diagram of another embodiment of the image acquisition system of the present invention that performs the labeling unification method.
[0014] Figure 4 is a schematic diagram of a display unit showing the image frame of an image segment acquired by the image acquisition system of the present invention, which performs the unified marking method.
[0015] Figures 5A to 5D are schematic diagrams showing the image frames of an image segment representing objects in different states.
[0016] Figure 6 is another flowchart of the marking unification method of the present invention.
[0017] Figures 7A to 7D are schematic diagrams showing the image frames of objects marked by the marking unification method of the present invention.
[0018] Figure 8 is another flowchart of the marking unification method of the present invention.
Claims
1. A labeling unification method based on multi-object tracking and voting, executed by a processing unit, comprising the following steps: receiving an image segment, wherein the image segment has multiple frames; tracking an object in the frames of the image segment using multi-object tracking (MOT); labeling the object with an inference label in each of the frames of the image segment; generating multiple counts corresponding to multiple categories; determining the category with the highest count as the corresponding unification label; and updating the inference label of the object in each frame to the unification label; wherein, A custom voting model is used to label the object with inference tags and generate counts; wherein the custom voting model is used to perform the following steps: generating multiple confidence values corresponding to categories in each frame of the image segment, and labeling the object with the inference tag according to the category with the highest confidence value; summing the confidence values corresponding to the same category to obtain the count corresponding to this category; wherein each count is the sum of the confidence values corresponding to the same category; when it is determined that the image segment corresponds to multiple highest counts, for each category, calculating the occurrence count of the inference tags in all frames of the image segment, and setting the category with the highest occurrence count as the corresponding unified tag.
2. The tagging uniformity method as described in request item 1, wherein, The object tracked using MOT is a vehicle, and the inferred labels are different models of the vehicle; wherein, the categories are also different models of the vehicle, and the uniform label is a uniformly identified model of the vehicle.
3. The tagging uniformity method as described in claim 2, wherein, The video clip was captured by a monitor, and the frames of the video clip were captured by the monitor from different angles, different brightness levels, or different degrees of blur.
4. The tagging uniformity method as described in request item 1, wherein, When these counts are generated, all the confidence values in these frames will be used to generate the counts corresponding to these categories.
5. The tagging uniformity method as described in request item 1, wherein, When these counts are generated, only the highest confidence value in each frame will be used to generate the counts corresponding to those categories.
6. The tagging uniformity method as described in request item 1, wherein, This custom voting model is used to further perform the following sub-step: when it is determined that multiple highest occurrences of the same value have occurred, an exception message is generated regarding the labels of those objects.
7. The tagging uniformity method as described in Request 1, wherein, The confidence values for these categories are generated based on an object recognition model, which is pre-trained using deep learning methods to recognize the object.
8. An image acquisition system, comprising: At least one camera unit is used to capture an image segment, wherein the image segment has multiple frames; a processing unit is connected to the at least one camera unit; a memory unit is electrically connected to the processing unit and stores a custom voting model; wherein the processing unit performs the following actions: receiving the image segment from the at least one camera unit; tracking an object in the frames of the image segment using multi-object tracking (MOT); labeling the object with an inference label in each frame of the image segment; generating multiple counts corresponding to multiple categories; determining the category with the highest count as the corresponding unified label; and updating the inference label of the object in each frame to the unified label; wherein the processing unit uses the custom voting model. The object is labeled with inference tags and counts are generated; wherein, the custom voting model is used to perform the following steps: in each frame of the image segment, multiple confidence values corresponding to the categories are generated, and the object is labeled with the inference tag according to the category with the highest confidence value; the confidence values corresponding to the same category are summed to obtain the count corresponding to this category; wherein each count is the sum of the confidence values corresponding to the same category; when it is determined that the image segment corresponds to multiple highest counts, for each category, the occurrence count of the inference tags in all frames of the image segment is calculated, and the category with the highest occurrence count is set as the corresponding unified tag.
9. The image acquisition system as described in claim 8, further comprising: A communication unit is electrically connected to the processing unit and wirelessly connected to the at least one camera unit; wherein the processing unit is connected to the at least one camera unit through the communication unit; wherein the communication unit is for connecting to an external device, and the processing unit outputs the image of the image segment to the external device through the communication unit.
10. The image acquisition system as described in claim 8, further comprising: A display unit is electrically connected to the processing unit; wherein the processing unit controls the display unit to display the image segment.
11. The image acquisition system as described in claim 8, wherein, The object tracked using MOT is a vehicle, and the inferred labels are different models of the vehicle; wherein, the categories are also different models of the vehicle, and the uniform label is a uniformly identified model of the vehicle.
12. The image acquisition system as described in claim 11, wherein, The at least one camera unit is a monitor, and the frames of the image segment are captured by the monitor from different angles, different brightness, or different degrees of blur of the vehicle.
13. The image acquisition system as described in claim 8, wherein, When these counts are generated, all the confidence values in these frames will be used to generate the counts corresponding to these categories.
14. The image acquisition system as described in claim 8, wherein, When these counts are generated, only the highest confidence value in each frame will be used to generate the counts corresponding to those categories.
15. The image acquisition system as described in claim 8, wherein, This custom voting model is used to further perform the following sub-step: when it is determined that multiple highest occurrences of the same value have occurred, an exception message is generated regarding the labels of those objects.
16. The image acquisition system as described in claim 8, wherein, The memory unit stores an object recognition model, and the confidence values corresponding to these categories are generated based on the object recognition model; wherein the object recognition model is pre-trained by a deep learning method to recognize the object.
Citation Information
Patent Citations
Vehicle detection and labeling method and system and electronic equipment
CN113033449A
Classified voting method based on vehicle track in video
CN114140717A
Access Control System Using Deep Learning for License Plate and Face Recognition
TWI621071B