Method for updating bounding box or key point in object detection model

By introducing feature compensation programs in the object detection model, and updating bounding boxes or key points with multiple previous frames and moving vectors, the problem of insufficient accuracy in early departure technology is solved, and more efficient and accurate object detection effects are achieved.

CN120236053APending Publication Date: 2025-07-01IND TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410252067.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-03-06
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing object detection models have problems with insufficient accuracy in early departure technology, especially when the confidence level does not exceed the threshold, making it difficult to maintain efficient and accurate object detection results.

Method used

By introducing feature compensation programs in the object detection model, the bounding box or key points are updated with multiple previous and current frames and moving vectors, including dynamic system prediction algorithms and interpolation methods, to improve detection accuracy when confidence does not meet the threshold.

Benefits of technology

The real-time inference accuracy of the object detection model is improved, and the fast and accurate effect is maintained. The inference speed is increased from 14 frames per second to 55 frames per second, and the average average accuracy is increased from 47.8 to 53.1.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236053A_ABST
    Figure CN120236053A_ABST
Patent Text Reader

Abstract

A method of updating a bounding box or keypoints in an object detection model, the method of updating a bounding box in an object detection model comprising, by an arithmetic device, inputting a video to the object detection model, the video comprising a plurality of previous frames and a current frame, the object detection model detecting an object in the current frame, and outputting a current bounding box and a confidence associated with the object, when the confidence is less than a threshold, updating the current bounding box according to the plurality of previous frames, the current frame, and a motion vector associated with one of the plurality of previous frames and the current frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to object recognition and artificial intelligence models, and in particular to a method for updating bounding boxes or key points in an object detection model. Background Art

[0002] Artificial Intelligence (AI) models for multi-object recognition or key point recognition have high complexity. Early Exit is a commonly used method to shorten the inference time of AI models. The confidence of an object is output at an intermediate layer of the model. If it exceeds the threshold, the result can be output early. Currently, there is a need for an object recognition model to improve the accuracy of real-time inference so that the object detection model has both fast and accurate effects. Summary of the Invention

[0003] A method for updating a bounding box in an object detection model according to an embodiment of the present invention includes executing by a computing device: inputting a video into the object detection model, where the video includes a plurality of previous frames and a current frame, the object detection model detects an object in the current frame, and outputs a current bounding box and a confidence level associated with the object; when the confidence level is less than the threshold, updating the current bounding box according to the plurality of previous frames, the current frame, and a motion vector, where the motion vector is associated with one of the plurality of previous frames and the current frame.

[0004] A method for updating key points in an object detection model according to an embodiment of the present invention includes executing by a computing device: inputting a video into the object detection model, where the video includes a plurality of previous frames and a current frame, the object detection model detects an object in the current frame, and outputs a plurality of key points and a plurality of confidence levels associated with the object, where a candidate point among the plurality of key points corresponds to a candidate confidence level among the plurality of confidence levels; when the candidate confidence level is less than the threshold, updating the candidate point according to the plurality of previous frames, the current frame, and a motion vector, where the motion vector is associated with one of the plurality of previous frames and the current frame.

[0005] The above description of the present disclosure and the following description of the embodiments are used to demonstrate and explain the spirit and principle of the present invention, and provide a further explanation of the claims of the present invention. Brief Description of the Drawings

[0006] Figure 1 is an architecture diagram of a feature compensation program and an object detection model;

[0007] Figure 2 is a flowchart of a method for updating a bounding box in an object detection model according to an embodiment of the present invention;

[0008] Figure 3It is a schematic diagram of a motion vector;

[0009] Figure 4 It is a flowchart of a feature compensation program illustrated according to one or more embodiments of the present invention;

[0010] Figure 5 It is a flowchart of a motion vector prediction program (for bounding boxes) illustrated according to an embodiment of the present invention;

[0011] Figure 6 It is a flowchart of a method for updating key points in an object detection model illustrated according to an embodiment of the present invention;

[0012] Figure 7 It is a flowchart of a motion vector prediction program (for key points) illustrated according to an embodiment of the present invention; and

[0013] Figure 8 It is a schematic diagram of an example of a moving box.

[0014]

Symbol Explanation

[0015] x1, x2, x3: Intermediate layer

[0016] y1, y2, y3: Output of the intermediate layer

[0017] 1 - 5, 50, 60 - 63, 70, 80 - 82, 90 - 99, 901 - 904, 11 - 15, 151 - 154: Steps

[0018] P: Candidate point

[0019] L: Specified length

[0020] MB: Moving box

[0021] v1 - v8: Motion vectors Detailed Embodiment

[0022] The detailed features and advantages of the present invention are described in detail in the following embodiments. The content is sufficient for any person skilled in the art to understand the technical content of the present invention and implement it accordingly. And according to the content disclosed in this specification, the scope of the patent application and the drawings, any person skilled in the art can easily understand the related purposes and advantages of the present invention. The following embodiments further illustrate the present invention in detail, but do not limit the scope of the present invention from any perspective.

[0023] The present invention proposes a method for updating bounding boxes in an object detection model, and this method performs multiple steps through an arithmetic device. In one embodiment, the arithmetic device may include, for example, a central processing unit (CPU), a graphic processing unit (GPU), a microcontroller (MCU), an application processor (AP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a digital signal processor (DSP), a system-on-a-chip (SOC), or a deep learning accelerator. However, the present invention is not limited to these above examples.

[0024] The method for updating bounding boxes in the object detection model proposed by the present invention adds a feature compensation program in the object detection model that adopts the Early Exit technology. Please refer to Figure 1 , Figure 1 which is the architecture diagram of the feature compensation program and the object detection model. As Figure 1 shown, the input of the object detection model can be a video captured by a camera in real time, a completed video file, or any video, video stream, or video file, and the present invention is not limited thereto. The object detection model adopts, for example, a neural network, which includes multiple intermediate layers x1, x2, x3. The output y1, y2, y3 of each intermediate layer includes multiple bounding boxes and the confidence levels of these bounding boxes. Each bounding box is used to enclose an object of a specific category, and the confidence level represents the probability that the object belongs to the specific category. In one embodiment, if the confidence levels of all the bounding boxes output by a certain intermediate layer (such as x1) exceed the threshold, these bounding boxes can be output to leave the object detection model early. In one embodiment, as long as the confidence level of one bounding box does not exceed the threshold, the object detection model is not left early.

[0025] In one embodiment, for the bounding boxes whose confidence levels do not exceed the threshold, a feature compensation program is added to correct the original bounding boxes, and the feature compensation program refers to, for example, the motion vector and the historical result. Please refer to Figure 2 , Figure 2 which is the flowchart of the method for updating bounding boxes in the object detection model illustrated according to an embodiment of the present invention.

[0026] Step 1: Input a video into an object detection model. The video includes multiple previous frames and a current frame.

[0027] Step 2: The object detection model detects an object in the current frame and outputs a current bounding box and a confidence level associated with this object. As described above, this step is implemented by one of the multiple intermediate layers of the object detection model.

[0028] Step 3: Determine whether the confidence level is less than a threshold. If the determination is no, proceed to Step 4 and directly output the current bounding box. If the determination is yes, proceed to Step 5 and execute a feature compensation program.

[0029] Step 5: When the confidence level is less than the threshold, update the current bounding box based on the multiple previous frames, the current frame, and a motion vector. In one embodiment, Step 5 is the aforementioned feature compensation program, whose input includes multiple previous frames, the current frame, and the motion vector, and whose output is the updated current bounding box.

[0030] The motion vector is associated with one of the multiple previous frames and the current frame. In one embodiment, assume that five frames in chronological order from earliest to latest are A, B, C, D, and E, where A, B, C, and D are multiple previous frames, E is the current frame, and D is the previous frame adjacent to the current frame. Then the motion vector is calculated based on D and E. Since the motion vector can be automatically generated during video decoding, no additional time cost is incurred. Please refer to Figure 3 , Figure 3 which is a schematic diagram of the motion vector.

[0031] The motion vector can be used for video compression coding. In video compression design, there are three types of frames: I-frame (Intra-coded picture), P-frame (Predicted picture), and B-frame (Bidirectional predicted picture or Bi-directional pictures). An I-frame is a complete frame and does not depend on the data of other frames. A P-frame is predicted based on one or more previous frames and uses motion compensation technology to describe the difference between itself and the reference frame. These differences are represented by the motion vector and residual data. A B-frame is predicted based on the previous and subsequent frames. Therefore, there are two motion vectors, one pointing to the previous reference frame and the other pointing to the subsequent reference frame.

[0032] As described above, motion vectors are automatically generated when the current frame is a P-frame or a B-frame. In one embodiment, if the current frame is an I-frame and a compensation process is required, the current frame and a previous frame can be input to extract the motion vector, for example, through the NVIDIA Optical Flow SDK, an Application Programming Interface (API).

[0033] In one embodiment, the input of the feature compensation process includes a plurality of previous frames A, B, C, D, the current frame E, and the motion vector. Depending on whether each previous frame includes a bounding box associated with an object, it can be divided into the six situations in Table 1 below:

[0034] Table 1

[0035] A B C D Situation 1 Yes Yes Yes Yes Situation 2 No Yes Yes Yes Situation 3 Yes Yes No Yes Situation 4 Yes No Yes Yes Situation 5 Yes Yes Yes No Situation 6 No No No Yes

[0036] In Table 1, the fields marked "Yes" indicate that the previous frame has a bounding box associated with the object, and the fields marked "No" indicate that the previous frame lacks a bounding box associated with the object.

[0037] Figure 4 is a flowchart of the feature compensation process illustrated according to one or more embodiments of the present invention. Executing the Figure 4 corresponding steps in can be an embodiment of the feature compensation process.

[0038] Step 50, determine whether the number of previous frames is greater than or equal to a specified number. If the determination is no, then execute Step 60. If the determination is yes, then execute Step 70. In one embodiment, the specified number is 3.

[0039] Step 60, determine whether the previous frame D has a bounding box associated with the object. If the determination is no, then execute Step 61. If the determination is yes, then execute Step 62. Step 61, end the feature compensation process because there are not enough historical results for feature compensation. Step 62, execute a motion vector prediction process based on the bounding box of the previous frame D and the motion vector to generate a first update result. Then execute Step 63, update the current bounding box according to the first update result.

[0040] Situation 6 in Table 1 can reach Step 63. Specifically, the plurality of previous frames includes a first frame D earlier than the current frame E. When the number of the plurality of previous frames is less than the specified number and the first frame D has a first bounding box associated with the object, execute a motion vector prediction process based on the first bounding box and the motion vector to update the current bounding box.

[0041] Step 70, determine whether the previous frame D has a bounding box associated with the object. If the determination is negative, execute step 80. If the determination is positive, execute step 90.

[0042] Step 80, input the bounding boxes of the previous frames A, B, and C respectively into a dynamic system prediction algorithm to generate the bounding box of the previous frame D. In one embodiment, the dynamic system prediction algorithm employs a Kalman filter. Then execute step 81, execute the dynamic system prediction algorithm based on the previous frames A, B, C, and D to produce a second update result. Then execute step 82, update the current bounding box according to the second update result.

[0043] Situation 5 in Table 1 can reach step 82. Specifically, the multiple previous frames include a first frame D earlier than the current frame E and multiple historical frames earlier than the first frame. The multiple historical frames include a second frame C earlier than the first frame D, a third frame B earlier than the second frame C, and a fourth frame A earlier than the third frame B, where the second frame C, the third frame B, and the fourth frame A each have a second bounding box, a third bounding box, and the fourth bounding box associated with the object. Updating the current bounding box according to the multiple previous frames, the current frame, and the motion vector includes: when the number of the multiple previous frames is not less than the specified number and the first frame D lacks the first bounding box associated with the object, execute the dynamic system prediction algorithm according to the multiple historical bounding boxes associated with the object in the multiple historical frames to update the current bounding box. Executing the dynamic system prediction algorithm according to the multiple historical bounding boxes associated with the object in the multiple historical frames to update the current bounding box includes: execute the dynamic system prediction algorithm according to the second bounding box, the third bounding box, and the fourth bounding box to generate a first bounding box; and execute the dynamic system prediction algorithm according to the first bounding box, the second bounding box, the third bounding box, and the fourth bounding box to update the current bounding box.

[0044] Step 90, execute a motion vector prediction program according to the previous frame D to generate a first update result and a statistical range.

[0045] Step 91, determine whether the number of previous frames is equal to the specified number + 1. In one embodiment, the specified number is 3. Therefore, in step 91, determine whether 4 previous frames are collected. If the determination is positive, execute step 92. If the determination is negative, execute step 93.

[0046] Step 92, execute the dynamic system prediction algorithm according to the previous frames A, B, C, and D to generate a second update result. According to situation 1, it can reach step 92. After step 92 is completed, execute step 97.

[0047] Step 93, determine whether the previous frame B or C lacks a bounding box. If the determination is negative, execute step 96. If the determination is positive, execute step 94.

[0048] Step 94, perform interpolation to generate the missing bounding box. If the previous frame with a missing bounding box is B, perform interpolation based on the previous frames A, C, and D to generate the bounding box corresponding to the previous frame B. If the previous frame with a missing bounding box is C, perform interpolation based on the previous frames A, B, and D to generate the bounding box corresponding to the previous frame B. After step 94 is completed, step 95 is executed, and a dynamic system prediction algorithm is performed based on the previous frames A, B, C, and D to generate a second update result.

[0049] Situation 2 can proceed to step 95. Specifically, the multiple historical frames include: a second frame C earlier than the first frame D, a third frame B earlier than the second frame C, and a fourth frame A earlier than the third frame B. Performing a dynamic system prediction algorithm based on the multiple historical bounding boxes associated with the object in the multiple historical frames to generate a second update result includes: when the second frame C has a second bounding box associated with the object and the third frame B has a third bounding box associated with the object, performing a dynamic system prediction algorithm based on the second bounding box, the third bounding box, and the fourth bounding box associated with the object in the fourth frame to generate a second update result.

[0050] Step 96, perform a dynamic system prediction algorithm based on the previous frames B, C, and D to generate a second update result.

[0051] Situation 3 or situation 4 can proceed to step 96. Specifically, updating the current bounding box based on the multiple previous frames, the current frame, and the motion vector includes: when the number of the multiple previous frames is not less than the specified number and the first frame D has a first bounding box associated with the object, performing a motion vector prediction program based on the first bounding box and the motion vector to generate a first update result and a statistical range; performing a dynamic system prediction algorithm based on the multiple historical bounding boxes associated with the object in the multiple historical frames to generate a second update result.

[0052] Performing a dynamic system prediction algorithm based on the multiple historical bounding boxes associated with the object in the multiple historical frames to generate a second update result includes: when one of the second frame C and the third frame B lacks a bounding box associated with the object, performing interpolation based on the first bounding box associated with the object in the first frame D, the second bounding box associated with the object in the second frame C or the third bounding box associated with the object in the third frame B, and the fourth bounding box associated with the object in the fourth frame A to generate a bounding box; performing a dynamic system prediction algorithm based on the first bounding box, the second bounding box, the third bounding box, and the fourth bounding box to generate a second update result.

[0053] Step 97, determine whether the second update result is outside the statistical range. If the determination is no, step 98 is executed to update the current bounding box based on the second update result. If the determination is yes, step 99 is executed to update the current bounding box based on the first update result.

[0054] In one embodiment, the second update result is a new bounding box. The statistical range includes at least a first interval and a second interval, and each interval includes a first range corresponding to the X-axis and a second range corresponding to the Y-axis. In one embodiment, if the coordinates of the upper left vertex of the second update result are outside the first interval (the X coordinate of the vertex is outside the first range and the Y coordinate of the vertex is outside the second range), and the coordinates of the lower right vertex of the second update result are outside the second interval, then it is judged as yes. On the contrary, if the coordinates of the upper left vertex of the second update result are within the first interval (the X coordinate of the vertex is within the first range and the Y coordinate of the vertex is within the second range), and the coordinates of the lower right vertex of the second update result are within the second interval, then it is judged as no.

[0055] The motion vector prediction program can refer to Figure 5 . Figure 5 FIG. is a flowchart of a motion vector prediction program (e.g., for a bounding box) according to an embodiment of the present invention.

[0056] Step 901, obtain the motion vectors within the bounding box. In one embodiment, the bounding box of the previous frame D is rectangular. According to the coordinates of the two vertices of the upper left and lower right of the bounding box, obtain the motion vectors belonging to the range of the bounding box in the motion vector map as shown in Figure 3 the motion vector map.

[0057] Step 902, divide the bounding box into multiple sub-boxes, and select at least two of the sub-boxes. In one embodiment, the division is performed in the form of 2×2, and the upper left sub-box and the lower right sub-box are selected. The present invention does not limit the number of sub-box divisions. For example, it can also be divided into the form of 3×3. The sub-box selection preferably adopts the diagonal form. In other embodiments, the lower left sub-box and the upper right sub-box can be selected.

[0058] Step 903, denoise the vector field in the selected at least two sub-boxes. The vector field corresponds to the motion vectors of all pixels in the at least two sub-boxes. In one embodiment, denoising refers to retaining the motion vectors of the average value plus or minus one standard deviation. However, the present invention does not limit the calculation method of denoising by the above example.

[0059] Step 904: Based on the denoised vector field, calculate the average value as the first update result, and calculate the first quartile and the third quartile as the statistical range. The first update result is a new bounding box. In one embodiment, the coordinates of the upper left vertex of the first update result are the average value of the remaining motion vectors in the upper left sub-box, and the coordinates of the lower right vertex of the first update result are the average value of the remaining motion vectors in the lower right sub-box. The statistical range consists of the first quartile (Q1) and the third quartile (Q3) calculated from the remaining motion vectors. However, the present invention is not limited to the average value or quartiles exemplified above. Additionally, the step of calculating the statistical range can be omitted in the motion vector prediction procedure executed in step 62.

[0060] The feature compensation procedure proposed by the present invention is applicable not only to bounding boxes but also to key points. In one embodiment, when the object to be detected by the object detection model is a human body, its torso and limbs can be represented in a skeleton form, and the starting point and the ending point of each line segment forming the skeleton can be the key points. Please refer to Figure 6 , Figure 6 is a flowchart of a method for updating key points in an object detection model according to an embodiment of the present invention. The method is executed by an arithmetic device Figure 6 shown in steps 11 to 15.

[0061] Step 11: Input a video into the object detection model. The video includes a plurality of previous frames and a current frame.

[0062] Step 12: The object detection model detects an object in the current frame and outputs a plurality of key points and a plurality of confidence levels associated with the object. The plurality of key points and the plurality of confidence levels have a one-to-one correspondence.

[0063] Step 13: Among the plurality of key points, determine whether there is a candidate confidence level of a candidate point that is less than a threshold. If the determination is no, then execute step 14. If the determination is yes, then execute step 15. The candidate point is used to represent one of the plurality of key points. The candidate confidence level is used to represent one of the plurality of confidence levels and corresponds to the candidate point.

[0064] If the candidate confidence level of each candidate point is not less than the threshold, then execute step 14 and output the plurality of key points.

[0065] If there is a candidate confidence level of a candidate point that is less than the threshold, then execute step 15 and update the candidate point based on the plurality of previous frames, the current frame, and a motion vector. In one embodiment, the Figure 4 adaptation of the bounding box in the process to key points can be performed to obtain a feature compensation procedure applicable to key points. As for the motion vector prediction procedure, it is modified as Figure 7 shown.

[0066] Figure 7 FIG. Figure 7 is a flowchart of a motion vector prediction process (for key points) according to an embodiment of the present invention.

[0067] Step 151: Centering on the candidate point, extend a specified length in the horizontal and vertical directions respectively to generate a Moving Box (MB). Figure 8 FIG. Figure 8 is a schematic diagram of an example of the moving box. Figure 8 Candidate point P and multiple key points (not labeled) around it are presented. In one embodiment, the specified length is 16 pixels. The moving box is a rectangle formed by extending 16 pixels in each of the up, down, left, and right directions centered on candidate point P, as shown by the moving box MB in Figure 8 FIG. Figure 8 .

[0068] Step 152: Denoise the vector field of the moving box. In one embodiment, calculate an average value and a standard deviation based on the motion vectors within the moving box MB. The motion vectors are as shown by the arrows v1 - v8 in Figure 8 FIG. Figure 8 . The calculation methods of the average value mean(dx,dy) and the standard deviation std(dx,dy) are shown in the following Method 1 and Equation 2 respectively, where N = 8.

[0069]

[0070]

[0071] Denoising means retaining the motion vectors of the average value plus or minus one standard deviation. In other words, excluding the motion vectors outside one standard deviation, as shown in the following Method 3, where mv represents the set of denoised motion vectors. Taking the example of Figure 8 FIG. Figure 8 , vector v8 will be excluded.

[0072] mv ∈ {(dx,dy)|mean(dx,dy) ± std(dx,dy)} (Equation 3)

[0073] Step 153: Calculate the average value mean’(dx,dy) based on the denoised vector field, as shown in the following Method 4:

[0074]

[0075] Step 154: Update the candidate point based on the sum of the average value mean(dx,dy) and the key point old_keypoint(x,y), as shown in the following Method 5, where new_keypoint(x,y) represents the updated candidate point.

[0076] new_keypoint(x,y) = old_keypoint(x,y) + mean(dx,dy) (Equation 5)

[0077] Figure 4 The multiple embodiments shown are based on a bounding box feature compensation program, and those of ordinary skill in the technical field to which the present invention pertains may replace the current bounding box mentioned in the Figure 4 embodiment with Figure 6 and Figure 8 the candidate points described in and replace the first to fourth bounding boxes with the first to fourth key points respectively, so as to implement a key point-based feature compensation program. Therefore, the feature compensation program in the method of updating key points in the object detection model will not be repeated here.

[0078] The operation speed of the present invention is fast, and in terms of accuracy, it shows stable performance, maintaining a high level of accuracy and being less sensitive to changes in the number of frames.

[0079] In summary, the present invention proposes a method for updating a bounding box or key points in an object detection model, which can improve the architecture of the early departure technology in the application of neural networks. The present invention is applicable to application fields such as multi-object detection or key point tracking. The feature compensation program proposed by the present invention can improve the accuracy of instant inference, enabling the object detection model to have both fast and accurate effects. In an experiment, the inference speed of the object detection model applying the present invention increased from the traditional 14 frames per second to 55 frames per second, and the mean average precision (MAP) increased from the original 47.8 to 53.1.

Claims

1. A method for updating a bounding box in an object detection model, comprising executing, by a computing device: Inputting a video to an object detection model, wherein the video includes a plurality of previous frames and a current frame; Detecting an object in the current frame using the object detection model, and outputting a current bounding box and a confidence level associated with the object; When the confidence level is less than a threshold, the current bounding box is updated according to the previous frames, the current frame, and a motion vector associated with one of the previous frames and the current frame.

2. The method for updating a bounding box in an object detection model as claimed in claim 1, wherein These previous frames include: The first frame before the current frame and updating the current bounding box according to the previous frames, the current frame and the motion vector comprises: When the number of the previous frames is less than the specified number and the first frame has a first bounding box associated with the object, a motion vector prediction process is performed according to the first bounding box and the motion vector to update the current bounding box.

3. The method for updating a bounding box in an object detection model as claimed in claim 1, wherein These previous frames include: A first frame earlier than the current frame and a plurality of historical frames earlier than the first frame; Updating the current bounding box according to the previous frames, the current frame and the motion vector comprises: When the number of the previous frames is not less than the specified number and the first frame lacks a first bounding box associated with the object, a dynamic system prediction algorithm is executed according to a plurality of historical bounding boxes associated with the object in the historical frames to update the current bounding box.

4. The method for updating a bounding box in an object detection model as claimed in claim 3, wherein These historical frames include: a second frame earlier than the first frame, a third frame earlier than the second frame, and a fourth frame earlier than the third frame, wherein the second frame, the third frame, and the fourth frame respectively have a second bounding box, a third bounding box, and the fourth bounding box associated with the object; Executing the dynamic system prediction algorithm according to a plurality of historical bounding boxes associated with the object in the historical frames to update the current bounding box includes: Executing the dynamic system prediction algorithm according to the second bounding box, the third bounding box and the fourth bounding box to generate the first bounding box; as well as The dynamic system prediction algorithm is executed according to the first bounding box, the second bounding box, the third bounding box and the fourth bounding box to update the current bounding box.

5. The method for updating a bounding box in an object detection model as claimed in claim 1, wherein These previous frames include: A first frame earlier than the current frame and a plurality of historical frames earlier than the first frame; Updating the current bounding box according to the previous frames, the current frame and the motion vector comprises: When the number of the previous frames is not less than a specified number and the first frame has a first bounding box associated with the object, performing a motion vector prediction procedure according to the first bounding box and the motion vector to generate a first update result and a statistical range; executing a dynamic system prediction algorithm according to a plurality of historical bounding boxes associated with the object in the historical frames to generate a second update result; When the second update result is outside the statistical range, updating the current bounding box with the first update result; as well as When the second update result is within the statistical range, the current bounding box is updated with the second update result.

6. The method for updating a bounding box in an object detection model as claimed in claim 5, wherein These historical frames include: a second frame earlier than the first frame, a third frame earlier than the second frame, and a fourth frame earlier than the third frame; Executing the dynamic system prediction algorithm according to a plurality of historical bounding boxes associated with the object in the historical frames to generate the second update result comprises: When one of the second frame and the third frame lacks a bounding box associated with the object, Performing interpolation based on a first bounding box associated with the object in the first frame, a second bounding box associated with the object in the second frame or a third bounding box associated with the object in the third frame, and a fourth bounding box associated with the object in the fourth frame to generate the bounding box; The dynamic system prediction algorithm is executed according to the first bounding box, the second bounding box, the third bounding box and the fourth bounding box to generate the second update result.

7. The method for updating a bounding box in an object detection model as claimed in claim 5, wherein These historical frames include: a second frame earlier than the first frame, a third frame earlier than the second frame, and a fourth frame earlier than the third frame; Executing the dynamic system prediction algorithm according to a plurality of historical bounding boxes associated with the object in the historical frames to generate the second update result comprises: When the second frame has a second bounding box associated with the object and the third frame has a third bounding box associated with the object, the dynamic system prediction algorithm is executed according to the second bounding box, the third bounding box and a fourth bounding box associated with the object in the fourth frame to generate the second update result.

8. The method for updating a bounding box in an object detection model as claimed in claim 5, wherein Executing a motion vector prediction procedure according to the first bounding box and the motion vector to generate a first update result and a statistical range includes: Obtaining the motion vector within the first bounding box; Splitting the first bounding box into a plurality of sub-boxes, and selecting at least two of the sub-boxes; Denoising the vector field of the at least two sub-frames, wherein the vector field corresponds to the motion vector of the at least two sub-frames; as well as According to the denoised vector field, an average value is calculated as a first update result, and a first quartile and a third quartile are calculated as the statistical range.

9. A method for updating key points in an object detection model, comprising executing, with a computing device: Inputting a video to an object detection model, wherein the video includes a plurality of previous frames and a current frame; Detecting an object in the current frame using the object detection model, and outputting a plurality of key points and a plurality of confidences associated with the object, wherein a candidate point among the key points corresponds to a candidate confidence among the confidences; When the candidate confidence is less than a threshold, the candidate point is updated according to the previous frames, the current frame and a motion vector, wherein the motion vector is associated with one of the previous frames and the current frame.

10. The method for updating key points in an object detection model as claimed in claim 9, wherein updating the candidate point according to the previous frames, the current frame and the motion vector comprises: Taking the candidate point as the center, extending in the horizontal direction and the vertical direction by a specified length respectively to generate a moving frame; Denoising a vector field of the moving frame, wherein the vector field corresponds to the moving vector in the moving frame; Calculating an average value according to the denoised vector field; and The candidate point is updated according to the sum of the average value and the key point.

11. The method for updating key points in an object detection model as claimed in claim 9, wherein These previous frames include: The first frame earlier than the current frame and updating the candidate point according to the previous frames, the current frame and the motion vector comprises: When the number of the previous frames is less than the specified number and the first frame has a first key point associated with the object, a motion vector prediction procedure is performed according to the first key point and the motion vector to update the candidate point.

12. The method for updating key points in an object detection model as claimed in claim 9, wherein These previous frames include: A first frame earlier than the current frame and a plurality of historical frames earlier than the first frame; Updating the candidate point according to the previous frames, the current frame and the motion vector comprises: When the number of the previous frames is not less than a specified number and the first frame lacks a first key point associated with the object, a dynamic system prediction algorithm is executed according to a plurality of historical key points associated with the object in the historical frames to update the candidate point.

13. The method for updating key points in an object detection model as claimed in claim 12, wherein These historical frames include: a second frame earlier than the first frame, a third frame earlier than the second frame, and a fourth frame earlier than the third frame, wherein the second frame, the third frame, and the fourth frame respectively have a second key point, a third key point, and the fourth key point associated with the object; Executing the dynamic system prediction algorithm according to the plurality of historical key points associated with the object in the historical frames to update the candidate point comprises: Executing the dynamic system prediction algorithm according to the second key point, the third key point and the fourth key point to generate the first key point; as well as The dynamic system prediction algorithm is executed according to the first key point, the second key point, the third key point and the fourth key point to update the candidate point.

14. The method for updating key points in an object detection model as claimed in claim 9, wherein These previous frames include: A first frame earlier than the current frame and a plurality of historical frames earlier than the first frame; Updating the candidate point according to the previous frames, the current frame and the motion vector comprises: When the number of the previous frames is not less than a specified number and the first frame has a first key point associated with the object, performing a motion vector prediction procedure according to the first key point and the motion vector to generate a first update result and a statistical range; Executing a dynamic system prediction algorithm according to a plurality of historical key points associated with the object in the historical frames to generate a second update result; When the second update result is outside the statistical range, updating the candidate point with the first update result; and When the second update result is within the statistical range, the candidate point is updated with the second update result.

15. The method for updating key points in an object detection model as claimed in claim 14, wherein These historical frames include: a second frame earlier than the first frame, a third frame earlier than the second frame, and a fourth frame earlier than the third frame; Executing the dynamic system prediction algorithm according to a plurality of historical key points associated with the object in the historical frames to generate the second update result includes: When one of the second frame and the third frame lacks a key point associated with the object, Performing interpolation to generate the key point according to a first key point associated with the object in the first frame, a second key point associated with the object in the second frame or a third key point associated with the object in the third frame, and a fourth key point associated with the object in the fourth frame; The dynamic system prediction algorithm is executed according to the first key point, the second key point, the third key point and the fourth key point to generate the second update result.

16. The method for updating key points in an object detection model as claimed in claim 14, wherein These historical frames include: a second frame earlier than the first frame, a third frame earlier than the second frame, and a fourth frame earlier than the third frame; Executing the dynamic system prediction algorithm according to a plurality of historical key points associated with the object in the historical frames to generate the second update result includes: When the second frame has a second key point associated with the object and the third frame has a third key point associated with the object, the dynamic system prediction algorithm is executed according to the second key point, the third key point and a fourth key point associated with the object in the fourth frame to generate the second update result.