Video target status detection method and device, electronic device and storage medium

By using multiple windows of different lengths to determine the forward confidence threshold in the video, the problem of insufficient response sensitivity in the video target state detection is solved, and more accurate and timely status detection results are achieved.

CN114332724BActive Publication Date: 2025-08-22NANJING HORIZON ROBOTICS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111670781.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-08-22
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

In the prior art, in the detection of target behavior status in video or images, the system output response sensitivity is insufficient, and false alarms and missed alarms are serious, especially in complex scenarios, which is difficult to achieve accurate status detection.

Method used

Using a pre-trained classification model, the forward confidence of the target in each frame of the video to be detected is determined, and the fusion forward confidence of the multi-frame images is determined through multiple windows of different lengths. Combined with the fusion forward confidence threshold, the state detection result of the target in the video is determined.

Benefits of technology

It improves the system's response sensitivity and accuracy of state detection, reduces false alarms and missed alarms, and ensures accurate output in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332724B_ABST
    Figure CN114332724B_ABST
Patent Text Reader

Abstract

The disclosed embodiments disclose a method and apparatus, electronic device, and storage medium for detecting the state of a video target. The method comprises: determining, based on a pre-trained classification model, a positive confidence score regarding the state of a target in each frame of the video to be detected; determining, based on multiple windows of different lengths, a fused positive confidence score regarding the state of the target for multiple frames of the video to be detected corresponding to each window; and determining a state detection result for the target in the video to be detected based on the fused positive confidence score for each window and a corresponding fused positive confidence threshold. The disclosed embodiments can improve the response sensitivity of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to image processing technology, and in particular to a method and device for detecting the state of a video target, an electronic device, and a storage medium. Background Art

[0002] In applications of behavioral state detection or environmental state detection of targets in videos or images, it is often necessary to use a deep learning model to determine the confidence level of the image regarding the behavioral state or environmental state of the target. Then, during the model post-processing process, the system can output the state detection results regarding the behavioral state or environmental state of the target based on the confidence level regarding the target state.

[0003] Regarding the model post-processing process, how to improve the system output response sensitivity is a technical issue worthy of attention. Summary of the Invention

[0004] In order to solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a method and apparatus for detecting the state of a video target, an electronic device, and a storage medium.

[0005] According to one aspect of an embodiment of the present disclosure, a method for detecting the state of a video target is provided, comprising: determining, based on a pre-trained classification model, a positive confidence in the state of a target in each frame image of a video to be detected; determining, based on a plurality of windows of different lengths, a fused positive confidence in the state of the target of multiple frames of images in the video to be detected corresponding to each window; determining, based on the fused positive confidence of each window and a corresponding fused positive confidence threshold, a state detection result of the target in the video to be detected; wherein, the fused positive confidence threshold is determined based on a given accuracy rate.

[0006] According to another aspect of an embodiment of the present disclosure, a state detection device for a video target is provided, comprising: a forward confidence determination module for determining the forward confidence of the state of the target in each frame image of the video to be detected based on a pre-trained classification model; a fused forward confidence determination module for determining, based on multiple windows of different lengths, the fused forward confidence of the state of the target in multiple frames of images in the video to be detected corresponding to each window; a detection result determination module for determining the state detection result of the target in the video to be detected based on the fused forward confidence of each window and the corresponding fused forward confidence threshold; wherein the fused positive confidence threshold is determined based on a given accuracy rate.

[0007] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and the computer program is used to execute the state detection method of the video target described in any of the above embodiments of the present disclosure.

[0008] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; and the processor for reading the executable instructions from the memory and executing the instructions to implement the video target status detection method described in any of the above embodiments of the present disclosure.

[0009] Based on the state detection method and device of the video target provided by the above-mentioned embodiment of the present disclosure, the fused positive confidence of the state of the target of the multiple frames of images corresponding to each window in a plurality of windows of different lengths can be determined respectively, and then based on the fused positive confidence of each window and the corresponding fused positive confidence threshold, the state detection result of the target in the video to be detected can be determined, and a state output strategy based on multiple time windows can be implemented, which has better flexibility than using a fixed time window to detect the state of the target in the video to be detected. In addition, the fused positive confidence threshold of each window is determined based on a given accuracy rate, and the detection process of the short window is short in time, which can help to output the state detection result of the target in the video to be detected in a timely manner while ensuring the accuracy of the state detection result, and can improve the response sensitivity of the system.

[0010] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other purposes, features, and advantages of the present disclosure will become more apparent through a more detailed description of the embodiments of the present disclosure in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and are not intended to limit the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.

[0012] Figure 1 is an exemplary framework diagram of a detection system to which the present disclosure is applicable;

[0013] Figure 2 is a flow chart of a method for detecting a state of a video target provided by an exemplary embodiment of the present disclosure;

[0014] Figure 3a-3b is a schematic diagram of multiple windows of different lengths disclosed herein;

[0015] Figure 4 is the accuracy curve of the present disclosure;

[0016] Figure 5 is a flow chart of a method for detecting a state of a video target provided by another exemplary embodiment of the present disclosure;

[0017] Figure 6 is a flowchart of a method for detecting a state of a video target provided by another exemplary embodiment of the present disclosure;

[0018] Figure 7 is a structural diagram of a video target state detection device provided by an exemplary embodiment of the present disclosure;

[0019] Figure 8 is a structural diagram of a video target state detection device provided by another exemplary embodiment of the present disclosure;

[0020] Figure 9 is a structural diagram of a video target state detection device provided by another exemplary embodiment of the present disclosure;

[0021] Figure 10 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] Below, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.

[0023] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.

[0024] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meanings, nor do they indicate a necessary logical order between them.

[0025] It should also be understood that in the embodiments of the present disclosure, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.

[0026] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.

[0027] In addition, the term "and / or" in this disclosure is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship.

[0028] It should also be understood that the description of the various embodiments in this disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced with each other. For the sake of brevity, they will not be described one by one.

[0029] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0030] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0031] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0032] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0033] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate in conjunction with numerous other general-purpose or specialized computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers, and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, among others.

[0034] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media, including storage devices.

[0035] Application Overview

[0036] In the process of realizing the present disclosure, the inventors found that in the process of detecting the behavior state or environmental state of a target in a video, the algorithm model based on deep learning usually uses a single frame image as input, and outputs the confidence level of the target state through the algorithm model. Due to the complexity and variability of the scene and the limitations of deep learning capabilities, the output of the algorithm model cannot achieve a high accuracy and recall rate, and false alarms and missed reports are inevitable. For example, in application scenarios such as detecting whether the driver is making a phone call, detecting whether the driver of a dangerous goods transport vehicle is smoking, and detecting whether there are obstacles on the road where the vehicle is traveling, it is impossible to obtain accurate state detection results based only on the confidence level of the target state based on a single frame image. Therefore, by selecting multiple consecutive frames of images on the timeline to run the algorithm model, multiple confidence levels of the target state are obtained, and then the fusion confidence level of the target state is determined through a fusion strategy, and then the target state detection result is determined based on the fusion confidence level, which can appropriately reduce false alarms and missed reports.

[0037] Existing methods typically take a fixed time window (such as 1 second), determine the confidence of multiple frames of images (such as 30 frames of images) within the time window, and use a fusion strategy to determine the fusion confidence of the multiple frames of images within the time window. The detection results of the behavioral state or environmental state are determined based on the fusion confidence. The fusion strategy can be to obtain the average confidence of the multiple frames of images within the time window, or it can be to take the weighted average confidence of the positive confidence of the multiple frames of images within the window. Based on this method, the longer the time window used, the more image frames are obtained, and the higher the judgment accuracy. However, the longer the time window, the later the reporting time of the behavioral state or environmental state, and the lower the sensitivity of the system.

[0038] Exemplary Systems

[0039] Figure 1 1 is a schematic diagram of an exemplary framework of a detection system applicable to the present disclosure. The detection system 100 may include a camera device 101 and a server 102. The camera device 101 communicates with the server 102 and transmits the captured video to be detected to the server 102 for storage and other processing. Taking a vehicle detection system as an example, the camera device 101 may be set at a preset position inside the vehicle, the preset position being determined based on the ability to capture the video to be detected of the driver in the driving seat. Alternatively, the camera device 101 may be set at a preset position outside the vehicle, the preset position being determined based on the ability to capture the video to be detected of the road section where the vehicle is traveling.

[0040] In the embodiment of the present disclosure, after capturing the video to be detected, the camera device 101 can first determine the positive confidence of the state of the target in each frame image of the video to be detected based on a pre-trained classification model, and then determine the fused positive confidence of the state of the target in multiple frames of images in the video to be detected corresponding to each window based on multiple windows of different lengths, and then determine the state detection result of the target in the video to be detected based on the fused positive confidence of each window and the corresponding fused positive confidence threshold.

[0041] In practical applications, the state detection of the video target to be detected can be achieved by adding a state detection unit between the image sensor and the image signal processor (Image Signal Processing, abbreviated as: ISP). The state detection unit can be implemented using a GPU (graphics processing unit) or an AI (Artificial Intelligence) chip. In addition, the state detection of the video target to be detected can also be implemented by the ISP. The specific settings can be based on actual needs and are not limited in the embodiments of the present disclosure. In some applications, the state detection of the video target to be detected can also be implemented by other electronic devices other than the camera device (such as terminal devices, computer systems, servers, etc.). By communicating with the camera device, the video to be detected shot by the camera device is transmitted to the other electronic devices, and the other electronic devices perform state detection of the video target to be detected and return it to the camera device.

[0042] Exemplary Methods

[0043] Figure 2 FIG. 1 is a flow chart of a method for detecting the state of a video target provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices, such as Figure 2 As shown, the following steps are included:

[0044] Step 201 : Based on a pre-trained classification model, determine the positive confidence level of the state of the target in each frame of the video to be detected.

[0045] In the embodiment of the present disclosure, a pre-trained classification model can be obtained in the following manner: constructing a classification model to be trained; obtaining multiple image samples from a training sample set; providing the multiple image samples as input to the classification model to be trained, performing state detection on each input image sample through the classification model to be trained, and obtaining the predicted positive confidence of the state of the target in each image sample based on the output of the classification model to be trained; adjusting the model parameters of the classification model to be trained based on the difference between the predicted positive confidence of the state of the target in each image sample and the positive confidence annotation information of each image sample, and iterating the above process until the preset training completion conditions are met to obtain the classification model.

[0046] In the embodiment of the present disclosure, the classification model to be trained may be a binary convolutional neural network model, and the output of the classification model to be trained may include positive confidence and negative confidence.

[0047] In the disclosed embodiments, the state of the target may be the target's behavioral state or environmental state. A positive confidence score for the target's state in an image may indicate the confidence that the target exists in that state in the image, while a negative confidence score for the target's state in an image may indicate the confidence that the target does not exist in that state in the image.

[0048] In an optional example, the target of the embodiment of the present disclosure may be a driver of a public transportation such as a taxi, an express train, a bus, or a ferry, and the corresponding status may be that the driver is not wearing a seat belt.

[0049] In another optional example, the target of the embodiment of the present disclosure may be a driver of a dangerous goods transport vehicle, and the corresponding state may be that the driver is smoking.

[0050] In yet another optional example, the target of the embodiment of the present disclosure may be workers at a construction site, and the corresponding status may be that the workers are not wearing safety helmets.

[0051] In yet another optional example, the target of the embodiment of the present disclosure may be the driving section of the vehicle, and the corresponding state may be that there is an obstacle on the driving section.

[0052] Step 202 : Based on a plurality of windows of different lengths, respectively determine the fused positive confidence of the state of the target in the plurality of frames of the video to be detected corresponding to each window.

[0053] In the embodiment of the present disclosure, the length of the window can be identified by the number of image frames or the detection time. As an example, assuming that the window T 10The length of the window is 10 frames, and the window can correspond to 10 frames of images in the video to be detected. As another example, assuming that the length of window T1 is 1 second, if the detection frame rate is 30 frames per second (i.e., 30fps), the window can correspond to 30 frames of images in the video to be detected. It should be noted that those skilled in the art can set the length and number of multiple windows according to actual needs, and the embodiment of the present disclosure does not limit the length and number of multiple windows.

[0054] In an optional example, a plurality of windows of different lengths may be determined based on a detection frame rate of a pre-trained classification model.

[0055] As an example, if the detection frame rate of the pre-trained classification model is set to 30 frames per second (ie, 30fps), multiple windows of different lengths can be set to T 30 、T 60 , you can also set multiple windows of different lengths, namely T9, T 30 、T 60 Etc. Based on the detection frame rate of 30fps, the time required to complete the detection of multiple frames of images corresponding to windows of different lengths is also different. For window T9, it takes 0.3s to complete the detection of multiple frames of images corresponding to the window; for window T 30 , the time required to complete the detection of multiple frames of images corresponding to the window is 1s; for window T 60 , the time required to complete the detection of multiple frames of images corresponding to the window is 2 seconds.

[0056] In the embodiment of the present disclosure, the video to be detected input into the detection system may be an image sequence captured in real time by a camera device, or may be an image sequence captured in advance by a camera device.

[0057] In an optional example, each window may correspond to a plurality of consecutive image frames in the video to be detected, and the last image frame in the plurality of consecutive image frames corresponding to each window is the last image frame in the video to be detected.

[0058] As an example, Figure 3a The video to be detected input into the detection system can be an image sequence with 70 frames, and the windows of different lengths are T 20 、T 25 、T 55 、T 60 Among them, T 20 The corresponding image frames are 20 frames, T 25 The corresponding image frames are 25 frames, T 55 The corresponding image frames are 55 frames, T 60The corresponding number of image frames is 60 frames, and the last frame image in the multi-frame image corresponding to each window is the last frame image in the image sequence currently acquired by the detection system, that is, the 70th frame image.

[0059] In an optional example, when a new image to be detected is input into the detection system, the number of image frames contained in the video to be detected will increase accordingly. By sliding multiple windows of different lengths, the last frame of the multiple frames corresponding to each window can be the last frame of the image currently acquired by the detection system, so as to realize the state detection of the target in the image sequence currently acquired by the detection system.

[0060] As an example, Figure 3b As described above, when there is a new image to be detected (such as Figure 3b When the 71st frame image shown in FIG is input into the detection system, the video to be detected becomes an image sequence with 71 frames, and multiple windows T of different lengths can be slid 20 、T 25 、T 55 and T 60 , so that the last frame image in the multiple frames corresponding to each window is the last frame image currently acquired by the detection system, that is, the 71st frame image, so as to realize the state detection of the target in the image sequence currently acquired by the detection system.

[0061] In one optional example, each window can correspond to a different fusion positive confidence threshold. When different windows have the same accuracy requirement, the size of the fusion positive confidence threshold corresponding to the window is negatively correlated with the length of the window. That is, the smaller the window length, the larger the corresponding fusion positive confidence threshold, and the larger the window length, the smaller the corresponding fusion positive confidence threshold.

[0062] As an example, Figure 3a As shown, the window T can be set separately 20 The corresponding fusion positive confidence threshold is 0.88 and the window T 25 The corresponding fusion positive confidence threshold is 0.85 and the window T 55 The corresponding fusion positive confidence threshold is 0.72 and the window T 60 The corresponding fusion positive confidence threshold is 0.65.

[0063] With window T 20 For example, the window T can be determined by the following operations 20 Corresponding fusion positive confidence threshold: select a test set (e.g. 1000 sample videos); set the target accuracy (e.g. 99%) and the initial fusion positive confidence threshold (e.g. 0.6); based on the window T 20, respectively determine the fusion positive confidence of each sample video in the test set, and determine the window T based on the fusion positive confidence of each sample video and the initial fusion positive confidence threshold 20 The test accuracy of the window T is continuously adjusted to increase the value of the fusion positive confidence threshold. 20 The test accuracy reaches the target accuracy, and the fusion positive confidence threshold corresponding to the target accuracy is the window T 20 The corresponding fusion positive confidence threshold that meets the accuracy requirements.

[0064] Based on this, different fusion positive confidence thresholds are corresponding to windows of different lengths, so that multiple windows of different lengths can meet the same accuracy requirements, thereby ensuring the accuracy of the state detection results.

[0065] In an optional example, when determining the fused positive confidence of the state of the target in the multiple frames of images to be detected corresponding to each window, the fused positive confidence of the state of the target in the multiple frames of images corresponding to each window can be determined in a preset manner based on the positive confidence of the state of the target in each frame of the multiple frames corresponding to each window.

[0066] In an optional example, the fused positive confidence may be an average of the positive confidences of multiple frames of images corresponding to the window.

[0067] In one optional example, the fused positive confidence score can be a weighted average of the positive confidence scores of multiple frames corresponding to the window. In practical applications, different weights can be set for each frame in the multiple frames, or the multiple frames can be divided into multiple image groups, and then different weights can be set for each image group.

[0068] In an optional example, for each of the multiple image groups, a different weight may be set for each image group according to the number of image frames contained in the image group and / or the time sequence of the image group.

[0069] Step 203 : Determine the state detection result of the target in the video to be detected based on the fused positive confidence of each window and the corresponding fused positive confidence threshold.

[0070] In an optional example, the state detection result of the target in the to-be-detected video may be that the target in the to-be-detected video exists in the above-mentioned state or that the target in the to-be-detected video does not exist in the above-mentioned state.

[0071] In an embodiment of the present disclosure, multiple windows of different lengths may correspond to different fusion forward confidence thresholds, respectively. Based on the fusion forward confidence of each window and the corresponding fusion forward confidence threshold, the state detection result of the video target in each window may be determined.

[0072] In an optional example, the state detection result of the video object in each window may be determined based on the relationship between the fused positive confidence of each window and the corresponding fused positive confidence threshold.

[0073] In an optional example, if the fused positive confidence is greater than the corresponding fused positive confidence threshold, it can be determined that the detection result of the video target is that the video target exists in the above-mentioned state; if the fused positive confidence is not greater than the corresponding fused positive confidence threshold, it can be determined that the detection result of the video target is that the video target does not exist in the above-mentioned state.

[0074] In an optional example, if the state detection result of the video target in a window of multiple windows of different lengths is that the video target exists in the above state, it can be determined that the state detection result of the target in the video to be detected is that the target in the video to be detected exists in the above state.

[0075] In an optional example, after determining that the state detection result of the target in the video to be detected is that the target in the video to be detected exists in the above-mentioned state, target state prompt information may be output to prompt that the target in the video exists in the above-mentioned state.

[0076] In an optional example, target status prompt information can also be output according to a preset period to periodically prompt that the target in the video exists in the above-mentioned state. For example, according to a preset period, the vehicle driver is reminded of dangerous driving behavior until the behavior ends.

[0077] The disclosed embodiment can respectively determine the fused positive confidence of the state of the target for each of the multiple frames of images corresponding to each window in a plurality of windows of different lengths, and then determine the state detection result of the target in the video to be detected based on the fused positive confidence of each window and the corresponding fused positive confidence threshold. This can implement a state output strategy based on multiple time windows, which is more flexible than using a fixed time window to detect the state of the target in the video to be detected. In addition, the fused positive confidence threshold of each window is determined based on a given accuracy rate, and the detection process of the short window is short in time. It can help to output the state detection result of the target in the video in a timely manner while ensuring the accuracy of the state detection result, and can improve the response sensitivity of the system.

[0078] In practical applications, the state detection method of the video target of the embodiment of the present disclosure is tested using a preset test sample set, and the following results can be obtained: Figure 4 The accuracy curve is shown.

[0079] exist Figure 4In the figure, the left figure is the precision-recall curve, which describes the relationship between the precision and recall of multiple windows of different lengths on the test sample set; the right figure is the precision-confidence curve, which describes the relationship between the fused positive confidence threshold and the precision corresponding to multiple windows of different lengths.

[0080] As can be seen in the figure on the right, to achieve the same accuracy, the smaller the window, the larger the fusion positive confidence threshold should be. Because smaller windows capture fewer image frames, only a higher fusion positive confidence threshold can ensure accuracy and reduce false positives of behavioral states.

[0081] As can be seen from the left figure, under the same accuracy conditions, a larger window size results in a higher recall rate, meaning fewer missed behavioral status reports. Therefore, if only a fixed window is used for behavioral status detection, a longer window should be selected to reduce false positives and improve recall. This significantly reduces system sensitivity and results in delayed behavioral status reporting.

[0082] If multiple windows of different lengths are used for behavioral state detection, the sensitivity and recall rate can be maximized while ensuring the accuracy. Figure 4 As shown, when the accuracy is 95%, the window T 20 The fusion positive confidence threshold is 0.88, and the corresponding recall rate is 38.5%, that is, 38.5% of the behavior states can be reported within 20 frames; the window T 25 The fusion positive confidence threshold is 0.85, and the recall rate is 53%. Similarly, under the condition of ensuring 95% accuracy, the proportion of behavior states detected within 20 frames is 38.5%, the proportion of behavior states detected within 25 frames is 53%, the proportion of behavior states detected within 55 frames is 72%, and the proportion of behavior states detected within 60 frames is 97%. Therefore, the behavior state output strategy based on multiple time windows can improve the system's response sensitivity while ensuring the accuracy and recall rate of behavior state output.

[0083] In an alternative example, Figure 5 FIG. 1 is a flow chart of a method for detecting a state of a video target provided by another exemplary embodiment of the present disclosure. Figure 5 As shown in the above Figure 2 Based on the embodiment shown, step 203 may include the following steps:

[0084] Step 203 - 1 : determining the detection order of multiple fused positive confidences corresponding to multiple windows of different lengths in ascending order of window lengths.

[0085] Step 203 - 2 : Based on the detection order of the multiple fused positive confidences, determine the magnitude relationship between the multiple fused positive confidences and the corresponding fused positive confidence thresholds in sequence.

[0086] Step 203 - 3 : Determine the state detection result of the target in the video to be detected based on the size relationship.

[0087] In an optional example, step 203-2 can be based on the detection order of multiple fused positive confidences, and when the size relationship between multiple fused positive confidences and the corresponding fused positive confidence thresholds is judged in turn, the first fused positive confidence (which can be called the first fused positive confidence) that is greater than the corresponding fused positive confidence threshold is determined from the multiple fused positive confidences.

[0088] In an optional example, when the first fused positive confidence is determined from multiple fused positive confidences in the above step 203-2, the detection result of the target in the video to be detected can be determined as the target in the video to be detected exists in the above-mentioned state, and status prompt information can be output to prompt that the target in the video to be detected exists in the above-mentioned state.

[0089] In an optional example, after determining that the target in the video to be detected exists in the above-mentioned state based on the first fused positive confidence and outputting the status prompt information, the state detection of the video to be detected can be ended, that is, the size relationship between other fused positive confidences except the first fused positive confidence and the corresponding fused positive confidence threshold will no longer be judged.

[0090] In an optional example, if the fused positive confidence of each window in multiple windows of different lengths is not greater than the corresponding fused positive confidence threshold, it can be determined that the detection result is that the target in the video to be detected does not exist in the above-mentioned state, and no status prompt information is output.

[0091] The embodiment of the present disclosure judges the size relationship between the multiple fused positive confidences and the corresponding fused positive confidence thresholds in turn based on the detection order of multiple fused positive confidences. If the current fused positive confidence is greater than the corresponding fused positive confidence threshold, it can be determined that the detection result is that the target in the video to be detected exists in the above-mentioned state, and then status prompt information can be output to prompt that the target in the video to be detected exists in the above-mentioned state. If a window with a smaller length can detect that the target in the video to be detected exists in the above-mentioned state, the status prompt information can be output in time, thereby effectively improving the response sensitivity of the system.

[0092] In an alternative example, Figure 6 FIG. 1 is a flow chart of a method for detecting a state of a video target provided by another exemplary embodiment of the present disclosure. Figure 6 As shown in the above Figure 2Based on the embodiment shown, the following steps may also be included:

[0093] Step 204 : Determine the negative confidence level of the state of the target in each frame of the video to be detected based on the pre-trained classification model.

[0094] Step 205 : Determine, from a plurality of windows of different lengths, a first window corresponding to a first fused positive confidence greater than a corresponding fused positive confidence threshold.

[0095] Step 206 : Based on the preset order of the multiple windows of different lengths, determine the fused negative confidence corresponding to each of the multiple windows before the first window.

[0096] Step 207 : Based on the magnitude relationship between the multiple fused negative confidences and the corresponding suppression thresholds, suppress the first fused positive confidence that is greater than the corresponding fused positive confidence threshold.

[0097] In an optional example, the preset order may be an order of window lengths from small to large, and the lengths of multiple windows before the first window are all smaller than the length of the first window.

[0098] In an optional example, the fused negative confidence may be an average of the negative confidences of multiple frames of images within the window.

[0099] In one optional example, the fused negative confidence score can be a weighted average of the negative confidence scores of multiple image frames corresponding to the window. In practical applications, different weights can be set for each image frame in the multiple image frames, or the multiple image frames can be divided into multiple image groups, and then different weights can be set for each image group.

[0100] In one optional example, for each of the multiple image groups, a different weight can be set for each image group based on the number of image frames contained in the image group and / or the time sequence of the image groups. In another optional example, multiple windows of different lengths can correspond to different suppression thresholds. Those skilled in the art can set the suppression thresholds corresponding to the multiple windows of different lengths as needed. The presently disclosed embodiments do not specifically limit the size of the suppression threshold.

[0101] In the embodiment of the present disclosure, suppressing the first fused positive confidence may be to reduce the value of the first fused positive confidence.

[0102] In an optional example, reducing the value of the first fused positive confidence may be setting the value of the first fused positive confidence to a difference between 1 and the first fused positive confidence.

[0103] In an optional example, reducing the value of the first fused positive confidence may be setting the value of the first fused positive confidence to any value that is smaller than a fused positive confidence threshold corresponding to the first window.

[0104] After determining the first fused positive confidence, the embodiment of the present disclosure suppresses the first fused positive confidence based on the fused negative confidence corresponding to each of the multiple windows before the first window. That is, the detection result of the first window is corrected according to the detection results of the multiple windows before the first window, thereby improving the accuracy of the detection results of the behavioral status of the target in the video, reducing the false alarm rate, and reducing the disturbance to the user.

[0105] In an optional example, in step 207, based on the size relationship between multiple fused negative confidences and corresponding suppression thresholds, when suppressing the first fused positive confidence that is greater than the corresponding fused positive confidence threshold, the detection order of the multiple fused negative confidences corresponding to each of the multiple windows before the first window can be determined according to the preset order of the window lengths. Then, based on the detection order of the multiple fused negative confidences, when the first fused negative confidence that is less than the corresponding suppression threshold is determined from the multiple fused negative confidences, the first fused positive confidence that is less than the corresponding suppression threshold can be used to replace the first fused positive confidence that is greater than the corresponding fused positive confidence threshold. Here, the preset order of the window lengths can be the order of the window lengths from small to large.

[0106] In an optional example, the above-mentioned fusion positive confidence threshold can be determined in the following manner: for each sample video in the sample video set, first determine the sample positive confidence of the state of the target in each frame sample image of the sample video based on the pre-trained classification model, and determine the sample fusion positive confidence of multiple frames of sample images in the sample video corresponding to the target window based on the sample positive confidence, and determine the state detection result of the target in the sample video based on the sample fusion positive confidence and the fusion positive confidence threshold of the target window, and then determine the accuracy of the target window based on the state detection result of the target in each sample video in the sample video set, and determine the functional relationship between the fusion positive confidence threshold of the target window and the accuracy, and then determine the fusion positive confidence threshold of the target window at a given accuracy based on the functional relationship as the above-mentioned fusion positive confidence threshold. It should be noted that the target window here can be any window among the above-mentioned multiple windows of different lengths, and the given accuracy can be set according to actual needs, which is not limited in the embodiments of the present disclosure.

[0107] The disclosed embodiment can determine the functional relationship between the sample fusion positive confidence threshold and the accuracy of each window based on a sample video set and a pre-trained classification model, and then, based on the functional relationship of each window, determine the sample fusion positive confidence threshold of each window at a given accuracy as the above-mentioned fusion positive confidence threshold, which helps to ensure the accuracy of the state detection results of the target in the video to be detected determined by the size relationship between the fusion positive confidence and the corresponding fusion positive confidence threshold, reduce the false alarm rate, and reduce the disturbance to the user.

[0108] In an optional example, when determining the above-mentioned fusion positive confidence threshold, for multiple windows of different lengths, when the length difference between any two adjacent windows is greater than one frame of image, the fusion positive confidence threshold of the window between the two adjacent windows can be determined using an interpolation formula based on the fusion positive confidence threshold corresponding to the two adjacent windows.

[0109] In an optional example, an interpolation formula can be determined based on the ratio of the difference between the lengths of two adjacent windows and the difference between the fusion forward confidence thresholds of the two adjacent windows, and the interpolation formula can be used to determine the fusion forward confidence threshold of the window between the two adjacent windows.

[0110] The embodiment of the present disclosure can use an interpolation formula to determine the fusion positive confidence threshold of the window between two adjacent windows. Compared with the method of determining the fusion positive confidence threshold of the window based on a sample video set and a pre-trained classification model, this method has a simple implementation process and a small amount of calculation, which helps to reduce the amount of calculation and save computing resources.

[0111] In an optional example, the above-mentioned suppression threshold is determined in the following manner: for each sample video in the sample video set, first determine the sample negative confidence of each frame sample image of the sample video based on the pre-trained classification model, and based on the sample negative confidence, determine the sample fusion negative confidence of multiple frames of sample images in the sample video corresponding to each target window, and determine the state detection result of the target in the sample video based on the sample negative confidence and the suppression threshold of the target window, and then determine the accuracy of the target window based on the state detection result of the target in each sample video in the sample video set, and determine the functional relationship between the suppression threshold of the target window and the accuracy, and then determine the suppression threshold of the target window at a given accuracy based on the functional relationship as the suppression threshold. It should be noted that the target window here can be any window among the above-mentioned multiple windows of different lengths, and the given accuracy can be set according to actual needs, which is not limited in the embodiments of the present disclosure.

[0112] The disclosed embodiment can determine the functional relationship between the suppression threshold and the accuracy of each window based on a sample video set and a pre-trained classification model, and then, based on the functional relationship of each window, determine the suppression threshold of each window at a given accuracy as the above-mentioned suppression threshold, which helps to ensure the accuracy of suppressing the first fused positive confidence based on the size relationship between multiple fused negative confidences and the corresponding suppression thresholds, thereby reducing the false alarm rate and reducing disturbance to users.

[0113] In an optional example, when determining the above-mentioned suppression threshold, for multiple windows of different lengths, when the length difference between any two adjacent windows is greater than one frame of image, the suppression threshold of the window between the two adjacent windows can be determined using an interpolation formula based on the suppression thresholds corresponding to the two adjacent windows.

[0114] The embodiment of the present disclosure can use an interpolation formula to determine the suppression threshold of a window between two adjacent windows. Compared with the method of determining the suppression threshold of a window based on a sample video set and a pre-trained classification model, this method has a simple implementation process and a small amount of calculation, which helps to reduce the amount of calculation and save computing resources.

[0115] Any of the video object status detection methods provided in the embodiments of the present disclosure can be executed by any appropriate device with data processing capabilities, including but not limited to terminal devices and servers. Alternatively, any of the video object status detection methods provided in the embodiments of the present disclosure can be executed by a processor, such as by invoking corresponding instructions stored in a memory to execute any of the video object status detection methods mentioned in the embodiments of the present disclosure. This will not be further described below.

[0116] Exemplary devices

[0117] Figure 7 FIG. 1 is a schematic diagram of a state detection device for a video target provided by an exemplary embodiment of the present disclosure. The device of this embodiment can be used to implement the corresponding method embodiment of the present disclosure. Figure 7 The device shown includes: a positive confidence determination module 301, a fused positive confidence determination module 302 and a detection result determination module 303.

[0118] The positive confidence determination module 301 is used to determine the positive confidence of the state of the target in each frame of the video to be detected based on the pre-trained classification model;

[0119] The fused positive confidence determination module 302 is used to determine the fused positive confidence of the state of the target of multiple frames of images in the video to be detected corresponding to each window based on multiple windows of different lengths;

[0120] The detection result determination module 303 is used to determine the state detection result of the target in the video to be detected based on the fused positive confidence of each window and the corresponding fused positive confidence threshold; wherein the fused positive confidence threshold is determined based on a given accuracy rate.

[0121] The disclosed embodiment can respectively determine the fused positive confidence of the state of the target for each of the multiple frames of images corresponding to each window in a plurality of windows of different lengths, and then determine the state detection result of the target in the video to be detected based on the fused positive confidence of each window and the corresponding fused positive confidence threshold. This can implement a state output strategy based on multiple time windows, which is more flexible than using a fixed time window to detect the state of the target in the video to be detected. In addition, the fused positive confidence threshold of each window is determined based on a given accuracy rate, and the detection process of the short window is short in time. It can help to output the state detection result of the target in the video to be detected in a timely manner while ensuring the accuracy of the state detection result, and can improve the response sensitivity of the system.

[0122] Figure 8 FIG. 1 is a schematic diagram of a state detection device for a video target provided by another exemplary embodiment of the present disclosure. Figure 8 As shown above Figure 7 Based on the embodiment shown, the above apparatus may further include: a negative confidence determination module 304, a first window determination module 305, a fused negative confidence determination module 306 and a suppression module 307.

[0123] The negative confidence determination module 304 is used to determine the negative confidence of the state of the target in each frame of the video to be detected based on the pre-trained classification model;

[0124] The first window determination module 305 is configured to determine, from a plurality of windows of different lengths, a first window corresponding to a first fused positive confidence greater than a corresponding fused positive confidence threshold;

[0125] The fused negative confidence determination module 306 is configured to determine the fused negative confidence corresponding to each of the multiple windows preceding the first window based on a preset order of the multiple windows of different lengths;

[0126] The suppression module 307 is configured to suppress the first fused positive confidence that is greater than the corresponding fused positive confidence threshold based on the magnitude relationship between the multiple fused negative confidences and the corresponding suppression thresholds.

[0127] In an optional example, under the same accuracy requirement, the corresponding fusion positive confidence threshold is negatively correlated with the length of each window.

[0128] In an optional example, the above-mentioned fused positive confidence determination module 302 can be used to determine the fused positive confidence of the state of the target in each window according to a preset method based on the positive confidence of the state of the target in each frame image in multiple images corresponding to each window.

[0129] In an alternative example, Figure 9 FIG. 1 is a structural diagram of a video target state detection device provided by another exemplary embodiment of the present disclosure. Figure 9 The detection result determination module 303 shown may include: a first detection order determination unit 303-1, a judgment unit 303-2 and a detection result determination unit 303-3.

[0130] The first detection order determination unit 303-1 is configured to determine the detection order of multiple fused positive confidences corresponding to multiple windows of different lengths in ascending order of window lengths;

[0131] The judging unit 303-2 is configured to sequentially judge the magnitude relationship between the multiple fused positive confidences and the corresponding fused positive confidence thresholds based on the detection order of the multiple fused positive confidences;

[0132] The detection result determination unit 303 - 3 is used to determine the state detection result of the target in the video to be detected based on the size relationship.

[0133] In an optional example, the above-mentioned judgment unit 303-2 can be used to determine the first fused positive confidence that is greater than the corresponding fused positive confidence threshold from multiple fused positive confidences based on the detection order of multiple fused negative confidences.

[0134] In an optional example, the above-mentioned detection result determination unit 303-2 can be used to determine that the state detection result of the target in the video to be detected is that the target in the video to be detected exists in the above-mentioned state when the first fused positive confidence that is greater than the corresponding fused positive confidence threshold is determined from multiple fused positive confidences.

[0135] In an optional example, the suppression module 307 may include: a second detection order determination unit and a replacement unit.

[0136] The second detection order determination unit is used to determine the detection order of multiple fused negative confidences corresponding to multiple windows before the first window according to the preset order of the window lengths;

[0137] The replacement unit is used to replace the first fused positive confidence greater than the corresponding fused positive confidence threshold with the first fused negative confidence less than the corresponding suppression threshold based on the detection order of multiple fused negative confidences, when the first fused negative confidence less than the corresponding suppression threshold is determined from the multiple fused negative confidences.

[0138] In an optional example, the above-mentioned fusion positive confidence threshold can be determined in the following manner: for each sample video in the sample video set, first determine the sample positive confidence of the state of the target in each frame sample image of the sample video based on the pre-trained classification model, and determine the sample fusion positive confidence of multiple frames of sample images in the sample video corresponding to the target window based on the sample positive confidence, and determine the state detection result of the target in the sample video based on the sample fusion positive confidence and the fusion positive confidence threshold of the target window, and then determine the accuracy of the target window based on the state detection result of the target in each sample video in the sample video set, and determine the functional relationship between the fusion positive confidence threshold of the target window and the accuracy, and then determine the fusion positive confidence threshold of the target window at a given accuracy based on the functional relationship as the above-mentioned fusion positive confidence threshold. It should be noted that the target window here can be any window among the above-mentioned multiple windows of different lengths, and the given accuracy can be set according to actual needs, which is not limited in the embodiments of the present disclosure.

[0139] In an optional example, the method for determining the above-mentioned fusion positive confidence threshold may also include: for multiple windows of different lengths, when the length difference between any two adjacent windows is greater than one frame of image, based on the fusion positive confidence threshold corresponding to the two adjacent windows, using an interpolation formula to determine the fusion positive confidence threshold of the window between the two adjacent windows.

[0140] In an optional example, the above-mentioned suppression threshold is determined in the following manner: for each sample video in the sample video set, first determine the sample negative confidence of each frame sample image of the sample video based on the pre-trained classification model, and based on the sample negative confidence, determine the sample fusion negative confidence of multiple frames of sample images in the sample video corresponding to each target window, and determine the state detection result of the target in the sample video based on the sample negative confidence and the suppression threshold of the target window, and then determine the accuracy of the target window based on the state detection result of the target in each sample video in the sample video set, and determine the functional relationship between the suppression threshold of the target window and the accuracy, and then determine the suppression threshold of the target window at a given accuracy based on the functional relationship as the suppression threshold. It should be noted that the target window here can be any window among the above-mentioned multiple windows of different lengths, and the given accuracy can be set according to actual needs, which is not limited in the embodiments of the present disclosure.

[0141] In an optional example, the method for determining the above-mentioned suppression threshold may also include: for multiple windows of different lengths, when the length difference between any two adjacent windows is greater than one frame of image, based on the suppression thresholds corresponding to the two adjacent windows, using an interpolation formula to determine the suppression threshold of the window between the two adjacent windows.

[0142] Exemplary electronic devices

[0143] Below, reference Figure 10 4. The electronic device according to an embodiment of the present disclosure is described as follows. The electronic device includes one or more processors 401 and a memory 402.

[0144] The processor 401 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0145] The memory 402 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 401 may execute the program instructions to implement the behavior state detection method of the video target of each embodiment of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. may also be stored in the computer-readable storage medium.

[0146] In one example, the electronic device may further include an input device 403 and an output device 404, which are interconnected via a bus system and / or other connection mechanisms (not shown). The input device 403 may be the aforementioned microphone or microphone array for capturing input signals from a sound source, or a communication network connector for receiving collected input signals from other electronic devices.

[0147] In addition, the input device 403 may also include, for example, a keyboard, a mouse, and the like.

[0148] The output device 404 can output various information to the outside, including determined distance information, direction information, etc. The output device 404 can include, for example, a display, a speaker, a printer, a communication network and its connected remote output device, etc.

[0149] Of course, to simplify, Figure 10 Only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.

[0150] Exemplary computer program products and computer-readable storage media

[0151] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the behavioral state detection method of a video target according to various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section of this specification.

[0152] The computer program product may be written in any combination of one or more programming languages ​​to implement the operations of the disclosed embodiments, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0153] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enable the processor to execute the steps of the state detection method of a video target according to various embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.

[0154] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0155] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0156] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. For system embodiments, since they largely correspond to method embodiments, their description is relatively simple. For relevant parts, references to the description of the method embodiments are sufficient.

[0157] The block diagrams of the devices, devices, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0158] The methods and apparatus of the present disclosure may be implemented in many ways. For example, the methods and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.

[0159] It should also be noted that in the apparatus, device, and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.

[0160] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0161] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A method for detecting the state of a video target, wherein: include: Based on the pre-trained classification model, determine the positive confidence of the state of the target in each frame of the video to be detected; Based on a plurality of windows of different lengths, respectively determining a fused positive confidence of a plurality of frames of images in the to-be-detected video corresponding to each window with respect to the state of the target; Based on the fused positive confidence of each window and the corresponding fused positive confidence threshold, the state detection result of the target in the video to be detected is determined; wherein, the fused positive confidence threshold is determined based on the functional relationship between the fused positive confidence threshold of each window and the accuracy, and the given accuracy, and the functional relationship is predetermined based on the sample video set and the pre-trained classification model.

2. The method according to claim 1, wherein Under the same accuracy requirement, the corresponding fusion positive confidence threshold is negatively correlated with the length of each window.

3. The method according to claim 2, wherein: The step of respectively determining the fused positive confidence of the multiple frames of images in the to-be-detected video corresponding to each window with respect to the state of the target includes: Based on the positive confidence of the state of the target in each frame of the multiple frames of images corresponding to each window, the fused positive confidence of the state of the target in the multiple frames of images corresponding to each window is determined in a preset manner.

4. The method according to claim 1, wherein Determining a state detection result of the target in the to-be-detected video based on the fused positive confidence of each window and a corresponding fused positive confidence threshold includes: Determining a detection order of multiple fused positive confidences corresponding to the multiple windows of different lengths in ascending order of window lengths; Based on the detection order of the multiple fused positive confidences, sequentially determining the magnitude relationship between the multiple fused positive confidences and the corresponding fused positive confidence thresholds; Based on the size relationship, a state detection result of the target in the video to be detected is determined.

5. The method according to claim 4, wherein The determining, based on the detection order, in sequence, the relationship between the multiple fused positive confidences and the corresponding fused positive confidence thresholds includes: Based on the detection order, a first fused positive confidence that is greater than a corresponding fused positive confidence threshold is determined from the multiple fused positive confidences.

6. The state detection method according to claim 5, wherein: The state detection method further includes: Determining a negative confidence score of a state of an object in each frame of the video based on the pre-trained classification model; Determining, from the multiple windows of different lengths, a first window corresponding to the first fused positive confidence greater than the corresponding fused positive confidence threshold; Determining, based on a preset order of the multiple windows of different lengths, fused negative confidences corresponding to each of the multiple windows preceding the first window; Based on the size relationship between the multiple fused negative confidences and the corresponding suppression thresholds, the first fused positive confidence that is greater than the corresponding fused positive confidence threshold is suppressed.

7. The method according to claim 1, wherein The corresponding fusion positive confidence threshold is determined as follows: For each sample video in the sample video set: determining, based on the pre-trained classification model, a sample positive confidence score of the state of the target in each frame of the sample image of the sample video; Based on the sample positive confidence, determine the sample fusion positive confidence of multiple frames of sample images in the sample video corresponding to the target window, and determine the state detection result of the target in the sample video based on the sample fusion positive confidence and the fusion positive confidence threshold of the target window; wherein the target window is any window among the multiple windows of different lengths; Determining the accuracy of the target window based on a state detection result of the target in each sample video in the sample video set; Determining a functional relationship between a fusion positive confidence threshold of the target window and the accuracy rate; Based on the functional relationship, the fused forward confidence threshold of the target window at the given accuracy is determined as the fused forward confidence threshold corresponding to the target window.

8. A state detection device for a video target, wherein: include: A positive confidence determination module is used to determine the positive confidence of the state of the target in each frame of the video to be detected based on a pre-trained classification model; A fused forward confidence determination module is used to determine, based on a plurality of windows of different lengths, a fused forward confidence of the state of the target for each of the plurality of frames of the image in the video to be detected corresponding to each window; A detection result determination module is used to determine the state detection result of the target in the video to be detected based on the fused positive confidence of each window and the corresponding fused positive confidence threshold; wherein the fused positive confidence threshold is determined based on the functional relationship between the fused positive confidence threshold of each window and the accuracy, and a given accuracy, and the functional relationship is predetermined based on the sample video set and the pre-trained classification model.

9. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the method for detecting the state of a video object according to any one of claims 1 to 7.

10. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the video target status detection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Confidence determination method based on trajectory analysis, roadside equipment and cloud control platform

    CN112528927A

  • Detecting a State of a Wearable Device

    US20160041048A1