Method and system for determining abnormal video frame, electronic equipment and computer medium

Through the "debate + referee" mechanism of the multi-agent decision-making architecture, the multi-modal large model is used for secondary verification to address the false alarm problem of traditional visual algorithms, which solves the problem of high false alarm rate in the video playback automation test, and realizes high-precision and low false alarm video quality detection.

CN120343326APending Publication Date: 2025-07-18ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510619237.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing video playback automation test, traditional visual algorithms have high false alarm rate when detecting abnormal video frames, resulting in large workload and low efficiency and insufficient accuracy.

Method used

The multi-agent decision-making architecture is adopted, and the multi-agent system composed of three multi-modal large models is used to perform secondary verification analysis, and the "debate + referee" secondary verification mechanism is introduced. The reasoning process information of the first and second agents is used to debate and refute, and the third agent makes the final conclusion to reduce the false positive rate.

Benefits of technology

It significantly reduces the false alarm rate of video playback quality detection, improves the accuracy and robustness of detection, reduces the workload of manual re-inspection, and improves the degree of automation and testing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343326A_ABST
    Figure CN120343326A_ABST
Patent Text Reader

Abstract

The invention provides a method for determining an abnormal video frame. The method comprises the following steps: acquiring a to-be-determined video frame and a to-be-determined abnormal type corresponding to the to-be-determined video frame; inputting the to-be-determined video frame and the to-be-determined anomaly type into a first intelligent agent, and obtaining first reasoning process information through the first intelligent agent; inputting the to-be-determined video frame and the to-be-determined anomaly type into a second intelligent agent, and obtaining second reasoning process information through the second intelligent agent; and inputting the to-be-determined video frame, the to-be-determined anomaly type, the first reasoning process information and the second reasoning process information into a third intelligent agent, and obtaining conclusion information about whether the to-be-determined video frame is an abnormal video frame or not through the third intelligent agent. The disclosure also provides a system, an electronic device, a computer readable medium, and a computer program product for determining abnormal video frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of machine vision technology, and in particular, to a method and system for determining abnormal video frames, an electronic device, a computer-readable medium, and a computer program product. Background Art

[0002] Some video playback automated testing works are based on traditional vision algorithms. For example, object detection model You Only Look Once (YOLO), image classification model Residual Network (ResNet), etc. have a high recall rate of abnormal video frames in automatically detecting abnormal frames during video playback, preventing the playback of abnormal video frames. However, limited by the lack of visual understanding ability, the precision is usually very low in the actual process.

[0003] Currently, it mainly relies on manual viewing to recheck whether abnormal frames (such as noise, vertical stripes, mosaics, blurs, etc.) appear during video playback. The workload of manual secondary recheck is extremely large, with low efficiency and insufficient accuracy. Summary of the Invention

[0004] The present disclosure provides a method and system for determining abnormal video frames, an electronic device, a computer-readable medium, and a computer program product.

[0005] According to an aspect of the present disclosure, a method for determining an abnormal video frame is provided, including: obtaining a video frame to be determined and a type of abnormality to be determined corresponding to the video frame to be determined; inputting the video frame to be determined and the type of abnormality to be determined into a first intelligent agent, and obtaining first inference process information through the first intelligent agent, where the first inference process information indicates that the video frame to be determined is not an abnormal video frame; inputting the video frame to be determined and the type of abnormality to be determined into a second intelligent agent, and obtaining second inference process information through the second intelligent agent, where the second inference process information indicates that the video frame to be determined is an abnormal video frame; inputting the video frame to be determined, the type of abnormality to be determined, the first inference process information, and the second inference process information into a third intelligent agent, and obtaining conclusion information on whether the video frame to be determined is an abnormal video frame through the third intelligent agent.

[0006] According to another aspect of the present disclosure, there is provided a system for determining abnormal video frames, including: a first agent configured to obtain first inference process information for a video frame to be determined and a corresponding abnormal type to be determined for the video frame to be determined, wherein the first inference process information indicates that the video frame to be determined is not an abnormal video frame; a second agent configured to obtain second inference process information for the video frame to be determined and the abnormal type to be determined, wherein the second inference process information indicates that the video frame to be determined is an abnormal video frame; and a third agent configured to obtain conclusion information as to whether the video frame to be determined is an abnormal video frame based on the video frame to be determined, the abnormal type to be determined, the first inference process information, and the second inference process information.

[0007] According to another aspect of the present disclosure, there is provided an electronic device including a memory and a processor, the memory storing a computer program executable by the processor, and when the computer program is executed by the processor, the processor is caused to execute the method for determining abnormal video frames according to the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is caused to execute the method for determining abnormal video frames according to the present disclosure.

[0009] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, wherein when the computer program is executed by a processor, the processor is caused to execute the method for determining abnormal video frames according to the present disclosure.

[0010] The method for determining abnormal video frames according to the present disclosure performs secondary verification and analysis on a video frame to be determined through multiple agents, provides a multi-agent decision-making architecture, and introduces a "debate-style" secondary verification mechanism for the video frame to be determined detected by a traditional video playback quality detection algorithm, reducing the false alarm rate of video playback quality detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In the drawings of the embodiments of the present disclosure:

[0012] Figures 1A to 1D Several common video frames that are easily misreported are shown;

[0013] Figure 2 A flowchart of the method for determining abnormal video frames according to an embodiment of the present disclosure is shown;

[0014] Figure 3 Another flowchart of the method for determining abnormal video frames according to an embodiment of the present disclosure is shown;

[0015] Figure 4 Shows another flowchart of a method for determining an abnormal video frame according to an embodiment of the present disclosure;

[0016] Figure 5 Shows another flowchart of a method for determining an abnormal video frame according to an embodiment of the present disclosure;

[0017] Figure 6 Shows Figure 5 The flowchart of step S200 shown;

[0018] Figure 7 Shows Figure 6 The flowchart of step S210 shown;

[0019] Figure 8 Shows Figure 6 The flowchart of step S220 shown;

[0020] Figures 9A to 9D Schematically shows an abnormal image obtained by an abnormal conversion algorithm;

[0021] Figure 10 Shows Figure 6 The flowchart of step S230 shown;

[0022] Figure 11 Shows a schematic structural diagram of a system for determining an abnormal video frame according to an embodiment of the present disclosure;

[0023] Figure 12 Is a block diagram of the composition of an electronic device according to an embodiment of the present disclosure;

[0024] Figure 13 Is a block diagram of the composition of a computer-readable medium according to an embodiment of the present disclosure. Detailed implementation manners

[0025] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0026] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, but the embodiments shown may be embodied in different forms and the present disclosure should not be construed as limited to the embodiments set forth below. On the contrary, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0027] The accompanying drawings of the embodiments of the present disclosure are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification. Together with the detailed embodiments, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. By describing the detailed embodiments with reference to the accompanying drawings, the above and other features and advantages will become more apparent to those skilled in the art.

[0028] The present disclosure may be described with reference to plan views and / or cross-sectional views by means of ideal schematic diagrams of the present disclosure. Therefore, the example illustrations may be modified according to manufacturing techniques and / or tolerances.

[0029] Without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0030] The terms used in the present disclosure are only used to describe specific embodiments and are not intended to limit the present disclosure. As used in the present disclosure, the term "and / or" includes any and all combinations of one or more related listed items. As used in the present disclosure, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. As used in the present disclosure, the terms "include" and "made of" specify the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their groups.

[0031] Unless otherwise defined, all terms (including technical terms and scientific terms) used in the present disclosure have the same meaning as commonly understood by those of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless the present disclosure clearly defines so.

[0032] The present disclosure is not limited to the embodiments shown in the accompanying drawings, but includes modifications to the configurations formed based on manufacturing processes. Therefore, the regions illustrated in the accompanying drawings have schematic attributes, and the shapes of the regions shown in the drawings illustrate the specific shapes of the regions, but are not intended to be restrictive.

[0033] In recent years, the detection of video playback quality has gradually developed towards automated intelligent detection, but there are still problems with high false alarm rates in traditional visual algorithm detections. Traditional visual algorithms, such as YOLO, image classification ResNet, etc., are very likely to misreport specific types of video images as abnormal video frames.

[0034] Figures 1A to 1D Several common video frames that are prone to false alarms are shown. As Figures 1A to 1D shown, video frames with more noise (for example, Figure 1AThe flower field shown) is misreported as an abnormal video frame of the "noise" type, and the picture with relatively dense vertical stripes (for example, Figure 1B The bookshelf shown) is misreported as an abnormal video frame of the "vertical stripe" type, and the picture with local squares (for example, Figure 1C The seat shown) is misreported as an abnormal video frame of the "mosaic" type, and the normal shot language (for example, Figure 1D The picture on the TV screen shown) is misreported as an abnormal video frame of the "blur" type.

[0035] In the face of a large number of abnormal false alarm frames, the workload of manual secondary recheck is extremely large, and real abnormal video frames may be ignored due to factors such as manual fatigue during recheck. Therefore, improving the accuracy of video playback testing plays a crucial role in reducing testing costs, improving automation levels, and testing accuracy, and is in line with the development trend of video playback testing automation.

[0036] Figure 2 The flowchart of the method for determining abnormal video frames according to an embodiment of the present disclosure is shown.

[0037] As Figure 2 shown, the method for determining abnormal video frames according to an embodiment of the present disclosure includes the following steps S110 to S140.

[0038] In step S110, obtain the video frame to be determined and the abnormal type to be determined corresponding to the video frame to be determined.

[0039] When existing video playback quality detection schemes detect abnormal video frames and determine the abnormal types corresponding to the detected abnormal video frames, since the false alarm rates of these video playback quality detection schemes are relatively high, it is necessary to perform secondary verification on the detected abnormal video frames to determine whether these video frames are abnormal video frames corresponding to the detected abnormal types. In this step, obtain the video frame to be determined that may be an abnormal video frame detected by the existing video playback quality detection scheme and the abnormal type to be determined corresponding to the video frame to be determined. The specific acquisition method is not limited in this application.

[0040] In step S120, input the video frame to be determined and the abnormal type to be determined into the first intelligent agent, and obtain the first inference process information through the first intelligent agent. The first inference process information indicates that the video frame to be determined is not an abnormal video frame.

[0041] In step S130, input the video frame to be determined and the abnormal type to be determined into the second intelligent agent, and obtain the second inference process information through the second intelligent agent. The second inference process information indicates that the video frame to be determined is an abnormal video frame.

[0042] After obtaining the video frame to be determined and the corresponding type of anomaly to be determined, the obtained video frame to be determined and the corresponding type of anomaly are input into the first agent and the second agent. Therefore, the first agent and the second agent are multimodal agents that can support multiple modalities (e.g., including images, text, etc.).

[0043] The present disclosure proposes a determination mode of "debate + referee". The first agent "plays" the role of the "proponent" in the "debate", while the second agent "plays" the role of the "opponent" in the "debate". The view supported by the first agent is that it is a false alarm to identify the video frame to be determined as an abnormal video frame, that is, the video frame to be determined is not an abnormal video frame; in contrast, the view supported by the second agent is that it is not a false alarm to identify the video frame to be determined as an abnormal video frame, that is, the video frame to be determined is an abnormal video frame.

[0044] According to an embodiment of the present disclosure, by inputting the first role definition prompt word into the first agent, the first agent can "play" the role of the "proponent" defined by the first role definition prompt word in the "debate" to output the first reasoning process information indicating that the video frame to be determined is not an abnormal video frame; in addition, by inputting the second role definition prompt word into the second agent, the second agent can "play" the role of the "opponent" defined by the second role definition prompt word in the "debate" to output the second reasoning process information indicating that the video frame to be determined is an abnormal video frame.

[0045] In the present disclosure, a "role" is a logical abstraction of an agent. An agent "role" can perform specific actions, have memory, be able to think and adopt action strategies.

[0046] For example, the first role definition prompt word input into the first agent can be set as: You are a proponent in a debate competition. Your core task is to defend and prove your view. Your view is that the input video frame to be determined is wrongly identified as an abnormal video frame corresponding to the input type of anomaly to be determined. You need to clearly state your view, emphasizing the rationality, necessity, and positivity of the view; use strict logical reasoning and rich factual basis to effectively support your view. Another example, the second role definition prompt word input into the second agent can be set as: You are an opponent in a debate competition. Your core task is to defend and prove your view. Your view is that the input video frame to be determined is an abnormal video frame corresponding to the input type of anomaly to be determined. You need to clearly state your view, emphasizing the rationality, necessity, and positivity of the view; use strict logical reasoning and rich factual basis to effectively support your view.

[0047] According to an embodiment of the present disclosure, steps S120 and S130 can be regarded as the processes of presenting arguments by the "pro side" role "played" by the first intelligent agent and the "con side" role "played" by the second intelligent agent during the debate. During the process of presenting arguments, the first intelligent agent puts forward the view that determining the video frame to be determined as an abnormal video frame is a false alarm based on the role it "plays", and provides the first reasoning process information that can support this view; the second intelligent agent puts forward the view that determining the video frame to be determined as an abnormal video frame is not a false alarm based on the role it "plays", and provides the second reasoning process information that can support this view.

[0048] In step S140, the video frame to be determined, the abnormal type to be determined, the first reasoning process information, and the second reasoning process information are input into the third intelligent agent, and the conclusion information on whether the video frame to be determined is an abnormal video frame is obtained through the third intelligent agent.

[0049] According to an embodiment of the present disclosure, the third intelligent agent can make a final conclusion by combining the video frame to be determined, the abnormal type to be determined, and the first reasoning process information and the second reasoning process information output by the first intelligent agent and the second intelligent agent, that is, the video frame to be determined is an abnormal video frame, or the video frame to be determined is not an abnormal video frame.

[0050] According to an embodiment of the present disclosure, the first intelligent agent, the second intelligent agent, and the third intelligent agent can be constructed based on a general domain large multimodal model (Large Multimodal Model, LMM), for example, Qwen-VL series models, InternVL series models, MiniCPM-V series models, DeepSeek-VL series models, etc.

[0051] According to an embodiment of the present disclosure, a method for improving the accuracy of video playback testing based on multimodality and multi-intelligent agents is proposed. For the abnormal video frame to be determined detected by the existing video playback quality detection scheme, a multi-agent system (Multi-Agent System, MAS) composed of three large multimodal models is used for secondary verification analysis. The first intelligent agent believes that the abnormal video frame to be determined is a false alarm, the second intelligent agent believes that the abnormal video frame to be determined is not a false alarm, and the third intelligent agent makes a final conclusion by combining the video frame to be determined, the abnormal type to be determined, and the reasoning processes of the first intelligent agent and the second intelligent agent. Through this multi-intelligent agent decision-making architecture, a secondary verification mechanism of "debate + referee" is introduced for the abnormal video frame to be determined detected by the existing video playback quality detection scheme, significantly reducing the false alarm rate of video playback quality detection. On the other hand, the large multimodal model integrates various modality information such as images and texts for comprehensive reasoning analysis, which can significantly improve the detection accuracy and improve the accuracy and robustness of video playback quality detection.

[0052] Figure 3 Another flowchart showing a method for determining abnormal video frames according to an embodiment of the present disclosure is presented.

[0053] As Figure 3 shown, the method for determining abnormal video frames according to an embodiment of the present disclosure may further include the following steps S131 to S132.

[0054] In step S131, the second inference process information is input into the first agent, and the first feedback information is obtained through the first agent.

[0055] In step S132, the first inference process information is input into the second agent, and the second feedback information is obtained through the second agent.

[0056] According to an embodiment of the present disclosure, steps S131 and S132 can be regarded as the processes of the "pro" role "played" by the first agent and the "con" role "played" by the second agent respectively refuting in the debate process. In the refutation process, the first agent refutes the view and its reasoning process proposed by the second agent, and the second agent refutes the view and its reasoning process proposed by the first agent.

[0057] As Figure 3 shown, obtaining the conclusion information on whether the video frame to be determined is an abnormal video frame through the third agent (i.e., step S140) includes step S140'.

[0058] In step S140', the video frame to be determined, the abnormal type to be determined, the first inference process information, the second inference process information, the first feedback information, and the second feedback information are input into the third agent, and the conclusion information on whether the video frame to be determined is an abnormal video frame is obtained through the third agent.

[0059] According to an embodiment of the present disclosure, the first feedback information obtained through the first agent and the second feedback information obtained through the second agent may represent the same or opposite conclusions. That is to say, in the refutation process, the first agent (or the second agent) may not be able to form a logically self-consistent refutation against the view and reasoning process proposed by the other party, so that in the refutation process, the first agent and the second agent form the same conclusion. In this case, the third agent can be made to output a conclusion with high confidence. On the other hand, in the refutation process, the first agent (or the second agent) can form a logically self-consistent refutation against the view and reasoning process proposed by the other party, thereby further supporting its own view.

[0060] Figure 4 Another flowchart showing a method for determining abnormal video frames according to an embodiment of the present disclosure is presented.

[0061] As Figure 4 shown, the method for determining an abnormal video frame according to an embodiment of the present disclosure may further include the following steps S133 to S134.

[0062] In step S133, the second feedback information is input into the first agent, and the third feedback information is obtained through the first agent.

[0063] In step S134, the first feedback information is input into the second agent, and the fourth feedback information is obtained through the second agent.

[0064] According to an embodiment of the present disclosure, steps S133 and S134 can be regarded as a process in which the "pro" role "played" by the first agent and the "con" role "played" by the second agent respectively defend themselves during a debate. During the defense process, the first agent defends its own view against the refutation proposed by the second agent, and the second agent defends its own view against the refutation proposed by the first agent.

[0065] As Figure 4 shown, obtaining the conclusion information on whether the video frame to be determined is an abnormal video frame through the third agent (i.e., step S140) includes step S140".

[0066] In step S140", the video frame to be determined, the abnormal type to be determined, the first inference process information, the second inference process information, the first feedback information, the second feedback information, the third feedback information, and the fourth feedback information are input into the third agent, and the conclusion information on whether the video frame to be determined is an abnormal video frame is obtained through the third agent.

[0067] According to an embodiment of the present disclosure, the third feedback information obtained through the first agent and the fourth feedback information obtained through the second agent may represent the same or opposite conclusions. That is to say, during the defense process, the first agent (or the second agent) may not be able to form a logically self-consistent support for its own view against the refutation proposed by the other party, so that during the defense process, the first agent and the second agent form the same conclusion. In this case, the third agent can be made to output a conclusion with high confidence. On the other hand, during the defense process, the first agent (or the second agent) can form a logically self-consistent support for its own view against the refutation proposed by the other party, thereby further improving its own view.

[0068] A method for determining abnormal video frames according to an embodiment of the present disclosure introduces three multi-modal agents with clear roles, analyzes and makes decisions on abnormal images through a debate mechanism. During the process of presenting arguments → refuting arguments → defending, the first agent and the second agent provide rich reasoning processes, refutations, and defenses as the basis for the third agent to make decisions, significantly improving the accuracy and confidence of the conclusions. During the debate process, if the first agent and the second agent reach a consistent conclusion, the third agent can output a conclusion with high confidence; if the first agent and the second agent cannot reach a consistent conclusion, the third agent can also comprehensively analyze and judge the video frame to be determined based on the reasoning processes, refutations, and defenses provided by both, improving the accuracy of the decision-making.

[0069] On the other hand, the present disclosure also proposes an automatic acquisition method for abnormal false alarm images. Since the method for determining abnormal video frames according to the present disclosure can be used as an optimization method for various video playback quality detection schemes, and the recognition capabilities of abnormal video frames of different original schemes are different, and it is also difficult to manually analyze each abnormal video frame detected by each original scheme one by one. Therefore, a corpus construction method with high generalization is required to train the agents.

[0070] Figure 5 Another flowchart of the method for determining abnormal video frames according to an embodiment of the present disclosure is shown.

[0071] As Figure 5 shown, the method for determining abnormal video frames according to an embodiment of the present disclosure may further include the following steps S200 to S300.

[0072] In step S200, a corpus for training the third agent is constructed.

[0073] In step S300, the third agent is trained using the corpus.

[0074] Figure 6 Shown is Figure 5 the flowchart of step S200 shown above.

[0075] As Figure 6 shown, according to an embodiment of the present disclosure, constructing a corpus for training the third agent (i.e., step S200) may include the following steps S210 to S240.

[0076] In step S210, a normal image library is constructed, and the normal image library includes normal images misreported as abnormal images.

[0077] In step S220, an abnormal image library is constructed, and the abnormal image library includes abnormal images determined to be abnormal.

[0078] In step S230, for each image in the normal image library and the abnormal image library, first debate information is obtained through a first agent, and second debate information is obtained through a second agent.

[0079] In step S240, a corpus is constructed based on the normal image library, the abnormal image library, and the first debate information and the second debate information corresponding to each image in the normal image library and the abnormal image library.

[0080] According to an embodiment of the present disclosure, a method for constructing a high-quality training corpus is proposed. A normal image library is constructed for "false alarm interception", and an abnormal image library is constructed for "non-false alarm release". By constructing the normal image library and the abnormal image library, the problems of low efficiency and deviation of traditional manual annotation can be effectively solved.

[0081] Figure 7 shows Figure 6 the flowchart of step S210 shown.

[0082] As Figure 7 shown, according to an embodiment of the present disclosure, constructing the normal image library (i.e., step S210) may include the following steps S211 to S215.

[0083] In step S211, normal image resources are acquired.

[0084] The normal image resources can be acquired by obtaining video frames or high-definition images in high-definition video resources. It should be recognized that these video frames or images are not abnormal. If the existing video playback quality detection scheme detects these normal video frames or images as abnormal images, it can be confirmed that this abnormal detection result is a false alarm. To ensure generalization, the video frames or high-definition images in the high-definition video resources should cover various types and should include various different image resolutions.

[0085] In step S212, the images detected as abnormal by a preset image detection algorithm in the normal image resources and the abnormal types corresponding to the images detected as abnormal are acquired.

[0086] The preset image detection algorithm can be various existing video playback quality detection schemes. For example, anomaly detection schemes based on the YOLO or GroundingDINO model, and anomaly image classification schemes based on the ResNet or MobileNet model. At this step, collect the anomaly images detected by the existing video playback quality detection schemes and the corresponding anomaly types. It should be recognized that at this time, all the detection results of the images detected as anomaly images are false positives. If these detection schemes detect high-quality video frames or images as anomalies, it means that these detection schemes have insufficient detection capabilities for these images, and this should be considered key when improving the false positive interception ability of the decision-making model in the future.

[0087] In step S213, count the number of images detected as anomalies corresponding to various anomaly types.

[0088] In step S214, determine whether the number of images detected as anomalies corresponding to various anomaly types is less than the threshold. If there is an anomaly type for which the number of images detected as anomalies is less than the threshold, return to step S211; otherwise, continue to execute step S215.

[0089] Count the number of images of various different anomaly types detected by the existing video playback quality detection schemes. If the number of images of one (or more) anomaly types is insufficient (i.e., less than the threshold), then repeat steps S211 to S213 until the number of images of all anomaly types reaches the threshold.

[0090] According to an embodiment of the present disclosure, the anomaly types may include (but are not limited to) noise, vertical stripes, mosaic, blur, resolution anomaly, black screen, etc.

[0091] In step S215, add each of all the images detected as anomalies to the normal image library and label it with the anomaly type corresponding to the image detected as an anomaly.

[0092] Save all the images misreported as anomalies by the existing video playback quality detection schemes and the corresponding anomaly types as the normal image library. After the above steps, a large-scale, high-quality normal image corpus for "false positive interception" can be constructed.

[0093] Figure 8 Shows Figure 6 The flowchart of step S220 shown.

[0094] As Figure 8 shown, according to an embodiment of the present disclosure, constructing the anomaly image library (i.e., step S220) may include the following steps S221 to S225.

[0095] In step S221, the normal image is converted into images of various abnormal types through an abnormal conversion algorithm.

[0096] Converting the normal image into an abnormal image through the abnormal conversion algorithm can stably obtain images of various abnormal types. The normal image can be obtained, for example, through the above-mentioned step S211. Subsequently, various types of abnormal conversion algorithms are performed on the normal image to obtain the abnormal image.

[0097] Figures 9A to 9D The abnormal image obtained through the abnormal conversion algorithm is schematically shown. Figure 9A The original normal image is shown, Figure 9B The image of the noise abnormal type constructed from the normal image is shown, Figure 9C The image of the vertical stripe abnormal type constructed from the normal image is shown, Figure 9D The image of the mosaic abnormal type constructed from the normal image is shown.

[0098] For example, an algorithm for constructing an image of the noise abnormal type includes: calculating the main colors of the picture through K-means clustering; generating a background color picture and overlaying rectangles; overlaying Gaussian noise, as Figure 9B shown. Another example, an algorithm for constructing an image of the vertical stripe abnormal type includes: randomly selecting image coordinates; extending the pixel values at the selected image coordinates upward or downward to cover the original pixel values, as Figure 9C shown. Still another example, an algorithm for constructing an image of the mosaic abnormal type includes: shrinking the input image; enlarging the shrunk image back to the original size to generate a mosaic in the image, as Figure 9D shown.

[0099] In step S222, the images detected as abnormal by a preset image detection algorithm in the converted image and the abnormal types corresponding to the images detected as abnormal are obtained.

[0100] Similar to step S212, in this step, the abnormal images detected by the existing video playback quality detection scheme and the corresponding abnormal types are collected. It should be recognized that at this time, all the detection results of the images detected as abnormal are not false positives, that is, all the images can be determined as abnormal images. In this way, abnormal images of various abnormal types can be stably obtained.

[0101] In step S223, the number of images detected as abnormal corresponding to various abnormal types is counted.

[0102] In step S224, it is determined whether the number of images detected as abnormal corresponding to various abnormal types is less than a threshold. If the number of images detected as abnormal corresponding to a specific abnormal type is less than the threshold, return to step S221; otherwise, continue to execute step S225.

[0103] Count the number of images of various different abnormal types detected by the existing video playback quality detection scheme. If the number of images of one (or more) abnormal type is insufficient (i.e., less than the threshold), then repeat steps S221 to S223 until the number of images of all abnormal types reaches the threshold.

[0104] In step S225, each of all the images detected as abnormal is added to the abnormal image library and labeled with the abnormal type corresponding to the image detected as abnormal.

[0105] Save all the images misreported as abnormal by the existing video playback quality detection scheme and the corresponding abnormal types as the abnormal image library. Through the above various steps, a large-scale, high-quality abnormal image corpus for "non-false-alarm release" can be constructed.

[0106] Figure 10 shows Figure 6 The flowchart of step S230 shown.

[0107] As Figure 10 shown, according to an embodiment of the present disclosure, for each image in the normal image library and the abnormal image library, obtaining first debate information through the first agent and obtaining second debate information through the second agent (i.e., step S230) may include the following steps S231 to S236.

[0108] In step S231, input the image and the abnormal type corresponding to the image into the first agent, and obtain third inference process information through the first agent. The third inference process information indicates that the image is not an abnormal image corresponding to the abnormal type.

[0109] In step S232, input the image and the abnormal type corresponding to the image into the second agent, and obtain fourth inference process information through the second agent. The fourth inference process information indicates that the image is an abnormal image corresponding to the abnormal type.

[0110] In step S233, input the fourth inference process information into the first agent, and obtain fifth feedback information through the first agent.

[0111] In step S234, input the third inference process information into the second agent, and obtain sixth feedback information through the second agent.

[0112] In step S235, the sixth feedback information is input into the first agent, and the seventh feedback information is obtained through the first agent.

[0113] In step S236, the fifth feedback information is input into the second agent, and the eighth feedback information is obtained through the second agent.

[0114] The first debate information includes the third inference process information, the fifth feedback information, and the seventh feedback information, and the second debate information includes the fourth inference process information, the sixth feedback information, and the eighth feedback information.

[0115] Figure 10 The processes shown require the participation of the first agent and the second agent. Similar to the "debate" process between the first agent and the second agent described with reference to Figures 2 to 4 In the processes shown, the first agent and the second agent debate each image in the normal image library constructed in step S210 and the abnormal image library constructed in step S220, and combine the first debate information output by the first agent and the second debate information output by the second agent with the corresponding images to construct a corpus for training the third agent. That is, the corpus for training the third agent includes each image in the normal image library and the abnormal image library, the abnormal type corresponding to the image, and the first debate information and the second debate information corresponding to the image. During training, after obtaining the conclusion information based on the image, the first debate information, and the second debate information, the third agent compares it with the correct answer. The correct answer can be provided through the label or annotation carried by the image, or the correct answer can be provided through the source of the image (i.e., from the normal image library or the abnormal image library), and the source of the image can be represented by the storage path information of the image. By comparing the conclusion information obtained by the third agent with the correct answer and feeding back the comparison result to the third agent, the accuracy of the third agent's judgment can be optimized. Figure 10 The final conclusion output by the third agent plays a key role in reducing the false alarm rate. Therefore, the third agent must be fine-tuned. Through the above process, a corpus for training the third agent is realized, simulating the inference interaction process between agents during the debate, generating training data for supervising or strengthening the fine-tuning model, and improving the inference and decision-making capabilities of the model.

[0116]

[0117] According to an embodiment of the present disclosure, training the third agent using the corpus (i.e., step S300) may include: training the third agent using the corpus by means of supervised fine-tuning or reinforcement fine-tuning.

[0118] ​According to embodiments of the present disclosure, supervised fine-tuning may include (but is not limited to), for example, the Supervised Fine-Tuning (SFT) algorithm, and reinforcement fine-tuning may include (but is not limited to), for example, the Group Relative Policy Optimization (GRPO) algorithm. The present disclosure does not limit the algorithms for supervised fine-tuning and reinforcement fine-tuning.

[0119] The third agent is trained using supervised fine-tuning or reinforcement fine-tuning, so that the trained third agent has the ability to judge the reasoning process in the "debate mode" scenario, in order to intercept false alarms and output truly abnormal video frames.

[0120] For example, the format of the training corpus is as follows:

[0121] {

[0122] "query": "It is known that you are the referee of this debate, and the topic of the debate is 'Is the xx anomaly in the image to be analyzed a false alarm?'. Please output the view of the winning side based on the main arguments and logical reasoning processes of both the affirmative and negative sides:\n# Affirmative argument: xx\n# Negative argument: xx\n# Affirmative rebuttal: xx\n# Negative rebuttal: xx\n# Affirmative defense: xx\n# Negative defense: xx",

[0123] "response": "Fill in 'yes / no' here",

[0124] "images": "Fill in the path of the image to be analyzed here"

[0125] }

[0126] According to embodiments of the present disclosure, since knowledge of judgment is added to model training, the decision-making ability of the model is significantly improved. When the third agent performs reasoning, as long as the same prompt words are followed, the third agent can output conclusions with high confidence.

[0127] It should be noted that the trained third agent already has the perspectives of the first agent and the second agent. Since the first agent, the second agent, and the corpus obtained through the debate mechanism are only used for the training of the third agent, in actual deployment, only deploying the third agent can also achieve the effect of reducing false alarms.

[0128] The method for determining abnormal video frames according to embodiments of the present disclosure effectively solves problems such as high false alarm rates and weak generalization abilities in existing video quality detection solutions, improves detection accuracy, reduces labor costs, and improves user experience.

[0129] Figure 11Shows a schematic structural diagram of a system for determining abnormal video frames according to an embodiment of the present disclosure.

[0130] As Figure 11 shown, the system for determining abnormal video frames according to an embodiment of the present disclosure includes a first agent 101, a second agent 102, and a third agent 103.

[0131] The first agent 101 is configured to obtain first inference process information for a video frame to be determined and a corresponding abnormal type to be determined for the video frame to be determined. The first inference process information indicates that the video frame to be determined is not an abnormal video frame. The second agent 102 is configured to obtain second inference process information for the video frame to be determined and the abnormal type to be determined. The second inference process information indicates that the video frame to be determined is an abnormal video frame. The third agent 103 is configured to obtain conclusion information as to whether the video frame to be determined is an abnormal video frame based on the video frame to be determined, the abnormal type to be determined, the first inference process information, and the second inference process information.

[0132] The system for determining abnormal video frames according to an embodiment of the present disclosure is a multi-agent system. For the abnormal video frames detected by the existing video playback quality detection scheme, a multi-agent system including three multi-modal large models according to an embodiment of the present disclosure is used for secondary verification and analysis. The first agent believes that the abnormal video frame to be determined is a false alarm, the second agent believes that the abnormal video frame to be determined is not a false alarm, and the third agent combines the video frame to be determined, the abnormal type to be determined, and the inference processes of the first agent and the second agent to make a final conclusion. Through this multi-agent decision-making architecture, a "debate + referee" secondary verification mechanism is introduced for the abnormal video frames to be determined detected by the existing video playback quality detection scheme, significantly reducing the false alarm rate of video playback quality detection. On the other hand, the multi-modal large model integrates various modal information such as images and texts for comprehensive reasoning and analysis, which can significantly improve the detection accuracy and enhance the precision and robustness of video playback quality detection.

[0133] According to an embodiment of the present disclosure, the first agent 101 is further configured to obtain first feedback information for the second inference process information. The second agent 102 is further configured to obtain second feedback information for the first inference process information. The third agent 103 is further configured to obtain conclusion information as to whether the video frame to be determined is an abnormal video frame based on the video frame to be determined, the abnormal type to be determined, the first inference process information, the second inference process information, the first feedback information, and the second feedback information.

[0134] According to an embodiment of the present disclosure, the first agent 101 is further configured to obtain third feedback information for the second feedback information. The second agent 102 is further configured to obtain fourth feedback information for the first feedback information. The third agent 103 is further configured to obtain conclusion information as to whether the video frame to be determined is an abnormal video frame based on the video frame to be determined, the abnormal type to be determined, the first inference process information, the second inference process information, the first feedback information, the second feedback information, the third feedback information, and the fourth feedback information.

[0135] According to an embodiment of the present disclosure, the first feedback information obtained by the first agent 101 and the second feedback information obtained by the second agent 102 may represent the same or opposite conclusions.

[0136] According to an embodiment of the present disclosure, the third feedback information obtained by the first agent 101 and the fourth feedback information obtained by the second agent 102 may represent the same or opposite conclusions.

[0137] As Figure 11 shown, the first agent 101 and the second agent 102 interact with each other to implement a "debate" process, and the first agent 101 and the second agent 102 provide the information of their respective output debate information to the third agent 103 so that the third agent 103 can obtain a conclusion with high confidence based on these debate information.

[0138] The interaction between multiple agents can be divided into four stages: "proposition → refutation → defense → decision-making". The input and output of each agent in each stage are shown in Table 1:

[0139] Table 1

[0140]

[0141]

[0142] According to an embodiment of the present disclosure, by inputting the first role definition prompt word into the first agent 101, the first agent 101 can output the first inference process information indicating that the video frame to be determined is not an abnormal video frame. By inputting the second role definition prompt word into the second agent 102, the second agent 102 can output the second inference process information indicating that the video frame to be determined is an abnormal video frame.

[0143] According to an embodiment of the present disclosure, the third intelligent agent 103 is trained using a constructed corpus, and the constructed corpus includes: constructing a normal image library; constructing an abnormal image library; for each image in the normal image library and the abnormal image library, obtaining first debate information through the first intelligent agent 101 and obtaining second debate information through the second intelligent agent 102; constructing a corpus based on the normal image library, the abnormal image library, and the first debate information and the second debate information corresponding to each image in the normal image library and the abnormal image library.

[0144] According to an embodiment of the present disclosure, constructing a normal image library includes: obtaining normal image resources; obtaining images in the normal image resources that are detected as abnormal by a preset image detection algorithm and the corresponding abnormal types of the images detected as abnormal; counting the number of images detected as abnormal corresponding to each type of abnormal type; in the case where the number of images detected as abnormal corresponding to a specific abnormal type is less than a threshold, repeatedly executing the steps of obtaining normal image resources and obtaining images in the normal image resources that are detected as abnormal and the corresponding abnormal types of the images detected as abnormal until the number of images detected as abnormal corresponding to all abnormal types is greater than or equal to the threshold; adding each of all the images detected as abnormal to the normal image library and marking it with the abnormal type corresponding to the image detected as abnormal.

[0145] According to an embodiment of the present disclosure, constructing an abnormal image library includes: converting normal images into images of various abnormal types through an abnormal conversion algorithm; obtaining images in the converted images that are detected as abnormal by a preset image detection algorithm and the corresponding abnormal types of the images detected as abnormal; counting the number of images detected as abnormal corresponding to each type of abnormal type; in the case where the number of images detected as abnormal corresponding to a specific abnormal type is less than a threshold, repeatedly executing the steps of converting normal images into images of various abnormal types through the abnormal conversion algorithm and obtaining images in the converted images that are detected as abnormal and the corresponding abnormal types of the images detected as abnormal until the number of images detected as abnormal corresponding to all abnormal types is greater than or equal to the threshold; adding each of all the images detected as abnormal to the abnormal image library and marking it with the abnormal type corresponding to the image detected as abnormal.

[0146] According to an embodiment of the present disclosure, for each image in the normal image library and the abnormal image library, obtaining first debate information through the first agent 101 and obtaining second debate information through the second agent 102 includes: inputting the image and the corresponding abnormal type of the image into the first agent 101, and obtaining third inference process information through the first agent 101, where the third inference process information indicates that the image is not an abnormal image corresponding to the abnormal type; inputting the image and the corresponding abnormal type of the image into the second agent 102, and obtaining fourth inference process information through the second agent 102, where the fourth inference process information indicates that the image is an abnormal image corresponding to the abnormal type; inputting the fourth inference process information into the first agent 101, and obtaining fifth feedback information through the first agent 101; inputting the third inference process information into the second agent 102, and obtaining sixth feedback information through the second agent 102; inputting the sixth feedback information into the first agent 101, and obtaining seventh feedback information through the first agent 101; inputting the fifth feedback information into the second agent 102, and obtaining eighth feedback information through the second agent 102. The first debate information includes the third inference process information, the fifth feedback information, and the seventh feedback information, and the second debate information includes the fourth inference process information, the sixth feedback information, and the eighth feedback information.

[0147] According to an embodiment of the present disclosure, the third agent 103 is trained by using supervised fine-tuning or reinforcement fine-tuning with the constructed corpus.

[0148] The system for determining abnormal video frames according to an embodiment of the present disclosure can be used for monitoring and optimizing the quality of online video platforms. Specifically, it can be used for automatic detection of the video content quality of video websites and streaming media platforms, accurately identifying and intercepting video playback problems such as abnormal pictures, freezes, black screens, and flower screens, and reducing the false alarm rate.

[0149] In addition, the system for determining abnormal video frames according to an embodiment of the present disclosure can also be used for testing the video playback of set-top boxes, reducing the workload of manual secondary re-inspection, improving the test accuracy of video playback, reducing the test cost, and improving the automation level and test accuracy.

[0150] Furthermore, the system for determining abnormal video frames according to an embodiment of the present disclosure can be applied to multiple application scenarios such as video surveillance and security systems, quality control in film and television post-production, video quality inspection in intelligent driving and autonomous driving systems, etc., and can be extended to quality detection in games and virtual reality (VR / AR), robot vision navigation systems, analysis of satellite remote sensing images and videos, intelligent education and remote examination systems, quality assurance in immersive scenarios of the Metaverse, video monitoring of aerospace and drones, etc.

[0151] Figure 12 It is a block diagram of the composition of an electronic device according to an embodiment of the present disclosure.

[0152] As Figure 12 shown, the electronic device according to an embodiment of the present disclosure includes a memory 1202 and a processor 1201. The memory 1202 stores a computer program executable by the processor 1201. When the computer program is executed by the processor 1201, the processor 1201 is caused to execute a method for determining an abnormal video frame according to various embodiments of the present disclosure.

[0153] The processor 1201 is a device with data processing capabilities, which includes but is not limited to a central processing unit (CPU), etc.; the memory 1202 is a device with data storage capabilities, which includes but is not limited to a random access memory (RAM, more specifically such as SDRAM, DDR, etc.), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (FLASH), etc.

[0154] In addition, the electronic device according to the present disclosure may further include an I / O interface (read / write interface) 1203, which is connected between the processor 1201 and the memory 1202 and can realize the information interaction between the memory 1202 and the processor 1201, and includes but is not limited to a data bus (Bus), etc.

[0155] In some embodiments, the processor 1201, the memory 1202, and the I / O interface 1203 are interconnected through a bus 1204 and further connected to other components of the computing device.

[0156] Figure 13 It is a block diagram of the composition of a computer-readable medium according to an embodiment of the present disclosure.

[0157] As Figure 13 shown, on the computer-readable medium according to an embodiment of the present disclosure, a computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute a method for determining an abnormal video frame according to various embodiments of the present disclosure.

[0158] The embodiment of the present disclosure further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the processor is caused to execute a method for determining an abnormal video frame according to various embodiments of the present disclosure.

[0159] It should be recognized that the electronic device, the computer-readable medium, and the computer program product according to the embodiments of the present disclosure are all used to implement the method for determining an abnormal video frame according to various embodiments of the present disclosure. Therefore, the detailed description of the above method embodiments will not be repeated here.

[0160] Those of ordinary skill in the art will understand that all or some of the steps, systems, and functional modules / units in the devices disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0161] In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components. For example, one physical component can have multiple functions, or one function or step can be executed by several physical components working together.

[0162] Some or all physical components can be implemented as software executed by a processor, such as a central processing unit (CPU), a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH), or other magnetic disk storage; compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical disc storage; magnetic cassette, tape, magnetic disk storage, or other magnetic storage; and any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0163] The present disclosure has disclosed exemplary embodiments, and although specific terms have been used, they are used only and should be construed only as having a general illustrative meaning and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly stated, the features, characteristics, and / or elements described in combination with a particular embodiment can be used alone or in combination with the features, characteristics, and / or elements described in combination with other embodiments. Accordingly, those skilled in the art will understand that various forms and details can be changed without departing from the scope of the present disclosure as set forth in the appended claims.

Claims

1. A method for determining an abnormal video frame, comprising: Obtaining a video frame to be determined and a type of abnormality to be determined corresponding to the video frame to be determined; Inputting the video frame to be determined and the type of abnormality to be determined into a first intelligent agent, and obtaining first inference process information through the first intelligent agent, wherein the first inference process information indicates that the video frame to be determined is not an abnormal video frame; Inputting the video frame to be determined and the type of abnormality to be determined into a second intelligent agent, and obtaining second inference process information through the second intelligent agent, wherein the second inference process information indicates that the video frame to be determined is an abnormal video frame; Inputting the video frame to be determined, the type of abnormality to be determined, the first inference process information, and the second inference process information into a third intelligent agent, and obtaining conclusion information on whether the video frame to be determined is an abnormal video frame through the third intelligent agent.

2. The method according to claim 1, further comprising: Inputting the second inference process information into the first intelligent agent, and obtaining first feedback information through the first intelligent agent; Inputting the first inference process information into the second intelligent agent, and obtaining second feedback information through the second intelligent agent, wherein the first feedback information and the second feedback information represent the same or opposite conclusions, obtaining the conclusion information on whether the video frame to be determined is an abnormal video frame through the third intelligent agent includes: Inputting the video frame to be determined, the type of abnormality to be determined, the first inference process information, the second inference process information, the first feedback information, and the second feedback information into the third intelligent agent, and obtaining the conclusion information on whether the video frame to be determined is an abnormal video frame through the third intelligent agent.

3. The method according to claim 2, further comprising: Inputting the second feedback information into the first intelligent agent, and obtaining third feedback information through the first intelligent agent; Inputting the first feedback information into the second intelligent agent, and obtaining fourth feedback information through the second intelligent agent, wherein the third feedback information and the fourth feedback information represent the same or opposite conclusions, obtaining the conclusion information on whether the video frame to be determined is an abnormal video frame through the third intelligent agent includes: Inputting the video frame to be determined, the type of abnormality to be determined, the first inference process information, the second inference process information, the first feedback information, the second feedback information, the third feedback information, and the fourth feedback information into the third intelligent agent, and obtaining the conclusion information on whether the video frame to be determined is an abnormal video frame through the third intelligent agent.

4. The method according to any one of claims 1 to 3, wherein, Before obtaining the first inference process information, the method further comprises: Inputting a first role definition prompt into the first intelligent agent, so that the first intelligent agent can output the first inference process information indicating that the video frame to be determined is not an abnormal video frame, Before obtaining the second inference process information, the method further comprises: Input the second role definition prompt into the second intelligent agent so that the second intelligent agent can output the second inference process information indicating that the video frame to be determined is an abnormal video frame.

5. The method according to claim 1, further comprising: Constructing a corpus for training the third intelligent agent; And Training the third intelligent agent using the corpus.

6. The method according to claim 5, wherein, Constructing a corpus for training the third intelligent agent includes: Constructing a normal image library, where the normal image library includes normal images misreported as abnormal images; Constructing an abnormal image library, where the abnormal image library includes abnormal images determined to be abnormal; For each image in the normal image library and the abnormal image library, obtaining first debate information through the first intelligent agent and obtaining second debate information through the second intelligent agent; Constructing the corpus based on the normal image library, the abnormal image library, and the first debate information and the second debate information corresponding to each image in the normal image library and the abnormal image library.

7. The method according to claim 6, wherein Constructing the normal image library includes: Obtaining normal image resources; Obtaining the images detected as abnormal by a preset image detection algorithm in the normal image resources and the corresponding abnormal types of the detected abnormal images; Counting the number of images detected as abnormal corresponding to each type of abnormal type; In the case where the number of images detected as abnormal corresponding to a specific abnormal type is less than the threshold, repeatedly execute the steps of obtaining normal image resources and obtaining the images detected as abnormal in the normal image resources and the corresponding abnormal types of the detected abnormal images until the number of images detected as abnormal corresponding to all abnormal types is greater than or equal to the threshold; Adding each of all the images detected as abnormal to the normal image library and marking it with the abnormal type corresponding to the detected abnormal image.

8. The method according to claim 7, wherein Constructing the abnormal image library includes: Converting normal images into images of various abnormal types through an abnormal conversion algorithm; Obtaining the images detected as abnormal by a preset image detection algorithm in the converted images and the corresponding abnormal types of the detected abnormal images; Counting the number of images detected as abnormal corresponding to each type of abnormal type; In the case where the number of images detected as abnormal corresponding to a specific abnormal type is less than the threshold, repeatedly execute the steps of converting normal images into images of various abnormal types through the abnormal conversion algorithm and obtaining the images detected as abnormal in the converted images and the corresponding abnormal types of the detected abnormal images until the number of images detected as abnormal corresponding to all abnormal types is greater than or equal to the threshold; Adding each of all the images detected as abnormal to the abnormal image library and marking it with the abnormal type corresponding to the detected abnormal image.

9. The method according to claim 7 or 8, wherein For each image in the normal image library and the abnormal image library, obtaining first debate information through the first intelligent agent and obtaining second debate information through the second intelligent agent includes: Input the image and the corresponding abnormal type of the image into the first intelligent agent, and obtain third inference process information through the first intelligent agent, where the third inference process information indicates that the image is not an abnormal image corresponding to the abnormal type; Input the image and the corresponding abnormal type of the image into the second intelligent agent, and obtain fourth inference process information through the second intelligent agent, where the fourth inference process information indicates that the image is an abnormal image corresponding to the abnormal type; Input the fourth inference process information into the first intelligent agent, and obtain fifth feedback information through the first intelligent agent; Input the third inference process information into the second intelligent agent, and obtain sixth feedback information through the second intelligent agent; Input the sixth feedback information into the first intelligent agent, and obtain seventh feedback information through the first intelligent agent; Input the fifth feedback information into the second intelligent agent, and obtain eighth feedback information through the second intelligent agent, where the first debate information includes the third inference process information, the fifth feedback information, and the seventh feedback information, and the second debate information includes the fourth inference process information, the sixth feedback information, and the eighth feedback information.

10. The method according to claim 6, wherein, Training the third intelligent agent using the corpus includes: Training the third intelligent agent using the corpus by means of supervised fine-tuning or reinforcement fine-tuning.

11. A system for determining abnormal video frames, comprising: A first intelligent agent configured to obtain first inference process information for a video frame to be determined and a corresponding abnormal type to be determined for the video frame to be determined, where the first inference process information indicates that the video frame to be determined is not an abnormal video frame; A second intelligent agent configured to obtain second inference process information for the video frame to be determined and the abnormal type to be determined, where the second inference process information indicates that the video frame to be determined is an abnormal video frame; A third intelligent agent configured to obtain conclusion information on whether the video frame to be determined is an abnormal video frame based on the video frame to be determined, the abnormal type to be determined, the first inference process information, and the second inference process information.

12. An electronic device, comprising a memory and a processor, The memory stores a computer program executable by the processor, When the computer program is executed by the processor, the processor is caused to execute the method for determining abnormal video frames according to any one of claims 1 to 10.

13. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, the processor is caused to execute the method for determining abnormal video frames according to any one of claims 1 to 10.

14. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, the processor is caused to execute the method for determining abnormal video frames according to any one of claims 1 to 10.