A Multi-Agent-Based Method and System for Short Video Content Classification and Analysis
Patent Information
- Application Number
- CN202610902371.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-06-23
AI Technical Summary
[0003]现有主流技术方案主要包括三类:一是基于目标检测的流水线方法,如YOLO系列模型,通过对视频逐帧或隔帧进行目标检测并聚合结果,该方法依赖大量标注数据,难以扩展至新类别,且无法理解深层语义;二是基于多模态大模型(如CLIP)的零样本分类方法,通过视觉-语言预训练模型实现零样本识别,但该类模型对图像局部细节感知能力弱,且难以有效处理长视频中的时序信息;三是音频分析的“切片-聚合”方案,将音频固定长度切片后独立分类并累加置信度,该方法存在固有缺陷:短暂关键事件易被持续背景噪声淹没,且缺乏上下文信息导致声学混淆
通过视频感知智能体并行提取全局视觉特征与基于目标检测的局部视觉特征并按时序融合,同时音频感知智能体采用动态事件检测与上下文扩展策略提取背景段与事件段的声学特征并通过注意力机制融合,有效解决了传统固定切片导致短暂关键事件被背景噪声淹没的问题,显著提升了音频事件检测的时序完整性与鲁棒性;通过推理智能体对联合视觉特征与最终音频特征进行跨模态融合,并结合知识管理智能体从动态知识库中检索候选类别描述生成初步分类结果,实现了视觉与音频信息的高效语义对齐与互补,大幅降低了单模态误判风险;进一步通过评估智能体对初步分类结果进行多模态一致性审查,并根据审查结果自适应确定最终分类结果,能够在无需人工干预的条件下自动识别并规避多模态之间的矛盾,显著提高了分类结果的可靠性、准确性与可解释性。该方法实现了对短视频内容的高鲁棒性、高时效性分类与分析。
Smart Images

Figure CN122435514B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a method and system for classifying and analyzing short video content based on multiple agents. Background Technology
[0002] With the rapid development of short video platforms, the rapid classification and analysis of massive amounts of internet short video content has significant application value in areas such as content review, public opinion monitoring, and security monitoring.
[0003] Current mainstream technical solutions mainly fall into three categories: First, pipelined methods based on object detection, such as the YOLO series models, which perform object detection frame by frame or every other frame in a video and aggregate the results. This method relies on a large amount of labeled data, is difficult to extend to new categories, and cannot understand deep semantics. Second, zero-shot classification methods based on multimodal large models (such as CLIP), which achieve zero-shot recognition through vision-language pre-trained models. However, these models have weak perception of local image details and are difficult to effectively process temporal information in long videos. Third, the "slice-aggregate" approach for audio analysis, which slices audio into fixed-length segments, classifies them independently, and accumulates confidence scores. This method has inherent defects: brief critical events are easily drowned out by continuous background noise, and the lack of contextual information leads to acoustic confusion. In addition, existing technologies generally suffer from limitations in single-modal information (ineffective fusion of vision and audio), lack of temporal context, high cost of long-tail category annotation, lack of interpretability, and difficulty in balancing processing efficiency and accuracy.
[0004] Therefore, there is an urgent need for a rapid classification and analysis method for short video content that can integrate visual and audio multimodal information, preserve the temporal context of key events, support zero-shot expansion, and possess high robustness and interpretability. Summary of the Invention
[0005] The problem solved by this invention is one or more of the aforementioned related technical problems.
[0006] To address the aforementioned problems, this invention provides a method and system for classifying and analyzing short video content based on multiple agents.
[0007] In a first aspect, the present invention provides a method for classifying and analyzing short video content based on multiple agents, including: Acquire short video data and separate the short video data into video streams and audio streams; The video stream is subjected to global feature extraction by a video perception agent to obtain global visual features. Local regions in the video stream are detected and cropped based on a preset target detection model, and local visual features are extracted. The global visual features and the local visual features are fused in a temporal sequence to generate joint visual features. The audio stream is segmented based on dynamic event detection by an audio-aware agent to obtain background segments and event segments. The event segments are then extended with context to generate extended event segments. The acoustic features of the background segments and the extended event segments are extracted and fused through an attention mechanism to generate the final audio features based on the acoustic features of the background segments and the extended event segments. The reasoning agent performs cross-modal fusion of the joint visual features and the final audio features to generate a multimodal joint representation; the knowledge management agent retrieves the most relevant candidate category descriptions from a preset dynamic knowledge base based on the multimodal joint representation; and based on the reasoning agent, a preliminary classification result is generated according to the multimodal joint representation and the candidate category descriptions. The preliminary classification results are evaluated by a multimodal consistency review of the agent to obtain the review results, and the final classification results are determined based on the review results.
[0008] Optionally, the step of fusing the global visual features and the local visual features sequentially to generate joint visual features includes: The global visual features at the same time point are weighted and fused with at least one local visual feature to generate a single-frame enhanced visual feature; wherein the weights of the weighted fusion are dynamically adjusted by the inference agent based on context information. The multiple single-frame enhanced visual features at different time points are arranged in temporal order to form a visual feature sequence; The visual feature sequence is superimposed with the temporal position code, and the superimposed feature sequence is then subjected to temporal modeling to obtain the joint visual feature.
[0009] Optionally, the step of segmenting the audio stream based on dynamic event detection using an audio-aware agent to obtain background segments and event segments, and then performing contextual expansion on the event segments to generate expanded event segments, includes: A short-time energy sequence is calculated from the audio stream in chronological order. The short-time energy sequence contains multiple short-time energy values, and each short-time energy value corresponds to a time point. If the short-time energy value is greater than the corresponding adaptive dynamic threshold, then the corresponding time point is marked as an event candidate point; otherwise, it is marked as a background candidate point. At least one event segment is divided based on all the event candidate points, and at least one background segment is divided based on all the background candidate points; A buffer window of preset duration is added before the start time and after the end time of the event segment to obtain the extended event segment.
[0010] Optionally, by using an attention mechanism to fuse the acoustic features of the background segment and the expanded event segment, a final audio feature is generated, including: The background segment and the extended event segment are converted into Mel spectrograms respectively, and then input into a preset acoustic feature extraction model to obtain the acoustic features of the background segment and the acoustic features of the extended event segment. The acoustic features of the extended event segment are input into a preset long short-term memory network for temporal encoding to obtain a hidden state sequence. Using the acoustic features of the background segment as the query vector and the hidden state sequence as the key vector and value vector, dot product attention calculation is performed to obtain the focused event features; The acoustic features of the background segment and the focused event features are fused together using residual connections and layer normalization to generate the final audio features.
[0011] Optionally, the joint visual features and the final audio features are fused across modally by an inference agent to generate a multimodal joint representation, including: Based on a bidirectional parallel cross-attention mechanism, the joint visual features are used as the query vector, and the final audio features are used as the key vector and value vector to calculate the output data of visual-guided audio attention; the final audio features are used as the query vector, and the joint visual features are used as the key vector and value vector to calculate the output data of audio-guided visual attention. The output data of the visually guided audio attention and the output data of the audio-guided visual attention are concatenated and fused to generate the multimodal joint representation.
[0012] Optionally, the review results include those with contradictions and those without; the review results obtained by evaluating the agent to perform multimodal consistency review on the preliminary classification results include: Obtain the predicted event category for each of the extended event segments, and the predicted local category for each sampled keyframe in the video stream; The highest confidence category in the preliminary classification results is determined, and the number of conflicting keyframes is determined based on the local categories and the highest confidence category; the number of conflicting event segments is determined based on the event categories and the highest confidence category. A comprehensive conflict score is obtained based on the number of conflict keyframes, the number of conflict event segments, the total number of keyframes, and the total number of expanded event segments. If the overall conflict score exceeds a preset threshold, a conflict is determined to exist; otherwise, no conflict is determined to exist.
[0013] Optionally, determining the final classification result based on the review result includes: When the review results are contradictory, a feedback adjustment loop is triggered; The feedback adjustment loop includes: The evaluation agent generates conflict information, which includes the timestamp of at least one sampled keyframe in the video stream where the conflict occurs, or the time interval of at least one of the extended event segments. Based on the reasoning agent, an adjustment instruction is generated according to the contradictory information, and the adjustment instruction includes at least one of the following operations: If the contradictory information originates from a conflict between the global category and the local category, then the weighted fusion weight of the global visual features and the local visual features in the video perception agent is adjusted, and the weight of the conflicting feature is reduced by a preset first adjustment value. If the contradictory information originates from a conflict between visual perception and audio perception, a mask matrix is introduced in the self-attention calculation of the temporal Transformer to reset the attention weight of the specified keyframe to zero, so that it does not participate in feature aggregation; or a mask is introduced in the audio attention fusion stage to reset the attention weight of the specified audio event segment features to zero. If the contradictory information originates from an abnormal target detection confidence level, then the detection confidence threshold of the target detection model in the video perception agent is adjusted by raising or lowering the threshold by a preset second adjustment value. In response to the adjustment command, the video sensing agent and / or the audio sensing agent perform corresponding parameter adjustments and re-execute the feature extraction and fusion steps to generate updated joint visual features and / or updated final audio features. The reasoning agent re-executes the cross-modal fusion, knowledge retrieval, and preliminary classification generation steps based on the updated joint visual features and / or the updated final audio features, and the evaluation agent performs a consistency review again. Repeat the above process until the review result is consistent with the previous result or the preset maximum number of iterations is reached, and then output the final classification result.
[0014] Optionally, the construction process of the dynamic knowledge base includes: Retrieve categories defined in natural language; Based on a pre-defined large language model, multimodal text descriptions are generated for the category under pre-defined output constraints; The multimodal text description is converted into a category description feature vector, and the dynamic knowledge base is constructed based on the category description feature vector; When the category confidence obtained by the reasoning agent matching the multimodal joint representation with the category description feature vector in the dynamic knowledge base is lower than a preset threshold, the agent calls an external knowledge source to obtain the latest information through retrieval enhancement generation technology, and generates or updates the category description feature vector. Furthermore, based on a preset time period, the category description feature vectors in the dynamic knowledge base are subjected to retrieval enhancement and generation updates to refresh outdated category descriptions.
[0015] Optionally, the multimodal consistency review of the preliminary classification results by evaluating the agent further includes a semantic rationality review; the process of the semantic rationality review includes: Text information is extracted from the keyframes corresponding to the category with the highest confidence in the preliminary classification results using an optical character recognition model; Based on the large language model, determine whether there is a logical conflict between the text information and the category of the preliminary classification result; If a conflict exists, the confidence score of the category is reduced, and the reason for the conflict is output.
[0016] Secondly, the present invention provides a short video content classification and analysis system based on multi-agent systems, including: An acquisition and segmentation unit is used to acquire short video data and separate the short video data into video streams and audio streams; The video processing unit is used to extract global features from the video stream using a video perception agent to obtain global visual features; to detect and crop local regions in the video stream based on a preset target detection model and extract local visual features; and to fuse the global visual features and the local visual features in a temporal sequence to generate joint visual features. An audio processing unit is configured to segment the audio stream based on dynamic event detection using an audio-sensing agent to obtain a background segment and an event segment, and to perform contextual expansion on the event segment to generate an expanded event segment. The unit also extracts the acoustic features of the background segment and the expanded event segment respectively, and fuses them through an attention mechanism to generate the final audio features based on the acoustic features of the background segment and the expanded event segment. The preliminary classification unit is used to perform cross-modal fusion of the joint visual features and the final audio features through an inference agent to generate a multimodal joint representation; to retrieve the most relevant candidate category description from a preset dynamic knowledge base based on the multimodal joint representation through a knowledge management agent; and to generate a preliminary classification result based on the inference agent, the multimodal joint representation, and the candidate category description. The final classification unit is used to perform multimodal consistency review on the preliminary classification results by the evaluation agent, obtain the review results, and determine the final classification results based on the review results.
[0017] The beneficial effects of the multi-agent-based short video content classification and analysis method and system of the present invention are: This method employs a video perception agent to extract global visual features and local visual features based on object detection in parallel and fuse them temporally. Simultaneously, an audio perception agent uses dynamic event detection and context expansion strategies to extract acoustic features from background and event segments and fuses them through an attention mechanism. This effectively solves the problem of brief critical events being drowned out by background noise in traditional fixed-slice methods, significantly improving the temporal integrity and robustness of audio event detection. An inference agent performs cross-modal fusion of joint visual features and final audio features, and a knowledge management agent retrieves candidate category descriptions from a dynamic knowledge base to generate preliminary classification results. This achieves efficient semantic alignment and complementarity between visual and audio information, significantly reducing the risk of single-modal misclassification. Furthermore, an evaluation agent performs multimodal consistency review on the preliminary classification results and adaptively determines the final classification result based on the review results. This method can automatically identify and avoid contradictions between multimodalities without human intervention, significantly improving the reliability, accuracy, and interpretability of the classification results. This approach achieves highly robust and timely classification and analysis of short video content. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a short video content classification and analysis method based on multi-agent technology according to an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the process of a short video content classification and analysis method based on multi-agent technology according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a short video content classification and analysis system based on multiple agents according to an embodiment of the present invention. Detailed Implementation
[0019] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0020] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0021] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0022] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0023] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0024] like Figure 1 As shown in the figure, an embodiment of the present invention provides a method for classifying and analyzing short video content based on multiple agents, including: Step S100: Acquire short video data and separate the short video data into video stream and audio stream.
[0025] Specifically, the system first receives short video data to be identified. Short video data refers to the original short video files to be analyzed, typically short in duration (e.g., a few seconds to a few minutes) containing synchronized video and audio tracks. This short video data can be user-uploaded, obtained from a network interface, or read from local storage, and its format includes, but is not limited to, common container formats such as MP4, AVI, and MOV. For example, the system receives a 15-second MP4 short video file from a content moderation platform.
[0026] Subsequently, the short video data undergoes demultiplexing to separate the video and audio tracks encapsulated together in the container, resulting in independent video and audio streams. The video stream refers to a sequence of undecoded or decoded video frames, containing consecutive image frames and their timestamps, used for subsequent visual feature extraction. The audio stream refers to a sequence of undecoded or decoded audio samples, containing consecutive audio samples and their timestamps, used for subsequent acoustic feature extraction. This separation operation can be implemented using existing multimedia processing tools (such as FFmpeg) or corresponding demultiplexers. The separated video and audio streams will serve as the foundational input data for subsequent parallel perceptual processing.
[0027] Step S100 achieves rapid decoupling of short video data, efficiently separating the mixed audiovisual information into independent video and audio streams, providing a data foundation for subsequent multimodal parallel processing. At the same time, the separated video and audio streams retain complete timestamp information, ensuring spatiotemporal alignment in the subsequent feature extraction and fusion process, avoiding temporal misalignment or information loss caused by mixed processing, thus laying a data foundation for the efficiency and accuracy of the entire classification and analysis system.
[0028] Step S200: The video perception agent extracts global features from the video stream to obtain global visual features. Based on a preset target detection model, the agent detects and crops local regions in the video stream and extracts local visual features. The global visual features and the local visual features are then fused sequentially to generate joint visual features.
[0029] Specifically, the video perception agent is an agent module responsible for processing video data, which has functions such as global feature extraction, local region detection and feature extraction, and time-series fusion.
[0030] The video-aware agent is invoked to execute the following two processing paths in parallel: Global Feature Extraction Path: Keyframe sampling is performed on the video stream obtained in step S100 (e.g., uniform sampling of 1 frame per second or adaptive sampling based on scene switching) to obtain several keyframe images. Each keyframe image is completely input into a pre-defined visual encoder (e.g., the visual end of a contrastive learning-based visual-language pre-trained model), and a high-dimensional vector is encoded and output as the global visual feature of that keyframe. This global visual feature represents the overall scene layout, distribution of major objects, and macroscopic semantic information of the current frame.
[0031] Local Feature Extraction Path: For the same keyframe, a pre-defined lightweight object detection model (such as the YOLO series, which only outputs bounding boxes and not categories) is invoked to detect objects in the image, obtaining the bounding box coordinates of several potential regions of interest. Based on each bounding box, corresponding local image patches are cropped from the original image. These local image patches are then fed into the same visual encoder that shares weights with the global path, encoding multiple local visual feature vectors. Each local visual feature represents detailed information within that local region, such as human body movements or small objects.
[0032] Temporal fusion: For the same keyframe at the same time point, its global visual features are fused with one or more local visual features extracted from that frame (e.g., weighted summation) to form a single-frame enhanced visual feature that includes both the global scene and highlights local details. Subsequently, the single-frame enhanced visual features from different time points are arranged in chronological order according to the original video, forming a visual feature sequence. Finally, this sequence is temporally integrated (e.g., by capturing inter-frame dynamic evolution through temporal modeling methods) to generate a joint visual feature. Joint visual feature: A feature representation obtained by temporally fusing global and local visual features from multiple time points, containing both spatial details and temporal evolution information. This joint visual feature serves as the final output of the video perception agent, used by subsequent inference agents.
[0033] By employing a dual-path parallel extraction strategy of global and local features through a video-aware agent, the strong semantic understanding of the global scene by the vision-language pre-trained model is preserved, while the lack of detail perception by the large model is compensated by cropping local regions using a lightweight object detection model. This significantly improves the ability to capture key information such as small targets and local actions. At the same time, the global and local features are fused temporally, so that the generated joint visual features can effectively encode the dynamic evolution process between video frames. This provides a high-quality visual representation rich in spatial details and temporal context for subsequent cross-modal classification, thereby greatly improving the robustness of recognition of brief critical events (such as knife-wielding actions and collision moments) while ensuring processing efficiency.
[0034] Step S300: The audio stream is segmented based on dynamic event detection by an audio perception agent to obtain background segments and event segments. The event segments are then extended with context to generate extended event segments. The acoustic features of the background segments and the extended event segments are extracted respectively and fused through an attention mechanism to generate the final audio features based on the acoustic features of the background segments and the extended event segments.
[0035] Specifically, the audio perception agent is an intelligent agent module responsible for processing audio data. It has functions such as dynamic event detection, context expansion, acoustic feature extraction, and attention fusion. The audio perception agent is invoked to analyze the audio stream obtained in step S100. The dynamic event detection method is used to replace the traditional fixed-length slices, and the audio stream is adaptively divided into event segments with clear semantics and background segments with no sound or noise.
[0036] Dynamic event detection segmentation: The audio-sensing agent analyzes the acoustic features of the audio stream in real time, such as energy changes and spectral characteristics, to dynamically detect the start and end times of sound events (such as collision sounds, alarm sounds, and human voices). When a significant change in energy or spectral pattern is detected, the corresponding time interval is designated as an event segment; the remaining continuous intervals where no significant events are detected are designated as background segments. Background segments typically include ambient noise, continuous wind sounds, engine sounds, etc., while event segments include brief or continuous sounds with clear semantics (such as braking sounds, glass breaking sounds, and screams).
[0037] Context expansion: For each detected event segment, the audio-aware agent adds a buffer window of preset duration (e.g., 0.5 seconds each) before its start time and after its end time, thus forming an expanded event segment. This expansion operation ensures that the event segment includes a background context before and after it, preserving the background sound before the event and the lingering sound after the event, avoiding event information truncation or spectrum leakage caused by hard cutting.
[0038] Acoustic Feature Extraction: The audio-sensing agent extracts acoustic features for each background segment and each extended event segment. Specifically, the time-domain signal of each audio segment is converted into a frequency-domain representation (e.g., a Mel spectrogram), and then mapped to a fixed-dimensional acoustic feature vector using a pre-defined acoustic coding model (e.g., an audio pre-trained model based on self-supervised learning). The features extracted from the background segment are called "background segment acoustic features," and the features extracted from the extended event segment are called "event segment acoustic features." The background segment acoustic features characterize the ambient noise floor and continuous sound information of the scene, while the event segment acoustic features characterize discriminative information such as timbre, intensity, and spectral structure of specific sound events.
[0039] Attention-based fusion: To generate final audio features that highlight key events while taking into account the background context, the audio perception agent employs an attention mechanism to weightedly fuse the acoustic features of the background segment and the event segment. The attention mechanism dynamically calculates the importance of the association between each event segment and the overall background, giving higher weights to acoustic information more valuable to the current classification task, while suppressing redundant or interfering information. The final audio features output after fusion retain the global acoustic environment of the background segment, enhance the local key semantics of the event segment, and the temporal relationship between event segments is indirectly reflected through the attention weights.
[0040] Dynamic event detection involves real-time analysis of the audio stream and adaptive segmentation of event segments and background segments based on energy or spectral changes, rather than fixed slices. Background segments are continuous intervals in the audio stream where no significant semantic events are detected, typically representing ambient noise, persistent interference, etc. Event segments are continuous intervals in the audio stream where brief or continuous sounds with clear semantics are detected, such as collision sounds, alarm sounds, and human voices. Extended event segments are new audio segments obtained by contextually expanding the original event segments, including the event itself and its surrounding information.
[0041] By employing a segmentation strategy based on dynamic event detection, the audio-sensing agent breaks through the limitations of traditional fixed-length slice isolation classification, enabling the complete capture of brief, critical acoustic events (such as a 0.2-second sound of breaking glass or a collision) without being drowned out by background noise. Contextual expansion of event segments preserves the complete evolution of sound events from onset to decay, as well as the temporal relationships between events, effectively avoiding acoustic confusion and information truncation. Furthermore, an attention mechanism adaptively fuses background and event segment features, ensuring that the final audio features highlight critical events while also considering the acoustic environment of the scene. This significantly improves the robustness and accuracy of audio event detection, providing a high-quality audio representation foundation for subsequent cross-modal alignment with visual features.
[0042] Step S400: The reasoning agent performs cross-modal fusion of the joint visual features and the final audio features to generate a multimodal joint representation; the knowledge management agent retrieves the most relevant candidate category description from a preset dynamic knowledge base based on the multimodal joint representation; and based on the reasoning agent, a preliminary classification result is generated according to the multimodal joint representation and the candidate category description.
[0043] Specifically, the reasoning agent is invoked as the fusion and decision-making hub to perform cross-modal fusion and preliminary classification.
[0044] Cross-modal fusion: The inference agent receives the joint visual features generated in step S200 and the final audio features generated in step S300. Since these two features originate from two heterogeneous modalities, visual and audio, respectively, their semantic spaces differ. The inference agent maps both features into a unified joint semantic space through cross-modal fusion operations (such as attention-based interaction or feature alignment), generating a comprehensive vector that integrates audiovisual bimodal information, i.e., a multimodal joint representation. This representation simultaneously includes visual information such as scenes, objects, and actions in the video, as well as corresponding auditory information such as sound events and ambient sounds, forming complementary audiovisual joint semantic features.
[0045] Dynamic Knowledge Base Retrieval: The reasoning agent sends the generated multimodal joint representation to the knowledge management agent and requests the retrieval of the most relevant category descriptions. The knowledge management agent maintains a dynamic knowledge base that stores category description feature vectors for multiple natural language categories (e.g., semantic descriptions of categories such as "traffic accident," "knife-wielding behavior," and "fire alarm," which have been pre-converted into vector form). The knowledge management agent calculates the similarity (e.g., cosine similarity) between the multimodal joint representation and the category description feature vectors in the knowledge base, selects the categories with the highest similarity as candidate categories, and returns the corresponding candidate category descriptions (including category name, natural language description, feature vectors, etc.) to the reasoning agent.
[0046] Generating preliminary classification results: After receiving the candidate category descriptions, the inference agent combines the original multimodal joint representation to further calculate the matching confidence score between the current video content and each candidate category (e.g., through similarity weighting or a classification network). Based on the confidence scores, the top N categories and their confidence scores are output as the preliminary classification results, sorted from highest to lowest. This result will serve as input for subsequent consistency review by the evaluation agent.
[0047] By using an inference agent to perform cross-modal fusion of visual and audio features, the generated multimodal joint representation can fully utilize complementary audiovisual information and effectively overcome semantic ambiguity in single-modality scenarios (for example, visual perception alone cannot distinguish between "real fire" and "fire drill," but combining audio alarms or instructions can improve the accuracy of the judgment). At the same time, by using a knowledge management agent to retrieve candidate category descriptions from a dynamic knowledge base, true zero-shot classification is achieved, eliminating the need to retrain the model for each new category and significantly reducing the annotation cost and iteration cycle for business expansion. The final preliminary classification results provide a foundation for subsequent consistency reviews and retain an interpretable decision path for the feedback adjustment mechanism.
[0048] Step S500: The preliminary classification results are reviewed by the evaluation agent to obtain the review results, and the final classification results are determined based on the review results.
[0049] Specifically, the evaluation agent is invoked to perform a multimodal consistency review on the preliminary classification results generated in step S400. The evaluation agent is responsible for verifying the rationality and self-consistency of the preliminary classification results from multiple perspectives.
[0050] Multimodal consistency review: The evaluation agent obtains preliminary classification results (including the top N candidate categories and their confidence scores), while simultaneously acquiring intermediate analysis cues (e.g., local categories of keyframes, predicted categories of event segments, etc.) from the video and audio perception agents. Based on this information, the evaluation agent performs a consistency review on the preliminary classification results. For example, the consistency review of the preliminary classification results includes: Visual-audio conflict detection: Determines whether there is a logical contradiction between the scene and actions presented in the video and the sound events identified in the audio. For example, if the visual classification is "traffic accident" but the main audio is cheerful music, it is determined to be a contradiction.
[0051] Local-Global Conflict Detection: Determine whether there is an unreasonable conflict between the local details of the keyframe (such as "holding a knife" detected in a local area) and the global scene (such as "classroom"). If there is a conflict, mark it as an anomaly.
[0052] Semantic rationality review: Using a language model, it is determined whether the classification result is consistent with the extracted text information (such as the screen text recognized by OCR). For example, if the classification result is "fire" but the screen OCR recognizes the text "fire drill", it is considered a contradiction.
[0053] Based on the review results, the evaluation agent outputs a review conclusion: if the preliminary classification result passes the review of each dimension (i.e., no significant contradictions are found), the preliminary classification result is deemed credible and directly output as the final classification result; if contradictions are found (e.g., visual-audio mismatch, local-global inconsistency), the evaluation agent will mark the contradiction and decide whether to trigger subsequent processing (such as feedback adjustment) or directly output a final result with a warning according to a preset strategy. Throughout the process, the evaluation agent ensures the reliability, interpretability, and logical consistency of the output results.
[0054] The final classification results are usually presented in a structured data format (such as JSON), and typically include the following core components: Overall category tags: One or more category tags to which the short video belongs (such as "traffic accident", "fire", "knife behavior", etc.), each tag is accompanied by a corresponding confidence score (between 0 and 1, reflecting the credibility of the prediction).
[0055] The highest confidence category: the category considered most likely and its confidence level, which is the primary classification conclusion.
[0056] Spatiotemporal positioning information: start and end timestamps of key events in the video (e.g., collision sound from 2.3 seconds to 2.8 seconds, knife-wielding action from 5 seconds to 7 seconds), and optional target bounding box coordinates (e.g., the position of people, vehicles).
[0057] Reasoning process recording (optional): Interactions and question-and-answer sessions between agents, conflict detection and correction processes, used to enhance interpretability.
[0058] Semantic description (optional): A short video content summary generated in natural language, such as "The video shows a traffic accident, followed by a verbal argument."
[0059] By evaluating the agent's independent multimodal consistency review of the preliminary classification results generated by the reasoning agent, misjudgments caused by unimodal noise, semantic ambiguity, or model illusion can be effectively identified and filtered, significantly improving the accuracy and robustness of the final classification results. At the same time, the review process can output specific contradiction types and location information, providing the system with interpretable decision-making basis, enhancing users' trust in the classification results, and providing a clear traceability path for possible subsequent manual review or model debugging.
[0060] In this embodiment, a video perception agent extracts global visual features and local visual features based on object detection in parallel and fuses them temporally. Simultaneously, an audio perception agent employs dynamic event detection and context expansion strategies to extract acoustic features from background and event segments and fuses them using an attention mechanism. This effectively solves the problem of brief critical events being submerged by background noise due to traditional fixed-slice methods, significantly improving the temporal integrity and robustness of audio event detection. An inference agent performs cross-modal fusion of joint visual features and final audio features, and a knowledge management agent retrieves candidate category descriptions from a dynamic knowledge base to generate preliminary classification results. This achieves efficient semantic alignment and complementarity between visual and audio information, significantly reducing the risk of single-modal misjudgment. Furthermore, an evaluation agent performs multimodal consistency review on the preliminary classification results and adaptively determines the final classification result based on the review results. This method can automatically identify and avoid contradictions between multimodalities without human intervention, significantly improving the reliability, accuracy, and interpretability of the classification results. This method achieves highly robust and timely classification and analysis of short video content.
[0061] Optionally, the step of fusing the global visual features and the local visual features sequentially to generate joint visual features includes: The global visual features at the same time point are weighted and fused with at least one local visual feature to generate a single-frame enhanced visual feature; wherein the weights of the weighted fusion are dynamically adjusted by the inference agent based on context information. The multiple single-frame enhanced visual features at different time points are arranged in temporal order to form a visual feature sequence; The visual feature sequence is superimposed with the temporal position code, and the superimposed feature sequence is then subjected to temporal modeling to obtain the joint visual feature.
[0062] Specifically, weighted fusion generates enhanced visual features for a single frame: for the same keyframe, the video perception agent acquires the global visual feature vector of that frame (denoted as...). ) and detected from that frame The local visual feature vector corresponding to each local region (denoted as ) To integrate global and local information, the system employs a weighted summation method: first, each local feature vector is multiplied by its corresponding weight. Then sum them up and multiply by a dynamic adjustment factor. Simultaneously multiply the global features by The final single-frame enhanced visual features The calculation formula is: ; in, and All adjustments are dynamically made by the inference agent based on the contextual information of the current video (such as scene complexity and the confidence level of local targets). For example, when local details in the image are crucial to classification (such as knife-wielding actions in dim lighting), the inference agent can increase... The value is increased to enhance the contribution of local features; when the global scene is more discriminative, the value is decreased. Value. Through dynamic weighted fusion, the enhanced visual features of each keyframe preserve both the macroscopic scene semantics and highlight key local details.
[0063] Constructing a visual feature sequence: Multiple enhanced visual features from single frames at different time points (e.g., second 0, second 1, ..., second T-1) in the same video stream are arranged in their original chronological order to form a visual feature sequence. This sequence preserves the sequential relationship between frames, providing a foundation for subsequent temporal modeling.
[0064] Temporal positional encoding and temporal modeling: To enable the model to perceive the evolutionary order of actions or events, the system injects a learnable temporal positional encoding (denoted as ) into each position in the visual feature sequence. The feature sequence superimposed with position encoding is input into a temporal modeling module (e.g., a Transformer-based encoder or an LSTM network), which captures the dependencies between frames, such as visual changes before and after a collision. After temporal modeling, the final joint visual features are output. This joint visual feature not only contains spatial information for each frame (global scene + local details), but also encodes the motion trajectory and state evolution of objects and people on the timeline, such as the complete process from "car driving normally" to "collision occurring" and then to "pedestrian falling to the ground".
[0065] By employing dynamically weighted global-local feature fusion for each keyframe, the contribution ratio of global and local information can be adaptively adjusted according to the video content, avoiding suboptimal fusion with fixed weights in complex scenes and significantly improving the ability to capture key details such as small targets and local actions. At the same time, by introducing temporal position coding and temporal modeling, the enhanced features of a single frame are organized into a sequence and the dynamic evolution between frames is captured. This makes the joint visual features not only contain rich spatial details but also contain complete temporal causal relationships, thus effectively solving the problem of insufficient recognition of dynamic events (such as collisions and knife swings) caused by traditional methods that rely solely on single frames or simple splicing. This provides a high-quality visual representation with both fine spatial perception and continuous temporal perception for subsequent cross-modal classification.
[0066] Optionally, the step of segmenting the audio stream based on dynamic event detection using an audio-aware agent to obtain background segments and event segments, and then performing contextual expansion on the event segments to generate expanded event segments, includes: A short-time energy sequence is calculated from the audio stream in chronological order. The short-time energy sequence contains multiple short-time energy values, and each short-time energy value corresponds to a time point. If the short-time energy value is greater than the corresponding adaptive dynamic threshold, then the corresponding time point is marked as an event candidate point; otherwise, it is marked as a background candidate point. At least one event segment is divided based on all the event candidate points, and at least one background segment is divided based on all the background candidate points; A buffer window of preset duration is added before the start time and after the end time of the event segment to obtain the extended event segment.
[0067] Specifically, short-time energy sequence calculation: The audio-sensing agent divides the input audio stream into frames according to a fixed time window (e.g., frame length 25ms, frame shift 10ms), and calculates the short-time energy (STE) for each frame. Let the audio signal of the j-th frame be... (k is the sampling point index), then the short-time energy of this frame. The calculation formula is: ; in, This represents the number of sampling points corresponding to the frame length. The calculation is performed sequentially for all frames to obtain the short-time energy sequence that varies with frame number j. Each It corresponds to a point in time (the center moment of the frame).
[0068] Adaptive Dynamic Threshold Comparison and Candidate Point Labeling: Instead of using a fixed energy threshold, the system tracks the noise floor and peak energy of the audio energy envelope in real time and dynamically calculates the adaptive threshold at each time point. The calculation formula is as follows: ; in, and For the preset scaling factor (e.g.) =1.5, =0.3), and This is obtained by tracking the recent sliding minimum and maximum energy values. The short-time energy value E[t] at each time point t is compared with the corresponding adaptive threshold. Comparison: If If the event occurs, the time point is marked as an event candidate point (representing a possible sudden acoustic event); otherwise, it is marked as a background candidate point (representing ambient noise or stable sound).
[0069] Background and event segments are divided: Based on consecutive time markers, adjacent and consecutive event candidate points are merged into one event segment, and adjacent and consecutive background candidate points are merged into one background segment. Thus, the original audio stream is segmented into multiple irregularly sized background and event segments, each with a clearly defined start and end time. This segmentation method accurately preserves the complete intervals of brief, sudden events, avoiding the disruption of event boundaries caused by fixed-length slices.
[0070] Context expansion generates expanded event segments: For each event segment, to preserve the acoustic context before and after the event (e.g., braking sound before a collision, echo after a collision), the system adds a buffer window of preset duration (e.g., 0.5 seconds before and after) before and after its start time, forming the expanded event segment. The expanded segment includes the original event body + preceding background + subsequent background, ensuring the integrity of the acoustic evolution process. The original background segment is not expanded and is directly retained.
[0071] By employing an event detection method based on short-time energy and adaptive dynamic thresholds through an audio-sensing agent, the detection sensitivity can be adjusted in real time according to the ambient noise level. Even in noisy backgrounds, it can accurately capture brief, sudden events such as collisions and screams, effectively avoiding the problems of false alarms in quiet environments and missed alarms in noisy environments caused by fixed thresholds. At the same time, through context extension operations, front and back buffer windows are added to each event segment, which fully preserves the onset-attenuation process of the acoustic event and its transition relationship with the background. This fundamentally overcomes the defects of traditional fixed-length slices, such as the truncation of key events and acoustic confusion. It provides high-quality audio units rich in temporal context for subsequent acoustic feature extraction and event classification, significantly improving the detection recall and classification accuracy of brief key audio events in short videos.
[0072] Optionally, by using an attention mechanism to fuse the acoustic features of the background segment and the expanded event segment, a final audio feature is generated, including: The background segment and the extended event segment are converted into Mel spectrograms respectively, and then input into a preset acoustic feature extraction model to obtain the acoustic features of the background segment and the acoustic features of the extended event segment. The acoustic features of the extended event segment are input into a preset long short-term memory network for temporal encoding to obtain a hidden state sequence. Using the acoustic features of the background segment as the query vector and the hidden state sequence as the key vector and value vector, dot product attention calculation is performed to obtain the focused event features; The acoustic features of the background segment and the focused event features are fused together using residual connections and layer normalization to generate the final audio features.
[0073] Specifically, acoustic feature extraction: The audio-aware agent converts each background segment and each expanded event segment into a Mel-spectrogram, and inputs them into a pre-defined acoustic feature extraction model (e.g., Wav2Vec2 or BEATs based on self-supervised learning). This model outputs a fixed-dimensional feature vector, where: Background acoustic characteristics: denoted as , representing the global acoustic context, such as environmental noise and continuous steady-state sound.
[0074] Extended acoustic characteristics of the event segments: Let the total number of event segments be... Its characteristic sequence is Each Each event segment corresponds to an expanded event segment, which encodes the core acoustic pattern of the event and its preceding and following transition information.
[0075] Long Short-Term Memory Network Temporal Encoding: Because there may be causal or temporal dependencies between event segments (e.g., "sudden braking sound" is often followed by "collision sound"), the system encodes the acoustic feature sequence of event segments. The input is fed into a Long Short-Term Memory (LSTM) network for temporal encoding. LSTM can capture long-range dependencies between event segments and output a sequence of hidden states. Each of them It not only contains information about the current event segment, but also incorporates the context of the preceding and following event segments.
[0076] Dot-product attention calculation focuses on event features: To highlight important event segments most relevant to the current acoustic environment, the system employs a dot-product attention mechanism. This mechanism incorporates background acoustic features... The query vector (Query, Q) is obtained through linear transformation, and the hidden state sequence H is transformed into the key vector (Key, K) and value vector (Value, V). Attention weights. The calculation is as follows: ; in, The dimension of the key vector is used as a scaling factor to prevent gradient vanishing. This represents the dot product of the query vector and the i-th key vector, used to calculate the relevance score between the i-th event segment and the global context. The Softmax function ensures that the sum of all weights is 1. Then, a weighted summation is performed on the value vectors to obtain the focused event features. : This feature emphasizes the event segments that are most in harmony with or stand out from the background environment (e.g., in a street background, collision sounds are given high weight, while irrelevant weak noises are given low weight).
[0077] Residual connectivity and layer normalization fusion: Finally, the acoustic features of the original background segment are merged. Features of the event after focusing The final audio features are generated by summing the residual connections and performing layer normalization. : Residual connections ensure that background information is not lost during attention focusing, while layer normalization stabilizes the feature distribution. Final audio features. It also includes the global acoustic environment and context-enhanced key event information.
[0078] By using LSTM to temporally encode the event segment sequence, the sequential dependencies and causal logic between different acoustic events are effectively captured, compensating for the lack of contextual association in isolated event features. Then, a dot product attention mechanism is used with background segment features as query vectors to dynamically focus on the most critical event segments, enabling brief but important acoustic events (such as collision sounds) to overcome the suppression of continuous background noise and gain dominance in the final audio features. Finally, by fusing background features through residual connections and layer normalization, global acoustic environment information is ensured not to be discarded, thereby generating an audio representation that takes into account both the global environment and local key events, with strong robustness and discriminative power, significantly improving the classification accuracy of short video audio events in complex acoustic scenarios.
[0079] Optionally, the joint visual features and the final audio features are fused across modally by an inference agent to generate a multimodal joint representation, including: Based on a bidirectional parallel cross-attention mechanism, the joint visual features are used as the query vector, and the final audio features are used as the key vector and value vector to calculate the output data of visual-guided audio attention; the final audio features are used as the query vector, and the joint visual features are used as the key vector and value vector to calculate the output data of audio-guided visual attention. The output data of the visually guided audio attention and the output data of the audio-guided visual attention are concatenated and fused to generate the multimodal joint representation.
[0080] Specifically, the bidirectional parallel cross-attention mechanism works as follows: the inference agent receives joint visual features (from the video-aware agent) and final audio features (from the audio-aware agent). To deeply fuse the two modalities, the system employs a bidirectional parallel cross-attention structure, executing two data streams simultaneously within the same layer: Visually Guided Audio Attention: Using Joint Visual Features As a query vector (Query, ), to the final audio characteristics As a key vector (Key, ) and value vector (Value, Cross-attention output of audio in computer vision This operation can filter out the audio segments that best match the current visual content. For example, when a "car collision" appears in the visuals, the attention mechanism will amplify the "collision sound" component in the audio. The calculation formula is as follows: ; in, The dimension of the key vector is used as a scaling factor to avoid excessive inner product that could lead to gradient saturation.
[0081] Audio-guided visual attention: in parallel, with final audio features As a query vector ( ), with combined visual features As a key vector ( ) and value vector ( ), calculate the cross-attention output of audio to vision This operation can enhance corresponding visual frames (such as flames and smoke) based on auditory events (such as "explosion sounds"). The calculation formula is as follows: ; Two data streams are computed in parallel and independently, enabling visual and audio signals to guide and enhance each other in both directions.
[0082] Concatenation and fusion to generate a multimodal joint representation: The attention outputs from two directions are concatenated along the feature dimension to obtain the fused multimodal joint representation. : This joint representation simultaneously includes visually guided audio features and audio-guided visual features, achieving deep semantic entanglement between audiovisual modalities.
[0083] By employing a bidirectional parallel cross-attention mechanism instead of traditional simple splicing or one-sided guided fusion, the inference agent enables visual and audio modalities to mutually enhance each other's key features through contextualization: vision provides spatiotemporal alignment indices for audio (e.g., locating the time point of a collision), while audio provides semantic verification for vision (e.g., confirming the consistency between auditory events and visual actions). This bidirectional interaction not only significantly improves the depth and semantic alignment accuracy of multimodal fusion but also effectively suppresses noise or ambiguity in a single modality (such as false actions without collision sounds or abnormal noises without visual evidence), thereby generating a more robust and compact multimodal joint representation, providing a high-quality semantic foundation for subsequent zero-shot classification.
[0084] Optionally, the review results include those with contradictions and those without; the review results obtained by evaluating the agent to perform multimodal consistency review on the preliminary classification results include: Obtain the predicted event category for each of the extended event segments, and the predicted local category for each sampled keyframe in the video stream; The highest confidence category in the preliminary classification results is determined, and the number of conflicting keyframes is determined based on the local categories and the highest confidence category; the number of conflicting event segments is determined based on the event categories and the highest confidence category. A comprehensive conflict score is obtained based on the number of conflict keyframes, the number of conflict event segments, the total number of keyframes, and the total number of expanded event segments. If the overall conflict score exceeds a preset threshold, a conflict is determined to exist; otherwise, no conflict is determined to exist.
[0085] Specifically, various predicted categories are obtained: the evaluation agent obtains the local category predicted for each sampled keyframe (e.g., "holding a knife," "car," "pedestrian," etc.) from the video perception agent, and the event category predicted for each extended event segment (e.g., "collision sound," "argument sound," "music," etc.) from the audio perception agent. These local categories and event categories are intermediate results obtained by the perception agents independently matching the dynamic knowledge base, reflecting modal information at different spatiotemporal granularities.
[0086] Determine the highest confidence category and the number of statistical conflicts: Evaluate the highest confidence category (denoted as ) in the preliminary classification results obtained by the agent from the inference agent's output. (e.g., "traffic accidents"). Then compare them item by item: For the local category of each keyframe, if it is related to... Semantic inconsistency (e.g., local category is "holding a knife") If the frame is a "traffic accident" and has no reasonable connection to the incident, then that frame is counted as a conflict keyframe, and the total number of conflict keyframes is recorded as follows: .
[0087] For each expanded event segment's event category, if it is related to... Semantic inconsistency (e.g., the event category is "upbeat music") If the event segment is classified as "traffic accident", then that event segment is counted as one conflict event segment, and the total number of conflict event segments is recorded as follows: .
[0088] At the same time, record the total number of keyframes. and the total number of expanded event segments .
[0089] Calculate the overall conflict score: The agent is evaluated according to the following formula. : ; The score is obtained by weighted summation of two parts: the first part measures the proportion of conflict between the visual modality and the top-level classification, and the second part measures the proportion of conflict between the audio modality and the top-level classification. The sum of the two parts means that a large number of conflicts in either modality will lead to a higher total score.
[0090] Determine if a contradiction exists: calculate the... With preset threshold (For example Compare with (=0.2). If If the initial classification results show a significant inconsistency across multiple modalities, then the initial classification results are considered credible. Otherwise, the initial classification results are considered reliable and no contradiction exists.
[0091] By evaluating the fine-grained local categories and event categories produced by the perceptual agent, and calculating a comprehensive conflict score based on a clear quantitative formula, the multimodal consistency review is transformed into a measurable numerical comparison, thus achieving automatic and objective conflict determination. This method avoids the tediousness of manually setting complex rules and can sensitively reflect the degree of consistency between the visual and audio dimensions and the top-level classification. By adjusting preset thresholds, the system's conservatism can be flexibly controlled (lower thresholds are more sensitive, higher thresholds are more lenient). Ultimately, this quantitative review mechanism effectively improves the credibility and robustness of the classification results, provides clear triggering conditions and conflict location information for subsequent possible feedback adjustments, and enhances the system's self-evaluation and self-correction capabilities.
[0092] Optionally, determining the final classification result based on the review result includes: When the review results are contradictory, a feedback adjustment loop is triggered; The feedback adjustment loop includes: The evaluation agent generates conflict information, which includes the timestamp of at least one sampled keyframe in the video stream where the conflict occurs, or the time interval of at least one of the extended event segments. Based on the reasoning agent, an adjustment instruction is generated according to the contradictory information, and the adjustment instruction includes at least one of the following operations: If the contradictory information originates from a conflict between the global category and the local category, then the weighted fusion weight of the global visual features and the local visual features in the video perception agent is adjusted, and the weight of the conflicting feature is reduced by a preset first adjustment value. If the contradictory information originates from a conflict between visual perception and audio perception, a mask matrix is introduced in the self-attention calculation of the temporal Transformer to reset the attention weight of the specified keyframe to zero, so that it does not participate in feature aggregation; or a mask is introduced in the audio attention fusion stage to reset the attention weight of the specified audio event segment features to zero. If the contradictory information originates from an abnormal target detection confidence level, then the detection confidence threshold of the target detection model in the video perception agent is adjusted by raising or lowering the threshold by a preset second adjustment value. In response to the adjustment command, the video sensing agent and / or the audio sensing agent perform corresponding parameter adjustments and re-execute the feature extraction and fusion steps to generate updated joint visual features and / or updated final audio features. The reasoning agent re-executes the cross-modal fusion, knowledge retrieval, and preliminary classification generation steps based on the updated joint visual features and / or the updated final audio features, and the evaluation agent performs a consistency review again. Repeat the above process until the review result is consistent with the previous result or the preset maximum number of iterations is reached, and then output the final classification result.
[0093] In some embodiments, when the evaluation agent's calculated overall conflict score exceeds a preset threshold (e.g., 0.2), it determines that the preliminary classification result contains multimodal contradictions, triggering a feedback adjustment loop. The evaluation agent generates contradiction information, which includes the keyframe timestamps where the contradiction occurs (e.g., a conflict between the local category "holding a knife" and the highest confidence category "traffic accident" in the 2.3-second keyframe) or the time interval of the extended event segment (e.g., a conflict between the event categories "upbeat music" and "traffic accident" in a certain event segment). The inference agent generates corresponding adjustment instructions based on the contradiction type. If the contradiction stems from a conflict between the global and local categories (e.g., the global scene is "fire" while the local detection is "flood"), the inference agent calls the Multimodal Large Language Model (MLLM) to make a judgment and then adjusts the weighted fusion weights of the global and local features: the weight of the feature corresponding to the conflict is reduced by a preset first adjustment value (e.g., reduced by 0.3). That is, if the local feature is unreliable, its weight is reduced by 0.3; otherwise, the weight of the global feature is reduced by 0.3.
[0094] If the contradiction stems from a conflict between visual perception and audio perception (e.g., a keyframe detects "traffic accident" but the audio event segment is "cheerful music"), then for the conflicting keyframe, a mask matrix is introduced in the self-attention calculation of the temporal Transformer to reset the attention weight of the frame to zero (i.e., the mask value is set to negative infinity, and the weight is 0 after Softmax), so that it does not participate in feature aggregation; or for the conflicting audio event segment, a mask is introduced in the dot product attention fusion stage of the audio background and the event segment to reset the attention weight of the event segment's features to zero, thereby logically ignoring the interfering information.
[0095] If the contradictory information originates from an abnormal target detection confidence level (e.g., false detection or missed detection due to environmental interference), then the detection confidence threshold of the target detection model in the video perception agent is adjusted by raising or lowering the threshold by a preset second adjustment value (e.g., raising it by 0.3 to filter low-confidence artifacts, or lowering it by 0.3 to recall missed targets).
[0096] The video perception agent and / or audio perception agent respond to adjustment instructions, performing corresponding parameter adjustments (such as modifying fusion weights, applying attention masks, and adjusting detection thresholds). Based on the adjusted parameters, they re-execute the feature extraction and fusion steps (without needing to re-detect the original audio and video), generating updated joint visual features and / or updated final audio features. The inference agent re-executes the cross-modal fusion, knowledge retrieval, and preliminary classification generation steps based on the updated features, and the evaluation agent performs a consistency review again. The above process is repeated until the inconsistencies are resolved or the preset iteration limit (e.g., 3 times) is reached, ultimately outputting a stable and reliable classification result.
[0097] This feedback adjustment mechanism constructs a complete "perception-questioning-adjustment-verification" closed loop, enabling differentiated and precise adjustment strategies for different types of contradictions (global-local conflict, audiovisual conflict, and abnormal detection confidence). These strategies include dynamically adjusting feature fusion weights, setting masking to zero to minimize interference attention, and adaptively modifying detection thresholds. This simulates the cognitive process of repeated deliberation by human experts, effectively overcoming the inherent problem of single-inference reasoning failing to self-correct when encountering noise, rare events, or modal conflicts. Without increasing the overhead of re-analyzing the original data, this mechanism significantly improves robustness to marginal cases and complex scenarios, while preserving the complete decision-making trajectory, enhancing the credibility and interpretability of classification results.
[0098] Optionally, the construction process of the dynamic knowledge base includes: Retrieve categories defined in natural language; Based on a pre-defined large language model, multimodal text descriptions are generated for the category under pre-defined output constraints; The multimodal text description is converted into a category description feature vector, and the dynamic knowledge base is constructed based on the category description feature vector; When the category confidence obtained by the reasoning agent matching the multimodal joint representation with the category description feature vector in the dynamic knowledge base is lower than a preset threshold, the agent calls an external knowledge source to obtain the latest information through retrieval enhancement generation technology, and generates or updates the category description feature vector. Furthermore, based on a preset time period, the category description feature vectors in the dynamic knowledge base are subjected to retrieval enhancement and generation updates to refresh outdated category descriptions.
[0099] Specifically, it allows users or administrators to define arbitrary categories using natural language without needing to label any training samples. For example, a user inputs "unauthorized drone flights" or "urban flooding caused by extreme weather," and these natural language phrases are directly received as new categories. This solves the pain point of traditional supervised learning, which requires a large amount of labeled data and cannot expand to new categories.
[0100] Large Language Model Generation of Multimodal Text Descriptions: The knowledge management agent inputs natural language categories into a pre-defined large language model (LLM, such as GPT-4, Qwen, etc.) and, under pre-defined three-dimensional output constraints, generates rich multimodal text descriptions for that category. Output constraints include: The existence of specific core entities (e.g., "a large area of water and a half-submerged car are visible in the picture"). Visual physical properties and interaction relationships (e.g., "the water surface is turbid and yellowish-brown, and pedestrians are having difficulty wading through the water"). Key acoustic features (e.g., "the audio is accompanied by a continuous sound of heavy rain or rushing water").
[0101] Through constraints, the structured descriptions generated by LLM can be precisely aligned with the perceptual capabilities of subsequent visual and audio encoders, avoiding overly abstract or hallucinatory content.
[0102] The generated multimodal text descriptions are fed into a text encoder (e.g., the text end of CLIP or SigLIP) that shares weights with the video perception agent, transforming them into fixed-dimensional category description feature vectors. These vectors reside in the same joint embedding space as the visual features extracted by the visual encoder. The knowledge management agent stores these vectors, along with their corresponding category names, natural language descriptions, and other information, in a dynamic knowledge base (typically a vector database), forming a searchable index.
[0103] During the online inference phase, when the inference agent matches the multimodal joint representation of the current video with the category description feature vector in the knowledge base, if the highest matching similarity (category confidence) is lower than a preset threshold (e.g., 0.5), it indicates that the existing knowledge base cannot adequately describe the current video content, possibly due to outdated category descriptions or a lack of up-to-date information. At this point, the knowledge management agent automatically invokes Retrieval Augmentation (RAG) technology to retrieve the latest image and text information related to the category in real time through reserved API interfaces (such as search engines, Wikipedia, and news databases). The retrieved information, along with the original description, is input into the large language model to generate an updated multimodal text description, which is then re-encoded into an updated category description feature vector and used to replace or supplement the dynamic knowledge base. This achieves online self-growth and dynamic updating of the knowledge base.
[0104] Furthermore, the knowledge base dynamically adds update timestamps to each category. Based on a preset time period (e.g., weekly or monthly), it proactively performs RAG updates on the category description feature vectors in the dynamic knowledge base. It utilizes the latest knowledge obtained from search engines, fused and refined through a large language model, to refresh outdated category descriptions (e.g., the description of "drone performance" has evolved from "multiple drones forming a pattern" to "a light show accompanied by music and lasers"). This regular proactive update mechanism ensures that the knowledge base always remains timely and accurate.
[0105] Through this dynamic knowledge base construction and update mechanism, users can define any category using natural language without any training samples. The large language model generates accurate multimodal text descriptions under three-dimensional constraints and stores them quantitatively, achieving true zero-sample classification and greatly reducing deployment costs and annotation overhead for new business scenarios. At the same time, it not only dynamically updates descriptions by retrieving the latest knowledge in real time when encountering low-confidence categories in online inference, but also actively refreshes outdated categories at preset time periods. This gives the knowledge base multiple self-growth capabilities: "definition is added to the database, low-confidence is updated, and periodic active refresh." It can continuously absorb newly emerging events, popular concepts, or domain terms, always maintaining sensitivity and accuracy to timely content, effectively overcoming the fundamental defect of traditional closed-set models that cannot adapt to dynamically changing business needs.
[0106] Optionally, the multimodal consistency review of the preliminary classification results by evaluating the agent further includes a semantic rationality review; the process of the semantic rationality review includes: Text information is extracted from the keyframes corresponding to the category with the highest confidence in the preliminary classification results using an optical character recognition model; Based on the large language model, determine whether there is a logical conflict between the text information and the category of the preliminary classification result; If a conflict exists, the confidence score of the category is reduced, and the reason for the conflict is output.
[0107] Specifically, the evaluation agent's multimodal consistency review of the preliminary classification results is a special form of semantic rationality review, which is mainly used to uncover logical conflicts between textual information in the image and the classification results.
[0108] The evaluation agent first extracts the category with the highest confidence (e.g., "fire") from the preliminary classification results and locates the keyframe most relevant to that category (usually the video frame that contributed the most to the prediction of that category). Then, it calls the Optical Character Recognition (OCR) model to perform text detection and recognition on the keyframe, extracting all text information appearing in the scene (such as slogans, signs, screen captions, banner text, etc.). For example, it can identify text such as "fire safety drill site" or "fire exercise" from the keyframe.
[0109] Large Language Model (LLM) logic conflict assessment: Extracted text information and the highest-confidence category from the initial classification results are input into a pre-defined large language model (LLM). The model is required to determine whether there is a logical conflict between the two. The criteria for judgment include common-sense knowledge and semantic consistency, for example: The fire was categorized as a "real fire," but the OCR system recognized text such as "fire drill," "exercise," and "simulation"—a logical inconsistency exists. The text is categorized as "traffic accident," but the OCR system recognizes phrases like "accident drill" and "road closure for construction"—there is a logical conflict. The category is "Celebration Activities" but the OCR identifies it as "Celebrating National Day" → No conflict.
[0110] Conflict handling: If the large language model determines that there is a logical conflict, the evaluation agent performs two operations: Lowering the confidence score for that category: This is typically done by subtracting a preset penalty value from the original confidence score (e.g., reducing it by 0.5). If the score drops below 0, it is set to 0. The penalty value can be dynamically adjusted based on the severity of the conflict, or it can be fixed at 0.5.
[0111] Output conflict reasons: The conflict explanations generated by the large language model (e.g., "The text 'fire drill' in the picture indicates that the event is a simulation exercise, not a real fire") are output as part of the review results for system recording or manual review.
[0112] If the LLM determines that there is no conflict, the confidence level is not adjusted, and the subsequent process continues.
[0113] By evaluating the semantic rationality review of the intelligent agent by combining optical character recognition with a large language model, the preliminary classification results can be logically verified using text information directly appearing in the video footage. This effectively prevents false alarms caused by visual ambiguity (such as a drill scene being misjudged as a real accident) or model illusion. At the same time, by reducing the confidence of conflict categories and outputting clear reasons, the interpretability and credibility of the classification results are enhanced, providing an intuitive basis for subsequent manual review or automated decision-making. This compensates for the shortcomings of pure visual-audio multimodal fusion in perceiving highly semantic textual cues, thereby significantly improving the accuracy and reliability of the system in scenarios sensitive to false alarms, such as security monitoring and content review.
[0114] like Figure 2 As shown, in some specific embodiments, taking a 15-second traffic accident short video with a frame rate of 25fps and a sampling rate of 16kHz as an example, the execution process of the short video content classification and analysis based on multi-agent intelligence includes: Step 1: Data Acquisition and Separation: Receive a 15-second MP4 short video file, demultiplex it using FFmpeg, and separate it into a video stream (375 original frames, timestamp 0~15 seconds) and an audio stream (mono, 16kHz sampling rate, 240k sampling points).
[0115] Step 2: Video perception process: Adaptive keyframe sampling and global-local feature extraction: The video-aware agent dynamically adjusts the sampling rate based on the intensity of motion in the scene. Seconds 0-2 (static street scene): The sampling rate is set to 1 frame / 2 seconds, and only the two frames at second 0 and second 2 are extracted.
[0116] Seconds 2-3 (collision occurs, intense motion): sampling rate increased to 5 frames / second, extracting 5 frames at seconds 2.0, 2.2, 2.4, 2.6, and 2.8.
[0117] Seconds 3-15 (argument and its aftermath, motion slows down): sampling rate is reduced to 2 frames / second, extracting 7 frames at seconds 3, 5, 7, 9, 11, 13, and 15.
[0118] The total number of keyframes extracted throughout the process is T=14 frames, which reduces the amount of computation compared to uniform sampling (15 frames) while fully preserving the details at the moment of collision.
[0119] For each keyframe: Global features: Input the complete frame into the CLIP visual encoder to obtain 512-dimensional global visual features.
[0120] Local features: The YOLOv11 model (which only outputs bounding boxes and does not classify) is called to detect potential regions of interest and crop them. For example, if two bounding boxes are detected, namely "car damage area" and "pedestrian arm area", they are cropped and fed into the same CLIP encoder to obtain two local visual features.
[0121] Weighted fusion: The reasoning agent dynamically sets a dynamic adjustment factor based on the context (local details are more important in the collision phase), and then calculates the single-frame enhanced features based on two local visual features.
[0122] The 14 single-frame augmentation features are arranged temporally into a sequence, injected with learnable positional codes, and input into a temporal Transformer encoder (the maximum number of input frames is preset to 64; in this example, T=14 is less than 64, so no compression is needed; if the long video exceeds 64 frames, it is compressed to 64 frames by calculating the cosine similarity of features from adjacent frames and performing average pooling fusion on the redundant frames with the highest similarity). The output is a joint visual feature that simultaneously encodes spatial details and the dynamic evolution of the "collision-dispute".
[0123] Step 3: Audio perception process: Dynamic event detection, context expansion, and attention fusion: The audio-aware agent first calculates a short-time energy sequence (frame length 25ms, frame shift 10ms), and then updates the adaptive threshold in real time based on the noise floor (see the formula above). The original audio is then segmented into: Background segment 1: 0~2.2 seconds (street ambient noise); Event segment A: 2.2~2.9 seconds (collision pulse); Event segment B: 2.9~10.2 seconds (arguing voices); Background segment 2: 10.2~15 seconds (traffic echoes); For each event segment, perform context expansion, adding a 0.5-second buffer window before and after, to obtain the expanded event segment: Extended segment A: 1.7~3.4 seconds (including pre-collision braking sound, collision pulse, and post-collision echo); Extended segment B: 2.4~10.7 seconds (including the prelude to the argument, the climax of the argument, and the lingering effects of the argument). The background segment and the expanded event segment are converted into Mel spectrograms, respectively, and input into the BEATs acoustic model to extract features, resulting in 512-dimensional acoustic features for the background segment and expanded acoustic features for the event segment. The expanded event segment acoustic features are then input into an LSTM to obtain the hidden state sequence. Using the background segment acoustic features as the query and the hidden state sequence as the key / value pair, dot product attention weights are calculated (collision sounds receive higher weights). The focused event features are then obtained, and the final audio features are generated through residual and layer normalization.
[0124] Step 4: Cross-modal fusion and preliminary classification: The reasoning agent receives joint visual features and final audio features, performs bidirectional parallel cross-attention, and obtains a multimodal joint representation. The knowledge management agent retrieves candidate categories from a dynamic knowledge base (the knowledge base has generated category description vectors such as "traffic accident," "crowd gathering," and "violent incident" through LLM), calculates cosine similarity, and recalls the top 3: "traffic accident" (0.85), "crowd gathering" (0.62), and "violent incident" (0.45). The reasoning agent outputs preliminary classification results, with the highest confidence category being "traffic accident" (0.85).
[0125] Step 5: Assessment and Feedback Adjustment The evaluation agent acquires the local category predicted by the video perception agent for each keyframe (e.g., "broken glass" instead of "holding a knife" at 2.4 seconds) and the event category predicted by the audio perception agent for each extended event segment ("collision sound," "argument sound"), and calculates a comprehensive conflict score. In this example, there is no conflict. If the value is less than the threshold of 0.2, the final classification result is output directly.
[0126] Suppose another scenario: In the 2.4-second keyframe, due to dim lighting, a local detection falsely detects "holding a knife" (confidence level only 0.4), which conflicts with "traffic accident." (Total number of keyframes) The value is still less than 0.2, so no feedback is triggered.
[0127] For demonstration purposes, assume there are more conflicting frames or "upbeat music" in the audio. When At that time, the evaluation agent generates contradictory information (including the 2.4-second timestamp). The inference agent issues an adjustment instruction: instructing the video perception agent to increase the confidence threshold for target detection in that frame from 0.3 to 0.6, and to re-detect. After re-detection, the "holding a knife" bounding box is filtered out, "broken glass" is correctly detected, local features are re-extracted and replaced with the original features, and weighted fusion and time-series modeling are re-executed to generate new joint visual features. After another cross-modal fusion, the confidence of "traffic accident" rises to 0.92, the evaluation agent reviews it again, and the contradiction is resolved.
[0128] Step 6: Semantic Reasonableness Review (Optional): If the OCR recognizes the text "fire drill" from the keyframe and initially classifies it as "fire" (confidence 0.9), and the large language model determines a logical conflict, then the evaluation agent will reduce the confidence of "fire" by 0.5 (to 0.4) and output the reason for the conflict: "The text 'fire drill' in the image indicates that this is a simulation training, not a real fire."
[0129] Step 7: Final classification result output: After analyzing the input short video of a traffic accident, the following conclusions were output: The highest confidence category for this video is "traffic accident," with a confidence level of 0.92. Two key events were detected in the video: the collision sound occurred between 2.2 and 2.9 seconds, and the argument sound occurred between 2.9 and 10.2 seconds. In the frame at 2.4 seconds, a "damaged car" was detected in pixel region (120, 200, 300, 180), and a "pedestrian" was detected in region (350, 220, 80, 180). The inference process record shows that the visual perception agent extracted the car damage and the pedestrian's arm-waving motion, while the audio perception agent identified the collision pulse and prolonged argument sound. Feedback adjustment eliminated local false detections, and the final comprehensive judgment was a traffic accident.
[0130] The above embodiments fully demonstrate the entire process of this invention, from video stream separation, adaptive sampling, global-local fusion, dynamic audio segmentation, context expansion, attention fusion, cross-modal bidirectional interaction, quantization consistency review, feedback adjustment (including threshold adjustment) to semantic rationality review (confidence penalty), achieving highly robust, zero-sample, and interpretable rapid classification of short video content.
[0131] like Figure 3 As shown in the figure, an embodiment of the present invention provides a short video content classification and analysis system based on multi-agent technology, comprising: An acquisition and segmentation unit is used to acquire short video data and separate the short video data into video streams and audio streams; The video processing unit is used to extract global features from the video stream using a video perception agent to obtain global visual features; to detect and crop local regions in the video stream based on a preset target detection model and extract local visual features; and to fuse the global visual features and the local visual features in a temporal sequence to generate joint visual features. An audio processing unit is configured to segment the audio stream based on dynamic event detection using an audio-sensing agent to obtain a background segment and an event segment, and to perform contextual expansion on the event segment to generate an expanded event segment. The unit also extracts the acoustic features of the background segment and the expanded event segment respectively, and fuses them through an attention mechanism to generate the final audio features based on the acoustic features of the background segment and the expanded event segment. The preliminary classification unit is used to perform cross-modal fusion of the joint visual features and the final audio features through an inference agent to generate a multimodal joint representation; to retrieve the most relevant candidate category description from a preset dynamic knowledge base based on the multimodal joint representation through a knowledge management agent; and to generate a preliminary classification result based on the inference agent, the multimodal joint representation, and the candidate category description. The final classification unit is used to perform multimodal consistency review on the preliminary classification results by the evaluation agent, obtain the review results, and determine the final classification results based on the review results.
[0132] This invention provides a short video content classification and analysis device based on multi-agent technology, comprising a memory and a processor; the memory is used to store a computer program; the processor is used to implement the short video content classification and analysis method based on multi-agent technology as described above when the computer program is executed.
[0133] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the multi-agent-based short video content classification and analysis method described above.
[0134] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A method for classifying and analyzing short video content based on multi-agent systems, characterized in that, include: Acquire short video data and separate the short video data into video streams and audio streams; The video stream is subjected to global feature extraction by a video perception agent to obtain global visual features. Local regions in the video stream are detected and cropped based on a preset target detection model, and local visual features are extracted. The global visual features and the local visual features are fused in a temporal sequence to generate joint visual features. The audio stream is segmented based on dynamic event detection by an audio-aware agent to obtain background segments and event segments. The event segments are then extended with context to generate extended event segments. The acoustic features of the background segments and the extended event segments are extracted and fused through an attention mechanism to generate the final audio features based on the acoustic features of the background segments and the extended event segments. The joint visual features and the final audio features are fused across modalities by an inference agent to generate a multimodal joint representation; The knowledge management agent retrieves the most relevant candidate category descriptions from a preset dynamic knowledge base based on the multimodal joint representation; and the reasoning agent generates preliminary classification results based on the multimodal joint representation and the candidate category descriptions. The preliminary classification results are evaluated by a multimodal consistency review of the agent to obtain the review results, and the final classification results are determined based on the review results. The step of generating final audio features through attention mechanism fusion based on the acoustic features of the background segment and the expanded event segment includes: The background segment and the extended event segment are converted into Mel spectrograms respectively, and then input into a preset acoustic feature extraction model to obtain the acoustic features of the background segment and the acoustic features of the extended event segment. The acoustic features of the extended event segment are input into a preset long short-term memory network for temporal encoding to obtain a hidden state sequence. Using the acoustic features of the background segment as the query vector and the hidden state sequence as the key vector and value vector, dot product attention calculation is performed to obtain the focused event features; The acoustic features of the background segment and the focused event features are fused together using residual connections and layer normalization to generate the final audio features.
2. The method for classifying and analyzing short video content based on multi-agent technology according to claim 1, characterized in that, The step of fusing the global visual features and the local visual features sequentially to generate joint visual features includes: The global visual features at the same time point are weighted and fused with at least one local visual feature to generate a single-frame enhanced visual feature; wherein the weights of the weighted fusion are dynamically adjusted by the inference agent based on context information. The multiple single-frame enhanced visual features at different time points are arranged in temporal order to form a visual feature sequence; The visual feature sequence is superimposed with the temporal position code, and the superimposed feature sequence is then subjected to temporal modeling to obtain the joint visual feature.
3. The method for classifying and analyzing short video content based on multi-agent technology according to claim 1, characterized in that, The step of segmenting the audio stream based on dynamic event detection using an audio-sensing agent to obtain background segments and event segments, and then performing contextual expansion on the event segments to generate expanded event segments, includes: A short-time energy sequence is calculated from the audio stream in chronological order. The short-time energy sequence contains multiple short-time energy values, and each short-time energy value corresponds to a time point. If the short-time energy value is greater than the corresponding adaptive dynamic threshold, then the corresponding time point is marked as an event candidate point; otherwise, it is marked as a background candidate point. At least one event segment is divided based on all the event candidate points, and at least one background segment is divided based on all the background candidate points; A buffer window of preset duration is added before the start time and after the end time of the event segment to obtain the extended event segment.
4. The method for classifying and analyzing short video content based on multi-agent technology according to claim 1, characterized in that, The step of fusing the joint visual features and the final audio features across modalities through an inference agent to generate a multimodal joint representation includes: Based on a bidirectional parallel cross-attention mechanism, the joint visual features are used as the query vector, and the final audio features are used as the key vector and value vector to calculate the output data of visual-guided audio attention; the final audio features are used as the query vector, and the joint visual features are used as the key vector and value vector to calculate the output data of audio-guided visual attention. The output data of the visually guided audio attention and the output data of the audio-guided visual attention are concatenated and fused to generate the multimodal joint representation.
5. The method for classifying and analyzing short video content based on multi-agent technology according to claim 1, characterized in that, The review results include those that contain contradictions and those that do not; The process of evaluating the agent's multimodal consistency of the preliminary classification results to obtain the evaluation results includes: Obtain the predicted event category for each of the extended event segments, and the predicted local category for each sampled keyframe in the video stream; The highest confidence category in the preliminary classification results is determined, and the number of conflicting keyframes is determined based on the local categories and the highest confidence category; the number of conflicting event segments is determined based on the event categories and the highest confidence category. A comprehensive conflict score is obtained based on the number of conflict keyframes, the number of conflict event segments, the total number of keyframes, and the total number of expanded event segments. If the overall conflict score exceeds a preset threshold, a conflict is determined to exist; otherwise, no conflict is determined to exist.
6. The method for classifying and analyzing short video content based on multi-agent technology according to claim 5, characterized in that, The step of determining the final classification result based on the review results includes: When the review results are contradictory, a feedback adjustment loop is triggered; The feedback adjustment loop includes: The evaluation agent generates conflict information, which includes the timestamp of at least one sampled keyframe in the video stream where the conflict occurs, or the time interval of at least one of the extended event segments. Based on the reasoning agent, an adjustment instruction is generated according to the contradictory information, and the adjustment instruction includes at least one of the following operations: If the contradictory information originates from a conflict between the global category and the local category, then the weighted fusion weight of the global visual features and the local visual features in the video perception agent is adjusted, and the weight of the conflicting feature is reduced by a preset first adjustment value. If the contradictory information originates from a conflict between visual perception and audio perception, a mask matrix is introduced in the self-attention calculation of the temporal Transformer to reset the attention weight of the specified keyframe to zero, so that it does not participate in feature aggregation; or a mask is introduced in the audio attention fusion stage to reset the attention weight of the specified audio event segment features to zero. If the contradictory information originates from an abnormal target detection confidence level, then the detection confidence threshold of the target detection model in the video perception agent is adjusted by raising or lowering the threshold by a preset second adjustment value. In response to the adjustment command, the video sensing agent and / or the audio sensing agent perform corresponding parameter adjustments and re-execute the feature extraction and fusion steps to generate updated joint visual features and / or updated final audio features. The reasoning agent re-executes the cross-modal fusion, knowledge retrieval, and preliminary classification generation steps based on the updated joint visual features and / or the updated final audio features, and the evaluation agent performs a consistency review again. Repeat the above process until the review result is consistent with the previous result or the preset maximum number of iterations is reached, and then output the final classification result.
7. The method for classifying and analyzing short video content based on multi-agent technology according to claim 1, characterized in that, The construction process of the dynamic knowledge base includes: Retrieve categories defined in natural language; Based on a pre-defined large language model, multimodal text descriptions are generated for the category under pre-defined output constraints; The multimodal text description is converted into a category description feature vector, and the dynamic knowledge base is constructed based on the category description feature vector; When the category confidence obtained by the reasoning agent matching the multimodal joint representation with the category description feature vector in the dynamic knowledge base is lower than a preset threshold, the agent calls an external knowledge source to obtain the latest information through retrieval enhancement generation technology, and generates or updates the category description feature vector. Furthermore, based on a preset time period, the category description feature vectors in the dynamic knowledge base are subjected to retrieval enhancement and generation updates to refresh outdated category descriptions.
8. The method for classifying and analyzing short video content based on multi-agent technology according to claim 6, characterized in that, The multimodal consistency review of the preliminary classification results by evaluating the intelligent agent also includes a semantic rationality review; The process of reviewing semantic reasonableness includes: Text information is extracted from the keyframes corresponding to the category with the highest confidence in the preliminary classification results using an optical character recognition model; Based on the large language model, determine whether there is a logical conflict between the text information and the category of the preliminary classification result; If a conflict exists, the confidence score of the category is reduced, and the reason for the conflict is output.
9. A short video content classification and analysis system based on multi-agent intelligence, characterized in that, The system, applied to the multi-agent-based short video content classification and analysis method as described in any one of claims 1 to 8, comprises: An acquisition and segmentation unit is used to acquire short video data and separate the short video data into video streams and audio streams; The video processing unit is used to extract global features from the video stream using a video perception agent to obtain global visual features; to detect and crop local regions in the video stream based on a preset target detection model and extract local visual features; and to fuse the global visual features and the local visual features in a temporal sequence to generate joint visual features. An audio processing unit is configured to segment the audio stream based on dynamic event detection using an audio-sensing agent to obtain a background segment and an event segment, and to perform contextual expansion on the event segment to generate an expanded event segment. The unit also extracts the acoustic features of the background segment and the expanded event segment respectively, and fuses them through an attention mechanism to generate the final audio features based on the acoustic features of the background segment and the expanded event segment. The preliminary classification unit is used to perform cross-modal fusion of the joint visual features and the final audio features through an inference agent to generate a multimodal joint representation; to retrieve the most relevant candidate category description from a preset dynamic knowledge base based on the multimodal joint representation through a knowledge management agent; and to generate a preliminary classification result based on the inference agent, the multimodal joint representation, and the candidate category description. The final classification unit is used to perform multimodal consistency review on the preliminary classification results by the evaluation agent, obtain the review results, and determine the final classification results based on the review results.
Citation Information
Patent Citations
Segmentation-based feature extraction for acoustic scene classification
CN111279414A
Dense video description method based on multi-modal memory knowledge
CN120318740A
Illegal content auditing method and device based on multi-modal data, equipment and medium
CN121278126A
Roadside monitoring traffic accident prediction method based on scale perception time domain aggregation
CN122135316A
Contextual advertising through multimodal content analysis
US20260059182A1