Audio and video multi-mode analysis system and analysis method based on large model collaboration
By adopting a modular and large-scale intelligent agent collaborative processing framework, combined with multimodal intelligent segmentation algorithms and keyframe intelligent extraction technology, the problem of insufficient multimodal collaborative analysis capabilities is solved, achieving efficient and secure audio and video multimodal parsing, and improving processing accuracy and real-time performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
- Filing Date
- 2026-01-05
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies lack multimodal collaborative analysis capabilities, making it difficult to achieve deep semantic alignment and collaborative analysis of multimodal information such as video, audio, and text. Furthermore, the lack of localized intelligent processing solutions results in high processing latency, significant data security risks, and insufficient depth in extracting structured metadata.
A modular and large-model intelligent agent collaborative processing framework is adopted, which combines multimodal intelligent segmentation algorithm and keyframe intelligent extraction technology. Through large-model collaborative optimization of audio and video multimodal parsing system, efficient fusion of multimodal information and extraction of structured metadata are achieved.
It enables efficient collaborative analysis of multimodal information, improves the accuracy of keyframe extraction and scene detection, ensures the integrity of multi-dimensional metadata extraction, reduces processing latency and cost, meets real-time requirements, and supports localized deployment and data security.
Smart Images

Figure CN121884233A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and multimedia information processing technology, and in particular to an audio and video multimodal parsing system, parsing method, computer-readable storage medium, and electronic device based on large model collaboration. Background Technology
[0002] I. Current Status of Technological Development 1. Traditional audio and video processing framework Technical features: It is built based on traditional computer vision libraries and uses a fixed algorithm process for processing.
[0003] Key shortcomings: Lack of semantic understanding capabilities, making it difficult to conduct in-depth content analysis; rigid processing procedures, unable to adapt to diverse business scenarios; metadata extraction is limited to the technical parameter level, lacking structured information representation at the semantic level.
[0004] 2. Single-modal intelligent analysis system Technical features: Focus on intelligent analysis of a single modality (video or audio).
[0005] Key drawbacks: information is isolated between modalities, making multimodal collaborative analysis impossible; lack of cross-modal semantic alignment capabilities; and one-sided analysis results, making it difficult to support comprehensive content understanding.
[0006] 3. Cloud-based large model service platform Technical features: Content analysis is achieved by relying on cloud-based large model APIs.
[0007] Main drawbacks: High data transmission latency, making it difficult to meet real-time processing requirements; data privacy and security risks exist; high cost makes it unsuitable for large-scale deployment.
[0008] II. Technical Bottlenecks of Existing Technologies Core Issue 1: Insufficient Multimodal Collaborative Analysis Capabilities. Existing technologies lack efficient multimodal information fusion mechanisms, making it difficult to achieve deep semantic alignment and collaborative analysis of multimodal information such as video, audio, and text.
[0009] The second core issue is the lack of localized intelligent processing capabilities. Over-reliance on cloud services leads to high processing latency, significant data security risks, and a lack of efficient localized intelligent processing solutions.
[0010] The third core issue is the insufficient depth of structured metadata extraction. Traditional methods primarily focus on extracting technical metadata and lack the ability to describe structured content based on deep semantic understanding. Summary of the Invention
[0011] To address the aforementioned problems, this application proposes a novel audio-visual multimodal parsing system and method based on large-model collaboration. More specifically, this invention provides an audio-visual multimodal parsing and structured metadata intelligent extraction system based on large-model collaboration. Through core technologies such as large-model collaborative optimization, intelligent scene detection, and keyframe semantic extraction, it overcomes the shortcomings of existing technologies in multimodal collaborative analysis capabilities, localized deployment of large models, and modular collaborative architecture design. This invention has significant practical value and broad application prospects in the fields of artificial intelligence and multimedia information processing.
[0012] To achieve the above objectives, the present invention employs the following technical strategies: 1. Modular and large-scale intelligent agent collaborative processing framework The traditional algorithm module and large model agent work together mechanism is adopted, including: (1) establishing a unified media processing coordination mechanism to manage the collaborative work of the algorithm module and the large model agent; (2) designing an intelligent task allocation algorithm to automatically select the algorithm module or the large model agent to process according to the task type; (3) constructing an asynchronous parallel processing architecture to support the efficient collaboration between the algorithm module and the large model agent.
[0013] 2. Multimodal intelligent segmentation algorithm The intelligent segmentation method based on content change detection and time slice clustering is constructed, including: (1) Content change detection mechanism: the inter-frame difference analysis algorithm is adopted to automatically identify scene transition boundaries based on video content change features; (2) Semantic boundary verification: the semantic rationality of the detected scene boundaries is verified by combining large model semantic understanding to ensure that the segmentation conforms to the content logic; (3) Adaptive parameter adjustment: the detection threshold (e.g., 0.3-0.8) and minimum scene length (e.g., 3-15 seconds) are dynamically adjusted according to the complexity of video content to achieve accurate segmentation; (4) Time slice clustering algorithm: for long scenes with a duration of >60 seconds, they are divided into 60-second time slices, and the semantically coherent sub-segment boundaries are identified based on content similarity clustering analysis; (5) Long and short scene processing strategy: short scenes with a duration of <3 seconds are intelligently merged, and long scenes with a duration of >60 seconds are segmented by time slice clustering; (6) Quality assurance mechanism: each segment is at least 20 seconds long, and segments are automatically merged when the interval is <5 seconds.
[0014] 3. Intelligent keyframe extraction technology A keyframe extraction and multi-dimensional deduplication optimization method based on scene transition points is constructed, including: (1) Scene transition point extraction mechanism: based on the content change boundary identified by the scene detection algorithm, representative keyframes are intelligently selected in each scene; (2) Multi-dimensional deduplication optimization: combining visual similarity detection (image feature comparison) and semantic importance evaluation (large model semantic understanding) to eliminate content redundancy; (3) Adaptive sampling strategy: dynamically adjust the number of keyframes (1-3 per scene) according to the scene duration and content complexity to achieve optimal content coverage; (4) Quality assurance mechanism: comprehensively consider image clarity, content richness and semantic representativeness to ensure the quality of keyframes.
[0015] Specifically, this application provides the following technical solutions: The first aspect of this application provides an audio and video multimodal parsing system based on large-model collaboration, such as... Figure 3 As shown, this system includes: A unified media processing agent is used to uniformly process audio and video files of various formats, and to achieve basic media information extraction and format standardization. The algorithm processing module cluster includes algorithm modules for video segmentation, audio extraction, and keyframe extraction, which are used for parallel analysis and processing of multimodal content. The large model intelligent agent cluster, which is communicatively coupled to the unified media processing intelligent agent, includes a multimodal content analysis intelligent agent, a text semantic analysis intelligent agent, a knowledge extraction intelligent agent, and a structured output intelligent agent, which respectively realize multimodal content analysis, text semantic processing, knowledge insight extraction, and formatted conversion of analysis results; The structured output module is used to generate standardized JSON metadata and Markdown analysis reports, which include technical metadata, content metadata, semantic metadata, and business metadata.
[0016] Furthermore, in the system of this application, the unified media processing agent adopts an asynchronous parallel architecture to coordinate the collaborative work of the algorithm processing module cluster and the large model agent cluster. The large model agent cluster feeds back the semantic analysis results to the unified media processing agent to optimize scene segmentation and key frame selection. The structured output module integrates the multimodal analysis results and semantic metadata and outputs structured description information.
[0017] Furthermore, in the system of this application, the unified media processing agent includes: The scene detection and segmentation module identifies scene transition boundaries based on the inter-frame difference analysis algorithm, and verifies semantic rationality by combining large model semantic understanding to achieve intelligent segmentation; The keyframe extraction and deduplication module extracts representative keyframes based on scene transition points and combines visual similarity detection and semantic importance assessment to achieve multi-dimensional deduplication. The audio extraction and recognition module enables multi-format speech recognition, audio separation, and audio feature extraction.
[0018] Furthermore, in this application's system, the scene detection and segmentation module employs the ContentDetector algorithm from the PySceneDetect library to calculate the differences in hue, saturation, and luminance components between adjacent frames in the HSV color space. When the difference score exceeds a preset threshold, it is determined to be a scene transition point; wherein, the difference score is calculated according to the following formula: content_val = (ΔHue × w_hue + ΔSat × w_sat + ΔLum × w_lum + ΔEdges × w_edges) / (w_hue + w_sat + w_lum + w_edges); In the formula, content_val is the difference score, Δ represents the difference between adjacent frames, w is the weight coefficient of each component, Hue is the hue, Sat is the saturation, Lum is the brightness, and Edges is the edge position.
[0019] Furthermore, in the system of this application, the keyframe extraction and deduplication module uses the pHash perceptual hash algorithm to generate a 64-bit visual fingerprint, performs visual similarity detection by calculating the Hamming distance, and determines visually similar frames when the Hamming distance is ≤5; and combines text semantic similarity calculation, uses the SequenceMatcher algorithm to compare the OCR recognition text or description text of the keyframe, and determines content duplication when the similarity ratio is ≥0.85; the multi-dimensional deduplication logic adopts "OR" logic decision, that is, if either visual similarity or text similarity condition is met, it is determined to be a duplicate keyframe.
[0020] Furthermore, in the system of this application, the keyframe extraction and deduplication module executes an adaptive sampling strategy in each scene: extracting the start frame and end frame for short scenes with a duration of less than 5 seconds, extracting additional intermediate frames for scenes with a duration of 5 to 15 seconds, and adding intermediate time point sampling for long scenes with a duration of more than 15 seconds; and comprehensively evaluating the image clarity, content richness, and semantic representativeness of candidate keyframes to determine the final keyframe set.
[0021] Furthermore, in the system of this application, the multimodal content analysis agent is based on the Qwen Omni multimodal large model to realize multimodal content analysis of video clips and keyframes; the text semantic analysis agent processes text summarization, content summarization and semantic analysis; the knowledge extraction agent extracts key knowledge and insights from multimodal information; and the structured output agent converts the analysis results into standardized JSON and Markdown formats.
[0022] Furthermore, in the system of this application, the unified media processing agent executes a long and short scene processing strategy: for short scenes with a duration of less than 3 seconds, intelligent merging is initiated; for long scenes with a duration of more than 60 seconds, time slice clustering is performed according to 60-second time slices; and a quality assurance mechanism is executed to ensure that the duration of each segment is not less than 20 seconds, and that adjacent segments are automatically merged when the interval is less than 5 seconds.
[0023] A second aspect of this application provides a method for audio and video multimodal parsing based on large-model collaboration, wherein the method is applied to the aforementioned system, such as... Figure 4 As shown, the method includes: S1. Media File Reception and Verification: Audio and video files are received through a unified media processing agent, and format verification and basic information extraction are performed. S2. Multimodal Parallel Analysis: Utilizes a cluster of algorithm processing modules to perform parallel analysis of multimodal content, including video stream analysis and keyframe extraction; S3. Semantic Fusion and Knowledge Extraction: Utilizing large-scale model agent clusters to perform semantic fusion of multimodal features and generate structured metadata; S4. Structured Output: Generates and outputs standardized JSON metadata and Markdown analysis reports.
[0024] Furthermore, in the method of this application, In step S1, the format verification includes file integrity check and format compatibility verification, supporting MP4, AVI, MOV, and MKV formats; the basic information includes file size, duration, resolution, frame rate, and encoding format. In step S2, the video stream analysis is based on the ContentDetector algorithm for scene detection, with a threshold of 30 and a minimum scene length of 15. The keyframe extraction is based on keyframes extracted from scene transition points, and the pHash perceptual hashing algorithm is used for visual deduplication, with a visual similarity threshold of 5 (Hamming distance ≤ 5). In step S4, the JSON metadata includes technical metadata, content metadata, semantic metadata, and business metadata; the Markdown analysis report includes an analysis summary, key findings, and visualization results, and supports integration with external systems through a standardized interface.
[0025] A third aspect of this application provides an electronic device, including: a memory and a processor; Memory: Used to store computer programs; Processor: Used to execute the computer program to implement the steps of the aforementioned audio and video multimodal parsing method based on large model collaboration.
[0026] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned audio-visual multimodal parsing method based on large-model collaboration.
[0027] In summary, compared with existing technologies, this invention has the following significant advantages in terms of architecture design, processing mechanism, and engineering implementation: 1. Technological innovation 1.1 Audio and Video Intelligent Agent Collaborative Architecture It pioneers a distributed processing framework based on multi-agent systems: each agent's task is decoupled and specialized, supporting dynamic expansion and plug-and-play integration; it is compatible with localized large-scale private deployment, avoiding the latency, cost, and data security risks of cloud services; and it adopts a modular service bus design to achieve loosely coupled communication and collaborative scheduling between agents.
[0028] 1.2 Multimodal Fusion and Alignment Mechanism A dual visual-semantic deduplication strategy is proposed: a 64-bit visual fingerprint is constructed based on the pHash perceptual hashing algorithm, and a joint fingerprint is formed by combining it with text semantic similarity calculation. By using an adaptive threshold combination of Hamming distance ≤ 5 and text similarity ≥ 0.85, a double OR logic verification is achieved to ensure deduplication robustness and recall. A cross-modal semantic alignment channel is established to map visual features and text descriptions to a unified semantic space, supporting the joint representation and collaborative reasoning of multi-source heterogeneous data.
[0029] 1.3 Asynchronous Parallel Processing Architecture Design a heterogeneous computing architecture that is parallel at both the task and data levels: asynchronous scheduling of tasks such as video stream parsing, scene segmentation, keyframe extraction, and audio processing, which improves throughput by 2-3 times compared to serial processes; integrates a dynamic resource allocation algorithm to intelligently adjust CPU / memory resources according to task load, with actual measurements showing an improvement in CPU utilization of over 40%; supports memory optimization and batch processing of large-scale video files.
[0030] 2. Technological Advantages 2.1 Accuracy Indicators Keyframe extraction accuracy ≥90%, scene detection accuracy ≥85%, multi-dimensional metadata extraction completeness ≥95%, covering a three-layer structure of technical parameters, content tags and semantic descriptions.
[0031] 2.2 Efficiency Optimization The asynchronous parallel architecture combined with multi-threaded fingerprint computing significantly reduces end-to-end processing latency; localized deployment eliminates network I / O overhead, meeting real-time business requirements; memory pooling and intelligent scheduling optimize resource usage.
[0032] 3. Practical value 3.1 Cost and Autonomy It adopts an open-source technology stack and a local inference engine, eliminating dependence on cloud APIs and reducing licensing and call costs; it supports deployment at the edge and in private data centers, meeting data sovereignty and compliance audit requirements.
[0033] 3.2 Scalability and Interoperability The standardized RESTful API and JSON metadata output format facilitate integration with third-party systems such as media asset management and search recommendation; the modular design supports hot-swapping of functional components and domain-adaptive fine-tuning.
[0034] 4. Commercial Value 4.1 Technological Barriers The core intellectual property consists of the intelligent agent collaboration framework and the localized large model deployment scheme; the algorithm has been verified by engineering and has technical feasibility and system stability; the dual deduplication and cross-modal alignment mechanism form a differentiated competitive advantage.
[0035] 4.2 Market adaptability It is suitable for vertical scenarios such as educational recording and broadcasting, media production, security monitoring, and digital archives; it supports flexible deployment from single nodes to clusters, taking into account the needs of small and medium-sized enterprises and large institutions; standardized interfaces are conducive to building a technology ecosystem and channel cooperation. Attached Figure Description
[0036] To more clearly illustrate the technical solution of this application, the accompanying drawings involved in the description of this application will be briefly introduced below. It should be noted that the drawings only show some embodiments of this application. For those skilled in the art, other related drawings can be derived from these drawings without creative effort.
[0037] Figure 1 This is a diagram showing the overall design architecture of the system of the present invention.
[0038] Figure 2 This is a timing diagram for asynchronous processing in the present invention.
[0039] Figure 3 This is a structural diagram of the audio and video multimodal parsing system based on large model collaboration in this application.
[0040] Figure 4 This is a flowchart illustrating the overall implementation of the audio and video multimodal parsing method based on large model collaboration in this application.
[0041] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be noted that the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.
[0043] In this document, the term "comprising" and any variations thereof (such as "including," "including," etc.) are open-ended expressions and should be understood as "including but not limited to," meaning that the listed content is not exhaustive and may include other content not explicitly mentioned. The term "based on" should be understood as "at least partially based on," meaning that the basis or condition referred to may not be the only factor and may involve other relevant factors. The term "one embodiment" should be understood as "at least one embodiment," meaning that the described embodiment is not the only possible implementation, and other similar embodiments may exist.
[0044] In this application, the terms "a" and "a plurality of" are used to modify related elements or features, and their expression is illustrative rather than restrictive. Unless otherwise expressly stated in the context, "a" should be understood as "at least one," and "a plurality of" should be understood as "at least two." Those skilled in the art should reasonably interpret these terms based on the semantic and logical relationships of the context to ensure that they cover the possibility of "one or more."
[0045] Example: An audio / video multimodal parsing system and method based on large-model collaboration This embodiment provides an audio and video multimodal parsing system based on a large model intelligent agent, which adopts a modular intelligent agent architecture to achieve efficient local processing.
[0046] Figure 1 The diagram shows the overall design architecture of the system of this invention. Architecture description: (1) Modular design: Each module has a clear responsibility, and the unified media processing agent acts as a coordinator to manage the processing flow.
[0047] (2) Algorithm and agent collaboration: Traditional algorithm modules work in collaboration with large model agents.
[0048] (3) Large model intelligent agent cluster: It includes four specialized intelligent agents: multimodal content analysis intelligent agent, text semantic analysis intelligent agent, knowledge extraction intelligent agent, and structured output intelligent agent, which realize specialized semantic analysis based on large model.
[0049] (4) Number of agents: The system includes a unified media processing agent (coordinator) and a large model agent cluster (4 specialized agents).
[0050] Figure 2 This is a timing diagram of asynchronous processing in the present invention, with timing description as follows: Participant definition: This system contains 6 core participants, including user audio and video file input, unified media processing agent (coordinator), multimodal content analysis agent, text semantic analysis agent, knowledge extraction agent, and structured output agent.
[0051] Parallel processing stage: The unified media processing agent performs video stream analysis (scene detection, keyframe extraction) and audio stream analysis (speech recognition) simultaneously.
[0052] Intelligent agent collaboration process: The multimodal content analysis intelligent agent and the text semantic analysis intelligent agent process in parallel, pass the results to the knowledge extraction intelligent agent for fusion, and finally generate a standardized report by the structured output intelligent agent.
[0053] Advantages of asynchronous processing: It supports the parallel execution of multimodal tasks, significantly improving processing efficiency.
[0054] Specifically, this system comprises four core components: I. Multimodal Media Processing Module 1. Technical features: The unified media processor supports the parsing and standardization of various audio and video formats.
[0055] Intelligent scene detection and segmentation algorithm automatically identifies scene boundaries based on content changes.
[0056] The keyframe intelligent extraction mechanism combines visual similarity and semantic importance for dual evaluation.
[0057] 2. Core sub-module functions: Video segmentation submodule: Enables high-quality video segmentation and scene boundary detection.
[0058] Audio extraction submodule: Supports multi-format audio separation and standardization processing.
[0059] Keyframe extraction submodule: applies advanced image processing algorithms to perform keyframe deduplication and optimization selection.
[0060] II. Large-scale collaborative analysis engine 1. Technical features: The multi-agent collaborative architecture includes a visual content parsing agent, a time-series content analysis agent, a knowledge extraction agent, and a metadata extraction agent.
[0061] An asynchronous parallel processing mechanism supports the simultaneous processing of multiple media segments.
[0062] A data transfer protocol between intelligent agents ensures the effective integration of processing results.
[0063] 2. Core Intelligent Agent Functions: Unified Media Processing Agent: Responsible for unified processing of audio and video files, scene detection, keyframe extraction, audio extraction and other basic processing.
[0064] Multimodal content analysis agent: Performs multimodal content analysis on video clips and keyframes based on the Qwen Omni multimodal large model.
[0065] Text semantic analysis agent: handles text processing tasks such as text summarization, content summarization, and semantic analysis.
[0066] Structured output agent: responsible for converting analysis results into standardized JSON and Markdown formats.
[0067] 3. Scene detection algorithm implementation: This system uses the ContentDetector algorithm to implement scene detection, specifically including: (1) Frame difference detection mechanism: The scene transition point is detected based on the color and intensity changes between adjacent frames. The algorithm calculates the inter-frame differences in the HSV color space and identifies scene transitions by comparing the changes in the hue, saturation and value components of adjacent frames.
[0068] (2) Adaptive Threshold Calculation: The threshold parameter is set to 30, and the minimum scene length is 15 frames. When the difference score between adjacent frames exceeds the threshold, the system determines it as a scene transition point. The formula for calculating the difference score is: content_val = (ΔHue × w_hue + ΔSat × w_sat + ΔLum × w_lum + ΔEdges × w_edges) / (w_hue + w_sat + w_lum + w_edges); Where Δ represents the difference between adjacent frames, and w is the weighting coefficient of each component.
[0069] (3) Multi-component weight configuration: Supports custom weight coefficients for each component. The default configuration is that the luminance component has a higher weight, and the edge detection component can be enabled optionally. By adjusting the weights, it can adapt to different types of video content.
[0070] (4) Minimum scene length constraint: Set the min_scene_len=15 parameter to ensure that the detected scene contains at least 15 frames to avoid false detections due to short-term screen changes.
[0071] (5) Scene boundary optimization: Post-processing of detected scenes, including short scene merging and long scene segmentation. When the scene duration is less than the minimum segment duration, it is automatically merged into an adjacent scene; when the scene duration exceeds the maximum block duration, it is segmented again according to a fixed duration.
[0072] (6) Anti-interference mechanism: The algorithm has a certain robustness to interference factors such as fast camera movement. By judging the continuity and consistency of inter-frame differences, the false detection rate is reduced.
[0073] 4. Keyframe extraction algorithm implementation: This system employs an intelligent keyframe extraction algorithm based on scene transition points, specifically including: (1) Multi-time point sampling strategy: In each detected scene, a three-time point sampling strategy is adopted to extract the start frame, middle frame and end frame as key frame candidates. This strategy ensures that the dynamic change process of scene content can be fully covered.
[0074] (2) Dynamic frame number adjustment mechanism: The number of key frames is intelligently adjusted according to the scene duration. For short scenes (duration < 5 seconds), only the start frame and the end frame are extracted; for medium-duration scenes (5-15 seconds), three key frames are extracted: start, middle and end; for long scenes (> 15 seconds), additional sampling is added at the middle time point.
[0075] (3) Scene boundary optimization processing: Post-processing of scene detection results, including short scene merging and long scene segmentation. When the scene duration is less than the minimum segment duration, it is automatically merged into an adjacent scene; when the scene duration exceeds the maximum block duration, it is segmented again according to a fixed duration.
[0076] (4) Degradation processing mechanism: When the scene detection algorithm fails, time uniform sampling is used as an alternative scheme. Key frames are extracted at fixed time intervals (such as every 10 seconds) to ensure that the system can work normally under various video content.
[0077] (5) Keyframe quality assessment: The extracted keyframes are screened for quality, including image sharpness assessment, content integrity check, and avoiding black frames or invalid frames, to ensure the representativeness and usability of the keyframes.
[0078] 5. Implementation of dual deduplication mechanism: This system employs a dual deduplication mechanism combining visual similarity and text similarity, specifically including: (1) pHash perceptual hashing algorithm: The image perceptual hashing algorithm (pHash) is used to calculate the visual fingerprint of the key frame.
[0079] The pHash algorithm is implemented through the following steps: Image scaling: Scaling the image to a uniform 32×32 pixels.
[0080] Grayscale conversion: Convert to a grayscale image.
[0081] Discrete Cosine Transform (DCT): Calculates the frequency domain features of an image.
[0082] Feature extraction: The 8×8 region in the upper left corner is selected as the main feature.
[0083] Hash generation: Generate 64-bit binary hash values based on feature values.
[0084] (2) Visual similarity threshold setting: Set keyframe_similarity_threshold=5 as the visual similarity threshold. Hamming distance calculation formula: hamming_distance = sum(hash1[i] != hash2[i]for i in range(64)) When the Hamming distance is ≤5, the frames are judged to be visually similar and deduplication is performed.
[0085] (3) Text similarity calculation: After performing OCR recognition and image description generation on the keyframes, the SequenceMatcher algorithm is used to calculate the text similarity: similarity_ratio = SequenceMatcher(None, text1, text2).ratio() Set text_similarity_threshold=0.85 as the text similarity threshold. When the similarity is ≥0.85, it is judged as content duplication.
[0086] (4) Dual deduplication logic: The "OR" logic is used for deduplication judgment, that is, as long as either visual similarity or text similarity is met, it is determined to be a duplicate keyframe. This dual guarantee mechanism effectively avoids the limitations of a single deduplication method.
[0087] (5) Multi-threaded parallel processing: Multi-threaded technology is used to process the deduplication calculation of key frames in parallel, which significantly improves the processing efficiency. Each key frame is independently hashed and analyzed, and finally the similarity comparison and deduplication decision are made in a unified manner.
[0088] (6) Deduplication result optimization: The deduplicated keyframe set is sorted and filtered to retain the most representative keyframes while avoiding content redundancy. Through this dual deduplication mechanism, the system can effectively eliminate duplicate content at both the visual and semantic levels, improving the accuracy and efficiency of the analysis results.
[0089] III. Module Collaboration and Optimization Mechanism 1. Technical features: Optimize data flow between modules to reduce redundant calculations and memory usage.
[0090] Intelligent caching mechanism improves the efficiency of handling duplicate content.
[0091] Error recovery mechanisms ensure stable system operation.
[0092] 2. Core Optimization Technology: Data preprocessing optimization: Automatically identify and skip invalid or duplicate content.
[0093] Process optimization: Dynamically adjust processing order and resource allocation.
[0094] Post-processing optimization of results: intelligent merging and deduplication of results.
[0095] IV. Configuration Management System 1. Technical features: The parameterized configuration system supports flexible adaptation to different business scenarios.
[0096] A performance monitoring mechanism tracks the system's operating status in real time.
[0097] Error handling and recovery strategies ensure system reliability.
[0098] 2. Intelligent audio and video segmentation algorithm Algorithm flow: (1) Media file parsing: Format verification and basic information extraction are performed using a unified media processor. Supported formats: MP4, AVI, MOV, MKV and other mainstream video formats.
[0099] Basic information: file size, duration, resolution, frame rate, encoding format, etc.
[0100] Verification mechanisms: file integrity check, format compatibility verification.
[0101] (2) Scene boundary detection: Calculate and identify content change points based on inter-frame difference. Core algorithm: The ContentDetector algorithm from the PySceneDetect library is used.
[0102] Threshold settings: threshold=30 (frame difference threshold), min_scene_len=15 (minimum scene length).
[0103] Difference calculation: Calculation of inter-frame difference based on RGB color space and brightness variation.
[0104] (3) Semantic segmentation optimization: optimize segmentation boundaries by combining semantic understanding of large model. Boundary verification: The semantic rationality of scene boundaries is verified using the Qwen Omni multimodal large model.
[0105] Segment merging: Intelligent merging of short scenes with a duration of less than 3 seconds.
[0106] Long scene segmentation: Long scenes with a duration of >60 seconds are clustered and segmented into 60-second time slices.
[0107] (4) Segment quality assessment: Ensure the semantic integrity and representativeness of each segment. Quality indicators: segment length (20-60 seconds), semantic coherence, and content completeness.
[0108] Boundary optimization: Automatic merging when the interval between segments is less than 5 seconds to ensure semantic coherence.
[0109] Technological innovation points: A dynamic threshold adjustment mechanism adapts to the characteristics of different video content.
[0110] Semantic consistency verification helps avoid excessive splitting or merging.
[0111] Real-time processing capabilities, supporting streaming audio and video analysis.
[0112] 3. Intelligent keyframe extraction technology Extraction strategy: (1) Uniform sampling within the scene: Extract three representative keyframes from the start, middle and end of each scene. Sampling strategy: Sampling is based on uniform distribution of scene time points.
[0113] Frame rate adjustment: Dynamically adjust the number of keyframes (1-3 frames) based on the scene duration.
[0114] Time point calculation: start frame (0%), middle frame (50%), end frame (100%).
[0115] (2) Visual deduplication: Image similarity detection algorithm is applied to identify similar frames. Core algorithm: pHash perceptual hashing algorithm.
[0116] Similarity threshold: keyframe_similarity_threshold=5 (Hamming distance ≤ 5).
[0117] Hash calculation: Generate a 64-bit visual fingerprint and calculate the Hamming distance.
[0118] (3) Semantic importance assessment: Combine large model analysis to assess the semantic representativeness of keyframes. Semantic analysis: The semantic importance of keyframes is analyzed using the Qwen Omni large model.
[0119] Importance score: based on a comprehensive score of content richness, semantic representativeness, and image quality.
[0120] Filtering mechanism: Retain the keyframes with the highest scores to ensure semantic representativeness.
[0121] (4) Quality optimization screening: Final selection is based on image clarity and content richness. Sharpness assessment: based on image gradient, contrast, and edge sharpness assessment.
[0122] Content richness: assessed based on image information entropy, color distribution, and texture complexity.
[0123] Final selection: The final selection was based on a combination of visual quality and semantic importance.
[0124] Technical advantages: A dual deduplication mechanism is used to avoid content redundancy.
[0125] Semantic-driven selection ensures the representativeness of keyframes.
[0126] Highly efficient processing algorithms support large-scale video analysis.
[0127] 4. Multimodal content analysis methods Analysis process: (1) Parallel feature extraction: Video, audio and text features are extracted in parallel.
[0128] (2) Agent collaborative analysis: Modality-specific analysis is performed by agents of different specialties. Multimodal content analysis agent: Processes multimodal analysis of video clips and keyframes.
[0129] Text semantic analysis agent: processes text summarization, content summarization, and semantic analysis.
[0130] Knowledge extraction agent: Extracts key knowledge and insights from multimodal information.
[0131] (3) Structured output generation: Generate standardized JSON and Markdown format reports. JSON format: contains technical metadata, content metadata, and semantic metadata.
[0132] Markdown report: Includes analysis summary, key findings, and visualization results.
[0133] Standardized interface: Supports seamless integration with other systems.
[0134] 5. Structured metadata extraction methods Extraction dimensions: Technical metadata: basic information such as file format, duration, and resolution.
[0135] Content metadata: scene segmentation, keyframe description, speech transcription, etc.
[0136] Semantic metadata: content summarization, topic identification, sentiment analysis, etc.
[0137] Business metadata: custom tags, category information, related data, etc.
[0138] 6. Algorithm Flow of Module Collaborative Processing Engine 1. Algorithm design principles: (1) Asynchronous parallel processing: non-blocking I / O operations are implemented using asyncio.
[0139] (2) Intelligent task scheduling: dynamically allocate computing resources based on task complexity.
[0140] (3) Fault tolerance mechanism: The failure of a single task does not affect the operation of the overall system.
[0141] 2. Detailed Implementation Process Core processing flow: (1) Media processor initialization: The system starts the unified media processor to prepare for processing the input audio and video files. Initialization parameters: set the maximum concurrency, memory limit, timeout, etc.
[0142] Resource pre-allocation: Pre-allocate computing resources to avoid runtime resource contention.
[0143] Status monitoring: Monitor system status in real time to ensure stable operation.
[0144] (2) Basic media information extraction: Perform format validation and basic information extraction on the input file. Format verification: Use the ffmpeg library to verify the file format and integrity.
[0145] Information extraction: Extract file size, duration, resolution, frame rate, encoding format, etc.
[0146] Quality assessment: Evaluate document quality and filter out damaged or low-quality documents.
[0147] (3) Multi-agent collaborative analysis: Multiple specialized agents process content of different modalities in parallel. Agent scheduling: Automatically selects the appropriate agent to handle the task based on its type.
[0148] Parallel processing: Video, audio, and text analysis tasks are executed in parallel.
[0149] Results aggregation: The analysis results of each agent are uniformly aggregated and fused.
[0150] (4) Structured output of results: Convert the analysis results into standardized JSON and Markdown formats. Data standardization: unifying data formats and coding standards.
[0151] Format conversion: Generate JSON metadata and Markdown reports.
[0152] Quality check: Perform quality verification and integrity checks on the output results.
[0153] In summary, the algorithmic features of this application are as follows: (1) Modular design: each component has a clear responsibility and standardized interfaces.
[0154] (2) Asynchronous processing: improves system throughput and supports high concurrency processing.
[0155] (3) Error isolation: Enhances system stability, and a single fault does not affect the overall operation.
[0156] (4) Intelligent scheduling: dynamically optimize resource allocation based on task complexity.
[0157] (5) Quality assurance: A multi-layered quality inspection mechanism ensures the reliability of the output results.
[0158] The flowcharts and block diagrams in the accompanying drawings illustrate possible implementations of systems, methods, and computer program products according to various embodiments of this application, including architecture, functionality, and operation. In these figures, each block may represent a module, program segment, or portion of code containing one or more executable instructions for implementing a specified logical function. Each block in the block diagrams and / or flowcharts, and combinations thereof, can be implemented using either a dedicated hardware-based system or a combination of dedicated hardware and computer instructions to achieve the specified function or operation.
[0159] like Figure 5 As shown in the illustration, an embodiment of this application also discloses an electronic device, including: a processor 310, a communication interface 320, a memory 330 for storing a processor-executable computer program, and a communication bus 340. The processor 310, communication interface 320, and memory 330 communicate with each other via the communication bus 340. The processor 310 executes the executable computer program to implement the steps of the aforementioned large-model collaborative audio-visual multimodal parsing method.
[0160] It is understood that, in addition to memory and a processor, this electronic device may also include input devices (such as a keyboard), output devices (such as a display), and other communication modules. These input devices, output devices, and other communication modules all communicate with the processor through I / O interfaces (i.e., input / output interfaces).
[0161] The operations described in this application can be implemented by writing computer program code using one or more programming languages or a combination thereof. The programming languages include, but are not limited to, the following types: Object-oriented programming languages, such as Java, Smalltalk, C++, etc. Conventional procedural programming languages, such as "C" or similar programming languages.
[0162] The execution methods of program code include, but are not limited to: It runs entirely on the user's computer; Part of it executes on the user's computer, and part of it executes on a remote computer; Execute as a standalone software package; It is executed entirely on a remote computer or server.
[0163] In scenarios involving remote computers, the remote computer can connect to the user's computer via any type of network, including but not limited to local area networks (LANs) or wide area networks (WANs). Furthermore, the remote computer can also connect to external computers through an internet service provider, for example, by utilizing the internet for connection.
[0164] Furthermore, this application also discloses a computer-readable storage medium, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform the various steps of the audio-visual multimodal parsing method based on large model collaboration disclosed in this application.
[0165] In the context of this application, a computer-readable storage medium refers to a tangible medium capable of storing computer program code and related data. Specific examples include, but are not limited to, the following: (1) Portable computer disk: such as floppy disks and other removable magnetic storage media.
[0166] (2) Hard disk: including mechanical hard disks and solid-state hard disks and other fixed storage devices.
[0167] (3) Random Access Memory (RAM): A volatile storage medium used for temporary storage of data and program code.
[0168] (4) Read-only memory (ROM): a non-volatile storage medium used to store fixed programs and data.
[0169] (5) Erasable programmable read-only memory (EPROM) or flash memory: non-volatile storage media that supports multiple erasures and reprogrammings.
[0170] (6) Fiber optic storage devices: storage media based on fiber optic technology.
[0171] (7) Portable compact disc read-only memory (CD-ROM): a read-only medium that stores data in the form of an optical disc.
[0172] (8) Optical storage devices: such as DVDs, Blu-ray discs and other storage media based on optical principles.
[0173] (9) Magnetic storage devices: such as magnetic tapes, disks and other storage media based on magnetic principles.
[0174] (10) Any suitable combination of the above: for example, combining multiple storage media to meet different storage needs.
[0175] These computer-readable storage media can be used to store the program code and related data described in this application to support program execution and persistent data storage.
[0176] Specifically, according to embodiments of this application, the processes described in the flowcharts can be implemented as computer software programs. For example, embodiments of this application relate to a computer program product comprising a computer program carried on a non-transitory computer-readable medium. This computer program contains program code for executing the large-model collaborative audio-visual multimodal parsing method disclosed in this application. When this computer program is executed by a processing system, it can achieve the functions defined in the embodiments of this application.
[0177] While the foregoing discussion contains several specific implementation details, these details should not be construed as limiting the scope of this application. The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this application is not limited to technical solutions formed by specific combinations of the above-described technical features. Furthermore, this application should also cover other technical solutions formed by any combination of the above-described technical features or their equivalents without departing from the foregoing disclosed concept.
[0178] Those skilled in the art should also understand that modifications can be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features, without departing from the spirit and scope of the technical solutions of the embodiments of this application. These modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the core spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A multimodal audio-visual parsing system based on large-model collaboration, characterized in that, The system includes: A unified media processing agent is used to uniformly process audio and video files of various formats, and to achieve basic media information extraction and format standardization. The algorithm processing module cluster includes algorithm modules for video segmentation, audio extraction, and keyframe extraction, which are used for parallel analysis and processing of multimodal content. The large model intelligent agent cluster, which is communicatively coupled to the unified media processing intelligent agent, includes a multimodal content analysis intelligent agent, a text semantic analysis intelligent agent, a knowledge extraction intelligent agent, and a structured output intelligent agent, which respectively realize multimodal content analysis, text semantic processing, knowledge insight extraction, and formatted conversion of analysis results; The structured output module is used to generate standardized JSON metadata and Markdown analysis reports, which include technical metadata, content metadata, semantic metadata, and business metadata.
2. The system according to claim 1, characterized in that, The unified media processing agent adopts an asynchronous parallel architecture to coordinate the collaborative work of the algorithm processing module cluster and the large model agent cluster. The large model agent cluster feeds back the semantic analysis results to the unified media processing agent to optimize scene segmentation and key frame selection. The structured output module integrates multimodal analysis results and semantic metadata and outputs structured description information.
3. The system according to claim 1, characterized in that, The unified media processing agent includes: The scene detection and segmentation module identifies scene transition boundaries based on the inter-frame difference analysis algorithm, and verifies semantic rationality by combining large model semantic understanding to achieve intelligent segmentation; The keyframe extraction and deduplication module extracts representative keyframes based on scene transition points and combines visual similarity detection and semantic importance assessment to achieve multi-dimensional deduplication. The audio extraction and recognition module enables multi-format speech recognition, audio separation, and audio feature extraction.
4. The system according to claim 3, characterized in that, The scene detection and segmentation module uses the ContentDetector algorithm from the PySceneDetect library to calculate the differences in hue, saturation, and luminance components between adjacent frames in the HSV color space. When the difference score exceeds a preset threshold, it is determined as a scene transition point. The difference score is calculated according to the following formula: content_val = (ΔHue × w_hue + ΔSat × w_sat + ΔLum × w_lum + ΔEdges× w_edges) / (w_hue + w_sat + w_lum + w_edges); In the formula, content_val is the difference score, Δ represents the difference between adjacent frames, w is the weight coefficient of each component, Hue is the hue, Sat is the saturation, Lum is the brightness, and Edges is the edge position.
5. The system according to claim 3, characterized in that, The keyframe extraction and deduplication module uses the pHash perceptual hash algorithm to generate a 64-bit visual fingerprint. Visual similarity is detected by calculating the Hamming distance. When the Hamming distance is ≤5, the keyframe is determined to be visually similar. Combined with text semantic similarity calculation, the SequenceMatcher algorithm is used to compare the OCR recognition text or description text of the keyframe. When the similarity ratio is ≥0.85, the keyframe is determined to be duplicated. The multi-dimensional deduplication logic uses "OR" logic decision, that is, if either visual similarity or text similarity condition is met, the keyframe is determined to be duplicated.
6. The system according to claim 3, characterized in that, The keyframe extraction and deduplication module executes an adaptive sampling strategy in each scene: extracting the start and end frames for short scenes with a duration of less than 5 seconds, extracting additional intermediate frames for scenes with a duration of 5 to 15 seconds, and adding intermediate time point sampling for long scenes with a duration of more than 15 seconds; and comprehensively evaluating the image clarity, content richness, and semantic representativeness of candidate keyframes to determine the final keyframe set.
7. The system according to claim 1, characterized in that, The multimodal content analysis agent is based on the QwenOmni multimodal large model to realize multimodal content analysis of video clips and keyframes; the text semantic analysis agent processes text summarization, content summarization and semantic analysis; the knowledge extraction agent extracts key knowledge and insights from multimodal information. The structured output agent converts the analysis results into standardized JSON and Markdown formats.
8. The system according to claim 1, characterized in that, The unified media processing agent executes a long and short scene processing strategy: it initiates intelligent merging for short scenes with a duration of less than 3 seconds, and performs time-slice clustering for long scenes with a duration of more than 60 seconds in 60-second time slices; and it executes a quality assurance mechanism to ensure that the duration of each segment is not less than 20 seconds, and automatically merges adjacent segments when the interval is less than 5 seconds.
9. A method for audio and video multimodal parsing based on large model collaboration, characterized in that, The method is applied to the system as described in any one of claims 1-8, and the method comprises: S1. Media File Reception and Verification: Audio and video files are received through a unified media processing agent, and format verification and basic information extraction are performed. S2. Multimodal Parallel Analysis: Utilizes a cluster of algorithm processing modules to perform parallel analysis of multimodal content, including video stream analysis and keyframe extraction; S3. Semantic Fusion and Knowledge Extraction: Utilizing large-scale model agent clusters to perform semantic fusion of multimodal features and generate structured metadata; S4. Structured Output: Generates and outputs standardized JSON metadata and Markdown analysis reports.
10. The method according to claim 9, characterized in that: In step S1, the format verification includes file integrity check and format compatibility verification, supporting MP4, AVI, MOV, and MKV formats; the basic information includes file size, duration, resolution, frame rate, and encoding format. In step S2, the video stream analysis is based on the ContentDetector algorithm for scene detection, with a threshold of 30 and a minimum scene length of 15. The keyframe extraction is based on keyframes extracted from scene transition points, and the pHash perceptual hashing algorithm is used for visual deduplication, with a visual similarity threshold of 5. In step S4, the JSON metadata includes technical metadata, content metadata, semantic metadata, and business metadata; the Markdown analysis report includes an analysis summary, key findings, and visualization results, and supports integration with external systems through a standardized interface.