A semantic timeline-based multi-modal video structured analysis and intelligent playing system and method
By using a multimodal video structured analysis system based on a semantic time axis, the system addresses the shortcomings of existing video data processing systems in terms of structured analysis capabilities and efficiency. It enables precise management and intelligent playback of video data, thereby improving data utilization efficiency and quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUJIAN POLYTECHNIC OF INFORMATION TECH
- Filing Date
- 2026-03-02
- Publication Date
- 2026-06-05
AI Technical Summary
Existing video data processing systems lack in-depth structured analysis capabilities, are unable to achieve intelligent chapter segmentation and accurate retrieval, have low processing efficiency, and the quality of the generated structured data is difficult to guarantee, resulting in low efficiency in video data utilization.
A multimodal video structured analysis system based on a semantic timeline is adopted, including a video access module, a multimodal analysis engine, a structured knowledge base, and an intelligent interactive player. Through speech recognition, semantic segmentation, subtitle generation, and summary extraction, structured knowledge data is generated, and a distributed asynchronous task architecture and quality control module are used to ensure data accuracy.
It enables precise structured management and intelligent playback of video data, improves processing efficiency, supports high-concurrency video processing, ensures data quality, provides accurate navigation and knowledge acquisition, and expands the application boundaries of video data.
Smart Images

Figure CN122153118A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer information processing technology, specifically to a multimodal video structured analysis and intelligent playback system and method based on a semantic time axis. Background Technology
[0002] Most existing video data exists in unstructured form, making it difficult for computers to effectively identify, parse, and mine its content information. This results in low utilization efficiency of video data, failing to fully realize its knowledge and application value. Furthermore, existing video processing and playback technologies still face numerous unresolved issues, specifically in the following aspects: Current mainstream video processing systems focus on basic processing such as audio and video decoding and transcoding, lacking the ability to perform in-depth structural analysis of the semantic level of video content. They cannot achieve intelligent segmentation of chapters based on the audio content and semantic logic of the video. Existing video chapters are mostly created manually, which not only consumes a lot of manpower and time, but also makes it difficult to unify production standards. Furthermore, manually divided chapters often lack precise timeline binding, making it impossible to achieve precise linkage with video playback. The development of video content retrieval technology is lagging behind. Existing retrieval methods are mostly based on shallow information such as video file names and manually added tags. They cannot perform accurate retrieval based on the semantic content of the video, nor can they achieve semantic retrieval of similar segments across videos. When users want to obtain specific knowledge content in a video, they need to watch the video frame by frame to search for it, which is extremely time-consuming and labor-intensive. The system architecture for video processing often adopts a serial processing mode, which executes the entire video processing flow as a single task. This results in low processing efficiency and a lack of effective result caching and reuse mechanisms. When faced with processing requests for the same video or videos with highly similar content, the entire processing flow must be executed repeatedly, resulting in a large waste of computing resources and failing to meet the needs of high concurrency and multi-batch video processing. Furthermore, existing technologies lack effective quality control mechanisms for machine-generated video-related structured data. Issues such as missing fields, timestamp logic errors, and semantic inconsistencies are prone to occur in generated chapter information, abstract keywords, etc. Moreover, there is no professional visual review interface for staff to make corrections, making it difficult to guarantee the accuracy and standardization of structured data, which further limits the intelligent utilization of video data.
[0003] In summary, there is an urgent need for a technical solution that can achieve structured analysis of video content at the semantic level and deeply integrate structured knowledge with video playback. This solution would address the problems of insufficient structured analysis capabilities, low retrieval efficiency, poor playback experience, and inefficient processing architecture in existing video processing and playback technologies, thereby improving the intelligent management and utilization of video data. Summary of the Invention
[0004] The purpose of this invention is to provide a multimodal video structured analysis and intelligent playback system and method based on a semantic time axis to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A multimodal video structured analysis and intelligent playback system based on a semantic time axis includes: Video access module, multimodal analysis engine, structured knowledge base, and intelligent interactive player; The video access module is used to receive and store raw video files and initiate processing tasks; The multimodal analysis engine is used to process the original video files and generate structured knowledge data based on a unified timeline. The structured knowledge data includes at least: a chapter directory based on semantic segmentation, subtitle text precisely aligned with the timeline, and video or chapter-level summaries and keywords. The structured knowledge base is used to store and index structured knowledge data, and to establish the relationship between video media files and structured knowledge data; The intelligent interactive player is used to load video files and their associated structured knowledge data, and realizes video playback, chapter navigation, subtitle synchronization and knowledge content linkage display based on the semantic timeline.
[0006] A method for multimodal video structured analysis and intelligent playback based on a semantic time axis, executed by a multimodal video structured analysis and intelligent playback system based on a semantic time axis, includes the following steps: S1: Receive the original video file, perform audio and video preprocessing, and extract standard audio track data; S2: Perform speech recognition on standard audio track data to generate an initial text sequence with timestamps; S3: Clean and semantically segment the initial text sequence to obtain a purified text stream; S4: Perform temporal semantic analysis on the purified text stream, combine semantic understanding and preset temporal constraint rules to determine the chapter segmentation points of the video, and generate structured chapter data; S5: Generates a standard subtitle file aligned with the video timeline based on the cleaned text stream and its timestamp; S6: Generate summary description information for the video based on structured chapter data; S7: Link and store the original video files, structured chapter data, subtitle files, and summary description information, and create an index.
[0007] As can be seen from the technical solution provided by the present invention above, the beneficial effects of the multimodal video structured analysis and intelligent playback system and method based on semantic time axis provided by the present invention are: The multimodal analysis engine transforms the original video into structured knowledge data containing semantically segmented chapter headings, time-aligned subtitles, multi-granular summaries, and keywords. This gives the originally unstructured video data a structured expression, significantly improving the intelligent management level of video data and laying the foundation for in-depth mining and utilization of video data. This invention precisely aligns all structured knowledge data with a unified video timeline at the millisecond level. From chapter segmentation to subtitle generation and knowledge summary display, data association is achieved around the timeline, providing core data support for the intelligent interactive player's precise chapter navigation, real-time subtitle synchronization, and linked display of knowledge cards, thus realizing the integrated fusion of video playback and knowledge acquisition. The system adopts a distributed asynchronous task architecture that breaks down the complex video processing flow into multiple asynchronous subtasks and executes them in a distributed parallel manner. Combined with progress management, it enables real-time monitoring of task status. The result caching and reuse module caches the results of processed videos or similar videos to avoid redundant calculations and effectively reduce the consumption of computing resources. This not only improves the efficiency of single video processing but also supports multiple batches of high-concurrency video processing requests, optimizing the overall computing and storage resource utilization of the system. The rule verification submodule of the quality control module realizes automated verification of data fields, formats and timestamps. Then, the consistency assessment submodule completes the semantic rationality detection. Combined with the visualization playback and correction function of the manual review workbench, it realizes dual control of machine verification and manual audit. At the same time, all correction operations are fully traceable, which improves the credibility of structured knowledge data and the traceability of the whole process. The intelligent interactive player offers precise chapter navigation, allowing users to quickly jump to video content of interest without cumbersome fast-forwarding or rewinding. The subtitle style can be customized to suit different users' viewing habits. Knowledge cards display the core knowledge of the current chapter or video in real time, allowing users to quickly grasp the key points without watching the video word by word, greatly improving the video viewing experience and the efficiency of knowledge acquisition. The inverted index and semantic vector index built by the structured knowledge base support both keyword retrieval and cross-video similar segment semantic retrieval, allowing users to quickly locate target video content. At the same time, it can generate shareable links pointing to specific time points in the video and embeddable content cards, combining the specific video content with structured knowledge to achieve lightweight sharing and multi-scenario reuse of video knowledge, thus expanding the application boundaries of video data. Attached Figure Description
[0008] Figure 1This is a schematic diagram of the structure of a multimodal video structured analysis and intelligent playback system based on a semantic time axis according to the present invention; Figure 2 This is a schematic diagram of the steps of a multimodal video structured analysis and intelligent playback method based on a semantic time axis according to the present invention. Detailed Implementation
[0009] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0010] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific embodiments.
[0011] like Figure 1-2 As shown, this embodiment of the invention provides a multimodal video structured analysis and intelligent playback system based on a semantic time axis, comprising: Video access module, multimodal analysis engine, structured knowledge base, and intelligent interactive player; The video access module is used to receive and store raw video files and initiate processing tasks; The multimodal analysis engine is used to process the original video files and generate structured knowledge data based on a unified timeline. The structured knowledge data includes at least: a chapter directory based on semantic segmentation, subtitle text precisely aligned with the timeline, and video or chapter-level summaries and keywords. The structured knowledge base is used to store and index structured knowledge data, and to establish the relationship between video media files and structured knowledge data; The intelligent interactive player is used to load video files and their associated structured knowledge data, and realizes video playback, chapter navigation, subtitle synchronization and knowledge content linkage display based on the semantic timeline.
[0012] In this embodiment, the video access module is the entry hub of the multimodal video structured analysis and intelligent playback system based on the semantic time axis. It provides a complete and standardized video data source for the subsequent multimodal analysis engine processing and intelligent interactive player display of the system by accurately starting and scheduling the entire process of receiving, storing and processing the original video files, ensuring that the system's process proceeds in an orderly and efficient manner from the data entry point. The video access module is primarily responsible for receiving raw video files of various formats with full compatibility, performing basic verification and standardized preprocessing of video data, classifying and uniquely identifying video files through a secure and efficient storage strategy, and generating corresponding video processing tasks according to system preset rules. It then accurately distributes task information to the system's task scheduling center to initiate subsequent multimodal analysis processes. The module also tracks the access status of video files and task distribution in real time, creating data records to provide a basis for video data traceability and system process control, ensuring seamless integration from data entry to analysis. The video access module includes: Video receiving unit: Multi-format compatible access: Supports access to raw video files in mainstream video container and encoding formats, covering both streaming video transmission access and local video file upload access. It can adapt to video data input needs in different scenarios, realize full reception of various video sources, and ensure a video data entry experience without format barriers. Basic video data verification: The received video data is verified for integrity and validity. The video file is checked for problems such as data corruption, missing file header, and interrupted bitstream. At the same time, the basic parameters of the video such as resolution, frame rate, and bitrate are verified to ensure that they are within the range that the system can process. Video data that fails the verification is marked as abnormal and fed back to the front end. Video standardization preprocessing: Basic standardization processing is performed on the verified video files to unify the video encapsulation format and encoding parameters, and extract basic metadata of the video, including video duration, total number of frames, file size, etc., to provide a uniform video source for the subsequent audio and video preprocessing unit and reduce the workload of format adaptation in the subsequent analysis stage. Video storage management unit: Tiered storage strategy execution: Based on the size and importance of video files and the priority of subsequent processing, video files are allocated to different storage media in the system. Large, high-priority video files are stored in high-speed read / write storage media to meet the needs of fast retrieval, while regular video files are stored in large-capacity storage media to achieve rational allocation and efficient utilization of storage resources. Video Unique Identifier Generation: A unique identifier is generated for each successfully accessed video file. This identifier is associated with the video's basic metadata and becomes the core identity identifier of the video file in the system. Subsequent structured knowledge data establishes a relationship with the original video file through this identifier, ensuring accurate matching between video data and analysis data. Video data archiving and retrieval: Video files are classified and archived according to their access time, file type, and processing status. A retrieval index based on the unique identifier of the video and basic metadata is established to support the system's quick retrieval and query of video files. At the same time, the stored video files are regularly checked for status to prevent data loss or damage. Task initiation and scheduling unit: Automatic task generation: Based on the unique identifier and basic metadata of the video, the corresponding video multimodal analysis and processing task is automatically generated according to the system's preset task generation rules. The task information includes the unique identifier of the video, storage address, processing requirements, and basic parameters, ensuring the completeness and accuracy of the task information. Intelligent task priority allocation: Based on the processing requirements of the video file access scenario and the current system load, the generated processing tasks are prioritized. Video tasks with urgent processing needs are assigned high priority, and regular video tasks are assigned normal priority, realizing intelligent scheduling of system task processing resources and ensuring the priority execution of high-priority tasks. Task information linkage distribution: The generated and priority-assigned processing task information is accurately distributed to the system's task scheduling center. At the same time, a real-time communication link is established between the module and the task scheduling center to synchronize the task distribution status, ensuring that the task information can be accurately received and orchestrated by the task scheduling center to start the subsequent asynchronous subtask processing flow. Real-time task status tracking: Through the communication link with the task scheduling center, the overall progress of video processing tasks is tracked in real time, the execution status of subtasks is received from the task scheduling center, the task status is associated with the unique identifier of the video and recorded, and the front end supports the visualization of video processing progress. Furthermore, the multi-format video parsing technology is based on the audio and video codec standard library and constructs a unified video format parsing framework. This framework can identify the file structure of different video encapsulation formats and the bitstream characteristics of encoding formats. Through corresponding parsing plugins, it can complete the decapsulation and bitstream extraction of various types of video data. The core technology lies in abstracting the common features and individual differences of different video formats to form a standardized parsing interface, realizing unified parsing and access of different video formats, and breaking down the compatibility barriers between video formats. Distributed video storage technology is based on the distributed architecture of the system and adopts block storage and replication mechanism to realize the storage management of video files. Large video files are divided into blocks according to a preset size, and different data blocks are allocated to different nodes of the distributed storage cluster. At the same time, multiple copies of each data block are generated and stored on different nodes to prevent data loss due to single node failure. A mapping relationship between data blocks and storage nodes is established through video unique identifiers to realize fast splicing and retrieval of video files, while taking into account storage security and read and write efficiency. The task triggering and scheduling technology is based on an event-driven architecture and a priority scheduling algorithm. Once a video file is stored and a unique identifier is generated, a task generation event is triggered, and the system automatically generates processing tasks according to preset rules. The priority scheduling algorithm constructs a task priority evaluation model, using the video processing demand access time and system load as evaluation dimensions to calculate a priority score for each task. The formula is: ,in, As a score for task priority, To handle demand weight values, reflecting the urgency of video processing needs, This is the access time coefficient; the later the access time, the larger the coefficient value. This is the system load factor; the lower the system load, the larger the factor value. , , Here are the weight coefficients for each dimension, and ; The system according to The tasks are sorted by their values, and the task information is sent to the task scheduling center according to the sorting results to achieve intelligent scheduling and orderly execution of tasks. The video data association and identification technology is based on a hash algorithm. It performs a hash operation on the basic metadata and storage address of the video to generate a unique hash value as the unique identifier of the video. This identifier serves as the core key value and is transmitted between various modules of the system. The structured knowledge data stored in the subsequent structured knowledge base is indexed by this identifier and associated with the original video file to achieve a one-to-one correspondence between video media files and structured knowledge data, ensuring the accuracy and traceability of data association.
[0013] In this embodiment, the multimodal analysis engine is the core processing module of a semantic timeline-based multimodal video structured analysis and intelligent playback system. It performs a full-process analysis of standardized video sources, including audio-video separation, speech recognition, text purification, semantic segmentation, subtitle generation, and summary extraction, transforming unstructured video data into structured knowledge data based on a unified timeline. It serves as the core hub connecting raw video data and structured knowledge data. The engine's processing results provide core data support for the storage index of the structured knowledge base and the interactive display of the intelligent interactive player. Simultaneously, its output is synchronized to the quality control module for verification and auditing, ensuring the accuracy and standardization of the structured knowledge data. The multimodal analysis engine receives standardized video files output by the video access module and performs multi-dimensional analysis and processing of the video through a series of interconnected functional units. Its core functions include audio track extraction and standardization, timestamped speech-to-text conversion, text stream cleanup and optimization, semantic-based video chapter segmentation, precise timeline-aligned subtitle generation, and video-level summary extraction. Ultimately, it generates structured knowledge data containing structured chapter data, standard subtitle files, an overall video summary, and key points. All generated structured knowledge data is precisely bound to the video's unified timeline and adheres to the system's preset format specifications, ensuring both semantic validity and meeting the needs of subsequent storage, retrieval, and playback. This is the core component for achieving structured video content analysis. The multimodal analysis engine includes: Audio and video preprocessing unit: Audio and video separation processing: The received standardized video file is subjected to audio and video stream separation operation to accurately extract pure audio track data from the video, remove the bit stream information related to the video screen, and retain only the audio data related to the speech content, so as to provide a pure audio track data source for subsequent speech recognition processing; Audio track standardization processing: The extracted audio track data is uniformly standardized and optimized, including unifying the audio sampling rate, merging the channels, normalizing the volume, and initially filtering audio noise. This process repairs problems such as cut-off distortion in the audio track, generates standardized audio track data with unified format and parameter specifications, and eliminates the impact of audio format differences on subsequent speech recognition. Audio track timeline calibration: The standardized audio track data is precisely calibrated with the timeline of the original video to ensure that the time dimension of the audio track data is completely synchronized with the video timeline, laying the foundation for the alignment of the subsequently generated text, subtitles, chapter data and the video timeline; Speech recognition and text processing unit: Timestamped speech recognition: Frame-by-frame speech recognition processing is performed on standardized audio track data to convert audio content into text information. At the same time, each segment of recognized text content is bound with corresponding timestamp information to generate an initial text sequence containing start timestamp, end timestamp and corresponding text, thereby realizing the conversion of speech content into text content and preserving time dimension information. Initial text cleaning: The generated initial text sequence is cleaned to remove invalid characters, misspelled words, and duplicate content generated during the recognition process, correct recognition errors caused by unclear speech, and filter out interjections and pauses without actual semantic meaning to improve the effectiveness and accuracy of the text content. Semantic sentence segmentation: Based on natural language processing technology, semantic sentence segmentation is performed on the cleaned text, breaking away from the sentence segmentation method based solely on audio pauses. By combining semantic logic to divide the text into paragraphs and sentences, the sentence segmentation results of the text stream conform to the expression habits of natural language, and finally generate a cleaned text stream with dual norms of timestamp and semantic logic. Temporal semantic segmentation unit: Semantic understanding and topic recognition: Deep natural language processing is performed on the cleaned text stream. Through topic modeling and key concept aggregation, semantic topic changes in the text stream are mined, potential semantic paragraph boundaries are identified, and key time nodes for topic switching in video content are determined, providing semantic basis for video chapter segmentation. Temporal constraint rule adaptation: The system's predefined temporal constraint rule set is retrieved, and the boundaries of the identified potential semantic paragraphs are matched with the rule set. The rule set includes the minimum duration threshold for a single chapter, the maximum duration threshold for a single chapter, the temporal continuity requirements between chapters, and the semantic paragraph coverage requirements of the entire text. Semantic paragraphs that initially meet the rules are then selected. Boundary optimization and structured chapter generation: The constraint optimization algorithm is applied to finely adjust the boundaries of the initially selected semantic segments so that the time boundaries of all chapters simultaneously meet the requirements of semantic coherence and temporal constraint rules. Finally, structured chapter data containing start timestamp, end timestamp, chapter title, chapter summary and chapter keywords are generated to realize semantic-based chapter segmentation of video content. Subtitle generation unit: Precise text-timeline mapping: Based on the purified text stream, the timestamp information corresponding to each sentence is extracted to achieve millisecond-level precise mapping between text content and video timeline, ensuring that the display and disappearance time of subtitles is completely synchronized with the audio content in the video; Standard subtitle format generation: According to the commonly used subtitle file format specifications, the mapped text and timestamp information are converted into standard format subtitle files, supporting sentence-by-sentence display of subtitles and timeline synchronization, meeting the subtitle rendering needs of intelligent interactive players; Subtitle text optimization: The generated subtitle text is optimized for lightweighting, simplifies overly long sentences, retains core semantic content, makes the display rhythm of the subtitle text conform to the user's viewing habits, and improves the readability of the subtitles; Summary of generated units: Chapter-level information extraction: Aggregate and organize chapter summaries and keywords in structured chapter data, extract the core information of each chapter, form a chapter-level knowledge point collection, and retain the core logic of segmentation of video content; Video-level summary generation: Based on the collection of chapter-level knowledge points, combined with the overall semantic logic of the video content, the core information of each chapter is integrated and refined to generate a video-level overall summary that can fully cover the core content of the video, while extracting the core key points that run through the entire video; Summary Information Standardization: The generated overall video summary and key points are standardized according to the system's preset format, and bound to the same video timeline and unique identifier with the structured chapter data and subtitle file to form a complete structured knowledge data package; Furthermore, the audio-video separation and standardization technology is based on audio-video encoding and decoding technology and digital signal processing technology. By analyzing the bitstream structure of the video encapsulation format, it separates the independent audio bitstream and video bitstream, and only performs subsequent processing on the audio bitstream. At the same time, digital signal processing algorithms are used to resample the audio data, merge channels, and suppress noise, unifying the audio parameters to the system's preset standard values, eliminating audio format differences between different video sources, and ensuring consistency of the processed audio track data. This provides high-quality input data for subsequent speech recognition and ensures the accuracy of speech recognition. The timestamped speech recognition technology is based on an end-to-end speech recognition model. It combines the time information of audio frames to realize the conversion of speech to text and bind timestamps. The technology divides standardized audio track data into audio frames according to a fixed duration, assigns a corresponding timestamp to each audio frame, the speech recognition model recognizes each audio frame and outputs the corresponding text, and then splices and semantically corrects the consecutive text frames to finally generate an initial text sequence with start and end timestamps in units of sentences. This achieves accurate binding of speech content with the time dimension and lays the foundation for subsequent timeline-aligned text and subtitle processing. Temporal semantic segmentation and constraint optimization are the core technologies of the multimodal analysis engine, integrating semantic segmentation techniques from natural language processing and constraint optimization algorithms from operations research. First, semantic segmentation of the cleaned text stream is performed using topic models and key concept clustering algorithms to identify the boundaries of topic changes and obtain preliminary semantic paragraphs. Then, a joint optimization model for semantic coherence and temporal constraints is constructed to maximize overall semantic coherence while satisfying all temporal constraint rules. The preliminary boundaries are optimized and adjusted. The objective function and constraints are as follows: ; ; in, For the overall semantic coherence of the video, The total number of chapters. For the first Internal semantic coherence of each chapter For the first The content weight of each chapter This is the minimum duration threshold for a single chapter. This is the threshold for the longest duration of a single chapter. For the first The start timestamp of each chapter. For the first The end timestamp of each chapter. The coverage of semantic paragraphs on the clean text flow. This represents the minimum coverage threshold. The optimal chapter time boundary is solved by this model to ensure that the generated structured chapter data not only conforms to semantic logic, but also meets the temporal constraints of the system. The timeline-aligned subtitle generation technology is based on the timestamp information of the cleaned text stream and the sentence segmentation rules of natural language. It segments the cleaned text stream according to semantic sentences, extracts the start and end timestamps of each sentence, and then encapsulates the sentences and timestamps according to the standard subtitle format specifications. The core technology is to achieve millisecond-level alignment between text sentences and audio timelines. By accurately mapping audio frames and text characters, it ensures that the display time of subtitles is completely synchronized with the audio playback time in the video. At the same time, it improves the display effect and readability of subtitles through lightweight text optimization. The multi-granularity text summary generation technology is based on a pre-trained natural language summarization model and adopts a process strategy of first separating and then merging to achieve multi-granularity summaries at the chapter and video levels. First, the structured data of each chapter is summarized and extracted to generate chapter summaries and keywords. Then, the summaries and keywords of all chapters are used as input and fused and refined by the summary model to generate video-level summaries and key points that can cover the overall content of the video. Through multi-granularity processing, this technology not only preserves the segmented core logic of the video content but also presents the overall content framework, achieving comprehensive knowledge extraction from the video content.
[0014] In this embodiment, the quality control module is the core of the quality control system for the multimodal video structured analysis and intelligent playback system based on the semantic time axis. It conducts full-dimensional automated verification and visual manual auditing of the structured knowledge data output by the multimodal analysis engine, accurately identifying problems such as format errors, semantic deviations, and logical loopholes in the data. It realizes the marking and correction of abnormal data and the release of compliant data, which is a key link to ensure the accuracy, standardization, and usability of the structured knowledge data entering the structured knowledge base. The verification results and correction records of this module can also provide data support for the parameter optimization and algorithm iteration of the multimodal analysis engine, and promote the continuous improvement of the system's analysis capabilities. The quality control module is primarily responsible for receiving structured knowledge data output from the multimodal analysis engine. Through automated rule validation and semantic consistency assessment, it performs full-scale inspection of data fields, including format, timestamps, and semantic content. Data that passes inspection is directly uploaded to the structured knowledge base, while data that fails inspection is marked as anomaly and pushed to the manual review stage. The module provides a visual manual review interface, allowing operators to accurately locate and correct abnormal data. It also records and tracks all verification and correction processes, creating a complete quality inspection archive. Furthermore, the module statistically analyzes the abnormal data generated during verification and review, extracts common problems, and feeds them back to the multimodal analysis engine, providing a basis for algorithm optimization and parameter adjustment. This achieves closed-loop quality control of the system. The quality control module includes: Rule validation submodule: Field integrity verification: All core fields in the structured knowledge data are checked one by one, including the start and end timestamps of the structured chapter data, chapter titles, chapter summaries, and chapter keywords; the text content timestamps of the subtitle files; and the overall summary and key points of the video summary information. This ensures that there are no missing, empty, or invalid values in each field, thus ensuring the integrity of the data structure. Timestamp order and legality verification: Logical and range verification is performed on all timestamp information in the structured knowledge data. The verification checks that the start timestamp of a chapter is less than the end timestamp, that the timestamps of each chapter are continuous without overlap or gaps, that the timestamps of the subtitle files match the overall video duration and are in a reasonable order, and that the timestamp format is consistent with the system's preset standard, to ensure the legality and logic of the timestamp information. Format compliance verification: According to the system's preset format standards, the file format, character encoding, field naming, etc. of the structured knowledge data are verified to ensure that the format of the subtitle file conforms to the general standard, the character encoding of chapter titles, abstracts and keywords is consistent, and the naming of each data field is consistent with the system specifications, so as to ensure that the data can be parsed and called normally by the structured knowledge base and intelligent interactive player. Consistency Assessment Submodule: Intra-chapter semantic consistency assessment: Semantic relevance analysis is performed on the title, abstract, and keywords within the same chapter to assess whether the three revolve around the same theme, whether the title can accurately summarize the core content of the chapter, whether the abstract can fully reflect the main idea of the chapter, and whether the keywords can accurately extract the core elements of the chapter, eliminating semantic inconsistencies such as the title and abstract being disconnected and the keywords being irrelevant to the content; Timeline and content matching consistency assessment: Check the time range of structured chapter data and the semantic matching degree of the corresponding content to confirm that the text content and core theme of a certain chapter are consistent with the video and audio content within that time period, and eliminate semantic deviations caused by misalignment between chapter content and timeline; Multi-granularity summary logical consistency assessment: assess the logical consistency between chapter-level summaries and video-level overall summaries, confirm that the video-level overall summary can fully integrate the core content of each chapter-level summary, that the content of each chapter-level summary is consistent and can support the main idea of the video-level overall summary, and ensure the logical coherence of multi-granularity summary information. Manual review workbench: Time-axis-based visual playback: Provides a visual playback interface that is precisely bound to the video timeline, and links the structured knowledge data generated by the multimodal analysis engine with the original video for display. Operators can view the chapter information subtitles and summary data at the corresponding time points while playing the video, and realize the visual location of abnormal data. Precise data correction: Operators can directly edit and correct marked abnormal data, including adjusting chapter timestamp boundaries, modifying chapter titles, abstract keywords, improving missing field content, and optimizing subtitle text and timeline matching. Correction operations take effect in real time and are synchronized to the data preview interface, making it convenient for operators to verify the correction effect immediately. Full-process operation record keeping: All review operations performed by operators are recorded in detail, including review time, operator information, original status of abnormal data, correction content, correction time, etc., forming an unalterable operation record to ensure the traceability of the review process; Data review and release: A second review process is set up for the revised data. After the operator completes the revision, the system automatically performs a second check on the revised data for rule verification and consistency assessment. If the check is passed, the data is automatically released to the structured knowledge base. If the check is still not passed, the data is returned to the review process. Furthermore, the rule engine verification technology builds an automated verification engine based on a pre-set business rule base. It breaks down the verification rules for structured knowledge data into executable logical judgment statements, forming a standardized set of verification rules. The rule engine performs field traversal and logical matching on the structured knowledge data, executing the judgment statements in the rule set one by one to quickly identify problems such as missing fields, format errors, and abnormal timestamps in the data. The core of this technology lies in the modularity and configurability of the rules. Verification rules can be flexibly added, modified, or deleted according to the system's business needs, achieving adaptability verification for different types of structured knowledge data and improving the flexibility and scalability of the verification engine. Semantic consistency calculation technology is the core technology of the consistency evaluation submodule. Based on a pre-trained natural language processing model, it transforms text content into high-dimensional semantic vectors. By calculating the similarity between semantic vectors, the semantic relevance between texts is evaluated. First, chapter titles, abstracts, and keywords are converted into semantic vectors respectively. Then, the similarity value between vectors is calculated using the cosine similarity algorithm to determine the semantic consistency among the three. For the matching consistency between chapter content and the timeline, the video speech-to-text at the corresponding timestamp and the chapter content are converted into semantic vectors respectively, and the evaluation is achieved through similarity calculation. The core calculation formula is as follows: ,in, This represents the similarity score between two semantic vectors, ranging from 0 to 1. A value closer to 1 indicates higher semantic consistency. This is the semantic vector of the first text. For the semantic vector of the second text, The dot product of two semantic vectors. for The length of the mold, for The length of the module.
[0015] The overall semantic consistency of keywords in the title summary within the same chapter is calculated using the average similarity between each element and the core semantic vector, as shown in the following formula: ,in, This represents the overall semantic consistency value within the chapter, ranging from 0 to 1. A value closer to 1 indicates higher overall semantic consistency. This represents the number of text elements within a chapter that participate in the consistency assessment. For the first The semantic vector of each text element. Semantic vectors of the core content of the chapter; The system presets a semantic consistency threshold. When the calculation result is lower than the threshold, it is judged as semantically inconsistent and marked as abnormal data. The timeline-based precise positioning technology establishes a precise mapping relationship between structured knowledge data and video frames based on a unified timeline of the video. It binds each timestamp in the structured knowledge data to a specific frame in the video. In the visual playback interface, when an operator clicks on the timestamp of a chapter or subtitle, the system can directly jump to the corresponding frame position in the video, achieving precise positioning of abnormal data in the video. This technology ensures that the timeline of structured knowledge data and video content is completely synchronized through millisecond-level parsing of timestamps and index matching of video frames, providing accurate visual references for manual review. Operation tracking and data version management technology, based on blockchain and database version control technology, fully records the original state of structured knowledge data and every correction operation. It generates a unique version identifier for each version of the dataset, which is associated with the operation record, enabling full lifecycle traceability of the data. The immutability of blockchain technology ensures that the operation record cannot be modified arbitrarily, while database version control technology allows operators to trace back and compare different versions of data, facilitating the investigation and correction of problems, and providing complete raw data for system quality analysis.
[0016] In this embodiment, the structured knowledge base is the core data storage and retrieval hub of a multimodal video structured analysis and intelligent playback system based on a semantic timeline. It receives structured knowledge data that has passed the quality control module's verification and establishes a precise association mapping with the original video files stored in the video access module. Through professional storage strategies and a multi-dimensional indexing system, it achieves secure data storage, intelligent indexing, efficient retrieval, and knowledge reuse. It provides real-time data retrieval support for video playback and knowledge linkage display in the intelligent interactive player, and also provides a data foundation for the system's knowledge mining and content reuse. This module is a key bridge connecting the data processing stage and the terminal interaction stage, ensuring that structured knowledge data can be used efficiently and called flexibly. The structured knowledge base is primarily responsible for receiving compliant structured knowledge data output from the quality control module, establishing a unique association between it and the original video files using exclusive identifiers, and employing a distributed storage strategy to achieve secure and categorized storage of video files and structured knowledge data. Based on the characteristics of the structured knowledge data, the module constructs a multi-dimensional index system, providing functions such as keyword search and cross-video semantic search, supporting rapid and accurate retrieval of stored data. Simultaneously, the module possesses knowledge retrieval and reuse capabilities, generating sharing links and content cards for specific video segments, enabling efficient reuse of knowledge content. Furthermore, the module manages and maintains all stored data throughout its entire lifecycle, including data status monitoring, updates, archiving, backup, and recovery, ensuring data integrity, relevance, and accessibility, providing stable data support for the normal operation of all system modules. The structured knowledge base includes: Structured data storage unit: Data association mapping establishment: Receive structured knowledge data transmitted by the quality control module, extract the unique video identifier from the data, associate the identifier with the original video file stored in the video access module, establish a one-to-one correspondence between structured knowledge data and original video file, and ensure that each piece of structured knowledge data can accurately match the corresponding video source. Hierarchical and categorized storage implementation: Based on the type, importance, and access frequency of video files and structured knowledge data, a hierarchical and categorized storage strategy is implemented. High-frequency access video clips and core structured knowledge data are stored in high-speed read / write storage media, while low-frequency access complete video files and archived data are stored in large-capacity storage media. At the same time, data is classified and stored according to video type and data category to improve data retrieval efficiency. Data integrity maintenance: Real-time integrity checks are performed on stored video files and structured knowledge data, and the storage status of data is checked regularly to prevent data loss, damage or tampering. Incomplete data detected is marked and a data recovery mechanism is triggered to ensure the integrity and availability of stored data. Knowledge index building block: Basic Feature Index Construction: Based on basic features such as keywords, tags, and unique video identifiers in structured knowledge data, a basic feature index is constructed according to the system's preset index specifications to enable basic and rapid data retrieval and provide index support for keyword retrieval functions; Inverted index building: For keywords and tags in structured chapter data, a professional inverted index is built, which associates keywords and tags with the corresponding unique chapter timestamps of the video, realizing a fast mapping from keywords to specific chapters of the video, and improving the accuracy and efficiency of keyword retrieval; Semantic Vector Index Construction: The semantic vectors of chapter summary text content in structured knowledge data are transformed into semantic vectors, the generated semantic vectors are stored and a semantic vector index is constructed, and the semantic vectors are associated with the corresponding video segments to provide index support for the semantic retrieval function of similar segments across videos; Multi-dimensional index integration: The basic feature index, inverted index, and semantic vector index are integrated to form a unified multi-dimensional index system. The association mapping between the indexes is established to ensure that different search methods can quickly locate the corresponding video files and structured knowledge data through the index system. Knowledge retrieval and reuse module: Keyword retrieval service: Receives keyword retrieval requests from the system, performs matching analysis on the search keywords based on the constructed inverted index and basic feature index, quickly locates the video files and specific chapters containing the keywords, feeds back the matching results to the requesting end, and returns the corresponding structured knowledge data and video access entry points; Cross-video semantic retrieval service: Receives semantic retrieval requests from the system, converts the retrieved text content into semantic vectors, calculates the similarity between the vector and all stored semantic vectors based on the semantic vector index, filters out video segments with similarity that meet the threshold requirements, and feeds back the cross-video similar segment results and related structured knowledge data to the requesting end. Knowledge content reuse and sharing: Based on the user's operation request, and using the video's timestamp and structured knowledge data, a sharing link pointing to a specific time point in the video is generated. At the same time, the structured knowledge data such as the chapter summary and keywords of the corresponding segment are integrated into an embeddable content card, supporting the reuse of sharing links and content cards in multiple scenarios, and realizing the efficient dissemination and reuse of video knowledge content. Data Management and Maintenance Unit: Real-time data status monitoring: Real-time monitoring of the operating status of storage media and data access status, recording data access frequency, retrieving storage location and other information, and providing timely warnings of abnormal access behavior and storage media failure status to ensure the security of data storage and access. Data update and archiving: Receives structured knowledge data optimized by the multimodal analysis engine, updates and replaces the original data, and archives data that has exceeded the preset usage period and has a low access frequency. The archived data is then migrated to a dedicated archiving storage medium, freeing up the resources of the main storage medium and improving the overall storage efficiency of the system. Data backup and recovery: According to the system's preset backup strategy, the system performs regular full and incremental backups of core video files and structured knowledge data, and stores the backup data on off-site storage media. When data in the main storage media is lost or damaged, data can be quickly recovered through the backup data, ensuring data security and continuity. Furthermore, the data association mapping technology is based on a unique identifier binding mechanism. It uses the unique identifier generated by the video access module for each video file as the core key value, embedding this identifier into the core fields of the structured knowledge data. The module parses the unique identifier in the structured knowledge data and precisely matches it with the identifier in the video file repository, establishing a bidirectional association mapping relationship between the structured knowledge data and the original video file. The core of this technology lies in the uniqueness and non-repeatability of the identifier, ensuring that each set of structured knowledge data can accurately match the corresponding original video file, achieving a one-to-one association between data, and providing accurate association basis for subsequent retrieval and retrieval. Inverted index construction technology is the core technology for achieving efficient keyword retrieval. This technology uses keywords and tags in structured chapter data as index items, and uses the index items and the corresponding video's unique identifier, such as the chapter's start and end timestamps, as index values to build a mapping relationship from index items to index values. Unlike traditional forward indexes, inverted indexes use keywords as the retrieval entry point, which can directly locate all videos and specific chapters containing the keyword without having to traverse all data. This significantly improves the efficiency and accuracy of keyword retrieval, while also supporting combined retrieval of multiple keywords to meet complex retrieval needs. Semantic vector retrieval technology, based on a pre-trained natural language processing model and similarity calculation algorithm, enables semantic retrieval of similar segments across videos. First, the pre-trained natural language processing model transforms chapter summaries and text content from structured knowledge data into high-dimensional semantic vectors, which accurately represent the semantic information of the text. When a semantic retrieval request is received, the retrieval text is also transformed into a semantic vector. The cosine similarity algorithm is then used to calculate the similarity value between the retrieval vector and all stored semantic vectors. The core calculation formula is as follows: ,in, To retrieve the similarity value between the vector and the stored vector, the value ranges from 0 to 1, with the closer the value is to 1, the higher the semantic similarity. To retrieve the semantic vector of the text, To store the semantic vector of the text, The dot product of two semantic vectors. for The length of the mold, for The modulus length; The system presets a similarity threshold and uses the video segments corresponding to the stored vectors with similarity values higher than the threshold as the search results, thereby realizing semantic-based cross-video similar segment retrieval. Distributed data storage technology is based on a clustered storage architecture. It divides video files and structured knowledge data into blocks according to preset rules. Different data blocks are allocated to different nodes in the distributed storage cluster for storage, and multiple copies of each data block are generated and stored on different nodes. This technology improves data storage capacity and retrieval efficiency through parallel storage and access of multiple nodes. At the same time, the failure of a single node will not lead to data loss, and normal data access can be achieved through the replica data of other nodes, thus improving the reliability and fault tolerance of data storage. Data backup and recovery technology is based on a strategy that combines incremental backup and full backup. Full backup completely copies and stores all core data in the system, while incremental backup only copies and stores data that has been updated or modified since the last backup. This technology reduces the storage space and backup time of backup data while ensuring the integrity of the backup through a reasonable backup strategy. Data recovery technology is based on the data block storage information and backup mapping relationship. When the master data fails, the data in the backup medium is used to splice and restore the data according to the block mapping relationship, quickly restoring the system data to the state before the failure, ensuring the continuity and availability of data.
[0017] In this embodiment, the intelligent interactive player is the core of the terminal interaction of the multimodal video structured analysis and intelligent playback system based on the semantic timeline. It is a key carrier connecting the structured knowledge base and user operations. It can accurately retrieve video files and their associated structured knowledge data from the structured knowledge base, and realize the deep linkage between video playback and structured knowledge content with the semantic timeline as the core link. It provides users with an integrated intelligent playback experience such as chapter navigation subtitles and synchronized knowledge card display. It transforms the structured analysis results of the system into intuitive terminal interactive functions, which is an important link in realizing the value of the system. The intelligent interactive player is primarily responsible for loading specified raw video files and corresponding structured knowledge data from a structured knowledge base. It precisely binds and visualizes structured data such as chapter headings, subtitles, text summaries, and keywords with the video timeline. The player responds to various user playback operations, enabling normal playback, pause, fast forward, and rewind. Based on a semantic timeline, it provides precise chapter navigation, real-time synchronized subtitle rendering, and linked display of knowledge content. Furthermore, the player allows users to customize subtitle styles and highlights the current chapter in real-time during video playback, enabling users to efficiently acquire core knowledge while watching videos. This dual experience of video playback and knowledge acquisition fully leverages the terminal value of the system's structured analysis. The intelligent interactive player includes: Semantic timeline rendering module: Chapter directory visualization: Read the structured chapter data from the structured knowledge data, and display the chapter titles in a hierarchical directory format in a designated area of the playback interface according to the chronological order of the chapters, clearly presenting the semantic chapter division structure of the video; Binding directory nodes to the timeline: Each chapter directory node is precisely bound to its corresponding start and end timestamps, establishing a one-to-one mapping relationship between directory nodes and the video timeline, ensuring that the directory nodes accurately correspond to the specific time periods of the video; Timeline interface rendering adaptation: Based on the size of the playback interface and the user's operating habits, the display style of the semantic timeline is adapted and rendered, including the display style of the timeline's scale markings and chapter nodes, so that the integrated display of the timeline and chapter table of contents is more in line with the user's visual experience. At the same time, it supports the zoom operation of the timeline to meet the user's needs for fine-grained viewing of chapter divisions. Linkage control module: Chapter navigation jump response: Real-time capture of user clicks on chapter directory nodes, and based on the timestamp information bound to the directory nodes, control the video player to quickly jump to the start time of the corresponding chapter and start playback, achieving precise chapter navigation; Real-time highlighting of the current chapter: During video playback, the current playback timestamp of the video is obtained in real time, and the timestamp is matched with the time range of each chapter. The chapter directory node to which the current timestamp belongs is automatically highlighted, so that users can clearly know the semantic chapter position of the current video playback. Playback status synchronized with timeline: The playback status of the video, such as pause, resume, fast forward, and rewind, is synchronized with the semantic timeline in real time. When the video playback status changes, the current progress indicator on the timeline is also updated synchronously to ensure the consistency between the timeline progress and the video playback progress. Subtitle rendering module: Precise subtitle text retrieval: The current playback timestamp of the video is obtained in real time, and the corresponding subtitle text content is accurately read from the standard subtitle file based on the timestamp, ensuring that the subtitle text and the audio content of the video are completely matched. Real-time synchronized subtitle rendering: The retrieved subtitle text is rendered and displayed in the subtitle display area of the playback interface in real time. When the video playback timestamp changes, the subtitle text is updated synchronously to achieve millisecond-level synchronization between subtitles and video audio. At the same time, the subtitle is automatically cleared and the next subtitle is loaded after the subtitle display is completed. Customizable subtitle style: Provides users with subtitle style adjustment functions, supporting custom settings for font size, color, transparency, and display position. User adjustments take effect in real time, meeting the subtitle viewing habits of different users; Knowledge card display module: Real-time display of chapter knowledge: During video playback, based on the currently highlighted chapter node, the summary and keywords of the corresponding chapter are retrieved from the structured knowledge data and presented in the form of knowledge cards in the knowledge display area of the playback interface, allowing users to acquire the core knowledge of the chapter while watching the chapter content. Video summary retrieval: Supports manual triggering by users. When users need to view the overall knowledge framework of a video, they can retrieve the overall summary and key points at the video level through specified operations, which are displayed in the form of knowledge cards to help users quickly grasp the core content of the video. Knowledge card display adaptation: The display style of knowledge cards is adaptively adjusted according to the overall layout and display size of the playback interface. When the size of the playback interface changes, the size and position of the knowledge cards will be adapted synchronously to ensure that the display of knowledge content will not affect the normal viewing of the video, while ensuring the readability of the knowledge content. Furthermore, semantic timeline binding technology is the core technology for the player to implement chapter navigation and linked display. This technology establishes a mapping model between directory node attributes and video timestamps, deeply binding the chapter identifier start and end timestamps in the structured chapter data with the chapter directory nodes. The core of the technology lies in achieving accurate mapping between directory nodes and the timeline. The matching degree of the mapping relationship is evaluated by the timestamp matching error, and the calculation formula is as follows: ,in, For timestamp matching error, This is the actual timestamp of the video playback. The theoretical timestamp bound to the directory node; The system presets a timestamp matching error threshold. When the value is less than the threshold, the mapping relationship is deemed valid, ensuring the binding accuracy between directory nodes and the timeline, and providing technical support for accurate chapter navigation. Real-time interactive response technology is implemented based on an event-driven architecture, which divides various user operations into different interactive events, including chapter click events, subtitle style adjustment events, and knowledge card retrieval events. During operation, the player listens for the triggering of various interactive events in real time. When an event is detected, the event scheduling engine quickly allocates processing resources and calls the corresponding functional modules to process the event. At the same time, an asynchronous processing method is adopted to ensure that the event processing does not affect the normal playback of the video, achieving millisecond-level response to user operations and improving the smoothness of interaction. The precise subtitle timeline rendering technology is based on frame-level timestamp matching. It splits each subtitle text segment in the standard subtitle file into frames, matching each subtitle character with its corresponding start and end timestamps, ensuring that each subtitle character matches its corresponding video playback frame. During video playback, the player precisely retrieves and renders the corresponding subtitle content based on the current playback frame position. Simultaneously, a dynamic buffering mechanism reduces subtitle synchronization deviations caused by video playback stuttering. The formula for calculating the subtitle synchronization rendering accuracy is as follows: ,in, To improve the accuracy of synchronized subtitle rendering, This refers to the number of subtitle renderings per unit of time. For the first Subtitle rendering timestamp matching error. This refers to the average display duration of the subtitles; The value ranges from 0 to 1. The closer the value is to 1, the higher the accuracy of the synchronized rendering of the subtitles, ensuring the synchronization between the subtitles and the video audio. The knowledge content linkage display technology is based on semantic chapter time range matching. This technology establishes an association index between video playback timestamps and structured knowledge data, using the time range of each chapter as the index condition. When the video playback timestamp falls within the time range of a certain chapter, the player quickly retrieves the summary and keywords of that chapter from the structured knowledge data through the association index. At the same time, a preloading mechanism is adopted to preload the knowledge content of subsequent chapters during video playback. When chapters switch, the knowledge card content can be seamlessly switched, avoiding delays in knowledge content display. The adaptive rendering technology is based on fluid layout and responsive design. The player divides the playback interface into multiple independent display areas, such as the video playback area, semantic timeline area, subtitle display area, and knowledge card area, and sets adaptive layout rules for each area. When the size of the playback interface changes, the system calculates the optimal display size for each area in real time, automatically adjusts the size and position of each area according to the layout rules, and simultaneously adapts the content style in each area to ensure that the display effect of the video timeline, subtitles, and knowledge cards is not affected by the change in interface size, thus improving the viewing experience on different display devices.
[0018] In this embodiment, the system adopts a distributed asynchronous task architecture. This architecture is the core of the underlying scheduling and resource management of a multimodal video structured analysis and intelligent playback system based on a semantic timeline. It forms the architectural foundation supporting the system's efficient processing of multiple batches and types of video processing tasks. This architecture provides asynchronous orchestration, distributed execution, real-time progress monitoring, and result caching and reuse capabilities for the entire video structured analysis task process. It decomposes the complex video processing workflow into independent, parallelizable asynchronous subtasks, achieving parallel processing through reasonable task scheduling and node allocation. Simultaneously, it optimizes system resource utilization by combining progress feedback and result caching mechanisms, ensuring stable and efficient system operation under high-concurrency video processing requests. This gives the system's video processing capabilities good scalability and adaptability, including: Task Scheduling Center: Request reception processing: Real-time reception of video structured analysis and processing requests initiated by the video access module, extraction of key information such as video unique identifier, processing requirements, and priority from the requests, format verification and validity determination of the requests, rejection of invalid requests and feedback of reasons to the initiating end, and unified registration and management of valid requests; Video processing task breakdown: Following the process steps of the system's video structured analysis, the complete video processing task is broken down into multiple independent asynchronous subtasks. The subtasks cover audio and video preprocessing, speech recognition and text processing, temporal semantic segmentation, subtitle generation, summary generation, quality control, index construction, etc. Each subtask has independent execution logic and input / output standards, and can be assigned to a computing node for execution independently. Subtask orchestration and node scheduling: Logically orchestrate the subtasks according to their execution dependencies to determine their execution order. Simultaneously, collect real-time information such as resource utilization, load status, and idle computing power of each computing node in the system. Allocate the orchestrated subtasks to idle computing nodes to achieve distributed parallel execution of subtasks and ensure maximum utilization of computing resources. Task lifecycle management: Manages the entire lifecycle of all subtasks and the overall task, covering the status of task creation, allocation, execution, pause, completion, and abnormal termination. Updates task status information in real time, collects the results of completed subtasks, and reassigns or triggers retry mechanisms for abnormally terminated subtasks to ensure the smooth completion of the overall task. Progress Management Module: Subtask status acquisition: Establish real-time communication links with each computing node to periodically collect the execution status of each asynchronous subtask, including status such as pending execution, in execution, completed, abnormal, paused, etc. At the same time, collect detailed information such as the progress ratio, time elapsed, and remaining estimated time of the subtask in execution to ensure the real-time and accuracy of subtask status data. Overall progress calculation and aggregation: Based on the execution weight and current progress of each subtask, the progress data of all subtasks are aggregated and calculated to obtain the real-time progress of the overall video processing task. At the same time, the overall execution time and the time of each stage are statistically analyzed to form a complete progress data system. Real-time progress feedback: The aggregated overall task progress, the execution status of each subtask, and detailed progress data are fed back to the front-end interface and relevant system modules in real time through standardized interfaces. It supports the visualization of progress data, allowing users and system modules to keep track of the progress of video processing in real time. Abnormal Status Warning and Handling: Real-time identification and graded warning of abnormal status of subtasks are performed. According to the type and severity of the abnormality, the corresponding handling mechanism is triggered. For minor abnormalities, the subtask retry mechanism is triggered. For serious abnormalities, the execution of related subtasks is suspended in time and feedback is sent to the task scheduling center. At the same time, the abnormal information and warning content are synchronized to the front end to realize the timely detection and handling of abnormalities. Result caching and reuse module: Standardized storage of cached data: Receive structured knowledge data that has passed the quality control module's verification, along with the corresponding unique video identifier, basic video metadata, processing result feature values, and other information. Store the data in a standardized manner according to the system's preset cache format, establish a precise association between cached data and video files, and assign a unique index identifier to the cached data for easy and quick retrieval. Video content similarity determination: When a new video processing request is received, the feature values of the video file in the request are extracted and similarity calculations are performed with the feature values of the videos already stored in the cache. This determines whether the video in the new request is the same video that has already been processed or a video with highly similar content, providing a basis for the reuse of cached results. Fast retrieval of cached results: If the similarity determination shows that the video requested by the new request is the same video or a video with highly similar content, the corresponding structured knowledge data is retrieved directly from the cache without having to initiate the distributed task execution process again. The cached results are quickly fed back to the request initiator, realizing the reuse of processing results. Cache lifecycle management: According to the system's preset caching strategy, the stored cached data is managed throughout its entire lifecycle, including cache creation, updating, eviction, and cleanup. Based on factors such as video access frequency, cache duration, and system cache resource utilization, a reasonable cache eviction mechanism is adopted to clean up invalid cached data, release system storage resources, and ensure efficient utilization of cache resources.
[0019] A method for multimodal video structured analysis and intelligent playback based on a semantic time axis, executed by a multimodal video structured analysis and intelligent playback system based on a semantic time axis, includes the following steps: S1: Receives the original video file, performs audio and video preprocessing, and extracts standard audio track data. Execution entity: The video access module and the audio / video preprocessing unit of the multimodal analysis engine work together to execute the operation; Specific processing steps: The video access module first completes the multi-format reception, integrity, and validity verification of the original video file. For the verified video file, it performs basic standardized encapsulation and generates a unique identifier. Then, it initiates the processing task and transmits the video file to the audio / video preprocessing unit of the multimodal analysis engine. The audio / video preprocessing unit separates the audio and video streams of the received video file, extracts the pure audio track data, and then performs standardization processing on the audio track data, including operations such as unifying the audio sampling rate, merging channels, normalizing volume, and reducing audio noise. Simultaneously, it performs precise calibration between the audio track data and the original video timeline, eliminating time deviations between the audio track and the video. Execution objective: To achieve compliant access to raw videos and standardized extraction of audio tracks, providing a uniform and high-quality pure audio track data source for subsequent speech recognition, and ensuring complete synchronization between the audio track data and the video timeline; Output: Standard audio track data precisely bound to the original video timeline, along with a unique video identifier and basic metadata; S2: Perform speech recognition on the standard audio track data to generate an initial text sequence with timestamps: Execution unit: Speech recognition and text processing unit of the multimodal analysis engine; Specific processing content: This unit performs frame-by-frame speech recognition processing on the standard audio track data output by S1. Relying on an end-to-end speech recognition model, it converts the audio content into text information. At the same time, it binds a corresponding start timestamp and end timestamp to each recognized text segment. The timestamp accuracy is consistent with the video timeline. After recognition is completed, all timestamped text segments are spliced together in the order of audio playback to form a continuous initial text sequence. Execution objective: To achieve the conversion of audio content into text content, while preserving the temporal dimension information corresponding to the text, audio track, and video, laying the foundation for subsequent text processing and semantic analysis; Output: A time-stamped initial text sequence precisely linked to the video timeline, with each text segment accompanied by its corresponding start and end timestamps; S3: Clean and semantically segment the initial text sequence to obtain a cleaned text stream: Execution unit: Speech recognition and text processing unit of the multimodal analysis engine; Specific processing steps: This unit first cleans the initial text sequence, removing invalid characters, typos, and duplicate content generated during the recognition process, correcting recognition errors caused by unclear speech and background noise, and filtering redundant information such as interjections and pauses that have no actual semantic meaning; then, it performs semantic sentence segmentation based on natural language processing technology, breaking away from the mechanical sentence segmentation method based solely on audio pauses, and combining semantic logic and natural language expression habits to divide text paragraphs and sentences, ultimately forming text data arranged in chronological order and standardized by both timestamps and semantic logic; Execution objective: To improve the accuracy and effectiveness of text data, to ensure that the sentence segmentation of the text stream conforms to semantic logic, and to provide high-quality text input for subsequent temporal semantic segmentation and subtitle generation; Output: A clean text stream with coherent semantics and logic, standardized timestamps, and no redundant information. Each statement in the text stream retains its corresponding timestamp information. S4: Perform temporal semantic analysis on the purified text stream, combining semantic understanding and preset temporal constraint rules to determine the chapter segmentation points of the video and generate structured chapter data, specifically including: S41: Semantic segmentation of the cleaned text stream is performed using a natural language processing model to obtain preliminary semantic paragraphs and their boundaries: A natural language processing model based on topic modeling and key concept aggregation is used to perform deep semantic understanding of the cleaned text stream, mine the semantic topic change patterns in the text stream, identify the key time nodes of topic switching, and divide potential semantic paragraphs based on this, and determine the preliminary time boundaries of each semantic paragraph to form preliminary chapter segmentation results. S42: Obtain a predefined set of constraint rules, including minimum chapter duration, maximum chapter duration, and boundary smoothness parameters. The system retrieves a predefined set of temporal constraint rules from the system configuration. This set sets hard standards for chapter segmentation, including the minimum duration threshold for a single chapter, the maximum duration threshold for a single chapter, and chapter boundary smoothness parameters. It also clarifies that chapters must maintain temporal continuity and that semantic paragraphs must achieve full coverage of the cleaned text stream, providing a constraint basis for optimizing the initial segmentation results. S43: To maximize semantic coherence and satisfy all constraint rules, optimize and adjust the initial semantic segment boundaries, and output the final chapter start and end timestamps: Apply constraint optimization algorithm to refine the initial semantic segment boundaries obtained in S41. The optimization goal is to maximize the overall semantic coherence of the video, while ensuring that all adjusted chapters satisfy all temporal constraint rules in S42. After completing the chapter boundary optimization, generate corresponding chapter titles, chapter summaries and chapter keywords for each chapter, and integrate all information to form complete structured chapter data. S5: Based on the cleaned text stream and its timestamps, generate a standard subtitle file aligned with the video timeline. Execution entity: The subtitle generation unit of the multimodal analysis engine; Specific processing steps: Based on the cleaned text stream output by S3, this unit extracts the start and end timestamps corresponding to each statement, achieving millisecond-level precise mapping between text content and video timeline; then, according to the commonly used subtitle file format specifications, it encapsulates the mapped text statements and timestamp information to generate standard format subtitle files; at the same time, it performs lightweight optimization on the subtitle text, simplifying overly long statements and retaining core semantics, so that the subtitle display rhythm conforms to the user's viewing habits; Execution objective: To generate standard subtitle files that are fully synchronized with the video and audio, ensuring that the display and disappearance times of the subtitles are accurately matched with the video and audio, and that the format conforms to common standards and can be directly parsed and rendered by intelligent interactive players; Output: A standard subtitle file precisely aligned with the video timeline, with text content completely synchronized with the video audio, and format meeting the system's parsing requirements; S6: Generate summary description information for the video based on structured chapter data: Execution entity: The summary generation unit of the multimodal analysis engine; Specific processing content: This unit first extracts chapter-level information from the structured chapter data output by S4, aggregates the summaries and keywords of each chapter to form a collection of chapter-level knowledge points, and retains the segmented core logic of the video content; then, relying on the pre-trained natural language summarization model, it merges, refines and integrates the collection of chapter-level knowledge points to generate a video-level overall summary that can fully cover the core content of the video, while extracting the core key points that run through the entire video. The overall summary and key points together constitute the summary description information of the video. Execution objective: To achieve multi-granularity knowledge extraction from video content, generate summary descriptive information that reflects the overall core content of the video, and provide a basis for users to quickly grasp the main idea of the video and for the system to realize knowledge linkage display; Output results include an overall video summary and a summary description of the core key points, which are semantically consistent with the structured chapter data and are related to the video timeline; S7: Link and store the original video files, structured chapter data, subtitle files, and summary description information, and create an index: Execution entity: The quality control module and the structured knowledge base work together to execute the task; The specific processing steps are as follows: The structured chapter data, subtitle files, and summary description information output by the multimodal analysis engine are first transmitted to the quality control module. After field, format, and timestamp verification by the rule validation submodule, and semantic consistency assessment by the consistency evaluation submodule, and if necessary, corrections are made through a manual review platform, resulting in qualified structured knowledge data. Subsequently, the structured knowledge base receives the original video files and the qualified structured knowledge data, establishing a one-to-one correspondence between the original video and the structured knowledge data through the video's unique identifier, achieving precise data association. Finally, the structured knowledge base employs a hierarchical and categorized storage strategy to ensure secure data storage. Based on the keywords, tags, semantic vectors, and other features of the structured knowledge data, a multi-dimensional indexing system, including inverted indexes and semantic vector indexes, is constructed to achieve efficient data retrieval and access. Execution objective: To achieve the associated storage of raw video and structured knowledge data, ensuring the integrity, relevance and security of the data, while establishing a multi-dimensional index to provide efficient data retrieval support for subsequent intelligent playback, knowledge retrieval and reuse; Output: A collection of video data and structured knowledge data that has been constructed with associated storage and multi-dimensional indexing. It can be retrieved in real time by the intelligent interactive player and supports functions such as keyword search and cross-video semantic search.
[0020] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multimodal video structured analysis and intelligent playback system based on a semantic time axis, characterized in that: include: Video access module, multimodal analysis engine, structured knowledge base, and intelligent interactive player; The video access module is used to receive and store the original video file and start the processing task; The multimodal analysis engine is used to process the original video files and generate structured knowledge data based on a unified timeline. The structured knowledge data includes at least: a chapter directory based on semantic segmentation, subtitle text precisely aligned with the timeline, and video or chapter-level summaries and keywords. The structured knowledge base is used to store and index structured knowledge data, and to establish the association between video media files and structured knowledge data; The intelligent interactive player is used to load video files and their associated structured knowledge data, and realizes video playback, chapter navigation, subtitle synchronization and knowledge content linkage display based on the semantic timeline.
2. The multimodal video structured analysis and intelligent playback system based on semantic time axis according to claim 1, characterized in that: The multimodal analysis engine comprises the following components connected in sequence: The audio and video preprocessing unit is used to extract the audio track from the video and generate standardized data to be processed. The speech recognition and text processing unit is used to recognize the audio track, generate initial text with timestamps, and clean and segment the initial text to obtain a cleaned text stream. The temporal semantic segmentation unit is used to perform semantic understanding on the cleaned text stream, identify topic boundaries, and optimize it in combination with predefined temporal constraint rules to generate structured chapter data containing start timestamp, end timestamp, chapter title, chapter summary and chapter keywords; The subtitle generation unit is used to generate standard format subtitle files based on the cleaned text stream and its timestamp; The summary generation unit is used to generate video-level overall summaries and key points based on structured chapter data.
3. The multimodal video structured analysis and intelligent playback system based on semantic time axis according to claim 2, characterized in that: The steps for semantic understanding and constraint optimization performed by the temporal semantic segmentation unit include: Topic modeling and key concept aggregation are performed on the cleaned text stream to identify potential semantic paragraphs; The constraint optimization algorithm is applied to adjust the boundaries of the initially identified semantic paragraphs. The constraint optimization algorithm considers at least the following rules: the shortest duration threshold of a single chapter, the longest duration threshold of a single chapter, the temporal continuity between chapters, and the coverage of the semantic paragraphs to the whole text. Output the final chapter time boundary and corresponding semantic content that satisfy all constraint rules.
4. The multimodal video structured analysis and intelligent playback system based on semantic time axis according to claim 2, characterized in that: It also includes a quality control module for verifying and auditing the output of the multimodal analysis engine, which includes: The rule validation submodule is used to validate the field integrity, timestamp order, and legality of structured chapter data; The consistency assessment submodule is used to evaluate the semantic consistency among headings, abstracts, and keywords within the same chapter. The manual review workbench provides a video playback interface based on timeline positioning, allowing users to visually verify and correct machine-generated chapter boundaries, titles, and abstracts, and record all corrections.
5. The multimodal video structured analysis and intelligent playback system based on semantic time axis according to claim 1, characterized in that: The structured knowledge base also includes a knowledge retrieval and reuse module, used for: Based on keywords and tags in structured chapter data, an inverted index is built to provide keyword search functionality; Based on semantic vectors from chapter summaries or text content, it provides semantic retrieval functionality for similar segments across videos. Generate shareable links or embeddable content cards that point to specific points in time in a video.
6. The multimodal video structured analysis and intelligent playback system based on semantic time axis according to claim 1, characterized in that: The intelligent interactive player includes: The semantic timeline rendering module is used to visually present the chapter directory on the playback interface and bind the directory nodes to the video timeline; The linkage control module is used to respond to the user's click operation on the chapter directory node, control the video to jump to the corresponding time point for playback; and during the video playback, the chapter node to which the current time point belongs is highlighted in real time; The subtitle rendering module is used to read and render the corresponding text from the subtitle file based on the current playback timestamp, and supports style adjustment; The knowledge card display module is used to display the summary, keywords, or overall video summary of the current chapter in the relevant area of the playback interface.
7. The multimodal video structured analysis and intelligent playback system based on semantic time axis according to claim 1, characterized in that: The system adopts a distributed asynchronous task architecture, including: The task scheduling center is used to receive processing requests, decompose the video processing flow into multiple asynchronous subtasks, and orchestrate them. The progress management module is used to track and provide real-time feedback to the front end on the execution status of each sub-task and the overall progress. The result caching and reuse module is used to cache processed video files and their corresponding structured knowledge data; when a request to process the same video file or video files with highly similar content is received, the cached result is returned directly.
8. A method for multimodal video structured analysis and intelligent playback based on a semantic time axis, characterized in that: The method is executed by the semantic time-axis-based multimodal video structured analysis and intelligent playback system according to any one of claims 1-7, and includes the following steps: S1: Receive the original video file, perform audio and video preprocessing, and extract standard audio track data; S2: Perform speech recognition on standard audio track data to generate an initial text sequence with timestamps; S3: Clean and semantically segment the initial text sequence to obtain a purified text stream; S4: Perform temporal semantic analysis on the purified text stream, combine semantic understanding and preset temporal constraint rules to determine the chapter segmentation points of the video, and generate structured chapter data; S5: Based on the purified text stream and its timestamp, generate a standard subtitle file aligned with the video timeline; S6: Based on the structured chapter data, generate summary description information for the video; S7: Link and store the original video files, structured chapter data, subtitle files, and summary description information, and create an index.
9. The multimodal video structured analysis and intelligent playback system and method based on semantic time axis according to claim 1, characterized in that: Step S4, which combines semantic understanding with preset temporal constraint rules to determine chapter segmentation points, specifically includes: S41: Semantic segmentation of the cleaned text stream is performed using a natural language processing model to obtain preliminary semantic paragraphs and their boundaries; S42: Obtain a predefined set of constraint rules, which includes minimum chapter duration, maximum chapter duration, and boundary smoothness parameters; S43: With the goal of maximizing semantic coherence and satisfying all constraint rules, optimize and adjust the initial semantic paragraph boundaries, and output the final chapter start and end timestamps.