Video content structured processing method and system fusing voice large model and visual semantic large model
By integrating large-scale speech models and large-scale visual semantic models for video content structuring, this method solves the problem of video content being difficult to structure in existing technologies, achieving efficient and accurate video content parsing and intelligent management, and is suitable for professional scenarios such as education and training.
Patent Information
- Application Number
- CN202511688596.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2025-12-19
AI Technical Summary
Existing technologies lack effective speech recognition and deep semantic analysis when processing video content, making it difficult to achieve structured and semantic processing. They perform poorly, especially in professional scenarios such as education and corporate training. Furthermore, existing solutions have shortcomings in recognition accuracy, multi-level semantic understanding, and structured output.
This method integrates large-scale speech models and large-scale visual semantic models to structure video content. Through steps such as audio extraction, speech recognition, semantic segmentation, and video cutting, it transforms video content into structured data. This includes audio stream standardization, generation of word-level and character-level timestamps, semantic summarization, and precise segmentation of video clips. Error correction is then performed using a large language model.
It enables automated and precise processing of video content, improves the organization efficiency and reuse rate of video resources, supports multiple information retrieval functions, is suitable for intelligent applications in education, training and other fields, and improves the synchronization of timestamps and the accuracy of output content.
Smart Images

Figure CN121166976A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video processing, in particular to a video content structured processing method and system fusing a voice large model and a visual semantic large model. BACKGROUND
[0002] With the development of the Internet and the lowering of the threshold for multimedia creation, online video resources are growing explosively - mainstream platforms add tens of thousands of hours of content every day, and the education, enterprise conference, and short video fields have also accumulated massive data. How to efficiently extract valuable information and achieve automated processing and knowledge organization has become an important direction in the field of artificial intelligence. However, traditional video understanding methods rely on manual annotation, keyword matching, and visual feature classification, and voice information is only used to generate subtitles without deep semantic analysis. When faced with complex content, it is difficult to meet the needs of structuring, semanticizing, and retrievability, and performs poorly in professional scenarios such as educational video analysis, enterprise training data organization, and knowledge base construction.
[0003] In recent years, deep learning has driven video understanding capabilities to leap forward: voice large models represented by Whisper can achieve high accuracy and multi-language speech-to-text, and large language models such as GPT and Qwen also have strong semantic modeling capabilities, providing new possibilities for video intelligent processing. However, existing solutions that combine speech recognition and large language models still lack a complete solution that takes into account recognition accuracy, multi-level semantic understanding, and structured output. There are obvious shortcomings in structured expression, automated processing, timestamp alignment accuracy, and recognition reliability. Therefore, the present application proposes a video content structured processing method and system fusing a voice large model and a visual semantic large model to solve these problems. SUMMARY
[0004] Technical problems to be solved To solve the problems in the background art, the present application provides a video content structured processing method and system fusing a voice large model and a visual semantic large model.
[0005] Technical scheme To achieve the above purpose, the present application realizes the following technical scheme: a video content structured processing method fusing a voice large model and a visual semantic large model, comprising: Step one: taking an original video file as input, using a professional audio and video processing tool to separate an audio stream from the original video file; performing standardization processing on the audio stream to obtain a standardized audio file; if the original video file has no audio track, triggering a preset exception processing mechanism and outputting an error prompt; Step 2: Input the standardized audio file into the speech model, perform efficient batch transcription, and generate associated data containing word-level timestamps and corresponding text; perform traditional-to-simplified conversion on the transcribed text and output it as simplified Chinese subtitle text; at the same time, optimize hardware resource usage through the video memory management mechanism, and finally obtain simplified Chinese subtitle text data with word-level timestamps. Step 3: Input the simplified Chinese subtitle text data with word-level timestamps into the visual semantic big model. The visual semantic big model will adaptively perform semantic segmentation according to the semantic logic of the content to obtain multiple complete sub-segments of themes. The word-level timestamps will be further refined into character-level timestamps, and a semantic summary will be generated for each sub-segment. Finally, semantic sub-segments with character-level timestamps and corresponding semantic summaries will be obtained. Step 4: Based on the semantic sub-segments with character-level timestamps, perform precise video segmentation to generate independent video segments with embedded timestamp information; construct a structured metadata object containing segment identifier, start and end timestamps, duration, corresponding subtitle text, semantic summary, and original video file name; after performing error correction processing on the subtitle text, integrate the independent video segments and the structured metadata object to finally output standard structured data.
[0006] Preferably, the specific steps for performing normalization processing on the audio stream are as follows: The multi-channel audio stream is synthesized into a mono, the mono audio is then uniformly resampled to a preset standard sampling rate, and finally converted into a preset standard audio format to obtain the standardized audio file.
[0007] Preferably, the specific steps for triggering the preset exception handling mechanism are as follows: Check if the original video file contains an audio track. If not, output an error message and terminate the current audio extraction process.
[0008] Preferably, the specific steps for performing efficient batch transcription are as follows: Load the large speech model, configure the corresponding language decoder according to the language type of the video to be processed, and then divide the standardized audio file into batch audio segments according to the preset duration, and input them into the large speech model in sequence.
[0009] Preferably, the specific steps of the video memory management mechanism are as follows: Monitor the memory usage of the large speech model during runtime. When the memory usage reaches a preset threshold, release the cache of audio segments that have completed transcription tasks, and control the concurrency of transcription tasks through task queue scheduling.
[0010] Preferably, the specific steps for the large visual semantic model to adaptively perform semantic segmentation are as follows: Based on the preset structured prompt template, the semantic discontinuity in the simplified Chinese subtitle text with word-level timestamps is identified, and a sub-clip with complete theme is generated by taking the semantic discontinuity as the segmentation boundary.
[0011] Preferably, the specific steps of further refining the word-level timestamps to character-level timestamps are as follows: The total duration corresponding to the word-level timestamps is calculated, the number of characters contained in the word is counted, the duration is allocated according to the character proportion, the start and end times of each character are determined, and the character-level timestamps are obtained.
[0012] Preferably, the specific steps of performing accurate video cutting are as follows: A preset professional audio and video processing tool is used to crop the original video according to the character-level timestamp interval of the semantic sub-clip, generate an independent video clip, and name the clip in the format of the original video file name plus the timestamp.
[0013] Preferably, the specific steps of performing error correction processing on the subtitle text are as follows: A large language model is called to perform context semantic verification on the subtitle text in combination with a preset field professional vocabulary, and the incorrect content in the text is corrected.
[0014] The video content structured processing system fusing the speech large model and the visual semantic large model comprises: An audio extraction module: taking an original video file as input, using a professional audio and video processing tool to separate an audio stream from the original video file; performing standardized processing on the audio stream to obtain a standardized audio file; if the original video file has no audio track, triggering a preset exception handling mechanism and outputting an error prompt; A speech recognition module: inputting the standardized audio file into a speech large model to perform efficient batch transcription and generate associated data containing word-level timestamps and corresponding text; performing traditional and simplified conversion on the transcribed text to output unified simplified Chinese subtitle text; at the same time, optimizing hardware resource usage through a video memory management mechanism to finally obtain simplified Chinese subtitle text data with word-level timestamps; A video semantic segmentation module: inputting the simplified Chinese subtitle text data with word-level timestamps into a visual semantic large model, and performing semantic segmentation by the visual semantic large model according to content semantic logic to obtain multiple sub-clips with complete theme; further refining the word-level timestamps to character-level timestamps, and generating a semantic abstract for each sub-clip to finally obtain semantic sub-clips with character-level timestamps and corresponding semantic abstracts; The structured output module: according to the semantic sub-fragments with character-level time stamps, performing accurate video cutting to generate independent video fragments embedded with timestamp information; constructing a structured metadata object containing fragment identification, start and end timestamps, duration, corresponding subtitle text, semantic abstract, and original video file name; after error correction processing of the subtitle text, integrating the independent video fragments and the structured metadata object, and finally outputting standard structured data.
[0015] Advantages The present application has the following advantages: (1) The video content structured processing method and system fusing a large speech model and a large visual semantic model can convert original unstructured video content into structured data with hierarchical semantic structure through an automatic chapter division mechanism based on semantic understanding, thereby greatly simplifying subsequent content retrieval, labeling, and reuse processes. Compared with traditional methods relying on manual labeling or fixed time slicing, the present application can dynamically generate structured nodes according to actual semantic changes of video content, significantly improving the organization efficiency and reusability of video resources.
[0016] (2) The video content structured processing method and system fusing a large speech model and a large visual semantic model realizes full-process automatic processing from video input to structured output, covering key links such as audio extraction, speech recognition, semantic segmentation, video fragment cutting, and multi-modal information integration, and can complete high-quality content analysis without human intervention. This feature not only greatly reduces the labor cost of video content processing, but also improves the scalable deployment capability of the system, making it suitable for centralized processing of large-scale video resources.
[0017] (3) The video content structured processing method and system fusing a large speech model and a large visual semantic model enables the output structured content to support keyword indexing, abstract matching, knowledge point labeling, and other information retrieval functions, facilitating the construction of intelligent applications such as video knowledge bases, question and answer systems, teaching resource platforms, and other fields for education, training, and knowledge management. Compared with traditional systems that only provide original subtitle text, the semantic-level structured output provided by the present application greatly enhances the discoverability and operability of the content, improving user interaction experience and information acquisition efficiency.
[0018] (4) The video content structured processing method and system fusing a large speech model and a large visual semantic model ensures high synchronization between video pictures and text information by using a word-level / word-level granularity timestamp alignment mechanism. This mechanism can achieve millisecond-level accurate positioning, especially suitable for professional scenarios such as teaching explanation, online courses, and meeting records that require high time control. In addition, high-precision timestamps also provide reliable data support for subsequent video editing, jumping, labeling, and other tasks based on the time axis.
[0019] (5) The video content structured processing method and system of the fusion of the large speech model and the large visual semantic model has good scalability and flexibility through the modular system architecture design, supports distributed deployment, batch task processing and external interface calling and various running modes. The design enables the system to flexibly adapt to different scales and types of video processing requirements, which can be used for lightweight local deployment and seamlessly integrated into large enterprise-level video management systems to meet diversified business expansion and engineering landing requirements.
[0020] (6) The video content structured processing method and system of the fusion of the large speech model and the large visual semantic model introduces a double error correction mechanism based on the built-in professional vocabulary and the large language model in the speech recognition post-processing stage, which effectively reduces the wrong words, missing words, misreading and other problems in the ASR recognition process. Especially in the professional fields such as finance, law and medicine, which have very high requirements for content accuracy, the mechanism can significantly improve the quality of the final output text, avoid information misunderstanding or decision deviation caused by recognition errors, and enhance the robustness and credibility of the system.
[0021] Of course, implementing any product of the present application does not necessarily require all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 The flowchart of the video content structured processing method of the fusion of the large speech model and the large visual semantic model of the present application; Figure 2 The structural diagram of the video content structured processing system of the fusion of the large speech model and the large visual semantic model of the present application. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0024] The embodiments of the present application provide a technical solution: a video content structured processing method of the fusion of the large speech model and the large visual semantic model, as shown in Figure 1 The steps include: Step one: Take the original video file as input, and use professional audio and video processing tools to separate the audio stream from the original video file. Selecting a professional audio and video processing tool is because the audio stream in the original video may contain multi-channel, different encoding format information, and professional tools can ensure complete separation of the audio stream, avoid audio data loss or distortion caused by the separation process, and provide high-quality audio basis for subsequent speech recognition.
[0025] Perform standardization processing on the separated audio stream. The specific process is to combine the multi-channel audio stream into a single channel, then uniformly resample the single-channel audio to a preset standard sampling rate, and finally convert it to a preset standard audio format, to obtain a standardized audio file. The core reason for this processing is that the multi-channel of the original audio may interfere with the recognition logic of the speech large model, and the audio sampling rate and format recorded by different devices are different, while the speech large model is usually adapted to a specific standard sampling rate (such as 16kHz) and format (such as wav format). After standardization, the model recognition error can be greatly reduced, ensuring the stability and accuracy of the subsequent transcription process.
[0026] At the same time, before processing, it will be detected whether the original video file has an audio track. If not, output an error prompt and terminate the current audio extraction process. The purpose of this design is to avoid the system continuing to perform subsequent invalid operations when there is no audio track, reduce the waste of hardware resources, and at the same time, timely feedback the problem to the user, facilitate the user to check the integrity of the original video file, and ensure the efficiency and controllability of the entire processing process.
[0027] Step two: input the obtained standardized audio file into the speech large model to perform efficient batch transcription. The specific process is to load the speech large model, configure the corresponding language decoder according to the language type of the video to be processed, then divide the standardized audio file into batch audio segments according to the preset time length, input them into the speech large model one by one, and finally generate associated data containing word-level timestamps and corresponding text. Configuring the corresponding language decoder is because the video content may involve different languages (such as Chinese, English, etc.), and matching the decoder can ensure the language accuracy of text transcription, avoiding character garbled or semantic deviation; dividing the audio segments by the preset time length is because inputting too long audio at a time will occupy a lot of hardware video memory, and dividing and processing them one by one can improve the transcription efficiency and reduce the risk of video memory overload, meeting the needs of subsequent video memory management.
[0028] Perform complex and simple conversion on the text obtained by transcription, and output it as a simplified Chinese subtitle text. This is to ensure that the subsequent semantic segmentation and structured output use a unified text format, avoid semantic understanding deviation caused by the mixed use of complex and simple, and adapt to the reading habits of most users in most scenarios (such as enterprise knowledge management and education video analysis, which often use simplified Chinese characters), ensuring the universality of text information.
[0029] In this process, the use of hardware resources is also optimized through the video memory management mechanism. Specifically, the video memory usage during the running of the voice large model is monitored. When the video memory usage reaches a preset threshold, the cache of the audio segment that has completed the transcription task is released, and the transcription task concurrency is controlled through task queue scheduling. Monitoring the video memory usage is to prevent the video memory from overflowing and causing the program to crash. Releasing the cache can timely release the hardware space for subsequent segment processing. Task queue scheduling can balance the processing speed and hardware load, avoid hardware overload caused by multiple tasks running simultaneously, and ultimately ensure the stability and efficiency of the entire transcription link. The simplified Chinese subtitle text data with word-level timestamps is obtained, providing accurate text and time basis for semantic segmentation in step three.
[0030] Step three: input the obtained simplified Chinese subtitle text data with word-level timestamps into the visual semantic large model. The visual semantic large model performs semantic segmentation based on the content semantic logic. Specifically, based on the preset structured prompt template, the semantic discontinuity in the simplified Chinese subtitle text with word-level timestamps is identified, and multiple theme-complete sub-segments are generated with the semantic discontinuity as the segmentation boundary. The core reason for choosing the visual semantic large model is that it can combine the semantic information of the subtitle text with the implicit association of the video picture (for example, when the subtitle mentions "the radius of a circle", the video picture synchronously displays the graph of the circle). By judging the theme boundary based on semantic logic, compared with traditional fixed-length segmentation, this method can ensure that the sub-segment is not only semantically complete, but also synchronously matched with the video picture content, avoiding the misplacement of text themes and picture content, such as mixing "radius of a circle explanation" and "circumference of a circle calculation" into the same sub-segment.
[0031] Subsequently, the word-level timestamps are further refined into character-level timestamps. The specific process is to calculate the total duration corresponding to the word-level timestamps, count the number of characters contained in the word, allocate the duration according to the character proportion, determine the start and end time of each character, and finally obtain the character-level timestamps. This is done to improve the accuracy of subsequent video cutting, subtitle and picture synchronization. For example, when a user watches an educational video and clicks on a character (such as the word "half") in the subtitle, they can accurately jump to the picture frame where the corresponding character is pronounced, avoiding the deviation caused by insufficient timestamp accuracy. At the same time, more detailed time basis is provided for precise video cutting in step four.
[0032] At the same time of generating sub-segments and refining timestamps, a semantic abstract is also generated for each sub-segment. The purpose of generating a semantic abstract is to provide intuitive segment theme information for subsequent structured output, making it easy for users to quickly understand the content of sub-segments. For example, the abstract of a sub-segment of a corporate meeting video is "project progress discussion", so users can know the core content without watching the entire segment, providing convenience for subsequent downstream tasks such as retrieval and knowledge management. Finally, semantic sub-segments with character-level timestamps and corresponding semantic abstracts are obtained.
[0033] The implementation process includes: { "input":{ "input_timestamped_subtitles":{ "type":"DataFrame or List[Dict]", "description":"Timestamped subtitle text output by the speech recognition module. Each entry contains:", "field_description":[ { "field_name":"timestamp", "type":"float", "description":"End timestamp of the subtitle segment (in seconds)" }, { "field_name":"simplified_chinese_subtitles", "type":"string", "description":"Simplified Chinese text content after conversion" } ] } } } Step four: Based on the obtained semantic sub-segments with character-level timestamps, perform accurate video cutting. Specifically, use a pre-set professional audio and video processing tool to crop the original video according to the character-level timestamp intervals of the semantic sub-segments, generate independent video segments, and name the segments in the format of the original video file name plus timestamp. Selecting a professional audio and video processing tool is because it can achieve frame-level cutting according to the precise character-level timestamp, avoiding problems such as video frame freezing and audio missing after cutting. Naming by original file name plus timestamp is to facilitate users to trace the original video source corresponding to the sub-segment, and quickly identify the time interval of the sub-segment, such as "math course_000310_000520.mp4". Users can directly know that the segment comes from the "math course" original video and corresponds to the time from 3 minutes and 10 seconds to 5 minutes and 20 seconds.
[0034] Then, build a structured metadata object containing segment identifier, start and end timestamps, duration, corresponding subtitle text, semantic summary, and original video file name. The segment identifier is used to uniquely distinguish each sub-segment to avoid confusion between multiple segments. The start and end timestamps and duration provide the basis for subsequent retrieval and playback progress control. The corresponding subtitle text and semantic summary supplement the textual information of the segment, supporting keyword retrieval and other functions. The original video file name associates the sub-segment with the original resource, making it easy for users to view the complete context of the segment.
[0035] The structured metadata object includes a clip id, a start timestamp, an end timestamp, a duration, corresponding subtitle text, a semantic summary, and a file name.
[0036] After that, the corresponding subtitle text is subjected to error correction processing, specifically, a large language model is called to perform context semantic verification on the subtitle text in combination with a preset field professional vocabulary to correct the errors in the text. This is because errors may occur in the speech recognition process due to factors such as accent and background noise (e.g., "Gougu theorem" is recognized as "Gougu theorem"). The large language model can judge the error through context semantics, and the professional vocabulary (such as mathematical terms in the education field and "litigation period" in the financial field) can ensure the professionalism of the error correction and avoid deviation of professional terms after correction - for example, "prescription drugs" in a medical video will not be mistakenly corrected to "over-the-counter drugs", ensuring the accuracy of the subtitle text and providing a reliable text basis for subsequent knowledge graph construction, question and answer systems and other downstream applications.
[0037] Finally, the independent video clip and the structured metadata object are integrated to output standard structured data. The integration of the two is to make the structured data contain not only video clips that can be directly watched, but also metadata information that is easy to retrieve and analyze. The standard structured data can ensure that different systems can be compatible and read, ultimately achieving the goal of transforming raw video into "watchable, searchable, and usable" structured content, improving video resource utilization and intelligent management level.
[0038] The video content structured processing system that integrates a speech large model and a visual semantic large model, as shown in Figure 2 , includes: An audio extraction module: taking an original video file as input, using a professional audio and video processing tool to separate an audio stream from the original video file; performing standardized processing on the audio stream to obtain a standardized audio file; if the original video file has no audio track, triggering a preset exception handling mechanism and outputting an error prompt; A speech recognition module: inputting the standardized audio file into a speech large model to perform efficient batch transcription and generate associated data containing word-level timestamps and corresponding text; performing complex and simple conversion on the transcribed text to output unified simple Chinese subtitle text; at the same time, optimizing hardware resource usage through a video memory management mechanism, and finally obtaining simple Chinese subtitle text data with word-level timestamps; The video semantic segmentation module inputs the simplified Chinese subtitle text data with word-level timestamps into a visual semantic large model, and performs semantic segmentation according to content semantic logic by the visual semantic large model to obtain a plurality of sub-clips with complete themes; the word-level timestamps are further refined into character-level timestamps, and a semantic abstract is generated for each sub-clip, and finally a semantic sub-clip with character-level timestamps and a corresponding semantic abstract are obtained; The structured output module performs accurate video cutting to generate independent video clips with embedded timestamp information according to the semantic sub-clip with character-level timestamps; a structured metadata object containing a clip identifier, start and end timestamps, a duration, corresponding subtitle text, a semantic abstract, and an original video file name is constructed; after error correction processing is performed on the subtitle text, the independent video clips and the structured metadata object are integrated, and finally standard structured data is output.
[0039] It should be noted that the relational terms herein such as first and second and the like are used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0040] The preferred embodiments of the application disclosed above are only used to help explain the application. The preferred embodiments do not describe all the details of the application, nor limit the application to the specific embodiments described. It is obvious that many modifications and variations can be made according to the content of the specification. The specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the application, so that those skilled in the art can well understand and utilize the application. The application is limited only by the claims and their full scope and equivalents.
Claims
1. A method for video content structured processing by fusing a speech large model and a visual semantic large model, characterized in that, The method comprises the following steps: Step 1: taking an original video file as input, using a professional audio and video processing tool to separate an audio stream from the original video file; Performing standardized processing on the audio stream to obtain a standardized audio file; if the original video file has no audio track, triggering a preset abnormal processing mechanism and outputting an error prompt; Step 2: inputting the standardized audio file into a speech large model to perform efficient batch transcription and generate associated data containing word-level timestamps and corresponding text; performing complex and simple conversion on the text obtained by transcription to uniformly output simplified Chinese subtitle text; at the same time, optimizing hardware resource use through a video memory management mechanism, and finally obtaining simplified Chinese subtitle text data with word-level timestamps; Step 3: inputting the simplified Chinese subtitle text data with word-level timestamps into a visual semantic large model, and performing semantic segmentation by the visual semantic large model according to content semantic logic to obtain multiple theme-complete sub-clips; Further refining the word-level timestamps into character-level timestamps, and generating a semantic abstract for each sub-clip to finally obtain semantic sub-clips with character-level timestamps and corresponding semantic abstracts; Step 4: performing accurate video cutting according to the semantic sub-clips with character-level timestamps to generate independent video clips embedded with timestamp information; Building a structured metadata object; After error correction processing on the subtitle text, integrating the independent video clips and the structured metadata object to finally output standard structured data.
2. The method of claim 1, wherein the method further comprises: The specific steps of performing standardized processing on the audio stream are as follows: Merging the multi-channel audio stream into a single-channel audio stream, uniformly resampling the single-channel audio to a preset standard sampling rate, and finally converting it to a preset standard audio format to obtain the standardized audio file.
3. The method of claim 1, wherein the method further comprises: The specific steps of triggering the preset abnormal processing mechanism are as follows: Detecting whether the original video file has an audio track, if not, outputting an error prompt and terminating the current audio extraction process.
4. The method of claim 1, wherein the method further comprises: The specific steps of performing efficient batch transcription are as follows: Loading the speech large model, configuring the corresponding language decoder according to the language type of the video to be processed, and then dividing the standardized audio file into batch audio clips according to the preset time length, and inputting them into the speech large model in turn.
5. The method of claim 1, wherein the method further comprises: The specific steps of the video memory management mechanism are as follows: Monitoring the video memory usage of the speech large model during operation, when the video memory usage reaches a preset threshold, releasing the audio clip cache that has completed the transcription task, and controlling the transcription task concurrency through task queue scheduling.
6. The method of claim 1, wherein the method further comprises: The specific steps of the visual semantic large model adaptively performing semantic segmentation are as follows: Based on a preset structured prompt template, identifying semantic discontinuities in the simplified Chinese subtitle text with word-level timestamps, taking the semantic discontinuities as the segmentation boundaries to generate theme-complete sub-clips.
7. The method of claim 1, wherein the method further comprises: The specific steps of further refining the word-level timestamps into character-level timestamps are as follows: Calculating the total duration corresponding to the word-level timestamps, counting the number of characters contained in the word, allocating the duration according to the character proportion, determining the start and end time of each character, and obtaining the character-level timestamps.
8. The method of claim 1, wherein the method further comprises: The specific steps of performing accurate video cutting are as follows: A preset professional audio and video processing tool is used to clip the original video according to the character-level timestamp interval of the semantic sub-fragment, generate an independent video segment, and name the segment in the format of the original video file name plus timestamp.
9. The method of claim 1, wherein the method further comprises: The specific steps of performing error correction processing on the subtitle text are as follows: A large language model is called to combine a preset field professional vocabulary library to perform context semantic checking on the subtitle text and correct the error content in the text.
10. A video content structured processing system fusing a speech large model and a visual semantic large model, used to implement the video content structured processing method fusing a speech large model and a visual semantic large model according to any one of claims 1-9, characterized in that, It includes: An audio extraction module: taking the original video file as input, using a professional audio and video processing tool to separate the audio stream from the original video file; Performing standardized processing on the audio stream to obtain a standardized audio file; if the original video file has no audio track, triggering a preset exception processing mechanism and outputting an error prompt; A speech recognition module: inputting the standardized audio file into a speech large model to perform efficient batch transcription and generate associated data containing word-level timestamps and corresponding text; performing complex and simple conversion on the text obtained by transcription to output unified simple Chinese subtitle text; at the same time, optimizing hardware resource use through a video memory management mechanism, and finally obtaining simple Chinese subtitle text data with word-level timestamps; A video semantic segmentation module: inputting the simple Chinese subtitle text data with word-level timestamps into a visual semantic large model, and performing semantic segmentation according to content semantic logic by the visual semantic large model to obtain multiple theme-complete sub-fragments; Further refining the word-level timestamp to a character-level timestamp, and generating a semantic abstract for each sub-fragment, finally obtaining semantic sub-fragments with character-level timestamps and corresponding semantic abstracts; A structured output module: performing precise video cutting according to the semantic sub-fragments with character-level timestamps to generate independent video segments embedded with timestamp information; Building a structured metadata object; After performing error correction processing on the subtitle text, integrating the independent video segments and the structured metadata object, and finally outputting standard structured data.
Citation Information
Patent Citations
News story segmentation method based on multi-feature fusion and random forest model
CN112633241A
Intelligent video editing and abstract generating method and device based on semantic segmentation
CN118381980A
Financial news editing method based on large language model
CN119906867A
Video multi-language conversion method and system based on semantic segmentation
CN120529106A
Video training data generation method based on multi-modal semantic alignment
CN120766057A