Key frame real-time extraction system and method for multi-modal large model

By screening key frames with high information density on the user side and combining them with a large multimodal model architecture on the cloud, the problems of heavy load and redundant information in multimodal large model video stream transmission are solved, achieving efficient video data processing and improved model output quality.

CN120711204AActive Publication Date: 2025-09-26BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510738814.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-26
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

Video streaming transmission of large multimodal models suffers from heavy load and excessive redundant information, which increases network load and interferes with the model's attention mechanism, reducing output accuracy.

Method used

By dynamically screening key frames with high information density on the user side, combining with the multimodal large model architecture on the cloud, and adopting the key frame selection algorithm and secretary algorithm based on adjacent frame differences, video data transmission is optimized and redundant information is reduced.

Benefits of technology

It significantly reduces the redundancy of video data transmission, improves model output accuracy and response speed, optimizes resource utilization, and enhances end-cloud collaboration efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120711204A_ABST
    Figure CN120711204A_ABST
Patent Text Reader

Abstract

The invention provides a key frame extraction method for a multi-modal large model. The multi-modal large model video data processing efficiency and generation quality are improved through cooperative computing of a user side and a cloud side. According to the technical scheme, the method comprises the following steps: 1) after a client application submits data, an interface service performs primary processing on the data, and video frames and audio frames are divided and merged into data pieces according to a time sequence; 2) dynamically extracting n key video frames in each data piece of the user side, and compressing, packaging and transmitting the n key video frames and corresponding audio frames to the cloud side; 3) the cloud performs coding fusion on the multi-modal data and then inputs a large model generation result; and 4) the cloud evaluates the generation quality, and the user side adaptively adjusts data preprocessing and key frame extraction parameter configuration according to the generation quality and the response time delay. By reducing redundant data transmission and dynamically optimizing calculation and transmission resources, the response speed and the output quality of the multi-mode large-model cloud service are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a terminal-cloud collaborative optimization technology for the Internet application layer, specifically a real-time key frame extraction system and method for multimodal large models. Background Art

[0002] With the deepening development of artificial intelligence technology, multimodal big models, by integrating multidimensional data such as video, audio, and text, have significantly enhanced perception and decision-making capabilities in complex scenarios. This has significantly improved models' ability to understand and process complex real-world problems, expanding the boundaries of business. They have demonstrated a significant driving force in a wide range of sectors, including education, healthcare, finance, and industry. With growing market demand and policy support, the scale of China's multimodal big model market continues to expand. According to statistics, the domestic multimodal big model industry market size was approximately 3.3 billion yuan in the first half of 2024. Numerous companies, universities, and research institutions have increased their investment and released numerous general-purpose big models. As the technology matures and costs decrease, the application of multimodal big models in more fields is becoming possible, further driving market growth. Research on multimodal big models is of great significance.

[0003] The service delivery of large multimodal models relies heavily on the computing power and resource scheduling efficiency of cloud platforms. Their computational intensity and real-time requirements pose a dual challenge to infrastructure. Core tasks such as video stream parsing and cross-modal feature fusion involve large-scale parallel computing, and a single inference consumes a significant amount of GPU resources, placing higher demands on the dynamic scheduling capabilities of the computing cluster. Regarding real-time performance, scenarios such as industrial quality inspection and autonomous driving require end-to-end responses to maintain millisecond-level accuracy, while applications such as medical diagnosis require strict time synchronization of multimodal data. The conflict between these high real-time demands and resource-intensive computing presents a key bottleneck. Regarding network transmission, the disparity between bursty bandwidth of video streams and the multi-level traffic generated by voice and sensor data can easily lead to network congestion. Traditional static resource allocation mechanisms can easily lead to contention for computing units during peak concurrency, significantly increasing service response fluctuations. Therefore, designing a resource-efficient cloud service architecture has become a core technical direction for ensuring service stability and responsiveness.

[0004] Current multimodal large-scale cloud-based services often rely on direct input of raw video streams, leading to issues such as high end-to-end data transmission redundancy and excessively long context windows. Statistics show that transmitting an unprocessed 1080P video stream can consume up to 4-8Mbps of bandwidth, with over 60% of consecutive frames containing semantically repetitive content. This not only increases network load but can also interfere with the model's attention mechanism due to redundant information, reducing output accuracy. Summary of the Invention

[0005] In response to the problems of heavy load and excessive redundant information in video stream transmission of existing multimodal large models, the present invention proposes a key frame extraction system and method for multimodal large models. By dynamically screening key frames with high information density, the video data flow is compressed, thereby reducing transmission overhead while optimizing model output quality.

[0006] The key frame extraction method for multimodal large models is implemented by following the steps below:

[0007] Step 1: Build a multimodal large-scale model cloud service architecture. Deploy service interfaces for preprocessing and extracting keyframes on end devices. Deploy multimodal large-scale models on cloud devices to perform data encoding, modality fusion, and inference tasks.

[0008] Step 2: The service interface of the terminal device receives the multimodal data stream and preprocesses it to obtain text data and data slices containing audio and video data;

[0009] Multimodal data streams include text data, video data, and audio data. The preprocessing process is as follows:

[0010] Text data: Text data directly received by the service interface is uploaded directly to the cloud device. Text data contained in video or audio data is extracted through voice recognition and OCR and uploaded to the cloud device together.

[0011] Video and audio data: Preset the time length of the data slice, split the video and audio data into video frames and audio frames of equal length according to the preset time length, and package them into data slices;

[0012] The amount of all data is initially reduced through parameter adjustment, sampling, compression, etc.

[0013] Step 3: Extract key frames from each video data slice using a video key frame extraction algorithm, and package the data slice containing only the key frames and the audio data slice and transmit them to the cloud device;

[0014] The key frame selection algorithm based on adjacent frame differences is as follows:

[0015] Step 301: For a certain data slice, initialize the number of key frames and the data slice boundary index list, recursively divide the data slice and allocate the number of key frames until only one key frame is required for the first sub-data slice;

[0016] The specific method of recursively partitioning data slices is:

[0017] For a certain data slice t, the total number of video frames n contained in the data slice is confirmed at the beginning t , and all video frames in the current data slice And obtain the number of key frames k that need to be selected in the data slice through the configuration algorithm t .

[0018] Then, the data slice is divided into two parts, and half of the key frames are selected in each part, that is, each part selects key frames; continue to divide the first half of the data slice equally, and evenly distribute the number of key frames to be selected in each newly divided data slice until only one key frame needs to be selected from the divided sub-data slices.

[0019] In this process, use the list key and list bd To record the number of key frames that need to be selected in each sub-data slice and the boundary index of the sub-data slice.

[0020] Step 302 , selecting a key frame from a sub-data slice where only one key frame needs to be selected based on the frame difference between adjacent frames using a secretary algorithm, and determining an initial threshold;

[0021] The process of selecting key frames through the secretary algorithm is:

[0022] First observe before The maximum frame difference between the two adjacent frames is recorded as the initial threshold, and the first frame whose frame difference with the previous frame exceeds the threshold is selected as the key frame in the subsequent frames. If no frame in the subsequent frames meets this condition, the last frame is selected as the key frame.

[0023] Step 303, according to the list key The remaining sub-data slices are processed in ascending order according to the number of key frames to be selected recorded in , and after the key frame selection is completed for each sub-data slice, the threshold is updated until all sub-data slices are processed;

[0024] For each sub-data slice, the frame differences between adjacent frames are calculated in turn, and the mechanism for triggering key frame selection is as follows: 1. If the frame difference value of the current frame exceeds the threshold and does not reach the number of key frames required for the interval, the current frame is selected as the key frame. When the required number of key frames is reached, the remaining frames are discarded; 2. If the number of remaining frames is equal to the number of key frames that have not yet been selected, if no frame is selected at this time, the number of key frames will be insufficient, then all remaining frames will be selected as key frames.

[0025] The threshold update mechanism is as follows: when the current sub-data slice completes the key frame selection, assuming that the next sub-data slice needs to select k′ key frames, the k′th largest value is selected from the frame differences of the processed frames as the new threshold.

[0026] Step 304, repeating steps 301 to 303 until all data slices have completed the selection of key frames;

[0027] In step 305 , after extracting key frames from each data slice, the key frames are used to replace the video frame portion of the original data slice, thereby obtaining a data slice with semantically repeated video frames eliminated.

[0028] In step 4, the end device sequentially uploads each data slice containing only key frames to the cloud device, and simultaneously inputs it into the multimodal large model for fusion with the audio data and text data. The generated content is then transmitted back to the end device. At the same time, the multimodal large model evaluates the generated content and transmits the evaluation results to the optimizer of the end device.

[0029] Step 5: The user-side optimizer comprehensively considers the response delay and generation quality, and adjusts the preprocessing configuration and the number of key frames specified in the key frame extraction algorithm;

[0030] Step 6: Return to step 2 according to the adjusted pre-processing configuration and the number of key frames, and repeat steps 2 to 5 until the data processing is completed.

[0031] A real-time keyframe extraction system for a multimodal large model includes a user end and a cloud device. The user end is equipped with a data acquisition module, a preprocessing module, a keyframe extraction module, and an optimizer module. First, the data acquisition module collects the multimodal data stream and sends it to the preprocessing module for preprocessing, dividing the multimodal data stream into text data, audio data, and video data. Then, the keyframe extraction module extracts the keyframes of the video data and replaces the original video data with the keyframes. Finally, the keyframe video data is packaged with the text data and audio data and submitted to the cloud device. The cloud device is equipped with a modal fusion module, a multimodal large model, and an evaluation module. The modal fusion module encodes and fuses the keyframe video data with the text data and audio data, transmits the fused data to the multimodal large model, and outputs the generated content. The generated content is simultaneously transmitted to the user end and the evaluation module. The evaluation module evaluates the quality of the generated content and feeds the evaluation results back to the optimizer module on the user end. The optimizer module optimizes the parameters of the preprocessing module and the keyframe extraction module.

[0032] The advantages of the present invention are:

[0033] (1) The present invention constructs an efficient cloud service processing framework for multimodal large models through a collaborative architecture between the user end and the cloud. It combines cloud-side quality feedback with dynamic adjustment of end-side parameters to balance transmission efficiency and generation quality. This framework effectively solves the problems of high data transmission redundancy and low resource utilization in multimodal large model services, and provides an efficient and reliable end-cloud collaborative solution for multimodal large model cloud service scenarios.

[0034] (2) After receiving the original multimodal data at the user end, the present invention uses compression coding technology to reduce transmission redundancy, accurately captures semantic key frames through the principle of maximizing frame difference, and screens key video frames online in real time based on the improved secretary algorithm, significantly reducing the amount of transmitted data and model context interference.

[0035] (3) After completing cross-modal fusion and large model inference on the cloud, the present invention drives the user end to adaptively optimize preprocessing parameters and key frame extraction strategies through quality assessment and real-time feedback, thereby optimizing the overall system efficiency and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is the multimodal large model cloud service framework proposed by the present invention;

[0037] Figure 2 It is a flow chart of the key frame selection algorithm proposed by the present invention. Specific implementation plan

[0038] The present invention proposes a real-time key frame extraction system and method for multimodal large models. The system selects key frames based on the difference between adjacent frames of the video modality to eliminate repeated semantic information in the original data. The system improves the processing efficiency and generation quality of multimodal large model video data through collaborative computing between the user end and the cloud. It effectively saves network resources while reducing the context length of the input large model, thereby improving the model accuracy. Its technical solution includes: 1) After the user end application submits the data, the interface service performs preliminary processing on the data, and divides the video frames and audio frames into data slices according to time sequence; 2) Dynamically extracts n key video frames from each data slice on the user end, and compresses and packages them with the corresponding audio frames and transmits them to the cloud; 3) The cloud encodes and fuses the multimodal data and inputs them into the large model to generate the result; 4) The cloud evaluates the generation quality, and the user end adaptively adjusts the data preprocessing and key frame extraction parameter configuration according to the generation quality and response delay. The present invention significantly improves the response speed and output quality of multimodal large model cloud services by reducing redundant data transmission and dynamically optimizing computing and transmission resources.

[0039] The overall structure of the system of the present invention is as follows Figure 1 As shown in the figure, the user-side and cloud-based devices jointly process different parts of the multimodal large model. The user-side is responsible for data acquisition and preprocessing, where a keyframe extraction algorithm is used to segment and extract keyframes from the video and submit the processed data to the cloud-based device. The cloud-based device is responsible for multimodal information encoding and fusion, as well as inference and generation quality assessment of the multimodal large model, and responds to user-side requests. The user-side ultimately adjusts its data preprocessing strategy based on the generation results and evaluation metrics returned by the cloud, such as adjusting the resolution or number of keyframes, to optimize overall system efficiency and user experience.

[0040] The key frame extraction method for multimodal large models is implemented in the following steps:

[0041] Step (1) building a multimodal large model cloud service architecture, wherein the end device deploys a service interface for preprocessing and extracting key frames, and the cloud device deploys a multimodal large model for performing data encoding, modality fusion and reasoning tasks;

[0042] Step (2) The terminal device receives the multimodal data stream submitted by the application and performs preprocessing, extracts text data through speech recognition and OCR, and uses compression and sampling to reduce the data volume, and packages the video frames and audio frames of equal length into data slices;

[0043] In order to reduce the data volume of metadata and process it into a data format suitable for subsequent frameworks, the service interface of the user-end device receives the video stream from the upper-layer application and pre-processes it. Specifically, the service interface uploads the text data directly to the cloud, while the other received video and audio data are divided into video frames and audio frames of equal length. Text is extracted from the video or video through voice recognition and OCR, and the data volume of all data is initially reduced through parameter adjustment, sampling, compression, etc.

[0044] Due to the multimodal large model's requirement for synchronization of video and audio data, the user-end device will package the video frames and audio frames within the same time period into data slices and upload them. Therefore, at this stage, it is necessary to perform preliminary segmentation and merging of the data frames and audio frames within the corresponding time according to the preset time length of the data slice.

[0045] Step (3) calls the key frame extraction algorithm based on the secretary algorithm for each data slice sequence to extract the key frames in each data slice, specifically including:

[0046] (3a) Initialize the number of key frames and the data slice boundary index list, recursively divide the data slice and assign the number of key frames until only one key frame is required for the first sub-data slice;

[0047] (3b) Frame difference method is used to calculate the differences in pixel, edge and regional features of adjacent frames, and dynamic threshold is combined to select key frames: in the first stage, the initial threshold is determined by the secretary algorithm, and in the subsequent stages, the threshold is updated according to the difference value of the selected key frames, and frames with difference values ​​exceeding the threshold are preferentially selected;

[0048] Traditional methods for selecting video keyframes typically select keyframes at uniform intervals within a data slice based on frame rate configuration, lacking a detailed understanding of video content. This approach can easily overlook important events or key changes, affecting the accuracy of video analysis. Furthermore, multimodal large models focus their understanding of visual information on the semantic content of the video rather than its smoothness, which is inconsistent with the actual requirements of multimodal large models for video streams. To address this issue, the present invention proposes a new keyframe selection algorithm based on adjacent frame differences. This algorithm extracts quantitative features from video frames to calculate the differences between adjacent frames to detect significant changes in video content. There are three main methods for quantitatively calculating video frame features with low computational complexity: pixel features compare the differences between pixels in adjacent video frames, edge features capture differences in object outlines, and area features measure the degree of change in the area segmented by edges. Within each video interval, the algorithm selects several frames with the greatest changes as keyframes. By prioritizing these key changes, the algorithm can more comprehensively understand the video content, thereby improving the accuracy of the analysis results.

[0049] Since key frame selection is an online problem, frames need to be processed sequentially and future frames cannot be predicted, so it is impossible to calculate the differences of all frames for evaluation at the same time. Therefore, the basic design idea of ​​the algorithm is to calculate the difference of the current frame, combine the previously processed frame difference information, and immediately decide whether to select the current frame as the key frame. To this end, the algorithm introduces the secretary problem and algorithm and optimizes it. The original secretary algorithm aims to select the largest video frame from n video frames online. The specific approach of the secretary algorithm is to first observe the previous frame. The maximum frame difference is recorded as a threshold, and the first frame exceeding the threshold is selected as the keyframe among the subsequent frames. If no frame meets this condition, the last frame is selected. This method can achieve near-optimal online keyframe selection under limited information conditions.

[0050] In the scenario of the present invention, a specific number k of key frames need to be selected, while the traditional secretary algorithm can only be used to select a single key frame. To solve this problem, the present invention proposes a key frame selection algorithm. At the beginning of each data slice t, the configuration algorithm obtains the number k of key frames that need to be selected in the data slice. t Then, in the initialization phase, the algorithm divides the data slice into two parts and selects half of the key frames in each part, that is, each part selects key frames, and then continue to divide the first half of the data slice equally, and evenly distribute the number of key frames to be selected in each newly divided data slice, until only one key frame is required for the divided data slice. In this process, the algorithm uses the list list key and list bd To record the number of key frames that need to be selected in each sub-data slice and the boundary index of the sub-data slice.

[0051] After initialization, the algorithm begins processing video frames. In the first phase, the algorithm processes the first sub-slice. Within this slice, each video frame is scored using the frame difference between it and the previous frame. A keyframe is selected within this interval using a traditional online secretary algorithm. After the first phase, the maximum frame difference in the first slice is recorded and used as a threshold for processing subsequent slices.

[0052] In the second phase, the algorithm processes the remaining data slices in sequence. Within each sub-data interval, there are two situations that trigger keyframe selection: 1. If the frame difference of the current frame exceeds the threshold and the number of keyframes required for the interval is less than the number of keyframes, the current frame will be selected as the keyframe; 2. If the number of remaining frames is equal to the number of keyframes that have not yet been selected, if no frame is selected, the number of keyframes will be insufficient, so all remaining frames must be selected as keyframes.

[0053] After each sub-slice is processed, the threshold is updated. Because the length of the next slice and the number of keyframes required are the same as the sum of the lengths and keyframes of all previously processed slices, the new slice can be selected using the frame differences of the previously processed frames. Assuming that a new sub-slice requires k′ keyframes, the k′th largest value among the frame differences of the previously processed frames is selected as the new threshold. The algorithm continues until all frames have been processed.

[0054] After extracting key frames from each data slice, the key frames are used to replace the video frame portion of the original data slice, thereby obtaining a data slice with semantically repeated video frames eliminated.

[0055] Specifically, if Figure 2 As shown, the key frame selection algorithm steps proposed by the present invention are as follows:

[0056] 1. Confirm the parameters and data required for this round of algorithm execution, including: the number of key frames k to be extracted in each data slice t , the total number of video frames in each data slice n t , and all video frames in the current data slice

[0057] 2. Create and maintain two lists: the number of key frames required to be selected for each sub-data slice key and the boundary index list of the sub-data slice bd .

[0058] 3. Add to list key , n t Add to list bd .

[0059] 4. When k t >1, Repeat step 3 until k t =1.

[0060] 5. Use the secretary algorithm based on the adjacent frame difference arrive Select a key frame and calculate the maximum frame difference dif max .

[0061] 6. Add keyframes to l t , and dif max Used as the subsequent frame difference threshold.

[0062] 7. Set temporary variable r=1 as list bd The loop index is used to indicate the sub-data slice currently being processed; the temporary variable c=0 records the number of key frames extracted from each data slice. Algorithm online processing and The contents of the sub-data slice.

[0063] 8. In each loop, calculate the frame and Frame difference dif i , if the frame difference is greater than the selected threshold (dif i >threshold) and the data slice also needs to select key frames (list key [r]>c), or the remaining number of frames is equal to the number of key frames to be selected (list key [r]-c==list bd [r]-i+1), then Select as keyframe and add to l t , and let c=c+1.

[0064] 9.If i <list bd [r] indicates that the current sub-data slice has been processed. Let r = r + 1, c = 0, and update the threshold threshold to the list of the processed video frame. key [r] large frame difference. If there are still unprocessed frames,

[0065] Repeat step 8.

[0066] Step (4) The end device uploads the data slice containing the key frame to the cloud. The cloud device integrates the multimodal data and inputs it into the large model for reasoning, generates output results, and uses technical indicators such as perplexity to make a preliminary judgment on the quality of model generation. The generation quality judgment evaluation indicators and generated content are transmitted back to the user end;

[0067] Step (5) The cloud feeds back the output results and quality indicators to the end device, and the end device dynamically adjusts the preprocessing parameters and the number of key frames according to the latency and quality requirements, and optimizes the subsequent data slice processing strategy;

[0068] The user-side optimizer comprehensively considers response latency and generation quality, adjusting the preprocessing configuration and the number of keyframes specified in the keyframe extraction algorithm to optimize the user experience. Specifically, the preprocessing configuration covers audio sampling rate, compression encoding method, video resolution, etc. The higher the configuration, the lower the degree of data distortion, and the more clearly the multimodal large model in the cloud can capture the semantic information in the data stream. However, the amount of data transmitted will also lead to increased transmission latency and computational latency. The number of keyframes affects the amount of data required to be transmitted, the context length, and the completeness of the video semantics. If there are too many keyframes, the data volume and context length will be too large, and the response latency and model generation quality will be negatively affected. If there are too few keyframes, the video semantics will be significantly missing, which will also affect the generation quality of the model.

[0069] Step (6) Repeat steps (2) to (5) until all data processing is completed.

[0070] In summary, the present invention proposes a key frame extraction method for multimodal large model cloud services, which realizes online screening of high-information-density key frames through dynamic frame difference threshold and recursive secretary algorithm, significantly reduces video data transmission redundancy and model context length, and improves the end-cloud collaboration efficiency while ensuring the quality of multimodal large model generation.

Claims

1. A key frame extraction method for a multimodal large model, characterized in that: The following steps are involved: Step 1: Build a multimodal large-scale model cloud service architecture. Deploy service interfaces for preprocessing and extracting keyframes on end devices. Deploy multimodal large-scale models on cloud devices to perform data encoding, modality fusion, and inference tasks. Step 2: The service interface of the terminal device receives the multimodal data stream and preprocesses it to obtain text data and data slices containing audio and video data; Step 3: Extract key frames from each video data slice using a video key frame extraction algorithm, and package the data slice containing only the key frames and the audio data slice and transmit them to the cloud device; The key frame selection algorithm based on adjacent frame differences is as follows: Step 301: For a certain data slice, initialize the number of key frames and the data slice boundary index list, recursively divide the data slice and assign the number of key frames until the first sub-data slice only needs to select a key frame, and use the list list key and list bd To record the number of key frames that need to be selected in each sub-data slice and the boundary index of the sub-data slice; Step 302 , selecting a key frame from a sub-data slice where only one key frame needs to be selected based on the frame difference between adjacent frames using a secretary algorithm, and determining an initial threshold; Step 303, according to the list key The remaining sub-data slices are processed in ascending order according to the number of key frames to be selected recorded in , and after the key frame selection is completed for each sub-data slice, the threshold is updated until all sub-data slices are processed; For each sub-data slice, the frame differences between adjacent frames are calculated in sequence. The mechanism for triggering key frame selection is as follows:

1. If the frame difference of the current frame exceeds the threshold and does not reach the required number of key frames in the interval, the current frame is selected as the key frame. When the required number of key frames is reached, the remaining frames are discarded; 2. If the number of remaining frames is equal to the number of key frames that have not yet been selected, and if no frames are selected, the number of key frames will be insufficient, then all remaining frames will be selected as key frames; The threshold update mechanism is as follows: when the key frame selection is completed for the current sub-data slice, assuming that k′ key frames need to be selected for the next sub-data slice, the k′th largest value is selected from the frame differences of the processed frames as the new threshold; Step 304, repeating steps 301 to 303 until all data slices have completed the selection of key frames; Step 305: After extracting key frames from each data slice, the key frames are used to replace the video frame portion of the original data slice, thereby obtaining a data slice with semantically repeated video frames eliminated. In step 4, the end device sequentially uploads each data slice containing only key frames to the cloud device, and simultaneously inputs it into the multimodal large model for fusion with the audio data and text data. The generated content is then transmitted back to the end device. At the same time, the multimodal large model evaluates the generated content and transmits the evaluation results to the optimizer of the end device. Step 5: The user-side optimizer comprehensively considers the response delay and generation quality, and adjusts the preprocessing configuration and the number of key frames specified in the key frame extraction algorithm; Step 6: Return to step 2 according to the adjusted pre-processing configuration and the number of key frames, and repeat steps 2 to 5 until the data processing is completed.

2. The key frame extraction method for multimodal large models according to claim 1 is characterized in that: The multimodal data stream includes text data, video data and audio data, and the process of preprocessing them respectively is as follows: Text data: Text data directly received by the service interface is uploaded directly to the cloud device. Text data contained in video or audio data is extracted through voice recognition and OCR and uploaded to the cloud device together. Video and audio data: Preset the time length of the data slice, split the video and audio data into video frames and audio frames of equal length according to the preset time length, and package them into data slices.

3. The key frame extraction method for multimodal large models according to claim 1, characterized in that: The specific method of recursively dividing the data slices is: For a certain data slice t, the total number of video frames n contained in the data slice is confirmed at the beginning t , and all video frames in the current data slice And obtain the number of key frames k that need to be selected in the data slice through the configuration algorithm t ; Then, the data slice is divided into two parts, and half of the key frames are selected in each part, that is, each part selects key frames; continue to divide the first half of the data slice equally, and evenly distribute the number of key frames to be selected in each newly divided data slice until only one key frame needs to be selected from the divided sub-data slices.

4. The key frame extraction method for multimodal large models according to claim 1, characterized in that: The process of selecting a key frame from a sub-data slice where only one key frame needs to be selected by the secretary algorithm is as follows: First observe before Frame, record the maximum adjacent frame difference as the initial threshold, and select the first frame whose frame difference with the previous frame exceeds the threshold as the key frame in the subsequent frames; if no frame in the subsequent frames meets this condition, select the last frame as the key frame.

5. A key frame extraction system for multimodal large models, characterized by: It includes a user end and a cloud device. The user end is equipped with a data acquisition module, a preprocessing module, a key frame extraction module, and an optimizer module. First, the data acquisition module collects multimodal data streams and sends them to the preprocessing module for preprocessing, which divides the multimodal data streams into text data, audio data, and video data. Then the key frame extraction module extracts the key frames of the video data and replaces the original video data with the key frames; finally, the key frame video data is packaged with the text data and audio data and submitted to the cloud device; the cloud device is equipped with a modal fusion module, a multimodal large model and an evaluation module. The modal fusion module encodes and fuses the key frame video data with the text data and audio data, transmits the fused data to the multimodal large model, and outputs the generated content; the generated content is transmitted to the user end and the evaluation module at the same time. The evaluation module evaluates the quality of the generated content and feeds back the evaluation results to the optimizer module on the user end. The optimizer module optimizes the parameters of the preprocessing module and the key frame extraction module.

Citation Information

Patent Citations

  • Fast high efficiency video coding method based on optimum stopping theory

    CN104301723A

  • Video keyframe extraction method based on linear dynamic system

    CN107027051A

  • Key frame extraction method based on inter-frame difference and color histogram difference

    CN112270247A

  • Lightweight method for extracting video key frame

    CN113691863A

  • Method for continuously monitoring change and damage of rock stratum in similar material model

    CN115205735A

Cited By

  • Video adaptive transmission and reconstruction method and device based on end-cloud collaboration

    CN121585319A

  • A video adaptive transmission and reconstruction method and device based on end-cloud cooperation

    CN121585319B