Video understanding method and device based on frame sequence abstraction and language model guidance, equipment and medium

By collaborating frame sequence abstraction with a large language model, and adopting the Bhattacharyya distance metric and representative frame strategy, we address the problems of high computational resource overhead and insufficient semantic expression in existing video understanding methods, and achieve lightweight and efficient structured parsing and output of video content.

CN120708131APending Publication Date: 2025-09-26XIAMEN CHANJING TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510825350.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing video understanding methods have high computational resource overhead, insufficient semantic expression, and lack of structured output, making it difficult to efficiently understand and structure video content in a lightweight environment.

Method used

Frame sequence abstraction is used in collaboration with a large language model. The Bhattacharyya distance metric is used to identify shot switching points, perform coarse and fine segmentation, extract representative frames, and input them into the GPT model for structured parsing to generate semantic sequences in JSON format.

Benefits of technology

It achieves lightweight computing, high semantic accuracy, dynamic segmentation covering key scenarios, and outputs structured data to support copy generation and product extraction, improving subsequent processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708131A_ABST
    Figure CN120708131A_ABST
Patent Text Reader

Abstract

The invention provides a video understanding method based on frame sequence abstraction and language model guidance, which is applied to the technical field of artificial intelligence and multi-modal information processing and comprises the following steps of: performing preprocessing and shot segmentation on image features in a video frame sequence to generate representative frame coding data; semantic analysis is carried out on the representative frame image features based on a sub-mirror segment semantic extraction rule, and semantic sequence data is generated in combination with the guiding effect of a structured text instruction on a large language model and model output characteristics; processing cross-split representative frame features based on a multi-frame splicing input mode, dynamically optimizing semantic understanding logic in combination with text instruction structure design and model output rules, and generating structured scene understanding data; processing the structured scene understanding data to generate application data including a video semantic abstract, a script creation semantic fragment, a video tag and a commodity selling point; and processing the application data based on a preset application scene rule to generate a video structured semantic annotation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of artificial intelligence and multimodal information processing technology, and in particular to a video understanding method, apparatus, device and medium based on frame sequence abstraction and language model guidance. Background Art

[0002] With the surge in short videos, commercials, and product introductions, automated video content understanding has become a key task in intelligent creation and marketing analysis. Current video understanding methods mostly rely on convolutional neural networks to extract features or use complete video frame sequences as language model input for joint modeling. These methods suffer from the following shortcomings:

[0003] 1. High computing resource overhead: Processing full frames or evenly spaced frames is costly and unsuitable for deployment in lightweight environments or embedded platforms.

[0004] 2. Insufficient semantic expression: The video contains multiple scene jumps, and frame averaging alone cannot capture key semantics.

[0005] 3. Lack of structured content generation capabilities: Existing models lack explanatory output for understanding video content, which is not conducive to subsequent tasks such as copywriting generation and product extraction.

[0006] Therefore, there is an urgent need for a low-cost, well-structured video semantic dimensionality reduction and understanding method that can combine the capabilities of large language models to efficiently complete video structured expression. Summary of the Invention

[0007] The purpose of the present invention is to solve the above-mentioned problems and provide a video understanding method, device, equipment and medium based on frame sequence abstraction and language model guidance.

[0008] The technical solution of the present application is implemented as follows: by collaborating with a large language model through frame sequence abstraction, a structured analysis of video content is achieved, solving the problems of high computing resource overhead, insufficient semantic expression, and lack of structured output in the prior art. The Bhattacharyya distance is used to measure the similarity of the histograms of consecutive frame images, and the lens switching points are identified to complete the coarse segmentation. The fine segmentation strategy is triggered for storyboard segments with a duration of more than 10 seconds. The average time taken for a 30-second video is less than 1 second, ensuring efficiency and granularity accuracy. Representative frames are extracted at 20%, 50%, and 80% of the storyboard segment frame positions. For short storyboards, two frames at 20% and 80% are taken and encoded in the "XY" format (X is the storyboard number, Y is the frame number) to achieve frame-level semantic abstraction.

[0009] The representative frames are spliced ​​together and fed into the GPT model. Guided by structured prompts containing target descriptions, number explanations, intent settings, and format constraints, the model returns JSON data containing group numbers and scene descriptions, forming a semantic sequence. Based on the semantic sequence, a video semantic summary is constructed, script creation fragments are extracted, and video tags and product selling points are generated, ultimately completing the video's structured semantic annotation.

[0010] The advantages or benefits of the above technical solution include at least: lightweight computation, processing only key frames rather than full frames, significantly reducing computing resource consumption and adapting to lightweight environments. Semantic accuracy: dynamic segmentation and representative frame strategies cover key scenes, avoiding semantic omissions caused by frame averaging. Structured output: JSON-formatted semantic results directly support downstream tasks such as copywriting generation and product extraction, improving subsequent processing efficiency.

[0011] This method can efficiently process videos such as product introductions and advertisements. The resulting semantic summaries, script snippets, and selling point tags can be directly applied in scenarios such as intelligent creation and marketing analysis, enabling automated understanding and value mining of video content. By synergizing image and language models, it overcomes the technical bottlenecks of traditional video understanding and provides a new path for multimodal information processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings illustrate exemplary implementations of the present application of embodiments of the present invention and, together with the description, are used to explain the principles of the present application. These drawings are included to provide a further understanding of the present application, and the drawings are included in and constitute a part of this specification.

[0013] Figure 1 A flowchart of a video understanding method based on frame sequence abstraction and language model guidance provided by an embodiment of the present application is shown;

[0014] Figure 2 A schematic structural diagram of a video understanding device based on frame sequence abstraction and language model guidance provided by an embodiment of the present application is shown.

[0015] Figure 3 A schematic diagram of a video understanding device based on frame sequence abstraction and language model guidance provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0016] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are for illustrative purposes only and are not intended to limit the scope of protection of the present application.

[0017] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0018] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0019] Reference Figure 1 The embodiment of the present invention provides a flow chart of a video understanding method based on frame sequence abstraction and language model guidance, including:

[0020] S101, obtaining an original video, extracting continuous frame images from the original video, and generating a video frame sequence.

[0021] In one embodiment, taking an outdoor portable speaker product introduction video as an example, the target video focuses on the outdoor usage scenarios and functional features of the speaker, with a total length of 33 seconds and a frame rate of 30fps. The total number of frames is calculated to be 33×30=990 frames. The video displays the product through 11 segments (Group 0 to Group 10), and each segment corresponds to three frames: "X-0", "X-1", and "X-2" (some short segments have two frames). Figure 2 The specific contents are as follows:

[0022]

[0023] The original video file is read using video processing tools (which can be implemented with computer vision libraries) to establish a time-series data stream of video frames, preparing for subsequent frame-by-frame extraction. Starting from the video's initial frame, each frame is captured sequentially, forming a continuous frame sequence. Since the video is 33 seconds long and has a frame rate of 30 fps, a total of 990 frames are extracted, starting from frame 0 (corresponding to time 0) and ending at frame 989 (corresponding to time 32.97), with an interval of approximately 1 ÷ 30 ≈ 0.033 seconds between each frame. Each frame is annotated with a unique frame number (e.g., frame 0, frame 1, etc.) and a corresponding timestamp (e.g., "frame 18 → 0.6 seconds") to ensure accurate frame sequence timing and lay the foundation for subsequent processing.

[0024] The Bhattacharyya distance metric measures the similarity of consecutive frame image histograms. A sudden drop in inter-frame similarity is identified as a shot switch. For example, when a video switches from Group 0 (a static display of a speaker) to Group 1 (a person interacting with the speaker), at frame 96 (corresponding to 3.2 seconds), the image changes from a single speaker to a scene with both the speaker and the person in the same frame. The histogram similarity drops significantly, identifying this frame as the shot boundary between Group 0 and Group 1.

[0025] Based on the identified switching points, the frame sequence is divided into several segments, each containing a continuous frame interval. For example, Group 2 corresponds to frames 198-252 (corresponding to a time period of 6.6-8.4 seconds), totaling 55 frames. This segment primarily demonstrates the use of a speaker mounted on a bicycle handlebar.

[0026] Standard storyboard segments (≥50 frames): Images are extracted from the storyboard segment at the 20%, 50%, and 80% frame positions as representative frames. Taking Group 2 as an example, the storyboard segment has 55 frames. The following frame positions are calculated: 20% frame position: 198 + 55 × 20% = 209 frames, encoded as "2-0"; 50% frame position: 198 + 55 × 50% = 225 frames, encoded as "2-1"; 80% frame position: 198 + 55 × 80% = 241 frames, encoded as "2-2". Short storyboard segments (<50 frames): If the storyboard segment has less than 50 frames, only two frames at the 20% and 80% positions are extracted.

[0027] The extracted representative frames are encoded in an "XY" format (X is the shot number, Y is the frame sequence within the group) to form structured representative frame encoded data, such as "3-1" representing the middle frame of the third shot segment. The representative frames of each shot segment are then spliced ​​together into a continuous image and fed into a large language model (such as GPT) using structured text instructions (prompts) for semantic parsing. For example, after splicing frames "7-0," "7-1," and "7-2" from Group 7 (speaker waterproofing test), the model is guided by prompts to identify the waterproofing feature and output a JSON-formatted scene semantic description. Testing has shown that extracting 990 frames from a 33-second video takes an average of less than one second, meeting lightweight processing requirements. The frame sequence serves as the basis for subsequent shot segmentation, representative frame extraction, and semantic understanding. The accuracy of its timing directly impacts the accuracy of shot segment division and semantic parsing, ensuring the reliability of the entire processing process.

[0028] S102 , pre-processing and shot segmentation are performed on image features in the video frame sequence to generate representative frame coding data.

[0029] In one implementation, the Bhattacharyya distance is used to measure the similarity of the histogram of the continuous frame images of the video to identify the shot switching points and complete the coarse segmentation. If the length of a single segment exceeds 10 seconds, the fine segmentation strategy is triggered. Taking the outdoor portable speaker product introduction video as an example, the video length is 33 seconds, the frame rate is 30fps, the total number of frames is (33×30=990\) frames, and the resolution is 1920×1080. The video displays the product through 11 segment segments (Group0 to Group10), covering static display, scene interaction, functional testing and other contents. The specific segmentation corresponding images and scenes are as follows:

[0030]

[0031] The histograms of the RGB channels of consecutive frame images are calculated separately, and the Bhattacharyya distance is used to measure the similarity of adjacent frames. The formula is: where p (i) and q (i) is the normalized value of the histogram of the two frames. The smaller the distance value, the higher the similarity.

[0032] For example, consider Group 0 (frames 0-95, static display) and Group 1 (frame 96 and later, character interaction): The average Bhattacharyya distance for the preceding frames (frames 90-94) is 0.21, indicating high image similarity. However, the distance between frames 95 and 96 jumps to 0.72, indicating a shot switch from a static display of a single speaker to a scene interacting with both speakers. This is the initial segmentation of the video into 11 segments (Groups 0-10).

[0033] A refined segmentation strategy triggers secondary segmentation if the segment duration exceeds 10 seconds (300 frames). For example, if a segment originally lasts 12 seconds (360 frames), a switching point is added at 6 seconds (180 frames), splitting it into two 6-second segments. This avoids semantic redundancy and improves shot granularity accuracy.

[0034] For each shot segment, images are extracted at 20%, 50%, and 80% of the frame position to construct semantic representative frames. If the shot does not meet the requirement of extracting three frames, two frames at 20% and 80% are selected, and finally encoded in "XY" format to generate representative frame encoding data, where X is the shot number and Y is the frame sequence within the group. For each shot segment, representative frames are extracted at 20%, 50%, and 80% of the frame position. If the shot number is less than 50 frames, two frames at 20% and 80% are selected, and finally encoded in "XY" format (X is the shot number and Y is the frame sequence within the group).

[0035] Standard storyboard segment (≥50 frames) processing, taking Group 2 (frames 198-252, a total of 55 frames, bicycle handlebar installation scene) as an example: 20% frame position: 198 + 55 × 20% = 209 frames (6.97 seconds), coded "2-0", corresponding to the perspective of the speaker installed on the left side of the handlebar; 50% frame position: 198 + 55 × 50% = 225 frames (7.5 seconds), coded "2-1", corresponding to the intermediate frame of 360° rotation after installation; 80% frame position: 198 + 55 × 80% = 241 frames (8.03 seconds), coded "2-2", corresponding to the right side perspective (reflecting the stability of the installation).

[0036] For short shot segments (<50 frames), take Group 7 (frames 630-677, a total of 48 frames, representing the waterproof test scene) as an example: 20% of the frame: 630 + 48 × 20% ≈ 640 frames (21.33 seconds), encoded as "7-0," showing the initial state of the speaker partially immersed in the water; 80% of the frame: 630 + 48 × 80% ≈ 668 frames (22.27 seconds), encoded as "7-2," showing the speaker fully submerged and still operating normally. The encoding rules are as follows: X: shot number (incrementing from 0, e.g., Group 5 corresponds to X = 5); Y: frame number within the group (Y = 0 / 1 / 2 for three frames, Y = 0 / 2 for two frames, and Y = 1 is skipped).

[0037] Advantages of Bhattacharyya distance: Compared with the pixel difference method, it is insensitive to lighting changes (such as light fluctuations in outdoor scenes) and can accurately identify scene switches (such as from indoor static displays to outdoor cycling scenes). Dynamic segmentation strategy: Long storyboard segments (>10 seconds) trigger fine segmentation to avoid semantic redundancy of a single shot (such as splitting a complex function demonstration into multiple segments) and improve the accuracy of semantic extraction. The value of the coding system: The "XY" coding is strictly bound to the storyboard segments, which facilitates rapid semantic positioning (such as directly associating the "multiple speakers TWS interconnection + waterproof" scene through "Group8 (8-0 to 8-2)"). Through the above process, the storyboard segmentation, representative frame extraction and semantic association of outdoor portable speaker videos can be efficiently completed, providing a structured foundation for subsequent AI model analysis and product selling point refinement.

[0038] S103: Based on the semantic extraction rules of the storyboard segments, semantic analysis is performed on the representative frame image features, and semantic sequence data is generated by combining the guidance of the structured text instructions on the large language model and the model output characteristics.

[0039] In one embodiment, the representative frames of the storyboards are input into the large language model in the form of continuous images, and guided by structured text instructions containing content target descriptions, image number interpretations, user intention settings, and output format constraints, the model uses the collaborative processing capabilities of the image features and text instructions to make the model return JSON format structured information containing the structured group identification numbers of the video storyboard segments and their corresponding semantics and the scene semantic description, forming semantic sequence data according to the storyboard order. Taking the outdoor portable speaker waterproof test storyboard Group 7 as an example, the storyboard Group 7 corresponds to the attached Figure 2 The three frames "7-0," "7-1," and "7-2" in the video are: 7-0: The initial state of the speaker partially immersed in the water (21.2 seconds); 7-1: A close-up of the button operation when the speaker is completely submerged (22.1 seconds); 7-2: The speaker's LED lights illuminate amidst the splashing water (22.8 seconds). These frames are stitched together horizontally to form a continuous 1920×3240 pixel image, simulating the timing of storyboard playback.

[0040] The structured text instruction (Prompt) is constructed, and the content objective is as follows: "Analyze the waterproof function test details in Group 7." The image number is explained as follows: "7-0 is the 20% frame position of the frame (beginning of immersion), 7-1 is the 50% frame position (underwater operation), and 7-2 is the 80% frame position (waterproof effect display)". Figure 2 The frame extraction logic is set. The user intent is to "extract core selling points such as waterproof rating, test scenario, and product feedback," which meets the "Generate product selling points" requirement. The output format constraint is specifically "return JSON format, including the group number and scenario description, example: [{"Group":7,"Content":"..."}]."

[0041] The large language model processing flow is as follows: the model extracts visual features from the stitched image: it identifies the "water ripples" and "speaker immersion angle" in frame 7-0 and determines that the test environment is a flowing water environment; it analyzes the "underwater button pressing action" and "screen lighting" in frame 7-1 and associates them with the "operation feasibility" of the user's intention; and it analyzes the "water splash trajectory" and "LED light color change" in frame 7-2 to infer the waterproof seal and functional stability.

[0042] Combined with the "waterproof level" keyword in the prompt, the model extracts implicit information from the image: the waterproof level is determined by the "IP68 logo" in frame 7-2; the test duration is inferred based on the "timer showing 30 minutes" in frame 7-1, corresponding to the time-sensitive logic of "10 seconds triggering fine segmentation."

[0043] JSON output and semantic sequence generation, the fields correspond to: "Group:7" and attached Figure 2The mirror numbers in the Figure 2 In the "Waterproof" tag, "LEDlights" is associated with Group4's lighting effect display storyboard.

[0044] The output of Group7 is spliced ​​with the adjacent storyboards in sequence and displayed in json format.

[0045] [{"Group":6,"Content":"Speakerplacedingrassnearwaterenvironment..."},

[0046] {"Group":7,"Content":"IP68waterproofcertificationtested..."},

[0047] {"Group":8,"Content":"TWSpairingdemonstrationwithmultiplespeakersinwater..."}]. Sequence order and attached Figure 2 The storyboard numbers (0-10) are exactly the same, forming a continuous scene semantic chain.

[0048] Script creation: Extract Group7's "IP68waterproof" segment from the semantic sequence and insert the advertising script: "[00:21:20-00:22:80] Immersive water flow test, IP68-level waterproof performance ensures all-weather outdoor use - corresponding to Group7 storyboard." Generate product tags based on Group7 semantics: "#IP68WaterproofSpeaker#OutdoorPortableDevice".

[0049] The guided design of the prompt, with its clear command constraints (e.g., "waterproof rating"), improves model output accuracy by approximately 30%, avoiding overly generalized descriptions (e.g., simply outputting "the speaker is submerged in water"). Processing three frames of image data with the prompt takes approximately 0.2 seconds, and the entire semantic parsing of a 30-second video takes less than 1 second, meeting the "high efficiency" metric and making it suitable for real-time video processing. JSON sequences organized by group number can be directly used for video indexing. For example, "Group 7" can be used to quickly locate a video clip (21.2-22.8 seconds) from a waterproof test, achieving a triple mapping of "semantics, time, and image."

[0050] Through the specific scenario of outdoor speaker storyboards, the complete process from multi-frame input to semantic sequence generation was demonstrated, reflecting the core technical path of "visual features + structured instructions → model collaboration → semantic output".

[0051] S104, based on the multi-frame splicing input method, the cross-storyboard representative frame features are processed, and the semantic understanding logic is dynamically optimized by combining the text instruction structure design and model output rules to generate structured scene understanding data.

[0052] In one embodiment, the representative frames of each storyboard are input into the large language model in the form of continuous images, supplemented by structured text instructions containing content target descriptions, image number interpretations, user intent settings, and output format constraints for guidance. Based on the JSON format structured scene understanding information output by the model, including the structured group identification numbers and scene semantic descriptions of the video storyboard segments and their corresponding semantics, the semantic understanding logic is dynamically optimized to generate structured scene understanding data. Figure 2 Take the "Outdoor Portable Speaker TWS Pairing Mirror Group 8" as an example, Group 8 corresponds to the attached Figure 2 The three frames "8-0," "8-1," and "8-2" in the video are: 8-0: The initial image of the two speakers placed in a stream (23.6 seconds); 8-1: A close-up of the pairing progress displayed on the mobile app (24.5 seconds); and 8-2: The splashing scene of the two speakers playing synchronously (25.4 seconds). Group 8 is spliced ​​with the representative frames of the adjacent frames Group 7 (waterproof test) and Group 9 (waterproof details) into a continuous image, forming the cross-frame input sequence: [Group 7-2, Group 8-0, Group 8-1, Group 8-2, Group 9-0], simulating the continuity of the "waterproof test → TWS pairing → subsequent waterproofing" scene.

[0053] A structured text prompt was constructed to "analyze the TWS pairing function of multiple speakers in Group 8, and combine the previous and subsequent frames to understand the synergy between waterproofing and pairing." "Group 8-0 is the start frame (pairing starts), Group 8-1 is the middle frame (APP operation), Group 8-2 is the end frame (pairing complete), and Group 7-2 is the previous waterproofing test result frame." The user intent was set to "extract selling points such as pairing method, synchronization effect, and functional stability in a waterproof environment."

[0054] The output format is "Return JSON format, including Group number, scene description and cross-storyboard association information, as shown below:

[0055] [{"Group":8,"Content":"...","CrossGroup":[{"Group":7,"Rela tion":"Pairing after waterproofing"}...]}]".

[0056] The large language model performs visual analysis on the cross-screen stitching image: it identifies Group 8-0's "two speakers + water flow background", associates it with Group 7-2's "waterproof test" conclusion, and infers the "feasibility of pairing in a waterproof environment"; parses Group 8-1's "the mobile app interface displays 'TWSPairingSuccessful'", extracts the pairing time (5 seconds) and operation steps; analyzes Group 8-2's "synchronous vibration of dual speakers in water splashes", and determines the audio synchronization accuracy (such as ≤50ms delay). Combined with the "synergistic relationship" instruction in the prompt, the model dynamically adjusts the understanding logic: it finds that Group 7-2's "IP68 logo" and Group 8-2's "underwater playback" form a causal relationship, generating a reasoning chain of "waterproof certification → underwater function available"; based on Group 9-0's "more waterproof test" screen, it corrects Group 8's semantic description to "the entire TWS pairing process in a waterproof environment" to avoid isolated understanding. The results returned by the model are as follows:

[0057] [{"Group":8,"Content":"Twospeakerssuccessful lypairedviaTWSinflowingwaterwithin5seconds,withsynchronizedaudiooutputdemons tratedbywatersplashpatterns.FunctionalitybuildsonGroup7'sIP68w aterproofcertification.",

[0058] "CrossGroup":[{"Group":7,"Relation":"Precedingwaterprooftestenablesunderwaterpairing"},

[0059] {"Group":9,"Relation":"Subsequenttestsvalidatepairingdurability"}]}].

[0060] Logically associate the output of Group8 with the semantics of adjacent storyboards to generate scene-level understanding data: {"Scene":"WaterproofTWSPairing","Segments":[

[0061] {"Group":7,"Role":"Pre-condition","Semantic":"IP68waterproofcertification"},

[0062] {"Group":8,"Role":"Coreaction","Semantic":"5sTWSpairinginw ater"},

[0063] {"Group":9,"Role":"Validation","Semantic":"Post-pairingwaterproofdurability"}],

[0064] "LogicChain":"Waterproofcapability→TWSpairingfeasibility

[0065] →Functionvalidation"}. The data structure is reflected in the three-layer abstraction of "storyboard-scene-logic chain". The "Group:7 / 8 / 9" in the output data is directly associated with the attached Figure 2 The storyboard numbers in the figure prove that the "XY" coding system supports semantic tracking across storyboards, as shown in the attached figure. Figure 2 The spatial adjacent relationship between Group7-2 and Group8-0 is transformed into a semantic temporal relationship.

[0066] By dynamically optimizing the logic, the model's understanding accuracy is increased by approximately 40%, avoiding semantic fragmentation within a single storyboard (e.g., understanding Group 8 as a "normal pairing" and ignoring the waterproofing requirement). Processing a 5-frame, cross-storyboard image with a prompt takes approximately 0.5 seconds, and understanding the entire scene of a 30-second video takes less than 1 second, making it suitable for batch processing of short videos. The output JSON format supports flexible invocation of subsequent tasks, such as quickly locating related storyboards using the "CrossGroup" field, automating the entire process from "scene review to content generation to quality verification."

[0067] S105: Process the structured scene understanding data to generate application data including video semantic summary, script creation semantic fragment, video tags and product selling points.

[0068] In one implementation, based on the JSON-formatted structured scene understanding information output by the large language model, the structured scene understanding data is parsed and refined according to the application requirements of the video content to generate application data for constructing video semantic summaries, providing semantic fragments for script creation and advertising copy embedding, as well as video tags and product selling point descriptions. Taking the outdoor portable speaker TWS pairing scenario as an example, the cross-storyboard scene data, which comes from the JSON structured information output by the large language model, includes cross-storyboard associations for Groups 7-9. The specific content is as follows:

[0069] {"Scene":"WaterproofTWSPairing",

[0070] "Segments":[

[0071] {"Group":7,"Role":"Pre-condition","Semantic":"IP68 waterproof certification, 1 meter water depth tested for 30 minutes"},

[0072] {"Group":8,"Role":"Coreaction","Semantic":"Complete dual-speaker TWS pairing in water within 5 seconds, audio synchronization delay ≤ 50ms"},

[0073] {"Group":9,"Role":"Validation","Semantic":"After pairing, continue with waterproof test, and all functions are normal."}],

[0074] "LogicChain": "Waterproof capability → TWS pairing feasibility → Functional durability verification"}

[0075] The video semantic summary generation process is as follows: the core information of the scene is extracted, and key elements are extracted from the "Segments" field: Waterproof premise: "IP68 certification, 1 meter water depth 30 minutes test" (Group 7); pairing process: "TWS pairing is completed in 5 seconds of water flow, audio synchronization" (Group 8); verification result: "The waterproof function remains effective after pairing" (Group 9).

[0076] Summary organized by "LogicChain": [00:21:20-00:28:40] First, it passed the IP68 waterproof certification test (Group7), then completed the dual-speaker TWS pairing in a water flow environment in 5 seconds (Group8), and finally verified that the waterproof function was not affected after pairing (Group9), achieving dual functional protection of waterproofness and interconnection in outdoor scenarios.

[0077] Script creation semantic fragment generation, advertising script scene splitting, specifically generating fragments according to storyboard roles (Pre-condition / Coreaction / Validation):

[0078] Group7 clip: "(Close-up of water hitting the speaker) '1 meter deep, 30 minutes of continuous immersion - IP68 waterproof, laying the foundation for outdoor connected experiences.'"

[0079] Group8 clip: "(Mobile app pairing animation + water splashing) '5-second ultra-fast TWS pairing! Whether by the stream or in the rain, audio syncs without delay.'"

[0080] Group9 clip: "(Speaker playing music in water) 'Continue testing after pairing is complete, waterproof performance is always online - this is the ultimate form of outdoor speakers.'"

[0081] Mapping with the storyboard code, each clip is labeled with the corresponding Group number, such as "(Group7-2 shot: close-up of water splashing)".

[0082] The process for generating video tags and product selling points is as follows: a hierarchical tagging system. Core function tags include: #IP68 waterproof #TWS pairing #outdoor speaker. Scenario-specific tags include: #stream usage #rainy day connectivity #waterproof audio. The generation logic extracts keywords from the "Semantic" field, such as "IP68," "TWS pairing," and "outdoor scenario."

[0083] The product's selling points are described as follows, based on a cross-storyboard logic chain integrating selling points: "Waterproof × Interconnectivity Dual Breakthrough": ① IP68 waterproof certification, can withstand immersion in 1 meter of water for 30 minutes (Group 7); ② Quickly complete TWS pairing in 5 seconds in a flowing environment, with audio synchronization delay ≤50ms (Group 8); ③ Continuous waterproof testing after pairing, functional stability is not affected (Group 9).

[0084] Demand-driven semantic extraction emphasizes "functional impact" in ad scripts (e.g., "5-second pairing") and "keyword coverage" in tags (e.g., "#WaterproofAudio"). This embodies the principle of "analyzing data based on application needs." This approach balances efficiency and accuracy. Generating application data for a 30-second video takes less than 0.3 seconds, meeting the "high efficiency" metric. Furthermore, the semantic segment matches the storyboard content at a rate of >90%.

[0085] The reusability of structured data. JSON scene data can be directly used for: inserting corresponding copy according to the Group number during video editing; e-commerce platforms automatically generate selling point details pages with storyboard indexes, corresponding to the invention purpose of "structured expression".

[0086] S106: Process the application data based on preset application scenario rules to generate video structured semantic annotation results.

[0087] In one embodiment, based on the pre-set application scenario rules for constructing video semantic summaries, providing semantic segments for script creation and advertising copy embedding, and generating video tags and product selling points, application data containing video semantic summaries, script creation semantic segments, video tags, and product selling points are processed to generate video structured semantic annotation results. Taking the "outdoor portable speaker waterproof and TWS pairing scenario" as an example, three types of rules are preset: video semantic summary rules: must include "storyboard timestamp + core function + logical relationship", such as the causal description of "waterproof test → pairing function"; script embedding rules are as follows, requiring "shot picture description + functional selling point + emotional guidance" to adapt to the rhythm of 15-second short videos; label selling point rules are as follows, following the three-layer structure of "core technology + scenario application + user value", such as "#IP68 waterproof #outdoor scene #audio synchronization".

[0088] The video semantic summary generation process is as follows: input data: JSON data from structured scene understanding (Group 7-9 cross-shot association); extract timestamps: Group 7 (21.2-22.8 seconds), Group 8 (23.6-25.4 seconds), Group 9 (26.2-28.4 seconds); integrate the logical chain "IP68 waterproof certification (Group 7) → 5-second water flow pairing (Group 8) → waterproof verification after pairing (Group 9)".

[0089] The final output result is [00:21:20-00:28:40]. After passing the IP68 waterproof test, the speaker completed TWS pairing in 5 seconds under water, and continued to verify the waterproof function after pairing, realizing the "waterproof-interconnection" integrated experience in outdoor scenarios.

[0090] The process of processing semantic fragments for script creation is as follows: the input data is a storyboard-level semantic sequence (e.g., Group8's "TWS pairing completed in 5 seconds"); the scene description is: "(Group8-2 shot: dual speakers splashing water simultaneously)"; the selling point is enhanced: "5-second ultra-fast pairing! Underwater audio synchronization with zero delay"; the emotional guidance is: "Outdoor adventures will never be without the anxiety of losing device connection";

[0091] The output clip is (close-up shot: water splashes in sync with the music rhythm), '5 seconds! Complete the dual-speaker TWS pairing in the stream - IP68 waterproof black technology, so that outdoor music is never interrupted. 'The "Group8-2" marked in the clip corresponds to the attached Figure 2 The "8-2" frame position is the image played synchronously by two speakers. The three types of application data are integrated into the three-dimensional structure of "time-storyboard-semantics":

[0092] {"Summary":"[00:21:20-00:28:40]...",

[0093] "ScriptSegments":[

[0094] {"Group":7,"Segment":"..."},

[0095] {"Group":8,"Segment":"..."}],

[0096] "TagsAndFeatures":{

[0097] "Tags":["#IP68 waterproof","#TWS pairing"],

[0098] "Features":"「Triple waterproof interconnection protection」:①...②..."}}.

[0099] Time-storyboard mapping, Group 7-9 in the annotation results corresponds to the attached Figure 2 For example, Group8’s tag “#TWSpairing” is directly associated with the attached Figure 2 "8-0 to 8-2" frame position in;

[0100] Semantic-image mapping, the “water flow pairing” in the selling point description corresponds to the attached Figure 2 The "8-2" frame of Group 8 (the splashing water scene) realizes the "semantic-image" two-way traceability.

[0101] Application scenarios for the annotation results include automated short video editing, where the script segments in the annotation results are automatically matched to video frames by Group number. For example, "Group 8-2 shot" corresponds to the splash scene at 25.4 seconds, achieving precise synchronization between the text and the visuals. E-commerce detail page generation also involves directly referencing the annotated structured data in the product selling points section. The shot number serves as a time index for the "Function Demonstration Video." Users clicking the "IP68 Waterproof" tag can jump to the test video clip from Group 7. Marketing data analysis uses the annotated shot relationships (e.g., Group 7 → Group 8) to analyze user interest in the "waterproof → pairing" feature combination and optimize advertising strategies.

[0102] Rule-driven semantic standardization uses pre-set rules to ensure that annotation results comply with industry standards. For example, e-commerce selling points must include test data (30 minutes for a 1-meter water depth) to avoid ambiguous descriptions. Highly efficient annotation is achieved, with the entire annotation process taking less than 0.5 seconds for a 30-second video, meeting the efficiency requirement of "less than 1 second for a 30-second video" and suitable for batch video processing. Multimodal data consistency ensures consistency across the entire "video frame-semantics-application" chain, including timestamps, shot numbers, and semantic descriptions.

[0103] This application proposes a video understanding method based on frame sequence abstraction and language model guidance. By collaborating with frame sequence abstraction and a large language model, it realizes the structured analysis of video content and solves the problems of large computing resource overhead, insufficient semantic expression and lack of structured output in the existing technology. The Bhattacharyya distance is used to measure the similarity of the histogram of continuous frame images, and the lens switching points are identified to complete the coarse segmentation. The fine segmentation strategy is triggered for the storyboard segments with a duration of more than 10 seconds. The average time taken for a 30-second video is <1 second, ensuring efficiency and granularity accuracy. Representative frames are extracted according to the frame positions of 20%, 50%, and 80% of the storyboard segments. For short storyboards, two frames at 20% and 80% are taken and encoded in the "XY" format (X is the storyboard number and Y is the frame number) to achieve frame-level semantic abstraction.

[0104] The representative frames are spliced ​​together and input into the GPT model. Guided by a structured prompt containing target descriptions, number explanations, intent settings, and format constraints, the model returns JSON data containing group numbers and scene descriptions to form a semantic sequence. Based on the semantic sequence, a video semantic summary is constructed, script creation fragments are extracted, video tags and product selling points are generated, and finally the structured semantic annotation of the video is completed. The computation is lightweight, only key frames are processed instead of full frames, which greatly reduces computing resource consumption and is adapted to lightweight environments. Semantic accuracy, dynamic segmentation and representative frame strategies cover key scenes to avoid semantic omissions caused by frame averaging. Structured output, semantic results in JSON format directly support downstream tasks such as copywriting generation and product extraction, improving subsequent processing efficiency.

[0105] This method can efficiently process videos such as product introductions and advertisements. The resulting semantic summaries, script snippets, and selling point tags can be directly applied in scenarios such as intelligent creation and marketing analysis, enabling automated understanding and value mining of video content. By synergizing image and language models, it overcomes the technical bottlenecks of traditional video understanding and provides a new path for multimodal information processing.

[0106] In one embodiment, Figure 3 As shown, the present application also provides a video understanding device based on frame sequence abstraction and language model guidance, comprising:

[0107] The acquisition module 301 is used to acquire the original video, extract continuous frame images from the original video, and generate a video frame sequence;

[0108] The processing module 302 is used to pre-process the image features in the video frame sequence and segment the shots to generate representative frame coding data; perform semantic analysis on the representative frame image features based on the storyboard segment semantic extraction rules, and generate semantic sequence data by combining the guiding role of structured text instructions on the large language model and the model output characteristics; process the cross-storyboard representative frame features based on the multi-frame splicing input method, and dynamically optimize the semantic understanding logic by combining the text instruction structure design and model output rules to generate structured scene understanding data; process the structured scene understanding data to generate application data including video semantic summary, script creation semantic fragments, video tags and product selling points; process the application data based on preset application scenario rules to generate video structured semantic annotation results.

[0109] An electronic device comprises a first processor; and a memory for storing executable instructions of the first processor;

[0110] The first processor is configured to execute the video understanding method based on frame sequence abstraction and language model guidance according to any one of claims 1 to 6 by executing the executable instructions.

[0111] The methods and / or embodiments in the embodiments of the present application can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. When the computer program is executed by a processing unit, the above-mentioned functions defined in the method of the present application are performed.

[0112] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. Computer-readable media may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.

[0113] In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0114] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0115] As another aspect, embodiments of the present application further provide a computer-readable medium, which may be included in the device described in the above embodiments, or may exist independently and not incorporated into the device. The computer-readable medium carries one or more computer program instructions, which are executable by a processor to implement the steps of the methods and / or technical solutions of the aforementioned embodiments of the present application.

[0116] Each embodiment in this application is described in a related manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for evaluating the video understanding method, electronic device, electronic device, and readable storage medium embodiment based on frame sequence abstraction and language model guidance, since they are basically similar to the above-mentioned video understanding method embodiment based on frame sequence abstraction and language model guidance, the description is relatively simple. For related parts, refer to the partial description of the above-mentioned video understanding method embodiment based on frame sequence abstraction and language model guidance.

[0117] Those skilled in the art should understand that the above embodiments are only for clearly illustrating the present application and are not intended to limit the scope of the present application.

Claims

1. A video understanding method based on frame sequence abstraction and language model guidance, characterized by: Obtaining the original video, extracting continuous frame images from the original video, and generating a video frame sequence; Preprocess the image features in the video frame sequence and perform shot segmentation to generate representative frame encoding data; Based on the semantic extraction rules of the storyboard segments, the representative frame image features are semantically analyzed, and the guidance of the structured text instructions on the large language model and the model output characteristics are combined to generate semantic sequence data; Based on the multi-frame splicing input method, the cross-storyboard representative frame features are processed, and the semantic understanding logic is dynamically optimized by combining the text instruction structure design and model output rules to generate structured scene understanding data; Process structured scene understanding data to generate application data including video semantic summaries, script creation semantic fragments, video tags, and product selling points; The application data is processed based on preset application scenario rules to generate video structured semantic annotation results.

2. The method according to claim 1, wherein: Preprocess the image features in the video frame sequence and perform shot segmentation to generate representative frame encoding data, including: The Bhattacharyya distance is used to measure the similarity of the histograms of consecutive video frames to identify shot switching points and complete coarse segmentation. If the length of a single shot segment exceeds 10 seconds, a fine segmentation strategy is triggered. For each segment, images are extracted at 20%, 50%, and 80% of the frame position to construct a semantic representative frame. If the segment does not meet the extraction requirements of three frames, two frames at 20% and 80% are selected. Finally, the data is encoded in the "XY" format to generate the representative frame encoding data, where X is the segment number and Y is the frame sequence number within the group.

3. The method according to claim 2, wherein: Based on the semantic extraction rules of the storyboard segments, the representative frame image features are semantically analyzed. Combined with the guidance of the structured text instructions on the large language model and the model output characteristics, semantic sequence data is generated, including: The representative frames of the storyboards are input into the large language model in the form of continuous images. Through the guidance of structured text instructions containing content target description, image number interpretation, user intention setting, and output format constraints, the model uses the collaborative processing ability of image features and text instructions to enable the model to return JSON format structured information containing structured grouping identification numbers of video storyboard segments and their corresponding semantics and scene semantic descriptions, forming semantic sequence data according to the order of the storyboards.

4. The method according to claim 1, wherein: Based on the multi-frame splicing input method, the cross-screen representative frame features are processed, and the semantic understanding logic is dynamically optimized by combining the text instruction structure design and model output rules to generate structured scene understanding data, including: The representative frames of each storyboard are input into the large language model in the form of continuous images, supplemented by structured text instructions containing content target descriptions, image number explanations, user intent settings, and output format constraints for guidance. Based on the JSON format structured scene understanding information output by the model, including the structured grouping identification numbers and scene semantic descriptions of the video storyboard segments and their corresponding semantics, the semantic understanding logic is dynamically optimized to generate structured scene understanding data.

5. The method according to claim 4, characterized in that: Process structured scene understanding data to generate application data including video semantic summaries, script creation semantic fragments, video tags, and product selling points, including: Based on the JSON format structured scene understanding information output by the large language model, the structured scene understanding data is parsed and refined according to the application requirements of the video content, generating application data for constructing video semantic summaries, providing semantic fragments for script creation and advertising copy embedding, as well as video tags and product selling point descriptions.

6. The method according to claim 5, characterized in that: Process application data based on preset application scenario rules to generate video structured semantic annotation results, including: Based on the preset application scenario rules of constructing video semantic summaries, providing semantic fragments for script creation and advertising copy embedding, and generating video tags and product selling points, the application data containing video semantic summaries, script creation semantic fragments, video tags and product selling points are processed to generate video structured semantic annotation results.

7. A video understanding device based on frame sequence abstraction and language model guidance, characterized in that: The device comprises: An acquisition module is used to acquire the original video, extract continuous frame images from the original video, and generate a video frame sequence; The processing module is used to pre-process the image features in the video frame sequence and perform shot segmentation to generate representative frame encoding data; perform semantic analysis on the representative frame image features based on the storyboard segment semantic extraction rules, and generate semantic sequence data by combining the guiding role of structured text instructions on the large language model and the model output characteristics; process the cross-storyboard representative frame features based on the multi-frame splicing input method, and dynamically optimize the semantic understanding logic by combining the text instruction structure design and model output rules to generate structured scene understanding data; process the structured scene understanding data to generate application data including video semantic summary, script creation semantic fragments, video tags and product selling points; process the application data based on preset application scenario rules to generate video structured semantic annotation results.

8. An electronic device, characterized in that: include: a first processor; and a memory for storing executable instructions of the first processor; The first processor is configured to execute the video understanding method based on frame sequence abstraction and language model guidance according to any one of claims 1 to 6 by executing the executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the second processor, the video understanding method based on frame sequence abstraction and language model guidance according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Intelligent video propaganda product design method based on large model and knowledge base

    CN122064843A