Prompt construction method and system of multi-mode large language model, computer equipment and medium
By using maximum semantic richness sampling and motion reconstruction techniques, keyframes are extracted from video streams and spatiotemporally correlated with each other, solving the efficiency and accuracy problems of spatial reasoning in dynamic scenes for multimodal large language models and achieving efficient and accurate spatial reasoning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-03-31
AI Technical Summary
Existing visual cues methods struggle to balance spatial reasoning accuracy and computational efficiency when processing pure visual video streams. They lack structured sampling capabilities for visual diversity and motion reconstruction mechanisms, resulting in insufficient spatiotemporal correlation in dynamic scenes.
By using maximum semantic richness sampling and motion reconstruction techniques, keyframes are extracted from the video stream to generate motion trajectory information, and spatiotemporal correlation coding is performed to construct a multimodal prompt input to a multimodal large language model.
It improves the spatial reasoning accuracy and computational efficiency of multimodal large language models in dynamic scenes, reduces the number of redundant frames, provides clear camera motion cues, and enhances the model's spatiotemporal understanding of dynamic scenes.
Smart Images

Figure CN121767993A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal large language model training technology, and in particular to a method, system, computer device and medium for constructing prompts for a multimodal large language model. Background Technology
[0002] With the rapid development of artificial intelligence technology, Multimodal Large Language Models (MLLM) have demonstrated enormous potential in understanding and generating cross-modal content, and have become a core tool in fields such as intelligent video analytics, autonomous driving, and robot navigation. Particularly in complex and high-risk environments such as ultra-high voltage (UHV) substations, there is an urgent practical need to utilize pure visual video streams for real-time spatial reasoning and safety risk assessment of maintenance personnel behavior. The operating scenarios in UHV substations are dynamically changing, and the behavior of maintenance personnel involves complex spatiotemporal relationships and spatial layout changes. The ability to accurately capture their movement trajectories, relative distances, and interaction relationships directly affects the stable operation of the power system and personnel safety. Therefore, improving the spatial intelligent reasoning capabilities of MLLM in pure visual video streams has become a key challenge in intelligent security and industrial automation.
[0003] In the field of MLLM, using prompting techniques to guide or enhance model capabilities is a common approach. Existing visual prompting methods have made some progress in multimodal tasks. For example, comparative document 1 (application publication number CN 119495128 A) discloses an efficient and adaptive image-to-person interaction detection method and system. This method utilizes a concept-guided memory module and a lightweight adapter to achieve person interaction detection in a training-free or fine-tuning mode using a pre-trained model (such as Contrastive Language-Image Pre-training, a pre-trained neural network model trained on a large number of image-text pairs through contrastive learning, enabling it to directly understand the semantic relationships between images and text, abbreviated as CLIP). This method performs well in static image recognition tasks, improving detection efficiency with a small number of samples by storing specific domain visual knowledge and universal domain semantic knowledge. However, this type of method mainly targets object recognition and interaction prediction in static images. Its design goals do not fully consider the temporal characteristics of video streams, making it difficult to directly transfer to spatial reasoning tasks in dynamic video scenes.
[0004] Most existing visual cueing technologies focus on solving the problems of recognition, localization, or interaction in static images, failing to effectively address the specific challenges of complex spatial intelligent reasoning in pure visual video streams. Specifically, existing methods have the following limitations: On the one hand, static object recognition cueing techniques based on masks or labels (such as mask sets or coarse correspondence methods), while enhancing the model's object localization ability in basic visual tasks, rely on static masks or coarse location labels and lack modeling support for the relationships between objects and scene-level spatial structure in dynamic scenes, making them unsuitable for spatiotemporal reasoning in video sequences. On the other hand, cueing methods focusing on other dimensions (such as visual attention cues or illusion mitigation techniques), while performing well in improving detail perception or reducing error output, do not provide mechanisms to enhance MLLM's understanding of spatial layout, temporal sequence, or self-motion, thus failing to solve the motion correlation problem in video streams. Although Comparative Document 1 achieves efficient adaptation in static human interaction detection, its technical solution is limited to the image domain and does not involve temporal processing of video streams. The method's memory module and adapter design focuses on object-level semantic features, but fails to address core issues such as visual homogeneity and unknown motion in long video sequences. For example, in monitoring videos of UHV substations, the actions of maintenance personnel involve continuous displacement and path changes. Existing technologies lack the ability to sample diverse keyframes of the video and cannot reconstruct camera motion trajectories, making it difficult for the model to capture spatial cues in dynamic scenes.
[0005] Therefore, existing visual cueing methods have limitations when processing purely visual video streams. On the one hand, due to the lack of structured sampling capabilities for visual diversity, existing methods typically employ a uniform frame sampling strategy, resulting in high redundancy of selected keyframe information (i.e., visual homogeneity), which limits the model's complete reconstruction of spatial layout. On the other hand, these methods fail to implement motion reconstruction mechanisms, failing to provide camera trajectories or spatiotemporal cues (i.e., unknown motion problems), making it difficult for the model to infer object displacement and temporal relationships. These shortcomings collectively make it difficult for Multimodal Large Language Models (MLLMs) to balance computational efficiency and inference accuracy in spatial reasoning tasks. Summary of the Invention
[0006] To address the aforementioned shortcomings or drawbacks, this invention provides a method, system, computer device, and medium for constructing prompts for a multimodal large language model, which can solve the technical problem that existing prompting methods struggle to balance spatial reasoning accuracy and computational efficiency.
[0007] This invention provides a method for constructing prompts in a multimodal large language model, comprising: Acquire the input video stream and extract a set of keyframes from the video stream.
[0008] Perform a motion reconstruction process on the video stream to generate motion trajectory information.
[0009] The motion trajectory information is visualized to generate a trajectory visualization map.
[0010] The keyframe set and motion trajectory information are spatiotemporally correlated and encoded to generate enhanced keyframes.
[0011] Based on enhanced keyframes and trajectory visualization, multimodal prompts are constructed by integrating visual and text inputs, and then the multimodal prompts are input into a pre-defined multimodal large language model.
[0012] According to a second aspect, this invention provides a prompting system for a multimodal large language model, comprising: The keyframe extraction module is used to acquire the input video stream and extract a set of keyframes from the video stream.
[0013] The motion trajectory generation module is used to perform motion reconstruction on the video stream and generate motion trajectory information.
[0014] The visualization generation module is used to visualize motion trajectory information and generate trajectory visualization maps.
[0015] The keyframe enhancement module is used to perform spatiotemporal correlation encoding of the keyframe set and motion trajectory information to generate enhanced keyframes.
[0016] The training prompt building module is used to construct multimodal prompts based on enhanced keyframes and trajectory visualizations by integrating visual and text inputs, and then input the multimodal prompts into a pre-defined multimodal large language model.
[0017] According to a third aspect, the present invention provides a computer device comprising: At least one processor; and The memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which enable the at least one processor to execute any multimodal large language model prompting construction method in the embodiments of the present invention.
[0018] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute a prompting construction method for any multimodal large language model in the embodiments of the present invention.
[0019] The present invention provides a method for constructing prompts for a multimodal large language model. This method is achieved collaboratively through steps such as video stream processing, keyframe extraction, motion reconstruction, spatiotemporal coding, and multimodal prompt construction. The method includes: first, acquiring an input video stream and extracting a set of keyframes from the video stream; then, performing motion reconstruction on the video stream to generate motion trajectory information; next, visualizing the motion trajectory information to generate a trajectory visualization map; then, spatiotemporally associating the keyframe set with the motion trajectory information to generate enhanced keyframes; finally, based on the enhanced keyframes and the trajectory visualization map, constructing a multimodal prompt input by integrating visual and text input, and inputting the multimodal prompt into a preset multimodal large language model for spatial reasoning.
[0020] In this technical solution, the present invention addresses the problem of high information redundancy caused by visual homogeneity as described in the background art. By extracting keyframe sets, it reduces the number of redundant frames in the video stream, solving the defects of low input signal-to-noise ratio and heavy computational burden under the traditional uniform sampling strategy. Addressing the problem of insufficient spatiotemporal correlation modeling caused by unknown motion, it generates motion trajectory information through motion reconstruction, providing clear camera motion cues and enhancing the model's spatiotemporal understanding of dynamic scenes. By fusing keyframes and motion trajectory information through spatiotemporal correlation encoding, it generates enhanced keyframes, further improving the accuracy of spatial reasoning. By constructing multimodal prompt input and integrating visual and textual information, the multimodal large language model can reason based on direct visual evidence, avoiding speculative output relying on common sense priors. Therefore, the technical solution of this invention solves the technical problem that existing prompting methods struggle to balance spatial reasoning accuracy and computational efficiency, enabling multimodal large language models to achieve efficient and accurate spatial reasoning, improving the model's computational efficiency, reasoning accuracy, and environmental adaptability. Attached Figure Description
[0021] Figure 1 This is a flowchart of a method for constructing prompts for a multimodal large language model according to an embodiment of the present invention; Figure 2 This table shows the comparison results of the complex spatial reasoning capabilities of various multimodal large language models of the present invention before and after SEE and TREK enhancement in the VSI-BENCH benchmark test; Figure 3 This table shows the comparison results of the spatiotemporal reasoning capabilities of various multimodal large language models of the present invention before and after SEE and TREK enhancement in the STI-BENCH benchmark test; Figure 4 This is a flowchart illustrating the prompting construction and system workflow of a multimodal large language model according to an embodiment of the present invention; Figure 5This is a schematic diagram of the structure of a prompting system for a multimodal large language model according to an embodiment of the present invention; Figure 6 This is a block diagram of a computer device for implementing embodiments of the present invention. Detailed Implementation
[0022] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0023] During the development of this invention, the inventors, through extensive experiments and data analysis, revealed the intrinsic relationship between visual homogeneity and unknown motion: traditional uniform sampling methods not only struggle to capture the semantic diversity in video streams but also introduce high information redundancy and missing motion cues, making it difficult for traditional methods to achieve accurate spatial reasoning under purely visual conditions. Based on this relationship, the inventors innovatively proposed this technical solution, utilizing Maximum Semantic Richness Sampling (MSRS) and motion reconstruction techniques. Keyframes are filtered through an equalization-TopK strategy (a multi-objective optimization algorithm for selecting keyframes from video streams), and frame-trajectory association is achieved by combining spatiotemporal coding. This enables training-free, plug-and-play multimodal cue construction, embodying the core concept of fusing structured perception and motion cues.
[0024] Specifically, through comparative experiments, the invention team discovered that the uniform frame sampling method used in traditional MLLM pipelines suffers from problems such as high visual homogeneity and lack of motion information. These technical defects lead to a decrease in the signal-to-noise ratio of the model input, insufficient spatiotemporal correlation modeling, and a lack of deep understanding of dynamic scenes in multimodal large language models. However, the maximum semantic richness sampling and motion reconstruction collaborative framework proposed in this invention can effectively improve the semantic diversity and motion trajectory integrity of keyframes, achieving a significant improvement in average accuracy on VSI-BENCH (Video Spatial Intelligence BENCHmark) and STI-BENCH (Spatiotemporal Intelligence BENCHmark) benchmarks.
[0025] Therefore, this invention provides a method for constructing prompts for a multimodal large language model based on the first aspect, which can be applied to an intelligent spatial reasoning system (hereinafter referred to as the "system"). This system can be deployed locally or run on edge computing devices or server platforms via cloud services to complete real-time spatial reasoning and multimodal prompt generation tasks for video streams.
[0026] Specifically, the system's physical devices include, but are not limited to, video surveillance cameras, embedded edge computing units, and GPU (Graphics Processing Unit) server clusters. These devices need to have high-performance video decoding capabilities, support real-time processing of computer vision algorithms, and multimodal data parallel computing, thereby supporting efficient analysis of large-scale video streams and ensuring that the system can operate stably in resource-constrained environments and achieve low-latency spatial inference.
[0027] like Figure 1 As shown, the method may include: Step S110: Obtain the input video stream and extract a set of keyframes from the video stream.
[0028] Here, a video stream refers to a sequential sequence of images arranged in chronological order, and a keyframe set refers to a set of representative frames selected from the video stream through a sampling strategy.
[0029] Specifically, the system can process the video stream using a maximum semantic richness sampling strategy, extract semantic information for each frame using a perceptual model (such as YOLO, a real-time object detection system), and filter key frames based on an equalization selection algorithm.
[0030] For example, the system monitors video streams from an UHV substation (resolution...) In a pixel (30 frames per second) frame, the object category (such as people, equipment, tools) in each frame is detected by a perception model. Ten key frames are selected using an equalization selection algorithm (such as prioritizing the frame containing the most object categories) to generate a key frame set (total size of five megabytes).
[0031] Step S120: Perform motion reconstruction on the video stream to generate motion trajectory information.
[0032] Among them, the motion reconstruction process refers to the method of estimating the camera motion trajectory through visual odometry technology, and the motion trajectory information refers to the parameter sequence describing the changes in camera position and attitude.
[0033] Specifically, the system can process consecutive video frames using feature extraction algorithms (such as Oriented FAST and Rotated BRIEF, a fast feature point detection and description algorithm, or ORB for short), match inter-frame feature points, and calculate the camera's relative motion parameters based on the essential matrix.
[0034] For example, the system performs motion reconstruction on a 30-second video stream (900 frames in total), extracts ORB feature points for each frame (about 1,000 feature points per frame), estimates the essential matrix using RANSAC (Random Sample Consensus), and generates motion trajectory information (containing 300 pose points, 1.2 megabytes of data).
[0035] Step S130: Visualize the motion trajectory information to generate a trajectory visualization map.
[0036] Visualization processing refers to the technology of converting motion trajectory data into intuitive images, and trajectory visualization refers to an image that displays the motion path in graphical form.
[0037] Specifically, the system can use a rendering engine to map motion trajectory information into a two-dimensional bird's-eye view or a three-dimensional trajectory map, and use a continuous color map (such as a rainbow color map) to color the trajectory points over time.
[0038] For example, the system inputs motion trajectory information into a rendering engine (such as Matplotlib, a Python plotting library) to generate a bird's-eye view (size...). The system outputs a 3D trajectory graph (containing XYZ coordinate axes) and a 2D trajectory plot (pixels). The trajectory points are colored in chronological order (from blue to red to indicate increasing time). The output trajectory visualization is in PNG format, Portable Network Graphics, a lossless compressed bitmap graphics format, and is 500 kilobytes in size.
[0039] Step S140: Spatiotemporally correlate the keyframe set with the motion trajectory information to generate enhanced keyframes.
[0040] Spatiotemporal correlation coding refers to the technique of superimposing motion trajectory markers on keyframes to establish spatiotemporal correlation, while enhanced keyframes refer to keyframes that contain trajectory perception information.
[0041] Specifically, the system can overlay trajectory-aware markers on keyframes using image processing libraries (such as OpenCV, an open-source computer vision library), including frame index text and color markers, and make the colors consistent with the colors of the corresponding time points in the trajectory visualization.
[0042] For example, the system selects a frame (timestamp 5 seconds) from the keyframe set, overlays the frame index text "T=5s" and a red circular marker (50 pixels in diameter), the color of which matches the color (red) of the trajectory point at the fifth second in the trajectory visualization, to generate an enhanced keyframe (resolution). (pixels, 200 kilobytes in size).
[0043] Step S150: Based on the enhanced keyframes and trajectory visualization, construct a multimodal prompt by integrating visual and text input, and input the multimodal prompt into the preset multimodal large language model.
[0044] Multimodal prompts refer to a mixed input format that includes visual and textual data. Visual input refers to image data, and text input refers to descriptive language data.
[0045] Specifically, the system can combine enhanced keyframes and trajectory visualizations into visual input through the data integration module, generate text prompts (such as instructions describing the meaning of trajectory colors), and finally submit them to a multimodal large language model (such as InternVL, a visual language model) through the API (Application Programming Interface).
[0046] For example, the system integrates five enhanced keyframes and a trajectory visualization as visual input, generates a text prompt "The red mark indicates a time point of five seconds. Please analyze the movement trajectory of the person". The text is submitted to the multimodal large language model through a single API (Application Programming Interface) call, and outputs spatial reasoning results (such as "The person moves from point A to point B, a distance of about ten meters").
[0047] In other embodiments, such as Figure 2This table shows a comparison of the complex spatial reasoning capabilities of various multimodal large language models before and after SEE and TREK enhancements in the VSI-BENCH benchmark test. Specifically, this table exemplifies the implementation effect of the technical solution of this invention (SEE and TREK enhancement method) on open-source multimodal large language models. Taking the LLaVA-NeXT-Video-7B model as an example, without applying the enhancement method of this invention, its average accuracy in the nine spatial reasoning tasks in the VSI-BENCH benchmark test was 32.5 percentage points; after applying the SEE and TREK enhancement method, its average accuracy increased to 33.8 percentage points, a relative improvement of 1.3 percentage points. In particular, in the relative direction recognition (Rel. Dir.) task, the enhanced model performance improved from 35.1 percentage points to 39.9 percentage points, an improvement of 4.8 percentage points; in the route planning (Route Plan) task, it improved from 34.0 percentage points to 36.6 percentage points, an improvement of 2.4 percentage points. For example, the InternVL3-8B model, after SEE and TREK enhancements, saw its average accuracy improve from 40.2 percentage points to 43.2 percentage points. Specifically, the absolute object distance estimation (Abs. Dist.) task improved by 7.4 percentage points (from 18.5% to 25.8%), and the relative distance judgment (Rel. Dist.) task improved by 10.0 percentage points (from 29.3% to 32.8%). These results demonstrate that the cue construction method of this invention, which utilizes maximum semantic richness sampling of keyframes, motion trajectory reconstruction, and spatiotemporal correlation encoding, can effectively improve the performance of various multimodal large language models in complex spatial reasoning tasks.
[0048] In other embodiments, such as Figure 3The table shows a comparison of the spatiotemporal reasoning capabilities of various multimodal large language models before and after enhancement with SEE and TREK in the STI-BENCH benchmark test. Specifically, the table exemplifies the implementation effect of the technical solution of this invention in dynamic understanding tasks, covering nine spatiotemporal reasoning sub-tasks, including three-dimensional spatial measurement, displacement and path length, and trajectory description. Taking the INTERNVL3-14B model as an example, without the enhancement method of this invention, its average accuracy in dynamic understanding-related tasks is 30.8 percentage points; after applying the SEE and TREK enhancement methods, the average accuracy increases to 32.2 percentage points. In particular, in the trajectory description task, the model performance significantly improves from 25.6 percentage points to 34.2 percentage points, a relative improvement of 8.6 percentage points; in the displacement and path length task, it improves from 19.1 percentage points to 22.4 percentage points, an improvement of 3.3 percentage points. Taking the QWEN2.5-VL-7B model as an example, the enhanced performance improved its average accuracy in dynamic understanding tasks from 35.6 percentage points to 36.9 percentage points. The most significant improvement was seen in velocity and acceleration estimation tasks, which increased from 41.6 percentage points to 57.8 percentage points, a rise of 16.2 percentage points. For the QWEN2.5-VL-32B model with a larger number of parameters, the enhanced performance in trajectory description tasks improved from 22.7 percentage points to 33.5 percentage points, a relative improvement of 10.8 percentage points. These results demonstrate that the explicit trajectory information provided by motion reconstruction technology and the multimodal cues constructed through spatiotemporal correlation coding in this invention can effectively enhance the ability of multimodal large language models to model spatiotemporal continuity, especially showing significant advantages in tasks requiring the understanding of dynamic scenarios such as motion trajectories and velocity changes.
[0049] Therefore, according to the above implementation method, the system first acquires the input video stream and extracts a set of keyframes from the video stream; then, it performs a motion reconstruction process on the video stream to generate motion trajectory information; next, it performs visualization processing on the motion trajectory information to generate a trajectory visualization map; then, it performs spatiotemporal correlation encoding on the set of keyframes and the motion trajectory information to generate enhanced keyframes; finally, based on the enhanced keyframes and the trajectory visualization map, it constructs a multimodal prompt input by integrating visual input and text input, and inputs the multimodal prompts into a preset multimodal large language model for spatial reasoning.
[0050] Specifically, in this implementation, the technical solution addresses the issue of high information redundancy caused by visual homogeneity as described in the background technology. By extracting keyframe sets, the number of redundant frames in the video stream is reduced, overcoming the shortcomings of low input signal-to-noise ratio and heavy computational burden under traditional uniform sampling strategies. Regarding the insufficient spatiotemporal correlation modeling caused by unknown motion, motion trajectory information is generated through motion reconstruction, providing clear camera motion cues and enhancing the model's spatiotemporal understanding of dynamic scenes. By fusing keyframes and motion trajectory information through spatiotemporal correlation encoding, enhanced keyframes are generated, further improving the accuracy of spatial reasoning. Furthermore, by constructing multimodal prompt inputs and integrating visual and textual information, the multimodal large language model can reason based on direct visual evidence, avoiding speculative outputs relying on common sense priors. Therefore, this implementation's technical solution solves the technical problem of existing prompting methods struggling to balance spatial reasoning accuracy and computational efficiency, enabling the multimodal large language model to achieve efficient and accurate spatial reasoning, improving the model's computational efficiency, reasoning accuracy, and environmental adaptability.
[0051] In some embodiments, the step of extracting a set of keyframes from a video stream includes: The video stream is filtered using a predefined maximum semantic richness sampling strategy to obtain multiple keyframes that form a keyframe set.
[0052] Among them, the maximum semantic richness sampling strategy is a method that analyzes the semantic information of video frames through a perceptual model and filters key frames based on an equalization selection algorithm, aiming to improve the semantic diversity and information richness of key frames.
[0053] Specifically, the system can perform object detection on each frame of the video stream by calling the perception model to obtain semantic information (such as a list of object categories), and then apply an equalization selection algorithm to dynamically select keyframes based on the semantic information.
[0054] For example, the system uses YOLO (You Only Look Once, a real-time object detection system) as a perception model to detect the object category (such as personnel, equipment, tools) in each frame from a video stream of monitoring a UHV substation (30 seconds long, 30 frames per second). Then, it selects ten key frames based on the equalization selection algorithm to form a key frame set (total size of five megabytes).
[0055] The steps for performing keyframe filtering on a video stream include: The semantic information of each frame is extracted from the video stream using the perceptual model specified in the maximum semantic richness sampling strategy. The perceptual model refers to the visual analysis model used to detect the object category in the video frame.
[0056] Semantic information refers to data such as the object categories and their confidence levels detected in each frame, which are used to characterize the semantic content of the frame.
[0057] Specifically, the system can load a pre-trained perceptual model (such as YOLOv5) to perform forward propagation calculations on each frame of the video stream and output the detected object bounding boxes, class labels, and confidence scores.
[0058] For example, the system assigns a frame resolution of The pixel image is processed, and the perception model outputs the detection results, which include three objects (person confidence 0.9, device confidence 0.8, and tool confidence 0.7), and generates a semantic information log (in JSON format, JavaScript Object Notation, a lightweight data exchange format).
[0059] Based on the semantic information of each frame, multiple keyframes are selected from the frame sequence of the video stream using the equalization selection algorithm specified in the maximum semantic richness sampling strategy.
[0060] Among them, the balanced selection algorithm is a strategy for filtering key frames by iteratively optimizing the semantic diversity and temporal distribution of frames. Its execution includes filtering conditions based on priority order.
[0061] Specifically, the algorithm first calculates the overlap (i.e. the number of common categories) between each frame and the selected keyframe category pool, and prioritizes the new frame with the smallest overlap. Based on minimizing the overlap, it selects the new frame with the largest self-detected category count. When semantic conditions (such as category count) are the same, it selects the new frame with the earliest timestamp.
[0062] For example, in video stream processing, the currently selected keyframe category pool contains {people, equipment}. The semantic information of new frame A is {people, tools}, with an overlap of two (common category "people"); the semantic information of new frame B is {vehicles, ground}, with an overlap of zero; therefore, frame B is selected first. If multiple frames have zero overlap, the frame with the largest category count is selected (e.g., frame C has four categories), ultimately filtering out multiple keyframes.
[0063] Therefore, according to the above implementation method, the system can efficiently extract semantically diverse keyframe sets from video streams, reduce information redundancy, and provide high-quality visual input for subsequent motion reconstruction and spatial reasoning.
[0064] In some embodiments, the step of selecting multiple keyframes from a frame sequence of a video stream using an equalization selection algorithm specified in a maximum semantic richness sampling strategy includes: The execution of the balanced selection algorithm includes filtering conditions in the following priority order.
[0065] Among them, the balanced selection algorithm is a screening method that balances information richness by iteratively optimizing the semantic diversity and temporal distribution of keyframes. Its core is to apply the conditions of minimizing overlap, maximizing class count, and earliest timestamp in sequence.
[0066] Specifically, the system can traverse the frame sequence of the video stream, calculate the overlap with the selected keyframe category pool, its own detected category count, and timestamp for each frame, and select candidate frames in priority order. For example, when processing a video stream of monitoring a UHV substation (900 frames in total), the system initializes the selected keyframe category pool to be empty, and then applies the balanced selection algorithm starting from the first frame, ultimately selecting ten keyframes.
[0067] Prioritize selecting new frames with the smallest overlap with the current selected keyframe category pool. The selected keyframe category pool refers to the set of all object categories contained in the selected keyframes. The overlap is determined by the number of common categories between the new frame and the selected keyframe category pool.
[0068] Minimizing overlap means prioritizing the selection of new frames that share the fewest common categories with the already selected category pool in order to maximize semantic diversity.
[0069] Specifically, the system can calculate the intersection of the semantic information (such as the object category list) of each new frame with the selected keyframe category pool. The smaller the intersection, the lower the overlap. For example, if the selected keyframe category pool contains {personnel, equipment}, the semantic information of new frame A is {personnel, tools}, the common category is {personnel}, and the overlap is one; the semantic information of new frame B is {vehicle, ground}, the common category is empty, and the overlap is zero; therefore, frame B is selected first.
[0070] Based on the condition of minimizing overlap, a new frame that satisfies the condition of maximizing the detection category count is selected. The detection category count is used to measure the total number of different object categories detected in the corresponding frame.
[0071] Among them, maximizing the detection category count means prioritizing the selection of new frames that contain the most object categories in order to increase the amount of information in a single frame.
[0072] Specifically, the system can count the number of unique object categories detected by the perceptual model in each new frame; a higher count indicates higher semantic richness. For example, in new frames with zero overlap, frame C detects three object categories (people, equipment, tools), with a detection category count of three; frame D detects two object categories (vehicles, ground), with a detection category count of two; therefore, frame C is preferred.
[0073] When semantic conditions are the same, select the new frame that satisfies the earliest timestamp condition.
[0074] Among them, "same semantic conditions" means that multiple new frames cannot be distinguished in terms of overlap and detection category count, and "earliest timestamp" means that the frame that appears earliest in time sequence is selected first.
[0075] Specifically, the system can compare the timestamps of new frames (such as frame number or time point) and select the frame with the smallest value. For example, frames E and F have zero overlap and three detection class counts, but frame E's timestamp is at the fifth second and frame F's timestamp is at the eighth second; therefore, frame E is selected first.
[0076] Frames that satisfy the conditions of minimizing overlap, maximizing class count, and earliest timestamp selection are selected as multiple keyframes.
[0077] This step refers to integrating the results of all filtering conditions and outputting the final set of keyframes.
[0078] Specifically, the system can apply priority conditions iteratively: first, it filters the set of frames with the smallest overlap, then selects the subset with the largest detection category count, and finally selects the frame with the earliest timestamp. For example, the system initially filters out five frames with zero overlap from the video stream, selects three frames with the largest detection category count (both counts are three), and finally selects the two frames with the earliest timestamps (such as the fifth and sixth seconds) as keyframes to add to the set.
[0079] Therefore, according to the above implementation method, the system can efficiently select semantically diverse and temporally distributed key frames from the video stream, reduce redundant information, and provide high-quality input for subsequent spatial reasoning.
[0080] In some embodiments, a motion reconstruction process is performed on the video stream to generate motion trajectory information, including: Visual odometry (VO) is a technique used to perform motion estimation on video streams. VO refers to a computer vision-based method for estimating camera motion trajectories.
[0081] Motion estimation processing refers to the process of calculating the relative motion parameters of the camera by analyzing the feature changes between consecutive frames in a video sequence.
[0082] Specifically, the system can establish motion constraints between adjacent frames by initializing camera pose parameters and then processing the video stream frame by frame. For example, the system performs motion estimation on a video of an UHV substation inspection (30 seconds long, 1920×1080 pixels in resolution), initializes the camera pose as an identity matrix, and then calculates the relative motion frame by frame.
[0083] Feature extraction and feature matching operations are performed on consecutive frames in the video stream to obtain the inter-frame feature correspondence.
[0084] Feature extraction and feature matching refer to the process of detecting significant feature points from video frames and establishing inter-frame correspondences. Inter-frame feature correspondences refer to the set of successfully matched feature point pairs in adjacent frames.
[0085] Specifically, the system can detect feature points in each frame using an ORB (Oriented Fast and Rotated BRIEF, a fast local feature detection and description algorithm) feature extractor, and calculate the Hamming distance between feature descriptors using a brute-force matching algorithm (an exhaustive search method for finding nearest neighbor matches in a feature point set) to establish matching relationships. For example, the system extracts features from two consecutive frames (frame numbers 100 and 101) in a video, obtaining 1000 and 1050 ORB feature points respectively. Through feature matching, it obtains 900 matching point pairs, generating inter-frame feature correspondence data.
[0086] Based on the inter-frame feature correspondence, the essential matrix is calculated using a robust estimation method. The essential matrix is used to describe the camera motion geometry relationship between adjacent video frames.
[0087] Among them, robust estimation method refers to parameter estimation technique that can effectively handle the influence of outliers, and the essential matrix is a three-by-three matrix that describes the basic geometric constraints of camera motion between two frames.
[0088] Specifically, the system can use RANSAC (Random Sample Consensus) to randomly sample five pairs of points from the matched point pairs to calculate candidate solutions for the essential matrix, and then select the optimal solution based on the reprojection error. For example, the system uses the RANSAC algorithm to randomly sample from nine hundred matched point pairs one hundred times, selecting five pairs of points each time to calculate the essential matrix, and finally selecting the matrix with the most interior points (e.g., 800 interior points) as the optimal essential matrix.
[0089] The essential matrix is decomposed to obtain the camera's relative rotation and relative translation parameters.
[0090] Among them, matrix decomposition refers to the process of obtaining camera motion parameters by performing singular value decomposition on the essential matrix.
[0091] Specifically, the system can decompose the essential matrix into the product of three matrices using the singular value decomposition algorithm, and then recover the relative rotation matrix and translation vector based on the camera intrinsic parameters.
[0092] For example, the system performs singular value decomposition on the essential matrix to obtain a relative rotation matrix (3×3 orthogonal matrix) and a relative translation vector (three-dimensional vector), and the translation vector is normalized.
[0093] The relative rotation parameters and relative translation parameters are integrated into global motion trajectory information through recursive accumulation.
[0094] Among them, the recursive accumulation method refers to the method of obtaining the global trajectory by sequentially accumulating the relative motion parameters of adjacent frames.
[0095] Specifically, the system can use a pose graph optimization method to associate the relative pose of each frame with the global coordinate system, and use Lie algebras (a mathematical tool for describing rigid body motion in 3D space) for pose updates and smoothing optimization. For example, starting from the first frame, the system accumulates the relative rotation and translation parameters of each frame into the global coordinate system, generating motion trajectory information containing 300 pose points (data format: (a sequence of homogeneous transformation matrices).
[0096] Therefore, according to the above implementation method, the system can accurately reconstruct the camera motion trajectory from the video stream, providing a precise motion information basis for subsequent spatiotemporal correlation coding.
[0097] In some embodiments, the motion trajectory information is visualized to generate a trajectory visualization graph, including: The motion trajectory information is rendered into a visual image, which includes a bird's-eye view and a 3D trajectory map. The 3D trajectory map refers to a three-dimensional image that displays the motion trajectory in three-dimensional space, while the bird's-eye view refers to a two-dimensional image that displays the motion trajectory from an overhead perspective.
[0098] Rendering refers to the technical process of converting motion trajectory data into visual images, including operations such as coordinate mapping, viewpoint transformation, and graphic drawing.
[0099] Specifically, the system can use computer graphics libraries (such as OpenGL, an open graphics library, a cross-platform two-dimensional and three-dimensional graphics application programming interface) to project the three-dimensional coordinate data of the motion trajectory onto a two-dimensional plane to generate a bird's-eye view, while maintaining the three-dimensional spatial relationship to generate a three-dimensional trajectory map.
[0100] For example, the system renders a segment of motion trajectory (containing 300 pose points) of an UHV substation inspection camera into a bird's-eye view (1920×1080 pixels resolution) and a 3D trajectory map (containing XYZ 3D coordinate axes), and both images are saved in PNG (Portable Network Graphics) format.
[0101] The trajectory points in the motion trajectory information are processed by time coloring using a continuous color map to generate time-colored trajectory points.
[0102] Among them, time coloring refers to a visualization technique that assigns different colors based on the timestamp information of trajectory points, while continuous color mapping refers to a color mapping scheme in which colors change continuously with numerical values.
[0103] Specifically, the system can use a color mapping function to map the timestamp (relative to the start time) of each trajectory point to a predefined color space (such as the HSV color space, hue, saturation, and brightness), generating a color-coded sequence of trajectory points. For example, for a 30-second motion trajectory, a rainbow color chart can be used for coloring: the start time (0 seconds) corresponds to blue, the middle time (15 seconds) corresponds to green, and the end time (30 seconds) corresponds to red, generating a time-colored sequence of trajectory points.
[0104] Based on motion trajectory information, bird's-eye view and 3D trajectory map are generated using image rendering technology.
[0105] Image rendering technology refers to computational methods that convert geometric data into raster images, including lighting calculations, texture mapping, and anti-aliasing.
[0106] Specifically, the system can load trajectory point data using a 3D rendering engine (such as Three.js, a JavaScript 3D graphics library based on WebGL), set the camera perspective (a top-down view for bird's-eye view and a perspective view for 3D trajectory maps), add coordinate axes and color-coded legends, and finally output a high-quality visualization image. For example, the system uses the Three.js engine to generate a 3D trajectory map: setting a perspective camera (60-degree field of view), adding a grid ground reference system, drawing gradient lines between trajectory points, and rendering an image file of 1200 by 800 pixels.
[0107] Therefore, according to the above implementation method, the system can transform abstract motion trajectory information into intuitive visual graphics, providing visual input with both spatial information and temporal clues for multimodal large language models.
[0108] In some embodiments, the keyframe set and motion trajectory information are spatiotemporally correlated and encoded to generate enhanced keyframes, including: Trajectory-aware markers are overlaid on each keyframe in the keyframe set. The trajectory-aware markers are a combination of markers that associate each keyframe with motion trajectory information through visual identifiers. The trajectory-aware markers include frame index markers and color markers. The frame index markers are used to indicate the temporal position of each keyframe in the video stream.
[0109] Overlay processing refers to the process of adding visual marker elements to keyframes using image processing techniques.
[0110] Specifically, the system can use the drawing functions of computer vision libraries (such as OpenCV, an open-source computer vision library) to add text-based frame index markers and graphic color markers at specified locations in keyframe images. For example, the system can add a frame index marker "T=5s" (30-pixel font size) to the upper left corner of a keyframe (5 seconds past the timestamp) in a UHV substation monitoring video, and a circular color marker with a diameter of 50 pixels to the lower right corner.
[0111] Set the color of the color marker to match the trajectory color at the corresponding time point in the motion trajectory information, so that each keyframe is explicitly associated with the motion trajectory.
[0112] Among them, color consistency setting refers to the matching process of determining the display color of the color marker on the key frame based on the color value of the corresponding time point in the motion trajectory visualization.
[0113] Specifically, the system can query the color value (such as RGB color value) of a specific time point in the trajectory visualization map through a color mapping table, and assign that color value to the corresponding color marker on the keyframe. For example, if the trajectory point at the fifth second in the motion trajectory visualization map is red (RGB value 255,0,0), the system will set the color marker of the corresponding keyframe to the same red, establishing a visual association.
[0114] Integrate trajectory-aware markers into each keyframe to generate enhanced keyframes.
[0115] Among them, integrated processing refers to the technical operation of combining trajectory-aware markers with the original keyframe images to form the final output image.
[0116] Specifically, the system can use image fusion algorithms (such as alpha fusion) to fuse the trajectory-aware marker layer with the keyframe base layer, maintaining the marker's recognizability without significantly affecting the original image content. For example, the system can overlay frame index markers and color markers onto the keyframe in a semi-transparent manner (alpha value 0.8) to generate enhanced keyframes (saved in JPEG format, a widely used lossy compressed digital image format with a quality factor of 85%).
[0117] Therefore, according to the above implementation method, the system can establish the spatiotemporal association between keyframes and motion trajectories through visual markers, generate enhanced keyframes containing rich contextual information, and provide enhanced input with both visual content and motion cues for multimodal large language models.
[0118] In some embodiments, based on enhanced keyframes and trajectory visualizations, multimodal cues are constructed by integrating visual and text inputs, and the multimodal cues are submitted to a multimodal large language model, including: The enhanced keyframes and trajectory visualizations are combined into a visual input set, which is a combination of data containing enhanced keyframe images and trajectory visualizations.
[0119] Combining visual input sets refers to the process of organizing multiple visual elements into an input data structure acceptable to the model according to a preset order and format.
[0120] Specifically, the system can arrange enhanced keyframes and trajectory visualizations into a grid layout using image stitching algorithms, or package multiple images into a single file using multiple image encapsulation formats (such as PDF, portable document format).
[0121] For example, the system will use five enhanced keyframes (each at a resolution) A set of visual inputs (total image size 4800×2800 pixels) is generated by arranging a trajectory visualization (bird's-eye view, resolution 1200×800 pixels) in a two-row, three-column grid.
[0122] Generate text prompts, which include explanations of the color consistency and time sequence of the trajectory-aware markers.
[0123] Among them, text prompts refer to natural language text that describes the visual input content and its relationships, and is used to guide the model to understand multimodal input.
[0124] Specifically, the system can generate structured text through template filling, including trajectory color mapping instructions (such as "red indicates start time, blue indicates end time") and time sequence indicators (such as "keyframes are arranged in ascending order of timestamps"). For example, the generated text prompt is: "The trajectory color changes from red to blue to indicate the time from start to end; the five keyframes are arranged in time sequence T=1s, 3s, 5s, 7s, 9s; the color markers are consistent with the trajectory point colors." The visual input set and text prompts are integrated into a multimodal prompt, which refers to a mixed data format used for input into a multimodal large language model.
[0125] Integration processing refers to the technical operation of combining and encapsulating visual data and text data according to the input specifications required by the model.
[0126] Specifically, the system can encode image data into pixel value tensors and text data into a sequence of labeled IDs using a multimodal data encapsulation library (such as Hugging Face's processor class), and add special delimiters required by the model. For example, the system uses the CLIP (Contrastive Language-Image Pre-trained) model processor to convert the visual input set into a four-dimensional tensor (batch size 6, number of channels 3, height 224, width 224), convert the text prompt content into a sequence of labeled IDs (length 128), and add [IMG] and [TXT] delimiters to generate multimodal prompt data.
[0127] Multimodal prompts are submitted to the multimodal large language model through a single forward propagation. A single forward propagation refers to the computational process of processing multimodal input data in one go.
[0128] In this context, a single forward propagation refers to the complete inference process of inputting multimodal input data into the model computation graph at once and obtaining the output.
[0129] Specifically, the system can input multimodal cue data into the model computation graph through the model inference interface of a deep learning framework (such as PyTorch, an open-source Python machine learning library), perform end-to-end forward computation, and obtain the model output. For example, the system inputs multimodal cue data into an InternVL (a visual language model) model, and through one forward propagation computation, obtains a spatial inference result (such as "a person moves from area A to area C via point B, with a total displacement of approximately 15 meters").
[0130] Therefore, according to the above implementation method, the system can construct multimodal prompts containing rich visual context and clear textual guidance, and achieve accurate spatial reasoning analysis through an efficient single-inference process.
[0131] In other embodiments, such as Figure 4The diagram illustrates the prompting construction and system workflow of a multimodal large language model. This diagram exemplifies the complete implementation process of the present invention's technical solution in an intelligent inspection scenario of an ultra-high voltage substation. Specifically, the system first acquires the input video stream (such as a 1080P resolution video collected by a substation inspection robot). Using a maximum semantic richness sampling strategy, it performs semantic analysis on the video frames based on the YOLO object detection model, filtering a set of keyframes from the continuous video stream at intervals of "every N frames." For example, from a 30-second inspection video (900 frames in total), the system uses an equalization selection algorithm to select 10 of the most semantically representative keyframes, covering various scenarios such as personnel operation and equipment status. In the motion reconstruction module, the system uses unsupervised feature matching technology to track and triangulate feature points between consecutive frames, reconstructing the camera's three-dimensional motion trajectory. As shown in the diagram of the three-dimensional coordinate system and feature point matching, the system gradually calculates the camera's pose changes through the feature correspondence between adjacent frames (such as feature point matching between Frame 1 and Frame 2), generating trajectory data containing translation and rotation parameters. Subsequently, the keyframes are associated with the motion trajectory through a spatiotemporal coding module. For example, trajectory color markers (such as red for the start point and blue for the end point) and frame index labels are overlaid on the keyframes to form enhanced keyframes. Finally, the system integrates the enhanced keyframes, 3D trajectory visualization, and text instructions (such as "the motion path of the analyst from device A to device B") to construct a multimodal prompt input, which is then submitted to a multimodal large language model for spatial reasoning. This flowchart clearly reveals the entire process from the original video input to the generation of multimodal prompts, demonstrating the core innovation of this invention in enhancing the model's spatial reasoning capabilities through structured sampling, motion reconstruction, and spatiotemporal coding.
[0132] Figure 5 This is a structural block diagram of a prompting system for a multimodal large language model according to an embodiment of the present invention.
[0133] like Figure 5 As shown, the prompting system for this multimodal large language model includes: The keyframe extraction module 210 is used to acquire the input video stream and extract a set of keyframes from the video stream.
[0134] The motion trajectory generation module 220 is used to perform a motion reconstruction process on the video stream and generate motion trajectory information.
[0135] The visualization generation module 230 is used to visualize motion trajectory information and generate trajectory visualization diagrams.
[0136] The keyframe enhancement module 240 is used to perform spatiotemporal correlation encoding of the keyframe set and motion trajectory information to generate enhanced keyframes.
[0137] The training prompt building module 250 is used to construct multimodal prompts based on enhanced keyframes and trajectory visualizations by integrating visual and text inputs, and input the multimodal prompts into a preset multimodal large language model.
[0138] The specific functions and examples of each module and submodule of the device in this embodiment of the invention can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0139] According to embodiments of the present invention, the above-described method of the present invention can be applied to a computer device and a readable storage medium.
[0140] Figure 6 A schematic block diagram of an example computer device 600 that can be used to implement embodiments of the present invention is shown. The computer device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The computer device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0141] like Figure 6 As shown, the computer device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the computer device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0142] Multiple components in computer device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows computer device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0143] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as a method for constructing prompts for a multimodal large language model. For example, in some embodiments, a method for constructing prompts for a multimodal large language model can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the computer device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the method for constructing prompts for a multimodal large language model described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured, by any other suitable means (e.g., by means of firmware), to perform a prompting method for a multimodal large language model.
[0144] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0145] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0146] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0148] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0149] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0150] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0151] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for constructing prompts in a multimodal large language model, characterized in that, include: Acquire the input video stream and extract a set of keyframes from the video stream; Perform a motion reconstruction process on the video stream to generate motion trajectory information; The motion trajectory information is visualized to generate a trajectory visualization map; The keyframe set and the motion trajectory information are spatiotemporally correlated and encoded to generate enhanced keyframes; Based on the enhanced keyframes and the trajectory visualization, a multimodal cue is constructed by integrating visual and text inputs, and the multimodal cue is submitted to the multimodal large language model.
2. The method according to claim 1, characterized in that, The step of extracting a set of keyframes from the video stream includes: The video stream is subjected to keyframe filtering using a predefined maximum semantic richness sampling strategy to obtain multiple keyframes that constitute the keyframe set. The step of performing keyframe filtering on the video stream includes: The semantic information of each frame is extracted from the video stream using the perceptual model specified in the maximum semantic richness sampling strategy. The perceptual model refers to a visual analysis model used to detect the object category in the video frame. Based on the semantic information of each frame, multiple keyframes are selected from the frame sequence of the video stream using the equalization selection algorithm specified in the maximum semantic richness sampling strategy.
3. The method according to claim 2, characterized in that, The step of selecting multiple keyframes from the frame sequence of the video stream using the equalization selection algorithm specified in the maximum semantic richness sampling strategy includes: The execution of the balanced selection algorithm includes the following priority order of filtering conditions: Prioritize selecting a new frame with the smallest overlap with the current selected keyframe category pool, where the selected keyframe category pool refers to the set of all object categories contained in the selected keyframes, and the overlap is determined by the number of common categories between the new frame and the selected keyframe category pool; Based on the condition of minimizing overlap, a new frame that satisfies the condition of maximizing the detection category count is selected. The detection category count is used to measure the total number of different object categories detected in the corresponding frame. When semantic conditions are the same, select the new frame that satisfies the earliest timestamp condition; The frames that satisfy the condition of minimizing overlap, the condition of maximizing class count, and the condition of selecting the earliest timestamp are selected as the multiple key frames.
4. The method according to claim 1, characterized in that, The motion reconstruction process performed on the video stream to generate motion trajectory information includes: Motion estimation processing is performed on the video stream using visual odometry, which refers to a camera motion trajectory estimation method based on computer vision. Perform feature extraction and feature matching operations on consecutive frames in the video stream to obtain the inter-frame feature correspondence; Based on the inter-frame feature correspondence, the essential matrix is calculated using a robust estimation method. The essential matrix is used to describe the camera motion geometric relationship between adjacent video frames. The essential matrix is decomposed to obtain the relative rotation parameters and relative translation parameters of the camera. The relative rotation parameters and the relative translation parameters are integrated into global motion trajectory information through a recursive accumulation method.
5. The method according to claim 1, characterized in that, The step of visualizing the motion trajectory information to generate a trajectory visualization map includes: The motion trajectory information is rendered into a visualization image, which includes a bird's-eye view and a three-dimensional trajectory map. The three-dimensional trajectory map refers to a stereoscopic image that displays the motion trajectory in three-dimensional space, and the bird's-eye view refers to a two-dimensional image that displays the motion trajectory from an overhead perspective. The trajectory points in the motion trajectory information are subjected to time coloring using a continuous color map to generate time-colored trajectory points. Based on the motion trajectory information, the bird's-eye view and the three-dimensional trajectory map are generated using image rendering technology.
6. The method according to claim 3, characterized in that, The step of spatiotemporally associating the keyframe set with the motion trajectory information to generate enhanced keyframes includes: A trajectory-aware marker is overlaid on each keyframe in the keyframe set. The trajectory-aware marker is a combination of markers that associate each keyframe with motion trajectory information through visual identifiers. The trajectory-aware marker includes a frame index marker and a color marker. The frame index marker is used to indicate the temporal position of each keyframe in the video stream. The color of the color marker is set to be consistent with the trajectory color at the corresponding time point in the motion trajectory information, so that each keyframe is explicitly associated with the motion trajectory; The trajectory-aware markers are integrated into each of the keyframes to generate the enhanced keyframes.
7. The method according to claim 6, characterized in that, The process of constructing multimodal cues based on the enhanced keyframes and the trajectory visualization map by integrating visual and text input, and submitting the multimodal cues to a multimodal large language model, includes: The enhanced keyframe and the trajectory visualization are combined into a visual input set, wherein the visual input set is a combination of data containing the enhanced keyframe image and the trajectory visualization image; Generate text prompts, which include a description of the color consistency and time sequence of the trajectory-aware markers; The visual input set and the text prompt content are integrated into a multimodal prompt, wherein the multimodal prompt refers to a mixed data format used for input into the multimodal large language model; The multimodal prompts are submitted to the multimodal large language model through a single forward propagation, where a single forward propagation refers to the computational process of processing multimodal input data in one go.
8. A prompting system for a multimodal large language model, characterized in that, include: A keyframe extraction module is used to acquire the input video stream and extract a set of keyframes from the video stream. The motion trajectory generation module is used to perform a motion reconstruction process on the video stream and generate motion trajectory information; The visualization generation module is used to visualize the motion trajectory information and generate a trajectory visualization graph. The keyframe enhancement module is used to perform spatiotemporal correlation encoding between the keyframe set and the motion trajectory information to generate enhanced keyframes. The training prompt construction module is used to construct multimodal prompts based on the enhanced keyframes and the trajectory visualization map by integrating visual and text inputs, and submit the multimodal prompts to the multimodal large language model.
9. A computer device, comprising: At least one processor; as well as The memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, Computer instructions are used to cause a computer to perform the method according to any one of claims 1-7.
Citation Information
Patent Citations
Efficient self-adaptive picture character interaction detection method and system
CN119495128A
Cited By
Multi-frame image space-time perception enhancement method, system and device
CN122048978A
A method, system and device for spatial-temporal perception enhancement of multi-frame images
CN122048978B
Camera-based iceberg instance segmentation identification detection method and system
CN122090177A