Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1337 results about "Key frame" patented technology

A keyframe in animation and filmmaking is a drawing that defines the starting and ending points of any smooth transition. The drawings are called "frames" because their position in time is measured in frames on a strip of film. A sequence of keyframes defines which movement the viewer will see, whereas the position of the keyframes on the film, video, or animation defines the timing of the movement. Because only two or three keyframes over the span of a second do not create the illusion of movement, the remaining frames are filled with inbetweens.

Video content semantic understanding and text description generation method based on deep learning

The invention discloses a video content semantic understanding and text description generation method based on deep learning, and relates to the technical field of multimedia information processing.The method comprises the steps that the semantic similarity of a text and a video frame is calculated through a CLIP model, related key frames are selected, and features are aggregated; respectively extracting audio, visual and semantic features; aligning different modal features by using self-attention, unifying dimensions of the LSTM, and then splicing and fusing; attention weights are calculated at a video level, a frame level and a channel level, and key information expression is enhanced; swin Transform encodes fusion features, and LSTM (Long Short Term Memory) decodes step by step to generate natural language description; and a text-video index database is constructed, and rapid retrieval is realized based on semantic similarity. According to the method, the mapping relation between the video features and the natural language is learned end to end through the deep learning model, dependence on a fixed template can be eliminated, and semantic description with various sentence patterns and coherent logic is generated.
Owner:CHINA UNIV OF MINING & TECH YINCHUAN COLLEGE

Autonomous lifelong SLAM method and system based on visual language model hidden space representation

The invention relates to an introspection lifelong SLAM method and system based on visual language model hidden space representation, and the method comprises the steps: extracting a semantic tag based on an RGB-D image through a semantic encoder, and generating a scene map and a semantic topological graph based on the RGB-D image and the semantic tag; generating a dynamic mask based on the scene map, obtaining a dynamic mask coverage rate, and screening key frames with high static confidence values based on the coverage rate; calculating camera pose estimation corresponding to the key frame in real time, sampling the key frame to realize layering of the key frame, and performing layering rendering by using a NeRF model to obtain a virtual view; the hidden space difference degree of the virtual view and the corresponding real image is calculated, whether error introspection needs to be carried out or not is judged based on the hidden space difference degree, and the system is used for achieving the method. Compared with the prior art, the method has the advantages that open semantic reasoning of VLM, high-precision reconstruction of NeRF and real-time positioning of SLAM are combined, and positioning and mapping accuracy is improved.
Owner:TONGJI UNIV

Task instruction generation method and device based on cross-modal fusion, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a task instruction generation method, device and equipment based on cross-modal fusion, and a medium, and the method comprises the steps: carrying out the decoding and noise reduction of an input video, generating a frame sequence, and recognizing a plurality of key frames based on the inter-frame similarity; extracting spatial features of the key frames to form a sequence, and generating video spatial-temporal features in combination with time features; performing semantic preprocessing on the input text to obtain text semantic features, and acquiring motion sensor signals to obtain motion features; fusing the video spatio-temporal features, the text semantic features and the action features to generate fused features; and generating a perception vector based on the fusion feature and outputting a task instruction. According to the method, multi-modal fusion is realized through key frame extraction and space-time fusion mechanisms in combination with text semantic features and action features, and the perception expression ability and the task instruction generation accuracy are improved by using time sequence information and multi-source perception input of the video.
Owner:PING AN TECH (SHENZHEN) CO LTD

Video stream-based attitude feature recognition method

The invention discloses a posture feature recognition method based on a video stream, and the method comprises the steps: carrying out the preprocessing of a continuous video stream, obtaining video frame training data, extracting a key frame and an adjacent frame in each frame of image, constructing a feature extraction module for a human body region, and obtaining a global frame, performing local extraction on the human body area by using adjacent frames on the left side and the right side to obtain local frames, and constructing semantic association information for the global frame through time sequence continuity between the local adjacent frames and the current key frame; acquiring enhanced feature representation by adopting a conditional feature aggregation algorithm; obtaining attitude sequence data through the attitude detail features; the method comprises the following steps: establishing three-dimensional coordinates, adaptively extracting posture change data by adopting a human body motion decoupling model, predicting human body posture characteristics through a smooth optimization strategy, and introducing a cross attention mechanism to realize deep fusion of spatio-temporal characteristics, so that the understanding ability of the model to a complex action mode is enhanced; and the attitude expression capability of the model in a sheltered or fuzzy region is obviously improved.
Owner:北京汇畅数宇科技发展有限公司

Long-time pedestrian re-identification method based on dual-path cooperation and key frame guided reconstruction

The invention discloses a long-time pedestrian re-identification method based on dual-path cooperation and key frame guided reconstruction. The method comprises the steps of firstly collecting a pedestrian video to be recognized, and extracting a video feature sequence; space and time position coding is introduced into the video feature sequence; capturing local fine-grained dynamic features through a local dynamic feature capturing path, and modeling long-range time sequence association through a cross-frame global feature modeling path; then, dual-path feature complementation is realized through bidirectional gating interaction; further screening out key frames, and realizing feature reconstruction through a full-frame attention propagation mechanism; and finally fusing the dual-path fusion features, the key frame guide reconstruction features and the refined features to generate pedestrian identity features. And processing pedestrian identity features to obtain standardized feature vectors, performing similarity comparison on the standardized feature vectors and pedestrian features in an image library, and returning a matching list. According to the method, video time sequence information is fully utilized, and the problem of insufficient robustness caused by appearance change in long-time pedestrian re-identification is effectively solved.
Owner:SHIJIAZHUANG TIEDAO UNIV

Panoramic video three-dimensional point cloud sparse reconstruction method and device integrating key frame screening

The invention provides a panoramic video three-dimensional point cloud sparse reconstruction method and device integrated with key frame screening, and the method comprises the steps: obtaining a panoramic video of a target scene, and carrying out the video quality evaluation of the panoramic video, and obtaining a video evaluation result; determining a target resolution of downsampling according to the video evaluation result; carrying out downsampling processing on the panoramic video based on the target resolution to obtain a downsampled second video; for the second video, dynamically adjusting a frame extraction interval based on a preset key frame selection mechanism, and extracting a frame index of the key frame from the second video according to the frame extraction interval to obtain a plurality of frame indexes; and obtaining a plurality of corresponding key frames in the panoramic video based on the plurality of frame indexes, and obtaining the three-dimensional scene sparse point cloud of the target scene based on the plurality of key frames, thereby solving the technical problem of low processing efficiency in the prior art.
Owner:TIANJIN FIRE SCI & TECH RES INST OF MEM

Animation video generation method and system based on video script

The invention discloses a cartoon video generation method and system based on a video script, and relates to the technical field of video synthesis, and the method comprises the steps: based on structured script data, combining a predefined lens rule base and a reinforcement learning model, determining the type, duration and lens operation effect of a lens, and generating a lens splitting sequence through a dynamic lens splitting automatic generation mechanism, based on the split mirror sequence, generating an animation style key frame image through a diffusion model, selecting action data matched with the emotion label through a predefined action library, generating a voice waveform matched with the emotion label through a text-to-voice model, selecting a background music audio matched with the emotion label through a music library, and generating a multi-modal content stream; according to the method, the lens type, the duration and the lens operation effect are dynamically optimized through the reinforcement learning model, the defects of a traditional lens scheduling method based on static rule mapping in time continuity and narrative continuity are overcome, and the narrative fluency and the dynamic adaptability of the lens division sequence are obviously improved.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

Non-inductive identity recognition and real-time tracking method based on personnel in video

The invention discloses a non-inductive identity recognition and real-time tracking method based on personnel in a video. The method comprises the following steps: S1, acquiring video stream data and extracting a key frame image; s2, key point coordinates are extracted, and a time sequence skeleton sequence is constructed; s3, decomposing gait, action and attitude features to generate standardized features; s4, inputting the standardized features into an improved multi-scale image convolutional neural network, extracting identity representation features, and generating a target identity code; s5, matching identity codes based on similarity measurement, calculating a consistency score, and establishing an identity tracking trajectory; s6, tracking is carried out in combination with identity codes and consistency scores, identity mapping is dynamically updated, and non-inductive recognition is achieved; and S7, updating a track in real time, adjusting a tracking state, and ensuring identity stability. According to the method, the improved graph convolutional network and the consistency measurement technology are combined, the identity recognition precision and tracking stability are improved, and cross-scene non-inductive identity recognition is achieved.
Owner:BEIJING LIYANG ZHIGUANG TECH CO LTD

Video generation method and device based on action coherence, equipment and medium

The invention relates to the technical field of data analysis, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a video generation method, device, equipment and medium based on action coherence, the method comprises the following steps: obtaining a video action material set, carrying out key frame analysis on the video action material set to obtain a key video frame sequence, extracting an inter-frame residual vector between consecutive frames in the key video frame sequence, adjusting a preset initial diffusion model by using the inter-frame residual vector to obtain an optimized diffusion model, performing cosine scaling on the inter-frame residual vector by using the optimized diffusion model to obtain a scaled residual vector, and obtaining a video adjustment text, and carrying out noise addition and splicing on the key video frame sequence by using the video adjustment text and the zoom residual vector to obtain a noise video, and carrying out noise reduction on the noise video to obtain a target action video. According to the method and the device, the action coherence in the customized generated video can be effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Generative video coding and decoding method based on multi-modal large model

The invention relates to the technical field of video coding and decoding, and discloses a multi-mode large model-based generative video coding and decoding, which comprises a key frame selection module for determining a key frame by analyzing the semantic and motion characteristics of a video frame; the multi-modal semantic description generation module is used for generating semantic description according to the key frame and the video clip; the key frame compression module is used for realizing efficient compression through latent variable modeling and entropy coding; the key frame reconstruction module reconstructs a key frame by using a conditional diffusion model in combination with the compressed data and the semantic description information; and the video generation module generates a non-key frame by using the semantic description and the key frame, and reconstructs a complete video. Through key frame screening combining semantic and motion information, key frame compression and reconstruction based on a conditional latent variable diffusion model, and frame supplementation and frame insertion generation based on semantic description, efficient compression and high-quality video reconstruction can be realized under a low code rate, and the video storage efficiency and the visual quality are effectively improved.
Owner:上海芯开技术有限公司

Key frame extraction method and device based on dynamic reinforcement learning, equipment and medium

The invention relates to the technical field of computer vision, can be applied to the medical field and the financial science and technology field, and discloses a key frame extraction method, device and equipment based on dynamic reinforcement learning and a medium, which are applied to electronic application in a high-frequency transaction abnormal behavior monitoring scene or can be applied to a medical operation key frame extraction scene. The method comprises the steps of obtaining an original video stream and performing preprocessing to generate a standardized video frame; performing feature extraction and feature splicing on the standardized video frame, and performing time sequence modeling on the generated pair frame-level mixed feature vector to generate video-level time sequence representation; generating an enhancement action instruction based on the video-level time sequence representation through the strategy network, and performing enhancement processing on the standardized video frame according to the enhancement action instruction to generate an enhanced video frame; performing optimization processing on the strategy network according to the enhanced video frame to generate an updated strategy network; and performing feature extraction and optimization on the enhanced video frame to generate a target key frame. According to the invention, the key frame extraction precision is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent conference video frame dynamic coding method based on multi-mode semantic understanding

The invention relates to the technical field of computer vision, in particular to an intelligent conference video frame dynamic coding method based on multi-modal semantic understanding, which comprises the following steps: acquiring a video stream sequence and a synchronous audio stream in a conference scene in real time; performing semantic analysis and decoupling on the video stream sequence, and extracting key frames and subsequent frames; extracting a sparse motion field from a subsequent frame, and segmenting a video frame into candidate visual areas including a face, a mouth shape and a background; extracting audio semantic features, executing cross-modal semantic correlation analysis, calculating semantic correlation between the sparse motion field distribution features and the audio semantic features, and positioning a pronunciation area highly related to the voice content; and calculating a quantization offset value of each candidate visual area according to the semantic relevancy, applying the quantization offset values in different areas, and packaging the quantization offset values into a variable-code-rate video code stream. According to the invention, the multi-mode semantic understanding model is constructed to carry out deep semantic analysis on the video frame content so as to realize the dynamic coding of the conference video frame.
Owner:SHENZHEN JIKEYUAN ELECTRONIC TECH CO LTD

Video text cross-modal retrieval method based on spatio-temporal feature fusion

The invention relates to the field of artificial intelligence cross-modal retrieval, and provides a video text cross-modal retrieval method and system based on spatio-temporal feature fusion. The method comprises the following steps: carrying out key frame sampling and time sequence partitioning on an input video, extracting static visual features through a spatial feature network, and extracting motion features through a time dynamic network; a self-adaptive gating fusion module is adopted to dynamically calculate spatial-temporal feature weights and perform weighted fusion; extracting text semantic features by using a pre-training language model; constructing a double-flow projection network to map video fusion features and text features to a unified measurement space, and optimizing a feature distance by adopting a contrast loss function containing difficult negative sample mining and intra-modal constraint; and outputting a retrieval result according to the cosine similarity sequence. The system comprises four units, wherein the gating fusion module is integrated with an FPGA acceleration circuit. According to the method, mAP (at) 10 is equal to 0.78 in a UCF-101 data set, the time sequence action retrieval accuracy rate is 92.8%, and the single video retrieval delay is 23 milliseconds.
Owner:ZHEJIANG UNIV

Prompt construction method and system of multi-mode large language model, computer equipment and medium

The invention relates to the technical field of multi-modal large language model training, in particular to a prompt construction method and system for a multi-modal large language model, computer equipment and a medium. The method comprises the following steps: extracting a key frame set from an input video stream; executing a motion reconstruction process on the video stream to generate motion track information; and performing visualization processing on the motion track information to generate a track visualization graph. Performing space-time correlation coding on the key frame set and the motion track information to generate an enhanced key frame; a multi-modal prompt is constructed in a mode of integrating visual input and text input, and the multi-modal prompt is input into a preset multi-modal large language model for spatial reasoning. Through the mode, the technical problem that an existing prompting method is difficult to give consideration to the spatial reasoning precision and the calculation efficiency is solved, efficient and accurate spatial reasoning of the multi-modal large language model is achieved, and the calculation efficiency, the reasoning precision and the environmental adaptability of the model are improved.
Owner:HONG KONG UNIV OF SCI & TECH (GUANGZHOU)

Automatic action trajectory labeling system and method based on visual language model

The invention provides an automatic action trajectory labeling system and method based on a visual language model, and relates to the field of visual technical models. The automatic action trajectory labeling system based on the visual language model comprises a key frame extraction module, a VLM semantic labeling module, a historical labeling database, a diffusion model optimization module and an reflection correction module, and the key frame extraction module extracts a key frame sequence from an input video through optical flow analysis and a TSN network scene change detection algorithm; and the VLM semantic annotation module is used for outputting a candidate segmentation point set containing action starting and ending points, the VLM semantic annotation module adopts a multi-modal VLM, takes a key frame image and a context text instruction as input, and generates a preliminary semantic tag and a corresponding vision-language embedding vector. According to the method, multiple bottlenecks of the existing action track labeling technology are solved, the labeling precision is remarkably improved, and the target of time boundary optimization is achieved.
Owner:高杨

Audio data selection for video matching using generative artificial intelligence model

A video editing system leverages a generative artificial intelligence (AI) model to identify songs to overlay on a video. The video editing system extracts a set of key frames from the video and prompts the generative AI model to generate a video narrative for the video. A video narrative is a text description of the plot, theme, feel, or other characteristics of the video. The video editing system uses the video narrative to prompt the generative AI model again to generate a set of descriptor tags for the video based on the video narrative. Descriptor tags are strings that represent themes, features, or characteristics of the song. The video editing system uses an audio tagging system to score a set of songs based on the set of descriptor tags and presents a selected subset of the set of songs based on the scores of the songs.
Owner:BEACON STREET TECHNOLOGIES LLC

Screen content coding method and device, equipment, storage medium and program product

The invention relates to the technical field of video coding, and provides a screen content coding method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring a to-be-coded screen video image, and extracting key frames of each picture group of the screen video image; edge detection is carried out on the key frame, and the complexity of the key frame is determined according to the detected number of edge pixel points of the key frame; and determining a coding rate according to the complexity, and carrying out coding processing on the screen video image based on the coding rate to obtain a coding stream. By calculating the complexity of the key frame, the calculation amount of the complexity is reduced, the number of edge pixel points is used as the complexity variable of the key frame, the encoding code stream is adaptively determined according to the complexity, the edge information of the video image is reserved, the compression quality of the encoded video is ensured, the encoding efficiency of the screen content is improved, and the user experience is improved. Efficient coding is realized, and the method is suitable for scenes with relatively high time delay requirements.
Owner:CHINA MOBILE ONLINE SERVICES CO LTD +1

Multi-modal fusion key frame extraction method and device, equipment and medium

The invention relates to the technical field of computers, and discloses a multi-modal fusion key frame extraction method and device, equipment and a medium, and the method comprises the steps: obtaining multi-modal input data, carrying out the modal feature coding, and obtaining a video modal feature, an audio modal feature and a text modal feature; carrying out attention fusion on the video modal features, the audio modal features and the text modal features to obtain fused cross-modal causal features; analyzing the cross-modal causal features through a causal reinforcement learning decision module in combination with a preset time sequence causal graph to obtain a fusion feature sequence and key frame probability distribution; and carrying out key frame selection operation on the time slice of the fusion feature sequence based on the key frame probability distribution to obtain a key frame set, and generating a space-time thermodynamic diagram and causal relationship visualization result corresponding to the key frame set. The multi-modal fusion key frame extraction method and device can be applied to financial science and technology or medical care service program systems, and the accuracy and interpretability of multi-modal fusion key frame extraction can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Underground pipeline intelligent detection and mapping method based on image recognition

The invention discloses an underground pipeline intelligent detection and mapping method based on image recognition, and the method comprises the following steps: 1, obtaining continuous video images, associating feature matching pairs of adjacent key frames, and forming a pose parameter set; 2, inputting the key frame into an improved YOLO-World detection network, and outputting a detection result set; 3, obtaining a geometric consistency matching set according to the detection result set; 4, performing multi-view triangularization on the geometric consistent matching set to form a weight factor; 5, introducing a weight factor, and executing incremental beam adjustment optimization on the pose parameter set and the three-dimensional sparse point set to obtain a sparse semantic point cloud; and step 6, outputting an underground pipe network topology map. According to the invention, high-precision and high-robustness intelligent identification and topological mapping in a complex underground pipeline environment are realized.
Owner:WUXI YIXING POWER TECH CO LTD

Long video understanding method capable of relieving time sequence illusion in video language large model

The invention provides a long video understanding method capable of relieving time sequence illusion in a video language large model. The long video understanding method is based on a static bias adaptive frame selection mechanism and a cross-modal feature fusion strategy. According to the static bias mechanism, inter-frame similarity is evaluated through a discriminator, redundant frames are identified, key frames are selected or a complete sequence is reserved, so that calculation overhead is reduced, and spatio-temporal information integrity is kept; a video frame and a text are mapped to a shared semantic space, the single-frame semantic understanding ability is enhanced, then an embedded sequence serves as a soft prompt to be input into a large language model, and a final answer is generated in an autoregression mode. According to the method, the efficiency and accuracy of long video understanding and video question and answer tasks can be remarkably improved; the problem of low training and reasoning efficiency caused by time sequence dependence redundancy and excessive computing resource consumption is effectively relieved; and through a dynamic multi-modal task processing framework and a space-time memory bank compression mechanism, the modeling capability and generalization performance of the model on a long video sequence are further improved.
Owner:LANZHOU UNIV

Dynamic scene robust visual SLAM method based on multi-feature collaborative optimization

The invention discloses a dynamic scene robust vision SLAM (Simultaneous Localization and Mapping) method based on multi-feature collaborative optimization, which comprises the following steps of: acquiring an image sequence, carrying out dynamic target detection and segmentation through an instance segmentation network, generating a segmentation mask containing a dynamic region mark, and identifying and separating a dynamic object and a static background; removing feature points corresponding to the dynamic object based on the segmentation mask to obtain static feature points; carrying out pose estimation based on the static feature points, and for the key frame, carrying out feature matching with the previous key frame by minimizing a re-projection error, and solving to obtain the camera pose of each key frame; for non-key frames, performing camera pose tracking and data association on the previous frame by adopting an optical flow algorithm, and accumulating solving results to obtain pose tracks of all the non-key frames; the key frames and the non-key frames are subjected to differential processing by fusing feature matching and an optical flow algorithm, so that the calculation efficiency is remarkably improved while the positioning precision is ensured, and the real-time performance is improved.
Owner:INNER MONGOLIA UNIVERSITY

A method, system, medium, device and terminal for monitoring video key frame extraction

The application belongs to the technical field of multimedia information processing, and discloses a kind of monitoring video key frame extraction method, system, medium, equipment and terminal, collect original video stream data, original video stream data is decomposed into image frame set;Image frame set obtained by decomposition is sampled, and the image frame result set obtained by sampling is filtered;The image frame set after filtering is adaptively clustered, and the result after clustering is collected to form a video abstract.In order to better utilize the memory space of storage medium and let the user better quickly browse the general content of original video stream, the application provides a key frame screening algorithm for original video stream by sampling, filtering and clustering.The key frame extraction method of the application screens out similar frames, redundant frames and fuzzy frames in original video stream through key frame screening algorithm for original video data, forms a video abstract storage of a section of original video, thereby greatly reduces the occupied storage space.
Owner:QINGDAO INST OF COMPUTING TECH XIDIAN UNIV

Video question-answering system and method based on iterative multi-mode

The invention provides a video question-answering system and method based on iterative multi-mode, and the method comprises the steps: carrying out the preprocessing of an input original video file and natural language query, extracting a key frame sequence, and generating an initial subtitle sequence; performing multi-granularity retrieval based on the preprocessed natural language query and the current subtitle sequence, and determining a candidate region; carrying out fine-grained frame selection in the candidate region by using a large language model, and identifying a key frame; based on the candidate area and natural language query, determining the type of the visual information to be supplemented and generating a multi-modal cue word corresponding to the type, and extracting the visual information by the visual language model according to the multi-modal cue word to update the subtitle sequence of the candidate area; generating a prediction answer by adopting a large language model and a visual language model; and judging the confidence of the generated predicted answer, and outputting a final answer. According to the method, the processing mode of video understanding can be optimized, and an accurate cross-modal coordination solution is provided through dynamic reasoning-sensing coordination.
Owner:EVALUATION & DEMONSTRATION RES CENT OF THE CHINESE PEOPLES LIBERATION ARMY ACAD OF MILITARY SCI

Video content textualization method and system based on multi-modal fusion

The invention provides a video content textualization method based on multi-modal fusion, which comprises the following steps of: 1, dynamically identifying effective modal information existing in a video, including subtitle information detection, audio information detection and key frame information sampling; step 2, carrying out subtitle extraction on the subtitle information by adopting an OCR (Optical Character Region) enhancement method based on regional clustering; generating a voice text for the audio information by adopting a multi-engine collaborative transcription and weight fusion strategy; generating a descriptive text for the key frame information; step 3, performing space-time alignment and semantic fusion on the subtitles, the voice text, the descriptive text and the video time axis; and feeding back and adjusting a strategy of key frame information sampling according to a fusion result, and forming self-adaptive feedback. According to the textualization method, multi-modal information can be adaptively fused, and video total elements are covered, so that the robustness and comprehensiveness of content extraction are improved.
Owner:WUHAN UNIV

Method, system and terminal for classifying echocardiography videos

The invention discloses an echocardiogram video classification method, system and terminal, and the method comprises the steps: constructing a classification model network which comprises a feature extraction module, a feature enhancement module and a feature aggregation module; obtaining an echocardiogram video, obtaining a plurality of standard section views according to the echocardiogram video, performing interpolation processing and feature extraction on the plurality of standard section views through a feature extraction module, and outputting a plurality of video features; inputting the plurality of video features into a feature enhancement module for aggregation enhancement of spatial features and time sequence features, and outputting a plurality of enhanced features; and inputting the plurality of enhanced features into a feature aggregation module for frame-level feature weighted fusion to obtain a plurality of key frame features, selecting related features from the plurality of key frame features, obtaining fusion features according to the related features, and classifying the fusion features to obtain a classification result of the echocardiogram video. According to the method, the classification accuracy of the echocardiogram videos is effectively improved.
Owner:SHENZHEN CHILDRENS HOSPITAL

Video generation method and device based on time sequence similarity, electronic equipment and medium

The invention relates to a video generation method based on time sequence similarity, and is applied to the field of video generation. Specifically, the video generation method based on the time sequence similarity comprises the following steps: acquiring generated video information; generating a key video frame sequence with frame rate information based on the text information, wherein the key video sequence is separated by a plurality of frame marks; performing recursive interpolation processing on the key video frame sequence, fusing front and back key frame features through a preset attention mechanism, and generating an initial intermediate frame between adjacent key frames; extracting visual features of a plurality of key frames and intermediate frames included in the key video frame sequence, calculating cosine similarity of the visual features of adjacent frames, and adjusting intermediate frame generation parameters based on the similarity to optimize inter-frame coherence so as to obtain an optimized intermediate frame; and combining the key video frame sequence with the optimized middle frame to generate a complete video clip which is consistent with the text information and has a coherent inter-frame time sequence.
Owner:ACADEMY OF BROADCASTING SCI STATE ADMINISTATION OF PRESS PUBLICATION RADIO FILM & TELEVISION

Cable three-dimensional modeling method and device based on video stream and key frame

The invention discloses a cable three-dimensional modeling method and device based on a video stream and a key frame, and belongs to the technical field of three-dimensional modeling of power system cable equipment, and the method comprises the following steps: constructing a dynamic truncation symbol field based on camera video stream data, and extracting a cable surface grid; triggering key frame capture, and establishing a key frame database; performing multi-modal fusion on the video stream features and the key frame features to generate a cable deformable template; obtaining a first-order deformable template model based on the cable deformable template; constructing a staged multi-scale adversarial discriminator, and deploying convolution kernel optimization at different resolution levels to obtain a second-order deformable template model; performing covariance optimal registration on the scene point cloud reconstructed by the video stream and the second-order template point cloud to obtain a third-order optimization template model; and constructing a density sensing gradient field, and dynamically fusing the video stream and the key frame data to generate a cable three-dimensional scene model. According to the invention, the precision and real-time performance of cable network modeling in a complex environment are improved.
Owner:STATE GRID SHANDONG ELECTRIC POWER CO

Dynamic scene SLAM optimization method based on improved YOLOv11 and geometric consistency constraint

In a dynamic environment, a visual SLAM (Simultaneous Localization and Mapping) system often causes the problems of large positioning error and inaccurate map construction due to dynamic target interference. In order to improve the robustness and precision of the system, the invention provides a dynamic scene SLAM optimization method based on improved YOLOv11 and geometric consistency constraint. Firstly, ORB features in a scene are extracted, and meanwhile a prior dynamic object and feature points on the prior dynamic object are removed through a YOLOv11 semantic segmentation model; secondly, eliminating feature points on the potential dynamic object by utilizing geometric consistency constraint, and recovering a background shielded by the dynamic object through a semantic perception Gaussian filter; and finally, selecting a high-quality key frame and applying the key frame to loopback detection and global optimization, constructing a basic Gaussian graph through a group of determined poses and point clouds, and finally fusing repair frame information to realize new view rendering and three-dimensional scene optimization.
Owner:KUNMING UNIV OF SCI & TECH

Camera state judgment method based on road surface covering and Hash comparison

The invention relates to the technical field of intelligent traffic, and discloses a camera state judgment method based on road surface covering and Hash comparison, and the method sequentially comprises the steps: extracting key frame pairs from a video stream at intervals; performing road surface region extraction on each frame of image to obtain a road surface mask; carrying out covering processing on the road surface area in each frame of image based on the mask to obtain a covered image; calculating a difference hash feature of each frame of covered image to obtain a hash code; calculating the Hamming distance between the Hash codes of the key frame pair; and judging whether the camera is in a rotating or stable state according to a comparison result of the Hamming distance and a threshold value. The method only depends on the video image data, does not need an external sensor, effectively eliminates the dynamic interference of the vehicle through covering the road surface, achieves the efficient and robust judgment of the state of the camera through combining the lightweight difference hash calculation, is low in calculation cost, and is suitable for edge equipment and complex road environments.
Owner:GUANGZHOU GUOJIAO RUNWAN TRAFFIC INFORMATION CO LTD

Video large model visual token dynamic adjustment method for long video understanding task

The invention discloses a video large model visual token dynamic adjustment method oriented to a long video understanding task, and belongs to the technical field of long video question answering of cross-modal video content understanding. The method comprises the steps that video information is extracted and sent to a video large model, key frames are identified through shallow cross attention scores, more computing resources are allocated to the key frames, meanwhile, non-key frame tokens are screened, and recombined visual tokens are obtained; in the middle-layer reasoning of the video large model, when the model is uncertain in the reasoning process, dynamically reintroducing the recombined visual token; in deep reasoning of the large video model, when cross-modal attention of the model does not pay attention to the recombined visual token, multi-modal fusion of the recombined visual token is stopped. According to the method, the original model does not need to be retrained, the reasoning efficiency of the video large model can be greatly improved by adopting a specific video token processing method in different stages of video large model reasoning, and the content understanding and reasoning capabilities of the video large model are improved at the same time.
Owner:NANJING UNIV OF POSTS & TELECOMM