Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1403 results about "Video sequence" patented technology

System and Method for Multi-Modal Hyperspectral Image Generation with Cross-Modal Attention and Adaptive Quality Assurance

A system and method are disclosed for generating hyperspectral images from multi-modal sensor data including RGB, LiDAR, thermal, and near-infrared inputs. Training data includes hyperspectral images and corresponding multi-modal measurements. Spectral band grouping is performed based on correlation coefficients. A multi-modal decomposition network with cross-modal attention mechanisms generate reconstructed hyperspectral images by fusing complementary sensor information. A fine-tuning network creates reconstructed RGB images. A comprehensive quality assurance system analyzes spectral consistency, cross-modal coherence, and fusion artifacts to generate quality metrics. Missing data compensation strategies handle corrupted sensor inputs using information from other modalities. The system includes temporal integration for video sequences and multi-resolution processing for different sensor resolutions. Quality metrics guide network weight adjustments to improve reconstruction accuracy while maintaining robustness to sensor failures and environmental variations.
Owner:ATOMBEAM TECH INC

Intelligent event identification method and system based on high-speed camera

The invention provides an intelligent event identification method and system based on a high-speed camera, and the method comprises the steps: setting the collection parameters of the high-speed camera, and triggering the camera to collect a target scene video stream. And hardware acceleration decoding processing is carried out on the collected original video data stream, real-time environment illumination information of the environment illumination sensor is obtained, and dynamic brightness equalization processing is executed. And performing motion adaptive denoising processing on the video sequence. Geometric distortion correction is carried out on the image sequence through camera calibration parameters, sub-pixel-level displacement vectors and dense optical flow field data of a moving object are extracted, and multi-scale morphological features are extracted. And the features are fused to generate motion feature data, the data are processed through a spatio-temporal joint event classification model, an event identification result is output, the result is bound with a high-precision timestamp, and event identification information is output to an industrial control system display device in real time. According to the invention, the accuracy and real-time performance of event identification can be improved.
Owner:广州思林杰科技股份有限公司

System and method for AI-powered narrative analysis of video content

A system, a method and a processor are for AI-powered generation and delivery of video clips. The processor is configured to: load a first video file of a first video content item, the first video file comprising video frames associated with timestamps; load a first subtitle file of the first video content item, the first subtitle file comprising subtitle text associated with the timestamps; execute a natural language processing (NLP) model with the subtitle text as input, the NLP model including language pre-processing steps for classifying words, names or phrases in the subtitle text and associating initial classifiers with the subtitle text, the NLP model including one or more of a recurrent neural network (RNN), a Bidirectional Encoder Representations from Transformers (BERT) model, or a generative pre-trained transformer (GPT) model for a dialogue analysis comprising processing sequences of dialogue in the subtitle text in view of the initial classifiers to associate one or more portions of the dialogue with one or more first classifiers of first narrative elements; execute an image recognition model with at least some of the video frames as input, the image recognition model including a convolutional neural network (CNN) for an object detection analysis and a facial recognition analysis comprising processing video sequences to associate one or more of the video frames with one or more second classifiers of second narrative elements; generate a narrative map of the first video content item by temporally aligning the first narrative elements with the second narrative elements based on the timestamps associated with the video frames and the first subtitle file; and generate a video clip including at least one segment of the first video content item, the at least one segment including selected video frames associated with at least one of the first or second narrative elements identified from the narrative map and selected for inclusion in the video clip.
Owner:PARAMOUNT GLOBAL INC

Guideboard surface deformation detection method and system based on dynamic video sequence

The invention discloses a guideboard surface deformation detection method and system based on a dynamic video sequence. The method comprises the following steps: acquiring road video data with a timestamp and position information; a traffic sign is automatically identified from the road video data, and geometric anomaly detection is carried out; a time sequence tracking sequence is established for the guideboard, and whether continuous abnormity exists in the guideboard is determined through multi-frame analysis; if yes, performing local texture consistency and local contour geometric analysis on the guideboard to obtain a fusion analysis result; performing state evaluation on supporting equipment corresponding to the guideboard to obtain a state evaluation result of the guideboard supporting equipment; and determining the final surface deformation state and the corresponding grade of the guideboard based on the fusion analysis result and the state evaluation result of the guideboard supporting equipment. By implementing the method provided by the invention, automatic detection and quantitative analysis of the bending or deformation condition of the road sign can be realized without manual intervention, and early warning information can be output in time.
Owner:WINTOO INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Activity-based person identification using biometric disentanglement

A system and method for person identification from video data by disentangling biometric identity features from non-biometric appearance and activity features are disclosed. The system processes RGB video sequences depicting individuals performing various activities to extract spatio-temporal features. These features are separated into distinct biometric identity representations and non-biometric features related to appearance and performed activities. To achieve this separation and minimize appearance bias, the system utilizes an auxiliary supervisory model. At least two implementations of this supervisory model are disclosed: one using semantic supervision via structured embeddings processed through a vision-language model, and another employing silhouette-based feature distillation from a silhouette-trained neural network. Joint training for biometric identification and activity classification ensures accurate identification of individuals independently of facial visibility, clothing differences, or activity variations.
Owner:UNIVERSITY OF CENTRAL FLORIDA RESEARCH FOUNDATION INC

Ultrasonic image discrimination perception pre-training method based on cooperative training framework

The invention provides an ultrasonic image discrimination perception pre-training method based on a cooperative training framework, and relates to the technical field of ultrasonic image analysis, and the method comprises the steps: obtaining an ultrasonic image sequence containing time sequence information; constructing a cooperative training architecture, wherein the two networks are both based on a visual converter and embedded with a time sequence anatomical attention module; inputting the original image into a teacher network, inputting the enhanced image into a student network, and respectively outputting global and local features; calculating an anatomical continuity measure based on the output features of the two-network time sequence anatomical attention module; constructing a total loss function; the regularization weight is dynamically adjusted according to the anatomical continuity measurement, and student network parameters are updated through gradient descent; dynamically calculating an index moving average coefficient according to the anatomical continuity measurement so as to update teacher network parameters; and iterative training is carried out until convergence. According to the method, the time sequence continuity and anatomical structure information contained in the ultrasonic video sequence are fully mined, so that the adaptive optimization of the model training process is realized.
Owner:XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV +1

Multi-mode collaborative video sequence segmentation method

The invention discloses a multi-modal collaborative video sequence segmentation method. The method comprises the following steps: obtaining a multi-scale local feature matrix and a multi-scale global feature matrix of an image sequence; obtaining a multi-scale text feature matrix of the text sequence; obtaining a multi-scale local-global fusion feature matrix of the multi-scale local feature matrix and the multi-scale global feature matrix; obtaining a multi-modal fusion feature matrix of the multi-scale local-global fusion feature matrix and the multi-scale text feature matrix; and utilizing a decoder of the pre-trained large model to predict and generate a segmentation mask, and outputting a semantic segmentation map. The video sequence segmentation method is stable in performance when facing complex and changeable scenes, does not need to depend on a large amount of labeled data, reduces the training cost, and is suitable for various practical application fields including intelligent monitoring, automatic driving, medical image analysis and the like.
Owner:HARBIN INST OF TECH AT WEIHAI +1

Weak supervision video anomaly detection method based on prompt learning knowledge enhancement

The invention discloses a weak supervision video anomaly detection method based on prompt learning knowledge enhancement, and belongs to the technical field of video intelligent analysis. A video side gives a section of abnormal scene video, video sequence features and audio sequence features are obtained through a feature extraction network, then a trained and complete feature aggregation network is input to carry out multi-modal feature aggregation, an abnormal score is obtained through a score prediction network, and text representation is carried out based on prompt learning. A prompt template is constructed for abnormal video tags through a knowledge graph, semantic expansion is performed on normal tags through a plurality of learnable parameters, cross-modal alignment is performed on the normal tags and a video side, so that features of the video side are close to different normal semantics, knowledge enhancement is performed by introducing external information, positive abnormal boundaries of the video are learned, and the detection performance is improved. And finally, multi-task joint optimization is carried out through different loss functions, and abnormal video clip positioning is carried out.
Owner:COMMUNICATION UNIVERSITY OF CHINA

Commodity video intelligent generation method based on multi-modal analysis and dynamic narrative architecture

The invention relates to the technical field of video generation, in particular to an intelligent commodity video generation method based on multi-modal analysis and a dynamic narrative architecture, which comprises the following steps: inputting original commodity data and a user behavior log, and outputting a dynamic selling point weight vector through a selling point value evaluation engine; s2, inputting the dynamic selling point weight vector output in S1 into a narrative flow generator, and executing the following steps: intercepting first K core selling points according to the weight vector to form a narrative trunk chain; based on the historical preference data of the user, inserting an emotion enhancement node to generate a branch enhancement narrative flow; generating a multi-version narrative flow instruction set in combination with the real-time network bandwidth data; retrieving a preset video clip library according to the instruction set to generate an initial video sequence; detecting semantic fault regions between adjacent segments to generate compensation animation parameters; and injecting compensation animation parameters and outputting a continuous narrative video stream. According to the invention, the jamming feeling and the frame skipping phenomenon caused by the change of a narrative structure or different sources of video clips are reduced, and the overall watching experience of a user is improved.
Owner:SHENZHEN YINGMENG INTELLIGENT TECHNOLOGY CO LTD

Video object segmentation method based on query adaptive attention and discriminative memory

The invention belongs to the technical field of computer vision and digital video processing, and discloses a video object segmentation method based on query adaptive attention and discriminative memory, which comprises the following steps: step 1, constructing a video data set and preprocessing data; step 2, constructing a video object segmentation model combining multi-scale semantic feature integration and query adaptive discriminant enhancement, and training the video object segmentation model; step 3, video object segmentation reasoning; performing target segmentation reasoning on the preprocessed video sequence, outputting a frame-by-frame target segmentation mask, and generating a time sequence tracking result; 4, exporting, deploying and applying the model; through lightweight network design and a discriminative memory optimization strategy, the model is deployed to edge equipment, and real-time segmentation and visualization are realized. According to the method, on one hand, comprehensive target representation is provided on multi-scale feature extraction, and on the basis of the characteristic that only high-confidence target features are stored, error propagation is avoided, and the long-term segmentation stability of the method is ensured.
Owner:NANTONG INST OF TECH

Water body color recognition regression method and system based on space-time causality and manifold learning

The invention belongs to the field of environment monitoring and computer vision, and particularly relates to a water body color recognition regression method and system based on space-time causality and manifold learning, and the method mainly comprises the steps: carrying out the detection of a current target water body video sequence, extracting a water body region, carrying out the high-dimensional feature dimension reduction of the water body region, and obtaining a water body color recognition result; and performing feature extraction on the water body region through a space-time causal feature learning model, fusing the flow shape learning features and the space-time causal features to obtain fused features, and outputting a finally predicted water body color value. According to the method, end-to-end assembly line design of preprocessing-segmentation-feature modeling-regression is adopted, manual intervention is not needed from video input to color prediction, and through cascade cooperation of five core modules (video preprocessing, water body segmentation, manifold learning, time sequence causal modeling and color recognition), the real-time performance of the system is improved. Full-link automation from environmental interference suppression, feature extraction to result output is realized, information loss of intermediate links is avoided, and recognition efficiency and robustness are improved.
Owner:CHINA TOWER CO LTD

Diffusion based end-to-end in-scene media generation

Embodiments of the present disclosure provide techniques for performing virtual object placement in a video sequence using generative artificial intelligence models. An example method generally includes receiving an input prompt specifying an object to insert into a scene depicted in an input image stream; decoding, using a generative artificial intelligence model, perspective and lighting information for the input image stream; determining, based on the decoded perspective and lighting information, a location in the scene in which the object is to be inserted; and generating, using the generative artificial intelligence model, an output image stream including the object into the scene at the determined location, wherein visual effects for the object are based on the perspective and lighting information for the input image stream.
Owner:REMBRAND INC

Al-based video content analysis method and system

The invention relates to the technical field of content recognition, in particular to an Al-based video content analysis method and system, and the method comprises the following steps: carrying out the image segmentation learning based on a video sequence through employing a U-Net convolutional neural network, analyzing scenes and elements in a video frame, recognizing and isolating key visual elements in a video through network learning, and carrying out the recognition of the key visual elements in the video. Comprising objects and figures. According to the method, image segmentation learning is carried out by adopting the U-Net convolutional neural network, key visual elements in the video can be identified and isolated more accurately, a clearer basis is provided for follow-up scene change and key event tracking, video content analysis is carried out by applying the graph neural network, the identification capability of dynamic scenes and events is enhanced, and the identification efficiency is improved. Deeper structured understanding is provided for video content indexing, the video quality is analyzed through a structural similarity index evaluation method, and the reason of visual quality reduction can be accurately recognized and improved.
Owner:CHENDA (GUANGZHOU) NETWORK TECH CO LTD +1

Target tracking method and system for deep twin network

The invention discloses a target tracking method and system for a deep twin network, and relates to the technical field of computer vision and artificial intelligence, and the target tracking method comprises the following steps: obtaining a to-be-tracked video sequence and a target template image in an initial frame, a target template image and a search area image of a current frame are input into a lightweight backbone network, multi-scale feature maps are extracted respectively, the lightweight backbone network is improved, a dynamic channel pruning module is embedded, cross-layer fusion is performed on the multi-scale feature maps, and a module is updated through a self-adaptive template. In combination with a dynamic channel pruning module, redundant channels can be closed in a self-adaptive manner, the model parameter quantity and the calculation quantity are reduced, a pruning threshold is dynamically optimized through meta-learning, and it is ensured that key feature information is reserved while complexity is reduced.
Owner:TIANJIN MODERN VOCATIONAL TECH COLLEGE

Audio and video depth forgery detection method based on quality perception and multi-scale alignment

The invention discloses an audio and video depth forgery detection method based on quality perception and multi-scale alignment, and the method comprises the following steps: coding a synchronous audio and video sequence, and obtaining a frame-level visual feature, a facial action unit and a phoneme-level voice representation; a visual quality evaluation module is introduced to generate a spatial reliability mask, and quality weighting is carried out on the visual features; designing a global-local multi-scale cross-modal alignment mechanism, performing bidirectional cross-attention modeling on voice and face dynamic synchronization globally, and performing physiological coupling alignment on phonemes and face action units locally; and an uncertainty perception reasoning and calibration scheme is provided, adaptive temperature scaling is carried out according to quality and consistency, and uncertainty calibration is carried out by self-supervision loss. According to the method, the problems of insufficient robustness and excessive self-confidence misjudgment of an existing method in a low-quality video and high-synchronization counterfeit scene are solved, and the cross-dataset generalization capability and the actual deployment reliability are remarkably improved.
Owner:NANJING UNIV OF SCI & TECH

Video snapshot compression imaging reconstruction method based on space-time deformable attention

The invention provides a video snapshot compression imaging reconstruction method based on spatio-temporal deformable attention, which improves the reconstruction quality and efficiency, and comprises the following steps: inputting a single frame compression measurement value and a measurement matrix into an initial reconstruction module to obtain an initial reconstruction video frame; inputting the initial reconstructed video frame into a feature extraction encoder, mapping the initial reconstructed video frame to a high-dimensional feature space through multi-layer 3D convolution, and outputting a feature map; the feature map is input into a plurality of stacked DenseRNet Blocks, and the number of the DenseRNet Blocks is one; the DenseRNet Block internally comprises a plurality of DeT Blocks, after the DenseRNet Block dynamically divides an input feature channel, grouping progressive processing and feature fusion are carried out through the plurality of DeT Blocks, and the DeT Blocks comprise a deformable space convolution branch used for modeling local deformation perception, a time self-attention branch used for modeling global time sequence dependence and a feature interaction module used for cross-channel information interaction; and the features processed by the DenseRNet Block are input into a video reconstruction decoder, and a reconstructed video sequence is output through up-sampling of transposition convolution and refining of multilayer 3D convolution.
Owner:DALIAN UNIV

Video content counterfeiting detection method and device, equipment and medium

The invention relates to the technical field of image detection, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a video content counterfeiting detection method, device, equipment and medium, and the method comprises the steps: obtaining a to-be-detected video stream, carrying out the frame sampling and standardization processing of the to-be-detected video stream, and generating a to-be-detected video sequence; performing visual double-branch feature extraction on the to-be-detected video sequence to obtain a universal visual feature, a local counterfeit feature and an audio feature; fusing the universal visual features, the local counterfeit features and the audio features to obtain audio and video consistency features; mapping the audio and video consistency feature to a low-dimensional decoupling space and carrying out feature decoupling to obtain a target counterfeit feature; and performing expansion convolution on the target counterfeiting feature to obtain a target counterfeiting probability sequence, and determining authenticity of the to-be-detected video stream according to the target counterfeiting probability sequence. According to the invention, the content counterfeiting detection efficiency and detection accuracy can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Video heart rate detection method based on deep learning

The invention relates to a video heart rate detection method based on deep learning, and belongs to the technical field of image processing. The method comprises the following steps: preprocessing a to-be-detected video to generate a preprocessed face video sequence; inputting the preprocessed face video sequence into a double-flow collaborative spatial-temporal feature enhancement module group to generate spatial-temporal features; inputting the spatial-temporal feature representation into a multi-scale spatial-temporal convolution module, and extracting spatial-temporal features of different scales in parallel; training the detection model by adopting a time-frequency domain joint constraint composite loss function; and inputting the spatiotemporal features of different scales into a trained detection model to obtain a video heart rate detection result output by the detection model. The invention aims to solve the technical problem that the anti-interference capability and the detection precision are difficult to guarantee in the prior art.
Owner:KUNMING UNIV OF SCI & TECH

Electric power operator behavior detection method and related equipment

The invention discloses a power worker behavior detection method and related equipment. The method comprises the following steps: acquiring first data in a scene where a worker is located through a camera in a working site; the first data comprises continuous frame images or a video sequence lasting for preset time; analyzing the first data to obtain a current first identification result of the operator; the first identification result is used for representing whether the operator lacks of wearing the protective equipment and / or carries illegal articles; determining the current posture of the operator from the first data; and fusing the first recognition result and the current posture to obtain the operation risk level of the operator. The method can be widely applied to the technical field of artificial intelligence.
Owner:FOSHAN UNIVERSITY

Reference video object segmentation method and system based on motion modeling and multi-modal interaction

The invention discloses a reference video object segmentation method and system based on motion modeling and multi-modal interaction, and the method comprises the steps: taking a video sequence and natural language description as input, and generating a preliminary segmentation mask through a text coding and mask decoder; a Kalman filtering motion modeling module is introduced to predict the motion trail of the target object, and time sequence consistency optimization is carried out on the preliminary segmentation mask; fusing the historical track of the object and the action semantics in the semantic features on the basis of a key action semantic coding module to realize action semantic alignment and mask dynamic correction; the segmentation quality of the current frame is subjected to multi-dimensional scoring based on a representative frame screening mechanism, the representative frame is screened out to update a memory bank, and the long-term tracking stability is improved. According to the method, the problems of target drift, insufficient semantic alignment and memory pollution in a complex dynamic scene in the prior art are effectively solved, and the segmentation precision, robustness and semantic consistency are remarkably improved while the light weight of the model is kept.
Owner:ZHONGKE (SHENZHEN) WIRELESS SEMICON CO LTD

Long video understanding method capable of relieving time sequence illusion in video language large model

The invention provides a long video understanding method capable of relieving time sequence illusion in a video language large model. The long video understanding method is based on a static bias adaptive frame selection mechanism and a cross-modal feature fusion strategy. According to the static bias mechanism, inter-frame similarity is evaluated through a discriminator, redundant frames are identified, key frames are selected or a complete sequence is reserved, so that calculation overhead is reduced, and spatio-temporal information integrity is kept; a video frame and a text are mapped to a shared semantic space, the single-frame semantic understanding ability is enhanced, then an embedded sequence serves as a soft prompt to be input into a large language model, and a final answer is generated in an autoregression mode. According to the method, the efficiency and accuracy of long video understanding and video question and answer tasks can be remarkably improved; the problem of low training and reasoning efficiency caused by time sequence dependence redundancy and excessive computing resource consumption is effectively relieved; and through a dynamic multi-modal task processing framework and a space-time memory bank compression mechanism, the modeling capability and generalization performance of the model on a long video sequence are further improved.
Owner:LANZHOU UNIV

High-fidelity dynamic scene video generation method based on single static image

The invention discloses a high-fidelity dynamic scene video generation method based on a single static image, and relates to the technical field of computer vision and video generation, and the method comprises the steps: obtaining key elements and potential dynamic information in an image through deep understanding and semantic deconstruction of a static scene; through dynamic representation and modeling, spatial-temporal feature decoupling and complex dynamic scene modeling are realized; according to the method, the dynamic complexity can be processed and the visual high fidelity can be ensured at the same time, the problems of dynamic incoherence, detail loss and the like when a video is generated from a single static image are solved, the method can be widely applied to the fields of film and television production, virtual reality and the like, and the method is suitable for popularization and application. According to the method, the rich details and the sense of reality of the original image can be reserved to the greatest extent while the complex dynamic state is generated, the space-time consistency is kept in the whole video sequence, the image quality is prevented from being sacrificed due to dynamic state generation, and the method has high use value.
Owner:SHENZHEN YINGSHI TECHNOLOGY CO LTD

Video snapshot compression imaging reconstruction method and system

The invention relates to a video snapshot compression imaging reconstruction method and system. The method comprises the following steps: inputting a video frame sequence and a time-varying mask set thereof into a measurement model to obtain initial estimation; constructing a reconstruction network which comprises a feature extraction module, a gating residual network module and a video reconstruction module; the feature extraction module comprises two three-dimensional convolution layers, each three-dimensional convolution layer is connected with an activation function, and the feature extraction module extracts initial features from the initial estimation; inputting the initial features into a gating residual network module, and outputting reconstruction information features; and the video reconstruction module fuses the reconstruction information features, and performs up-sampling and detail refining to reconstruct a video sequence. According to the method, on the premise that parameters and computing power are hardly increased, ghosting and flickering are effectively restrained, the stability of long-time reconstruction is improved, and an effective scheme is provided for SCI reconstruction with the high compression ratio, the super-definition resolution ratio and the long sequence.
Owner:GUANGDONG UNIV OF TECH

Remote sensing image target detection method and system based on multi-modal feature fusion

The invention discloses a remote sensing image target detection method and system based on multi-modal feature fusion, and relates to the field of artificial intelligence, and the method comprises the steps: obtaining a first modal remote sensing video sequence and a corresponding second modal remote sensing video sequence; performing feature analysis on the first modal remote sensing image to obtain first modal features of one or more channels, and performing feature analysis on the second modal remote sensing video sequence to obtain second modal features; calculating a cross-modal correlation value of the first modal feature and the second modal feature under each pixel coordinate to indicate correlation intensity under the same pixel coordinate; integrating the first modal feature and the second modal feature based on the cross-modal correlation value to obtain a multi-modal integrated feature; and positioning and identifying a remote sensing target in the first modal remote sensing image according to the multi-modal integration feature and the first modal feature. According to the method, dynamic feature fusion is realized through pixel-level cross-modal association intensity quantification, the limitation of a single mode is overcome, and the precision of remote sensing target detection is improved.
Owner:GUOHUA (FENGNING MANCHU AUTONOMOUS COUNTY) NEW ENERGY CO LTD

Monocular camera and concentric annulus-based structure three-dimensional displacement monitoring method and device

The invention discloses a structure three-dimensional displacement monitoring method and device based on a monocular camera and a concentric annulus, and relates to the field of computer vision and structure health monitoring, a concentric annulus target is fixed at a to-be-monitored structure monitoring point, and the monocular camera carries out positioning and lens parameter adjustment; shooting a checkerboard image and calibrating the checkerboard image by adopting a camera calibration algorithm to obtain a distortion parameter and an internal reference matrix; collecting a video sequence containing a structure displacement process of the target; image distortion correction is carried out based on the distortion parameters and the internal reference matrix, a target is identified through a target detection algorithm, and target parameters are calculated through least square circle fitting; realizing multi-target tracking by a target point topology identification algorithm based on Y-X threshold sorting; calculating scale factors according to target parameters and actual physical sizes, and calculating in-plane orthogonal direction displacement and out-of-plane displacement to form three-dimensional displacement information of the structure; according to the invention, the monocular camera is adopted to realize the accurate measurement of the three-dimensional displacement of the structure and reduce the complexity and calibration difficulty of the system.
Owner:INST OF ENG MECHANICS CHINA EARTHQUAKE ADMINISTRATION

Video generation method and device based on time sequence similarity, electronic equipment and medium

The invention relates to a video generation method based on time sequence similarity, and is applied to the field of video generation. Specifically, the video generation method based on the time sequence similarity comprises the following steps: acquiring generated video information; generating a key video frame sequence with frame rate information based on the text information, wherein the key video sequence is separated by a plurality of frame marks; performing recursive interpolation processing on the key video frame sequence, fusing front and back key frame features through a preset attention mechanism, and generating an initial intermediate frame between adjacent key frames; extracting visual features of a plurality of key frames and intermediate frames included in the key video frame sequence, calculating cosine similarity of the visual features of adjacent frames, and adjusting intermediate frame generation parameters based on the similarity to optimize inter-frame coherence so as to obtain an optimized intermediate frame; and combining the key video frame sequence with the optimized middle frame to generate a complete video clip which is consistent with the text information and has a coherent inter-frame time sequence.
Owner:ACADEMY OF BROADCASTING SCI STATE ADMINISTATION OF PRESS PUBLICATION RADIO FILM & TELEVISION

Dosage prediction method and system based on multi-mode flotation froth digital characterization

The invention provides a dosage prediction method and system based on multi-mode flotation froth digital characterization, and relates to the technical field of froth flotation. The method comprises the following steps: collecting multi-modal data of a foam state in a metal ore flotation process, wherein the multi-modal data comprises an RGB image, a video sequence and a three-dimensional point cloud; constructing three foam feature extraction networks, and respectively extracting RGB features, video dynamic features and three-dimensional point cloud features of the multi-modal data; fusing the extracted RGB features, the video dynamic features and the three-dimensional point cloud features based on a cross-modal attention mechanism to generate multi-modal features; constructing an expert knowledge base, wherein the expert knowledge base comprises structured rules and historical cases; and constructing a dosage decision model based on an expert knowledge base and a large language model, converting the multi-modal features into standardized vector representation, retrieving the expert knowledge base by utilizing similarity matching, and generating a dosage adjustment strategy. According to the invention, accurate prediction from foam multi-mode characteristics to dosage adjustment can be realized.
Owner:UNIV OF SCI & TECH BEIJING

Multi-target tracking method and device, electronic equipment and computer readable storage medium

The invention provides a multi-target tracking method and device, electronic equipment and a computer readable storage medium, and the method comprises the steps: carrying out the target detection of a video sequence, and obtaining a detection frame of a current frame; under the condition that the detection frame of the current frame is successfully matched with the prediction frame, calculating a track consistency score; for the target trajectory of which the trajectory consistency score exceeds a set score, extracting feature representation of a historical frame by using a time sequence attention mechanism; and calculating feature similarity according to the feature representation and a set model, and determining the identity of the target trajectory. According to the method, the video frame is detected through the target detection technology, the corresponding high-confidence detection box is screened out, the purposes of accurately positioning multiple targets and reducing false detection can be achieved, and reliable input is provided for follow-up track association. The target position is predicted through the trajectory prediction technology, matching of a detection frame and a prediction frame is achieved, the purpose of maintaining trajectory continuity in a shielding or rapid motion scene is achieved, and the target identifier switching rate is remarkably reduced.
Owner:SHANGHAI JIDOU TECH CO LTD

Video processing method and device, storage medium and electronic equipment

The invention discloses a video processing method and device, a storage medium and electronic equipment, and relates to the technical field of image processing, in a processing mode corresponding to a processing demand, a diffusion model generation graph is used as a reference graph, and super-resolution processing is performed on a single-scene video clip and a reference image through a video super-resolution model, so that the video processing efficiency is improved. Due to the fact that super-resolution processing is an image processing mode that detail information in the reference image is extracted and supplemented to super-resolution, the detail richness of the image is improved, time consumption and instability caused by frame-by-frame generation of a diffusion model are effectively avoided, video production is carried out according to the reconstructed resolution video sequence, and the video quality is improved. The purpose of improving the resolution of video production is achieved.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Two-stage road traffic abnormal event identification method and system based on visual large model

The invention relates to a two-stage road traffic abnormal event identification method and system based on a visual large model, and the method comprises the steps: collecting the monitoring image and video data of an expressway and an urban expressway, and building a static image semantic data set and a dynamic video traffic semantic data set; utilizing the static image semantic data set to train a visual large model to obtain a first visual large model; constructing a same-preference data pair, and performing direct preference optimization training of the first visual large model by using the same-preference data pair to obtain a second visual large model; intercepting an abnormal video key frame based on the dynamic video traffic semantic data set, and performing parameter fine tuning on the second visual large model based on the abnormal video key frame to obtain a road traffic abnormal event recognition model; and collecting a monitoring image or video sequence in real time, and performing abnormal event identification by using the road traffic abnormal event identification model. Compared with the prior art, the traffic abnormal event identification method provided by the invention can effectively combine dynamic and static characteristics of data and is efficient.
Owner:TONGJI UNIV