Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1893 results about "Video sequence" patented technology

System and Method for Multi-Modal Hyperspectral Image Generation with Cross-Modal Attention and Adaptive Quality Assurance

A system and method are disclosed for generating hyperspectral images from multi-modal sensor data including RGB, LiDAR, thermal, and near-infrared inputs. Training data includes hyperspectral images and corresponding multi-modal measurements. Spectral band grouping is performed based on correlation coefficients. A multi-modal decomposition network with cross-modal attention mechanisms generate reconstructed hyperspectral images by fusing complementary sensor information. A fine-tuning network creates reconstructed RGB images. A comprehensive quality assurance system analyzes spectral consistency, cross-modal coherence, and fusion artifacts to generate quality metrics. Missing data compensation strategies handle corrupted sensor inputs using information from other modalities. The system includes temporal integration for video sequences and multi-resolution processing for different sensor resolutions. Quality metrics guide network weight adjustments to improve reconstruction accuracy while maintaining robustness to sensor failures and environmental variations.
Owner:ATOMBEAM TECH INC

Historical block scene three-dimensional reconstruction method and system based on Gaussian sputtering

The invention discloses a historical block scene three-dimensional reconstruction method and system based on Gaussian sputtering, and the method comprises the steps: collecting a historical block video sequence through a lightweight panorama camera, extracting a multi-frame panorama, and generating an image data set through a projection converter; monocular depth and normal estimation is carried out through a pre-training visual model, and a prior depth and normal graph data set of a historical block scene is constructed; sparse reconstruction is carried out on the multi-view image data set based on the SfM technology, initial point cloud and camera pose information are acquired, and a Gaussian ellipsoid is optimized in combination with prior depth and normal information; dynamically calculating the geometric width estimation value of the street, and guiding and adjusting the adaptive density of the Gaussian ellipsoids of the vertical surfaces on the two sides; rendering the optimized Gaussian ellipsoid through an improved rasterization renderer; and designing a multi-modal loss function and a regularization mechanism to optimize reconstruction and rendering results. According to the method, high-fidelity three-dimensional reconstruction of the historical block scene is realized, and the adaptability of the Gaussian sputtering method to the complex block scene is enhanced.
Owner:BEIJING UNIV OF CIVIL ENG & ARCHITECTURE

Regional abnormal condition real-time early warning method based on high-point panoramic intelligent inspection

The invention relates to the technical field of intelligent inspection, and discloses a regional abnormal condition real-time early warning method based on high-point panoramic intelligent inspection, which comprises the following steps: collecting visible light and infrared thermal imaging video streams of a target region to form a panoramic video sequence, and carrying out intelligent analysis, feature extraction and analysis, construction of a spatio-temporal topological graph and detection of an abnormal behavior mode. Predicting environmental risks, constructing a spatio-temporal evolution model, and further generating graded early warning information of regional abnormal conditions; the method effectively solves the problems of large panoramic inspection data volume, exception complexity, easy environmental influence on target detection and the like in a large-scale scene, significantly improves the accuracy and real-time performance of exception early warning, and guarantees the regional safety.
Owner:CHN ENERGY SUQIAN POWER GENERATION CO LTD

Methods For Generating Advertisement Videos Consistent With The Context And Storyline Of A Primary Video Stream

Embodiments include methods for generating advertisement videos for insertion into a video stream to promote a product, service, or brand in a manner that is consistent with the context and storyline of the video stream before and at the time of ad insertion. Methods may include capturing an image from the video stream and generating caption text using an image-to-text description model. A product, service, or brand that is consistent with the context and storyline of the captured image is selected and ad video sequence description text is generated that includes descriptions and a storyline blending descriptions of the selected product, service, or brand with the context and storyline of the primary video stream. The ad video sequence description text is used to prompt a text-to-video generation model that generates a new advertisement video clip, which is inserted into the primary video stream before distribution to video content rendering devices.
Owner:CHARTER COMM OPERATING LLC

Intelligent event identification method and system based on high-speed camera

The invention provides an intelligent event identification method and system based on a high-speed camera, and the method comprises the steps: setting the collection parameters of the high-speed camera, and triggering the camera to collect a target scene video stream. And hardware acceleration decoding processing is carried out on the collected original video data stream, real-time environment illumination information of the environment illumination sensor is obtained, and dynamic brightness equalization processing is executed. And performing motion adaptive denoising processing on the video sequence. Geometric distortion correction is carried out on the image sequence through camera calibration parameters, sub-pixel-level displacement vectors and dense optical flow field data of a moving object are extracted, and multi-scale morphological features are extracted. And the features are fused to generate motion feature data, the data are processed through a spatio-temporal joint event classification model, an event identification result is output, the result is bound with a high-precision timestamp, and event identification information is output to an industrial control system display device in real time. According to the invention, the accuracy and real-time performance of event identification can be improved.
Owner:广州思林杰科技股份有限公司

Multi-mode collaborative awareness power station high-risk operation inspection method and system

The invention provides a multi-mode cooperative sensing power station high-risk operation inspection method and system, and relates to the technical field of video recognition, and the method comprises the steps: activating an unmanned plane and a quadruped robot after a power station operation task is started; starting a video acquisition unit, and establishing a synchronous video stream; carrying out fusion alignment with the global reference coordinate system through an external synchronization signal; inputting the video sequence of the fused view angle into a multi-view angle action behavior recognition network, and establishing a dangerous behavior grade score; auditory data and olfactory data of the quadruped robot are obtained, and linkage abnormity is established; and polling abnormity is reported according to linkage abnormity and dangerous behavior grade scores. Through the method and the device, the technical problem of low inspection efficiency caused by difficulty in comprehensively identifying potential risks in a dynamic environment due to limitation of a single sensing mode is solved, and the inspection efficiency of a power station is improved by fusing multi-mode data, timely finding and processing the potential risks and improving the accuracy of video and audio identification.
Owner:BEIJING HUADIAN TIANREN ELECTRIC POWER CONTROL TECH

Non-contact heart rate detection method, system and device based on visual Transform and multi-scale feature aggregation and medium

The invention discloses a non-contact heart rate detection method, system and device based on visual Transform and multi-scale feature aggregation and a medium. The method comprises the steps that a face visible light video is obtained, the unified video frame rate is re-sampled through the frame rate, face key point positioning and region division are carried out on each frame of image, and a face visible light video sequence is obtained; performing time migration operation on the video sequence to obtain a difference frame sequence, performing channel fusion on the video sequence and a difference frame, and performing down-sampling on a spatial dimension to obtain a low-resolution feature tensor; constructing a non-contact heart rate extraction model, inputting the low-resolution feature tensor into the non-contact heart rate extraction model, and outputting a heart rate value of the target; the non-contact heart rate extraction model comprises a feature enhancement module, a multi-scale mask feature aggregation module, a Transform time sequence modeling module and an rPPG predictor, the system, the device and the medium are used for achieving the non-contact heart rate detection method based on visual Transform and multi-scale feature aggregation, and the precision of rPPG signal extraction is improved.
Owner:NORTHWEST UNIV

System and method for AI-powered narrative analysis of video content

A system, a method and a processor are for AI-powered generation and delivery of video clips. The processor is configured to: load a first video file of a first video content item, the first video file comprising video frames associated with timestamps; load a first subtitle file of the first video content item, the first subtitle file comprising subtitle text associated with the timestamps; execute a natural language processing (NLP) model with the subtitle text as input, the NLP model including language pre-processing steps for classifying words, names or phrases in the subtitle text and associating initial classifiers with the subtitle text, the NLP model including one or more of a recurrent neural network (RNN), a Bidirectional Encoder Representations from Transformers (BERT) model, or a generative pre-trained transformer (GPT) model for a dialogue analysis comprising processing sequences of dialogue in the subtitle text in view of the initial classifiers to associate one or more portions of the dialogue with one or more first classifiers of first narrative elements; execute an image recognition model with at least some of the video frames as input, the image recognition model including a convolutional neural network (CNN) for an object detection analysis and a facial recognition analysis comprising processing video sequences to associate one or more of the video frames with one or more second classifiers of second narrative elements; generate a narrative map of the first video content item by temporally aligning the first narrative elements with the second narrative elements based on the timestamps associated with the video frames and the first subtitle file; and generate a video clip including at least one segment of the first video content item, the at least one segment including selected video frames associated with at least one of the first or second narrative elements identified from the narrative map and selected for inclusion in the video clip.
Owner:PARAMOUNT GLOBAL INC

Guideboard surface deformation detection method and system based on dynamic video sequence

The invention discloses a guideboard surface deformation detection method and system based on a dynamic video sequence. The method comprises the following steps: acquiring road video data with a timestamp and position information; a traffic sign is automatically identified from the road video data, and geometric anomaly detection is carried out; a time sequence tracking sequence is established for the guideboard, and whether continuous abnormity exists in the guideboard is determined through multi-frame analysis; if yes, performing local texture consistency and local contour geometric analysis on the guideboard to obtain a fusion analysis result; performing state evaluation on supporting equipment corresponding to the guideboard to obtain a state evaluation result of the guideboard supporting equipment; and determining the final surface deformation state and the corresponding grade of the guideboard based on the fusion analysis result and the state evaluation result of the guideboard supporting equipment. By implementing the method provided by the invention, automatic detection and quantitative analysis of the bending or deformation condition of the road sign can be realized without manual intervention, and early warning information can be output in time.
Owner:WINTOO INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Activity-based person identification using biometric disentanglement

A system and method for person identification from video data by disentangling biometric identity features from non-biometric appearance and activity features are disclosed. The system processes RGB video sequences depicting individuals performing various activities to extract spatio-temporal features. These features are separated into distinct biometric identity representations and non-biometric features related to appearance and performed activities. To achieve this separation and minimize appearance bias, the system utilizes an auxiliary supervisory model. At least two implementations of this supervisory model are disclosed: one using semantic supervision via structured embeddings processed through a vision-language model, and another employing silhouette-based feature distillation from a silhouette-trained neural network. Joint training for biometric identification and activity classification ensures accurate identification of individuals independently of facial visibility, clothing differences, or activity variations.
Owner:UNIVERSITY OF CENTRAL FLORIDA RESEARCH FOUNDATION INC

Monitoring video enhancement method for farm

The invention belongs to the technical field of video processing, and particularly relates to a monitoring video enhancement method for a farm, which aims to solve the technical problem of low quality of an enhanced video in the prior art, and comprises the following steps: S1, processing each frame of image in a monitoring video sequence frame by frame; s2, distinguishing a target animal area from a background area, and identifying and generating an artifact mask; s3, aiming at the background area, carrying out key smoothing processing on the artifact position to inhibit the artifact; s4, for the image of the target animal area, performing adaptive nonlinear enhancement on the brightness component, and performing color correction on the chrominance component; s5, performing pixel-level fusion on the enhanced target animal area and the background area; and S6, spreading the information of the previous frame to the current frame by using the forward optical flow field, and carrying out weighted fusion on the information of the previous frame and the current frame. According to the method, the target bred animals, the background areas and the artifacts are accurately distinguished, so that refined and differentiated processing of pictures is realized.
Owner:EGG NO 1 FOOD CO LTD

Craniofacial dynamic reconstruction method and system based on multi-modal data fusion

The invention relates to the technical field of medical image processing, and discloses a craniofacial dynamic reconstruction method and system based on multi-modal data fusion, and the method comprises the steps: arranging a multi-modal data collection device in a target craniofacial region, and obtaining a static CT image, a static MRI image, a dynamic expression video sequence and a surface electromyogram signal; preprocessing the static CT image, the static MRI image, the dynamic expression video sequence and the surface electromyogram signal; inputting the preprocessed data into a multi-scale finite element model, and simulating a coupling relationship between muscle contraction force and skin deformation by adopting a biomechanical driving strategy to generate a dynamic craniofacial model; and fusing the geometric error and the motion consistency score of the real data by adopting a linear regression method, and outputting a comprehensive reconstruction quality index. According to the method, the problems of low craniofacial dynamic modeling accuracy and poor robustness in the prior art can be solved.
Owner:青峰宇

Ultrasonic image discrimination perception pre-training method based on cooperative training framework

The invention provides an ultrasonic image discrimination perception pre-training method based on a cooperative training framework, and relates to the technical field of ultrasonic image analysis, and the method comprises the steps: obtaining an ultrasonic image sequence containing time sequence information; constructing a cooperative training architecture, wherein the two networks are both based on a visual converter and embedded with a time sequence anatomical attention module; inputting the original image into a teacher network, inputting the enhanced image into a student network, and respectively outputting global and local features; calculating an anatomical continuity measure based on the output features of the two-network time sequence anatomical attention module; constructing a total loss function; the regularization weight is dynamically adjusted according to the anatomical continuity measurement, and student network parameters are updated through gradient descent; dynamically calculating an index moving average coefficient according to the anatomical continuity measurement so as to update teacher network parameters; and iterative training is carried out until convergence. According to the method, the time sequence continuity and anatomical structure information contained in the ultrasonic video sequence are fully mined, so that the adaptive optimization of the model training process is realized.
Owner:XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV +1

Robot medical image segmentation and feature extraction method for precise operation

The invention relates to the field of medical image segmentation, and discloses a precision surgery-oriented robot medical image segmentation and feature extraction method, which comprises the steps of constructing a dynamic segmentation network model of a bidirectional attention architecture, and inputting a preprocessing module for feature extraction and generating a multi-scale feature pyramid; the Transform coding branch is used for time sequence feature modeling and outputting a time sequence enhancement feature; the convolutional coding branch is used for enhancing anatomical features and surgical instrument features and outputting spatial enhancement features; the multi-stage feature fusion unit is used for performing multi-stage iterative fusion and outputting final fusion features; the decoding output module is used for decoding and generating pixel-level segmentation masks of the anatomical structure and the surgical instrument; training the dynamic segmentation network model; and performing medical image segmentation and feature extraction based on a medical image video sequence input in real time by using the trained dynamic segmentation network model. Accurate segmentation of anatomical tissues and dynamic instruments in an operation scene is realized.
Owner:BEIJING JISHUITAN HOSPITAL

Long-term video understanding system based on adaptive sparse memory and language model

The invention provides a long-term video understanding system based on adaptive sparse memory and a language model. The long-term video understanding system comprises a visual encoder used for extracting visual features from a long video; the memory bank module is used for storing and retrieving visual features of historical video contents; and the sparse adaptive module is used for dynamically managing the memory bank module, and the memory bank module interacts with the multi-modal large language model by querying a converter Q-Former and is used for incrementally processing video data and mapping visual features to a language space. By introducing an adaptive sparse memory mechanism, a long-term video sequence can be effectively processed, redundant features can be dynamically compressed, and key information can be reserved, so that efficient analysis of a long video is realized; compared with the prior art, the method has high accuracy in multiple tasks, and can dynamically manage the memory bank and reduce processing of redundant features through a sparse adaptive mechanism, so that the calculation overhead is reduced, and the overall efficiency of the system is improved.
Owner:JINAN UNIVERSITY

Multi-mode collaborative video sequence segmentation method

The invention discloses a multi-modal collaborative video sequence segmentation method. The method comprises the following steps: obtaining a multi-scale local feature matrix and a multi-scale global feature matrix of an image sequence; obtaining a multi-scale text feature matrix of the text sequence; obtaining a multi-scale local-global fusion feature matrix of the multi-scale local feature matrix and the multi-scale global feature matrix; obtaining a multi-modal fusion feature matrix of the multi-scale local-global fusion feature matrix and the multi-scale text feature matrix; and utilizing a decoder of the pre-trained large model to predict and generate a segmentation mask, and outputting a semantic segmentation map. The video sequence segmentation method is stable in performance when facing complex and changeable scenes, does not need to depend on a large amount of labeled data, reduces the training cost, and is suitable for various practical application fields including intelligent monitoring, automatic driving, medical image analysis and the like.
Owner:HARBIN INST OF TECH AT WEIHAI +1

Weak supervision video anomaly detection method based on prompt learning knowledge enhancement

The invention discloses a weak supervision video anomaly detection method based on prompt learning knowledge enhancement, and belongs to the technical field of video intelligent analysis. A video side gives a section of abnormal scene video, video sequence features and audio sequence features are obtained through a feature extraction network, then a trained and complete feature aggregation network is input to carry out multi-modal feature aggregation, an abnormal score is obtained through a score prediction network, and text representation is carried out based on prompt learning. A prompt template is constructed for abnormal video tags through a knowledge graph, semantic expansion is performed on normal tags through a plurality of learnable parameters, cross-modal alignment is performed on the normal tags and a video side, so that features of the video side are close to different normal semantics, knowledge enhancement is performed by introducing external information, positive abnormal boundaries of the video are learned, and the detection performance is improved. And finally, multi-task joint optimization is carried out through different loss functions, and abnormal video clip positioning is carried out.
Owner:COMMUNICATION UNIVERSITY OF CHINA

OLS for multi-view scalability

To provide a video coding mechanism.SOLUTION: A video coding mechanism includes receiving a bitstream comprising an output layer set (OLS) and a video parameter set (VPS). The OLS includes one or more layers of coded pictures and the VPS includes an OLS mode identification code (ols_mode_idc) specifying that for each OLS, all layers in each OLS are output layers. The output layers are determined on the basis of the ols_mode_idc in the VPS. The coded picture from the output layers is decoded to produce a decoded picture. The decoded picture is forwarded for display as a part of a decoded video sequence.SELECTED DRAWING: Figure 7
Owner:HUAWEI TECH CO LTD

Dynamic object suppression video image stabilization method based on depth information and global consistency

The invention discloses a dynamic object suppression video image stabilization method based on depth information and global consistency, and the method comprises the steps: extracting key points in a video, carrying out the optical flow estimation, and obtaining the depth information of the key points; a static weight and a motion vector residual error are calculated based on the motion information, the depth information and the global homography matrix of each key point, a dynamic weight is obtained based on an attention mechanism, the dynamic weight and the static weight are fused to obtain a comprehensive weight, and a weighted motion residual error is calculated based on the comprehensive weight and the motion vector residual error; spreading the motion information to grid vertexes, obtaining dense residual motion based on weighted motion residual prediction, and accumulating the dense residual motion to obtain a track; and carrying out smoothing and motion compensation on the track and outputting a stable video sequence. The problem that in a complex scene, motion of a dynamic object causes distortion and inaccuracy of a stabilization result is solved, interference of the dynamic object can be effectively suppressed, and the effect and robustness of video image stabilization are improved.
Owner:WUHAN UNIV OF SCI & TECH

Multi-target fruit tracking detection and counting method based on complex orchard

The invention discloses a multi-target fruit tracking detection and counting method based on a complex orchard. The method comprises the following steps: acquiring a to-be-detected video sequence; the to-be-detected video sequence is input to a fruit detection model, a detection result is obtained, the fruit detection model is constructed through an improved YOLOv8n network model and is obtained through training of a training set, and the training set is fruit image data; inputting the fruit image data and the detection result into a trajectory prediction model to obtain a prediction result of the fruit, the trajectory prediction model being obtained by introducing a dynamic Kalman filtering algorithm of a variable forgetting factor; and performing data association on the detection result and the prediction result to obtain fruit position information and a counting result. According to the invention, the prediction precision can be further improved, noise accumulation and prediction errors are reduced, and continuous tracking of fruits in a video sequence is realized.
Owner:南宁桂电电子科技研究院有限公司 +1

Commodity video intelligent generation method based on multi-modal analysis and dynamic narrative architecture

The invention relates to the technical field of video generation, in particular to an intelligent commodity video generation method based on multi-modal analysis and a dynamic narrative architecture, which comprises the following steps: inputting original commodity data and a user behavior log, and outputting a dynamic selling point weight vector through a selling point value evaluation engine; s2, inputting the dynamic selling point weight vector output in S1 into a narrative flow generator, and executing the following steps: intercepting first K core selling points according to the weight vector to form a narrative trunk chain; based on the historical preference data of the user, inserting an emotion enhancement node to generate a branch enhancement narrative flow; generating a multi-version narrative flow instruction set in combination with the real-time network bandwidth data; retrieving a preset video clip library according to the instruction set to generate an initial video sequence; detecting semantic fault regions between adjacent segments to generate compensation animation parameters; and injecting compensation animation parameters and outputting a continuous narrative video stream. According to the invention, the jamming feeling and the frame skipping phenomenon caused by the change of a narrative structure or different sources of video clips are reduced, and the overall watching experience of a user is improved.
Owner:SHENZHEN YINGMENG INTELLIGENT TECHNOLOGY CO LTD

Video object segmentation method based on query adaptive attention and discriminative memory

The invention belongs to the technical field of computer vision and digital video processing, and discloses a video object segmentation method based on query adaptive attention and discriminative memory, which comprises the following steps: step 1, constructing a video data set and preprocessing data; step 2, constructing a video object segmentation model combining multi-scale semantic feature integration and query adaptive discriminant enhancement, and training the video object segmentation model; step 3, video object segmentation reasoning; performing target segmentation reasoning on the preprocessed video sequence, outputting a frame-by-frame target segmentation mask, and generating a time sequence tracking result; 4, exporting, deploying and applying the model; through lightweight network design and a discriminative memory optimization strategy, the model is deployed to edge equipment, and real-time segmentation and visualization are realized. According to the method, on one hand, comprehensive target representation is provided on multi-scale feature extraction, and on the basis of the characteristic that only high-confidence target features are stored, error propagation is avoided, and the long-term segmentation stability of the method is ensured.
Owner:NANTONG INST OF TECH

Video editing and artistic creation system based on AI intelligence

The invention discloses a video editing and artistic creation system based on AI intelligence, which relates to the technical field of video editing and comprises a data acquisition module, a feature analysis module, a feature library construction module, a rule matching module and a split mirror generation module, the data acquisition module is used for acquiring traditional opera stage video data and corresponding audience feedback audio waveform data, and the feature analysis module is used for extracting spatial-temporal feature vectors of stylized actions from the traditional opera stage video data. The feature library construction module is used for constructing an action feature library containing action type codes and music beat deviation level codes according to the spatio-temporal feature vectors; the method has the beneficial effects that dynamic association analysis is performed by constructing the action feature library and combining audience feedback audio data, the short video sequence conforming to the art rule is automatically generated, and the method has the advantage of improving the Chinese opera performance video editing efficiency and the art expressive force.
Owner:PACO VIDEO TECH (HANGZHOU) CO LTD

Water body color recognition regression method and system based on space-time causality and manifold learning

The invention belongs to the field of environment monitoring and computer vision, and particularly relates to a water body color recognition regression method and system based on space-time causality and manifold learning, and the method mainly comprises the steps: carrying out the detection of a current target water body video sequence, extracting a water body region, carrying out the high-dimensional feature dimension reduction of the water body region, and obtaining a water body color recognition result; and performing feature extraction on the water body region through a space-time causal feature learning model, fusing the flow shape learning features and the space-time causal features to obtain fused features, and outputting a finally predicted water body color value. According to the method, end-to-end assembly line design of preprocessing-segmentation-feature modeling-regression is adopted, manual intervention is not needed from video input to color prediction, and through cascade cooperation of five core modules (video preprocessing, water body segmentation, manifold learning, time sequence causal modeling and color recognition), the real-time performance of the system is improved. Full-link automation from environmental interference suppression, feature extraction to result output is realized, information loss of intermediate links is avoided, and recognition efficiency and robustness are improved.
Owner:CHINA TOWER CO LTD

Micro-expression recognition method based on staged adaptive course learning

The invention relates to a micro-expression recognition method based on staged adaptive course learning. The method comprises the following steps: A, preprocessing a micro-expression video sequence and a macro-expression video sequence; b, constructing a spatial-temporal feature fusion model, performing deep feature extraction on the micro-expression data set obtained by preprocessing, and pre-training a macro-expression recognition teacher model; c, constructing a deep learning algorithm based on staged adaptive curriculum learning, and introducing a micro-expression recognition-oriented deep learning algorithm based on staged adaptive curriculum learning in the training process of the constructed spatio-temporal feature fusion model to optimize the training process; and D, carrying out classification identification on the macro expression identification teacher model obtained by training on a test set. According to the method, more effective and discriminative micro-expression features are obtained, the generalization ability of the model is improved, and the problems that in the existing micro-expression recognition field, available data sets are lacked, redundant information contained in the data sets is large, and the recognition accuracy is not high are further solved.
Owner:SHANDONG UNIV +1

Diffusion based end-to-end in-scene media generation

Embodiments of the present disclosure provide techniques for performing virtual object placement in a video sequence using generative artificial intelligence models. An example method generally includes receiving an input prompt specifying an object to insert into a scene depicted in an input image stream; decoding, using a generative artificial intelligence model, perspective and lighting information for the input image stream; determining, based on the decoded perspective and lighting information, a location in the scene in which the object is to be inserted; and generating, using the generative artificial intelligence model, an output image stream including the object into the scene at the determined location, wherein visual effects for the object are based on the perspective and lighting information for the input image stream.
Owner:REMBRAND INC

Unsupervised region-growing network for object segmentation in atmospheric turbulence

An unsupervised region-growing network (RGN) is trained to perform object segmentation on video data degraded by atmospheric turbulence. The method includes obtaining input data containing turbulence-degraded video, extracting a video frame sequence, and training the RGN using a selected algorithm incorporating a region-growing algorithm and a grouping loss function. A bidirectional optical flow sequence is computed for multiple reference frames within the video sequence. Pixel-level masks are generated for detected moving objects, followed by applying the region-growing algorithm to create coarse masks. A grouping loss function refines these masks to ensure consistency across consecutive frames. The trained RGN outputs refined masks as object segmentation data for the received video, improving segmentation accuracy in turbulent environments. This approach enables robust object detection and segmentation without requiring prior video restoration, maintaining fidelity to the original turbulence-distorted input.
Owner:CLEMSON UNIV RES FOUND +2

Al-based video content analysis method and system

The invention relates to the technical field of content recognition, in particular to an Al-based video content analysis method and system, and the method comprises the following steps: carrying out the image segmentation learning based on a video sequence through employing a U-Net convolutional neural network, analyzing scenes and elements in a video frame, recognizing and isolating key visual elements in a video through network learning, and carrying out the recognition of the key visual elements in the video. Comprising objects and figures. According to the method, image segmentation learning is carried out by adopting the U-Net convolutional neural network, key visual elements in the video can be identified and isolated more accurately, a clearer basis is provided for follow-up scene change and key event tracking, video content analysis is carried out by applying the graph neural network, the identification capability of dynamic scenes and events is enhanced, and the identification efficiency is improved. Deeper structured understanding is provided for video content indexing, the video quality is analyzed through a structural similarity index evaluation method, and the reason of visual quality reduction can be accurately recognized and improved.
Owner:CHENDA (GUANGZHOU) NETWORK TECH CO LTD +1

Target tracking method and system for deep twin network

The invention discloses a target tracking method and system for a deep twin network, and relates to the technical field of computer vision and artificial intelligence, and the target tracking method comprises the following steps: obtaining a to-be-tracked video sequence and a target template image in an initial frame, a target template image and a search area image of a current frame are input into a lightweight backbone network, multi-scale feature maps are extracted respectively, the lightweight backbone network is improved, a dynamic channel pruning module is embedded, cross-layer fusion is performed on the multi-scale feature maps, and a module is updated through a self-adaptive template. In combination with a dynamic channel pruning module, redundant channels can be closed in a self-adaptive manner, the model parameter quantity and the calculation quantity are reduced, a pruning threshold is dynamically optimized through meta-learning, and it is ensured that key feature information is reserved while complexity is reduced.
Owner:TIANJIN MODERN VOCATIONAL TECH COLLEGE

Audio and video depth forgery detection method based on quality perception and multi-scale alignment

The invention discloses an audio and video depth forgery detection method based on quality perception and multi-scale alignment, and the method comprises the following steps: coding a synchronous audio and video sequence, and obtaining a frame-level visual feature, a facial action unit and a phoneme-level voice representation; a visual quality evaluation module is introduced to generate a spatial reliability mask, and quality weighting is carried out on the visual features; designing a global-local multi-scale cross-modal alignment mechanism, performing bidirectional cross-attention modeling on voice and face dynamic synchronization globally, and performing physiological coupling alignment on phonemes and face action units locally; and an uncertainty perception reasoning and calibration scheme is provided, adaptive temperature scaling is carried out according to quality and consistency, and uncertainty calibration is carried out by self-supervision loss. According to the method, the problems of insufficient robustness and excessive self-confidence misjudgment of an existing method in a low-quality video and high-synchronization counterfeit scene are solved, and the cross-dataset generalization capability and the actual deployment reliability are remarkably improved.
Owner:NANJING UNIV OF SCI & TECH