Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

51 results about "Video prediction" patented technology

Video encoding and decoding using deep learning based inter prediction

A video encoding or decoding apparatus and method perform existing inter prediction on a current block to generate a motion vector and predicted samples. The video encoding or decoding apparatus and method generate enhanced predicted samples for the current block using a deep learning-based video prediction network (VPN) on the basis of the motion vector, the reference samples, the predicted samples, and the like to improve encoding efficiency.
Owner:HYUNDAI MOTOR CO LTD +2

Video sequence prediction method and system based on object segmentation guidance

The invention relates to the field of video prediction, in particular to a video sequence prediction method and system based on object segmentation guidance, and the prediction method comprises the steps: receiving a historical video frame sequence, and carrying out the video object segmentation and tracking processing of the historical video frame sequence, generating structural representation information of each object in each frame and distributing a continuous and unique tracking ID for each object; encoding the structural representation information and the tracking ID into an object-level structured feature sequence; inputting the object-level structured feature sequence into a conditional diffusion model as a guide condition, and generating a potential spatial intermediate feature representing a future video frame through an iterative denoising process; the potential spatial intermediate features are decoded into a pixel-level sequence of future video frames. According to the method, the problems of the existing video prediction technology in the aspects of object consistency, physical authenticity, error accumulation and the like are solved, the application potential in a complex scene is expanded, and support is provided for video prediction in the fields of automatic driving, robot perception, content creation and the like.
Owner:JILIN HUAQIAO FOREIGN LANGUAGES INST

Lemon health analysis method

The invention discloses a lemon health analysis method, and relates to the technical field of intelligent agricultural monitoring and plant health diagnosis, lemon plant images at continuous moments are collected through a hyperspectral camera, and the images are input into a dynamic differential spectrum segmentation model for change detection. And carrying out sequence modeling on the pixel-level multiband spectral data. The self-adaptive motion perception scheduler is used for automatically inhibiting loss when the pixel proportion of the motion change mask exceeds a preset dizziness threshold value so as to avoid background misjudgment, and a pixel-level dynamic change map is generated. And compressing into a pathological dynamic state vector through a variational auto-encoder, and performing end-to-end training by taking action condition video prediction as a training target. In the deduction stage, the cyclic dynamic model is used for executing multi-step imagination prediction based on the candidate action sequence, the future evolution trajectory of the pathological dynamic state vector is output, and an outbreak risk level quantitative early warning or decision simulation report is generated. The system further comprises change type classification and model closed-loop optimization functions.
Owner:WEISHAN JUFENG AGRI TECH CO LTD

Virtual reality video adjustment method and device, equipment and storage medium

The invention provides a virtual reality video adjusting method and device, equipment and a storage medium, and the method comprises the steps: in a process that at least two users watch the same virtual reality video, adjusting the virtual reality video based on the user features of each user and a target video which is not played in the virtual reality video; predicting an emotional stimulation value corresponding to each moment when each user watches the target video; determining a target moment based on the emotional stimulus value corresponding to each moment when each user watches the target video; determining an inter-cut video based on the emotional stimulus value corresponding to at least one target user in the at least two users at the target moment, the inter-cut video being used for adjusting the emotional stimulus value of the at least one user; and inserting the inter-cut video after the target moment in the virtual reality video to obtain an adjusted virtual reality video. According to the invention, the flexibility of playing the virtual reality video can be improved, and the experience feeling of each user who participates in watching the same virtual reality video is improved.
Owner:XIAN UNIVIEW INFORMATION TECH CO LTD

Efficient video prediction using motion graph

A video prediction technique generates a motion graph based on given video frames. The motion graph includes spatial edges and temporal edges. Each spatial edge describes a same-frame semantic relationship between two graph nodes that are associated with a same video frame. Each temporal edge describes an interframe relationship between two graph nodes of temporally neighboring frames. The temporal edges include backward temporal edges and forward temporal edges. The technique further includes generating initial motion feature information associated with the graph nodes in the plural given video frames, and updating the motion feature information by performing message-passing operations. The technique decodes the motion feature information into dynamic vector information. The technique then predicts and synthesizes a subsequent video frame based on the given video frames and the dynamic vector information.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Efficient Video Prediction using Motion Graph

A video prediction technique generates a motion graph based on given video frames. The motion graph includes spatial edges and temporal edges. Each spatial edge describes a same-frame semantic relationship between two graph nodes that are associated with a same video frame. Each temporal edge describes an interframe relationship between two graph nodes of temporally neighboring frames. The temporal edges include backward temporal edges and forward temporal edges. The technique further includes generating initial motion feature information associated with the graph nodes in the plural given video frames, and updating the motion feature information by performing message-passing operations. The technique decodes the motion feature information into dynamic vector information. The technique then predicts and synthesizes a subsequent video frame based on the given video frames and the dynamic vector information.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Method and apparatus for video predictive coding

Provided is a method for video predictive coding. The method includes: determining, according to an executed mode, information of the executed mode in a decision making process of a best mode of a current prediction unit in inter-frame prediction, wherein the information of the executed mode includes a temporary best mode and a cost of the temporary best mode; and determining, based on the information of the executed mode, whether to skip an intra-frame prediction mode of the decision making process.
Owner:BIGO TECH PTE LTD

Parallel processing method and system for video decoding and video prediction

The invention discloses a parallel processing method and system for video decoding and video prediction, and belongs to the technical field of computer vision. The method comprises the following steps: acquiring a compressed code stream; reconstructing the current frame image in combination with the compressed code stream and a previously decoded video frame to obtain a currently decoded video frame; a video frame image at a future moment is recursively predicted based on a previously decoded video frame and a current decoded video frame. According to the method, the system time delay and the calculation overhead can be reduced, the requirements of low-time-delay video transmission and prediction are met, and the method can be widely applied to automatic driving, remote cooperative control, robot navigation and other low-delay-requirement scenes.
Owner:PEKING UNIV

Video code rate determination method and apparatus, electronic device, medium, and product

PendingCN122457831AVideo encodingSimulation
The embodiments of the present disclosure disclose a video code rate determination method, device, electronic equipment, storage medium and product. The method comprises: in response to triggering of a video transcoding event, determining a target transcoded video; predicting a predicted playing probability of the target transcoded video on a terminal device corresponding to each device performance level; determining a candidate video transcoding code rate combination according to a preset video code rate level; determining a comprehensive playing performance analysis result of a video transcoding result corresponding to each candidate video transcoding code rate combination on the terminal device corresponding to each device performance level; performing optimal comprehensive playing performance analysis on the predicted playing probability and the comprehensive playing performance analysis result of the terminal device corresponding to each device performance level; and determining a target video transcoding code rate combination according to the optimal comprehensive playing performance analysis result. The technical scheme of the embodiments of the present disclosure can determine a coding code rate that maximizes the business value of video coding.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Speaking head video synthesis method and device, computer equipment and storage medium

The embodiment of the invention provides a speaking head video synthesis method and device, computer equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring training audio features and training image features of a training object; performing expression prediction on the training audio features through a preset speaking head synthesis model to obtain predicted expression parameters; performing speaking head video construction on the predicted expression parameter and the training image through a preset speaking head synthesis model to obtain a predicted speaking head video; according to the training image features, the predicted speaking head video and the predicted expression parameters, performing parameter adjustment on a preset speaking head synthesis model to obtain a target speaking head synthesis model; and obtaining a target audio feature, and performing talking head video synthesis on the target audio feature through the target talking head synthesis model to obtain a target talking head video. The method and the device can be applied to business systems needing a large amount of data, such as financial science and technology and health medical treatment, and speaking head videos capable of synthesizing expression changes and improving user experience can be synthesized.
Owner:PING AN TECH (SHENZHEN) CO LTD

General online AI streaming methods, devices, equipment, and media

This application provides a general online AI streaming processing method, apparatus, device, and medium. The method includes: querying the cached results in the video AI buffer of an AI server based on a video prediction request sent by a user; if no cached result corresponding to the URL address is found in the video AI buffer, the video AI processing thread of the AI ​​server checks whether a playback instruction sent by the user has been received within a preset time threshold from the first current time; if a playback instruction sent by the user is confirmed to have been received, the video AI processing thread retrieves the video image from the media server pointed to by the URL address, and sends the AI ​​prediction result of the video image as the cached result corresponding to the URL address to the video AI buffer. This method allows a wide range of internet users to find different media servers based on different URL addresses and perform online AI prediction on the video images from the media servers, which is very convenient.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Road network traffic state video prediction method and device and electronic equipment

The embodiment of the invention provides a road network traffic state video prediction method and device and electronic equipment, and relates to the field of intelligent traffic and computer vision, and the method comprises the steps: obtaining a to-be-processed video, determining to-be-processed text information corresponding to the to-be-processed video, and transmitting the to-be-processed text information to a server; and taking the to-be-processed video and the to-be-processed text information corresponding to the to-be-processed video as inputs of the trained U-Net model to obtain a denoised first video, taking the first video and the to-be-processed text information corresponding to the to-be-processed video as inputs of the trained video prediction model to obtain a predicted video corresponding to the first video, and performing video prediction on the predicted video. And de-noising the to-be-processed video based on the trained U-Net model, and performing guidance generation on the predicted video based on the to-be-processed text information corresponding to the to-be-processed video, thereby improving the accuracy of traffic prediction.
Owner:BEIHANG UNIV

Rate-distortion prediction based method and system for rate control of depth video encoder

The application provides a rate distortion prediction-based deep video encoder code rate control method and system, which comprises the following steps: step 1, training a prediction module; step 2, inputting a video frame into the prediction module to obtain a prediction point set; step 3, fitting a code rate and quality model according to the prediction point set; step 4, obtaining a frame-level code rate allocation ratio through a code rate control algorithm; and step 5, determining the corresponding encoding parameters of each frame and inputting the encoding parameters into an encoder for encoding. The application directly utilizes a neural network and an original video to predict the code rate model and the quality model of each frame for the first time, without pre-encoding; the video frame is down-sampled to a fixed small resolution before being inputted into the neural network, so that the efficiency is improved and the generalization is enhanced; and the application realizes code rate control at a mini-GOP level for the first time. Compared with the existing code rate control methods, the application can realize the same code rate control accuracy and finer code rate control granularity at a faster speed.
Owner:NANJING UNIV

Small sample class incremental video action recognition method and device

The application discloses a small sample class incremental video action recognition method and device, and belongs to the field of computer vision, the method comprises the following steps: for each video, by adopting visual soft prompt and time sequence soft prompt, video features of fusing space-time information are acquired, video features with prior knowledge are acquired at the same time, and the two kinds of video features are fused to acquire final video features; secondly, a text prototype of a class is extracted; finally, the similarity between the above-mentioned video features and the text prototype is calculated, and the input video is predicted as the class with the maximum similarity. The application can effectively capture the space-time features of the input video, improve the recognition accuracy of the video action, and the method is simple and flexible, which significantly improves the prediction accuracy of the new class, and can effectively alleviate the catastrophic forgetting phenomenon of the model on the old class.
Owner:ZHEJIANG LAB

A variable frame rate video generation method based on optical flow estimation

ActiveCN116708869Bquality improvementEncoder decoderVariable frame rate
A variable frame rate video generation method based on optical flow estimation, which introduces optical flow supervision information into an OpFode-Net model, the OpFode-Net model comprising an encoder-decoder structure; the encoder uses an ODE-ConvGRU to embed input video sequence X T into a hidden state h T ; wherein the ODE-ConvGRU uses a ConvGRU as a node of a neural ODE and embeds it into the neural ODE to realize dynamic modeling of the video sequence; the decoder starts from h T , and uses an ODE solver to generate a new video frame at any time step S, which can realize more accurate prediction results and achieve optimal performance in video interpolation and video prediction tasks.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Pixel-level video prediction with improved performance and efficiency

One aspect provides a machine-learned video prediction model configured to receive and process one or more previous video frames to generate one or more predicted subsequent video frames, wherein the machine-learned video prediction model comprises a convolutional variational auto encoder, and wherein the convolutional variational auto encoder comprises an encoder portion comprising one or more encoding cells and a decoder portion comprising one or more decoding cells.
Owner:GOOGLE LLC

Quantum auto-encoder model, video prediction method and related device

The invention discloses a quantum auto-encoder model, a video prediction method and a related device, and belongs to the technical field of quantum computing, the quantum auto-encoder model comprises an encoder, an intermediate module and a decoder which are sequentially connected in series, and the encoder, the intermediate module and the decoder all comprise different variable component sub-circuits; the encoder is used for extracting features of an input target image frame, and the target image frame is an image frame corresponding to a moment before a target moment; the intermediate module is used for generating an initial prediction image frame by using the features extracted by the encoder and a relationship between adjacent image frames constructed during prediction at a moment before a target moment; and the decoder is used for reconstructing the initial prediction image frame to obtain a target prediction image frame at a target moment. By applying the embodiment of the invention, the accuracy of video prediction is improved.
Owner:BENYUAN TIANGONG (ZHENGZHOU) QUANTUM TECH CO LTD

Video prediction model based on operator learning

PendingCN121887987AFlexible frame rate predictionBreak through limitsImage codingDigital video signal modificationPattern recognitionTemporal information
The invention relates to a video prediction model based on operator learning. The video prediction model comprises a coding layer, an operator layer and a decoding layer, the coding layer receives an externally input video, the video is a series of continuous picture sequences, gridding labeling is carried out on the video, and time information and the originally input picture sequences are fused; carrying out downsampling processing on the fused picture sequence, and sending a processing result to an operator layer; the operator layer adopts a self-adaptive Fourier neural operator to carry out Fourier transform on a result processed by the coding layer, and extracts spatial-temporal characteristics; and the decoding layer obtains a predicted picture sequence according to the spatial-temporal characteristics provided by the operator layer. According to the method, the application requirement of space-time flexibility of the prediction task can be met under the conditions of data missing, low input data frame rate and the like, and meanwhile, the precision in the video prediction task is relatively high.
Owner:CHINA ACAD OF AEROSPACE SCI & TECH INNOVATION

Method for training autonomous driving model, electronic device, and storage medium

Provided method for training an autonomous driving model including a video prediction model, and the method including: determining, according to at least one of an initial video frame collected by a target vehicle or scenario description metadata of an initial video frame, a scenario context of the initial video frame; determining a vehicle movement instruction of the target vehicle according to at least one of the initial video frame or trajectory data of the target vehicle corresponding to the initial video frame; and training an initial model using the initial video frame and a control text corresponding to the initial video frame, to obtain the video prediction model, where the control text comprises the scenario context and the vehicle movement instruction, and the video prediction model is configured to output a predicted video frame.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Video prediction method for diffusion model of unmanned aerial vehicle

The invention relates to the technical field of video processing, and discloses a video prediction method for a diffusion model of an unmanned aerial vehicle, which is suitable for frame missing, inter-frame interpolation and key frame completion scenes in a video shot by the unmanned aerial vehicle. Comprising the following steps: video compression and frame loss degradation modeling: compressing an original video sequence by using a vector quantization generative adversarial network, and simulating the degradation process in a potential space; after a video diffusion converter model is constructed, random sampling is carried out from standard normal distribution, and potential representation of initial noise is obtained; iterative optimization is carried out, and observation consistency correction is carried out; and restoring the potential representation after iterative optimization into a complete video frame through a decoder, and completing prediction and completion of the missing video frame. The time-space relationship is automatically learned through the self-attention mechanism of the video diffusion converter, prediction distortion caused by alignment errors is avoided without depending on traditional alignment modules such as optical flow, and frame missing at any position in the video can be directly processed through airborne hardware equipment of the unmanned aerial vehicle.
Owner:TOPXGUN (NAN JING) ROBOTICS CO LTD

Optical flow guided knowledge distillation video prediction model compression method and system

PendingCN122179582ASolve the problem of missing motion featuresefficient migrationBiological modelsDigital video signal modificationVisual technologyOptical flow
This invention discloses a video prediction model compression method and system based on optical flow-guided knowledge distillation, belonging to the field of computer vision technology. The method includes: generating optical flow distillation loss by constraining the optical flow prediction distribution of the student network to be consistent with that of the teacher network; generating channel alignment loss by aligning the channel attention distributions of the teacher and student networks using divergence; generating pixel-level reconstruction loss based on the difference in pixel values ​​between the generated image predicted by the student network and the real image; generating perceptual loss based on the difference in high-level semantic feature space between the generated image predicted by the student network and the real image; and obtaining the trained student network based on the optical flow distillation loss, channel alignment loss, pixel-level reconstruction loss, and perceptual loss. This invention can solve the performance degradation problem caused by neglecting spatiotemporal characteristics in existing compression methods for video prediction tasks.
Owner:PEKING UNIV

A method and system for video pre-distribution

The embodiment of the application provides a video pre-distribution method and system, which is used for making a client predict a video based on data with higher real-time performance, and improving the accuracy of the pre-distributed video. The method comprises the following steps: a client sends video playing behavior data to a server; the server generates video features based on the video playing behavior data; the server recalls a to-be-watched video list based on the video playing behavior data; the server sends the video features and the to-be-watched video list to the client; the client acquires local user behavior and video playing records; the client determines a local to-be-played list according to the video playing records; the client determines to-be-predicted video features according to the video features, the local user behavior, the to-be-watched video list and the local to-be-played list; the client determines a first to-be-cached video according to the to-be-predicted video features; and the client caches the first to-be-cached video.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Tennis self-training system

A tennis self-training system comprising a control device comprising recording unit configured to record a tennis game and a processor configured to analysis the tennis game based on a video obtained from the recording unit; and a ball machine unit configured to move and launch a ball according to the instructions of the control device; wherein the control device is configured to: determine the position of a player and the position of the ball machine unit based on the video, predict a falling position of the ball hit by the player based on the video, calculate a ball launch position and a ball arrival position of the ball machine unit based on the position of the player and the falling position of the ball, generate a control signal related to the ball launch position and the ball arrival position, transmit the control signal to the ball machine unit.
Owner:CURINGINNOS INC

An audio and video parsing method based on noise label learning

The application belongs to the technical field of deep learning, and relates to an audio and video parsing method based on noise label learning, which comprises the following steps: preprocessing original audio and video to obtain a segment-level input sequence; constructing a mutual learning noise-resistant double-flow network; training the mutual learning noise-resistant double-flow network according to a training set; comparing the validation set indicators of two sub-networks in the trained mutual learning noise-resistant double-flow network, and taking the sub-network with the larger validation set indicator as an audio and video parsing model; and parsing through the audio and video parsing model according to a test set to obtain a video prediction result. The mutual learning noise-resistant double-flow network is composed of two sub-networks with the same structure but different initializations, a cross filtering mechanism is executed according to the clean masks generated by the two sub-networks during the training of the mutual learning noise-resistant double-flow network, and the dynamic confidence ratio is gradually reduced through a cosine strategy, so that the problems of high pseudo-label noise rate and easy overfitting noise in the existing audio and video parsing task are solved.
Owner:UNIV OF ELECTRONIC SCI & TECH OF CHINA CHENGDU COLLEGE

Video prediction method based on multi-scale stream fusion

A video prediction method based on multi-scale stream fusion belongs to the technical field of computer vision, and comprises the following steps: applying a scale selector to an input continuous video frame to obtain an appropriate processing scale factor; respectively extracting inter-frame motion features by using an improved residual structure under a selected scale path and an original scale path; fully interacting motion features of two scales by means of a multi-scale feature fusion network; estimating optical flow and shielding conditions between each current frame and a future frame; reversely distorting the current frame by using the optical flow, and generating a prediction frame in combination with the shielding condition; an existing estimation result is continuously refined in an iteration mode; and combining the reconstruction loss and the perception loss to optimize the method model. According to the method, the scale selector and the multi-scale feature fusion network are provided, so that multi-scale moving objects can be processed, a prediction algorithm is enabled to adapt to a high-resolution video, and the application universality is improved.
Owner:JILIN UNIVERSITY +1

Predicting health or disease from user captured images or videos

A method can include capturing an image, images, or a video of a user with the image, images or video capturing system of an exercise device and detecting a health or disease indicator from the images or video. An apparatus includes a camera, a processor, and computer readable medium containing programming instructions that, when executed, will cause the processor to use a machine learning model and one or more images or videos of a subject captured by the camera to detect a health or disease indicator.
Owner:ADVANCED HEALTH INTELLIGENCE LTD

Noise model based compression

Techniques are disclosed for performing residual image compression techniques used in conjunction with image and / or video predictors. The techniques utilize a compression scheme that implements a noise model to estimate noise values of pixels in an originally acquired image. These noise value estimates are then used to perform residual image compression more efficiently by performing a non-uniform reduction in resolution of the residual image. The resolution reduction includes dropping least significant bits (LSBs) used to encode each pixel on a pixel-by-pixel basis based upon the noise value estimates of the originally acquired image.
Owner:MOBILEYE VISION TECH LTD

Video prediction method based on deep learning non-autoregression model

The invention provides a video prediction method based on a deep learning non-autoregression model, and relates to the field of man-machine interaction, climate prediction and automatic driving. According to the method, how to accurately predict a future video sequence according to an existing video sequence is researched, firstly, Encoder is used for decoding video frames, two-dimensional convolution calculation, LayerNorm normalization processing and SILU activation are carried out, and space information between the video frames can be extracted. And secondly, position codes are added in the model, the capability of extracting time information is increased on the basis of extracting the space information, and the time information and the space information are well fused, so that the problem that the space-time information is difficult to obtain by a non-autoregression model is solved. Besides, a previous attention mechanism is improved, a time-space attention mechanism with higher pertinence is designed, the acquisition of time-space information is further improved, and the accuracy of a prediction result is improved.
Owner:HOHAI UNIV

Video sequence prediction method and system based on object segmentation guidance

The present application relates to the field of video prediction, and more particularly to a video sequence prediction method and system based on object segmentation guidance, the prediction method comprising receiving a historical video frame sequence, and performing video object segmentation and tracking processing on the historical video frame sequence to generate structural representation information of each object in each frame and assign a unique tracking ID to each object; encoding the structural representation information and the tracking ID into an object-level structured feature sequence; inputting the object-level structured feature sequence as a guidance condition into a conditional diffusion model to generate latent space intermediate features representing future video frames through an iterative denoising process; and decoding the latent space intermediate features into a pixel-level future video frame sequence. The present application solves the problems of existing video prediction techniques in terms of object consistency, physical reality and error accumulation, and also expands the application potential in complex scenarios, providing support for video prediction in the fields of autonomous driving, robot perception, content creation and the like.
Owner:JILIN HUAQIAO FOREIGN LANGUAGES INST

Video prediction method and system, computer equipment and storage medium

The invention provides a video prediction method and system, computer equipment and a storage medium, and belongs to the field of computer vision, and the method comprises the steps: obtaining a continuous video sequence, and extracting a time feature, a bottom-layer texture, a middle-layer contour and a high-layer semantic multi-dimensional feature from the continuous video sequence; calculating attention based on the front theta-layer spatial memory to obtain the current time step spatial memory; based on the previous time step hidden state, the previous tau step time memory and the current time feature, generating the time memory of the current time step, fusing the time memory, the spatial memory and the previous time step hidden state, and outputting a top layer hidden feature; and carrying out weighted fusion on the layered spatial features by using weights, adding the layered spatial features with top hidden features, then carrying out up-sampling, and finally generating a prediction frame. Dynamic changes and detail structures are considered through multi-dimensional feature extraction; the spatial-temporal characteristic coherence is enhanced by using an attention mechanism; through feature fusion and iteration, historical and current information is integrated, video frame detail pictures are greatly restored, and the video precision is improved.
Owner:BEIJING JIAOTONG UNIV