Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

23 results about "Video prediction" patented technology

Efficient video prediction using motion graph

A video prediction technique generates a motion graph based on given video frames. The motion graph includes spatial edges and temporal edges. Each spatial edge describes a same-frame semantic relationship between two graph nodes that are associated with a same video frame. Each temporal edge describes an interframe relationship between two graph nodes of temporally neighboring frames. The temporal edges include backward temporal edges and forward temporal edges. The technique further includes generating initial motion feature information associated with the graph nodes in the plural given video frames, and updating the motion feature information by performing message-passing operations. The technique decodes the motion feature information into dynamic vector information. The technique then predicts and synthesizes a subsequent video frame based on the given video frames and the dynamic vector information.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Efficient Video Prediction using Motion Graph

A video prediction technique generates a motion graph based on given video frames. The motion graph includes spatial edges and temporal edges. Each spatial edge describes a same-frame semantic relationship between two graph nodes that are associated with a same video frame. Each temporal edge describes an interframe relationship between two graph nodes of temporally neighboring frames. The temporal edges include backward temporal edges and forward temporal edges. The technique further includes generating initial motion feature information associated with the graph nodes in the plural given video frames, and updating the motion feature information by performing message-passing operations. The technique decodes the motion feature information into dynamic vector information. The technique then predicts and synthesizes a subsequent video frame based on the given video frames and the dynamic vector information.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Parallel processing method and system for video decoding and video prediction

The invention discloses a parallel processing method and system for video decoding and video prediction, and belongs to the technical field of computer vision. The method comprises the following steps: acquiring a compressed code stream; reconstructing the current frame image in combination with the compressed code stream and a previously decoded video frame to obtain a currently decoded video frame; a video frame image at a future moment is recursively predicted based on a previously decoded video frame and a current decoded video frame. According to the method, the system time delay and the calculation overhead can be reduced, the requirements of low-time-delay video transmission and prediction are met, and the method can be widely applied to automatic driving, remote cooperative control, robot navigation and other low-delay-requirement scenes.
Owner:PEKING UNIV

Video code rate determination method and apparatus, electronic device, medium, and product

PendingCN122457831AVideo encodingSimulation
The embodiments of the present disclosure disclose a video code rate determination method, device, electronic equipment, storage medium and product. The method comprises: in response to triggering of a video transcoding event, determining a target transcoded video; predicting a predicted playing probability of the target transcoded video on a terminal device corresponding to each device performance level; determining a candidate video transcoding code rate combination according to a preset video code rate level; determining a comprehensive playing performance analysis result of a video transcoding result corresponding to each candidate video transcoding code rate combination on the terminal device corresponding to each device performance level; performing optimal comprehensive playing performance analysis on the predicted playing probability and the comprehensive playing performance analysis result of the terminal device corresponding to each device performance level; and determining a target video transcoding code rate combination according to the optimal comprehensive playing performance analysis result. The technical scheme of the embodiments of the present disclosure can determine a coding code rate that maximizes the business value of video coding.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

General online AI streaming methods, devices, equipment, and media

This application provides a general online AI streaming processing method, apparatus, device, and medium. The method includes: querying the cached results in the video AI buffer of an AI server based on a video prediction request sent by a user; if no cached result corresponding to the URL address is found in the video AI buffer, the video AI processing thread of the AI ​​server checks whether a playback instruction sent by the user has been received within a preset time threshold from the first current time; if a playback instruction sent by the user is confirmed to have been received, the video AI processing thread retrieves the video image from the media server pointed to by the URL address, and sends the AI ​​prediction result of the video image as the cached result corresponding to the URL address to the video AI buffer. This method allows a wide range of internet users to find different media servers based on different URL addresses and perform online AI prediction on the video images from the media servers, which is very convenient.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

A variable frame rate video generation method based on optical flow estimation

ActiveCN116708869Bquality improvementEncoder decoderVariable frame rate
A variable frame rate video generation method based on optical flow estimation, which introduces optical flow supervision information into an OpFode-Net model, the OpFode-Net model comprising an encoder-decoder structure; the encoder uses an ODE-ConvGRU to embed input video sequence X T into a hidden state h T ; wherein the ODE-ConvGRU uses a ConvGRU as a node of a neural ODE and embeds it into the neural ODE to realize dynamic modeling of the video sequence; the decoder starts from h T , and uses an ODE solver to generate a new video frame at any time step S, which can realize more accurate prediction results and achieve optimal performance in video interpolation and video prediction tasks.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Pixel-level video prediction with improved performance and efficiency

One aspect provides a machine-learned video prediction model configured to receive and process one or more previous video frames to generate one or more predicted subsequent video frames, wherein the machine-learned video prediction model comprises a convolutional variational auto encoder, and wherein the convolutional variational auto encoder comprises an encoder portion comprising one or more encoding cells and a decoder portion comprising one or more decoding cells.
Owner:GOOGLE LLC

Quantum auto-encoder model, video prediction method and related device

PendingCN121998114AFast multi-scale analysiseasy to captureQuantum computersBiological modelsQuantum circuitVideo prediction
The invention discloses a quantum auto-encoder model, a video prediction method and a related device, and belongs to the technical field of quantum computing, the quantum auto-encoder model comprises an encoder, an intermediate module and a decoder which are sequentially connected in series, and the encoder, the intermediate module and the decoder all comprise different variable component sub-circuits; the encoder is used for extracting features of an input target image frame, and the target image frame is an image frame corresponding to a moment before a target moment; the intermediate module is used for generating an initial prediction image frame by using the features extracted by the encoder and a relationship between adjacent image frames constructed during prediction at a moment before a target moment; and the decoder is used for reconstructing the initial prediction image frame to obtain a target prediction image frame at a target moment. By applying the embodiment of the invention, the accuracy of video prediction is improved.
Owner:BENYUAN TIANGONG (ZHENGZHOU) QUANTUM TECH CO LTD

Video prediction model based on operator learning

PendingCN121887987AFlexible frame rate predictionBreak through limitsImage codingDigital video signal modificationPattern recognitionTemporal information
The invention relates to a video prediction model based on operator learning. The video prediction model comprises a coding layer, an operator layer and a decoding layer, the coding layer receives an externally input video, the video is a series of continuous picture sequences, gridding labeling is carried out on the video, and time information and the originally input picture sequences are fused; carrying out downsampling processing on the fused picture sequence, and sending a processing result to an operator layer; the operator layer adopts a self-adaptive Fourier neural operator to carry out Fourier transform on a result processed by the coding layer, and extracts spatial-temporal characteristics; and the decoding layer obtains a predicted picture sequence according to the spatial-temporal characteristics provided by the operator layer. According to the method, the application requirement of space-time flexibility of the prediction task can be met under the conditions of data missing, low input data frame rate and the like, and meanwhile, the precision in the video prediction task is relatively high.
Owner:CHINA ACAD OF AEROSPACE SCI & TECH INNOVATION

Method for training autonomous driving model, electronic device, and storage medium

Provided method for training an autonomous driving model including a video prediction model, and the method including: determining, according to at least one of an initial video frame collected by a target vehicle or scenario description metadata of an initial video frame, a scenario context of the initial video frame; determining a vehicle movement instruction of the target vehicle according to at least one of the initial video frame or trajectory data of the target vehicle corresponding to the initial video frame; and training an initial model using the initial video frame and a control text corresponding to the initial video frame, to obtain the video prediction model, where the control text comprises the scenario context and the vehicle movement instruction, and the video prediction model is configured to output a predicted video frame.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Optical flow guided knowledge distillation video prediction model compression method and system

PendingCN122179582ASolve the problem of missing motion featuresefficient migrationBiological modelsDigital video signal modificationVisual technologyOptical flow
This invention discloses a video prediction model compression method and system based on optical flow-guided knowledge distillation, belonging to the field of computer vision technology. The method includes: generating optical flow distillation loss by constraining the optical flow prediction distribution of the student network to be consistent with that of the teacher network; generating channel alignment loss by aligning the channel attention distributions of the teacher and student networks using divergence; generating pixel-level reconstruction loss based on the difference in pixel values ​​between the generated image predicted by the student network and the real image; generating perceptual loss based on the difference in high-level semantic feature space between the generated image predicted by the student network and the real image; and obtaining the trained student network based on the optical flow distillation loss, channel alignment loss, pixel-level reconstruction loss, and perceptual loss. This invention can solve the performance degradation problem caused by neglecting spatiotemporal characteristics in existing compression methods for video prediction tasks.
Owner:PEKING UNIV

A method and system for video pre-distribution

ActiveCN116389803BSelective content distributionVideo predictionVideo based
The embodiment of the application provides a video pre-distribution method and system, which is used for making a client predict a video based on data with higher real-time performance, and improving the accuracy of the pre-distributed video. The method comprises the following steps: a client sends video playing behavior data to a server; the server generates video features based on the video playing behavior data; the server recalls a to-be-watched video list based on the video playing behavior data; the server sends the video features and the to-be-watched video list to the client; the client acquires local user behavior and video playing records; the client determines a local to-be-played list according to the video playing records; the client determines to-be-predicted video features according to the video features, the local user behavior, the to-be-watched video list and the local to-be-played list; the client determines a first to-be-cached video according to the to-be-predicted video features; and the client caches the first to-be-cached video.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Tennis self-training system

A tennis self-training system comprising a control device comprising recording unit configured to record a tennis game and a processor configured to analysis the tennis game based on a video obtained from the recording unit; and a ball machine unit configured to move and launch a ball according to the instructions of the control device; wherein the control device is configured to: determine the position of a player and the position of the ball machine unit based on the video, predict a falling position of the ball hit by the player based on the video, calculate a ball launch position and a ball arrival position of the ball machine unit based on the position of the player and the falling position of the ball, generate a control signal related to the ball launch position and the ball arrival position, transmit the control signal to the ball machine unit.
Owner:CURINGINNOS INC

An audio and video parsing method based on noise label learning

The application belongs to the technical field of deep learning, and relates to an audio and video parsing method based on noise label learning, which comprises the following steps: preprocessing original audio and video to obtain a segment-level input sequence; constructing a mutual learning noise-resistant double-flow network; training the mutual learning noise-resistant double-flow network according to a training set; comparing the validation set indicators of two sub-networks in the trained mutual learning noise-resistant double-flow network, and taking the sub-network with the larger validation set indicator as an audio and video parsing model; and parsing through the audio and video parsing model according to a test set to obtain a video prediction result. The mutual learning noise-resistant double-flow network is composed of two sub-networks with the same structure but different initializations, a cross filtering mechanism is executed according to the clean masks generated by the two sub-networks during the training of the mutual learning noise-resistant double-flow network, and the dynamic confidence ratio is gradually reduced through a cosine strategy, so that the problems of high pseudo-label noise rate and easy overfitting noise in the existing audio and video parsing task are solved.
Owner:UNIV OF ELECTRONIC SCI & TECH OF CHINA CHENGDU COLLEGE

System and method for predicting pedestrian safety information based on video

A method and apparatus for predicting pedestrian safety information based on video is disclosed. Pedestrian trajectory prediction data based on video input is first generated. Then, pedestrian behavior prediction data based on the video input is generated. A potential risk to pedestrian safety is estimated based on the pedestrian trajectory prediction data, the pedestrian behavior prediction data, and a surface classification data in the video data.
Owner:ELECTRONICS & TELECOMM RES INST

A video prediction method and system based on a dynamic diffusion model

The present application relates to computer vision technology, aiming at providing a video prediction method and system based on dynamic diffusion model. It includes: a feature extraction network formed by stacking a plurality of serial feature extraction modules, and the activation state is controlled by a routing module; each pair of input frames is split into parallel spatial feature branches and motion feature branches, and after branch separation and scale transformation processing, the fusion features with spatial details and motion correlation are output, and then the last frame is encoded and fused as the conditional information input into the diffusion model; the next moment prediction optical flow field is generated by step-by-step denoising, the position of the current pixel point in the next moment is calculated through forward warping operation, and the final prediction frame is synthesized; the prediction frame is connected with the historical video frame sequence through repeated operation, and the continuous prediction task is realized. The present application can adaptively learn the features of the current and past driving environment, and use the randomness of the diffusion model to generate the prediction optical flow, so as to achieve the purpose of video frame prediction.
Owner:ZHEJIANG UNIV

A video prediction method based on an automatic driving remote control system

The application provides a video prediction method based on an automatic driving remote control system, including the following steps: acquiring and decoding original images through camera shooting; collecting original image data acquired by the camera to generate a training set; training an MCNet model using the training set and performing data processing on the training set through the MCNet model; taking the position information of a spreader in a predicted image as an item in a loss function, adding a plurality of fully connected layers to the backbone network of the MCNet model to output the position information of the spreader, and adding a spreader regression loss function when the MCNet model is updated; copying the currently used MCNet model to obtain a mirrored MCNet model; performing small-step training using the original image of a previous historical frame as training data; performing performance verification on the original image of the second-to-last historical frame as verification data of the training data, if the accuracy of the verification data is higher, the MCNet model is updated and optimized; otherwise, the weights of the network of the MCNet model remain unchanged; and video streaming of the predicted image is generated and published.
Owner:FOCUSIGHT TECH (JIANGSU) CO LTD

Video predictive coding method and apparatus

Provided is a method for video predictive coding. The method includes: acquiring decision information related to a current coding block in inter-frame predictive coding, wherein the decision information includes one of: pre-analysis information determined by a coder in a lightweight video coding pre-analysis, coding information of a plurality of sub-blocks acquired by recursively coding the current coding block, and information of an executed mode determined based on the executed mode; and determining, based on the decision information, whether to skip motion estimation (ME) coding of the current coding block.
Owner:BIGO TECH PTE LTD

Audio and video analysis method based on noise label learning

The invention belongs to the technical field of deep learning, and relates to an audio and video analysis method based on noise label learning, which comprises the following steps: preprocessing an original audio and video to obtain a fragment-level input sequence; constructing a mutual learning anti-noise double-current network; training the mutual learning anti-noise double-flow network according to the training set; comparing the verification set indexes of the two sub-networks in the trained mutual learning anti-noise double-current network, and taking the sub-network with the verification set index greater than that of the other sub-network as an audio and video analysis model; and according to the test set, analyzing through the audio and video analysis model to obtain a video prediction result. According to the invention, through a mutual learning anti-noise double-flow network formed by two sub-networks with the same structure and different initializations, a cross filtering mechanism is executed according to clean masks generated by the two sub-networks in the training of the mutual learning anti-noise double-flow network, and a dynamic confidence proportion is gradually reduced through a cosine strategy. The problems that in an existing audio and video analysis task, the false label noise rate is high, and noise overfitting is prone to occurring are solved.
Owner:UNIV OF ELECTRONIC SCI & TECH OF CHINA CHENGDU COLLEGE

A video prediction method and apparatus for reducing inter-frame error accumulation

The application provides a video prediction method and device for reducing inter-frame error accumulation, wherein the method comprises: acquiring an image frame at a current time and m continuous image frames before the current time; processing the image frame at the current time and the m continuous image frames before the current time by using a space-time prediction model to obtain implicit states of k continuous image frames after the current time; the space-time prediction model adopts a parallel long short-term memory model with implicit state decoupling; processing the implicit states of the k continuous image frames after the current time by using a reconstruction layer to obtain the predicted k continuous image frames. The application can map the collected multiple observation videos to multiple prediction videos at one time, and avoid error accumulation caused by frame-by-frame recursive prediction.
Owner:BEIJING UNIV OF CHEM TECH

Model training method and device, video prediction method and device, equipment and medium

The invention discloses a model training method and device, a video prediction method and device, equipment and a medium, and the method comprises the steps: obtaining reasoning tasks and task data for processing different modal data, converting the task data, obtaining the task data of a video modal, dividing the task data of the video modal according to the task original data and the task result data, and obtaining a video model; and obtaining condition frame data and result frame data, associating the result frame data as a label with the condition frame data to form a sample, and training the initial video prediction model to obtain a target video prediction model. And converting to-be-predicted data of the target reasoning task to obtain to-be-predicted data of a video mode, and inputting the to-be-predicted data of the video mode into the target video prediction model to obtain a prediction result. The method can be applied to the field of financial science and technology, the multi-modal learning normal form is optimized, and the reasoning accuracy of the multi-modal reasoning task is improved while the model complexity is reduced.
Owner:PING AN TECH (SHENZHEN) CO LTD

Video prediction method and system based on dynamic diffusion model

The invention relates to a computer vision technology, and aims to provide a video prediction method and system based on a dynamic diffusion model. Comprising the following steps: stacking a plurality of serial feature extraction modules to form a feature extraction network, and controlling an activation state by using a routing module; splitting each pair of input frames into parallel spatial feature branches and motion feature branches, outputting fusion features associated with spatial details and motion after branch separation and scale transformation processing, and performing coding fusion with the last frame as conditional information to be input into a diffusion model; generating a prediction optical flow field at the next moment through step-by-step denoising, calculating the position where the current pixel point appears at the next moment through forward warping operation, and synthesizing a final prediction frame; and repeating the operation to join the prediction frame with the historical video frame sequence to realize a continuous prediction task. According to the method, the characteristics of the current and past driving environments can be adaptively learned, and the prediction optical flow is generated by using the randomness of the diffusion model, so that the purpose of video frame prediction is achieved.
Owner:ZHEJIANG UNIV