Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

7 results about "Video prediction" patented technology

Video code rate determination method and apparatus, electronic device, medium, and product

PendingCN122457831AVideo encodingSimulation
The embodiments of the present disclosure disclose a video code rate determination method, device, electronic equipment, storage medium and product. The method comprises: in response to triggering of a video transcoding event, determining a target transcoded video; predicting a predicted playing probability of the target transcoded video on a terminal device corresponding to each device performance level; determining a candidate video transcoding code rate combination according to a preset video code rate level; determining a comprehensive playing performance analysis result of a video transcoding result corresponding to each candidate video transcoding code rate combination on the terminal device corresponding to each device performance level; performing optimal comprehensive playing performance analysis on the predicted playing probability and the comprehensive playing performance analysis result of the terminal device corresponding to each device performance level; and determining a target video transcoding code rate combination according to the optimal comprehensive playing performance analysis result. The technical scheme of the embodiments of the present disclosure can determine a coding code rate that maximizes the business value of video coding.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

A variable frame rate video generation method based on optical flow estimation

ActiveCN116708869Bquality improvementEncoder decoderVariable frame rate
A variable frame rate video generation method based on optical flow estimation, which introduces optical flow supervision information into an OpFode-Net model, the OpFode-Net model comprising an encoder-decoder structure; the encoder uses an ODE-ConvGRU to embed input video sequence X T into a hidden state h T ; wherein the ODE-ConvGRU uses a ConvGRU as a node of a neural ODE and embeds it into the neural ODE to realize dynamic modeling of the video sequence; the decoder starts from h T , and uses an ODE solver to generate a new video frame at any time step S, which can realize more accurate prediction results and achieve optimal performance in video interpolation and video prediction tasks.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Optical flow guided knowledge distillation video prediction model compression method and system

PendingCN122179582ASolve the problem of missing motion featuresefficient migrationBiological modelsDigital video signal modificationVisual technologyOptical flow
This invention discloses a video prediction model compression method and system based on optical flow-guided knowledge distillation, belonging to the field of computer vision technology. The method includes: generating optical flow distillation loss by constraining the optical flow prediction distribution of the student network to be consistent with that of the teacher network; generating channel alignment loss by aligning the channel attention distributions of the teacher and student networks using divergence; generating pixel-level reconstruction loss based on the difference in pixel values ​​between the generated image predicted by the student network and the real image; generating perceptual loss based on the difference in high-level semantic feature space between the generated image predicted by the student network and the real image; and obtaining the trained student network based on the optical flow distillation loss, channel alignment loss, pixel-level reconstruction loss, and perceptual loss. This invention can solve the performance degradation problem caused by neglecting spatiotemporal characteristics in existing compression methods for video prediction tasks.
Owner:PEKING UNIV

An audio and video parsing method based on noise label learning

ActiveCN121682445BVideo data clustering/classificationSpeech analysisNoise (video)Noise
The application belongs to the technical field of deep learning, and relates to an audio and video parsing method based on noise label learning, which comprises the following steps: preprocessing original audio and video to obtain a segment-level input sequence; constructing a mutual learning noise-resistant double-flow network; training the mutual learning noise-resistant double-flow network according to a training set; comparing the validation set indicators of two sub-networks in the trained mutual learning noise-resistant double-flow network, and taking the sub-network with the larger validation set indicator as an audio and video parsing model; and parsing through the audio and video parsing model according to a test set to obtain a video prediction result. The mutual learning noise-resistant double-flow network is composed of two sub-networks with the same structure but different initializations, a cross filtering mechanism is executed according to the clean masks generated by the two sub-networks during the training of the mutual learning noise-resistant double-flow network, and the dynamic confidence ratio is gradually reduced through a cosine strategy, so that the problems of high pseudo-label noise rate and easy overfitting noise in the existing audio and video parsing task are solved.
Owner:UNIV OF ELECTRONIC SCI & TECH OF CHINA CHENGDU COLLEGE

System and method for predicting pedestrian safety information based on video

A method and apparatus for predicting pedestrian safety information based on video is disclosed. Pedestrian trajectory prediction data based on video input is first generated. Then, pedestrian behavior prediction data based on the video input is generated. A potential risk to pedestrian safety is estimated based on the pedestrian trajectory prediction data, the pedestrian behavior prediction data, and a surface classification data in the video data.
Owner:ELECTRONICS & TELECOMM RES INST

A video prediction method and system based on a dynamic diffusion model

The present application relates to computer vision technology, aiming at providing a video prediction method and system based on dynamic diffusion model. It includes: a feature extraction network formed by stacking a plurality of serial feature extraction modules, and the activation state is controlled by a routing module; each pair of input frames is split into parallel spatial feature branches and motion feature branches, and after branch separation and scale transformation processing, the fusion features with spatial details and motion correlation are output, and then the last frame is encoded and fused as the conditional information input into the diffusion model; the next moment prediction optical flow field is generated by step-by-step denoising, the position of the current pixel point in the next moment is calculated through forward warping operation, and the final prediction frame is synthesized; the prediction frame is connected with the historical video frame sequence through repeated operation, and the continuous prediction task is realized. The present application can adaptively learn the features of the current and past driving environment, and use the randomness of the diffusion model to generate the prediction optical flow, so as to achieve the purpose of video frame prediction.
Owner:ZHEJIANG UNIV

A video prediction method based on an automatic driving remote control system

The application provides a video prediction method based on an automatic driving remote control system, including the following steps: acquiring and decoding original images through camera shooting; collecting original image data acquired by the camera to generate a training set; training an MCNet model using the training set and performing data processing on the training set through the MCNet model; taking the position information of a spreader in a predicted image as an item in a loss function, adding a plurality of fully connected layers to the backbone network of the MCNet model to output the position information of the spreader, and adding a spreader regression loss function when the MCNet model is updated; copying the currently used MCNet model to obtain a mirrored MCNet model; performing small-step training using the original image of a previous historical frame as training data; performing performance verification on the original image of the second-to-last historical frame as verification data of the training data, if the accuracy of the verification data is higher, the MCNet model is updated and optimized; otherwise, the weights of the network of the MCNet model remain unchanged; and video streaming of the predicted image is generated and published.
Owner:FOCUSIGHT TECH (JIANGSU) CO LTD