Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

613 results about "Noise (video)" patented technology

Noise, in analog video and television, is a random dot pixel pattern of static displayed when no transmission signal is obtained by the antenna receiver of television sets and other display devices. The random pattern superimposed on the picture, visible as a random flicker of "dots" or "snow", is the result of electronic noise and radiated electromagnetic noise accidentally picked up by the antenna. This effect is most commonly seen with analog TV sets or blank VHS tapes.

Panoramic video frame insertion method based on potential diffusion model

The invention discloses a panoramic video frame insertion method based on a potential diffusion model. The method comprises the following steps: compressing an input image to a potential space by using a panoramic perception vector quantization variational auto-encoder to obtain potential features; constructing initial noise, and fusing the motion features extracted by the panoramic optical flow adapter to obtain condition information; performing iterative denoising operation on the potential features of the intermediate frame to obtain final potential features of the intermediate frame; and restoring the final potential features of the intermediate frame into a frame insertion image of a pixel space through a condition decoder. According to the panoramic video frame interpolation method provided by the invention, the potential diffusion model and the panoramic characteristic enhancement technology are combined, so that the perception quality of a frame interpolation result can be remarkably improved, details and complex textures of a pole region can be better reserved, and a high-quality time interpolation solution is provided for immersive panoramic video application.
Owner:HANGZHOU DIANZI UNIV

Video generation method and related device

PCT designated stage expiredWO2025119059A1Image enhancementImage analysisConvertersNoise (video)
The present invention provides a video generation method, comprising: encoding an image comprising a reference figure to obtain a feature representation of the reference figure; encoding a video comprising a reference action to obtain a feature representation of the reference action; injecting the feature representation of the reference figure and the feature representation of the reference action into a latent diffusion network in a video latent diffusion model by means of at least one timing hold module in a trained video latent diffusion model, wherein the latent diffusion network comprises a plurality of cross-attention and denoising units, one timing hold module is connected in series between at least two cross-attention and denoising units, and each timing hold module comprises a time converter and a space converter; using the latent diffusion network to denoise an inputted noise sequence to obtain a denoised latent vector sequence; and decoding the latent vector sequence to obtain an action video of a target object. The present invention further provides a video generation apparatus, an electronic device, a storage medium, and a program product.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Video generation method and device based on action coherence, equipment and medium

The invention relates to the technical field of data analysis, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a video generation method, device, equipment and medium based on action coherence, the method comprises the following steps: obtaining a video action material set, carrying out key frame analysis on the video action material set to obtain a key video frame sequence, extracting an inter-frame residual vector between consecutive frames in the key video frame sequence, adjusting a preset initial diffusion model by using the inter-frame residual vector to obtain an optimized diffusion model, performing cosine scaling on the inter-frame residual vector by using the optimized diffusion model to obtain a scaled residual vector, and obtaining a video adjustment text, and carrying out noise addition and splicing on the key video frame sequence by using the video adjustment text and the zoom residual vector to obtain a noise video, and carrying out noise reduction on the noise video to obtain a target action video. According to the method and the device, the action coherence in the customized generated video can be effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Video generation method and apparatus, device, and medium

Embodiments of the present application provide a video generation method and apparatus, a device, and a medium. The method can be applied to the technical field of video content generation, and is used for improving the video generation quality. The method comprises: acquiring a sample video frame sequence from a sample video, determining a first step count and sample original noise, and performing data noise addition processing on the sample video frame sequence to obtain video input data; inputting the video input data, a sample text encoded feature corresponding to sample description text, and first embedding information corresponding to the first step count into an initial generation model; and performing noise prediction on the sample video frame sequence by means of M spatiotemporal residual components and M spatiotemporal attention components in the initial generation model to obtain sample predicted noise, correcting a network parameter in the initial generation model on the basis of the sample original noise and the sample predicted noise, and determining the initial generation model comprising the corrected network parameter as a video generation model.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Non-contact physiological signal extraction method and system based on frequency self-adaption and illumination noise perception

The invention relates to the technical field of biomedical engineering and computer vision, in particular to a non-contact physiological signal extraction method and system based on frequency self-adaption and illumination noise perception.The method comprises the following steps of multi-mode video stream collection and spatio-temporal data preprocessing, illumination-noise perception mask generation and feature filtering, multi-mode video stream collection and spatio-temporal data preprocessing, illumination-noise perception mask generation and feature filtering, and non-contact physiological signal extraction. Frequency adaptive gating and frequency domain feature enhancement, depth time attention feature re-calibration, physiological signal regression and closed loop optimization; the method has the beneficial effects that a lightweight end-to-end deep learning network architecture is constructed by systematically fusing three core modules of illumination-noise perception mask, frequency adaptive gating and depth time attention, and the defects that a traditional physical model depends on artificial prior and is poor in anti-interference performance and high in reliability are overcome. And the one-sidedness caused by high calculation complexity and difficulty in distinguishing the signal and noise of the existing deep learning model is avoided, and the weak physiological signal can be recovered from the face video more accurately and robustly.
Owner:CENT SOUTH UNIV

Temporally consistent and semantics guided text-based video editing generative artificial intelligence (AI) model with improved initialization

A processor-implemented method performed for text-based video editing includes receiving a video input and a text prompt. The video input includes a sequence of video frames. Features of the video input are extracted to generate a latent representation of the video input. Noise is injected to the latent representation of the video input to generate a noise injected latent. The noise is conditioned on the video input. An artificial neural network (ANN) model processes the noise injected latent based on the text prompt to adapt the video input according to the text prompt.
Owner:QUALCOMM INC

Anti-compression coding robust video watermark generation method based on adversarial neural network

The invention is suitable for the field of digital watermarking, and provides an anti-compression coding robust video watermark generation method based on an adversarial neural network, and the method comprises the steps: constructing an MSCA-GAN model; the model comprises an encoder, a decoder, a distortion layer, a discriminator network and an opponent network, and the specific steps are as follows: step S1, the encoder extracts different scale features of a video frame by using a multi-scale convolution attention mechanism, calculates attention weights, efficiently hides and embeds binary watermark information into the video, generates a watermark-containing video, and transmits the watermark-containing video to the decoder; performing confrontation optimization on an embedding strategy with a discriminator in training; according to the method, for H.264 compression layer special training, the multi-scale convolution attention mechanism and the depth separable convolution are combined, and the anti-compression robustness of the watermark under the H.264 standard is effectively improved; through common attack training such as noise layer simulation cutting and zooming, the watermark can still keep high extraction accuracy and robustness in a complex environment.
Owner:ENG UNIV OF THE CHINESE PEOPLES ARMED POLICE FORCE

Multi-modal fusion green port digital management and control system and method

The invention discloses a multi-modal fusion green port digital management and control system and method, and relates to the technical field of port management and control, and the method comprises the following steps: collecting the signal-to-noise ratio data of a radio channel of a video link, and generating a port area electromagnetic interference thermodynamic diagram through spatial interpolation; and dynamically adjusting a synchronous clock of each camera according to the electromagnetic interference thermodynamic diagram, mapping the frame triggering offset into a pixel compensation matrix, and performing sub-pixel-level space-time correction on the image frame to obtain a corrected video frame sequence. According to the method, target credibility judgment and path optimization under multi-modal perception are realized through construction of an interference thermodynamic diagram, space-time correction, artifact recognition and point cloud fusion, an interference model and a recognition threshold are dynamically regulated and controlled through closed-loop feedback, and an end-to-end self-adaptive port digital management and control method is constructed. The artifact identification accuracy and scheduling stability are significantly improved, and the anti-interference and green efficient operation capabilities of the port management and control system are enhanced.
Owner:TIANJIN RES INST FOR WATER TRANSPORT ENG M O T

Deeply-forged face video frame-level positioning method and system based on weak supervised learning

The invention discloses a deeply-forged face video frame-level positioning method and system based on weak supervised learning, and the method comprises the steps: firstly constructing a training set with a video as a unit, carrying out the data enhancement of a frame-level sample, and generating an enhanced view pair; secondly, splicing the enhanced view pair and inputting the spliced enhanced view pair into a depth forgery detection model to obtain and generate fusion enhanced frame-level features; then, intra-class contrast learning loss, time sequence consistency constraint loss and frame weight loss are constructed, and a deep forgery detection model is trained based on fusion-enhanced frame-level feature joint optimization. And finally, inputting a video to be detected into the deep counterfeiting detection model to output the frame-level confidence, judging whether the video is a forged video or not, and realizing frame-level counterfeiting positioning. According to the method, video-level detection and frame-level positioning are effectively realized, meanwhile, the influence of label noise in part of forged videos is relieved, and the generalization and robustness of the system are improved.
Owner:HANGZHOU DIANZI UNIV

AI-driven smooth video-to-video generation

Provided are systems and methods for artificial intelligence (AI)-driven smooth video-to-video generation. An example method includes receiving a first video including first frames; acquiring a text including instructions for transforming the first video; encoding the text into text embeddings corresponding to the first frames; encoding the first frames into image latents; generating initial noise vectors and adding the initial noise vectors to the image latents to obtain noisy image latents; providing the text to a pretrained motion model to generate animation parameters corresponding to the first frames; providing the noisy image latents, the text embeddings, and the animation parameters to a neural network to generate second noise vectors for the image latents; removing the second noise vectors from the noisy image latents to obtain denoised image latents; and decoding the denoised image latents into second frames of a second video.
Owner:GLAM LABS INC

Running state monitoring and fault diagnosis method for loom control system based on machine vision

The invention relates to the technical field of industrial vision and intelligent monitoring, and discloses a loom control system operation state monitoring and fault diagnosis method based on machine vision, which comprises the following steps: acquiring a video stream in a loom shed area and constructing a two-dimensional space-time slice tensor; performing global motion compensation processing on the space-time slice tensor by using a homography transformation matrix, mapping a compensated dynamic texture feature sequence to a three-dimensional phase space by using a time delay embedding algorithm, and reconstructing a closed phase space trajectory representing periodic operation logic of the loom; the discrete Frechet distance between the phase space trajectory of the current operation cycle and the preset reference trajectory is calculated, and a control instruction is generated. The health degree of the sequential logic of the system is directly quantified on the premise that specific components are not recognized by using the invariant characteristic of the phase space manifold topology; the technical problems that small phase lag is difficult to perceive and nonlinear faults cannot be early warned in a strong noise environment are solved.
Owner:HU ZHOU XIN NAN HAI ZHI ZAO CHANG

Video editing model based on common editing of text and image and construction method thereof

The invention provides a video editing model based on text and image common editing and a construction method thereof. The video editing model introduces an optical flow guide mask fusion module and a multi-modal feature recognition and segmentation module into a denoising diffusion implicit model; the method comprises the following steps: inputting an original video into a denoising diffusion implicit model to carry out forward diffusion noise addition to obtain a multi-frame submerged space generation frame, and inputting the multi-frame submerged space generation frame into an optical flow guide mask fusion module to carry out inter-frame feature alignment to obtain a time consistency submerged space generation frame; a text prompt, an image prompt and an original video are input into a multi-modal feature recognition and segmentation module to be aligned and positioned to obtain a condition vector, the condition vector and a time consistency submerged space generation frame are subjected to iterative denoising to generate a target editing video, and dynamic feature modulation is performed in each denoising process; a video editing model based on text and image common editing efficiently edits a video under combined guidance of text and image prompts.
Owner:HANGZHOU GISWAY INFORMATION TECH CO LTD

Video identification and analysis method based on physical characteristics

The invention relates to the technical field of video recognition and analysis, and discloses a video recognition and analysis method based on physical characteristics. Video frame pixels are mapped to a two-dimensional coordinate system with the upper left corner as an original point, and mirror image expansion and median filtering are carried out on a gray level image; constructing a binary image based on a gray threshold value, and analyzing and extracting a target region by using a four-neighborhood connected domain; using neighborhood search and polar angle sorting to close the tracking contour, and generating equidistant re-sampling points based on Euclidean distance and an interpolation method; the curvature of the re-sampling points is estimated through a three-point difference algorithm, and zero denominator is avoided through numerical protection; performing discrete Fourier transform on the curvature sequence to extract a frequency spectrum, and normalizing an amplitude to form a standardized feature vector; and finally, inter-frame similarity is calculated based on the feature vector, and the most similar frame is automatically retrieved. By processing unified data standards in stages, edge noise is suppressed, sampling uniformity is ensured, and feature stability and cross-frame comparability are improved.
Owner:BEIJING SIHAI TONGDA TECH CO LTD

Encrypted video identification method in Tor environment

The invention discloses an encrypted video identification method in a Tor environment. The method comprises the four steps of collecting flow, extracting an AU-burst length sequence, extracting upstream features and classifying downstream tasks. Firstly, real-time Tor traffic is captured at a network information service center of a local area network entrance, then features are extracted from the real-time Tor traffic by using a corresponding TREFS i T segmentation strategy according to Tor video traffic characteristics to obtain an encrypted ADU-burst length sequence, then noise filtering and secondary feature extraction are performed by using a 1DCNN model, and finally classification is performed by using a random forest. According to the method for identifying the encrypted video in the Tor complex environment, the encrypted video does not need to be decrypted, the video played by the client through the Tor network can be identified through the video traffic characteristics so as to supervise harmful videos, and the method has wide application scenes and good supervision effects.
Owner:SOUTHEAST UNIV

Systems and methods for motion-controllable video diffusion

Methods for motion-controllable video diffusion include extracting optical flow fields from an input video and computing warped noise by iteratively warping noise between consecutive frames using the optical flow fields. The iteratively warping includes (i) re-Gaussianizing expanded pixel regions by sampling fresh Gaussian noise, and (ii) aggregating contracted pixel regions by merging noise particles and renormalizing variance to preserve spatial Gaussianity. An output video is generated by initializing a diffusion process with the warped noise and iteratively denoising to produce temporally coherent output frames. Various other methods, systems, and computer-readable media are also disclosed.
Owner:NETFLIX INC

Cross-modal semantic attention collaborative enhancement video subtitle generation method and system

The invention provides a cross-modal semantic attention collaborative enhancement video subtitle generation method and system, and belongs to the field of video subtitle generation. The method aims at solving the problems that an existing video description model mostly stays in first-order relation modeling on the attention mechanism level, high-order semantic dependence is difficult to capture, and noise is easy to introduce in multi-modal feature fusion. The invention provides a cross-modal semantic attention collaborative enhancement module, the module comprises two key components of context semantic modulation of attention enhancement and cross-modal structure alignment, and the fine modeling ability of the generative model for visual and text semantics is effectively improved by dynamically modulating attention weight and optimizing a modal alignment structure. And integration is carried out based on a non-autoregression coarse-to-fine video description model. Experimental results show that the method can significantly improve the accuracy and diversity of video description generation on the premise of keeping the model scale and the calculation overhead basically unchanged.
Owner:HARBIN ENG UNIV

Real-time drowning detection method and system based on multi-module integration

The invention discloses a real-time drowning detection method and system based on multi-module integration, and belongs to the technical field of computer vision and artificial intelligence. Aiming at the problems that a traditional drowning detection method is sensitive in environment interference, insufficient in real-time performance, low in small target detection precision and the like, an innovative scheme is provided through multi-technology collaborative optimization: a lightweight backbone network is constructed by adopting DWConv, so that the calculation complexity is reduced; a dynamic receptive field attention module is designed, and noise interference is suppressed by means of RFAConv; an SEAM module is fused to enhance sheltered target feature response; biFPN is integrated, and the robustness of long-distance and small target detection is improved. Experiments show that the mAP at 0.5 of the improved DRSB-YOLO model reaches 86.253%, the model volume is 4.39 MB, the reasoning speed of a Windows platform is 56.1 + / -5.1 FPS, and the real-time processing of 1080P video streams is supported. The system supports cross-scene adaptive switching and local / cloud dual-mode data transmission, can be widely applied to the fields of intelligent lifesaving and the like, and improves the accuracy and timeliness of drowning early warning.
Owner:ROCKET FORCE UNIV OF ENG

Video user data security analysis and prediction method based on privacy calculation

The invention discloses a video user data security analysis and prediction method based on privacy calculation, and the method comprises the steps: S1, collecting video user data, and constructing a data set; s2, performing privacy budget distribution on different processing stages of the video user data through multi-level privacy budget division; s3, defining an adjustment formula of a noise injection amount by adopting a self-adaptive noise injection mechanism, and obtaining preprocessed video user data; s4, constructing a privacy calculation framework, and combining the balance objective function and the graph reconstruction loss function to set a loss function; s5, training a privacy computing framework by using the preprocessed video user data; and S6, executing video user data analysis and prediction through a privacy computing framework, and outputting a prediction result. According to the method, an efficient and scientific optimization scheme can be provided in video user data security analysis and prediction, and remarkable technical values and economic benefits are brought to practical application.
Owner:XIAOYUAN PERCEPTION (HULUDAO) TECH CO LTD

Video generation method and device, medium, equipment and computer program product

The invention discloses a video generation method and device, a medium, equipment and a computer program product. The method comprises the following steps: acquiring a control image, a depth map corresponding to the control image and target camera track information used for video generation; and according to the control image, the depth map, the target camera track information and a video generation model, performing de-noising processing on noise image features, and determining a target video corresponding to the control image and target video depth information corresponding to the target video. Therefore, in the process of performing video generation through the image, video generation can be performed in combination with the camera track information and the depth information for controlling the image, so that the target video and the video depth information thereof under the control of the camera track information are obtained, and a 3D scene can be obtained through direct rendering based on the target video and the video depth information. Therefore, the matching degree between the target video and the camera track information can be improved, and the 3D scene can be quickly obtained without complex processes and additional operations.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Video generation method and device, equipment, medium and program product

The invention discloses a video generation method and device, equipment, a medium and a program product, and relates to the field of artificial intelligence. The method comprises the following steps: denoising noise feature representation based on a first text, an object movement track and a camera attitude parameter to obtain video feature representation; and performing video prediction based on the video feature representation to obtain a first video. In addition to the first text, additionally acquiring an object movement track to determine video object movement in the generated first video, namely controlling local movement in the generated first video through the object movement track; in the embodiment of the invention, the camera attitude parameter is additionally acquired to control the camera motion in the determined first video, that is, the global motion in the generated first video is controlled through the camera attitude parameter, so that local motion control and global motion control of the generated first video are realized, and the motion diversity in the first video is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Video content enhancement method for low-light environment

The invention provides a video content enhancement method for a low-illumination environment, and the method comprises the steps: achieving the data preprocessing based on an original low-illumination video frame sequence through frame synchronization, color space conversion and local brightness analysis, generating a noise sensitivity thermodynamic diagram through multi-feature unsupervised learning, and constructing a noise perception gating mechanism through the combination of affine transformation. Dynamic modulation of the characteristic channel is realized; in the multi-scale network structure, a channel attention module is used for carrying out layer-by-layer self-adaptive adjustment on a noise sensitive area; a basic illumination image and an edge enhancement image are generated through double-branch decoding, and then weighted fusion is carried out in combination with a noise thermodynamic diagram, so that brightness balance and detail enhancement are realized; a noise smoothing regular term is introduced during end-to-end training, so that the network achieves dynamic balance between an enhancement effect and noise control.
Owner:GUANGZHOU CHENXI NETWORK TECH CO LTD

Target-oriented video semantic communication system based on visual model

The invention provides a target-oriented video semantic communication system based on a visual model, and the system comprises a semantic extractor which is based on an SAM2 model and is used for processing an original video, generating a segmentation mask and extracting semantic information; the ViMama encoder is used for carrying out channel encoding on the output of the semantic extractor; the channel adaptation module is used for optimizing a coding sequence ViMama decoder based on the signal-to-noise ratio information of a physical channel, and is used for carrying out channel decoding to obtain a feature sequence; and the semantic reconstruction device is used for performing semantic reconstruction based on the feature sequence, recovering data and outputting a target video. According to the method, the problems of large redundant semantic information interference, insufficient deep semantic coding capability and poor communication robustness in a complex channel environment in a video are solved, efficient compression and robust transmission of video data are realized on the premise of ensuring semantic integrity, and the overall performance and adaptability of semantic communication are greatly improved.
Owner:湖南工商大学

Video restoration method and device, equipment and storage medium

The invention provides a video restoration method and device, equipment and a storage medium. The method comprises the following steps: acquiring an original video frame sequence and a mask; performing frame-level compression on the original video frame sequence, and mapping the original video frame sequence into a compact potential space representation; generating a description text related to the scene according to the original video frame sequence; fusing the noise of each time step and the code of the description text; fusing the codes of the potential space representation and the description text; and generating a repaired video frame sequence according to the mask and the fusion result. According to the method of the invention, the scene-related description text generated through the original video frame sequence can ensure the naturalness and coordination of the restoration area, and at the same time, the noise of each time step and the coding of the description text are fused, and the potential space representation and the coding of the description text are fused. The repair area obtained according to the mask and the two fusion results is more natural and harmonious.
Owner:SHANGHAI ZHIXIANG FUTURE COMPUTER TECHNOLOGY CO LTD

Camera track length-controllable video generation method and system based on video diffusion model

The invention discloses a method and a system for generating a long video with a controllable camera track based on a video diffusion model. Comprising a camera track and initial frame preparation stage, a point cloud construction and multi-view image generation stage, a scale factor alignment optimization stage, a camera motion prior injection and noise initialization stage, a diffusion inversion generation stage and a sliding window time consistency fusion stage. Three-dimensional camera track modeling, projection reconstruction and diffusion processes are combined, a video content generation process is explicitly guided to be aligned with a track path set by a user, and long video generation which is reasonable in structure, natural in vision and continuous in time is achieved.
Owner:ZHEJIANG UNIV

Video generation method and related equipment

The invention provides a video generation method. The video generation method comprises the following steps: encoding an image containing a reference image to obtain a feature representation of the reference image; coding the video containing the reference action to obtain the feature representation of the reference action; injecting the feature representation of the reference image and the feature representation of the reference action into an implicit diffusion network in the video implicit diffusion model through at least one time sequence maintaining module in the trained video implicit diffusion model; wherein the implicit diffusion network comprises a plurality of cross attention and denoising units; a time sequence maintaining module is connected in series between the at least two cross attention and denoising units; the time sequence holding module comprises a time converter and a space converter; performing denoising processing on an input noise sequence by using the implicit diffusion network to obtain a denoised implicit vector sequence; and decoding the implicit vector sequence to obtain an action video of the target object. The invention further provides a video generation device, electronic equipment, a storage medium and a program product.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Method of detecting video anomaly on basis of multimodal diffusion and device therefor

Proposed are a method of detecting a video anomaly on the basis of multimodal diffusion, and the method includes a step of obtaining video data including a plurality of frames, a step of detecting an object included in each of the plurality of frames, a step of extracting a multimodal feature vector including a visual feature vector, a text feature vector, and a motion feature vector for the detected object, a step of generating a noise vector by injecting noise into the visual feature vector, a step of generating a restoration vector with the noise removed by inputting the noise vector into a diffusion model and by using the text feature vector and the motion feature vector as conditions, and a step of performing anomaly detection on the video data by comparing the visual feature vector and the restoration vector.
Owner:UI (UNIVERSITY IND FOUNDATION) YONSEI UNIVERSITY

Video super-resolution method and device for complex operation scene based on space-time consistency and medium

The invention discloses a time-space consistency-based video super-resolution method and device for a complex operation scene, and a medium, belongs to the field of image and video processing, and aims at solving the problem that a stable, clear and continuous-structure video sequence is difficult to generate in an existing method, and the time-space consistency-based video super-resolution method for the complex operation scene based on video prior is provided. Comprising the following steps: utilizing submerged space modeling, mapping an input low-resolution video to a submerged space, and obtaining a corresponding submerged variable representation; initializing a random noise tensor with the same size as the latent variable expression as an initial noise state; in each step of sampling, dividing a noise state and latent variable representation into a plurality of blocks according to space and time dimensions; denoising is carried out on each tile block; the initial noise is converted into high-quality latent variable representation; and a high-resolution video is reconstructed. According to the method, the power inspection video with high resolution and high consistency can be generated, the generation stability is kept, and the video texture detail expression capability is improved.
Owner:ELECTRIC POWER RES INST OF STATE GRID ZHEJIANG ELECTRIC POWER COMAPNY

Dynamic scene re-operation mirror video generation method and system based on diffusion model

The invention discloses a dynamic scene re-operation mirror video generation method and system based on a diffusion model, and belongs to the technical field of computer vision and video generation. A diffusion generation architecture with a control branch is adopted, and the core is composed of an embedded layer, a main branch and a control branch. In the control branch, the output of each sub-block is added with the output of the corresponding block of the main branch after being processed by the zero initial linear layer, and the sum is input into the next block of the main branch. During training, generating a rendered video by using the target video and the reference video in the same scene; and inputting the target video latent variable after noise addition into the control branch, inputting the splicing result of the target video, the reference video and the rendering video latent variable into the main branch, and simultaneously providing the text latent variable of the reference video for the two branches as a condition. During generation, the model finally generates a target video latent variable through step-by-step denoising and decodes the target video latent variable into a target track video, and it is ensured that the motion of a moving object in a scene of a generated video and a reference video is consistent at the same time.
Owner:ZHEJIANG UNIV

Intelligent image selecting and cutting system fusing visual features and quality scores

The invention relates to the technical field of industrial visual intelligence, in particular to an intelligent image selection and switching system fusing visual features and quality scores, which comprises the following steps: receiving multiple paths of video coding streams, inter-frame motion vectors and camera parameters; generating a macro block activeness distribution map based on the video coding stream and the inter-frame motion vector, and performing local window positioning and feature reconstruction on the video coding stream to generate an enhanced video vector; calculating a confidence coefficient mean value and a consistency score of the video coding stream according to the enhanced video vector, and fusing the confidence coefficient mean value and the consistency score with a channel transmission signal-to-noise ratio and a quantization noise increment to generate a video quality score; constructing a multi-criterion optimization model, and setting a feature representation vector for the multi-criterion optimization model; mapping the viewpoint weight based on the inner product of the feature representation vector, and calculating with the video quality score to generate a video switching score; and performing priority ranking based on the video switching score, and triggering a mapping switching instruction. And realizing video image scheduling by fusing the visual features and the quality score.
Owner:XINAOTE (NANJING) VIDEO TECH CO LTD