Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

427 results about "Noise (video)" patented technology

Noise, in analog video and television, is a random dot pixel pattern of static displayed when no transmission signal is obtained by the antenna receiver of television sets and other display devices. The random pattern superimposed on the picture, visible as a random flicker of "dots" or "snow", is the result of electronic noise and radiated electromagnetic noise accidentally picked up by the antenna. This effect is most commonly seen with analog TV sets or blank VHS tapes.

Non-contact physiological signal extraction method and system based on frequency self-adaption and illumination noise perception

The invention relates to the technical field of biomedical engineering and computer vision, in particular to a non-contact physiological signal extraction method and system based on frequency self-adaption and illumination noise perception.The method comprises the following steps of multi-mode video stream collection and spatio-temporal data preprocessing, illumination-noise perception mask generation and feature filtering, multi-mode video stream collection and spatio-temporal data preprocessing, illumination-noise perception mask generation and feature filtering, and non-contact physiological signal extraction. Frequency adaptive gating and frequency domain feature enhancement, depth time attention feature re-calibration, physiological signal regression and closed loop optimization; the method has the beneficial effects that a lightweight end-to-end deep learning network architecture is constructed by systematically fusing three core modules of illumination-noise perception mask, frequency adaptive gating and depth time attention, and the defects that a traditional physical model depends on artificial prior and is poor in anti-interference performance and high in reliability are overcome. And the one-sidedness caused by high calculation complexity and difficulty in distinguishing the signal and noise of the existing deep learning model is avoided, and the weak physiological signal can be recovered from the face video more accurately and robustly.
Owner:CENT SOUTH UNIV

Anti-compression coding robust video watermark generation method based on adversarial neural network

The invention is suitable for the field of digital watermarking, and provides an anti-compression coding robust video watermark generation method based on an adversarial neural network, and the method comprises the steps: constructing an MSCA-GAN model; the model comprises an encoder, a decoder, a distortion layer, a discriminator network and an opponent network, and the specific steps are as follows: step S1, the encoder extracts different scale features of a video frame by using a multi-scale convolution attention mechanism, calculates attention weights, efficiently hides and embeds binary watermark information into the video, generates a watermark-containing video, and transmits the watermark-containing video to the decoder; performing confrontation optimization on an embedding strategy with a discriminator in training; according to the method, for H.264 compression layer special training, the multi-scale convolution attention mechanism and the depth separable convolution are combined, and the anti-compression robustness of the watermark under the H.264 standard is effectively improved; through common attack training such as noise layer simulation cutting and zooming, the watermark can still keep high extraction accuracy and robustness in a complex environment.
Owner:ENG UNIV OF THE CHINESE PEOPLES ARMED POLICE FORCE

Deeply-forged face video frame-level positioning method and system based on weak supervised learning

The invention discloses a deeply-forged face video frame-level positioning method and system based on weak supervised learning, and the method comprises the steps: firstly constructing a training set with a video as a unit, carrying out the data enhancement of a frame-level sample, and generating an enhanced view pair; secondly, splicing the enhanced view pair and inputting the spliced enhanced view pair into a depth forgery detection model to obtain and generate fusion enhanced frame-level features; then, intra-class contrast learning loss, time sequence consistency constraint loss and frame weight loss are constructed, and a deep forgery detection model is trained based on fusion-enhanced frame-level feature joint optimization. And finally, inputting a video to be detected into the deep counterfeiting detection model to output the frame-level confidence, judging whether the video is a forged video or not, and realizing frame-level counterfeiting positioning. According to the method, video-level detection and frame-level positioning are effectively realized, meanwhile, the influence of label noise in part of forged videos is relieved, and the generalization and robustness of the system are improved.
Owner:HANGZHOU DIANZI UNIV

Running state monitoring and fault diagnosis method for loom control system based on machine vision

The invention relates to the technical field of industrial vision and intelligent monitoring, and discloses a loom control system operation state monitoring and fault diagnosis method based on machine vision, which comprises the following steps: acquiring a video stream in a loom shed area and constructing a two-dimensional space-time slice tensor; performing global motion compensation processing on the space-time slice tensor by using a homography transformation matrix, mapping a compensated dynamic texture feature sequence to a three-dimensional phase space by using a time delay embedding algorithm, and reconstructing a closed phase space trajectory representing periodic operation logic of the loom; the discrete Frechet distance between the phase space trajectory of the current operation cycle and the preset reference trajectory is calculated, and a control instruction is generated. The health degree of the sequential logic of the system is directly quantified on the premise that specific components are not recognized by using the invariant characteristic of the phase space manifold topology; the technical problems that small phase lag is difficult to perceive and nonlinear faults cannot be early warned in a strong noise environment are solved.
Owner:HU ZHOU XIN NAN HAI ZHI ZAO CHANG

Video identification and analysis method based on physical characteristics

The invention relates to the technical field of video recognition and analysis, and discloses a video recognition and analysis method based on physical characteristics. Video frame pixels are mapped to a two-dimensional coordinate system with the upper left corner as an original point, and mirror image expansion and median filtering are carried out on a gray level image; constructing a binary image based on a gray threshold value, and analyzing and extracting a target region by using a four-neighborhood connected domain; using neighborhood search and polar angle sorting to close the tracking contour, and generating equidistant re-sampling points based on Euclidean distance and an interpolation method; the curvature of the re-sampling points is estimated through a three-point difference algorithm, and zero denominator is avoided through numerical protection; performing discrete Fourier transform on the curvature sequence to extract a frequency spectrum, and normalizing an amplitude to form a standardized feature vector; and finally, inter-frame similarity is calculated based on the feature vector, and the most similar frame is automatically retrieved. By processing unified data standards in stages, edge noise is suppressed, sampling uniformity is ensured, and feature stability and cross-frame comparability are improved.
Owner:BEIJING SIHAI TONGDA TECH CO LTD

Systems and methods for motion-controllable video diffusion

Methods for motion-controllable video diffusion include extracting optical flow fields from an input video and computing warped noise by iteratively warping noise between consecutive frames using the optical flow fields. The iteratively warping includes (i) re-Gaussianizing expanded pixel regions by sampling fresh Gaussian noise, and (ii) aggregating contracted pixel regions by merging noise particles and renormalizing variance to preserve spatial Gaussianity. An output video is generated by initializing a diffusion process with the warped noise and iteratively denoising to produce temporally coherent output frames. Various other methods, systems, and computer-readable media are also disclosed.
Owner:NETFLIX INC

Cross-modal semantic attention collaborative enhancement video subtitle generation method and system

The invention provides a cross-modal semantic attention collaborative enhancement video subtitle generation method and system, and belongs to the field of video subtitle generation. The method aims at solving the problems that an existing video description model mostly stays in first-order relation modeling on the attention mechanism level, high-order semantic dependence is difficult to capture, and noise is easy to introduce in multi-modal feature fusion. The invention provides a cross-modal semantic attention collaborative enhancement module, the module comprises two key components of context semantic modulation of attention enhancement and cross-modal structure alignment, and the fine modeling ability of the generative model for visual and text semantics is effectively improved by dynamically modulating attention weight and optimizing a modal alignment structure. And integration is carried out based on a non-autoregression coarse-to-fine video description model. Experimental results show that the method can significantly improve the accuracy and diversity of video description generation on the premise of keeping the model scale and the calculation overhead basically unchanged.
Owner:HARBIN ENG UNIV

Video content enhancement method for low-light environment

The invention provides a video content enhancement method for a low-illumination environment, and the method comprises the steps: achieving the data preprocessing based on an original low-illumination video frame sequence through frame synchronization, color space conversion and local brightness analysis, generating a noise sensitivity thermodynamic diagram through multi-feature unsupervised learning, and constructing a noise perception gating mechanism through the combination of affine transformation. Dynamic modulation of the characteristic channel is realized; in the multi-scale network structure, a channel attention module is used for carrying out layer-by-layer self-adaptive adjustment on a noise sensitive area; a basic illumination image and an edge enhancement image are generated through double-branch decoding, and then weighted fusion is carried out in combination with a noise thermodynamic diagram, so that brightness balance and detail enhancement are realized; a noise smoothing regular term is introduced during end-to-end training, so that the network achieves dynamic balance between an enhancement effect and noise control.
Owner:GUANGZHOU CHENXI NETWORK TECH CO LTD

Camera track length-controllable video generation method and system based on video diffusion model

The invention discloses a method and a system for generating a long video with a controllable camera track based on a video diffusion model. Comprising a camera track and initial frame preparation stage, a point cloud construction and multi-view image generation stage, a scale factor alignment optimization stage, a camera motion prior injection and noise initialization stage, a diffusion inversion generation stage and a sliding window time consistency fusion stage. Three-dimensional camera track modeling, projection reconstruction and diffusion processes are combined, a video content generation process is explicitly guided to be aligned with a track path set by a user, and long video generation which is reasonable in structure, natural in vision and continuous in time is achieved.
Owner:ZHEJIANG UNIV

Method of detecting video anomaly on basis of multimodal diffusion and device therefor

Proposed are a method of detecting a video anomaly on the basis of multimodal diffusion, and the method includes a step of obtaining video data including a plurality of frames, a step of detecting an object included in each of the plurality of frames, a step of extracting a multimodal feature vector including a visual feature vector, a text feature vector, and a motion feature vector for the detected object, a step of generating a noise vector by injecting noise into the visual feature vector, a step of generating a restoration vector with the noise removed by inputting the noise vector into a diffusion model and by using the text feature vector and the motion feature vector as conditions, and a step of performing anomaly detection on the video data by comparing the visual feature vector and the restoration vector.
Owner:UI (UNIVERSITY IND FOUNDATION) YONSEI UNIVERSITY

Video super-resolution method and device for complex operation scene based on space-time consistency and medium

The invention discloses a time-space consistency-based video super-resolution method and device for a complex operation scene, and a medium, belongs to the field of image and video processing, and aims at solving the problem that a stable, clear and continuous-structure video sequence is difficult to generate in an existing method, and the time-space consistency-based video super-resolution method for the complex operation scene based on video prior is provided. Comprising the following steps: utilizing submerged space modeling, mapping an input low-resolution video to a submerged space, and obtaining a corresponding submerged variable representation; initializing a random noise tensor with the same size as the latent variable expression as an initial noise state; in each step of sampling, dividing a noise state and latent variable representation into a plurality of blocks according to space and time dimensions; denoising is carried out on each tile block; the initial noise is converted into high-quality latent variable representation; and a high-resolution video is reconstructed. According to the method, the power inspection video with high resolution and high consistency can be generated, the generation stability is kept, and the video texture detail expression capability is improved.
Owner:ELECTRIC POWER RES INST OF STATE GRID ZHEJIANG ELECTRIC POWER COMAPNY

Dynamic scene re-operation mirror video generation method and system based on diffusion model

The invention discloses a dynamic scene re-operation mirror video generation method and system based on a diffusion model, and belongs to the technical field of computer vision and video generation. A diffusion generation architecture with a control branch is adopted, and the core is composed of an embedded layer, a main branch and a control branch. In the control branch, the output of each sub-block is added with the output of the corresponding block of the main branch after being processed by the zero initial linear layer, and the sum is input into the next block of the main branch. During training, generating a rendered video by using the target video and the reference video in the same scene; and inputting the target video latent variable after noise addition into the control branch, inputting the splicing result of the target video, the reference video and the rendering video latent variable into the main branch, and simultaneously providing the text latent variable of the reference video for the two branches as a condition. During generation, the model finally generates a target video latent variable through step-by-step denoising and decodes the target video latent variable into a target track video, and it is ensured that the motion of a moving object in a scene of a generated video and a reference video is consistent at the same time.
Owner:ZHEJIANG UNIV

Intelligent image selecting and cutting system fusing visual features and quality scores

The invention relates to the technical field of industrial visual intelligence, in particular to an intelligent image selection and switching system fusing visual features and quality scores, which comprises the following steps: receiving multiple paths of video coding streams, inter-frame motion vectors and camera parameters; generating a macro block activeness distribution map based on the video coding stream and the inter-frame motion vector, and performing local window positioning and feature reconstruction on the video coding stream to generate an enhanced video vector; calculating a confidence coefficient mean value and a consistency score of the video coding stream according to the enhanced video vector, and fusing the confidence coefficient mean value and the consistency score with a channel transmission signal-to-noise ratio and a quantization noise increment to generate a video quality score; constructing a multi-criterion optimization model, and setting a feature representation vector for the multi-criterion optimization model; mapping the viewpoint weight based on the inner product of the feature representation vector, and calculating with the video quality score to generate a video switching score; and performing priority ranking based on the video switching score, and triggering a mapping switching instruction. And realizing video image scheduling by fusing the visual features and the quality score.
Owner:XINAOTE (NANJING) VIDEO TECH CO LTD

Video target detection method and system based on multi-scale perception diffusion

The invention provides a video target detection method and system based on multi-scale perception diffusion, and relates to the field of target detection. The method comprises the steps of obtaining a plurality of frame images of a to-be-detected video; inputting the plurality of frame images into a trained target detection model, wherein the target detection model comprises a diffusion query module and a multi-scale sensing module which are parallel to each other; a diffusion query module takes the frame feature and the noise frame as input to generate an initial query feature, and gradual refinement is carried out on a diffusion time step length through an iterative optimizer to obtain a diffusion query feature related to semantics; a multi-scale sensing module extracts scale coding features based on a scale branch and an attention branch which are arranged in parallel; decoding the diffusion query feature and the scale coding feature based on a space-time Transform decoder to obtain a decoding feature of the frame image; and identifying the decoding features to obtain a target detection result of the frame image. And through multi-module collaborative design, the accuracy and robustness of video target detection are effectively improved.
Owner:QINGDAO UNIV OF SCI & TECH

Methods and apparatus for frame denoising

Systems, apparatus, and methods for post-processing video e.g. frame denoising. Noise reduction techniques may be employed to improve the quality of digital video. Frames may be extracted from a video. Synthetic frames may be created using motion data between the extracted frames. Synthetic frames may be masked to exclude pixels from the composite frame. Thresholds used in masking may vary based on the temporal distance of the extracted frame used to create the synthetic frame and the extracted frame. Masking may be based on frame differences between extracted and synthetic frames (e.g., sub-pixel / luminance differences), areas of lower quality motion data (e.g., occlusions), or edge detection in the extracted frames. Synthetic and extracted frames may be composited generating frames having less noise. The composited frame may be based on averaging pixel values across the synthetic and extracted frames. Composited frames may be compiled and encoded into denoised video.
Owner:GOPRO INC

Video slow-action frame insertion playback system based on generative adversarial network

The invention relates to the technical field of image communication, in particular to a video slow-action frame insertion playback system based on a generative adversarial network, which comprises a video signal acquisition module, a signal processing module, a frame rate conversion module and a playback code output module. The acquisition module outputs an original video frame sequence through an annular buffer; the processing module separates the denoised detail components and the edge gradient features in parallel; the frame rate conversion module uses a geometric structure as a boundary locking condition to guide detail components to execute nonlinear motion compensation and intermediate frame reconstruction; and the playback module executes time base remapping and distributed coding transmission. According to the invention, through a depth generation architecture of structural constraint textures, in combination with space-time consistency verification and a nearest neighbor pixel backfilling mechanism and a virtual time base rate decoupling technology, real-time slow-action video redisk with high signal-to-noise ratio and no artifacts for a high-speed moving target is realized.
Owner:XINAOTE (NANJING) VIDEO TECH CO LTD

Screen content video quality evaluation method and device based on frequency-space complementation and semantics

The invention discloses a screen content video quality evaluation method and device based on frequency-space complementation and semanteme, and relates to the field of computer vision, and the method comprises the steps: S1, extracting a video block and a key frame of a screen content video, and inputting the key frame into a high-frequency structure texture information extraction branch to obtain high-frequency structure texture information; s2, inputting the key frame into a noise sensing module to obtain a noise sensing feature; s3, inputting the noise perception features into a self-adaptive time sequence embedding module to obtain noise and semantic information; s4, splicing the high-frequency structure texture information and the noise and semantic information, performing quality regression to obtain a quality score of a single-frame key frame, and summing and averaging to obtain a spatial domain video quality score; s5, inputting the video blocks into a Fast-VQA-based quality evaluation branch to obtain a time sequence distortion perception degradation score; and S6, dynamically fusing the spatial domain video quality score and the time sequence distortion perception degradation score to obtain a final video quality score. According to the method provided by the invention, the screen content video quality is effectively evaluated.
Owner:XIAMEN UNIV OF TECH +1

Video generation method and apparatus, and electronic device, computer-readable storage medium and computer program product

A video generation method and apparatus, and an electronic device, a computer-readable storage medium and a computer program product. The method comprises: acquiring a target trajectory map, and mapping the target trajectory map on the basis of a target rank decomposition matrix configured for a motion control network, so as to obtain a target motion feature, wherein the target rank decomposition matrix is used for controlling the target motion feature to constrain a motion trajectory of a video lens and / or a motion trajectory of a video display object during a video generation process; acquiring target reference content required during the video generation process, and extracting a first content feature of the target reference content; and acquiring a target noise image stack, and by means of a diffusion model, using the target motion feature and the first content feature as constraint conditions to perform denoising processing on the target noise image stack, so as to obtain a target video in which the motion trajectory of the video lens and / or the motion trajectory of the video display object is constrained.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Ultrahigh-definition video noise reduction method

The invention relates to the technical field of video noise reduction, in particular to an ultra-high-definition video noise reduction method. The method comprises the following steps: acquiring video data and noise variance; obtaining a gray threshold and a gradient threshold based on the gray value and the gradient value of each frame of image pixel point of the video data, obtaining the region of each frame of image, and then determining a growth criterion through the gray difference and the gradient difference to block the image; determining a gray feature weight and a texture feature weight based on the difference between the gray value and the gray limit value and the difference between the gradient value and the gradient limit value; improving the original DCT coefficient based on the feature weight and the texture feature weight to construct an exclusive dictionary, and denoising, decoding and reconstructing the image based on the exclusive dictionary; and synthesizing each frame of reconstructed image into a complete video. According to the invention, optimal noise reduction of the video image frame is realized and detail information of the video is reserved to the greatest extent.
Owner:GUANGDONG TUSHENG ULTRA HD INNOVATION CENT CO LTD

Space-time consistent video depth completion method under zero sample unified diffusion framework

The invention discloses a space-time consistent video depth completion method under a zero sample unified diffusion framework. The method comprises the following steps: constructing a depth completion model comprising a variational auto-encoder, a semantic coding network and a space-time diffusion generation network; preparing training data, and generating a frame-level semantic feature vector and a conditional latent variable fusing an original depth and a relative depth for a video frame; training the space-time diffusion generation network in stages by taking the conditional latent variable sequence as input and the semantic features as conditions; in the inference stage, a video sequence to be complemented is processed through a sliding window fusion mechanism, a de-noising depth latent variable is obtained through a trained network, and finally a complemented depth sequence is output through decoding and scale recovery of a variational auto-encoder. The method has the advantages that a depth sequence with measurement consistency, structural integrity and time stability can be generated when depth completion is performed on a sparse, noisy or structurally damaged long sequence video.
Owner:浙江大学宁波国际科创中心

Video generation methods, electronic device, and computer-readable storage medium

The present disclosure relates to the technical field of computers and video processing. Disclosed are video generation methods, an electronic device, and a computer-readable storage medium. A method comprises: acquiring a generation condition and a control condition, the generation condition being used for providing a video material for a target video to be generated, the control condition being used for guiding generation of video content matched with the generation condition, and the duration of the target video being greater than a preset duration; and, on the basis of the generation condition, the control condition and target noise corresponding to the target video, generating the target video, the target noise being the same initial noise added to a plurality of target image frames contained in the generated target video, and the target noise being used for controlling the smooth transition of the target video between the different image frames. The present disclosure solves the technical problems of bad video rationality, video continuity and video transition of long videos generated by long video generation methods in the prior art.
Owner:ALIBABA (CHINA) CO LTD

Processor and system to encode sequence data in neural networks

Apparatuses, systems, and techniques to encode sequence data in one or more neural networks. In at least one embodiment, a video frame sequence is generated using a neural network to map noise frames to video frames.
Owner:NVIDIA CORP

Three-dimensional model sequence generation method and related equipment

The embodiment of the invention discloses a three-dimensional model sequence generation method and related equipment. The related equipment can comprise a three-dimensional model sequence generation device, electronic equipment, a computer program product and a computer readable storage medium. According to the embodiment of the invention, feature extraction is carried out on video frames in a monocular video to obtain image features, an initial noise sequence corresponding to a three-dimensional model of a target object is generated, denoising is carried out on the initial noise sequence according to the image features to obtain a feature sequence set, and based on the frame positions of the video frames and the time distance between the video frames, the target object is obtained. Screening at least one reference feature block associated with the feature block from the feature sequence set, denoising the feature block according to the reference feature block to obtain a target feature sequence of the video frame, and generating a three-dimensional model sequence of the target object based on the target feature sequence; according to the scheme, the reference feature blocks can be screened to perform block-level cross-frame information interaction, so that the generation quality of the three-dimensional model sequence can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Lip synchronization method, device and equipment based on diffusion Transformer, medium and program product

The invention provides a lip synchronization method and device based on diffusion Transform, equipment, a medium and a program product, and relates to the technical field of artificial intelligence processing. The method comprises the following steps: acquiring a source video frame, a target text and a target audio; performing noise adding processing on the source video frame by using a stream matching noise generator to generate a noisy video frame; performing global lip shape synchronization processing based on the noisy video frame, the target text, the target audio and the trained lip synchronization diffusion model, and outputting a lip synchronization video frame; wherein the trained lip synchronous diffusion model is obtained by maskless training based on a video sample, a text sample and an audio sample. According to the embodiment of the invention, noise is added through the stream matching noise generator, the inter-frame jitter phenomenon is reduced, the lip synchronization video frame can be smoothly generated through the trained lip synchronization diffusion model, and the cross-language adaptation capability is improved.
Owner:BEIJING XIAOBING YUEDONG TECHNOLOGY CO LTD

Method, electronic device, and computer program product for generating video

A method includes obtaining a reference image and a reference speech, the reference image specifying a head of a target object in the video, and the reference speech specifying a voice of the target object; and generating, based on the reference image and the reference speech, a fusion vector by combining a feature of the head and a feature of the voice. The method further includes generating, based on the fusion vector, a plurality of video frames in a video that represents the target object speaking in a timbre of the reference speech by denoising a plurality of initial frames including noise; and generating the video based on the plurality of video frames. In embodiments of the present disclosure, a video in which a semantic feature and a speaking style of the target object are merged can be generated, and the resolution and quality of the generated video are enhanced.
Owner:DELL PROD LP

Target image recognition and target detection method based on video enhancement algorithm

The invention discloses a target image recognition and target detection method based on a video enhancement algorithm, and relates to the technical field of image processing. The method comprises the following steps: firstly, receiving a rain, snow and fog scene video stream through a visual sensor, extracting a video frame target image, performing video enhancement processing, eliminating rain and snow shielding, fog blurring and noise, and generating an effectively enhanced image which is complete in target contour, clear in details and adaptive to subsequent detection; inputting the image into an improved YOLO model of the rain, snow and fog scene, completing feature extraction and category recognition through an optimized feature extraction network, outputting a preliminary target bounding box and a category label, and judging whether a target to be detected and a specific category exist or not; and finally, if the target exists, counting the detection data and carrying out validity verification, thereby realizing high precision, low misjudgment and strong real-time performance of target detection in severe weather of rain, snow and fog, and further effectively solving the problem of high detection result misjudgment rate caused by parameter adjustment lag of adaptive filtering in the prior art.
Owner:BEIJING LISIDA NEW TECH CO LTD

Structural vibration displacement identification method, device and equipment and storage medium

The invention relates to the technical field of bridge structures and vision measurement, and discloses a structure vibration displacement recognition method, device and equipment and a storage medium, and the method comprises the steps: obtaining a vibration video of a to-be-recognized region, and carrying out the preprocessing of the vibration video, and obtaining an image sequence and a displacement proportion; deblurring the image sequence by adopting a space-time coupling method to obtain a clear image sequence and an optimized optical flow matrix; performing intermediate frame interpolation based on the clear image sequence and the optimized optical flow matrix to obtain a target image sequence; and performing structure vibration displacement identification based on the target image sequence and the displacement proportion to obtain vibration displacement. According to the method, deblurring is carried out through a space-time coupling method, motion blurring is effectively eliminated, noise is suppressed, a clear image sequence is used for frame insertion, the frame density of the image sequence is increased, the problem of large inter-frame displacement caused by insufficient video frame rate is relieved, vibration recognition is carried out on the target image sequence in combination with the displacement proportion, and physical displacement of structural vibration is obtained. And the displacement identification precision under complex conditions is improved.
Owner:CENT SOUTH UNIV +1

Video pulse wave extraction method and system based on automatic noise recognition

The invention discloses a video pulse wave extraction method and system based on automatic noise recognition. The method comprises the steps that a face area is divided into a plurality of sub-areas; calculating the signal-to-noise ratio of each sub-region in the pulse frequency band, and performing weighted fusion on the original pulse signals of each sub-region according to the signal-to-noise ratio; performing ensemble empirical mode decomposition on the fusion signal; according to the peak-to-peak interval variation coefficient, the amplitude stability index and the spectrum purity, determining a noise component dominated by the motion artifact; performing adaptive filtering processing on the fused signal by using an adaptive filter and taking a reference noise signal as input; carrying out distortion detection on the filtered fusion signal, and carrying out signal restoration processing on a detected distortion region to obtain an extracted pulse signal; in order to solve the problem that in the prior art, when video pulse signals are extracted, motion artifacts and physiological signal frequency spectrums are overlapped, and consequently signal separation of a traditional frequency domain filtering method fails, effective signal separation and accurate pulse signal extraction are achieved.
Owner:ANHUI UNIV

Positional embedding and training techniques for a diffusion model

The present disclosure relates to systems, methods, and non-transitory computer-readable media that generates spatial-temporal positional encodings. For example, the disclosed systems generate a noised token from adding noise to an embedding of a frame of a video. Moreover, the disclosed systems generate a spatial embedding for a token using a centered two-dimensional coordinate map. Further, the disclosed systems generate temporal embeddings for the token from a timestamp of the token in the video. Further, the disclosed systems generate a denoised token by removing noise from the noised token according to spatial-temporal positional encodings that include the spatial embedding and the temporal embedding via a diffusion model. Additionally, the disclosed systems modify parameters of the diffusion model based on a comparison of the denoised token and the token.
Owner:ADOBE INC