Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

504 results about "Reference frame" patented technology

Reference frames are frames of a compressed video that are used to define future frames. As such, they are only used in inter-frame compression techniques. In older video encoding standards, such as MPEG-2, only one reference frame – the previous frame – was used for P-frames. Two reference frames (one past and one future) were used for B-frames.

Super-resolution imaging method based on focal plane splicing and adaptive fusion

The invention relates to the field of digital image processing, in particular to a super-resolution imaging method based on focal plane splicing and adaptive fusion. According to the method, sub-pixel offset among nine CCDs is preset through hardware, and nine frames of low-resolution image sequences with accurate displacement are obtained in push-broom. A central image is taken as a reference frame, high-precision mapping is realized based on hardware offset, motion estimation errors are avoided, effective pixels are screened by calculating robustness weight, an anisotropic Gaussian kernel function with a self-adaptive local structure is constructed so as to maintain image edge and detail features, and each frame is accumulated to a high-resolution grid in a weighting mode, so that a high-resolution image is obtained. And a sample compensation mechanism based on cumulative robustness is introduced, a fusion strategy is adaptively adjusted in an information insufficient area, and finally a high-resolution image is generated through normalization. The method significantly improves the imaging quality, suppresses artifacts and noise, and is suitable for the field of satellite remote sensing.
Owner:XIANGTAN UNIV

Video monitoring data transmission and storage method based on narrow bandwidth

The invention relates to the technical field of video compression, in particular to a narrow-bandwidth-based video monitoring data transmission and storage method, which comprises the following steps of: for an input video monitoring data frame, calculating a difference absolute value sum between a current frame and a reference frame; according to the method, the pixel difference of the current frame and the reference frame is calculated, and the average amplitude of the motion vector field is fused to form a comprehensive index of the dynamic degree of the quantized content, so that the image group structure is not fixed or periodic any more, and the dynamic degree of the quantized content can be calculated according to the score of scene activity and the change intensity. When a monitoring picture is static, a super-long image group is established to limit a compression code rate, when the picture is suddenly changed, the super-long image group is quickly switched to a short image group to ensure instant refreshing and definition of key information, meanwhile, non-uniform redistribution is performed on limited total code rate budget, bit resources can be intelligently inclined to frames with violent content change and large information amount, and the real-time refreshing and definition of the key information are ensured. Therefore, the subjective visual quality of the key dynamic moments is improved under the narrow bandwidth.
Owner:THE FIRST MONITORING AND APPLICATION CENTER CHINA EARTHQUAKE ADMINISTRATION +1

Generative video compression with a transformer-based discriminator

A method, an apparatus, and a non-transitory computer-readable storage medium for video compression using a generative adversarial network (GAN) are provided. The method includes obtaining, by a generator of the GAN, a reconstructed target frame based on a reference frame and a raw target frame to be reconstructed; concatenating, by a transformer-based discriminator of the GAN, the reference frame, the raw target frame and the reconstructed target frame to obtain a paired data; determining, by the transformer-based discriminator of the GAN, whether the paired data is real or fake to guide reconstruction of the raw target frame; and determining a generator loss and a transformer-based discriminator loss, and performing gradient back propagation and updating network parameters of the GAN based on the generator loss and the transformer-based discriminator loss.
Owner:SANTA CLARA UNIVERSITY +1

Anti-unmanned aerial vehicle low-altitude small target detection method based on RTDETR

The invention belongs to the technical field of computer vision and target detection, and discloses an anti-unmanned aerial vehicle low-altitude small target detection method based on RTDETR, and the method comprises the steps: extracting the features of an input image through employing a PVM-based backbone network, and obtaining the multi-scale features; processing the highest level features in the multi-scale features by using a self-attention-based space gating unit in the hybrid encoder to obtain gating features; other features except the highest-level feature in the multi-scale features are taken and combined with the gating features to be input into an attention fusion module based on content guidance, and fusion features are output; extracting a candidate frame with the minimum uncertainty of the fusion features as an initial reference frame, and taking a corresponding feature vector as an initial object for query; and taking the initial object query and the initial reference frame as a decoder based on the Transform, and enabling the output of the decoder based on the Transform to pass through a prediction head to obtain a final target detection result. According to the method, the small target detection performance can be remarkably improved while the real-time performance is kept.
Owner:ZHEJIANG UNIV OF TECH

Intelligent edge flame and smoke identification method based on deep learning

The invention discloses an edge flame and smoke intelligent identification method based on deep learning, and the method comprises the steps: collecting continuous video frames, carrying out the normalization, correction and noise reduction, and generating a preprocessing video sequence; constructing a background static reference frame, and carrying out pixel difference on the background static reference frame and the current frame to generate geometric refraction potential field codes; performing refraction phase mapping on the geometric refraction potential field code to form a refraction phase disturbance tensor; inputting an improved SlowFast model, dynamically adjusting three-branch sampling, and outputting a preliminary candidate region; extracting refraction, phase and energy evolution sequences, constructing a coupling sequence and correcting an identification result; and calculating a risk level, marking a high-risk area, and outputting fire early warning at edge equipment. According to the invention, through constructing the refraction potential field features and the phase disturbance features and combining the improved multi-branch SlowFast deep learning model, rapid, accurate and stable edge side intelligent identification and early warning of flames and smog are realized.
Owner:ZHONGLANG INFORMATION TECH CO LTD

Video subtitle erasing method and device, equipment and storage medium

The invention discloses a video subtitle erasing method, device and equipment and a storage medium, and relates to the field of digital image processing, and the method comprises the steps: detecting subtitles of a subtitle video to be erased, merging timestamps of the same subtitles in the subtitle video to be erased, and determining a subtitle fragment set and a subtitle-free fragment set; splitting the subtitle segment into independent shot segments by using a preset lens splitting algorithm, and analyzing video frames of the independent shot segments to obtain a first frame and a tail frame; determining a target reference frame based on the first frame and the tail frame, generating an expanded independent shot segment according to the target reference frame and the independent shot segment, and segmenting a target character mask; and generating an erased independent lens segment according to the expanded independent lens segment and the target character mask through a preset erasure algorithm, performing a preset post-processing optimization operation on the erased independent lens segment to obtain a target independent lens segment, and integrating the target independent lens segment and the subtitle-free segment set to generate a target video. According to the method and the device, the video subtitles can be accurately erased.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Performance improvement method for double-independent quantum key distribution of actual reference system measurement equipment

The invention discloses a method for improving the performance of double-independent quantum key distribution of actual reference system measurement equipment, belongs to the technical field of quantum key distribution (QKD), and aims to improve the performance of the double-independent quantum key distribution when non-ideal factors such as a post-pulse effect of a detector and finite sample statistical fluctuation are considered. The invention relates to a method and a system for improving the performance of a reference system measurement device by introducing an advantage extraction technology into a double independent quantum key distribution (RFI-MDI-QKD) protocol of the reference system measurement device. Simulation results show that when the post-pulse effect and the statistical fluctuation effect are considered, the advantage extraction method can effectively improve the security key rate and the security transmission distance of the RFI-MDI-QKD protocol on the premise of not changing optical hardware. According to the invention, a valuable reference technology can be provided for practical research of the RFI-MDI-QKD system.
Owner:NANJING UNIV OF POSTS & TELECOMM

Two-stage cascaded video focus detection method and system

The invention discloses a two-stage cascaded video focus detection method and system, and relates to the technical field of image target recognition, and the method comprises the steps: obtaining an endoscope image; selecting a reference frame; carrying out iterative calculation on the feature embedding of the reference frame to obtain deep space-time representation; generating a time attention weight based on the deep spatio-temporal representation, and performing weighted fusion on the deep spatio-temporal representation to obtain an enhanced feature; performing dimension reduction on the enhanced features to obtain video prompt information; and extracting feature embedding of the target reasoning frame and reasoning frame features in the video prompt information, performing iterative calculation on the reasoning frame features to obtain deep reasoning features, and performing focus detection on the deep reasoning features to obtain a detection result. According to the method, a two-stage cascade Transform architecture is adopted, the quality degradation phenomena of dynamic blur, exposure imbalance, reflection artifacts and the like of inference frames are effectively relieved, and efficient joint modeling of spatial-temporal characteristics is achieved.
Owner:XIAN UNIV OF POSTS & TELECOMM

Video compression platform based on AI

The invention relates to the technical field of video compression, in particular to an AI-based video compression platform, which comprises a key target detection module, an image region division module, a coding level setting module, a parameter adjustment and analysis module and a coding result integration module. According to the method, key region extraction is completed by collecting pixel color combination and contour change information in a video frame image, background and non-background region boundaries are delimited, a region attribute labeling result is constructed, a region distribution classification relation mapping coding level is established, and reference frame and prediction interval configuration is extracted. Real-time adjustment of compression levels is realized by performing normalized combination analysis on textures and motion change trends of different regions, and region fragments under different compression levels are recombined and subjected to unified code stream processing, so that coding resources can be dynamically allocated on the basis of accurately identifying contents in a compression process; the problems of inaccurate target positioning, fixed compression configuration, delayed content response and the like in an existing compression system are effectively solved.
Owner:HUNAN SANLI INFORMATION TECHNOLOGY CO LTD

Video code rate dynamic allocation compression method based on content complexity prediction

The invention discloses a video code rate dynamic allocation compression method based on content complexity prediction, which comprises the following steps of: S1, acquiring a video frame sequence, and preprocessing; s2, inputting the frame-level feature tensor into a gated residual convolutional network, and outputting a complexity prediction sequence; s3, constructing a frame priority queue, and calculating the complexity jump amplitude between adjacent frames; s4, a compression area is divided, and a corresponding code rate resource scale factor is allocated; s5, configuring a reference frame structure, a prediction interval and an initial quantization step size for each compression region, and calculating a region target bit number; s6, distributing regional bits to each frame in a compression coding process, and dynamically adjusting a frame-level quantization parameter and an entropy coding strategy; and S7, after compression is completed, reversely updating convolution prediction network parameters through bit distribution errors. According to the invention, fine code rate dynamic allocation and adaptive compression control based on content complexity are realized, and the video compression quality and bit utilization efficiency are effectively improved.
Owner:HANGZHOU DIGITAL AMBER TECHNOLOGY CO LTD

Multi-frame edge-enhanced deghosting

A method includes selecting a reference frame and a non-reference frame from a plurality of image frames. The method also includes generating a reference edge map identifying edges in the reference frame and a non-reference edge map identifying edges in the non-reference frame. The method further includes generating a moving edge map based on one or more movements between one or more of the edges of the reference edge map and one or more of the edges of the non-reference edge map. The method also includes generating a blend map based on the reference and non-reference frames. The method further includes modifying the blend map based on one or more indications of movement of corresponding pixels in the moving edge map to generate a modified blend map. In addition, the method includes blending the reference and non-reference frames based on the modified blend map to generate an output image.
Owner:SAMSUNG ELECTRONICS CO LTD

Video decoding with lossy reference frame

A device for decoding video data includes an integrated circuit (IC) comprising a video decoder, and a memory that is external to the IC and coupled to the IC. The video decoder is configured to in a first mode, decode a first frame based on a first reference frame stored in the memory, and in a second mode, decode a second frame based on a second lossy reference frame, wherein the second lossy reference frame is generated based on decompression of a lossy compressed reference frame.
Owner:QUALCOMM INC

System and method for AI segmentation-based registration for multi-frame processing

A method includes obtaining a reference frame from among multiple image frames of a scene. The method also includes generating a segmentation mask using the reference frame, where the segmentation mask contains information for separation of foreground and background in the scene. The method further includes applying the segmentation mask to each of the multiple image frames to generate foreground image frames and background image frames. The method also includes performing multi-frame registration on each of the foreground image frames to generate registered foreground image frames. The method further includes performing multi-frame registration on each of the background image frames to generate registered background image frames. In addition, the method includes combining the registered foreground image frames and the registered background image frames to generate a combined registered multi-frame image of the scene.
Owner:SAMSUNG ELECTRONICS CO LTD

Cross-channel distributed video coding method and system based on multi-dimensional attention

This disclosure provides a cross-channel distributed video encoding and decoding method and system based on multidimensional attention, which can be applied to the field of video encoding and decoding technology. The method includes: acquiring a group of video images to be transmitted, the group of video images including a first key reference frame, a second key reference frame, and at least one intermediate video frame; encoding the first key reference frame and the second key reference frame into first encoded data and second encoded data, respectively; encoding each intermediate video frame into third encoded data based on a frame encoder; extracting multi-scale features from each third encoded data based on a multidimensional attention mechanism; processing the multi-scale features into fourth encoded data based on a multi-channel feature extraction mechanism; processing at least one fourth encoded data into a first bitstream; and converting the first encoded data and the second encoded data into a second bitstream, so as to transmit the video image group to the decoder via the first bitstream and the second bitstream.
Owner:INST OF MEDICAL ROBOTICS & INTELLIGENT SYST TIANJIN UNIV

Gaussian neural field dynamic scene reconstruction system based on depth consistency constraint

The invention provides a Gaussian neural field dynamic scene reconstruction system based on depth consistency constraint, and relates to the technical field of computer graphics, and the system comprises an estimation module which generates a target frame initial depth map; the calculation module reconstructs the point cloud and obtains a point cloud normal direction and a pixel normal direction; the optimization module is used for iteratively correcting the initial depth based on the two types of normal consistency to obtain an optimized depth map; the alignment module is used for determining a scale parameter through regression by taking the first target frame as a reference, and carrying out scale transformation on the depths of other frames to form a consistent depth sequence; the reconstruction module is used for constructing or training a Gaussian neural field based on the sequence and outputting a three-dimensional representation; in addition, the calculation module can contain multi-dimensional wavelets and sparse reconstruction and is used for multi-scale noise suppression and direction weighted fitting. Reference frame selection is based on frame-level quality, geometry and scale stability indexes; according to the system, the intra-frame geometric credibility and the cross-frame scale consistency are improved, ghosting and tearing are reduced, and the stability and integrity of dynamic scene reconstruction are enhanced.
Owner:LISHUI RES INST OF HANGZHOU UNIV OF ELECTRONIC SCI & TECH

Expressway electromechanical operation and maintenance management method based on time series data of Internet of Things

The invention relates to the field of Internet of Things terminal management, in particular to a highway electromechanical operation and maintenance management method based on Internet of Things time series data, which comprises the following steps of: selectively decoding a coded data stream, and marking background interference frames according to a plurality of decoded frames so as to analyze the influence of background interference on coded data; the method comprises the steps of setting an interference identifier for coded data of an observation time domain segment, identifying abnormal coded data for unidentified coded data, synchronously determining a target reference time domain segment where the abnormal coded data is located, decoding coded data in the target reference time domain segment, obtaining a plurality of reference frames, and processing the reference frames. According to the method, the difference of the coded data collected by the Internet of Things terminals is considered, the rule constraint portraits aiming at different Internet of Things terminals are constructed, the multi-dimensional coded data are monitored adaptively, on the premise of ensuring the reliability, only partial decoding is carried out, and when the method is oriented to the large-range Internet of Things terminals, the computing resource consumption is reduced, and the monitoring efficiency is improved.
Owner:CHINA MERCHANTS XINZHI TECH CO LTD +1

Coding and decoding method, encoder, decoder, code stream and storage medium

An encoding and decoding method, an encoder, a decoder, a code stream and a storage medium, the decoding method being applied to the decoder, the decoding method comprising: parsing the code stream, determining first syntax element information of a current decoding unit of a current frame, the first syntax element information representing a type of the decoding unit; the reference information of the reference frames corresponding to different types of decoding units is different; and predicting the current node based on the type of the current decoding unit indicated by the first syntax element information to obtain an attribute prediction value of the current node.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Quality management system and method for engineering construction

The invention relates to the field of engineering management, and particularly discloses a quality management system and method for engineering construction, and the method comprises the steps: firstly carrying out the preprocessing of a BIM model, and constructing a reference frame integrating visual features and spatial geometric information; when a problem image shot by a mobile terminal at a construction site is received, a computer vision and space calculation technology is introduced, a two-dimensional field image is analyzed into an accurate pose of a shooting camera in a three-dimensional BIM space, and sight line tracing is automatically performed in a BIM digital model based on the accurate pose. And intelligently inferring and retrieving candidate components which are highly matched with the perspective of the problem. And finally, on-site unstructured problem data and the unique identifier of the BIM component are firmly bound through lightweight interaction confirmation with the user, and a standardized structured problem record is obtained. Therefore, the problems of low manual positioning efficiency and unstable association can be solved, and the whole-process intelligent closed loop of quality problem data from acquisition to structured storage is realized.
Owner:国家能源集团雄安能源有限公司

Image processing method, device and equipment

The invention provides an image processing method, device and equipment. The method comprises the following steps: dividing all P frames in a GOP sequence into a first type of P frames, a second type of P frames and a third type of P frames; wherein the first type of P frames cannot serve as reference frames, the second type of P frames can only serve as the reference frames of the first type of P frames, and the third type of P frames can serve as the reference frames of the first type of P frames, can serve as the reference frames of the second type of P frames and can serve as the reference frames of the third type of P frames; and if it is determined that the P frames in the GOP sequence need to be subjected to frame extraction, performing frame extraction on the first type of P frames, or performing frame extraction on the first type of P frames and the second type of P frames, or performing frame extraction on the first type of P frames, the second type of P frames and the third type of P frames. Through the technical scheme of the invention, a part of P frames are discarded in some special scenes to achieve the purpose of saving bandwidth, and the characteristics of decoding multiplication paths, low-bandwidth transmission, low-code-rate storage and the like are realized.
Owner:HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Projected motion field hole filling for motion vector reference

Projected motion field hole filling includes determining a motion field of a current frame using a block true subset of the current frame, the motion field including, for each block in the block true subset, a respective motion vector projected onto a reference frame. For a block in the current frame that does not have a motion vector in the motion field, a respective motion vector of a spatial domain neighboring block of the proper subset of the block is reused as a projected motion vector of the block within the motion field. A list of motion vector candidates may be determined using the motion field. A reference motion vector for the current block may be selected from the motion vector candidate list. The current block may be encoded into an encoded bitstream or decoded from an encoded bitstream using the reference motion vector.
Owner:GOOGLE LLC

Joint correction method and system for GNSS (Global Navigation Satellite System) non-structural deformation and storage medium

The invention discloses a joint correction method and system for GNSS (Global Navigation Satellite System) non-structural deformation and a storage medium. According to the method, GNSS pseudo-range and carrier phase observation data and precise products such as a precise orbit and clock correction are obtained, a three-component displacement time sequence caused by a non-tidal atmospheric load, a non-tidal ocean load and a land hydrological load and a temperature-driven thermoelastic displacement time sequence are constructed, and the three-component displacement time sequence and the temperature-driven thermoelastic displacement time sequence are unified to a reference frame and a time system consistent with GNSS calculation. After interpolating the non-structural displacement vector to an observation epoch according to time, converting the non-structural displacement vector to a geocentric rectangular coordinate system from a site local coordinate system ENU, and projecting the non-structural displacement vector to a sight line direction from a satellite to an observation station reference point to form an equivalent geometric distance correction item; the correction terms are respectively applied to pseudo-range observation and carrier phase equivalent meter domain observation in the form of the same domain as observation, and then GNSS daily arc precise calculation is performed to output a coordinate time sequence, and preferably, the method can be used for a precise single-point positioning integer ambiguity fixing technology.
Owner:WUHAN UNIV

Video encoding apparatus with limited reconstruction buffer and associated video encoding method

A video encoding apparatus including a reconstruction buffer with fixed size and / or bandwidth limitations, and an associated video encoding method. Specifically, a video encoding apparatus includes a data buffer and video encoding circuitry. Encoding of a first frame includes: deriving reference pixels of a reference frame from the reconstructed pixels of the first frame, and storing the reference pixel data in the data buffer for inter-frame prediction, wherein the reference pixel data includes information about the pixel values ​​of the reference pixels. Encoding of a second frame includes: performing prediction on coding units in the second frame to determine a target prediction factor for the coding unit. The prediction step performed on the coding unit includes: determining the target prediction factor of the coding unit based on whether a search range for the prediction factor of the coding unit for the reference frame includes at least one reference pixel of the reference frame that cannot be accessed by the video encoding circuitry.
Owner:MEDIATEK INC

A motion estimation based video recognition acceleration method

This invention provides a motion estimation-based method for accelerating video recognition, comprising: identifying keyframes and non-keyframes in a video sequence; extracting features from keyframes using a Bayer domain basic model to obtain perceptual features; calculating motion vectors between non-keyframes and a reference frame (the frame preceding the non-keyframe) using a fast motion estimation module, wherein the fast motion estimation module employs a pyramid block structure and performs multi-level matching search from coarse to fine under a GPU parallel architecture; deforming the features of the reference frame using the motion vectors to obtain propagation features; predicting the perceptual residual of the current frame using a perceptual residual correction network, numerically correcting the propagation features using the residual, and outputting the corrected propagation features, wherein the perceptual residual correction network is a lightweight network structure; and performing video recognition based on the corrected propagation features. This invention achieves faster and more efficient video recognition.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Improvements on signaling inter prediction mode

The various implementations described herein include methods and systems for coding video. In one aspect, a method includes receiving a video bitstream comprising a plurality of blocks including a current block. The method includes determining that the current block is encoded using motion information from a first reference block and a second reference block. The method includes (i) when the first and second reference blocks are in different reference frames, selecting a compound inter prediction mode for the current block from a first set of compound inter prediction modes, and (ii) when the first and second reference blocks are in a same reference frame, selecting the compound inter prediction mode for the current block from a second set of compound inter prediction modes. The method includes reconstructing the current block using the compound inter prediction mode and the motion information from the first and second reference blocks.
Owner:TENCENT AMERICA LLC

Determining a device location on a body part

According to an aspect, there is provided a method of determining a location of a device on a surface of a body part of a subject that is treated by the device during a treatment operation. The device is for performing the treatment operation on the body part, and the device comprises one or more orientation sensors for measuring the orientation of the device in a reference frame of the device. The method comprises obtaining a three dimensional, 3D, representation of the body part, the 3D representation comprising normal vectors for respective positions on the surface of the body part; receiving a plurality of orientation measurements from the one or more orientation sensors representing orientation of the device during the treatment operation on the body part; processing the received orientation measurements to determine a sequence of orientations of the device during the treatment operation; and comparing the determined sequence of orientations of the device to the normal vectors and respective positions of the normal vectors to determine the location of the device on the surface of the body part during the treatment operation.
Owner:KONINKLIJKE PHILIPS NV

Visual axis stable video generation method and device based on dynamic fuzzy integral

The invention provides a visual axis stable video generation method and device based on dynamic fuzzy integration, and the method comprises the steps: collecting inertial pointing data of a servo system in an inertial stability model, and reading a gray value of an initial scene image as a reference frame; performing interpolation processing on the acquired azimuth angle, pitch angle and roll angle data to generate a high-density pointing angle sequence, and calculating azimuth stability precision and pitch stability precision; reversely rotating the reference frame to generate imaging slices, superposing all corresponding imaging slices pixel by pixel in single-frame exposure time, and normalizing a result; and overlapping character information including a field angle, stability precision and an imaging parameter on each frame of image to generate a visual axis stabilization effect video file. In the embodiment of the invention, the visual axis stability precision effect video is generated according to the pointing data and the imaging sensor parameters, the intuition of stability precision effect display in a design stage can be improved, and the dependence of a servo system on an optoelectronic equipment image link in a physical test stage is reduced.
Owner:CENT CHINA OPTOELECTRONICS TECH RES INST (CHINA STATE SHIPBUILDING CORP 717TH RES INST)

Warped motion compensation with explicitly signaled extended rotations

Video coding using warped motion compensation is described. Extended rotations for the warped motion compensation can be explicitly signaled. For example, motion parameters for predicting the current block and a rotation angle can be decoded. A warping matrix is obtained using the motion parameters and the rotation angle, and a prediction block is obtained by projecting the current block to a quadrilateral in a reference frame. Also described is determining a prediction model of the current block and obtaining a prediction block by projecting the current block to a quadrilateral in a reference frame. Determining the prediction model can include determining whether to predict the current block using a motion vector, a local warping model, or a global motion model, obtaining motion parameters of the prediction model, decoding a rotation angle, and obtaining a warping matrix using the motion parameters and the rotation angle.
Owner:GOOGLE LLC

Video temporal variable sampling method, device and medium based on reversible network

The application discloses a video time variable sampling method and device based on a reversible network and a medium, belongs to the technical field of video coding, and solves the problem of serious loss of information in the down-sampling and coding stages during video time variable sampling coding; the reversible time domain wavelet transform structure of multiple input and multiple output is adopted, input high frame rate video is decomposed into odd and even sequence frames, a bidirectional reference frame mechanism is introduced, a time domain low frequency subband and a high frequency subband are generated, high-quality video time frequency band decomposition can be realized without transmitting explicit motion information between forward transformation and inverse transformation, a gradient agent network simulating the compression behavior of a traditional encoder is introduced, motion vectors and quantization parameters are taken as input, a reversible affine coupling transform model encoder perceiving coding quality is adopted to code the distortion of the time domain low frequency, and the defects that the existing agent network depends on motion vector transmission or fixed coding parameters and is difficult to adapt to actual variable application scenarios can be overcome.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST

Defining a search range for motion estimation for each scenario frame set

A video motion estimation method including obtaining a plurality of image frames in a video, and performing scenario classification processing on the plurality of image frames to obtain a plurality of image frame sets. The method further includes extracting a contour feature and a color feature of a foreground object of each image frame, and determining a search range corresponding to each image frame set. The method further includes determining a starting search point in each predicted frame. The method further includes, for each image frame set, performing motion estimation processing in a search region corresponding to the search range of each predicted frame set based on the starting search point of the respective predicted frame, a reference block in at least one reference frame of the respective image frame set, and the color feature of the foreground object, to obtain a motion vector corresponding to the reference block.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Mesh decoding device, mesh decoding method, and program

A displacement decoding unit (205) of a mesh decoding device (200) according to the present invention includes: a bypass arithmetic decoding unit (205A) configured to generate a coefficient level value by performing bypass arithmetic decoding on a displacement bit stream; an inverse quantization unit (205B) configured to generate a first transformed coefficient by performing inverse quantization on the coefficient level value; an adder (205D) configured to generate a second transformed coefficient by adding a prediction transformed coefficient and a prediction residual; an inter prediction unit (205F) configured to generate the prediction transformed coefficient by performing inter prediction by using the second transformed coefficient of a reference frame read from the frame buffer; and a second inverse transform unit (205G) configured to generate a decoded displacement by performing second inverse transform on the second transformed coefficient.
Owner:KDDI CORP