Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

27826results about "Digital video signal modification" patented technology

Multi-Scale Temporal Attention Processing System for Multimodal Deep Learning with Vector-Quantized Variational Autoencoder

A system and method for multi-scale temporal attention processing in multimodal technology deep learning systems. This system processes time-series, textual, sentiment, and structured tabular data across three hierarchically-organized temporal streams—quarterly, weekly, and intraday levels—with bidirectional cross-temporal information flow. Scale-specific attention mechanisms are optimized for respective temporal granularities, while an adaptive controller dynamically weights each temporal level based on real-time market volatility indicators. A multi-scale fusion processor integrates attention-weighted representations to generate temporally unified representations preserving both short-term market dynamics and long-term trends. This approach enables superior forecasting and risk assessment by leveraging temporal correlations across multiple time scales while automatically adapting to changing market conditions. The system facilitates interpretable AI analysis through attention visualization and enables synthetic scenario generation for model testing.
Owner:ATOMBEAM TECH INC

Methods for delta-QP signaling for decoder parallelization in hevc

ActiveUS20120183049A1Color television with pulse code modulationColor television with bandwidth reductionComputer architectureCoded block flag
By implementing a new bitstream for a Delta-Quantization Parameter (DQP), a decoder is able to implement parallel decoding of multiple coding units within a largest coding unit. In some embodiments, the DQP is placed immediately after the mode information of the first coding unit. In some embodiments, the DQP is placed after the mode information of the first non-skipped coding unit. In some embodiments, the DQP is placed after the first non-zero coded block flag.
Owner:SONY GROUP CORP

Variable-bit-rate image compression method and system, apparatus, terminal, and storage medium

The present disclosure provides a variable-bit-rate image compression method and system, an apparatus, a terminal, and a storage medium. The variable-bit-rate image compression method includes: obtaining an initial feature map from a to-be-encoded image; quantizing the initial feature map by a dead-zone quantizer; performing entropy encoding on the quantized feature map and hyper-prior information to obtain a compressed bit-stream; performing entropy decoding on the compressed bit-stream, and recovering quantized hyper-prior information and the quantized feature map; performing inverse quantization on the quantized feature map to obtain a reconstructed feature map; obtaining a reconstructed image from the reconstructed feature map; and adjusting quantization and inverse quantization parameters according to a target bit-rate or target distortion. The present disclosure provides a precise bit-rate control solution, makes the bit-rate of the compressed bit-stream better adapt to the dynamic change of a network bandwidth, and has an extremely high actual application value.
Owner:SHANGHAI JIAOTONG UNIV

Large-scale scene multi-level-of-detail cloud rendering processing method and device based on 3DGS

The invention provides a large-scale scene multi-level-of-detail cloud rendering processing method and device based on 3DGS, and relates to the technical field of three-dimensional modeling, and the method comprises the steps: dividing a target modeling scene into a plurality of sub-blocks; performing particle redundancy reconstruction and overlapping region marking on boundary regions between adjacent sub-blocks of each sub-block to obtain processed sub-blocks; performing multi-detail level division on each processing sub-block to generate a particle level set; determining a current visual area according to the user motion data, and scheduling a target hierarchy of a particle hierarchy set in the current visual area; the rendering tasks of all the processing sub-blocks of the target hierarchy are distributed to a plurality of rendering nodes to execute real-time rendering operation; and performing video stream coding on pictures rendered by each rendering node, decoding and displaying received video stream data, and performing particle level updating and re-rendering operation according to an interaction instruction. According to the invention, high-quality detail rendering can be realized for a model of a large scene.
Owner:MOBILE BROADCASTING & INFORMATION SERVICE IND INNOVATION RES INST (WUHAN) CO LTD

Data compression transmission method and system applied to ferry inspection images

The invention discloses a data compression transmission method and system applied to a ferry inspection image, and the method comprises the steps: collecting and obtaining the ferry inspection image in real time, recognizing a key inspection target region in the image, and carrying out the segmentation and partitioning of the image; compressing the key inspection target area based on lossless compression coding; the quantization step size is dynamically adjusted by comparing statistical variances of background pixels between continuous frames, and lossy compression coding is carried out on a background area; based on the boundary distance between the key inspection target area and the background area, adaptive compression coding is carried out on the transition area; constructing a hierarchical data packet; and constructing a data transmission optimization model, dynamically allocating data transmission links, and obtaining a transmission scheme with the highest total transmission. The method has the advantages that efficient data compression transmission is realized by accurately segmenting the image area and adopting a lossless, lossy and adaptive compression technology, the overall transmission efficiency is improved through intelligent transmission optimization, and the definition and real-time performance of the inspection image are ensured.
Owner:JIANGSU ZHENYANG QIDU CO LTD

Adaptive intelligent multi-modal media processing and delivery system

This invention introduces an adaptive system for multi-modal media processing and delivery, addressing challenges in modern digital content distribution. The technology dynamically analyzes and processes media content in real-time, optimizing delivery across diverse devices, networks, and content types. Key features include adaptive processing that adjusts compression, encoding, and delivery protocols based on content characteristics and delivery constraints. The system incorporates artificial intelligence for continuous improvement, learning from historical data and user feedback. It addresses network variability and device diversity, adapting to changing conditions and optimizing content for different platforms. Security and personalization features enable protected content distribution and tailored user experiences. The invention's cross-media optimization approach allows efficient handling of various media formats within a unified framework. Its scalable, modular design suits applications from consumer streaming to enterprise-level distribution. This comprehensive solution aims to enhance content distribution efficiency and user experience in the complex, evolving digital media landscape.
Owner:QOMPLX INC

Multi-scale semantic guidance image compression method and system and storage medium

The invention discloses a multi-scale semantic guidance image compression method and system and a storage medium, and the method comprises the following steps: obtaining input image data, carrying out the preprocessing of an input image, and obtaining standardized image data; inputting the standardized image data into a pre-trained semantic segmentation network to generate a multi-scale semantic feature map and a semantic weight map corresponding to the multi-scale semantic feature map; a three-stage pyramid encoder is constructed, and the standardized image data is subjected to the following steps of: sampling under depth separable convolution to generate multi-scale features; the reversible neural network carries out nonlinear transformation on the multi-scale features; the multi-scale feature subjected to nonlinear transformation is decomposed into a low-frequency sub-band and a high-frequency sub-band through adaptive discrete wavelet transformation, dynamic selective state space modeling is executed on the high-frequency sub-band based on a semantic weight map, and a compressed code stream is generated; and inputting the compressed code stream into a decoder, decoding based on a lightweight Mama module, and reconstructing an image in combination with inverse wavelet transform and a semantic weight map.
Owner:XIANGJIANG LAB

Cross-modal joint source-channel coding and decoding method adaptable to changeable scenarios

PCT designated stageWO2026031415A1Internal combustion piston enginesBiological modelsChannel decoderImage signal
The present invention relates to the technical field of cross-modal image signal reconstruction. Disclosed is a cross-modal joint source-channel coding and decoding method adaptable to changeable scenarios. The method comprises: first designing a Transformer encoder-based cross-modal channel coding and decoding optimization solution, so as to achieve the performance improvement and robustness of a channel encoder and a channel decoder; then designing a cross-modal source coding and decoding optimization solution for haptic-to-image generation based on a latent diffusion model, so that under an image signal loss scenario, haptic information is used to guide image generation; and finally, incorporating transfer learning technology, so as to reduce additional training costs caused by a system facing changeable cross-modal communication scenarios such as a changeable channel signal-to-noise ratio and different transmission tasks. Under cross-modal changeable communication scenarios, the joint source-channel coding and decoding method provided in the present invention can solve the problems of the inability of a receiving end to well complete image reconstruction, and additional model training costs caused by changeable channel environments and scenarios.
Owner:NANJING UNIV OF POSTS & TELECOMM

Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method

A point cloud data transmission method according to embodiments may comprise the steps of: encoding point cloud data; and transmitting a bitstream comprising the point cloud data. A point cloud data reception method according to embodiments may comprise the steps of: receiving a bitstream comprising point cloud data; and decoding the point cloud data.
Owner:LG ELECTRONICS INC

Unmanned aerial vehicle multichannel image transmission optimization system based on link quality perception

The invention provides an unmanned aerial vehicle multi-channel image transmission optimization system based on link quality perception so as to improve the transmission stability and the image quality guarantee capability of image data in a complex wireless environment. The prediction module constructs a time sequence model based on the historical link quality index of the wireless channel, and outputs a future link stability prediction value; the modeling module analyzes texture change and the like between the image frames, generates an image frame evolution vector, and calculates the aging sensitivity and reconstruction importance score of the image frames according to the image frame evolution vector; the decision-making module fuses the information and generates an optimal matching relation between the image frame and the channel and a transmission priority parameter; the grouping module performs image frame clustering according to the similarity between the priority parameters and the evolution vectors, constructs compression groups and generates corresponding compression configuration files and compression image data; the transmission module schedules the compressed image data to a corresponding wireless channel for transmission; the system can be widely applied to an unmanned aerial vehicle image transmission task with relatively high requirements on real-time performance and image quality in a dynamic environment.
Owner:SHENZHEN RUIWO MOBILE CO LTD

Video stream processing method for dynamic Gaussian compression and adaptive code rate regulation

The invention discloses a video stream processing method for dynamic Gaussian compression and adaptive code rate regulation, which is suitable for scenes such as virtual reality, augmented reality and three-dimensional video, and comprises the following steps: S1, Gaussian attribute modeling and initialization; s2, constructing a binary hash grid; s3, constructing a deformation prediction network; s4, designing a mask pruning mechanism; s5, entropy modeling and arithmetic coding and decoding module design; s6, model training; and S7, video stream transmission under multiple code rates. According to the method, a unified scheme combining Gaussian volume cloud coding and adaptive video transmission is proposed for the first time, the video data storage and transmission cost is remarkably reduced, and the comprehensive performance superior to that of an existing method is obtained on multiple real and synthetic data sets.
Owner:THE CHINESE UNIV OF HONG KONG (SHENZHEN) FUTURE NETWORK OF INTELLIGENCE INST +1

Method and apparatus of encoding / decoding point cloud geometry data captured by a spinning sensors head

There is provided methods and apparatus of encoding / decoding a point cloud representing a physical object. Points are captured by a spinning sensors head and are represented by sensor indices associated with sensors that captured the points, azimuthal angles representing capture angles of said sensors, and radius values of spherical coordinates of the point. Points are ordered based order indices obtained from the azimuthal angles and the sensor indices. Order index differences are encoded. An order index difference represents a difference between order indices associated with two consecutive ordered points. Optionally, the method encodes radius values, residual azimuthal angles associated with ordered points and residuals of three-dimensional cartesian coordinates of ordered points based on their three-dimensional cartesian coordinates, decoded azimuthal angles based on azimuthal angles, decoded radius values and sensor indices.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

Lossless and lossy automatic hardware compression in graphics-to-graphics network links

An apparatus to facilitate lossless and lossy automatic hardware compression in graphics-to-graphics network links is disclosed. The apparatus includes compressor / decompressor circuitry (CDC) integrated with physical layer (PHY) intellectual property (IP) hardware circuitry for a graphics processor unit (GPU)-to-GPU communication link communicably coupling a first GPU to one or more other GPUs, the CDC to: receive a data message from the first GPU, wherein the data message is in an uncompressed format; determine that a compression process is to be applied to the data message; apply the compression process to the data message to generate a compressed data message; and cause a GPU link IP hardware circuitry that comprises the PHY IP hardware circuitry to transmit the compressed data message over the GPU-to-GPU communication link.
Owner:INTEL CORP

Video display and storage method and system based on wayland protocol and storage medium

The invention belongs to the technical field of computers, and provides a video display and storage method and system based on a wayland protocol and a storage medium in order to solve the problems that an existing wayland video system is high in rendering delay, high in CPU load, low in storage efficiency and poor in cross-platform compatibility. Video frames are directly written into an Overlay layer through DRM-KMS and mixed by bypassing a synthesizer, meanwhile, CPU memory amplitude overhead is eliminated through zero copy transmission, an ISP-GPU-VPU direct transmission assembly line is constructed, full-link DMA-BUF direct transmission is achieved, and delay is greatly reduced; iSP preprocessing, GPU shader scaling and VPU coding are all executed by special hardware, so that the load is reduced, and the multi-path processing capability is improved; vPU dynamic code rate compression is utilized, segmented storage is carried out according to events / time, metadata is synchronized to a database, redundant storage is avoided, and space is saved.
Owner:CHANGSHA YINGBEIDI ELECTRONIC TECH CO LTD

Visual encoding method and apparatus, and visual encoding model training method and apparatus

The present application relates to the field of computer vision. Provided are a visual encoding method and apparatus, and a visual encoding model training method and apparatus, which are used for using the same visual encoding model to encode images of different resolutions, and are applied to encoding scenarios for images of more sizes. The visual encoding method comprises: first, acquiring an input image, wherein the input image may be a high-resolution image and may also be a low-resolution image; and then inputting the input image into a visual encoding model, so as to output visual encoding data, wherein the visual encoding model is used for dividing the input image into a plurality of image blocks according to positional embedding, extracting features from each image block, and outputting visual encoding data on the basis of the features of each image block and corresponding positional encoding, the positional embedding is obtained by means of adjusting initial positional embedding on the basis of the difference between the input image and a preset resolution, and the positional embedding may specifically comprise a matrix corresponding to the division of the input image
Owner:HUAWEI TECH CO LTD

Video signal processing method using dependent quantization and device therefor

A video signal decoding device comprises a processor which: determines a particular quantizer for reconstructing a first quantized transform coefficient, the particular quantizer being one of a first quantizer and a second quantizer which are different from each other, the particular quantizer being determined on the basis of the state of the first quantized transform coefficient; reconstructs the first quantized transform coefficient on the basis of the particular quantizer to obtain a reconstructed transform coefficient; and updates the state of a second quantized transform coefficient that is reconstructed after the first quantized transform coefficient, wherein the first quantized transform coefficient and the second quantized transform coefficient are transform coefficients in the current block.
Owner:WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC

Visual call information processing method and system based on 5G

The invention relates to the field of data processing, and provides a 5G-based video call information processing method and system, and the method comprises the steps: continuously obtaining a real-time video frame sequence and 5G network environment perception data in a video call scene, carrying out the multi-dimensional state mapping processing of the 5G network environment perception data, constructing a network transmission adaption model, and carrying out the real-time video frame sequence and 5G network environment perception data. Generating a video coding control instruction based on the network transmission adaptation model, performing content-aware coding conversion on the real-time video frame sequence, and outputting a coding optimization stream; in the transmission process of the coding optimization stream, link state fluctuation information is obtained through a 5G network feedback channel, transmission strategy dynamic calibration is performed on the coding optimization stream according to the link state fluctuation information, and a calibration transmission stream is obtained; and carrying out decoding time sequence alignment processing on the calibration transport stream, generating a visual call output sequence which is synchronous with the time of the original video stream unit, and pushing the visual call output sequence to a receiving end presentation device.
Owner:CHENGDU IKE IND CO LTD

EVTOL multi-camera cooperative video compression coding method based on multi-source perception and intelligent partitioning

The invention discloses an eVTOL multi-camera cooperative video compression coding method based on multi-source perception and intelligent partitioning. The method comprises the following steps: constructing a scene three-dimensional perception model by fusing multi-source data of visible light, infrared and depth sensors; the method comprises the following steps: realizing dynamic video partitioning based on motion vectors and semantic analysis, and dividing a picture into a core region, a secondary region and a background region; establishing a parallax compensation motion prediction model by adopting a cross-camera reference frame sharing mechanism; and high-fidelity compression of the key area is realized through layered entropy coding and a dynamic quantization parameter distribution strategy. And the decoding end reversely executes multi-source data fusion and partition reconstruction according to the coding metadata. The method is suitable for eVTOL multi-camera video real-time transmission scenes such as polling, surveying and mapping.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Method, device, and recording medium for image encoding / decoding

Disclosed herein are a method, an apparatus and a storage medium for image encoding / decoding. In typical image encoding / decoding methods, a decoder-side motion information derivation method may be limitedly used. Therefore, the improvement of encoding efficiency attributable to the decoder-side motion information derivation method may also be limited. In embodiments, a motion information search method used in an inter-prediction mode, an intra block copy mode and an intra template matching prediction mode is disclosed. With the use of various motion search methods, encoding efficiency may be improved.
Owner:ELECTRONICS & TELECOMM RES INST

Layered multi-context space adaptive image compression method

The invention relates to the technical field of image compression, in particular to a hierarchical multi-context space adaptive image compression method. The method comprises the steps of obtaining a training set and a test set; constructing a spatial adaptive feature modulation network, wherein the spatial adaptive feature modulation network comprises an encoder, a decoder, a super-prior encoder and a super-prior decoder; constructing a multi-context joint entropy estimation model, wherein the multi-context joint entropy estimation model comprises a local context model, a channel context model, an in-layer anchor point context model and a cross-layer global context model; obtaining an image compression network in combination with the spatial adaptive feature modulation network and the multi-context joint entropy estimation model; respectively training and testing an image compression network by using the training set and the test set to obtain a trained image compression network; and inputting a to-be-processed image into the trained image compression network to obtain a reconstructed image. According to the invention, the image compression effect is improved.
Owner:HENAN UNIVERSITY

Low-delay video stream real-time processing method and device

The invention relates to the technical field of computer video processing, and discloses a low-delay video stream real-time processing method and device, and the method comprises the steps: obtaining original video stream data, and processing the original video stream data through employing a lightweight motion prediction method; processing the macro block data set and the predicted coding configuration parameter by adopting multi-thread assembly line coding to obtain a coded data block; establishing a data transmission mechanism to perform data flow control on the unified memory access interface; a heterogeneous task scheduling strategy is adopted to distribute task division results; a lightweight neural network is adopted to carry out parameter adaptive adjustment, and an optimized video stream processing result is obtained; according to the method, a zero-copy data transmission technology is adopted, and optimal configuration and efficient utilization of computing resources are achieved.
Owner:HUNAN BEICHUANG INTELLIGENT TECHNOLOGY CO LTD

Template based CCLM / MMLM slope adjustment

Systems, methods, and instrumentalities are disclosed for template based cross component linear model / multimode linear model (CCLM / MMLM) adjustment. In an example, a device, such as a video decoding device, or a video encoding device, may obtain a prediction model for predicting a coding block. The device may select an adjustment model, from multiple adjustment models, for adjusting the prediction model. The device may adjust the prediction model based on the selected adjustment model. The device may process (e.g., encode and / or decode) the coding block based on the adjusted prediction model.
Owner:INTERDIGITAL CE PATENT HOLDINGS SAS

Dynamic mesh geometry refinement component adaptive coding

Computer-implemented methods and systems for processing geometry replacements are disclosed. The methods include decoding / encoding a syntax element associated with a coding mode from / into a bitstream associated with geometry displacements; and reconstructing / converting, based on a coefficient configuration associated with the coding mode, a plurality of quantized transform coefficients from / to a plurality of zero-run length codes.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Extremely compressed video coding method based on intelligent reference frame

The invention discloses an extreme compressed video coding method based on an intelligent reference frame, comprising the following steps: S1, acquiring an original video stream, dividing the original video stream into frame groups according to a time sequence, each group comprising a current frame and a candidate reference frame; s2, carrying out region positioning on the candidate reference frame, extracting motion, edge and background regions, and generating a mapping graph; s3, calculating a quality score based on the mapping graph according to a motion vector, gradient change and a region overlapping rate; s4, selecting three frames with highest scores to form a reference set, and establishing an index structure; s5, predicting the current frame by using the reference set to generate a predicted frame and a residual error; s6, multi-path coding cost is calculated, and a path with the minimum cost is selected; and S7, entropy coding is carried out on the reference index, the motion vector, the residual error and the control parameter, a code stream is output, and calling information is recorded. According to the method, the compression ratio and prediction precision of video coding are improved, the image quality is maintained while the code rate is reduced, and the method is suitable for efficient transmission and storage of high-resolution videos.
Owner:NINGXIA ANYING INFORMATION TECHNOLOGY SERVICE CO LTD

End-to-end learning-based point cloud coding framework

In one implementation, point cloud data for a point cloud is decoded. The decoder obtains features representing voxels in a tree structure, where feature for a current voxel is representative of at least a set of voxels that are still to be reconstructed. The decoder then determines an occupancy probability of the current voxel based on the feature, and decodes occupancy information of voxels in the tree structure, where whether a current voxel is occupied or not is decoded based on the occupancy probability for the current voxel. The point cloud can be reconstructed based on the occupancy information. On the encoder side, the feature for the current voxel is obtained from the voxels that are still to be encoded and encoded into a bitstream.
Owner:INTERDIGITAL VC HOLDINGS INC

Neural network codec with hybrid entropy model and flexible quantization

Innovations in systems, methods, and software for features of a neural image or video codec are described herein. For example, a neural video encoder can receive a current video frame, encode the current video frame to produce encoded data, and output the encoded data as part of a bitstream. As part of the encoding, the encoder can determine a current latent representation for the current video frame, and encode the current latent representation using an entropy model network that includes one or more convolutional layers. As part of the encoding the current latent representation, the encoder can estimate statistical characteristics of a quantized version of the current latent representation based at least in part on a previous latent representation for a previous video frame, and entropy code the quantized version of the current latent representation based at least in part on the estimated statistical characteristics.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Visual encoding and decoding of 3D gaussian splats

An encoder projects a scene represented by 3D gaussian splats into 2D representation(s), where the 2D representation(s) store different parameters of the 3D gaussian splats, along with associated metadata describing transformation from 3D space into 2D representations. The encoder forms the 2D representation(s) and the associated metadata into bitstream(s). The encoder outputs the bitstreams. The decoder receives the bitstream(s) encoding 2D representation(s) and associated metadata. The decoder applies decoding to corresponding individual bitstreams to form the 2D representation(s) and the associated metadata. The decoder reprojects, as described by the associated metadata, the 2D representations) into the scene represented by the 3D gaussian splats. The decoder outputs the scene for rendering, or renders the scene, to a viewer.
Owner:NOKIA TECHNOLOGIES OY

End-to-end learning-based dynamic point cloud coding framework

Some embodiments of a method may include: decoding a motion feature by accessing a motion bitstream; predicting a predicted feature based on the motion feature and one or more reference point cloud frames; decoding a first feature representing an occupancy status of a child level voxel; predicting a second feature based on the first feature and the predicted feature; and decoding a tree voxel occupancy status of the child level voxel via the second feature.
Owner:INTERDIGITAL VC HOLDINGS INC