Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

220 results about "Visual media" patented technology

“Visual Media” is a colloquial expression used to designate things like TV, movies, photography, painting and so on . But it is highly inexact and misleading. All the so called visual media turn out, on closer inspection, to involve the other senses (especially touch and hearing.)

Utilizing interactive deep learning to select objects in digital visual media

Systems and methods are disclosed for selecting target objects within digital images utilizing a multi-modal object selection neural network trained to accommodate multiple input modalities. In particular, in one or more embodiments, the disclosed systems and methods generate a trained neural network based on training digital images and training indicators corresponding to various input modalities. Moreover, one or more embodiments of the disclosed systems and methods utilize a trained neural network and iterative user inputs corresponding to different input modalities to select target objects in digital images. Specifically, the disclosed systems and methods can transform user inputs into distance maps that can be utilized in conjunction with color channels and a trained neural network to identify pixels that reflect the target object.
Owner:ADOBE INC

Attitude trajectory optimization method and system based on adaptive sliding window

The invention provides an attitude trajectory optimization method and system based on a self-adaptive sliding window. The method comprises the following steps: acquiring visual media data and decoding the visual media data into a time sequence image frame sequence; performing attitude estimation on each frame of image to obtain a coordinate set of each attitude feature point; caching a coordinate set of the historical attitude feature points to form a variable-length dynamic trajectory cache queue; constructing a sliding window according to the coordinate sets of the attitude feature points of the current frame and the previous frame; performing abnormal data elimination on each coordinate track; performing B spline fitting on each coordinate track; sequentially executing multi-stage optimization on the coordinate track; and updating the sliding window frame by frame to obtain a sequential sequence of the optimized attitude feature points. The method has the beneficial effects that the window radius is dynamically adjusted based on the motion intensity, the window is expanded at low speed to enhance noise suppression, and the window is reduced at high speed to reduce delay; and nonparametric regression fitting is carried out by adopting a cubic B spline, so that the modeling capability of a nonlinear motion track is enhanced, and the track smoothness is improved.
Owner:SHANGHAI BAISHU YOUFANG EDUCATIONAL EQUIPMENT CO LTD

Indications of processing orders of post-processing filters

A mechanism for processing video data is disclosed. The mechanism includes determining to signal a processing order or a preferred processing order of different post-processing filters, including zero or more neural-network post-filters (NNPFs) and zero or more non-NNPF post-processing filters, in a supplemental enhancement information (SEI) processing order SEI message. A conversion is performed between a visual media data and a bitstream based on the SEI processing order SEI message.
Owner:BYTEDANCE INC

Advanced bilateral filter in video coding

A mechanism for processing video data is disclosed. The mechanism determines to apply a bilateral filter and a cross component sample adaptive offset (CCSAO) filter to samples in a current block of a current picture. The bilateral filter includes filter weights that vary based on a distance between surrounding samples and a central sample and differences in intensities of the surrounding samples and the central sample. A conversion is performed between a visual media data and a bitstream based on the bilateral filter and the CCSAO filter.
Owner:DOUYIN VISION CO LTD +1

Conditional filter shape switch for adaptive loop filter in video coding

A mechanism for processing video data is disclosed. The mechanism includes determining to use at least one extended tap in an adaptive loop filter (ALF). A conversion can then be performed between a visual media data and a bitstream based on the ALF. The ALF may also employ a conditional filter shape switch.
Owner:BYTEDANCE INC +1

Neural-network post-filter on value ranges and coding methods of syntax elements

A mechanism for processing video data is disclosed. The mechanism includes determining that a value of a neural-network post-filter characteristics (NNPFC) input format indicator (nnpfc_inp_format_inc) is coded as a u(N) coded syntax element where N is an integer greater than 0. A conversion is performed between a visual media data and a bitstream based on the NNPFC input format indicator.
Owner:BYTEDANCE INC

Neural-network post-filter characteristics SEI message and the neural-network post-filter activation SEI message

A mechanism for processing video data is disclosed. The mechanism includes determining that nnpfc_absent_input_pic_zero_flag equal to 1 indicates that a neural-network post-filter (NNPF) expects an input picture that is not present in a candidate input picture list to be represented by sample arrays with sample values equal to 0. A conversion is performed between a visual media data and a bitstream based on the NNPFC SEI message.
Owner:BYTEDANCE INC

Refining item descriptions and recommendations using visual media inputs

Technologies are described herein for refining, using visual media inputs, generative artificial intelligence (AI) model outputs that include item descriptions. In some implementations, a method includes receiving, from a first device, a component list for an item that is to be included in a menu of items, the list including multiple components. Using a first generative AI model, a text natural language response is generated that includes a description for the item based on the component list. The text natural language response and visual media data of the item are provided to a second generative AI model that modifies the description for the item in the text natural language response based on detection of at least one component in the visual media data. The modified description is provided to the first device for inclusion in the menu of items.
Owner:BLOCK INC

Efficient neural network architecture for loop filtering in video coding

A mechanism for processing video data is disclosed. The mechanism determines to process video information with a high operating point (HOP) filter that includes a deep convolutional layer or a packet convolutional layer. Conversion is performed between the visual media data and the bitstream based on the HOP filter.
Owner:DOUYIN CO LTD

Boundary strength termination for deblocking filters in video processing

In an exemplary aspect, a method for visual media processing includes determining whether a pair of adjacent blocks of visual media data are both intra block copy (IBC) coded; and selectively applying, based on the determination, a deblocking filter (DB) process by identifying a boundary at a vertical edge and / or a horizontal edge of the pair of adjacent blocks, calculating a boundary strength of a filter, deciding whether to turn on or off the filter, and selecting a strength of the filter in case the filter is turned on, wherein the boundary strength of the filter is dependent on a motion vector difference between motion vectors associated with the pair of adjacent blocks.
Owner:DOUYIN VISION CO LTD +1

Visual media search method and electronic device

The application provides a visual media search method and an electronic device, and relates to the technical field of image processing. After receiving a search statement input by a user, the electronic device determines a text feature vector of the search statement. Then, the electronic device inputs the text feature vector of the search statement and an image feature vector of a visual media stored locally by the electronic device into a text-image matching model, so as to determine whether the search statement matches the visual media by using the text-image matching model, thereby realizing the search of the visual media. The text-image matching model is trained by using positive and negative samples, the positive and negative samples are determined by distinguishing sample text elements corresponding to sample images, the sample text elements corresponding to the sample images include one text element in a text set corresponding to each sample image, the text set corresponding to the sample images represents a set of description contents corresponding to the sample images, the training samples of the text-image matching model are determined quickly, and therefore the training efficiency of the text-image matching model can be improved.
Owner:HONOR DEVICE CO LTD

Monitoring and three dimensional georegistration of visual media published on social media

Monitoring and three dimensional georegistration of visual media published on social media may be provided. First, a search criteria may be received. Then social media flows may be monitored for a combination of video imagery and the search criteria. In response to monitoring the social media flows, the video imagery may be obtained from a social media flow when the social media flow contains the video imagery and the search criteria. A broad matching process may be performed for an area of interest associated with the search criteria. A registration process may be performed for matches of the pairs having ones of the plurality of probability values meeting a predetermined criteria until a matching pair is determined by the registration process.
Owner:MAXAR INT SWEDEN AB

Use of the neural-network post-filter characteristics and neural-network post-filter activation SEI message in a bitstream

A mechanism for processing video data is disclosed. The mechanism includes determining that a variable representing a number of inferences (numInferences) is derived based on whether either of two conditions is true. One of the two conditions is that a corresponding coded picture of a current picture (currPic) is a last picture of the bitstream in output order that has a network abstraction layer unit header layer identifier (nuh_layer_id) equal to a current layer identifier (currLayerId). Another of the two conditions is that a corresponding coded picture of a current picture (currPic) is a last picture in a coded layer video sequence (CLVS) in output order and a neural-network post-filter activation (NNPFA) no following CLVS flag (nnpfa_no_foll_clvs_flag) is equal to 1. A conversion is performed between a visual media data and a bitstream based on the number of inferences.
Owner:BYTEDANCE INC

Identifying a digital watermark in an image / video / audio stream where the image has been converted to a different format

A visual media is received. For example, the received visual media may be a digital image, a video file, or a video stream. A plurality of colors in the visual media are identified. In response to identifying the plurality of colors in the visual media, one or more colors not in the visual media are identified. A watermark is placed in the visual media to produce a watermarked visual media. The watermark comprises at least one of the identified colors not in the visual media. The watermarked visual media is verified using image processing.
Owner:MICRO FOCUS LLC

Enhancement of supplemental enhancement information signaling in video bitstream

A mechanism for processing video data is disclosed. The mechanism includes determining one or more bytes of data from a supplemental enhancement information (SEI) original byte sequence payload (RBSP) header, wherein the SEI RBSP header is included in the SEI RBSP and precedes one or more SEI messages carried in the SEI RBSP. Conversion between the visual media data and the bitstream is performed based on the SEI RBSP header.
Owner:DOUYIN CO LTD

Handling of a processing chain with processing stage 0 being film grain processing

A mechanism for processing video data is disclosed. The mechanism includes determining to crop a processed picture to obtain a cropped processed picture in a same manner as a decoded output picture is cropped from a corresponding decoded picture and to replace a picture in a candidate input picture list with the cropped processed picture when a supplemental enhancement information (SEI) message is a film grain characteristics (FGC) SEI message and has an index value of a processing order supplemental enhancement information (SEI) type. A conversion is performed between a visual media data and a bitstream based on the candidate input picture list.
Owner:BYTEDANCE INC

Signaling enhancement of SEI processing order in video bitstream

A mechanism for processing video data is disclosed. The mechanism processes sequential SEI messages for supplemental enhancement information (SEI), determines that an SEI prefix indication (in the presence) is signaled in units of bits, where the SEI prefix indication of a particular SEI payloadType is a bit string following an SEI load syntax corresponding to a value of the payloadType, and the SEI prefix indication of the payloadType corresponds to the SEI load syntax of the value of the payloadType. And the bit string contains a number of complete syntax elements starting from the first syntax element in the SEI load. A conversion between the visual media data and the bitstream is performed based on the SEI processing sequence SEI message.
Owner:DOUYIN CO LTD

CONTEXT CODING FOR TRANSFORM JUMP MODE

Devices, systems, and methods for coefficient coding in transform-hopping mode are described. An example method for video processing includes determining, for encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, a maximum allowable dimension up to which a current video block of the one or more video blocks is permitted to be encoded using a transform-hopping mode so that a residue of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform; and including a syntax element indicative of the maximum allowable dimension in the bitstream representation.
Owner:BYTEDANCE INC

Value range of neural network post-processing filter about syntax element and encoding and decoding method

A mechanism for processing video data is disclosed. The mechanism includes determining that a value of a neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfcinformatinc) is within a range from 0 to N (including a boundary value), where N is a positive integer. Conversion between the visual media data and the bitstream is performed based on the NNPFC input format indicator.
Owner:DOUYIN VISION CO LTD

Bitstream conformance constraints for intra block copy in video coding

A method of processing visual media includes performing a conversion between a current video block of a current picture of a visual media data and a bitstream representation of the visual media data using a buffer comprising reference samples from the current picture for derivation of a prediction block of the current video block. The conversion is based according to rule which specifies that, for the bitstream representation to conform the rule, a reference sample in the buffer is to satisfy a bitstream conformance constraint.
Owner:BYTEDANCE INC +1

Value range of neural network post-processing filter about syntax element and encoding and decoding method

A mechanism for processing video data is disclosed. The mechanism includes determining a value of a neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfcinformatinc) to be encoded as a syntax element for u (N) encoding, where N is an integer greater than 0. Conversion between the visual media data and the bitstream is performed based on the NNPFC input format indicator.
Owner:DOUYIN CO LTD

Adaptive scaling filter selection in dynamic resolution coding

A mechanism for processing video data is disclosed. The mechanism includes determining to select one or more adaptive scaling filters. A conversion is performed between a visual media data and a bitstream based on the adaptive scaling filters.
Owner:BYTEDANCE INC

Cross component model for loop-filters in video coding

A mechanism for processing video data is disclosed. The mechanism includes determining to apply offset scaling in an adaptive loop filter (ALF) when parameters are reused from an adaptation parameter set (APS) in a bitstream. A conversion is performed between a visual media data and the bitstream based on the ALF.
Owner:DOUYIN VISION CO LTD +1

Specifying processing stages and picture-level SEI messages for a processing chain

A mechanism for processing video data is disclosed. The mechanism includes determining that, for a picture, a processing order supplemental enhancement information (SEI) list (PoSeiList) contains a list of SEI messages associated with SEI message types in a processing chain indicated by an SEI processing order (SPO) SEI message that may be applied to the picture. A conversion is performed between a visual media data and a bitstream based on the SPO SEI message.
Owner:BYTEDANCE INC

User interfaces for altering visual media

The present disclosure generally relates to user interfaces for altering visual media. In some embodiments, user interfaces capturing visual media (e.g., via a synthetic depth-of-field effect), playing back visual media (e.g., via a synthetic depth-of-field effect), editing visual media (e.g., that has a synthetic depth-of-field effect applied), and / or managing media capture.
Owner:APPLE INC

Subpicture track referencing and handling

This application relates to sub-picture track referencing and processing. Systems, methods, and apparatuses for processing visual media data are described. An example method includes performing a conversion between visual media data and a visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the visual media file comprises a base track referencing one or more sub-picture tracks storing coded information of one or more sub-pictures of the visual media data, and wherein the format rule specifies a process for reconstructing a video unit from samples in the one or more sub-picture tracks and the base track.
Owner:FACE CUTE CO LTD

Visual media data backup method and device, electronic equipment and storage medium

The invention relates to the technical field of visual media data backup, and discloses a visual media data backup method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining the information of each high-speed write-in port and a common port through main equipment, and interacting with a specified terminal through a user-defined protocol, and then, on the basis of a write-in scheme generated by the reading instruction, a personalized processing scheme is formulated for each piece of to-be-backed-up data to ensure that the data is backed up. The method has the advantages that the data backup and transmission efficiency is improved through the user-defined protocol, the increasing image and video data requirements are met, compared with a traditional manual backup mode, the automatic data processing flow is provided, and the data loss or damage risk caused by manual operation is remarkably reduced.
Owner:SHENZHEN ZHONGXIN CHUANGZHAN TECH CO LTD

Inter prediction in region adaptive hierarchical transform coding

A mechanism for processing video data is disclosed. In one example, when utilizing regional adaptive hierarchical transform (RAHT) in geometry-based point cloud compression (G-PCC), the mechanism includes determining to disable alternating current (AC) inter prediction based on direct current (DC) of a reference node, DC of a current node, and one or more thresholds. Conversion between the visual media data and the bitstream can then be performed with AC inter prediction disabled.
Owner:DOUYIN VISION CO LTD +1

Improvements on the SEI processing order SEI message

A mechanism for processing video data is disclosed. The mechanism includes determining a supplemental enhancement information (SEI) processing order (SPO) SEI message is included in a bitstream, wherein the SPO SEI message includes a mth processing order SEI payload type (po_sei_payload_type[ m ]) and a processing order number of SEI messages minus 2 plus 1 (po_num_sei_messages_minus2 + 1), and wherein for any integer value of m in the range of 0 to po_num_sei_messages_minus2 + 1, inclusive, when po_sei_payload_type[ m ] indicates a region-wise packing (RWP) SEI message, there shall be a different integer value of n in a same range such that po_sei_payload_type[ n ] indicates an SEI message that indicates a type of projection. A conversion is performed between a visual media data and the bitstream based on the SPO SEI message.
Owner:BYTEDANCE INC