Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

395 results about "Visual media" patented technology

“Visual Media” is a colloquial expression used to designate things like TV, movies, photography, painting and so on . But it is highly inexact and misleading. All the so called visual media turn out, on closer inspection, to involve the other senses (especially touch and hearing.)

User interfaces for generating automatically-generated content

In some embodiments, an electronic device generates an automatically-generated visual media using one or more recognized concepts extracted from a prompt inputted by a user. The recognized concepts include personalized template subjects and / or prompt suggestions. While displaying the user interface including the recognized concepts, the electronic device receives one or more inputs to modify the recognized concepts. The electronic device generates multiple variants of the automatically-generated visual content using the one or more recognized concepts. The electronic device adds an automatically-generated visual content to a content entry field of an application, different than the automatically-generated visual media application, without opening the automatically-generated visual media application. The electronic device applies a visual effect to content that is generated using an artificial intelligence model. The electronic device displays visual information corresponding to an artificial intelligence model. The electronic device displays an animation including displaying a user interface with high dynamic range luminance.
Owner:APPLE INC

Visual media-based multimodal chatbot

Example embodiments of the present disclosure relate to a visual media-based multimodal chatbot. According to example embodiments, a method for operating a multimodal chatbot may include receiving a user input via a chatbot interface. The user input may include at least one of: a text, an audio, a first image, and a first video. The method may further include obtaining a visual media associated with the user input. The visual media may include at least one of: a second image, a second video, and an avatar associated with a person. The method may further include outputting the visual media via the chatbot interface.
Owner:YONUX LLC

Method and apparatus processing visual media

An apparatus and a method for processing an image are provided. The method includes obtaining one or more frames from at least one input visual media, for at least one frame of the one or more frames, detecting one or more features from the at least one frame based on feature detection model, determining at least one cropping window based on the one or more detected features and information regarding an aspect ratio of a display, obtaining one or more cropped frames based on the at least one cropping window, selecting one or more overlays based on one or more cropped out features, text, picture-in-picture display, and spaces left in the display, and generating one or more reframed frames by situating one or more selected overlays on the one or more cropped frame.
Owner:SAMSUNG ELECTRONICS CO LTD

Jointly coding of texture and displacement data in dynamic mesh coding

A mechanism for processing video data is disclosed. The mechanism includes determining that the texture data and the displacement data are included in a single bitstream and use different coding methods. A conversion is performed between the visual media data and the single bitstream based on the different coding methods of the texture data and the displacement data.
Owner:BYTEDANCE INC +1

Signaling in transform skip mode

Devices, systems and methods for coefficient coding in transform skip mode are described. An exemplary method for visual media processing includes: for encoding a current video block in a video region of a visual media data into a bitstream representation of the visual media data, identifying usage of a coding mode and / or an intra prediction mode and / or a set of allowable intra prediction modes; and upon identifying the usage, making a decision of whether to include or exclude, in the bitstream representation, a syntax element indicative of selectively applying a transform skip mode to the current video block, wherein, in the transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transformation.
Owner:BYTEDANCE INC

Minimizing initialization delay in live streaming

A method for processing media data includes identifying in a media presentation description (MPD) an indication of a Tuning-In Media Segment. The Tuning-In Media Segment comprises a latest media data for a client device to start with when tuning into an ongoing live streaming service. The latest media data is selected from either a current media segment that is being generated by the ongoing live streaming service or a previous media segment generated by the ongoing live streaming service based on a length of the current media segment. The MPD is stored by the client device prior to the client device tuning into the ongoing live streaming service. The method further includes performing a conversion between a visual media data and a bitstream according to the MPD.
Owner:BYTEDANCE INC +1

Adaptive filter for decoder-side intra mode derivation

Methods and apparatuses for video decoding and video encoding and a method of processing visual media data are disclosed. The apparatus for video decoding includes processing circuitry that receives coded information indicating that a current block in a current picture is coded with a decoder-side intra mode derivation (DIMD) mode. A template of the current block includes reconstructed samples in the current picture and is adjacent to the current block. The template includes one of a left template and a top template. The processing circuitry determines a filter type from a plurality of filter types associated with the one of the left template and the top template, applies the DIMD mode to the template based on the determined filter type to determine one or more intra prediction modes for the current block, and reconstructs the current block according to the one or more intra prediction modes.
Owner:TENCENT AMERICA LLC

Utilizing interactive deep learning to select objects in digital visual media

Systems and methods are disclosed for selecting target objects within digital images utilizing a multi-modal object selection neural network trained to accommodate multiple input modalities. In particular, in one or more embodiments, the disclosed systems and methods generate a trained neural network based on training digital images and training indicators corresponding to various input modalities. Moreover, one or more embodiments of the disclosed systems and methods utilize a trained neural network and iterative user inputs corresponding to different input modalities to select target objects in digital images. Specifically, the disclosed systems and methods can transform user inputs into distance maps that can be utilized in conjunction with color channels and a trained neural network to identify pixels that reflect the target object.
Owner:ADOBE INC

Attitude trajectory optimization method and system based on adaptive sliding window

The invention provides an attitude trajectory optimization method and system based on a self-adaptive sliding window. The method comprises the following steps: acquiring visual media data and decoding the visual media data into a time sequence image frame sequence; performing attitude estimation on each frame of image to obtain a coordinate set of each attitude feature point; caching a coordinate set of the historical attitude feature points to form a variable-length dynamic trajectory cache queue; constructing a sliding window according to the coordinate sets of the attitude feature points of the current frame and the previous frame; performing abnormal data elimination on each coordinate track; performing B spline fitting on each coordinate track; sequentially executing multi-stage optimization on the coordinate track; and updating the sliding window frame by frame to obtain a sequential sequence of the optimized attitude feature points. The method has the beneficial effects that the window radius is dynamically adjusted based on the motion intensity, the window is expanded at low speed to enhance noise suppression, and the window is reduced at high speed to reduce delay; and nonparametric regression fitting is carried out by adopting a cubic B spline, so that the modeling capability of a nonlinear motion track is enhanced, and the track smoothness is improved.
Owner:SHANGHAI BAISHU YOUFANG EDUCATIONAL EQUIPMENT CO LTD

Multiple input sources based extended taps for adaptive loop filter in video coding

A mechanism for processing video data is disclosed. The mechanism includes determining to apply an adaptive looper filter (ALF) with an extended tap to a picture in a video. An intermediate filtering result of a second filter is used as input for the extended tap. A conversion is performed between a visual media data and a bitstream based on the ALF.
Owner:BYTEDANCE INC +1

Neural-network post-filter purposes with picture rate upsampling

A mechanism for processing video data is disclosed. The mechanism includes determining a neural-network post-filter (NNPF) purpose based on a neural-network post-filter characteristics (NNPFC) supplemental enhancement information (SEI) message. A conversion is performed between a visual media data and a bitstream based on the NNPF purpose. Multiple input pictures are used for the NNPF purpose, and the NNPF is enabled to selectively generate output pictures for some input picture(s) and not to generate output pictures for other input picture(s).
Owner:BYTEDANCE INC +1

Indications of processing orders of post-processing filters

A mechanism for processing video data is disclosed. The mechanism includes determining to signal a processing order or a preferred processing order of different post-processing filters, including zero or more neural-network post-filters (NNPFs) and zero or more non-NNPF post-processing filters, in a supplemental enhancement information (SEI) processing order SEI message. A conversion is performed between a visual media data and a bitstream based on the SEI processing order SEI message.
Owner:BYTEDANCE INC

Quantization of point cloud attribute transform domain coefficients

A mechanism for processing video data is disclosed. The mechanism includes determining a quantization based on region-adaptive hierarchical transform (RAHT) weight. A conversion is performed between a visual media data and a bitstream based on the quantization.
Owner:BYTEDANCE INC +1

Deployment scenarios of content steering in 5g media streaming

There is provided a method and apparatus including computer code to cause a processor or processors to generate, by a trusted distribution network of a mobile network operator (MNO), at least one 5G Media Downlink Streaming instance configured to deliver at least part of a content to a user equipment including a 5GMSd-aware application and a 5GMSd client, the at least part of the content comprising video data, obtain obtaining, by the trusted DN, a presentation manifest published by a 5GMSd Application Provider of an external DN that is external to both the trusted DN and the UE, and perform a conversion, between a visual media file of the content and a bitstream of a visual media data of the content according to a format rule, and at least one of streaming the content to the UE and playing the content by the UE based on the conversion.
Owner:TENCENT AMERICA LLC

Advanced bilateral filter in video coding

A mechanism for processing video data is disclosed. The mechanism determines to apply a bilateral filter and a cross component sample adaptive offset (CCSAO) filter to samples in a current block of a current picture. The bilateral filter includes filter weights that vary based on a distance between surrounding samples and a central sample and differences in intensities of the surrounding samples and the central sample. A conversion is performed between a visual media data and a bitstream based on the bilateral filter and the CCSAO filter.
Owner:DOUYIN VISION CO LTD +1

Conditional filter shape switch for adaptive loop filter in video coding

A mechanism for processing video data is disclosed. The mechanism includes determining to use at least one extended tap in an adaptive loop filter (ALF). A conversion can then be performed between a visual media data and a bitstream based on the ALF. The ALF may also employ a conditional filter shape switch.
Owner:BYTEDANCE INC +1

Neural-network post-filter on value ranges and coding methods of syntax elements

A mechanism for processing video data is disclosed. The mechanism includes determining that a value of a neural-network post-filter characteristics (NNPFC) input format indicator (nnpfc_inp_format_inc) is coded as a u(N) coded syntax element where N is an integer greater than 0. A conversion is performed between a visual media data and a bitstream based on the NNPFC input format indicator.
Owner:BYTEDANCE INC

Neural-network post-filter characteristics SEI message and the neural-network post-filter activation SEI message

A mechanism for processing video data is disclosed. The mechanism includes determining that nnpfc_absent_input_pic_zero_flag equal to 1 indicates that a neural-network post-filter (NNPF) expects an input picture that is not present in a candidate input picture list to be represented by sample arrays with sample values equal to 0. A conversion is performed between a visual media data and a bitstream based on the NNPFC SEI message.
Owner:BYTEDANCE INC

Neural-network post-filter purposes with picture rate upsampling

A mechanism for processing video data is disclosed. The mechanism includes determining a neural-network post-filter (NNPF) purpose based on a neural-network post-filter characteristics (NNPFC) supplemental enhancement information (SEI) message, wherein the NNPF purpose includes two or more types of post-filter operations. A conversion is performed between a visual media data and a bitstream based on the NNPF purpose.
Owner:DOUYIN VISION CO LTD +1

Signaling and filtering of template and flexible partition split

Methods and apparatuses for video decoding and video encoding and methods of processing visual media data are provided. A method for video decoding includes receiving coded information of a current block in a current picture. Whether to apply a filter to a neighboring reconstructed area of the current block is determined. The neighboring reconstructed area is adjacent to the current block and includes reconstructed samples in the current picture. The filter is applied to the neighboring reconstructed area of the current block. The current block is reconstructed according to the filtered neighboring reconstructed area.
Owner:TENCENT AMERICA LLC

Refining item descriptions and recommendations using visual media inputs

Technologies are described herein for refining, using visual media inputs, generative artificial intelligence (AI) model outputs that include item descriptions. In some implementations, a method includes receiving, from a first device, a component list for an item that is to be included in a menu of items, the list including multiple components. Using a first generative AI model, a text natural language response is generated that includes a description for the item based on the component list. The text natural language response and visual media data of the item are provided to a second generative AI model that modifies the description for the item in the text natural language response based on detection of at least one component in the visual media data. The modified description is provided to the first device for inclusion in the menu of items.
Owner:BLOCK INC

Template-based intra mode coding and decoding

Aspects of the present disclosure include methods and apparatus for video decoding and encoding, and methods of processing visual media data. A method for video decoding includes receiving encoded information in a bitstream, the encoded information indicating that a first intra mode coding method is not enabled for a current block, the first intra mode coding method using one of (i) template-based intra prediction mode derivation (TIMD) and (ii) decoder-side intra mode derivation (DIMD). The method for video decoding includes: determining at least one template-based intra mode using a first intra mode coding and decoding method using one of TIMD and DIMD; and when the at least one template-based intra mode is excluded from the second intra mode codec method, determining an intra prediction mode of the current block using the second intra mode codec method excluding the at least one template-based intra mode, and reconstructing the current block using the intra prediction mode.
Owner:TENCENT AMERICA LLC

Video decoder initialization information signaling

A mechanism for processing video data is disclosed. A parent block is partitioned with an Extended Ternary-Tree (ETT) partition to create three sub-blocks. At least one of the sub-blocks includes a side with a measurement that is not a power of two. A conversion is performed between a visual media data and a bitstream based on the sub-blocks.
Owner:DOUYIN VISION CO LTD +1

Efficient neural network architecture for loop filtering in video coding

A mechanism for processing video data is disclosed. The mechanism determines to process video information with a high operating point (HOP) filter that includes a deep convolutional layer or a packet convolutional layer. Conversion is performed between the visual media data and the bitstream based on the HOP filter.
Owner:DOUYIN CO LTD

Video usability information related indications and miscellaneous items in neural-network post-processing filter SEI messages

A mechanism for processing video data is disclosed. The mechanism includes determining to signal a syntax element to indicate whether neural-network post-filter (NNPF) output pictures are in a full range when a color space of the NNPF output pictures and a color space of decoded pictures or a color space of cropped decoded output pictures are different. A conversion is performed between a visual media data and a bitstream based on the NNPF.
Owner:BYTEDANCE INC

Boundary strength termination for deblocking filters in video processing

In an exemplary aspect, a method for visual media processing includes determining whether a pair of adjacent blocks of visual media data are both intra block copy (IBC) coded; and selectively applying, based on the determination, a deblocking filter (DB) process by identifying a boundary at a vertical edge and / or a horizontal edge of the pair of adjacent blocks, calculating a boundary strength of a filter, deciding whether to turn on or off the filter, and selecting a strength of the filter in case the filter is turned on, wherein the boundary strength of the filter is dependent on a motion vector difference between motion vectors associated with the pair of adjacent blocks.
Owner:DOUYIN VISION CO LTD +1

Systems, methods, and apparatuses for dynamic content extraction in visual media content

Various embodiments are directed to apparatuses, methods, computer readable media, computer program products, and systems related to dynamic content extraction in visual media content. In some embodiments the system for dynamic content extraction in visual media content may comprise one or more processors and at least one non-transitory memory comprising instructions that, with the one or more processors, cause the system to receive a segment selection indication associated with visual media content; identify a segment of the visual media content based on temporal indicator associated with the segment selection indication; extract a content data object from at least one portion of the segment of the visual media content; generate a relevance data object based on the content data object; and cause display of the relevance data object to a user.
Owner:ASSURANT INC +1

Visual media search method and electronic device

The application provides a visual media search method and an electronic device, and relates to the technical field of image processing. After receiving a search statement input by a user, the electronic device determines a text feature vector of the search statement. Then, the electronic device inputs the text feature vector of the search statement and an image feature vector of a visual media stored locally by the electronic device into a text-image matching model, so as to determine whether the search statement matches the visual media by using the text-image matching model, thereby realizing the search of the visual media. The text-image matching model is trained by using positive and negative samples, the positive and negative samples are determined by distinguishing sample text elements corresponding to sample images, the sample text elements corresponding to the sample images include one text element in a text set corresponding to each sample image, the text set corresponding to the sample images represents a set of description contents corresponding to the sample images, the training samples of the text-image matching model are determined quickly, and therefore the training efficiency of the text-image matching model can be improved.
Owner:HONOR DEVICE CO LTD

Cross-component prediction and geometric partition weight adaption in geometric partition mode

Methods and apparatuses for video decoding and video encoding and methods of processing visual media data are provided. A method for video decoding includes receiving coded information indicating that a current block is coded with a geometric partition mode (GPM) with multiple blending width sets, determining a blending width set from the multiple blending width sets to be applied to the current block based on block size information and GPM information of the current block, determining a blending width from the determined blending width set, and reconstructing the current block according to the GPM and the determined blending width.
Owner:TENCENT AMERICA LLC