Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

553 results about "Visual media" patented technology

“Visual Media” is a colloquial expression used to designate things like TV, movies, photography, painting and so on . But it is highly inexact and misleading. All the so called visual media turn out, on closer inspection, to involve the other senses (especially touch and hearing.)

Neural-network post-filter purposes with downsampling capabilities

A method for processing media data is disclosed. In an embodiment, the method includes determining a value of a neural-network post-filter characteristics (NNPFC) purpose (nnpfc_purpose) in a NNPFC supplemental enhancement information (SEI) message, wherein the nnpfc_purpose is configured to be set to include output picture size downsampling, and wherein at least one of (1) an output picture width is not equal to an input picture width and (2) an output picture height is not equal to an input picture height; and performing a conversion between a visual media data and a bitstream based on the nnpfc_purpose.
Owner:BYTEDANCE INC

User interfaces for generating automatically-generated content

In some embodiments, an electronic device generates an automatically-generated visual media using one or more recognized concepts extracted from a prompt inputted by a user. The recognized concepts include personalized template subjects and / or prompt suggestions. While displaying the user interface including the recognized concepts, the electronic device receives one or more inputs to modify the recognized concepts. The electronic device generates multiple variants of the automatically-generated visual content using the one or more recognized concepts. The electronic device adds an automatically-generated visual content to a content entry field of an application, different than the automatically-generated visual media application, without opening the automatically-generated visual media application. The electronic device applies a visual effect to content that is generated using an artificial intelligence model. The electronic device displays visual information corresponding to an artificial intelligence model. The electronic device displays an animation including displaying a user interface with high dynamic range luminance.
Owner:APPLE INC

Prompt editor for use with a visual media generative response engine

The present technology pertains to a prompt editor for use with a visual media generative response engine, where a user inputs a text prompt describing visual media to be generated by the visual media generative response engine. Upon receiving a command to generate the visual media, the present technology determines at least one of the duration, resolution, or aspect ratio for the media prior to generation. The visual media generative response engine creates the visual media based on the input prompt and the specified and determined characteristics. The generated visual media is received having the specified attributes.
Owner:OPENAI OPCO LLC

Generative video engine capable of outputting videos in a variety of durations, resolutions, and aspect ratios

The present technology pertains to a visual media generative response engine that can create visual media from prompts. The visual media generative response engine can generate visual media in a variety of durations, aspect ratios, and resolutions. Further, the visual media generative response engine is capable of receiving both visual media and text as prompts. Additionally, the present technology pertains to a variety of user interfaces to enable more influence over the output of the visual media generative response engine.
Owner:OPENAI OPCO LLC

Methods and processors for executing adaptive frame generation

A method and processor for real-time frame processing of a sequence of frames, focusing on motion vector determination, decision metric calculation, and selective triggering of modes for frame processing. It involves calculating motion vectors and their metrics based on pixel displacement and gradient magnitude between consecutive frames to decide on the most appropriate processing mode-copying, generating via neural networks, or rendering using GPUs. The method adapts over time by adjusting thresholds for motion vector metrics, employing decay functions to refine the decision-making process for subsequent frames. This adaptive approach aims to optimize frame generation and rendering quality in dynamic sequences, providing a sophisticated method for managing and enhancing visual media in real-time applications. The processor is configured to execute these steps, adjusting its operations as it processes more frames, ensuring efficient and high-quality visual outputs.
Owner:HUAWEI TECH CO LTD

Visual media-based multimodal chatbot

Example embodiments of the present disclosure relate to a visual media-based multimodal chatbot. According to example embodiments, a method for operating a multimodal chatbot may include receiving a user input via a chatbot interface. The user input may include at least one of: a text, an audio, a first image, and a first video. The method may further include obtaining a visual media associated with the user input. The visual media may include at least one of: a second image, a second video, and an avatar associated with a person. The method may further include outputting the visual media via the chatbot interface.
Owner:YONUX LLC

Generating interactive vehicle inspection interfaces using multi-model artificial intelligence and anchor-based spatial tracking

A system and method for generating an interactive user interface for inspection visualization. The system includes multiple imaging devices positioned along a inspection passage and at least one processor that executes instructions to: obtain multiple sets of images of vehicle surface segments captured during relative movement between the vehicle and imaging devices; stitch the images into a dataset record mapping vehicle parts and surface anomalies; transform the image data into a moving visual media object using a first generative AI model; compute a mapping record between segmented vehicle parts and target frame areas; and transform the mapping record and visual media object into an interactive interface using a second generative AI model. The interface displays user-selectable markers synchronized with media playback, indicating anomaly locations from multiple viewing angles, and performs data retrieval and display actions based on user selection of anomalies.
Owner:UVEYE LTD

Method and apparatus processing visual media

An apparatus and a method for processing an image are provided. The method includes obtaining one or more frames from at least one input visual media, for at least one frame of the one or more frames, detecting one or more features from the at least one frame based on feature detection model, determining at least one cropping window based on the one or more detected features and information regarding an aspect ratio of a display, obtaining one or more cropped frames based on the at least one cropping window, selecting one or more overlays based on one or more cropped out features, text, picture-in-picture display, and spaces left in the display, and generating one or more reframed frames by situating one or more selected overlays on the one or more cropped frame.
Owner:SAMSUNG ELECTRONICS CO LTD

Jointly coding of texture and displacement data in dynamic mesh coding

A mechanism for processing video data is disclosed. The mechanism includes determining that the texture data and the displacement data are included in a single bitstream and use different coding methods. A conversion is performed between the visual media data and the single bitstream based on the different coding methods of the texture data and the displacement data.
Owner:BYTEDANCE INC +1

Signaling of neural-network post-filter output picture resolution

A method, apparatus, and system for processing media data are disclosed. An example method for processing media data includes obtaining a first parameter or a second parameter used to determine a ratio of a neural-network post-filter characteristics (NNPFC) picture relative to a cropped width or a cropped width, where a value of the ratio is constrained to a range with an endpoint, and where the endpoint of the range is based on a value of 16; and performing a conversion between a visual media data and a bitstream based on the ratio.
Owner:BYTEDANCE INC

Cross random access point signaling in video coding

A mechanism for processing video data is disclosed. An indication is signaled. The indication indicates whether a picture following a dependent random access point (DRAP) picture in decoding order and preceding the DRAP picture in output order is permitted to refer to a reference picture positioned prior to the DRAP picture in decoding order for inter prediction. A conversion is performed between a visual media data and a bitstream based on the indication.
Owner:DOUYIN VISION CO LTD +1

Signaling in transform skip mode

Devices, systems and methods for coefficient coding in transform skip mode are described. An exemplary method for visual media processing includes: for encoding a current video block in a video region of a visual media data into a bitstream representation of the visual media data, identifying usage of a coding mode and / or an intra prediction mode and / or a set of allowable intra prediction modes; and upon identifying the usage, making a decision of whether to include or exclude, in the bitstream representation, a syntax element indicative of selectively applying a transform skip mode to the current video block, wherein, in the transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transformation.
Owner:BYTEDANCE INC

Minimizing initialization delay in live streaming

A method for processing media data includes identifying in a media presentation description (MPD) an indication of a Tuning-In Media Segment. The Tuning-In Media Segment comprises a latest media data for a client device to start with when tuning into an ongoing live streaming service. The latest media data is selected from either a current media segment that is being generated by the ongoing live streaming service or a previous media segment generated by the ongoing live streaming service based on a length of the current media segment. The MPD is stored by the client device prior to the client device tuning into the ongoing live streaming service. The method further includes performing a conversion between a visual media data and a bitstream according to the MPD.
Owner:BYTEDANCE INC +1

Adaptive filter for decoder-side intra mode derivation

Methods and apparatuses for video decoding and video encoding and a method of processing visual media data are disclosed. The apparatus for video decoding includes processing circuitry that receives coded information indicating that a current block in a current picture is coded with a decoder-side intra mode derivation (DIMD) mode. A template of the current block includes reconstructed samples in the current picture and is adjacent to the current block. The template includes one of a left template and a top template. The processing circuitry determines a filter type from a plurality of filter types associated with the one of the left template and the top template, applies the DIMD mode to the template based on the determined filter type to determine one or more intra prediction modes for the current block, and reconstructs the current block according to the one or more intra prediction modes.
Owner:TENCENT AMERICA LLC

Multi-source based extended taps for adaptive loop filter in video coding

A mechanism for processing video data is disclosed. The mechanism includes determining to apply an adaptive loop filter (ALF) to a first component of a video unit. The ALF includes one or more extended taps. The one or more extended taps utilize an input source other than spatial neighbor samples of the first component. A conversion is performed between a visual media data and a bitstream based on the ALF.
Owner:DOUYIN VISION CO LTD +1

Utilizing interactive deep learning to select objects in digital visual media

Systems and methods are disclosed for selecting target objects within digital images utilizing a multi-modal object selection neural network trained to accommodate multiple input modalities. In particular, in one or more embodiments, the disclosed systems and methods generate a trained neural network based on training digital images and training indicators corresponding to various input modalities. Moreover, one or more embodiments of the disclosed systems and methods utilize a trained neural network and iterative user inputs corresponding to different input modalities to select target objects in digital images. Specifically, the disclosed systems and methods can transform user inputs into distance maps that can be utilized in conjunction with color channels and a trained neural network to identify pixels that reflect the target object.
Owner:ADOBE INC

Inter-prediction on non-dyadic blocks

A mechanism for processing video data implemented by a video coding apparatus is disclosed. The mechanism determines whether a block is dyadic or non-dyadic. The mechanism also enables a coding tool associated with inter prediction when the block is determined to be dyadic. The mechanism also disables the coding tool when the block is determined to be non-dyadic. A conversion between a visual media data and a bitstream is performed by applying inter prediction to the block.
Owner:BYTEDANCE INC

Attitude trajectory optimization method and system based on adaptive sliding window

The invention provides an attitude trajectory optimization method and system based on a self-adaptive sliding window. The method comprises the following steps: acquiring visual media data and decoding the visual media data into a time sequence image frame sequence; performing attitude estimation on each frame of image to obtain a coordinate set of each attitude feature point; caching a coordinate set of the historical attitude feature points to form a variable-length dynamic trajectory cache queue; constructing a sliding window according to the coordinate sets of the attitude feature points of the current frame and the previous frame; performing abnormal data elimination on each coordinate track; performing B spline fitting on each coordinate track; sequentially executing multi-stage optimization on the coordinate track; and updating the sliding window frame by frame to obtain a sequential sequence of the optimized attitude feature points. The method has the beneficial effects that the window radius is dynamically adjusted based on the motion intensity, the window is expanded at low speed to enhance noise suppression, and the window is reduced at high speed to reduce delay; and nonparametric regression fitting is carried out by adopting a cubic B spline, so that the modeling capability of a nonlinear motion track is enhanced, and the track smoothness is improved.
Owner:SHANGHAI BAISHU YOUFANG EDUCATIONAL EQUIPMENT CO LTD

Systems and Methods for Generating Simulated Motion from Static Images Using Machine Learning

A system and method for enhancing static photographs to simulate motion using machine learning models. The system comprises one or more processors and a non-transitory computer readable medium storing instructions that, when executed, cause the system to receive a static photograph, apply a machine learning model to generate a sequence of modified images creating an illusion of motion, and display the sequence to simulate a moving video. The machine learning model, trained on a dataset of static photographs and corresponding video sequences, extracts features using a convolutional neural network, processes the features with a recurrent neural network to generate motion vectors, and applies the vectors to create the modified images. The simulated motion may include tilting, vibrating, shaking, zooming, panning, and rotating. A user interface allows specifying the desired type or intensity of motion. The method enables creating video-like effects from static images, enhancing expressiveness and engagement of visual media.
Owner:LAMPERT DAVID +2

Multiple input sources based extended taps for adaptive loop filter in video coding

A mechanism for processing video data is disclosed. The mechanism includes determining to apply an adaptive looper filter (ALF) with an extended tap to a picture in a video. An intermediate filtering result of a second filter is used as input for the extended tap. A conversion is performed between a visual media data and a bitstream based on the ALF.
Owner:BYTEDANCE INC +1

Intra prediction based on extrapolation filter

Aspects of the present disclosure include methods and apparatus for video decoding and video encoding, and methods of processing visual media data. An apparatus for video decoding includes processing circuitry configured to: receive prediction information indicating that a current block in a current picture is predicted using an extrapolation filter-based intra prediction (EIP) mode; determining gradient information associated with a current sample in the current block; determining a prediction value of the current sample based on an initial prediction value predicted using the EIP mode and additional information including gradient information; and reconstructing the current sample according to the predicted value of the current sample.
Owner:TENCENT AMERICA LLC

Presence and relative decoding order of neural-network post-filter SEI messages

A mechanism for processing video data is disclosed. The mechanism includes performing a conversion between a visual media data and a bitstream based on a rule. The rule specifies that a neural-network post-filter activation (NNPFA) Supplemental Enhancement Information (SEI) message with a first particular value of an NNPFA identifier is only present in a current Picture Unit (PU) when one or both of the following conditions are met. First, a current Coded Layer Video Sequence (CLVS) contains a neural-network post-filter characteristics (NNPFC) SEI message with a NNPFC identifier (nnpfc_id) equal to the first particular value of the NNPFA identifier in a preceding PU that precedes the current PU in decoding order. Second, an NNPFC SEI message with nnpfc_id equal to the first particular value of the NNPFA identifier is contained in the current PU.
Owner:BYTEDANCE INC

Switchable input sources based extended taps for adaptive loop filter in video coding

A mechanism for processing video data is disclosed. The mechanism includes determining to apply an extended tap in an adaptive loop filter (ALF). The ALF may also include a spatial tap, and the extended tap may be different from the spatial tap. A conversion is performed between a visual media data and a bitstream based on the ALF.
Owner:DOUYIN VISION CO LTD +1

Fine-grained intra prediction fusion

Aspects of the disclosure includes methods and apparatuses for video decoding and encoding and a method of processing visual media data. The method for video decoding includes receiving coded information in a bitstream indicating that a current block is predicted based on a combination of a plurality of intra prediction modes. The method includes determining a plurality of intra predictions of the current block based on the respective intra prediction modes, determining a fused prediction of the current block based on a weighted summation of the plurality of intra predictions where the weighted summation is according to respective weights associated with the plurality of intra predictions, and reconstructing the current block based on the fused prediction. Each weight is based on one of a plurality of weighting functions that depends on a sample location (x, y) and the intra prediction mode of the intra prediction associated with the respective weight.
Owner:TENCENT AMERICA LLC

Visual media searching method and device and storage medium

The embodiment of the invention provides a visual media searching method and device and a storage medium. The method is suitable for the electronic equipment and comprises the following steps: displaying a first interface; the first interface comprises a search box; receiving a search operation of a search statement input to the search box; according to a first sentence semantic vector of the search statement, determining M candidate visual media files of which the visual semantic vectors are matched with the first sentence semantic vector from a plurality of visual media; the visual semantic vector of each visual media file is obtained by carrying out semantic understanding on an image or an image frame of the visual media file by utilizing a natural picture understanding model; determining N to-be-matched dimensions according to a semantic subject related to the visual content in the search statement; determining a search result of the search statement according to the matching degree of the visual contents of the M candidate visual media files on the N to-be-matched dimensions; and displaying the search result. According to the technical scheme provided by the embodiment of the invention, the search accuracy can be improved, and the user search experience is improved.
Owner:HONOR DEVICE CO LTD

Storyboard graphical user interface to a visual media generative response engine

The present technology pertains to presenting a storyboard user interface that includes a visual media timeline and a representation of a first frame on the timeline along with a prompt to generate visual media. The input prompt is then sent to a visual media generative response engine, which generates output visual media in response to the input prompt.
Owner:OPENAI OPCO LLC

Neural-network post-filter purposes with picture rate upsampling

A mechanism for processing video data is disclosed. The mechanism includes determining a neural-network post-filter (NNPF) purpose based on a neural-network post-filter characteristics (NNPFC) supplemental enhancement information (SEI) message. A conversion is performed between a visual media data and a bitstream based on the NNPF purpose. Multiple input pictures are used for the NNPF purpose, and the NNPF is enabled to selectively generate output pictures for some input picture(s) and not to generate output pictures for other input picture(s).
Owner:BYTEDANCE INC +1

Indications of processing orders of post-processing filters

A mechanism for processing video data is disclosed. The mechanism includes determining to signal a processing order or a preferred processing order of different post-processing filters, including zero or more neural-network post-filters (NNPFs) and zero or more non-NNPF post-processing filters, in a supplemental enhancement information (SEI) processing order SEI message. A conversion is performed between a visual media data and a bitstream based on the SEI processing order SEI message.
Owner:BYTEDANCE INC

Quantization of point cloud attribute transform domain coefficients

A mechanism for processing video data is disclosed. The mechanism includes determining a quantization based on region-adaptive hierarchical transform (RAHT) weight. A conversion is performed between a visual media data and a bitstream based on the quantization.
Owner:BYTEDANCE INC +1

Blending user interface for blending visual media using a visual media generative response engine

The present technology pertains to influencing the blending of two visual media inputs by first receiving them through a prompt editor. A blending interface is presented, displaying at least one frame from each of the first and second input visual media. The blending is adjusted in response to user input by modifying a blend curve that represents the relative influence of the first visual media compared to the second visual media over time.
Owner:OPENAI OPCO LLC