Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

250results about "Using detectable carrier information" patented technology

Processing monocular videos using three-dimensional gaussian splatting

The present disclosure describes techniques for processing monocular videos using three-dimensional gaussian splatting (3DGS). Spatial decomposition and temporal decomposition are performed on a monocular video to generate a plurality of clips. A first set of 3DGS representing foreground objects in each of the plurality of clips are initialized and optimized. A second set of 3DGS representing background in each of the plurality of clips are initialized and optimized. Two images are generated for each frame comprised in each of the plurality of clips based on the first set of 3DGS and the second set of 3DGS, respectively. Two images are merged to generate a resulting image for each frame in each of the plurality of clips. The resulting image accurately represents a corresponding frame in the monocular video.
Owner:LEMON INC(GB)

Audio data selection for video matching using generative artificial intelligence model

A video editing system leverages a generative artificial intelligence (AI) model to identify songs to overlay on a video. The video editing system extracts a set of key frames from the video and prompts the generative AI model to generate a video narrative for the video. A video narrative is a text description of the plot, theme, feel, or other characteristics of the video. The video editing system uses the video narrative to prompt the generative AI model again to generate a set of descriptor tags for the video based on the video narrative. Descriptor tags are strings that represent themes, features, or characteristics of the song. The video editing system uses an audio tagging system to score a set of songs based on the set of descriptor tags and presents a selected subset of the set of songs based on the scores of the songs.
Owner:BEACON STREET TECHNOLOGIES LLC

Systems, devices, and methods for dynamic synchronization of a prerecorded vocal backing track to a live vocal performance

Disclosed are systems, methods, and devices, that overcome timing and self-expression limitations experienced by vocalists when using prerecorded vocal backing tracks to enhance live performances. The disclosed system, devices, and methods, dynamically synchronizes prerecorded vocal backing tracks with a live vocal stream by extracting vocal elements, such as phonemes, vector embeddings, or vocal audio spectra, from the live vocal performance in real-time. These extracted vocal elements are matched against corresponding timestamped vocal elements previously derived from the prerecorded vocal backing track, enabling precise real-time adjustment and alignment of the backing track timing to the live performance. Additionally, the system enhances expressive performance by identifying prosody factors, such as pitch, vibrato, accent, stress, dynamics, and level, in the live vocal performance, and dynamically adjusting corresponding prerecorded prosody factors within predefined ranges. This maintains naturalness and spontaneity in the vocalist's live performance, overcoming traditional limitations associated with prerecorded vocal backing tracks.
Owner:EIDOL CORP

Systems and methods for processing video elements

Described herein is a computer implemented method for automatically generating a trimmed video clip from a video content item. The method includes: receiving a trim request from a user device, the trim request including the video content item; determining trim parameters for the trimmed video clip, the trim parameters including a trim start time and a trim end time; generating the trimmed video clip based on the trim parameters; and causing display of the trimmed video clip on the user device.
Owner:CANVA PTY LTD

Contextual advertising through multimodal content analysis

A system and method for contextual advertising that analyzes video content through multimodal examination of visual, audio, and textual elements to create detailed contextual understanding of individual scenes. The system segments video content into discrete scenes and simultaneously processes each scene to extract contextual characteristics including objects, settings, dialogue, music, and emotional tone. These characteristics are classified according to advertising industry taxonomies and converted into numerical embeddings that enable semantic similarity matching. During video playback, when advertisement opportunities occur, the system identifies the current scene context, analyzes available advertisements using similar techniques, computes similarity scores between scene and advertisement characteristics, and selects contextually appropriate advertisements for seamless integration. This approach enables privacy-compliant advertising that matches advertisement content with scene context rather than relying solely on user behavioral data, improving advertisement relevance and viewer experience.
Owner:TUBI INC

Artificial intelligence-based video content creation with predetermined styles

A method generates AI-based video content with a style by capturing scenes in various formats, applying alterations, and training AI with feedback for authenticity. A system includes processors and memory to capture scenes, apply alterations, construct datasets, and train AI for generating styled video content. A non-transitory computer-readable medium has instructions for capturing scenes, applying post-production alterations, and training AI to generate video content with a predetermined style.
Owner:INTERPOSITIVE LLC

Automatically cropping of landscape videos

Example systems, computer readable medium and methods for automatically generating portrait videos from landscape videos include receiving a landscaped video, identifying one or more objects to track, and tracking the one or more objects by moving a cropping window. A user interface is presented that provides options for a user to automatically convert a landscape video into a portrait video. The options include selecting one or more objects in the landscape video to track and tracking based on a flight plan used to capture the video. The automatic tracking uses a hierarchy of a person for tracking where a face is used if the person is close, an upper body is used if the person is farther away, and a whole body is used if the person is even farther away. A hierarchy of a person includes indications of which parts of the hierarchy to exclude first.
Owner:SNAP INC

Systems and methods for generating video content using natural language

A computer implemented method for generating video content based on natural language input is disclosed. The method includes receiving a natural language instruction describing one or more desired characteristics of a video. A structured script file comprising at least one story beat is generated using a natural language processing engine. A storyboard comprising one or more storyboard frames is created based on the structured script file. One or more virtual components are generated based on the storyboard. An intermediate video sequence comprising a visual component and an auditory component is created using virtual components and the storyboard. The intermediate video sequence is then refined to produce a modified video sequence by applying one or more post-processing effects.
Owner:RITUAL ADS INC

Systems, methods, and user interface for navigating media playback using scrollable text

A mobile computing device can be configured, with an improved user interface, to synchronously play audio (or video) and text associated thereto, such as text stored in a synchronization index. Using the synchronization index, the device can periodically compare the current track time with that time associated to a word or range of words, such as a line (or segment) in a plurality of lines of text (or segments of text). Improved navigability of content using an improved mobile computing device and user interface is provided, because a user of the device can scroll through the lines of text associated with the audio or video to find a target word or range of words. If the user selects a particular word or range of words, by making a gesture on the mobile computing device, the device can identify a start time for the selected text. The device can then play the audio or video at the identified start time of the selected text. The improved user interface is a practical application for navigating audio (or video) and text associated thereto on a mobile computing device, providing bimodal reading on mobile computing devices and ease of navigability.
Owner:EVANS CURTIS

Audio alignment systems and techniques

A system is configurable to: after receiving first user input, (i) initiate playback of selected audio content at one or more playback components and (ii) initiate audio recording at one or more recording components to obtain recorded audio content, the recorded audio content being recorded at least partially during the playback of the selected audio content; process the selected audio content and the recorded audio content as inputs to an alignment module to generate a temporal offset value, wherein the alignment module is configured to generate the temporal offset value by correlating feature frames of the selected audio content with feature frames of the recorded audio content; and after receiving second user input, initiate synchronized playback of the selected audio content and the recorded audio content, wherein the recorded audio content is synchronized with the selected audio content using the temporal offset value.
Owner:MOISES SYSTEMS INC

Long form video to short clips

A method, apparatus, non-transitory computer readable medium, and system for generating a video clip includes obtaining an input video and a transcript of the input video, wherein the transcript comprises a plurality of sentences. In some cases, a language model is configured to generate a subset of the plurality of sentences based on the transcript. Additionally, the language model generates a video clip based on the input video and a subset of the plurality of sentences, wherein the video clip comprises a portion of the input video.
Owner:ADOBE INC

Machine-Learned Model for Generating an Output Based on Image Frames Adaptively Extracted from a Video

A computing device for generating content includes one or more memories to store instructions and one or more processors to execute the instructions to perform operations, the operations including: receiving a video; receiving an input prompt associated with the video; processing the video by adaptively extracting a plurality of image frames from the video at irregular intervals, based on content of the video; and implementing one or more machine-learned models to generate an output responsive to the input prompt, based on the input prompt and the plurality of image frames adaptively extracted from the video.
Owner:GOOGLE LLC

Machine narration

A technique for generating and inserting voice narration about action in audio-video (AV) content such as a movie or computer game includes generating the narration, e.g., from dialog in the AV content, and determining how and when to insert portions of the narration into the AV content so as not to interfere with dialog in the content.
Owner:SONY INTERACTIVE ENTERTAINMENT LLC

Performance recording data based on temporal analysis

Systems and methods are described for creating dynamic displays of live events in real-time. A real-time display of a live performance will typically have a static presentation because there is no time for video editing to take place before the event is displayed to an end user. This static display may be considered less entertaining to viewers and cause them to view an event as uninteresting. The present disclosure enables a system to analyze the live event in real-time detect objects and people within the frame, associate sounds from the recording with the detected people and objects and generate video commands on the fly to implement on the video as it is displayed to an end user. This will enhance the presentation and make it more dynamic. There is a manual override for the generated commands if a user prefers to watch the more static display.
Owner:NAJAFI HAMID

Automatic generation of clips of captured video of an event

A device obtains video of an event captured by an image capture device and detects one or more objects within frames of the video. Tracking data for detected objects across frames of the video is also generated for detected objects. Based on the tracking data, the device generates a graph representing detections of objects in different frames and selects an optimal path traversing the graph. The device selects a set of key frames based on nodes along the optimal path and applies one or more reframing methods to the set of key frames to generate a clip comprising a subset of the video. The clip may be distributed to user devices or to a backend server for distribution.
Owner:KLUTCHSHOTS INC

Method for analyzing dialogue content

The present disclosure improves techniques for analyzing dialog content. A method executed by an information processing device (1) provided with a control unit (10), an imaging unit (11), and an input unit (12), in which the control unit (10) executes an operation including: acquiring, via the input unit (12), an agreement of a customer with respect to a recording corresponding to the customer; after the agreement is obtained, recording is started; a worker detection unit (11) that detects a worker from the image of the imaging unit (11); and interrupting the recording when a predetermined condition is satisfied, the predetermined condition including a first condition in which the worker disappears from the image of the imaging unit (11).
Owner:TOYOTA JIDOSHA KK

Apparatus and method for dynamic range transforming of images

An image processing apparatus comprises a receiver (201) for receiving an image signal which comprises at least an encoded image and a target display reference. The target display reference is indicative of a dynamic range of a target display for which the encoded image is encoded. A dynamic range processor (203) generates an output image by applying a dynamic range transform to the encoded image in response to the target display reference. An output (205) then outputs an output image signal comprising the output image, e.g. to a suitable display. The dynamic range transform may furthermore be performed in response to a display dynamic range indication received from a display. The invention may be used to generate an improved High Dynamic Range (HIDR) image from e.g. a Low Dynamic Range (LDR) image, or vice versa.
Owner:KONINKLIJKE PHILIPS NV

Machine learning model training with parameter variation in filmmaking

A method trains machine learning models for filmmaking by varying a single camera parameter, capturing shots, and processing these to adjust the Al model. A system includes processors and memory to perform single-parameter variation, capture and process shots, and enhance Al model understanding. A non-transitory computer-readable medium has instructions for configuring a multi-camera environment, varying a camera parameter, and processing shots to improve an Al model.
Owner:FIN BONE LLC

Information processing device and information processing program

To provide an information processing device and an information processing program which can display a video at a time point at which a user wants a video to be stopped from a time point at which a stopping instruction of the video is actually received as a static image.SOLUTION: An information processing device comprises a processor. The processor records a movement of a user to a time point at which a video stopping instruction for stopping a displayed video is performed, and displays a static image of the video at a second time point that is before the first time point at which the video stopping instruction is received, that is when acceleration following the movement of the user exceeds a predetermined reference for the first time and that is immediately after the first time point.SELECTED DRAWING: Figure 6
Owner:FUJIFILM BUSINESS INNOVATION CORP

On demand visual recall of objects / places

Aspects of the subject disclosure may include, for example, observing a plurality of objects viewed through a smart lens, wherein the plurality of objects are in a frame of an image viewed by the smart lens, determining an identification for an object of the plurality of objects, assigning tag information for the object based on the identification, storing the tag information for the object and the frame in which the object was observed, receiving a recall request for the object, retrieving the tag information for the object and the frame responsive to the receiving the recall request, and displaying the tag information and the frame. Other embodiments are disclosed.
Owner:AT&T INTELLECTUAL PROPERTY I L P

PROCEDURE

A procedure executed by an end device comprising a control unit, an imager, and an input interface involves the control unit performing operations, including capturing a customer's consent to record customer interaction audio via the input interface, starting audio recording after consent has been captured, and stopping audio recording in a case where a predetermined condition is met.
Owner:TOYOTA JIDOSHA KK

Machine-learned model for generating an output based on image frames adaptively extracted from a video

A computing device for generating content includes one or more memories to store instructions and one or more processors to execute the instructions to perform operations, the operations including: receiving a video; receiving an input prompt associated with the video; processing the video by adaptively extracting a plurality of image frames from the video at irregular intervals, based on content of the video; and implementing one or more machine-learned models to generate an output responsive to the input prompt, based on the input prompt and the plurality of image frames adaptively extracted from the video.
Owner:GOOGLE LLC

Video editing support device, video editing support method, and recording medium

A video editing support device includes: a captured video acquirer that acquires a captured video capturing a state of a sport being played, in which the captured video includes an object as a subject; a stock video storage unit that stores a stock video based on the captured video; a condition storage unit that stores multiple attention mark assignment conditions for the object; an attention scene detector that detects an attention scene that satisfies an attention mark assignment condition in the stock video; an attention mark assignment unit that relates, to a detected attention scene, an attention mark corresponding to the type of an attention mark assignment condition associated with the detection; and a stock video display controller that displays the stock video on a certain display device and that displays, on the stock video, an attention mark related to an attention scene in the stock video, together with playback time information.
Owner:ASICS CORP

Image processing method and apparatus, device, and medium

An image processing method is provided. In the method, a target video frame set is acquired from video data of a plurality of video frames. The target video frame set includes a subset of the video frames that is selected based on characteristics of the subset of the video frames. A global color feature of a reference video frame is acquired. An image semantic feature of the reference video frame is acquired. An enhancement parameter of the reference video frame is acquired for each of at least one image information dimension according to the global color feature and the image semantic feature. Image enhancement is separately performed on the video frames in the target video frame set according to each enhancement parameter of the reference video frame to obtain target image data for each of the video frames in the target video frame set.
Owner:TENCENT CLOUD COMPUTING (BEIJING) CO LTD