Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2252 results about "Media content" patented technology

Content (media) In publishing, art, and communication, content is the information and experiences that are directed toward an end-user or audience. Content is "something that is to be expressed through some medium, as speech, writing or any of various arts".

Audio and video player control method based on voice instruction

The invention relates to the technical field of audio and video control, and discloses an audio and video player control method based on a voice instruction. The method comprises the steps that an original voice instruction stream of a user is collected, the instruction stream comprises a time domain audio signal sequence, an environment noise spectrum and user pronunciation characteristic parameters, and voice information can be comprehensively captured; multi-modal instruction analysis processing is carried out on the original voice instruction stream, a structured control instruction set containing acoustic control intention identification, semantic operation object description and context correlation parameters is generated, and the analysis precision is improved; then executing player state adaptation based on the set, generating a dynamic control response sequence containing an equipment state adjustment command, a media content positioning parameter and an interface interaction logic identifier, driving a player to execute a multi-dimensional control operation and generating real-time play control effect feedback data; and finally, multi-modal analysis parameters are optimized according to feedback data, a self-adaptive instruction analysis strategy is generated, and the control experience of a user on the audio and video player is optimized.
Owner:ONWAY TECH LTD

Knowledge-intensive visual question and answer automatic data generation method and device

The invention relates to a knowledge-intensive visual question and answer automatic data generation method and device, and the method comprises the steps: constructing an original visual data set containing the professional knowledge of a target domain according to a static image, a video stream and multimedia content; extracting a representative frame sequence, converting the audio information into text information, and extracting character information in the static image to construct a structured visual instance database; according to the prompt text meeting the preset professional depth condition, establishing a three-level prompt system containing domain knowledge, an evaluation standard and a generation specification; generating a corresponding visual question and answer pair data set according to the dynamic cooperation of the main agent and the domain expert agent; generating a multi-agent quality evaluation system according to the quality evaluation result; and designing a difficulty grading mechanism according to the negative example sample. According to the method, the professionality, the accuracy and the diversity of the visual question and answer data are remarkably improved, and reliable data support is provided for training and evaluation of a multi-modal large model.
Owner:TSINGHUA UNIVERSITY

Segmentation of media content using vision language models

Disclosed are apparatuses, systems, and techniques for efficient instance segmentation with vision language models (VLMs). In an embodiment, the techniques include processing an input into the VLM to generate a segmentation map of a media item. The input includes the media item, which includes a plurality of media item units (e.g., pixels, groups of pixels), and further includes a prompt associated with the media item. The segmentation map includes identification of media item units associated with individual objects of one or more objects in the media item, and the VLM includes a dynamic portion having parameters that are determined in view of the media item.
Owner:NVIDIA CORP

Cross-platform information interaction method and device, equipment and medium

The invention relates to the technical field of data processing, can be applied to business scenes such as financial science and technology and medical health, and discloses a cross-platform information interaction method, device and equipment and a medium, and the method comprises the steps: obtaining original information sent by a source platform, and analyzing the original information into a standardized data structure according to a preset data format; extracting multimedia content elements from the structure and performing format conversion to generate standardized multimedia content; packaging the structure and the content into a to-be-transmitted data packet; acquiring network delay and bandwidth parameters of the target platform, and selecting a transmission path based on the parameters; compressing the to-be-transmitted data packet and sending the compressed to-be-transmitted data packet to the target platform through the selected path; and the target platform receives and analyzes the compressed data packet and presents the original information content. Information format compatibility is achieved through a standardized structure and content packaging, transmission efficiency is improved through network parameter perception and path selection, complete presentation of information is guaranteed in combination with compression processing and terminal adaptation, and stability and consistency of cross-platform interaction are enhanced.
Owner:PING AN TECH (SHENZHEN) CO LTD

Engagement-based collaboration recommendations

A recommendation system is described, which identifies engaged fans for an artist and requests input from the engaged fans regarding a collaboration by the artist with at least one different artist. In implementations, engaged fans are identified as having user profiles on a media content platform that satisfy at least one threshold engagement criteria based on consumption of at least one media content item associated with the artist. The recommendation system presents a user interface that includes at least one prompt for feedback that enables engaged fans to recommend how the artist collaborate with others. In some implementations, the user interface includes controls that are selectable to define artist characteristics to feature in a collaboration and the recommendation system is configured to generate a synthesized collaboration by automatically combining different artists' characteristics using a trained machine learning model. Recommendations based on engaged fan feedback are then provided to artists.
Owner:BLOCK INC

Artificial intelligence systems for automated social media content generation and trend integration

Certain aspects of the disclosure provide artificial intelligence (AI) methods and systems for generating personalized social media content with trend integration. A method generally includes retrieving data from data sources that includes customer interactions with a business, and inventory data of the business, determining trending-product pairs that increase engagement of the customers with products recorded in the inventory data of the business based on the retrieved data. A generative artificial intelligence (AI) model is used to generate one or more of a caption, a hashtag, and a promotional image that are personalized to each of the customers in response to receiving prompts that contain information about the customers, information about trending-product pairs, and social media platforms of the customers. The method sends one or more of the captions, the hashtags, and the promotional images that are personalized to the customers to social media platforms of the customers.
Owner:INTUIT INC

Resource Allocation Based on Media Content Engagement

A technique for resource allocation estimation for media content items is described. In accordance with the described techniques engagement by a set of user accounts with respective media content items of at least one media content service provider system is obtained. The media content service provider system and / or a payment service system generates historical streaming data for the respective media content items based on the engagement of the set of user accounts. An estimated streaming count of a media content item over a time period based on the historical streaming data for the respective media content items is determined. An estimated resource allocation for the artist is determined based on the estimated streaming count and an advance of funds is facilitated based on the estimated resource allocation to an account of the artist during the time period.
Owner:BLOCK INC

Private network content copyright monitoring and evidence obtaining system based on AI and large model

The invention discloses a private network content copyright monitoring and evidence obtaining system based on AI and a large model, which utilizes AI and large model technologies to carry out copyright monitoring and evidence obtaining on multimedia content in a private network environment, and carries out deep semantic understanding and cross-modal feature extraction through a large model processor to generate unified semantic representation. The method comprises the steps that a copyright content database is established and stored in a copyright content knowledge base, the copyright content knowledge base is used for storing metadata of original content protected by copyright, unified semantic representation and copyright declarations in a natural language form provided by a copyright party, and the copyright declarations are converted into query vectors through a semantic understanding technology; the method comprises the following steps: acquiring a semantic representation of a multimedia content, performing similarity calculation with the semantic representation of the multimedia content, identifying infringement content, when the infringement content is identified, recording an original source, publishing time and publisher information of the infringement content, performing differentiation analysis, generating an evidence chain, displaying a copyright monitoring result through a user interface, generating infringement alarm information, and presenting details of the evidence chain.
Owner:BEIJING LIUJINSUIYUE TECH CO LTD

Commodity multimedia recommendation method and system combining RPA and AI

The invention provides a commodity multimedia recommendation method and system combined with RPA and AI.The method comprises the steps that firstly, a current interaction behavior flow of a user and a commodity multimedia interface is recorded in real time through an RPA interaction capture module, the current interaction behavior flow comprises an operation triggering time sequence and an attention staying track, then an intention evolution track is extracted based on a preset historical interaction mode library, and the intent evolution track is extracted; the method comprises the following steps of: generating a dynamic matching rule set according to an intention evolution track, including an association constraint condition and a priority ranking logic, inputting the dynamic matching rule set into a pre-trained AI recommendation model, performing rule matching screening on a candidate commodity multimedia content set, generating a screening result, and outputting the screening result to a user. And finally, a recommendation sequence is rendered in real time according to a screening result through an RPA display arrangement module, and the display size and the text typesetting style are adjusted, so that the individuation degree of commodity multimedia recommendation and the user experience are improved.
Owner:QIANFENG HIGH ENERGY ARTIFICIAL INTELLIGENCE TECH (CHENGDU) CO LTD

Automated system and method for creating structured data objects for a media-based electronic document

A system including a media data optimization engine (MDOE) and a method for automatically creating structured data objects for media content rendered in one or more languages in an electronic document of a business entity are provided. The MDOE identifies non-textual objects including media content rendered in one or more languages in the electronic document and generates textual objects in the corresponding language(s) therefrom. The MDOE transforms the textual objects into structured data objects based on configurable criteria and generates a dynamic index-oriented object for the structured data objects specific to the business entity. The MDOE connects the structured data objects to the dynamic index-oriented object by creating linked data nodes therefrom with the dynamic index-oriented object as a core. The MDOE connects the dynamic index-oriented object with the linked data nodes to the electronic document, thereby facilitating dynamic changes to the electronic document and dynamically optimizing the electronic document.
Owner:MEHTA JATIN V +1

Artificial intelligence semantic processing system and method for digital media creation

The invention provides an artificial intelligence semantic processing system and method oriented to digital media creation, and relates to the technical field of artificial intelligence semantic process.The artificial intelligence semantic processing method comprises the steps that predicate argument relation pairs of language texts are extracted, object space relation pairs of sketch images are extracted at the same time, and a basic semantic unit set is constructed; the integrity and accuracy of cross-modal semantic understanding are ensured, further, semantic units are clustered by using a dynamic routing algorithm, a semantic concept cluster with a clear importance weight is generated, deep mining and structured representation of creation intentions are realized, and the creation intentions are quickly and accurately understood. An initial semantic relation graph is constructed, a graph attention network is used for dynamic reweighting, finally, an enhanced dynamic semantic graph is generated, complex association and a hierarchical structure between semantic concepts are effectively captured, finally, hierarchical analysis is carried out on the semantic graph, and a structured semantic blueprint is output, so that the dynamic semantic graph is obtained. And a reliable semantic processing technology is provided for creation of high-quality digital media contents.
Owner:HUNAN INST OF INFORMATION TECH

Media editing method and device, equipment and storage medium

The embodiment of the invention provides a media editing method and device, equipment and a storage medium. The method comprises the following steps: displaying a first sub-lens component in a sub-lens editing area of an editing interface; displaying a media generation area in the editing interface in response to a preset operation received in the first sub-lens assembly; on the basis of parameter information acquired in the media generation area, generating a first sub-lens section corresponding to the first sub-lens assembly; and displaying the first sub-lens segment in a content preview area of the editing interface. In this way, according to the embodiment of the invention, the corresponding segment of the media content can be efficiently edited through the sub-mirror component, so that the media editing efficiency is improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Auto trimming for augmented reality content in messaging systems

The subject technology receives frames of a source media content. The subject technology detects from the frames of the source media content, a first gesture indicating a cut point at a particular frame of the source media content, the cut point associated with a trimming operation to be performed on the source media content. The subject technology selects a starting frame and an ending frame from the frames based at least in part on the cut point at the particular frame. The subject technology performs the trimming operation based on the starting frame and the ending frame. The subject technology generates a second media content using the third set of frames. The subject technology provides for display at least a portion of the third set of frames of the second media content.
Owner:SNAP INC

Adaptive Streaming Content Selection for Playback Groups

A playback device is configured to (i) operate as part of a synchrony group including at least one other group member, (ii) obtain a respective indication of each group member's capability to play back media content, (iii) based on the respective indications, determine a group capability to play back media content, (iv) transmit, to a cloud-based computing system, a request for a media item, (v) receive, from the cloud-based computing system, a list of different renditions of the requested media item, the list including a respective media item identifier usable to obtain each different rendition, (vi) select a rendition of the requested media item that corresponds to the determined group capability, (vii) use a media item identifier corresponding to the selected rendition to retrieve the selected rendition of the requested media item, and (viii) play back the selected rendition in synchrony with the at least one other group member.
Owner:SONOS INC

Passive and continuous multi-speaker voice biometrics

Embodiments described herein provide for a voice biometrics system execute machine-learning architectures capable of passive, active, continuous, or static operations, or a combination thereof. Systems passively and / or continuously, in some cases in addition to actively and / or statically, enrolling speakers as the speakers speak into or around an edge device (e.g., car, television, radio, phone). The system identifies users on the fly without requiring a new speaker to mirror prompted utterances for reconfiguring operations. The system manages speaker profiles as speakers provide utterances to the system. Machine-learning architectures implement a passive and continuous voice biometrics system, possibly without knowledge of speaker identities. The system creates identities in an unsupervised manner, sometimes passively enrolling and recognizing known or unknown speakers. The system offers personalization and security across a wide range of applications, including media content for over-the-top services and IoT devices (e.g., personal assistants, vehicles), and call centers.
Owner:PINDROP SECURITY INC

Image processing method, electronic device and readable storage medium

The present disclosure provides an image processing method, an electronic device and a readable storage medium. The method includes: obtaining identification information by identifying an identification pattern in a media image; obtaining virtual information corresponding to media content displayed in the media image according to the identification information; obtaining an image acquired in real time; and obtaining a three-dimensional image based on the virtual information and the image acquired in real time.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Multimedia content preloading method based on vehicle cloud cooperation

The invention discloses a multimedia content preloading method based on vehicle cloud collaboration, and particularly relates to the technical field of multimedia content preloading, and the method comprises the steps: carrying out the modeling through a multi-dimensional space-time driving path, and combining with historical road communication coverage features; a network availability prediction result and a user historical consumption mode are deeply mapped, and a weight fusion mechanism of content timeliness, capacity characteristics and scene correlation is introduced, so that a generated multi-level priority sequence better meets the instant requirements of vehicles at different time and different positions; through dynamic comparison of an initial plan, a real-time path, a network and user interaction data, a superposition out-of-control situation of prediction deviation can be rapidly identified, a scheduling priority offset mode under multiple tasks is identified in combination with vehicle calculation, caching and bandwidth occupation states, and then a pre-loading dislocation cooperative imbalance fault network is constructed through bidirectional correlation analysis. And visual diagnosis and quantitative evaluation of the imbalance state of the prediction layer and the execution layer are realized.
Owner:深圳市鼎微科技有限公司

Multimedia object tracking and merging

In multimedia object tracking and merging of tracked objects, an object is tracked through frames of multimedia content until a frame appears in which the tracked object is not detected. A first track is designated as one or more consecutive frames in which the tracked object is detected, the first track ending at the first frame. Tracking continues to try to detect the tracked object in a second frame subsequent to the first frame. If the tracked object is not again detected, information about the first track is output. If the tracked object is detected subsequently, a second track of consecutive tracked object detection is designated. The tracked objects in the two tracks are then compared with the aid of trained data models, and a matching score is determined to reflect the degree of match. If the matching score meets or exceeds a first threshold, the compared tracks are merged using the same identifier assigned to both tracks. If the matching score does not exceed a second threshold that is less than the first threshold, the tracks may be discarded as showing no match. If the matching score falls between the first and second thresholds, an indication is output for further analysis of the compared tracked objects.
Owner:GETAC TECH CORP +1

Systems and methods for simulating future asset performance based on consumable media content

PendingUS20250363560A1FinanceUser deviceData stream
Systems, apparatuses, methods, and computer program products are disclosed for simulating future asset performance based on consumable media content. An example method includes monitoring a user device for receipt of a data stream comprising media content. The example method further includes receiving a simulation request requesting a prediction model for an asset of a user portfolio based on the media content. The example method further includes generating a prediction model output indicating future performance of the asset of the user portfolio based on the media content and historical data. The example method further includes generating a natural language report representative of the future performance of the asset of the user portfolio. The example method may further include transmitting the natural language report to the user device.
Owner:WELLS FARGO BANK NA

Multi-selection shutter camera app that selectively sends images to different artificial intelligence and innovative platforms that allow for fast sharing and informational purposes

A multi-selection shutter camera application and method selectively sends images to different artificial intelligence and innovative platforms for fast sharing and informational purposes. An electronic device with a touch sensitive display and a processor is utilized. A capture screen on the display includes a multi-shutter view (a live view and at least two selective capture buttons (e.g., shutters)) for capturing media content. The method may include receiving input from the buttons to capture the content and direct it to an artificial intelligence platform; processing the content information based on the selected button; analyzing the content through parameters and show options to the user on the same screen that displays the content; and presenting a plurality of selectable options related to the user's selected shutter and intent of capturing the content, including, but not limited to uses related to at least one of the following: discovery, shopping and sharing functionality alternatives.
Owner:YAE LLC

Audio-lip movement correlation measurement for dubbed content

Methods and apparatus are described for evaluating dubbing of media content. Phonemes in dubbed audio are extracted and mapped to visemes. Lip poses in video frames of the media content corresponding to the phonemes of the dubbed audio are compared to the visemes determined from the dubbed audio. A notification may be generated based on the comparison that indicates synchronization of the dubbed audio to lip poses of the video.
Owner:AMAZON TECH INC

Digital media content element accurate screening method based on artificial intelligence image recognition

The invention discloses a digital media content element accurate screening method based on artificial intelligence image recognition, and relates to the technical field of digital media, and the method comprises the steps: building a distributed capture network to form a dynamic content pool, and building a metadata index database; calling a multi-modal image perception engine to generate a double-layer characteristic spectrum containing dominant and recessive elements; constructing a distributed recognition model cluster based on federated learning; converting the user demand into a screening parameter set and generating a decision tree; screening and secondarily verifying an output result through a double-path matching mechanism; and constructing a reinforcement learning reward function based on user behaviors, and driving the model and the decision tree to co-evolve. Multi-source heterogeneous content full-dimension analysis is achieved, the recognition comprehensiveness and depth are improved, knowledge barriers and privacy risks are solved, the screening accuracy and flexibility are improved, the system is endowed with the continuous optimization capacity, and the method is suitable for efficient and accurate digital media content screening scenes.
Owner:XIAMEN HUAXIA UNIV

Adaptive video recap of media content episodes in an electronic device

A computing system, a method and a computer program product for presenting a determined optimal duration of video recap of media content. The method includes detecting, via a processor of a computing system, selection of a current episode of media content for playback. In response to detecting selection of the current episode of the media content for playback, the method includes determining a first time difference between a current time and a previous viewing time of prior episodes. The method includes determining, based on the first time difference, a first time duration for a first video recap of the prior episodes of the media content. The method includes streaming the first time duration of the first video recap of the media content for presentation on an electronic device as a preview presented prior to streaming the current episode of the media content for presentation on the electronic device.
Owner:MOTOROLA MOBILITY LLC

Intelligent media content association propagation and influence analysis method based on knowledge graph

The invention discloses an intelligent media content association propagation and influence analysis method based on a knowledge graph, and the method comprises the following steps: collecting multi-source media data from social media, a news website, a video platform and a forum, and constructing and forming a heterogeneous knowledge graph; performing feature initialization and embedding on nodes in the heterogeneous knowledge graph to generate initial node embedding representation; constructing a dynamic heterogeneous graph attention network based on the node initial embedding representation to obtain a node dynamic propagation state representation; calculating a propagation influence score based on the node dynamic propagation state representation to obtain a node influence sorting result and a core propagation path; and introducing a causal consistency training mechanism based on the core propagation path, optimizing parameters of the dynamic heterogeneous graph attention network, and outputting a propagation influence evaluation result. According to the method, the dynamic heterogeneous graph attention network is adopted, and intelligent analysis of the media content propagation influence is realized.
Owner:ZHUHAI COLLEGE OF JILIN UNIV

Customizable latency for automatic speech recognition

Techniques for customizable latency, from the customer's side, for automatic speech recognition (ASR) are described. In particular, the customer may specify a parameter that controls how fast or how slow the customer's media content will be streamed or processed. Slower processing means higher accuracy, with near real-time latency, while faster processing means lower accuracy, but offers much lower latency (e.g., less than 600 ms). Enabling tuning of the latency-versus-accuracy tradeoff of the ASR system offers customers the flexibility to meet varying needs for different ASR applications.
Owner:AMAZON TECH INC

Advertisement creativity matching method based on multi-modal content generation

The invention discloses an advertisement creativity matching method based on multi-modal content generation, and relates to the technical field of digital media content generation, and the method comprises the following steps: building a cross-modal time anchoring belt facing advertisement creativity matching, carrying out metaphor level decomposition on input text information, marking a symbol axis for image information, and carrying out data processing on the image information; obtaining an initial semantic boundary list; and constructing a culture fingerprint database according to the initial semantic boundary list, and mapping the territory taboo information and the brand symbol information into constraint tags to obtain a semantic guardrail set. According to the method, through cross-modal time anchoring and semantic boundary control, accurate correspondence of the text and the image in time and semantic levels is achieved, and it is ensured that generated content is clear in semantic meaning and adaptive in culture. In combination with breathing type phase traction and cultural fingerprint dynamic adjustment, multi-modal content rhythm and emotion are coordinated and unified, brand expression is kept stable, and the overall consistency and propagation effect of advertisement creativity are improved.
Owner:大根控股股份有限公司

Method and system for volume control

A method performed by a first electronic device, the method includes, while engaged in a call with a second electronic device, initiating a joint media playback session in which the first and second electronic devices independently stream media content for synchronous playback; driving a speaker with a mix of a downlink signal of the call and an audio signal of the media content at an overall volume level; receiving a user-adjustment at a single volume control for the first electronic device to reduce the overall volume level; in response to the user adjustment, applying a first gain adjustment to the downlink signal and a second gain adjustment to the audio signal; and driving the speaker with a mix of the downlink signal and the audio signal at the reduced volume level.
Owner:APPLE INC

Media content for any request using LLMs

Described are systems and processes for identifying relevant media content for users in response to natural language requests. The media content may be formed as a playlist of music. Requests may be sent to various models to obtain results, such as a large language model (LLM) and a local search model. Results may be selected from one or both models. The results may be used to fetch media content and provide user interfaces with a playlist with the media content to a user that submitted the request. User interfaces may include selectable search terms suggested for the user and animations during processing of the request.
Owner:AMAZON TECH INC

Systems and methods for automatic media generation for game sessions

In some aspects, the disclosure is directed to systems and methods for automatic text generation for game sessions. A system may include one or more processors that are configured to store an identification of a game session and data corresponding to one or more events of the game session, the game session corresponding to one or more characteristics of individuals participating in the game session; determine, based on the one or more characteristics of the individuals participating in the game session, a set of criteria for selecting a set of media segments, wherein different criteria of the set of criteria correspond to different media segments; determine which of the set of criteria is satisfied; select one or more media segments based on the satisfied criteria; and generate a media content item for at least a portion of the game session based on the selected one or more media segments.
Owner:GAMECHANGER MEDIA