Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

106results about "Video data querying" patented technology

Shopping interface and method

A media sharing and communication system, including a recording mechanism that records a desired portion of media upon activation by a first individual user, a first user transmitter / receiver that transmits the portion of media and a message generated by the first individual user regarding the portion of media to a second individual user and is capable of transmitting a message to a second individual user, a confirmation mechanism that confirms that the second individual user is authorized to view the portion of media, a notification mechanism that notifies the first individual user if the second individual user is not authorized to receive the portion of media, a second user transmitter / receiver that receives the portion of media and voice message upon authorization of the second individual user, a search mechanism, and a video recording mechanism, an online betting module, and an online food ordering module.
Owner:TAYLOR DAVID A

Internet of Things video monitoring big data privacy protection and efficient retrieval method based on artificial intelligence

The invention discloses an Internet of Things video monitoring big data privacy protection and efficient retrieval method based on artificial intelligence, belongs to the crossing field of artificial intelligence, Internet of Things and information security, and is suitable for vehicle-mounted, park, battery swap stations and other scenes. The method comprises the following steps: firstly, establishing a self-adaptive acquisition framework, and realizing multi-protocol switching and video preprocessing; secondly, extracting privacy information through an improved YOLO algorithm and a Graph-Cut technology, and combining reversible watermark embedding; constructing a multi-dimensional privacy level model for hierarchical encryption, and matching two-level storage and three-level index; then, two-factor authentication authorization is performed to extract privacy and optimize retrieval; and finally, the system state is monitored in real time and self-adaptive optimization is performed. According to the method, the problems of protocol heterogeneity, insufficient privacy protection, low retrieval efficiency and the like can be solved, and privacy security, storage overhead and retrieval efficiency are balanced.
Owner:ZHEJIANG HAISHI HUAYUE DIGITAL TECHNOLOGY CO LTD

Editing strategy scheduling method and device, electronic device and storage medium

The invention relates to an editing strategy scheduling method and device, an electronic device and a storage medium, and the method comprises the steps: receiving an original material, a copywriting and an editing style instruction inputted by a user, and generating a corresponding dubbing audio according to the copywriting; performing multi-dimensional analysis on the original material to generate a lens-level structured index; performing semantic analysis on the copywriting to obtain copywriting semantic features; performing similarity retrieval based on the copywriting semantic features and the multi-modal semantic features of the materials to obtain a candidate shot set matched with the copywriting semantic features; generating a global style vector and a target rhythm curve through the first agent; and taking the global style vector and the target rhythm curve as control signals, driving a second agent to select and trim shots from the candidate shot set, recombining the editing sequence, and outputting a final editing sequence and an editing jump point structure. According to the style vector, the rhythm target curve and reinforcement learning, the editing style is met, and the intelligent agent editing stylized presentation is achieved.
Owner:ZHEJIANG HUAZHI WANXIANG TECHNOLOGY CO LTD

Media content memory retrieval

Aspects of the subject disclosure may include, for example, a media consumption database that stores data elements describing conditions under which electronic media content is consumed by a user on an electronic device. A search of the media consumption database based on at least a portion of the conditions may result in at least a portion of the electronic media content to be re-presented to an electronic device of the user Other embodiments are disclosed.
Owner:AT&T INTELLECTUAL PROPERTY I L P

End-to-end multi-task video retrieval with cross attention

A method includes obtaining a video and a relational spatio-temporal query, and identifying at least one type of the relational spatio-temporal query. The at least one type of identification of the relational spatio-temporal query represents at least one of: an activity type, an object type, or a temporal type. The method further includes learning correlations between activities, objects, and time in the video using one or more cross attention models. The method further includes obtaining one or more predictions generated using one or more outputs of the one or more cross attention models based on the identified at least one type of relational spatiotemporal query. Further, the method includes generating a response to the relational spatiotemporal query based on the one or more predictions.
Owner:SAMSUNG ELECTRONICS CO LTD

system

We provide the system. [Solution] A means for automatically identifying the voice of a specific person, adding a timestamp, and recording it as audio data, A means of recognizing an object held by a specific person as image data and recording it with a timestamp, A means for editing and generating a natural language record based on the above audio data and image data, A means for automatically transmitting this record to an external communication application via a communication network, A means for performing behavior recognition within a specific environment and generating messages corresponding to specific behaviors, A means of transmitting the generated message to an information terminal, A system that includes this.
Owner:SOFTBANK GROUP CORP

Camera gun video resource positioning method and related equipment

The invention discloses a camera gun video resource positioning method and related equipment. The method comprises the following steps: acquiring camera gun information of each financial network; obtaining structured risk model data for the daily recovery high-risk model troubleshooting index; determining a target financial network point corresponding to each piece of risk model data in the structured risk model data; according to the structured risk model data, the target financial network point corresponding to each piece of risk model data, the camera gun information of each financial network point and the current popularity value of each camera gun, determining a plurality of current retrieval camera guns; according to the current user operation log, updating the current popularity value corresponding to each camera gun; and according to the updated current popularity value corresponding to each camera gun, determining a plurality of updated current retrieval camera guns. The method can achieve the precise pushing of the video resources needed by the daily redisk of the website, improves the video verification efficiency, and can be widely applied to the technical field of artificial intelligence.
Owner:GUANGDONG BRANCH OF CHINA POST GRP CO LTD

Video-based omnibearing remote rehabilitation training system and method, and medium

The invention discloses a video-based omni-directional remote rehabilitation training system, a video-based omni-directional remote rehabilitation training method and a medium, which are characterized in that a digital standardized training video library classified according to four stages of PT physical therapy is established, and more than 700 digital standardized training videos such as joint activity, balance training, gait correction and the like are covered; it is ensured that the training content is scientific and covers the whole period and all directions of rehabilitation training; rehabilitators generate personalized training schemes in combination with rehabilitation scene types (hospitalization / home / community / old-age care institutions) and clinical data and in combination with remote video evaluation of the rehabilitators, and training requirements in different environments are met. By capturing actions of limb joint angles, joint point motion trails and muscle force changes, a training action deviation rate is calculated and is fed back and output to a rehabilitation teacher in a grading manner, so that the training standardability is remarkably improved; rehabilitators can manage multiple patients online at the same time, remotely check training percentage data, adjust schemes and generate rehabilitation notes, and efficient, accurate and personalized rehabilitation guidance services are provided for the patients.
Owner:FALCON HEALTH TECHNOLOGY (SHANGHAI) CO LTD

Multimodal data processing for content retrieval systems and applications

In various examples, multimodal data processing for content retrieval systems and applications is described herein. Systems and methods described herein may convert different modalities of data into a common type of modality. For instance, content data representing a video may be separated into audio data representing sound corresponding to the video—such as speech—along with video data representing frames of the video. The audio data may then be processed using one or more models to generate first text corresponding to a transcript of the speech. Additionally, the video data may be processed to identify specific keyframes that provide important information associated with the video. The keyframes may then be processed using one or more models to generate second text describing the keyframes. The systems and methods may then combine the text from the different modalities and generate data for storage in one or more databases.
Owner:NVIDIA CORP

A method and system for weakly supervised location of video clips based on a large-scale video corpus

The present invention relates to the technical field of video data recognition. For the obtained training dataset, self-supervised learning is used to extract common semantic information between text and video, and based on the semantic information, steps are taken to obtain fused semantic video features; for the fused semantic video features and corresponding text features, multi-scale contrast learning is performed using a weak-supervision method to determine the spatial mapping relationship between video features and text features, map them to a metric space, and obtain a trained metric space; steps are taken to obtain a search query, search for text features similar to the search query in the trained metric space, and use the video clip corresponding to the text feature with the highest similarity as the video positioning result. The present invention provides a weak-supervision positioning method and system for video clips based on a large-scale video corpus. The positioning method of the present invention can realize directly determining the position of a video clip accurately and quickly from a large-scale video database.
Owner:SHANDONG JIANZHU UNIV

Neural network retraining based on image classification feedback

Described herein are systems and methods that highlight target objects in media content. The detection system trains a neural network to identify objects within an image. The detection system receives a text input requesting a search within the image and applies a search language model to the text input, which identifies a target object associated with the requested search. The detection system applies the neural network to the image to identify instances of the target object. The detection system modifies a user interface to include the image and modifies the image to highlight the identified instances of the target object. The detection system receives feedback that modifies the highlighted instances of the target object within the user interface and retrains the neural network based on the modified highlighted instances.
Owner:MATROID INC

Autonomous activity monitoring system and method

A system for automatically monitoring activity on an athletic activity area is provided. The network device includes artificial intelligence configured to automatically identify objects and gestures from video received from cameras disposed at the activity area. The artificial intelligence automatically edits the video based on objects and gestures identified from the video and generates a video file including predetermined objects and gestures. The artificial intelligence may generate a video including each trial completed by a particular athlete and may provide the video automatically to the player or a third party.
Owner:HOLE IN ONE MEDIA INC

Display device and operating method thereof

A display device according to an embodiment of the present disclosure comprises: a controller for receiving a speech command and obtaining an official content name mapped to an uttered content name included in the speech command; and a display for outputting an operation result corresponding to a speech command obtained by changing the uttered content name to the official content name. Thus, a display device in which a result intended by a user is provided even when a content name in a speech command is not accurate, and an operating method thereof, may be provided.
Owner:LG ELECTRONICS INC

Distributed video storage and search with edge computing

Systems and methods are provided for distributed video storage and search with edge computing. The method may comprise caching a first portion of data on a first device. The method may further comprise determining, at a second device, whether the first device has the first portion of data. The determining may be based on whether the first piece of data satisfies a specified criterion. The method may further comprise sending the data, or a portion of the data, and / or a representation of the data from the first device to a third device.
Owner:NETRADYNE INC

Electronic device and operation method thereof

The present disclosure relates to an artificial intelligence (AI) system and application thereof, which use a machine learning algorithm. An electronic device according to the present disclosure may include memory storing one or more instructions, and one or more processors configured to execute the one or more instructions stored in the memory, wherein the one or more processors are configured to transmit, to a server, request information that is obtained from at least one of situation information and metadata corresponding to content, the request information including input conversational text information, and receive, from the server, a recommendation result based on the request information.
Owner:SAMSUNG ELECTRONICS CO LTD

Method and apparatus for providing similar content in a content streaming system

This disclosure relates to a method and apparatus for providing similarity content in a content streaming system, the operation method of the server in the content streaming system may include the steps of: acquiring first sequence-type text data including information contained in first metadata of a first content item; acquiring second sequence-type text data including information contained in second metadata of a second content item; determining a first vector corresponding to the first sequence-type text data and a second vector corresponding to the second sequence-type text data using a language model learned based on synopsis information contained in the metadata of the content item; determining the similarity between the first content item and the second content item using the first vector and the second vector; and providing a content list including at least one content item, including the second content item selected based on the similarity.
Owner:TVING CO LTD

Method of processing video, method of quering video, and method of training model

The present application provides a method of processing a video, a method of querying a video, and a method of training a video processing model. A specific implementation solution of the method of processing the video includes: extracting, for a video to be processed, a plurality of video features under a plurality of receptive fields; extracting a local feature of the video to be processed according to a video feature under a target receptive field in the plurality of receptive fields; obtaining a global feature of the video to be processed according to a video feature under a largest receptive field in the plurality of receptive fields; and merging the local feature and the global feature to obtain a target feature of the video to be processed.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

A video big data retrieval method and system based on semantic tags

ActiveCN121681871BSolve the problem of occupying storage resourcesImplement semantic-level deduplicationVideo data indexingVideo data queryingSignal-to-noise ratioMultimedia information retrieval
This invention relates to the field of video data processing and multimedia information retrieval technology, specifically a video big data retrieval method and system based on semantic tags. The method includes: a data mapping step: mapping a video stream into a sequence of feature vectors using a multimodal encoder; a difference calculation step: calculating the rate of change of differences between adjacent feature vectors to generate a semantic velocity vector; an anchor point extraction step: extracting semantic abrupt change event anchor points based on a comparison of the velocity vector magnitude with a threshold; an index construction step: constructing a differential manifold index with semantic transitions as nodes and time spans as edges; a dynamic dimensionality reduction step: constructing a dynamic mask matrix based on the retrieval request, projecting nodes to a low-dimensional subspace; and a matching output step: calculating the matching degree in the subspace and outputting the result. This invention significantly improves retrieval speed and signal-to-noise ratio, and substantially reduces storage volume through index event anchor points and dynamic subspace folding.
Owner:XIAMEN HUAMEI YUNHAI TECH CO LTD +1

Video search device, video search method, and program

In order to solve the problem of improving video search accuracy even in a case where the accuracy or amount of information concerning video is not sufficient, a video search device (1) comprises: a generation unit (11) that generates explanation information for each video image stored in a video storage device; an acquisition unit (12) that acquires a search query; a search unit (13) that searches the video storage device for a video image by using the search query and explanation information; an output unit (14) that outputs a search result by the search unit (13); an input unit (15) that receives input of a determination result of a user with respect to the search result; and an update unit (16) that updates the explanation information on the basis of the determination result and the search query.
Owner:NEC CORP

Indexing fingerprints

Example methods and systems for indexing fingerprints are described. Fingerprints may be made up of sub-fingerprints, each of which corresponds to a frame of the media, which is a smaller unit of time than the fingerprint. In some example embodiments, multiple passes are performed. For example, a first pass may be performed that compares the sub-fingerprints of the query fingerprint with every thirty-second sub-fingerprint of the reference material to identify likely matches. In this example, a second pass is performed that compares the sub-fingerprints of the query fingerprint with every fourth sub-fingerprint of the likely matches to provide a greater degree of confidence. A third pass may be performed that uses every sub-fingerprint of the most likely matches, to help distinguish between similar references or to identify with greater precision the timing of the match. Each of these passes is amenable to parallelization.
Owner:GRACENOTE INC

Information processing system, information processing method, and program

To appropriately detect start of work SOLUTION: An information processing apparatus comprising: an acquisition unit configured to acquire an image captured by an image capturing apparatus; a detection unit configured to detect start work by a worker by analyzing the image acquired by the acquisition unit; and a recording unit configured to record that predetermined work is started when it is detected that the start work continues for a predetermined time.SELECTED DRAWING: Figure 1
Owner:CANON MARKETING JAPAN INC +1

Matching video content with podcast episodes

A system and method are provided for matching videos and podcast episodes. A data store comprising podcast episode identifiers is accessed. The podcast episode identifiers are associated with one or more podcast episode attributes. A video content item is identified. The video content item includes one or more video content item attributes. A matching podcast episode identifier that matches the video content item is determined based on the one or more podcast episode attributes and the one or more video content item attributes. A ranking of one of the video content items or the matching podcast episode identifiers is adjusted to reflect a correspondence between the video content item and the matching podcast episode identifier. Information associated with the matching podcast episode identifier is provided to a first user device.
Owner:GOOGLE LLC

Smart automated assistant for television user interaction

A system and method for controlling television user interactions using a virtual assistant is disclosed. The virtual assistant can interact with a television set-top box to control content displayed on a television. Voice input for the virtual assistant can be received from a device having a microphone. A user intent can be determined from the voice input, and the virtual assistant can perform a task in accordance with the user intent, including causing media to be played back on the television. Virtual assistant interactions can be displayed on the television in an interface that expands or contracts to occupy a minimal amount of space while conveying desired information. A user intent can be determined from the voice input, and information conveyed to the user using multiple devices associated with multiple displays. In some examples, virtual assistant query suggestions can be provided to the user based on media content displayed on the displays.
Owner:APPLE INC

Automatically generating descriptions of augmented reality effects

The computer system accesses a first image and a second image. The second image is generated by applying augmented reality (AR) effects to the first image. The computer system provides the first image, the second image, and a cue to a visual semantic machine learning model to obtain an output describing at least one feature of the AR effect. The computer system generates a description of the AR effect based on the output of the visual semantic machine learning model. The computer system stores the description of the AR effect in association with an identifier for the AR effect.
Owner:SNAP INC

Editing strategy scheduling methods, devices, electronic devices, and storage media

ActiveCN121619466BAchieve stylized presentation of clipsVideo data indexingVideo data queryingComputer graphics (images)Control signal
This application relates to an editing strategy scheduling method, apparatus, electronic device, and storage medium. The editing strategy scheduling method includes: receiving user-input raw materials, script, and editing style instructions, and generating corresponding dubbing audio based on the script; performing multi-dimensional analysis on the raw materials to generate a shot-level structured index; performing semantic analysis on the script to obtain script semantic features; performing similarity retrieval based on the script semantic features and the multimodal semantic features of the materials to obtain a set of candidate shots matching the script semantic features; generating a global style vector and a target rhythm curve through a first intelligent agent; using the global style vector and target rhythm curve as control signals to drive a second intelligent agent to select and trim shots from the candidate shot set, and reorganizing the editing sequence to output the final editing sequence and editing jump point structure. Through style vectors, target rhythm curves, and reinforcement learning to conform to the editing style, the intelligent agent achieves stylized editing presentation.
Owner:ZHEJIANG HUAZHI WANXIANG TECHNOLOGY CO LTD

Entropy-aware robot control method and system based on expert demonstration

The application provides an entropy-aware robot control method and system based on expert demonstration, and relates to the field of embodied intelligence technology, which comprises obtaining a task instruction input by a user, observation data and ontology perception data of a robot at a current time; cleaning and feature extracting the data to obtain multiple features, and retrieving a reference video in an expert reference video database; encoding and splicing the multiple features to generate a multi-modal feature representation; feature extracting the reference video to obtain reference video features; processing the multi-modal feature representation and the reference video features through a skill generation model to obtain a discrete skill codebook index probability; calculating an information entropy of the probability, and determining a sampling candidate number of a current candidate skill based on the information entropy; generating a skill token sequence based on the discrete skill codebook index probability and the sampling candidate number; and finally decoding the skill token sequence to obtain a robot action sequence. The control method provided by the application can guarantee accurate control of a robot in a complex environment.
Owner:HEFEI UNIV OF TECH