Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

181results about "Video data querying" patented technology

Character recognition model training method and apparatus, character recognition method and apparatus, device and storage medium

The present disclosure provides a character recognition model training method and apparatus, a character recognition method and apparatus, a device and a medium, relating to the technical field of artificial intelligence, and specifically to the technical fields of deep learning, image processing and computer vision, which can be applied to scenarios such as character detection and recognition technology. The specific implementing solution is: partitioning an untagged training sample into at least two sub-sample images; dividing the at least two sub-sample images into a first training set and a second training set; where the first training set includes a first sub-sample image with a visible attribute, and the second training set includes a second sub-sample image with an invisible attribute; performing self-supervised training on a to-be-trained encoder by taking the second training set as a tag of the first training set, to obtain a target encoder.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Long video content information acquisition method based on OCR (Optical Character Recognition) and voice recognition technology

The invention discloses a long video content information acquisition method based on an OCR and voice recognition technology. The method comprises the following steps: S1, carrying out preprocessing n on input long video data to extract an image frame sequence and an audio stream; s2, inputting the image frame sequence into an OCR recognition module, inputting the audio stream into an ASR recognition module, and obtaining a preliminary recognition result; s3, constructing a multi-target fitness function, and optimizing an OCR and ASR parameter combination by using a Kanglizard optimization algorithm; s4, respectively applying the optimal parameter group to an OCR identification module and an ASR identification module to obtain an optimized identification result; s5, constructing a fusion factor graph, executing edge message passing by adopting a belief propagation algorithm, and generating a multi-modal semantic block set; and S6, processing the multi-modal semantic block set to generate a unified multi-modal content information set. According to the invention, through fusion of the horny lizard optimization algorithm and the belief propagation mechanism, high-precision recognition and multi-modal semantic consistency extraction of the image text and the voice information in the long video are realized.
Owner:华电(海西)新能源有限公司

Video space-time retrieval method and device based on grid coding

The invention discloses a grid-coding-based video space-time retrieval method and device, the grid-coding-based video space-time retrieval method is executed in computing equipment, and the method comprises the following steps: mapping a coordinate range of a target area into a corresponding grid code set according to a Beidou grid standard on the basis of the coordinate range of the target area; acquiring video data with an intersection or inclusion relationship between the spatial range of the shot object and the target area; combining the shooting time of any video data with each grid code to generate a space-time grid code set; constructing an index structure corresponding to the video data by taking each time-space grid code in the time-space grid code set as a main key; and in response to query information of the user, analyzing a space-time range in the query of the user, mapping the space-time range into a corresponding space-time grid code range, executing retrieval based on the range, and returning hit video data. According to the method, rapid positioning and accurate retrieval of large-scale aerial photography video data are realized.
Owner:BEIJING ZHIWANG YILIAN TECH CO LTD

Method and system for text search capability of live or recorded video content streamed over a distributed communication network

A server receives and rebroadcasts live streaming video content from a video capture device, such as a mobile phone or unmanned surveillance vehicle. The server includes a media server configured to stream selected video content to a client device, a video analysis system configured to analyze the live video content and generate object detection data, a storage system configured to store the generated object detection data and an identifier of the associated live video content, and a search engine configured to receive a text-based search request, search the object detection data stored in the storage system for relevant search results, and generate a list of live and stored video content associated with the relevant search results.
Owner:AERYON LABS

Shopping interface and method

A media sharing and communication system, including a recording mechanism that records a desired portion of media upon activation by a first individual user, a first user transmitter / receiver that transmits the portion of media and a message generated by the first individual user regarding the portion of media to a second individual user and is capable of transmitting a message to a second individual user, a confirmation mechanism that confirms that the second individual user is authorized to view the portion of media, a notification mechanism that notifies the first individual user if the second individual user is not authorized to receive the portion of media, a second user transmitter / receiver that receives the portion of media and voice message upon authorization of the second individual user, a search mechanism, and a video recording mechanism, an online betting module, and an online food ordering module.
Owner:TAYLOR DAVID A

Privacy Controls for Sharing Embeddings for Searching and Indexing Media Content

This document describes techniques and systems that enable privacy controls for sharing embeddings for searching and indexing media content. A set of images of a user's face are obtained and a machine-learned model is applied to the set of images to generate a user-specific dataset of face embeddings for the user. Media content stored in a media storage is indexed by applying the machine-learned model to the media content to provide indexed media information identifying one or more faces shown in the media content. Access to the indexed media information by another user querying the media content for images or videos depicting the user is controlled based on a digital key shared by the user with the other user, where the digital key is associated with the user-specific dataset and the user-specific dataset is usable to identify the images or videos depicting the user.
Owner:GOOGLE LLC

Internet of Things video monitoring big data privacy protection and efficient retrieval method based on artificial intelligence

The invention discloses an Internet of Things video monitoring big data privacy protection and efficient retrieval method based on artificial intelligence, belongs to the crossing field of artificial intelligence, Internet of Things and information security, and is suitable for vehicle-mounted, park, battery swap stations and other scenes. The method comprises the following steps: firstly, establishing a self-adaptive acquisition framework, and realizing multi-protocol switching and video preprocessing; secondly, extracting privacy information through an improved YOLO algorithm and a Graph-Cut technology, and combining reversible watermark embedding; constructing a multi-dimensional privacy level model for hierarchical encryption, and matching two-level storage and three-level index; then, two-factor authentication authorization is performed to extract privacy and optimize retrieval; and finally, the system state is monitored in real time and self-adaptive optimization is performed. According to the method, the problems of protocol heterogeneity, insufficient privacy protection, low retrieval efficiency and the like can be solved, and privacy security, storage overhead and retrieval efficiency are balanced.
Owner:ZHEJIANG HAISHI HUAYUE DIGITAL TECHNOLOGY CO LTD

Method and system for three-dimensional modeling of chemical industry park

The invention discloses a three-dimensional modeling method and system for a chemical industrial park, and the method comprises the steps: carrying a multi-lens inclined camera and a monitoring sensor through employing an unmanned plane, and obtaining the multi-view image data, positioning and attitude determination data of the chemical industrial park, and the real-time operation data of chemical equipment; performing data processing on the multi-view image data to generate sparse point clouds, and generating three-dimensional point clouds through dense matching and depth estimation; performing block processing on the three-dimensional point cloud, and performing splicing by adopting a point cloud registration algorithm to obtain global point cloud data; performing surface reconstruction on the global point cloud data to generate a three-dimensional grid model, and completing model texture mapping through a texture mapping algorithm; a three-dimensional visual management platform is constructed, a three-dimensional grid model and real-time monitoring data are integrated, and park global visualization, risk visualization and intelligent management of chemical engineering devices, equipment and facilities are achieved.
Owner:BEIJING UNIV OF CHEM TECH

Systems and methods for few-shot new action recognition

A method includes: (i) receiving a query video including performance of an action; (ii) receiving a predetermined number of support videos including performance of actions, respectively, the predetermined number of support videos being less than 100 support videos; (iii) determining a similarity matrix based on a comparison of temporally ordered images of the query video with temporally ordered images of one of the support videos, respectively; (iv) determining a similarity value for the one of the support videos based on the similarity matrix; (v) repeating (iii) and (iv) for each of the support videos; (vi) identifying the highest one of the similarity values and the one of the support videos associated with the highest one of the similarity values; and (vii) setting a first indicator of the action in the query video to the same as a second indicator of the action performed in the one of the support videos.
Owner:CZECH TECH UNIV IN PRAGUE +1

Editing strategy scheduling method and device, electronic device and storage medium

The invention relates to an editing strategy scheduling method and device, an electronic device and a storage medium, and the method comprises the steps: receiving an original material, a copywriting and an editing style instruction inputted by a user, and generating a corresponding dubbing audio according to the copywriting; performing multi-dimensional analysis on the original material to generate a lens-level structured index; performing semantic analysis on the copywriting to obtain copywriting semantic features; performing similarity retrieval based on the copywriting semantic features and the multi-modal semantic features of the materials to obtain a candidate shot set matched with the copywriting semantic features; generating a global style vector and a target rhythm curve through the first agent; and taking the global style vector and the target rhythm curve as control signals, driving a second agent to select and trim shots from the candidate shot set, recombining the editing sequence, and outputting a final editing sequence and an editing jump point structure. According to the style vector, the rhythm target curve and reinforcement learning, the editing style is met, and the intelligent agent editing stylized presentation is achieved.
Owner:ZHEJIANG HUAZHI WANXIANG TECHNOLOGY CO LTD

Media content memory retrieval

Aspects of the subject disclosure may include, for example, a media consumption database that stores data elements describing conditions under which electronic media content is consumed by a user on an electronic device. A search of the media consumption database based on at least a portion of the conditions may result in at least a portion of the electronic media content to be re-presented to an electronic device of the user Other embodiments are disclosed.
Owner:AT&T INTELLECTUAL PROPERTY I L P

End-to-end multi-task video retrieval with cross attention

A method includes obtaining a video and a relational spatio-temporal query, and identifying at least one type of the relational spatio-temporal query. The at least one type of identification of the relational spatio-temporal query represents at least one of: an activity type, an object type, or a temporal type. The method further includes learning correlations between activities, objects, and time in the video using one or more cross attention models. The method further includes obtaining one or more predictions generated using one or more outputs of the one or more cross attention models based on the identified at least one type of relational spatiotemporal query. Further, the method includes generating a response to the relational spatiotemporal query based on the one or more predictions.
Owner:SAMSUNG ELECTRONICS CO LTD

system

We provide the system. [Solution] A means for automatically identifying the voice of a specific person, adding a timestamp, and recording it as audio data, A means of recognizing an object held by a specific person as image data and recording it with a timestamp, A means for editing and generating a natural language record based on the above audio data and image data, A means for automatically transmitting this record to an external communication application via a communication network, A means for performing behavior recognition within a specific environment and generating messages corresponding to specific behaviors, A means of transmitting the generated message to an information terminal, A system that includes this.
Owner:SOFTBANK GROUP CORP

Camera gun video resource positioning method and related equipment

The invention discloses a camera gun video resource positioning method and related equipment. The method comprises the following steps: acquiring camera gun information of each financial network; obtaining structured risk model data for the daily recovery high-risk model troubleshooting index; determining a target financial network point corresponding to each piece of risk model data in the structured risk model data; according to the structured risk model data, the target financial network point corresponding to each piece of risk model data, the camera gun information of each financial network point and the current popularity value of each camera gun, determining a plurality of current retrieval camera guns; according to the current user operation log, updating the current popularity value corresponding to each camera gun; and according to the updated current popularity value corresponding to each camera gun, determining a plurality of updated current retrieval camera guns. The method can achieve the precise pushing of the video resources needed by the daily redisk of the website, improves the video verification efficiency, and can be widely applied to the technical field of artificial intelligence.
Owner:GUANGDONG BRANCH OF CHINA POST GRP CO LTD

Video-based omnibearing remote rehabilitation training system and method, and medium

The invention discloses a video-based omni-directional remote rehabilitation training system, a video-based omni-directional remote rehabilitation training method and a medium, which are characterized in that a digital standardized training video library classified according to four stages of PT physical therapy is established, and more than 700 digital standardized training videos such as joint activity, balance training, gait correction and the like are covered; it is ensured that the training content is scientific and covers the whole period and all directions of rehabilitation training; rehabilitators generate personalized training schemes in combination with rehabilitation scene types (hospitalization / home / community / old-age care institutions) and clinical data and in combination with remote video evaluation of the rehabilitators, and training requirements in different environments are met. By capturing actions of limb joint angles, joint point motion trails and muscle force changes, a training action deviation rate is calculated and is fed back and output to a rehabilitation teacher in a grading manner, so that the training standardability is remarkably improved; rehabilitators can manage multiple patients online at the same time, remotely check training percentage data, adjust schemes and generate rehabilitation notes, and efficient, accurate and personalized rehabilitation guidance services are provided for the patients.
Owner:FALCON HEALTH TECHNOLOGY (SHANGHAI) CO LTD

A video management method, device, storage medium and system

This invention discloses a video management method, apparatus, storage medium, and system. The method includes: a main electronic device acquiring video storage information sent by various slave electronic devices in real time; the main electronic device determining video stream information corresponding to each monitoring device based on the video storage information; wherein the video stream information includes time distribution information of at least one video segment and storage block information of the target storage block where the at least one video segment is located in the slave electronic device, the at least one video segment is acquired by the same monitoring device, and the at least one video segment is stored in at least one slave electronic device; when a video stream query command sent by a client is detected, the main electronic device acquires the target video stream information corresponding to the video stream query command and feeds back the target video stream information to the client. The technical solution provided by this invention can effectively ensure the integrity and availability of the videos collected by the various monitoring devices managed by the video management system.
Owner:ZHEJIANG UNIVIEW TECH CO LTD

Multimodal data processing for content retrieval systems and applications

In various examples, multimodal data processing for content retrieval systems and applications is described herein. Systems and methods described herein may convert different modalities of data into a common type of modality. For instance, content data representing a video may be separated into audio data representing sound corresponding to the video—such as speech—along with video data representing frames of the video. The audio data may then be processed using one or more models to generate first text corresponding to a transcript of the speech. Additionally, the video data may be processed to identify specific keyframes that provide important information associated with the video. The keyframes may then be processed using one or more models to generate second text describing the keyframes. The systems and methods may then combine the text from the different modalities and generate data for storage in one or more databases.
Owner:NVIDIA CORP

Systems and methods for improved searching and categorization of media content items based on their destination - Patents.com

To provide systems and methods for improved searching and categorizing of media content items on the basis of a destination for the media content items.SOLUTION: A method includes: a step 702 of receiving, by a user computing device, data that describes a destination for a media content item including the media content item and a digital location such as website and social networking page; a step 704 of selecting one or more media content items on the basis of the data that describes the destination for the media content item; and a step 706 of displaying the selected media content item in a dynamic keyboard interface by the user computing device.SELECTED DRAWING: Figure 7
Owner:GOOGLE LLC

A method and system for weakly supervised location of video clips based on a large-scale video corpus

The present invention relates to the technical field of video data recognition. For the obtained training dataset, self-supervised learning is used to extract common semantic information between text and video, and based on the semantic information, steps are taken to obtain fused semantic video features; for the fused semantic video features and corresponding text features, multi-scale contrast learning is performed using a weak-supervision method to determine the spatial mapping relationship between video features and text features, map them to a metric space, and obtain a trained metric space; steps are taken to obtain a search query, search for text features similar to the search query in the trained metric space, and use the video clip corresponding to the text feature with the highest similarity as the video positioning result. The present invention provides a weak-supervision positioning method and system for video clips based on a large-scale video corpus. The positioning method of the present invention can realize directly determining the position of a video clip accurately and quickly from a large-scale video database.
Owner:SHANDONG JIANZHU UNIV

Neural network retraining based on image classification feedback

Described herein are systems and methods that highlight target objects in media content. The detection system trains a neural network to identify objects within an image. The detection system receives a text input requesting a search within the image and applies a search language model to the text input, which identifies a target object associated with the requested search. The detection system applies the neural network to the image to identify instances of the target object. The detection system modifies a user interface to include the image and modifies the image to highlight the identified instances of the target object. The detection system receives feedback that modifies the highlighted instances of the target object within the user interface and retrains the neural network based on the modified highlighted instances.
Owner:MATROID INC

Autonomous activity monitoring system and method

A system for automatically monitoring activity on an athletic activity area is provided. The network device includes artificial intelligence configured to automatically identify objects and gestures from video received from cameras disposed at the activity area. The artificial intelligence automatically edits the video based on objects and gestures identified from the video and generates a video file including predetermined objects and gestures. The artificial intelligence may generate a video including each trial completed by a particular athlete and may provide the video automatically to the player or a third party.
Owner:HOLE IN ONE MEDIA INC

Display device and operating method thereof

A display device according to an embodiment of the present disclosure comprises: a controller for receiving a speech command and obtaining an official content name mapped to an uttered content name included in the speech command; and a display for outputting an operation result corresponding to a speech command obtained by changing the uttered content name to the official content name. Thus, a display device in which a result intended by a user is provided even when a content name in a speech command is not accurate, and an operating method thereof, may be provided.
Owner:LG ELECTRONICS INC

Video identification method and apparatus, computer equipment, and computer program

A video identification method executed by a computer device, comprising: a step (202) of acquiring a target video and a video set reference video in a video series video set, where the video series video set includes videos belonging to the same series; a step (204) of identifying a video set local similar segment in the target video to the video set reference video based on a first matching result obtained by video frame matching between the target video and the video set reference video; a step (206) of acquiring a platform reference video from a video platform to which the target video belongs; a step (208) of identifying a platform global similar segment in the target video to the platform reference video based on a second matching result obtained by video frame matching between the target video and the platform reference video; and a step (210) of determining an overall similar segment in the target video to the video set reference video and the platform reference video based on the positions in the target video of the video set local similar segment and the platform global similar segment, respectively.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Distributed video storage and search with edge computing

Systems and methods are provided for distributed video storage and search with edge computing. The method may comprise caching a first portion of data on a first device. The method may further comprise determining, at a second device, whether the first device has the first portion of data. The determining may be based on whether the first piece of data satisfies a specified criterion. The method may further comprise sending the data, or a portion of the data, and / or a representation of the data from the first device to a third device.
Owner:NETRADYNE INC