Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

519results about "Video data indexing" patented technology

Systems and Methods for Processing, Analyzing, and Visualizing Complex Object Sets

Disclosed are methods, systems, and non-transitory computer readable memory for processing, analyzing, and visualizing complex digital object sets using large language models. For instance, a method may include receiving a set of digital objects; indexing the received set of digital objects to generate indexed data; generating a timeline prompt based on the indexed data; processing the timeline prompt to generate a timeline response; and outputting the timeline response as an interactive timeline of based on the set of digital objects.
Owner:TRANQUILITY AI INC

Systems and methods for multimodal indexing of video using machine learning

Systems, methods, and computer-readable media are disclosed for systems and methods multimodal indexing of video using machine learning. An example method may include deceiving, by a video encoder of an audio-video transformer neural network comprising one or more computer processors coupled to memory, a first frame and a second frame associated with a first segment of a video. The example method may also include receiving, by an audio encoder of the audio-video transformer neural network, an audio spectrogram comprising first audio data associated with the first segment of the video. generating, by the video encoder, a first video embedding. The example method may also include generating, by the audio encoder, a first audio embedding. The example method may also include determining a fusion of the first video embedding and the first audio embedding using a multimodal bottleneck token. The example method may also include determining an output including the first video embedding and the first audio embedding. The example method may also include determining a classification of the first portion of the video based on the output.
Owner:AMAZON TECH INC

Intelligent database video retrieval method based on big data technology

The invention relates to the technical field of intelligent information retrieval, in particular to an intelligent database video retrieval method based on a big data technology, which comprises the following steps: S1, carrying out space-time slicing processing on an original video stream; s2, extracting a visual feature vector, an audio waveform vector and a text description vector; s3, constructing a cross-modal incidence matrix; s4, constructing a layered mixed index structure which comprises a real-time updating layer and a static storage layer; s5, after a user retrieval request is received, candidate video set screening is carried out based on the layered mixed index structure, and a retrieval result is generated; and S6, updating the cross-modal incidence matrix, and synchronously updating the weight parameter of the layer. According to the method, the retrieval accuracy is improved through cross-modal feature fusion, the data storage and retrieval efficiency is optimized by adopting a hierarchical mixed index structure, and the index weight is dynamically adjusted in combination with user behavior feedback, so that efficient, accurate and intelligent video retrieval is realized.
Owner:HEBEI ZHENGTONG ARCHIVES MANAGEMENT CO LTD

Dual-stream video management

An Internet of Things (IoT) or vehicle dash cam may store both a high-resolution and low-resolution video stream on a device. The video streams are selectively accessible by remote devices. Because of the relatively smaller storage requirements of low-resolution video files, retaining of additional video data on the vehicle device (beyond what would be possible with only high-resolution video) is possible. The user may be provided an option to adjust the amount of low-resolution and high-resolution video to store on the device. A combined media file may be generated by a device to include time-synced high-resolution video, low-resolution video, and / or metadata for a particular time period.
Owner:SAMSARA INC

Factory real-time digital monitoring method and system based on Internet of Things, and storage medium

The invention relates to the technical field of industrial monitoring, and discloses a factory real-time digital monitoring method and system based on the Internet of Things, and a storage medium. The method comprises the following steps: carrying out encryption access on heterogeneous Internet of Things equipment through dual identity authentication to obtain a security authentication equipment identifier pool; performing multi-dimensional data acquisition on the equipment identification pool based on Hash verification to obtain an encrypted time sequence data chain; performing dynamic identification on the factory visual image by adopting a confidence adaptive threshold to obtain a multi-task fusion feature code; performing heterogeneous fusion on the time sequence data chain and the feature code by using a sliding window weight distribution algorithm to obtain a factory digital state fingerprint; and performing knowledge graph reasoning on the state fingerprint through three-level early warning threshold judgment to obtain a real-time risk level early warning code. The technical problems that heterogeneous Internet of Things equipment cannot be uniformly accessed and authenticated, multi-source data lacks safety fusion processing, and the factory state cannot be intelligently reasoned and pre-warned are solved.
Owner:BEIJING TIANYUAN 3D TECH CO LTD

Systems and Methods for Entity, Relationship, and Timeline Generation from Complex Object Sets

Disclosed are methods, systems, and non-transitory computer readable memory for processing, analyzing, and visualizing complex digital object sets using large language models. For instance, a method may include receiving a set of digital objects; indexing the received set of digital objects to generate indexed data; generating a timeline prompt based on the indexed data; processing the timeline prompt to generate a timeline response; and outputting the timeline response as an interactive timeline of based on the set of digital objects.
Owner:TRANQUILITY AI INC

Al-based video content analysis method and system

The invention relates to the technical field of content recognition, in particular to an Al-based video content analysis method and system, and the method comprises the following steps: carrying out the image segmentation learning based on a video sequence through employing a U-Net convolutional neural network, analyzing scenes and elements in a video frame, recognizing and isolating key visual elements in a video through network learning, and carrying out the recognition of the key visual elements in the video. Comprising objects and figures. According to the method, image segmentation learning is carried out by adopting the U-Net convolutional neural network, key visual elements in the video can be identified and isolated more accurately, a clearer basis is provided for follow-up scene change and key event tracking, video content analysis is carried out by applying the graph neural network, the identification capability of dynamic scenes and events is enhanced, and the identification efficiency is improved. Deeper structured understanding is provided for video content indexing, the video quality is analyzed through a structural similarity index evaluation method, and the reason of visual quality reduction can be accurately recognized and improved.
Owner:CHENDA (GUANGZHOU) NETWORK TECH CO LTD +1

Generation-Augmented Latent Navigation for Continuous Spatiotemporal Zoom and Rotation in Immersive Environments

A system and method for generation-augmented latent hyperspace navigation in spatiotemporal media using hierarchical and Lorentzian autoencoders. The system compresses media into latent representations while preserving geometric, temporal, and semantic relationships. A latent hyperspace manager organizes compressed data as geodesic trajectories, and a geodesic trajectory mapper computes navigation paths. Symbolic anchors provide persistent reference points, while spatiotemporal routing coordinates decisions across multiple scales. A strategy caching system preserves successful navigation patterns for reuse as procedural memory. A synthetic content generator including latent diffusion models, neural radiance fields, and context-aware refinement produces augmentation for continuous zoom, bidirectional traversal, and rotational reorientation. A user input interface and zoom controller enable interactive exploration and reconstruction, supporting applications in immersive media, visualization, and surveillance.
Owner:ATOMBEAM TECH INC

Multi-modal data pairing method and system based on deep learning

The invention provides a multi-modal data pairing method and system based on deep learning, and relates to the technical field of data processing, and the method comprises the steps: obtaining a video multi-frame sequence and a target text, and respectively extracting an overlapped frame group set and a standardized text sequence; performing spatio-temporal feature extraction and text dependency relationship coding to obtain a video time sequence vector sequence and a text vector sequence; executing cross-modal alignment search, and constructing a monotonic matching path set; calculating a semantic and action entity relationship consistency score of the paired elements on the path to obtain a comprehensive score; and determining an alignment relationship between the video and the text based on the optimal path. According to the method, accurate matching of the video and the text is realized, and the cross-modal retrieval efficiency is improved.
Owner:BEIJING YIZHUANG INTELLIGENT CITY RES INST GRP CO LTD

Multimodal ai-based search for digital assets

Embodiments of the present disclosure relate to multimodal AI-based search for digital assets via an indexing and / or search pipeline. With respect to the indexing pipeline, some embodiments obtain first data and second data associated with a first digital asset. Such data represents different data types or modalities of the same digital asset. After obtaining the first and second data, some embodiments then generate a composite index. After the composite index is built such index can then be used to execute a query via the search pipeline. To execute the query some embodiments compute a relevance score for each digital asset, of multiple digital assets, based at least in part on a measure in which each digital asset satisfies one or more parameters or conditions for two or more data types of the query. Various embodiments then rank each digital asset and present one or more associated indicators.
Owner:NVIDIA CORP

Media information processing method and device, storage medium and electronic equipment

The invention discloses a media information processing method and device, a storage medium and electronic equipment. The method comprises the steps that to-be-retrieved source media information and prompt template information are obtained, the source media information and the prompt template information are input into a target retrieval model, target media information is determined, and the target retrieval model represents a model obtained through joint training according to at least two retrieval task types; a loss function used in the training process of the target retrieval model is a joint loss function, the joint loss function is determined by at least two sub-loss functions in one-to-one correspondence with the at least two retrieval task types, and the target media information represents media information obtained by retrieving the source media information according to the prompt template information. According to the method and the device, the technical problem of relatively low retrieval efficiency of the media information caused by single multi-modal retrieval task is solved.
Owner:TENCENT TECH (BEIJING) CO LTD

Intelligent conference recording and recording method and system

The invention discloses an intelligent conference recording and recording method and system. According to the invention, cloud service equipment receives multi-modal information sent by edge equipment; processing the multi-modal information, and generating conference record information in combination with context sensing error correction, content segmentation, a dynamic segmentation strategy and a preset storage strategy; carrying out encryption processing on the conference record information based on the participant role information, and generating target format conference information; processing the target format conference information to generate initial conference scene mode information including a conference outline, a PPT and a mind map; processing the target format conference information and the initial conference scene mode information to generate conference information data after priority ranking and strategy processing; performing semantic analysis, feature fusion and cross-modal association processing on the conference information data to generate target conference scene mode information; and determining a target application program based on the target conference scene mode information, and sending the target format instruction information to the target application program.
Owner:XIAMEN RGBLINK SCI & TECH CO LTD

Video retrieval generation method and device based on sparse representation and reordering

The invention discloses a video retrieval generation method and device based on sparse representation and reordering, and relates to the technical field of cross-modal video retrieval. The method comprises the following steps: acquiring a query text, a candidate video sequence and a video retrieval generation model; obtaining a query text dense representation and a candidate video sequence dense representation according to the encoder module, the query text and the candidate video sequence; through a sparse representation module, text sparse representation is generated according to the query text dense representation, candidate video sparse representation is generated according to the candidate video sequence dense representation, and reverse index screening is performed on the candidate video sequence to obtain part of candidate videos; determining a global score of each video in the partial candidate videos through a cross attention module, and sorting the partial candidate videos; and through a generation module, generating a target text corresponding to the query text according to the sorted candidate videos and the query text. By adopting the method and the device, the retrieval efficiency is improved in a sparse representation and reordering mode.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Latent Geodesic Traversal Across Multi-Axis Hyperspaces for Real-Time Video Reconstruction and Augmentation

A system and method for latent geodesic traversal across multi-axis hyperspaces for real-time video reconstruction and augmentation. Spatiotemporal video data are compressed into navigable latent representations using hierarchical and Lorentzian autoencoders that preserve geometric and temporal structure. A geodesic traversal engine computes paths across spatial, temporal, spectral, and semantic axes, guided by symbolic anchors and spatiotemporal routing protocols. A correlation network restores fine detail, while an augmentation generator synthesizes additional or counterfactual content to enable infinite zoom, continuous multi-scale exploration, and temporally coherent augmentation. A strategy caching system preserves successful traversal patterns for reuse, supporting persistent learning and adaptive real-time performance.
Owner:ATOMBEAM TECH INC

Multi-dimensional time sequence data compression and rapid retrieval method and system

The invention discloses a multi-dimensional time sequence data compression and rapid retrieval method and system, and relates to the technical field of large data compression retrieval. The multi-dimensional time sequence data compression and rapid retrieval method comprises the following steps: S1, collecting and preprocessing multi-source data in a monitoring abstract video stream, and constructing a standardized time sequence behavior data set; s2, analyzing behavior characteristics of frame segments in the sliding window, and dynamically adjusting anchor point labeling and compression strategies; s3, evaluating the coverage integrity of anchor point information in a compression section, and driving generation of an index path; s4, comprehensively evaluating the path behavior association strength and dynamically adjusting a loading decision; and S5, verifying the matching integrity of the behavior anchor point field and the jump pointer, and guaranteeing the localizability and jump stability of the event in the compression structure. The problems that in an intelligent monitoring scheme, video abstract compression is not bound with behavior semantics, a fine-grained positioning mechanism based on behavior labels is lacked, and key action segments cannot be directly positioned during user playback are solved.
Owner:TIANRUI TECHNOLOGY (TIANJIN) CO LTD

Methods and systems for segmenting video content based on speech data and for retreiving video segments to generate videos

A method includes receiving a series of video segments and providing the series of video segments as input to a first machine learning model to produce text data. The text data is provided as input to a second machine learning model to produce categorized text data that includes a classification indication. The classification indication is added to metadata of the video segment, and the categorized text data is provided as input to a third machine learning model to produce a semantic vector. The method also includes causing the video segment and the metadata that includes the classification indication to be stored at a location of a database based on the semantic vector, the database being configured to be searched based on a search query associated with the semantic vector.
Owner:VIDEOFORCEAI INC

Multi-modal accident scene library construction method for end-to-end automatic driving test

The invention belongs to the technical field of automatic driving test, and particularly relates to a multi-mode accident scene library construction method for end-to-end automatic driving test. The method specifically comprises the following steps: step 1, extracting text information based on an MCAT-BiLSTM-CRF algorithm; 2, designing the ontology architecture, and storing accident scene information by using a knowledge graph to obtain an accident text knowledge base; step 3, derivative expansion is carried out on the accident scene element combination based on a HyCon-Sg-Net algorithm; step 4, performing enhanced fine tuning on the Open-Sora model to realize modal conversion from the text knowledge base to the accident video database so as to construct a multi-modal accident scene library; the method can be used for testing the performance of an end-to-end automatic driving automobile in an extreme accident scene, and the adaptive capacity of an automatic driving algorithm to the extreme accident scene is remarkably improved.
Owner:JILIN UNIVERSITY

Long video content information acquisition method based on OCR (Optical Character Recognition) and voice recognition technology

The invention discloses a long video content information acquisition method based on an OCR and voice recognition technology. The method comprises the following steps: S1, carrying out preprocessing n on input long video data to extract an image frame sequence and an audio stream; s2, inputting the image frame sequence into an OCR recognition module, inputting the audio stream into an ASR recognition module, and obtaining a preliminary recognition result; s3, constructing a multi-target fitness function, and optimizing an OCR and ASR parameter combination by using a Kanglizard optimization algorithm; s4, respectively applying the optimal parameter group to an OCR identification module and an ASR identification module to obtain an optimized identification result; s5, constructing a fusion factor graph, executing edge message passing by adopting a belief propagation algorithm, and generating a multi-modal semantic block set; and S6, processing the multi-modal semantic block set to generate a unified multi-modal content information set. According to the invention, through fusion of the horny lizard optimization algorithm and the belief propagation mechanism, high-precision recognition and multi-modal semantic consistency extraction of the image text and the voice information in the long video are realized.
Owner:华电(海西)新能源有限公司

Knowledge graph-based long video key frame retrieval method and device

The invention relates to the technical field of multi-mode intelligent video understanding, and provides a long video key frame retrieval method and device based on a knowledge graph. Through the processes of frame-level subtitle generation, frame-level knowledge graph construction, similarity video segmentation and fragment and abstract generation, a long video knowledge graph construction assembly line is constructed, and structured modeling of long video semantic content is realized. By setting a two-stage retrieval mechanism, higher retrieval precision can be obtained while the efficiency is ensured. Vector matching and multi-hop neighbor extension are carried out on a unified knowledge graph, and a node set strongly related to the problem is positioned, so that the semantic gap between the natural language problem and a structured graph is reduced, and the accuracy of key frame selection is improved. By setting an iterative retrieval mechanism, the retrieved key frame can be used as a basis for answering a question text to the maximum extent.
Owner:NAT UNIV OF DEFENSE TECH

Video processing method and system based on big data technology

The invention discloses a video processing method and system based on a big data technology, and constructs an intelligent video processing system with semantic driving, man-machine collaboration and continuous evolution by fusing advanced technologies such as big data processing, bimodal AI analysis, knowledge graph, natural language understanding and feedback learning. Compared with a traditional video monitoring system, the scheme has the advantages that the retrieval efficiency, the semantic understanding depth, the event association analysis capability, the user interaction experience, the system self-optimization capability and the like are remarkably improved, and the problems of incomplete seeing, difficulty in finding, inaccuracy in judgment and poor use are effectively solved.
Owner:HANGZHOU LINPIN SECURITY TECH CO LTD

Semantic-based navigation of temporally sequenced content

A method and system for semantic-based navigation of temporally sequenced content such as videos interprets the image and audio-based content by applying computer-implemented neural networks and performs multi-modal inferences of temporally aligned content. The multi-modal inferences may be performed by means of the application of vectorized embeddings and / or by application of semantic chaining techniques. The multi-modal inferences are applied to generate navigational indicators and / or responses to user inputs that comprise natural language or images. The navigational indicators and responses to user inputs may be personalized based upon user behaviors.
Owner:MANYWORLDS INC

Video indexing system using parallel decoding

A video analysis system performs indexing of one or more videos in parallel by using pipelines executed by task processors. A task processor may include compute resources configured on a cloud infrastructure or on-premise compute resources. The video analysis system receives requests to index one or more videos. In one instance, the requests may be from users of client devices with requests to index the videos, such that the indexed information can be used to perform downstream applications, such as search query-based retrieval, and the like. The video analysis system performs the indexing process in parallel, so that significant bottlenecks can be eliminated compared to existing methods.
Owner:TWELVE LABS INC

Sparse auto-encoder-based lexical element construction method and system

The invention discloses a lexical element construction method and system based on a sparse auto-encoder. The method comprises the following steps: acquiring a semantic embedding vector of an article; a weight sharing strategy is adopted, and the similarity between a codebook vector and original embedding is calculated through a weight-shared codebook; performing top-K sparsification on the similarity, retaining K codebook vectors with the maximum similarity value, and performing zero setting on the rest to generate sparse representation; generating a discrete lexical element sequence according to positions and numerical values of non-zero elements in sparse representation; reconstructing semantic embedding through a decoder; and combining the reconstruction loss, the orthogonal constraint and the diversity regularization optimization model. According to the method, through sparse representation and orthogonal constraint optimization, the quality and recommendation performance of article representation are remarkably improved. According to the method, a sparse self-encoder is combined with a trainable codebook to directly learn sparse representation of object semantic features, and independence and semantic uniqueness of codebook vectors are ensured through orthogonal constraints, so that the problems of training imbalance and embedding collapse are effectively relieved.
Owner:UNIV OF SCI & TECH OF CHINA

Video space-time retrieval method and device based on grid coding

The invention discloses a grid-coding-based video space-time retrieval method and device, the grid-coding-based video space-time retrieval method is executed in computing equipment, and the method comprises the following steps: mapping a coordinate range of a target area into a corresponding grid code set according to a Beidou grid standard on the basis of the coordinate range of the target area; acquiring video data with an intersection or inclusion relationship between the spatial range of the shot object and the target area; combining the shooting time of any video data with each grid code to generate a space-time grid code set; constructing an index structure corresponding to the video data by taking each time-space grid code in the time-space grid code set as a main key; and in response to query information of the user, analyzing a space-time range in the query of the user, mapping the space-time range into a corresponding space-time grid code range, executing retrieval based on the range, and returning hit video data. According to the method, rapid positioning and accurate retrieval of large-scale aerial photography video data are realized.
Owner:BEIJING ZHIWANG YILIAN TECH CO LTD

Traffic comprehensive law enforcement situation awareness, study and judgment system based on big data and cloud computing

The invention discloses a traffic comprehensive law enforcement situation awareness, study and judgment system based on big data and cloud computing, and relates to the technical field of traffic law enforcement management. The problem that a traffic law enforcement platform based on edge calculation has calculation and storage bottlenecks and cannot timely and accurately recognize novel and rare traffic illegal behaviors when dealing with extremely complex traffic scenes and mass data is solved. According to the invention, the law enforcement scene and the vehicle abnormal condition are identified through data collected by various devices, illegal behaviors are identified and early warned in time by combining audio and video analysis, the law enforcement accuracy and efficiency are improved, the road safety and the market order are guaranteed, the traffic law enforcement situation model carries out clustering analysis on traffic data, a visual situation map is generated, and law enforcement decision is assisted. Resource configuration is optimized, illegal trend is predicted, a targeted scheme is formulated, law enforcement scientificity is improved, the system is optimized by analyzing feedback information of law enforcement officers, key information is pushed in real time, collaborative law enforcement is promoted, and comprehensive law enforcement efficiency is improved.
Owner:JINLING INST OF TECH

Automated audio description system and method

An audio description system includes a memory and a processor. The memory stores source media comprising frames positioned within the source media according to a time index. The processor is configured to generate, using an image-to-text model, a textual description of each frame; identify intervals within the time index, each interval encompassing one or more positions of one or more frames; identify placement periods within the time index, each placement period being temporally proximal to an interval; generate a summary description based on at least one textual description of at least one frame positioned within a selected interval temporally proximal to a placement period; and associate the summary description with the placement period.
Owner:3PLAY MEDIA

Multi-modal intelligent detection method for practical training and practical operation behaviors

The embodiment of the invention discloses a multi-mode intelligent detection method for practical training and practical operation behaviors, and relates to the technical field of artificial intelligence. The method comprises the steps of loading a plurality of sub-task processes in response to a training task selected by a student; traversing the plurality of sub-task processes in sequence, and respectively executing the following operations: triggering each piece of audio and video equipment associated with the current sub-task process; traversing each action to be detected in the current sub-task process in sequence, prompting a student to execute the current action and collecting an action video; carrying out picture detection on the action video by utilizing an association algorithm to obtain a plurality of target pictures of which the current action is detected; and determining a time period when the student completes the current action according to each target picture, extracting an audio of the time period from the action video, and analyzing the similarity between the audio and a standard term by using a large model so as to judge the audio completion degree of the current action. The embodiment of the invention can fully detect the practical training effect and standardize the practical operation process.
Owner:知学云(北京)科技股份有限公司

Information processing system and methods for clinical video retrieval

The present disclosure generally relates to an integrated approach for retrieving biomedical information from clinical video presentations. In particular, the present disclosure is directed to video retrieval systems and methods of text-video retrieval from clinical video presentations.
Owner:THE CURATORS OF THE UNIVERSITY OF MISSOURI

Traditional Chinese medicine specimen collection and informatization processing method based on image machine learning

The invention belongs to the technical field of image recognition, and discloses a traditional Chinese medicine specimen collection and informatization processing method based on image machine learning, and the method comprises the steps: carrying out the multispectral collection of a traditional Chinese medicine specimen through a special collection platform, and guaranteeing the data accuracy through a black and white bottom plate and a graduated scale; carrying out geometric correction, color correction and enhancement processing on the acquired image by using an image processing technology; extracting multi-dimensional features based on an improved SIFT algorithm and a deep convolutional neural network, and constructing a specimen feature map; performing specimen attribute identification in combination with a traditional Chinese medicine knowledge base, and establishing a traditional Chinese medicine specimen informatization model; specimen recognition is realized through multi-dimensional feature matching; evaluating the growth state of the specimen in the collection environment based on the ecological adaptability model, and verifying the accuracy of the recognition result; and finally generating a traditional Chinese medicine specimen digital file containing comprehensive information. According to the invention, digitization, intelligentization and standardization of collection, identification and management of the traditional Chinese medicine specimens are realized, and the accuracy and efficiency of identification of the traditional Chinese medicine specimens are improved.
Owner:NAT INST FOR FOOD & DRUG CONTROL