Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

348results about "Video data indexing" patented technology

Dual-stream video management

An Internet of Things (IoT) or vehicle dash cam may store both a high-resolution and low-resolution video stream on a device. The video streams are selectively accessible by remote devices. Because of the relatively smaller storage requirements of low-resolution video files, retaining of additional video data on the vehicle device (beyond what would be possible with only high-resolution video) is possible. The user may be provided an option to adjust the amount of low-resolution and high-resolution video to store on the device. A combined media file may be generated by a device to include time-synced high-resolution video, low-resolution video, and / or metadata for a particular time period.
Owner:SAMSARA INC

Factory real-time digital monitoring method and system based on Internet of Things, and storage medium

The invention relates to the technical field of industrial monitoring, and discloses a factory real-time digital monitoring method and system based on the Internet of Things, and a storage medium. The method comprises the following steps: carrying out encryption access on heterogeneous Internet of Things equipment through dual identity authentication to obtain a security authentication equipment identifier pool; performing multi-dimensional data acquisition on the equipment identification pool based on Hash verification to obtain an encrypted time sequence data chain; performing dynamic identification on the factory visual image by adopting a confidence adaptive threshold to obtain a multi-task fusion feature code; performing heterogeneous fusion on the time sequence data chain and the feature code by using a sliding window weight distribution algorithm to obtain a factory digital state fingerprint; and performing knowledge graph reasoning on the state fingerprint through three-level early warning threshold judgment to obtain a real-time risk level early warning code. The technical problems that heterogeneous Internet of Things equipment cannot be uniformly accessed and authenticated, multi-source data lacks safety fusion processing, and the factory state cannot be intelligently reasoned and pre-warned are solved.
Owner:BEIJING TIANYUAN 3D TECH CO LTD

Generation-Augmented Latent Navigation for Continuous Spatiotemporal Zoom and Rotation in Immersive Environments

A system and method for generation-augmented latent hyperspace navigation in spatiotemporal media using hierarchical and Lorentzian autoencoders. The system compresses media into latent representations while preserving geometric, temporal, and semantic relationships. A latent hyperspace manager organizes compressed data as geodesic trajectories, and a geodesic trajectory mapper computes navigation paths. Symbolic anchors provide persistent reference points, while spatiotemporal routing coordinates decisions across multiple scales. A strategy caching system preserves successful navigation patterns for reuse as procedural memory. A synthetic content generator including latent diffusion models, neural radiance fields, and context-aware refinement produces augmentation for continuous zoom, bidirectional traversal, and rotational reorientation. A user input interface and zoom controller enable interactive exploration and reconstruction, supporting applications in immersive media, visualization, and surveillance.
Owner:ATOMBEAM TECH INC

Multi-modal data pairing method and system based on deep learning

The invention provides a multi-modal data pairing method and system based on deep learning, and relates to the technical field of data processing, and the method comprises the steps: obtaining a video multi-frame sequence and a target text, and respectively extracting an overlapped frame group set and a standardized text sequence; performing spatio-temporal feature extraction and text dependency relationship coding to obtain a video time sequence vector sequence and a text vector sequence; executing cross-modal alignment search, and constructing a monotonic matching path set; calculating a semantic and action entity relationship consistency score of the paired elements on the path to obtain a comprehensive score; and determining an alignment relationship between the video and the text based on the optimal path. According to the method, accurate matching of the video and the text is realized, and the cross-modal retrieval efficiency is improved.
Owner:BEIJING YIZHUANG INTELLIGENT CITY RES INST GRP CO LTD

Multimodal ai-based search for digital assets

Embodiments of the present disclosure relate to multimodal AI-based search for digital assets via an indexing and / or search pipeline. With respect to the indexing pipeline, some embodiments obtain first data and second data associated with a first digital asset. Such data represents different data types or modalities of the same digital asset. After obtaining the first and second data, some embodiments then generate a composite index. After the composite index is built such index can then be used to execute a query via the search pipeline. To execute the query some embodiments compute a relevance score for each digital asset, of multiple digital assets, based at least in part on a measure in which each digital asset satisfies one or more parameters or conditions for two or more data types of the query. Various embodiments then rank each digital asset and present one or more associated indicators.
Owner:NVIDIA CORP

Latent Geodesic Traversal Across Multi-Axis Hyperspaces for Real-Time Video Reconstruction and Augmentation

A system and method for latent geodesic traversal across multi-axis hyperspaces for real-time video reconstruction and augmentation. Spatiotemporal video data are compressed into navigable latent representations using hierarchical and Lorentzian autoencoders that preserve geometric and temporal structure. A geodesic traversal engine computes paths across spatial, temporal, spectral, and semantic axes, guided by symbolic anchors and spatiotemporal routing protocols. A correlation network restores fine detail, while an augmentation generator synthesizes additional or counterfactual content to enable infinite zoom, continuous multi-scale exploration, and temporally coherent augmentation. A strategy caching system preserves successful traversal patterns for reuse, supporting persistent learning and adaptive real-time performance.
Owner:ATOMBEAM TECH INC

Multi-dimensional time sequence data compression and rapid retrieval method and system

The invention discloses a multi-dimensional time sequence data compression and rapid retrieval method and system, and relates to the technical field of large data compression retrieval. The multi-dimensional time sequence data compression and rapid retrieval method comprises the following steps: S1, collecting and preprocessing multi-source data in a monitoring abstract video stream, and constructing a standardized time sequence behavior data set; s2, analyzing behavior characteristics of frame segments in the sliding window, and dynamically adjusting anchor point labeling and compression strategies; s3, evaluating the coverage integrity of anchor point information in a compression section, and driving generation of an index path; s4, comprehensively evaluating the path behavior association strength and dynamically adjusting a loading decision; and S5, verifying the matching integrity of the behavior anchor point field and the jump pointer, and guaranteeing the localizability and jump stability of the event in the compression structure. The problems that in an intelligent monitoring scheme, video abstract compression is not bound with behavior semantics, a fine-grained positioning mechanism based on behavior labels is lacked, and key action segments cannot be directly positioned during user playback are solved.
Owner:TIANRUI TECHNOLOGY (TIANJIN) CO LTD

Multi-modal accident scene library construction method for end-to-end automatic driving test

The invention belongs to the technical field of automatic driving test, and particularly relates to a multi-mode accident scene library construction method for end-to-end automatic driving test. The method specifically comprises the following steps: step 1, extracting text information based on an MCAT-BiLSTM-CRF algorithm; 2, designing the ontology architecture, and storing accident scene information by using a knowledge graph to obtain an accident text knowledge base; step 3, derivative expansion is carried out on the accident scene element combination based on a HyCon-Sg-Net algorithm; step 4, performing enhanced fine tuning on the Open-Sora model to realize modal conversion from the text knowledge base to the accident video database so as to construct a multi-modal accident scene library; the method can be used for testing the performance of an end-to-end automatic driving automobile in an extreme accident scene, and the adaptive capacity of an automatic driving algorithm to the extreme accident scene is remarkably improved.
Owner:JILIN UNIVERSITY

Knowledge graph-based long video key frame retrieval method and device

The invention relates to the technical field of multi-mode intelligent video understanding, and provides a long video key frame retrieval method and device based on a knowledge graph. Through the processes of frame-level subtitle generation, frame-level knowledge graph construction, similarity video segmentation and fragment and abstract generation, a long video knowledge graph construction assembly line is constructed, and structured modeling of long video semantic content is realized. By setting a two-stage retrieval mechanism, higher retrieval precision can be obtained while the efficiency is ensured. Vector matching and multi-hop neighbor extension are carried out on a unified knowledge graph, and a node set strongly related to the problem is positioned, so that the semantic gap between the natural language problem and a structured graph is reduced, and the accuracy of key frame selection is improved. By setting an iterative retrieval mechanism, the retrieved key frame can be used as a basis for answering a question text to the maximum extent.
Owner:NAT UNIV OF DEFENSE TECH

Video processing method and system based on big data technology

The invention discloses a video processing method and system based on a big data technology, and constructs an intelligent video processing system with semantic driving, man-machine collaboration and continuous evolution by fusing advanced technologies such as big data processing, bimodal AI analysis, knowledge graph, natural language understanding and feedback learning. Compared with a traditional video monitoring system, the scheme has the advantages that the retrieval efficiency, the semantic understanding depth, the event association analysis capability, the user interaction experience, the system self-optimization capability and the like are remarkably improved, and the problems of incomplete seeing, difficulty in finding, inaccuracy in judgment and poor use are effectively solved.
Owner:HANGZHOU LINPIN SECURITY TECH CO LTD

Video space-time retrieval method and device based on grid coding

The invention discloses a grid-coding-based video space-time retrieval method and device, the grid-coding-based video space-time retrieval method is executed in computing equipment, and the method comprises the following steps: mapping a coordinate range of a target area into a corresponding grid code set according to a Beidou grid standard on the basis of the coordinate range of the target area; acquiring video data with an intersection or inclusion relationship between the spatial range of the shot object and the target area; combining the shooting time of any video data with each grid code to generate a space-time grid code set; constructing an index structure corresponding to the video data by taking each time-space grid code in the time-space grid code set as a main key; and in response to query information of the user, analyzing a space-time range in the query of the user, mapping the space-time range into a corresponding space-time grid code range, executing retrieval based on the range, and returning hit video data. According to the method, rapid positioning and accurate retrieval of large-scale aerial photography video data are realized.
Owner:BEIJING ZHIWANG YILIAN TECH CO LTD

Traffic comprehensive law enforcement situation awareness, study and judgment system based on big data and cloud computing

The invention discloses a traffic comprehensive law enforcement situation awareness, study and judgment system based on big data and cloud computing, and relates to the technical field of traffic law enforcement management. The problem that a traffic law enforcement platform based on edge calculation has calculation and storage bottlenecks and cannot timely and accurately recognize novel and rare traffic illegal behaviors when dealing with extremely complex traffic scenes and mass data is solved. According to the invention, the law enforcement scene and the vehicle abnormal condition are identified through data collected by various devices, illegal behaviors are identified and early warned in time by combining audio and video analysis, the law enforcement accuracy and efficiency are improved, the road safety and the market order are guaranteed, the traffic law enforcement situation model carries out clustering analysis on traffic data, a visual situation map is generated, and law enforcement decision is assisted. Resource configuration is optimized, illegal trend is predicted, a targeted scheme is formulated, law enforcement scientificity is improved, the system is optimized by analyzing feedback information of law enforcement officers, key information is pushed in real time, collaborative law enforcement is promoted, and comprehensive law enforcement efficiency is improved.
Owner:JINLING INST OF TECH

Automated audio description system and method

An audio description system includes a memory and a processor. The memory stores source media comprising frames positioned within the source media according to a time index. The processor is configured to generate, using an image-to-text model, a textual description of each frame; identify intervals within the time index, each interval encompassing one or more positions of one or more frames; identify placement periods within the time index, each placement period being temporally proximal to an interval; generate a summary description based on at least one textual description of at least one frame positioned within a selected interval temporally proximal to a placement period; and associate the summary description with the placement period.
Owner:3PLAY MEDIA

Information processing system and methods for clinical video retrieval

The present disclosure generally relates to an integrated approach for retrieving biomedical information from clinical video presentations. In particular, the present disclosure is directed to video retrieval systems and methods of text-video retrieval from clinical video presentations.
Owner:THE CURATORS OF THE UNIVERSITY OF MISSOURI

Internet of Things video monitoring big data privacy protection and efficient retrieval method based on artificial intelligence

The invention discloses an Internet of Things video monitoring big data privacy protection and efficient retrieval method based on artificial intelligence, belongs to the crossing field of artificial intelligence, Internet of Things and information security, and is suitable for vehicle-mounted, park, battery swap stations and other scenes. The method comprises the following steps: firstly, establishing a self-adaptive acquisition framework, and realizing multi-protocol switching and video preprocessing; secondly, extracting privacy information through an improved YOLO algorithm and a Graph-Cut technology, and combining reversible watermark embedding; constructing a multi-dimensional privacy level model for hierarchical encryption, and matching two-level storage and three-level index; then, two-factor authentication authorization is performed to extract privacy and optimize retrieval; and finally, the system state is monitored in real time and self-adaptive optimization is performed. According to the method, the problems of protocol heterogeneity, insufficient privacy protection, low retrieval efficiency and the like can be solved, and privacy security, storage overhead and retrieval efficiency are balanced.
Owner:ZHEJIANG HAISHI HUAYUE DIGITAL TECHNOLOGY CO LTD

Multi-level alignment video text retrieval method and system based on multi-modal large model

The invention provides a multi-level alignment video text retrieval method and system based on a multi-modal large model. The method comprises the steps that S1, sparse sampling is carried out to obtain a video frame sequence needing to be used for retrieval; s2, extracting video feature representation and text feature representation; s3, obtaining a video-text global level similarity; s4, selecting a plurality of video-text expert multi-level alignment networks through a routing network; s5, the total loss is obtained through calculation, the model is adjusted and updated, and a target video text retrieval model is obtained; s6, according to the video text retrieval model obtained through training, similarity calculation is conducted on videos and texts used for retrieval and video text data in a database, and videos or texts which are related to the videos and texts and have the highest similarity are obtained; by applying the technical scheme, sparse and refined execution path selection of the retrieval task is realized, and the accuracy of video text retrieval is improved through multi-level semantic alignment.
Owner:FUZHOU UNIV

Virtual behavior processing method and system based on semantic decoupling and elastic coupling

The invention discloses a virtual character behavior processing method and system based on semantic decoupling and elastic coupling and a computer readable storage medium. According to the method, unstructured action data is decomposed into physical layer intention descriptors, measurable style descriptors and high-dimensional semantic layer features by utilizing a parallel physical calculation engine and an AI semantic analysis engine through a double-track feature decoupling mechanism. Besides, an elastic coupling mechanism between intentions and styles is introduced, style parameters are dynamically clamped based on a physical priority principle, and physical topology collapse caused by style overload is prevented. According to the method, a physical-semantic dual index system is constructed, database autonomous evolution based on manifold density analysis is supported, and the problems that in the prior art, virtual character action generation is poor in physical controllability, semantic understanding is lacked, and cross-scene generalization ability is weak are effectively solved.
Owner:谢云

File format for selective streaming of data

Provided are systems and methods for selectively streaming content using a new binary file format. In one example, a method may include storing a plurality of binary files, establishing a network communication session between the computing system and a computing terminal via a network, receiving identifiers of one or more intervals of time from the computing terminal via the network communication session, identifying a subset of data within a data section of a binary file stored in memory which is mapped to the identifiers of the one or more intervals of time based on an index within the binary file, and transmitting a stream including the identified subset of data to the computing terminal via the network communication session.
Owner:EMBARK TRUCKS INC

An intelligent clipping application method and device through picture recognition, equipment and medium

The application relates to an intelligent clipping application method through picture recognition, which comprises the following steps: acquiring medical image video data to be clipped; performing grouped image recognition on the medical image video data to be clipped to obtain a key frame index; and performing clipping on the medical image video data to be clipped based on the key frame index. The application does not require a clipping personnel with medical knowledge to perform clipping, and the clipping personnel can complete the video clipping operation without watching the whole content of the video or repeatedly watching the video for multiple times, so that the efficiency of the video clipping is improved. The application also relates to an intelligent clipping application device through picture recognition, a storage medium and equipment.
Owner:北京泽桥数智科技有限公司

Editing strategy scheduling method and device, electronic device and storage medium

The invention relates to an editing strategy scheduling method and device, an electronic device and a storage medium, and the method comprises the steps: receiving an original material, a copywriting and an editing style instruction inputted by a user, and generating a corresponding dubbing audio according to the copywriting; performing multi-dimensional analysis on the original material to generate a lens-level structured index; performing semantic analysis on the copywriting to obtain copywriting semantic features; performing similarity retrieval based on the copywriting semantic features and the multi-modal semantic features of the materials to obtain a candidate shot set matched with the copywriting semantic features; generating a global style vector and a target rhythm curve through the first agent; and taking the global style vector and the target rhythm curve as control signals, driving a second agent to select and trim shots from the candidate shot set, recombining the editing sequence, and outputting a final editing sequence and an editing jump point structure. According to the style vector, the rhythm target curve and reinforcement learning, the editing style is met, and the intelligent agent editing stylized presentation is achieved.
Owner:ZHEJIANG HUAZHI WANXIANG TECHNOLOGY CO LTD

Machine-Learned Model for Generating an Output Based on Image Frames Adaptively Extracted from a Video

A computing device for generating content includes one or more memories to store instructions and one or more processors to execute the instructions to perform operations, the operations including: receiving a video; receiving an input prompt associated with the video; processing the video by adaptively extracting a plurality of image frames from the video at irregular intervals, based on content of the video; and implementing one or more machine-learned models to generate an output responsive to the input prompt, based on the input prompt and the plurality of image frames adaptively extracted from the video.
Owner:GOOGLE LLC

Pre-Training a Model Using Unlabeled Videos

Systems and methods for performing captioning for image or video data are described herein. The method can include receiving unlabeled multimedia data, and outputting, from a machine learning model, one or more captions for the multimedia data. Training the machine learning model to create these outputs can include inputting a subset of video frames and a first utterance into the machine learning model, using the machine learning model to predict a predicted utterance based on the subset of video frames and the first utterance, and updating one or more parameters of the machine learning model based on a loss function that compares the predicted utterance with the second utterance.
Owner:GOOGLE LLC

Training method of video time positioning model, video time positioning method, equipment and medium

The invention relates to the technical field of computer vision, particularly provides a training method of a video time positioning model, a video time positioning method, equipment and a medium, and aims to solve the problem of large video time positioning error. In order to achieve the purpose, the model training method comprises the steps that multiple frames of images are sampled from a training video to serve as training data, the training data and a first preset query text are coded to obtain visual features and text query features, and a video time positioning model is trained based on the visual features and the text query features, obtaining a plurality of candidate answers, respectively calculating the relative advantage value of each candidate answer based on the real answer of the first preset query text and the plurality of candidate answers, adjusting the parameters of the video time positioning model based on each relative advantage value, and continuing to execute the step of sampling multiple frames of images from the training video. Therefore, the performance and accuracy of video time positioning can be improved.
Owner:PEKING UNIV +1

Video content retrieval method, device and terminal based on voice interaction of television system

The invention discloses a video content retrieval method and device based on television system voice interaction and a terminal, and relates to the technical field of video processing, and the method comprises the steps: when a video is played for the first time, extracting a picture frame from the video at a preset frequency, converting the picture frame into a multi-dimensional image feature vector and a corresponding video timestamp, and carrying out hierarchical storage in a database, constructing vectorized data containing visual semantic information; obtaining a voice retrieval instruction, performing intention recognition and semantic understanding, extracting a detection keyword, and generating a multi-dimensional retrieval feature vector; calculating a matching degree between the multi-dimensional retrieval feature vector and a multi-dimensional image feature vector of a video picture frame stored in a database, and screening out picture frames of which the similarity is higher than a preset similarity threshold to form a retrieval candidate matching set; and determining matched picture playing. The video content retrieval method is efficient, accurate and high in interactivity, and retrieval experience and operation efficiency of the user in the video watching process are remarkably improved.
Owner:SHENZHEN COOCAA NETWORK TECH CO LTD

An aircraft engine defect identification system based on real-time analysis of borehole exploration images

The application provides an aircraft engine defect identification system based on borehole exploration image real-time analysis. It is characterized by including: a borehole video acquisition module that acquires video images of the internal structure of aircraft engine equipment components and transmits them to an image processing engine in real time; the image processing engine automatically identifies cracks, pits, burns, notches, deformations, corrosion, material loss and other defects in the aircraft engine equipment components in each frame of video image through a neural network image recognition analysis algorithm, and transmits the defect identification analysis results to a data background; the data background is used to match the defect identification results found by the image processing engine with the records in the aircraft engine defect index database, confirm and record the defects and trends; the system realizes real-time tracking and monitoring of the defect state of aircraft components, automatically learns and accumulates a defect feature database, accurately identifies and quickly verifies defects, and promotes the application and development of intelligent maintenance and inspection technology in the field of aircraft operation and maintenance.
Owner:GUANGDONG HAOYUN INTELLIGENT TECH CO LTD

A homomorphic encryption-based intelligent retrieval method and system for video streams

The application discloses a homomorphic encryption-based intelligent video stream retrieval method and system. The method homomorphically encrypts video frames by CTU at the acquisition end, removes redundancy through plaintext thumbnail frame difference pre-screening, and only uploads significant ciphertext frames and primary features. The cloud end performs homomorphic reasoning using preset encryption model weights, obtains 512-dimensional compressed ciphertext features through encrypted principal component projection, and further constructs an encrypted inverted product quantization index. When a user searches, the cloud end converts plaintext query features into encrypted query vectors, performs asymmetric distance calculation in the ciphertext index, and returns a ciphertext ranking result. The user end uses a private key to decrypt the top-K plaintext frames. The method realizes millisecond-level retrieval without decryption, balances high throughput, low latency and strong privacy, and can be widely used in sensitive video scenes such as security, medical treatment and industrial vision.
Owner:BEIJING FUSION HSBC TECH CO LTD

Intelligent power station mass monitoring video foreign matter real-time identification and distributed alarm system

The invention relates to the technical field of intelligent power station monitoring, and particularly discloses an intelligent power station mass monitoring video foreign matter real-time identification and distributed alarm system, which comprises a video clip segmentation module for segmenting each piece of monitoring video data in a video database into a plurality of video clips, and storing the plurality of video clips into a video clip database; a foreign matter judgment module judges whether each video clip is a suspected foreign matter clip or not, and if yes, the suspected foreign matter clips in a video clip database are added into a suspected foreign matter clip database; the foreign matter recognition module judges whether each suspected foreign matter segment in the suspected foreign matter segment database is a real foreign matter segment or not, and if yes, the real foreign matter segments in the suspected foreign matter segment database are removed into an alarm database; the alarm module gives an alarm according to the real foreign matter fragments in the alarm database. According to the intelligent power station mass monitoring video foreign matter real-time identification and distributed alarm system provided by the invention, the accuracy and stability of foreign matter detection are improved.
Owner:ALPHA ESS CO LTD

Storing and retrieving unused advertisements

The exemplary embodiments relate to implementing a mechanism that is configured to select and insert a video advertisement into a video stream that is to be provided to a user device by a streaming service. This may include receiving a request for a video stream from a user device. In response to the request, transmitting a first portion of the video stream to the user device and determining that second a portion of the video stream is to include multiple video advertisements. One or more video advertisements may be selected from a database that includes a set of video advertisements that were previously removed from a further video stream. The one or more video advertisements may then be inserted into the video stream. The second portion of the video stream is then transmitted to the user device.
Owner:PARAMOUNT GLOBAL INC

Traffic incident intelligent detection method and device based on visual big language model

The invention provides a traffic incident intelligent detection method and device based on a visual big language model, and relates to the technical field of computer vision and natural language processing, and the method comprises the steps: obtaining video image data, and constructing a multi-level analysis cue word; performing low-frequency sampling on the video image data, and determining first target time period information corresponding to a target first sampling image with an abnormal traffic event in the first sampling image based on a second hierarchy analysis prompt word; performing medium-frequency sampling on the video image data, and determining a key moment of an abnormal traffic event based on a second hierarchy analysis prompt word; extracting a target video clip according to the key moment; and performing high-frequency sampling on the target video clip, and detecting a third sampling image based on the third hierarchy analysis prompt word to obtain a traffic event detection result. By adopting the traffic incident intelligent detection method and device based on the visual large language model, the processing delay is reduced, and the real-time traffic monitoring requirement is met.
Owner:ZHEJIANG ZHIJIANG INTELLIGENT TRANSPORTATION TECH CO LTD +1

Systems and Methods for Efficient Video Storage and Retrieval via Reverse Retrieval-Augmented Generation

Systems and methods are provided for selectively storing and retrieving video data by converting video segments into textual descriptions associated with key visual features and incrementally building a vector database of feature embeddings. The feature embeddings can be used in combination with text-to-video generation models to reconstruct video segments on demand. Selective storage of feature embeddings enables optimized, accurate, and contextually relevant text-to-text and text-to-video generation and video searching functionalities.
Owner:ADEIA IMAGING LLC