Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

214results about "Video data indexing" patented technology

Multimodal ai-based search for digital assets

Embodiments of the present disclosure relate to multimodal AI-based search for digital assets via an indexing and / or search pipeline. With respect to the indexing pipeline, some embodiments obtain first data and second data associated with a first digital asset. Such data represents different data types or modalities of the same digital asset. After obtaining the first and second data, some embodiments then generate a composite index. After the composite index is built such index can then be used to execute a query via the search pipeline. To execute the query some embodiments compute a relevance score for each digital asset, of multiple digital assets, based at least in part on a measure in which each digital asset satisfies one or more parameters or conditions for two or more data types of the query. Various embodiments then rank each digital asset and present one or more associated indicators.
Owner:NVIDIA CORP

Knowledge graph-based long video key frame retrieval method and device

The invention relates to the technical field of multi-mode intelligent video understanding, and provides a long video key frame retrieval method and device based on a knowledge graph. Through the processes of frame-level subtitle generation, frame-level knowledge graph construction, similarity video segmentation and fragment and abstract generation, a long video knowledge graph construction assembly line is constructed, and structured modeling of long video semantic content is realized. By setting a two-stage retrieval mechanism, higher retrieval precision can be obtained while the efficiency is ensured. Vector matching and multi-hop neighbor extension are carried out on a unified knowledge graph, and a node set strongly related to the problem is positioned, so that the semantic gap between the natural language problem and a structured graph is reduced, and the accuracy of key frame selection is improved. By setting an iterative retrieval mechanism, the retrieved key frame can be used as a basis for answering a question text to the maximum extent.
Owner:NAT UNIV OF DEFENSE TECH

Information processing system and methods for clinical video retrieval

The present disclosure generally relates to an integrated approach for retrieving biomedical information from clinical video presentations. In particular, the present disclosure is directed to video retrieval systems and methods of text-video retrieval from clinical video presentations.
Owner:THE CURATORS OF THE UNIVERSITY OF MISSOURI

Internet of Things video monitoring big data privacy protection and efficient retrieval method based on artificial intelligence

The invention discloses an Internet of Things video monitoring big data privacy protection and efficient retrieval method based on artificial intelligence, belongs to the crossing field of artificial intelligence, Internet of Things and information security, and is suitable for vehicle-mounted, park, battery swap stations and other scenes. The method comprises the following steps: firstly, establishing a self-adaptive acquisition framework, and realizing multi-protocol switching and video preprocessing; secondly, extracting privacy information through an improved YOLO algorithm and a Graph-Cut technology, and combining reversible watermark embedding; constructing a multi-dimensional privacy level model for hierarchical encryption, and matching two-level storage and three-level index; then, two-factor authentication authorization is performed to extract privacy and optimize retrieval; and finally, the system state is monitored in real time and self-adaptive optimization is performed. According to the method, the problems of protocol heterogeneity, insufficient privacy protection, low retrieval efficiency and the like can be solved, and privacy security, storage overhead and retrieval efficiency are balanced.
Owner:ZHEJIANG HAISHI HUAYUE DIGITAL TECHNOLOGY CO LTD

Multi-level alignment video text retrieval method and system based on multi-modal large model

The invention provides a multi-level alignment video text retrieval method and system based on a multi-modal large model. The method comprises the steps that S1, sparse sampling is carried out to obtain a video frame sequence needing to be used for retrieval; s2, extracting video feature representation and text feature representation; s3, obtaining a video-text global level similarity; s4, selecting a plurality of video-text expert multi-level alignment networks through a routing network; s5, the total loss is obtained through calculation, the model is adjusted and updated, and a target video text retrieval model is obtained; s6, according to the video text retrieval model obtained through training, similarity calculation is conducted on videos and texts used for retrieval and video text data in a database, and videos or texts which are related to the videos and texts and have the highest similarity are obtained; by applying the technical scheme, sparse and refined execution path selection of the retrieval task is realized, and the accuracy of video text retrieval is improved through multi-level semantic alignment.
Owner:FUZHOU UNIV

Virtual behavior processing method and system based on semantic decoupling and elastic coupling

The invention discloses a virtual character behavior processing method and system based on semantic decoupling and elastic coupling and a computer readable storage medium. According to the method, unstructured action data is decomposed into physical layer intention descriptors, measurable style descriptors and high-dimensional semantic layer features by utilizing a parallel physical calculation engine and an AI semantic analysis engine through a double-track feature decoupling mechanism. Besides, an elastic coupling mechanism between intentions and styles is introduced, style parameters are dynamically clamped based on a physical priority principle, and physical topology collapse caused by style overload is prevented. According to the method, a physical-semantic dual index system is constructed, database autonomous evolution based on manifold density analysis is supported, and the problems that in the prior art, virtual character action generation is poor in physical controllability, semantic understanding is lacked, and cross-scene generalization ability is weak are effectively solved.
Owner:谢云

Editing strategy scheduling method and device, electronic device and storage medium

The invention relates to an editing strategy scheduling method and device, an electronic device and a storage medium, and the method comprises the steps: receiving an original material, a copywriting and an editing style instruction inputted by a user, and generating a corresponding dubbing audio according to the copywriting; performing multi-dimensional analysis on the original material to generate a lens-level structured index; performing semantic analysis on the copywriting to obtain copywriting semantic features; performing similarity retrieval based on the copywriting semantic features and the multi-modal semantic features of the materials to obtain a candidate shot set matched with the copywriting semantic features; generating a global style vector and a target rhythm curve through the first agent; and taking the global style vector and the target rhythm curve as control signals, driving a second agent to select and trim shots from the candidate shot set, recombining the editing sequence, and outputting a final editing sequence and an editing jump point structure. According to the style vector, the rhythm target curve and reinforcement learning, the editing style is met, and the intelligent agent editing stylized presentation is achieved.
Owner:ZHEJIANG HUAZHI WANXIANG TECHNOLOGY CO LTD

Machine-Learned Model for Generating an Output Based on Image Frames Adaptively Extracted from a Video

A computing device for generating content includes one or more memories to store instructions and one or more processors to execute the instructions to perform operations, the operations including: receiving a video; receiving an input prompt associated with the video; processing the video by adaptively extracting a plurality of image frames from the video at irregular intervals, based on content of the video; and implementing one or more machine-learned models to generate an output responsive to the input prompt, based on the input prompt and the plurality of image frames adaptively extracted from the video.
Owner:GOOGLE LLC

Training method of video time positioning model, video time positioning method, equipment and medium

The invention relates to the technical field of computer vision, particularly provides a training method of a video time positioning model, a video time positioning method, equipment and a medium, and aims to solve the problem of large video time positioning error. In order to achieve the purpose, the model training method comprises the steps that multiple frames of images are sampled from a training video to serve as training data, the training data and a first preset query text are coded to obtain visual features and text query features, and a video time positioning model is trained based on the visual features and the text query features, obtaining a plurality of candidate answers, respectively calculating the relative advantage value of each candidate answer based on the real answer of the first preset query text and the plurality of candidate answers, adjusting the parameters of the video time positioning model based on each relative advantage value, and continuing to execute the step of sampling multiple frames of images from the training video. Therefore, the performance and accuracy of video time positioning can be improved.
Owner:PEKING UNIV +1

Video content retrieval method, device and terminal based on voice interaction of television system

The invention discloses a video content retrieval method and device based on television system voice interaction and a terminal, and relates to the technical field of video processing, and the method comprises the steps: when a video is played for the first time, extracting a picture frame from the video at a preset frequency, converting the picture frame into a multi-dimensional image feature vector and a corresponding video timestamp, and carrying out hierarchical storage in a database, constructing vectorized data containing visual semantic information; obtaining a voice retrieval instruction, performing intention recognition and semantic understanding, extracting a detection keyword, and generating a multi-dimensional retrieval feature vector; calculating a matching degree between the multi-dimensional retrieval feature vector and a multi-dimensional image feature vector of a video picture frame stored in a database, and screening out picture frames of which the similarity is higher than a preset similarity threshold to form a retrieval candidate matching set; and determining matched picture playing. The video content retrieval method is efficient, accurate and high in interactivity, and retrieval experience and operation efficiency of the user in the video watching process are remarkably improved.
Owner:SHENZHEN COOCAA NETWORK TECH CO LTD

An aircraft engine defect identification system based on real-time analysis of borehole exploration images

The application provides an aircraft engine defect identification system based on borehole exploration image real-time analysis. It is characterized by including: a borehole video acquisition module that acquires video images of the internal structure of aircraft engine equipment components and transmits them to an image processing engine in real time; the image processing engine automatically identifies cracks, pits, burns, notches, deformations, corrosion, material loss and other defects in the aircraft engine equipment components in each frame of video image through a neural network image recognition analysis algorithm, and transmits the defect identification analysis results to a data background; the data background is used to match the defect identification results found by the image processing engine with the records in the aircraft engine defect index database, confirm and record the defects and trends; the system realizes real-time tracking and monitoring of the defect state of aircraft components, automatically learns and accumulates a defect feature database, accurately identifies and quickly verifies defects, and promotes the application and development of intelligent maintenance and inspection technology in the field of aircraft operation and maintenance.
Owner:GUANGDONG HAOYUN INTELLIGENT TECH CO LTD

A homomorphic encryption-based intelligent retrieval method and system for video streams

The application discloses a homomorphic encryption-based intelligent video stream retrieval method and system. The method homomorphically encrypts video frames by CTU at the acquisition end, removes redundancy through plaintext thumbnail frame difference pre-screening, and only uploads significant ciphertext frames and primary features. The cloud end performs homomorphic reasoning using preset encryption model weights, obtains 512-dimensional compressed ciphertext features through encrypted principal component projection, and further constructs an encrypted inverted product quantization index. When a user searches, the cloud end converts plaintext query features into encrypted query vectors, performs asymmetric distance calculation in the ciphertext index, and returns a ciphertext ranking result. The user end uses a private key to decrypt the top-K plaintext frames. The method realizes millisecond-level retrieval without decryption, balances high throughput, low latency and strong privacy, and can be widely used in sensitive video scenes such as security, medical treatment and industrial vision.
Owner:BEIJING FUSION HSBC TECH CO LTD

Storing and retrieving unused advertisements

The exemplary embodiments relate to implementing a mechanism that is configured to select and insert a video advertisement into a video stream that is to be provided to a user device by a streaming service. This may include receiving a request for a video stream from a user device. In response to the request, transmitting a first portion of the video stream to the user device and determining that second a portion of the video stream is to include multiple video advertisements. One or more video advertisements may be selected from a database that includes a set of video advertisements that were previously removed from a further video stream. The one or more video advertisements may then be inserted into the video stream. The second portion of the video stream is then transmitted to the user device.
Owner:PARAMOUNT GLOBAL INC

Traffic incident intelligent detection method and device based on visual big language model

The invention provides a traffic incident intelligent detection method and device based on a visual big language model, and relates to the technical field of computer vision and natural language processing, and the method comprises the steps: obtaining video image data, and constructing a multi-level analysis cue word; performing low-frequency sampling on the video image data, and determining first target time period information corresponding to a target first sampling image with an abnormal traffic event in the first sampling image based on a second hierarchy analysis prompt word; performing medium-frequency sampling on the video image data, and determining a key moment of an abnormal traffic event based on a second hierarchy analysis prompt word; extracting a target video clip according to the key moment; and performing high-frequency sampling on the target video clip, and detecting a third sampling image based on the third hierarchy analysis prompt word to obtain a traffic event detection result. By adopting the traffic incident intelligent detection method and device based on the visual large language model, the processing delay is reduced, and the real-time traffic monitoring requirement is met.
Owner:ZHEJIANG ZHIJIANG INTELLIGENT TRANSPORTATION TECH CO LTD +1

Systems and Methods for Efficient Video Storage and Retrieval via Reverse Retrieval-Augmented Generation

Systems and methods are provided for selectively storing and retrieving video data by converting video segments into textual descriptions associated with key visual features and incrementally building a vector database of feature embeddings. The feature embeddings can be used in combination with text-to-video generation models to reconstruct video segments on demand. Selective storage of feature embeddings enables optimized, accurate, and contextually relevant text-to-text and text-to-video generation and video searching functionalities.
Owner:ADEIA IMAGING LLC

Machine-learned model for generating an output based on image frames adaptively extracted from a video

A computing device for generating content includes one or more memories to store instructions and one or more processors to execute the instructions to perform operations, the operations including: receiving a video; receiving an input prompt associated with the video; processing the video by adaptively extracting a plurality of image frames from the video at irregular intervals, based on content of the video; and implementing one or more machine-learned models to generate an output responsive to the input prompt, based on the input prompt and the plurality of image frames adaptively extracted from the video.
Owner:GOOGLE LLC

A method, device and system for storing and synchronously triggering augmented reality events in a panoramic video

This invention provides a method for storing augmented reality events in panoramic videos, a method for synchronously triggering events, and an apparatus. The method uses time-triggered special effects or footage in the panoramic video as trigger events. It uses the ID of the trigger event as the key and the content of the trigger event as the value to create an index for a linked list. Each trigger event is stored in the linked list in chronological order. Simultaneously, a balanced tree is built using the trigger time of each event as the key and the ID of each trigger event as the value. To insert or delete a trigger event, the method searches the balanced tree to obtain the IDs of time-adjacent trigger events, then finds the corresponding position in the linked list based on that ID for insertion or deletion. It also allows for timed triggering of each event at specific times. This invention enables convenient and quick insertion, deletion, and synchronous triggering of events in panoramic videos. By using a balanced tree to store events, it greatly reduces the computational requirements during event lookup and improves the synchronization during panoramic video playback.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Video-based omnibearing remote rehabilitation training system and method, and medium

The invention discloses a video-based omni-directional remote rehabilitation training system, a video-based omni-directional remote rehabilitation training method and a medium, which are characterized in that a digital standardized training video library classified according to four stages of PT physical therapy is established, and more than 700 digital standardized training videos such as joint activity, balance training, gait correction and the like are covered; it is ensured that the training content is scientific and covers the whole period and all directions of rehabilitation training; rehabilitators generate personalized training schemes in combination with rehabilitation scene types (hospitalization / home / community / old-age care institutions) and clinical data and in combination with remote video evaluation of the rehabilitators, and training requirements in different environments are met. By capturing actions of limb joint angles, joint point motion trails and muscle force changes, a training action deviation rate is calculated and is fed back and output to a rehabilitation teacher in a grading manner, so that the training standardability is remarkably improved; rehabilitators can manage multiple patients online at the same time, remotely check training percentage data, adjust schemes and generate rehabilitation notes, and efficient, accurate and personalized rehabilitation guidance services are provided for the patients.
Owner:FALCON HEALTH TECHNOLOGY (SHANGHAI) CO LTD

Artificial neural network based search engine circuitry

Method (140, 200) and apparatus (120, 270) for characterizing digital content (124) using an artificial neural network (ANN) engine (122, 274). Computer data sets (126, 128, 130, 132, 134, 160, 202, 232, 242, 302, 332) from a library store (124) are processed to generate a corresponding sequence of multi-dimensional embedding vectors (162, 172, 182) in a latent space (170, 180). The embedding vectors are grouped into intervals or segments (166A, 168A, 228A) of the data sets based on movement metrics (164, 166, 168, 228) associated with the embedding vectors. A representative vector, RV (174A, 184B, 210, 276) is selected for each group. Thereafter, in response to a query input (272), selected intervals among the various computer data sets are identified and output based on a similarity measure (278) between the RVs and a search vector derived from the query input (150). Further embodiments provide a transformation model (322, 334) that transforms the embedding vectors and / or the RVs from a first latent space based on a first embedding model (304, 314, 332) to a different, second latent space based on a second embedding model (316, 338).
Owner:OBVIOUSFUTURE GMBH

Image retrieval method, device and equipment for target image information

The embodiment of the invention provides an image retrieval method, device and equipment for target image information. The method comprises the following steps: acquiring to-be-retrieved target image information; analyzing and extracting the target image information to obtain target image feature information; matching the target image feature information in a preset image data structured index database, and retrieving to obtain a retrieval result containing matched target image data and structured description information associated with the target image data; outputting a retrieval result; wherein the image data structured index database is obtained through the following processes: obtaining multi-modal original data containing images, videos and text descriptions of the same target object; performing content analysis and feature extraction on the multi-modal original data to obtain corresponding structured description information; and storing the structured description information as an index structure supporting feature-based matching retrieval according to a preset storage format. According to the embodiment of the invention, the deep semantic content of the image can be efficiently understood and accurately matched.
Owner:CHINA AI MEDIA&ENTERTAINMENT TECH CO LTD

Video and audio multimodal searching system

A multimodal search system using a video query is described. The system can receive video data captured by a camera of a user device. The video data can have a sequence of image frames. Additionally, the system can receive audio data associated with the video data captured by the user device. Moreover, the system can process, using one or more machine-learned models, the sequence of image frames to generate video embeddings related to the sequence of the image frames. The video embeddings can have a plurality of image embeddings associated with the sequence of image frames. Furthermore, the system can determine one or more video results based on the video embeddings and the audio data. Subsequently, the system can transmit, to the user device, the one or more video results.
Owner:GOOGLE LLC

A risk deployment method, device, equipment and medium of a honghome device

This invention discloses a risk control method, apparatus, device, and medium for HarmonyOS devices. The method includes: acquiring a current risk event and determining event information of the current risk event, wherein the event information includes event content and event type; selecting a target control scheme from candidate control schemes to handle the current risk event based on the event information; wherein the target control scheme includes the target device type of the target control device required to handle the current risk event; determining the target HarmonyOS device based on the target device type and event content; networking the target HarmonyOS devices based on a distributed soft bus and sending control commands to each target HarmonyOS device, wherein the control commands are used to control the target HarmonyOS devices to complete device actions to respond to the current risk event. The technical solution provided by this invention can improve the timeliness of risk control response and improve the efficiency of risk event handling.
Owner:BEIJING HONGHU SHUAN TECHNOLOGY DEVELOPMENT CO LTD

Real-time screenshot splicing-based training room desktop operation key step playback method

The application discloses a real-time screenshot splicing-based operation key step playback method in a practical training room, relates to the technical field of data processing, and balances integrity and storage economy by aligning desktop screenshots with multi-dimensional event timestamps and combining operation strength adaptive sampling, which not only completely captures operation details, but also avoids invalid screenshot waste; relying on screenshot splicing and incremental storage technology, a globally consistent operation context is constructed to solve the information loss problem under window switching and complex operation; through intelligent screening and atlas indexing of key steps, accurate positioning of operation steps is realized, and playback step-level time compression and event superposition presentation are matched to greatly improve the backtracking efficiency; the playback consistency verification and rollback correction mechanism guarantee the accuracy of reproduction, and finally, the operation review effect in practical training teaching is optimized, the students are helped to quickly master core skills, and the training quality is significantly improved.
Owner:SHANDONG PANLONG INFORMATION TECH CO LTD

A method and system for processing a face image

The application relates to a face image processing method and system, and belongs to the technical field of computers, wherein the method comprises the following steps: acquiring a face image of a property owner uploaded by a terminal; measuring a face length, a face width and a position of an intersection of a straight line where the face length and the face width are located of the face image of the property owner; calculating face proportion data of the property owner based on the face length, the face width and the position of the intersection and storing the face proportion data of the property owner into a preset face database of the property owner; before playing a target recording video, calculating face proportion data of all to-be-measured face images based on face lengths, face widths and positions of intersections of all to-be-measured face images in the target recording video; and performing fuzzy processing on to-be-measured face images in the target recording video, which satisfy the face proportion data of the face database of the property owner. The application has the effect of reducing the operation cost of a property.
Owner:SHENZHEN XIAOZHOU TECH

Video labeling method, device, equipment and computer program product

The invention provides a video labeling method, device and equipment and a computer program product, and relates to the technical field of video processing. The video labeling method comprises the following steps: acquiring a target video; labeling the target video based on a preliminary labeling strategy to obtain video labeling data, the preliminary labeling strategy comprising a labeling tool and / or configuration information for performing video labeling; performing quality evaluation on the video annotation data to determine a quality evaluation result; analyzing problems existing in the video annotation data based on the quality evaluation result, and generating an optimized annotation strategy based on the problems existing in the video annotation data, the optimized annotation strategy comprising an optimized annotation tool and / or configuration information; and marking the target video based on the optimized marking strategy, and updating the video marking data.
Owner:SHANGHAI BILIBILI TECH CO LTD

A method and system for weakly supervised location of video clips based on a large-scale video corpus

The present invention relates to the technical field of video data recognition. For the obtained training dataset, self-supervised learning is used to extract common semantic information between text and video, and based on the semantic information, steps are taken to obtain fused semantic video features; for the fused semantic video features and corresponding text features, multi-scale contrast learning is performed using a weak-supervision method to determine the spatial mapping relationship between video features and text features, map them to a metric space, and obtain a trained metric space; steps are taken to obtain a search query, search for text features similar to the search query in the trained metric space, and use the video clip corresponding to the text feature with the highest similarity as the video positioning result. The present invention provides a weak-supervision positioning method and system for video clips based on a large-scale video corpus. The positioning method of the present invention can realize directly determining the position of a video clip accurately and quickly from a large-scale video database.
Owner:SHANDONG JIANZHU UNIV

Autonomous activity monitoring system and method

A system for automatically monitoring activity on an athletic activity area is provided. The network device includes artificial intelligence configured to automatically identify objects and gestures from video received from cameras disposed at the activity area. The artificial intelligence automatically edits the video based on objects and gestures identified from the video and generates a video file including predetermined objects and gestures. The artificial intelligence may generate a video including each trial completed by a particular athlete and may provide the video automatically to the player or a third party.
Owner:HOLE IN ONE MEDIA INC

Practical training room desktop operation key step playback method based on real-time screenshot splicing

The invention discloses a practical training room desktop operation key step playback method based on real-time screenshot splicing, relates to the technical field of data processing, and aims to completely capture operation details, avoid invalid screenshot waste and balance integrity and storage economy through alignment of desktop screenshots and multi-dimensional event timestamps and adaptive sampling in combination with operation intensity. On the basis of screenshot splicing and incremental storage technologies, a globally consistent operation context is constructed, and the problem of information missing under window switching and complex operation is solved; through intelligent screening of key steps and map indexing, accurate positioning of operation steps is realized, step-level time compression and event superposition presentation of a playback end are matched, and the backtracking efficiency is greatly improved; a playback consistency verification and rollback correction mechanism guarantees reproduction accuracy, and finally, operation reproduction and efficient backtracking experience are performed with low storage transmission cost and high integrity, an operation replay effect in practical training teaching is optimized, students are assisted to quickly master core skills, and training quality is remarkably improved.
Owner:SHANDONG PANLONG INFORMATION TECH CO LTD

Image recognition-based leather product comparison identification method and system

The application relates to the technical field of image recognition, and particularly discloses a leather product comparison and identification method and system based on image recognition. The application extracts a multi-dimensional standardized feature vector reflecting inherent physical and mechanical properties by collecting dynamic deformation videos of a to-be-inspected leather and a reference sample under controlled micro force, receives identification task description parameters, maps core discriminant features from a physical and mechanical feature knowledge base, dynamically instantiates a self-adaptive identification model through a meta-learning model, generates a feature weighting scheme and a dynamic decision threshold, calculates a mechanical feature matching degree and generates a visual preliminary report, adjusts the scheme in combination with user interactive correction instructions, updates the result and feeds back data optimization meta-models. The application upgrades the identification basis to essential mechanical properties, realizes task self-adaptive decision and man-machine collaborative optimization, can improve identification reliability and scene adaptability, makes the decision process transparent and interpretable, has a continuous optimization capability, and is suitable for multi-class requirements such as authenticity identification and traceability.
Owner:海宁中国皮革城网络科技有限公司