Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

25 results about "Video retrieval" patented technology

A long video text period retrieval method for power maintenance scene

PendingCN122412645APattern recognitionVideo retrieval
The application discloses a long video text period retrieval method for a power maintenance scene, first divides the power maintenance long video into multiple candidate video clips, extracts a frame-level video feature sequence and an alignment target feature sequence of each candidate video clip, fuses the frame-level video feature sequence and the alignment target feature sequence, and obtains a fused frame-level video feature sequence; inputs a text feature sequence after encoding of a query text and the fused frame-level video feature sequence of each candidate video clip for cross-modal interaction calculation, obtains a clip-level context feature and a frame-level content feature sequence, and calculates a clip-level correlation score and a retrieval enhancement feature sequence based on the clip-level context feature and the frame-level content feature sequence; and finally, sorts and optimizes the clip-level correlation scores of all the candidate video clips, obtains a target retrieval enhancement feature sequence, and inputs the target retrieval enhancement feature sequence into a time boundary predictor to predict a video retrieval result corresponding to the input query text. The application enhances the expression capability of fine-grained content of the power maintenance scene and improves the retrieval efficiency in a multi-query scene.
Owner:STATE GRID ANHUI ULTRA HIGH VOLTAGE CO +1

An intelligent home system based on security technology

PendingCN122362909AVideo retrievalAutomatic control
This application discloses a smart home system based on security technology, relating to the fields of smart home and smart elderly care technology, including: a terminal acquisition layer, a network edge computing layer, a cloud platform layer, and an interactive application layer; the terminal acquisition layer integrates elderly care wristbands, environmental sensors, fall radar, smart cameras, door magnets, positioning tags, and companion robots; the cloud platform layer deploys a data fusion engine, an AI big data model, a video analysis module, and a positioning service engine. This invention achieves functions such as health anomaly monitoring and medication reminders, environmental safety early warning and automatic control, anti-wandering voice reminders, dangerous area protection, real-time positioning of people and objects, multimodal fall detection, voice control of home appliances and monitoring of abnormal electricity use, intelligent object search, AI big data model video retrieval, and community-linked missing person search. Through multimodal data fusion and AI big data model-driven interaction, it achieves proactive smart monitoring of the elderly's lives, improving the level of intelligence in elderly care monitoring and the user experience.
Owner:ZHEJIANG COLLEGE OF SECURITY TECH

Video frame processing method and system based on semantic guidance, and storage medium

PendingCN122112302AOvercome the shortcoming of easily selecting redundant informationFilter out background noiseSemantic analysisVideo data clustering/classificationPattern recognitionVideo retrieval
The application discloses a kind of based on semantic guide's video frame processing method, system and storage medium, comprising: obtaining original video, obtains semantic guide text;The time position information of each frame is encoded after fusion with the visual features of this frame, form time sequence enhancement features, according to time sequence enhancement features to video frame grouping, select multiple representative and diversity key frames from pre-sampling frame;Multiple candidate local cropping regions are generated for selected key frame, the semantic similarity between each candidate region and semantic guide text is calculated respectively, and the candidate region with the highest similarity to semantic guide text is cropped out. Through the semantic perception of key frame selection in time dimension, and adaptive frame cropping in spatial dimension, the fine-grained refinement and alignment of video content are realized, thereby the performance of cross-modal text-video retrieval is significantly improved.
Owner:XIANGTAN UNIV

Big data-based city-level video monitoring resource invocation management method and system

PendingCN122368114AVideo retrievalVideo monitoring
This application belongs to the field of video image processing technology, specifically relating to a city-level video surveillance resource retrieval management method and system based on big data. The method includes: acquiring raw video streams from multiple monitoring points and performing time alignment and imaging correction to obtain a corrected video sequence; performing moving target detection and feature extraction on the corrected video sequence to obtain target feature vectors; determining initial target association pairs based on the target feature vectors, the topological relationship of monitoring points, the target migration time window, and the consistency of the target migration direction, constructing a target association graph and obtaining a target tracking path; generating resource scheduling instructions based on the target tracking path to form a linked video stream; and retrieving video segments from backup monitoring points for continuity verification when the trajectory is interrupted to obtain a complete tracking result. This application can improve the continuity of cross-point target tracking and the accuracy of video retrieval.
Owner:SHANXI LONGHAI LUTONG INTELLIGENT TECH CO LTD

An elevator video retrieval and recognition method and system based on an improved neural network

PendingCN122346565AVideo retrievalData acquisition
The application belongs to the technical field of fault monitoring of feature equipment, and relates to an elevator video retrieval and identification method and system based on an improved neural network, and the technical points are as follows: real-time original video data is collected through a monitoring device, preprocessed, and video frame data is obtained; in combination with physical constraint conditions, feature screening is performed on the video frame data, the screened effective feature area is input into a twin liquid mixed neural network, feature extraction, similarity matching and time sequence optimization are completed, and the final retrieval and identification result is obtained. The application designs a device-specific physical constraint feature screening mechanism for the feature equipment such as the elevator, and eliminates interference features; a mixed neural network model is used to realize accurate extraction of device features, similarity matching and time sequence correlation capture; a complete video retrieval and identification system is constructed to form a complete process technical scheme from data collection to result application, and high-precision and high-efficiency retrieval and identification of elevator videos are realized.
Owner:SICHUAN SPECIAL EQUIP INSPECTION & RES INST

Long video retrieval method and device based on multi-scale multi-example similarity learning

The application discloses a long video retrieval method and device based on multi-scale multi-example similarity learning. The method acquires video and text preliminary features; uses coarse-to-fine coding mode to extract information of different time granularities from video segment scale and frame scale; based on video representation of two scales, uses segment scale similarity learning branch to filter out video segments most relevant to the text and obtain segment scale similarity; uses frame scale similarity learning branch to aggregate video features guided by the filtered most relevant video segments to obtain more detailed video information, and after similarity calculation with the text, frame scale similarity is obtained; a common space learning algorithm is used to learn multi-scale similarity between long videos and texts, and a model is trained in an end-to-end manner to realize text-to-long video retrieval. The application uses the idea of multi-scale multi-example learning, and can effectively solve the text-to-long video retrieval task.
Owner:ZHEJIANG GONGSHANG UNIVERSITY +2

Training-free video corpus time instant retrieval method based on adaptive calibration mechanism

The present application relates to a training-free video corpus time moment retrieval method based on an adaptive calibration mechanism, and belongs to the technical field of computer vision and multi-modal information processing. It comprises: based on the constructed query event chain and video event chain, calculating the event level similarity score, and combining the mean-variance joint scoring mechanism for cross-modal retrieval to obtain candidate proposals; in the time positioning stage, a boundary level association is established among the video candidate proposals through a cooperative mechanism, and a profit-loss dynamic feedback strategy is used to perform adaptation and iteratively optimize the time boundary of the candidate proposals. The present application realizes a closed-loop retrieval process from text semantic analysis to video accurate matching. Compared with the prior art, the present application does not need large-scale labeled data for training or parameter updating, effectively solves the generalization bottleneck and deployment cost problem of video retrieval in open domain scenarios, and significantly improves the accuracy and robustness of time positioning under zero training conditions.
Owner:KUNMING UNIV OF SCI & TECH

Video playing method and device based on voice control, electronic equipment and medium

PendingCN122157672ABiological modelsSpeech recognitionVideo retrievalNoise (video)
Embodiments of the present disclosure disclose a voice control-based video playing method and device, electronic equipment and a medium. A specific implementation of the method comprises: collecting a video playing voice signal in a noisy environment to obtain a video playing audio signal; filtering noise from the video playing audio signal to obtain a filtered audio signal; performing noise separation processing on the filtered audio signal to obtain a target audio signal; performing amplitude normalization processing on the target audio signal to obtain a normalized audio signal; performing text conversion processing on the normalized audio signal to obtain audio text information; generating an audio text vector according to the audio text information; performing video retrieval on a target video corresponding to the audio text vector to obtain a video retrieval result; and generating a video playing control instruction according to the audio text information and the video retrieval result to play the video. The implementation can reduce the time for retrieving and playing the video and the error operation for playing the target video.
Owner:ZIYAO (BEIJING) TECH CO LTD

A Video Retrieval Method and Server Based on Semantic Embedding and Video Memory Coding

This application provides a video retrieval method and server based on semantic embedding and video memory coding. The method performs visual content parsing on consecutive frames of the input video using a pre-trained visual language model, mapping the visual content of each frame to a semantic embedding vector. A video memory coding module is constructed based on the temporal series features of the semantic embedding vectors, dynamically encoding the semantic embedding vectors of consecutive frames in chronological order to generate a video memory feature sequence containing temporal correlation information. Semantic parsing is performed on the query text based on this visual language model, transforming the query text into a query semantic embedding vector in the same spatial dimension as the semantic embedding vector. Using this query semantic embedding vector as the retrieval basis, similarity matching is performed in the video memory feature sequence, and a candidate set of related video segments is determined through feature correlation analysis within a temporal window. Temporal coherence verification is performed on the candidate set, and combined with the temporal correlation degree of the semantic embedding vectors, the matching video segments are output.
Owner:HUAYUNSHENGDA(BEIJING)METEROLOGICAL TECH CO LTD

A video hash large model, a video hash model construction method and device

This invention relates to the field of video retrieval technology, specifically to a large video hash model, a method for constructing the video hash model, and an apparatus. In this invention, the base model utilizes Video-LlaMA to extract deep features from both visual and audio data, ensuring semantic quality. The cross-modal attention layer in the hash model innovatively achieves semantic-level interaction and fusion of audio and video features, making the generated fused features more discriminative. Finally, after further integration by the multimodal encoder and binarization processing by the hash layer, a compact hash code is output. Thus, the large video hash model, including the base model and the hash model, not only significantly improves the semantic representation accuracy of video content, thereby achieving higher accuracy in large-scale retrieval, but also greatly improves storage and retrieval efficiency through the generated binary hash code, providing a reliable technical foundation for accurate and rapid video retrieval and analysis in resource-constrained environments.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Video retrieval inference method and related apparatus

PendingCN122332607APattern recognitionVideo retrieval
The application discloses a video retrieval reasoning method and related device, and relates to the technical field of video processing, which comprises the following steps: obtaining a user query, retrieving a candidate video segment matched with the semantic intention of the user query, performing spatio-temporal joint up-sampling on each video frame of the candidate video segment to construct a spatio-temporal zoom-in tensor, encoding the spatio-temporal zoom-in tensor into an original visual token, performing query-aware attention dimension reduction compression on the original visual token according to the user query to obtain a compressed visual token, and generating a query reply according to the compressed visual token. The application combines spatio-temporal joint up-sampling with query-aware attention dimension reduction compression, breaks the mutual restriction bottleneck that retaining micro targets will increase background redundancy false detection and filtering background noise will lose micro target details, realizes improvement in both micro target missing detection suppression and background redundancy false detection suppression, and improves the accuracy and reliability of video retrieval reasoning.
Owner:IFLYTEK CO LTD

Video retrieval system and method based on event semantic keyframes and text queries

The present application belongs to the technical field of computer vision or artificial intelligence or intelligent video monitoring, and particularly relates to a video retrieval system and method based on event semantic key frame and text query. The system comprises video preprocessing and event detection module, key frame extraction and semantic annotation module, multi-modal embedding and index construction module connected in sequence; further comprises text query interface, result output module; the method comprises: (1) video preprocessing and event detection; (2) key frame extraction and semantic annotation; (3) multi-modal embedding and index construction; (4) text-driven cross-modal retrieval and result display. The present application realizes rapid, accurate and interpretable positioning from natural language keywords to high-value video evidence without human intervention.
Owner:CHANGFENG DIGITAL TECH (SHANDONG) CO LTD

Video aggregation method, aggregation system thereof, and storage medium

PendingCN122372702AVideo retrievalVideo storage
This invention relates to a video aggregation method, system, and storage medium, comprising the following steps: S1, constructing a video aggregation rule chain, a permission confirmation and identification rule chain, and a video storage retrieval rule chain for each region; S2, performing unified protocol adaptation and access management for video surveillance devices from different regions, manufacturers, and communication protocols, without changing the original network configuration and IP address of the devices; S3, mapping the original device accounts, platform accounts, and permission information of each region to a unified permission management system; S4, performing unified scheduling and policy management of video storage resources in each region, and supporting cross-regional video retrieval and playback; S5, coordinating the execution between the video aggregation rule chain, the permission confirmation and identification rule chain, and the video storage retrieval rule chain through a linkage control mechanism. The advantages of this invention are: retaining the original IP address and account of the devices, avoiding the high costs and configuration error risks associated with large-scale network upgrades and device resets.
Owner:SHANDONG HUAFANGYUN ENERGY SAVING INTEGRATION CO LTD

A method and system for multi-modal feature fusion and clustering for ski videos

PendingCN122289747AVideo retrievalEngineering
This invention relates to the fields of computer vision and video analysis technology, and discloses a multimodal feature fusion and clustering method and system for ski videos. It aims to solve the problems of low accuracy and poor efficiency in existing ski resort video retrieval schemes under complex scenarios with strong occlusion, high clothing similarity, and motion blur. This invention first extracts the skier's appearance features, motion features, and equipment semantic features from captured ski video clips. After projecting these three types of features onto a unified dimension, the fusion weights of each feature are adaptively adjusted based on the skier's movement speed to obtain a fused feature vector. Then, a time decay weight is introduced to adjust feature similarity, and a density clustering algorithm is used to aggregate video clips of the same skier. Simultaneously, model parameters are optimized through user feedback. This invention significantly improves the accuracy and efficiency of ski video retrieval and is suitable for ski resort video retrieval and operation service scenarios.
Owner:HUIPAI INTELLIGENT COMPUTING (HANGZHOU) TECHNOLOGY CO LTD

A video retrieval method and system for long difficult events

The application discloses a video retrieval method and system for long and difficult events, and relates to video retrieval. Video frames of a target video are extracted according to different time intervals; a to-be-retrieved event is split according to a preset causal logic split condition; event text is encoded by a text encoder; frame segment sequences are encoded by a video encoder; cosine similarity between a video feature sequence and an event vector sequence is calculated to obtain a similarity score sequence; sliding window analysis is performed on the similarity score sequence, a time range corresponding to a sliding position with the largest similarity score is selected as a target time window; a starting boundary and an ending boundary of the target time window are expanded to obtain a time positioning result of the to-be-retrieved event in the target video; and the positioning accuracy is improved.
Owner:NANJING UNIV

A data processing system for video retrieval

PendingCN122309803AVideo retrievalData processing system
This application discloses a data processing system for video retrieval, which processes text data information and image data information to be processed to obtain first information; and determines target image data information based on the first information and the image data information to be processed.
Owner:SHENZHEN TCL HIGH TECH DEVELOPMENT CO LTD

Video clip summarization system and method

The application provides a video clip abstract generation system and method, the system comprises a video content analysis module, an abstract generation module, a user interaction module and a retrieval cooperation module; the method comprises the following sub-steps: S1, video content analysis is performed and video features are extracted; S2, a clip abstract reflecting the core content of the video is generated; S3, user interaction and dynamic adjustment; S4, the abstract clip is integrated with a video retrieval system; the method improves the degree of automation through feature extraction of the video content analysis module, no longer relies on manual segment-by-segment editing, significantly improves the efficiency of abstract generation, reduces labor costs, shortens processing time and is suitable for processing massive video content; by setting a user feedback entrance, allowing users to provide real-time feedback through text or check the label, the system automatically records the feedback, adjusts the feature weight according to the user's demand, continuously improves the generated abstract according to the user's demand, and improves the intelligent level and flexibility of the system.
Owner:NANJING COMPREHENSIVE SAFETY CONSULTING CO LTD

A method for building a smart factory by integrating simulated digital AI video

PendingCN122368347AVideo retrievalDigital video
This invention relates to the field of digital twin technology, specifically to a method for constructing a smart factory by integrating simulated digital AI videos. The method includes: constructing a joint index structure comprising a physical video spatiotemporal index tree and a simulated video spatiotemporal index tree; extracting the optical flow features of the physical video and the kinematic matrices of the corresponding rigid body nodes in the simulated video; inputting both into a self-attention alignment network based on physics engine constraints, outputting a spatiotemporal deviation compensation vector; performing an affine transformation on the leaf node coordinates of the physical video spatiotemporal index tree according to the compensation vector, so that physical video frames and simulated video frames share the same global three-dimensional coordinate node in the joint index structure; and extracting the matching global three-dimensional coordinate node when a video retrieval command is received. This invention compensates for dynamic spatial errors caused by temporal jitter, aligning the physical video and simulated video in spatial dimensions, and eliminating ghosting and screen tearing in the integrated video.
Owner:GUANGDONG JIUYUN INFORMATION TECHNOLOGY CO LTD

Video retrieval method and apparatus, electronic device, and storage medium

PendingCN122309809Aeasy to understandVideo retrievalTime information
This disclosure provides a video retrieval method, apparatus, electronic device, and storage medium, relating to the field of computer technology, and particularly to the field of multimodal model technology and video processing technology. The specific implementation scheme is as follows: At least one semantic unit for each query event is determined from the retrieval text, wherein the semantic unit includes at least one of time information, subject information, and behavioral description information; based on the semantic unit, target video frames of the query event are retrieved from a video library, wherein the video library includes multiple video frames, each video frame being associated with a timestamp, subject tag, and visual feature vector; based on the semantics of the retrieval text, the target video frames are analyzed to generate retrieval results.
Owner:SHANGHAI XIAODU TECHNOLOGY CO LTD

A method for verifying long clinical videos based on evidence retrieval and tool-enhanced trajectory

The application discloses a long clinical video verification method based on evidence retrieval and tool enhanced trajectory, a decision model based on reinforcement learning is constructed, the long clinical video verification is modeled as a multi-round state-action interaction process, intermediate verification reasoning information is generated in each verification round, and verification actions are adaptively output to call a video retrieval tool to obtain local video evidence. The video retrieval tool adopts a hierarchical retrieval mechanism combining a segment level and a frame level, and realizes coarse-to-fine evidence positioning and fine verification. The local video evidence returned by the tool is continuously fed back to the verification state, forming an auditable verification trajectory of reasoning and evidence closed-loop updating. In the model training stage, an evidence center data suite and a reinforcement learning strategy based on evidence alignment are introduced, which can ensure verification accuracy, improve evidence dependency, reasoning reliability and result safety, and are suitable for intelligent medical video analysis and clinical decision support scenes.
Owner:RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE

A key frame extraction method, device, equipment and medium

The application relates to the technical field of artificial intelligence, in particular to a key frame extraction method and device, equipment and medium. Applied to a medical scene, in the application, the advantages of multi-modal information such as images, audio and text in a video are fully combined, the contribution degrees of various modes to key frame prediction are dynamically calculated, and adaptive weighted fusion of features is realized. The fusion mode based on the contribution degree avoids the problems of information redundancy or weakening of key features that may be caused by traditional fixed weight fusion, so that the feature vector after fusion can more accurately reflect the importance of the video frame. Compared with a single mode key frame extraction method, the method comprehensively considers the rich information of the video content in the visual, auditory and semantic levels, thereby effectively improving the accuracy and robustness of the key frame extraction result, and better meeting the demand of video retrieval, content summary, intelligent analysis and other practical application scenarios for high-quality key frames.
Owner:PING AN TECH (SHENZHEN) CO LTD