Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

158 results about "Video retrieval" patented technology

Universal scene retrieval analysis method and system based on multi-modal feature fusion

The invention discloses a universal scene retrieval analysis method and system based on multi-modal feature fusion, the method comprises a video analysis step and an application service step, and the application service step comprises the steps of receiving a user input request, describing a multi-dimensional standardized video tag based on a video summary, and obtaining a multi-dimensional standardized video tag; the steps of cross-modal video retrieval, dynamic knowledge enhancement question answering and interactive enhancement analysis can be executed, efficient video preprocessing is achieved by constructing an offline feature library, the retrieval precision is improved by adopting a cross-modal feature fusion technology, the analysis authority is enhanced in combination with a dynamic knowledge base, and the interactive enhancement analysis is supported to achieve abnormal early warning. The method has the advantages that the offline video processing efficiency is improved, cross-modal feature fusion retrieval is realized, and the authority of an analysis result is enhanced.
Owner:SHENZHEN KAOLA YOURAN TECHNOLOGY CO LTD

Context-aware video retrieval and inference system

Various examples, systems, and methods are disclosed relating to an agentic curation pipeline. One system can process questions and other inquiries about video content by using a combination of models and stored information. The system can receive a query related to an event in a video, selects relevant portions of the video using embeddings, and apply the selected video data and a related sub-query to a video model. The output from the video model can be used by a language model, along with stored context, to generate an answer to the original query. The system can returns the answer to the requester.
Owner:NVIDIA CORP

Video text cross-modal retrieval method based on spatio-temporal feature fusion

The invention relates to the field of artificial intelligence cross-modal retrieval, and provides a video text cross-modal retrieval method and system based on spatio-temporal feature fusion. The method comprises the following steps: carrying out key frame sampling and time sequence partitioning on an input video, extracting static visual features through a spatial feature network, and extracting motion features through a time dynamic network; a self-adaptive gating fusion module is adopted to dynamically calculate spatial-temporal feature weights and perform weighted fusion; extracting text semantic features by using a pre-training language model; constructing a double-flow projection network to map video fusion features and text features to a unified measurement space, and optimizing a feature distance by adopting a contrast loss function containing difficult negative sample mining and intra-modal constraint; and outputting a retrieval result according to the cosine similarity sequence. The system comprises four units, wherein the gating fusion module is integrated with an FPGA acceleration circuit. According to the method, mAP (at) 10 is equal to 0.78 in a UCF-101 data set, the time sequence action retrieval accuracy rate is 92.8%, and the single video retrieval delay is 23 milliseconds.
Owner:ZHEJIANG UNIV

Image video retrieval method based on domain fine-tuning large language model

The invention provides an image video retrieval method based on a domain fine-tuning large language model, which comprises the following steps: performing fine-tuning on a pre-training model to obtain a fine-tuning pre-training model for intention classification and keyword extraction; performing dynamic iteration screening on an optimal prompt template through Monte Carlo tree search in combination with a hidden Markov model (HMM); performing noise filtering on the keyword list, and predicting category labels of the filtered keywords through a conditional random field model to obtain a keyword enhancement set; combining with the user intention to generate a query condition, and obtaining a candidate resource set; and according to the similarity between the user query text and the candidate resource set, and in combination with the optimal prompt template, obtaining the resource path with the highest matching score between the user query and the candidate resource, and obtaining the retrieved image or video, so that the identification deviation possibly occurring when a general model processes proper nouns and terminologies can be effectively solved, and the user experience is improved. And the retrieval accuracy and response speed are improved, so that the retrieval accuracy and professional adaptability are improved.
Owner:HUBEI ZHONGKE NETWORK ENG

Video retrieval method and server based on semantic embedding and video memory coding

The invention provides a video retrieval method and server based on semantic embedding and video memory coding, and the method comprises the steps: executing visual content analysis on continuous frames of an input video through a pre-trained visual language large model, and mapping the visual content of each frame into a semantic embedding vector; a video memory coding module is constructed based on the time sequence features of the semantic embedding vectors, and the continuous frame semantic embedding vectors are dynamically coded according to the time sequence to generate a video memory feature sequence containing time sequence association information; performing semantic analysis on the query text based on the visual language large model, and converting the query text into a query semantic embedding vector which is in the same spatial dimension as the semantic embedding vector; taking the query semantic embedding vector as a retrieval basis, performing similarity matching in a video memory feature sequence, and determining an associated video clip candidate set through feature association analysis in a time sequence window; and performing time sequence coherence verification on the candidate set, and outputting a matched video clip in combination with the semantic embedding vector time sequence correlation degree.
Owner:HUAYUNSHENGDA(BEIJING)METEROLOGICAL TECH CO LTD

Video space-time retrieval method and device based on grid coding

The invention discloses a grid-coding-based video space-time retrieval method and device, the grid-coding-based video space-time retrieval method is executed in computing equipment, and the method comprises the following steps: mapping a coordinate range of a target area into a corresponding grid code set according to a Beidou grid standard on the basis of the coordinate range of the target area; acquiring video data with an intersection or inclusion relationship between the spatial range of the shot object and the target area; combining the shooting time of any video data with each grid code to generate a space-time grid code set; constructing an index structure corresponding to the video data by taking each time-space grid code in the time-space grid code set as a main key; and in response to query information of the user, analyzing a space-time range in the query of the user, mapping the space-time range into a corresponding space-time grid code range, executing retrieval based on the range, and returning hit video data. According to the method, rapid positioning and accurate retrieval of large-scale aerial photography video data are realized.
Owner:BEIJING ZHIWANG YILIAN TECH CO LTD

Information processing system and methods for clinical video retrieval

The present disclosure generally relates to an integrated approach for retrieving biomedical information from clinical video presentations. In particular, the present disclosure is directed to video retrieval systems and methods of text-video retrieval from clinical video presentations.
Owner:THE CURATORS OF THE UNIVERSITY OF MISSOURI

Text video retrieval method of fine-grained relation learning network based on energy perception

The invention provides a text video retrieval method of a fine-grained relation learning network based on energy perception. The method comprises the following steps: giving a query text and a video clip; inputting the query text into a text encoder of the CLIP, inputting the video clip into a text encoder image encoder of the CLIP, and extracting to obtain text embedding and frame embedding; inputting the text embedding and the frame embedding into a fine-grained relation learning network for text enhancement operation to obtain enhanced text embedding; taking the enhanced text embedding as a frame fusion condition to carry out frame fusion operation on the frame embedding to obtain video embedding; calculating the similarity of a text-video pair formed by the query text and the video embedding based on a cosine similarity function; and selecting the text-video pair with the highest similarity as a retrieval output result of text video retrieval. According to the method, the problem of randomness of the random text of single sampling is solved, so that semantic information of text coding is better expanded, and the final retrieval effect is improved.
Owner:SUN YAT SEN UNIVERSITY SHENZHEN +1

Text-video retrieval method based on multi-granularity attention

The invention discloses a text-video retrieval method based on multi-granularity attention, which comprises the following steps of: firstly, aiming at a given text-video pair, generating disturbance text characteristics through a learnable semantic preserving strategy, and constructing a text with semantic key conflicts as a difficult negative sample to carry out comparative learning so as to strengthen semantic discrimination capability; secondly, short-time actions, medium-time semantics and long-time features are extracted from the video, video content is represented in an omnibearing mode, text features serve as query signals, multi-granularity video features are dynamically weighted and fused through a multi-granularity attention mechanism, and accurate alignment of the text and the video features is achieved; and finally, calculating a similarity score of the query text and the fused video features to retrieve a video matched with the query text. According to the method, multi-granularity video features in a text-video retrieval method are considered, adaptive association of text and video semantics is established, and accurate matching is realized to improve the retrieval effect.
Owner:ZHEJIANG UNIV OF TECH

Dual-granularity alignment efficient partial correlation video retrieval based on implicit fragment modeling and semantic decomposition

The invention provides double-granularity alignment efficient partial correlation video retrieval based on implicit fragment modeling and semantic decomposition, and aims to solve the problems of information redundancy and low efficiency in an existing video modeling method, the problem of granularity mismatching between sentence representation and video frame features and the problem that alignment between a text and a video is not refined enough. According to the method, the expression and modeling of video data are optimized by introducing a structure combining a Gaussian mixture model and Transform, and multi-scale local details with short time span are adaptively integrated by introducing a window attention mechanism and a cross attention mechanism, so that finer features are obtained, and cross-modal similarity score calculation of texts and videos is facilitated. The method comprises the following steps: data preprocessing and frame segmentation: preprocessing and segmenting an input video into frames, and preparing for feature extraction after each frame of image is subjected to standardization processing; implicit modeling is carried out on the fragment-level features, uniformly sampled frame-level visual features are input into a Gaussian mixture modeling module, adjacent frame focusing modeling is carried out through a multi-scale Gaussian attention mechanism, and fragment-level representation of different receptive fields is implicitly formed. Enhancing local details of the frame-level features, performing convolution operation on each frame by adopting windows with different scales, and calculating a feature relationship in each local window; local features of all scales are fused through a cross attention mechanism, semantic expression of frame features is enhanced, and it is ensured that each video frame can capture important local details.
Owner:NANJING TECH UNIV

Vehicle image-text retrieval method and system based on multistage semantic graph alignment and attribute enhancement

The invention belongs to the technical field of intelligent traffic, and relates to a vehicle image-text retrieval method and system based on multistage semantic graph alignment and attribute enhancement. According to the method, fusion visual embedding is obtained according to the vehicle video, and fusion text embedding is obtained according to the text data; obtaining an updated visual attribute embedding vector and an updated text attribute embedding vector based on fusion visual embedding and fusion text embedding; performing semantic matching on the updated visual attribute embedding vector and the updated text attribute embedding vector, and aligning vehicle cross-modal semantic attributes from three different levels by utilizing multi-granularity semantics to obtain final visual attribute embedding and final text attribute embedding; and according to the final visual attribute embedding and the final text attribute embedding, obtaining the similarity between the vehicle video and the artificial text description, wherein the vehicle video with the highest similarity is a vehicle video retrieval result. According to the method, the target vehicle can be quickly and accurately positioned, and higher robustness and higher retrieval success rate are shown.
Owner:CHANGAN UNIV

Audio and video retrieval method and device, electronic equipment and storage medium

This invention provides an audio / video retrieval method, apparatus, electronic device, and storage medium. The method includes: obtaining current search conditions; retrieving each secondary data body based on a primary search table; each secondary data body corresponds to a secondary search table, and each secondary data body includes at least one data block; each data block includes at least one audio / video data, vehicle information corresponding to each audio / video data, and statistical information corresponding to the vehicle information; retrieving statistical information from each data block in the corresponding secondary data body based on each secondary search table; and when target statistical information that satisfies the current search conditions is retrieved, and it is determined that the target data block corresponding to the target statistical information contains target vehicle information corresponding to the current search conditions, the audio / video data corresponding to the target vehicle information is determined as the target audio / video data corresponding to the current search conditions. This invention eliminates the need to traverse all audio / video data, reducing retrieval time and improving retrieval speed.
Owner:HANGZHOU HOPECHART

End-to-end multi-task video retrieval with cross attention

A method includes obtaining a video and a relational spatio-temporal query, and identifying at least one type of the relational spatio-temporal query. The at least one type of identification of the relational spatio-temporal query represents at least one of: an activity type, an object type, or a temporal type. The method further includes learning correlations between activities, objects, and time in the video using one or more cross attention models. The method further includes obtaining one or more predictions generated using one or more outputs of the one or more cross attention models based on the identified at least one type of relational spatiotemporal query. Further, the method includes generating a response to the relational spatiotemporal query based on the one or more predictions.
Owner:SAMSUNG ELECTRONICS CO LTD

A Fake News Video Detection Method Based on Retrieval Enhancement and Prototype Alignment Techniques

This invention discloses a method for detecting fake news videos based on retrieval enhancement and prototype alignment techniques. First, the target video is processed to construct its multimodal information. Then, the multimodal information is integrated under a large-scale language model to generate a unified text-centric query, and video retrieval is performed to obtain real and fake video samples semantically related to the target video. Next, a dual prototype alignment mechanism is used to construct prototype representations for real and fake categories using a graph attention network, and a final operation-aware representation is generated through prototype alignment learning. Finally, the final operation-aware representation is integrated with an existing model to complete the detection of the target video content. This method can effectively identify subtle differences between real and fake news videos, achieving enhanced representation learning for video content detection and improving the accuracy and robustness of fake news video detection.
Owner:郑州埃文科技有限公司

Video summarization method based on multi-dimensional features and fine-grained hierarchical modeling

The application provides a video summarization method based on multi-dimensional features and fine-grained hierarchical modeling, and relates to the technical field of video processing. In practical application, the video summarization technology can facilitate large-scale video retrieval and browsing. The method comprises the following steps: firstly, frame extraction is performed on an input video to obtain a frame sequence, and a multi-dimensional feature extraction network composed of a 2D network and a 3D network is used to extract multi-dimensional features; then, hierarchical temporal modeling is performed to complete the modeling process of the temporal dependence of the entire video sequence; finally, a regression network is used to obtain the importance score of each frame and generate a video summary. The application further explores the influence of the spatiotemporal features extracted by 3D feature extractors with different spatiotemporal complexities on the video summary result. The application shows excellent performance on the video summary datasets SumMe and TVSum. Whether from the application scene or the performance index, the application has strong practical value.
Owner:SHANDONG UNIV

Unsupervised video clip retrieval method based on time sequence anchor point mining and semantic alignment

The invention discloses an unsupervised video clip retrieval method based on time sequence anchor point mining and semantic alignment, which comprises a video retrieval training system based on time sequence anchor point mining and point supervised learning, and the system comprises a key anchor point extraction module, a semantic alignment description generation module and a point supervised learning enhancement module. Key time anchor points are extracted from an original unlabeled video sequence through a key anchor point extraction module, then a pseudo-label triple is constructed through a semantic alignment description generation module, a weak supervision video clip retrieval model is trained accordingly, and in the training process, the pseudo-label triple is extracted from the original unlabeled video sequence. Constructing a point supervised contrast learning target through a point supervised learning enhancement module to further optimize the model, so that the finally trained weak supervised video clip retrieval model outputs a clip starting and ending time boundary related to query statement semantics; the problems of low pseudo tag quality, fuzzy positioning boundary and inconsistent semantic space in an unsupervised scene are effectively solved, and the positioning precision and generalization ability of video clip retrieval are remarkably improved.
Owner:SOUTH CHINA UNIV OF TECH

Intelligent monitoring video retrieval method and device based on RAG enhanced retrieval

The invention provides an intelligent monitoring video retrieval method and device based on RAG enhanced retrieval. According to the intelligent monitoring video retrieval method and device based on the RAG enhanced retrieval, the external knowledge base can be dynamically integrated through an RAG mechanism, the natural language complex query accurate response to the monitoring video is achieved, minute-level knowledge updating is supported, and the retrieval efficiency and the system adaptability are remarkably improved.
Owner:CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD

Audio and video positioning method and device, equipment and storage medium

The embodiment of the invention relates to the field of artificial intelligence, medical health and financial science and technology, and discloses an audio and video positioning method and device, equipment and a storage medium. The method comprises the following steps: collecting conference records of a service conference according to service information to obtain audios and videos corresponding to the service conference; performing voice recognition processing on the audio and video corresponding to the service conference to obtain text information corresponding to the audio and video; establishing an association relationship between the text information and the audio and video moments according to the service information; performing link mapping processing on the text information based on the association relationship to generate audio and video retrieval text information; determining a target audio / video and positioning key information corresponding to the target audio / video according to the audio / video positioning requirement; and traversing the audio and video retrieval text information, determining a target audio and video moment corresponding to the positioning key information, and skipping and positioning the target audio and video to the target audio and video moment for playing. The objective of the invention is to solve the problems of low accuracy and low efficiency of audio and video positioning in the prior art.
Owner:PING AN HEALTH CLOUD CO LTD

A long video text period retrieval method for power maintenance scene

The application discloses a long video text period retrieval method for a power maintenance scene, first divides the power maintenance long video into multiple candidate video clips, extracts a frame-level video feature sequence and an alignment target feature sequence of each candidate video clip, fuses the frame-level video feature sequence and the alignment target feature sequence, and obtains a fused frame-level video feature sequence; inputs a text feature sequence after encoding of a query text and the fused frame-level video feature sequence of each candidate video clip for cross-modal interaction calculation, obtains a clip-level context feature and a frame-level content feature sequence, and calculates a clip-level correlation score and a retrieval enhancement feature sequence based on the clip-level context feature and the frame-level content feature sequence; and finally, sorts and optimizes the clip-level correlation scores of all the candidate video clips, obtains a target retrieval enhancement feature sequence, and inputs the target retrieval enhancement feature sequence into a time boundary predictor to predict a video retrieval result corresponding to the input query text. The application enhances the expression capability of fine-grained content of the power maintenance scene and improves the retrieval efficiency in a multi-query scene.
Owner:STATE GRID ANHUI ULTRA HIGH VOLTAGE CO +1

Video retrieval method based on three-dimensional convolutional neural network of fusion feature gate

A video retrieval method based on a three-dimensional convolutional neural network of a fusion feature gate, which is composed of the steps of dataset preprocessing, dataset division, three-dimensional convolutional neural network construction, three-dimensional convolutional neural network training and three-dimensional convolutional neural network testing. The three-dimensional convolutional neural network is optimized and improved, the technical problem of low retrieval accuracy in the prior art is solved, and a different conception scheme is provided for solving similar problems. Three feature gates are adopted, the feature gate is composed of a reset gate and an update gate, the technical problem of video information redundancy in the prior art is solved, the video feature information can be more accurately extracted, and the foundation for further retrieval is laid. The present application has the advantages of high retrieval accuracy, fast retrieval speed and good retrieval effect, and can be used for video image retrieval.
Owner:XIAN UNIV OF POSTS & TELECOMM

Video retrieval method and device, electronic equipment and storage medium

The invention discloses a video retrieval method and device, electronic equipment and a storage medium, and belongs to the technical field of retrieval. The method comprises the steps of obtaining at least one type of query statement feature based on a target query statement; for each candidate video in the video library, obtaining a dynamic visual feature and at least one type of static visual feature of the candidate video; the number of the dynamic visual features is at least one; on the basis of at least one type of query statement feature, the dynamic visual feature and at least one type of static visual feature, obtaining a comprehensive similarity score of a candidate video; and based on the comprehensive similarity score of each candidate video, determining a retrieval result of the target query statement. According to the video retrieval method disclosed by the invention, the problem that the accuracy of a video retrieval result is not high is solved.
Owner:GRG BANKING EQUIPMENT CO LTD

Video retrieval method and apparatus, electronic device, and storage medium

The application discloses a video retrieval method and device, electronic equipment and storage medium, and belongs to the technical field of retrieval. The method comprises the following steps: acquiring at least one type of query sentence feature based on a target query sentence; acquiring at least one type of visual feature, a first subtitle feature and a second subtitle feature of each candidate video in a video library; the first subtitle feature is an overall subtitle feature of the candidate video, the second subtitle feature is a subtitle feature of each frame in the candidate video, and the number of the second subtitle feature is at least 1; acquiring a comprehensive similarity score of the candidate video based on the query sentence feature, the visual feature, the first subtitle feature and the second subtitle feature; and determining a retrieval result of the target query sentence based on the comprehensive similarity score of each candidate video. The video retrieval method disclosed by the application solves the problem that the accuracy of the video retrieval result is not high.
Owner:GRG BANKING EQUIPMENT CO LTD

Combined video retrieval method and system based on sharing and difference semantic enhancement

The invention relates to a combined video retrieval method and system based on sharing and difference semantic enhancement. The method comprises the following steps: inputting a reference video, a target video and a modified text into a REFINE model to realize combined video retrieval; the method specifically comprises the following steps: carrying out shared semantic enhancement, inter-frame difference decoupling and associated object aggregation on an REFINE model, and carrying out multi-modal query combination on an input reference video and a modified text to obtain a combination feature; and calculating cosine similarity of the combined features and candidate video marks in WebVid-CoVR, taking the cosine similarity as similarity scores, carrying out descending sort on the similarity scores, and selecting target videos with the similarity scores ranking the top K as required to complete combined video retrieval. According to the method, the target video of the user is effectively retrieved.
Owner:SHANDONG UNIV +1

VideoRAG using Natural Language as Intermediate Representation in Multi-Camera, Closed-Domain Applications

A Video Retrieval Augmented Generation (VideoRAG) system for closed-domain applications that uses natural language text as an intermediate representation between video content and query systems. A unified vision-language model (VLM) processes video frames and generates structured JSON text descriptions conforming to domain-specific event schemas, while simultaneously answering natural language queries through retrieval-augmented generation. The natural language intermediate representation provides substantial storage efficiency improvements over embedding-based approaches, human-interpretable analytics capabilities, and cross-camera entity tracking. The architecture supports a closed-domain applications including but not limited to retail analytics, healthcare monitoring, and industrial safety operations.
Owner:SHOBDO LLC

Alarm short video retrieval method and system

The application provides a retrieval method and system for alarm short videos, and relates to alarm short video retrieval technology. The method comprises the following steps: receiving a retrieval signal, extracting a retrieval alarm time in the retrieval signal; judging whether a retrieval alarm start time in the retrieval alarm time is equal to a retrieval alarm end time; if yes, performing jitter processing on the retrieval alarm time to obtain a retrieval alarm jitter time interval, and obtaining an alarm short video associated with the retrieval alarm jitter time interval; if no, obtaining an alarm short video associated with the retrieval alarm time; calculating an association weight of the alarm short video according to the retrieval alarm time and an alarm time of the alarm short video; sorting the alarm short video according to the association weight to obtain an alarm short video list, and sending a response signal.
Owner:UNIVERSAL UBIQUITOUS TECH CO LTD

Video recall method

The present invention relates to the field of natural language processing technology and discloses a video recall method, which aims to solve the problem of inaccurate video retrieval in existing systems. The method mainly includes: performing pronunciation preprocessing on all film title texts in a film and television database and extracting pinyin features; creating a pinyin full-value recall database and a character-by-character pinyin recall database based on the extracted pinyin features; upon receiving a user-input voice text, extracting a text to be corrected that may be a film title from the voice text and performing the same pronunciation preprocessing and pinyin feature extraction; performing full-value feature recall and character-by-character feature recall based on the pinyin full-value recall database and the character-by-character pinyin recall database; if a film title is included in the full-value feature recall result, the film title is used as the video recall result; otherwise, the video recall result is determined based on the similarity of each film title in the character-by-character feature recall result. The present invention improves the accuracy of video retrieval and is suitable for smart TVs with voice recognition.
Owner:SICHUAN CHANGHONG ELECTRIC CO LTD

Devices, systems, and methods for video retrieval

Methods and systems provided. A system may include an application program. The application program may be configured enable data captured by a remote device at a remote location to be viewable via an end-user device without transmitting the data to the end-user device. Further, after enabling the data to be viewable and in response to an input, the application program may be configured to cause a previously captured video associated with the data to be sent from the remote device via a metered connection.
Owner:LIVEVIEW TECHNOLOGIES LLC

A method for identifying complex targets in massive videos based on human-computer collaboration

The application discloses a kind of based on man-machine cooperation's mass video complex target retrieval method.Currently, the generalization ability of intelligent system based on machine vision is weak, when environment changes, target exists shielding or presents camouflage state, still need a lot of manpower intervention in mass video retrieval target, time-consuming and laborious.The application first according to the coarse-grained feature of target uses retrieval model in video library and carries out pre-screening, and the target candidate set obtained by screening is made into brain-eye cooperation RSVP paradigm and presented to subject.Subject determines target, and its specific features are input into retrieval model, so as to realize the fast positioning and tracking of target in mass video.The application combines the generalization reasoning ability of human and the rapid retrieval ability of machine, can effectively improve the efficiency and accuracy of video investigation, has strong scientific significance and social significance.
Owner:HANGZHOU DIANZI UNIV

A video retrieval method, apparatus, device, storage medium, and program product.

This application provides a video retrieval method, apparatus, device, computer-readable storage medium, and computer program product. The method includes: performing multi-task extraction processing on a video to be retrieved to obtain global features, content category, and channel category of the video to be retrieved; retrieving multiple reference videos from a first video set based on the global features and content category of the video to be retrieved to form a second video set; determining the modal features of the video to be retrieved based on the channel category of the video to be retrieved, and retrieving a target video from the multiple reference videos in the second video set based on the modal features of the video to be retrieved, so as to use the target video as the retrieval result of the video to be retrieved. This application can improve the efficiency and accuracy of video retrieval.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A distance constraint-based video spatio-temporal quantization range retrieval method and system

The present application relates to the technical field of video retrieval, and especially relates to a video space-time quantization range retrieval method and system based on distance constraint. The method comprises the following steps: acquiring video data; performing data preprocessing on the acquired video data; constructing a space-time joint index fusing distance constraint based on the preprocessed data; performing space-time range retrieval based on the constructed space-time joint index fusing distance constraint; and performing object consistency judgment and result output based on the retrieval result. The present application quantitatively represents the fine-grained space-time relationship of video objects in the same video or multiple videos, converts the space-time sensitive video data query semantics into accurate quantitative range queries in the time and space dimensions of the video objects, i.e. the spatial range queries among the same intra-frame video objects and the continuous time length range queries on the continuous video frames.
Owner:YANTAI UNIV