Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

239 results about "Video retrieval" patented technology

Intelligent database video retrieval method based on big data technology

The invention relates to the technical field of intelligent information retrieval, in particular to an intelligent database video retrieval method based on a big data technology, which comprises the following steps: S1, carrying out space-time slicing processing on an original video stream; s2, extracting a visual feature vector, an audio waveform vector and a text description vector; s3, constructing a cross-modal incidence matrix; s4, constructing a layered mixed index structure which comprises a real-time updating layer and a static storage layer; s5, after a user retrieval request is received, candidate video set screening is carried out based on the layered mixed index structure, and a retrieval result is generated; and S6, updating the cross-modal incidence matrix, and synchronously updating the weight parameter of the layer. According to the method, the retrieval accuracy is improved through cross-modal feature fusion, the data storage and retrieval efficiency is optimized by adopting a hierarchical mixed index structure, and the index weight is dynamically adjusted in combination with user behavior feedback, so that efficient, accurate and intelligent video retrieval is realized.
Owner:HEBEI ZHENGTONG ARCHIVES MANAGEMENT CO LTD

Universal scene retrieval analysis method and system based on multi-modal feature fusion

The invention discloses a universal scene retrieval analysis method and system based on multi-modal feature fusion, the method comprises a video analysis step and an application service step, and the application service step comprises the steps of receiving a user input request, describing a multi-dimensional standardized video tag based on a video summary, and obtaining a multi-dimensional standardized video tag; the steps of cross-modal video retrieval, dynamic knowledge enhancement question answering and interactive enhancement analysis can be executed, efficient video preprocessing is achieved by constructing an offline feature library, the retrieval precision is improved by adopting a cross-modal feature fusion technology, the analysis authority is enhanced in combination with a dynamic knowledge base, and the interactive enhancement analysis is supported to achieve abnormal early warning. The method has the advantages that the offline video processing efficiency is improved, cross-modal feature fusion retrieval is realized, and the authority of an analysis result is enhanced.
Owner:SHENZHEN KAOLA YOURAN TECHNOLOGY CO LTD

Context-aware video retrieval and inference system

Various examples, systems, and methods are disclosed relating to an agentic curation pipeline. One system can process questions and other inquiries about video content by using a combination of models and stored information. The system can receive a query related to an event in a video, selects relevant portions of the video using embeddings, and apply the selected video data and a related sub-query to a video model. The output from the video model can be used by a language model, along with stored context, to generate an answer to the original query. The system can returns the answer to the requester.
Owner:NVIDIA CORP

Video text cross-modal retrieval method based on spatio-temporal feature fusion

The invention relates to the field of artificial intelligence cross-modal retrieval, and provides a video text cross-modal retrieval method and system based on spatio-temporal feature fusion. The method comprises the following steps: carrying out key frame sampling and time sequence partitioning on an input video, extracting static visual features through a spatial feature network, and extracting motion features through a time dynamic network; a self-adaptive gating fusion module is adopted to dynamically calculate spatial-temporal feature weights and perform weighted fusion; extracting text semantic features by using a pre-training language model; constructing a double-flow projection network to map video fusion features and text features to a unified measurement space, and optimizing a feature distance by adopting a contrast loss function containing difficult negative sample mining and intra-modal constraint; and outputting a retrieval result according to the cosine similarity sequence. The system comprises four units, wherein the gating fusion module is integrated with an FPGA acceleration circuit. According to the method, mAP (at) 10 is equal to 0.78 in a UCF-101 data set, the time sequence action retrieval accuracy rate is 92.8%, and the single video retrieval delay is 23 milliseconds.
Owner:ZHEJIANG UNIV

Multi-granularity video retrieval method and device based on multi-modal large model, computer equipment and readable storage medium

The invention discloses a multi-granularity video retrieval method and device based on a multi-modal large model, computer equipment and a readable storage medium, and the method comprises the steps: firstly obtaining video query information input by a user, carrying out the intention recognition to obtain a query field, rewriting the query information and the field to obtain a video query vector, and carrying out the retrieval of the video query vector; and retrieving in a preset retrieval video knowledge base according to the vector and the field, and finally obtaining the target retrieval video content, so as to improve the efficiency and accuracy of video retrieval and adapt to multi-field retrieval requirements.
Owner:DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD

Video retrieval generation method and device based on sparse representation and reordering

The invention discloses a video retrieval generation method and device based on sparse representation and reordering, and relates to the technical field of cross-modal video retrieval. The method comprises the following steps: acquiring a query text, a candidate video sequence and a video retrieval generation model; obtaining a query text dense representation and a candidate video sequence dense representation according to the encoder module, the query text and the candidate video sequence; through a sparse representation module, text sparse representation is generated according to the query text dense representation, candidate video sparse representation is generated according to the candidate video sequence dense representation, and reverse index screening is performed on the candidate video sequence to obtain part of candidate videos; determining a global score of each video in the partial candidate videos through a cross attention module, and sorting the partial candidate videos; and through a generation module, generating a target text corresponding to the query text according to the sorted candidate videos and the query text. By adopting the method and the device, the retrieval efficiency is improved in a sparse representation and reordering mode.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Video retrieval method and device, electronic equipment and storage medium

The invention discloses a video retrieval method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring each current to-be-retrieved video; dividing each video to be retrieved into a plurality of split videos; for each target retrieval dimension, obtaining a text description of the target retrieval dimension of each split video; wherein the target retrieval dimension comprises a subtitle dimension, an audio dimension and a visual dimension; the text description of the visual dimension of one split video is combined by the context text description of the split video and the text description of each key frame of the split video; combining the text description of the target retrieval dimension of each split video with a retrieval problem, and performing retrieval enhancement to obtain a plurality of retrieval videos based on the target retrieval dimension; weighting the text description of each target retrieval dimension of each retrieval video to obtain a plurality of sub-mirror multi-dimensional text descriptions; and combining the multi-dimensional text description of each sub-mirror with a retrieval problem to carry out retrieval enhancement to obtain a final retrieval result.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Video ROI retrieval method and system based on multi-modal neural network

The invention discloses a video ROI retrieval method and system based on a multi-mode neural network, and the method comprises the steps: carrying out the key frame extraction of an input video, and obtaining a key frame; performing visual feature extraction, audio feature extraction and text feature extraction on the key frame to obtain a multi-modal feature of the key frame; inputting the multi-modal features of the key frames and the query text into a multi-modal neural network, and outputting timestamps of candidate video frames containing the region of interest and diagonal coordinates of candidate frames; arranging the timestamps of the candidate video frames according to a time sequence, and determining at least one interception interval of the input video according to a difference value of the timestamps of the adjacent candidate video frames; and intercepting the input video by using the interception interval, generating an ROI frame according to the diagonal coordinates of the candidate frame, and covering the ROI frame on the corresponding video clip to obtain a video ROI retrieval result. The invention relates to the technical field of artificial intelligence, and solves the technical problem of insufficient retrieval accuracy of a video region of interest ROI in the prior art.
Owner:ANHUI RUIJI INTELLIGENT TECH CO LTD

Image video retrieval method based on domain fine-tuning large language model

The invention provides an image video retrieval method based on a domain fine-tuning large language model, which comprises the following steps: performing fine-tuning on a pre-training model to obtain a fine-tuning pre-training model for intention classification and keyword extraction; performing dynamic iteration screening on an optimal prompt template through Monte Carlo tree search in combination with a hidden Markov model (HMM); performing noise filtering on the keyword list, and predicting category labels of the filtered keywords through a conditional random field model to obtain a keyword enhancement set; combining with the user intention to generate a query condition, and obtaining a candidate resource set; and according to the similarity between the user query text and the candidate resource set, and in combination with the optimal prompt template, obtaining the resource path with the highest matching score between the user query and the candidate resource, and obtaining the retrieved image or video, so that the identification deviation possibly occurring when a general model processes proper nouns and terminologies can be effectively solved, and the user experience is improved. And the retrieval accuracy and response speed are improved, so that the retrieval accuracy and professional adaptability are improved.
Owner:HUBEI ZHONGKE NETWORK ENG

Video retrieval method and server based on semantic embedding and video memory coding

The invention provides a video retrieval method and server based on semantic embedding and video memory coding, and the method comprises the steps: executing visual content analysis on continuous frames of an input video through a pre-trained visual language large model, and mapping the visual content of each frame into a semantic embedding vector; a video memory coding module is constructed based on the time sequence features of the semantic embedding vectors, and the continuous frame semantic embedding vectors are dynamically coded according to the time sequence to generate a video memory feature sequence containing time sequence association information; performing semantic analysis on the query text based on the visual language large model, and converting the query text into a query semantic embedding vector which is in the same spatial dimension as the semantic embedding vector; taking the query semantic embedding vector as a retrieval basis, performing similarity matching in a video memory feature sequence, and determining an associated video clip candidate set through feature association analysis in a time sequence window; and performing time sequence coherence verification on the candidate set, and outputting a matched video clip in combination with the semantic embedding vector time sequence correlation degree.
Owner:HUAYUNSHENGDA(BEIJING)METEROLOGICAL TECH CO LTD

Video space-time retrieval method and device based on grid coding

The invention discloses a grid-coding-based video space-time retrieval method and device, the grid-coding-based video space-time retrieval method is executed in computing equipment, and the method comprises the following steps: mapping a coordinate range of a target area into a corresponding grid code set according to a Beidou grid standard on the basis of the coordinate range of the target area; acquiring video data with an intersection or inclusion relationship between the spatial range of the shot object and the target area; combining the shooting time of any video data with each grid code to generate a space-time grid code set; constructing an index structure corresponding to the video data by taking each time-space grid code in the time-space grid code set as a main key; and in response to query information of the user, analyzing a space-time range in the query of the user, mapping the space-time range into a corresponding space-time grid code range, executing retrieval based on the range, and returning hit video data. According to the method, rapid positioning and accurate retrieval of large-scale aerial photography video data are realized.
Owner:BEIJING ZHIWANG YILIAN TECH CO LTD

Text-video retrieval method and system based on multi-granularity learnable interaction

The invention discloses a text-video retrieval method and system based on multi-granularity learnable interaction, and relates to the field of data retrieval, and the method comprises the following steps: encoding a query text and an unpaired video in to-be-retrieved data, and extracting query text features and video features; obtaining a plurality of learnable vectors, carrying out cross-modal interaction on the learnable vectors and fine-grained video features, carrying out cross-grained alignment on query text features and learnable variables, carrying out fine-grained alignment on each word feature and the learnable vector, and aligning the query text features and video features; and capturing fine-granularity alignment information and coarse-granularity alignment information, capturing weights of different similar vectors among modals, summing multiple similarity scores based on the weights to obtain final multi-granularity similarity scores, and sorting the final multi-granularity similarity scores to obtain a final retrieval result. According to the method, cross-modal interaction and other operations are carried out through the plurality of learnable vectors and the fine-grained visual features, the fine-grained visual information is utilized, and the retrieval accuracy is improved.
Owner:NINGXIA UNIVERSITY

Multi-mode combined video retrieval method and device

The embodiment of the invention provides a multi-mode combined video retrieval method and device. The method comprises the following steps: acquiring text information and visual information; extracting character features from the character information; extracting visual features from the visual information; extracting visual semantic features from the visual information according to the character features; extracting common features and difference features between the character features and the visual semantic features from the character features and the visual semantic features; querying a preset video information base according to the visual features and the common features to obtain a plurality of video retrieval results matched with the visual features and the common features; and screening the plurality of video retrieval results according to the difference characteristics to obtain a screened video retrieval result. According to the method, effective information of multi-modal information can be fused, the real intention of the user can be accurately understood, and the accuracy of multi-modal combined video retrieval is improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Intelligent video retrieval method and system

The invention relates to an intelligent video retrieval method and system. The method comprises the following steps: receiving a video retrieval keyword input by a user; historical video retrieval data of the user is called, and video retrieval behavior analysis is carried out on the user; according to the video retrieval behavior analysis result of the user, performing semantic analysis on the video retrieval keyword input by the user to obtain a natural language query statement; mapping the natural language query statement to a unified semantic space, and in the unified semantic space, calculating semantic similarity between the natural language query statement and the multi-modal fusion feature of each video to obtain a video similarity sorting result; historical video retrieval data of the user is called, the video similarity sorting result is adjusted according to historical feedback data of the user and the natural language query statement, and a video retrieval result is obtained; and outputting a video retrieval result.
Owner:SHANGHAI LINKE ZHIHUA DIGITAL TECHNOLOGY CO LTD

Information processing system and methods for clinical video retrieval

The present disclosure generally relates to an integrated approach for retrieving biomedical information from clinical video presentations. In particular, the present disclosure is directed to video retrieval systems and methods of text-video retrieval from clinical video presentations.
Owner:THE CURATORS OF THE UNIVERSITY OF MISSOURI

Surveillance video retrieval method, device, equipment and program product

The invention relates to the field of monitoring, in particular to a monitoring video retrieval method and device, equipment and a program product. The method comprises the following steps: receiving a surveillance video retrieval request, wherein the surveillance video retrieval request comprises a first time range and a type range; performing preliminary retrieval according to the first time range, and determining an index ID range of a first monitoring video event index set associated with the first time range; and performing second retrieval in the monitoring video event corresponding to the index ID range, determining a second monitoring video event index set associated with the type range in batches, and sequentially returning the second monitoring video event index set to the retrieval terminal according to the batches. The index ID range obtained through preliminary retrieval only occupies a small cache, and the second monitoring video event index set generated in batches is written into the cache, so that the monitoring video event and video loss illusion caused by insufficient cache can be effectively avoided on the premise of not increasing the cache, and the user experience is improved.
Owner:TP-LINK

False news video detection method based on retrieval enhancement and prototype alignment technology

The invention discloses a false news video detection method based on a retrieval enhancement and prototype alignment technology, which comprises the following steps of: processing a target video, constructing multi-modal information of the target video, and integrating the multi-modal information under a large language model to generate a unified text center query; performing video retrieval to obtain real and false video samples related to target video semantics; then respectively constructing prototype representations of real and false categories by using a graph attention network through a double prototype alignment mechanism, and generating a final operation perception representation through prototype alignment learning; and integrating the final operation perception representation with the existing model to complete the detection of the target video content. The subtle difference between the real news video and the false news video can be effectively identified, the detection of the video content is enhanced, the representation learning is realized, and the accuracy and robustness of the detection of the false news video are improved.
Owner:郑州埃文科技有限公司

Intelligent event video retrieval data storage and display method

The invention relates to an intelligent event video retrieval data storage and display method, which comprises the following steps: acquiring video data of each video source, and inserting a track area number of a track area where a target appears in a video frame in which the target is identified to exist; according to the target identifier of the target, the time information and the track area number, track data are obtained and stored; based on the track area number and the retrieval condition of the time information, retrieving from the track data to obtain a target identifier of the corresponding target; corresponding target video data are obtained through integration based on the target identifier; and playing the target video data corresponding to the target identifier in a display window, and displaying the corresponding target attribute. By means of the video retrieval method and device, the track area number of the target can be inserted into the video frame with the target, the track data are formed and stored, the target video data of the needed target can be rapidly obtained and played based on the track area number and the time information during retrieval, and the problem that the video retrieval efficiency is low is solved.
Owner:ZHEJIANG DAHUA TECH CO LTD +1

Text video retrieval method of fine-grained relation learning network based on energy perception

The invention provides a text video retrieval method of a fine-grained relation learning network based on energy perception. The method comprises the following steps: giving a query text and a video clip; inputting the query text into a text encoder of the CLIP, inputting the video clip into a text encoder image encoder of the CLIP, and extracting to obtain text embedding and frame embedding; inputting the text embedding and the frame embedding into a fine-grained relation learning network for text enhancement operation to obtain enhanced text embedding; taking the enhanced text embedding as a frame fusion condition to carry out frame fusion operation on the frame embedding to obtain video embedding; calculating the similarity of a text-video pair formed by the query text and the video embedding based on a cosine similarity function; and selecting the text-video pair with the highest similarity as a retrieval output result of text video retrieval. According to the method, the problem of randomness of the random text of single sampling is solved, so that semantic information of text coding is better expanded, and the final retrieval effect is improved.
Owner:SUN YAT SEN UNIVERSITY SHENZHEN +1

Text-video retrieval method based on multi-granularity attention

The invention discloses a text-video retrieval method based on multi-granularity attention, which comprises the following steps of: firstly, aiming at a given text-video pair, generating disturbance text characteristics through a learnable semantic preserving strategy, and constructing a text with semantic key conflicts as a difficult negative sample to carry out comparative learning so as to strengthen semantic discrimination capability; secondly, short-time actions, medium-time semantics and long-time features are extracted from the video, video content is represented in an omnibearing mode, text features serve as query signals, multi-granularity video features are dynamically weighted and fused through a multi-granularity attention mechanism, and accurate alignment of the text and the video features is achieved; and finally, calculating a similarity score of the query text and the fused video features to retrieve a video matched with the query text. According to the method, multi-granularity video features in a text-video retrieval method are considered, adaptive association of text and video semantics is established, and accurate matching is realized to improve the retrieval effect.
Owner:ZHEJIANG UNIV OF TECH

Dual-granularity alignment efficient partial correlation video retrieval based on implicit fragment modeling and semantic decomposition

The invention provides double-granularity alignment efficient partial correlation video retrieval based on implicit fragment modeling and semantic decomposition, and aims to solve the problems of information redundancy and low efficiency in an existing video modeling method, the problem of granularity mismatching between sentence representation and video frame features and the problem that alignment between a text and a video is not refined enough. According to the method, the expression and modeling of video data are optimized by introducing a structure combining a Gaussian mixture model and Transform, and multi-scale local details with short time span are adaptively integrated by introducing a window attention mechanism and a cross attention mechanism, so that finer features are obtained, and cross-modal similarity score calculation of texts and videos is facilitated. The method comprises the following steps: data preprocessing and frame segmentation: preprocessing and segmenting an input video into frames, and preparing for feature extraction after each frame of image is subjected to standardization processing; implicit modeling is carried out on the fragment-level features, uniformly sampled frame-level visual features are input into a Gaussian mixture modeling module, adjacent frame focusing modeling is carried out through a multi-scale Gaussian attention mechanism, and fragment-level representation of different receptive fields is implicitly formed. Enhancing local details of the frame-level features, performing convolution operation on each frame by adopting windows with different scales, and calculating a feature relationship in each local window; local features of all scales are fused through a cross attention mechanism, semantic expression of frame features is enhanced, and it is ensured that each video frame can capture important local details.
Owner:NANJING TECH UNIV

Vehicle image-text retrieval method and system based on multistage semantic graph alignment and attribute enhancement

The invention belongs to the technical field of intelligent traffic, and relates to a vehicle image-text retrieval method and system based on multistage semantic graph alignment and attribute enhancement. According to the method, fusion visual embedding is obtained according to the vehicle video, and fusion text embedding is obtained according to the text data; obtaining an updated visual attribute embedding vector and an updated text attribute embedding vector based on fusion visual embedding and fusion text embedding; performing semantic matching on the updated visual attribute embedding vector and the updated text attribute embedding vector, and aligning vehicle cross-modal semantic attributes from three different levels by utilizing multi-granularity semantics to obtain final visual attribute embedding and final text attribute embedding; and according to the final visual attribute embedding and the final text attribute embedding, obtaining the similarity between the vehicle video and the artificial text description, wherein the vehicle video with the highest similarity is a vehicle video retrieval result. According to the method, the target vehicle can be quickly and accurately positioned, and higher robustness and higher retrieval success rate are shown.
Owner:CHANGAN UNIV

A video surveillance storage system and method based on distributed cloud storage

This application relates to the field of cloud storage technology. Specifically, it discloses a video surveillance storage system and method based on distributed cloud storage. It collects surveillance video streams through cameras, performs time segmentation on the surveillance video streams based on the video retrieval query status of users to form multiple video data blocks, and then constructs a video distributed cloud storage architecture. It uses the consistent hashing algorithm to determine the storage locations of each video data block in the distributed cloud storage, and monitors the usage frequency of each video data block in real time to dynamically adjust the storage priorities of each video data block. This application can not only effectively avoid the risk of single point of failure, improve the reliability and fault tolerance of storage, but also dynamically adjust the storage priorities of video data according to their usage frequencies, realizing efficient management and access of data.
Owner:浙江幸福轨道交通运营管理有限公司

Audio and video retrieval method and device, electronic equipment and storage medium

This invention provides an audio / video retrieval method, apparatus, electronic device, and storage medium. The method includes: obtaining current search conditions; retrieving each secondary data body based on a primary search table; each secondary data body corresponds to a secondary search table, and each secondary data body includes at least one data block; each data block includes at least one audio / video data, vehicle information corresponding to each audio / video data, and statistical information corresponding to the vehicle information; retrieving statistical information from each data block in the corresponding secondary data body based on each secondary search table; and when target statistical information that satisfies the current search conditions is retrieved, and it is determined that the target data block corresponding to the target statistical information contains target vehicle information corresponding to the current search conditions, the audio / video data corresponding to the target vehicle information is determined as the target audio / video data corresponding to the current search conditions. This invention eliminates the need to traverse all audio / video data, reducing retrieval time and improving retrieval speed.
Owner:HANGZHOU HOPECHART

Video retrieval method and device

The embodiment of the invention provides a video retrieval method and device, computer equipment, a computer readable storage medium and a computer program product, and relates to the technical field of video processing. The video retrieval method comprises the following steps: displaying a video editing interface, wherein a search entry is configured on the video editing interface; receiving target sub-mirror description information through the search entry; according to the sub-mirror description information, retrieving a corresponding target video slice from a preset local material retrieval table; wherein the local material retrieval table comprises video retrieval information of each video slice output based on the large language model. According to the technical scheme provided by the embodiment of the invention, the required video slice can be accurately retrieved from the local video material.
Owner:SHANGHAI BILIBILI TECH CO LTD

Video retrieval feature extraction and retrieval positioning method, electronic equipment, storage medium and program product

The invention provides a video retrieval feature extraction and retrieval positioning method, electronic equipment, a storage medium and a program product. In particular to a two-stage zero video retrieval and fragment positioning method based on dense video text description, which comprises the following steps: in a construction stage, constructing dense video text description at a video fragment level; in a retrieval stage, related videos and fragments are positioned by using text similarity of multiple granularities of sentences and keywords. On one hand, labeling of a fragment level or a video level is not needed, and manpower and material resources are saved; and on the other hand, the interpretability is high.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

End-to-end multi-task video retrieval with cross attention

A method includes obtaining a video and a relational spatio-temporal query, and identifying at least one type of the relational spatio-temporal query. The at least one type of identification of the relational spatio-temporal query represents at least one of: an activity type, an object type, or a temporal type. The method further includes learning correlations between activities, objects, and time in the video using one or more cross attention models. The method further includes obtaining one or more predictions generated using one or more outputs of the one or more cross attention models based on the identified at least one type of relational spatiotemporal query. Further, the method includes generating a response to the relational spatiotemporal query based on the one or more predictions.
Owner:SAMSUNG ELECTRONICS CO LTD

A Fake News Video Detection Method Based on Retrieval Enhancement and Prototype Alignment Techniques

This invention discloses a method for detecting fake news videos based on retrieval enhancement and prototype alignment techniques. First, the target video is processed to construct its multimodal information. Then, the multimodal information is integrated under a large-scale language model to generate a unified text-centric query, and video retrieval is performed to obtain real and fake video samples semantically related to the target video. Next, a dual prototype alignment mechanism is used to construct prototype representations for real and fake categories using a graph attention network, and a final operation-aware representation is generated through prototype alignment learning. Finally, the final operation-aware representation is integrated with an existing model to complete the detection of the target video content. This method can effectively identify subtle differences between real and fake news videos, achieving enhanced representation learning for video content detection and improving the accuracy and robustness of fake news video detection.
Owner:郑州埃文科技有限公司

Video summarization method based on multi-dimensional features and fine-grained hierarchical modeling

The application provides a video summarization method based on multi-dimensional features and fine-grained hierarchical modeling, and relates to the technical field of video processing. In practical application, the video summarization technology can facilitate large-scale video retrieval and browsing. The method comprises the following steps: firstly, frame extraction is performed on an input video to obtain a frame sequence, and a multi-dimensional feature extraction network composed of a 2D network and a 3D network is used to extract multi-dimensional features; then, hierarchical temporal modeling is performed to complete the modeling process of the temporal dependence of the entire video sequence; finally, a regression network is used to obtain the importance score of each frame and generate a video summary. The application further explores the influence of the spatiotemporal features extracted by 3D feature extractors with different spatiotemporal complexities on the video summary result. The application shows excellent performance on the video summary datasets SumMe and TVSum. Whether from the application scene or the performance index, the application has strong practical value.
Owner:SHANDONG UNIV

Video retrieval method based on deep neural network model and multiple example learning

The application relates to the field of computer vision processing, in particular to a video retrieval method based on a deep neural network model and multi-example learning, which comprises the following steps: obtaining initial features by pre-training a query text, extracting I 3D-RGB features, ROI features and connection features from a video; updating frame-level visual features and word-level text features; constructing a graph for training, learning word-level text features by using a graph attention network; calculating the residual error of the word-level text features and the word-level text features, and taking the mean value of the residual error as a sentence-level text feature; performing segment dimension average operation on the frame-level visual features to obtain pipeline-level visual features; calculating the alignment score of the sentence-level text features and the pipeline-level visual features, constructing positive sample pairs and negative sample pairs, and training a video retrieval network; and the application constructs a graph neural network by acquiring discriminative features in multiple query texts through deep learning features, so as to provide text features with more representation meanings and multi-modal alignment supervision signals under weak supervision settings.
Owner:BEIJING INST OF TECH +2