Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

106 results about "Video Library" patented technology

Video Library was a publicly traded video rental shop based in San Diego, California. It had 43 corporate stores from 1979 through 1989 before they were acquired and converted into Blockbuster Video in 1989.

Interactive teaching method and system based on teaching video

The embodiment of the invention relates to the technical field of information, in particular to an interactive teaching method and system based on teaching videos. The method comprises the following steps: carrying out content segmentation on a historical teaching video, identifying explanation fragments of knowledge points in the historical teaching video, carrying out semantic annotation and label annotation, and constructing a structured knowledge point explanation video library; constructing a teaching agent, wherein the teaching agent integrates a natural language processing module, a learning state recognition module and a video fusion module; identifying a current learning demand and a target knowledge point of the student, and automatically retrieving a matched explanation fragment according to the target knowledge point; the teaching agent analyzes the content of the retrieved explanation segment, extracts the explanation content and the teaching logic, and fuses the explanation content and the teaching logic into a knowledge base of the agent; and the teaching agent provides explanation, question answering, practice recommendation and feedback evaluation for students in a dialogue or interaction mode based on the fused knowledge base, so that interactive teaching is realized.
Owner:HANGZHOU EXPOLANG XINZHI EDUCATION TECHNOLOGY CO LTD

Method and apparatus for recommending short video, electronic device, and storage medium

The present application belongs to the technical field of video recommendation, and particularly relates to a method and apparatus for recommending a short video, an electronic device, and a storage medium. The method comprises: step 1, obtaining historical viewing data of a user, determining a first short video identifier information list and a second short video identifier information list, and generating a first feature vector and a second feature vector; step 2, obtaining a short video identifier information list to be expanded, and, on the basis of the short video identifier information list to be expanded, generating a third feature vector; step 3, on the basis of the first feature vector, the second feature vector, and the third feature vector, calculating a fourth feature vector; and step 4, obtaining a short video library to be recommended, extracting a short video identifier information list to be recommended, generating a fifth feature vector, calculating the similarity between the fifth feature vector and the fourth feature vector, and, on the basis of the similarity, filtering out a short video to be recommended. According to the present application, short video identifier information is expanded, thereby improving the diversity and freshness of short video recommendation.
Owner:BEIJING FENGPING INTELLIGENT TECHNOLOGY CO LTD

Calligraphy teaching digital system and method

The invention relates to a calligraphy teaching digital system and method, and belongs to the field of computer education software, and the system comprises a multi-modal data input module which comprises a binocular camera, a pressure induction pen and a gyroscope sensor, the binocular camera collects a writing video, the pressure induction pen and the gyroscope sensor collect pen wielding track, force and angle data, and the multi-modal data input module is used for inputting the writing video; the data is preprocessed; the intelligent analysis module is used for constructing a calligraphy evaluation model based on an attention mechanism and carrying out feature extraction and fusion on the multi-modal data input module to obtain calligraphy practice evaluation; and the resource management module comprises a video library and a copybook library. Various teaching decomposition videos and copybooks are stored; and the personalized learning module is used for analyzing the calligraphy practice evaluation result, associating the corresponding teaching decomposition video and copybook according to the analysis result, and carrying out personalized learning customization. Through the full-closed-loop teaching process of data acquisition, intelligent analysis and real-time feedback adjustment, the writing level of a writer can be quickly improved.
Owner:JILIN NORMAL UNIV

Backboard video-based interactive digital human presentation method and related device

The invention provides an interactive digital human presentation method based on a backplane video and a related device, and relates to the technical field of artificial intelligence such as human-computer interaction, digital human, end-cloud integration and large models. The method comprises the following steps: extracting a real-time inquiry from a real-time interaction request initiated by a user for a digital human offline video in a playing state, and sending the real-time inquiry and historical interaction content to a cloud server; the cloud server is controlled to generate reply voice according to the real-time inquiry and historical interaction content and select a target digital person bottom plate video matched with the reply voice from a digital person bottom plate video library, and different digital person bottom plate videos correspond to digital persons showing different actions respectively; controlling the cloud server to generate a digital person online video for answering the real-time inquiry according to the target digital person bottom plate video and the reply voice; and linking and playing the digital human online video stream-pushed by the cloud server in the digital human offline video in the playing state in a manner of inserting the waiting transition video.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Teaching-oriented multi-modal interactive digital human teaching assistant generation method

The invention discloses a teaching-oriented multi-modal interactive digital human teaching-assisted generation method, and belongs to the technical field of artificial intelligence education, and the method comprises the steps: generating a structured answer through multi-modal input (voice, text and portrait graph) in combination with a semantic enhancement question and answer model (SE-QA); generating personalized voice by using an emotion adaptive voice synthesis technology; constructing a teaching action video library, extracting action features by using a space-time diagram convolutional network (ST-GCN), generating a video through a time sequence convolutional network (TCN), and optimizing audio-lip synchronization and micro expressions; and a'generation-evaluation-optimization 'closed loop is realized through a multi-modal evaluation and reinforcement learning optimization generation process. According to the teaching-oriented multi-mode interactive digital human teaching-assistant generation method, the technical bottlenecks of semantic-action mismatch, single emotion expression and the like of a traditional digital human system are broken through, the knowledge transmission efficiency and the interaction reality sense can be remarkably improved, and an innovative solution is provided for an intelligent education tool.
Owner:CHINA UNIV OF MINING & TECH

Video pushing method and system based on vehicle and storage medium

The invention discloses a vehicle-based video pushing method and system and a storage medium, and belongs to the technical field of data processing. When the vehicle accesses the WiFi network, the server obtains an interest tag of a driver of the vehicle; the server searches for a push video matched with the interest tag in a network video library, and pushes a video clip of a predetermined duration in the push video to the vehicle through a WiFi network; the vehicle obtains vehicle state information, environment information and driver state information, and generates a safety level according to the vehicle state information, the environment information and the driver state information; the vehicle obtains visual attention information and user interaction behavior information, and an attention weight is generated according to the visual attention information and the user interaction behavior information; and the vehicle determines a playing mode according to the safety level and the attention weight, and controls playing of the video clip according to the playing mode. According to the invention, the video is pushed in advance when the WiFi is accessed, the pushing flow is saved, and the driving safety and the watching experience can also be considered.
Owner:ZERON AUTOMOBILE TECHNOLOGY CO LTD

A Video Retrieval Method Based on Frame Index and Cross-Modal Representation

The present invention relates to a video retrieval method based on frame index and cross-modal representation. The method includes the following steps: preprocessing a massive video library by using the distributed computing power of a hadoop cluster to construct a video retrieval database; performing frame segmentation on the video segment to be retrieved; using the proposed cross-modal representation method enhanced by a graph structure under multi-task optimization to map the two-modal data of video frame images and text obtained in step S1 to a unified cross-modal feature space; using the Frobenius norm similarity algorithm to find the position of the first frame that meets the similarity threshold in the video library and record the time sequence of the frame in the video where it is located; calculating the similarity of each video by using the similarity algorithm in step S3 and sorting according to the similarity of the videos, and selecting the top ten videos as the final retrieval results. The present invention solves the problem that it is difficult to quickly and stably retrieve the target video in a massive video source based on the existing title and keyword retrieval strategies in video retrieval.
Owner:SHANDONG BAIMENG INFORMATION TECH CO LTD

Video-based omnibearing remote rehabilitation training system and method, and medium

The invention discloses a video-based omni-directional remote rehabilitation training system, a video-based omni-directional remote rehabilitation training method and a medium, which are characterized in that a digital standardized training video library classified according to four stages of PT physical therapy is established, and more than 700 digital standardized training videos such as joint activity, balance training, gait correction and the like are covered; it is ensured that the training content is scientific and covers the whole period and all directions of rehabilitation training; rehabilitators generate personalized training schemes in combination with rehabilitation scene types (hospitalization / home / community / old-age care institutions) and clinical data and in combination with remote video evaluation of the rehabilitators, and training requirements in different environments are met. By capturing actions of limb joint angles, joint point motion trails and muscle force changes, a training action deviation rate is calculated and is fed back and output to a rehabilitation teacher in a grading manner, so that the training standardability is remarkably improved; rehabilitators can manage multiple patients online at the same time, remotely check training percentage data, adjust schemes and generate rehabilitation notes, and efficient, accurate and personalized rehabilitation guidance services are provided for the patients.
Owner:FALCON HEALTH TECHNOLOGY (SHANGHAI) CO LTD

Information display method, device, computer equipment and storage medium

The present disclosure provides an information display method, apparatus, computer equipment and storage medium, wherein the method comprises: generating question information in response to a question operation regarding a parsing step in a question parsing of a question to be answered; the question parsing comprises a plurality of pre-split consecutive parsing steps for answering the question to be answered; the question information comprises at least one selected target parsing step; determining whether there is a recorded explanation video related to the target parsing step in a preset video library based on a pre-established association relationship between the parsing step and the explanation video; if not, sending the question information to a teacher end to obtain a target explanation video for the question information fed back by the teacher end; displaying the target explanation video, and storing the target explanation video in association with the target parsing step in the video library.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

A method and system for evaluating the training effect of tactical decision-making based on human-computer interaction

The present invention relates to a method and system for evaluating the training effect of tactical decision-making based on human-computer interaction. The method includes: obtaining a training video set corresponding to the initial training plan in a preset training video library according to the initial training plan of the trainee; receiving a target training video selected from the training video set, playing the target training video to the trainee, and based on a human-computer interaction device, obtaining the first tactical decisions selected by the trainee in each tactical scenario of the target training video, and obtaining the tactical decision-making data of the trainee according to all the first tactical decisions; substituting the tactical decision-making data of the trainee and the initial training plan into a preset training effect evaluation model to generate an initial evaluation result of the trainee. By means of human-computer interaction, the present invention improves the trainee's observation ability of the game scenario and instant tactical decision-making ability, and also tests the training effect of the trainee's tactical decision-making.
Owner:BEIJING SPORT UNIV

A method for simulating head movement in a three-dimensional avatar articulatory process

The application provides a three-dimensional image pronunciation process head action simulation method, and belongs to the technical field of three-dimensional virtual images. The three-dimensional image pronunciation process head action simulation method obtains a human face video and corresponding audio from a video library, aligns video frames and audio frames, extracts multiple frames of human face images, head posture parameters and mel spectra as training samples; pre-processes the human face images to generate face images after erasing the mouth; establishes a three-dimensional image head model and trains the three-dimensional image head model by using the training samples. The three-dimensional image head model comprises an audio feature extraction module, a lip shape synchronization module, a mouth generation module, a head posture module and a fusion module. The trained three-dimensional image head model is used to generate a three-dimensional image head model for specific audio. The method greatly reduces the calculation amount, simultaneously enables good linkage between the head posture and pronunciation, and avoids the stiff phenomenon of the three-dimensional image pronunciation process.
Owner:JINDONG CULTURE TECHNOLOGY CO LTD

Video search method, device and computer readable storage medium

The application discloses a video search method and device and a computer readable storage medium. The method comprises the following steps: obtaining a video search text, and performing text feature extraction on the video search text to obtain first text features; mapping the first text features to a video feature space corresponding to video information to obtain second text features; performing video feature extraction on each candidate video in a candidate video library to obtain first video features of each candidate video; mapping the first video features to a text feature space corresponding to text information to obtain second video features of each candidate video; and searching for a target candidate video corresponding to the video search text in the candidate video library based on a first similarity relationship between the first text features and the second video features and a second similarity relationship between the first video features and the second text features. The method can greatly improve the accuracy of searching for videos based on text.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Video code rate control method, device and equipment

The embodiment of the invention discloses a video code rate control method, device and equipment. According to the scheme, when videos in a video library are coded, a code rate value upper limit and a code rate value lower limit for coding each target video are determined according to a target scene category to which each target video in the video library belongs, and then the target videos are coded based on the code rate value upper limit and the code rate value lower limit.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Video retrieval method and device, electronic equipment and storage medium

The invention discloses a video retrieval method and device, electronic equipment and a storage medium, and belongs to the technical field of retrieval. The method comprises the steps of obtaining at least one type of query statement feature based on a target query statement; for each candidate video in the video library, obtaining a dynamic visual feature and at least one type of static visual feature of the candidate video; the number of the dynamic visual features is at least one; on the basis of at least one type of query statement feature, the dynamic visual feature and at least one type of static visual feature, obtaining a comprehensive similarity score of a candidate video; and based on the comprehensive similarity score of each candidate video, determining a retrieval result of the target query statement. According to the video retrieval method disclosed by the invention, the problem that the accuracy of a video retrieval result is not high is solved.
Owner:GRG BANKING EQUIPMENT CO LTD

Video retrieval method and apparatus, electronic device, and storage medium

The application discloses a video retrieval method and device, electronic equipment and storage medium, and belongs to the technical field of retrieval. The method comprises the following steps: acquiring at least one type of query sentence feature based on a target query sentence; acquiring at least one type of visual feature, a first subtitle feature and a second subtitle feature of each candidate video in a video library; the first subtitle feature is an overall subtitle feature of the candidate video, the second subtitle feature is a subtitle feature of each frame in the candidate video, and the number of the second subtitle feature is at least 1; acquiring a comprehensive similarity score of the candidate video based on the query sentence feature, the visual feature, the first subtitle feature and the second subtitle feature; and determining a retrieval result of the target query sentence based on the comprehensive similarity score of each candidate video. The video retrieval method disclosed by the application solves the problem that the accuracy of the video retrieval result is not high.
Owner:GRG BANKING EQUIPMENT CO LTD

An intelligent system capable of realizing personalized makeup teaching and application thereof

The application discloses an intelligent system capable of realizing individualized makeup teaching and application thereof, and the intelligent system comprises a video library provided with a plurality of classified makeup teaching video items with semantic classification labels, a classified makeup image library composed of a plurality of video frame images from the classified makeup teaching video items in the video library and with corresponding semantic classification labels, a face image acquisition module, a face feature analysis module, an adaptive makeup analysis module, a virtual makeup module and a video teaching module. The application can not only realize customization of individualized makeup, but also ensure that the customized individualized makeup has accurate corresponding teaching videos, and the customized individualized makeup can be effectively implemented, and the customized makeup effect is highly consistent with the learned makeup effect, so that the problem of difference between the actual makeup and the recommended makeup can be avoided, and people can easily and quickly obtain individualized makeup guidance suitable for themselves, and the application has a significant application prospect.
Owner:FUDAN UNIVERSITY

Large-model-driven method for real-time interaction with digital human

The present invention relates to the technical field of digital humans. Disclosed is a large-model-driven method for real-time interaction with a digital human, comprising the following steps: S1, inputting a question text by means of a human-machine dialog frontend and sending the question text to a large model backend; S2, the large model backend performing large model inference on the basis of the question text and a knowledge base to generate an answer summary and a full answer and respectively sending the answer summary and the full answer to a digital human backend and the human-machine dialog frontend for display; S3, a classifier generating a plurality of classification labels on the basis of the answer summary, and the digital human backend performing retrieval and matching on a preset video library on the basis of the classification labels to obtain a target preset video and sending the target preset video to the human-machine dialog frontend; and S4, a digital human processing engine generating a digital human summary video on the basis of the answer summary and pushing the digital human summary video to a player, and the player sequentially playing the target preset video and the digital human summary video. Smoother experience of real-time question-and-answer interaction with a digital human is finally achieved, thereby reducing loss of realism in the experience.
Owner:WEIKE ZHIJIAN (FOSHAN) TECHNOLOGY CO LTD

Video generation method and apparatus, electronic device, and storage medium

The present disclosure relates to a video generation method and device, an electronic device and a storage medium. The method comprises: obtaining a target single-shot video in a preset video library as a first single-shot video of a multi-shot video to be generated, and the preset shot library comprising a plurality of single-shot videos; determining at least one group of video sequences starting from the first single-shot video according to visual correlation features of the first single-shot video and visual correlation features of a first group of candidate videos, the visual correlation features representing contextual correlation of the single-shot videos, and the first group of candidate videos comprising other single-shot videos in the preset video library except the first single-shot video; and splicing the at least one group of video sequences to obtain at least one multi-shot video. The embodiments of the present disclosure can improve the video production efficiency.
Owner:SENSETIME GRP LTD

Video analysis method and device, equipment and medium

The embodiment of the invention relates to a video analysis method and device, equipment and a medium, and the method comprises the steps: obtaining screening information in response to a screening input operation on a first analysis page, inputting the screening information into a video analysis model, extracting a plurality of second video sets corresponding to the screening information from a video library through the video analysis model, and switching the displayed multiple first video sets into multiple second video sets in the first analysis page, and displaying multiple third video sets corresponding to a first question tag in the multiple second video sets in response to a selection operation on the first question tag in the multiple question tags. According to the technical scheme, the video sets corresponding to the screening information are extracted from the video library through the video analysis model and displayed, the multi-video analysis efficiency is effectively improved, the video sets corresponding to the problem labels are further screened and displayed by selecting the problem labels, and the video analysis efficiency is improved. And the difficulty and the cost of common problem analysis are reduced.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Street view and satellite image matching method and equipment

The invention provides a street view and satellite image matching method and equipment. The method comprises the following steps: acquiring a street view image, a satellite image and an unmanned aerial vehicle video library; extracting a first visual feature of the satellite image and a second visual feature of the unmanned aerial vehicle video key frame; calculating a first similarity score between the satellite image and the unmanned aerial vehicle video key frame by using the first visual feature and the second visual feature; performing visual feature comparison and semantic description association on the streetscape image and the unmanned aerial vehicle video key frame, and calculating a second similarity score between the streetscape image and the unmanned aerial vehicle video; and based on the first similarity score and the second similarity score, bridging calculation is carried out by using the unmanned aerial vehicle video as transfer to obtain an indirect matching score of the streetscape image and the satellite image, and a cross-view matching result is obtained in combination with the direct visual feature matching score. Based on the method, the invention also provides a street view and satellite image matching device. According to the method, the accuracy and practicability of cross-view image retrieval are remarkably improved.
Owner:HARBIN INST OF TECH AT WEIHAI

Video generation method and apparatus, electronic device, and medium

This application discloses a video generation method, apparatus, electronic device, and medium, belonging to the field of artificial intelligence technology. The video generation method includes: extracting behavioral descriptive words and visual descriptive words from a first text; determining target video segments matching the behavioral descriptive words from N first videos, and determining target video frames matching the visual descriptive words from the N first videos; generating a target video based on the target video segments and the target video frames; wherein the N first videos are N videos in a video library similar to the first text; and N is an integer greater than 1.
Owner:VIVO MOBILE COMM CO LTD

A method and system for family audio-video imprint aggregation and biographical film generation based on bloodline narration

The application discloses a kind of family audio-video imprint automatic aggregation and personal biography film generation method and system based on genealogy figure. First, structured family knowledge base including family member node, blood relationship edge and associated multimedia data is constructed. In response to biography generation request, start multi-modal figure imprint aggregation: fusion face recognition and blood relationship auxiliary matching, automatically identify and extract all image fragments of target figure appearance from audio-video library, form cross-modal material set. Further, blood narrative guided intelligent editing is carried out: based on age estimation, material time sequence segmentation is carried out, narrative structure is automatically generated in combination with blood event node, and material enhancement technology is used to output into film. The system automatically generates the "family imprint" column in personal data card, and aggregates and displays related audio-video materials. Through integrated path, scattered family memories are gathered into personal image biography, which greatly reduces the production threshold, realizes the automatic aggregation and generation of family memories.
Owner:BEIJING AIHE INFORMATION TECHNOLOGY CO LTD

A method for identifying complex targets in massive videos based on human-computer collaboration

The application discloses a kind of based on man-machine cooperation's mass video complex target retrieval method.Currently, the generalization ability of intelligent system based on machine vision is weak, when environment changes, target exists shielding or presents camouflage state, still need a lot of manpower intervention in mass video retrieval target, time-consuming and laborious.The application first according to the coarse-grained feature of target uses retrieval model in video library and carries out pre-screening, and the target candidate set obtained by screening is made into brain-eye cooperation RSVP paradigm and presented to subject.Subject determines target, and its specific features are input into retrieval model, so as to realize the fast positioning and tracking of target in mass video.The application combines the generalization reasoning ability of human and the rapid retrieval ability of machine, can effectively improve the efficiency and accuracy of video investigation, has strong scientific significance and social significance.
Owner:HANGZHOU DIANZI UNIV

Method and system for automatically generating video content

The present invention provides a method and system for automatically generating video content. The method comprises the following steps: obtaining a hot event and its corresponding original video, extracting multiple event features describing the hot event; querying a plurality of target candidate videos whose video content matches each event feature from a video library; querying a plurality of target candidate songs whose lyrics match each event feature from a song library; combining the target candidate videos corresponding to each event feature with the corresponding target candidate songs to generate a video clip corresponding to each event feature; determining the video node corresponding to each event feature in the original video, and inserting the video clip corresponding to each event feature into the video node corresponding to the original video to generate the target video. The present invention improves the efficiency of short video creation and combines hot events with content display to provide users with a richer, more interactive, and personalized content experience.
Owner:GUANGZHOU FANYU NETWORK TECH CO LTD

Video text combined retrieval method, device, electronic device and storage medium

The present application discloses a video text combined retrieval method, device, electronic device and storage medium, and relates to the field of video retrieval technology. By obtaining the original video and the retrieval text, the original video frame is encoded to obtain high-level visual features and mid-level visual features, and the retrieval text is encoded to obtain retrieval text features. A high-level branch is set to extract high-level retention features in high-level vision and high-level difference features in retrieval text features based on temporal information, and fuse them to obtain high-level fusion features. A mid-level branch is set to use the attention mechanism to extract finer-grained spatiotemporal features based on mid-level retention features and mid-level difference features, and fuse them to obtain mid-level fusion features. Finally, the target fusion features are obtained by performing hierarchical multi-fusion on each feature, thereby retrieving the preset video library to obtain the target video. In this way, the user's visual needs are described from different granularities, which effectively improves the accuracy of video retrieval and accurately finds the target video that meets the user's needs.
Owner:PENG CHENG LAB

Method and apparatus for recommending short video, electronic device, and storage medium

The present application belongs to the technical field of video recommendation, and particularly relates to a method and apparatus for recommending a short video, an electronic device, and a storage medium. The method comprises: step 1, obtaining historical viewing data of a user, determining a first short video identifier information list and a second short video identifier information list, and generating a first feature vector and a second feature vector; step 2, obtaining a short video identifier information list to be expanded, and, on the basis of the short video identifier information list to be expanded, generating a third feature vector; step 3, on the basis of the first feature vector, the second feature vector, and the third feature vector, calculating a fourth feature vector; and step 4, obtaining a short video library to be recommended, extracting a short video identifier information list to be recommended, generating a fifth feature vector, calculating the similarity between the fifth feature vector and the fourth feature vector, and, on the basis of the similarity, filtering out a short video to be recommended. According to the present application, short video identifier information is expanded, thereby improving the diversity and freshness of short video recommendation.
Owner:BEIJING FENGPING INTELLIGENT TECHNOLOGY CO LTD

Cross-camera video pedestrian search method, terminal device and computer-readable storage medium based on representation learning

The present invention discloses a method for searching pedestrians across camera videos based on representation learning, a terminal device, and a computer-readable storage medium, comprising the following steps: obtaining pedestrian data across camera videos, constructing a query set and a candidate video library corresponding to the query set; inputting the video into a target detection network to learn each pedestrian bounding box, the confidence score of the bounding box, and the features of each pedestrian; inputting the pedestrian bounding box and features into a multi-target tracking network based on temporal feature fusion, performing data association on the video frames, obtaining the trajectory of each pedestrian and the temporal features of each trajectory; calculating the similarity between the pedestrian trajectory feature vector in the query set and all the pedestrian trajectory feature vectors of the pedestrian in the candidate video library, and performing accuracy calculation. Aiming at real-world video surveillance scenarios, the present invention realizes the search for target pedestrians in cross-camera scenarios by extracting pedestrian features and data association using long-term temporal features in the video.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Video retrieval method based on video-text double alignment and ETVA model

The invention discloses a video retrieval method based on video-text dual alignment and an ETVA model, and the method comprises the steps: constructing and training the ETVA model, capturing local text features and global text features described by a text input by a user, retrieving a video matched with the global text feature description from a massive video library based on the global text features, and carrying out video retrieval based on the video-text dual alignment. And retrieving a video matched with the local text feature description from the videos matched with the global text feature description based on the local text feature. According to the method, the global features and the fine-grained features are extracted from the video and the text through the ETVA model, and accurate alignment is carried out. By means of the double alignment mode, the model can deeply mine and capture more complex semantic association between the video and the text, interaction between local representation of the video and local representation of the text is enhanced, the model can more accurately understand video content and text description, and then the accuracy and efficiency of video retrieval are remarkably improved.
Owner:HEFEI UNIV OF TECH

360-degree heavy rail circular track ring video synthetic aperture radar jamming method and system

This invention discloses a method and system for jamming 360-degree heavy-orbit circular-track video synthetic aperture radar (SAR). Specifically, it involves: analyzing intercepted enemy single-track signals to obtain enemy radar signal parameters and platform motion parameters; guiding multi-track analysis based on enemy parameters to obtain parameters for different tracks; using a self-deception template to solve for false point coefficients and controlling the size and orientation of the video template to obtain a deception jamming modulation coefficient video library; and generating a heavy-orbit-specific deception strategy based on the reconnaissance parameters of different tracks; finally, using the deception jamming modulation coefficient video library and the heavy-orbit-specific deception strategy to generate and modulate a deception signal, obtaining a track-by-track forwarding azimuth-elevation joint deception jamming signal to jam the enemy track by track, achieving the deception effect. This invention achieves the purpose of deception jamming of a 360-degree heavy-orbit circular-track videoSAR system, with the advantage of simultaneously performing heavy-orbit jamming and azimuth-elevation joint deception.
Owner:NANJING UNIV OF SCI & TECH

Video library clip retrieval method and device for subject-text joint query

The invention provides a video library fragment retrieval method and device for subject-text joint query, and relates to the technical field of computer vision. The method comprises the following steps: acquiring a first video representation sequence of each video in a video library; performing representation extraction on the subject query and the text query in the subject-text joint query to obtain a first query representation sequence; performing interactive calculation on the first video representation sequence to obtain a second video representation sequence fused with context semantics, and performing interactive calculation on the first query representation sequence to obtain a second query representation sequence; according to the semantic representation in the second query representation sequence and the second video representation sequence, calculating the similarity between the subject-text joint query and each video, and taking the video corresponding to the maximum similarity as a retrieval video; and according to the second query representation sequence and the starting and ending timestamps of a second video representation sequence prediction fragment corresponding to the retrieval video, obtaining a retrieval video fragment corresponding to the subject-text joint query.
Owner:TSINGHUA UNIVERSITY