Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

67 results about "Open domain" patented technology

Open domain-oriented adaptive public opinion data classification method and system

The invention discloses a self-adaptive public opinion data classification method and system for an open domain, and the method comprises the steps: collecting original text data from a plurality of data sources, carrying out the preprocessing of the original text data, obtaining a plain text list, converting the plain text list into high-dimensional semantic vectors in batches, and enabling all high-dimensional semantic vectors to form an embedded matrix; calculating a minimum clustering number and a maximum clustering number according to the number of texts in the plain text list, generating various clustering schemes corresponding to the embedded matrix through all clustering thresholds in a clustering range, calculating a comprehensive score of each clustering scheme, and selecting the clustering scheme with the highest comprehensive score as an optimal scheme; generating subject terms of all clustering clusters based on a large model; and integrating the optimal parameters, the texts in each cluster and the subject terms in each cluster, and outputting a structured classification result. The open domain public opinion data can be efficiently, intelligently and interpretably classified.
Owner:CHENGDU SPACEON IND CO LTD

Customized image generation method and device based on main body consistency, equipment and storage medium

The invention discloses a customized image generation method and device based on main body consistency, equipment and a storage medium, and the method comprises the steps: screening images and texts from a public data set, and constructing an open domain image and text data set; generating a target map by means of an existing model, and constructing a good and bad preference comparison data set; extracting an image by using a pre-trained coding model, and embedding and associatively storing the image and the text to establish a multi-modal feature library; fusing and embedding through an attention mechanism to generate a multi-modal fusion feature; diffusion and noise addition are carried out on the good and bad target images to serve as training supervision signals; and generating a subject consistent target map by using a noise distribution training model. According to the method, through signal supervision of subject consistency provided by a target, model learning is enabled to balance subject reservation and instruction following, test fine tuning is not needed, generalization and generation efficiency is improved, the problems of cooperative control and insufficient generalization efficiency of an existing method are solved, and the method can be widely applied to customized and personalized image generation scenes.
Owner:XINJIANG TECH INST OF PHYSICS & CHEM CHINESE ACAD OF SCI

Defense method and device for graph neural network backdoor attack, equipment and medium

The invention relates to the technical field of machine learning, in particular to a defense method and device for graph neural network backdoor attacks, equipment and a medium, when an open domain node classification task is received, the open domain node classification task is input into a preset trigger detection model, the model can determine unknown class nodes and cut edges of the unknown class nodes to obtain an initial defense sub-graph, and the initial defense sub-graph is used for defending the open domain node classification task. And performing importance score calculation on the target defense nodes in the initial defense subgraph to form a final defense subgraph, inputting the final defense subgraph into a preset dynamic classifier, and outputting a classification result of the target defense nodes. According to the method, the backdoor attack problem faced by the graph neural network in an open domain scene is effectively solved, and the classification accuracy and security are improved.
Owner:SHENZHEN UNIV

Large language model retrieval enhancement generation method based on adaptive rewriting selection

The invention provides a large language model retrieval enhancement generation method based on adaptive rewriting selection, and is suitable for the field of natural language processing and information retrieval. According to the method, a pre-trained large language model is introduced to automatically generate diversified rewriting queries, and a self-supervised rewriting sequencer is combined to perform correlation evaluation and sequencing on candidate rewriting statements. Through a context multi-arm bandit selector, the optimal rewriting number is dynamically determined according to query semantics, a high-quality rewriting subset is selected in a self-adaptive mode, and the coverage degree and precision of information retrieval are effectively improved. And the rewriting-driven knowledge retrieval module utilizes a plurality of high-quality rewriting, integration and deduplication related knowledge blocks in parallel, and continuously optimizes a Bandit strategy based on a feedback signal to realize online self-learning. Different from a traditional RAG system depending on fixed parameters and static rewriting, the method can intelligently adjust the retrieval process for complex or variable queries, and the accuracy and practicability of a retrieval enhancement generation system in an open domain and a multi-hop reasoning scene are remarkably improved. According to the method, efficient and flexible technical support is provided for intelligent question answering and knowledge discovery in a complex environment, and the application effect and popularization value of the large language model are greatly enhanced.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Structured dialog segmentation and state tracking

Systems and methods for open domain dialog segmentation and state tracking are provided. Specifically, a computing device may acquire and analyze a dialog in near real-time, generate a structured cue template for a state prediction model based on the dialog, and generate a structured output using the state prediction model based on the structured cue template. The structured output includes a round summary and a state tag for each round of conversation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Large-model-driven real-time interactive digital human driving method, system and equipment

The invention belongs to the technical field of digital humans, and relates to a large-model-driven real-time interaction digital human driving method, system and device and a storage medium, and the method comprises the steps: 1) obtaining the dialogue input of a user, and inserting a dialogue background and a historical question and answer pair for the dialogue input to form an input text of a large language model; 2) inputting the input text of the large language model into the large language model, and generating answer content by the large language model; 3) generating voice data based on the answer content; 4) generating facial animation and limb animation data based on the voice data; 5) playing the voice data, and driving the digital person by using the facial animation data and the limb animation data; and 6) outputting the real-time dialogue content of the digital person in the video stream. According to the method, the problems of insufficient semantic level synchronization precision and low interaction naturalness during digital human face action driving are solved, open domain real-time question answering is supported, and the natural performance of digital human real-time interaction is realized.
Owner:BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD

Teacher-query guide type compression optimization training method and system and query method and system based on open domain questions and answers

The invention provides a teacher-query guide type compression optimization training method and system based on open domain questions and answers and a query method and system. The compression optimization training method comprises the steps that candidate paragraphs are obtained according to user queries, each query is matched with the corresponding candidate paragraphs, the candidate paragraphs are segmented into sentence sets, and training units are formed; the teacher model performs data distillation on the training unit based on a preset requirement to generate a distillation sample set; the method comprises the following steps: respectively coding queries and sentences in a distillation sample set into vector representations through a predefined coding function, measuring the similarity of the queries and sentences, and training a student model by taking the similarity as a supervision signal to obtain a trained student model; through a teacher model knowledge migration and query focusing mechanism, the accuracy and efficiency of a student model in an extraction type compression task are improved, redundant information interference is reduced, more compact and highly related context input is provided for a retrieval enhancement generation framework, and the method is suitable for an efficiency-sensitive open domain question and answer scene.
Owner:MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY

Open domain question and answer task processing method and comprehensive processing framework system

The invention relates to the technical field of natural language processing and artificial intelligence, in particular to an open domain question and answer task processing method and a comprehensive processing framework system. A complete, reliable and explainable reasoning chain from an original question to an explicit reasoning basis and then to a final answer is constructed. According to the framework, the accuracy and robustness of a question answering system are remarkably improved, and the illusion problem of a large language model in an open domain scene is effectively relieved. According to the method, through three innovation mechanisms of question enhancement, collaborative evidence collection and multi-stage abstract integration, substantial technical progress is made in the aspects of accuracy, robustness, interpretability, information integration capability and the like, and an effective solution is provided for constructing a high-reliability large language model open domain question-answering system.
Owner:BEIJING TECH & BUSINESS UNIV

A planning calibration method and system for an information analysis intelligent assistant

The application discloses a planning calibration method and system for an information analysis intelligent assistant, and comprises the following steps: in each planning round, a candidate strategy set is generated by a planning agent, the candidate strategy set contains an information retrieval target, a retrieval keyword combination and an information processing flow; the candidate strategy set is evaluated by multiple agents with heterogeneous private memories, a strategy selection module based on the agents is constructed for information consistency, and a strategy with the most stable evaluation result is selected as an optimal strategy of the round; a difference between a self-selected strategy of the planning agent and the selected optimal strategy of the round is recorded, a calibration constraint module guided by consistency is constructed, and the difference is converted into a cognitive calibration constraint and stored in the memory of the planning agent; strategy selection optimization is realized in single-round planning by relying on the strategy selection module based on information consistency; and the cognitive calibration constraint generated by the calibration constraint module guided by consistency is combined, so that the constraint is transmitted between different rounds and gradually accumulated, and a constraint is formed on strategy generation of the planning agent; and the trained agent is applied to an open domain webpage retrieval task to generate accurate retrieval results.
Owner:TIANJIN UNIV

Intelligent question-answering system and method based on multi-modal document view

The application discloses a multi-modal document view-based intelligent question answering system and method, and relates to the field of natural language processing. In order to solve the problems that a traditional open domain question answering system is complex in layout for different types of document views and it is difficult to uniformly model all objects in the prior art, the technical scheme provided by the application is as follows: a multi-modal document view-based intelligent question answering system, which comprises a multi-modal document view analysis module, a multi-modal document view retrieval module and a multi-modal document view question answering module; the multi-modal document view analysis module is used for extracting text information; the multi-modal document view retrieval module is used for retrieving document views related to preset information in the text and performing priority arrangement; and the multi-modal document view question answering module is used for extracting the document views with higher priority and outputting. The application is suitable for application in intelligent question answering of multi-modal document views.
Owner:HARBIN INST OF TECH

Role-based large language model to enable security and accuracy

A first query having a first privacy status is received. A response to the first query is obtained based on output(s) of one or more machine learning (ML) models of an open domain dialog system. Each ML model is trained to predict responses to queries having the first privacy status. Data associated with the first query and the response is provided as training data for the ML models, in view of the first privacy status. A second query having a second privacy status is received. A closed domain dialog system associated with a context of the second query and having a privacy status corresponding to the second privacy status is identified. The second query is forwarded to the closed domain dialog system for obtaining a response to the second query. Data associated with the second query is not provided to train the ML models of the open domain dialog system.
Owner:NVIDIA CORP

Open-domain large model reasoning enhancement method based on probability process supervision

The application discloses an open field large model reasoning enhancement method based on probability process supervision. The application adopts reinforcement learning to perform multi-round iterative optimization on a strategy model: in each round, a dynamic golden thinking chain is generated based on a question and an answer as a reference, a plurality of exploratory paths are sampled, a step-level path loyalty reward is calculated through alignment to quantify logical effectiveness, a layered reward mechanism is constructed by combining a weighted process reward and answer correctness, and a final reward signal is used to update the model. The method does not require an external reward model and artificial labeling, realizes fine-grained process supervision through self-supervision, dynamically updates the supervision signal to adaptively improve the model, is highly versatile, and is suitable for various open reasoning tasks.
Owner:ZHEJIANG UNIV

A medical text classification method and device based on prompt learning

The application provides a medical text classification method and device based on prompt learning. The method comprises the following steps: obtaining prompt information for primary classification from original medical text based on event prior information and knowledge prior information, wherein the primary classification comprises a department category; filtering the original medical text, and inputting the filtered text and the prompt information into a large language generation model after integration; calculating the similarity between the result sequence output by the large language generation model and each secondary classification label representing a disease category under the primary classification, and taking the label category corresponding to the maximum value of the similarity as the secondary category output by the large language generation model. The application can realize primary classification based on the department category, and can also realize secondary classification based on the disease category under the primary classification, which is more consistent with the general cognition in the medical field, so that the classification result is more standardized; meanwhile, the classification label can be unfixed, so that multi-level text classification in an open domain can be effectively realized.
Owner:BEIJING SHENRUI BOLIAN TECH CO LTD +1

Knowledge boundary perception-based search enhancement generation method and system, electronic device, and storage medium

The application discloses a retrieval enhancement generation method and system based on knowledge boundary perception, an electronic device and a storage medium, and belongs to the technical field of natural language processing. The method comprises the following steps: generating a high-quality supervised track by using a teacher model, and learning the ability of gap planning and answers by instruction fine-tuning of a weak model; paired samples reflecting overconfidence and over-conservatism are constructed, and a DPO algorithm is used for confidence calibration; gap planning is generated by a student model during actual prediction, and it is accurately determined whether each knowledge point needs retrieval according to cognitive information labels, and accurate retrieval is triggered only for the knowledge points with knowledge gaps. The application can be widely applied to open domain question answering, dialogue systems and knowledge-intensive tasks by explicitly identifying knowledge boundaries, dynamically adjusting thresholds and fine-grained on-demand retrieval, while ensuring the accuracy of answers, significantly reducing the consumption of computing resources and response delay.
Owner:DALIAN UNIV OF TECH

A reliable reasoning and question answering method and apparatus that combines large-scale models and knowledge graphs

This invention discloses a reliable reasoning and question-answering method and apparatus based on the collaboration of a large model and a knowledge graph, belonging to the fields of artificial intelligence and natural language processing. It extracts keywords through semantic guidance and links them to entities in the knowledge graph to determine the initial entities; then, it explores paths based on a decision-evaluation collaborative mechanism, expanding paths by combining hard constraints and semantic soft guidance, and dynamically pruning based on the difference in immediate rewards and patience thresholds; finally, it outputs the answer and a traceable reasoning path. This invention solves the problems of rule-dependent entity recognition, black-box reasoning process, and poor interpretability in open-domain question answering, balancing semantic understanding generalization ability with controllable and reliable reasoning process, improving question-answering accuracy, reasoning efficiency, and result reliability, and its overall performance is superior to traditional rule-based, heuristic, or pure neural network-based methods.
Owner:XIAN INT STUDIES UNIV

Open-domain natural language reasoning question answering system and method driven by large language model

The application provides a large language model driven open domain natural language reasoning question and answer system and method, a question rewriting module rewrites a user question into a rewritten question; a center calculation and management module manages calculation and knowledge resources of a large language model, and outputs the calculation and knowledge resources of the large language model required by a question core engine module to one or more sub-question and answer modules in a question and answer core engine module according to the type of the rewritten question; the question and answer core engine module reasons to obtain one or more candidate answers of the rewritten question and explainability information of the candidate answers according to the rewritten question and the calculation and knowledge resources of the large language model; and an aggregation reasoning module aggregates and reasons to obtain a final answer of the rewritten question and explainability information of the final answer according to the one or more candidate answers of the rewritten question and the explainability information of the candidate answers, and supports comprehensive question types, is easy to expand, is explainable, and has strong universality by using a large language model.
Owner:TSINGHUA UNIVERSITY

An open domain corpus relation joint extraction method

An open domain corpus relationship joint extraction method comprises the following steps: S1, extracting the feature vector of characters in the corpus; S2, performing feature fusion in the graph attention network; S3, extracting the relationship phrase in the corpus; S4, extracting the entity pair phrase in the corpus; S5, according to the relationship phrase extracted in step S3 and the corresponding entity pair phrase extracted in step S4, the three tuples are formed, and the confidence of the three tuples is determined, if the confidence is greater than or equal to the set confidence threshold, then the three tuples are taken as the open domain relationship three tuples of the input corpus. Through the above scheme, the problems of redundant relationship three tuple sequence, overlapping relationship three tuple, and low relationship three tuple extraction accuracy in open domain relationship extraction are solved.
Owner:JINLING INST OF TECH

Insurance intelligent customer service method and device based on large language model

The invention provides an insurance intelligent customer service method and device based on a large language model. The device comprises an insurance knowledge base module, a data access module, a data processing module, an insurance knowledge graph module, a multi-modal fusion module, a large language model module and a response generation module. By introducing an optimized GPT-NEOX large language model and fusing an InsuranceQA professional corpus and an InsurKG insurance knowledge graph, diversified insurance problems proposed by users can be dynamically understood, and natural, accurate and interpretable answers are generated. Compared with a traditional static answer extraction mode based on BERT and other models, the method has the remarkable advantages in the aspects of context understanding, open domain question and answer adaptability, deployment efficiency and professional domain knowledge coverage, and the response quality and user experience of insurance intelligent customer service are effectively improved.
Owner:PICC INFORMATION TECH CO LTD +1

Open domain long text classification method and apparatus based on topic analysis

The application relates to an open domain long text classification method and device based on theme analysis. The method comprises the following steps: constructing a field self-adaptive word table; preprocessing original long text to obtain purified text; using an LDA model to construct a latent theme space; extracting a core sentence group from the long text through semantic clustering to generate an abstract; mapping the abstract to the theme space to calculate a correlation degree score and output a theme identification; and converting the theme identification into a business classification label by querying a theme-label mapping table. The application effectively solves the problems of complex long text semantics and dynamic label system changes, and improves classification accuracy and expansibility.
Owner:NAT UNIV OF DEFENSE TECH +1

Chinese multi-turn dialogue model

The application provides a Chinese multi-round dialogue model, taking GPT-2 as a main framework, fusing a dialogue history keyword extraction model and a dialogue theme switching discrimination model, so that the Chinese multi-round dialogue model can discriminate whether to switch the chat theme when generating a reply sentence, and consider whether to copy a word from the dialogue history or select a word from a global word table according to a generation mechanism, thereby improving the multi-round dialogue effect, enhancing the relevance of the context, and overcoming the difficulty of answering irrelevant questions in the open domain dialogue.
Owner:SHANGHAI MARITIME UNIVERSITY

An open-domain scientific knowledge discovery method and device based on a pre-trained language model

The application discloses an open domain scientific knowledge discovery method and device based on a pre-trained language model, constructs an input template comprising a head entity, a first prompt, a second prompt and a tail entity mask, fills the pre-trained embedding of the head entity of each triple containing a target relation, the discrete tokens of the first prompt corresponding to the target relation and the second prompt tokens into the input template, processes the tail entity mask, forms input sample data, constructs a single pre-trained language model for each target relation, trains the pre-trained language model using the input sample data corresponding to the target relation, optimizes the embedding representation of the first prompt and the second prompt, and predicts the missing entity in the triple using the optimized embedding representation of the first prompt and the second prompt and the pre-trained language model, thereby improving the efficiency and accuracy of the pre-trained language model in discovering knowledge.
Owner:ZHEJIANG UNIV +1

Automation method oriented to large language model instruction tuning data optimization

The invention discloses an instruction tuning data automatic optimization method oriented to a large language model. According to the method, the problems of answer errors, information missing, expression redundancy, instruction-output inconsistency and the like existing in existing instruction data are solved through the end-to-end process of hierarchical data selection, iterative data rewriting and data fusion and model optimization, and the model instruction following ability and downstream task performance are remarkably improved while the labor cost is reduced. According to the method, data quality division is carried out through subjective scoring and objective indexes, iteration data rewriting of four steps is combined, quality and diversity are considered, semantic drift is effectively inhibited, and generalization performance is improved. The method is generally used for open domain instruction data construction and optimization.
Owner:EAST CHINA UNIV OF SCI & TECH

A large model routing data synthesis method with diversified expression styles

The application discloses a large model routing data synthesis method with diversified expression styles, constructs a user agent to simulate the language style and preference of a real user, generates query questions with diversified styles and clear semantics, inputs the query questions into a candidate language model for reasoning, and generates a benchmark answer by a large language model with superior performance; based on the comparison between the output of the candidate model and the benchmark answer, the content quality is automatically evaluated, the user satisfaction evaluation is simulated in combination with the user agent features, and multi-dimensional routing labels are comprehensively generated; finally, the question, the candidate model output and the routing label are packaged to form a routing data set, which is used for training or evaluating a large model router. The application can automatically synthesize a routing data set covering multiple expression styles, multiple question types and performance differences of multiple models, significantly improves the generalization ability and practicability of a routing model in a real open domain question and answer scene, and does not require high artificial labeling cost.
Owner:ZHEJIANG UNIV

Steganographic text steganalysis resistance enhancement method based on deep learning

The application belongs to the technical field of information hiding in information transmission, and discloses a steganographic text anti-steganalysis capability enhancement method based on deep learning. The method learns the characteristics of the data set in the open domain environment by using the deep learning method, generates steganographic text with the statistical distribution of the original data set while embedding secret information, and finally recommends the corresponding emoticons for the steganographic text according to the distribution of the emotional tendency and emoticons of the steganographic text in the real environment, so as to further strengthen the anti-steganalysis capability of the steganographic text. The application solves the problem of missing emotional clues in the generated text steganography, so that the anti-attack capability of the steganographic text is stronger when the steganographic text is transmitted in the public channel.
Owner:ZHEJIANG GONGSHANG UNIVERSITY

Natural language dialogue method and device, electronic equipment, medium and product

The invention discloses a natural language dialogue method and device, electronic equipment, a medium and a product, and relates to the technical field of artificial intelligence, the natural language dialogue method comprises the following steps: inputting dialogue input content and a preset cue word template into a preset large model; guiding the preset large model to generate a concept feature set related to the dialogue input content through the preset cue word template; and guiding the preset large model to generate each language segment through the concept features of different concept dimensions in the concept feature set, and determining reply content of the dialogue input content based on each language segment. According to the method, a hierarchical dialogue generation framework for guiding the preset large model to gradually think is generated, a systematic thinking mode of human is successfully simulated, the limitation of the large model in an open domain dialogue scene is broken through, and the dialogue reply quality is effectively improved.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Open domain Chinese event pattern induction method and system based on multi-dimensional feature fusion

The invention discloses an open domain Chinese event pattern induction method and system based on multi-dimensional feature fusion, and the method comprises the steps: S100: extracting and fusing the multi-dimensional features of event texts in a corpus, and obtaining an event representation vector based on the multi-dimensional features; the multi-dimensional features comprise semantic features, simplified short text features and structured event features; s200, the event representation vectors are clustered in the spherical potential space, and event clusters are obtained; and S300, generating a normalized event mode corresponding to each event cluster according to the event clusters. According to the automatic Chinese event pattern induction method and system, the discrimination capability of event representation and the clustering precision of event clusters can be enhanced, more condensed and more standard event patterns can be inducted, and the method and system are suitable for open domain scenes.
Owner:BEIJING INFORMATION SCI & TECH UNIV

A semantic-based open domain webpage knowledge extraction method and system

The application provides a semantic-based open domain webpage knowledge extraction method, which comprises the following steps: obtaining a skeleton tree of an open domain webpage, splitting a skeleton node of the skeleton tree to obtain skeleton sub-nodes of the skeleton node, and generating a skeleton sub-node sequence; labeling a classification label for the skeleton sub-nodes and the skeleton node, performing relation extraction on the skeleton tree according to the classification label, obtaining a relation sub-node sequence of an extraction task, and generating a relation fragment; performing object extraction on the skeleton tree based on the relation fragment, taking the extracted skeleton sub-node sequence as an object fragment; and taking the relation fragment and the corresponding object fragment as an extraction result of the extraction task. The application also provides a semantic-based open domain webpage knowledge extraction system and a data processing device for open domain webpage knowledge extraction.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Visual analytics system for interpreting open-domain question answering models

The present application belongs to the technical field of open domain question answering model analysis, and particularly relates to a visual analysis system for explaining an open domain question answering model.The system comprises an explanation engine module, a process analysis module and a view module; the explanation engine module uses an attribution method to attribute the final output and implicit output of each module of the OpenQA model at global and local levels; the process analysis module visualizes the model information, data and explainability data generated via the explanation engine in the VEQA as various views of a user analysis interface, and the user performs multi-level exploration in the order of data set, subset, single instance and single paragraph according to a linear workflow; the view module comprises a summary view, a context view, an instance view and a tree view, and is used for visual analysis; the system can help understand the decision reasons of the OpenQA model and provide insights for model improvement; and the system also supports fine-grained exploration of the decision process within a single module.
Owner:FUDAN UNIVERSITY

Video non-main body element stripping method and system and storage medium

The invention relates to the technical field of video processing and artificial intelligence, in particular to a video non-main element stripping method and system and a storage medium, and the method comprises the steps: carrying out the audio-video separation of an advertisement video, and obtaining a video frame sequence; inputting the video frame sequence into a multi-modal large model, identifying non-main-body elements in corresponding video frames, and generating text description of the non-main-body elements; based on the text description and the video frame sequence, positioning a bounding box of a non-main body element by using an open domain target detection model; taking the bounding box as a prompt input segmentation model, generating a mask of a non-main body element, tracking the non-main body element in a subsequent video frame by using a time consistency mechanism of the segmentation model, and generating a continuous video mask sequence; and background filling is carried out on the missing area after the non-main body elements are removed in the video frame through the video diffusion repair model, and decoupled video output is obtained. According to the scheme of the invention, the accuracy and editing efficiency of video non-main body element stripping are effectively improved.
Owner:GUANGZHOU TAIDONG TECH CO LTD

Training-free video corpus time instant retrieval method based on adaptive calibration mechanism

The present application relates to a training-free video corpus time moment retrieval method based on an adaptive calibration mechanism, and belongs to the technical field of computer vision and multi-modal information processing. It comprises: based on the constructed query event chain and video event chain, calculating the event level similarity score, and combining the mean-variance joint scoring mechanism for cross-modal retrieval to obtain candidate proposals; in the time positioning stage, a boundary level association is established among the video candidate proposals through a cooperative mechanism, and a profit-loss dynamic feedback strategy is used to perform adaptation and iteratively optimize the time boundary of the candidate proposals. The present application realizes a closed-loop retrieval process from text semantic analysis to video accurate matching. Compared with the prior art, the present application does not need large-scale labeled data for training or parameter updating, effectively solves the generalization bottleneck and deployment cost problem of video retrieval in open domain scenarios, and significantly improves the accuracy and robustness of time positioning under zero training conditions.
Owner:KUNMING UNIV OF SCI & TECH