Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

252 results about "Sample classification" patented technology

Collaborative classification method and system fusing advantages of large and small models

The invention provides a collaborative classification method and system fusing advantages of large and small models, and the method comprises the steps: inputting to-be-classified data into a trained zero-sample classification small model, and outputting candidate label domain screening information and an initial classification result; inputting the to-be-classified data, the candidate label domain screening information and the initial classification result into the trained large model, and outputting a final classification result; the training process of the small model comprises the following steps: inputting training data into the zero sample classification small model to obtain a preliminary prediction result; screening and obtaining pseudo label data based on an active learning strategy; checking and re-marking the pseudo-label data by using the large model to obtain a modified pseudo-label data set; and carrying out iterative training on the zero sample classification small model by utilizing the modified pseudo label data set. The method has the advantages that a small model has classification performance close to that of a large model while keeping lightweight calculation characteristics; the calculation burden of a large model is reduced, and the accuracy of classification decision is improved.
Owner:MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY

Multimodal context selection for large language model based resolutions addressing technical issues

A method for technical issue resolution. The method includes: receiving, from a user, a text query concerning a technical issue; obtaining query-related context relevant to the text query; and processing, through a large language model (LLM), the text query and the query-related context to produce a multimodal query response used by the user to address the technical issue. More specifically, embodiments described herein utilize text topic and zero shot classification models to translate multimodal technical documentation (e.g., including text and images) into topic relevant metadata; and process queries, pertaining to technical issues, using a multimodal LLM provided with query-related text and image context derived from said topic relevant metadata.
Owner:DELL PROD LP

Backdoor attack method and device, processing equipment, program product and medium

The embodiment of the invention provides a backdoor attack method and device, processing equipment, a program product and a medium, and is applied to the technical field of artificial intelligence. The method comprises the steps that target attention distribution corresponding to a target vision converter model is constructed based on a sample set of a target category, and the target vision converter model is trained to be capable of classifying all samples in the sample set into the target category; based on the target attention distribution and the to-be-processed source image, generating a trigger corresponding to the to-be-processed source image; based on a dual-objective loss function, parameters of the trigger are iteratively optimized, a target trigger is obtained, and the dual-objective loss function is constructed based on Wasserstein distance and classification loss; and generating a poisoning sample corresponding to the to-be-processed source image according to the target trigger and the to-be-processed source image. By the adoption of the method, the problem that an existing backdoor attack method is difficult to adapt to a self-attention mechanism, and consequently efficient hidden attack on a visual converter model is difficult to achieve is solved.
Owner:CHINA MOBILE SHANGHAI ICT CO LTD +2

Semi-automatic labeling method and system for rail transit engineering construction video images

The invention discloses a semi-automatic labeling method and system for rail transit engineering construction video images, and the method comprises the steps: removing redundant frames from a video stream, extracting key frames, and forming a block index set; and driving the multi-modal large model to pre-annotate the image by using the constructed cue word, and carrying out binarization processing on an annotation result. Based on active learning, migrating unmarked samples to a marked set for multiple times and training a key sample classifier CD until the scale of the marked set reaches the standard; meanwhile, a confidence classifier CC is trained by comparing manual and pre-labeling results. And for residual samples in the unmarked set, after the residual samples reach the standard through a CC precision test, marking tasks are divided according to a threshold value theta: high-confidence samples are pre-marked, and low-confidence samples are manually marked. And finally, updating the set of the manually labeled samples, and determining whether to retrain the model or not according to category distribution. On the premise of ensuring the labeling quality, the blindness of manual labeling is effectively reduced, and the efficiency of the labeling process is improved.
Owner:BEIJING URBAN CONSTRUCTION DESIGN & DEVELOPMENT GROUP CO LIMITED

Classification of tickets in building automation using a large language model

For ticket classification in a building automation system, a large language model (LLM) is used to classify. In one approach, a prompt is generated for zero-shot classification, and a prompt is generated for few-shot classification. In another approach, a hybrid annotation provides corrections (review) by an expert to correct LLM classification for sample tickets to be used as examples in the few-shot classification. The LLM may operate on a diverse and complex range of tickets in an efficient and scalable manner.
Owner:SIEMENS SCHWEIZ AG

Auxiliary rock core geological logging method based on drilling parameters

PendingCN121858905AImproved lithology prediction accuracyImprove catalogingKnowledge representationNeural learning methodsLithologyRelational model
The invention provides an auxiliary rock core geological logging method based on drilling parameters, and relates to the technical field of geological investigation, the method is applied to a logging control terminal, and the method mainly comprises the following steps: obtaining regional typical rock core samples and five types of drilling parameter signals, classifying and extracting static characteristics of the rock core samples, and carrying out parallel noise reduction analysis on the parameter signals; obtaining a matching relationship between the rock core and the parameter signal, and constructing a lithologic parameter bidirectional association reference library; on the basis of the matching relation of the lithologic parameter bidirectional association reference library, the model is embedded into an incremental learning module to update the weight, and a subsection lithologic pre-judgment result is output; and calling a preset catalog template to load a pre-judgment result, generating a catalog report after man-machine interaction recheck and correction, collecting correction data, reversely transmitting the correction data back to the reference library to update a matching relationship, and triggering incremental learning of the model to complete a closed loop. The technical problems that in traditional rock core logging, regional lithology adaptation is poor, parameter interference is large, a model is not dynamically optimized, the deep operation risk is high, and efficiency is low are effectively solved.
Owner:CHANGCHUN GOLD RES INST

Text classification method and device, electronic equipment and storage medium

The invention provides a text classification method and device, electronic equipment and a storage medium, and relates to the technical field of natural language process.The method comprises the steps that a training sample set is obtained, and the training sample set comprises sample texts and sample classification labels and sample reasoning reasons corresponding to the sample texts; performing fine tuning on a first pre-trained large language model through the training sample set to obtain a text classification model; in response to a text classification request, through the text classification model, based on a preset reasoning constraint parameter, only performing classification processing on a to-be-classified text to obtain a prediction classification label; wherein the preset reasoning constraint parameters comprise an output length limiting parameter and a logit probability intervention parameter. According to the method, the text classification efficiency can be improved while the accuracy and the reliability of a text classification result are improved.
Owner:IFLYTEK CO LTD +1

Tool calling and tool calling model training method and data generation model training method

The embodiment of the invention provides a tool calling method, a tool calling model training method and a data generation model training method. The tool calling model training method comprises the steps that a target intention and target tool calling parameters corresponding to the target intention are determined; the target intention is input into a data generation model, target query data corresponding to the target intention is obtained, the data generation model is obtained by training an initial data generation model through a positive and negative sample pair composed of an intention sample and a query data sample, and the positive and negative sample pair is determined in a preset sample classification mode; the query data sample is determined by performing data enhancement on the intention sample; inputting the target query data into the initial tool calling model to obtain a prediction tool calling parameter; according to the prediction tool calling parameter and the target tool calling parameter, training an initial tool calling model to obtain a tool calling model; the tool calling model is obtained by training the high-quality target query data, so that the accuracy and efficiency of processing the query data by the tool calling model are improved.
Owner:ALIBABA (CHINA) CO LTD

Power system transient stability assessment method

The invention discloses a power system transient stability assessment method, which comprises the steps of constructing a training sample set for transient stability assessment of a target power system, and marking real sample tags of the training sample set to generate an original sample set with tags; dividing samples in the original sample set with the labels into boundary samples, edge samples and internal samples; respectively executing differential weight extraction operation on the boundary sample, the edge sample and the internal sample to generate a model training subset; training the initial transient stability evaluation model based on the model training subset until a trained transient stability evaluation model is obtained; and inputting the collected real-time operation data of the target power system into the trained transient stability evaluation model to output a transient stability evaluation result of the power system. According to the invention, through sample classification and differential weighted training, the evaluation precision, generalization ability and decision reliability of the model in a critical state are significantly improved.
Owner:ZHEJIANG ZHENENG LANXI POWER GENERATION CO LTD +1

A multi-modal medical information fusion fetal auxiliary diagnosis system

The application discloses a multi-modal medical information fusion fetal auxiliary diagnosis system, which comprises a signal preprocessing module, an image generation module, a text extraction module, a multi-modal fusion module and an output module; wherein the multi-modal fusion module utilizes an image-text feature fusion network ITFN to perform feature extraction on image data acquired by the image generation module and text data acquired by the text extraction module, to obtain an image feature vector M i and a text feature vector M t ; the image feature vector M i and the text feature vector M t are weighted and fused to obtain a multi-modal feature vector Z f ; and the output module uses a full connection layer to reduce dimensions of the multi-modal feature vector Z f , and outputs a result of normal and pathological sample classification. The application adopts different models to learn local semantics and acquire more relevant knowledge information, realizes effective fetal distress pathological diagnosis and reduces a misdiagnosis rate, and has wide practical application value.
Owner:HANGZHOU DIANZI UNIV +1

Sample classification method and device, model training method and device, equipment and storage medium

The invention provides a sample classification method and device, a model training method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the field of generative models and the like. According to the specific implementation scheme, for any training sample in a sample set, a first reasoning result of a first model on the training sample and a second reasoning result of a second model on the training sample are obtained; wherein the first model is obtained by training a sample set, the sample set comprises a clean sample set and other sample sets except the clean sample set, the second model is obtained by training the clean sample set, and the clean sample set comprises training samples of which labels are considered to be correct; and determining the sample category of the training sample according to the first reasoning result, the second reasoning result and the label of the training sample.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

A mental illness recognition system based on visual sensor collected optical flow features

The mental illness recognition system based on the optical flow features collected by a visual sensor comprises a mental illness expert consultation video data preprocessing module, a facial stress unit extraction module, an optical flow change feature unit construction module, a classification model construction module and a patient mental illness category recognition module connected in sequence. The mental illness expert consultation video data preprocessing module feeds the patient facial picture to the facial stress unit extraction module. The facial stress unit extraction module feeds the optical flow calculation method and the patient facial picture sequence to the optical flow change feature unit construction module. The optical flow change feature unit construction module feeds the optical flow change feature unit to the classification model construction module and the patient mental illness category recognition module respectively. The classification model construction module feeds the classification model to the patient mental illness category recognition module. The mental illness and normal sample classification recognition is realized under the condition that the user facial video has good clarity.
Owner:ZHEJIANG UNIV OF TECH

Small sample motion radiation source individual identification method and system based on migration network

The invention relates to a migration network-based small sample motion radiation source individual identification method and system, belongs to the technical field of communication radiation source individual identification, and solves the problem of poor identification effect caused by limited motion radiation data in the prior art. The method comprises the following steps: constructing a radiation source individual identification model, wherein the radiation source individual identification model comprises a feature extraction network and a sample classification layer; pre-training the radiation source individual identification model by adopting static radiation source signal sample data; for the pre-trained radiation source individual identification model, retaining model parameters of the feature extraction network, and initializing model parameters of a sample classification layer; training the migrated radiation source individual identification model by adopting the moving radiation source signal sample data; and inputting a to-be-identified motion radiation source signal into the trained radiation source individual identification model for identification. According to the invention, the identification effect of motion radiation source individual identification under the small sample condition is improved.
Owner:36TH RES INST OF CETC

Methods, apparatus, media, devices, and products for classifying wine samples

Embodiments of the present application provide a method, device, medium, equipment and product for classifying wine samples, relating to the technical field of artificial intelligence. The method comprises: contacting odor molecules of a wine sample with at least one recombinant cell expressing an olfactory receptor and a reporter protein, the reporter protein generating a detectable signal after the receptor binds to the odor molecules; obtaining time series data of the signal, calculating the maximum value and baseline value thereof, and determining a response characteristic value of each olfactory receptor according to the same; and inputting the characteristic value into a trained machine learning model to obtain a wine sample classification result. The present method accurately extracts a stable characteristic value of the response of an olfactory receptor to a wine sample, thereby eliminating interference caused by initial value fluctuations and improving the accuracy of model classification.
Owner:HANVON CORP

Postoperative intelligent follow-up visit monitoring system for cataract

The invention provides an intelligent follow-up visit monitoring system after cataract operation, and relates to the technical field of image processing. The intelligent follow-up visit monitoring system after the cataract operation comprises an acquisition module used for acquiring a plurality of image samples after the cataract operation; the classification module is used for inputting the plurality of image samples into a pre-trained image classification model and outputting probability values of classification results of various complications after the cataract operation corresponding to the image samples; the determining module is used for determining the uncertain value of each image sample according to the probability value of the classification result of various complications after the cataract operation corresponding to each image sample; the screening module is used for screening out a target image sample according to the uncertain value of each image sample; and the training module is used for training an image classification model according to the target image sample and the labeling result of the target image sample, and the trained image classification model is used for assisting in determining a post-cataract follow-up visit scheme. The method can improve the follow-up visit efficiency after the cataract operation.
Owner:NING BO EYE HOSPITAL

Three-dimensional point cloud zero-sample classification method and device based on mutex cooperation network

The application discloses a three-dimensional point cloud zero sample classification method and device based on a mutex cooperation network, which comprises the following steps: data set preprocessing, dividing the data set into visible category and invisible category samples, and obtaining semantic feature vectors of each category; constructing a mutex cooperation neural network, including a point cloud encoder, a mutex learning module, a double-branch learning module and a classification module; in the training stage, a mutex loss, a classification loss and a regularization loss function are used to update the gradient propagation, and the network is trained in an end-to-end manner; in the test stage, the double-branch features are fused, and are matched with the category semantic feature vectors to obtain a classification result. The application learns 3D features divergently through mutex learning, so that more discriminative parts of an object can be found clearly, and in order to enhance the learning features of these different parts of the object, further design is made through the utilization of a graph to propagate cooperation information between activation regions, so that the performance and accuracy of three-dimensional point cloud zero sample classification are improved.
Owner:SOUTH CHINA UNIV OF TECH

Enterprise-level digital clone running method and device based on personalized increment, equipment and medium

The application discloses an enterprise-level digital clone running method and device based on personalized increment, equipment and medium, relates to the technical field of artificial intelligence, including: under the condition of meeting the preset authorization, preprocessing the original file of the target personnel, and classifying the sample; the preset authorization condition is that the authorization of the target personnel for the distillation operation has been previously obtained; according to the classification result, the working skills and personality style of the target personnel are extracted and sorted respectively to obtain the working skills draft and personality style draft; the draft is checked and the personal increment package draft to be audited is generated; the audit result is obtained by feeding back the personal increment package draft, if the audit is passed, the target personal increment package is constructed and stored based on the personal increment package draft, so that when the enterprise-level digital clone system runs, the runtime control package is constructed by using the target personal increment package, the preset system public capability package and the preset business field capability package, and the enterprise business service is provided.
Owner:FOUNDER SECURITIES CO LTD

X-ray single projection imaging wood type intelligent identification method and system and medium

The invention provides an X-ray single projection imaging wood type intelligent identification method and system and a medium, and the method comprises the steps: carrying out the multi-time translation through an X-ray phase contrast imaging system, and obtaining single projection images of different positions of a wood sample, and taking the single projection images as an image data set; preprocessing the single projection image to obtain a preprocessed image data set; training a deep learning model by using the preprocessed image data set; the deep learning model is an improved OfficientNetV2 model, and the depth learning model is an improved EfficientNetV2 model; and obtaining a to-be-tested sample, shooting a single projection image, preprocessing the single projection image, and inputting the preprocessed image into the trained deep learning model for sample classification. According to the invention, based on phase contrast X-ray single projection imaging, through improving a deep learning network model, high-accuracy and high-efficiency intelligent identification of a wood sample is realized.
Owner:SHANGHAI ADVANCED RES INST CHINESE ACADEMY OF SCI

A digital image encoding method based on metabolomics mass spectrometry data

The application discloses a digital image coding method based on metabolomics mass spectrum data. The method groups the mass-to-charge ratios in a preset mass range in the obtained first liquid chromatography-tandem mass spectrometry data according to preset division conditions; generates a full metabolome profile image according to the first mass-to-charge ratios and the first scan indexes in each group after grouping, and cuts and stacks the image to obtain a first multi-channel image; and screens a first target block based on the pooling signal intensity and the image entropy corresponding to the first multi-channel image, and the second multi-channel image stacked by the first target block is used to train a deep learning model for classifying biological samples. Thus, the metabolomics mass spectrum data can be converted into an image thereof through the digital image coding mode, the image can accurately retain the mass spectrum structure of the mass spectrum data of the metabolite species and level in the analytical metabolomics, and the deep learning model trained thereby can comprehensively reflect the original state of the metabolomics in LC-MS.
Owner:SHANGHAI INST OF ORGANIC CHEM CHINESE ACAD OF SCI

Loop filtering method based on adaptive pixel classification criteria

ActiveCN116233422BLoop filterComputer vision
Disclosed is a loop filter method based on an adaptive pixel classification criterion. The loop filter method based on an adaptive pixel classification criterion in an image decoding device includes a step of classifying a restored sample according to an absolute classification criterion or a relative classification criterion, a step of obtaining offset information according to a result of classifying the restored sample, a step of adding the offset value to the restored sample with reference to the obtained offset information, and a step of outputting the restored sample to which the offset value is added. Thus, errors of the restored sample can be corrected.
Owner:INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY

Sequence detection device and sequence detection method

PendingCN122053300ATransmitter/receiver shaping networksAlgorithmFeedforward filter
The invention provides a sequence detection device which comprises a feed-forward filter and a sequence detection circuit. The feed-forward filter is used for processing a received signal to generate an equalized signal. The sequence detection circuit is used for carrying out sequence detection on the equalization signal to generate and output a symbol sequence. The sequence detection circuit comprises a region estimation circuit and a grid selection circuit. The region estimation circuit is configured to classify each of a plurality of samples included in the equalized signal into one of a plurality of regions. The grid selection circuit is used for selecting one of a plurality of grid schemes to perform branch metric calculation according to the area estimation results of two of the plurality of samples output by the area estimation circuit.
Owner:AIROHA TECHNOLOGY CORPORATION

Multi-modal colony sample classification and identification method and system based on diffusion model

The application discloses a kind of multi-modal bacterial colony sample autonomous learning classification identification method and system based on diffusion model, it is related to computer vision technical field, by being input into multi-modal bacterial colony sample classification model to multi-modal bacterial colony sample, obtain classification result;First, the application is by the feature fusion of multi-modal sample input, the correlation between different modalities is in-depth mined, and the semantic information of multi-modal is fully utilized.Secondly, the application combines diffusion model, and the multi-modal bacterial colony sample is classified after being denoised using diffusion model, which is conducive to eliminating inherent noise in multi-modal data samples and interference introduced by human operation errors, and is more conducive to learning.Finally, the application combines autonomous learning, and the model can complete the automatic labeling of new multi-modal samples, can select the most suitable multi-modal sample for training and learning, and continuously improve the robustness of multi-modal bacterial colony sample classification model.
Owner:SICHUAN UNIV

Multimodal context selection for large language model based resolutions addressing technical issues

A method for technical issue resolution. The method includes: receiving, from a user, a text query concerning a technical issue; obtaining query-related context relevant to the text query; and processing, through a large language model (LLM), the text query and the query-related context to produce a multimodal query response used by the user to address the technical issue. More specifically, embodiments described herein utilize text topic and zero shot classification models to translate multimodal technical documentation (e.g., including text and images) into topic relevant metadata; and process queries, pertaining to technical issues, using a multimodal LLM provided with query-related text and image context derived from said topic relevant metadata.
Owner:DELL PROD LP

Classification method of fully strong weathered granite

The invention relates to a classification method of fully strong weathered granite. The method is suitable for the technical fields of solid waste utilization and ore processing. The technical problem to be solved by the invention is to provide the classification method for the fully strong weathered granite. According to the technical scheme, the classification method for the fully strong weathered granite comprises the steps that the washing difficulty index of a granite sample is determined based on the clay mineral content and the elutriation rate of the granite sample; determining a harmful substance content index of the granite sample based on the harmful substance content of the granite sample; determining the product quality index of the granite sample based on the crushing values of the sand samples of multiple particle sizes screened from the granite sample; based on the washing difficulty index, the harmful substance content index and the product quality index of the granite sample, classifying the granite sample in combination with a preset sample classification rule; and based on granite sample classification, in combination with a preset classification process suggestion library, determining a process suggestion for the granite sample.
Owner:POWERCHINA HUADONG ENG CORP LTD +1

Plant image classification method

The invention relates to the field of computer vision and image processing, discloses a plant image classification method, and solves the problems that the existing plant image classification technology is poor in generalization ability and low in classification precision, and lacks self-diagnosis classification uncertainty and a continuous optimization mechanism, so that difficult sample classification errors and repetition occur. According to the scheme, the method comprises the steps of obtaining a to-be-classified plant image and performing preprocessing; performing feature extraction on the plant image through a pre-trained feature extraction network to obtain a plurality of image features including different abstract levels; fusing the plurality of image features to obtain a fused feature, and according to the fused feature, generating a classification probability that the plant image belongs to each of the plurality of plant categories; determining an uncertainty metric value of the classification result according to the classification probability; and if the uncertainty metric value satisfies a preset condition, sending the plant image to a labeling terminal for requesting manual labeling of a label, and updating the feature extraction network according to the received manual labeling label and the plant image.
Owner:中国雅江集团有限公司 +1

Sample classification device for water quality component detection

The utility model relates to the technical field of water quality sample classification, and discloses a sample classification device for water quality component detection, which comprises a box body, the top of the box body is rotatably connected with a box cover, an inner container is fixedly connected in the box body, a partition plate is arranged in the inner container, one side of the partition plate is fixedly connected with a support, and the support is fixedly connected with a water inlet. A fixing assembly is arranged on one side of the support. The fixing assembly comprises a shell, one end of the shell is fixedly connected to one side of the support, and a connecting column is slidably connected into the shell. According to the water quality sample fixing device, a water quality sample is placed between the clamping plates, after the connecting plates are loosened, the connecting plates are driven to reset through the connecting columns under tension supporting of the first springs, then the clamping plates are driven to reset to be attached to the outer wall of the sample, and the effects that the sample is conveniently fixed and better attached to a bottle body are achieved; the problem that a water quality sample is directly placed in the box body of the sample classification device and easily topples over during carrying is solved, and the sample storage stability of the sample classification device is improved.
Owner:HEBEI RES INST OF INVESTIGATION & DESIGN OF WATER CONSERVANCY & HYDROPOWER

Multifunctional spectrograph sample classified storage device

The utility model discloses a multifunctional spectrograph sample classified storage device, which relates to the technical field of sample storage and comprises a shell, a storage component is arranged above the shell, a sponge sleeve is fixedly mounted on the inner surface of the shell, and a partition plate is fixedly mounted on the upper surface of an inner cavity of the shell. A plurality of vertical plates are integrally formed on the front side and the rear side of the partition plate, a plurality of fixing assemblies are slidably connected to the front side and the rear side of the upper surface of the shell, and a box cover is detachably installed on the upper surface of the shell. The interior of the shell is divided into a plurality of independent storage spaces through the partition plates and the vertical plates, the purpose of storing spectrograph samples in a classified mode is achieved, the samples are marked through the label plate, the needed samples can be conveniently and rapidly found, the samples can be clamped and fixed through the electric push rod, the first T-shaped plate and the clamping plate, and the sample storage box is convenient to use. It is ensured that the samples cannot move or incline in the storage or moving process, meanwhile, the size of the storage space can be adjusted, and storage of the samples is facilitated.
Owner:BEIJING JINGKE RUIDA TECH CO LTD

Sample classification method and device for first arrival wave deep learning

The application provides a sample classification method and device for first arrival wave deep learning, and relates to the technical field of oil exploration. The method comprises the following steps: acquiring training sample data and classification information to which the training sample data belongs; the training sample data comprises first arrival sample data and first arrival label data corresponding to the first arrival sample data; a classification network model is trained by using the training sample data and the classification information to obtain a training result; and the training result is used for classification prediction processing on to-be-classified first arrival sample data to obtain a first arrival sample classification result. The application can accurately and efficiently classify and screen the first arrival sample data, enriches the type feature information of the first arrival sample data, and improves the pertinence and effectiveness of first arrival wave sample deep learning.
Owner:CHINA NAT PETROLEUM CORP +1

Zero-shot video classification method based on test time visual proxy tuning

The application provides a zero-shot video classification method based on test time visual proxy tuning, which realizes zero-shot action recognition of video actions by constructing a visual proxy using a support set and simultaneously fine-tuning visual and text prompts, and avoids the semantic gap problem between the two modalities of video and text in the video classification task. The visual proxy construction module and the dual-modal prompt collaborative tuning module constitute a zero-shot learning framework TPC. In the visual proxy construction module, a pre-trained video encoder is used to extract support set video features to construct a visual proxy, and a learnable visual prompt is added to the sampled support set video to make the visual proxy adjustable. In the dual-modal prompt collaborative tuning module, the learnable visual and text prompts are fine-tuned by minimizing the KL divergence between the prediction probability distribution of the visual proxy and the text proxy, the information of the text and visual modalities is used to optimize the visual proxy, and the zero-shot classification performance of the visual proxy is improved.
Owner:NANJING UNIV OF SCI & TECH