Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

80 results about "Triplet loss" patented technology

Triplet loss is a loss function for artificial neural networks where a baseline (anchor) input is compared to a positive (truthy) input and a negative (falsy) input. The distance from the baseline (anchor) input to the positive (truthy) input is minimized, and the distance from the baseline (anchor) input to the negative (falsy) input is maximized.

Water turbine fault classification diagnosis method based on multi-modal fusion and meta learning

The invention discloses a water turbine fault classification diagnosis method based on multi-modal fusion and meta-learning. The method comprises the following steps: step 1, respectively extracting time domain features and Mel-language spectrogram features of monitoring noise signals of a water turbine through a feature extraction module; step 2, performing cross-modal attention mechanism fusion on the time domain features and the voiceprint features through a multi-modal fusion module, and adjusting a fusion weight based on a dynamic weight distribution mechanism; and step 3, performing small sample training optimization on the fused features through a meta-learning module, and improving the classification capability of the model for new fault types in combination with a Triplet Loss-KNN algorithm and twin network pre-training. The fault diagnosis method based on the time domain-voiceprint fusion network and meta learning has the advantages of being rapid in diagnosis, accurate in classification, high in generalization ability and the like, and the fault diagnosis precision of the water turbine based on noise signals can be effectively improved.
Owner:CHINA THREE GORGES UNIV

Aero-engine bearing fault diagnosis method based on clustering enhancement domain generalization

The invention belongs to the technical field of intelligent fault diagnosis, and discloses an aero-engine bearing fault diagnosis method based on clustering enhancement domain generalization. Firstly, a time domain vibration signal is converted into frequency domain representation through fast Fourier transform, a two-stage convolutional neural network architecture comprising a domain alignment encoder and a classification encoder is constructed, and cross-working-condition feature alignment under adversarial training is realized by using a domain discriminator; introducing a multi-source domain maximum mean difference statistical alignment strategy and a clustering enhanced triple loss mechanism, and synchronously optimizing intra-class feature compactness and inter-class separability through pseudo label clustering center anchoring and dynamic interval constraint; and gradient inversion adversarial training and multi-modal feature fusion are combined. The method provided by the invention can effectively improve the accuracy and robustness of generalization judgment when the aero-engine bearing fault diagnosis lacks target domain data.
Owner:DALIAN UNIV OF TECH

Small sample cross-domain fault diagnosis method based on multistage feature alignment

The invention provides a small sample cross-domain fault diagnosis method based on multistage feature alignment, and the method comprises the steps: collecting rotating machinery monitoring data in a source domain and a target domain, and constructing a cross-domain small sample fault diagnosis task data set with scarce target domain samples; constructing a diffusion model combining a multi-head self-attention mechanism and a normalization strategy, and generating an enhanced sample with target domain feature distribution; a shared feature extraction network fusing a source domain, a target domain and a generated sample is built, and unified modeling of diagnosis features is achieved; and designing a joint loss function which comprises source domain classification loss, dual-MMD distribution alignment loss, cross-domain triple loss and enhanced consistency regular terms, constructing a multi-stage feature alignment mechanism, and realizing fault feature migration and accurate diagnosis under the condition of small samples of a target domain. According to the method, source domain diagnosis knowledge can be effectively migrated under the condition of target domain data scarcity, the fault feature discrimination capability and the cross-domain adaptability of the model are enhanced, and the method is suitable for intelligent diagnosis tasks of the rotating machinery under complex working conditions.
Owner:NORTHEASTERN UNIV CHINA

Cross-view-angle image geographic positioning method based on dynamic threshold value pseudo label self-training learning

The invention discloses a cross-view image geographic positioning method based on dynamic threshold pseudo tag self-training learning, and the method specifically comprises the following steps: introducing a difficult sample feature mining method, dynamically adjusting the loss weight of a sample according to the change of similarity, and building a dynamic difficult sample triple loss model; the method comprises the following steps: dynamically adjusting a confidence threshold value of a sample by adopting an index moving average weighting method, iteratively training and screening an unlabeled sample, namely a pseudo label, establishing a pseudo label self-training mechanism of a dynamic threshold value, mining and utilizing non-paired data, and solving the problem of high manual labeling cost; a reference image most similar to a query image is found through image retrieval, and the offset of a query position is predicted. Experiments on CVUSA and CVACT data sets show that as the distance threshold increases, the accuracy of the cross-view image geographic positioning method based on dynamic threshold pseudo tag self-training learning presents a stable rising trend, and the cross-view image geographic positioning method based on dynamic threshold pseudo tag self-training learning is superior to other methods under the same threshold condition.
Owner:HENAN UNIVERSITY

Retrieval enhancement generation method and system in dual-carbon field

The invention provides a retrieval enhancement generation method and system in the dual-carbon field, and relates to the field of data processing. According to the method, multi-source unstructured data in the dual-carbon field is collected, after data preprocessing is carried out, a multi-granularity query problem set is formed, a dual-carbon field knowledge base is obtained, and a dual-carbon field-oriented embedding model CEMBING and a reordering model CReranker are constructed. The CEMBING model adopts a semantic partitioning method and a joint training strategy, so that semantic information of a dual-carbon field text can be effectively captured; the CReranker model adopts a negative example mining strategy and a triple loss function, candidate documents can be accurately sorted, the problems of knowledge limitation and insufficient timeliness of LLMs in the application of the dialogue system in the dual-carbon field are effectively solved, and the accuracy and efficiency of retrieval enhancement generation of the dialogue system in the dual-carbon field are improved.
Owner:CHINA THREE GORGES UNIV

Text washing detection method based on deep learning

The invention discloses a text washing detection method based on deep learning. The method comprises the following steps: S1, carrying out feature coding on an input text through a BERT model; s2, grouping and fusing the features of the draft washing text extracted in the above step to obtain higher-quality feature representation of the draft washing text; s3, performing multi-statement feature fusion on the text features in the above step, integrating multi-view context information, and capturing deep text semantic features; s4, optimizing the neural network model through triple loss and a positive example enhancement loss function; and S5, designing and constructing a manuscript washing detection data set through a data set generation module based on a comparative learning thought. According to the method disclosed by the invention, the method has good robustness on retouching and multi-round translation manuscript washing, while the text features are efficiently extracted, the text manuscript washing behavior is accurately detected, and the method is suitable for the field of text manuscript washing detection and copyright protection.
Owner:UNIV OF SHANGHAI FOR SCI & TECH

Cross-modal hash retrieval model training method and device and cross-modal hash retrieval method and device

The invention relates to the technical field of information retrieval, in particular to a cross-modal hash retrieval model training method and device and a retrieval method and device. In the method, adaptive gradient triple loss can distribute adaptive gradients for triples with different difficulties by introducing a constraint term based on an included angle between an anchor point and a negative sample, so as to maintain the consistency of a neighborhood relationship in an original space, promote intra-class compactness and inter-class separability of a heterogeneous mode, and improve the robustness of the heterogeneous mode. Samples which are semantically similar to query samples can be retrieved; the step-by-step quantization loss decouples representation learning and binary code generation, the two-stage design retains a semantic structure in an embedding stage, and the subsequent quantization loss is minimized; therefore, the cross-modal hash retrieval model has good retrieval performance.
Owner:WEIFANG UNIVERSITY

Motor bearing fault detection system and method based on robust deep learning

The invention discloses a motor bearing fault detection system and method based on robust deep learning, and belongs to the technical field of mechanical fault detection and intelligent perception. Feature extraction is carried out on an original vibration signal with a label based on a supervised learning branch network, and the original vibration signal is used as a reference sample; the samples with the same fault category and different fault categories as the reference samples are positive samples and negative samples, and inter-class separation and intra-class aggregation relations in a triple loss optimization embedding feature space are introduced to generate embedding representation; based on an unsupervised learning branch network, encoding the original vibration signal after time domain and frequency domain artificial feature extraction, and introducing triple loss to carry out unsupervised embedding learning to generate high-level feature embedding representation; and the embedded representations output by the two branch networks are fused, dual loss of triple loss and center loss is introduced for training, and a bearing fault detection model after training is completed is used for bearing fault detection.
Owner:ZHEJIANG GONGSHANG UNIVERSITY

Satellite signal authentication method and device based on complex valued neural network, and storage medium

The invention discloses a satellite signal authentication method and device based on a complex valued neural network and a storage medium. The satellite signal authentication method comprises the following steps: acquiring a complex valued data sequence of a satellite signal to be authenticated; inputting the complex-valued data sequence of the to-be-authenticated satellite signal into a trained complex-valued neural network model to enable the trained complex-valued neural network model to output the radio frequency feature vector of the to-be-authenticated satellite signal, the trained complex-valued neural network model being obtained by training based on a preset triple loss function and a back propagation algorithm; calculating a target similarity score between the radio frequency feature vector of the satellite signal to be authenticated and the corresponding target anchor point sample, wherein the score is used for representing an average angular distance between the radio frequency feature vector and the target anchor point sample; and classifying and authenticating the satellite signal to be authenticated based on a preset score threshold and the target similarity score. According to the method, the recognition complexity can be reduced while the recognition precision of the satellite signal is improved.
Owner:XIDIAN UNIV

Ballistocardiogram signal data enhancement method and system based on BiLSTM-GAN

The invention belongs to the technical field of medical signal processing and deep learning, and particularly relates to a ballistocardiogram signal data enhancement method and system based on BiLSTM-GAN, and the method comprises the steps: S1, carrying out the Butterworth band-pass filtering of a preset frequency band on an original ballistocardiogram signal, and outputting a first signal; s2, performing down-sampling processing on the first signal to a preset sampling rate, and outputting a second signal; s3, segmenting the second signal into segments with a preset length, labeling categories, and outputting BCG signal segments subjected to band-pass filtering; s4, constructing a BiLSTM-GAN model on the basis of the BCG signal fragments subjected to the band-pass filtering; s5, taking the BCG signal segments subjected to band-pass filtering as training data, and training the BiLSTM-GAN model based on a triple loss function including reconstruction loss, supervision loss and adversarial loss; s6, generating an enhanced BCG signal fragment; and S7, generating equivalent enhanced BCG signal fragments, and outputting a balanced data set. The technical problem of sample scarcity caused by class imbalance in medical BCG signals can be solved.
Owner:GUANGZHOU INST OF RAILWAY TECH

System and methods for classification of image data from synthetic aperture radar images and electro-optical images

Systems and methods are disclosed for image classification of electro-optical images and synthetic aperture radar images using training techniques that can include appearance labeling and triplet mining to train a neural network system. The training data can include image pairs of electro-optical images and synthetic radar aperture images. The training data can include anchor, positive, and negative images. The neural network can be trained using triplet loss and cross-entropy loss. The trained neural network can be used for object classification such as automatic target recognition of aerial images.
Owner:ATOMBEAM TECH INC

A zero-shot sketch retrieval method based on mask and matching

The present invention belongs to the technical field of zero-shot sketch retrieval, and discloses a zero-shot sketch retrieval method based on mask and matching. First, a visual-cross-language sampler is designed. This module uses semantic labels to generate masks, shielding semantic information in the image that is irrelevant to the sketch target, and optimizing the semantic alignment effect of cross-domain matching. Then, a purification mask matching module is designed, which includes two parts: feature reconstruction and semantic interaction. It suppresses redundant semantics by forcing the image encoder to reconstruct masked features, and uses a transformer decoder to promote cross-domain interaction between sketch and image features, thereby achieving more refined semantic matching. Finally, the training mechanism of triplet loss, reconstruction loss and interaction loss is combined to enable the model to significantly improve the retrieval accuracy in the purified semantic space. This method masks the interference elements in the image through masks and achieves pure sketch-image matching, thereby effectively solving the problem of semantic differences between sketches and natural images.
Owner:ROBOTICS RESEARCH CENTER OF YUYAO CITY +1

Rolling bearing fault diagnosis method and system based on comparison decoupling single source domain generalization

The invention discloses a rolling bearing fault diagnosis method and system based on comparison decoupling single-source domain generalization, relates to the field of rolling bearing fault diagnosis, and aims to solve the problem that in the prior art, a method for effectively solving the problem that the model generalization performance is reduced due to insufficient diversity of single-source data in bearing fault diagnosis does not exist. The method is technically characterized by comprising the following steps: step 1, adopting rolling bearing time domain data, and performing two data enhancement processing on the rolling bearing time domain data to obtain two groups of enhanced data; step 2, the model construction is based on a causal decoupling network and a comparative learning BYOL framework, the causal decoupling network is used for extracting causal features and non-causal features of data, and the rank of the causal features is constrained through adaptive threshold weighted nuclear norm regularization; the contrast learning BYOL framework is used for pulling in causal features extracted from the two groups of enhanced data to obtain causal domain invariant features of each sample; step 3, using a prototype Gaussian triple loss function to constrain causal domain invariant features, and improving intra-class compactness and inter-class separability of training samples; and step 4, finally inputting a target domain sample to the trained comparison decoupling network model to realize diagnosis.
Owner:HARBIN UNIV OF SCI & TECH

Medical image classification method and equipment based on improved triple loss

The invention discloses a medical image classification method and equipment based on improved triple loss, and the method comprises the steps: constructing a medical image classification model based on a neural network and a triple loss function, introducing the angle loss of a gradient direction into the triple loss function, enabling the loss function to reduce the deviation of the gradient direction when the gradient direction is calculated, and enabling the calculation precision of the gradient direction to be improved. Convergence of the training process is accelerated. According to the method, the training efficiency can be improved, by introducing the negative group distance, the angle loss and the adaptive Margin calculation strategy, the improved triple loss function converges faster in the training process, and the model training efficiency and stability are improved. According to the invention, the classification accuracy can be improved, and different types of medical images can be identified more accurately.
Owner:TIANJIN UNIV

Natural language processing systems and methods for intent classification of speech transcription

Aspects of the subject disclosure may include, for example, generating a natural language processing model by training an automatic speech recognition (ASR) encoder with manual transcription. The training is performed by correcting and adjusting relevant factors of the ASR encoder based on determined triplet loss, classification loss and Kullback-Leibler divergence loss. In response to an ASR utterance, the trained natural language processing model generates a predicted intent associated with the ASR utterance with improved accuracy. Other embodiments are disclosed.
Owner:JPMORGAN CHASE BANK NA

An EEG emotional state classification method based on multi-source domain adaptation of knowledge distillation

The application discloses a multi-source domain adaptation EEG emotion state classification method based on knowledge distillation. First, data is acquired for band pass filtering, and independent component analysis technology is used to remove artifacts. Second, the electroencephalogram feature is extracted through the differential entropy method, and the three-dimensional electroencephalogram time sequence is converted into a two-dimensional sample matrix. Then, the training set and the test set are respectively demarcated under two task scenarios, and it is ensured that they do not coincide. The application adopts the pseudo-label triplet loss based on marginal sampling combined with the maximum mean difference. The application learns knowledge from different source domains to maximize the use of multiple single-source models and realize a more powerful model with less time consumption. Finally, the classification accuracy is used to evaluate the performance of the model under two task scenarios. The application combines the triplet loss and the maximum mean difference, which can not only realize unbiased alignment between each pair of source domains and the target domain at the domain level, but also consider the correlation at the data pair level.
Owner:HANGZHOU DIANZI UNIV

A fault diagnosis method for polyester esterification stage

This invention discloses a fault diagnosis method for the polyester esterification stage. This method combines global supervised learning and contextual metric meta-learning. Using the attribute information of a single sample and similarity information from a sample group, it first uses variational modal decomposition to obtain multi-scale data through global supervised training. Multi-scale components with fault characteristics are extracted, and multi-scale feature fusion learning is performed. Triplet loss is used to learn finer, more subtle features. A fixed multi-scale feature fusion module is then used for task meta-learning training to learn a single feature, converting the raw data of the meta-task into a basic feature space. Finally, a dimensional variational prototype module is used to adaptively measure the feature similarity of sample pairs. The statistical method of variational inference automatically learns metric scaling parameters to transform the embedding space. The method is simple and solves the fault diagnosis problem in scenarios with limited data and a fully open set.
Owner:DONGHUA UNIV

A gait recognition method and system based on unsupervised domain adaptation

The present invention relates to a gait recognition method and system based on unsupervised domain adaptation. First, the original gait silhouette sequence is obtained. The unsupervised adaptive gait recognition network is used to assign pseudo labels to the unlabeled data, and the labeled source domain data and the unlabeled target domain data are sampled as input data respectively. The backbone network is used to extract the input sequence features to obtain the fine-grained spatiotemporal features of the source domain and the target domain. For each local feature, the global environment aggregation is used to establish the association between each feature in the time dimension, and the global feature is extracted. The global motion pattern of each part is constrained by the triplet loss, and the triplet loss is calculated by weighted summation. Each normalized feature is used to initialize and update the hybrid memory unit, and supervision is provided for the unlabeled data during the training process. The present invention can learn knowledge from the labeled source domain data and transfer it to the target domain, effectively realizing high-precision gait recognition on unlabeled data.
Owner:BEIJING INST OF TECH

Ceramic authenticity detection method based on deep learning

The invention relates to the technical field of ceramic detection, in particular to a ceramic authenticity detection method and system based on deep learning, and the method comprises the steps: S1, obtaining and preprocessing a ceramic image; s2, constructing triple training data according to the preprocessed ceramic image; s3, constructing a ceramic authenticity detection model based on a deep metric learning framework, inputting the triple training data into the ceramic authenticity detection model, and performing optimization training on the ceramic authenticity detection model by using a triple loss function; and S4, evaluating the trained ceramic authenticity detection model to obtain a trained ceramic authenticity detection model. According to the invention, the trained ceramic authenticity detection model is constructed to realize automatic, high-efficiency and high-precision detection of ceramic authenticity.
Owner:HUBEI POST TELECOMM PLANNING DESIGN

A Small Sample Fault Diagnosis Method Based on Multi-Scale Feature Learning and Domain Adaptive Optimization

This invention relates to a small-sample fault diagnosis method based on multi-scale feature learning and domain adaptive optimization, comprising: acquiring historical heterogeneous fault data of several mechanical devices as a source domain dataset, acquiring historical heterogeneous fault data of a target mechanical device as a target domain dataset, and defining label spaces for the source and target domain datasets respectively; randomly sampling and combining the source and target domain datasets to construct a triplet dataset and a domain adaptive dataset, inputting them into a dual-branch feature extraction subnetwork, measuring the triplet feature embedding distance, and iteratively training using a joint loss function composed of triplet loss and domain adaptive loss to obtain a dual-branch fault diagnosis model; collecting current heterogeneous fault data of the target mechanical device and inputting it into the dual-branch fault diagnosis model, outputting the fault classification of the vibration signal of the target mechanical device. This invention can solve the problem of model performance degradation caused by small sample data conditions and differences in data distribution.
Owner:GUANGDONG UNIV OF TECH

A pedestrian re-identification network based on high-altitude visual field

The present invention discloses a high-altitude visual field pedestrian re-identification network, characterized by a multi-branch network including global branches and local branches; the global branch constructs a feature pyramid based on the outputs of the four stages {R2, R3, R4, R5} of ResNet50 to obtain feature semantic information in each stage of ResNet50; the local branch evenly divides the feature vector output by Conv4_1 of the R4 stage into five parts, and adopts different combination strategies to construct feature extraction branches for these five parts to learn local semantic features of different coarse and fine granularity. The present invention uses cross entropy and triplet loss functions to jointly optimize the multi-branch network, combining the advantages of both, while ensuring the network feature extraction effect, reducing the intra-class spacing and improving the inter-class spacing.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Model optimization method, electronic equipment and computer readable storage medium

The invention discloses a model optimization method, electronic equipment and a computer readable storage medium, and relates to the technical field of knock detection. The method comprises the following steps: acquiring a knocking signal sample set which comprises a plurality of anchor point samples, a plurality of positive class samples and a plurality of negative class samples; according to the envelope curve of each anchor point sample and each negative sample, determining the anchor point similarity of each negative sample, and based on the anchor point similarity of each negative sample, determining the triple loss boundary of each negative sample, the larger the anchor point similarity is, the larger the triple loss boundary is; inputting the knocking signal sample set into a to-be-optimized feature extraction model to obtain feature vectors of each anchor point sample, each positive class sample and each negative class sample; and optimizing a feature extraction model based on the feature vector of each anchor point sample, the feature vector of each positive class sample, the feature vector of each negative class sample and the triple loss boundary. According to the method, the distinguishing capability of a knocking detection algorithm on high-similarity interference signals can be improved in a complex real environment.
Owner:GOERTEK MICROELECTRONICS CO LTD

A method for boundary refinement of hand-woven cashmere images based on weakly supervised learning and multi-stage segmentation

The present invention discloses a method for fine-tuning the boundaries of hand-woven cashmere images based on weakly supervised learning and multi-stage segmentation. First, a coarse-grained trimap map is obtained through a basic weakly supervised teacher-student network, and a cashmere dense alpha map, edge area and sparse area are obtained through segmentation. Then, the cashmere edge area is expanded using a global cold loss and local ant loss algorithm to obtain the expanded edge area, and the expanded area is input into a collaborative parallel multi-resolution convolutional network for fine cashmere edge segmentation. A lightweight triple loss network is used to segment the sparse area to obtain a coarse-grained cashmere sparse area alpha map. Finally, through continuity judgment, a fine cashmere sparse area alpha map is synthesized, and the three alpha maps are combined for batch normalization and splicing to output the final overall alpha map. This method can effectively improve the accuracy of cashmere area segmentation, especially the robustness of fine-grained and edge processing, solves the problem that the length of existing hand-woven cashmere can only be detected manually, and reduces the annotation cost of pixel-level images.
Owner:INNER MONGOLIA UNIV OF TECH

Pedestrian re-identification method based on prototype guided knowledge propagation and adaptive learning

The invention discloses a pedestrian re-identification method, system and product based on prototype guided knowledge propagation and adaptive learning, belongs to the technical field of image retrieval in computer vision, and adopts a non-sample method based on prototype guided knowledge propagation and adaptive learning to carry out lifelong pedestrian re-identification. The core contribution of the method is that a prototype is constructed through triple loss constraints to relieve representation deviation of similar identity images in new and old tasks, and meanwhile, dynamic parameter selection and fusion strategy optimization model integration are designed, so that the model performance is remarkably improved on the premise of protecting privacy. Besides, according to the method, storage of historical samples is completely avoided, old knowledge is represented through a prototype, the privacy disclosure risk is fundamentally eliminated, an efficient and safe solution is provided for continuous learning in a dynamic network, and collaborative optimization of privacy protection and model performance is promoted. The effectiveness of the algorithm provided by the invention is proved on data sets such as Market1501, MSMT17, cuhksysu and the like.
Owner:WUHAN UNIV

Target detection countermeasure method

The application discloses a target detection anti-defense method, comprising the following steps: a multi-scale feature extraction module is used to obtain a detection feature representation of a to-be-detected image; a double-path attention mechanism is constructed to generate attention signals from an original image and the detection feature, and the comprehensive perception ability for spatial position disturbance is improved; a dynamic prediction module is constructed to perform weighted calculation on a plurality of convolution kernel parameters under the guidance of a spatial attention map, and different convolution parameters are adaptively allocated for different spatial positions; a network optimization mechanism based on a multi-loss function is constructed, wherein a dense triplet loss is used to ensure the discrimination ability of the spatial attention map for the local regions of two kinds of samples, and a target detection loss enables the model to have the prediction ability for target positions and categories; a target detection anti-defense model is trained, the target detection model is optimized in an end-to-end mode by using the network optimization mechanism of the multi-loss function, and target detection anti-defense is realized based on the trained target detection model.
Owner:TIANJIN UNIV

Software defect detection method based on Word2Vec and auto-encoder triple network

The invention provides a software defect detection method based on Word2Vec and an auto-encoder triple network, which belongs to the field of software defect detection and comprises the following steps: acquiring an original code snippet and adding a label; processing an original code snippet with a label by utilizing Word2Vec, and weighting to obtain an embedded vector identifier of the code snippet; learning features of normal codes and defect codes through an auto-encoder; and constructing sample data containing anchor point samples, positive samples and negative samples to train the triple network, and optimizing an embedding space through a triple loss function to obtain a defect detection model. And calculating an embedding vector and a classification distance of the new code snippets based on the defect detection model, and judging whether the code snippets are defect codes or not. While the complexity of feature engineering is reduced, the problems of data imbalance and no sample detection are effectively solved, and the method has wide practical application value.
Owner:HUBEI UNIV

Cross-modal-based medicine logistics retrieval method and system, terminal and storage medium

This invention relates to the field of logistics management technology, and discloses a cross-modal pharmaceutical logistics retrieval method, system, terminal, and storage medium. The method includes: based on an encoder, introducing a feature fusion mechanism to extract visual representations of textual information within the fused image, and achieving semantic enhancement through a cross-modal attention mechanism guided by a tag graph; designing a semantic neighborhood-aware contrastive hash code, a tag distribution-aware semantic alignment loss, and a cosine triplet loss to enhance the discriminative power of the hash code. This invention can improve the semantic alignment of multimodal features, enhance the model's generalization ability, and significantly improve the retrieval accuracy of pharmaceutical logistics.
Owner:GUANGDONG UNIV OF TECH

Multi-modal image-text retrieval method fusing fine-grained local semantics and global semantics

The invention discloses a multi-modal image-text retrieval method fusing fine-grained local semantics and global semantics, which comprises the following steps of: firstly, extracting original features of an image and a text, and performing regional relation reasoning to obtain relation-enhanced local features; and then semantic interaction is carried out on different samples in the same mode by utilizing an attention mechanism, so that the incidence relation between the samples in each mode is fully learned, and semantic enhanced image text embedding is obtained. And finally, training the whole model by adopting a triad loss function improved by triangular constraint. According to the cross-modal image-text retrieval method, semantic similarity and difference are fully mined, the capability of distinguishing semantic fuzzy samples by the model is enhanced, and the problems that in existing cross-modal image-text retrieval, the recognition accuracy of subtle differences in difficult-to-distinguish samples with subtle differences or semantic fuzzy is low and the like are solved.
Owner:HUNAN UNIV OF TECH

A small sample cross-domain fault diagnosis method based on multi-level feature alignment

The application provides a kind of small sample cross-domain fault diagnosis method based on multi-level feature alignment, comprising: collecting rotating machinery monitoring data under source domain and target domain, constructing cross-domain small sample fault diagnosis task data set of target domain sample scarcity;Diffusion model combined with multi-head self-attention mechanism and normalization strategy is constructed, and enhanced sample with target domain feature distribution is generated;Build a shared feature extraction network that integrates source domain, target domain and generated samples, and realize unified modeling of diagnostic features;Design a joint loss function, including source domain classification loss, double MMD distribution alignment loss, cross-domain triplet loss and enhanced consistency regularization term, construct a multi-level feature alignment mechanism, and realize fault feature transfer and accurate diagnosis under target domain small sample condition.The method of the application can effectively transfer the diagnostic knowledge of the source domain under the condition of the scarcity of data in the target domain, enhance the discriminability of the fault features and the cross-domain adaptability of the model, and is suitable for intelligent diagnosis tasks of rotating machinery under complex working conditions.
Owner:NORTHEASTERN UNIV CHINA

Verification code identification method and device based on large model, and related equipment

The embodiment of the invention discloses a verification code identification method and device based on a large model and related equipment. The method comprises the following steps: acquiring different types of sample verification codes; preprocessing different types of sample verification codes, and extracting multi-modal features of the sample verification codes; obtaining a verification strategy and an interference feature of each sample verification code, inputting each multi-modal feature and the corresponding verification strategy and interference feature into the initial multi-modal large model for identification training, and identifying the sample verification code based on an output result of the initial multi-modal large model and a corresponding real verification code type identifier. Carrying out loss calculation according to the triple loss function to obtain model loss, carrying out back propagation according to the model loss, and optimizing model parameters of the multi-modal large model to obtain an optimized multi-modal large model; and obtaining a target verification code needing to be identified at present, and outputting the target verification code to the multi-modal large model for identification to obtain an identification result. According to the method, the recognition accuracy and success rate of the cross-type verification code are improved.
Owner:BEIJING TAIXIN TIANCHENG TECHNOLOGY CO LTD