Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

152 results about "Text alignment" patented technology

Data analysis problem generation method based on image input and large model combination

The invention discloses a data analysis problem generation method based on image input and large model combination, and the method comprises the following steps: S1, obtaining original image data, and carrying out the preprocessing of the original image data; s2, constructing a visual language model based on a Qwen-VL architecture, and performing bidirectional alignment to generate a visual feature vector; s3, constructing a cue word template, and fusing the cue word template through a cross attention mechanism to generate structured text description; s4, constructing a large language model based on a Transform architecture, and performing supervision and fine tuning by adopting a LoRA method to generate a data analysis problem candidate sequence; s5, semantic consistency detection and structural rule matching are carried out, and sequences which do not meet semantic specifications or structural constraints are removed; and S6, constructing an image-text alignment triple, and writing the image-text alignment triple into the data structure in the JSON format for coding and storage. According to the method, the image can be converted into a data analysis problem, and the text generation quality and efficiency are remarkably improved.
Owner:HUAZHONG NORMAL UNIV

Cross-language text fusion intelligent alignment method and system

The invention relates to the technical field of cross-language information processing, and provides a cross-language text fusion intelligent alignment method and system.The cross-language text fusion intelligent alignment method comprises the steps that a hierarchical alignment model is constructed through preprocessing and label recognition of multi-coding-type texts, deep semantic feature extraction and labeling of a multi-language pre-training model and text semantic and format information analysis; transform is taken as a core, cross-language semantic association is enhanced through a multi-head attention mechanism, character-level and paragraph-level format collaboration is realized through label weight allocation and condition constraint, the format alignment accuracy is obviously improved in a multi-language mixed typesetting scene, and the multi-language mixed typesetting efficiency is improved. The analysis capability of the model on the structured information can be extended to layout relation processing of texts, images and tables, so that the comprehensive alignment efficiency in a multi-modal fusion scene is remarkably improved. The accuracy of text and label fusion is guaranteed, the problem of confusion of messy codes and labels is avoided, and high-precision cross-language text alignment from semantics to formats, from single mode to multiple modes and from semantics to formats is achieved.
Owner:SHANGHAI MEGALIN SOFTWARE TECH CO LTD

Text generation image diffusion model enhancement method based on multi-target preference optimization

The invention discloses a text generation image diffusion model enhancement method based on multi-target preference optimization. The method comprises the following steps: firstly, determining a plurality of reward models, and constructing a sample pair training set comprising positive and negative samples; and then, generating a loss weight of each sample pair in the sample pair training set, performing fine tuning training on the text map diffusion model by using the sample pair training set, and in the fine tuning training process, calculating a loss function value of each sample pair in combination with the loss weight of each sample pair until the training is completed, thereby obtaining an aligned text map diffusion model. According to the method provided by the invention, manual data annotation is not needed, the problems of preference inconsistency and over-optimization in a multi-reward scene are effectively solved, and the image quality, the text alignment capability and the multi-target optimization performance of the text-to-image generation model are remarkably improved. The method is superior to an existing optimization method under single-reward and multi-reward setting, shows higher generation quality and robustness, and can be seamlessly applied to various picture generation models.
Owner:ZHEJIANG UNIV

Landslide classification method and system based on visual language model and cross attention mechanism

The invention provides a landslide classification method and system based on a visual language model and a cross attention mechanism. The method and the system specifically comprise the following steps: data preprocessing: carrying out Canny edge detection on an RGB image, and calculating terrain attributes such as a gradient and a slope direction for a DEM (Digital Elevation Model); feature extraction: capturing local features by adopting a reflection filling convolution layer and multi-scale residual connection; a visual language model is introduced, wherein semantic enhancement features are extracted through image-text alignment by means of the visual language model; cross self-attention fusion: capturing a global context through self-attention, and focusing heterogenous data complementary information by cross attention; and classifying and outputting: outputting a result by using global average pooling and a linear classifier. Through the visual language model and the cross self-attention mechanism, the landslide recognition capability under the complex terrain is effectively improved, an efficient and reliable technical means is provided for geological disaster monitoring, and the method can be widely applied to the fields of landslide recognition, risk assessment and the like.
Owner:福州海洋研究院 +3

Intelligent decision-making method based on multi-modal large model and related equipment

The invention provides an intelligent decision-making method based on a multi-modal large model and related equipment, which are applied to an artificial intelligence technology, and the method comprises the steps: collecting a current field image of a to-be-decided region, and carrying out the preprocessing of the current field image, and obtaining a target field image; inputting the target field image into a pre-trained vision-text alignment model, and performing prediction by the vision-text alignment model by using the target field image to obtain a disease and pest diagnosis scheme corresponding to the to-be-decided region; dividing the to-be-decided region into a plurality of partitions, and processing current data of each mode in current multi-mode data of the partitions to obtain a space-time sequence of each mode; inputting the space-time sequence of each modal of the subarea into a pre-trained irrigation decision model, and performing prediction by the irrigation decision model according to the space-time sequence of each modal to obtain irrigation data of the subarea; the irrigation strategy of each subarea is generated according to the disease and pest diagnosis scheme and the irrigation data of each subarea. The cost can be reduced, and the generation efficiency and accuracy of the irrigation strategy can be improved.
Owner:INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD

Cross-modal image-text retrieval method and device based on pulse fusion

The invention discloses a cross-modal image-text retrieval method based on pulse fusion, and belongs to the technical field of image-text retrieval. The method comprises the following steps: inputting a target image set into a target detection network to obtain an image floating point code, and inputting a target text set into a word segmentation device to obtain a word floating point code; inputting the image floating point code and the word floating point code into a pulse encoder to obtain a first image pulse code and a first text pulse code; inputting the first image pulse code and the first text pulse code into a pulse cross attention fusion module to obtain a second image pulse code and a second text pulse code; respectively carrying out weighted accumulation and average pooling on the second image pulse code and the second text pulse code to obtain an image floating point feature vector set and a text floating point feature vector set, and calculating cosine similarity to obtain an image-text alignment result; and the text retrieval result of each image in the target image set is obtained based on the image-text alignment result, so that the accuracy and efficiency of image-text retrieval are improved.
Owner:WUHAN UNIV OF TECH

Multi-modal data pairing method and system based on deep learning

The invention provides a multi-modal data pairing method and system based on deep learning, and relates to the technical field of data processing, and the method comprises the steps: obtaining a video multi-frame sequence and a target text, and respectively extracting an overlapped frame group set and a standardized text sequence; performing spatio-temporal feature extraction and text dependency relationship coding to obtain a video time sequence vector sequence and a text vector sequence; executing cross-modal alignment search, and constructing a monotonic matching path set; calculating a semantic and action entity relationship consistency score of the paired elements on the path to obtain a comprehensive score; and determining an alignment relationship between the video and the text based on the optimal path. According to the method, accurate matching of the video and the text is realized, and the cross-modal retrieval efficiency is improved.
Owner:BEIJING YIZHUANG INTELLIGENT CITY RES INST GRP CO LTD

Open vocabulary semantic segmentation method and device based on feature interaction and multi-modal data fusion

The invention provides an open vocabulary semantic segmentation method and device based on feature interaction and multi-modal data fusion. An SAM encoder of freezing parameters is adopted to extract RGB image and Mask image features in parallel, and boundary information enhancement is carried out through a feature fusion module. A CLIP image encoder of freezing parameters is used for carrying out multi-layer feature extraction on an RGB image, and multi-scale semantic information optimization is carried out through a feature enhancement module. And fusing image features extracted by the SAM branch and the CLIP branch through a feature interaction module to realize cross-network feature complementation. In a prediction stage, a category integration strategy combined with a temperature scaling operation is introduced to optimize prediction effects of known and unknown categories. According to the method, the advantages of SAM in the aspect of image boundary extraction and the image-text alignment capability of CLIP are combined, and the generalization capability and robustness of the model in an open vocabulary semantic segmentation task are effectively improved through multi-modal feature fusion and interaction.
Owner:BEIJING TECH & BUSINESS UNIV

Vision-language model training method and device and related equipment

The invention provides a visual-language model training method and device and related equipment, and the method comprises the steps: constructing a training data set comprising an image-text alignment task, a text-driven visual positioning task and a plain text reasoning task through determining a visual encoder and a heterogeneous language model which form a multi-modal heterogeneous recognition model; carrying out dimension dynamic alignment adaptation on the heterogeneous language model, enabling a parameter architecture of the heterogeneous language model to be matched with a visual feature dimension output by a visual encoder, and carrying out supervision and fine adjustment on the visual-language model by adopting a freezing-unfreezing two-stage training strategy based on a training data set; and performing parameter training on a cross-modal alignment module connected with the visual encoder and the heterogeneous language model so as to obtain a trained visual-language model. According to the training method, time consumed by model training is saved, and the convergence speed is increased. Through dimension dynamic alignment and hierarchical weight mapping, the pre-training language ability is reserved to the maximum extent, the plain text task performance loss is reduced, and disastrous forgetting of the model is avoided.
Owner:AISINO CORPORATION

Image-text alignment model training method and device, computer equipment and storage medium

The invention discloses an image-text alignment model training method and device, computer equipment and a storage medium, and the method comprises the steps: carrying out the coding of input image data and text data, and obtaining an image block feature sequence and a text token feature sequence; performing bidirectional feature interaction on the image block feature sequence and the text token feature sequence to obtain fused image features and fused text prompt features; inputting the fused image features, the fused text prompt features and a preset output token into an SAM mask decoder to generate an image mask; filtering the fused image features based on the image mask, extracting image global features related to the text description, and obtaining text global features; constructing a joint loss function; and performing end-to-end training on the model based on the joint loss function until the model converges. According to the method, the limitation of coarse-grained alignment of a traditional model is effectively solved, and the pertinence and matching precision of image-text features are improved.
Owner:SHENZHEN TVT DIGITAL TECH CO LTD

End-to-end automatic driving long tail identification method based on comparative learning pre-training

The invention relates to the technical field of automatic driving end-to-end perception, in particular to an end-to-end automatic driving long tail recognition method based on comparative learning pre-training, and the method comprises the steps: firstly generating synthetic image data with long tail distribution characteristics through a conditional diffusion model; a fine-grained scene classifier is adopted to carry out systematic arrangement and semantic annotation on the generated samples, and a structured multi-modal image-text alignment data set is constructed; and finally, fusing the enhanced data set with the original training set, and optimizing a vision-language joint embedding space through a multi-task contrast loss function to realize parameter updating of the pre-training model. According to the method, a closed-loop optimization mechanism of a generative data enhancement and contrast learning framework is creatively established, the problem of data scarcity in a long-tail distribution scene is effectively relieved, and the cross-modal representation capability and downstream task generalization performance of the model on low resource categories are remarkably improved.
Owner:JIANGSU UNIV

Multi-modal fine-grained semantic alignment method and device based on graph neural network

The invention discloses a multi-modal fine-grained semantic alignment method and device based on a graph neural network, and relates to the field of multi-modal deep learning. Firstly, deep feature extraction is performed on input multi-modal original data, and then word-level text features and local image features are constructed into a cross-modal graph structure. And performing weighted aggregation on node neighborhood information of the cross-modal graph structure through the graph attention network. And finally, carrying out weighted fusion on the text alignment features and the image alignment features. According to the cross-modal feature fusion method, the word-level text features and the local image features are uniformly abstracted into the graph structure nodes for refined alignment, a more accurate cross-modal semantic corresponding relation can be captured, the heterogeneity problem in expression modes and semantic structures is effectively relieved, and the accuracy and reliability of cross-modal feature fusion are improved. The graph attention network can adaptively adjust the weight distribution of information propagation, highlights the effect of key features in the alignment process, and ensures that the model makes full use of important semantic relationships.
Owner:ZHENGZHOU NORMAL UNIV +1

Bidding and cross-bidding early warning and blocking method and system based on multi-modal large model

The invention relates to the technical field of artificial intelligence, in particular to a multi-modal large model-based bidding and cross-bidding early warning and blocking method and system, and the method comprises the following steps: receiving and preprocessing various qualification files, including business licenses, industry licenses, auditing reports and bidding files, submitted by suppliers; applying a document layout analysis model based on joint training of a visual coding network and a text alignment network; the method has the beneficial effects that a scanned certificate, an audit report and a bidding file are automatically analyzed into structured semantic fragments; performing same semantic space alignment on the fragments, supplier basic information, historical records and real-time captured data of an authoritative authentication website; any contradiction or gap among different sources of key information such as registration capital, permission items, financial indexes and the like can be quickly found, and interpretable risk scores and processing suggestions can be given; on the premise of ensuring the private security of the data, the qualification auditing efficiency and accuracy are improved, and the workload of manual checking is reduced.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Vision and text alignment method and system

The invention discloses a vision and text alignment method and system, and belongs to the technical field of artificial intelligence and multi-modal semantic understanding. In order to solve the problem of insufficient vision and language deep fusion in the existing multi-mode question answering, the method mainly comprises the following steps: mapping visual features to a self-attention input space of a language model through a perceptron network, and introducing a fusion attention mechanism into each layer of decoder of the language model to realize layer-by-layer interaction processing of vision and texts. According to the method, deep alignment and fusion of visual information and text semantics can be realized, and the understanding and generation capabilities of a multi-modal question-answering system are improved.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Cross-modal image-text pedestrian retrieval method based on generative model

The invention relates to a cross-modal retrieval technology, in particular to a cross-modal image-text pedestrian retrieval method based on a generative model, which comprises the following steps: acquiring text description information and a corresponding original pedestrian image; calling a diffusion generation model to generate an intermediate image based on the text description; original image features and text features are extracted through an image encoder and a text encoder respectively, and semantic-fused text feature representation is constructed; calculating the similarity between the image feature representation and the text feature representation to obtain an image-text matching score; introducing and generating an intermediate image as a semantic bridge, and realizing multi-modal fusion among the image, the text and the intermediate image through a cross attention mechanism to obtain fused image and text features; and using the fusion features to train a recognition model, and finally outputting a pedestrian image retrieval result corresponding to the text description. According to the method, the image-text alignment precision can be improved, the robustness of a scene with incomplete text description is enhanced, and the recognition accuracy in a cross-modal retrieval task is improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Multi-modal power sample feature migration method and system based on dual cross-modal information decoupling, electronic equipment and storage medium

The invention discloses a multi-modal power sample feature migration method and system based on dual cross-modal information decoupling, electronic equipment and a storage medium, which are applied to the technical field of smart power grids, and comprise the following steps: carrying out hierarchical alignment of visual features and language features by adopting a fine-grained image-text alignment strategy, and capturing distinctive multi-modal representation; based on a double-flow attention mechanism, performing enhanced feature extraction on the multi-modal data; wherein the double-flow attention mechanism comprises an image self-adaptive attention mechanism and a text guide attention mechanism; the method comprises the following steps: performing cross-modal information decoupling processing on multi-modal data, separating to obtain semantic features representing general semantics and modal specific features representing modal features, and performing fine-grained alignment recombination processing to obtain cross-modal fusion features; and migrating the deeply fused multi-modal power sample features to a downstream task of the power system. According to the method, the understanding of the multi-modal features and the further improvement of the migration capability and the efficient operation of the power system tasks are realized.
Owner:STATE GRID INFORMATION & TELECOMM BRANCH +2

Cross-modal semantic alignment method, device and equipment based on fine-grained semantic decoupling

The invention provides a cross-modal semantic alignment method, device and equipment based on fine-grained semantic decoupling, and the method comprises the steps: inputting a to-be-aligned image into a pre-training image encoder of a cross-modal semantic alignment model, and obtaining a global semantic feature, meanwhile, a semantic feature extraction module is used in a high-level encoding block of the image encoder and based on a preset query vector, high-level visual feature extraction is conducted, visual fine-grained semantic features are obtained and input into a semantic feature decoupling module to be decoupled, and fine-grained semantic features of attributes, objects and combinations are obtained; inputting prompt words corresponding to attributes, objects and combinations in the to-be-aligned image into a text encoder for text feature extraction to obtain soft prompt text features of the attributes, the objects and the combinations; and inputting the fine-grained semantic features and the soft prompt text features corresponding to the global semantic features, the attributes, the objects and the combinations into a loss calculation module for image-text alignment, and performing linear processing on an alignment result to obtain an alignment prediction result.
Owner:GUIZHOU UNIV

Image text alignment method based on multi-modal large language model

PendingCN121542772AText alignmentFeature vector
The invention relates to the technical field of image texts, in particular to an image text alignment method based on a multi-modal large language model, which comprises the following steps of: segmenting an image into local areas and encoding the local areas into visual feature vectors by adopting a pre-trained visual encoder; visual and text features are mapped to the same dimension space through a learning projection layer, an InfoNCE loss function is adopted to maximize the similarity of a positive sample pair, dynamic interaction between text query and an image area is achieved through a cross attention mechanism, language module parameters are frozen to pre-train a visual encoder, an overall module is combined to be fine-tuned, and text query is completed. An error case is analyzed, and a verified and improved alignment module is obtained through a data enhancement optimization module. The method solves the problems that a traditional image text alignment method is incomplete in manual feature extraction and low in precision, an existing deep learning model is limited in cross-modal semantic understanding ability, and the generalization ability is poor when facing long-tail data.
Owner:GUANGDONG UNIVERSITY OF FOREIGN STUDIES

Illusion phenomenon mitigation method and system of visual language large model, terminal and medium

The invention discloses an illusion phenomenon mitigation method and system for a visual language large model, a terminal and a medium, and the method comprises the steps: obtaining input image-text data, processing the image-text data to obtain a text token and an image token, and aligning the text token with the image token to obtain an input sequence of the image-text data; inputting the input sequence into the visual language large model, calculating attention distribution through a cross-modal attention mechanism, and obtaining initial vocabulary probability distribution of the current time step based on the attention distribution; and based on the initial vocabulary probability distribution and the attention distribution, performing multi-scale enhancement and fusion reasoning to obtain target vocabulary probability distribution of the current time step. According to the method, through key technologies such as image-text alignment, attention guidance, multi-scale enhancement and fusion reasoning, the understanding and generation capability of the visual language model in a complex image-text scene is effectively improved, and the illusion problem possibly occurring in the cross-modal reasoning process of the large visual language model is relieved.
Owner:SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Time sequence prediction method and device for performing multi-level text alignment by using large model

The invention provides a time sequence prediction method and device for performing multi-level text alignment by using a large model, and belongs to the technical field of time sequence prediction of a rail transit system based on the large model. Comprising the following steps of splitting multivariable time sequence input into a plurality of univariate time sequences according to feature dimensions, performing additive decomposition on each univariate time sequence, and performing fragmentation processing on each decomposed time sequence component; embedding the fragments into a text embedding space of a pre-training language model and aligning the fragments with the text embedding space; combining the structured prompt with the aligned time sequence representation to form input of a large model; and feeding the input combining the prompt and the alignment representation into the frozen large language model, obtaining an output representation of the model, and mapping the output representation into a final prediction result through a linear projection layer. According to the method, the time sequence data and the natural language modality are effectively aligned and fused, and the prediction accuracy and interpretability are remarkably improved.
Owner:CRRC CHANGCHUN RAILWAY VEHICLES CO LTD

Weak supervision video anomaly detection method and system based on prompt learning

The invention provides a weak supervision video anomaly detection method and system based on prompt learning, and belongs to the technical field of abnormal event detection based on computer vision, and the method comprises the steps: obtaining to-be-processed video data; and processing the acquired to-be-processed video data by using a pre-trained anomaly detection model to obtain a specific classification result of the abnormal events in the video. According to the invention, a video local and global adaptive time modeling module is introduced to capture local and global dependency relationships at the same time, and the relationship between the demand of detailed time modeling and the calculation efficiency is balanced; by utilizing an external knowledge base, the distinguishing capability of the model on different categories is improved; according to the method, a text-video comparison loss function is designed, the similarity of a correctly matched text-video pair is enhanced, the similarity of wrong matching is reduced, and too high similarity of a negative sample is effectively inhibited, so that the distinguishing capability of the model is improved, the matching of the text and the video is more accurate, and the video and text alignment capability of the model is enhanced.
Owner:BEIJING JIAOTONG UNIV

Text-driven video generation method based on single video fine tuning

The invention relates to the technical field of video generation and machine learning, and provides a text-driven video generation method based on single video fine tuning, which comprises the following steps: encoding an input video into potential feature representation through a pre-training encoder, and constructing input features; generating a video frame by adopting a video generation network based on a diffusion model, and optimizing the generation process of the video frame by adopting a pre-trained CLIP model and a pre-trained VGG model; and based on the dynamic loss weight adjustment strategy, adaptively balancing the weight of the semantic and texture constraint, and optimizing the training process of the CLIP model and the VGG model. According to the method provided by the invention, a multi-scale space adapter, semantic features based on CLIP and texture features based on VGG are combined. Through a progressive training strategy and dynamic loss weight balance, the quality and consistency of generated frames are improved, text alignment, time coherence and generalization ability are remarkably improved, and the development of the text to video generation field is promoted.
Owner:CHONGQING UNIV OF TECH

Open vocabulary action recognition method and system based on heterogeneous skeleton

The invention discloses a heterogeneous skeleton-based open vocabulary action recognition method and system. The method comprises the following steps of: constructing a heterogeneous open vocabulary skeleton data set; unifying the skeleton representation in the heterogeneous open vocabulary skeleton data set, and defining a unified space structure by establishing the maximum joint number and member number; a skeleton motion encoder model based on a Transform architecture is constructed, the skeleton motion encoder model comprises feature embedding, space-time encoding and a projection layer used for cross-modal alignment, and global time features, global space features and global visual features output by space-time encoding are subjected to cross-modal alignment through three parallel projection networks of the projection layer. Mapping to a semantic space aligned with text embedding from the pre-trained language model; training loss is constructed based on a multi-granularity motion-text alignment strategy including global instance alignment, stream specific alignment and fine granularity alignment, and a motion encoder model is trained. A wide range of experiments on a popular benchmark with heterogeneous skeleton data prove the effectiveness and generalization ability of the proposed method.
Owner:SOUTHEAST UNIV

Speech synthesis method and system for controllable latent variable modeling based on semantic distillation

The invention relates to the technical field of speech synthesis, and particularly discloses a speech synthesis method and system for controllable latent variable modeling based on semantic distillation, and the method comprises the steps: converting a Mel spectrum into continuous latent variable distribution through a speech coding module, generating continuous latent variables through re-parameterization sampling, introducing a self-supervised model for semantic distillation, and carrying out the semantic distillation. According to the method, alignment of latent variables and semantic features is constrained through marginal cosine similarity and distance matrix structure loss, a text encoder maps a phoneme sequence into latent variable distribution, time sequence alignment of a text and the latent variables is achieved in combination with monotonic alignment search, and a decoder reconstructs the latent variables into a Mel spectrum. According to the method, waveform synthesis through a vocoder and total loss function joint optimization reconstruction, KL divergence, distillation, text alignment and confrontation loss are carried out, discrete information loss is avoided through continuous latent variable modeling, semantic consistency and text alignment efficiency are enhanced, the naturalness, coherence and real-time performance of synthesized voice are improved, and the method is suitable for scenes such as voice assistants and virtual anchors.
Owner:BEIJING TIMES RUILANG TECH CO LTD

Machine-learned text alignment prediction for providing an augmented-reality translation interface

Systems and methods for providing an augmented-reality translation interface can include obtaining an image, processing the image to generate an image representation and one or more paragraph bounding boxes, and processing the image representation and the one or more paragraph bounding boxes with a machine-learned alignment classification model to generate one or more alignment classifications. In parallel or in series, text from the image can be determined and translated. The image, the translated text, and the one or more alignment classifications can then be processed to generate and provide the translation in an augmented-reality interface.
Owner:GOOGLE LLC

Multi-modal fine-grained emotion recognition method oriented to human-computer interaction and based on large model

According to the man-machine interaction-oriented multi-modal fine-grained emotion recognition method based on the large model provided by the invention, cross-modal alignment from coarse granularity to fine granularity is realized through an attention pairing interaction module (APIM) on the basis of an aspect-driven vision-text alignment and fusion network (AVTAF); emotion-related visual features (such as facial expressions and gestures) in a robot scene can be accurately captured, and environmental noise is inhibited; meanwhile, the RD-GAT is enhanced, and the reasoning ability of a large model on multi-modal emotion semantics is improved by integrating external emotion knowledge (such as SenticNet). The technology provides a new normal form for intelligent upgrading of robot emotion interaction and multi-modal understanding of a large model, and is expected to promote breakthrough application in the fields of family service robots, medical accompanying assistants, multi-modal content generation and the like.
Owner:BEIJING INST OF TECH

Method, system and medium for constructing a variant character dictionary of ancient chinese medical books and text alignment

The present application belongs to the technical field of natural language processing for traditional Chinese medicine ancient books, and particularly relates to a method and system for constructing a variant character dictionary and text alignment of traditional Chinese medicine ancient books, and a medium. The present application combines the recognition of variant characters and the construction of a variant character dictionary to achieve a text alignment method for traditional Chinese medicine ancient books. Specifically, the present application uses deep learning and natural language processing technology to automatically extract variant character features, significantly improving the coverage range and recognition accuracy; through dynamic programming, semantic similarity calculation and knowledge graph fusion, the multi-modal features are comprehensively considered to significantly improve the alignment accuracy. At the same time, the model can dynamically adapt to new texts and variant characters, and has stronger expansibility and adaptability; and the knowledge graph is used to optimize the alignment result, improving the accuracy and efficiency of text processing. The final generated result is the aligned text sequence, in which the variant characters are correctly recognized and mapped to standard characters. The present application has good application prospects in the digitization of traditional Chinese medicine ancient books.
Owner:CHENGDU UNIV OF TRADITIONAL CHINESE MEDICINE

Small sample learning and out-of-distribution detection method and system based on global and local image-text alignment

The invention discloses a small sample learning and out-of-distribution detection method and system based on global and local image-text alignment, and the method comprises the steps: collecting image data and corresponding class labels, obtaining a data set, processing the data set, and dividing the data set into a training set, an in-distribution test set, and an out-of-distribution test set; an out-of-distribution detection model is constructed and trained, the training comprises two stages, in the pre-training stage, text description is carried out on various image data of a training set, text description is obtained, in the training stage, global and local image-text features are extracted, and related and unrelated local image features are screened out; performing local supervised contrast learning on the fine-tuned text and related local image features, calculating global and local matching scores by using the image features and the fine-tuned text features, and completing model training; and performing classification prediction probability detection on the trained model. According to the invention, the performance of small sample image classification and small sample target detection is improved at the same time.
Owner:SUN YAT SEN UNIV

Small sample action recognition method based on semantic-time sequence adaptive representation learning

The invention relates to a small sample action recognition method based on semantic-time sequence adaptive representation learning, which comprises the following steps of: randomly sampling action categories, and constructing a task sample pair containing a support set and a query set; obtaining temporal and spatial features of category semantic embedding and support samples and query samples; based on an explicit semantic guidance focusing mechanism, the spatial-temporal features of the support samples and corresponding category semantic embedding are aligned to obtain semantic enhanced spatial-temporal features, and multi-frequency sub-sequence construction, one-way modeling and two-way modeling are sequentially carried out on the semantic enhanced spatial-temporal features and the spatial-temporal features of the query samples to obtain the spatial-temporal features of the query samples. Focusing time sequence prototype representation is obtained, and prototype matching is carried out; performing fine-grained video-text alignment on the spatial-temporal characteristics and category semantic embedding of the support sample by using a frame-level cross-modal attention mechanism; prototype matching is carried out on the to-be-recognized sample, and recognition of the query set action category under the condition of few samples is achieved. Compared with the prior art, the method has the advantages of effectively relieving insufficient characterization capability under the condition of few samples and the like.
Owner:TONGJI UNIV

Learable retrieval enhancement-based radiology report generation method for visual text alignment and fusion

The invention discloses a learnable radiology report generation method based on visual text alignment and fusion of retrieval enhancement. The method comprises the steps of collecting and respectively constructing a model training data set and an auxiliary data set, then constructing a visual text alignment and fusion model based on retrieval enhancement, and inputting the model training data set and the auxiliary data set into the visual text alignment and fusion model based on retrieval enhancement together for training. Constructing an inference model according to the trained visual text alignment and fusion model based on retrieval enhancement; and inputting the to-be-detected medical image and the auxiliary data set into the reasoning model for processing to obtain a radiology report corresponding to the to-be-detected medical image. According to the method, retrieval correlation is enhanced, meanwhile, a fine-grained vision-text alignment and fusion method is adopted to align and fuse features, and the problem that fine-grained region-sentence alignment is difficult due to weak supervision of an image report level in the medical report generation process is solved.
Owner:ZHEJIANG UNIV