Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

111 results about "Text alignment" patented technology

End-to-end automatic driving long tail identification method based on comparative learning pre-training

The invention relates to the technical field of automatic driving end-to-end perception, in particular to an end-to-end automatic driving long tail recognition method based on comparative learning pre-training, and the method comprises the steps: firstly generating synthetic image data with long tail distribution characteristics through a conditional diffusion model; a fine-grained scene classifier is adopted to carry out systematic arrangement and semantic annotation on the generated samples, and a structured multi-modal image-text alignment data set is constructed; and finally, fusing the enhanced data set with the original training set, and optimizing a vision-language joint embedding space through a multi-task contrast loss function to realize parameter updating of the pre-training model. According to the method, a closed-loop optimization mechanism of a generative data enhancement and contrast learning framework is creatively established, the problem of data scarcity in a long-tail distribution scene is effectively relieved, and the cross-modal representation capability and downstream task generalization performance of the model on low resource categories are remarkably improved.
Owner:JIANGSU UNIV

Multi-modal fine-grained semantic alignment method and device based on graph neural network

The invention discloses a multi-modal fine-grained semantic alignment method and device based on a graph neural network, and relates to the field of multi-modal deep learning. Firstly, deep feature extraction is performed on input multi-modal original data, and then word-level text features and local image features are constructed into a cross-modal graph structure. And performing weighted aggregation on node neighborhood information of the cross-modal graph structure through the graph attention network. And finally, carrying out weighted fusion on the text alignment features and the image alignment features. According to the cross-modal feature fusion method, the word-level text features and the local image features are uniformly abstracted into the graph structure nodes for refined alignment, a more accurate cross-modal semantic corresponding relation can be captured, the heterogeneity problem in expression modes and semantic structures is effectively relieved, and the accuracy and reliability of cross-modal feature fusion are improved. The graph attention network can adaptively adjust the weight distribution of information propagation, highlights the effect of key features in the alignment process, and ensures that the model makes full use of important semantic relationships.
Owner:ZHENGZHOU NORMAL UNIV +1

Cross-modal semantic alignment method, device and equipment based on fine-grained semantic decoupling

The invention provides a cross-modal semantic alignment method, device and equipment based on fine-grained semantic decoupling, and the method comprises the steps: inputting a to-be-aligned image into a pre-training image encoder of a cross-modal semantic alignment model, and obtaining a global semantic feature, meanwhile, a semantic feature extraction module is used in a high-level encoding block of the image encoder and based on a preset query vector, high-level visual feature extraction is conducted, visual fine-grained semantic features are obtained and input into a semantic feature decoupling module to be decoupled, and fine-grained semantic features of attributes, objects and combinations are obtained; inputting prompt words corresponding to attributes, objects and combinations in the to-be-aligned image into a text encoder for text feature extraction to obtain soft prompt text features of the attributes, the objects and the combinations; and inputting the fine-grained semantic features and the soft prompt text features corresponding to the global semantic features, the attributes, the objects and the combinations into a loss calculation module for image-text alignment, and performing linear processing on an alignment result to obtain an alignment prediction result.
Owner:GUIZHOU UNIV

Image text alignment method based on multi-modal large language model

PendingCN121542772AText alignmentFeature vector
The invention relates to the technical field of image texts, in particular to an image text alignment method based on a multi-modal large language model, which comprises the following steps of: segmenting an image into local areas and encoding the local areas into visual feature vectors by adopting a pre-trained visual encoder; visual and text features are mapped to the same dimension space through a learning projection layer, an InfoNCE loss function is adopted to maximize the similarity of a positive sample pair, dynamic interaction between text query and an image area is achieved through a cross attention mechanism, language module parameters are frozen to pre-train a visual encoder, an overall module is combined to be fine-tuned, and text query is completed. An error case is analyzed, and a verified and improved alignment module is obtained through a data enhancement optimization module. The method solves the problems that a traditional image text alignment method is incomplete in manual feature extraction and low in precision, an existing deep learning model is limited in cross-modal semantic understanding ability, and the generalization ability is poor when facing long-tail data.
Owner:GUANGDONG UNIVERSITY OF FOREIGN STUDIES

Illusion phenomenon mitigation method and system of visual language large model, terminal and medium

The invention discloses an illusion phenomenon mitigation method and system for a visual language large model, a terminal and a medium, and the method comprises the steps: obtaining input image-text data, processing the image-text data to obtain a text token and an image token, and aligning the text token with the image token to obtain an input sequence of the image-text data; inputting the input sequence into the visual language large model, calculating attention distribution through a cross-modal attention mechanism, and obtaining initial vocabulary probability distribution of the current time step based on the attention distribution; and based on the initial vocabulary probability distribution and the attention distribution, performing multi-scale enhancement and fusion reasoning to obtain target vocabulary probability distribution of the current time step. According to the method, through key technologies such as image-text alignment, attention guidance, multi-scale enhancement and fusion reasoning, the understanding and generation capability of the visual language model in a complex image-text scene is effectively improved, and the illusion problem possibly occurring in the cross-modal reasoning process of the large visual language model is relieved.
Owner:SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Time sequence prediction method and device for performing multi-level text alignment by using large model

The invention provides a time sequence prediction method and device for performing multi-level text alignment by using a large model, and belongs to the technical field of time sequence prediction of a rail transit system based on the large model. Comprising the following steps of splitting multivariable time sequence input into a plurality of univariate time sequences according to feature dimensions, performing additive decomposition on each univariate time sequence, and performing fragmentation processing on each decomposed time sequence component; embedding the fragments into a text embedding space of a pre-training language model and aligning the fragments with the text embedding space; combining the structured prompt with the aligned time sequence representation to form input of a large model; and feeding the input combining the prompt and the alignment representation into the frozen large language model, obtaining an output representation of the model, and mapping the output representation into a final prediction result through a linear projection layer. According to the method, the time sequence data and the natural language modality are effectively aligned and fused, and the prediction accuracy and interpretability are remarkably improved.
Owner:CRRC CHANGCHUN RAILWAY VEHICLES CO LTD

Text-driven video generation method based on single video fine tuning

The invention relates to the technical field of video generation and machine learning, and provides a text-driven video generation method based on single video fine tuning, which comprises the following steps: encoding an input video into potential feature representation through a pre-training encoder, and constructing input features; generating a video frame by adopting a video generation network based on a diffusion model, and optimizing the generation process of the video frame by adopting a pre-trained CLIP model and a pre-trained VGG model; and based on the dynamic loss weight adjustment strategy, adaptively balancing the weight of the semantic and texture constraint, and optimizing the training process of the CLIP model and the VGG model. According to the method provided by the invention, a multi-scale space adapter, semantic features based on CLIP and texture features based on VGG are combined. Through a progressive training strategy and dynamic loss weight balance, the quality and consistency of generated frames are improved, text alignment, time coherence and generalization ability are remarkably improved, and the development of the text to video generation field is promoted.
Owner:CHONGQING UNIV OF TECH

Multi-modal fine-grained emotion recognition method oriented to human-computer interaction and based on large model

According to the man-machine interaction-oriented multi-modal fine-grained emotion recognition method based on the large model provided by the invention, cross-modal alignment from coarse granularity to fine granularity is realized through an attention pairing interaction module (APIM) on the basis of an aspect-driven vision-text alignment and fusion network (AVTAF); emotion-related visual features (such as facial expressions and gestures) in a robot scene can be accurately captured, and environmental noise is inhibited; meanwhile, the RD-GAT is enhanced, and the reasoning ability of a large model on multi-modal emotion semantics is improved by integrating external emotion knowledge (such as SenticNet). The technology provides a new normal form for intelligent upgrading of robot emotion interaction and multi-modal understanding of a large model, and is expected to promote breakthrough application in the fields of family service robots, medical accompanying assistants, multi-modal content generation and the like.
Owner:BEIJING INST OF TECH

Method, system and medium for constructing a variant character dictionary of ancient chinese medical books and text alignment

The present application belongs to the technical field of natural language processing for traditional Chinese medicine ancient books, and particularly relates to a method and system for constructing a variant character dictionary and text alignment of traditional Chinese medicine ancient books, and a medium. The present application combines the recognition of variant characters and the construction of a variant character dictionary to achieve a text alignment method for traditional Chinese medicine ancient books. Specifically, the present application uses deep learning and natural language processing technology to automatically extract variant character features, significantly improving the coverage range and recognition accuracy; through dynamic programming, semantic similarity calculation and knowledge graph fusion, the multi-modal features are comprehensively considered to significantly improve the alignment accuracy. At the same time, the model can dynamically adapt to new texts and variant characters, and has stronger expansibility and adaptability; and the knowledge graph is used to optimize the alignment result, improving the accuracy and efficiency of text processing. The final generated result is the aligned text sequence, in which the variant characters are correctly recognized and mapped to standard characters. The present application has good application prospects in the digitization of traditional Chinese medicine ancient books.
Owner:CHENGDU UNIV OF TRADITIONAL CHINESE MEDICINE

Learable retrieval enhancement-based radiology report generation method for visual text alignment and fusion

The invention discloses a learnable radiology report generation method based on visual text alignment and fusion of retrieval enhancement. The method comprises the steps of collecting and respectively constructing a model training data set and an auxiliary data set, then constructing a visual text alignment and fusion model based on retrieval enhancement, and inputting the model training data set and the auxiliary data set into the visual text alignment and fusion model based on retrieval enhancement together for training. Constructing an inference model according to the trained visual text alignment and fusion model based on retrieval enhancement; and inputting the to-be-detected medical image and the auxiliary data set into the reasoning model for processing to obtain a radiology report corresponding to the to-be-detected medical image. According to the method, retrieval correlation is enhanced, meanwhile, a fine-grained vision-text alignment and fusion method is adopted to align and fuse features, and the problem that fine-grained region-sentence alignment is difficult due to weak supervision of an image report level in the medical report generation process is solved.
Owner:ZHEJIANG UNIV

Speech recognition method and device and computer program product

The invention discloses a speech recognition method and device, and a computer program product. The method comprises the following steps: firstly, extracting a first acoustic feature of a target speech by using a pre-trained convolutional layer; inputting the first acoustic feature into a pre-trained Dynamic Mask module, and adaptively generating a first mask matrix of the target voice in a recognition mode (such as a real-time recognition mode or a non-real-time recognition mode) to which the target voice belongs; inputting the first mask matrix into a pre-trained Transform layer for feature extraction to obtain a second acoustic feature of the target voice; according to the embodiment of the invention, the second acoustic feature of the target voice and the second mask matrix aligned with the voice text corresponding to the target voice are generated, and then the second acoustic feature of the target voice and the second mask matrix are input into the pre-trained text decoder for decoding, so that the text recognition result of the target voice can be obtained through more accurate decoding. Therefore, the accuracy of the target speech recognition result is effectively improved, the cost is reduced, and an ideal speech recognition effect is achieved.
Owner:IFLYTEK CO LTD

Image generation and image-text alignment methods, apparatus, terminal devices and storage media

This invention discloses an image generation and image-text alignment method, apparatus, terminal device, and storage medium. In the training process of a visual language model, training images and training text are input into the visual language model. The visual language model simultaneously performs image generation and image-text alignment tasks to obtain a unified visual language model, which can realize image-text alignment and multimodal generation functions, solving the technical problem that the prior art cannot achieve a unified discrimination and generation model.
Owner:SHANGHAI TIANTU YUNRUI INTELLIGENT TECH CO LTD

Dynamic news abstract generation method based on voice and visual perception

The invention discloses a dynamic news abstract generation method based on voice and visual perception, which comprises the steps of voice recognition text extraction, audio-text alignment, prompt word generation and final news abstract generation. The objective is achieved through an innovative process comprising a plurality of key optimization links. A robust long audio processing and recognition mechanism, an optimized audio-text alignment method, an elaborately designed cue word framework and an efficient fine-tuning multi-modal large language model are integrated. According to the method, deep understanding of multi-modal information and high-quality text generation are jointly realized, and the quality of four core dimensions including content coverage, semantic accuracy, logic continuity and language fluency of the abstract is comprehensively improved.
Owner:NANJING INFORMATION HIGH-SPEED RAILWAY RES INST OF SCI AND TECH

Multimodal chart question and answer large model construction method, electronic device and storage medium

The application provides a multimodal graph table question and answer large model construction method, an electronic device and a storage medium, comprising: training a graph-text alignment model based on a first sample data set to obtain a trained graph-text feature alignment model; wherein the first sample data set comprises image samples and corresponding text content; training a multimodal graph table question and answer large model with the trained graph-text feature alignment model based on a second sample data set to obtain a trained multimodal graph table question and answer large model as a final multimodal graph table question and answer large model, and the second sample data set comprises context representation information of a graph table sample, an image and question and answer pair data. The multimodal graph table question and answer large model obtained by the application can further improve the graph table question and answer capability of the existing multimodal graph table question and answer large model, and has strong Chinese understanding capability.
Owner:北京中科闻歌科技股份有限公司

A multi-directional text alignment method

The application provides a multi-directional text comparison method, comprising: exporting a packaging design drawing into a PDF format and parsing text content and corresponding position information; splitting the text parsed from the PDF according to intervals, and judging whether the split text is consistent; calculating 2-gram word frequency by using a large amount of Chinese corpus, calculating the probability of text in a normal order and a reverse order, and taking the probability as a basis to judge whether the text is in a reverse order; processing all the text according to position coordinates, judging the direction of the text block according to the normal order and the reverse order of the text, then sorting and merging the text in the text block; comparing the PDF text content with the standard text content for examination, matching similar lines, and marking the differences of the similar lines. The processing result can greatly reduce the structured difference between the parsed text and the actual text, improve the detection accuracy of the parsed text and the actual text, reduce the workload of manual intervention, thereby reducing the packaging design cost and precision.
Owner:PU HUA KE JI YOU XIAN GONG SI

Combined zero sample recognition method based on shared learnable soft query vector and related equipment

The invention discloses a combined zero sample recognition method based on a shared learnable soft query vector and related equipment. The method comprises the following steps: acquiring an initial visual feature, an initial text attribute feature, an initial text object feature and an initial text combination feature; performing visual alignment and decoupling by sharing the soft query vector to obtain a target visual attribute feature, a target visual object feature and a target visual combination feature; performing text alignment and decoupling by sharing the soft query vector to obtain a target text attribute feature, a target text object feature and a target text combination feature; obtaining the similarity between the target visual attribute feature and the target text attribute feature, the similarity between the target visual object feature and the target text object feature, and the similarity between the target visual combination feature and the target text combination feature; and carrying out weighted fusion on the similarity to obtain a combined zero sample recognition result. The method can improve the recognition capability of unseen combinations, and can be widely applied to the technical field of computer vision.
Owner:GUANGZHOU UNIVERSITY

program

To provide a program with label editing functionality that reduces the effort required to adjust the object frame of a text object. [Solution] When the text alignment of a text object 60 placed on the editing screen 50 by the label editing application 41 is left-aligned, the width W1 of the text object 60 is changed by manipulating the right edge 62d of the object frame 62, causing the length W11 of the right margin 65d to change. If the length W1 of the right margin 65d is greater than or equal to a first threshold, the PC1 displays the four edges 62a, 62b, 62c, and 62d of the object frame 62 on the display 13a in C1 color. If the length W1 of the right margin 65d is not greater than or equal to the first threshold, the manipulated right edge 62d of the object frame 62 is displayed on the display 13a in a different C2 color from C1 color, and the other edges 62a, 62b, and 62c are displayed on the display 13a in C1 color.
Owner:BROTHER KOGYO KK

Intelligent recognition method for CAD drawing

The invention relates to the technical field of drawing recognition, and particularly discloses a CAD drawing intelligent recognition method which comprises the steps of drawing preprocessing, primary scheme drawing recognition, character recognition, character content standardization, table structure recognition and image-text alignment and output. Deformation, block shielding and multi-layer table recognition are realized through a multi-vertical-domain special small model, and text information standardization is completed in combination with a vertical domain large model subjected to electrical knowledge fine adjustment and an updatable AI knowledge base. An assembly line framework is adopted, independent deployment and flexible expansion of a multi-product line model are supported, and the newly-added model does not affect an existing system. The whole recognition error rate is reduced, unified specification of multi-source drawings is achieved, seamless joint of recognition results and an enterprise production system is guaranteed, and services can be provided for the outside through an API interface.
Owner:BEIJING YIYANG VISION INFORMATION TECHNOLOGY CO LTD

Knowledge graph construction method and device for intellectual property retrieval and storage medium

The invention provides an intellectual property retrieval-oriented knowledge graph construction method and device and a storage medium, and the method comprises the steps: reading a target intellectual property text data set and a parallel corpus and a reference associated text in the target intellectual property text data set, mining synonymous mapping clues, reference traceability clues and technical theme associated clues implied in the texts, and constructing the intellectual property retrieval-oriented knowledge graph by the synonymous mapping clues, the reference traceability clues and the technical theme associated clues; sorting to obtain a weak supervision signal set, extracting synonymous expression pairs in parallel corpora and semantic association pairs in a reference association text, carrying out cross validation and duplicate removal to obtain a text alignment reference set, and carrying out global association matching on the text alignment reference set and the target intellectual property text data set to obtain an initial structured knowledge unit set; the method comprises the steps of obtaining a standardized knowledge unit set, performing credibility regularization to obtain a standardized knowledge unit set, performing entity classification and relation association organization according to hierarchical requirements of intellectual property retrieval, and constructing to obtain the intellectual property retrieval-oriented knowledge graph. The intellectual property retrieval accuracy and response efficiency can be effectively improved through the intellectual property retrieval method and device.
Owner:HENAN UNIV OF ANIMAL HUSBANDRY & ECONOMY

A multi-layer multi-modal semantic knowledge graph generation method and system for active interaction

The application discloses a kind of multi-layer multi-modal semantic knowledge graph generation method and system for active interaction, belong to artificial intelligence and natural language processing technical field, it includes the multi-modal feature of extraction terminal device document, fusion generation graphic-text alignment similarity matrix;Concatenate prompt word sequence and extract initial static knowledge graph by combining large language model;The time sequence correlation of dynamic running log is used to adjust edge weight to output time sequence update knowledge graph;To feature node, graph layer division is executed to build multi-layer multi-modal semantic knowledge graph;Digital twin intelligent agent is instantiated in edge side and dynamically configured polling trigger frequency;Combined with environmental sensing data flow and knowledge subgraph, generate terminal action control instruction set.The application adopts multi-source data fusion and digital twin closed-loop control mechanism, can realize the dynamic time sequence evolution and intelligent decision of graph, improve the active interaction and adaptive precision control ability of multi-device cooperation in intelligent space.
Owner:CHINA ACADEMY OF MACHINERY SCIENCE & TECHNOLOGY

Multi-modal feature alignment method and device for heterogeneous data and medium

The invention discloses a multi-modal feature alignment method and device for heterogeneous data and a medium, and relates to the field of intelligent retrieve.The method comprises the steps that an input multi-modal document is analyzed, and text elements and non-text elements are recognized; extracting original text features, original image features and original structured text features; constructing a shared semantic space, mapping the original image features to the shared semantic space through an image mapping network, and mapping the original text features and the original structured text features to the shared semantic space through a text mapping network; based on the mapping features in the shared semantic space, through a predefined multi-objective optimization function, executing multi-objective joint optimization to drive feature alignment; and generating a query vector based on the aligned image alignment feature and text alignment feature. By constructing a unified shared semantic space and utilizing multi-target joint optimization to drive feature alignment, it is ensured that original information of non-text elements is completely reserved in the feature mapping process.
Owner:山东浪潮智慧建筑科技有限公司

A multi-source remote sensing data-based visual language model cross-modal alignment method

PendingCN122435397ASensing dataText alignment
The embodiment of the application discloses a visual language model cross-modal alignment method based on multi-source remote sensing data, which can improve the completion effect of remote sensing intelligent tasks. The method adopts a position coding technology to extract visual features of visible light images and invisible light images in multi-modal remote sensing data, then performs cross-modal feature fusion based on a cross attention mechanism, and then performs image-text alignment processing through a projection layer to obtain target features, and takes the target features as input data of a large language model to support the large language model to complete remote sensing intelligent tasks.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Multi-modal reasoning method based on multi-modal knowledge graph

The invention discloses a multi-modal reasoning method based on a multi-modal knowledge graph, and relates to the technical field of natural language processing. The method aims to improve the accuracy and robustness of a large language model in cross-modal understanding and generation tasks. The method specifically comprises the following steps: independently embedding an input text into an LLM embedding layer through a language encoder; an input picture is independently embedded into an embedding layer of an LLM through a visual encoder, a multi-modal knowledge graph is independently embedded into the embedding layer of the LLM through a KG encoder, the embedding space of the visual encoder and the embedding space of the KG encoder are aligned with the text embedding space of the LLM through a visual and knowledge adapter, and image-text alignment is improved through a cross-modal alignment module. According to the method, external knowledge from the multi-modal large model is utilized and injected, and the multi-modal reasoning capability of the LLM is expanded and strengthened, so that the accuracy and robustness of the model in cross-modal understanding and task generation are improved.
Owner:贵州省通信产业服务有限公司

Operation and construction illegal behavior identification method and system

The invention discloses an operation and construction violation behavior identification method and system, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a plurality of video clips of operation and construction violation behaviors; performing semantic alignment between the video clips and the dangerous behavior text on each video clip by using a pre-trained video text alignment model to obtain a semantic alignment data set; outputting a window text similarity score according to the semantic alignment data set by utilizing a pre-trained video text alignment model; generating an enhanced training set comprising positive samples and difficult negative samples according to the window text similarity score and the semantic alignment data set; using an o-LoRA module, a large-scale video understanding data set and an enhanced training set to train a multi-modal large model; and inputting monitoring videos of operation and construction into the trained multi-modal large model, and identifying to obtain illegal behaviors. According to the method, the operation and construction violation behaviors can be efficiently and accurately identified.
Owner:SUN YAT SEN UNIV

A method for re-identifying a dressed pedestrian based on causal representation learning

This application provides a method for pedestrian re-identification under clothing changes based on causal representation learning. Applying the field of computer vision technology, this method acquires the pedestrian image to be identified and the corresponding identity prompt text. After preprocessing the pedestrian image and the corresponding identity prompt text, the data is input into a trained pedestrian re-identification model under clothing changes for analysis and processing, resulting in a clothing-independent pedestrian identity feature representation. This model includes: an image encoder, a text encoder, an image-text alignment module, a clothing texture causal intervention module, a causal consistency constraint module, and a clothing-invariant fine-grained representation module. The obtained clothing-independent pedestrian identity feature representation is matched with the pedestrian identity features of pedestrian images stored in the database to obtain the pedestrian re-identification result. This method improves the pedestrian re-identification capability in scenarios where clothing texture and style change.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Time series prediction method and device for multi-level text alignment using large models

The application provides a time series prediction method and device for multi-level text alignment using a large model, belonging to the technical field of time series prediction based on large models in rail transit systems. The method comprises the following steps: inputting a multivariate time series into a plurality of single-variable time series according to feature dimensions, performing additive decomposition on each single-variable time series, and performing fragmentation processing on each time series component after decomposition; aligning the fragments with the text embedding space of a pre-trained language model; combining the structured prompts with the aligned time series representation to form the input of the large model; feeding the input combined with the prompt and the aligned representation into a frozen large language model to obtain the output representation of the model, and mapping the output representation to the final prediction result through a linear projection layer. The application effectively aligns and fuses time series data and natural language modalities, significantly improving the accuracy and interpretability of the prediction.
Owner:CRRC CHANGCHUN RAILWAY VEHICLES CO LTD

Preference-based long text alignment method

This invention belongs to the field of long text modeling in large language models, and relates to a preference-optimized long text alignment method to improve the text understanding and information alignment capabilities of large language models. The method includes: collecting seed data of long texts; performing data length upsampling to increase the proportion of long texts; formatting each long text to obtain triples including the question, long contextual information, and prediction results; segmenting the long contextual information into equal-length text blocks; scoring the critical path for each text block; synthesizing preferred and unpreferred data based on the critical path scores; constructing positional indices based on continuous and sparse blocks respectively, and synthesizing fine-tuning data; and training and fine-tuning the large language model using a low-rank adaptation method based on the fine-tuning data to obtain the fine-tuned large language model. This invention further optimizes model performance through specific training objective design, avoids unpreferred predictions, and improves the efficiency and quality of long text alignment.
Owner:MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY

Cross-modal data semantic understanding and alignment method and device based on large model

The embodiment of the invention provides a cross-modal data semantic comprehension and alignment method and device based on a large model, and the method comprises the steps: designing a multi-modal hybrid expert MoE encoder, and employing the MoE encoder to carry out the integration of multi-modal data through sharing a bottom representation space; and designing a region text alignment module and a semantic graph reasoning network through the shared bottom layer representation space, and performing modal alignment of the multi-modal data through the region text alignment module and the semantic graph reasoning network.
Owner:CHINA NAT BUILDING MATERIALS TECH CO LTD +3

Power grid bidding document review method and system based on large language model assistance

The invention provides a power grid bidding document review method and system based on large language model assistance. The method comprises the steps of obtaining a power grid bidding document PDF and a review index explanatory file, performing structured analysis on the bidding document, and generating a structured text with aligned pictures; the method comprises the following steps: establishing a text three-level index system, partitioning a structured text, performing semantic vectorization, storing vectorized data into an ES database, and establishing a new index supporting semantic retrieval; analyzing the fuzzy indexes of the review system based on LLM, performing multi-layer disassembly, and establishing association with bidding document content; performing parallel mixed retrieval by adopting keyword and vector retrieval, and obtaining each disassembly problem support text block in combination with a new index; integrating a scoring interval rule, a disassembling problem and a support text block, and constructing an LLM review cue word; the same batch of bidding documents are processed in parallel, the previous steps are executed for each part, sub-scores are obtained through multiple rounds of iterative optimization and then weighted aggregation is performed to form a review conclusion, and the power grid bidding auditing efficiency and the review accuracy and fairness are improved; image-text alignment fusion is realized, and the review result is more comprehensive and reliable.
Owner:XI AN JIAOTONG UNIV +1