Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1099 results about "Text graph" patented technology

In natural language processing (NLP), a text graph is a graph representation of a text item (document, passage or sentence). It is typically created as a preprocessing step to support NLP tasks such as text condensation term disambiguation (topic-based) text summarization, relation extraction and textual entailment.

Semantic comprehension driven cross-modal information fusion and retrieval method and system

The invention discloses a cross-modal information fusion and retrieval method and system driven by semantic comprehension, and the method comprises the steps: obtaining text, image and audio original data, and extracting an initial feature set of each modal through a deep neural network; dynamically distributing each modal weight coefficient based on an attention mechanism, and performing weighted fusion on the initial feature set to obtain cross-modal fusion feature representation; through a cross-modal semantic association analysis model, high-dimensional semantic association features are extracted from the fusion feature representation, and semantic enhancement feature vectors are generated; constructing a cross-modal semantic graph network based on the vector, complementing missing modal features, and generating an optimized multi-modal feature set; and inputting the optimized feature set and the query sample into a contrast learning model, calculating a semantic similarity score, and generating a cross-modal retrieval result sorting list according to the score.
Owner:SHANGHAI CIVIL AVIATION VOCATIONAL & TECH COLLEGE

Cross-modal knowledge reasoning method based on multi-modal large model

The invention relates to a cross-modal knowledge reasoning method based on a multi-modal large model. In a cross-modal knowledge reasoning process, an existing model is usually limited by single-modal information extraction and shallow feature fusion, so that deep semantic association among data such as texts, images and videos is difficult to fully capture. In order to solve the problem, the invention provides a model for fusing multi-modal information such as texts, images, videos, documents and the like, and processing of multi-modal data is converted into unified feature extraction, interaction and deep reasoning tasks by fully utilizing a supervision fine tuning strategy, a self-adaptive attention mechanism and a cross-language processing technology. The model adopts a modular design, integrates multi-source data complementary analysis, spatial-temporal feature modeling and emotional semantic analysis, and realizes multi-modal collaborative interaction, dynamic scene understanding, long video key event analysis and man-machine co-emotional response. Through sufficient training, the multi-modal large model shows excellent logical reasoning ability and emotion understanding ability in a complex cognitive task, and a brand new solution is provided for efficient extraction, deep semantic analysis and intelligent response of cross-modal information.
Owner:SHENYANG INST OF COMPUTING TECH CO LTD THE CHINESE ACAD OF SCI

Multi-modal bill processing method based on dynamic knowledge enhancement

The invention discloses a multi-modal bill processing method based on dynamic knowledge enhancement. The multi-modal bill processing method comprises the following steps: S1, constructing a dynamic knowledge base containing an aging weight; s2, synchronously processing text, image and format features of the bill by adopting a multi-modal feature fusion network to generate a composite feature vector; s3, semantic-level, format-level and timeliness three-level fusion retrieval is carried out based on the composite feature vector, and a three-level fusion retrieval engine comprises dynamic weighted sorting with timeliness attenuation, a difference degree triggered artificial review mechanism and a policy sensitive slope adjustment algorithm; s4, setting a multi-expert cooperative verification system, wherein the multi-expert cooperative verification system comprises cooperative work of a rule engine, a large language model and a logical reasoning module; s5, implementing a dynamic knowledge updating mechanism, and automatically triggering incremental learning of the knowledge base when policy change or format update is detected; and S6, outputting structured data, and synchronously generating an auditing traceability chain containing a decision path. According to the method, the key field identification accuracy can be improved, and auditing traceability and non-perceptual increment updating in the whole process are realized.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY

Document retrieval method based on multistage index and feature clustering

The invention relates to the technical field of document retrieval and information processing in the data processing technology, in particular to a document retrieval method based on multistage indexing and feature clustering, which comprises the following steps: performing high-dimensional space mapping on multi-modal features such as texts and images through a quantum embedding layer to generate cross-modal joint feature representation; a first-level index of a multi-level index architecture is dynamically initialized based on a meta-clustering algorithm, and semantic blocks of a second-level index are divided in combination with a multi-head self-attention mechanism. And an optimal transmission matrix is generated by using a Sinkhorn algorithm to align cross-node feature distribution. The multi-target mixed retrieval strategy is fused with vector retrieval, keyword retrieval and graph retrieval results, and weight distribution is dynamically adjusted. Through collaborative optimization of quantum calculation, federated learning and causal reasoning, a closed-loop technical architecture from feature analysis to dynamic index construction is formed, the problems of insufficient cross-modal fusion, static clustering deviation and semantic association deficiency are solved, and the precision, efficiency and dynamic adaptability of heterogeneous document retrieval are improved.
Owner:TAIJI COMPUTER CORPORATION LIMITED

Text detail map-based method for supervising end-to-end text detection and recognition

A text detail map-based method for supervising end-to-end text detection and recognition, pertaining to the field of text processing. The method comprises the following steps: given an input image containing text in any shape, processing the input image by means of two separate processing branches; designing a text attention head (TAH), and designing a feature pyramid enhancement fusion module (FPEFM); the FPEFM performing feature self-enhancement at different sizes, fusing text image local features and global text position information extracted by a TAH module, and fusing features extracted by the TAH from feature maps of different sizes; stacking a plurality of FPEFMs to continuously enhance the feature representation capability of a model and the depth of the model; and sampling the feature maps to a unified size to obtain a final enhanced feature map.
Owner:CHONGQING UNIV OF TECH

LLM-driven complex report OCR error self-correction method and system

The invention discloses an LLM-driven complex report OCR error self-correction method and system, and the method comprises the following steps: S1, obtaining complex report image data, executing OCR processing, and constructing an original field data set; s2, extracting context information, identifying semantic contradiction fields, and generating a to-be-corrected field set; s3, the pointer generation network generates a plurality of field correction candidates to form a candidate field set; s4, constructing a dobby machine model, selecting an optimal field correction result, and forming a correction field output set; s5, executing format analysis, and extracting a chart title field, a legend field and a data region text; s6, generating a corrected field result of a chart title field according to a chart structure semantic consistency mechanism; and S7, performing field restoration and format reconstruction, and outputting structured report data. According to the method, intelligent error correction and structured reconstruction of fields in a complex report are realized by fusing a large language model, a pointer generation network and a multi-arm machine mechanism.
Owner:ZHEJIANG FULIN TECH CO LTD

Multi-agent dynamic arrangement method based on multi-modal analysis and adaptive retrieval

The invention discloses a multi-agent dynamic arrangement method based on multi-modal analysis and adaptive retrieval, and relates to the technical field of artificial intelligence and information retrieval. Comprising the steps of S1, converting a text, an image, structured data and voice content input by a user into a unified multi-mode semantic representation, S2, converting the unified multi-mode semantic representation into a specific execution process, and S3, automatically scheduling a reasoning agent, a knowledge obtaining agent and an execution agent according to DAG nodes, task elements and available resources, and obtaining the task elements and the execution agent according to the reasoning agent, the knowledge obtaining agent and the execution agent. S4, after task process construction and agent arrangement are completed, dynamic retrieval, evidence convergence and strategy optimization are carried out on information requirements related to a user task, so that a reasoning agent obtains complete knowledge support with consistent context, and S5, knowledge evidence is combined with a task process, so that the task process is completed. The method comprises the following steps: step S6, implementing problem solving, strategy generation and task closed-loop execution through a reasoning agent, step S6, performing actual operation on a target task by an execution agent according to an executable instruction sequence output by the reasoning agent, and outputting a result, and step S7, performing result verification according to an output result returned by the execution agent, and the correctness, integrity and consistency of an output result are examined through rule verification, model evaluation and evidence alignment.
Owner:INSPUR GROUP CO LTD +1

Long-tail image recognition method based on multi-modal semantic generation and image-text fusion

The invention discloses a long-tail image recognition method based on multi-modal semantic generation and image-text fusion. The method comprises the following steps: extracting structured semantic description from a tail image; carrying out semantic rewriting and enhancement based on a multi-modal visual language model, and generating an image semantic extension description; the image semantic extension description is optimized based on semantic duplicate judgment and a style alignment mechanism, and an optimized text description set is obtained; inputting the optimized text description set into a text graph model, generating a tail class image sample, performing semantic and visual quality screening, and constructing to obtain an enhanced image set for training; constructing a training data set based on the original long-tail data set and the enhanced image set, and training an image-text fusion classification model; and inputting a to-be-identified image into the trained image-text fusion classification model, and outputting classification results of all categories. According to the method, the discrimination capability in a long-tail distribution scene is enhanced, and the method has a stronger generalization characteristic.
Owner:SOUTH CHINA UNIV OF TECH

Artificial intelligence-powered large-scale content generator

An AI-powered content generation system that creates consistent, coherent, and engaging multi-modal content by integrating multiple specialized AI components. The system analyzes user input, identifies key elements, and maintains continuity throughout the generation process. It incorporates a feedback loop to learn and adapt based on user preferences, enabling personalized content experiences. The modular architecture allows for seamless integration of AI components focusing on text, images, audio, and interactive elements. The system ensures consistency across modalities and over extended periods, while managing rights, licenses, and royalties using blockchain technology. This advanced platform revolutionizes content creation, consumption, and management in the digital age.
Owner:QOMPLX INC

Large model Agent intelligent decision-making method and system fusing multi-modal data

The invention discloses a multi-modal data fused large model Agent intelligent decision-making method and system, belongs to the technical field of artificial intelligence, multi-modal data processing, deep learning, reinforcement learning and intelligent decision-making, and aims to solve the technical problem of how to improve the performance and adaptability of intelligent decision-making in processing complex tasks and dynamic environments. According to the technical scheme, the method comprises the steps of multi-modal data fusion, wherein text, image and audio data from different modals are integrated, and unified feature representation is generated through feature extraction and feature fusion technologies; intelligent decision-making: decision-making reasoning is carried out based on the fused feature representation, and a final decision-making result is generated by adopting a deep learning model and a reinforcement learning algorithm; adaptive learning: monitoring data changes and decision-making effects in real time, and dynamically adjusting deep learning model parameters and strategies; and feedback optimization: further optimizing the performance of the deep learning model by collecting the feedback information of the decision result.
Owner:浪潮智慧城市科技有限公司

Multi-modal natural language understanding and generating system and method

The invention discloses a multi-modal natural language understanding and generating system and method. The method comprises the following steps: constructing a cross-modal pre-training module, training a multi-modal encoder, and establishing a cross-modal association mapping space; mixing prompt fine tuning is carried out, and a complete blank filling template is constructed; according to the intention reasoning network, extracting multi-round dialogue intention representation of the user, and retrieving an external knowledge base for fine-grained reasoning; constructing a unified semantic representation framework, embedding the text, the image and the voice into a unified space, and generating a query vector of multi-modal intention perception; and the knowledge query module based on key value memory generates entity-level multi-modal replies and optimizes the semantic comprehension and generation capability of the dialogue model. According to the method, the multi-modal information understanding and generating capacity is improved, deep association and understanding of image and text information are achieved, downstream task adaptability is enhanced, task completion accuracy and efficiency are improved, unified semantic representation of the multi-modal information is achieved, and support is provided for information retrieval and utilization.
Owner:UNIV OF ELECTRONIC SCI & TECH OF CHINA CHENGDU COLLEGE

Multi-modal heterogeneous model retrieval enhancement method and system

The invention provides a multi-modal heterogeneous model retrieval enhancement method and system, and the method comprises the steps: building a knowledge and application example double-corpus based on user multi-modal query, and designing a joint retrieval mechanism to obtain a result set; mapping and scheduling to obtain feature representation through special processing channels for texts, images and audios and a Spiking neural network with a segmented trapezoidal topological structure; constructing a three-stage cascade architecture of a basic model, an advanced model and human experts, and obtaining a decision path and answer candidate set in combination with a recursive and discarding decision mechanism; a Hamiltonian graph network is used for representing a multi-modal relation, and a gradient-free descent method is used for rapidly training and optimizing model parameters; an enhanced retrieval result is obtained through cross-modal semantic alignment and dynamic retrieval window adjustment; and high-quality response is obtained through context-aware sorting and retrieval enhanced reasoning. According to the method, the multi-modal information retrieval processing efficiency and the heterogeneous model reasoning response quality are improved.
Owner:贵州中汇科技发展有限公司

Traffic large model construction and decision-making method and device based on multi-modal two-way map reasoning

The invention discloses a traffic large model construction and decision-making method and device based on multi-modal two-way map reasoning, and the method comprises the steps: constructing a multi-modal data set of a text, an image and a track, generating fusion features through spatial-temporal clustering and cross-modal Transform coding, carrying out the two-way map reasoning in combination with a traffic knowledge map, and carrying out the decision-making of the traffic large model. The method comprises the following steps: generating an embedded representation through a forward graph neural network, reversely mapping a decision scheme generated by a language model to a graph to verify consistency, outputting knowledge to enhance embedding, fusing multi-modal features and knowledge embedding by adopting an LoRA multi-task joint fine tuning technology, adapting to traffic field tasks, deploying a real-time inference engine, and carrying out real-time inference on the traffic field. And processing the dynamic data flow through an aging perception attention mechanism, and outputting traffic event identification, path planning and scene question and answer results in parallel. Compared with the prior art, the method has the advantages that the problems of insufficient multi-source heterogeneous data fusion, low knowledge utilization efficiency and poor real-time decision consistency can be solved, and the semantic understanding and decision accuracy of the traffic large model is effectively improved.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY

Multi-modal data fusion method and system based on large model agent

The invention discloses a multi-modal data fusion method and system based on a large-model intelligent agent, belongs to the technical field of artificial intelligence, big data processing and intelligent agents, and aims to solve the technical problem of how to effectively integrate data of different modalities such as texts, images and voices by using the large-model intelligent agent. According to the technical scheme, the method comprises the steps that data collection and preprocessing are conducted, specifically, various modal data of texts, images and voice are collected through a web crawler, an API interface, a camera and microphone equipment, the collected data are preprocessed, the preprocessed multi-modal data are obtained, and the data quality is ensured; feature extraction and mapping: extracting corresponding modal features from the preprocessed multi-modal data through CNN and Transform models, mapping the different modal features to the same space, and combining the aligned features to form comprehensive feature representation; carrying out multi-modal fusion processing; and performing intelligent decision and feedback.
Owner:浪潮智慧城市科技有限公司

Multi-modal fusion rumor detection method and system based on dynamic graph convolutional neural network

The invention discloses a multi-modal fusion rumor detection method and system based on a dynamic graph convolutional neural network. According to the method, a dynamic feature graph of a language propagation path is constructed, and potential features in the language propagation process are extracted and analyzed by utilizing time sequence changes and key node relations between nodes in a propagation graph. A neural network is adopted to extract and enhance image data, text semantic features are extracted in combination with a text feature modeling network, text feature vectorization expression is achieved based on a BERT model, and rich semantic information is obtained. And a gating mechanism is introduced to dynamically adjust fusion weights of different modal features, and an information fusion strategy is optimized. A collaborative attention mechanism is further adopted for deep fusion, interactive learning of text, image and propagation path features is enhanced, and the relevance of cross-modal and time series data is improved. And finally, inputting the fused feature vectors into a classifier for accurate classification, thereby realizing accurate detection of the social media rumors. According to the method, the multi-modal features are effectively integrated, and the false information identification efficiency is remarkably improved.
Owner:CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY +1

Multi-modal fusion intelligent question answering and knowledge retrieval method and system

The invention discloses a multi-modal fused intelligent question answering and knowledge retrieval method and system, and the method comprises the steps: building a multi-modal data index model oriented to a heterogeneous knowledge source, carrying out the feature mapping of text, image, table, chart, audio and video contents through a unified semantic embedding space, and generating a cross-modal index set; after a query request is received, performing semantic matching and structure matching on the cross-modal index set by using a multi-channel retriever to obtain candidate evidence fragments; and based on an evidence granularity decomposition strategy, performing minimum evidence unit division on text statements, table units, chart data points and multimedia frame contents in the candidate evidence fragments, and establishing a semantic consistency graph among the units. According to the method, high-credibility traceable generation of question and answer results is realized through multi-modal fusion and space-time consistency constraint, and the retrieval precision and interpretation transparency in a complex knowledge scene are remarkably improved.
Owner:NANJING CHUANGLIAN INTELLIGENT SOFT INFORMATION TECH CO LTD

Government information consultation system based on large language model

PendingCN120653787ASemantic analysisKnowledge representationConsultation systemEngineering
The invention belongs to the technical field of artificial intelligence, and discloses a government information consultation system based on a large language model, which comprises a user interaction module, a large language model core engine, a knowledge base integration module, a multi-level authority management module, a feedback optimization mechanism and a risk control module. The comprehensive intelligent government affair service system is constructed through the six core modules, remarkable advantages are shown in government affair service digital transformation, the system innovatively adopts multi-mode interactive design, multiple input and output modes of voice, text and images are supported, an intelligent authority management mechanism is matched, and the intelligent authority management mechanism is matched with the intelligent authority management mechanism. According to the technical scheme, the accessibility and convenience of government affair services are greatly improved, precise services for different user groups are achieved, it is guaranteed that sensitive data are safe and controllable while wide spreading of government affair information is guaranteed, and a large language model of a system core is subjected to professional government affair scene optimization training and is combined with a dynamically-updated knowledge graph technology.
Owner:JIANGXI YUANREN ENTERPRISE MANAGEMENT CO LTD

False information detection method based on cross-modal confrontation and progressive training

The invention discloses a false information detection method based on cross-modal confrontation and progressive training, and belongs to the technical field of artificial intelligence and information security. Mainly aiming at a false information detection task in a social media image-text contradiction form, a progressive adversarial training framework with multi-modal consistency constraint is provided. The core content of the method comprises: multi-modal data acquisition and processing: acquiring aligned social media multi-modal data (text, image / video); multi-modal adversarial sample generation: generating an adversarial sample based on text semantic disturbance and a visual contradiction scene; performing cross-modal progressive training, and optimizing model robustness by combining cross-modal cross attention fusion and a progressive three-stage dynamic training strategy; and generating an interpretability analysis result, and outputting an interpretability thermodynamic diagram to position cross-modal logic conflicts. According to the method, the accuracy and robustness of false information detection of the model in a multi-modal scene can be improved.
Owner:SOUTHEAST UNIV

Expert question and answer technical method, system and equipment based on local geological knowledge graph semantic reasoning

The invention discloses an expert question and answer technical method, system and equipment based on local geological knowledge graph semantic reasoning, and the method comprises the following steps: collecting and integrating multi-source heterogeneous geological data, including texts, images and remote sensing, and constructing a unified geological knowledge base covering multi-modal data; on the basis of a pre-training language model, semantic analysis is performed on a geological text, and'entity-relationship-entity 'structured knowledge is automatically generated by utilizing a triple extraction module. The method has field breadth and multidisciplinary fusion, is different from a knowledge graph technology focusing on single fields of mineralogy, geophysics and the like in the prior art, innovatively constructs a large-scale comprehensive knowledge graph system covering the whole geological disciplinary, supports cross-field knowledge extraction and reasoning, and is high in practicability. And complex interdisciplinary geological problems can be handled.
Owner:JIANGSU PROVINCIAL GEOLOGICAL BUREAU BIG DATA CENTER

Multi-modal data driven general report generation method and system based on large model

The invention discloses a multi-modal data driven general report generation method and system based on a large model, and belongs to the technical field of intelligent report generation. Firstly, texts, images and sensor data related to a report theme are obtained and subjected to standardized preprocessing; analyzing the report generation instruction, and matching and querying a task modal mapping library matching modal configuration scheme according to a task demand; quantitatively evaluating the data quality of each modal, dynamically calculating the final decision weight of each modal in combination with the basic weight, and distributing the final decision weight to a corresponding processing path to form dominant, supplementary and reference data; inputting the dominant data and the supplementary data into a multi-modal model for analysis to obtain a preliminary conclusion with confidence score, performing consistency judgment, if no conflict exists, performing fusion to form a comprehensive conclusion, and if the conflict exists, combining a quality evaluation result and an arbitration rule to complete conflict judgment; and finally, inputting the comprehensive conclusion and the reference data into a large language model to generate a report text, and outputting a complete report after typesetting and proofreading.
Owner:NANJING ANCIENT NETWORK TECH CO LTD

Method and System for Optimizing Use of Retrieval Augmented Generation Pipelines in Generative Artificial Intelligence Applications

Systems and methods for dynamic knowledge integration in LLM systems including receiving multimodal input data comprising text, image, audio, video, and / or code data, extracting information by processing the multimodal input data through a document processor, storing the extracted information in a dynamic knowledge base, receiving a user query at a query processor, identifying knowledge domains related to the user query using domain-specific agents, retrieving real-time information from the dynamic knowledge base responsive to the identified knowledge domains, integrating the real-time information into the processing of an LLM by a dynamic knowledge integrator, and generating a response using the LLM with the real-time information.
Owner:MADISETTI VIJAY

Document processing method and system based on text content extraction

The invention relates to a document processing method and system based on text content extraction. The method comprises the steps that an original document containing text, image and format information is received, the encoding format of the document is automatically detected, character set conversion is executed, and hierarchical indexes including page numbers, paragraphs and tables are established for an unstructured document; the method comprises the following steps: synchronously processing text content and visual layout through a pre-trained visual-language model, extracting word-level and sentence-level semantic features by a text stream embedding layer, analyzing spatial distribution features of document elements by a visual encoder, and fusing text and visual features through a cross-modal attention mechanism; and loading the domain knowledge graph matched with the document type, and executing entity linking to associate the text mentions to the knowledge nodes. According to the document processing method and system based on text content extraction, through the synergistic effect of vision-text joint coding and knowledge enhancement, the accuracy of financial contract key clause recognition tasks is improved, the error rate is lower than that of industry benchmark products, and the semantic understanding precision is remarkably improved.
Owner:WIN THE BID HUIKANG TECH CO LTD

Multi-modal heterogeneous knowledge fusion construction and semantic enhancement retrieval system based on large model

The invention relates to the technical field of multi-modal data processing and semantic retrieval, in particular to a multi-modal heterogeneous knowledge fusion construction and semantic enhancement retrieval system based on a large model, which comprises a data acquisition module, a semantic analysis module, a knowledge fusion module and a retrieval optimization module. Multi-modal data such as texts, images and audios are uniformly expressed and deeply analyzed by introducing a large model technology, a knowledge graph is dynamically constructed, a structure is optimized in combination with a user query intention, and meanwhile accurate sorting and screening are achieved through a semantic enhancement algorithm. According to the method, the semantic comprehension capability and the intelligent level of the system can be improved, the real-time and diversified scene requirements are met, and the accuracy and the adaptability of a retrieval result are remarkably enhanced.
Owner:ZHONGYU SOFTCOM (CHONGQING) INFORMATION TECH CO LTD

Relation extraction method and system based on graph neural network

The invention discloses a relation extraction method and system based on a graph neural network, and belongs to the technical field of natural language processing. A target text is obtained, word segmentation, part-of-speech tagging and named entity recognition are carried out, and an entity set is extracted; constructing a text graph structure containing multiple edge types based on the entity set; performing feature coding on nodes in the graph to generate an initial feature vector fusing semantic, part-of-speech and position information; inputting the graph into the graph neural network model, and obtaining high-order node representation through multi-layer message passing and aggregation; modeling the entity pair in combination with the structure path and the context information, and inputting a multi-channel classification network to predict the relationship type of the multi-channel classification network; and finally, outputting an entity relationship triple according to a prediction result. The method has stronger semantic modeling ability and structure expression ability in a relation extraction task, and is suitable for scenes such as knowledge graph construction and information extraction systems.
Owner:CHANGCHUN GUANGHUA UNIV

Financial bill auditing and decision-making method and system, terminal and medium

The invention relates to the field of bill auditing, and particularly provides a financial bill auditing decision-making method and system, a terminal and a medium, and the method comprises the steps: collecting multi-mode finance and tax data including texts, images and structured data, and carrying out the preprocessing and alignment; carrying out feature extraction on the preprocessed data by using a multi-modal large model, carrying out feature fusion by using a dynamic weight distribution algorithm based on the credibility of each modal, the service priority and historical feedback, and obtaining a risk probability through a risk identification neural network; the risk probability is compared with a preset threshold value and rule in an auditing rule base, automatic passing is triggered, after risk abnormity is recorded, passing is conducted, and auditing actions such as manual auditing or starting of a high-risk emergency plan are pushed; and finally, adjusting model weight parameters according to manual feedback information to realize system self-optimization. The auditing decision-making efficiency is improved, and the accuracy and the service adaptability are improved.
Owner:INSPUR GENERSOFT CO LTD

Digital human generation method based on multi-modal large model

The invention provides a digital human generation method based on a multi-modal large model. The method comprises the following steps: constructing a digital human basic model; generating a structured training set; generating a question and answer model supporting multi-channel interaction; semantic answers of the user questions are output, text emotional tendencies of the semantic answers are extracted, and emotional intensity parameters are output; generating facial muscle movement track data, and performing real-time rendering on the digital human basic model according to the facial muscle movement track data to output a digital human three-dimensional image with emotion expression. According to the embodiment of the invention, cross-modal alignment is carried out on text, image and audio data, and a multi-modal large model containing visual, voice and knowledge models is optimized by using a joint training method, so that more natural and smoother multi-channel interaction experience is realized; in addition, by introducing an emotion recognition model and a face interaction model, the emotion tendency contained in the semantic answer can be captured and reflected more accurately, so that a digital human three-dimensional image with real emotion expression is output.
Owner:CHINA NAT BUILDING MATERIALS TECH CO LTD +2

Intelligent archive opening identification method based on large model

The invention discloses an intelligent archive opening and identifying method based on a large model, particularly relates to the technical field of archive data auditing, and is used for solving the problems of insufficient cross-modal data analysis capability, lagging rule updating and low man-machine cooperation efficiency in the prior art. Fusing cross-modal features of texts, images and metadata through a hybrid expert model to generate multi-modal feature vectors, and dynamically allocating the multi-modal feature vectors to a rule network, a semantic network and a domain network for cooperative processing based on attention weights; the rule network parameters are optimized through gradient projection constraint, and regulation-driven real-time adaptation is achieved; matching sensitive data in combination with a multi-dimensional feature matrix of auditing personnel, and optimizing task allocation accuracy; removing redundant links by utilizing value flow analysis to generate a lightweight process, and recording as a tamper-proof evidence chain through a block chain evidence storage solidification operation; the auditing efficiency and accuracy are improved, and the compliance traceability is guaranteed.
Owner:CHONGQING SHIJI KEYI TECH DEV CO LTD

Intelligent agent reasoning method based on knowledge graph

The invention discloses an agent reasoning method based on a knowledge graph, and belongs to the technical field of knowledge graphs, and the method comprises the steps: integrating user query, the knowledge graph and multi-modal data through an NLP model and an entity linking technology, and extracting a target entity, a relation constraint and a structured feature vector. The LLM can synthesize more information to generate a more comprehensive conclusion, the knowledge embedding model maps an entity relationship into a geometric relationship in a vector space, the LLM is assisted to verify reasonability of reasoning, the system can comprehensively generate confidence through the LLM output probability, knowledge embedding similarity and data quality, and the reliability of a result can be explained through confidence score. The LLM generation capability is combined with the vector reasoning capability of knowledge embedding, and the limitation of a single model is made up.
Owner:SUZHOU LAPLACE ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system

The invention relates to document verification, in particular to a heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system, which is used for heterogeneous document input, supports multi-format document input and comprises multi-modal elements including texts, pictures, tables and charts. The document analysis module is used for carrying out structured extraction on document contents; extracting multi-modal elements, identifying and classifying various elements in the document, and establishing position and type labels of a foundation; the knowledge graph construction module is used for uniformly modeling heterogeneous elements into a multi-modal knowledge graph; the graph neural network semantic alignment module is used for realizing accurate cross-modal semantic alignment by using a specially designed graph neural network based on the multi-modal knowledge graph; the hybrid consistency verification engine is used for performing logic consistency verification in combination with a symbol logic verification mechanism and a semantic consistency verification mechanism; according to the method, the defect that accurate cross-modal semantic alignment and logic consistency verification are difficult to carry out on professional documents with multi-modal elements can be effectively overcome.
Owner:ANHUI GAOSHAN TECH CO LTD

Financial data intelligent quality inspection method and system

The invention provides a financial data intelligent quality inspection method and system, and the method comprises the steps: S1, accessing a real-time transaction data flow through a dynamic rule engine, and enabling the dynamic rule engine to dynamically adjust the rule weight through a Bayesian network and reinforcement learning hybrid model; s2, calling a multi-modal LLM verification framework, performing joint semantic analysis on the text, the image and the time series data, and generating a risk early warning signal; s3, identifying a cross-entity risk path based on the financial knowledge graph, and converting the identified risk path into a structured risk report; and S4, a closed loop iteration system is formed according to the weight of the early warning feedback optimization rule. According to the method, full-life-cycle quality management and control of financial transactions can be realized through technical collaboration of real-time data stream processing, multi-dimensional semantic verification and cross-entity risk tracking.
Owner:AACAT TECHNOLOGY LTD