Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

20787 results about "Multi modality" patented technology

Multimodality is a theory which looks at the many different modes that people use to communicate with each other and to express themselves.

Log aggregation fault diagnosis method and system based on artificial intelligence

The invention relates to the field of log fault analysis, in particular to a log aggregation fault diagnosis method and system based on artificial intelligence. The method comprises the following steps: collecting a multi-modal heterogeneous log, carrying out sliding time sequence slicing processing, carrying out time sequence association sequence reconstruction, and constructing a time sequence reconstruction log data stream; log event deep semantic analysis is carried out on the time sequence reconstruction log data stream, event semantic topological evolution is carried out, and a multi-dimensional event topological representation matrix is constructed; performing routine event behavior analysis and abnormal fault mode inference based on the multi-dimensional event topology representation matrix, and marking abnormal fault points; and the occurrence timestamp and the abnormal propagation rate of the abnormal fault point are calculated, fault space-time diffusion evolution is carried out, and a dynamic fault propagation path map is constructed. Through efficient and accurate fault traceability analysis, the fault diagnosis efficiency is greatly improved, and the stability and reliability of log data are improved.
Owner:SHANGHAI FEIWEI INFORMATION TECH CO LTD +2

Multi-modal medical image data intelligent processing system

The invention discloses a multi-modal medical image data intelligent processing system, relates to the field of medical image analysis, and is applied to multi-modal medical image whole-process analysis of CT, MRI, PET, ultrasound and the like. According to the system, different modal image features are extracted and fused through a cross-modal manifold fusion network; a semantic guidance dynamic registration engine optimizes registration parameters to ensure that the registration error is less than or equal to 1.5 mm; the multi-task collaborative diagnosis network realizes multiple tasks such as disease classification; the clinical knowledge embedding and interpretable module generates a structured report and is in butt joint with an HIS system. Meanwhile, the model is optimized through a federated learning architecture, the adaptability of newly added data is improved by more than or equal to 20%, and intelligent processing and analysis of multi-modal medical images are realized.
Owner:SHANDONG JUNKANGLIN MEDICAL TECHNOLOGY CO LTD

Multi-modal data processing method and apparatus, electronic device, computer-readable storage medium, and computer program product

Disclosed in the present application are a multi-modal data processing method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring a reference image and a reference text; extracting a reference visual feature of the reference image; by means of a multi-modal large language model, determining an embedding of the reference text, an embedding of a start mark of the reference visual feature, an embedding of the reference visual feature, and an embedding of an end mark of the reference visual feature; on the basis of the multi-modal large language model, splicing the embedding of the reference text, the embedding of the start mark, the embedding of the reference visual feature, and the embedding of the end mark into a target embedding sequence, performing attention processing on the basis of the embedding of the start mark, the embedding of the end mark, and an embedding selected by a sliding window in the target embedding sequence, and outputting a predicted sequence; and generating a predicted image and a predicted text on the basis of the predicted sequence.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Multi-modal document generation method and device based on multi-agent collaboration

The invention belongs to the field of natural language processing, particularly relates to a multi-modal document generation method and device based on multi-agent collaboration, and aims to solve the problems that an existing method is low in intention recognition accuracy, limited in retrieval range and not professional enough in content generation. The method comprises the following steps: generating a structured template; analyzing the text input by the user to identify a writing intention, and determining a target template; vectorizing each candidate resource feature to obtain a corresponding sparse vector, a dense vector and a knowledge vector, and performing semantic alignment; extracting context features of the input text, respectively performing multi-path retrieval recall, evaluating and sorting recall results, and screening out target features; and constructing a thinking chain in combination with the knowledge graph, and generating a multi-modal document according to the target template. According to the method, the outline structure can be extracted, the picture / table style can be recognized, the templates adaptive to different document types can be dynamically generated, and full-process automation from user input to document output is achieved.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Multi-modal visual fusion complex scene small target detection tracking method and system

The invention discloses a multi-modal visual fusion complex scene small target detection tracking method and system, and relates to the technical field of unmanned aerial vehicle target tracking, and the method comprises the steps: employing a visible light camera, an infrared thermal imager and a laser radar sensor which are carried on an unmanned aerial vehicle platform, and synchronously collecting RGB images, thermal infrared images and point cloud data; the consistency of the multi-modal data is ensured through data preprocessing and space-time alignment; constructing a lightweight double-branch network to extract multi-scale features, generating a fusion feature map by adopting adaptive weighted fusion, and generating depth information by utilizing point cloud to assist in scale estimation; a small target detection head is designed based on the fusion feature map, and precise detection is realized in combination with a feature pyramid network, adaptive scale prediction and a context awareness suppression mechanism; furthermore, through multi-mode cooperative tracking, including target association, spatio-temporal context modeling, trajectory prediction and a re-detection mechanism, tracking continuity is ensured.
Owner:BEIJING INSTITUTE OF GRAPHIC COMMUNICATION

Perception collaborative decision-making method and system based on multi-modal heterogeneous data fusion

The invention provides a perception collaborative decision-making method and system based on multi-modal heterogeneous data fusion, and relates to the technical field of artificial intelligence, and the method comprises the steps: inputting a global environment situation perception graph into a pre-trained multi-target collaborative decision-making model; the multi-target collaborative decision-making model forms a multi-target decision-making feature set by analyzing the resource entities and the incidence relation in the graph; based on the multi-target decision feature set, decision optimization is carried out to obtain a comprehensive collaborative scheduling scheme; performing instruction analysis and packaging on the comprehensive collaborative scheduling scheme to obtain an executable instruction sequence; and issuing the executable instruction sequence to a corresponding decision node and a control terminal in parallel through a distributed communication architecture to complete real-time scheduling of resources and collaborative issuing of control instructions. According to the invention, by constructing a linkage mechanism of multi-modal data fusion, dynamic environment perception and collaborative decision execution, intelligent perception and quick response to a complex environment are realized.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Multi-agent dynamic arrangement method based on multi-modal analysis and adaptive retrieval

The invention discloses a multi-agent dynamic arrangement method based on multi-modal analysis and adaptive retrieval, and relates to the technical field of artificial intelligence and information retrieval. Comprising the steps of S1, converting a text, an image, structured data and voice content input by a user into a unified multi-mode semantic representation, S2, converting the unified multi-mode semantic representation into a specific execution process, and S3, automatically scheduling a reasoning agent, a knowledge obtaining agent and an execution agent according to DAG nodes, task elements and available resources, and obtaining the task elements and the execution agent according to the reasoning agent, the knowledge obtaining agent and the execution agent. S4, after task process construction and agent arrangement are completed, dynamic retrieval, evidence convergence and strategy optimization are carried out on information requirements related to a user task, so that a reasoning agent obtains complete knowledge support with consistent context, and S5, knowledge evidence is combined with a task process, so that the task process is completed. The method comprises the following steps: step S6, implementing problem solving, strategy generation and task closed-loop execution through a reasoning agent, step S6, performing actual operation on a target task by an execution agent according to an executable instruction sequence output by the reasoning agent, and outputting a result, and step S7, performing result verification according to an output result returned by the execution agent, and the correctness, integrity and consistency of an output result are examined through rule verification, model evaluation and evidence alignment.
Owner:INSPUR GROUP CO LTD +1

Mine abnormal event real-time identification method and system based on time sequence characteristics

The invention provides a mine abnormal event real-time identification method and system based on time sequence characteristics, and relates to the technical field of mode identification, and the method comprises the steps: carrying out the time-space alignment and semantic annotation of multi-modal monitoring data, and constructing a time sequence knowledge graph; calculating a dynamic association weight between entities, and analyzing a risk propagation path; predicting a risk situation based on a historical evolution rule; and dynamically generating a differential early warning strategy and establishing a closed-loop tracking system. According to the invention, early identification, accurate prediction and efficient disposal of mine safety risks can be realized, and the mine safety management level is improved.
Owner:BEIJING YANGGUANG JINLI TECH DEV

Artificial intelligence-powered large-scale content generator

An AI-powered content generation system that creates consistent, coherent, and engaging multi-modal content by integrating multiple specialized AI components. The system analyzes user input, identifies key elements, and maintains continuity throughout the generation process. It incorporates a feedback loop to learn and adapt based on user preferences, enabling personalized content experiences. The modular architecture allows for seamless integration of AI components focusing on text, images, audio, and interactive elements. The system ensures consistency across modalities and over extended periods, while managing rights, licenses, and royalties using blockchain technology. This advanced platform revolutionizes content creation, consumption, and management in the digital age.
Owner:QOMPLX INC

Multi-mode large model interpretable diagnosis method and system for wind turbine generator

The invention discloses a multi-modal large model interpretable diagnosis method and system for a wind turbine generator, and relates to the technical field of wind turbine generator fault diagnosis, comprising the step of combining multi-modal data (vibration, time sequence, image and text) and topological information to realize fault diagnosis through cross-modal contrast learning and topological modeling. The method comprises the steps of multi-modal feature extraction, standardization and alignment, and feature fusion through topology embedding optimization and a cross-modal attention mechanism. In the fault diagnosis process, dynamic correction and path reliability evaluation are introduced by using a regular Agent and a topology consistent Agent, weighted fusion is performed on each modal feature and a topology structure, and finally an accurate fault type and a component positioning result are output. Through combination of knowledge retrieval and a multi-Agent decision model, the adaptability and precision of fault diagnosis are improved, especially in a complex environment, the fault mode of the wind turbine generator can be effectively identified, and the system reliability is improved.
Owner:BEIJING INST OF TECH

AI generation content detection and review method and device, equipment and storage medium

The invention discloses an AI generation content detection and review method, device and equipment and a storage medium, and the method comprises the steps: receiving multi-modal input data which comprises text data, image data and video data; calling a large language model to carry out compliance analysis on the text data to obtain an analysis result; when the analysis result is compliance, calling a multi-modal model to identify and detect the image data and the video data, and respectively obtaining an image detection result and a video detection result; and aggregating the analysis result, the image detection result and the video detection result to obtain a content detection result. According to the method, the processing paths are automatically allocated according to the content types, redundant calculation is avoided, hardware consumption is remarkably reduced, the multi-modal detection results are aggregated into structured output, the large-batch content processing efficiency is remarkably improved, calculation resource occupation is greatly reduced, and seamless integration of the detection results and a downstream service system is achieved.
Owner:深圳市维卓数字营销有限公司

Multi-agent-based gas insulated switchgear fault diagnosis method and system

The invention discloses a multi-agent-based gas insulated switchgear fault diagnosis method and system, and relates to the technical field of intelligent operation and maintenance of power equipment, and the method comprises the steps: obtaining signal data of target equipment, carrying out the feature extraction of the signal data, and constructing a multi-modal feature matrix; time delay features of acoustic and electromagnetic signals are extracted from the multi-modal feature matrix, a GIS propagation model is established, and the space coordinate position of a liberated power source is solved through a wave field inversion algorithm; combining the space coordinate position and the multi-modal feature matrix into a complete fusion feature vector, inputting the fusion feature vector into a dynamic Bayesian model, and outputting a fault type label and a corresponding confidence coefficient; migrating the dynamic Bayesian model based on a migration learning mechanism, and dynamically updating a classification threshold value; inputting the diagnosis history sequence into a time sequence prediction model, and predicting a future operation state; through multi-modal fusion and intelligent reasoning, GIS fault accurate positioning and prediction are realized, and the problems of low precision and poor adaptability of traditional diagnosis are solved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Dam safety perception fusion association method based on multi-modal space-time diagram neural network

The invention provides a dam safety perception fusion association method based on a multi-modal space-time diagram neural network. The method comprises the following steps: dividing a dam into a plurality of structural units, and mapping various data into a three-dimensional coordinate system; a heterogeneous graph structure is defined, and a dynamic adjacency matrix is calculated based on the real-time stress gradient so as to reflect physical connection, mechanical conduction and geological association relationships among nodes; carrying out fusion modeling on multi-source data in the heterogeneous graph structure by utilizing a multi-modal space-time diagram neural network, constructing a causal inference engine based on an output result of the multi-modal space-time diagram neural network, and updating a three-level modeling system through structural equation modeling, anti-factual inference and dynamic weight to obtain the heterogeneous graph structure. According to the method, the dynamic coupling rule among the dam structure, geology and material states is excavated, cross-modal space-time fusion of manual inspection and sensor monitoring data can be realized, the early recognition capability and early warning accuracy of dam potential safety hazards are improved, and the problems of data islands and insufficient relevance in a traditional monitoring method are effectively solved.
Owner:HUANENG SICHUAN HYDROPOWER CO LTD +2

CAD automatic generation system and method based on intelligent model selection and application

The invention discloses a CAD automatic generation system and method based on intelligent model selection and application, and aims at achieving automatic modeling under the multi-modal design requirement. The system comprises a user interaction module for receiving multi-modal input such as natural language, sketch and voice; the intelligent demand analysis module is used for combining an industrial large language model and a product knowledge graph, combining semantic analysis and generating a structured demand; the intelligent model selection calculation module is used for matching the optimal parameter combination and the component list based on a multi-objective optimization algorithm; the CAD automatic generation module calls a parametric modeling engine to generate an editable three-dimensional model; the constraint solving module is used for processing hard constraints and soft constraints in real time and dynamically adjusting model parameters; and the model output and interaction module feeds back a design state and supports user iteration. The system realizes full-process automation from the design intention to the CAD model, improves the design efficiency and accuracy, and is suitable for the fields of mechanical design, intelligent manufacturing and the like.
Owner:HOFMANN (BEIJING) ENG TECH CO LTD

Electric power engineering multi-mode RAG system based on knowledge graph and multi-Agent cooperation

The invention relates to the technical field of electric power engineering, and discloses an electric power engineering multi-modal RAG system based on a knowledge graph and multi-Agent collaboration, and the system comprises a multi-modal dynamic knowledge base construction module which is configured to carry out the structural processing, multi-dimensional knowledge organization and dynamic optimization of electric power engineering multi-modal data; the self-adaptive retrieval strategy engine module is configured to construct a weight decision network based on deep reinforcement learning and execute multi-channel parallel retrieval and result fusion; the iterative self-reflection reasoning module is configured to generate a reasoning path in combination with the retrieval result and verify evidence validity from multiple dimensions; the MCP tool intelligent calling module is configured to integrate multiple types of standardized MCP tools; and the multi-Agent collaborative framework is configured to provide multiple types of Agents which are specific in function and have a cross-module interaction capability. According to the method, the question and answer accuracy, the reasoning depth and the result interpretability in the complex multi-modal scene of the electric power engineering can be remarkably improved.
Owner:SOUTHWEST ELECTRIC POWER DESIGN INST OF CHINA POWER ENG CONSULTING GROUP CORP

Engineering document index consistency proofreading method and system based on multi-modal large model

The invention relates to an engineering document index consistency proofreading method and system based on a multi-modal large model, and the method comprises the steps: Q1. OCR detection and recognition: carrying out the optical character recognition and format analysis of a source document, converting an uploaded PDF document into a processable text message in a Markdown format, and carrying out the format discrimination of a table, a formula and a plain text; and Q2, table and formula processing: adopting a hierarchical processing strategy, intelligently selecting an optimal processing mode according to the complexity of the table, and converting table information into a descriptive long text through a language large model and cue words. According to the method, accurate, reliable and efficient document index checking service can be provided for a user, the quality and efficiency of professional document processing are remarkably improved, the efficiency and quality of knowledge graph construction are remarkably improved, a knowledge verification system capable of being evolved continuously is established, and the method is suitable for popularization and application. And a reliable technical support is provided for knowledge management and professional decision-making in a complex field.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Robot multi-modal fusion autonomous decision-making method and system based on large language model

The invention relates to the technical field of robot decision making, and provides a robot multi-modal fusion autonomous decision making method and system based on a large language model.The method comprises the steps that a robot obtains multi-modal environment information through a visual sensor, a touch sensor, an auditory sensor and a laser radar which are carried by the robot; performing preliminary filtering and noise reduction processing on the original sensor data, and synchronously recording all the sensor data by timestamps; performing space-time semantic alignment on the preprocessed multi-modal data, mapping pixel coordinates of a target in a visual target coordinate quantization original image to a robot coordinate system, performing uncertainty evaluation on a multi-modal signal through a dynamic Bayesian network, and taking entropy or variance as an uncertainty quantitative evaluation index. According to the method, the information quality is improved from a data fusion source, accurate and reliable basic support is provided for subsequent decision making, and decision making errors caused by data deviation are greatly reduced.
Owner:ANHUI UNIV +1

Intelligent substation safety measure checking method and system

The invention discloses an intelligent substation safety measure checking method and system, and the method comprises the steps: obtaining a secondary system topological structure, equipment information and historical safety measure ticket data of a substation, and constructing a quaternary knowledge graph; according to the quaternary knowledge graph, extracting multi-modal features of the maintenance task and identifying the type of a maintenance scene by combining a memory guide reflection decision reasoning mechanism, optimizing a rule reasoning process through a distributed guide local search algorithm, and generating an optimal safety measure operation set for a specific maintenance scene; automatically identifying a correlation loop and determining a minimum safety isolation range through a memory guide decision reasoning mechanism, dynamically generating a minimum safety operation set according to the real-time state of the equipment, and adaptively adjusting operation steps; and an operation dependency relationship model is constructed in combination with unwrapping variational multi-graph representation learning and a distributed guide local search algorithm, and operation sequence compliance verification, risk level assessment and dynamic visual early warning are realized. According to the invention, accurate formulation, dynamic adjustment and risk early warning of safety measures of the intelligent substation are realized.
Owner:GUIZHOU ANRONG TECH DEV CO LTD +2

Multi-modal dynamic fusion and incremental learning fault diagnosis method for deep vertical shaft equipment

The invention discloses a multi-modal dynamic fusion and incremental learning fault diagnosis method for deep vertical shaft equipment, which belongs to the technical field of industrial equipment fault diagnosis, and comprises the following four steps of: constructing a pre-training large model to perform feature extraction, and relying on a multi-layer Transformer encoder and a dual loss function, establishing a multi-modal dynamic fusion and incremental learning fault diagnosis model; mining cross-modal universal fault features from vibration, temperature and current multi-modal time sequence data; according to the method, multi-modal features are fused, multi-modal association is constructed, modal weights are dynamically adjusted through a modal gating unit and a time delay compensation attention mechanism to adapt to signal quality changes, and meanwhile time sequence deviation is corrected to achieve accurate association; incremental learning is realized by using a decoupling projection layer, and a lightweight projection module is designed for a newly added fault task to suppress disastrous forgetting; network training is optimized, pre-training loss, incremental learning loss and attention regularization loss are integrated through a multi-objective loss function, and model stability and diagnosis precision are improved. The method has the advantage that the model stability and the diagnosis precision are improved.
Owner:CHINA COAL NO 5 CONSTR +1

Domain intelligent question-answering method and system based on multi-modal knowledge graph and RAG

The invention relates to the technical field of intelligent questioning and answering, in particular to a domain intelligent questioning and answering method and system based on a multi-modal knowledge graph and RAG, and the method comprises the steps: constructing a concept layer knowledge graph based on a directory structure of a domain multi-modal document, and constructing an instance layer knowledge graph based on document content; obtaining a user question, pruning and positioning the user question in combination with the concept layer knowledge graph and the thinking chain, and determining a target chapter; splitting the question into sub-questions through intention analysis, and performing semantic retrieval in the instance layer knowledge graph corresponding to the target chapter to obtain a graph retrieval result; optimizing the original problem based on the atlas retrieval result, and executing semantic retrieval in a vector database to obtain a vector retrieval result; and fusing the atlas retrieval result and the vector retrieval result to generate a preliminary answer, and performing iterative optimization until a final answer is generated. According to the method, the semantic coverage, the expression accuracy and the response efficiency of the vertical domain question-answering system are remarkably improved by constructing the multi-modal knowledge graph and optimizing the retrieval process.
Owner:HENAN UNIVERSITY

Remote sensing image semantic segmentation method based on CNN-Transform-SAM dynamic collaboration and scene adaptation

The invention discloses a remote sensing image semantic segmentation method based on CNN-Transform-SAM dynamic collaboration and scene adaptation, and a constructed remote sensing image segmentation network comprises a scene attribute analysis module, a dynamic backbone decision module, a CNN-Transform expert sub-network, a cross-modal feature calibration module, a multi-modal prompt generator and an SAM adaptive general sub-network. And all the modules realize dynamic collaboration through data interaction. Wherein the scene attribute analysis module analyzes image resolution, spectrum and target scale attributes, the dynamic backbone decision-making module matches the optimal feature extractor according to the image resolution, spectrum and target scale attributes, the CNN-Transform expert sub-network generates small target enhanced adaptive masks through multi-scale interaction and up-sampling refinement, the cross-modal feature calibration module optimizes the masks and semantic distribution to generate alignment masks, and the cross-modal feature calibration module outputs the alignment masks. And the multi-modal prompt generator generates a multi-modal optimization prompt set based on the alignment mask, and guides the SAM adaptive universal sub-network to complete segmentation. The method effectively solves the problems of poor small target segmentation, fuzzy boundary and lack of remote sensing exclusive semantic priori in the prior art.
Owner:HOHAI UNIV

Large model illusion suppression method, system and equipment based on dynamic knowledge base and multi-modal consistency constraint

The invention belongs to the field of artificial intelligence, particularly relates to a large model illusion suppression method, system and equipment based on a dynamic knowledge base and multi-modal consistency constraint, and aims at solving the problem that factual illusion is likely to occur when an existing large language model generates content. The method comprises the steps that a knowledge base of multi-source heterogeneous data is constructed and dynamically maintained, and a dynamic credibility weight fusing data source authority, knowledge timeliness and multi-modal consistency is calculated for each piece of knowledge in the knowledge base; when the content is generated by the model, high-credibility related knowledge is retrieved from the knowledge base according to the current context; in the decoding stage of the model, a constraint loss item is designed, and the generation probability is adjusted in real time by calculating the similarity between the currently generated content and the retrieval knowledge in the feature space. According to the method, the multi-modal knowledge base for dynamic credibility evaluation is introduced, and the real-time consistency constraint is applied in the generation and decoding link, so that the accuracy and the reliability of the generated content are remarkably improved.
Owner:ZIGUANG HENGYUE TECH CO LTD +1

Large-scene monitoring video abnormal event early warning method based on multi-modal large model

The invention relates to the technical field of abnormal event early warning, and provides a large-scene monitoring video abnormal event early warning method based on a multi-mode large model. According to the invention, the problems of delay, low accuracy and limited coverage range of abnormal event early warning of large-scene monitoring videos in the prior art are solved. According to the main scheme, multiple paths of high-resolution monitoring videos are spliced and preprocessed to generate a panoramic video; synchronously acquiring and preprocessing audio and sensor data to construct a multi-modal data set; video key frames are extracted by adopting a traditional small model, and the video key frames and multi-modal data are jointly input into a multi-modal large model based on a Transform architecture for deep feature fusion; abnormal events such as tumble, congestion and fight are identified based on the fusion features; triggering an early warning mechanism to send event type and position information in real time; and storing the full-dimensional data of the abnormal event for tracing analysis. The real-time processing performance is optimized through edge calculation, the complex scene understanding ability is enhanced in combination with a multi-modal large model, and the detection precision and the response speed are remarkably improved.
Owner:PEKING UNIV (TIANJIN BINHAI) NEW GENERATION INFORMATION TECH RES INST +1

Retrieval enhancement method based on multi-modal data fusion and modal perception

The invention relates to the technical field of information retrieval and generation, in particular to a retrieval enhancement method based on multi-modal data fusion and modal perception. According to the method, firstly, a dual-channel architecture is adopted to perform feature extraction and coding on a text and an image respectively, and mutually independent embedded representation spaces are constructed, so that high-quality collaboration and matching of cross-modal representation are realized; and a pseudo-pairing generation mechanism is introduced to effectively mine and reconstruct the existing non-paired data in the knowledge base. And designing a query modal perception and dynamic weighting mechanism for accurately controlling the fusion proportion of the image-text bimodal information in the retrieval stage so as to match the modal demand difference of different query contents. And further executing aggregation retrieval and reordering of the cross-modal information by using dynamic weighted fusion retrieval to generate a candidate set of multi-modal responses. According to the method, accurate matching and dynamic weight adjustment of the image-text content are realized, and the accuracy and expression integrity of the generated content are improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Multi-modal remote sensing semantic segmentation method and system for learning frequency domain fusion

The invention discloses a multi-modal remote sensing semantic segmentation method and system for learning frequency domain fusion. The method comprises the following steps: respectively extracting multi-scale features of two modal input images by adopting a double-branch encoder; sequentially executing frequency domain decoupling and fusion, mutual information constraint-based feature optimization and low-frequency guided cross-modal fusion processing on each scale feature to generate a fused semantic feature; and performing up-sampling and feature refining on the fused features through a decoder, and outputting a full-resolution segmentation prediction map. According to the multi-modal remote sensing image semantic segmentation method, modal sharing information and specific details are effectively separated through frequency domain decoupling, feature representation is optimized through mutual information constraint, adaptive feature fusion is achieved in combination with an attention mechanism, and the accuracy and robustness of multi-modal remote sensing image semantic segmentation are remarkably improved.
Owner:NORTHEAST FORESTRY UNIV

Multi-mode ultrasonic fusion pressure vessel welding seam defect nondestructive testing method and multi-mode ultrasonic fusion pressure vessel welding seam defect nondestructive testing system

The invention provides a multi-mode ultrasonic fusion pressure vessel weld defect nondestructive testing method and system, and relates to the technical field of nondestructive testing. According to the method, geometric parameters of a welding seam are obtained through three-dimensional laser scanning, and an optimal scanning parameter set is generated; driving ultrasonic phased array equipment to scan for one time and synchronously acquire shear wave full-matrix capture and longitudinal wave linear scanning data; performing energy flow angular spectrum analysis and envelope analysis on the bimodal data, extracting defect feature parameters and constructing a three-dimensional feature tensor; carrying out multi-dimensional feature fusion by adopting Tucker decomposition, and enhancing a core tensor through physical modeling; generating three types of defect indication diagrams including a defect existence possibility diagram, a defect relative scale diagram and a defect space orientation diagram from the enhanced feature tensor; and the three types of indication diagrams are visually presented for comprehensive interpretation of detection personnel. Through multi-modal data fusion and physical modeling enhancement, the defect identification accuracy and detection efficiency are remarkably improved, the false alarm rate is reduced, and reliable technical support is provided for pressure vessel welding seam safety detection.
Owner:YUNNAN SPECIAL EQUIP SAFETY TESTING RES INST

AI-based composite insulator internal defect ultrasonic detection method

The invention relates to the technical field of artificial intelligence, and discloses an AI-based composite insulator internal defect ultrasonic detection method, which comprises a multi-mode ultrasonic probe array module, a signal preprocessing module, an AI defect analysis module, a dynamic parameter optimization module, an edge calculation module and a visual report module, the method comprises the following steps: acquiring a full-dimensional signal through a multi-modal ultrasonic probe array, and inputting the full-dimensional signal into a deep space-time convolutional neural network for defect recognition after adaptive noise reduction and feature fusion; the detection precision is improved by combining dynamic waveform matching and multi-physics coupling analysis; model lightweight and real-time processing are realized by adopting transfer learning and edge calculation. The system integrates the functions of parameter adaptive optimization, three-dimensional visualization and Internet of Things cooperation, solves the problems of low efficiency and high false detection rate of a traditional detection method, and improves the intelligent level and engineering applicability of composite insulator defect detection.
Owner:超创数能科技有限公司 +2

Automobile body innovative design system based on multi-modal knowledge

The invention discloses a multi-modal knowledge-based automotive body innovative design system, which comprises a multi-modal data fusion module, a multi-modal data fusion module, a multi-modal data fusion module, a multi-modal data fusion module and a multi-modal data fusion module, wherein the multi-modal data fusion module is used for receiving text data, picture data, a three-dimensional CAD (Computer Aided Design) model file and an engineering symbol expression from an automotive body design process and is used for carrying out feature extraction and semantic alignment on input data of four modals; generating a unified semantic vector representation; the cross-modal knowledge mining and graph construction module is used for extracting multi-level entities and relationships from the unified semantic vector and constructing a dynamically weighted multi-modal knowledge graph; the large-model-driven multi-hop collaborative reasoning module is used for analyzing the multi-modal design requirement of a user, carrying out multi-hop reasoning on a knowledge graph, and outputting design parameter recommendation and an interpretable reasoning chain. The method aims at breaking through the limitation of an existing design system in the aspects of multi-modal processing and shallow semantic understanding, and deep fusion and intelligent application of multi-source heterogeneous design data are achieved.
Owner:CHONGQING UNIV

Multi-modal large language model fine tuning method, system, equipment and medium

The invention relates to a multi-mode large language model fine tuning method, system and device and a medium, and belongs to the technical field of artificial intelligence and computer vision crossing. The fine tuning method comprises the steps that an original business scene image is acquired and preprocessed, and a preprocessed image is obtained; performing bounding box coordinate labeling and semantic label definition on the entity target in the preprocessed image through a labeling tool, and outputting a structured labeling file; based on the preprocessed image and the structured annotation file, constructing a training sample set comprising multiple rounds of image-text dialogues; loading the pre-trained multi-modal large language model, configuring low-rank matrix decomposition parameters, and generating a fine tuning instruction set; and inputting the training sample set into a pre-trained multi-modal large language model, carrying out joint training operation based on the fine tuning instruction set, and outputting the fine-tuned multi-modal large language model. According to the method, the identification accuracy, the interaction capability and the system availability of the visual question-answering system in an actual application scene are improved.
Owner:GOLDEN TIMES CULTURE COMM

Power operation risk identification method, system and device based on multi-modal data fusion and storage medium

The invention relates to the technical field of power grid monitoring, in particular to a power operation risk identification method, system and device based on multi-modal data fusion and a storage medium. In order to solve the problems of multi-modal data splitting, topological constraint missing and the like in traditional power disturbance analysis, an improved BERT model is constructed, and electrical signal time-frequency features and text semantic information are mapped to a unified vector space through a multi-modal embedding mechanism; a time sequence attention mechanism is adopted to establish a time dependency relationship between signals and texts, and a graph attention network is combined to realize risk propagation modeling under power grid topology constraints; and collaborative optimization of disturbance classification, risk prediction and trend analysis is carried out through a multi-task learning framework. The technical problems that heterogeneous data fusion is difficult and risk identification precision is insufficient are effectively solved, accurate identification and intelligent early warning of electric power operation risks are achieved, and the safe operation level of a power grid is improved.
Owner:YUNNAN POWER GRID CO LTD KUNMING POWER SUPPLY BUREAU