Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

4228 results about "Multi modal fusion" patented technology

High-performance loosely-coupled multi-modal data fusion system for smart driving environmental perception system and vehicle-mounted device

Disclosed are a high-performance loosely-coupled multi-modal data fusion system for a smart driving environmental perception system and a vehicle-mounted device, comprising: a fusion detection model based on a modality-independent feature interaction strategy, which is configured for converting a LiDAR point cloud, a camera image, and a millimeter-wave radar point cloud into a unified bird's-eye view representation, and performing multi-modal fusion; and a fusion tracking model based on a motion-appearance feature cascaded coupling data association strategy, which is configured for performing subsequent trajectory tracking and matching according to multi-modal fusion feature information. A VoD data set and a K-Radar data set are selected for training, verifying, and testing the comprehensive performance of the models, and a TensorRT accelerated inference model is applied, then quantized, and deployed to a vehicle-mounted computational testing platform. The present invention is compatible with mainstream sensor deployment solutions, and achieves the efficient complementary fusion of multi-source heterogeneous sensor information, significantly improving the reliability, accuracy, and adaptability of vehicle-mounted perception systems, thereby effectively responding to extreme operating conditions such as complex traffic scenarios and inclement weather.
Owner:JIANGSU UNIV

LED display defect prediction and process adjustment method and system based on multi-modal fusion

The invention relates to the technical field of LED display, solves the problem that the existing LED display defect detection and parameter adjustment technology is lack of multi-modal information fusion and intelligent process control capability and is difficult to meet the quality control requirement of a high-precision display product, and provides an LED display defect prediction and process adjustment method and system based on multi-modal fusion. The method comprises the following steps: performing multi-modal data fusion processing on optical image data, electrical test data and thermal infrared imaging data corresponding to a to-be-tested LED display screen to obtain fused data; inputting the fused data into a pre-trained defect recognition model to obtain a defect recognition result; according to a process parameter adjustment strategy corresponding to the defect identification result, adjusting the original process parameter to obtain a target process parameter; and according to the target process parameters, process flow correction processing is carried out, and a qualified LED display screen is produced. According to the method, the defect identification precision is improved, and the quality control requirement of high-precision LED display screen production is met.
Owner:XIAMEN PROD QUALITY SUPERVISION & INSPECTION INST +1

Rare disease knowledge graph construction method based on modal injection and multi-modal fusion

The invention relates to the technical field of medical artificial intelligence and knowledge graph construction, in particular to a rare disease knowledge graph construction method based on modal injection and multi-modal fusion. Comprising the following steps: S1, collecting multi-modal medical information including texts, images and genes; s2, standardization processing is carried out, and a three-layer metadata structure is constructed; s3, complementing missing modal data, and performing feature extraction and unified dimension conversion on the modal data to realize representation alignment in a shared semantic space; s4, performing multi-level semantic fusion to obtain a unified fusion semantic vector; and S5, constructing a double-layer structure system rare disease knowledge graph comprising an ontology layer and an instance layer. According to the method, multi-modal medical information of texts, images and genes is selected to construct the knowledge graph of the rare disease, the application range, coverage and accuracy of the knowledge graph are improved, correspondence adaptation of rare cases during clinical diagnosis and treatment of the rare disease can be achieved, and the method has high recognition capacity.
Owner:湖南工商大学

Submarine cable risk dynamic assessment method and system based on multi-modal deep learning

The invention discloses a submarine cable risk dynamic assessment method and system based on multi-modal deep learning, and belongs to the field of marine infrastructure operation and maintenance. Aiming at the problems of incomplete data coverage, unreal generated scene, low evaluation reliability and the like in the prior art, the method comprises the following steps of: 1) constructing a multi-source heterogeneous data set containing six types of data including geology, ocean, ships, biology and the like, and realizing data alignment by adopting space-time grid coding; 2) designing a physical constraint generative adversarial network, and generating risk scene data conforming to a fluid mechanics law through a Navier-Stokes equation constraint; 3) creating a hierarchical space-time fusion network (HST-Transform), and combining CNN spatial feature extraction, a time sequence attention mechanism and a dynamic memory module to realize multi-modal fusion; according to the method, the detection rate of rare risk events is increased by 62%, the evaluation accuracy rate reaches 91.7%, the false alarm rate is reduced by 34% compared with a traditional method, and submarine cable breakage accidents can be effectively prevented.
Owner:GUANGDONG POWER GRID CO LTD

Distributed intelligent authentication method based on dynamic multi-modal fusion

A distributed intelligent authentication method based on dynamic multi-modal fusion relates to the field of network security, and adopts an alliance chain + DAG hybrid block chain architecture, combines a threshold signature to realize secret key fragment management, and switches among PBFT, Raft and probabilistic algorithms through a dynamic consensus mechanism to improve authentication efficiency. The multi-mode authentication module is based on a dynamic weight distribution algorithm, integrates biological characteristics, behavior analysis, equipment fingerprints and environmental factors, and combines an LSTM-GAN model and a quantum random number driven challenge-response mechanism to realize zero-trust verification under environmental perception. The session management module generates a session key by using a chaotic mapping algorithm. In the aspect of privacy protection, CKKS homomorphic encryption, zero-knowledge proof and attribute-based encryption are fused. According to the method, the block chain technology, the secure multi-party computing technology, the machine learning technology and the quantum cryptography technology are fused, and a high-performance, high-security and strong-privacy-protection distributed authentication solution is provided.
Owner:JINLING INST OF TECH

Adaptive teaching real-time feedback method based on multi-modal fusion

The invention discloses an adaptive teaching real-time feedback method based on multi-modal fusion, and the method comprises the following steps: synchronously collecting text modal information, voice modal information and image modal information generated by students in a teaching process, forming multi-modal original data information, and extracting historical student interaction behavior data; preprocessing the multi-modal original data information, and respectively generating corresponding text, voice and image sequence features; a visual feature encoder and a sequence feature encoder are adopted to encode each modal sequence feature to obtain a high-dimensional feature; inputting the modal high-dimensional features into a cross-modal fusion network for deep fusion; parameters of the feedback model are optimized through a model-independent element learning feedback regulation and control algorithm, and a personalized feedback strategy is generated; generating comprehensive feature representation according to the fusion features, and outputting personalized teaching feedback; and the interaction information is updated based on the feedback behavior data to realize closed-loop optimization.
Owner:JIANGSU LINGSHU YOUZHI TECHNOLOGY CO LTD

Intelligent analysis method based on medical document structure perception and multi-modal fusion

An intelligent analysis method based on medical document structure perception and multi-modal fusion comprises the following steps: carrying out structure topology modeling on a medical document, extracting visual layout, text meta-information, space coordinates and semantic keyword features, constructing a semantic topological graph and dynamically shielding irrelevant contents; selecting an extraction path according to a document type, performing deep semantic analysis and entity recognition on a text-type document, and performing visual enhancement OCR recognition on a scanning-type document; the features are injected into a medical knowledge graph, and feature fusion, semantic verification, relation reasoning and information completion are achieved through a graph neural network; a three-stage strategy optimization model of basic pre-training, domain adaptation and online reinforcement learning is adopted; and large-scale processing is realized through a dynamically aggregated distributed architecture. The method is used for intelligent analysis and structured conversion of documents of hospitals, medical insurance and medical scientific research. The problems that heterogeneous medical document analysis adaptability is poor, multi-modal fusion is difficult, medical knowledge utilization is insufficient, and large-scale processing efficiency is low are solved.
Owner:NORTHWEST UNIV

Multi-modal fusion and reinforcement learning collaborative retrieval enhancement generation method and system

The invention relates to the technical field of information retrieval, and discloses a multi-modal fusion and reinforcement learning collaborative retrieval enhancement generation method and system. The method comprises the following steps: receiving an original query input by a user, and generating a sub-query based on a large language model in combination with a multi-modal context of a current iteration step; forming a current state in combination with the sub-query and the multi-modal context, modeling a retrieval enhancement generation task as a Markov decision process, and adaptively selecting an optimal action from a predefined action set in the current state by utilizing a large language model according to a decision strategy; executing a corresponding multi-modal retrieval operation according to the optimal action, fusing the obtained multi-modal information, generating an intermediate answer or a final answer of the sub-query, and updating a multi-modal context by using the intermediate answer; off-line training optimization is carried out on the large language model through imitation learning and a calibration chain, and decision strategies and sub-queries are inferred online through the model after fine adjustment. According to the invention, more efficient and accurate complex query processing is realized.
Owner:DATA SPACE RES INST

Infrared thermal imaging building facade defect intelligent diagnosis method based on multi-modal fusion

The invention provides an infrared thermal imaging building facade defect intelligent diagnosis method based on multi-modal fusion, and relates to the technical field of building detection.The method comprises the steps that infrared thermal imaging, visible light images and three-dimensional point cloud data are synchronously collected to construct a multi-modal data set; segmenting a hot spot region by adopting an improved morphological watershed algorithm and extracting contour and temperature features; recognizing a surface crack and peeling area based on a double-branch attention network to generate a texture defect feature map; curvature distribution and thermal deformation gradient are calculated through space registration constrained by a heat conduction equation; multi-source features are fused to calculate a hot spot form dispersion TSMD and a structure risk quantification factor SRQF; constructing a defect risk decision matrix to output defect types, positions and risk levels; and superposing the diagnosis result to a BIM model to generate a three-dimensional visual report and predicting a thermodynamic evolution trend. The multi-modal data collaborative analysis is realized, the defect risk is accurately quantified, and the problems of poor anti-interference performance, inaccurate segmentation and large registration error of a traditional method are solved.
Owner:SHAOXING MUNICIPAL DESIGN INST

Thyroid cancer electronic medical record system based on multi-modal data fusion

The invention relates to the field of medical informatization. The invention discloses a thyroid cancer electronic medical record system based on multi-modal data fusion. The thyroid cancer electronic medical record system comprises a multi-modal data acquisition module which acquires patient texts, ultrasonic images, genes, biochemical indexes and clinical data and performs standardized calibration to generate standard data; the multi-modal feature extraction module extracts semantic, structure, mutation, change and fluctuation features of each standard data through multiple technologies; the single-mode prediction model construction module constructs single-mode prediction models of texts, images and the like based on the features and outputs results; and the multi-modal fusion prediction module fuses the single-modal model based on the deep learning framework to output a multi-modal fusion prediction result. According to the invention, multi-modal data are integrated, and the accuracy and comprehensiveness of thyroid cancer diagnosis are improved. The system ensures consistency through standardized data processing, and assists doctors to accurately judge pathological types, recommend therapeutic schedules and evaluate prognosis by means of a multi-modal feature extraction and fusion mechanism.
Owner:ZHEJIANG CANCER HOSPITAL

Exhibition hall three-dimensional modeling intelligent optimization system based on multi-modal data fusion

The invention relates to the technical field of computer vision and three-dimensional reconstruction, and discloses an exhibition hall three-dimensional modeling intelligent optimization system based on multi-modal data fusion, and the system comprises a data collection module which is configured to synchronously obtain laser radar point cloud data, a multispectral image sequence and inertial measurement unit data; the preprocessing module is used for receiving the output of the data acquisition module, aligning a multi-source sensor coordinate system through a space-time calibration algorithm, and separating a static scene from a dynamic interference element by using a dynamic segmentation network; and the multi-modal fusion module is used for receiving the preprocessed data and carrying out adaptive weighted fusion on the geometric features of the laser radar and the visual texture features through a cross-modal attention mechanism. According to the invention, through multi-modal data fusion and a dynamic scene adaptive mechanism, the modeling precision and the real-time updating capability in a complex exhibition hall environment are significantly improved.
Owner:SHANDONG BAITE EXHIBITION ENG CO LTD

Intelligent enterprise data asset analysis method and system based on AI identification

The invention discloses an enterprise data asset intelligent analysis method and system based on AI recognition, and the method comprises the steps: receiving an enterprise multi-source heterogeneous data stream, carrying out the joint feature extraction and semantic alignment through a pre-trained multi-modal fusion recognition model, and generating a structured data asset recognition result; constructing a dynamic enterprise data asset atlas according to the structured data asset identification result in combination with the data access trajectory and authority metadata collected in real time; performing spatio-temporal evolution analysis on the dynamic enterprise data asset map, and extracting potential data value density features and risk exposure features; inputting the data value density features and the risk exposure features into a self-organizing mapping network to generate a data asset grading topological graph; and based on the data asset grading topological graph, through strategy constraint reinforcement learning, generating an executable data governance action sequence. According to the embodiment of the invention, the identification precision and real-time analysis capability of special assets of enterprises can be improved.
Owner:WUPO DIGITAL TECHNOLOGY (HANGZHOU) GROUP CO LTD

Dissimilar metal laser welding device based on swing light beam and molten pool state online monitoring

The invention discloses a dissimilar metal laser welding device based on swing light beams and molten pool state monitoring. The dissimilar metal laser welding device aims at improving the welding quality and the joint stability. The device comprises a laser welding head with a light beam swinging function, and the laser welding head can implement nonlinear energy scanning in a welding area according to a preset track type, frequency and amplitude; the acquisition module can synchronously acquire a visual image, an excitation spectrum and an infrared thermal imaging signal of the molten pool at a high frame rate, and extracts interface diffusion and metal mixing characteristics through multi-modal fusion; the calculation module performs dynamic feature modeling according to the collected physical feature information and the swing parameters, and quantifies multi-dimensional state indexes of the welding quality; the control module carries out combined adjustment on the laser power, the welding speed and the swing parameters according to the state indexes, closed-loop feedback control is constructed, and therefore real-time stable regulation and control and defect suppression in the dissimilar metal welding process are achieved.
Owner:SHENZHEN JUXIN AURORA TECH CO LTD

Multi-modal fusion tunnel structure apparent disease identification and risk assessment system

PendingCN121256709AData synchronizationDisease
The invention relates to the technical field of civil engineering tunnel structure safety monitoring and intelligent detection, in particular to a multi-modal fusion tunnel structure apparent disease identification and risk assessment system, which comprises an image acquisition module used for acquiring continuous images of the inner wall of a tunnel lining; a laser point cloud acquisition module; a structure sensor acquisition module; a data synchronization and preprocessing module; the multi-modal feature extraction module is used for performing depth feature extraction on the image, the point cloud and the sensor data; the heterogeneous feature fusion and disease identification module is used for fusing each modal feature and outputting a disease type identification result; and the risk assessment module is used for carrying out size estimation and parameterized expression on the identified diseases. The problems that in an existing tunnel inspection technology, the detection means is single, appearance and internal information cannot be considered, and the disease size is difficult to quantify automatically are solved.
Owner:HUAZHONG UNIV OF SCI & TECH

Robot sensing and decision-making method based on lightweight multi-modal large model

The invention relates to a robot sensing and decision-making method based on a lightweight multi-modal large model. The method comprises the following steps: constructing a semantic voxel map; collecting multi-modal data based on the semantic voxel map and preprocessing the multi-modal data, wherein the multi-modal data comprises visual data, point cloud data and a target semantic tag; performing feature extraction on the preprocessed multi-modal data, and performing dynamic cross-modal attention fusion to obtain multi-modal fusion features; inputting the multi-modal data into a lightweight multi-modal large model at the same time, and performing semantic analysis to obtain global space object semantic description; and based on the global space object semantic description, the target and direction embedding vector and the multi-modal fusion feature, a reinforcement learning algorithm is adopted to carry out hierarchical navigation decision making to obtain a target decision, and the target and direction embedding vector is a preprocessed target semantic tag. And the accuracy, timeliness and adaptability of robot perception and navigation decision making in a complex scene are improved.
Owner:SOUTHWEST JIAOTONG UNIV

Multi-modal retrieval enhanced generation government affair intelligent system

The invention relates to the technical field of government affair service intellectualization, in particular to a multi-modal retrieval enhanced generation government affair intelligent system which comprises multi-level government affair knowledge base construction, retrieval enhanced agent collaborative fine adjustment, a credible generation mechanism of government affair agents and multi-modal fused personalized interaction application. Multi-source heterogeneous data semantic unification is realized through a dynamic cross-modal alignment framework and a knowledge graph; a dual-path parameter efficient fine tuning architecture is adopted to improve the retrieval and generation capability; the tool chain is verified by integrating logic, law and ethics to ensure the compliance of the generated content; and guiding and optimizing complex scene demand understanding by using distributed user portraits and progressive dialogues. According to the application, the intelligent level and the user experience of government affair services can be remarkably improved, and the government affair services are promoted to transition from passive response to active service.
Owner:WEIHAI CHIYUN NETWORK TECH CO LTD

Multi-label electrocardiogram classification method based on self-supervised pre-training and multi-modal semantic alignment

The invention discloses a multi-label electrocardiogram classification method based on self-supervised pre-training and multi-modal semantic alignment, which belongs to the technical field of artificial intelligence, and comprises the following steps: realizing self-supervised pre-training of unlabeled data through a single-modal contrast enhancement network, generating global and local contrast views by adopting a multi-scale random cutting strategy, and classifying the global and local contrast views in a multi-scale random cutting mode; in combination with a teacher-student network architecture, the potential invariance features of the ECG signals are learned while negative sample dependence is avoided, the problem of annotation data scarcity is effectively relieved, and the feature robustness is improved. A multi-modal fusion mechanism based on label semantic guidance is provided, a time domain signal and a frequency domain time-frequency graph are mapped to a unified semantic space through fine-grained semantic alignment, local feature enhancement and cross-modal complementary information fusion are realized by using a cross attention mechanism, and the problem of semantic difference caused by modal heterogeneity in a traditional method is overcome. A multi-label comparison loss function based on a disease co-occurrence relation is proposed, a category discrimination boundary is dynamically optimized by modeling a label co-occurrence probability, the feature separability of a tail category is improved while the head category discrimination ability is enhanced, and the problem of sample category imbalance in a multi-label scene is remarkably relieved.
Owner:YANSHAN UNIV

Multi-agent-based gas insulated switchgear fault diagnosis method and system

The invention discloses a multi-agent-based gas insulated switchgear fault diagnosis method and system, and relates to the technical field of intelligent operation and maintenance of power equipment, and the method comprises the steps: obtaining signal data of target equipment, carrying out the feature extraction of the signal data, and constructing a multi-modal feature matrix; time delay features of acoustic and electromagnetic signals are extracted from the multi-modal feature matrix, a GIS propagation model is established, and the space coordinate position of a liberated power source is solved through a wave field inversion algorithm; combining the space coordinate position and the multi-modal feature matrix into a complete fusion feature vector, inputting the fusion feature vector into a dynamic Bayesian model, and outputting a fault type label and a corresponding confidence coefficient; migrating the dynamic Bayesian model based on a migration learning mechanism, and dynamically updating a classification threshold value; inputting the diagnosis history sequence into a time sequence prediction model, and predicting a future operation state; through multi-modal fusion and intelligent reasoning, GIS fault accurate positioning and prediction are realized, and the problems of low precision and poor adaptability of traditional diagnosis are solved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Robot control method, system and equipment based on multi-modal large model and medium

The invention relates to the technical field of robot control, and discloses a robot control method, system, equipment and medium based on a multi-modal large model, and the method comprises the steps: collecting the multi-source modal data of a scene where an operation task is located, and carrying out the processing through a machine learning model, obtaining a multi-modal feature, and carrying out the position coding and Transform fusion processing, multi-modal fusion features are obtained, the multi-modal fusion features and the constructed job task knowledge base are input into a large language model to decompose a target job task, a human-in-the-loop mechanism is introduced to optimize a decomposition result, and a sub-task sequence is obtained; according to a subtask type in the subtask sequence, processing the subtask sequence through a visual language action model or a reinforcement learning model, and generating a motion instruction to enable the robot to start an execution process of the target operation task; live-line work tasks are processed through the multi-modal large models LLM, VLA and the like, and the work efficiency of the autonomous distribution network live-line work robot is improved.
Owner:WENZHOU ELECTRIC POWER BUREAU +2

Project risk monitoring method and system based on large language model

The invention relates to the technical field of project risk management, in particular to a project risk monitoring method and system based on a large language model, and aims to guide a language model to complete risk identification in a professional context by analyzing a natural language supervision request of a user, identifying a task field, matching a corresponding knowledge graph and a rule base, generating a reasoning configuration set and guiding the language model to complete risk identification in a professional context. Through a multi-modal fusion mechanism, unstructured data such as contract texts, drawing images and progress logs are coded in a unified mode, context modeling and rule reasoning of cross-modal information are achieved in combination with a large language model guided by a strategy, hidden risks needing image-text linkage judgment are effectively recognized, the analysis capacity for complex semantic association is improved, and the method is suitable for large-scale popularization and application. And furthermore, through a reinforcement learning mechanism, a supervision sample is constructed according to user feedback, a reward signal is generated, language model strategy parameters are optimized in real time, and continuous evolution and self-adaptive updating of a risk monitoring model are realized.
Owner:GUANGZHOU SAIBAO LIANRUI INFORMATION TECH

Bill voucher information extraction method, system and equipment based on multi-mode and OCR model fusion

The invention relates to a bill voucher information extraction method based on multi-mode and OCR model fusion. The method comprises the following steps: S1, obtaining an image of a bill voucher; s2, preprocessing the image; s3, identifying the preprocessed image by using an OCR engine to obtain the text content and the corresponding two-dimensional coordinates of each text block; s4, taking the recognized text segments and the original image as input, performing joint coding by using a pre-trained multi-modal model, evaluating and outputting the matching degree of each text segment and a predefined field category by the model, and determining candidate texts of each field and confidence of the candidate texts; s5, accurately positioning and extracting the key field, and verifying the consistency of the OCR output and the semantic result; s6, if the verification result conflicts or the identification reliability of a certain field is lower than a threshold value, error correction operation is carried out; and S7, outputting the structured bill voucher information. Through multi-modal fusion and iterative correction, the error rate of non-standard voucher information extraction is effectively reduced, and the method is suitable for various voucher formats and complex scenes.
Owner:ZHIWEI (SUZHOU) INFORMATION TECH CO LTD

Multi-modal fusion rumor detection method and system based on dynamic graph convolutional neural network

The invention discloses a multi-modal fusion rumor detection method and system based on a dynamic graph convolutional neural network. According to the method, a dynamic feature graph of a language propagation path is constructed, and potential features in the language propagation process are extracted and analyzed by utilizing time sequence changes and key node relations between nodes in a propagation graph. A neural network is adopted to extract and enhance image data, text semantic features are extracted in combination with a text feature modeling network, text feature vectorization expression is achieved based on a BERT model, and rich semantic information is obtained. And a gating mechanism is introduced to dynamically adjust fusion weights of different modal features, and an information fusion strategy is optimized. A collaborative attention mechanism is further adopted for deep fusion, interactive learning of text, image and propagation path features is enhanced, and the relevance of cross-modal and time series data is improved. And finally, inputting the fused feature vectors into a classifier for accurate classification, thereby realizing accurate detection of the social media rumors. According to the method, the multi-modal features are effectively integrated, and the false information identification efficiency is remarkably improved.
Owner:CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY +1

Respiratory system risk prediction method and system based on graph neural network

The invention relates to the technical field of respiratory system risk prediction, and provides a respiratory system risk prediction method and system based on a graph neural network, and the method comprises the steps: collecting the multi-modal medical data of a patient, and constructing a multilayer heterogeneous graph based on the multi-modal medical data; constructing a weighted adjacency matrix and a node feature vector through the multi-layer heterogeneous graph; matrix product operation and convolution operation are carried out based on the weighted adjacent matrix and the node feature vector, splicing combination with historical moment state information is carried out, graph state representation is obtained, weighted aggregation of time dimensions is carried out, and time sequence attention features are obtained; performing coding processing based on the clinical examination data to obtain multi-modal fusion features; and inputting the multi-modal fusion features into a risk classifier for classification calculation to obtain a respiratory system risk level prediction result, generating a risk assessment report, and outputting respiratory risk early warning information. The accuracy and clinical practicability of respiratory system risk prediction are improved.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

Multi-modal environment sensing method and system of low-altitude medical unmanned aerial vehicle

The invention relates to the field of unmanned aerial vehicle environment perception, in particular to a multi-mode environment perception method and system for a low-altitude medical unmanned aerial vehicle. The method comprises the following steps: collecting a multi-modal data stream, carrying out adaptive data optimization processing, and constructing a multi-modal fusion data set; performing environment multi-level obstacle identification and evaluation on the multi-modal fusion data set to generate a threat mapping environment map; multi-dimensional environment parameters are collected based on the unmanned aerial vehicle, wind field time-varying prediction and safe flight area calculation are performed based on the threat mapping environment map, and a flight area map is constructed; performing multi-position collision risk assessment based on the multi-modal fusion data set to generate collision risk coefficients of different positions; and carrying out safe flight constraint analysis on the flight area map according to the collision risk coefficient, and extracting an optimal flight path. According to the invention, in combination with real-time environment data, comprehensive flight path planning is provided, and the flight safety and task completion efficiency of the unmanned aerial vehicle are improved.
Owner:GUANGZHOU XIAOWEI TECH CO LTD

Multi-modal fusion intelligent question answering and knowledge retrieval method and system

The invention discloses a multi-modal fused intelligent question answering and knowledge retrieval method and system, and the method comprises the steps: building a multi-modal data index model oriented to a heterogeneous knowledge source, carrying out the feature mapping of text, image, table, chart, audio and video contents through a unified semantic embedding space, and generating a cross-modal index set; after a query request is received, performing semantic matching and structure matching on the cross-modal index set by using a multi-channel retriever to obtain candidate evidence fragments; and based on an evidence granularity decomposition strategy, performing minimum evidence unit division on text statements, table units, chart data points and multimedia frame contents in the candidate evidence fragments, and establishing a semantic consistency graph among the units. According to the method, high-credibility traceable generation of question and answer results is realized through multi-modal fusion and space-time consistency constraint, and the retrieval precision and interpretation transparency in a complex knowledge scene are remarkably improved.
Owner:NANJING CHUANGLIAN INTELLIGENT SOFT INFORMATION TECH CO LTD

Dynamic calibration method and system of vehicle-mounted emotion recognition system

The invention provides a dynamic calibration method and system for a vehicle-mounted emotion recognition system, and the method comprises the steps: S1, obtaining multi-source data which comprises a facial image, a voice signal and a physiological signal; the obtained multi-source data are preprocessed, and preprocessed multi-source data are obtained; s2, performing feature extraction based on the preprocessed facial image, the voice signal and the physiological signal to obtain a facial expression feature vector, an audio feature vector and a physiological state feature vector; s3, evaluating the current environment credibility based on an environment credibility evaluation function; s4, dynamically distributing the weight of the multi-source data according to the credibility of the current environment and the real-time scene; and S5, constructing a multi-modal fusion vector based on the dynamically distributed weight of the multi-source data, the facial expression feature vector, the audio feature vector and the physiological state feature vector, and performing emotion recognition by using the constructed emotion recognition model based on the multi-modal fusion vector.
Owner:SHANGHAI PUFAFEN ELECTRONIC TECH CO LTD

Image recognition and analysis system based on AI

The invention relates to the technical field of image processing, and discloses an image recognition and analysis system based on AI. The system comprises a data acquisition module, a feature extraction module, a model training module, a multi-modal fusion module, a dynamic optimization module and the like. The method comprises the steps of collecting real-time image data by a multi-source sensor, extracting features by a cascade convolutional neural network, generating an adversarial network training model, integrating multi-source data by multi-modal fusion, optimizing feature vectors by an improved genetic algorithm, and constructing a classification decision tree. In addition, an anomaly detection module, a real-time reasoning module, a data enhancement module and a visualization module are further arranged. The system can accurately identify and analyze images, improve the model performance and generalization ability, meet the real-time requirement of edge computing equipment, generate an interpretable report to assist decision making, and have wide application prospects in the fields of security, medical treatment, automatic driving and the like.
Owner:ZHUHAI WANDU TECHNOLOGY CO LTD

Outer wall hollowing microwave reflection detection method based on multi-modal fusion

The invention belongs to the technical field of microwave measurement, and discloses an outer wall hollowing microwave reflection detection method based on multi-modal fusion, which comprises the following steps: carrying out multi-modal scanning on a building outer wall to be detected, obtaining visible light image data and infrared temperature distribution data of the outer wall surface, and carrying out space registration and coordinate mapping; a unified multi-modal fusion data set is formed; performing anomaly screening on the multi-modal fusion data set, and identifying a thermal anomaly region by analyzing infrared temperature distribution data; detecting a bump or crack area in combination with texture and morphology anomaly features of the visible light image data; performing information fusion on the thermal anomaly region and the bump or crack region, and extracting candidate detection regions of suspected hollowing; high-precision recognition and quantitative evaluation of the outer wall hollowing are achieved, and the precision and stability of outer wall hollowing detection are improved.
Owner:HEFEI HUIXIAO ROBOT TECHNOLOGY CO LTD

Method and system for generating official document key abstract based on multi-modal feature extraction

The invention provides an official document key abstract generation method and system based on multi-modal feature extraction, and relates to the technical field of multi-modal artificial intelligence generation, and the method comprises the steps: firstly obtaining text modal data, image modal data and table modal data of a to-be-processed official document, then carrying out the hierarchical semantic analysis processing of the text modal data, and obtaining a to-be-processed official document key abstract; the method comprises the following steps: generating a text semantic feature set, performing visual element extraction processing on image modal data, generating an image feature set, performing structured analysis processing on table modal data, generating a table feature set, and performing cross-modal alignment processing on the text semantic feature set, the image feature set and the table feature set. According to the method, the dynamic association feature sets among the text semantics, the image elements and the table elements are determined according to the text semantics, the image elements and the table elements, the three feature sets are subjected to multi-modal fusion processing according to the feature sets, the target abstract content of the to-be-processed official document is generated, complementarity of multi-modal information in the official document is fully utilized, and the generated abstract is more complete, accurate and targeted.
Owner:STATE GRID SHANDONG ELECTRIC POWER COMPANY WEIFANG POWER SUPPLY

Multi-modal fusion perception robot dog inspection slope disaster risk assessment method and device and storage medium

The invention provides a multi-modal fusion sensing robot dog inspection slope disaster risk assessment method and device and a storage medium, and relates to the field of slope disaster assessment, and the method comprises the steps: obtaining multi-modal sensing data; performing space-time alignment processing on the multi-modal sensing data and then converting the multi-modal sensing data into a voxel coordinate system; performing feature extraction on the multi-modal sensing data after time-space alignment, and mapping each extracted feature to a unified voxel unit to form a three-dimensional voxel structure containing each extracted feature; constructing a three-dimensional multi-modal data fusion model; identifying disaster types based on the three-dimensional multi-modal data fusion model, wherein the disaster types comprise the ground surface crack length, the underground cavity volume, the water seepage point number, the local collapse and bulging area, the vegetation degradation area and the slope gradient; and calculating a risk index based on the identified disaster type in combination with the association degree of the disaster type. By adopting the evaluation method provided by the invention, rapid and efficient evaluation of slope disasters can be realized, and the method has relatively good accuracy.
Owner:CHANGAN UNIV +1