Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3156 results about "Multi modal fusion" patented technology

Intelligent analysis method based on medical document structure perception and multi-modal fusion

An intelligent analysis method based on medical document structure perception and multi-modal fusion comprises the following steps: carrying out structure topology modeling on a medical document, extracting visual layout, text meta-information, space coordinates and semantic keyword features, constructing a semantic topological graph and dynamically shielding irrelevant contents; selecting an extraction path according to a document type, performing deep semantic analysis and entity recognition on a text-type document, and performing visual enhancement OCR recognition on a scanning-type document; the features are injected into a medical knowledge graph, and feature fusion, semantic verification, relation reasoning and information completion are achieved through a graph neural network; a three-stage strategy optimization model of basic pre-training, domain adaptation and online reinforcement learning is adopted; and large-scale processing is realized through a dynamically aggregated distributed architecture. The method is used for intelligent analysis and structured conversion of documents of hospitals, medical insurance and medical scientific research. The problems that heterogeneous medical document analysis adaptability is poor, multi-modal fusion is difficult, medical knowledge utilization is insufficient, and large-scale processing efficiency is low are solved.
Owner:NORTHWEST UNIV

Multi-modal fusion tunnel structure apparent disease identification and risk assessment system

PendingCN121256709AData synchronizationDisease
The invention relates to the technical field of civil engineering tunnel structure safety monitoring and intelligent detection, in particular to a multi-modal fusion tunnel structure apparent disease identification and risk assessment system, which comprises an image acquisition module used for acquiring continuous images of the inner wall of a tunnel lining; a laser point cloud acquisition module; a structure sensor acquisition module; a data synchronization and preprocessing module; the multi-modal feature extraction module is used for performing depth feature extraction on the image, the point cloud and the sensor data; the heterogeneous feature fusion and disease identification module is used for fusing each modal feature and outputting a disease type identification result; and the risk assessment module is used for carrying out size estimation and parameterized expression on the identified diseases. The problems that in an existing tunnel inspection technology, the detection means is single, appearance and internal information cannot be considered, and the disease size is difficult to quantify automatically are solved.
Owner:HUAZHONG UNIV OF SCI & TECH

Multi-agent-based gas insulated switchgear fault diagnosis method and system

The invention discloses a multi-agent-based gas insulated switchgear fault diagnosis method and system, and relates to the technical field of intelligent operation and maintenance of power equipment, and the method comprises the steps: obtaining signal data of target equipment, carrying out the feature extraction of the signal data, and constructing a multi-modal feature matrix; time delay features of acoustic and electromagnetic signals are extracted from the multi-modal feature matrix, a GIS propagation model is established, and the space coordinate position of a liberated power source is solved through a wave field inversion algorithm; combining the space coordinate position and the multi-modal feature matrix into a complete fusion feature vector, inputting the fusion feature vector into a dynamic Bayesian model, and outputting a fault type label and a corresponding confidence coefficient; migrating the dynamic Bayesian model based on a migration learning mechanism, and dynamically updating a classification threshold value; inputting the diagnosis history sequence into a time sequence prediction model, and predicting a future operation state; through multi-modal fusion and intelligent reasoning, GIS fault accurate positioning and prediction are realized, and the problems of low precision and poor adaptability of traditional diagnosis are solved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Multi-modal fusion intelligent question answering and knowledge retrieval method and system

The invention discloses a multi-modal fused intelligent question answering and knowledge retrieval method and system, and the method comprises the steps: building a multi-modal data index model oriented to a heterogeneous knowledge source, carrying out the feature mapping of text, image, table, chart, audio and video contents through a unified semantic embedding space, and generating a cross-modal index set; after a query request is received, performing semantic matching and structure matching on the cross-modal index set by using a multi-channel retriever to obtain candidate evidence fragments; and based on an evidence granularity decomposition strategy, performing minimum evidence unit division on text statements, table units, chart data points and multimedia frame contents in the candidate evidence fragments, and establishing a semantic consistency graph among the units. According to the method, high-credibility traceable generation of question and answer results is realized through multi-modal fusion and space-time consistency constraint, and the retrieval precision and interpretation transparency in a complex knowledge scene are remarkably improved.
Owner:NANJING CHUANGLIAN INTELLIGENT SOFT INFORMATION TECH CO LTD

Outer wall hollowing microwave reflection detection method based on multi-modal fusion

The invention belongs to the technical field of microwave measurement, and discloses an outer wall hollowing microwave reflection detection method based on multi-modal fusion, which comprises the following steps: carrying out multi-modal scanning on a building outer wall to be detected, obtaining visible light image data and infrared temperature distribution data of the outer wall surface, and carrying out space registration and coordinate mapping; a unified multi-modal fusion data set is formed; performing anomaly screening on the multi-modal fusion data set, and identifying a thermal anomaly region by analyzing infrared temperature distribution data; detecting a bump or crack area in combination with texture and morphology anomaly features of the visible light image data; performing information fusion on the thermal anomaly region and the bump or crack region, and extracting candidate detection regions of suspected hollowing; high-precision recognition and quantitative evaluation of the outer wall hollowing are achieved, and the precision and stability of outer wall hollowing detection are improved.
Owner:HEFEI HUIXIAO ROBOT TECHNOLOGY CO LTD

Mineral resource intelligent prediction method and system based on multi-source heterogeneous data fusion and deep learning

The invention discloses a mineral resource intelligent prediction method and system based on multi-source heterogeneous data fusion and deep learning, and the method comprises the steps: collecting and preprocessing multi-source heterogeneous data, and carrying out the standardization processing to form a structured data set; multi-source heterogeneous data fusion: realizing data layer space registration and feature layer weight dynamic allocation through an attention mechanism multi-modal fusion module, and outputting a high-dimensional metallogenic feature vector; constructing a CNN-LSTM mixed deep learning model and completing initialization training, and outputting an initial mineralization probability graph; and establishing a dynamic updating engine, performing model increment training based on transfer learning, correcting the mineralization probability through positive and negative sample reinforcement learning in combination with a newly added data type, and outputting a time sequence dynamic mineralization probability graph. According to the method, mineralization probability dynamic evaluation and risk quantitative updating are realized, the prediction precision and the model updating efficiency are improved, the method is adaptive to a multi-stage exploration scene, and accurate real-time support is provided for exploration decision making.
Owner:EAST CHINA UNIV OF TECH

End-to-end automatic driving method based on dynamic multi-modal fusion in complex scene

The invention discloses an end-to-end automatic driving method based on dynamic multi-modal fusion in a complex scene, and belongs to the technical field of automatic driving. In order to solve the problems of sensor perception deficiency, cross-modal feature mismatching, unstable trajectory planning and the like easily occurring in night, low-illumination and complex dynamic environments in the existing end-to-end automatic driving method, texture details of a camera mode and geometric structure features of a laser radar mode are respectively enhanced through a double-flow feature refining mechanism; the characteristic difference between different modes is relieved; an information-driven dynamic fusion strategy is designed, the fusion weight is adaptively adjusted according to scene factors such as environment illumination and obstacle density, and the scene sensitivity and discrimination ability of the model are improved; asymmetric convolution and a low-rank-sparse decoupling technology are introduced, multi-order reconstruction of key channels is carried out on the multi-modal features, and the path change modeling capability is enhanced; and in combination with time sequence dependence of waypoints, outputting a future trajectory through an autoregression decoder to realize high-precision trajectory prediction and stable decision control.
Owner:ZHONGBEI UNIV

Multi-modal fusion perception smoke and fire identification system and method

The invention relates to the technical field of fire safety monitoring, in particular to a firework identification system and method based on multi-modal fusion perception, and the core of the scheme is a visible light, multispectral and temperature three-modal framework: feature extraction optimization of each modal, improved YOLOv8s for visible light branches, dynamic background modeling and flame color screening, and multi-modal fusion perception. False positive is rejected by a multispectral branch depending on a waveband ratio and an index, an error compensation algorithm is introduced into a thermopile branch, and finally, a final smoke and fire area and the confidence coefficient thereof are determined by associating a three-mode area through collaborative decision. According to the scheme, the complex environment adaptability and the recognition reliability can be improved, the false alarm risk is reduced, the extremely-early smoke and fire detection capability is enhanced, good real-time performance and deployment flexibility are achieved, and the method is suitable for various types of fire safety monitoring scenes.
Owner:SHENZHEN HOT WHEELS TECHNOLOGY CO LTD

Dam leakage intelligent identification method based on multi-modal fusion and knowledge enhancement

The invention provides a dam leakage intelligent identification method based on multi-modal fusion and knowledge enhancement, and the method comprises the steps: collecting real-time data of a multi-source sensor disposed at a key part of a dam in a preset monitoring time period, and generating seepage characteristic data; identifying a seepage form entity based on the seepage characteristic data and extracting an instantaneous characteristic entity, and associating the entity into a structured knowledge unit according to a space-time proximity principle; knowledge units are classified according to spatial positions and influence ranges, association rules of the knowledge units are complemented, and knowledge graph construction is achieved; and then, a map inference engine is triggered in a real-time feature matching mode, and graded early warning is implemented. According to the method, physical enhanced seepage characteristics are constructed, seepage forms, dynamic characteristics and inducements are deeply associated by utilizing a knowledge graph technology, accurate diagnosis and reasoning from data abnormity to seepage types, causes and risk levels are realized, and finally, the seepage characteristics are analyzed and analyzed through a dynamic conflict resolution and self-evolution mechanism. And a reliable dam leakage intelligent identification and decision-making system is formed.
Owner:ANHUI DANFENGYUAN TECH CO LTD

Intelligent hardware dynamic interaction system based on voice semantic fusion and multi-mode perception

The invention relates to the field of intelligent interaction, and discloses an intelligent hardware dynamic interaction system based on voice semantic fusion and multi-modal perception, which comprises the following steps of: constructing a context model of continuous operation by collecting continuous voice instructions, gesture actions and expression information of a user; semantic analysis and feature fusion are carried out on currently collected voice, gesture and expression features, meanwhile, credibility indexes of all modes are calculated through a weighting or deep learning model, weighting correction is carried out on a fusion result, a real-time feedback algorithm is adopted for weight adjustment for continuous optimization, the next operation intention of a user is predicted through deep learning, and the user experience is improved. And in combination with historical interaction data, online feedback and prediction errors, context management, modal weight and intention prediction strategies are adaptively optimized, and the updated strategies are used for next-round context acquisition and multi-modal fusion. The method has the advantage of improving the recognition accuracy in the continuous interaction scene.
Owner:华欧同惠(苏州)科技有限公司

Weld joint quality intelligent diagnosis system based on deep learning

The invention discloses a weld quality intelligent diagnosis system based on deep learning, and relates to the technical field of welding quality detection, and the weld quality intelligent diagnosis system comprises an image quality evaluation module, a feature alignment module, a deviation detection module, a path reconstruction module, a prior enhancement module and a defect identification module, identifying an area of which the signal-to-noise ratio is lower than a preset threshold value, and constructing a noise interference distribution diagram; and the feature alignment module executes a deformable convolution feature alignment operation with a confidence factor adjustment mechanism based on the noise interference distribution diagram to generate an initial space mapping result. Through mechanisms such as image quality perception, robust alignment, deviation detection, self-adaptive reconstruction and prior enhancement, a closed-loop weld joint intelligent diagnosis process is constructed, false alignment errors are effectively inhibited, the multi-modal fusion stability and the defect recognition precision are improved, and the reliability and the intelligent level of the system under complex working conditions are enhanced.
Owner:ZHEJIANG ELECTRIC POWER CONSTR CO LTD +1

Power distribution network data intelligent analysis method based on data consanguinity and multi-modal fusion learning

The invention relates to a power distribution network data intelligent analysis method based on data consanguinity and multi-modal fusion learning. The method comprises the following steps: S1, constructing a dynamically evolved data consanguinity topological graph; s2, designing a label-guided graph neural network architecture, embedding historical abnormal knowledge into a graph learning process, and outputting a deep semantic feature vector; s3, constructing a multi-modal fusion analysis framework, performing multi-dimensional feature fusion and data quality analysis, and identifying abnormal nodes; s4, designing a semi-supervised and incremental learning combined mixed training normal form, and performing model training and strategy optimization; and S5, based on the dynamic consanguinity topology constructed in the step S1 and the identified abnormal nodes, constructing a probabilistic reasoning framework, and fusing the model parameters obtained by optimization in the step S4 to realize quality abnormality root positioning and full-link visualization so as to form a complete data intelligent analysis scheme. According to the invention, efficient and accurate management of the topological data quality of the power distribution network is realized.
Owner:STATE GRID TIANJIN ELECTRIC POWER COMPANY +1

Knowledge graph agent construction method and system based on multi-modal fusion

The invention relates to a knowledge graph agent construction method and system based on multi-modal fusion, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring multi-modal data, and mapping different modal data to a unified semantic space through a cross-modal embedding technology to obtain multi-modal reference data; extracting features of the multi-modal reference data through an attention mechanism to obtain multi-modal joint features; constructing and updating a dynamic knowledge graph based on the multi-modal joint features to obtain a time sequence dynamic dependency knowledge graph; and according to the time sequence dynamic dependency knowledge graph, context-aware reasoning is carried out through a cross-modal reasoning engine to obtain a user input mode, and an interaction strategy is dynamically adjusted based on the user input mode to generate a multi-modal response. And the knowledge graph agent construction based on multi-modal fusion is realized.
Owner:SHANGHAI YUTA TECH CO LTD

Vehicle driving safety early warning method and system fused with meteorological data

The invention relates to the technical field of safety early warning, and particularly discloses a vehicle driving safety early warning method fused with meteorological data, which comprises the following steps: acquiring real-time multi-modal data of meteorological, traffic flow and vehicle state of a target road area, performing exception handling, space-time alignment and standardization to form a standardized data sequence, then constructing a multi-modal fusion tensor, and finally performing data fusion on the multi-modal fusion tensor. Extracting each modal dynamic mode, fusing cross-modal features, outputting a joint feature vector, inputting the joint feature vector into a safety risk prediction model to calculate a dynamic safety risk value, combining digital twin simulation risk conduction, generating graded early warning according to a preset threshold value, and performing management and control through vehicle-road collaborative network publishing and high-risk scene linkage traffic facilities. And finally, collecting feedback data evaluation effects, associating decision data to generate hash records, recording the hash records in the block chain, and carrying out federated learning incremental training optimization model based on feedback. According to the invention, accurate early warning under multi-factor coupling can be realized, data privacy is guaranteed, closed-loop optimization is formed, and road traffic safety and stability are improved.
Owner:XINYOUXI TRAVEL TECHNOLOGY (HANGZHOU) CO LTD

Slope intelligent inspection early warning method and system based on time-space-multi-mode fusion

The invention provides a side slope intelligent inspection early warning method and system based on time-space-multi-mode fusion, and relates to the technical field of side slope monitoring. Through preprocessing of multi-source heterogeneous data, extraction and fusion of multi-modal features, physically constrained neural network prediction and risk-driven inspection adjustment, a complete slope intelligent inspection early warning process is constructed, complementary information of monitoring nodes, images and point cloud data is effectively integrated, the limitation of limited coverage of a single sensor is avoided, and the intelligent slope inspection early warning method is suitable for intelligent slope inspection early warning. The comprehensive perception and dynamic early warning of the overall state of the slope are realized, so that the technical problems that the manual patrol efficiency is low and the overall state is difficult to grasp by a point sensor are solved, and the problems caused by insufficient spatio-temporal evolution law mining and multi-modal data shallow fusion are systematically solved.
Owner:HEFEI UNIV OF TECH

Tunnel lining disease automatic identification method and system based on multi-source data fusion

The invention discloses a tunnel lining disease automatic identification method and system based on multi-source data fusion, and relates to the technical field of facility detection, and the method comprises the steps: collecting multi-modal time sequence data, carrying out the time-space alignment, and obtaining a time sequence multi-source data set; reconstructing a tunnel center line based on a vehicle pose and constructing a lining structure consistency coordinate framework, and performing structured projection and distortion correction on alignment data to obtain a multi-modal fusion data set; dividing a two-dimensional structure grid under the coordinate framework, extracting and fusing geometric, texture, depth and energy features, calculating a structure consistency damage index, and extracting a suspected disease area; and calculating a disease credibility index and judging a disease type in combination with multi-modal physical evidence, mapping a suspected disease region back to a three-dimensional space, completing disease boundary extraction and geometric quantization, and outputting structured disease information. According to the method, the structure expression and the structured alignment of the cross-modal data under the unified geometric reference are realized by constructing the consistent coordinate framework of the lining structure.
Owner:EAST CHINA JIAOTONG UNIVERSITY

Open vocabulary industrial defect detection method based on multi-modal prior prompt

The invention provides an open vocabulary industrial defect detection method based on multi-modal prior prompt, which comprises the following steps: S1, collecting and sorting defect images of industrial products to be detected, and constructing a large industrial reference data set; s2, building a visual priori prompt pool, and injecting refined and fine-tuned priori into multi-scale industrial features; s3, designing a significance Gaussian distribution modeling mechanism, and capturing a complex spatial mode of position prior, so that fine-grained prior has better expression ability, and the generalization of prior is improved; s4, establishing a decoupling LoRA text encoder, and extracting an industrial semantic basis in the industrial text template through a hierarchical prompt template and a hierarchical decoupling LoRA mechanism; and S5, constructing a visual text fusion unit, and performing multi-modal fusion on text prompt embedding and visual features. According to the method, high-precision and strong-generalization defect detection is realized by fusing visual and text prior prompts, and the robustness and adaptability in a complex open environment are remarkably improved.
Owner:CENT SOUTH UNIV

Dexterous hand based on multi-modal fusion control

The invention provides a dexterous hand based on multi-modal fusion control, and relates to the technical field of automatic control, the dexterous hand comprises a palm module, a palm center vision sensing module, a finger module, a fingertip force sensing module and a multi-modal control module; the palm center visual sensing module is used for acquiring visual information of the dexterous hand in a working area; the finger module comprises at least two fingers; the fingertip force sensing module is used for detecting information of contact force applied to the fingertip position; and the multi-mode control module is used for controlling the dexterous hand to grab a target object based on the visual information and the contact force information. According to the dexterous hand based on multi-mode fusion control, data of two modes of visual information and contact force information are fused, so that the dexterous hand can adapt to different grabbing environments, intelligent closed-loop control of visual guidance and force sense regulation and control is realized, and the environmental adaptability and flexibility of the dexterous hand are improved.
Owner:HUBEI JINGCHU HUMANOID ROBOT CO LTD

Small target identification method and system for multi-modal fusion image in complex environment

The invention discloses a small target recognition method and system for a multi-modal fusion image in a complex environment, and belongs to the technical field of computer vision and image recognition, and the method comprises the steps: obtaining a visible light image, an infrared image and environment sensor data; image registration is carried out on visible light and infrared images, and a multi-scale image feature pyramid is constructed. And respectively extracting visible light and infrared image features to obtain visible light and infrared imaging feature data. And performing multi-modal data fusion on the visible light and infrared imaging feature data based on a cross-modal attention mechanism, and adaptively adjusting a fusion weight based on environmental sensor data to generate fusion features. And performing space-time enhancement processing on the fusion feature to obtain an enhanced fusion feature. And performing target tracking detection on the small target, and outputting position and category information of the small target. According to the method, the small target recognition capability in a severe environment is remarkably improved, and high precision and robustness can still be kept in a foggy, low-visibility and dark scene.
Owner:CHINA TOWER CO LTD +1

Emotion prediction and disease derivation method and system based on multi-modal fusion

The invention discloses an emotion prediction and disease derivation method and system based on multi-modal fusion, and the system comprises a data collection and preprocessing module, an emotion fusion module, an abnormal condition detection and cloud uploading module, and a disease possibility derivation module. The data acquisition and preprocessing module comprises a video part, a text part and an audio part, and the video part comprises face emotion recognition and prediction and human motion recognition and prediction; the text part comprises text content emotion recognition and prediction; the audio part comprises voice-to-text and voice tone emotion recognition and prediction, the system comprehensively captures an emotion state by fusing multi-mode information such as video, text and voice, and the accuracy and prediction capability of emotion recognition are improved; and by predicting the future emotion trend, the abnormal condition is warned in advance, and the response timeliness is improved.
Owner:JIANGSU UNIV OF SCI & TECH IND TECH RES INST OF ZHANGJIAGANG

Intelligent ammeter clock calibration method and system based on multi-modal fusion

The invention relates to the technical field of electronic timer calibration, in particular to an ammeter clock intelligent calibration method and system based on multi-modal fusion, and the method comprises the steps: firstly obtaining ammeter clock voltage and environment data, carrying out the preprocessing, and generating an instantaneous disturbance sequence and a steady-state voltage time sequence; respectively calculating a voltage chronic influence factor, a voltage acute influence factor and an environment acceleration factor; determining the influence coefficient of each factor through historical data regression of similar ammeters, carrying out fusion calculation on total life loss, and carrying out iteration to generate a dynamic correction period adaptive to individual damage; and finally, constructing a same-batch and same-environment electric meter reference group, screening similar clusters through clustering, comparing corrected time consistency, judging unconventional abnormity, and outputting a targeted calibration strategy. According to the method, the multi-dimensional error factors are quantified through multi-modal fusion, the differentiation model is constructed, the dynamic correction period is generated through iteration, the traditional fixed period calibration defect is overcome, and the long-term timing precision stability of the electric meter is improved.
Owner:SHENZHEN FRIENDCOM TECH DEV

Crop disease diffusion prediction method and system based on multi-modal fusion

The invention discloses a crop disease diffusion prediction method and system based on multi-modal fusion, and the method comprises the following steps: S1, collecting and preprocessing an RGB image sequence and a sensor data sequence of a crop growth environment, and generating an RGB image time sequence difference result and a sensor difference result through time difference processing; s2, mapping the RGB image time sequence difference result and the sensor difference result to a shared time sequence space through a time alignment algorithm, and generating a sensor alignment result and an RGB alignment result; s3, an FD-ViT prediction model is constructed; inputting the sensor alignment result and the RGB alignment result into an FD-ViT prediction model for prediction, and generating a prediction result; and S4, generating a disease diffusion thermodynamic diagram and early warning information according to a prediction result. According to the method, RGB image data and sensor network data are fused, a Transform-based time sequence prediction model is constructed, and early recognition and diffusion trend prediction of crop diseases are realized.
Owner:HANGZHOU DIANZI UNIV

Tunnel fire inversion method based on multi-modal fusion and physical constraint

The invention provides a tunnel fire behavior inversion method based on multi-modal fusion and physical constraint, and aims to solve the problem that the position and power of a fire source are difficult to invert accurately in real time in traditional fire behavior monitoring. According to the method, data, temperature distribution images and physical parameters of distributed temperature sensors in a tunnel are collected, and a spatial-temporal feature fusion network is constructed; designing an information leakage prevention training mechanism, and adopting a modal random discarding and sensor shielding technology to avoid overfitting of the model to a specific input mode; a physical constraint loss function is introduced, and the temperature gradient, the maximum temperature rise and the longitudinal attenuation law are combined to ensure the physical rationality of an inversion result; a coarse-fine two-stage position prediction framework is adopted, and meter-scale precision positioning of the position of a fire source is achieved. According to the invention, the power and position of the fire source can be inversed accurately in real time, and reliable technical support is provided for intelligent monitoring and emergency rescue of tunnel fire.
Owner:CHINA UNIV OF MINING & TECH

Intelligent extraction and indexing system for file metadata

The invention relates to the technical field of archive information management, and discloses an archive metadata intelligent extraction and indexing system, which comprises a multi-modal preprocessing module for obtaining and preprocessing original multi-modal archive data; the context entity recognition module is used for performing entity recognition and standardization according to the context vector; the cross-modal fusion module is used for carrying out confidence weighted multi-modal fusion and logic verification; the archive association module is used for carrying out association identification and consistency detection between archives; the intelligent indexing module is used for carrying out hierarchical intelligent indexing and quality feedback on the consistency constrained file metadata set; the quality evaluation module is used for carrying out metadata quality evaluation and active repair on the standardized indexing result; the knowledge graph module is used for constructing a time sequence knowledge graph and intelligent retrieval service; according to the method, collaborative extraction of cross-modal information is realized by constructing a confidence-weighted multi-modal fusion model and a bidirectional attention mechanism.
Owner:SHANDONG ZHENGTU INFORMATION POLYTRON TECH INC

Semi-supervised target detection method for visible light-infrared multi-mode fusion scene

The invention provides a semi-supervised target detection method for a visible light-infrared multi-mode fusion scene. The method comprises the following steps: constructing a semi-supervised visible light-infrared multi-modal image data set based on an LLVIP data set; on the basis of a YOLOv11 model architecture, constructing a target detection model oriented to multi-modal image feature fusion, and training the target detection model by using a semi-supervised visible light-infrared multi-modal image data set in a deep learning end-to-end mode to obtain a trained target detection model; and inputting a to-be-detected multi-modal image into the trained target detection model, and outputting a target detection result of the to-be-detected multi-modal image by the trained target detection model. The method is based on a semi-supervised learning normal form, so that the precision and robustness of target detection in a multi-modal fusion scene are improved, and the requirements for high efficiency and reliability of target recognition in practical application scenes such as intelligent traffic and intelligent security and protection are met.
Owner:BEIJING JIAOTONG UNIV

Unmanned aerial vehicle outdoor inspection method based on multi-modal fusion

The invention discloses an unmanned aerial vehicle outdoor inspection method based on multi-modal fusion, and the method comprises the following steps: obtaining the inspection data of an unmanned aerial vehicle, and carrying out the feature extraction and space-time alignment; calling a GNPDE algorithm, and constructing a multi-modal field mapping layer; constructing a structure-thermal field joint graph structure, and establishing a double-branch coupling solution structure; constructing a topological adaptive edge weight adjustment module, and dynamically modulating the edge weight by adopting a physical modulation function; introducing an energy conservation constraint layer, and executing constraint solution; a multi-scale PDE evolution algorithm subset is quoted and combined, and a scale weight sharing mechanism is adopted to complete cross-scale joint optimization; constructing an abnormal residual reasoning module, and generating an abnormal significance map; and performing spatial registration and superposition on the abnormal saliency map and unmanned aerial vehicle inspection data, and outputting an equipment-level inspection report and a risk level conclusion. According to the invention, the inspection abnormity identification precision and the structure-thermal field reasoning stability are improved.
Owner:TIANJIN HONGBANG TECH CO LTD

Old people cognitive ability evaluation system and method based on multi-modal fusion

The invention discloses an old people cognitive ability assessment system and method based on multi-modal fusion, and relates to the field of old people cognitive assessment, and the method comprises the steps: firstly obtaining the current multi-modal data of a user, and fusing the current multi-modal data into a current feature vector; then, instead of being compared with the universality standard, the personal baseline portrait of the user is called, and the score of the difference degree between the current state and the historical baseline of the user is calculated; particularly, the core index of the drift speed is introduced, and the difference degree and the historical change rate are combined, so that quantitative modeling is carried out on the changed speed. By analyzing the amplitude and the speed of the change at the same time, normal aging with gentle change and low speed and pathological recession with violent change and high speed can be effectively distinguished, so that the key technical problem that the normal aging and the pathological recession are easy to be confused in the background technology is accurately solved.
Owner:ZHEJIANG FUBAO INTELLIGENT TECH CO LTD

Urban space intelligent processing method based on multi-modal fusion

The invention provides an urban space intelligent processing method based on multi-modal fusion, and the method comprises the steps: taking multi-modal data as input, and constructing a unified data stream processing and feature alignment mechanism; a physical space is used as a core framework, and the multi-modal data is converted into a space behavior graph with space-time position semantics; constructing an entity attribute-relation type-influence weight ternary interaction model on the basis of an interaction layer of the spatial behavior map, and analyzing a human, object and environment ternary interaction relation in the city and the park based on the ternary interaction model; designing a space intelligent engine with time sequence modeling and dynamic prediction capabilities; and constructing a task processing system. According to the invention, through deep combination of multi-modal fusion and a space intelligent technology, full-link upgrading of urban space from data perception to intelligent decision is realized, and powerful technical support is provided for fine management, efficient operation and safety guarantee of complex space scenes.
Owner:SHANGHAI ELECTRIC SMART CITY INFORMATION TECH CO LTD

Electric power operation risk early warning method and system based on knowledge enhancement and multi-modal fusion

The invention discloses an electric power operation risk early warning method and system based on knowledge enhancement and multi-modal fusion. The method comprises the steps that video monitoring data, sensor monitoring data and service system data are collected in real time through multi-source sensing equipment deployed on an electric power operation site; the method comprises the following steps of: extracting entities and relationships from unstructured texts such as regulation documents and job logs by utilizing a natural language processing technology based on deep learning, extracting behavior characteristics from video streams by adopting a computer vision algorithm, and constructing an electric power security knowledge graph with dynamic updating capability; designing a multi-modal feature fusion algorithm based on an attention mechanism, and effectively integrating visual features, text features and sensor data; a graph neural network is adopted to train a dynamic risk prediction model to carry out risk prediction, intelligent research and judgment of electric power operation risks are realized, accurate management and control of the risks are realized through a grading early warning mechanism, and closed-loop management from risk perception to early warning treatment is formed.
Owner:FUJIAN YIRONG INFORMATION TECH

Visual classification processing method and device based on large model and multi-modal data fusion

The invention relates to the field of visual processing, and provides a visual classification processing method and device based on large model and multi-modal data fusion. The method comprises the following steps: inputting a to-be-classified input image and a corresponding category text description into a text encoder for multi-level feature extraction to obtain global text features and local text features; performing fine-grained cross-modal alignment on the local visual features and the local text features, calculating association weights between the image regions and the text phrases through a bidirectional cross attention mechanism, and generating aligned intermediate features; splicing and fusing the aligned middle features and the global visual features, and inhibiting background noise in a fusion result and reinforcing discriminative features in the fusion result through a feature mask algorithm in combination with the global text features to obtain multi-modal fusion features; and synchronously inputting the multi-modal fusion features into a multi-space classifier to generate respective classification results, and adaptively outputting an image classification result according to a confidence threshold in combination with a dynamic routing mechanism.
Owner:SUZHOU YINPO TECHNOLOGY DEVELOPMENT CO LTD