Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2587 results about "Multimodal data" patented technology

A multimodel database is a data processing platform that supports multiple data models, which define the parameters for how the information in a database is organized and arranged. Being able to incorporate multiple models into a single database lets information technology (IT) teams and other users meet various application requirements without...

System and method for causality-augmented generative intelligence to discover non-obvious insights from heterogeneous data sources

The present invention provides a system and method for causality-augmented generative intelligence capable of autonomously discovering non-obvious actionable insights from heterogeneous and multimodal data sources. The system integrates a data ingestion unit for semantic and temporal harmonization of structured and unstructured datasets, a causal inference processor for constructing a dynamically evolving directed causal knowledge representation using perturbation-based validation, a latent representation processor that combines multimodal semantic embeddings with causal parameters to generate fused latent vectors, and a generative insight processor utilizing causally constrained generative reasoning to synthesize hypotheses anchored to verified cause-effect dependencies. A validation processor performs counterfactual assessment and observational verification to ensure retention of only those insights that remain consistent with causal ground truth.
Owner:MIA MD TOFAYEL GONEE MANIK

Multi-modal causal reasoning and explaining method, device, equipment and medium

PendingCN120952184ABiological modelsInference methodsCausal strengthCausal reasoning
The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a multi-modal causal reasoning and interpretation method, device, equipment and medium, and the method comprises the steps: obtaining original data streams of at least two different modals, and extracting modal features; a cross-modal attention mechanism is utilized to fuse modal features, and causal features are extracted through feature distillation; constructing a dynamic causal graph based on causal features, and updating an edge weight through a causal intensity function; identifying the causal relationship in the dynamic causal graph and performing anti-factual reasoning verification to evaluate the reliability of the causal relationship; and generating a causal interpretation result in combination with the dynamic causal graph and the causal relationship reliability. According to the method, the multi-modal data are fused, the causal features are extracted, and dynamic causal graph updating and anti-factual reasoning verification are combined, so that reliable modeling and explanation of the causal relationship in a complex scene are realized, the defects of single modal or simple fusion in the prior art are overcome, and the accuracy and interpretability of causal reasoning are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Electric power design knowledge base construction method fusing multi-modal data and RAG technology

The invention relates to a multi-modal data and RAG technology fused power design knowledge base construction method, and belongs to the technical field of power software development. The method comprises the following steps: carrying out collection and information extraction on multi-source heterogeneous original data; the method comprises the following steps of: constructing a multi-dimensional knowledge element structure containing parameters, specifications and case relationships by carrying out classification, specialized and precise processing and cross-modal association on data; based on a vector, graph and relational database mixed storage architecture, semantic vector efficient retrieval, knowledge graph relation management and business data synchronization are achieved respectively; and a dynamic optimization result is subjected to hybrid retrieval, a dual-drive reasoning mechanism outputs compliance conclusions and bases, and a retrieval enhancement generation service ensures that output contents conform to specifications. And systematic management and intelligent application of the electric power design knowledge are realized.
Owner:常州常供电力设计院有限公司

Multi-modal visual fusion complex scene small target detection tracking method and system

The invention discloses a multi-modal visual fusion complex scene small target detection tracking method and system, and relates to the technical field of unmanned aerial vehicle target tracking, and the method comprises the steps: employing a visible light camera, an infrared thermal imager and a laser radar sensor which are carried on an unmanned aerial vehicle platform, and synchronously collecting RGB images, thermal infrared images and point cloud data; the consistency of the multi-modal data is ensured through data preprocessing and space-time alignment; constructing a lightweight double-branch network to extract multi-scale features, generating a fusion feature map by adopting adaptive weighted fusion, and generating depth information by utilizing point cloud to assist in scale estimation; a small target detection head is designed based on the fusion feature map, and precise detection is realized in combination with a feature pyramid network, adaptive scale prediction and a context awareness suppression mechanism; furthermore, through multi-mode cooperative tracking, including target association, spatio-temporal context modeling, trajectory prediction and a re-detection mechanism, tracking continuity is ensured.
Owner:BEIJING INSTITUTE OF GRAPHIC COMMUNICATION

Robot multi-modal fusion autonomous decision-making method and system based on large language model

The invention relates to the technical field of robot decision making, and provides a robot multi-modal fusion autonomous decision making method and system based on a large language model.The method comprises the steps that a robot obtains multi-modal environment information through a visual sensor, a touch sensor, an auditory sensor and a laser radar which are carried by the robot; performing preliminary filtering and noise reduction processing on the original sensor data, and synchronously recording all the sensor data by timestamps; performing space-time semantic alignment on the preprocessed multi-modal data, mapping pixel coordinates of a target in a visual target coordinate quantization original image to a robot coordinate system, performing uncertainty evaluation on a multi-modal signal through a dynamic Bayesian network, and taking entropy or variance as an uncertainty quantitative evaluation index. According to the method, the information quality is improved from a data fusion source, accurate and reliable basic support is provided for subsequent decision making, and decision making errors caused by data deviation are greatly reduced.
Owner:ANHUI UNIV +1

Lower limb weight-bearing gait rehabilitation training system

The invention relates to the technical field of medical rehabilitation, and discloses a lower limb weight-bearing gait rehabilitation training system which comprises a data acquisition module, a data processing and analysis module, a patient individualized modeling module, an intelligent decision and control module, a rehabilitation execution module and a man-machine interaction and medical information interface module which are in communication connection through a network. The data acquisition module is used for acquiring multi-modal data of a patient in real time, and the multi-modal data comprises static sign data, dynamic physiological parameters, kinematics and dynamics parameters and non-motion physiological and psychological state data; and the data processing and analysis module is used for carrying out preprocessing, feature extraction and deep analysis on the original data, and outputting a structured patient individualized feature vector and an evaluation result. According to the invention, a patient three-dimensional skeletal muscle digital twinborn model is constructed through the patient individualized modeling module, and in combination with a continuous learning intelligent model library, body sign differences of different patients can be accurately adapted.
Owner:SHANGHAI TIANYOU HOSPITAL CO LTD

Multi-modal metadata alignment fusion method and device, equipment and storage medium

The invention discloses a multi-modal metadata alignment fusion method and device, equipment and a storage medium, and the method comprises the steps: carrying out the metadata extraction of structured data, unstructured text data and image data, and generating multi-modal metadata with semantic annotations; establishing a shared semantic embedding space, and mapping the multi-modal metadata to the shared semantic embedding space for coding to obtain a unified spatial vector; determining an alignment candidate pair from the unified spatial vector through similarity calculation, and identifying a semantic relationship of the alignment candidate pair; and performing conflict detection on the aligned candidate pairs and the corresponding semantic relationships, and resolving conflicts based on weight weighting fusion to obtain a unified metadata system. According to the method, improvement and optimization are carried out from multiple aspects of multi-modal data processing, semantic understanding, alignment accuracy, conflict resolution and the like, the defects in the prior art are overcome, and a more accurate and comprehensive multi-modal metadata alignment fusion result can be provided.
Owner:CHINESE PEOPLES LIBERATION ARMY INFORMATION SUPPORT CORPS ENGINEERING UNIVERSITY

Persistent Cognitive Machine with Temporally Synchronized Multimodal Processing and Typed Latent Entity Management

A system and method for persistent cognitive computation with temporally synchronized multimodal processing implements a geometric approach to artificial intelligence through typed latent entities within a dynamic manifold substrate. The system maintains a latent manifold incorporating heterogeneous data modalities where local curvature reflects semantic density and typed entities are stratified according to structural properties. Temporal synchronization coordinates asynchronous multimodal data streams through generation of temporal alignment fields within the manifold that preserve semantic coherence across modal boundaries. Type-aware geometric operations enforce operation legality based on entity type and local manifold geometry, enabling structured recombination, compression, and traversal while preventing semantic distortion. The system executes synchronized manifold reorganization during idle periods through coordinated optimization operations including perturbation analysis and topological surgery. This architecture enables persistent memory through geometric encoding where frequently accessed concepts develop high-curvature regions and cognitive patterns emerge from usage-based manifold evolution.
Owner:ATOMBEAM TECH INC

Multi-modal data fused LSTM photovoltaic power generation power prediction method and system

The invention discloses an LSTM photovoltaic generation power prediction method and system fused with multi-modal data, and the method comprises the steps: generating a time-space consistent multi-modal training feature matrix according to historical meteorological data, historical photovoltaic monitoring data and corresponding historical generation power data; based on the multi-modal training feature matrix, constructing a multi-input-channel LSTM prediction model, and outputting a cross-regional generalization pre-training model; according to the real-time meteorological data of the target area, the current photovoltaic monitoring data and the electricity price fluctuation curve, updating the pre-training model on line by adopting a dual reward reinforcement learning strategy, and outputting a dynamically optimized power prediction model; and inputting the meteorological data and the photovoltaic monitoring data at the current moment into the dynamically optimized power prediction model to generate a photovoltaic power generation power prediction sequence in the future preset time. According to the embodiment of the invention, high-precision and high-adaptability photovoltaic power generation power prediction can be realized.
Owner:ZHEJIANG POST & TELECOMM

Mine ecological restoration dynamic optimization method and system based on multi-modal data fusion

The invention provides a mine ecological restoration dynamic optimization method based on multi-modal data fusion, and the method comprises the steps: obtaining multi-modal data of a to-be-restored mining area, and carrying out the preprocessing and space-time alignment of the multi-modal data; a dynamic knowledge graph of the to-be-restored mining area is constructed, and the dynamic knowledge graph takes ecological elements and engineering entities as nodes, takes incidence relations between the nodes as edges, and takes the multi-modal data as dynamic attributes of the nodes and the edges; based on the dynamic knowledge graph and the fused multi-modal data, ecological risk diagnosis and stability prediction are carried out, and a diagnosis and prediction result is generated; based on the diagnosis and prediction result, performing restoration scheme simulation deduction in the digital twin environment by using an optimization algorithm, and outputting a dynamically optimized restoration scheme; and iteratively updating the dynamic knowledge graph according to the repair scheme execution effect. The method has the technical effect of improving the early warning timeliness and reliability.
Owner:SHENZHEN TIANJING YUHONG TECHNOLOGY CO LTD

Water resource predictive analysis method based on artificial intelligence

The invention relates to the technical field of water resource analysis, and discloses a water resource predictive analysis method based on artificial intelligence. The method relates to the technical field of water resource analysis, and comprises the following steps: acquiring an original hydrological data set including rainfall intensity, river flow and the like through a sensing terminal, and performing multi-modal data alignment to generate a hydrological space-time tensor; constructing a dynamic water level threshold response mechanism in combination with watershed topographic features to obtain a partition water level calibration matrix; inputting the hydrological feature map into a spatial-temporal feature coupling network containing a long-short-term memory module and a spatial self-attention module to generate a hydrological feature map; constructing a multi-dimensional abnormal association tensor based on the multi-dimensional abnormal association tensor, and identifying rainfall flood event nodes by using an adaptive sliding window detection algorithm; and an optimal hydrological parameter set is obtained through genetic algorithm optimization, and the three-dimensional hydrological dynamic model is driven to establish a mapping relation chain. The method can effectively fuse hydrological data spatio-temporal characteristics, and improves the accuracy and efficiency of water resource prediction analysis.
Owner:盱眙县水资源管理所

Complex scene-oriented end-to-end multi-modal content unified perception method and system

The invention belongs to the technical field of multi-modal data processing, and discloses a complex scene-oriented end-to-end multi-modal content unified perception method and system. The method comprises the following steps: performing intelligent sensing, identification, acquisition, screening and standardization processing on multi-modal data content, and outputting structured and standardized multi-modal content data; inputting a feature extraction model in parallel, and performing multi-modal content data feature unified modeling and preliminary fusion by adopting a multi-modal unified encoder which is internally integrated with a cross-modal attention layer and is based on a Transform architecture; cascade fusion high-order semantic representation is extracted step by step through a multi-stage and multi-level cascade cross-modal fusion structure; and performing deep semantic analysis on the extracted cascaded fusion high-order semantic representation by adopting a pre-trained semantic understanding model. According to the method, the fusion depth and perception precision of the multi-modal information in a complex scene are improved, and efficient and accurate understanding and interactive response of the multi-modal content are facilitated.
Owner:SHENZHEN WANGLIAN ANRUI NETWORK TECH CO LTD

Domestic big language model retrieval enhancement generation method for customs

The invention discloses a customs-used domestic large language model retrieval enhancement generation method, which comprises the following steps of S1, constructing a GraphRAG, and implementing dynamic relationship deconstruction on multi-modal data such as customs announcements, enterprise customs declarations and international provisions based on the deep semantic understanding capability of a large language model; s2, dividing laws and regulations sub-communities based on a Leiden algorithm, combining a Leiden community discovery algorithm with laws and regulations effectiveness analysis, and constructing a triad effectiveness map of laws and regulations clauses-revision events-effective areas; s3, hierarchically retrieving a framework, and decomposing a retrieval process into three-level probability decisions; and S4, establishing a knowledge graph full-stack system architecture of the multi-source heterogeneous data. According to the method, a GraphRAG technology is taken as a core carrier, and breakthrough of a customs complex knowledge scene is realized through double innovation paths: firstly, a graph structure retrieval enhancement model adaptive to customs business characteristics is constructed, and secondly, a dynamic governance mechanism of a Leiden community discovery algorithm optimization regulation system is introduced, so that the decision reliability of a customs intelligent supervision system is improved.
Owner:HUANGPU CUSTOMS DISTRICT OF PEOPLES REPUBLIC OF CHINA

Silicon carbide part stress distribution monitoring and crack risk prediction method

The invention relates to the technical field of deep learning, in particular to a stress distribution monitoring and crack risk prediction method for a silicon carbide part, which realizes comprehensive sensing of the stress state of the silicon carbide part, accurate positioning of a risk area and advanced early warning of a crack fault. The method comprises the following steps: synchronously acquiring multi-modal data through multiple types of sensors, and realizing cross-modal time sequence synchronization through feature alignment; designing a crack risk multi-branch feature extraction module, and respectively extracting general depth features and risk features oriented to thermal stress mismatch, microcrack evolution and structural instability through a shared backbone network and a special branch network; constructing a stress nephogram generation and risk area positioning module based on a graph neural network, and realizing visual reasoning and risk area marking from discrete features to full-field stress distribution; and designing a crack risk comprehensive prediction module based on multi-dimensional risk feature fusion, fusing an instantaneous state and an evolution trend, outputting a multi-risk confidence vector and triggering graded early warning.
Owner:EVIC SEMICONDUCTOR TECHNOLOGY (SHANGHAI) CO LTD

Gradienter attitude real-time calibration method based on multi-modal data fusion

The invention relates to the technical field of attitude measurement, and discloses a gradienter attitude real-time calibration method based on multi-modal data fusion, which comprises the following steps: acquiring multi-modal original data and completing unified preprocessing to obtain multi-modal data; reconstructing a liquid surface form in a physical domain neural operator layer, and outputting a physical domain attitude and a residual error; noise and drift are deduced in a sensing domain neural operator layer, and a sensing domain attitude estimation value, an offset parameter, a scale parameter and a residual error are output; establishing a deviation memory bank, updating by using residual errors and historical results, generating a long-term drift compensation amount, and superposing a sensing domain result; inputting a physical domain and a compensated sensing domain result into a dynamic constraint reversible transformation model, and outputting a fusion attitude and uncertainty; and executing slow variable refining compensation on the updated parameters, and outputting final real-time calibration attitude and quality information. According to the invention, by introducing multi-modal data fusion and double-layer reversible neural operator modeling, real-time, high-precision and long-term stable calibration of the attitude of the gradienter is realized.
Owner:NANTONG DIO AMP PHOTOELECTRIC TECH CO LTD

Multi-modal information fusion body-equipped intelligent robot control method

The invention discloses a control method for a multi-modal information fusion intelligent robot with a body. The control method comprises the following steps: initializing a system, and collecting surrounding physical environment and object state information and a natural language instruction of a user; performing scene understanding and task analysis, processing data through a multi-modal information fusion mechanism and a cross-modal attention module, and generating unified multi-modal data; task planning and priority ranking are carried out, complex tasks are decomposed into subtask sequences, and a priority ranking layer dynamically adjusts the execution sequence; performing action execution and feedback adjustment, and generating a control instruction through a self-adaptive operation control algorithm; continuous learning and strategy verification are carried out, integrated execution is realized by using a hybrid AI system, and the robustness of an operation strategy is verified through a simulation environment; and closed-loop iteration is carried out to realize real-time response of the intelligent robot with the body. According to the method, the perception understanding precision and the task execution efficiency of the intelligent robot with the body are improved, the operation precision adaptability and the system robustness flexibility are guaranteed, and the method is suitable for multiple scenes.
Owner:ROSIWIT TECHNOLOGY CO LTD +1

Multi-modal data fusion air conditioner optimization control method and system

The invention discloses a multi-modal data fusion air conditioner optimization control method and system, and belongs to the technical field of intelligent building equipment control. According to the method, temperature, humidity, energy consumption, user behaviors and meteorological data are collected through a multi-source sensor, the data are subjected to dynamic window standardization processing and then input into a gated convolution LSTM network to predict the system state, an optimization strategy is generated in combination with an online reinforcement learning algorithm, multi-modal instructions are dynamically weighted and fused, and execution parameters are adjusted in real time through a feedback correction mechanism. The system comprises a multi-source sensing array, an edge computing unit, a strategy optimization engine and an intelligent execution controller. According to the method, through collaborative optimization of multi-modal spatial-temporal feature fusion and deep reinforcement learning, the problem of unbalance of energy efficiency and comfort is solved, the energy efficiency level and the user comfort of the air conditioning system are remarkably improved, and the method has the characteristics of real-time response and high stability, is suitable for intelligent air conditioning control of modern buildings and has wide application prospects.
Owner:TIANJIN CONSTR ENG GRP ARCHITECTURAL DESIGN CO LTD

Space-time decoupling sentiment analysis method and system based on multi-modal data

The invention discloses a space-time decoupling sentiment analysis method and system based on multi-modal data, and relates to the technical field of multi-modal sentiment analys.The method comprises the steps that text, audio and video initial features are input into a space-time decoupling and language focusing fusion model to be processed, and corresponding feature extraction is conducted on the enhanced text, audio and video features; obtaining specific text, audio and video features; performing shared feature extraction on the enhanced text, audio and video features to obtain shared text, audio and video features; the specific text, audio and video features and the shared text, audio and video features are input into a language focusing attractor module for multi-level feature extraction, and low-level features, middle-level features and high-level features are obtained respectively; emotion prediction is carried out based on the low-level features, the middle-level features and the high-level features, emotion prediction results of the features of all the levels are fused, a final emotion prediction result is obtained, and decoupling and fusion in multi-modal emotion analysis are achieved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

GNSS interference analysis method and system based on big data

The invention relates to the technical field of GNSS interference analysis, in particular to a GNSS interference analysis method and system based on big data. The method comprises the following steps: deploying a multi-source signal acquisition sensor to acquire GNSS signals, and carrying out standardization processing to generate a multi-modal standard data set; performing non-uniform sampling alignment on the multi-modal standard data set to obtain a GNSS synchronization spatio-temporal data stream; mapping the GNSS synchronization spatio-temporal data stream to a three-dimensional space grid to carry out multi-dimensional tensor construction, and obtaining a GNSS interference detection feature tensor; therefore, through multi-source multi-modal data fusion and three-dimensional space propagation modeling, the problems of data asynchronization, low space resolution and single alarm strategy in traditional GNSS interference detection are solved, and the accuracy of interference detection and the intelligent level of response are improved.
Owner:CHANGSHA TECH RES INST OF BEIDOU IND SAFETY CO LTD

Cerebral stroke multi-mode early screening intelligent evaluation system based on large model

The invention discloses a cerebral apoplexy multi-mode early screening intelligent evaluation system based on a large model, and relates to the technical field of medical health information, the cerebral apoplexy multi-mode early screening intelligent evaluation system comprises an intelligent management platform, and the intelligent management platform is in communication connection with the following modules: a multi-source heterogeneous data fusion engine, the data integration module is used for integrating multi-modal data including clinical data and terminal health data and constructing a health portrait of a patient; and the cerebral apoplexy knowledge graph construction platform is used for constructing a cerebral apoplexy domain knowledge graph in combination with evidence-based medical knowledge. By combining the digital twinning technology and the intelligent risk assessment engine, the influence of different intervention schemes on the cerebral apoplexy risk can be simulated, personalized intervention suggestions are generated, a patient is helped to reduce the cerebral apoplexy risk and change from passive prediction to active intervention, the patient is helped to take effective measures earlier, the health condition is improved, and the patient experience is improved. The occurrence of cerebral apoplexy is prevented, so that the disability rate and the death rate caused by cerebral apoplexy are reduced.
Owner:GUILIN MEDICAL UNIVERSITY +1

Method for measuring residual stress of metal component

The invention relates to the technical field of material mechanical property testing, in particular to a method for measuring residual stress of a metal component. According to the technical scheme, the method for measuring the residual stress of the metal component comprises a working process for measuring the residual stress of the metal component; according to the method, a technical path of multi-modal data assimilation is adopted, and a mechanical release mechanism of a drilling method and an ultrasonic stress measurement technology are deeply fused to calibrate and correct mass sound velocity data obtained by carrying out rapid ultrasonic scanning on the whole component on site, so that the precise measurement of the residual stress state of the metal component is realized, and the measurement precision is improved. According to the method, data authority and high credibility of a drilling method in a key area are reserved, the advantages of high efficiency and wide coverage of full-field measurement of an ultrasonic method are brought into full play, and efficient and accurate stress information is provided on the premise that components are not excessively damaged.
Owner:AVIC METAL MATERIAL PHYSICAL & CHEM TESTING TECH CO LTD

Multi-mode fusion extra-high voltage converter transformer fault diagnosis method and system

The invention relates to the technical field of intelligent diagnosis of power equipment, and provides a multi-mode fusion extra-high voltage converter transformer fault diagnosis method and system, and the method comprises the steps: collecting a multi-mode signal from a sensor of an extra-high voltage converter transformer, and classifying the signal into a plurality of modes; performing intra-modal feature reinforcement learning on the multi-modal data by adopting a SimCLR framework to obtain a discriminative representation feature vector zm of each modal; inputting the zm into a Transform branch encoder of a corresponding mode, and obtaining context feature vector enhancement representation hm of each mode; mapping the hm of different modalities to the same dimension in a unified manner, and performing weighted attention fusion to obtain feature vectors zf of all modalities after fusion; and inputting a multi-layer perceptron classifier (MLP) to obtain a fault category prediction result of the extra-high voltage converter transformer. According to the method, the multi-mode signals are processed and fused in parallel, and the fault recognition capability of the model on the extra-high voltage converter transformer in the complex operation state is improved.
Owner:STATE GRID ANHUI ULTRA HIGH VOLTAGE CO

Medical data management method and system based on multi-modal fusion and privacy protection

The invention discloses a medical data management method and system based on multi-modal fusion and privacy protection, and the method comprises the steps: receiving multi-modal medical data, and executing cross-modal embedding learning to generate a joint embedding vector; constructing a patient multi-modal graph based on the joint embedded vector, and detecting abnormal inconsistency between the cross-modal data by using a graph attention network to generate an abnormal detection result; a privacy protection strategy is dynamically adjusted according to the anomaly detection result and the access context information, and desensitization processing is conducted on the medical data; and constructing a global model through security aggregation based on the desensitized medical data to generate an optimal governance action and apply the optimal governance action to a data management process. Through organic cooperation of multi-modal fusion, anomaly detection, adaptive privacy protection, security aggregation and dynamic strategy optimization, a closed-loop medical data governance platform is successfully constructed, and the problems of multi-modal data fragmentation, privacy compliance conflict and dynamic governance deficiency are effectively solved.
Owner:BEIJING CHANGCHANGJIA INFORMATION TECH CO LTD

Carotid plaque stability and stroke risk prediction system fusing multiple modes

The invention relates to the technical field of image recognition, and discloses a carotid plaque stability and stroke risk prediction system fusing multiple modes, and the system comprises a data collection and fusion module, a feature extraction and association module, a risk assessment and layering module, an intervention decision module, and a report generation and feedback module. When carotid plaque stability and cerebral apoplexy risk prediction is carried out, plasma proteomics, radiomics and clinical data are synergistically integrated by establishing a multi-modal data fusion analysis framework, so that the limitation that a traditional method depends on a single data source is overcome; the plaque risk can be evaluated from multiple dimensions of biological activity and morphological features, the comprehensiveness and accuracy of risk prediction are improved, a more reliable diagnosis basis is provided for clinic, interaction and sensitivity influence among different modal features can be adaptively quantified by introducing dynamic feature correlation modeling and a real-time weight calibration mechanism, and the risk prediction accuracy is improved. And objectivity and consistency of risk assessment results are ensured.
Owner:LINFEN CENT HOSPITAL (THE FOURTH PEOPLES HOSPITAL OF LINFEN)

Multi-modal large model incremental training data screening method

The invention provides a multi-modal large model incremental training data screening method, and relates to the technical field of data processing, and the method comprises the steps: executing modal structure analysis on newly added multi-modal data, extracting each modal vector, calculating a semantic matching degree, and removing samples lower than a preset first threshold value; calculating a multi-level semantic distance between a sample embedding vector and a historical clustering center in a unified semantic space, and dividing a core semantic region sample, a boundary semantic region sample and a discrete semantic region sample according to the change rate of the multi-level semantic distance; performing semantic fine-grained alignment on the boundary semantic region samples, when multimodal unstable distribution is detected, executing local context reconstruction to repair semantic deviation, and if the multimodal unstable distribution is still unstable, removing the semantic deviation; performing multiple rounds of small-batch reasoning, calculating a semantic stability coefficient based on a semantic prediction result, and when the semantic stability coefficient is lower than a preset second threshold value, determining that the sample is a potential drift sample and removing the potential drift sample; constructing an incremental training data set; according to the method, the autonomy and accuracy of incremental training data screening are improved.
Owner:ZHONGSHU (XIAMEN) INFORMATION TECH CO LTD +1

System and method for artificial intelligence based field service assistance for telecommunications operations

A system and method for field service assistance for telecommunications operations are described, which utilize a data acquisition module configured to receive multimodal data inputs including structured and unstructured data from field operations. A preprocessing module normalizes the multimodal data inputs to generate pre-processed data. A vectorization module transforms the pre-processed data into numerical vector representations using domain-specific embedding models trained on telecom equipment data, implementing convolutional neural networks for image feature extraction and transformer-based encoders for text vectorization. A contextual retrieval module retrieves contextually relevant historical data from a vector database by computing similarity metrics between current job vectors and stored job completion vectors. A response generation module processes the numerical vector representations and retrieved contextual data using an evolutionary algorithm engine to generate structured job summaries and real-time field recommendations.
Owner:ANAND PAWAN +2

Underground equipment fault real-time diagnosis method and system based on edge calculation

The invention provides an underground equipment fault real-time diagnosis method and system based on edge calculation, and relates to the technical field of coal mine safety production, and the method comprises the steps: collecting multi-modal data through a distributed sensor network, extracting multi-scale time sequence features, projecting the features to a Lie group manifold space, constructing a coupling mapping relation matrix, obtaining fusion features, and carrying out the real-time diagnosis of an underground equipment fault; and constructing a causal directed acyclic graph based on a topological connection relationship and a Granger causal coefficient, executing Bayesian probabilistic reasoning, determining an execution strategy in combination with entropy similarity matching, and performing deep time-frequency analysis and causal chain verification. High-precision real-time diagnosis of equipment faults in an underground complex environment is realized, and the fault early warning accuracy is improved.
Owner:BEIJING YANGGUANG JINLI TECH DEV

Power network attack chain dynamic deduction and intelligent response process method, system and device based on deep reinforcement learning, and medium

The invention discloses a power network attack chain dynamic deduction and intelligent response process method, system and device based on deep reinforcement learning and a medium, and belongs to the technical field of network security and power system protection. Multi-modal data is aligned and normalized, an event view cache is constructed, and the generalization detection capability on process camouflage and memory injection attacks is improved through a federated learning collaborative detection mechanism; based on the event view cache and historical threat intelligence, generating a dynamic attack knowledge graph, constructing a deep reinforcement learning model taking the attack knowledge graph as an environment, calculating an attack influence index by using a Bayesian network, and generating a differentiated security response instruction; and realizing attack path backtracking and attack source positioning based on the attack knowledge graph. According to the method, multi-modal data fusion analysis and strategy adaptive updating are realized, and the attack chain identification accuracy and evidence chain construction integrity are remarkably improved.
Owner:GUANGXI POWER GRID CORP

Multimodal Data Ingestion And Retrieval For Agent Systems

Techniques for multimodal document retrieval are disclosed herein. Multimodal documents that include both textual and graphical components are retrieved from a knowledge base by a multimodal retrieval augmented generation (RAG) agent in response to a query. The documents and / or components or chunks thereof are retrievable by the RAG agent from the knowledge base using the semantic summaries and / or vector search of embeddings in the knowledge base that are generated from text extracted from processing non-textual components of the data. The RAG agent classifies the query type to determine whether to use a semantic match for text or image summaries, full text semantic search, vector cosine similarity search, and / or other multimodal vector search. The RAG agent performs types of searches selected based on the modality used to generate the response to the query.
Owner:ORACLE INT CORP

Multi-modal time sequence anomaly analysis method and device, equipment and medium

The invention relates to the technical field of data analysis, and discloses a multi-modal time sequence anomaly analysis method, device, equipment and medium, and the method comprises the steps: collecting visual data, audio data and process text data, carrying out the preprocessing and standardization processing of different types of data, constructing a multi-modal data set with aligned timestamps, and storing the multi-modal data set in a database; the method comprises the following steps: extracting a time-frequency dynamic feature and a semantic vector feature, extracting a map structure feature, a time-frequency dynamic feature and a semantic vector feature, fusing the features by using a cross-modal attention mechanism to generate a fused feature vector, finally performing analysis processing based on the fused feature vector, and outputting an analysis result. According to the method, the multi-modal data set with consistent time is constructed, the structural features of various modals are extracted, and the cross-modal attention mechanism is introduced to realize deep fusion of the feature level, so that the problems of single information utilization and weak feature relevance of the existing detection means are solved, and the comprehensiveness of defect detection and the accuracy of fault diagnosis are improved.
Owner:SUN YAT SEN UNIV