Intelligent decision and risk management system for multi-modal semantic alignment
The intelligent decision-making and risk management system with multimodal semantic alignment generates multiple highly reliable inference chains and makes dynamic decisions by combining knowledge conflicts and historical risks. This solves the problem of insufficient cross-modal semantic association in traditional systems and realizes refined and differentiated risk management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2026-03-20
AI Technical Summary
Traditional systems have shortcomings in multimodal data processing, reasoning and risk assessment, and response methods. They are unable to achieve deep semantic association across modalities, multidimensional risk assessment, and dynamic response, resulting in incomplete risk judgment and waste of resources.
The intelligent decision-making and risk management system adopts multimodal semantic alignment. Through the semantic alignment association representation module, the possible reasoning chain generation module, and the risk decision analysis module, it generates multiple highly reliable reasoning chains and combines knowledge conflicts and historical risks to make dynamic decision responses.
It enables multi-dimensional and multi-angle analysis of potential risks, dynamically adjusts response methods, ensures timely handling in high-risk scenarios and resource optimization in low-risk scenarios, and achieves refined and differentiated risk management.
Smart Images

Figure CN120975241B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of semantic decision management, more specifically, it relates to an intelligent decision and risk management system for multi-modal semantic alignment. BACKGROUND
[0002] In the field of semantic decision management, traditional systems have obvious shortcomings in processing multi-modal data. They are often limited to the analysis of single modal information, making it difficult to effectively fuse different types of data such as text, images, audio, and sensors, resulting in the fragmentation of deep semantic associations across modalities, making it difficult to fully capture complete information in complex scenarios, and thus affecting accurate judgment of the overall situation, creating hidden dangers for subsequent decision-making.
[0003] Existing systems also have shortcomings in reasoning and risk assessment. On the one hand, the generated reasoning paths are relatively single, lacking coverage of multiple potential possibilities, and the reasoning process has weak explainability, making it difficult to trace the formation logic of the conclusion; on the other hand, the risk assessment dimensions are limited, relying mainly on a few indicators for judgment, failing to combine multi-dimensional features such as knowledge conflicts and historical risks, and lacking effective assessment of the correlation between different reasoning paths, resulting in an incomplete understanding of potential risks.
[0004] The response mode of traditional risk management mechanisms is relatively rigid, often using fixed threshold standards and uniform disposal strategies. This model is difficult to adapt to the needs of dynamic changing scenarios, and may cause losses due to delayed response in high-risk situations, while in low-risk scenarios, it may cause resource waste due to excessive disposal, failing to achieve a dynamic balance between risk prevention and execution efficiency, and failing to meet the actual needs of fine and differentiated management.
[0005] Based on the above, the present application proposes an intelligent decision and risk management system for multi-modal semantic alignment. SUMMARY
[0006] In view of the shortcomings of the prior art, the purpose of the present application is to provide an intelligent decision and risk management system for multi-modal semantic alignment.
[0007] To achieve the above purpose, the present application provides the following technical solutions:
[0008] The intelligent decision and risk management system for multi-modal semantic alignment comprises a semantic alignment correlation representation module, which collects the semantic features of each modal data in real time, processes the semantic features of each modal data for semantic alignment, and obtains cross-modal fusion semantic correlation representation;
[0009] A possible reasoning chain generation module generates multiple possible reasoning chains based on the cross-modal fusion semantic correlation representation;
[0010] a risk decision analysis module configured to determine a decision response value of each possible reasoning chain; a determination process of the decision response value of the possible reasoning chain: obtaining a risk fingerprint value and a reasoning overlap value of a possible reasoning chain, and calculating the decision response value of the possible reasoning chain based on the risk fingerprint value and the reasoning overlap value;
[0011] a risk fingerprint value of the possible reasoning chain: selecting a possible reasoning chain, obtaining a knowledge conflict risk value and a historical risk density value of the possible reasoning chain, and calculating the risk fingerprint value of the possible reasoning chain based on the knowledge conflict risk value and the historical risk density value;
[0012] a risk decision control module configured to determine a decision response mode of each possible reasoning chain according to the decision response value, and to execute a decision response according to the decision response mode of each possible reasoning chain.
[0013] Further, the process of generating a plurality of possible reasoning chains is as follows: step 1: inputting the cross-modal semantic association representation into a generative engine, and the generative engine analyzing the core association between the features through a cross-attention mechanism; step 2: setting diversified decoding parameters to guide multi-path exploration; and step 3: generating a plurality of possible reasoning chains based on the semantic association.
[0014] Further, the process of obtaining the knowledge conflict risk value is as follows: selecting a possible reasoning chain, extracting key entities and relationships in the possible reasoning chain, calculating the total matching degree of the possible reasoning chain and the "entity-relation" in the knowledge graph, determining an anchoring coefficient, and calculating the knowledge conflict risk value based on the anchoring coefficient.
[0015] Further, the process of obtaining the historical risk density value is as follows: selecting a possible reasoning chain, counting the occurrence frequency proportion of the same type of possible reasoning chain in history, calculating the loss intensity by the ratio of the average single loss and the maximum loss, and calculating the historical risk density value based on the occurrence frequency proportion and the loss intensity.
[0016] Further, the process of obtaining the reasoning overlap value of the possible reasoning chain is as follows: selecting a possible reasoning chain, labeling the remaining possible reasoning chains as comparison reasoning chains, comparing each comparison reasoning chain with the possible reasoning chain, further obtaining a comprehensive overlap value of each comparison reasoning chain, and calculating the reasoning overlap value of the possible reasoning chain by summing and averaging the comprehensive overlap values of each comparison reasoning chain.
[0017] Further, the process of obtaining the comprehensive overlap value of the comparison reasoning chain is as follows: selecting a comparison reasoning chain, obtaining a feature overlap value and a conclusion overlap value between the comparison reasoning chain and the possible reasoning chain, and calculating the comprehensive overlap value of the comparison reasoning chain by summing the feature overlap value and the conclusion overlap value.
[0018] Further, the way to obtain the feature overlap value of the compared reasoning chain and the possible reasoning chain: obtain the intersection quantity and the union quantity of the compared reasoning chain and the possible reasoning chain, and calculate the feature overlap value by comparing the intersection quantity with the union quantity.
[0019] Further, the way to obtain the conclusion overlap value of the compared reasoning chain and the possible reasoning chain: obtain the fault type similarity and the solution similarity of the compared reasoning chain and the possible reasoning chain, and calculate the conclusion overlap value by summing and averaging the fault type similarity and the solution similarity.
[0020] Further, the decision response mode of each possible reasoning chain is determined according to the decision response value: set the decision response upper value and the decision response lower value, when the decision response value of the possible reasoning chain is greater than or equal to the decision response upper value, trigger an emergency response to the possible reasoning chain, when the decision response value of the possible reasoning chain is less than or equal to the decision response lower value, trigger a monitoring response to the possible reasoning chain, and when the decision response value of the possible reasoning chain is between the decision response upper value and the decision response lower value, trigger a pre-warning response to the possible reasoning chain.
[0021] Compared with the prior art, the present application has the following beneficial effects:
[0022] The system of the present application comprehensively and accurately captures cross-modal semantic association through multi-modal semantic alignment processing, and then generates multiple possible reasoning chains with reliability and explainability, and simultaneously calculates reasoning overlap values to evaluate the correlation between chains, realizing multi-dimensional and multi-angle analysis of potential risks. The risk fingerprint value and the reasoning overlap value are included in the quantitative calculation framework of the decision response value, and the decision response threshold is dynamically set, so that the system can flexibly trigger emergency, pre-warning or monitoring response according to different risk levels, realizing the refinement and differentiation of risk control. This mechanism not only ensures timely disposal in high-risk scenarios, but also avoids resource waste in low-risk scenarios, effectively balancing risk prevention and control and execution efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0023] Fig. 1 The principle block diagram of the multi-modal semantic alignment intelligent decision and risk control system;
[0024] Fig. 2 The flowchart for obtaining the decision response value of the possible reasoning chain;
[0025] Fig. 3 The flowchart for obtaining the risk fingerprint value of the possible reasoning chain. DETAILED DESCRIPTION
[0026] REFERENCE Figs. 1 to 3The application discloses a multimodal semantic alignment intelligent decision and risk management system, which comprises a semantic alignment correlation representation module, a possible reasoning chain generation module, a risk decision analysis module and a risk decision management module.
[0027] The semantic alignment correlation representation module collects semantic features of each modality data in real time, performs semantic alignment processing on the semantic features of each modality data, and obtains cross-modality fusion semantic correlation representation.
[0028] The types of the modality data include but are not limited to a text modality, an image modality, an audio modality and a sensor modality.
[0029] The semantic feature collection process of the text modality is as follows:
[0030] Step 1, text modality data collection of a data source; the following is an example of the data source, real-time log stream (such as system log, transaction note);
[0031] Step 2, lightweight preprocessing; the following is an example of the lightweight preprocessing, real-time word segmentation, stop word removal (preset high-frequency meaningless word table), and redundant calculation is avoided; a sliding window truncation (such as every 512 characters) is adopted for long text (such as real-time news stream), and the semantic integrity and processing speed are balanced;
[0032] Step 3, efficient semantic coding; the following is an example of the efficient semantic coding, a distilled version of a pre-training model (such as DistilBERT, TinyBERT and ALBERT) is adopted, the parameter quantity is reduced to 1 / 3-1 / 10 of that of an original model, the inference speed is increased by 2-3 times, and the core semantic extraction capability (such as word vector and sentence vector) is retained; an incremental coding is adopted for a streaming text (such as real-time dialogue), that is, based on the coding result of previous text, only the features of a newly added segment are updated, and repeated calculation is avoided (similar to the buffer mechanism of Transformer-XL).
[0033] Step 4, semantic feature extraction; the following is an example of the semantic feature, a sentence-level vector (used for topic classification, such as “complaint” and “consultation”) is extracted, a keyword weight (such as TF-IDF real-time calculation, and risk words such as “default” and “overdue” are identified), and a sentiment polarity (such as a lightweight model VADER is used to output positive and negative sentiment scores in real time).
[0034] The semantic feature collection process of the image modality is as follows:
[0035] Step 1, image modality data collection of a data source; the following is an example of the data source, a camera real-time stream (such as a monitoring picture and an industrial quality inspection lens);
[0036] Step 2, dynamic pre-processing; the following is an example of dynamic pre-processing, real-time resize (such as fixed to 224x224 pixels), normalization (subtract mean value / divide standard deviation), avoid redundant calculation of high-resolution images; for static scenes (such as fixed monitoring), use inter-frame sampling (such as processing 1 frame every 3 frames), for dynamic scenes (such as traffic monitoring), keep high frame rate (15-30 frames / second), balance efficiency and semantic integrity;
[0037] Step 3, lightweight visual coding; the following is an example of lightweight visual coding, using mobile-optimized models: MobileNetV3 (parameter quantity 3.2M), EfficientNet-Lite (inference speed is 4 times faster than ResNet), MobileViT (combination of CNN and Transformer lightweight design), extract high-level semantic features (such as "pedestrian", "vehicle", "defect area"); in industrial scenarios, use knowledge distillation to customize models: such as distilling the quality inspection knowledge of ResNet50 to ShuffleNet, retaining the ability to recognize specific semantics such as "welding defects" and "scratches", and reducing inference delay to <20ms.
[0038] Step 4, semantic feature extraction; the following is an example of semantic feature extraction, lightweight object detection model (such as YOLOv8-nano, PP-YOLOE-tiny) outputs object categories (such as "safety helmet", "open flame") and position coordinates (for spatial semantic association); scene classification features (such as using MobileNetV3 to output "workshop", "office" and other scene vectors), assist multi-modal environment understanding.
[0039] The semantic feature acquisition process of the audio modality is as follows:
[0040] Step 1, audio modality data acquisition of data source; the following is an example of data source, environmental sound sensors (such as industrial noise, abnormal sound);
[0041] Step 2, real-time pre-processing; the following is an example of real-time pre-processing, frame windowing (such as 20ms per frame, 50% overlap), noise reduction (using spectral subtraction to remove steady-state noise), voice activity detection (VAD, such as WebRTC-VAD), only processing segments containing valid sound (filtering silence / noise); convert to Mel-spectrum (Mel-spectrum), retain human ear-sensitive frequency features, dimension compressed to 80x100 (time x frequency);
[0042] Step 3, efficient audio coding; the following is an example of efficient audio coding, speech semantics: extract speech content features (such as "stop running", "temperature too high") with a light version of Wav2Vec2.0 (such as HuBERT-base-distilled), or output speaker emotion (such as "angry", "anxious") with Tiny-WavLM;
[0043] Step 4, semantic feature extraction; the following is an example of semantic feature extraction, speech-to-text (ASR) synchronous output of text features (combined with text modal processing), while retaining prosodic features (such as speech rate, pause) for sentiment analysis; extract voiceprint features (such as abnormal fluctuations in device vibration frequency) in industrial scenarios to identify "bearing wear" and other fault semantics through peak changes in the mel spectrum.
[0044] The semantic feature collection process of the sensor modal is as follows:
[0045] Step 1, sensor modal data collection of data sources; the following is an example of data sources, industrial sensors (vibration, temperature, pressure);
[0046] Step 2, real-time signal preprocessing; the following is an example of real-time signal preprocessing, filtering (such as using a Butterworth filter to remove high-frequency noise), normalization (to eliminate the range differences of different sensors); down-sampling: down-sample high-frequency signals (such as 1kHz sampling of vibration signals) to 200Hz, retaining key frequency components (such as device resonance frequency);
[0047] Step 3, time series semantic coding; the following is an example of time series semantic coding, frequency domain features: use fast Fourier transform (FFT) to extract spectral peaks (such as abnormal frequencies of motor vibration) in real time, or wavelet transform to extract time-frequency features (such as QRS complexes of electrocardiogram); time series model: use light-weight LSTM (such as LSTM-Lite) or Temporal Convolutional Network (TCN) to extract trend semantics (such as "temperature continues to rise", "pressure drops sharply");
[0048] Step 4, semantic feature extraction; the following is an example of semantic feature extraction, abnormal threshold features (such as temperature > 80℃ triggering "overheating risk"); pattern matching features (such as the difference between "atrial fibrillation waveform" and normal waveform of electrocardiogram).
[0049] The semantic feature of each modal data is processed as follows:
[0050] I. Multi-modal semantic alignment framework;
[0051] 1. Feature normalization and preprocessing;
[0052] Dimensional unification: Project the features of each modality onto a unified 512-dimensional vector space; use a fully connected layer + BatchNorm to achieve non-linear mapping, preserving modality specificity while ensuring spatial compatibility;
[0053] Timestamp alignment: Based on microsecond-level timestamps synchronized by NTP, a sliding time window (±50ms) is constructed; for asynchronously arriving features, nearest neighbor interpolation or bidirectional linear interpolation is used for time calibration;
[0054] 2. Shared semantic space mapping;
[0055] Cross-modal encoder: A shared encoder is built using the Transformer architecture;
[0056] Contrastive learning optimization: Constructing a triplet loss function: Anchor: Feature vector of any modality; Positive sample: Feature vector of other modalities of the same event; Negative sample: Random feature vector of different events; Using the InfoNCE loss function to maximize the mutual information of positive sample pairs and minimize the mutual information of negative sample pairs;
[0057] Knowledge-guided mapping: Pre-training stage: Self-supervised learning is performed using large-scale multimodal data (such as LAION-5B); Fine-tuning stage: Domain knowledge graphs (such as industrial equipment graphs) are introduced to constrain the mapping space;
[0058] 3. Dynamic association and weight allocation;
[0059] Intermodal attention mechanism: Calculate the intermodal attention matrix to reflect the importance of each modality to the current task; use a multi-head attention mechanism to capture multi-dimensional semantic associations; implement a gating mechanism to dynamically suppress the influence of noisy modalities;
[0060] Temporal attention mechanism: For temporal data (such as sensor streams, audio streams), a causal attention mask is used; a time decay factor is designed to reduce the weight of historical information (such as...). );
[0061] Context-aware weight allocation: Automatically adjust modality weights based on the current task type (such as fault diagnosis and sentiment analysis); achieve uncertainty-aware weight allocation to reduce the impact of low-confidence modalities;
[0062] 4. Alignment verification; consistency check: calculate the cosine similarity between modalities, and alignments below a threshold (e.g., 0.3) are considered invalid;
[0063] II. Alignment quality assessment indicator setting and alignment implementation;
[0064] Cross-modal retrieval accuracy: taking one modality as query, retrieving relevant samples in other modalities; calculating top-k accuracy (such as R@1, R@5, R@10);
[0065] Mutual information estimation: estimating mutual information between modalities using MINE (Mutual Information Neural Estimator); the higher the mutual information, the better the alignment quality;
[0066] Consistency score: calculating the consistency of cross-modal prediction results (such as the consistency of text classification and image classification results); using Cohen's kappa coefficient to evaluate the consistency;
[0067] When cross-modal retrieval accuracy, mutual information estimation, and consistency score all meet the set standards, the semantic alignment of semantic features of each modality data is completed;
[0068] Possible reasoning chain generation module, generating multiple possible reasoning chains according to the semantic association representation of cross-modal fusion.
[0069] The process of generating multiple possible reasoning chains is as follows:
[0070] Step 1: Input the cross-modal semantic association representation into the generative engine, which analyzes the core association between features through a cross-attention mechanism (the building scheme of the generative engine is 1. Cross-modal association encoder (analyzing core association); input layer processing: shared semantic vector: 512-dimensional shared vector of text, image, sensor, etc. modalities as basic input; association matrix embedding: convert the association strength (0-1) between modalities into an "association weight matrix" as an attention bias (e.g. "vibration-temperature" association degree 0.8, the attention weight of the corresponding position is initialized to 0.8); time chain encoding: for time sequence semantic chain (such as the sequence of vibration changing over time), extract time-dependent features (such as the trend of "vibration 120Hz→180Hz") through causal convolution (CausalConv); cross-attention mechanism design: multi-head association attention: set 8 attention heads to capture different dimensions of association (e.g. head 1 focuses on "vibration-temperature", head 2 focuses on "vibration-text description"); gating association fusion: filter noise association (e.g. features with an association degree <0.3 are suppressed) for the output of each attention head using a sigmoid gating mechanism; output: generate an "association feature map" (128-dimensional vector + association type label, such as "strong association-causal type", "moderate association-accompanying type"); 2. Knowledge-enhanced decoder; knowledge embedding: domain knowledge graph projection: convert "entity-relation" in the knowledge graph (e.g. "bearing wear → reason → insufficient lubrication") into a knowledge vector (128-dimensional), and concatenate it with the association feature map; rule constraint injection: introduce domain rules (e.g. "vibration frequency > 150Hz may be a mechanical failure") through prompts as the decoder's generation prefix; diversified decoder; basic architecture: use an autoregressive Transformer decoder (6 layers, 8 heads per layer) to output natural language reasoning chains; diversity control unit: dynamic temperature adjustment: adjust the temperature coefficient according to the fuzziness of the association features (e.g. entropy value of the association matrix) (high entropy value → temperature 0.8, encouraging diversity; low entropy value → temperature 0.5, ensuring accuracy); path memory mechanism: maintain a vector library of "generated reasoning chains", and the cosine similarity between the newly generated chain and the library vectors must be <0.7 (to ensure difference); beam search + random sampling hybrid strategy: first generate candidate chains through beam search (beam width = 5), then randomly select 3-5 most different chains from them as output. 3. Reasoning chain evaluator; association consistency: the matching degree of the feature association mentioned in the reasoning chain (e.g. "vibration → insufficient lubrication") with the input association matrix (>70% is qualified); knowledge matching degree: the matching degree of the reasoning chain with the "entity-relation" in the domain knowledge graph (e.g. whether "bearing wear" is associated with "insufficient lubrication", >80% is qualified); semantic fluency: calculate the perplexity of the reasoning chain through a pre-trained language model (Perplexity <50 is qualified).
[0071] Step 2: Set diverse decoding parameters to guide multi-path exploration; e.g., Temperature: set to 0.8 (higher than the default 0.5) to increase output randomness and encourage the model to explore non-optimal but reasonable reasoning paths; Top-K sampling: select K=5, i.e., randomly select from the top 5 candidate words at each generation step to avoid a single path; Beam search variants: set beam width=3 (generate 3 reasoning chains) and introduce a diversity penalty factor to ensure that the semantic difference of each chain is >30% (calculated by cosine similarity).
[0072] Step 3: Generate multiple possible reasoning chains based on semantic association.
[0073] Taking the "industrial equipment vibration anomaly" scenario as an example, the cross-modal semantic association representation includes: sensor "vibration 180Hz + temperature 85℃", image "bearing shell micro-deformation", text "not lubricated last week", and audio "high-frequency abnormal sound". Possible reasoning chain instances are as follows: Possible reasoning chain 1: vibration anomaly (180Hz) + temperature 85℃ + high-frequency abnormal sound → bearing wear (friction exacerbated due to lack of lubrication) → need to shut down and replace the bearing; Possible reasoning chain 2: vibration anomaly + bearing shell micro-deformation → loose installation bolts (causing bearing offset) → need to tighten the bolts and recalibrate; Possible reasoning chain 3: vibration anomaly + environmental temperature 35℃ (implicit feature) → insufficient equipment heat dissipation (motor overheating driving vibration) → need to clean the heat dissipation holes.
[0074] Risk decision analysis module: obtain the risk fingerprint value and reasoning overlap value of each possible reasoning chain, and then determine the decision response value of each possible reasoning chain.
[0075] Risk decision control module: determine the decision response mode of each possible reasoning chain according to the decision response value, and execute the decision response according to the decision response mode of each possible reasoning chain.
[0076] Determine the decision response mode of each possible reasoning chain according to the decision response value: set the decision response upper value and the decision response lower value (the decision response upper value is higher than the decision response lower value). When the decision response value of the possible reasoning chain ≥ the decision response upper value, trigger an emergency response (such as immediate shutdown and bearing replacement) for the possible reasoning chain. When the decision response value of the possible reasoning chain ≤ the decision response lower value, trigger a monitoring response (increase temperature patrol frequency) for the possible reasoning chain. When the decision response value of the possible reasoning chain is between the decision response upper value and the decision response lower value, trigger a warning response (such as running at reduced load, preparing spare parts) for the possible reasoning chain.
[0077] Determination process of the decision response value of the possible reasoning chain: obtain the risk fingerprint value and reasoning overlap value of a possible reasoning chain, and determine the decision response value of the possible reasoning chain by calculate the decision response value of the possible reasoning chain , wherein P1 is a risk fingerprint coefficient, P2 is an inference overlap coefficient, when it is considered that the risk fingerprint value is more important in decision-making, the value of P1 is 0.7, and the value of P2 is 0.3.
[0078] The acquisition process of the risk fingerprint value of the possible reasoning chain is: selecting a possible reasoning chain, and acquiring the knowledge conflict risk value of the possible reasoning chain , a historical risk density value , by calculating the risk fingerprint value of the possible reasoning chain , wherein B1 is a knowledge conflict risk coefficient, and B2 is a historical risk density coefficient, when an expert considers that the knowledge conflict risk is more critical in the current reasoning chain risk assessment, the value of the knowledge conflict risk coefficient can be 0.6, and the value of the historical risk density coefficient can be 0.4.
[0079] The acquisition process of the inference overlap value of the possible reasoning chain is: selecting a possible reasoning chain, marking the remaining (unselected) possible reasoning chains as comparison reasoning chains, comparing each comparison reasoning chain with the possible reasoning chain, further acquiring the comprehensive overlap value of each comparison reasoning chain, and performing sum-mean value calculation on the comprehensive overlap values of each comparison reasoning chain to calculate the inference overlap value of the possible reasoning chain .
[0080] The acquisition process of the comprehensive overlap value of the comparison reasoning chain is: selecting a comparison reasoning chain, acquiring the feature overlap value and the conclusion overlap value of the comparison reasoning chain and the possible reasoning chain, performing sum value calculation on the feature overlap value and the conclusion overlap value to calculate the comprehensive overlap value of the comparison reasoning chain.
[0081] The acquisition method of the feature overlap value of the comparison reasoning chain and the possible reasoning chain is: acquiring the intersection quantity and the union quantity of the comparison reasoning chain and the possible reasoning chain, performing ratio calculation on the intersection quantity and the union quantity to calculate the feature overlap value.
[0082] The acquisition method of the conclusion overlap value of the comparison reasoning chain and the possible reasoning chain is: acquiring the fault type similarity and the solution similarity of the comparison reasoning chain and the possible reasoning chain, performing sum-mean value calculation on the fault type similarity and the solution similarity to calculate the conclusion overlap value.
[0083] Taking possible reasoning chain 1 and possible reasoning chain 2 as examples, the calculation processes of the feature overlap value and the conclusion overlap value of the two are as follows:
[0084] Feature set of possible reasoning chain 1: {vibration anomaly (180 Hz), temperature 85℃, high-frequency abnormal sound, no lubrication, friction aggravation, bearing wear, stop and replace bearing}; (7 features in total, covering original perception, intermediate cause, decision action);
[0085] Feature set of possible reasoning chain 2: {vibration anomaly, bearing shell micro-deformation, installation bolt loosening, bearing deviation, fastening bolt, recalibration}; (6 features in total, covering original perception, intermediate cause, decision action);
[0086] Intersection feature: only “vibration anomaly” (the “vibration anomaly (180 Hz)” in reasoning chain 1 and the “vibration anomaly” in reasoning chain 2 are the same core feature, and the difference in specific parameters does not affect the feature attribution) -> intersection number = 1.
[0087] Union feature (after removing duplicates from all features of the two chains): vibration anomaly, temperature 85℃, high-frequency abnormal sound, no lubrication, friction aggravation, bearing wear, stop and replace bearing, bearing shell micro-deformation, installation bolt loosening, bearing deviation, fastening bolt, recalibration -> union number = 12.
[0088] Feature overlap value = intersection number / union number = 1 / 12 ≈ 8.3%.
[0089] Core conclusion of possible reasoning chain 1: fault type: bearing wear (caused by no lubrication leading to friction aggravation); solution: stop and replace bearing;
[0090] Core conclusion of possible reasoning chain 2: fault type: installation bolt loosening (leading to bearing deviation and shell micro-deformation); solution: fasten bolt and recalibrate;
[0091] Fault type similarity: “bearing wear” in possible reasoning chain 1 and “bolt loosening” in possible reasoning chain 2 belong to completely different fault types (mechanical wear vs. connection loosening), with no semantic association -> similarity = 0%.
[0092] Solution similarity: “stop and replace bearing” in possible reasoning chain 1 and “fasten bolt and recalibrate” in possible reasoning chain 2 are completely different operations (replace components vs. tighten and adjust), with no shared actions -> similarity = 0%.
[0093] Conclusion overlap value = (fault type similarity + solution similarity) / 2 = 0%.
[0094] Knowledge conflict risk value Acquisition process: select a possible reasoning chain, extract key entities and relationships in the possible reasoning chain, calculate the total matching degree of “entity-relation” in the possible reasoning chain and the knowledge graph, and then determine the anchor coefficient, knowledge conflict risk value ;
[0095] Possible reasoning chain 1 knowledge conflict risk value Acquisition process: acquire the core "entity-relation" library of the industrial equipment knowledge graph and the key "entity-relation" of the possible reasoning chain 1. The core "entity-relation" library of the industrial equipment knowledge graph and the key "entity-relation" of the possible reasoning chain 1 are shown in Table 1 and Table 2, respectively.
[0096] Table 1. Core "entity-relation" library of industrial equipment knowledge graph
[0097]
[0098] Table 2. Key "entity-relation" of possible reasoning chain 1
[0099]
[0100] C1 and R1: complete match (vibration + high temperature + abnormal sound are all bearing wear characteristics) → matching degree 1.0; C2 and R2: complete match (unlubricated → friction aggravation is an explicit relationship in the knowledge graph) → matching degree 1.0; C3 and R2: complete match (friction aggravation → bearing wear belongs to the latter part of R2) → matching degree 1.0; C4 and R3: complete match (the solution to bearing wear is replacement, consistent with R3) → matching degree 1.0; C5 and R2: partial match (unlubricated as an upstream cause, consistent with the logical chain of R2, but the knowledge graph does not explicitly emphasize the direct association of "unlubricated → vibration anomaly", only implicit) → matching degree 0.6, and the matching degree of the remaining partial match is 0.6.
[0101] Total matching degree of possible reasoning chain 1 and "entity-relation" in the knowledge graph: 1.0+1.0+1.0+1.0+0.6=4.6 (only for the scenario in this embodiment, the full score of five corresponding relationships is 5, and the full score of n corresponding relationships is 5); anchoring coefficient: 4.6 / 5=0.92; knowledge conflict risk value (0-10 points, the lower the score, the smaller the knowledge conflict).
[0102] Historical risk density value Acquisition process: select a possible reasoning chain, count the proportion of the occurrence times of the same type of possible reasoning chain in history, calculate the loss intensity by the ratio of the average single loss and the maximum loss, and the historical risk density value = 10 × (occurrence times proportion × loss intensity).
[0103] Historical risk density value of possible reasoning chain 1 Acquisition process: same kind of reasoning chain definition: feature combination: vibration anomaly (180Hz) + temperature 85℃ + high frequency abnormal sound → bearing wear → need to stop and replace the bearing. In the historical cases, there are 20 cases with the characteristics of "vibration + temperature + abnormal sound" and diagnosed as "bearing wear" (accounting for 20% of all cases 100 times), the occurrence frequency ratio = 20 / 100 = 0.2, the average single loss:
[0104] The average loss of bearing wear cases is 120,000 yuan (from historical data), the maximum loss in bearing wear cases is 500,000 yuan (from historical data, which may be caused by production line chain shutdown), loss intensity = average single loss / maximum loss = 12 / 50 = 0.24, historical risk density value = 10 x (0.2 x 0.24) = 0.48 points (0-10 points, the higher the score, the higher the historical risk).
[0105] The above system comprehensively and accurately captures cross-modal semantic association through multi-modal semantic alignment processing, and then generates multiple possible reasoning chains with reliability and explainability, and combines knowledge conflict, historical risk and other dimensions, while calculating reasoning overlap value to evaluate the correlation between chains, realizing multi-dimensional and multi-angle analysis of potential risks. The risk fingerprint value and the reasoning overlap value are included in the quantitative calculation framework of the decision response value, and the decision response threshold is dynamically set, so that the system can flexibly trigger emergency, early warning or monitoring response according to different risk levels, realizing the refinement and differentiation of risk control. This mechanism not only ensures timely disposal in high-risk scenarios, but also avoids resource waste in low-risk scenarios, effectively balancing risk prevention and control and execution efficiency.
[0106] The above formulas are dimensionless to calculate their numerical values, and the preset parameters in the formulas are set by those skilled in the art according to actual conditions.
[0107] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A multimodal semantic alignment-based intelligent decision-making and risk management system, characterized in that, It includes a semantic alignment and association representation module, which is used to collect the semantic features of each modality data in real time, perform semantic alignment processing on the semantic features of each modality data, and obtain a cross-modal fusion semantic association representation; the types of modality data include text modality, image modality, audio modality and sensor modality; The possible inference chain generation module is used to generate multiple possible inference chains based on the semantic association representation fused across modalities; The risk decision analysis module is used to determine the decision response value for each possible reasoning chain; The process for determining the decision response value of a possible inference chain includes: obtaining the risk fingerprint value and inference overlap value of a possible inference chain, and calculating the decision response value of the possible inference chain based on the risk fingerprint value and inference overlap value; The process of obtaining the inference overlap value of possible inference chains includes: selecting one possible inference chain, marking the remaining possible inference chains as comparison inference chains, comparing each comparison inference chain with the possible inference chain, obtaining the comprehensive overlap value of each comparison inference chain, summing and averaging the comprehensive overlap values of each comparison inference chain, and calculating the inference overlap value of the possible inference chain. The process of obtaining the comprehensive overlap value of the comparison inference chain includes: selecting a comparison inference chain, obtaining the feature overlap value and conclusion overlap value between the comparison inference chain and the possible inference chain, summing the feature overlap value and the conclusion overlap value, and calculating the comprehensive overlap value of the comparison inference chain. The methods for obtaining the feature overlap value between the comparison inference chain and the possible inference chain include: obtaining the number of intersections and the number of unions between the comparison inference chain and the possible inference chain, calculating the ratio between the number of intersections and the number of unions, and calculating the feature overlap value. The methods for obtaining the conclusion overlap value between the comparison inference chain and the possible inference chain include: obtaining the fault type similarity and solution similarity between the comparison inference chain and the possible inference chain, summing the fault type similarity and solution similarity and calculating the mean to obtain the conclusion overlap value; The process of obtaining the risk fingerprint value of a possible inference chain includes: selecting a possible inference chain, obtaining the knowledge conflict risk value and historical risk density value of the possible inference chain, and calculating the risk fingerprint value of the possible inference chain based on the knowledge conflict risk value and historical risk density value. The process of obtaining the knowledge conflict risk value includes: selecting a possible reasoning chain, extracting the key entities and relationships in the possible reasoning chain, calculating the total matching degree between the possible reasoning chain and the "entity-relationship" in the knowledge graph, and calculating the knowledge conflict risk value. The process of obtaining historical risk density values includes: selecting a possible inference chain, statistically analyzing the proportion of occurrences of the same type of possible inference chain in history, calculating the loss intensity by the ratio of average single loss to maximum loss, and calculating the historical risk density value based on the proportion of occurrences and the loss intensity. The risk decision control module is used to determine the decision response method for each possible inference chain based on the decision response value, and to execute the decision response according to the decision response method for each possible inference chain.
2. The intelligent decision-making and risk management system with multimodal semantic alignment according to claim 1, characterized in that, The process of generating multiple possible inference chains is as follows: Step 1: Input the cross-modal semantic association representation into the generative engine. The generative engine parses the core association between features through a cross-attention mechanism; Step 2: Set diverse decoding parameters to guide multi-path exploration; Step 3: Generate multiple possible inference chains based on semantic association.
3. The intelligent decision-making and risk management system with multimodal semantic alignment according to claim 1, characterized in that, The decision response method for each possible inference chain is determined based on the decision response value, including: setting an upper and lower decision response value; triggering an emergency response for the possible inference chain when the decision response value of the possible inference chain is greater than or equal to the upper decision response value; triggering a monitoring response for the possible inference chain when the decision response value of the possible inference chain is less than or equal to the lower decision response value; and triggering an early warning response for the possible inference chain when the decision response value of the possible inference chain is between the upper and lower decision response values.
Citation Information
Patent Citations
Cross-modal knowledge reasoning method and device for industrial quality inspection and medium
CN120069096A