Intelligent decision-making and risk management and control system based on multi-modal semantic alignment

The intelligent decision-making and risk management system with multimodal semantic alignment solves the problems of insufficient multimodal data fusion and dynamic response in traditional systems, realizes multi-dimensional risk analysis and dynamic response, and improves the accuracy of risk assessment and resource utilization efficiency.

CN120975241AActive Publication Date: 2025-11-18HEBEI DENGPU INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511127198.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-18
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Traditional semantic decision-making systems struggle to effectively integrate multimodal data and lack multidimensional risk assessment and dynamic response mechanisms, leading to inaccurate risk judgments and wasted resources.

Method used

The intelligent decision-making and risk management system adopts multimodal semantic alignment. Through the semantic alignment association representation module, the possible reasoning chain generation module, and the risk decision analysis module, it generates multiple highly interpretable reasoning chains and combines knowledge conflicts and historical risks to make dynamic decision responses.

Benefits of technology

It achieves comprehensive capture of cross-modal semantic associations, generates reliable multiple inference chains, enables multi-dimensional risk analysis and dynamic response, and improves the accuracy of risk assessment and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975241A_ABST
    Figure CN120975241A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of semantic decision management and control, in particular to a multi-modal semantic alignment intelligent decision and risk management and control system, which comprehensively and accurately captures cross-modal semantic association through multi-modal semantic alignment processing so as to generate a plurality of possible reasoning chains with reliability and interpretability. And combining dimensions such as knowledge conflicts and historical risks, calculating reasoning overlapping values to evaluate inter-chain association, realizing multi-dimensional and multi-angle analysis of potential risks, bringing risk fingerprint values and the reasoning overlapping values into a quantitative calculation framework of decision response values, dynamically setting a decision response threshold value, and realizing quantitative calculation of the risk fingerprint values and the reasoning overlapping values. The system can flexibly trigger emergency, early warning or monitoring response according to different risk levels, and refinement and differentiation of risk management and control are realized. The mechanism not only ensures timely disposal in a high-risk scene, but also avoids resource waste in a low-risk scene, and effectively balances risk prevention and control and execution efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semantic decision-making and control technology, and more specifically, to an intelligent decision-making and risk management system based on multimodal semantic alignment. Background Technology

[0002] In the field of semantic decision-making and control technology, traditional systems have significant shortcomings in their ability to process multimodal data. They are often limited to the analysis of single-modal information and struggle to effectively integrate different types of data such as text, images, audio, and sensor data. This results in the fragmentation of deep semantic relationships across modalities, making it impossible to fully capture complete information in complex scenarios. Consequently, this affects the accurate judgment of the overall situation and creates potential risks for subsequent decision-making.

[0003] The existing system also has shortcomings in the reasoning and risk assessment stages. On the one hand, the generated reasoning paths are relatively simple, lacking coverage of multiple potential possibilities, and the interpretability of the reasoning process is weak, making it difficult to trace the logic behind the conclusions. On the other hand, the risk assessment dimensions are limited, relying heavily on a few indicators for judgment, failing to combine multi-dimensional characteristics such as knowledge conflicts and historical risks, and also lacking an effective assessment of the relationships between different reasoning paths, resulting in an incomplete understanding of potential risks.

[0004] Traditional risk management mechanisms are relatively rigid in their response, often employing fixed threshold standards and uniform handling strategies. This model is ill-suited to dynamically changing scenarios. In high-risk situations, delayed responses may lead to losses, while in low-risk scenarios, over-handling can result in resource waste. It fails to achieve a dynamic balance between risk control and execution efficiency, and is ill-suited to the practical needs of refined and differentiated management.

[0005] Based on the above, this invention proposes an intelligent decision-making and risk management system with multimodal semantic alignment. Summary of the Invention

[0006] In view of the shortcomings of existing technologies, the purpose of this invention is to provide an intelligent decision-making and risk management system with multimodal semantic alignment.

[0007] To achieve the above objectives, the present invention provides the following technical solution: A multimodal semantic alignment-based intelligent decision-making and risk management system includes a semantic alignment and association representation module, which collects the semantic features of data from each modality in real time, performs semantic alignment processing on the semantic features of data from each modality, and obtains a cross-modal fusion semantic association representation; The possible inference chain generation module generates multiple possible inference chains based on the semantic association representation fused across modalities; The risk decision analysis module determines the decision response value for each possible inference chain. The process for determining the decision response value of a possible inference chain is as follows: obtain the risk fingerprint value and inference overlap value of a possible inference chain, and calculate the decision response value of the possible inference chain based on the risk fingerprint value and inference overlap value. The process for obtaining the risk fingerprint value of a possible inference chain is as follows: Select a possible inference chain, obtain the knowledge conflict risk value and historical risk density value of the possible inference chain, and calculate the risk fingerprint value of the possible inference chain based on the knowledge conflict risk value and historical risk density value. The risk decision control module determines the decision response method for each possible inference chain based on the decision response value, and executes the decision response according to the decision response method for each possible inference chain.

[0008] Furthermore, the process of generating multiple possible inference chains is as follows: Step 1: Input the cross-modal semantic association representation into the generative engine, and the generative engine parses the core association between features through a cross-attention mechanism; Step 2: Set diverse decoding parameters to guide multi-path exploration; Step 3: Generate multiple possible inference chains based on semantic association.

[0009] Furthermore, the process for obtaining the knowledge conflict risk value is as follows: Select a possible reasoning chain, extract the key entities and relationships in the possible reasoning chain, calculate the total matching degree between the possible reasoning chain and the "entity-relationship" in the knowledge graph, and then determine the anchoring coefficient. Based on the anchoring coefficient, the knowledge conflict risk value is calculated.

[0010] Furthermore, the process for obtaining historical risk density values ​​is as follows: Select a possible inference chain, count the percentage of occurrences of similar possible inference chains in history, calculate the loss intensity by the ratio of average single loss to maximum loss, and calculate the historical risk density value based on the percentage of occurrences and the loss intensity.

[0011] Further, the process for obtaining the inference overlap value of possible inference chains is as follows: Select one possible inference chain, mark the remaining possible inference chains as comparison inference chains, compare each comparison inference chain with the possible inference chain, further obtain the comprehensive overlap value of each comparison inference chain, sum and average the comprehensive overlap values ​​of each comparison inference chain, and calculate the inference overlap value of the possible inference chain.

[0012] Furthermore, the process for obtaining the comprehensive overlap value of the comparison inference chain is as follows: Select a comparison inference chain, obtain the feature overlap value and conclusion overlap value between the comparison inference chain and the possible inference chain, sum the feature overlap value and conclusion overlap value to calculate the comprehensive overlap value of the comparison inference chain.

[0013] Furthermore, the method for obtaining the feature overlap value between the comparison inference chain and the possible inference chain is as follows: obtain the number of intersections and the number of unions between the comparison inference chain and the possible inference chain, calculate the ratio between the number of intersections and the number of unions, and calculate the feature overlap value.

[0014] Furthermore, the method for obtaining the conclusion overlap value between the comparison inference chain and the possible inference chain is as follows: obtain the fault type similarity and solution similarity between the comparison inference chain and the possible inference chain, sum the fault type similarity and solution similarity and calculate the mean to obtain the conclusion overlap value.

[0015] Furthermore, the decision response method for each possible inference chain is determined based on the decision response value: an upper decision response value and a lower decision response value are set. When the decision response value of a possible inference chain is greater than or equal to the upper decision response value, an emergency response is triggered for the possible inference chain. When the decision response value of a possible inference chain is less than or equal to the lower decision response value, a monitoring response is triggered for the possible inference chain. When the decision response value of a possible inference chain is between the upper and lower decision response values, an early warning response is triggered for the possible inference chain.

[0016] Compared with the prior art, the present invention has the following beneficial effects: The system of this invention comprehensively and accurately captures cross-modal semantic relationships through multimodal semantic alignment processing, thereby generating multiple possible reasoning chains with reliability and interpretability. It also incorporates dimensions such as knowledge conflict and historical risk, and calculates reasoning overlap values ​​to assess inter-chain relationships. This enables multi-dimensional and multi-angle analysis of potential risks. By integrating risk fingerprint values ​​and reasoning overlap values ​​into the quantitative calculation framework of decision response values, and dynamically setting decision response thresholds, the system can flexibly trigger emergency, early warning, or monitoring responses based on different risk levels, achieving refined and differentiated risk management. This mechanism ensures timely handling in high-risk scenarios while avoiding resource waste in low-risk scenarios, effectively balancing risk prevention and execution efficiency. Attached Figure Description

[0017] Figure 1 A schematic diagram of the principle of an intelligent decision-making and risk management system with multimodal semantic alignment; Figure 2 A flowchart for obtaining the decision response value of a possible reasoning chain; Figure 3 This is a flowchart illustrating the process of obtaining risk fingerprint values ​​for possible inference chains. Detailed Implementation

[0018] Reference Figures 1 to 3 A multimodal semantic alignment-based intelligent decision-making and risk management system, comprising a semantic alignment association representation module, a possible reasoning chain generation module, a risk decision analysis module, and a risk decision management module.

[0019] The semantic alignment and association representation module collects the semantic features of each modality in real time, performs semantic alignment processing on the semantic features of each modality, and obtains a cross-modal fusion semantic association representation.

[0020] Modal data types include, but are not limited to, text modalities, image modalities, audio modalities, and sensor modalities.

[0021] The semantic feature acquisition process for text modality is as follows: Step 1: Text modal data acquisition from the data source; the following is an example of a data source: real-time log stream (such as system logs, transaction notes). Step 2, lightweight preprocessing; The following is an example of lightweight preprocessing: real-time word segmentation and stop word removal (pre-set high-frequency meaningless word list) to avoid redundant calculations; for long texts (such as real-time news feeds), sliding window truncation (e.g., in units of 512 characters) is used to balance semantic integrity and processing speed. Step 3, Efficient Semantic Encoding; The following is an example of efficient semantic encoding, using a distilled version of the pre-trained model (such as DistilBERT, TinyBERT, ALBERT), the number of parameters is reduced to 1 / 3 to 1 / 10 of the original model, the inference speed is improved by 2 to 3 times, while retaining the core semantic extraction capabilities (such as word vectors and sentence vectors); for streaming text (such as real-time dialogue), incremental encoding is used: based on the encoding results of the preceding text, only the features of the newly added segments are updated to avoid repeated calculations (similar to the caching mechanism of Transformer-XL).

[0022] Step 4, semantic feature extraction; the following is an example of semantic features, extracting sentence-level vectors (for topic classification, such as "complaint" and "consultation"), keyword weights (such as TF-IDF real-time calculation to identify risk words such as "default" and "overdue"), and sentiment polarity (such as using the lightweight model VADER to output positive and negative sentiment scores in real time).

[0023] The semantic feature acquisition process for image modalities is as follows: Step 1: Acquisition of image modal data from the data source; the following is an example of a data source: real-time camera stream (such as surveillance footage, industrial quality inspection footage). Step 2, Dynamic Preprocessing; The following is an example of dynamic preprocessing: real-time resizing (e.g., fixed at 224×224 pixels), normalization (subtracting the mean / dividing the standard deviation) to avoid redundant calculations for high-resolution images; frame interval sampling (e.g., processing 1 frame every 3 frames) is used for static scenes (e.g., fixed monitoring), while high frame rates (15-30 frames / second) are retained for dynamic scenes (e.g., traffic monitoring) to balance efficiency and semantic integrity; Step 3, Lightweight Visual Encoding; The following are examples of lightweight visual encoding, using mobile-optimized models: MobileNetV3 (3.2M parameters), EfficientNet-Lite (inference speed 4 times faster than ResNet), and MobileViT (a lightweight design combining CNN and Transformer) to extract high-level semantic features (such as "pedestrians", "vehicles", and "defect areas"); In industrial scenarios, knowledge distillation is used to customize models: for example, the quality inspection knowledge of ResNet50 is distilled into ShuffleNet, retaining the recognition ability of specific semantics such as "weld joint defects" and "scratches", and the inference latency is reduced to <20ms.

[0024] Step 4, semantic feature extraction; the following are examples of semantic features: lightweight object detection models (such as YOLOv8-nano, PP-YOLOE-tiny) output object categories (such as "safety helmet", "open flame") and location coordinates in real time (for spatial semantic association); scene classification features (such as using MobileNetV3 to output scene vectors such as "workshop" and "office") to assist in multimodal environment understanding.

[0025] The semantic feature acquisition process for audio modalities is as follows: Step 1: Acquisition of audio modal data from the data source; the following is an example of a data source: ambient sound sensor (such as industrial noise, abnormal sounds). Step 2, Real-time Preprocessing; The following is an example of real-time preprocessing: frame-by-frame windowing (e.g., 20ms per frame, 50% overlap), noise reduction (using spectral subtraction to remove steady-state noise), speech activity detection (VAD, such as WebRTC-VAD), processing only segments containing valid sound (filtering silence / noise); converting to Mel-Spectrogram, preserving frequency features sensitive to the human ear, and compressing the dimension to 80×100 (time×frequency). Step 3, Efficient Audio Coding; The following are examples of efficient audio coding: Speech Semantics: Use a lightweight version of Wav2Vec2.0 (such as Hubert-base-distilled) to extract speech content features (such as "stopped running", "overheating"), or use Tiny-WavLM to output speaker emotions (such as "anger", "anxiety"); Ambient Sound Semantics: Use YAMNet (3.7M parameters) to identify event types in real time (such as "glass breakage", "alarm sound", "abnormal equipment noise"), and output the probability distribution of 521 event categories; Step 4, semantic feature extraction; the following are examples of semantic features: speech-to-text (ASR) outputs text features simultaneously (combined with text modality processing), while retaining prosodic features (such as speech rate and pauses) for sentiment analysis; voiceprint features (such as abnormal fluctuations in equipment vibration frequency) are extracted in industrial scenarios, and fault semantics such as "bearing wear" are identified by the peak changes in the Mel spectrum.

[0026] The semantic feature acquisition process for sensor modalities is as follows: Step 1: Acquisition of sensor modal data from the data source; the following is an example of a data source: industrial sensors (vibration, temperature, pressure). Step 2, real-time signal preprocessing; the following is an example of real-time signal preprocessing: filtering (e.g., using a Butterworth filter to remove high-frequency noise), normalization (eliminating range differences between different sensors); downsampling: downsampling high-frequency signals (e.g., vibration signals sampled at 1kHz) to 200Hz, retaining key frequency components (e.g., the resonant frequency of the equipment). Step 3, Temporal Semantic Encoding; The following are examples of temporal semantic encoding: Frequency Domain Features: Use Fast Fourier Transform (FFT) to extract spectral peaks in real time (such as the abnormal frequency of motor vibration), or use wavelet transform to extract time-frequency features (such as the QRS complex of an electrocardiogram); Temporal Model: Use lightweight LSTM (such as LSTM-Lite) or Temporal Convolutional Network (TCN) to extract trend semantics (such as "temperature continues to rise" or "pressure drops sharply"). Step 4, semantic feature extraction; the following are examples of semantic features, such as abnormal threshold features (e.g., temperature > 80℃ triggers "overheating risk"); pattern matching features (e.g., the difference between "atrial fibrillation waveform" and normal waveform in an electrocardiogram).

[0027] The following scheme is used to perform semantic alignment processing on the semantic features of each modality of data: I. Multimodal semantic alignment framework; 1. Feature normalization and preprocessing; Dimensional unification: Project the features of each modality onto a unified 512-dimensional vector space; use a fully connected layer + BatchNorm to achieve non-linear mapping, preserving modality specificity while ensuring spatial compatibility; Timestamp alignment: Based on microsecond-level timestamps synchronized by NTP, a sliding time window (±50ms) is constructed; for asynchronously arriving features, nearest neighbor interpolation or bidirectional linear interpolation is used for time calibration; 2. Shared semantic space mapping; Cross-modal encoder: A shared encoder is built using the Transformer architecture; Contrastive learning optimization: Constructing a triplet loss function: Anchor: Feature vector of any modality; Positive sample: Feature vector of other modalities of the same event; Negative sample: Random feature vector of different events; Using the InfoNCE loss function to maximize the mutual information of positive sample pairs and minimize the mutual information of negative sample pairs; Knowledge-guided mapping: Pre-training stage: Self-supervised learning is performed using large-scale multimodal data (such as LAION-5B); Fine-tuning stage: Domain knowledge graphs (such as industrial equipment graphs) are introduced to constrain the mapping space; 3. Dynamic association and weight allocation; Intermodal attention mechanism: Calculate the intermodal attention matrix to reflect the importance of each modality to the current task; use a multi-head attention mechanism to capture multi-dimensional semantic associations; implement a gating mechanism to dynamically suppress the influence of noisy modalities; Temporal attention mechanism: For temporal data (such as sensor streams, audio streams), a causal attention mask is used; a time decay factor is designed to reduce the weight of historical information (such as...). ); Context-aware weight allocation: Automatically adjust modality weights based on the current task type (such as fault diagnosis and sentiment analysis); achieve uncertainty-aware weight allocation to reduce the impact of low-confidence modalities; 4. Alignment verification; consistency check: calculate the cosine similarity between modalities, and alignments below a threshold (e.g., 0.3) are considered invalid; II. Alignment quality assessment indicator setting and alignment implementation; Cross-modal retrieval accuracy: Using one modality as the query, retrieve relevant samples in other modalities; calculate top-k accuracy (e.g., R@1, R@5, R@10). Mutual information estimation: The mutual information between modes is estimated using MINE (Mutual Information Neural Estimator); the higher the mutual information, the better the alignment quality. Consistency score: Calculate the consistency of prediction results across modalities (e.g., consistency between text classification and image classification results); evaluate consistency using the Cohen's skappa coefficient; When the cross-modal retrieval accuracy, mutual information estimation, and consistency score all meet the set criteria, semantic alignment of the semantic features of the data in each modality is completed. The possible inference chain generation module generates multiple possible inference chains based on the semantic association representation fused across modalities.

[0028] The process for generating multiple possible reasoning chains is as follows: Step 1: Input the cross-modal semantic association representation into the generative engine. The generative engine parses the core associations between features through a cross-attention mechanism (the generative engine construction scheme is as follows: 1. Cross-modal association encoder (parses core associations); Input layer processing: Shared semantic vector: Use 512-dimensional shared vectors from modalities such as text, image, and sensor as the basic input; Association matrix embedding: Transform the inter-modal association strength (0-1) into an "association weight matrix" as an attention bias (e.g., if the "vibration-temperature" association degree is 0.8, then the attention weight at the corresponding position is initialized to 0.8); Temporal chain encoding: Extract time-dependent features from the temporal semantic chain (e.g., the sequence of vibration changes over time) through causal convolution (CausalConv). Features (e.g., the trend of "vibration 120Hz→180Hz"); Cross-attention mechanism design: Multi-head association attention: Set up 8 attention heads, each capturing associations in different dimensions (e.g., head 1 focuses on "vibration-temperature", head 2 focuses on "vibration-text description"); Gated association fusion: For the output of each attention head, use sigmoid gating (Gating Mechanism) to filter noisy associations (e.g., features with association degree <0.3 are suppressed); Output: Generate "association feature map" (128-dimensional vector + association type label, such as "strong association-causal type", "medium association-adjoint type"); 2. Knowledge enhancement decoder; Knowledge embedding: Domain knowledge graph projection: Project the "entity-relationship" in the knowledge graph (e.g., "bearing wear → cause → insufficient lubrication") is transformed into a knowledge vector (128 dimensions) and concatenated with the associated feature map; rule constraint injection: domain rules (e.g., "vibration frequency > 150Hz may be a mechanical fault") are introduced through prompts as a prefix for the decoder's generation; diversified decoder; infrastructure: an autoregressive Transformer decoder (6 layers, 8 attention heads per layer) is used to output a natural language inference chain; diversity control unit: dynamic temperature adjustment: the temperature coefficient is adjusted according to the ambiguity of the associated features (e.g., the entropy value of the association matrix) (high entropy value → temperature 0.8, encouraging diversity; low entropy value → temperature 0.5, ensuring accuracy); path memory mechanism: maintains "generated inferences" The vector library for "chains" requires newly generated chains to have a cosine similarity of <0.7 with vectors in the library (ensuring differences); a hybrid strategy of beam search + random sampling: first, candidate chains are generated through beam search (beam width=5), and then 3-5 chains with the greatest differences are randomly selected as outputs. 3. Inference chain evaluator; Association consistency: the matching degree between the feature associations mentioned in the inference chain (e.g., "vibration → insufficient lubrication") and the input association matrix (>70% is acceptable); Knowledge matching degree: the matching degree between the inference chain and the "entity-relationship" of the domain knowledge graph (e.g., whether "bearing wear" is associated with "insufficient lubrication", >80% is acceptable); Semantic fluency: the perplexity of the inference chain is calculated through a pre-trained language model (Perplexity<50 is acceptable)). Step 2: Set diverse decoding parameters to guide multi-path exploration; for example, temperature coefficient: set to 0.8 (higher than the default 0.5) to increase output randomness and encourage the model to explore non-optimal but reasonable inference paths; Top-K sampling: select K=5, that is, randomly select from the 5 candidate words with the highest probability at each step to avoid a single path; bundle search variant: set bundle width=3 (generate 3 inference chains) and introduce a diversity penalty factor to ensure that the semantic difference of each chain is >30% (calculated by cosine similarity).

[0029] Step 3: Generate multiple possible reasoning chains based on semantic associations; Taking the scenario of "abnormal vibration of industrial equipment" as an example, the cross-modal semantic association representation includes: sensor "vibration 180Hz + temperature 85℃", image "bearing housing micro-deformation", text "no lubrication last week", and audio "high-frequency abnormal noise". The following are three possible inference chain examples that can be generated: Possible inference chain 1: abnormal vibration (180Hz) + temperature 85℃ + high-frequency abnormal noise → bearing wear (due to increased friction caused by lack of lubrication) → the bearing needs to be replaced after shutdown; Possible inference chain 2: abnormal vibration + micro-deformation of bearing housing → loose mounting bolts (causing bearing misalignment) → the bolts need to be tightened and recalibrated; Possible inference chain 3: abnormal vibration + ambient temperature 35℃ (implicit feature) → insufficient heat dissipation of equipment (motor overheating and vibration) → the heat dissipation holes need to be cleaned.

[0030] The risk decision analysis module obtains the risk fingerprint value and inference overlap value of each possible inference chain, and then determines the decision response value of each possible inference chain.

[0031] The risk decision control module determines the decision response method for each possible inference chain based on the decision response value, and executes the decision response according to the decision response method for each possible inference chain.

[0032] The decision response method for each possible inference chain is determined based on the decision response value: an upper decision response value and a lower decision response value are set (the upper decision response value is higher than the lower decision response value). When the decision response value of a possible inference chain is greater than or equal to the upper decision response value, an emergency response is triggered for the possible inference chain (such as immediately stopping the machine and replacing the bearing). When the decision response value of a possible inference chain is less than or equal to the lower decision response value, a monitoring response is triggered for the possible inference chain (increasing the frequency of temperature inspection). When the decision response value of a possible inference chain is between the upper and lower decision response values, an early warning response is triggered for the possible inference chain (such as reducing the load and preparing spare parts).

[0033] The process for determining the decision response value of a possible inference chain: Obtain the risk fingerprint value of a possible inference chain. and inference overlap value ,pass Calculate the decision response value of the possible inference chain. Where P1 is the risk fingerprint coefficient and P2 is the inference overlap coefficient. When the risk fingerprint value is considered to be more important in decision-making, the value of P1 is 0.7 and the value of P2 is 0.3.

[0034] The process for obtaining the risk fingerprint value of a possible inference chain is as follows: Select a possible inference chain and obtain the knowledge conflict risk value of that possible inference chain. Historical risk density value ,pass Calculate the risk fingerprint value of the possible inference chain. Where B1 is the knowledge conflict risk coefficient and B2 is the historical risk density coefficient. When experts believe that the knowledge conflict risk is more critical in the current reasoning chain risk assessment, the knowledge conflict risk coefficient can be 0.6 and the historical risk density coefficient can be 0.4.

[0035] The process for obtaining the inference overlap value of possible inference chains is as follows: Select one possible inference chain, mark the remaining (unselected) possible inference chains as comparison inference chains, compare each comparison inference chain with the possible inference chain, further obtain the comprehensive overlap value of each comparison inference chain, sum and average the comprehensive overlap values ​​of each comparison inference chain, and calculate the inference overlap value of that possible inference chain. .

[0036] The process for obtaining the comprehensive overlap value of the comparison inference chain is as follows: Select a comparison inference chain, obtain the feature overlap value and conclusion overlap value between the comparison inference chain and the possible inference chains, calculate the sum of the feature overlap value and the conclusion overlap value, and calculate the comprehensive overlap value of the comparison inference chain.

[0037] The method for obtaining the feature overlap value between the comparison inference chain and the possible inference chain is as follows: obtain the number of intersections and the number of unions between the comparison inference chain and the possible inference chain, calculate the ratio between the number of intersections and the number of unions, and calculate the feature overlap value.

[0038] The method for obtaining the conclusion overlap value between the comparison inference chain and the possible inference chain is as follows: obtain the fault type similarity and solution similarity between the comparison inference chain and the possible inference chain, sum and average the fault type similarity and solution similarity, and calculate the conclusion overlap value.

[0039] Taking possible reasoning chain 1 and possible reasoning chain 2 as examples, the calculation process of their feature overlap value and conclusion overlap value is as follows: The feature set of possible inference chain 1: {abnormal vibration (180Hz), temperature 85℃, high-frequency abnormal noise, lack of lubrication, increased friction, bearing wear, shutdown to replace bearing}; (a total of 7 features, covering original perception, intermediate causes, and decision-making actions); The feature set of possible inference chain 2: {abnormal vibration, slight deformation of bearing housing, loose mounting bolts, bearing misalignment, tightening bolts, recalibration}; (a total of 6 features, covering original perception, intermediate causes, and decision-making actions); Intersection features: Only "vibration anomaly" (the "vibration anomaly (180Hz)" in reasoning chain 1 and the "vibration anomaly" in reasoning chain 2 are the same core feature, and the difference in specific parameters does not affect the feature attribution) → number of intersections = 1.

[0040] Union features (after removing duplicate features from both chains): abnormal vibration, temperature 85℃, high-frequency abnormal noise, lack of lubrication, increased friction, bearing wear, stop to replace bearing, slight deformation of bearing housing, loose mounting bolts, bearing misalignment, tighten bolts, recalibrate → number of unions = 12.

[0041] Feature overlap value = number of intersections / number of unions = 1 / 12 ≈ 8.3%.

[0042] The core conclusion of possible reasoning chain 1: Fault type: Bearing wear (due to increased friction caused by lack of lubrication); Solution: Shut down and replace the bearing; Possible core conclusion of reasoning chain 2: Fault type: Loose mounting bolts (causing bearing misalignment and slight deformation of the housing); Solution: Tighten the bolts and recalibrate; Fault type similarity: "bearing wear" in possible reasoning chain 1 and "loose bolts" in possible reasoning chain 2 belong to completely different fault types (mechanical wear vs. loose connection), with no semantic association → similarity = 0%.

[0043] Solution similarity: The "stop and replace bearing" in possible reasoning chain 1 and the "tighten bolts and calibrate" in possible reasoning chain 2 are completely different operations (replacing parts vs. tightening and adjusting), with no shared actions → similarity = 0%.

[0044] Conclusion: Overlap value = (fault type similarity + solution similarity) / 2 = 0%.

[0045] Knowledge Conflict Risk Value The acquisition process involves: selecting a possible reasoning chain, extracting the key entities and relationships within that chain, calculating the overall matching degree between the possible reasoning chain and the "entity-relationship" in the knowledge graph, and then determining the anchoring coefficient and the knowledge conflict risk value. ; Knowledge conflict risk value of possible inference chain 1 Acquisition Process: Acquire the core "entity-relationship" library of the industrial equipment knowledge graph and the key "entity-relationship" of possible reasoning chain 1. The comparison table of the core "entity-relationship" library of the industrial equipment knowledge graph and the comparison table of the key "entity-relationship" of possible reasoning chain 1 are shown in Table 1 and Table 2, respectively. Table 1. Comparison of the core "entity-relationship" database for the industrial equipment knowledge graph.

[0046] Table 2. Key "Entity-Relationship" Comparison Table for Possible Reasoning Chain 1

[0047] C1 and R1: Complete match (vibration + high temperature + abnormal noise are all characteristics of bearing wear) → Match degree 1.0; C2 and R2: Complete match (unlubricated → increased friction is an explicit relationship in the knowledge graph) → Match degree 1.0; C3 and R2: Complete match (increased friction → bearing wear belongs to the latter half of the relationship in R2) → Match degree 1.0; C4 and R3: Complete match (the solution for bearing wear is replacement, consistent with R3) → Match degree 1.0; C5 and R2: Partial match (unlubricated as an upstream cause, conforms to the logical chain of R2, but the knowledge graph does not explicitly emphasize the direct association of "unlubricated → abnormal vibration", only implies it) → Match degree 0.6. Except for the complete match case, the match degree of the other partial match cases is 0.6.

[0048] The total matching degree between possible reasoning chain 1 and the "entity-relationship" in the knowledge graph is: 1.0 + 1.0 + 1.0 + 1.0 + 0.6 = 4.6 (only for the scenario in this implementation, five sets of corresponding relationships have a maximum score of 5 points, and n sets of corresponding relationships have a maximum score of 5 points); anchoring coefficient: 4.6 / 5 = 0.92; knowledge conflict risk value. (0-10 points, the lower the score, the less conflict with the knowledge).

[0049] Historical risk density value The acquisition process is as follows: Select a possible inference chain, count the percentage of occurrences of the same possible inference chain in history, and calculate the loss intensity by the ratio of average single loss to maximum loss. Historical risk density value = 10 × (percentage of occurrences × loss intensity).

[0050] Historical risk density value of possible inference chain 1 Acquisition Process: Definition of Similar Inference Chains: Feature Combination: Abnormal Vibration (180Hz) + Temperature 85℃ + High-Frequency Abnormal Noise → Bearing Wear → Requires Shutdown and Bearing Replacement. In historical cases, there were 20 cases (20% of all 100 cases) that simultaneously exhibited the characteristics of "vibration + temperature + abnormal noise" and were diagnosed as "bearing wear." The percentage of occurrences = 20 / 100 = 0.2, with an average loss per incident. The average loss from bearing wear cases is 120,000 yuan (from historical data), and the maximum loss from bearing wear cases is 500,000 yuan (from historical data, possibly due to a production line shutdown). The loss intensity = average single loss / maximum loss = 120,000 / 500,000 = 0.24. The historical risk density value = 10 × (0.2 × 0.24) = 0.48 points (0-10 points, the higher the score, the higher the historical risk).

[0051] The aforementioned system comprehensively and accurately captures cross-modal semantic relationships through multimodal semantic alignment processing, thereby generating multiple possible reasoning chains with reliability and interpretability. It also incorporates dimensions such as knowledge conflict and historical risk, and calculates reasoning overlap values ​​to assess inter-chain relationships. This enables multi-dimensional and multi-faceted analysis of potential risks. By integrating risk fingerprint values ​​and reasoning overlap values ​​into the quantitative calculation framework of decision response values, and dynamically setting decision response thresholds, the system can flexibly trigger emergency, early warning, or monitoring responses based on different risk levels, achieving refined and differentiated risk management. This mechanism ensures timely handling in high-risk scenarios while avoiding resource waste in low-risk scenarios, effectively balancing risk prevention and execution efficiency.

[0052] The above formulas are all dimensionless calculations, and the preset parameters in the formulas should be set by those skilled in the art according to the actual situation.

[0053] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A multimodal semantic alignment-based intelligent decision-making and risk management system, characterized in that, It includes a semantic alignment and association representation module, which is used to collect the semantic features of each modality data in real time, perform semantic alignment processing on the semantic features of each modality data, and obtain a cross-modal fusion semantic association representation; The possible inference chain generation module is used to generate multiple possible inference chains based on the semantic association representation fused across modalities; The risk decision analysis module is used to determine the decision response value for each possible reasoning chain; The process for determining the decision response value of a possible inference chain includes: obtaining the risk fingerprint value and inference overlap value of a possible inference chain, and calculating the decision response value of the possible inference chain based on the risk fingerprint value and inference overlap value; The process of obtaining the risk fingerprint value of a possible inference chain includes: selecting a possible inference chain, obtaining the knowledge conflict risk value and historical risk density value of the possible inference chain, and calculating the risk fingerprint value of the possible inference chain based on the knowledge conflict risk value and historical risk density value. The risk decision control module is used to determine the decision response method for each possible inference chain based on the decision response value, and to execute the decision response according to the decision response method for each possible inference chain.

2. The intelligent decision-making and risk management system with multimodal semantic alignment according to claim 1, characterized in that, The process of generating multiple possible inference chains is as follows: Step 1: Input the cross-modal semantic association representation into the generative engine. The generative engine parses the core association between features through a cross-attention mechanism; Step 2: Set diverse decoding parameters to guide multi-path exploration; Step 3: Generate multiple possible inference chains based on semantic association.

3. The intelligent decision-making and risk management system with multimodal semantic alignment according to claim 1, characterized in that, The process of obtaining the knowledge conflict risk value includes: selecting a possible reasoning chain, extracting the key entities and relationships in the possible reasoning chain, calculating the total matching degree between the possible reasoning chain and the "entity-relationship" in the knowledge graph, and then determining the anchoring coefficient, and calculating the knowledge conflict risk value based on the anchoring coefficient.

4. The intelligent decision-making and risk management system with multimodal semantic alignment according to claim 1, characterized in that, The process of obtaining historical risk density values ​​includes: selecting a possible inference chain, statistically analyzing the proportion of occurrences of similar possible inference chains in history, calculating the loss intensity by the ratio of average single loss to maximum loss, and calculating the historical risk density value based on the proportion of occurrences and the loss intensity.

5. The intelligent decision-making and risk management system with multimodal semantic alignment according to claim 1, characterized in that, The process for obtaining the inference overlap value of a possible inference chain includes: selecting a possible inference chain, marking the remaining possible inference chains as comparison inference chains, comparing each comparison inference chain with the possible inference chain, further obtaining the comprehensive overlap value of each comparison inference chain, summing and averaging the comprehensive overlap values ​​of each comparison inference chain, and calculating the inference overlap value of the possible inference chain.

6. The intelligent decision-making and risk management system with multimodal semantic alignment according to claim 5, characterized in that, The process of obtaining the comprehensive overlap value of the comparison inference chain includes: selecting a comparison inference chain, obtaining the feature overlap value and conclusion overlap value between the comparison inference chain and the possible inference chain, summing the feature overlap value and the conclusion overlap value, and calculating the comprehensive overlap value of the comparison inference chain.

7. The intelligent decision-making and risk management system with multimodal semantic alignment according to claim 6, characterized in that, The methods for obtaining the feature overlap value between the comparison inference chain and the possible inference chain include: obtaining the number of intersections and the number of unions between the comparison inference chain and the possible inference chain, calculating the ratio between the number of intersections and the number of unions, and calculating the feature overlap value.

8. The intelligent decision-making and risk management system with multimodal semantic alignment according to claim 6, characterized in that, The method for obtaining the conclusion overlap value between the comparison inference chain and the possible inference chain includes: obtaining the fault type similarity and solution similarity between the comparison inference chain and the possible inference chain, summing the fault type similarity and solution similarity and calculating the mean to obtain the conclusion overlap value.

9. The intelligent decision-making and risk management system with multimodal semantic alignment according to claim 1, characterized in that, The decision response method for each possible inference chain is determined based on the decision response value, including: setting an upper and lower decision response value; triggering an emergency response for the possible inference chain when the decision response value of the possible inference chain is greater than or equal to the upper decision response value; triggering a monitoring response for the possible inference chain when the decision response value of the possible inference chain is less than or equal to the lower decision response value; and triggering an early warning response for the possible inference chain when the decision response value of the possible inference chain is between the upper and lower decision response values.

Citation Information

Patent Citations

  • Operation and maintenance decision driving method and system based on multi-modal data knowledge graph

    CN119722037A

  • Cross-modal knowledge reasoning method and device for industrial quality inspection and medium

    CN120069096A

  • Intelligent security management system based on AI large model

    CN120145312A

  • Auditing decision support system and method based on dynamic knowledge graph

    CN120387671A

  • Water conservancy multi-modal intelligent decision-making method and system based on model association protocol

    CN120410258A

Cited By

  • Block chain-based cross-border transaction method and device, computer equipment and storage medium

    CN121599573A

  • Deep learning-fused exploration scene monitoring illegal behavior automatic identification method and system

    CN121659076A

  • Office automation cooperation system based on multi-mode artificial intelligence

    CN121766938A

  • Risk map construction method based on animal metaphor

    CN122198056A

  • Method for constructing a risk map based on animal metaphor

    CN122198056B