Industrial simulation optimization method for multi-modal feature fusion
By employing a multimodal feature fusion mechanism and a cross-modal attention weighting algorithm, the problems of data correlation and weight allocation in multimodal feature fusion are solved, thereby improving the accuracy and reliability of industrial simulation and making it suitable for simulation optimization in intelligent manufacturing and process industries.
Patent Information
- Application Number
- CN202511891409.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-03-20
AI Technical Summary
Existing industrial simulation technologies suffer from insufficient multi-source data correlation mining, rigid cross-modal feature weight allocation, and weak adaptability to dynamic scenarios, resulting in insufficient simulation accuracy and high feature redundancy, making it difficult to achieve stable improvement in simulation accuracy.
A multimodal feature fusion mechanism is adopted. Through the collaborative design of the attention query matrix and the cross-modal attention weight matrix, the parameters are optimized by combining gradient descent and reinforcement learning strategies to generate the cross-modal attention weight matrix. Then, feature correlation evaluation and weighted fusion are performed by GNN and spectral clustering to generate the fused feature vector, and finally the optimized configuration parameters are generated.
It achieves dynamic weight allocation and feature fusion of multi-source industrial data, improving simulation accuracy and reliability. It is suitable for simulation optimization tasks in the fields of intelligent manufacturing and process industries, ensuring the stability and adaptability of simulation results.
Smart Images

Figure CN121706571A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of industrial simulation, and in particular to a multi-modal feature fusion industrial simulation optimization method. BACKGROUND
[0002] In the technical field of industrial simulation, the existing solutions related to the multi-modal feature fusion industrial simulation optimization method generally achieve simulation parameter generation through single modal feature extraction or static weighted fusion, which has limitations such as insufficient multi-source data correlation mining, rigid cross-modal feature weight distribution, and weak dynamic scene adaptability. Existing methods mostly rely on preset rules or fixed thresholds to linearly superimpose multi-modal features, which may cause high feature redundancy and key information masking when matching the dynamic optimization needs of industrial simulation scenarios, making it difficult to achieve stable improvement of simulation accuracy. For the joint processing of standardized feature vectors in the multi-modal feature library, existing technologies generally lack the ability to dynamically adjust weights guided by cross-modal attention mechanisms, making it difficult to form a continuous process of feature input-weight optimization-fusion parameter generation-verification feedback in the industrial simulation scenario, resulting in insufficient simulation accuracy caused by insufficient multi-modal feature fusion. SUMMARY
[0003] The present application provides a multi-modal feature fusion industrial simulation optimization method to solve the problem of how to improve simulation accuracy in industrial simulation scenarios based on standardized feature vectors from a multi-modal feature library through a multi-modal feature fusion mechanism and a cross-modal attention weighting algorithm.
[0004] To solve the above technical problems, the present application provides a multi-modal feature fusion industrial simulation optimization method, comprising: Obtain multi-source industrial data, standardize it through data synchronization and timestamp alignment strategies, and use modal-specific algorithms such as time-frequency analysis, CNN, or NLP for feature extraction and vectorization operations to generate feature mapping relationships expressed in graph structures or matrix forms; Obtain modal features from the feature mapping relationships, normalize them using min-max normalization or mean-variance standardization methods, and use PCA, LDA, or auto-encoder algorithms for dimensionality reduction and model construction based on CNN, RNN, or Transformer architecture to generate a multi-modal feature library; Obtain feature inputs from the multi-modal feature library, construct an attention mechanism based on the Transformer query mechanism or the improved Shannon entropy model, and use a strategy combining gradient descent and reinforcement learning for parameter optimization to generate a cross-modal attention weight matrix; Based on the attention weight matrix and the multi-modal feature library, the correlation evaluation, screening and fusion operations of features are performed by using the correlation evaluation based on GNN and spectral clustering, the dynamic threshold screening based on Bayesian optimization and the fusion method based on weighting and MLP nonlinear transformation, to generate a fusion feature vector; Based on the fusion feature vector, the simulation parameter conversion is performed according to the simulation parameter mapping rule, the physical engine environment covering multiple physical fields is loaded, and the genetic algorithm or particle swarm optimization is executed, to generate an original simulation result. Based on the original simulation result, the time-space alignment processing such as timestamp correction and spatial coordinate unification is performed, the verification and evaluation such as time accuracy and spatial consistency are calculated, and the feature fusion parameters are adjusted by using the Bayesian optimization method, to generate an optimized configuration parameter.
[0005] Further, the process of generating the feature mapping relationship expressed in the form of a graph structure or a matrix further includes: The multi-modal industrial data including physical quantity data collected by sensors, device operation logs, image and video data and control instructions are obtained, and the standardized processing of missing value filling, abnormal value detection and elimination, noise filtering and format unification is performed, to obtain preprocessed data.
[0006] The modal features are extracted from the preprocessed data, and the vectorization processing specific to each mode such as time-frequency analysis, wavelet transform, CNN or NLP is performed, to generate an original feature vector.
[0007] The original feature vector is stored in a structured manner by constructing a multi-dimensional index structure and an association mapping mechanism, to generate a feature mapping relationship expressed in the form of a graph structure or a matrix.
[0008] Further, the multi-source industrial data includes: The multi-source industrial data refers to a set of original data collected through multiple heterogeneous channels in the industrial simulation optimization scenario, which has different physical meanings, data structures and space-time characteristics; Specifically, it includes physical quantity data collected by sensors, device operation logs, image and video data and control instructions multi-modal industrial data; The multi-modal industrial data covers structured data, semi-structured data and unstructured data, and involves time series signals, discrete event records and multi-dimensional spatial information.
[0009] Further, the process of generating the multi-modal feature library further includes: The modal features are obtained from the feature mapping relationship, and the normalization processing such as minimum-maximum normalization, mean-variance standardization or distribution transformation is performed, to generate standardized features.
[0010] The standardized features are processed by PCA, LDA, auto-encoder or nonlinear dimensionality reduction based on graph embedding and manifold learning, to generate a reduced feature sequence.
[0011] constructing a feature representation model based on CNN, RNN or Transformer based on the reduced feature sequence, and generating a multi-modal feature library.
[0012] Further, the process of generating the cross-modal attention weight matrix further comprises: obtaining a feature input from the multi-modal feature library, constructing an attention query matrix based on the query, key, value mechanism in the Transformer architecture or the improved Shannon entropy model, and generating an initial attention distribution.
[0013] Based on the initial attention distribution, the parameter optimization is performed based on the gradient descent, adaptive learning rate, regularization term or reinforcement learning strategy, and the dynamic weight parameter is generated.
[0014] The dynamic weight parameter is normalized by maximum-minimum normalization, L1 norm normalization or Softmax normalization, and the cross-modal attention weight matrix stored in a sparse matrix structure is generated.
[0015] Further, the process of generating the fusion feature vector further comprises: The attention weight matrix and the multi-modal feature library are combined, the feature correlation evaluation based on the graph neural network and the spectral clustering algorithm is performed, and the candidate feature set is generated.
[0016] The candidate feature set is combined, the dynamic screening threshold based on the mixed strategy of Bayesian optimization and genetic algorithm is set, and the effective feature subset is generated.
[0017] The effective feature subset is weighted and fused based on the cross-modal attention weight weighting and MLP nonlinear transformation, and the fusion feature vector is generated.
[0018] Further, the process of generating the original simulation result further comprises: The fusion feature vector is converted into a simulation parameter format according to the simulation parameter mapping rule, and the initial simulation configuration is generated.
[0019] Based on the initial simulation configuration, a physical engine covering dynamics, thermodynamics, fluid or control system simulation is loaded, and a simulation computing environment is generated.
[0020] An optimization algorithm including genetic algorithm, particle swarm optimization or gradient descent method is executed in the simulation computing environment, and an original simulation result is generated.
[0021] Further, the process of generating the optimized configuration parameter further comprises: The original simulation result is processed by time stamp correction, time window segmentation, spatial coordinate unification and data interpolation completion, and a standardized verification input is generated.
[0022] Based on the standardized verification input, the key parameters including time accuracy, spatial consistency and performance stability are calculated to generate a verification evaluation report.
[0023] According to the verification evaluation report, the feature fusion parameters including the weighting coefficient, the dynamic screening threshold, the fusion model hyperparameter or the attention weight adjustment factor are adjusted to generate the optimized configuration parameters.
[0024] Further, the step of generating the cross-modal attention weight matrix specifically comprises: Performing effective value screening to eliminate abnormal weights and noise interference, and adopting a statistical threshold and an anomaly detection algorithm to determine the screening boundary; According to the characteristics and fusion requirements of the industrial simulation task, dynamically selecting the maximum-minimum normalization, L1 norm normalization or Softmax normalization method for normalization processing; The normalized weight parameters are constructed into a cross-modal attention weight matrix stored in a sparse matrix structure, and consistency verification and positive definiteness detection are performed.
[0025] Further, the step of generating the fusion feature vector specifically comprises: Based on the weight distribution in the cross-modal attention weight matrix, each feature vector in the effective feature subset is linearly weighted; A multi-layer perceptron is used to perform nonlinear transformation on the weighted features to capture complex interaction relationships between the features; During the linear weighting and nonlinear transformation process, L2 regularization and sparse constraints are introduced to prevent overfitting; A multi-head fusion strategy is used to process different modal feature subsets through multiple independent weighting channels in parallel, and the fusion results are summarized and normalized; In the fusion process, batch normalization and residual connection technology are combined; Real-time monitoring of the statistical distribution and information entropy of the fusion features, and the adjustment mechanism is triggered by abnormal fluctuations; Supporting online updating and incremental fusion to respond to the dynamic changes of data and models in the industrial simulation environment.
[0026] The key innovations of the present application include: (1) Introducing a multi-modal feature fusion mechanism, through the collaborative design of attention query matrix and cross-modal attention weight matrix, realizing dynamic weight distribution and fusion of multi-source industrial data.
[0027] (2) Designing a cross-modal attention weighting algorithm, using initial attention distribution and dynamic weight parameters for feature correlation evaluation and weighted fusion processing to generate a fusion feature vector.
[0028] (3) Construct a dynamic feature selection module, set a dynamic screening threshold to generate an effective feature subset, and generate simulation parameters through weighted fusion to realize multi-dimensional verification of simulation results.
[0029] The following are its main benefits: (1) Through the multi-modal feature fusion mechanism, multi-source industrial data can be effectively integrated in the simulation scene, improving simulation accuracy and reliability, and being suitable for simulation optimization tasks in intelligent manufacturing and process industry fields.
[0030] (2) The implementation of the cross-modal attention weighting algorithm makes the feature correlation evaluation and weighted fusion processing more accurate, ensuring that the generated fusion feature vector can meet the dynamic needs of simulation calculation and adapt to complex industrial simulation scenarios.
[0031] (3) The construction of the dynamic feature selection module makes the generation of simulation parameters and the verification of simulation results form a continuous process, ensuring stable improvement of simulation accuracy and being suitable for joint processing of standardized feature vectors in the multi-modal feature library. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 A flowchart of a multi-modal feature fusion industrial simulation optimization method provided by the embodiments of the present application. DETAILED DESCRIPTION
[0033] Embodiment one: refer to Figure 1 is a flowchart of a multi-modal feature fusion industrial simulation optimization method provided by the embodiments of the present application. The flowchart can at least include steps S100-S600: S100, acquire multi-source industrial data, perform standardized processing through data synchronization and timestamp alignment strategy, and perform feature extraction and vectorization operation by using time-frequency analysis, CNN or NLP modal specific algorithm, to generate feature mapping relationship expressed in graph structure or matrix form; S200, acquire modal features from the feature mapping relationship, perform normalization processing by using minimum-maximum normalization or mean-variance standardization method, and perform dimension reduction and model construction processing based on CNN, RNN or Transformer architecture by using PCA, LDA or automatic encoder algorithm, to generate a multi-modal feature library; S300, acquire feature input from the multi-modal feature library, construct attention mechanism based on Transformer query mechanism or improved Shannon entropy model, and perform parameter optimization processing by using the strategy combining gradient descent and reinforcement learning, to generate a cross-modal attention weight matrix; S400, based on the attention weight matrix and the multi-modal feature library, performing feature correlation evaluation, screening and fusion operations by using correlation evaluation based on GNN and spectral clustering, dynamic threshold screening based on Bayesian optimization, and fusion method based on weighting and MLP nonlinear transformation, to generate a fusion feature vector; S500, based on the fusion feature vector, performing simulation parameter conversion according to the simulation parameter mapping rule, loading a physical engine environment covering multiple physical fields and executing an optimization algorithm such as genetic algorithm or particle swarm optimization, to generate an original simulation result; S600, based on the original simulation result, performing time stamp correction and space coordinate unification for space-time alignment processing, calculating time precision and space consistency for verification and evaluation, and adjusting feature fusion parameters by using Bayesian optimization method, to generate optimized configuration parameters.
[0034] Step S100 includes at least steps S110-S130: S110, obtaining multi-source industrial data, performing standardization processing to obtain pre-processed data; Specifically, the multi-source industrial data refers to a set of original data collected through multiple heterogeneous channels in the industrial simulation optimization scenario of the method, having different physical meanings, data structures and space-time characteristics, specifically including but not limited to physical quantity data collected by sensors, device operation logs, image and video data and control instructions, etc. Multi-modal industrial data covers structured data, semi-structured data and unstructured data, involving time series signals, discrete event records and multi-dimensional spatial information. Its specific composition is closely related to the processing logic of the patent full text, mainly including: Sensor measurement data: time series signals collected by physical sensors deployed in industrial equipment and environment, such as temperature, pressure, vibration, current, flow, etc. Analog or digital signal, is a structured data subject reflecting the state change of the physical world.
[0035] Device operation and state data: discrete event records and state snapshots from manufacturing execution system (MES), supervisory control and data acquisition (SCADA) system and device controller, such as device start-stop, fault code, alarm information, working mode switching, etc., often in the form of semi-structured logs or messages.
[0036] Visual perception data: image and video stream captured by industrial camera and video monitoring system, used to obtain multi-dimensional spatial information such as product surface quality, device behavior, personnel operation, environment condition, etc., which belongs to typical unstructured data.
[0037] Control instructions and parameter data: derived from the set value, control logic, process parameter formula, etc. issued by programmable logic controller (PLC), distributed control system (DCS) and other devices, which are the core structured data to drive the operation of the equipment.
[0038] These data sources have significant differences in modality (numeric / text / image), structure (structured / semi-structured / unstructured), timing (high-speed streaming / discrete events), and spatial dimensions, which collectively form the raw data basis for subsequent steps of standardization, feature extraction, cross-modal fusion, and simulation optimization.
[0039] Firstly, the system accesses the multi-source industrial data in real time or batch through the pre-set interface protocol and data adaptation module, ensuring the completeness and timeliness of data acquisition. During the acquisition process, according to the characteristics of various data sources, specially designed data synchronization mechanisms and timestamp alignment strategies are applied to solve the asynchronous problem of multi-modal data in the time dimension and ensure the spatio-temporal consistency of subsequent processing.
[0040] In the data preprocessing link, first, the multi-modal raw data is executed for data cleaning, including missing value filling, outlier detection and elimination, noise filtering and format unification processing. The missing value filling adopts a multi-strategy fusion method based on statistical inference and machine learning, selects the optimal filling scheme for different modalities and data properties, and ensures data integrity and accuracy. The outlier detection combines rule engine and adaptive threshold algorithm to dynamically identify abnormal sampling points and record abnormal event logs, facilitating subsequent tracking and diagnosis. Noise filtering uses modality-specific signal processing techniques, such as filter design for time series data and denoising algorithms for image data, to improve data quality. Further, the cleaned data is standardized, specifically through normalization, mean-variance standardization and distribution transformation methods, to adjust different modal data to a unified numerical scale and distribution range, eliminating the influence of dimension and facilitating subsequent fusion and analysis. During the standardization process, the system dynamically adjusts parameters according to pre-set industrial standards and application requirements, supporting the switching and combined application of multiple standardization schemes.
[0041] Further, the preprocessing module is built-in with data quality evaluation mechanism, which monitors the data completeness, accuracy and consistency indicators in real time, and triggers automatic alarm and records detailed logs in abnormal cases. The preprocessing process adopts distributed computing architecture, supports efficient parallel processing of large-scale industrial data, and guarantees the real-time and stability of processing. After processing, a structured preprocessed dataset is formed, containing multi-modal data in unified format and its cleaning and standardization state identification. The preprocessed data is passed to the "raw features" processing module in the next step S120 as the output field name, for feature extraction and vectorization processing. The integrity and standardization of the preprocessed data lay a foundation for accurate extraction of subsequent multi-modal features, and support cross-step data calling and reuse, ensuring data flow closed loop and consistency of the entire industrial simulation optimization process.
[0042] S120, extracting each modal feature from the preprocessed data, performing vectorization processing to generate raw feature vectors; Specifically, each modal feature is extracted from the preprocessed data and vectorized. First, the system automatically identifies and separates different modal information in the preprocessed data according to pre-defined modal classification rules, such as sensor signal modal, image modal, text log modal, etc. For each modal, a specially designed feature extraction algorithm is used to fully exploit its internal information. For time series sensor data, time-frequency analysis, wavelet transform and statistical feature extraction methods are applied to generate a feature set representing dynamic changes; for image data, deep learning models such as convolutional neural network (CNN) are used to extract spatial features; for text log data, word embedding and semantic analysis methods in natural language processing (NLP) are used to extract semantic feature vectors.
[0043] During feature vector generation, the system normalizes each modal feature to ensure consistent feature scale and avoid dimensional differences affecting fusion. Further, to improve the compactness and effectiveness of feature expression, feature selection and dimensionality reduction techniques such as principal component analysis (PCA), linear discriminant analysis (LDA) and autoencoder are used to select a representative and low-redundancy feature subset. During this process, the system automatically records the extraction parameters and quality indicators of each feature, supporting subsequent tracing and optimization.
[0044] Further, the system uniformly encapsulates the modal feature vectors in a format to form a set of original feature vectors as the output field name of this step. The original feature vectors are passed to the "feature mapping" module of the next step S130 for constructing a structured feature mapping relationship. It is worth noting that the feature extraction process supports dynamic adjustment and expansion, and can flexibly adapt to different modalities and feature types according to the needs of industrial simulation, ensuring the diversity and expression ability of the feature vectors, and providing a solid foundation for cross-modal fusion. Further, the generation and management of the original feature vectors follow uniform data specifications and interface standards, ensuring seamless integration and efficient calling of subsequent modules.
[0045] S130, structurally storing the original feature vectors to generate a feature mapping relationship; Specifically, the aforementioned original feature vectors are taken as inputs to complete the systematic organization of features by constructing a multi-dimensional index structure and an associated mapping mechanism. First, the system establishes a multi-level, multi-dimensional index system according to the metadata of the features, such as modal categories, timestamps, and spatial identifiers, to realize fast positioning and access of the feature vectors. Further, for the feature correlation between different modalities, a feature mapping relationship is constructed to clearly define the corresponding relationship, dependency relationship, and mutual influence path between the features of different modalities. The mapping relationship is expressed in the form of a graph structure or a matrix, supporting dynamic updating and expansion, and facilitating the description of complex inter-modal interaction and fusion requirements.
[0046] During the mapping relationship construction process, the system sets mapping rules and constraints in combination with the specific needs of the industrial simulation scene to ensure the rationality and effectiveness of the mapping. The mapping process includes feature classification, label assignment, and correlation strength calculation, and statistical analysis and machine learning methods are used to evaluate the correlation and complementarity between features. The system manages the changes to the mapping relationship and records the key parameters and state information during the mapping construction process, ensuring the traceability and consistency of the mapping data. The feature mapping relationship is passed to the "modal feature" processing module of the next step S210 as the output field name, for further feature normalization and dimensionality reduction processing.
[0047] Further, the structured storage of the feature mapping relationship uses an efficient database system or a distributed storage architecture to support the storage and access of large-scale feature data, meeting the strict requirements of industrial simulation on data volume and access speed. The mapping relationship not only provides basic data support for subsequent cross-modal attention weighting calculation, but also provides key inputs for the dynamic feature selection and fusion module, ensuring the accuracy and efficiency of multi-modal feature fusion.
[0048] Step S200 includes at least steps S210-S230: S210, acquire modal features from the feature mapping relationship, perform normalization processing, and generate standardized features; Specifically, the feature mapping relationship from the S130 step is taken as input, and the feature mapping relationship contains pre-structured stored multi-modal original feature vectors and their corresponding indexes and associated information. The system first parses the metadata in the mapping relationship, such as modal category, timestamp, and spatial identifier, and uniformly processes each modal feature according to a predefined normalization strategy. The normalization processing includes but is not limited to minimum-maximum normalization, mean-variance standardization, and distribution transformation, and the specific selection is dynamically determined based on the numerical range, distribution characteristics, and industrial simulation requirements of the modal features. During the normalization process, the system performs a preset fault-tolerant mechanism for abnormal values and missing data, such as abnormal value truncation and interpolation filling, to ensure the stability and consistency of the normalization results.
[0049] Further, the normalization processing is performed in parallel using a distributed computing framework, which realizes efficient batch processing for large-scale multi-modal feature data, and monitors the change trend of the normalization parameters in real time during the processing process, dynamically adjusts the normalization scheme to adapt to the time-varying nature of data distribution. The system evaluates the quality of the normalization results, calculates the statistical indicators and distribution consistency of the normalized features, and marks and logs the features that are abnormal or deviate from the expected values for subsequent analysis and optimization. The standardized features are taken as the output field name of this step and are passed to the "feature sequence" module of the next step S220 for subsequent dimensionality reduction processing and feature simplification. The generation of the standardized features ensures the comparability and fusion feasibility of different modal features in numerical scale, and provides basic data support for the effective integration of subsequent multi-modal features.
[0050] S220, perform dimensionality reduction processing on the standardized features to generate a simplified feature sequence; Specifically, the standardized features output from the S210 step are taken as input, and the system uses various dimensionality reduction algorithms to process the standardized features based on the internal structure and redundancy of multi-modal features. The dimensionality reduction processing includes principal component analysis (PCA), linear discriminant analysis (LDA), automatic encoder (Autoencoder), and nonlinear dimensionality reduction methods based on graph embedding and manifold learning. The system dynamically selects suitable dimensionality reduction algorithms or combination strategies according to the characteristics and simulation requirements of different modal features, and the triggering conditions include feature dimension threshold, computational resource limitation, and dimensionality reduction effect index.
[0051] In the dimensionality reduction process, the system first performs statistical analysis on the standardized features of the input, identifies the correlation and redundancy between the features, and combines or eliminates highly correlated or redundant feature groups. The dimensionality reduction algorithm maps high-dimensional features to a reduced feature subspace by constructing a low-dimensional space mapping function, retaining the main information components and discriminant ability. The system performs reconstruction error evaluation and information retention analysis on the dimensionality reduction results, and automatically adjusts or replaces abnormal or severely information loss dimensionality reduction schemes. The parameter configuration, calculation process and results in the dimensionality reduction process are recorded in detail, supporting subsequent optimization and traceability.
[0052] The reduced feature sequence is passed to the "feature representation" module of the next step S230 as the output field name of this step, for use in constructing the multi-modal feature library. This reduced feature sequence reduces the data dimension and computational complexity, improves the compactness and effectiveness of feature expression, and provides an optimized input data structure for subsequent cross-modal fusion and attention weighting calculation.
[0053] S230, constructing a feature representation model based on the reduced feature sequence to generate a multi-modal feature library; Specifically, the reduced feature sequence from the S220 step is taken as input, and the system encodes the feature sequence in depth through the construction of a feature representation model to form a unified multi-modal feature library. The feature representation model uses a multi-layer neural network architecture, including but not limited to convolutional neural network (CNN), recurrent neural network (RNN), transformer and its variants, combined with self-attention mechanism to encode and fuse the feature sequence. The model training process is based on labeled data in industrial simulation scenarios or unsupervised learning strategies, using gradient descent and other optimization algorithms to adjust model parameters, improving the discriminability and generalization ability of feature representation.
[0054] In the feature representation construction process, the system combines the spatiotemporal characteristics of multi-modal features to form multi-dimensional and multi-scale feature representations by fusing time series information and spatial distribution features. The feature representation model supports dynamic updating and online learning, and can adapt to changes in industrial field data and the introduction of new modalities. The system monitors and records the model performance indicators, training process parameters and output feature vectors during the construction process to ensure the stability and usability of the feature library. The multi-modal feature library is stored in an efficient data structure to support fast retrieval and batch access, meeting the needs of the subsequent cross-modal attention weighting calculation module for feature input.
[0055] The multi-modal feature library is taken as an output field name of the step, and is transmitted to the "feature input" module of the next step S310 for cross-modal attention weighting calculation. The multi-modal feature library provides a rich and structured feature basis for subsequent weighting calculation and dynamic feature selection, and promotes deep fusion of multi-modal information and fine processing of industrial simulation optimization.
[0056] The step S300 at least includes steps S310-S330: S310, obtaining feature input from the multi-modal feature library, constructing an attention query matrix, and generating an initial attention distribution; Specifically, the multi-modal feature library from the step S230 is taken as input, which includes a multi-dimensional and multi-scale feature vector set constructed by deep feature representation, covering time series features, spatial distribution features, and semantic information. The feature input is first filtered by a feature selection module to obtain a subset meeting the current simulation task requirements, and is weighted and adjusted according to preset modal weights and feature importance indexes to form a preliminary input feature matrix. Further, the system constructs an attention query matrix (Query Matrix) according to the design of the cross-modal attention weighting algorithm, the structure of which is based on the query mechanism in the Transformer architecture. Specifically, the query matrix is obtained by linear transformation of the feature input, and is used to capture the mutual dependence between modal features.
[0057] During the construction process, the system simultaneously generates a key matrix (Key Matrix) and a value matrix (Value Matrix), which correspond to different linear mappings of the feature input respectively, and support the implementation of the multi-head attention mechanism. The query matrix, the key matrix, and the value matrix are all parameterized by using the multilayer perceptron (MLP, Multilayer Perceptron) and the attention weight parameter initialization strategy, and the parameters include weight matrices and bias terms, the initial values of which are preset according to historical training data and simulation scenarios. The system performs parallel calculation of the query, key, and value vectors of the modal features through a batch processing mechanism, ensuring the efficiency and real-time performance of data processing. Further, the system performs dot product operation to perform matrix multiplication between the query matrix and the key matrix, and obtains an attention score matrix reflecting the correlation strength between modal features.
[0058] To avoid gradient vanishing or explosion caused by excessively large values, the attention score matrix is scaled, specifically by dividing by the square root of the query vector dimension. Further, the system applies a Softmax function to the scaled score matrix to generate an initial attention distribution. The Softmax function converts the scores into a probability distribution, representing the relative attention of each modality feature to the target task. This initial attention distribution is accompanied by a confidence evaluation mechanism, which quantifies the certainty of the distribution using statistical confidence intervals and entropy calculations. Abnormal distributions and extreme weight values are recorded in the log system for future parameter optimization reference. The initial attention distribution serves as the output field name of this step and is passed to the "weight parameter" module in the next step S320 for dynamic weight parameter optimization.
[0059] S320, parameter optimization based on initial attention distribution to generate dynamic weight parameters; Specifically, taking the initial attention distribution output from S310 as input, the system first performs quality evaluation on the distribution using multiple indicators including distribution smoothness, weight concentration, and inter-modal weight difference to determine the optimization direction and adjustment strategy. The parameter optimization process uses an iterative algorithm combined with Gradient Descent and an adaptive learning rate mechanism to dynamically adjust the attention weight parameters, improving the accuracy and robustness of the weighting effect. During optimization, the system introduces a regularization term to avoid overfitting. Regularization strategies include a weighted combination of L1 and L2 norms to constrain the weight adjustment amplitude for different modality features.
[0060] Further, the system combines feedback information from industrial simulation scenarios, such as simulation error indicators and historical optimization results, to adopt a Reinforcement Learning strategy and construct a reward function to guide the update of weight parameters. The reward function considers simulation accuracy improvement, computational resource consumption, and model stability to achieve optimal adjustment of weight parameters through policy gradient methods. During optimization, the system monitors the update trajectory of dynamic weight parameters, and abnormal fluctuations or sudden weight changes trigger an abnormal handling mechanism to perform parameter rollback or re-initialization, ensuring the stability of the optimization process. The dynamic weight parameters are generated in parallel through a distributed computing framework, supporting real-time processing requirements for large-scale multi-modal features.
[0061] Furthermore, the system implements version control and parameter snapshot functions for the dynamic weight parameters, recording the parameter status and related indicators for each optimization iteration, facilitating subsequent auditing and model tuning. The dynamic weight parameters, as the output field names of this step, are passed to the "Adjustment Coefficient" module in the next step S330 for normalization and weight matrix generation, supporting subsequent dynamic feature selection and fusion. This dynamic weight parameter generation process enables fine-tuning of the initial attention distribution, improving the adaptability and expressiveness of cross-modal weighting.
[0062] S330. Normalize the dynamic weight parameters to generate a cross-modal attention weight matrix; Specifically, taking the dynamic weight parameters from step S320 as input, the system first filters the weight parameters for valid values, eliminating abnormal weights and noise interference, and uses statistical thresholds and anomaly detection algorithms to determine the filtering boundaries. Further, for the filtered weight parameters, the system performs normalization processing, employing various normalization methods, including but not limited to max-min normalization, L1 norm normalization, and Softmax normalization, dynamically selecting the appropriate scheme based on the characteristics and fusion requirements of the industrial simulation task. The normalization process is achieved through matrix operations, adjusting the weight parameters to a uniform numerical range to ensure the comparability and stability of weights from different modalities.
[0063] Furthermore, the system constructs a cross-modal attention weight matrix, where the rows and columns correspond to different modalities and their feature dimensions, and the matrix elements represent the weight relationships and mutual influence between modalities. The weight matrix is stored using a sparse matrix structure to optimize storage efficiency and computational performance, supporting fast indexing and updates. The system combines the weight matrix with consistency checks and positive definiteness checks to ensure that the matrix's mathematical properties meet the requirements of subsequent fusion algorithms. Abnormal matrix elements trigger an automatic correction mechanism, using interpolation and neighborhood smoothing techniques to adjust abnormal weights, maintaining the overall smoothness and rationality of the matrix.
[0064] During normalization and matrix construction, the system records the statistical characteristics and trends of the weight parameters in real time, generating a weight parameter log to support model performance analysis and parameter tracking. The cross-modal attention weight matrix, as the output field name of this step, is passed to the "fusion parameter" module in the next step S410 for use by the dynamic feature selection and fusion module, enabling weighted fusion and optimization of multimodal features. The generation of this weight matrix marks the completion of cross-modal attention weighted calculation, providing accurate weighting basis for subsequent feature fusion.
[0065] In another embodiment, in S310, standardized feature vectors stored in a multimodal feature library are first input, and sensor signals, images, and text logs are separated into three modal features using preset modality classification rules. The system constructs a multimodal attention calculation model based on improved Shannon entropy, and formula ① defines the cross-modal information entropy weights: in: For cross-modal information entropy weights; Indicates the first Modal feature vectors; For summation index variables; The upper limit for summation represents the total number of dimensions of the feature or the total number of data points; The feature probability distribution comes from a multimodal feature library; For cross-modal reference distribution; This is the time synchronization offset; This is the time decay factor (values range from 0.01 to 0.1).
[0066] Data source mapping: generated from physical quantity data collected by sensors. Control command data generation Device operation logs provide Formula ① Calculation result Row vector elements used to construct the attention query matrix.
[0067] Formula ② calculates the multi-head attention score: in: The formula output is the first Attention score matrix for each attention head; For the first The query matrix of each attention head is given by formula ①. Generated through linear transformation; The key matrix is derived from standardized features from a multimodal feature library; It is a normalized exponential function; Represents the matrix transpose operation; The query vector dimension (values 64-512).
[0068] Data source mapping: Query matrix elements originate from sensor signal modal features, and key matrix elements originate from image modal features. The attention score matrix output by Formula ② is concatenated to form the initial attention distribution, which is consumed by the weight parameter module of S320.
[0069] The output field name "Initial Attention Distribution" contains the weight distribution probability of each modality feature, which is passed to the weight parameter optimization module of S320 through the distributed computing framework.
[0070] Furthermore, in S320, the initial attention distribution from S310 is first adopted, and a modified quadratic cost function is used for parameter optimization. Equation ③ defines the dynamic weight adjustment amount: in: This represents the weight adjustment for the m-th mode; The index of the target mode specifies the mode in which the weight adjustment is currently being calculated. ; The total number of modes; The gradient learning rate controls the strength of the influence of the gradient descent term on the weight adjustment. ; Let be the gradient of the weights of the m-th mode, and be the cross-modal information entropy in Equation ①. Weights The partial derivatives of the loss function represent the partial derivatives of the loss function. exist Rate of change in direction; The coefficients are used to control the influence of feature-similarity-based collaborative optimization terms on the total adjustment. ; The index of the source mode is used to iterate through the indices of all modes during the summation operation. ; The current weight parameters for the k-th mode; The current weight parameters for the m-th mode; This is a feature similarity scaling factor that controls the sensitivity of the exponential term to feature distance. The larger the value, the faster the similarity decays. ; It is the L2 norm, used to calculate the distance between two eigenvectors in space; Let be the standardized feature vector of the m-th mode, representing the overall representation of all features of that mode; Let be the standardized feature vector of the k-th mode.
[0071] Data source mapping: Gradient calculation data comes from historical optimization records in the device's operation logs, feature vectors From a multimodal feature library. The output of formula ③ is... Used for iteratively updating dynamic weight parameters.
[0072] Formula ④ establishes the parameter optimization constraints: in: To optimize the constraints of the problem, it means that the optimization of the weight parameters must be carried out under the premise of satisfying this inequality; Here, the optimized weight parameters for the m-th mode are... The adjustment amount is calculated using formula ③. The result after iterative update; is the stability index of the m-th modal feature, which is a value that quantifies the reliability of the data of that modality (e.g., calculated based on data integrity rate, signal-to-noise ratio, etc.). ; A preset global constraint threshold limits the upper limit of the sum of all weight-stability products. .
[0073] Data source mapping: The completeness rate index is calculated from the data quality assessment module. The optimization results under the constraints of Formula ④ generate dynamic weight parameters, which are consumed by the adjustment coefficient module of S330.
[0074] The output field name "Dynamic Weight Parameters" contains the collaboratively optimized modal weight coefficients, which are passed to the normalization processing module of S330 through the version control interface.
[0075] The S330 receives dynamic weight parameters and performs normalization based on the improved L2 norm. Formula ⑤ defines the normalization mapping function: in: This represents the normalized weights for the m-th modality. This is the output of the formula, containing the elements of the attention weight matrix. This is the moving average of the weights for the m-th modality. It is used for centering and is calculated based on historical data of this weight parameter. This represents the standard deviation of the weights for the m-th modality. It is used for scaling and is calculated based on historical data for this weight parameter. For symbolic functions, defined as follows: This is used here to preserve the original polarity (positive or negative) of the weights. The sparsity coefficient controls the intensity of exponential decay. The larger the value, the greater the penalty for weights with larger absolute values, thus encouraging the weight vector to become sparse (more values close to zero). .
[0076] Data source mapping: and The weights are obtained from historical weight data of the multimodal feature library. The calculation results of Formula ⑤ are matrixed and reorganized to form a cross-modal attention weight matrix.
[0077] Formula ⑥ verifies the matrix orthogonality constraint: in: , These are the row and column indices of the weight matrix; The total number of modalities (number of rows); The total number of dimensions (columns) of the features; The element in the m-th row and n-th column of the weight matrix represents the contribution weight of the n-th dimension feature of the m-th modality to the final fusion result. The preset orthogonality threshold constrains the upper limit of the sum of squares of all elements in the matrix (the square of the Frobenius norm), and is used to control the overall energy or sparsity of the matrix (values range from 1.0 to 1.5).
[0078] Data source mapping: The threshold parameter is derived from the simulation scenario configuration table. The weight matrix verified by formula ⑥ is used as a valid output and consumed by the fusion parameter module of S410.
[0079] The output field name "Attention Weight Matrix" is stored in a distributed sparse matrix system. Its row and column indices correspond to different modal feature dimensions, and the matrix elements represent the normalized cross-modal influence weights. In summary, this section describes the technical effect: By using an improved information entropy model and a quadratic optimization algorithm, a spatiotemporal collaborative dynamic weight calculation system is constructed. The generated normalized weight matrix provides a quantitative basis for cross-modal feature fusion.
[0080] Step S400 includes at least steps S410-S430: S410. Combine the attention weight matrix with the multimodal feature library to evaluate feature relevance and generate a candidate feature set; Specifically, step S410 first obtains multi-dimensional feature vectors from the multimodal feature library, which contains time-series features, spatial distribution features, and semantic information constructed using deep feature representation. The attention weight matrix, output by step S330, represents a weighted relational matrix between modalities and feature dimensions, with its elements reflecting the degree and importance of mutual influence between different modal features. The system interfaces the cross-modal attention weight matrix with the multimodal feature library through an interface module, ensuring consistency in the dimensional and index correspondence between the weight matrix and the feature library, supporting subsequent fusion calculations. Specifically, based on the weight values in the weight matrix, the system performs a correlation assessment on each modal feature in the multimodal feature library. This correlation assessment includes calculating the covariance matrix, correlation coefficient, and mutual information index of the weighted features, used to quantify the linear and non-linear dependencies between features.
[0081] Furthermore, the feature correlation evaluation module combines statistical analysis and machine learning methods, employing a hybrid model based on Graph Neural Network (GNN) and spectral clustering algorithms to calculate the correlation strength of feature nodes in the multimodal feature library. Based on the weight distribution among different modalities in the weight matrix, the system dynamically adjusts the edge weights in the graph structure to construct a feature correlation graph between modalities. This feature correlation graph, through an iterative propagation mechanism, identifies the centrality index of feature nodes and the importance of edge weights, forming a candidate feature set. This candidate feature set contains a subset of features with high relevance and information content after weighted evaluation. During the evaluation process, the system incorporates an anomaly detection mechanism to mark feature nodes with abnormal weights or correlations and records anomaly logs to support subsequent analysis and adjustments.
[0082] Furthermore, the system, considering the specific needs of the industrial simulation scenario, sets feature correlation thresholds and constraints to filter out features that meet the specified correlation criteria, forming a preliminary candidate feature set. This candidate feature set, as the output field name of this step, is passed to the "Filtering Threshold" module in the next step S420 for dynamic threshold setting and generation of effective feature subsets. This candidate feature set is used in subsequent modules for further filtering and fusion, promoting the efficient utilization and optimization of multimodal features. Step S410 achieves a deep integration of cross-modal weights and multimodal features, constructing a key data foundation for feature correlation evaluation and supporting refined processing of dynamic feature selection and fusion.
[0083] S420. Combine the candidate feature set, set a dynamic filtering threshold, and generate an effective feature subset; Specifically, based on the simulation scenario requirements and task parameters, feature selection thresholds are dynamically set. These thresholds are defined as numerical limits used to determine the importance and effectiveness of features, encompassing multiple dimensions such as relevance thresholds, information entropy thresholds, and weight contribution thresholds. The system first obtains the parameter settings for the current simulation scenario from the simulation task configuration module, including but not limited to simulation objectives, accuracy requirements, computational resource limitations, and real-time requirements. Based on these parameters, the system dynamically selects the range and weights of the thresholds. The dynamic threshold setting module employs a hybrid optimization strategy based on Bayesian optimization and genetic algorithms to iteratively calculate the optimal threshold combination, balancing the accuracy of feature selection with computational efficiency.
[0084] Furthermore, the system inputs the feature importance index of the candidate feature set into the threshold determination module, and filters and determines each feature based on a dynamic threshold. During the filtering process, the system performs multiple rounds of iterative filtering, combining statistical significance testing, redundancy analysis, and intermodal complementarity evaluation of features to eliminate features below the threshold standard. The filtering operation is implemented through a parallel computing framework, supporting efficient processing of large-scale feature data. For boundary features, the system adopts a fuzzy determination mechanism, combining historical simulation feedback and expert rules for manual intervention or automatic adjustment to ensure the rationality of the filtering results. Abnormal features and threshold adjustment records during the filtering process are fully stored in the log system for easy subsequent tracking and optimization.
[0085] The effective feature subset, output by the filtering module, contains multimodal features that meet the dynamic threshold conditions. This effective feature subset serves as the output field name for this step and is passed to the "fusion object" module in the next step, S430, for weighted fusion processing. This effective feature subset reflects the most representative and contributing feature set in the simulation scenario, providing an optimized input basis for subsequent fusion. Step S420, through dynamic threshold setting and a multi-dimensional filtering mechanism, achieves adaptive and flexible adjustment of feature selection, enhancing the targeting and efficiency of multimodal feature fusion.
[0086] S430. Perform weighted fusion processing on the effective feature subset to generate a fused feature vector; Specifically, the effective feature subset undergoes weighted fusion processing to generate a fused feature vector. This weighted fusion processing is based on the weight distribution in the cross-modal attention weight matrix, combining feature importance indices and inter-modal complementary information, and achieves feature fusion through a combination of linear weighting and non-linear mapping. Specifically, the system first multiplies each feature vector in the effective feature subset by its corresponding weight coefficient. This weight coefficient originates from the cross-modal attention weight matrix output in step S330, reflecting the feature's contribution to the overall fusion. Further, the system employs a multilayer perceptron (MLP) to perform a non-linear transformation on the weighted features, capturing complex interactions and implicit patterns between features.
[0087] Furthermore, during the fusion process, the system introduces a feature regularization mechanism, including L2 regularization and sparsity constraints, to prevent feature overfitting and the accumulation of redundant information. The fusion module supports a multi-head fusion strategy, processing different modal feature subsets in parallel through multiple independent weighted channels. The fusion results are summarized and normalized in the final stage to form a unified fused feature vector. The fusion process combines batch normalization and residual connection techniques to improve the stability and expressive power of the fused features. The system monitors the statistical distribution and information entropy of the fused features in real time during the fusion computation, triggering an adjustment mechanism based on abnormal fluctuations to ensure the quality of the fusion results.
[0088] Furthermore, the fusion module supports online updates and incremental fusion, enabling it to respond to dynamic changes in data and models within the industrial simulation environment and maintain the timeliness and adaptability of the fused features. The fused feature vector, as the output field name of this step, is passed to the "Parameter Input" module in the next step S510 for use in simulation parameter generation and execution. Step S430 implements weighted integration and information enhancement of multimodal features, providing a high-quality fused feature foundation for the accurate generation of subsequent simulation parameters.
[0089] In another embodiment, the input comes from the attention weight matrix of S330 and the multimodal feature library of S230, and a feature association model is constructed through a graph neural network. Formula ③ defines the feature node centrality index: in: The centrality index of node v quantifies the importance of this feature node in the entire feature graph; the larger the value, the more important it is. The index of the current target node; Let u be the set of neighboring nodes of node v, and u be each neighboring node in the set. The neighbor relationships are defined based on the attention weight matrix. For example, nodes with a weight exceeding a certain threshold are considered neighbors; Let v be the set of neighbors of node v; The edge weights between nodes u and v are directly derived from the cross-modal attention weight matrix output by S330. The corresponding element in the matrix represents the correlation strength between the two features; The degree of the node; Let v be the degree of node v; The feature space distance; This is the distance attenuation coefficient (value range: 0.1-0.3).
[0090] Data source mapping: Weight data comes from the attention weight matrix, and feature distance is obtained by vector calculation from the multimodal feature library. Formula ③ outputs... As the basis for generating candidate feature sets.
[0091] Formula ④ enables dynamic selection of candidate features: in: The candidate feature set is a collection of feature vectors that have been initially selected. Let be the i-th feature vector, which is a specific feature sample in the multimodal feature library; Using the index of the feature vector, iterate through all features in the multimodal feature library; It is a linear rectified function. This is used to filter out features with centrality below the basic threshold (setting the input values of these features to 0); is the centrality index of the i-th feature node (calculated from formula ③); The basic threshold (values 0.4-0.6); As a characteristic stability factor; This is the noise suppression coefficient (value range: 0.05-0.1).
[0092] Data source mapping: The stability factor is derived from the accuracy metric of the data quality assessment module. The candidate feature set generated by formula ④ is consumed by the screening threshold module of S420.
[0093] The output field name "candidate feature set" is stored as an edge weight subgraph in the graph structure database, containing feature nodes and their association strength information, and is passed to the dynamic threshold module of S420 through the parallel computing interface.
[0094] The S420 receives a set of candidate features and uses an improved Bayesian optimization algorithm to set a dynamic threshold. Equation ⑤ defines the threshold adaptive function: in: This is a dynamic threshold, ultimately used to filter the centrality threshold for effective features; This represents the historical average threshold. Standard deviation; The inverse error function, the error function It is the integral transform of the standard normal distribution probability, and its inverse function is... Used to find the quantile (z-score) of the standard normal distribution based on the target probability. Target selection probability (value 0.7-0.9).
[0095] Data source mapping: The mean and variance parameters are obtained from historical data statistics in the validation and evaluation report. The result of formula ⑤ serves as the dynamic threshold benchmark.
[0096] Formula ⑥ constructs multi-constraint screening conditions: in: An effective feature subset contains all feature vectors that simultaneously satisfy the centrality and stability requirements; This represents the minimum stability requirement (value 0.6). Let j be the j-th feature vector, which comes from the candidate feature set. ; Let be the centrality index of the j-th feature; The dynamic threshold (calculated using formula ⑤); Let be the stability factor of the j-th feature; This is the minimum stability requirement; features must meet a minimum stability standard. Features falling below this value are considered unreliable and are discarded. .
[0097] Data source mapping: The stability threshold is derived from the simulation scenario configuration file. The effective feature subset output by Equation ⑥ is consumed by the fusion object module of S430.
[0098] The output field name "Valid Feature Subset" is stored through a feature index mapping table, which contains feature vector indexes that meet the dynamic threshold conditions and their associated metadata.
[0099] Technical effect in this section: Based on feature association analysis and Bayesian optimization threshold setting using graph neural networks, an adaptive feature selection mechanism for industrial scenarios is realized, providing high-quality input for multimodal fusion.
[0100] Step 500 includes at least steps S510-S530: S510. Convert the fused feature vector into simulation parameter format to generate the initial simulation configuration; Specifically, the fused feature vector output from step S430 is used as input. This fused feature vector contains multimodal industrial feature information that has undergone weighted fusion processing, reflecting the comprehensive expression of each modality's features and the results of dynamic weight adjustment. The system first performs format parsing and attribute recognition on the fused feature vector, and then maps the high-dimensional fused features to the simulation parameter space according to preset simulation parameter mapping rules. The simulation parameter format definition includes attributes such as parameter name, type, value range, and unit. The mapping rules are dynamically configured according to the interface standards of the industrial simulation system and the requirements of the simulation model, supporting adaptation to multiple simulation platforms.
[0101] During the mapping process, the system employs a combination of mapping matrices and transformation functions to convert the numerical features in the fused feature vector into specific simulation parameter values. The mapping matrix is constructed based on the correspondence between feature dimensions and simulation parameter dimensions. The transformation functions include various forms such as linear mapping, nonlinear activation, and threshold truncation, adapting to the physical meaning and constraints of different simulation parameters. Specifically, the system performs a normalization inverse transformation on each element in the fused feature vector, adjusting the numerical scale to the effective range of the simulation parameters to prevent parameter overflow or invalidation. The mapping process supports batch processing and parallel computing, meeting the real-time and efficiency requirements of industrial-grade simulation tasks.
[0102] Furthermore, the system performs integrity and consistency checks on the generated simulation parameters, including parameter value range detection, data type verification, and parameter dependency checks. Abnormal parameters trigger an early warning mechanism, automatically recording exception logs and executing parameter corrections or invoking alternative solutions according to preset rules. The system supports version management and change tracking of simulation parameters, recording key steps in the mapping process and the status of the parameter mapping matrix for easy subsequent auditing and optimization. The initial simulation configuration is stored in a unified data structure, containing a complete set of simulation parameters and their metadata, supporting rapid loading and invocation of the simulation engine.
[0103] The initial simulation configuration, as the output field name of this step, is passed to the "Environmental Parameters" module in the next step S520 for configuring and loading the simulation environment. This initial simulation configuration, which incorporates the results of multimodal feature fusion, is a key data carrier for the transition from the simulation parameter generation and execution module to the simulation computing environment, ensuring the accuracy and applicability of the simulation parameters.
[0104] S520: Loads the physics engine based on the initial simulation configuration to generate the simulation computing environment; Specifically, the initial simulation configuration from step S510 is taken as input. This initial simulation configuration includes a set of mapped simulation parameters and their related attribute information. The system first parses the simulation parameters and, based on parameter type and function classification, calls the corresponding physics engine module for environment configuration. The physics engine is a core component of the industrial simulation platform, covering various physical models such as dynamics simulation, thermodynamics simulation, fluid simulation, and control system simulation, and supports modular loading and parameterized customization.
[0105] During environment configuration, the system automatically adjusts the model parameters, boundary conditions, and initial states of the physics engine based on the parameter values in the initial simulation configuration. These adjustments include, but are not limited to, assigning physical properties, constructing the simulation scene, and setting the simulation time step. The configuration process is dynamically optimized in conjunction with simulation task requirements and resource constraints, supporting multi-threaded and distributed computing architectures. The system performs configuration integrity checks, verifies interface compatibility and parameter consistency between physics engine modules, and triggers logging and error recovery mechanisms for abnormal configurations to ensure the stability and availability of the simulation environment.
[0106] Furthermore, the system loads the auxiliary resources required for the simulation computing environment, including mesh generation data, material property libraries, and control strategy scripts, ensuring that the simulation environment has complete operating conditions. The simulation computing environment is deployed using containerization and virtualization technologies, supporting rapid environment reconstruction and version switching. The system monitors performance and schedules resources during the environment loading process, optimizing computational load allocation and improving the response speed and throughput of the simulation computing.
[0107] The simulation computing environment, as the output field name of this step, is passed to the "Execution Instruction" module of the next step S530 for execution of simulation calculations and optimization algorithms. This simulation computing environment builds a bridge between simulation parameters and physical simulation models, and is the core operating platform of the simulation parameter generation and execution module, carrying out the execution of subsequent simulation calculation tasks.
[0108] S530. Execute the optimization algorithm in the simulation computing environment to generate the original simulation results; Specifically, the simulation computing environment from step S520 is taken as input, and the simulation computing environment includes the loaded physics engine module and its configuration parameters. The system starts the simulation computing process according to the preset optimization algorithm strategy. The optimization algorithm includes, but is not limited to, genetic algorithm, particle swarm optimization, gradient descent and its variants, and performs parameter search and adjustment in combination with the specific objectives and constraints of the industrial simulation task.
[0109] During execution, the system inputs simulation parameters into the physics engine, driving the simulation model to perform time-step calculations and dynamically simulate the physical behavior and system response of industrial processes. The simulation calculations encompass multiphysics coupling, nonlinear dynamics, and complex boundary condition handling, employing high-precision numerical solution methods and parallel computing techniques to improve computational efficiency. The optimization algorithm iteratively adjusts simulation parameters, evaluates the merits of parameter combinations based on performance indicators from the simulation results, and guides the next round of parameter updates. The system, combined with a real-time monitoring module, tracks key variables and status information during the simulation process, identifies computational anomalies and convergence conditions, and triggers anomaly handling and interruption mechanisms to ensure the continuity and stability of the simulation calculations.
[0110] Furthermore, the system preprocesses the simulation results data, including data cleaning, noise filtering, and format conversion, to generate a structured raw simulation results dataset. The raw simulation results include time-series data, spatial distribution data, and key performance indicators, supporting multi-dimensional analysis and subsequent verification. The system implements version management and storage strategies for the results data, recording simulation parameter configurations, calculation process logs, and optimization algorithm iteration history, facilitating result traceability and performance evaluation.
[0111] The original simulation result, as the output field name of this step, is passed to the "Verification Input" module of the next step S610 for use in multi-dimensional verification and optimization output. This original simulation result carries the final calculation result of the simulation parameter generation and execution module and is a direct reflection of the multi-modal feature fusion effect in the industrial simulation optimization process.
[0112] Step S600 includes at least steps S610-S630: S610. Perform spatiotemporal alignment processing on the original simulation results to generate standardized verification input; Specifically, the raw simulation results from step S530 are used as input. These raw simulation results include multidimensional time-series data, spatial distribution data, and key performance indicators generated during the industrial simulation calculation. The raw simulation results are first connected to the spatiotemporal alignment processing unit via a data interface module. Based on preset spatiotemporal alignment rules and multimodal data timestamp information, the time-series data and spatial data in the simulation results are synchronized and standardized. The spatiotemporal alignment processing includes timestamp correction, time window segmentation, spatial coordinate unification, and data interpolation completion, aiming to transform asynchronous, non-uniformly sampled simulation data into a standardized data set within a unified spatiotemporal reference framework.
[0113] Specifically, the system first parses the timestamp information in the original simulation results and, combined with the time synchronization mechanism during multi-source industrial data acquisition, corrects the time dimension to resolve time offset issues caused by computation delays or data transmission. The time window segmentation module, designed according to simulation task requirements and verification indicators, sets a fixed or adaptive time window length to divide the continuous time series into several time periods, facilitating subsequent indicator calculations and comparative analysis. In the spatial dimension, the system performs coordinate transformation and unification of the spatial data based on the spatial reference coordinate system of the simulation model, and uses spatial interpolation algorithms to complete missing or irregular spatial points, ensuring the integrity and consistency of the spatial data.
[0114] During processing, the system, in conjunction with an anomaly detection module, performs anomaly removal or correction operations on abnormal fluctuations in the time series and outliers in the spatial data. Anomaly events and correction records are written to the log system for subsequent tracking and analysis. The spatiotemporal alignment processing employs a distributed computing architecture, supporting parallel processing of large-scale simulation results data and improving processing efficiency. After processing, a structured, standardized verification input dataset is formed, containing spatiotemporal aligned data within a unified time window and corresponding performance metric summaries.
[0115] The standardized verification input, as the output field name of this step, is passed to the "Verification Indicators" module in the next step S620 for use in calculating key indicators of the multi-dimensional verification framework. This standardized verification input provides a unified and standardized data foundation for the multi-dimensional verification and optimization output module, promoting the accurate evaluation of simulation results and the formulation of subsequent optimization strategies.
[0116] S620: Based on standardized verification inputs, calculate key parameters and generate a verification evaluation report; Specifically, the standardized verification input from step S610 is used as input. This standardized verification input includes multi-dimensional simulation data and performance indicators that have undergone spatiotemporal alignment processing. The system first loads a preset multi-dimensional verification indicator system, which covers multiple dimensions such as temporal accuracy indicators, spatial consistency indicators, performance stability indicators, and statistical characteristic indicators of simulation results. The indicator definitions are combined with industrial simulation application scenarios and quality control standards to reflect the comprehensive evaluation requirements of simulation results.
[0117] The system calculates indicators for each dimension based on the standardized verification input. The time accuracy indicator assesses the accuracy of the simulation results' time response by calculating the error distribution between the simulation results and reference data (such as historical real data or a high-precision simulation baseline) over time, including mean squared error (MSE) and dynamic time warping (DTW) distance. The spatial consistency indicator evaluates the rationality and stability of the simulation results in the spatial dimension by statistically analyzing the similarity and deviation of spatially distributed data, using spatial correlation coefficients, spatial variograms, and spatial clustering consistency measures. The performance stability indicator combines the fluctuation range and trend analysis of key parameters during the simulation process, using analysis of variance and spectral density estimation methods to reflect the stability and robustness of the simulation results.
[0118] During the index calculation process, the system performs outlier detection and fault tolerance handling. For outlier data points in the index calculation, neighborhood interpolation and weighted averaging methods are used for correction, and anomalies are recorded in the verification log. The multi-dimensional verification framework supports distributed computing and parallel processing, meeting the needs of efficient analysis of large-scale simulation data. After calculation, the system summarizes the results of each index, combines the weight allocation strategy and the multi-index comprehensive evaluation model, and generates a structured verification evaluation report. The verification evaluation report includes detailed index values, trend analysis, anomaly warnings, and a comprehensive score, stored in a unified data format, supporting subsequent querying and display.
[0119] The verification and evaluation report, as an output field name for this step, is passed to the "Optimization Strategy" module in the next step S630, allowing it to adjust the feature fusion parameters and simulation configuration based on the evaluation results. This verification and evaluation report provides comprehensive simulation result quality feedback to the multidimensional verification and optimization output module, promoting closed-loop control and continuous improvement of the simulation optimization process.
[0120] S630. Adjust the feature fusion parameters according to the verification and evaluation report to generate optimized configuration parameters; Specifically, the verification and evaluation report from step S620 is used as input. This report includes the calculation results of multi-dimensional verification indicators and comprehensive evaluation information. The system first analyzes the key indicators in the report and, based on preset optimization rules and strategies, determines the adjustment direction and magnitude of the feature fusion parameters. These feature fusion parameters include the adjustment range of weighting coefficients, dynamic screening thresholds, hyperparameters of the fusion model, and adjustment factors for attention weights, reflecting the key control variables in the multimodal feature fusion process.
[0121] The system combines historical optimization data and machine learning models, employing a parameter adjustment algorithm based on gradient estimation and heuristic search to dynamically generate optimized configuration parameters. Specifically, the system models the response curves between validation metrics and feature fusion parameters, iteratively updating the parameter combinations using Bayesian optimization to balance simulation accuracy and computational complexity. During adjustment, the system performs parameter boundary checks and constraint verification to prevent parameters from exceeding reasonable ranges or causing instability in the fusion process. Abnormal parameter adjustments trigger a rollback mechanism, restoring the previous stable configuration and recording the adjustment anomaly log.
[0122] Furthermore, the system matches and verifies the optimized configuration parameters with the current multimodal feature library and cross-modal attention weight matrix to ensure consistency between parameter adjustments and feature structures. The optimized configuration parameters are distributed to the feature fusion module via a unified interface module, specifically to the "feature input" module in step S310, serving as one of the inputs for subsequent cross-modal attention weighting calculations, thus achieving optimization closure. The system performs version management and change tracking of the optimized configuration parameter generation process, recording the time, basis indicators, and adjustment results of parameter adjustments to support subsequent auditing and performance analysis.
[0123] The optimized configuration parameters, as the output field names of this step, are passed to the "Feature Input" module of S310 for use by the cross-modal attention weighting calculation module, promoting the dynamic adjustment and optimization of the multimodal feature fusion strategy. These optimized configuration parameters achieve an effective mapping from multidimensional verification results to feature fusion parameters, constructing a feedback loop in the industrial simulation optimization process.
Claims
1. A multimodal feature fusion method for industrial simulation optimization, characterized in that, include: Acquire multi-source industrial data, standardize it through data synchronization and timestamp alignment strategies, and use modality-specific algorithms such as time-frequency analysis, CNN or NLP for feature extraction and vectorization to generate feature mapping relationships expressed in graph structure or matrix form; Modal features are obtained from feature mapping relationships, normalized using minimum-maximum normalization or mean-variance standardization methods, and dimensionality reduction and model building based on CNN, RNN or Transformer architectures are performed using PCA, LDA or autoencoder algorithms to generate a multimodal feature library. Feature inputs are obtained from a multimodal feature library. An attention mechanism is constructed based on the Transformer query mechanism or an improved Shannon entropy model. A strategy combining gradient descent and reinforcement learning is used to optimize the parameters and generate a cross-modal attention weight matrix. Based on the attention weight matrix and multimodal feature library, we use correlation evaluation based on GNN and spectral clustering, dynamic threshold screening based on Bayesian optimization, and fusion method based on weighted and MLP nonlinear transformation to perform feature correlation evaluation, screening and fusion operations to generate fused feature vectors. Based on the fused feature vector, the simulation parameters are transformed according to the simulation parameter mapping rules, a physics engine environment covering multiple physics fields is loaded, and optimization algorithms such as genetic algorithm or particle swarm optimization are executed to generate the original simulation results. Based on the original simulation results, spatiotemporal alignment processing such as timestamp correction and spatial coordinate unification is performed, and verification and evaluation of time accuracy and spatial consistency are carried out. The Bayesian optimization method is used to adjust the feature fusion parameters and generate optimized configuration parameters.
2. The method according to claim 1, characterized in that, The process of generating feature mapping relationships expressed in graph structure or matrix form also includes: Acquire multimodal industrial data including sensor physical quantity data, equipment operation logs, image and video data, and control commands, and perform standardized processing such as missing value imputation, outlier detection and removal, noise filtering, and format unification to obtain preprocessed data; Extract modal features from preprocessed data, and use modality-specific vectorization processing such as time-frequency analysis, wavelet transform, CNN or NLP to generate original feature vectors; The original feature vectors are structured and stored by constructing a multi-dimensional index structure and an association mapping mechanism, generating feature mapping relationships expressed in graph structure or matrix form.
3. The method according to claim 1, characterized in that, Multi-source industrial data includes: Multi-source industrial data refers to the collection of raw data with different physical meanings, data structures, and spatiotemporal characteristics, collected through various heterogeneous channels in industrial simulation and optimization scenarios. Specifically, this includes physical quantity data collected by sensors, equipment operation logs, image and video data, and multimodal industrial data of control commands; Multimodal industrial data encompasses structured, semi-structured, and unstructured data, involving time-series signals, discrete event records, and multi-dimensional spatial information.
4. The method according to claim 1, characterized in that, The process of generating a multimodal feature library also includes: Modal features are obtained from the feature mapping relationship, and normalization processes such as min-max normalization, mean-variance standardization, or distribution transformation are performed to generate standardized features; The standardized features are subjected to PCA, LDA, autoencoder, or nonlinear dimensionality reduction based on graph embedding and manifold learning to generate a simplified feature sequence. Based on simplified feature sequences, construct feature representation models based on CNN, RNN, or Transformer to generate a multimodal feature library.
5. The method according to claim 1, characterized in that, The process of generating the cross-modal attention weight matrix also includes: Feature inputs are obtained from a multimodal feature library. An attention query matrix is constructed based on the query, key, and value mechanism in the Transformer architecture or an improved Shannon entropy model to generate an initial attention distribution. Based on the initial attention distribution, parameter optimization is performed by combining gradient descent, adaptive learning rate, regularization term or reinforcement learning strategy to generate dynamic weight parameters; The dynamic weight parameters are normalized by methods such as max-min normalization, L1 norm normalization, or Softmax normalization to generate a cross-modal attention weight matrix stored in a sparse matrix structure.
6. The method according to claim 1, characterized in that, The process of generating fused feature vectors also includes: By combining the attention weight matrix with a multimodal feature library, feature correlation evaluation based on graph neural networks and spectral clustering algorithms is performed to generate candidate feature sets; By combining the candidate feature set, a dynamic screening threshold based on a hybrid strategy of Bayesian optimization and genetic algorithm is set to generate an effective feature subset; The effective feature subset is subjected to weighted fusion processing based on cross-modal attention weights and MLP nonlinear transformation to generate a fused feature vector.
7. The method according to claim 1, characterized in that, The process of generating the original simulation results also includes: The fused feature vectors are converted into simulation parameter format according to simulation parameter mapping rules to generate the initial simulation configuration. Based on the initial simulation configuration, load the physics engine covering dynamics, thermodynamics, fluid or control system simulation to generate the simulation computing environment; Execute optimization algorithms, including genetic algorithms, particle swarm optimization, or gradient descent, in the simulation computing environment to generate raw simulation results.
8. The method according to claim 1, characterized in that, The process of generating optimized configuration parameters also includes: The original simulation results are subjected to spatiotemporal alignment processing, including timestamp correction, time window segmentation, spatial coordinate unification, and data interpolation completion, to generate standardized verification input; Based on standardized verification inputs, key parameters including time accuracy, spatial consistency and performance stability are calculated, and a verification evaluation report is generated. Based on the verification and evaluation report, adjust the feature fusion parameters, including weighting coefficients, dynamic screening thresholds, fusion model hyperparameters, or attention weight adjustment factors, to generate optimized configuration parameters.
9. The method according to claim 1, characterized in that, The specific steps for generating the cross-modal attention weight matrix include: Valid values are screened to remove abnormal weights and noise interference, and statistical thresholds and anomaly detection algorithms are used to determine the screening boundaries. Based on the characteristics and fusion requirements of industrial simulation tasks, the maximum-minimum normalization, L1 norm normalization, or Softmax normalization method is dynamically selected for normalization processing. The normalized weight parameters are constructed into a cross-modal attention weight matrix stored in a sparse matrix structure, and consistency verification and positive definiteness detection are performed.
10. The method according to claim 1, characterized in that, The specific steps for generating the fused feature vector include: Based on the weight distribution in the cross-modal attention weight matrix, each feature vector in the effective feature subset is linearly weighted. A multilayer perceptron is used to perform a nonlinear transformation on the weighted features in order to capture the complex interaction relationships between features; In the process of linear weighting and nonlinear transformation, L2 regularization and sparse constraints are introduced to prevent overfitting; A multi-head fusion strategy is adopted to process different modal feature subsets in parallel through multiple independent weighted channels, and the fusion results are summarized and normalized. Batch normalization and residual connection techniques are combined during the fusion process; Real-time monitoring of the statistical distribution and information entropy of fused features, with abnormal fluctuations triggering adjustment mechanisms; It supports online updates and incremental fusion to respond to dynamic changes in data and models in industrial simulation environments.
Citation Information
Cited By
Transformer insulation aging assessment method based on full feature quantity fingerprint library
CN122330624A