Method and system for establishing data base model
By using multimodal data acquisition and feature decoupling processing, a data foundation model is established, which solves the problem of incomplete data acquisition in existing technologies. This enables efficient and real-time feature processing and dynamic model updates for complex scenarios, improving the model's adaptability and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CIVIL AVIATION FLIGHT UNIV OF CHINA
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-12
AI Technical Summary
Existing data foundation models struggle to comprehensively cover the multi-dimensional parameters of target objects when facing complex scenarios. They suffer from insufficient real-time data collection, a mix of static and dynamic features, a lack of dynamic adjustment capabilities, low efficiency in utilizing historical data, and inadequate model adaptability and accuracy.
The system acquires basic operating parameters through a multimodal heterogeneous data acquisition interface, performs feature decoupling processing to separate static feature vectors and dynamic feature sequences, constructs a feature topology structure, establishes a virtual mapping model, loads a historical operating mode library, generates a dynamic strategy library, and iteratively updates the model during anomaly monitoring.
It achieves comprehensive data capture of target objects, clearly separates feature relationships, enhances the model's mapping accuracy and adaptability, ensures that the model is synchronized with the actual state, and improves data utilization efficiency and decision accuracy.
Smart Images

Figure CN122020384A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data modeling technology, specifically to a method and system for establishing a data foundation model. Background Technology
[0002] In the current digital transformation process, the data generated by various systems exhibits significant characteristics of being multi-sourced, heterogeneous, and dynamic. As the core architecture supporting efficient system operation and decision analysis, the quality of data infrastructure construction directly affects the mining and utilization of data value. However, existing data infrastructure model building methods are gradually revealing many problems when facing complex scenarios.
[0003] Traditional data acquisition methods often rely on single or limited types of interfaces, making it difficult to cover multi-dimensional parameters such as the physical environment and logical behavior of the target object. This results in incomplete basic data that fails to fully reflect the object's true operational status. Furthermore, the lack of real-time performance in data acquisition often leads to delays, causing subsequent data-driven analysis and decision-making to lag behind actual operational needs.
[0004] In the data feature processing stage, existing methods lack an effective decoupling mechanism for the collected basic parameter set, resulting in a mixture of static and dynamic features that makes it difficult to form a clear feature structure. This masks the correlation between features, making it impossible to construct a stable feature topology with clear hierarchical relationships, thereby affecting the accuracy and reliability of subsequent models.
[0005] Virtual mapping, as a key means of connecting physical entities and digital space, currently relies heavily on static mapping to associate virtual models with actual features, lacking dynamic adjustment capabilities. When the target object's operating state changes, virtual space nodes cannot respond promptly to these dynamic changes, leading to discrepancies between the virtual mapping model and the physical entity's operating state, thus reducing the model's reference value.
[0006] Furthermore, the inefficient utilization of historical operational data is a significant problem. Most methods fail to effectively integrate historical operational patterns with real-time data, lacking a matching-based rule engine activation mechanism, making it difficult for historical experience to effectively guide real-time operational decisions. Simultaneously, the existing model's correction mechanism is rather passive in the face of anomalies, often requiring manual intervention after an anomaly occurs. The lack of automated dynamic strategy libraries leads to slow model iteration and updates, making it difficult to adapt to continuously changing operating environments. These issues collectively constrain the adaptability, accuracy, and real-time performance of data-driven models, failing to meet the high data support requirements of complex systems. Summary of the Invention
[0007] The purpose of this invention is to provide a method and system for establishing a data foundation model to solve the problems mentioned in the background art.
[0008] To achieve the above objectives, the present invention provides a method for establishing a data foundation model, the method comprising: Create a data foundation model ontology, and obtain the basic operating parameter set of the target object in real time through a multimodal heterogeneous data acquisition interface. The basic operating parameter set includes physical environment parameters and logical behavior parameters. Perform feature decoupling processing on the basic operating parameter set to separate static feature vectors and dynamic feature sequences, and construct the feature topology structure of the data base model ontology; A virtual mapping model is established based on the aforementioned feature topology, and the static feature vector and dynamic feature sequence are associated with virtual space nodes through a dynamic topology mapping algorithm; The historical operation mode library is loaded into the virtual mapping model, and the corresponding operation rule engine is activated according to the matching degree between the real-time collected basic operation parameter set and the historical operation mode library. The runtime rule engine generates a dynamic strategy library, which includes data reconstruction strategies, anomaly monitoring thresholds, and model correction instructions. When the set of basic operating parameters triggers the anomaly monitoring threshold, the model correction instruction is invoked to iteratively update the parameters of the virtual mapping model, and the updated mapping relationship is synchronized to the data base model body.
[0009] Preferably, the feature topology structure for constructing the data foundation model ontology includes: Perform spatial dimensionality reduction on the static feature vector to extract key dimensional features and generate dimension identifiers; The dynamic feature sequence is segmented in the time domain, and the fluctuation entropy value in each time period is calculated. The node connection relationships of the feature topology are constructed based on the correlation between the dimension identifier and the fluctuation entropy value.
[0010] Preferably, establishing a virtual mapping model based on the feature topology includes: Generate a spatial topology vector set based on the node connection relationships; Set spatial weight allocation rules based on the pattern matching results in the historical operation mode library; The spatial topology vector set is weighted and fused according to the spatial weight allocation rule to generate a coordinate mapping table of virtual spatial nodes.
[0011] Preferably, the generation of the dynamic strategy library includes the following operations: Perform a neighborhood scan on the coordinate mapping table to identify high-density node clusters and sparse node regions; A data reconstruction strategy is set based on the distribution characteristics of the high-density node cluster; The anomaly monitoring threshold is calculated based on the offset of the sparse node region; The data reconstruction strategy is bound to the anomaly monitoring threshold to generate a model correction instruction set.
[0012] Preferably, the step of calling the model correction instruction to iteratively update the parameters of the virtual mapping model includes the following operations: Extract the offset vector of the sparse node region that triggers the anomaly monitoring threshold; Calculate the directional deviation between the offset vector and the reference vector in the historical operation mode library; The weighting coefficients in the spatial weight allocation rule are adjusted according to the directional deviation value; The coordinate mapping table is regenerated using the adjusted weighting coefficients.
[0013] Preferably, the configuration of the multimodal heterogeneous data acquisition interface includes the following operations: Deploy an environmental sensor array at the physical layer of the target object to collect temperature gradient and vibration spectrum data in real time; Deploy a behavior capture agent at the logic layer to continuously acquire operation command sequences and state transition logs; The temperature gradient, vibration spectrum data, operation command sequence, and status switching log are aligned by timestamp and then merged into the basic operating parameter set.
[0014] Preferably, the feature decoupling process includes the following operations: Physical feature extraction is performed on the temperature gradient and vibration spectrum data to generate a physical feature matrix; The operation instruction sequence and state transition log are analyzed for behavioral patterns to generate a behavioral state transition diagram; The physical feature matrix is mapped to a static feature vector, and the behavioral state transition diagram is transformed into a dynamic feature sequence.
[0015] Preferably, the construction of the historical operating mode library includes the following operations: Collect the physical feature matrix and behavioral state transition diagram of the target object during its historical operation cycle; Cluster analysis was performed on the physical feature matrix to divide it into multiple physical feature clusters; Path mining is performed on the behavioral state transition graph to extract high-frequency state transition chains; The correspondence between the physical feature clusters and the high-frequency state transition chains is stored as a historical operating mode library.
[0016] Preferably, activating the corresponding operation rule engine based on the matching degree between the real-time collected basic operation parameter set and the historical operation mode library includes: Calculate the similarity index between the real-time physical feature matrix and the historical physical feature cluster; Detect the overlap between the real-time behavioral state transition graph and the high-frequency state transition chain; When the similarity index and overlap both meet the preset conditions, the running rule engine bound to the corresponding physical feature cluster is activated.
[0017] Preferably, the present invention further includes a data foundation model building system for implementing the above-described data foundation model building method, the system comprising: The multimodal data acquisition module is deployed in the physical and logical layers of the target object to obtain a set of basic operating parameters; A feature decoupling engine, connected to the multimodal data acquisition module, is used to generate static feature vectors and dynamic feature sequences; The virtual modeling core receives the static feature vector and dynamic feature sequence, and is used to construct the feature topology and virtual mapping model; The strategy library generator loads the historical running mode library and connects to the virtual modeling core to generate dynamic strategy libraries; The model iteration controller receives anomaly monitoring signals from the dynamic policy library and is used to trigger parameter updates of the virtual mapping model. The data synchronization agent connects the model iteration controller and the data base model body to synchronize the updated mapping relationship.
[0018] Compared with the prior art, the beneficial effects of the present invention are: This data foundation model building method acquires the basic operational parameter set of the target object in real time through a multimodal heterogeneous data acquisition interface. This set covers both physical environment parameters and logical behavior parameters, comprehensively capturing various key data during the object's operation. This avoids the limitations of traditional single-interface data acquisition, providing richer and more realistic basic data support for the construction of the data foundation model itself. This comprehensive data acquisition approach enables the model to perceive the object's operational status from multiple dimensions, laying a solid data foundation for subsequent feature processing and model construction.
[0019] In the feature processing stage, feature decoupling is performed on the basic operating parameter set to separate static feature vectors and dynamic feature sequences. These are then used to construct a feature topology, effectively clarifying the relationships between different types of features. Static feature vectors reflect the inherent attributes of an object, while dynamic feature sequences reflect the real-time changing trends of the object. This clear separation makes the logical connections between features more explicit, resulting in a feature topology with stronger structure and stability. This facilitates the accurate construction of the subsequent virtual mapping model and reduces model errors caused by feature confounding.
[0020] When establishing a virtual mapping model based on feature topology, a dynamic topology mapping algorithm is used to associate static feature vectors and dynamic feature sequences with virtual space nodes, achieving a precise correspondence between physical features and virtual nodes. This dynamic association method can adjust the mapping relationship in real time as features change, ensuring that the node distribution in the virtual space remains consistent with the actual feature state. This allows the virtual mapping model to realistically reflect the operation of objects, enhancing the mapping accuracy of the virtual model to physical entities and giving the model better intuitiveness and operability.
[0021] By loading a historical operation pattern library into the virtual mapping model and activating the operation rule engine based on the matching degree between real-time parameters and the historical pattern library, historical operation experience can be fully utilized to guide real-time decision-making. The historical operation pattern library contains the operation rules of objects in different scenarios. The method of activating the rule engine based on the matching degree makes the application of rules more targeted, avoids the problem of blindly applying rules, and makes the generated dynamic strategies more in line with the actual needs of the current operation scenario.
[0022] The dynamic strategy library includes data reconstruction strategies, anomaly monitoring thresholds, and model correction instructions, providing multifaceted support for model operation and maintenance. Data reconstruction strategies can optimize the collected data and improve data quality; anomaly monitoring thresholds provide clear standards for judging whether the object's operating status is normal; and model correction instructions provide specific basis for iterative updates of the model. These three elements work together to form a complete strategy support system, enabling the model to cope with various situations during operation.
[0023] When the set of basic operating parameters triggers the anomaly monitoring threshold, the model correction command is invoked to iteratively update the parameters of the virtual mapping model and synchronize it to the data foundation model itself, thus achieving dynamic optimization of the model. This real-time iteration mechanism can respond promptly to abnormal changes during the object's operation, and by adjusting the model parameters, it keeps the model synchronized with the actual operating state of the object, avoiding the problem of the model becoming out of touch with reality due to long-term operation, and ensuring the effectiveness and reliability of the data foundation model in long-term use. Attached Figure Description
[0024] Figure 1 This is a schematic diagram illustrating the working principle of the data foundation model establishment method described in this invention. Figure 2 A flowchart for establishing a virtual mapping model; Figure 3 A flowchart for iteratively updating the parameters of the virtual mapping model; Figure 4 A flowchart for building a library of historical operating modes. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Please see Figure 1 This invention provides a method and system for establishing a data foundation model, the method comprising: The core architecture is a data foundation model ontology. A multimodal heterogeneous data acquisition interface collects the basic operational parameter set of the target object in real time, including physical environment parameters and logical behavior parameters. These parameters are transmitted to the processing unit via the interface. Feature decoupling processing is performed on the basic operational parameter set, decomposing it into static feature vectors and dynamic feature sequences. Static feature vectors represent stable state attributes, while dynamic feature sequences reflect temporal behavioral changes. A feature topology structure is constructed based on the decoupled features, using nodes and edges to represent the relationships between features. A virtual mapping model is then established based on this feature topology structure. A dynamic topology mapping algorithm associates the static feature vectors and dynamic feature sequences with virtual space nodes, simulating real system behavior in the virtual space. A predefined historical operational mode library is loaded into the virtual mapping model, storing past operational modes. The corresponding operational rule engine is activated based on the pattern matching degree between the real-time collected basic operational parameter set and the historical operational mode library. The operational rule engine analyzes the matching results and generates a dynamic strategy library containing data reconstruction strategies, anomaly monitoring thresholds, and model correction instructions. When the set of basic operating parameters triggers the anomaly monitoring threshold in the dynamic strategy library, a model correction instruction is invoked to iteratively update the parameters of the virtual mapping model. The update process adjusts the model's internal weights and mapping relationships. Finally, the updated mapping relationships are synchronized back to the data base model body via a data synchronization mechanism, achieving real-time alignment between the model and the entity system.
[0027] Example 1: See Figure 2 When constructing the feature topology of the data foundation model ontology, the static feature vectors first undergo dimensionality reduction. This process analyzes the correlation between feature dimensions, identifies and removes redundant or low-contribution dimensions, and retains key dimension features. Each key dimension is assigned a unique identifier, which encodes the feature type, location, and weight information in the overall structure. The dimensionality-reduced feature set forms a compact representation, reducing subsequent computational complexity while retaining the main information of the original data. Dynamic feature sequences employ time windowing techniques to divide continuous time-series data into equal-length time periods. Data points within each time period are evaluated for volatility using entropy calculations, quantifying the randomness and stability of the sequence. Time periods with higher entropy values indicate drastic dynamic changes, while time periods with lower entropy values reflect relatively stable behavioral patterns.
[0028] Dimension identifiers are analyzed in conjunction with fluctuation entropy values to establish correlations between features. A correlation matrix evaluates the statistical relationship between dimension identifiers and entropy values, identifying highly correlated feature combinations. These correlations are mapped to nodes and edges in the feature topology; nodes represent key dimensions or time-period entropy values, and edges represent the strength of dependence or influence between features. Node connections are generated using graph theory algorithms, calculating distances or similarities between nodes and determining whether to establish a connection based on a preset threshold. The weights of edges reflect the strength of the correlation; edges with higher weights indicate stronger dependencies between features.
[0029] The generated feature topology serves as the foundational input to the virtual mapping model. The spatial topology vector set is constructed based on node connectivity, containing node coordinate vectors and connection weights. Node coordinate vectors are mapped to a lower-dimensional space using a dimensionality reduction algorithm, facilitating subsequent visualization or computational processing. Connection weights reflect the degree of interaction between features, influencing the node distribution in the virtual space. Pattern samples from the historical running pattern library are input into the matching module to calculate the similarity between the current topology vector set and historical patterns. Similarity evaluation employs distance metrics or pattern matching algorithms to identify the historical running pattern closest to the current state.
[0030] Based on similarity results, spatial weight allocation rules dynamically adjust the weight factors of nodes in the virtual space. These weight factors influence the positional distribution of nodes in the virtual mapping model; high-weight nodes tend to cluster in the core region, while low-weight nodes are distributed at the edges. The weighted fusion operation multiplies and accumulates the elements of the spatial topological vector set with their corresponding weights to generate a fused feature representation. This representation is input to the coordinate mapping table generation module to calculate the precise coordinates of the nodes in the virtual space. The coordinate mapping table stores the virtual positions and connectivity relationships of the nodes, supporting subsequent model calculations and virtual space simulation.
[0031] The coordinate mapping table of the virtual mapping model is used for dynamic strategy generation and anomaly monitoring. The distribution of nodes in the virtual space reflects the system's operating state; dense regions represent stable or common behavior, while sparse regions may indicate anomalies or rare events. The coordinate mapping table's update mechanism ensures the model can adapt to system changes, improving prediction and monitoring accuracy through iterative optimization. The entire implementation process emphasizes modeling the correlations between features and dynamically adjusting weights, enabling the virtual mapping model to effectively reflect the behavioral patterns of the real system.
[0032] The construction of the feature topology relies not only on the dimensionality analysis of static features but also on the time-varying characteristics of dynamic sequences. Dimensionality reduction of static features reduces data redundancy, while segmentation and entropy calculation of dynamic sequences capture the fluctuation patterns of temporal behavior. The correlation analysis between the two establishes a comprehensive feature representation, supporting subsequent node mapping in the virtual space. The generation of node connections is based on statistical and graph theory methods, ensuring that the topology accurately represents the dependencies between features.
[0033] Dynamic adjustment of spatial weight allocation rules enables the model to adapt to different operating modes. Historical data similarity matching provides a reference benchmark, guiding the optimization of weight factors. Weighted fusion operations integrate multi-dimensional feature information to generate a unified virtual spatial representation. The generation and updating mechanism of the coordinate mapping table ensures that the virtual mapping model remains synchronized with the real system, supporting real-time monitoring and strategy adjustments. The entire process, through feature extraction, correlation modeling, and dynamic mapping, achieves efficient construction and continuous optimization of the data foundation model.
[0034] The establishment of a virtual mapping model involves not only technical implementation but also consideration of computational efficiency and scalability. Dimensionality reduction and segmentation reduce data volume and improve processing speed. Correlation matrices and graph theory algorithms optimize the generation efficiency of node connection relationships. Weighted fusion and coordinate mapping calculations employ parallel or distributed processing to adapt to large-scale data scenarios. The indexing and retrieval mechanism of the historical pattern library ensures rapid matching and supports real-time decision-making. The entire implementation process balances accuracy and performance, enabling the data foundation model to operate stably in practical applications.
[0035] A dynamic update mechanism for the feature topology ensures the model can adapt to system evolution. Continuous input of new data triggers feature decoupling and topology reconstruction, maintaining the model's timeliness. Anomaly monitoring and strategy adjustments are fed back to model updates, forming a closed-loop optimization. A data synchronization mechanism ensures real-time consistency between the virtual mapping model and the data-based model, avoiding information lag or deviation. The entire system achieves adaptive model maintenance through automated processes, reducing the need for manual intervention.
[0036] Example 2: See Figure 3 The generation process of the dynamic strategy library is based on the coordinate mapping table in the virtual mapping model. This table stores the location information and connection relationships of nodes in the virtual space, serving as input data for the neighborhood scanning operation. The scanning algorithm employs density-based clustering to analyze the distribution of nodes in the virtual space. The distance between nodes is calculated using spatial coordinates, identifying areas with high clustering and sparse distribution. The boundaries of high-density node clusters are automatically defined using a density threshold, with nodes within the cluster exhibiting similar characteristic patterns and operational states. Sparse node regions are determined based on the average distance between nodes; areas exceeding a specific threshold are marked as sparsely distributed.
[0037] The distribution characteristics analysis of high-density node clusters includes geometric shape and topological properties. The geometric parameters of the clusters include coverage area, center point location, and boundary morphology; these parameters reflect the system's behavior under steady-state conditions. Based on the cluster distribution characteristics, data reconstruction strategies are dynamically generated. These strategies involve adjusting data sampling frequency, reallocating feature weights, and optimizing computational resource allocation. For highly dense node cluster regions, the system may employ downsampling to reduce computational load while preserving core feature information. Monitoring of sparse node regions focuses on node displacement changes, calculating the offset vector of each node relative to its historical reference position. The calculation of the offset vector comprehensively considers distance and direction components to form a complete displacement description.
[0038] The process of setting the anomaly monitoring threshold incorporates the offset characteristics of sparse node regions. The system analyzes the positional fluctuation range of nodes under historical normal operating conditions to establish benchmark reference values. The comparison between the current offset and the historical benchmark generates a deviation index, which is used to dynamically adjust the anomaly monitoring threshold. The threshold is not a fixed value but adapts to the system's operating status to avoid false alarms or missed alarms. The data reconstruction strategy and the anomaly monitoring threshold are bound together by logical rules to form a model correction instruction set. The instruction set includes conditional judgments and execution operations; when the monitored data meets specific conditions, the corresponding model adjustment command is triggered.
[0039] When an anomaly detection threshold is triggered, the system extracts offset vector data from the relevant sparse node regions. The analysis of the offset vectors includes directional consistency checks and magnitude assessment. The directional deviation value is calculated using a reference vector stored in the historical operating mode library, quantifying the difference between the current offset direction and typical patterns through an angle measurement method. Large directional deviations may indicate abnormal changes in system behavior, requiring close monitoring. Based on the directional deviation value, the weighting coefficients in the spatial weight allocation rules are readjusted. The adjustment process follows a pre-defined optimization algorithm, giving higher weights to dimensions sensitive to abnormal directions, thus enhancing the model's ability to identify abnormal patterns.
[0040] After adjusting the weighting coefficients, the virtual mapping model regenerates the coordinate mapping table. The new coordinate mapping table reflects the adjusted spatial weight allocation, and node positions are recalculated based on the latest weights. This process enables the model to dynamically adapt to changes in system state and promptly capture abnormal behavior characteristics. The regenerated coordinate mapping table synchronizes the updated node distribution information to subsequent processing modules, maintaining the consistency of data across the entire system. The execution of model correction instructions forms a closed-loop feedback mechanism, continuously iterating and optimizing to improve the model's adaptability and monitoring accuracy.
[0041] The coordinate mapping table update mechanism ensures that the virtual mapping model always reflects the latest system state. Each iteration update is based on rigorous data analysis and logical judgment, avoiding model instability caused by arbitrary modifications. The historical operating mode library provides a reliable benchmark reference, making the current state assessment objective. The direction analysis of the offset vector enhances the understanding of system behavior patterns, focusing not only on numerical changes but also on the consistency of the changing trends. The optimized adjustment of the weighting coefficients balances the influence of each feature dimension, preventing any single dimension from excessively dominating the model behavior.
[0042] The parameter settings of the neighborhood scanning algorithm consider the dimension and scale of the virtual space to ensure a balance between scanning accuracy and computational efficiency. The data reconstruction strategy generation algorithm needs to be compatible with different types of data features, maintaining the strategy's universality and specificity. The adaptive adjustment algorithm for the anomaly detection threshold needs stable convergence characteristics to avoid misjudgments caused by threshold oscillations. The logical design of the model correction instructions needs to cover various possible anomaly scenarios while maintaining a concise and efficient instruction set. The update mechanism of the spatial weight allocation rules requires mathematical rigor to ensure that the adjusted weight distribution is reasonable and effective.
[0043] The distribution changes of virtual space nodes are continuously monitored, serving as the basis for system behavior analysis. The effectiveness of the data reconstruction strategy is verified through subsequent monitoring data, forming a feedback loop for strategy optimization. Adaptive adjustments to anomaly monitoring thresholds are based on comparative analysis of historical data and the current state, achieving intelligent evolution of the thresholds. The execution results of model correction instructions are evaluated through a new round of neighborhood scanning to confirm the update effect and guide subsequent adjustments. This closed-loop operation mode enables the system to continuously learn and self-optimize, gradually improving its adaptability to complex operating environments.
[0044] Raw monitoring data undergoes rigorous verification and cleaning to ensure input quality. Intermediate calculation results employ redundant verification to prevent error accumulation during processing. Model update operations are logged comprehensively, supporting issue tracing and analysis. Key parameter adjustments retain version history, supporting rollback operations when necessary. The entire system's data flow design incorporates fault tolerance, ensuring basic functionality even when some data is abnormal. These mechanisms collectively form the foundation of the system's robustness, ensuring long-term stable operation.
[0045] Example 3: The configuration of the multimodal heterogeneous data acquisition interface involves the coordinated deployment of the physical and logical layers. At the physical layer of the target object, the environmental sensor group includes a temperature sensor array and a vibration sensing unit, continuously acquiring environmental parameters at a fixed sampling frequency. Temperature gradient data records the rate of temperature change at different spatial locations, forming a time-series temperature field distribution. Vibration spectrum data is acquired through an accelerometer and processed by a Fast Fourier Transform to obtain the frequency domain energy distribution. The deployment locations of the physical layer sensors are optimized to cover key monitoring areas and avoid data acquisition blind spots. The network topology of the sensor nodes adopts a redundant configuration to ensure that basic monitoring functions are maintained even if some nodes fail.
[0046] The behavior capture agent in the logic layer is embedded in the system instruction processing pipeline to intercept the operation instruction stream in real time. The operation instruction sequence records the command code and parameters executed by the system, preserving complete timing relationships. The state transition log tracks the changes in the system's internal state machine, recording state transition events and their triggering conditions. The behavior capture agent adopts a non-intrusive design, without affecting the normal operation of the original system. Accurate timestamps are appended to the collected logical data, achieving millisecond-level precision, providing a benchmark for multi-source data alignment. The timestamp synchronization mechanism uses network time protocol correction to eliminate clock deviations between different acquisition nodes.
[0047] Time alignment processing for multimodal data establishes a unified time reference system. The alignment algorithm identifies key time markers in each data stream and compensates for data point mismatches caused by different sampling rates through interpolation. The aligned data is merged into a basic runtime parameter set according to time windows, maintaining temporal consistency between the physical and logical layers. The structured storage of the basic runtime parameter set employs a hierarchical design, storing original sampled data and derived features separately for easy access by subsequent processing modules. A data caching mechanism balances real-time and integrity requirements, maintaining stable output under bursty traffic conditions.
[0048] Feature decoupling processing executes parallel analysis paths on the basic operating parameter set. Temperature gradient and vibration spectrum data are input into the physical feature extraction process. The spatial variation pattern of the temperature field is calculated using the grid difference method to identify heat conduction characteristics. The energy concentration band of the vibration signal is determined through spectral analysis to extract the dominant vibration mode. The physical feature matrix is constructed using the following normalization process:
[0049] in, This represents the elements of the standardized physical feature matrix. These are the original eigenvalues. It is the mean of all samples in this feature dimension. This corresponds to the standard deviation. Row index of the matrix. Corresponding time point, column index Different physical characteristic types are identified. Standardization eliminates the influence of different physical dimensions, ensuring that subsequent analysis is not affected by differences in numerical scale.
[0050] The process for parsing operation command sequences and state transition log input behavior patterns is as follows: Syntactic analysis of command sequences identifies command structure and parameter patterns, establishing an operation semantic graph. The state transition graph is constructed using a probabilistic finite state automaton model, where nodes represent system states and edges are labeled with transition conditions and probabilities. An optimization algorithm for the behavior state transition graph merges equivalent states, simplifying the graph structure while preserving key transition paths. The parsing process retains statistical information on state dwell time, reflecting the system's behavioral inertia under different states.
[0051] The generation of static feature vectors is achieved through dimensionality compression of the physical feature matrix. Principal component analysis (PCA) identifies key linear combinations in the feature matrix, generating low-dimensional representation vectors. Each dimension of the vector corresponds to the main direction of change of the original feature, and the dimension value reflects the projection intensity of the current state in that direction. The update cycle of the static feature vectors is synchronized with the generation of the physical feature matrix to ensure temporal consistency. The transformation of dynamic feature sequences is based on path traversal of the behavioral state transition graph. A depth-first search algorithm extracts typical transition sequences, and the sequence length is adaptively determined according to the state complexity. Sequence encoding uses permutations and combinations of state identifiers to preserve the temporal logic of state transitions.
[0052] A cross-modal feature mapping is established through correlation analysis between the physical feature matrix and the behavioral state transition diagram. A correspondence is established between the matrix row vectors and the nodes of the state diagram, reflecting the joint distribution of physical features and system states at a specific time point. A similarity metric is used to establish the mapping relationship, building a bridge between the feature space and the state space. This cross-modal correlation enhances the overall understanding of system behavior and provides richer feature inputs for subsequent virtual modeling.
[0053] The reliability design of the data acquisition system is reflected in multiple layers. The physical layer sensors have self-diagnostic capabilities, periodically reporting their operational status and automatically calibrating. Data from abnormal sensor nodes is marked and excluded from subsequent processing. The logical layer capture agent performs data integrity verification, checking and validating the correctness of data packets. The data transmission channel employs encryption and redundant coding to prevent data loss or tampering. The acquisition system's resource management module dynamically adjusts the sampling frequency and data accuracy to adapt to different operating load conditions.
[0054] The quality control mechanism of the feature decoupling process ensures the reliability of the output features. During the physical feature extraction stage, a signal quality threshold is set; low signal-to-noise ratio data segments trigger re-acquisition or special labeling. In the behavior pattern parsing stage, syntax verification is implemented; illegal instruction sequences trigger an exception handling process. The generation process of the feature matrix and state transition diagram records detailed metadata, including data processing parameters and algorithm version, supporting the repeatability verification of the results. The exception data processing path is separated from the normal process to avoid contaminating the mainstream feature output.
[0055] The time alignment accuracy of multimodal data directly affects the effectiveness of subsequent analysis. Timestamp correction algorithms compensate for clock drift from different acquisition devices, maintaining sub-millisecond synchronization accuracy. Data interpolation methods select appropriate interpolation kernel functions based on signal characteristics, balancing computational complexity and interpolation accuracy. Aligned data undergoes a time-domain consistency check to ensure that the physical meaning of different modal data corresponds at the same time point. A sliding time window mechanism allows for partial overlap, preventing feature truncation caused by critical events falling at the window edges.
[0056] Standardization of the physical feature matrix enhances data comparability under different operating conditions. A moving window statistical method calculates the mean and standard deviation, adapting to slow changes in system parameters. The update strategy for standardized parameters considers the decay weighting of historical data, giving higher importance to new observations. Outlier detection is performed on outlier points before inclusion in statistical calculations to prevent extreme values from distorting the standardized baseline. The scaling range of matrix elements is reasonably limited to avoid numerical overflow or precision loss.
[0057] The probabilistic annotations in the behavioral state transition graph reflect the statistical regularities of system operation. Probability estimation employs smoothing techniques to handle sparse data, avoiding path breaks caused by zero probabilities. The logical expressions of state transition conditions are standardized to unify the descriptions of conditions from different sources. The graph's visualization layout algorithm optimizes node arrangement, making high-frequency transition paths more visually apparent. Graph version management records the history of structural changes, supporting comparative analysis of behavioral patterns across different periods.
[0058] The selection of dimensions for static feature vectors balances information preservation and computational efficiency. The eigenvalue decay curve from principal component analysis guides the determination of the number of dimensions, preserving sufficient information while controlling vector size. Sparse vector representation techniques are employed in specific scenarios to further compress feature storage space. The static feature update triggering mechanism supports both event-driven and timed polling modes to adapt to different application requirements. The similarity measurement function for feature vectors is configurable, allowing selection of an appropriate distance calculation method based on the specific task.
[0059] The algorithm for generating dynamic feature sequences considers the complexity of states in actual operation. A sliding window mechanism captures behavioral patterns at different time scales. The sequence encoding scheme is designed to balance expressive power and storage efficiency, employing variable-length encoding to handle state paths of varying complexity. The sequence matching algorithm implements fault-tolerant comparison, allowing for a certain degree of state transitions or omissions. Dynamic sequence timeliness management automatically eliminates outdated patterns, maintaining the timeliness of the sequence library.
[0060] Cross-modal feature mapping is implemented using a supervised learning method. The labeled dataset is constructed based on known system operating scenarios, covering typical states and abnormal situations. Regularization techniques are employed during the training process of the mapping model to prevent overfitting and achieve better generalization ability with limited labeled data. The interpretability of the mapping results is enhanced through feature importance analysis, identifying the physical feature dimensions most influential on state judgments. An online learning mechanism enables the mapping model to adapt to gradual changes in system behavior.
[0061] Example 4: See Figure 4 The construction process of the historical operation pattern database is illustrated using an industrial centrifugal compressor system as a specific example. Over a three-month continuous operation cycle, the system's physical feature matrix recorded 12 key parameters, including bearing temperature, rotor vibration, and lubricating oil pressure, with data collected every minute to form time-series data. The behavioral state transition diagram recorded the compressor's switching operations between different speed levels, as well as abnormal protection trigger events. After cleaning and preprocessing, the raw data formed a standardized historical dataset for pattern mining.
[0062] Cluster analysis of the physical feature matrix employs an improved k-medoids algorithm to process high-dimensional time series data. The algorithm first determines the optimal number of clusters using silhouette coefficients, which is set to 5 physical feature clusters in this compressor case. The centroid vector of each feature cluster represents the typical operating state characteristics of that category. Refer to Table 1, which shows partial parameters of the centroid vectors for three typical feature clusters; the values have been normalized.
[0063] Table 1: Partial parameters of the centroid vectors of three typical feature clusters.
[0064]
[0065] Path mining of the behavioral state transition graph focuses on the state transition sequences in the compressor control logic. High-frequency state transition chains extracted from the operation log show that during normal operation, the system mainly follows a gradual switching path of "start-up -> low speed -> medium speed -> high speed," accounting for 83% of the total transitions. Abnormal state transitions mainly exhibit two patterns: the "high speed -> alarm -> shutdown" sequence triggered by insufficient oil pressure accounts for 67% of abnormal situations, while the "medium speed -> speed reduction -> inspection" sequence caused by excessive vibration accounts for 29%. The extraction process of the state transition chains uses a prefix tree structure to efficiently statistically analyze sequence frequencies and sets a minimum support threshold to filter out random transitions.
[0066] Association rule mining between physical feature clusters and high-frequency state transition chains revealed stable correspondences. Feature cluster C-1 primarily associates with asymptotic transition chains from low to medium speed, feature cluster C-2 corresponds to the standard transition from medium to high speed, and feature cluster C-3 is highly correlated with two abnormal transition modes. These associations were verified using confidence metrics, with the rule "If a physical feature belongs to class C-3, the probability of a subsequent abnormal oil pressure transition chain" reaching 0.79. The association rules are stored in a key-value database, with the physical feature cluster ID as the key and the associated state transition chain and its confidence level as the value, supporting fast retrieval.
[0067] The incremental update mechanism of the historical operation pattern library is executed weekly in this compressor system. Newly collected operation data undergoes the same feature extraction and clustering process, and is matched with existing feature clusters for similarity. When the distance between a newly added feature point and the centroid of an existing cluster exceeds a threshold, a new cluster generation or cluster split operation is triggered. The statistics of behavioral state transition chains are also updated periodically, dynamically adjusting the composition and order of high-frequency chains. The pattern library version control system records the content changes of each update, allowing rollback to historical versions when necessary.
[0068] Actual operating data from the compressor system shows that the distribution of physical feature clusters changes regularly over time. The frequency of feature cluster C-2 is 18% higher in spring than in winter, which is related to differences in cooling efficiency caused by changes in ambient temperature. The seasonal variation in the state transition chain is that the transition time from "low speed to medium speed" is on average 23 seconds longer in winter than in summer. These temporal patterns are supplemented and recorded in the metadata of the pattern library, enhancing the context-awareness of state judgment.
[0069] The pattern library's query interface enables multi-level access. Basic queries return the most matching feature clusters and associated state chains based on physical feature vectors. Advanced queries support combined conditions, such as filtering state transition patterns within a specific time period or under the control of a specific operator. The query results are displayed using a visual design, with feature cluster distribution presented as a 3D scatter plot and state transition chains displayed as a directed graph, facilitating an intuitive understanding of the system's behavioral patterns.
[0070] The learning mechanism for abnormal patterns demonstrated particular value in the compressor case. When the system first exhibited a novel transition sequence of "high speed -> abnormal vibration -> deceleration -> successful secondary acceleration," the anomaly detection module in the pattern library marked it as a pattern to be verified. After seven repetitions over three weeks, the sequence was formally incorporated into the abnormal pattern library and associated with precursors to changes in physical characteristics. This progressive learning approach avoids prematurely solidifying random events into patterns while simultaneously capturing the real evolution of system behavior in a timely manner.
[0071] The boundary management of physical feature clusters employs fuzzy set theory. Approximately 15% of the cases in the compressor operation data lie within the boundary regions of feature clusters; these cases are assigned membership degrees to multiple clusters. State transition prediction comprehensively considers the association rules of relevant clusters, using weighted voting to arrive at the final judgment. This approach improves the sensitivity of identifying transitional states and the initial stages of anomalies.
[0072] The distributed storage design of the pattern library is suitable for large-scale industrial scenarios. Multiple parallel units of the compressor unit share the core pattern library, while each unit retains its local operating mode. The central node periodically synchronizes newly discovered patterns from each unit, and updates the global library after consistency verification. This architecture ensures both the integrity of the pattern library and respects the uniqueness of individual devices.
[0073] The compressor maintenance personnel's interface integrates real-time prompts from the pattern library. When system operating characteristics begin to deviate from the associated typical state transition chain, the interface displays warning prompts and suggested inspection items. After the maintenance personnel confirm the handling results, the feedback information is recorded and used to improve the case library of the pattern library. This interactive mechanism forms a closed loop for continuous optimization of the pattern library.
[0074] Historical data retrospective analysis helps understand the evolution of system behavior. By comparing the distribution of feature clusters before and after a compressor overhaul, it was found that bearing replacement reduced the frequency of feature cluster C-3 by 40%. This analysis provides a quantitative basis for preventative maintenance and helps determine the remaining service life of critical components. The time series analysis module of the pattern library automatically detects these long-term trend changes and generates equipment health status evolution reports.
[0075] In the process of upgrading compressor control systems, the model library plays an important reference role. Before the new control algorithm is introduced, its potential state transition chains are simulated and tested in the model library to assess compatibility with historical safety models. After actual deployment, the new model library specifically records the characteristic changes during the transition period between the old and new systems, establishing a comparative reference benchmark. This application approach reduces the operational risks brought about by system upgrades.
[0076] The schema library's access control system differentiates between different levels of access requirements. On-site operators can only query basic status association rules, maintenance engineers can view details of abnormal patterns, and system designers have the authority to modify core clustering algorithm parameters. Audit logs record all modifications to the schema library, ensuring the traceability of changes. This security management protects core intellectual assets without hindering normal use.
[0077] In compressor unit energy efficiency optimization projects, the pattern library helps identify the optimal operating range. Analysis of energy consumption indicators corresponding to different feature clusters revealed that the energy consumption per unit output in feature cluster C-2 was 12% lower than the average. Based on this, the control strategy was adjusted to keep the system operating as close to C-2 as possible, thus improving energy efficiency. This type of derivative application of the pattern library expands its value scope.
[0078] In fault diagnosis scenarios, the pattern library supports case-based reasoning. When a new type of anomaly occurs in the system, the most similar physical characteristic change history is retrieved from the pattern library, and the corresponding handling solution is referenced. Newly accumulated experience during the diagnosis process is fed back into the pattern library, forming a virtuous cycle of knowledge accumulation. This application method significantly shortens fault diagnosis time and improves processing accuracy.
[0079] Automated testing of compressor systems utilizes a pattern library to generate test cases. Typical operational sequences are generated based on historical state transition chains, with special coverage of all abnormal transition paths. Test results are automatically compared to the expected behavior in the pattern library, and any deviations are immediately flagged for review. This testing method provides a more comprehensive reflection of the system's actual operating condition than traditional boundary value testing.
[0080] The schema library's maintenance tools include a data quality monitoring module. It periodically scans the feature clusters and state chains in the library to detect potential data anomalies, such as outlier feature clusters or low-frequency state chains being mislabeled as high-frequency ones. Upon detecting a problem, it initiates a data repair process, including re-clustering or verifying association rules. This autonomous function maintains the schema library's inherent consistency and reduces manual maintenance workload.
[0081] Example 5: The matching process between the real-time acquired basic operating parameter set and the historical operating mode library adopts a multi-dimensional parallel analysis strategy. The similarity calculation of the physical feature matrix first performs the same standardization preprocessing on the real-time data stream as on historical data to eliminate the effects of dimensional differences and baseline drift. The distance measurement between the row vectors of the matrix and the centroid vectors of historical physical feature clusters uses an improved similarity algorithm. This algorithm assigns dynamic weights to each feature dimension, reflecting the changes in the importance of each parameter at different operating stages. The numerical range of the similarity index is normalized to a uniform scale between 0 and 1, facilitating the setting of general threshold conditions.
[0082] The overlap detection of behavioral state transition graphs employs a hybrid approach combining graph isomorphism and path matching. The real-time generated state transition graph is first topologically simplified, merging functionally equivalent intermediate state nodes while retaining key decision points. The simplified graph structure is then compared layer-by-layer with historical high-frequency state transition chains, recursively detecting matching sub-paths starting from the initial state. The overlap score comprehensively considers the consistency of the state sequence and the similarity of transition conditions, using a weighted average to process matching results at different levels. The matching of path branch points has a higher influence on the final score, reflecting the degree of consistency in the system's decision-making logic.
[0083] The verification logic for the preset conditions adopts a hierarchical judgment structure. The physical feature similarity index first filters out obviously mismatched historical patterns through a primary threshold. Candidate patterns that pass the initial screening proceed to the second layer of verification, checking the overlap between their associated behavioral state transition chains and the real-time state graph. A tolerance mechanism is incorporated into the dual-condition judgment, allowing individual indicators to fluctuate within a specific range, as long as the overall score meets the activation criteria. This design avoids misjudgments due to transient anomalies in a single parameter, improving the robustness of the rule engine's activation decisions.
[0084] The activation process of the running rule engines achieves a smooth transition. When real-time data matches multiple historical patterns, the system allocates activation weights based on the matching score, enabling multiple rule engines to work collaboratively. Newly activated engines initially run in observer mode, and their output suggestions are only gradually taken over control after verification. The currently executing engine continuously monitors changes in matching score, and when its score falls below a maintenance threshold, it initiates an exit process, transferring its decision weights to a more matching engine instance. This gradual switching mechanism ensures a smooth transition when the system's control logic changes, avoiding disturbances caused by sudden switching.
[0085] The runtime rule engine, bound to physical feature clusters, adopts a modular design. Each engine instance encapsulates the control strategies and decision-making logic for a specific operating mode, including parameter adjustment rules, exception handling procedures, and optimization target settings. The engine's input interface receives real-time feature data and status information, while its output interface generates control commands and warning signals. The internal logic implementation uses extensible rule templates, supporting online modification and strategy optimization. Engine instances exchange context information through shared memory to maintain the consistency of control decisions.
[0086] The real-time performance of matching evaluation is guaranteed by a streaming computing framework. Data stream processing nodes are deployed in a computing cluster, employing a micro-batch processing mode to balance processing latency and throughput requirements. Similarity calculation and overlap detection tasks are assigned to different computing units for parallel execution, and the result aggregation module combines the outputs of each unit to generate the final score. A caching mechanism stores intermediate calculation results to avoid repeatedly processing the same data segments. A system resource dynamic allocation module adjusts the load on computing nodes based on data traffic to maintain stable processing performance.
[0087] The matching logic under abnormal conditions has a special processing flow. When the similarity between real-time features and all historical physical feature clusters is below a threshold, the system initiates the abnormal pattern recognition process. This process relaxes the matching requirements of the state transition chain, focusing on detecting the feature change trend and state deviation path in the early stages of the anomaly. Partially matched abnormal patterns are granted temporary activation permissions, with control strategies emphasizing security protection and fault prediction. Simultaneously, the system records the detailed evolution process of abnormal features, providing material for subsequent new pattern learning.
[0088] The cumulative effect of matching degree over time is incorporated into activation decisions. The system maintains a matching degree trajectory within a sliding time window, identifying trends of continuous improvement or deterioration in matching. Short-term fluctuations do not immediately trigger engine switching, while patterns of continuous improvement in matching degree gain a progressively stronger influence. The trend analysis algorithm distinguishes between normal fluctuations and substantial changes, avoiding overreaction to temporary fluctuations. This time-context-based decision-making mechanism enhances the system's ability to identify changes in its operational state.
[0089] Version management of the historical pattern library affects the matching process. When the pattern library is updated, the system automatically reassesses the current running state and its compatibility with the new version. During the version switching transition, a hybrid matching strategy is used, simultaneously calculating the match degree with both the old and new version patterns, and gradually migrating to the new version library. Major version updates trigger a system self-check process to verify the compatibility of the core matching algorithm and adjust parameter settings if necessary. A version rollback mechanism ensures that a previously stable pattern library state can be quickly restored in the event of matching anomalies.
[0090] Environmental context information enhances matching accuracy. The system integrates external parameters such as ambient temperature and load requirements as correction factors for matching degree calculation. The same physical characteristics may correspond to different optimal operating modes under different environmental conditions; the correction factor quantifies this environmental dependence. The context-aware module continuously tracks changes in external parameters and dynamically adjusts the sensitivity parameters of the matching algorithm to ensure that pattern recognition remains consistent with environmental conditions.
[0091] The operator feedback mechanism refines the matching results. The control interface provides a visual display of matching accuracy information, allowing experienced operators to confirm or correct automatically identified results. Human feedback signals are used to calibrate the internal parameters of the matching algorithm, reducing systematic bias. Disputed cases automatically trigger detailed logging for subsequent algorithm optimization and analysis. This human-machine collaborative mechanism overcomes the limitations of purely automatic matching and improves the reliability of critical decisions.
[0092] The audit trail for the matching process achieves full-cycle recording. The system records in detail the input data, intermediate results, and final decision for each matching degree calculation, including information such as timestamps, calculation nodes, and algorithm versions. Audit logs support post-event analysis and troubleshooting, reproducing the matching decision process at a specific moment. The log compression and archiving strategy balances storage costs and traceability requirements, ensuring that complete data from key decision points is stored long-term.
[0093] Cross-subsystem matching and coordination handles complex scenarios. For devices containing multiple interactive subsystems, the matching results of each subsystem are input into the coordination and arbitration module. This module analyzes the matching consistency between subsystems, resolves cross-system conflicts, and generates a globally optimal control strategy. The coordination mechanism considers the physical coupling and functional dependencies between subsystems, prioritizing the matching needs of critical subsystems. This holistic optimization perspective avoids suboptimal global decisions caused by local matching.
[0094] The matching algorithm's online learning capability continuously optimizes its performance. The system automatically collects actual performance data on matching decisions and adjusts algorithm parameters using a reinforcement learning framework. Performance evaluation metrics include multi-dimensional objectives such as control stability, energy efficiency, and anomaly detection rate. The learning process employs a small-step, incremental strategy, controlling the magnitude of each adjustment to ensure smooth system behavior evolution. Abnormal performance fluctuations trigger a learning pause mechanism, leading to manual analysis and intervention.
[0095] The real-time matching system is resource-aware and adaptable. When computing resources are limited, the system automatically simplifies the matching algorithm process, focusing on key feature comparisons. The resource monitoring module predicts peak processing loads and pre-allocates backup computing power. This flexible design ensures that basic matching functions can be maintained even in the event of hardware failure or sudden load increases, preventing system control interruptions.
[0096] User-defined matching strategies meet specific needs. The system provides an open interface for configuring matching strategies, allowing experienced users to adjust matching weights and thresholds for specific application scenarios. Customized strategies are stored in an independent library after security verification, without affecting the standard matching process. Strategy versions are bound to device serial numbers to ensure that dedicated matching logic for a specific device is not misused on other devices. This flexibility expands the system's applicability in special operating conditions.
[0097] The uncertainty of matching results is quantified to aid decision-making. In addition to matching score, the system calculates the confidence interval and risk estimate for the matching decision. Low-confidence matches trigger additional verification processes, and high-risk decisions require multiple levels of approval. Uncertainty analysis considers factors such as input data quality, algorithm limitations, and pattern library coverage, providing a multi-dimensional assessment of decision reliability. This transparent risk assessment mechanism enhances the rational use of automatic matching results.
[0098] Cross-timescale pattern matching enables comprehensive monitoring. The system simultaneously runs matching processes at different time granularities: second-level matching for real-time control, minute-level matching to track slow-changing trends, and hour-level matching to identify long-term patterns. Multi-scale matching results are fused using algorithms to generate a comprehensive judgment, capturing the full spectrum of features from transient fluctuations to long-term drift. This design allows the system to respond quickly to sudden changes while adapting to gradual evolution.
[0099] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0100] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for establishing a data foundation model, characterized in that, Includes the following steps: Create a data foundation model ontology, and obtain the basic operating parameter set of the target object in real time through a multimodal heterogeneous data acquisition interface. The basic operating parameter set includes physical environment parameters and logical behavior parameters. Perform feature decoupling processing on the basic operating parameter set to separate static feature vectors and dynamic feature sequences, and construct the feature topology structure of the data base model ontology; A virtual mapping model is established based on the aforementioned feature topology, and the static feature vector and dynamic feature sequence are associated with virtual space nodes through a dynamic topology mapping algorithm; The historical operation mode library is loaded into the virtual mapping model, and the corresponding operation rule engine is activated according to the matching degree between the real-time collected basic operation parameter set and the historical operation mode library. The runtime rule engine generates a dynamic strategy library, which includes data reconstruction strategies, anomaly monitoring thresholds, and model correction instructions. When the set of basic operating parameters triggers the anomaly monitoring threshold, the model correction instruction is invoked to iteratively update the parameters of the virtual mapping model, and the updated mapping relationship is synchronized to the data base model body.
2. The data foundation model establishment method as described in claim 1, characterized in that, The feature topology structure for constructing the data foundation model ontology includes: Perform spatial dimensionality reduction on the static feature vector to extract key dimensional features and generate dimension identifiers; The dynamic feature sequence is segmented in the time domain, and the fluctuation entropy value in each time period is calculated. The node connection relationships of the feature topology are constructed based on the correlation between the dimension identifier and the fluctuation entropy value.
3. The data foundation model establishment method as described in claim 2, characterized in that, The establishment of the virtual mapping model based on the feature topology includes: Generate a spatial topology vector set based on the node connection relationships; Set spatial weight allocation rules based on the pattern matching results in the historical operation mode library; The spatial topology vector set is weighted and fused according to the spatial weight allocation rule to generate a coordinate mapping table of virtual spatial nodes.
4. The data foundation model establishment method as described in claim 3, characterized in that, The generation of the dynamic strategy library includes the following operations: Perform a neighborhood scan on the coordinate mapping table to identify high-density node clusters and sparse node regions; A data reconstruction strategy is set based on the distribution characteristics of the high-density node cluster; The anomaly monitoring threshold is calculated based on the offset of the sparse node region; The data reconstruction strategy is bound to the anomaly monitoring threshold to generate a model correction instruction set.
5. The data foundation model establishment method as described in claim 4, characterized in that, The invocation of the model correction instruction to iteratively update the parameters of the virtual mapping model includes the following operations: Extract the offset vector of the sparse node region that triggers the anomaly monitoring threshold; Calculate the directional deviation between the offset vector and the reference vector in the historical operation mode library; The weighting coefficients in the spatial weight allocation rule are adjusted according to the directional deviation value; The coordinate mapping table is regenerated using the adjusted weighting coefficients.
6. The data foundation model establishment method as described in claim 1, characterized in that, The configuration of the multimodal heterogeneous data acquisition interface includes the following operations: Deploy an environmental sensor array at the physical layer of the target object to collect temperature gradient and vibration spectrum data in real time; Deploy a behavior capture agent at the logic layer to continuously acquire operation command sequences and state transition logs; The temperature gradient, vibration spectrum data, operation command sequence, and status switching log are aligned by timestamp and then merged into the basic operating parameter set.
7. The data foundation model establishment method as described in claim 6, characterized in that, The feature decoupling process includes the following operations: Physical feature extraction is performed on the temperature gradient and vibration spectrum data to generate a physical feature matrix; The operation instruction sequence and state transition log are analyzed for behavioral patterns to generate a behavioral state transition diagram; The physical feature matrix is mapped to a static feature vector, and the behavioral state transition diagram is transformed into a dynamic feature sequence.
8. The data foundation model establishment method as described in claim 7, characterized in that, The construction of the historical operation mode library includes the following operations: Collect the physical feature matrix and behavioral state transition diagram of the target object during its historical operation cycle; Cluster analysis was performed on the physical feature matrix to divide it into multiple physical feature clusters; Path mining is performed on the behavioral state transition graph to extract high-frequency state transition chains; The correspondence between the physical feature clusters and the high-frequency state transition chains is stored as a historical operating mode library.
9. The data foundation model establishment method as described in claim 8, characterized in that, The step of activating the corresponding operation rule engine based on the matching degree between the real-time collected basic operation parameter set and the historical operation mode library includes: Calculate the similarity index between the real-time physical feature matrix and the historical physical feature cluster; Detect the overlap between the real-time behavioral state transition graph and the high-frequency state transition chain; When the similarity index and overlap both meet the preset conditions, the running rule engine bound to the corresponding physical feature cluster is activated.
10. A data foundation model building system for implementing the method according to any one of claims 1-9, characterized in that... include: The multimodal data acquisition module is deployed in the physical and logical layers of the target object to obtain a set of basic operating parameters; A feature decoupling engine, connected to the multimodal data acquisition module, is used to generate static feature vectors and dynamic feature sequences; The virtual modeling core receives the static feature vector and dynamic feature sequence, and is used to construct the feature topology and virtual mapping model; The strategy library generator loads the historical running mode library and connects to the virtual modeling core to generate dynamic strategy libraries; The model iteration controller receives anomaly monitoring signals from the dynamic policy library and is used to trigger parameter updates of the virtual mapping model. The data synchronization agent connects the model iteration controller and the data base model body to synchronize the updated mapping relationship.