Intelligent data mining system
By constructing a multi-source data sensing node and a dynamic dimension fusion network, and combining conceptual topology modeling and knowledge vector clustering, a core decision-making benchmark framework is generated, which solves the shortcomings of traditional data mining systems in heterogeneous data processing and achieves efficient and accurate data parsing and decision support.
Patent Information
- Application Number
- CN202510679842.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-05-26
AI Technical Summary
Traditional data mining systems suffer from limitations when processing heterogeneous data, including limited data acquisition dimensions, insufficient accuracy in semantic association parsing, and poor dynamic adaptability of decision-making models. These limitations make it difficult to meet the real-time, accuracy, and comprehensive requirements of feasibility study scenarios.
It employs a heterogeneous data acquisition module, a semantic association parsing module, and a decision graph generation module, including multi-source data perception nodes, dynamic dimension fusion networks, concept topology modeling units, knowledge vector clustering units, policy optimization nodes, and a rule reasoning engine. It captures data features in real time, dynamically constructs topology structures, generates a core decision benchmark framework, and achieves closed-loop feedback through time-series feature correction and knowledge distillation optimization modules.
It achieves comprehensive perception and dynamic analysis of multi-source heterogeneous data, improves the accuracy and flexibility of semantic association, ensures the practicality and relevance of decision graphs, enhances the robustness and adaptability of the system, and can effectively cope with data anomalies and noise interference.
Smart Images

Figure CN120542437B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data mining technology, specifically to an intelligent mining system for feasibility study data. Background Technology
[0002] In modern data-driven decision-making environments, the mining of feasibility study (feasibility study) data faces complex challenges across multiple dimensions and domains. Traditional data mining systems, when processing heterogeneous data, generally suffer from problems such as single data collection dimensions, insufficient accuracy in semantic association parsing, and poor dynamic adaptability of decision models, making it difficult to meet the real-time, accuracy, and comprehensiveness requirements of feasibility study scenarios.
[0003] From a data acquisition perspective, feasibility study data typically includes structured data (such as numerical statistics and tabular data) and unstructured data (such as text reports, image information, and voice recordings). The formats, semantic rules, and update frequencies of these different data types vary significantly. Traditional systems often employ fixed data acquisition patterns, failing to dynamically perceive the meta-attribute characteristics and latent semantic tags of multi-source data. This results in incomplete data feature extraction and difficulty in constructing a topological structure that reflects the inherent relationships within the data. For example, in financial feasibility studies, there are implicit semantic relationships between structured financial data and unstructured industry news texts. Traditional acquisition methods struggle to effectively capture these relationships, impacting the accuracy of subsequent analysis.
[0004] In semantic association parsing, existing technologies largely rely on static concept modeling and single clustering algorithms, lacking the ability to respond to dynamic changes in data semantics. As data volume increases and domain knowledge updates, semantic associations between data exhibit temporal fluctuations and noise interference. Traditional systems cannot correct and compensate for these dynamic changes in real time, leading to increased concept mapping bias rates and decreased reliability of knowledge vector clustering results. For example, in medical feasibility studies, the semantic associations between newly emerging disease terms and existing diagnostic standards need to be updated in real time. Traditional parsing methods struggle to adapt quickly to such changes, potentially resulting in lagging or biased decision-making.
[0005] In the decision graph generation stage, traditional systems typically employ fixed rule engines and strategy frameworks, lacking the ability to dynamically adjust the weights of multi-source data and the mechanism for integrating cross-domain knowledge. When faced with complex feasibility study problems, they cannot generate flexible decision benchmark frameworks based on the real-time status and correlation density of the data, nor can they dynamically optimize information extraction paths, resulting in insufficient practicality and relevance of the decision graph. For example, in urban planning feasibility studies, which involve data from multiple fields such as economy, environment, and society, traditional systems struggle to effectively integrate cross-domain knowledge to generate decision graphs that address multi-dimensional needs, potentially leading to one-sided planning solutions.
[0006] Furthermore, existing systems lack effective closed-loop feedback mechanisms and adaptive optimization capabilities when dealing with data anomalies and noise interference. When abnormal fluctuations or noise interference occur in the data, traditional systems cannot detect and correct them in a timely manner, which may lead to the accumulation of errors throughout the mining process and affect the reliability of the final decision. For example, in industrial production feasibility studies, if sudden noise in sensor data cannot be compensated for in a timely manner, it may lead to misjudgments of the production process, thereby affecting the formulation of production plans. Summary of the Invention
[0007] The purpose of this invention is to provide an intelligent data mining system for feasibility studies, in order to solve the problems mentioned in the background section.
[0008] To achieve the above objectives, the present invention provides the following technical solution: an intelligent data mining system for feasibility studies, the system comprising:
[0009] Heterogeneous data acquisition module, semantic association parsing module, and decision graph generation module;
[0010] The heterogeneous data acquisition module includes multi-source data sensing nodes and a dynamic dimension fusion network; the semantic association parsing module includes a concept topology modeling unit and a knowledge vector clustering unit; the decision graph generation module includes a strategy optimization node and a rule reasoning engine; the multi-source data sensing nodes are used to capture the meta-attribute features of structured data and the latent semantic labels of unstructured data in real time; based on the data update frequency and entity association density, a multi-layer data feature topology is constructed through the dynamic dimension fusion network; the dynamic dimension fusion network consists of M fusion sub-networks connected in series based on domain thresholds; based on the real-time semantic response characteristics of the M fusion sub-networks, a core decision benchmark framework is generated through the strategy optimization node, and combined with the association constraint signals output by the knowledge vector clustering unit, the information extraction path is dynamically weighted and sorted through the rule reasoning engine.
[0011] Preferably, the system further includes a time-series feature correction module and a knowledge distillation optimization module; the time-series feature correction module includes a periodic fluctuation suppression unit and a noise interference compensation unit;
[0012] Based on the generated core decision-making benchmark framework and dynamic weight ranking results, the data semantic association range is corrected in time series by the periodic fluctuation suppression unit, and the vector parameters of the semantic association parsing module are dynamically smoothed by the noise interference compensation unit. At the same time, the concept mapping deviation rate is detected in real time by the entity relationship evaluation network constructed by the knowledge distillation optimization module, and the detection results are fed back to the policy optimization node and rule reasoning engine to perform closed-loop iteration of the decision framework.
[0013] Preferably, the periodic fluctuation suppression unit is composed of feature transfer channels corresponding to the fusion subnets connected in series in the dynamic dimension fusion network; the construction steps of the multi-layer data feature topology include:
[0014] Capture meta-attribute features and semantic label fluctuation signals, input them into the dynamic dimension projection model configured in each level of the fusion subnet, generate data correlation compensation coefficients and store them in the corresponding fusion subnet;
[0015] Based on the relevance compensation coefficients stored in each level of the fusion subnet, and combined with the instantaneous change rate of entity relationships, the information extraction weights are allocated to the M fusion subnets through a feature importance evaluation algorithm, forming a dynamic feature matching topology.
[0016] Preferably, the construction conditions for the fusion subnet configured by the multi-source data sensing nodes include: the meta-attribute features are within a preset domain range, the entity recognition rate corresponding to the fusion subnet is greater than a threshold, and the semantic association density is less than the allowable deviation range;
[0017] The conceptual topology modeling unit is also configured with a multimodal switching link; the multimodal switching link includes a steady-state correlation sub-link and an anomaly blocking sub-link; the steady-state correlation sub-link is generated based on the matching degree between the entity relationship signal of the semantic correlation parsing module and the core decision benchmark framework; the anomaly blocking sub-link is constructed based on the correlation between the instantaneous conflict risk of the information extraction path and the data noise event.
[0018] Preferably, the strategy optimization node is configured with a cross-domain knowledge fusion model; the cross-domain knowledge fusion model includes an attribute-semantic mapping table corresponding to the heterogeneous data acquisition module and an anomaly blocking parameter library of the semantic association parsing module; the steps for generating the core decision benchmark framework include:
[0019] Based on the real-time attribute data and weight ranking results of the heterogeneous data acquisition module, the theoretical upper limit of the association range of each level of the fusion subnet in the steady-state association sublink is calculated.
[0020] By inputting the upper limit of the theoretical correlation range and transient data noise into the cross-domain knowledge fusion model, a decision logic reference graph is generated. The reference graph is then corrected for conflict risks by blocking abnormal sub-links, forming a core decision benchmark framework.
[0021] Preferably, the parameter update steps of the cross-domain knowledge fusion model include:
[0022] When the meta-attribute features are detected to exceed the preset domain range or the concept mapping deviation rate exceeds the threshold, the first abnormal blocking instruction is triggered, that is, the alternative knowledge unit is started and the parsing granularity of the semantic association parsing module is adjusted.
[0023] If the associated range still does not recover to the allowed range after the first abnormal blocking instruction is executed, the second abnormal blocking instruction is triggered, that is, the abnormal blocking sub-link is switched through the multi-modal switching link, and the topology weight of the information extraction path is reallocated based on the duration of the data noise event.
[0024] Preferably, the parameter update step of the cross-domain knowledge fusion model further includes:
[0025] When the entity recognition rate of a certain level of fusion subnet in the heterogeneous data acquisition module is lower than the threshold, the third abnormal blocking instruction is triggered, that is, the associated output channel of the subnet is blocked, and the corresponding data stream is transferred to other fusion subnets through the periodic fluctuation suppression unit.
[0026] If overload of other fusion subnets is detected during the execution of the third abnormal blocking instruction, the fourth abnormal blocking instruction is triggered, which calls the knowledge distillation optimization module to downgrade the output strategy of the rule inference engine and restrict the association requirements of the edge domain.
[0027] Preferably, the degradation processing logic of the knowledge distillation optimization module includes:
[0028] Based on the deviation rate data and knowledge vector clustering index output by the entity relationship assessment network, a domain association security level table is constructed.
[0029] When the fourth abnormal blocking instruction is triggered, the priority of the associated path is dynamically downgraded according to the security level table, and the parsing frequency of the semantic association parsing module is adjusted synchronously to match the decision strategy after the downgrade.
[0030] Preferably, the construction step of the converged subnet further includes:
[0031] Based on the data clustering results of meta-attribute features and the fluctuation range of historical semantic labels, a domain gradient adaptation index table is generated.
[0032] Based on the differences in the correlation compensation coefficients of each fusion subnet in the index table, the topological weights of the feature transmission channels of the M fusion subnets are allocated using a semantic weighted regression algorithm, and redundant correlation links between the fusion subnets are established.
[0033] Preferably, the generation process of the decision logic reference map includes:
[0034] After inputting the upper limit of the theoretical correlation range and transient data noise into the cross-domain knowledge fusion model, the preset knowledge graph database and data anomaly interference layer are loaded simultaneously.
[0035] Based on the ontology relation strength data output by the conceptual topology modeling unit, the conflict threshold of the information extraction path in the reference map is corrected, and the data sparsity parameter is superimposed on the abnormal blocking sub-link for secondary calibration.
[0036] Compared with the prior art, the beneficial effects of the present invention are:
[0037] In heterogeneous data acquisition, multi-source data sensing nodes can capture the meta-attribute features of structured data and the potential semantic labels of unstructured data in real time, breaking through the limitations of traditional systems in acquiring single data types and achieving comprehensive perception of multi-source data. The dynamic dimension fusion network consists of M interconnected fusion subnetworks based on domain thresholds, which can construct multi-layer data feature topologies according to data update frequency and entity association density. This structural design not only adapts to the differences in characteristics of different data types but also dynamically adjusts the fusion method of data features through real-time semantic response characteristics, ensuring that the acquired data can comprehensively and accurately reflect the complex relationships in the actual scenario, providing a rich and reliable data foundation for subsequent semantic parsing and decision generation.
[0038] The semantic association parsing module achieves deep analysis of data semantic associations through the collaborative work of the concept topology modeling unit and the knowledge vector clustering unit. The multimodal switching links configured in the concept topology modeling unit include steady-state association sub-links and anomaly blocking sub-links. These can dynamically switch parsing modes based on the matching degree between entity relationship signals and the core decision-making benchmark framework, as well as the correlation between instantaneous conflict risks in information extraction paths and data noise events, effectively addressing dynamic changes and anomalies in data semantics. The association constraint signals output by the knowledge vector clustering unit, combined with the real-time semantic response characteristics of the dynamic dimension fusion network, can dynamically weight and rank information extraction paths, improving the accuracy and flexibility of semantic parsing. Furthermore, the periodic fluctuation suppression unit and noise interference compensation unit in the temporal feature correction module can perform time-series correction of the data semantic association range and dynamically smooth vector parameters, further reducing the impact of temporal fluctuations and noise interference on semantic parsing, ensuring the stability and reliability of the semantic association parsing results.
[0039] The decision graph generation module achieves intelligent generation and optimization of the decision framework through innovative design of strategy optimization nodes and rule inference engines. The cross-domain knowledge fusion model configured in the strategy optimization nodes can integrate cross-domain knowledge such as the attribute-semantic mapping table from the heterogeneous data acquisition module and the anomaly blocking parameter library from the semantic association parsing module to generate a core decision benchmark framework. This framework calculates the theoretical upper limit of the association range of each level of the fusion subnet in the steady-state association sub-link and combines it with the conflict risk correction of transient data noise and anomaly blocking sub-links to ensure that the decision framework can accurately reflect the real-time state and potential associations of the data. The rule inference engine optimizes the information extraction path based on dynamic weight ranking results and association constraint signals, enabling the decision graph to dynamically adjust the focus and direction of information extraction according to actual needs, thus improving the practicality and relevance of the decision graph.
[0040] The system's temporal feature correction module and knowledge distillation optimization module construct a robust closed-loop feedback mechanism. The temporal feature correction module can detect temporal fluctuations and noise interference in data semantic associations in real time, and perform dynamic correction and compensation. The knowledge distillation optimization module uses an entity relationship evaluation network to detect the concept mapping deviation rate in real time and feeds the results back to the policy optimization node and rule inference engine, achieving closed-loop iteration of the decision framework. This feedback mechanism enables the system to self-optimize and adjust based on the real-time status of the data and the processing results, effectively responding to data anomalies and noise interference, and improving the system's robustness and adaptability.
[0041] In the construction steps of the fusion subnet, the domain gradient adaptation index table generated based on the data clustering results of meta-attribute features and the fluctuation range of historical semantic labels, as well as the topological weight allocation and redundant association link establishment through semantic weighted regression algorithm, further improve the efficiency and reliability of data feature fusion. This design enables the system to dynamically adjust the working mode and weight allocation of the fusion subnet according to the inherent characteristics and historical patterns of the data, ensuring that effective information is retained and utilized to the maximum extent during data acquisition and fusion, and reducing the risk of data loss and misjudgment.
[0042] The parameter update mechanism of the cross-domain knowledge fusion model and the hierarchical triggering strategy of anomaly blocking commands enable the system to handle data anomalies and parsing deviations in a timely and effective manner. Through various means such as activating alternative knowledge units, adjusting parsing granularity, switching parsing links, transferring data streams, and degradation processing, the system can maintain stable operation under varying degrees of anomalies, ensuring the continuity of the data mining process and the reliability of decision results. This multi-layered anomaly handling mechanism significantly improves the system's fault tolerance and anti-interference capabilities in complex environments. Attached Figure Description
[0043] Figure 1 This is a schematic diagram illustrating the working principle of the intelligent data mining system for feasibility studies described in this invention. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] Please see Figure 1 This invention provides an intelligent data mining system for feasibility studies, which includes a heterogeneous data acquisition module, a semantic association parsing module, and a decision graph generation module. The specific structure and implementation of each module are described in detail below.
[0046] The heterogeneous data acquisition module includes multi-source data sensing nodes and a dynamic dimension fusion network. The multi-source data sensing nodes are used to capture the meta-attribute features of structured data and the latent semantic tags of unstructured data in real time. Specifically, for structured data, it connects to databases and APIs through preset data interfaces to extract meta-attribute features such as field names, data types, and value ranges. For unstructured data, such as text, images, and audio, it uses natural language processing, image recognition, and speech recognition technologies to parse the semantic information in the data and generate latent semantic tags, such as extracting keywords and entity names from text, and object categories and color features from images.
[0047] After capturing data features, a multi-layered data feature topology is constructed using a dynamic dimension fusion network based on the data update frequency and entity association density. The dynamic dimension fusion network consists of M interconnected fusion subnets based on domain thresholds, each corresponding to a different domain threshold range, used for hierarchical processing of data features. The specific construction steps are as follows: First, meta-attribute features and semantic label fluctuation signals are captured and input into the dynamic dimension projection model configured in each level of the fusion subnet. This model analyzes the input signals using a preset algorithm, generates data correlation compensation coefficients, and stores these coefficients in the corresponding fusion subnet. Then, based on the correlation compensation coefficients stored in each level of the fusion subnet, combined with the instantaneous change rate of entity relationships, a feature importance evaluation algorithm is used to calculate the importance of each fusion subnet for information extraction. The information extraction weights are then allocated to the M fusion subnets, forming a dynamic feature matching topology, achieving multi-level, dynamic fusion processing of data features.
[0048] The semantic association parsing module includes a concept topology modeling unit and a knowledge vector clustering unit. The concept topology modeling unit constructs the topological structure of concept associations between data. By analyzing the meta-attribute features and semantic labels of the data, it establishes the relationships between concepts, forming a concept network. The knowledge vector clustering unit then performs clustering processing on the concept vectors output by the concept topology modeling unit, grouping concepts with similar semantics into one category and outputting association constraint signals for subsequent decision graph generation.
[0049] The decision graph generation module includes a strategy optimization node and a rule inference engine. The strategy optimization node generates a core decision benchmark framework based on the real-time semantic response characteristics of M fusion subnets. Specifically, the process involves: calculating the theoretical association range upper limit of each level of fusion subnet in the steady-state association sublink based on the real-time attribute data and weight ranking results from the heterogeneous data acquisition module; inputting the theoretical association range upper limit and transient data noise into the cross-domain knowledge fusion model, which includes the attribute-semantic mapping table corresponding to the heterogeneous data acquisition module and the abnormal blocking parameter library of the semantic association parsing module; generating a decision logic reference graph through model analysis; and correcting the conflict risk of the reference graph through abnormal blocking sublinks to form the core decision benchmark framework. The rule inference engine, combined with the association constraint signals output by the knowledge vector clustering unit, dynamically weights and ranks the information extraction paths to optimize the information extraction process.
[0050] Example 1:
[0051] This embodiment relates to the system's extended modules and time-series feature correction function. In addition to the heterogeneous data acquisition module, semantic association parsing module, and decision graph generation module, the system also includes a time-series feature correction module and a knowledge distillation optimization module. The time-series feature correction module further includes a periodic fluctuation suppression unit and a noise interference compensation unit. Each module, through a specific logical architecture and data interaction mechanism, achieves time-series correction of data semantic associations, noise processing, and closed-loop optimization of the decision framework.
[0052] I. Structure and Functional Implementation of the Timing Feature Correction Module
[0053] The periodic fluctuation suppression unit consists of feature transfer channels corresponding to the fusion subnets connected in series within the dynamic dimension fusion network. During system operation, the M fusion subnets of the dynamic dimension fusion network are connected in a preset order. Each fusion subnet corresponds to a specific domain threshold range, used for hierarchical processing of data features. The feature transfer channel, as the data transmission path between fusion subnets, not only undertakes data flow functions but also integrates a time-series feature perception mechanism. Specifically, when the multi-source data perception node captures the meta-attribute features of structured data (such as field update timestamps and numerical fluctuation frequencies) and the latent semantic tags of unstructured data (such as time adverbs implied in text and image capture time metadata), these time-dimensional feature signals enter the dynamic dimension fusion network along with the data. The periodic fluctuation suppression unit monitors the time-series characteristics of the output data of each fusion subnet in real time through the feature transfer channel, such as the periodic fluctuations of the data semantic association range over time (such as regular changes presented by day, week, and month).
[0054] To address detected periodic fluctuations, the periodic fluctuation suppression unit executes the following processing steps: First, it acquires the feature data and timestamps output by each level of the fusion subnet through the feature propagation channel. Then, it uses time-series analysis algorithms (such as the Autoregressive Moving Average (ARMA) model) to model the fluctuation patterns of the data's semantic association range, identifying periodic components (such as sine waves and square waves) and non-periodic noise. Next, it generates an adjustment signal based on the modeling results. This signal is fed back to the corresponding fusion subnet through the feature propagation channel, adjusting the processing parameters within the fusion subnet (such as the weight coefficients of the dynamic dimension projection model and the threshold of the feature importance evaluation algorithm) to suppress the impact of periodic fluctuations on the data's semantic association analysis. For example, if significant fluctuations are found in the semantic association range of a certain type of data between 9:00 AM and 11:00 AM daily, the periodic fluctuation suppression unit can send parameter adjustment instructions to the corresponding fusion subnet through the feature propagation channel, enabling the subnet to automatically enhance the smoothing of the fluctuating features during that period.
[0055] The noise interference compensation unit dynamically smooths the vector parameters of the semantic association parsing module. When constructing the concept association topology, the concept topology modeling unit in the semantic association parsing module converts data features into concept vectors in a high-dimensional vector space. The knowledge vector clustering unit then clusters these vectors. During this process, noise in the data (such as outliers and random interference signals) may cause errors in the concept vectors, affecting the accuracy of the clustering results. The noise interference compensation unit achieves dynamic smoothing through the following mechanism: First, it collects the concept vectors output by the concept topology modeling unit in real time and uses statistical methods (such as the Z-score algorithm) to detect outliers and identify noisy data points. Then, for the detected noise points, it uses a neighborhood smoothing algorithm (such as the K-nearest neighbor algorithm) to generate corrected vector values based on the distribution characteristics of normal vectors around the noise point, replacing the original noise vectors. Simultaneously, the noise interference compensation unit tracks the changing trends of the vector parameters and dynamically filters them using exponential smoothing to reduce the disturbance of high-frequency noise to the vector space structure, ensuring that the association constraint signals output by the knowledge vector clustering unit have high stability and accuracy.
[0056] II. Entity Relationship Evaluation and Feedback Mechanism of the Knowledge Distillation Optimization Module
[0057] The core component of the knowledge distillation optimization module is the entity relationship evaluation network, which performs real-time detection of the concept mapping deviation rate. The entity relationship evaluation network is built on a deep learning architecture. Its inputs are the concept association topology structure output by the concept topology modeling unit and the clustering results output by the knowledge vector clustering unit. The output is the concept mapping deviation rate. The specific implementation process is as follows: First, the concept nodes and their associated edges in the concept association topology are converted into numerical features (such as node degree and edge weight), and the clustering results are converted into label vectors (such as the cluster category to which each concept belongs). Then, the neural network model in the entity relationship evaluation network (as shown in the figure, the convolutional network GCN) jointly analyzes these features and labels, calculates the deviation value between the mapping position of each concept in the clustering space and the ideal position, and thus obtains the overall concept mapping deviation rate.
[0058] The detection results of the entity relationship assessment network are transmitted to the policy optimization node and rule inference engine via feedback links. For the policy optimization node, when it receives a signal that the concept mapping deviation rate exceeds a preset threshold, it triggers the adjustment process of the core decision-making benchmark framework. Specifically, the policy optimization node will re-invoke the cross-domain knowledge fusion model, generate a new decision logic reference graph based on the current deviation data, and correct conflict risks by abnormally blocking sub-links, thereby updating the core decision-making benchmark framework. For example, if a high concept mapping deviation rate is detected in a certain domain, the policy optimization node may adjust the upper limit of the theoretical association range of the corresponding fusion sub-network in the steady-state association sub-link to adapt to the actual changes in concept association.
[0059] For the rule-based reasoning engine, feedback signals from the entity relationship evaluation network are used to adjust the dynamic weight ranking of information extraction paths. The rule-based reasoning engine identifies information extraction paths containing concepts with mapping biases based on the bias rate data. It then reduces the weights of these paths using a weight adjustment algorithm (such as a priority queue algorithm) while increasing the weights of other accurate paths, thereby optimizing the overall efficiency of information extraction. Furthermore, the knowledge distillation optimization module establishes a linkage mechanism with the cross-domain knowledge fusion model. When a persistently high bias rate is detected, it triggers the parameter update process of the cross-domain knowledge fusion model (such as the anomaly blocking instruction mechanism described in Example 4), achieving closed-loop iterative optimization of the overall system processing logic.
[0060] III. Collaborative Workflow of Temporal Feature Correction and Knowledge Distillation Optimization
[0061] During system operation, the temporal feature correction module and the knowledge distillation optimization module form a collaborative processing mechanism. The specific process is as follows: First, the data captured by the heterogeneous data acquisition module is processed by the dynamic dimension fusion network to generate a multi-layer data feature topology with time-series characteristics. This topology structure simultaneously contains the spatial correlation features and temporal evolution features of the data. The periodic fluctuation suppression unit of the temporal feature correction module analyzes the temporal features in the topology through the feature transmission channel to suppress periodic fluctuations, while the noise interference compensation unit smooths the vector parameters of the semantic association parsing module to reduce the impact of noise.
[0062] After time-series correction and noise processing, the data features enter the semantic association parsing module. The concept topology modeling unit constructs an updated concept association topology, and the knowledge vector clustering unit outputs optimized association constraint signals. The entity relationship evaluation network of the knowledge distillation optimization module monitors the mapping accuracy of the concept association topology in real time. If an abnormal deviation rate is detected, the signal is immediately fed back to the policy optimization node and the rule inference engine. The policy optimization node adjusts the core decision benchmark framework based on the feedback signal, the rule inference engine adjusts the information extraction path weights, and simultaneously triggers the parameter update mechanism of the cross-domain knowledge fusion model, forming a closed-loop process of "detection-feedback-adjustment".
[0063] In this collaborative process, the modules interact through standardized data interfaces to ensure the accuracy and real-time performance of data transmission. For example, the adjustment signals sent by the periodic fluctuation suppression unit to the fusion subnet through the feature transmission channel, and the deviation rate data sent by the entity relationship evaluation network to the policy optimization node, all follow preset data formats and communication protocols. Furthermore, the system is equipped with a fault-tolerance mechanism; when one module fails, other modules can continue to operate through redundant links, ensuring the stability and reliability of the overall system.
[0064] IV. Key Technical Details and Parameter Settings
[0065] The time series analysis algorithm for the periodic fluctuation suppression unit can be dynamically selected based on the data type. For structured data (such as numerical time series), models such as ARIMA and SARIMA are preferred; for unstructured data (such as temporal semantics in text), the Time Expression Recognition (TimeML) technique from Natural Language Processing combined with a Hidden Markov Model (HMM) is used for time series analysis. The number of feature transmission channels is consistent with the number M of fusion subnets, and each channel corresponds to an independent data transmission link to ensure the accurate transmission of time series adjustment signals.
[0066] In the neighborhood smoothing algorithm of the noise interference compensation unit, the value of K is dynamically adjusted according to the dimension of the concept vector and the cluster density, and is usually in the range of 3-10. The smoothing coefficient α of the exponential smoothing method is determined by an adaptive algorithm, with an initial value of 0.3, and is adjusted in real time according to the noise intensity (range 0.1-0.6).
[0067] The graph convolutional network model for the entity relationship evaluation network employs two GCN layers, with 256 hidden units in each layer. The activation function is ReLU, and the loss function is mean squared error (MSE). Parameter updates are performed using the backpropagation algorithm. The threshold for the concept mapping deviation rate is set according to the business needs of different domains, typically between 0.1 and 0.3, and can be manually adjusted through the system management interface.
[0068] The feedback delay between the knowledge distillation optimization module and the strategy optimization node does not exceed 50ms, ensuring the system's real-time response to deviation data. The parameter update cycle of the cross-domain knowledge fusion model is consistent with the detection frequency of the entity relationship evaluation network, both being once per second, to ensure the system can adapt to dynamic changes in data characteristics in a timely manner.
[0069] Example 2:
[0070] This embodiment describes the construction conditions of the fusion subnet and the multimodal switching link of the conceptual topology modeling unit. The system ensures the effectiveness of the dynamic dimensional fusion network in the heterogeneous data acquisition module by setting strict construction conditions for the fusion subnet; simultaneously, the conceptual topology modeling unit achieves adaptive switching of data processing paths under steady-state and abnormal scenarios through a multimodal switching link, ensuring the accuracy and robustness of semantic association parsing.
[0071] I. Construction Conditions and Implementation Logic of the Converged Subnet
[0072] The construction of a fusion subnet configured with multi-source data sensing nodes must meet three conditions: the meta-attribute features are within a preset domain range, the entity recognition rate corresponding to the fusion subnet is greater than a threshold, and the semantic association density is less than the allowable deviation range. The specific implementation is as follows:
[0073] When structured data's meta-attribute features (such as the value range of numeric fields and the set of valid values for enumerated fields) are integrated into the system, they are first validated using preset domain range verification rules. For example, the preset range for the "transaction amount" field in financial data is [0, 106]. If the real-time captured meta-attribute feature value exceeds this range (such as -500 or 1.2 × 107), it is determined to be an invalid feature, and the corresponding fusion subnet will not be constructed. The latent semantic tags of unstructured data (such as entity type tags like "person's name" or "organization name" in text) need to match a preset domain tag set. If the tag belongs to a domain-independent category (such as the "weather" tag appearing in medical data), it is considered to be outside the domain range, and the creation of a fusion subnet will not be triggered.
[0074] The entity recognition rate of the fusion subnet is achieved through technologies such as Natural Language Processing (NLP) and Computer Vision (CV). Taking text data as an example, the entity recognition module uses the BERT-BiLSTM-CRF model to perform named entity recognition (NER) on the input text and output entity labels such as person names, place names, and organization names. The entity recognition rate is calculated using the following formula:
[0075]
[0076] When the recognition rate is greater than a preset threshold (e.g., 85%), the fusion subnet is considered to have effective processing capabilities and can be built; if it is lower than the threshold, it indicates that the model's ability to parse entities in the current data is insufficient, and it needs to be optimized through data augmentation, model fine-tuning, etc., until the conditions are met.
[0077] Semantic association density is measured by calculating the sum of edge weights between concept nodes in the conceptual topology modeling unit, reflecting the tightness of semantic associations in the data. The allowable deviation range is pre-set based on domain characteristics; for example, the semantic association density of medical data is typically higher than that of social media data. When the real-time calculated semantic association density exceeds the upper limit of the allowable deviation (e.g., 1.5 times the average of the normal range), it may indicate abnormal data clustering (e.g., noise interference or data errors), and in this case, a fusion subnet is not constructed. If it is below the lower limit of the allowable deviation (e.g., 0.5 times the average of the normal range), the data is considered to have too sparse semantic associations, and constructing a fusion subnet would be of limited significance, and it will also not be created.
[0078] The construction process of the fusion subnet is as follows: After the multi-source data sensing nodes capture data features in real time, they first perform domain interval verification of meta-attribute features. After passing the verification, the entity recognition rate is calculated. If the recognition rate meets the standard, the semantic association density deviation is verified. When all three conditions are met, the dynamic dimension fusion network creates a new fusion subnet and connects it to the existing network to participate in data processing in the order of domain threshold.
[0079] II. Multimodal switching links in conceptual topology modeling units
[0080] The multimodal switching link includes a steady-state correlation sub-link and an abnormal blocking sub-link, which achieve adaptive switching of the data processing path through different triggering mechanisms and processing logics.
[0081] The steady-state associated sub-links are generated based on the matching degree between the entity relationship signals from the semantic association parsing module and the core decision-making benchmark framework. The specific process is as follows: the concept topology modeling unit converts data features into concept nodes and associated edges, forming entity relationship signals (such as the association strength and type between nodes); the core decision-making benchmark framework generated by the strategy optimization nodes includes preset association ranges, weight allocation rules, etc. The matching degree is calculated using the cosine similarity algorithm, with the formula:
[0082]
[0083] When the matching degree is greater than or equal to a preset threshold (e.g., 0.8), it is determined to be a steady-state scenario, and the steady-state association sub-link is activated. This sub-link adopts forward propagation logic, with data sequentially passing through various levels of fusion sub-networks of the dynamic dimension fusion network, the concept topology modeling unit, and the knowledge vector clustering unit, finally outputting to the decision graph generation module to form a complete semantic association parsing path. During this process, the steady-state association sub-link monitors the output data of each module in real time. If an anomaly is detected (e.g., a sudden increase in the variance of the output features of a certain fusion sub-network), a warning signal is sent to the anomaly blocking sub-link, but the path is not immediately switched to avoid misjudgment.
[0084] The abnormal blocking sub-link is constructed based on the correlation between the instantaneous conflict risk and data noise events in the information extraction path. The instantaneous conflict risk is identified by calculating the weight difference of each node in the information extraction path, using the following formula:
[0085]
[0086] When the difference exceeds a threshold (e.g., 0.6), a transient conflict risk is identified. Data noise events are detected using a sliding window algorithm; for example, if more than three abnormal data markers are received consecutively within 5 seconds, a noise event is triggered.
[0087] The triggering conditions for abnormal blocking of sub-links are: the matching degree is lower than the threshold (e.g., 0.5) and the difference in instantaneous conflict risk is greater than 0.6; a data noise event is detected and the duration exceeds 1 second; the cumulative number of steady-state associated sub-link warning signals exceeds 3.
[0088] After triggering an abnormal blocking of a sub-link, the system performs the following operations:
[0089] Path switching: Disconnect data transmission from the steady-state associated sub-link and redirect data to the abnormally blocked sub-link;
[0090] Conflict resolution: By using preset rules in the abnormal blocking parameter library, the node weights of conflict paths are redistributed, for example, reducing the weight of high-difference nodes and increasing the proportion of low-difference nodes.
[0091] Noise filtering: The dynamic dimensional projection model is used to perform secondary feature extraction on the data to filter out the feature components corresponding to noise. For example, principal component analysis (PCA) is used to remove principal components containing noise.
[0092] Path verification: After switching, monitor the output data of the abnormally blocked sub-link in real time. If the matching degree rises to above 0.7 for 5 consecutive cycles (cycle duration is configurable, default is 100ms), then gradually switch back to the steady-state associated sub-link; if the standard is not met, maintain the abnormal blocking mode and trigger the parameter update process of the cross-domain knowledge fusion model (as described in Example 4).
[0093] III. Coordination Mechanism and Technical Details of Multimodal Switching Links
[0094] The steady-state associated sub-link and the abnormal blocking sub-link are mutually exclusively managed through a switching control module. This module adopts a finite state machine (FSM) design, and the states include: steady-state operation, abnormal warning, blocking switching, and recovery verification. The state transition logic is as follows:
[0095] Steady-state operation → Anomaly warning: When the matching degree drops below 0.8 but above 0.5, or the instantaneous conflict risk difference is between 0.4 and 0.6, anomaly warning state is entered, and lightweight data verification is initiated (such as sampling and checking 10% of data features).
[0096] Anomaly Warning → Blocking Switch: If two consecutive verifications fail under warning status (e.g., the sampling data error rate exceeds 15%), or the matching degree is lower than 0.5, blocking switch is triggered, and the abnormal blocking sub-link is entered.
[0097] Blocking Switching → Recovery Verification: After the abnormally blocked sub-link has run for 5 cycles, it will automatically enter the recovery verification state to check the matching degree and conflict risk indicators;
[0098] Resume verification → Steady-state operation / blocking switch: If the indicators meet the requirements, return to steady-state operation; otherwise, maintain the blocking switch state and record the exception log.
[0099] In terms of technical implementation, the multimodal switching link uses a lock-free queue to achieve zero-copy switching of the data path, ensuring a switching latency of less than 50 microseconds. Steady-state correlated sub-links and abnormal blocking sub-links are configured with independent computing resources (such as GPU cores and memory space) to avoid processing latency caused by resource contention. The conceptual topology modeling unit is shared by two links, but differentiated processing is achieved through different parameter configurations (e.g., high-precision analytical parameters for steady-state links and anti-interference parameters for abnormal links).
[0100] IV. Application Scenarios of Converged Subnet Construction and Multimodal Switching
[0101] Taking power system feasibility study data processing as an example:
[0102] Construction of the fusion subnet: When capturing real-time current data from smart meters, first verify the meta-attribute features (such as current value range of 0-100A), entity recognition rate (device number recognition rate must be ≥90%), and semantic association density (the association density between current and voltage, power must be within the normal operating range). After meeting the conditions, create a "current feature fusion subnet" and connect it to the dynamic dimension fusion network to handle feature extraction and fusion of this type of data.
[0103] Multi-mode switching: When a power grid fault causes a large amount of impulse noise in the data, the matching degree of the steady-state associated sub-links rapidly drops to 0.4, while the difference in instantaneous conflict risk rises to 0.7, triggering abnormal blocking of sub-links. After the system switches to abnormal mode, it enhances the filtering of impulse noise through fault data processing rules in the abnormal blocking parameter library, and reallocates the weights of information extraction paths, prioritizing the parsing of fault-related key features (such as voltage sag nodes) to ensure the accuracy of fault diagnosis decisions.
[0104] Example 3:
[0105] This embodiment details the steps for generating the cross-domain knowledge fusion model and core decision-making benchmark framework for the strategy optimization node. The cross-domain knowledge fusion model configured in the strategy optimization node integrates attribute-semantic mapping relationships and anomaly blocking parameters from heterogeneous data to achieve unified representation and analysis of multi-source data. The generation of the core decision-making benchmark framework is based on the dynamic fusion of real-time data and historical knowledge, ensuring the accuracy and robustness of the decision-making logic through multi-stage calculation and correction.
[0106] The cross-domain knowledge fusion model, as a core component of the strategy optimization node, mainly consists of an attribute-semantic mapping table, an anomaly blocking parameter library, and a fusion inference engine. The attribute-semantic mapping table employs a multi-level index structure to associate and map the meta-attribute features of heterogeneous data with semantic labels. For example, in the financial field, the "transaction amount" attribute in structured data can be mapped to the semantic labels "large transaction" and "small transaction," with the mapping rule based on a preset threshold range (e.g., an amount > 1 million is considered a large transaction). Textual descriptions in unstructured data (such as "buying stocks" and "redeeming funds") are extracted using natural language processing techniques and mapped to the semantic category of "investment behavior." The mapping table supports dynamic updates; when new attribute features or semantic labels are detected, the mapping relationships are automatically expanded through an incremental learning algorithm.
[0107] The anomaly blocking parameter library stores the anomaly handling experience accumulated during historical operation, including characteristic patterns, blocking strategies, and parameter configurations for various anomaly scenarios. For example, when abnormal fluctuations are detected in the data (such as a surge in transaction frequency within a short period), the corresponding anomaly handling strategy in the parameter library will be activated. This strategy may include operations such as reducing the weight of relevant information extraction paths and increasing the data verification frequency. The parameter library adopts a hierarchical storage structure, categorized and indexed according to anomaly type (such as data noise, logical conflicts) and severity (such as minor, moderate, and severe) for rapid retrieval and application.
[0108] The fusion inference engine is implemented based on a graph neural network architecture, using an attribute-semantic mapping table and an anomaly blocking parameter library as prior knowledge to perform joint inference on the input real-time data. Specifically, the engine converts heterogeneous data into a knowledge graph representation, where nodes represent entities (such as users and transaction records) and edges represent relationships between entities (such as "user-initiator-transaction"). Feature extraction and propagation are performed on the knowledge graph through a graph convolutional network (GCN) to learn the latent semantic relationships between entities. Simultaneously, the fusion inference engine introduces an attention mechanism to dynamically adjust the level of attention given to different entities and relationships based on information from the anomaly blocking parameter library, enhancing robustness to anomalous data.
[0109] The core decision-making benchmark framework is generated based on the output of a cross-domain knowledge fusion model, and is achieved through multi-stage calculation and correction. The specific steps are as follows:
[0110] Step 1: Calculate the theoretical upper limit of the associative range of the steady-state associative sub-links. Based on the real-time attribute data and weight ranking results of the heterogeneous data acquisition module, calculate the theoretical upper limit of the associative range of each level of the fusion subnet in the steady-state associative sub-links. This upper limit reflects the degree of data association that the fusion subnet can effectively handle under normal business logic. The calculation formula is:
[0111]
[0112] in, Indicates the first The theoretical upper limit of the associativity range of the tiered converged subnet. This represents the maximum correlation value in the historical processed data of this subnet. This is the historical average correlation value. An adjustment coefficient (ranging from 0.6 to 0.8) is used to balance the impact of extreme cases and normal conditions. This formula dynamically determines the reasonable boundary of the correlation range by combining historical statistical data with real-time adjustments.
[0113] Step 2: Generate a Decision Logic Reference Graph. The theoretical correlation range upper limit calculated in Step 1 and transient data noise are input into the cross-domain knowledge fusion model. The fusion inference engine, based on the attribute-semantic mapping table and the anomaly blocking parameter library, performs joint analysis on the input data to generate a decision logic reference graph. This graph is represented by a graph structure, where nodes are decision elements (such as business indicators and risk factors), and edges are the logical relationships between elements (such as causal relationships and dependency relationships). The edge weights in the graph are dynamically adjusted through an attention mechanism to reflect the importance of different elements in the current decision-making scenario. For example, in a financial risk assessment scenario, when anomalies in market fluctuations are detected, the edge weights related to market risk will increase significantly, highlighting their impact on decision-making.
[0114] Step 3: Conflict Risk Correction of Anomaly Blocking Sub-links. The decision logic reference graph is corrected for conflict risks using anomaly blocking sub-links. Specifically, the system first checks for conflicting paths in the reference graph (e.g., two paths leading to contradictory decision conclusions). If such paths exist, corrections are made according to conflict resolution strategies in the anomaly blocking parameter library. For example, when a user's transaction behavior triggers both "high-risk" and "normal" decision paths simultaneously, the system retrieves the priority rules for transaction risk assessment from the anomaly blocking parameter library, prioritizing the more reliable risk indicator path and adjusting the reference graph accordingly. Furthermore, the anomaly blocking sub-links also consider the impact of data noise events, attenuating the weights of decision elements significantly affected by noise to ensure the high reliability of the final core decision benchmark framework.
[0115] The cross-domain knowledge fusion model and the core decision-making benchmark framework generation process collaborate through data exchange and feedback mechanisms. Regarding data exchange, the output of the cross-domain knowledge fusion model (such as the decision logic reference graph) directly serves as input for the generation of the core decision-making benchmark framework, ensuring that the decision logic is based on comprehensive knowledge representation. Simultaneously, new anomaly patterns or parameter adjustment requirements discovered during the core decision-making benchmark framework generation process are fed back to the cross-domain knowledge fusion model, prompting it to update its attribute-semantic mapping table and anomaly blocking parameter library, forming a closed-loop optimization.
[0116] Regarding the feedback mechanism, the system employs a multi-level feedback channel. When deviations occur in the core decision-making benchmark framework during practical application, the first-level feedback is triggered, performing local optimization by adjusting the strategy parameters in the anomaly blocking parameter library. If the deviation persists, the second-level feedback is triggered, initiating the attribute-semantic mapping table update process and expanding or correcting the mapping relationship through incremental learning algorithms. Furthermore, the system also incorporates a periodic global feedback mechanism, regularly conducting comprehensive evaluations and updates to the cross-domain knowledge fusion model and the core decision-making benchmark framework to ensure the system remains adaptable to the ever-changing data environment.
[0117] The graph neural network of the cross-domain knowledge fusion model employs a multi-head attention mechanism with 8 heads to capture semantic relationships across different dimensions. Model training combines supervised and self-supervised learning, utilizing historical labeled data for supervised training while contrastive learning is used to mine latent semantic structures within the data. The Adam optimizer is used during training with a learning rate of 0.001 and a batch size of 64. The number of training epochs is dynamically adjusted based on the data size.
[0118] Adjustment coefficients in the core decision-making benchmark framework generation step An adaptive adjustment strategy is adopted, with an initial value set to 0.7. The system monitors the matching between the theoretical upper limit of the correlation range and the actual correlation degree in real time. When the matching error exceeds 10% for five consecutive periods, fine-tuning is performed using the gradient descent algorithm. The value is adjusted with a step size of 0.01 until the matching error is reduced to an acceptable range.
[0119] The collision risk correction for abnormally blocked sub-links employs a heuristic search algorithm with a search depth of 3 layers to balance correction effectiveness and computational efficiency. During the correction process, the system maintains a collision resolution log, recording the reason, strategy, and effect of each correction for subsequent analysis and optimization.
[0120] Example 4:
[0121] In the production management system of a smart manufacturing enterprise, the anomaly blocking mechanism of the cross-domain knowledge fusion model plays a crucial role. The system constructs a decision graph to optimize production processes by collecting multi-source heterogeneous information such as production equipment data, material inventory data, and process parameter data in real time. When the system detects data anomalies, it triggers different levels of anomaly blocking commands according to preset rules, ensuring the accuracy of decision-making and the stability of the system.
[0122] Triggering and response to the first abnormal blocking command:
[0123] At 10:00 AM one day, while analyzing equipment operation data, the system discovered that the "temperature" meta-attribute characteristic value of a key production equipment reached 180℃, exceeding the preset range [50℃, 150℃]. Simultaneously, the concept mapping deviation rate output by the concept topology modeling unit was 0.35, exceeding the threshold of 0.25. With both conditions met, the first anomaly blocking command was immediately triggered.
[0124] The system quickly activates an alternative knowledge unit, which stores backup handling strategies for equipment malfunctions. Simultaneously, the granularity of the semantic association parsing module is adjusted from "minute-level" to "second-level" to more accurately capture changes in equipment status. Specifically, the system temporarily suspends the knowledge model based on normal operating conditions and instead invokes the equipment malfunction diagnostic model, which is specifically designed to analyze equipment overheating. The semantic association parsing module begins analyzing equipment sensor data on a second-by-second basis, identifying temperature change trends and fluctuations in related parameters.
[0125] Over the next five minutes, the system continuously monitored the equipment temperature by executing the first anomaly blocking command. It found that although the temperature exceeded the normal range, it remained stable and did not continue to rise. At this point, the system determined that the equipment might be in a temporary overload state, but had not yet constituted a serious malfunction. Therefore, it maintained the current production plan but increased the monitoring frequency of the equipment.
[0126] Triggering and execution of the second abnormal blocking command:
[0127] If, after the first abnormal blocking command is executed, the device temperature fails to return to the allowable range within 10 minutes and instead continues to rise to 200°C, the system will trigger the second abnormal blocking command. At this time, the multi-mode switching link immediately switches from the steady-state associated sub-link to the abnormal blocking sub-link, and the topology weights of the information extraction path are reallocated.
[0128] The system dynamically adjusts the priority of each information extraction path based on the duration of data noise events (10 minutes in this case). The weight of the "production progress" path, which originally had a higher weight, is reduced, while the weight of the "equipment safety" related paths is significantly increased. For example, during the generation of the decision graph, information about equipment maintenance and safety warnings is prioritized, while information about production plan adjustments is temporarily shelved.
[0129] Meanwhile, based on the characteristics of data noise events, the system adjusts the parameters of the abnormally blocked sub-links. In this case, because the equipment temperature was abnormal, the system increased the sampling frequency of the temperature sensor data from once per second to five times per second, and adjusted the correlation analysis weights of other relevant parameters (such as pressure and vibration frequency) to more comprehensively assess the equipment status.
[0130] Triggering and handling of the third abnormal blocking command:
[0131] When the entity recognition rate of a certain level of the fusion subnet in the heterogeneous data acquisition module falls below a threshold, a third anomaly blocking command will be triggered. For example, when analyzing raw material quality inspection data, if the entity recognition rate of the fusion subnet responsible for identifying "metal impurity content" drops to 75%, below the 85% threshold, the system will immediately block the associated output channel of that subnet to prevent erroneous data from affecting decision-making.
[0132] Simultaneously, the periodic fluctuation suppression unit transfers the corresponding data stream to other fusion subnets. In this case, the detection data for "metal impurity content" is rerouted to the backup "material composition analysis" fusion subnet. This backup subnet performs secondary analysis on the same batch of raw materials using different algorithms and feature extraction methods, ensuring data accuracy.
[0133] During the execution of the third abnormal blocking command, the system continuously monitors the load of each converged subnet. If other converged subnets are found to be overloaded due to handling additional data flows (e.g., processing latency exceeding 200ms), the fourth abnormal blocking command is triggered.
[0134] Implementation and downgrade handling of the fourth abnormal blocking command:
[0135] In the above case, when the backup "Material Composition Analysis" fusion subnet became overloaded due to processing additional data streams, the system triggered a fourth anomaly blocking instruction. The knowledge distillation optimization module immediately downgraded the output strategy of the rule inference engine and limited the correlation requirements of the edge domain.
[0136] The system first constructs a domain-related security level table, dividing the various domains in the production management system into core domains (such as equipment safety and product quality) and peripheral domains (such as energy consumption statistics and personnel scheduling). Under resource constraints, priority is given to ensuring the related needs of core domains. In this case, "equipment safety" and "product quality" are listed as core domains, while "energy consumption statistics" is listed as a peripheral domain.
[0137] The rule-based reasoning engine dynamically downgrades the priority of associated paths based on security levels. The weight of previously high-priority "energy optimization" paths is reduced to ensure resources are concentrated on analyzing equipment anomalies and raw material quality issues. Simultaneously, the semantic association parsing module's parsing frequency is reduced from 10 times per second to 3 times per second to match the downgraded decision-making strategy and alleviate system load.
[0138] The system also activated a resource scheduling mechanism, temporarily allocating additional computing resources (such as GPU cores and memory space) to key fusion subnets to ensure they could efficiently process data in core domains. In this case, the "material composition analysis" fusion subnet received additional computing resource support, enabling it to maintain a high entity recognition rate even under increased load.
[0139] Coordination and recovery mechanism of abnormal blocking commands
[0140] Throughout the anomaly handling process, the anomaly blocking commands at each level do not operate independently, but rather coordinate with each other to form a hierarchical response system. For example, after the fourth anomaly blocking command is triggered, the system will continue to monitor equipment temperature and raw material quality data. If the situation is found to be deteriorating, the handling strategy may be further adjusted, such as suspending the relevant production process.
[0141] Once the abnormal situation is alleviated, the system will gradually release the blocking commands according to the preset recovery mechanism. Taking an abnormal device temperature as an example, after the temperature returns to the normal range and stabilizes for 15 minutes, the system first releases the second abnormal blocking command and switches back to the steady-state associated sub-link. Subsequently, after a period of observation and confirmation that the device is operating normally, the first abnormal blocking command is released, restoring the normal knowledge model and parsing granularity.
[0142] Throughout the anomaly handling process, the system records detailed logs, including anomaly type, trigger time, handling measures, and recovery time. This log data is used for subsequent system optimization and problem tracing. By analyzing historical anomaly cases, the anomaly blocking parameter library and handling strategies are continuously improved, enhancing the system's robustness and adaptability.
[0143] Through this multi-level anomaly blocking mechanism, the intelligent manufacturing enterprise's production management system can effectively cope with various data anomalies, ensuring the accuracy of decision-making and the stability of production, and maintaining efficient operation even in complex and ever-changing production environments.
[0144] Example 5:
[0145] This embodiment supplements the construction steps of the fused subnet and describes the process of generating the decision logic reference graph. The system optimizes the topology of the fused subnet through a domain gradient adaptation index table and a semantic weighted regression algorithm, and generates an accurate decision logic reference graph by combining a knowledge graph database and anomaly interference layers. The following describes the specific implementation method in detail using a product data mining scenario in the e-commerce field.
[0146] In the construction of the M fusion subnets of the dynamic dimensional fusion network, in addition to basic feature capture and weight allocation, the system generates a domain gradient adaptation index table based on the data clustering results of meta-attribute features and the historical semantic label fluctuation range. Taking e-commerce product data as an example, meta-attribute features include "product category," "price range," and "sales level," etc., and the product data is divided into high-end, medium-end, and low-end market clusters using the K-means clustering algorithm. The historical semantic label fluctuation range is determined by analyzing the frequency and time distribution of sentiment words (such as "high cost performance" and "poor quality") in user reviews. For example, the "cooling effect" label of the "air conditioner" category fluctuates significantly higher in summer than in other seasons.
[0147] The domain gradient adaptation index table stores the clustering results and fluctuation characteristics corresponding to each fusion subnet in tabular form, for example:
[0148]
[0149] Based on the differences in relevance compensation coefficients among the fusion subnets in the index table, the system assigns topological weights to the feature transmission channels of the M fusion subnets using a semantic weighted regression algorithm. For example, subnet A (high-end market) has the highest relevance compensation coefficient for the "brand" tag, and its feature transmission channel is given higher weight when processing the "brand-user loyalty" association path, ensuring that this type of high-value semantic information is transmitted preferentially. Simultaneously, the system establishes redundant association links between fusion subnets. For instance, a "cross-level association channel" is set up between subnet A and subnet B. When mid-range market characteristics appear in the high-end market data (such as a high-end brand launching a budget-friendly series), the redundant link is automatically activated, avoiding semantic gaps caused by data isolation between subnets.
[0150] When generating the decision logic reference graph, the system first inputs the upper limit of the theoretical correlation range and transient data noise into the cross-domain knowledge fusion model, and simultaneously loads the preset knowledge graph database and data anomaly interference layer. Taking the e-commerce promotion decision scenario as an example, the knowledge graph database contains historical correlation data of "user-product-promotion activity", such as the combination of "young users + beauty products + holiday promotions" usually has a high conversion rate; the data anomaly interference layer records abnormal patterns that have appeared in historical promotions, such as "high discount + low sales", which may be caused by insufficient inventory or inadequate promotion.
[0151] The system uses ontology relation strength data output by the conceptual topology modeling unit (e.g., the correlation strength between "product sales volume" and "promotion intensity" is 0.92) to correct the conflict threshold of information extraction paths in the reference graph. For example, when there is a conflict between the "promotion intensity" and "profit margin" paths (high discounts may lead to low profits), the system automatically increases the conflict threshold of the "profit margin" path based on the ontology relation strength, prioritizing the protection of the company's profit goals. In abnormal blocking sub-links, the system performs secondary calibration by overlaying data sparsity parameters. If user feedback data for a promotional activity is sparse (e.g., number of participants < 100), the system reduces the decision weight of that path and marks it as "data needs to be supplemented" to avoid decision bias caused by insufficient data.
[0152] Specific example: Fusion subnet and graph generation for summer air conditioner category promotion decisions.
[0153] E-commerce platforms need to develop promotional strategies for air conditioners during the summer. The data involved includes prices, sales volume, user reviews, and inventory status of air conditioners from various brands. The data types cover both structured (price, sales volume) and unstructured (user reviews, product descriptions).
[0154] The air conditioner data was divided into three clusters using the K-means algorithm: high-end (unit price > 5000 yuan), mid-range (3000-5000 yuan), and low-end (< 3000 yuan), generating subnets A, B, and C respectively. Historical semantic tag analysis showed that the tags "quiet operation" and "intelligent control" for high-end air conditioners fluctuated by 40% during the summer months of June to August, the "cost-effectiveness" tag for mid-range air conditioners fluctuated by 25%, and the "promotional price" tag for low-end air conditioners fluctuated by 35%. Based on this, the "fluctuation period" and "correlation compensation coefficient" fields of the domain gradient adaptation index table were populated.
[0155] According to the index table, the correlation compensation coefficient of the "silent effect" label of subnet A is 0.88, and its feature transmission channel is weighted at 0.9 when processing the "user reviews-product selling points" path; a redundant link is established between subnet B and subnet C. When a backlog of mid-range air conditioner inventory is detected, some data can be diverted to the "promotional activities" processing path of subnet C to try to clear inventory through a low-price strategy.
[0156] Based on historical data, the theoretical upper limit of the "promotional intensity" range for subnet A is set at 15% (high-end brands typically offer discounts of no more than 15%), for subnet B it is 25%, and for subnet C it is 35%. If an outlier of 20% discount for subnet A appears in the real-time data, the system will mark it as transient data noise.
[0157] The knowledge graph database shows that the combination of "high temperature warning area + smart air conditioner" has a conversion rate of 30% in historical promotions, which is significantly higher than the average level. The data anomaly interference layer indicates that there were user complaints about "high-discount air conditioners being out of stock" during the same period last year, and the corresponding path needs to set a conflict threshold.
[0158] Conflict threshold correction and secondary calibration: The ontology relation strength of "Inventory Status - Promotion Feasibility" output by the conceptual topology modeling unit is 0.95. The system increases the conflict threshold of this path from the default 20% to 30%, meaning that the promotion plan can be maintained even when the inventory deviation is within 30%. Since the current inventory data is complete (sparseness < 5%), the abnormal blocking sub-links are not subject to weight adjustment, and the reference graph directly enters the core decision benchmark framework generation stage.
[0159] The domain gradient adaptation index table is updated in sync with the data clustering results, performing a full update by default every day at midnight. When a sudden change in data distribution is detected (such as the launch of a new product category), a real-time incremental update is triggered. The semantic weighted regression algorithm uses an iterative weighting method. After processing each batch of data, the weight coefficients are automatically adjusted based on the path processing effect, with the adjustment not exceeding 10% of the current weight to avoid excessive fluctuations.
[0160] The knowledge graph database of the decision logic reference graph supports multi-source data access, including data from internal business systems and third-party market data (such as weather data and competitor prices), which are cleaned and standardized through a unified data platform. The data anomaly interference layer is constructed using a combination of crowdsourced labeling and automatic detection. Operations personnel can manually mark typical anomaly cases, and the system automatically identifies new anomaly patterns through machine learning algorithms (such as Isolation Forest).
[0161] Redundant links between converged subnets are in a dormant state by default, and are only activated when the main link processing latency exceeds a threshold (e.g., 500ms) or the data feature matching degree is lower than a preset value (e.g., 30%). After activation, the redundant links adopt an asynchronous processing mode to avoid preempting the main link's computing resources and ensure the overall system processing efficiency.
[0162] By using a domain gradient adaptation index table and a semantic weighted regression algorithm, the system achieves dynamic optimization of the fusion subnet topology, improving the processing accuracy of subdivided domain data. The decision logic reference graph integrates multi-source knowledge and anomaly experience to ensure that the generated decision framework conforms to historical patterns and can adapt to real-time data changes, providing a technical solution that combines accuracy and flexibility for feasibility study data mining in fields such as e-commerce.
[0163] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0164] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A feasibility study data intelligent mining system, characterized in that, include: Heterogeneous data acquisition module, semantic association parsing module, and decision graph generation module; The heterogeneous data acquisition module includes multi-source data sensing nodes and a dynamic dimension fusion network; The semantic association parsing module includes a concept topology modeling unit and a knowledge vector clustering unit; The decision graph generation module includes a strategy optimization node and a rule inference engine; the multi-source data perception node is used to capture the meta-attribute features of structured data and the latent semantic labels of unstructured data in real time; based on the data update frequency and entity association density, a multi-layer data feature topology is constructed through the dynamic dimension fusion network; the dynamic dimension fusion network consists of M fusion sub-networks connected in series based on domain thresholds; based on the real-time semantic response characteristics of the M fusion sub-networks, a core decision benchmark framework is generated through the strategy optimization node, and combined with the association constraint signals output by the knowledge vector clustering unit, the information extraction path is dynamically weighted and sorted through the rule inference engine; The system also includes a time-series feature correction module and a knowledge distillation optimization module; The timing feature correction module includes a periodic fluctuation suppression unit and a noise interference compensation unit; Based on the generated core decision-making benchmark framework and dynamic weight ranking results, the data semantic association range is corrected in time series by the periodic fluctuation suppression unit, and the vector parameters of the semantic association parsing module are dynamically smoothed by the noise interference compensation unit. At the same time, the concept mapping deviation rate is detected in real time by the entity relationship evaluation network constructed by the knowledge distillation optimization module, and the detection results are fed back to the strategy optimization node and rule reasoning engine to perform closed-loop iteration of the decision framework. The periodic fluctuation suppression unit is composed of the feature transmission channels corresponding to the fusion subnets connected in series in the dynamic dimension fusion network; The steps for constructing the multi-layer data feature topology include: Capture meta-attribute features and semantic label fluctuation signals, input them into the dynamic dimension projection model configured in each level of the fusion subnet, generate data correlation compensation coefficients and store them in the corresponding fusion subnet; Based on the relevance compensation coefficients stored in each level of the fusion subnet, and combined with the instantaneous change rate of entity relationships, the information extraction weights are allocated to the M fusion subnets through a feature importance evaluation algorithm, forming a dynamic feature matching topology.
2. The intelligent data mining system for feasibility studies as described in claim 1, characterized in that, The conditions for constructing the fusion subnet configured by the multi-source data sensing nodes include: the meta-attribute features are within a preset domain range, the entity recognition rate of the fusion subnet is greater than a threshold, and the semantic association density is less than the allowable deviation range. The conceptual topology modeling unit is also configured with a multimodal switching link; the multimodal switching link includes a steady-state correlation sub-link and an anomaly blocking sub-link; the steady-state correlation sub-link is generated based on the matching degree between the entity relationship signal of the semantic correlation parsing module and the core decision benchmark framework; the anomaly blocking sub-link is constructed based on the correlation between the instantaneous conflict risk of the information extraction path and the data noise event.
3. The intelligent data mining system for feasibility studies as described in claim 2, characterized in that, The strategy optimization node is configured with a cross-domain knowledge fusion model; the cross-domain knowledge fusion model includes an attribute-semantic mapping table corresponding to the heterogeneous data acquisition module and an anomaly blocking parameter library for the semantic association parsing module; The steps involved in generating the core decision-making benchmark framework include: Based on the real-time attribute data and weight ranking results of the heterogeneous data acquisition module, the theoretical upper limit of the association range of each level of fusion subnet in the steady-state association sublink is calculated; By inputting the upper limit of the theoretical correlation range and transient data noise into the cross-domain knowledge fusion model, a decision logic reference graph is generated. The reference graph is then corrected for conflict risks by blocking abnormal sub-links, forming a core decision benchmark framework.
4. The intelligent data mining system for feasibility studies as described in claim 3, characterized in that, The parameter update steps of the cross-domain knowledge fusion model include: When the meta-attribute features are detected to exceed the preset domain range or the concept mapping deviation rate exceeds the threshold, the first abnormal blocking instruction is triggered, that is, the alternative knowledge unit is started and the parsing granularity of the semantic association parsing module is adjusted. If the associated range still does not recover to the allowed range after the first abnormal blocking instruction is executed, the second abnormal blocking instruction is triggered, that is, the abnormal blocking sub-link is switched through the multi-modal switching link, and the topology weight of the information extraction path is reallocated based on the duration of the data noise event.
5. The intelligent data mining system for feasibility studies as described in claim 4, characterized in that, The parameter update steps of the cross-domain knowledge fusion model also include: When the entity recognition rate of a certain level of fusion subnet in the heterogeneous data acquisition module is lower than the threshold, the third abnormal blocking instruction is triggered, that is, the associated output channel of the subnet is blocked, and the corresponding data stream is transferred to other fusion subnets through the periodic fluctuation suppression unit. If overload of other fusion subnets is detected during the execution of the third abnormal blocking instruction, the fourth abnormal blocking instruction is triggered, which calls the knowledge distillation optimization module to downgrade the output strategy of the rule inference engine and restrict the association requirements of the edge domain.
6. The intelligent data mining system for feasibility studies as described in claim 5, characterized in that, The degradation processing logic of the knowledge distillation optimization module includes: Based on the deviation rate data and knowledge vector clustering index output by the entity relationship assessment network, a domain association security level table is constructed. When the fourth abnormal blocking instruction is triggered, the priority of the associated path is dynamically downgraded according to the security level table, and the parsing frequency of the semantic association parsing module is adjusted synchronously to match the decision strategy after the downgrade.
7. The intelligent data mining system for feasibility studies as described in claim 1, characterized in that, The construction steps of the converged subnet also include: Based on the data clustering results of meta-attribute features and the fluctuation range of historical semantic labels, a domain gradient adaptation index table is generated. Based on the differences in the correlation compensation coefficients of each fusion subnet in the index table, the topological weights of the feature transmission channels of the M fusion subnets are allocated using a semantic weighted regression algorithm, and redundant association links between the fusion subnets are established.
8. The intelligent data mining system for feasibility studies as described in claim 3, characterized in that, The generation process of the decision logic reference map includes: After inputting the upper limit of the theoretical correlation range and transient data noise into the cross-domain knowledge fusion model, the preset knowledge graph database and data anomaly interference layer are loaded simultaneously. Based on the ontology relation strength data output by the conceptual topology modeling unit, the conflict threshold of the information extraction path in the reference map is corrected, and the data sparsity parameter is superimposed on the abnormal blocking sub-link for secondary calibration.
Citation Information
Patent Citations
Knowledge graph-oriented knowledge reasoning method based on online graph processing technology
CN117972111A
Automatic filling and examination and knowledge base integrated optimization method based on dynamic rule engine
CN119849619A