Industrial large and small model collaboration framework, reasoning method, training method, device and equipment

Through the collaboration framework of small models and large models, combined with confidence evaluation, dynamically select models for feature extraction and inference of industrial time series data, solving the problem of low reliability of traditional models in complex scenarios and achieving efficient and reliable inference results.

CN120494113APending Publication Date: 2025-08-15BEIHANG UNIV

Patent Information

Application Number
CN202510948832.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When processing industrial time series data, traditional models are difficult to capture complex nonlinear relationships and multi-scale coupling features, and lack quantitative evaluation of model uncertainty, resulting in low reliability of inference results, especially in key decision scenarios, which is difficult to trust model output.

Method used

The small model and large model collaborative framework is adopted to extract the confidence scores through the small model initial feature and determine the confidence score. If the confidence level of the small model is insufficient, the big model is called for deep feature extraction and confidence evaluation, and finally select the appropriate model based on the confidence score for inference results.

Benefits of technology

It realizes the optimization balance between inference accuracy, efficiency and stability in dynamic and high reliability scenarios of industrial time series data, and improves the reliability and computing efficiency of model inference results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494113A_ABST
    Figure CN120494113A_ABST
Patent Text Reader

Abstract

The invention provides an industrial large and small model collaboration framework, a reasoning method, a training method, a device and equipment. Relates to the technical field of industrial control. The method comprises the following steps: when to-be-processed time series data output by an industrial system is obtained, performing feature extraction on the time series data through a small model to obtain a first feature matrix of the time series data, and determining a first confidence score of the first feature matrix; when the first confidence score is smaller than a preset threshold value, performing feature extraction on the time sequence data through a large model to obtain a second feature matrix of the time sequence data, and determining a second confidence score of the second feature matrix; when the second confidence score is greater than the first confidence score, taking a reasoning result obtained by the large model based on the second feature matrix as a target reasoning result; and when the second confidence score is smaller than or equal to the first confidence score, determining a target reasoning result through a small model. According to the method, the reliability of the model reasoning result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of industrial control technology, and in particular to an industrial large-scale model collaborative framework, reasoning method, training method, device and equipment. Background Art

[0002] With the development of industrial intelligence, data-driven models (such as machine learning and deep learning models) have become the mainstream solutions for processing time series data, and are widely used in scenarios such as fault warning, production capacity forecasting, and process optimization.

[0003] Conventional machine learning models or shallow neural networks are commonly used to process industrial time series data. These methods extract time series features through manual feature engineering (such as sliding window statistics and frequency domain transformations), and then establish a mapping between input features and output targets through model training.

[0004] However, due to the particularity of time series data, the above process has the defect of low reliability of inference results. Summary of the Invention

[0005] This application provides an industrial-scale model collaboration framework, reasoning method, training method, device and equipment to improve the reliability of model reasoning results.

[0006] In a first aspect, the present application provides an industrial-scale model collaborative reasoning method, the method comprising:

[0007] When obtaining time series data to be processed output by the industrial system, extracting features of the time series data using the small model to obtain a first feature matrix of the time series data, and determining a first confidence score of the first feature matrix;

[0008] When the first confidence score is less than a preset threshold, extracting features from the time series data using a large model to obtain a second feature matrix of the time series data, and determining a second confidence score of the second feature matrix;

[0009] When the second confidence score is greater than the first confidence score, using the inference result obtained by the large model based on the second feature matrix as the target inference result;

[0010] When the second confidence score is less than or equal to the first confidence score, the target inference result is determined by the small model.

[0011] In an optional implementation, determining a first confidence score of the first feature matrix includes:

[0012] Performing dimensionality decomposition on the first feature matrix to obtain multiple feature dimensions;

[0013] Performing fuzzy processing on each feature dimension to obtain a fuzzy feature matrix corresponding to the first feature matrix;

[0014] The fuzzy feature matrix is input into a trained fuzzy neural network model to obtain a first confidence score of the first feature matrix; the fuzzy neural network model takes learning a first mapping relationship between the fuzzy feature matrix and the first confidence score as a training goal; the first confidence score is used to characterize the reliability of the small model inference result, and the first mapping relationship is determined based on the inference error of the small model.

[0015] In an optional implementation manner, the first mapping relationship is obtained through the following process:

[0016] Calculating a first error between the inference result of the small model and the corresponding true label using a preset loss function, and normalizing the first error to obtain an inference error of the small model;

[0017] Using a nonlinear activation function, by minimizing the loss function between the first inference confidence score and the first target confidence score, the inference error of the small model is mapped to the interval [0,1] to obtain the corresponding first confidence score, so as to obtain the first mapping relationship.

[0018] In an optional implementation, performing fuzzification processing on each feature dimension to obtain a fuzzy feature matrix corresponding to the first feature matrix includes:

[0019] For each feature dimension, a corresponding Gaussian membership function is used to map the feature dimension to a preset fuzzy membership range; wherein the Gaussian membership function corresponding to each feature dimension is obtained by the following process: during the training process of the fuzzy neural network model, the fuzzy mean and fuzzy variance of the Gaussian membership function are set as learnable parameters;

[0020] The fuzzy membership of each of the feature dimensions is combined into the fuzzy feature matrix.

[0021] In an optional implementation, determining a second confidence score of the second feature matrix includes:

[0022] Converting the second feature matrix into a one-dimensional feature vector along the time dimension;

[0023] The one-dimensional feature vector is input into the self-reflective model to obtain a second confidence score of the second feature matrix; the self-reflective model is implemented using a fully connected network, and the self-reflective model takes learning a second mapping relationship between the one-dimensional feature vector and the second confidence score as a training goal; the second confidence score is used to characterize the reliability of the inference result of the large model, and the second mapping relationship is determined based on the inference error of the large model.

[0024] In an optional implementation manner, the second mapping relationship is obtained through the following process:

[0025] Calculating a second error between the inference result of the large model and the corresponding true label using a preset loss function, and normalizing the second error to obtain an inference error of the large model;

[0026] Using a nonlinear activation function, by minimizing the loss function between the second inference confidence score and the second target confidence score, the inference error of the large model is mapped to the interval [0,1] to obtain the corresponding second confidence score, so as to obtain the second mapping relationship.

[0027] In an optional implementation, determining the target inference result by using the small model includes:

[0028] Using the inference result obtained by the small model based on the first feature matrix as the target inference result;

[0029] Alternatively, the target inference result is obtained by averaging the inference result obtained by the small model based on the first feature matrix and the inference result obtained by the large model based on the second feature matrix;

[0030] Alternatively, the target inference result is obtained by performing weighted averaging processing on the inference result obtained by the small model based on the first feature matrix and the inference result obtained by the large model based on the second feature matrix.

[0031] In a second aspect, the present application provides an industrial-scale model collaborative reasoning device, the device comprising:

[0032] a first processing module configured to, when acquiring time series data to be processed output by the industrial system, extract features of the time series data using a small model to obtain a first feature matrix of the time series data, and determine a first confidence score of the first feature matrix;

[0033] a second processing module, configured to, when the first confidence score is less than a preset threshold, perform feature extraction on the time series data using a large model to obtain a second feature matrix of the time series data, and determine a second confidence score of the second feature matrix;

[0034] a third processing module, configured to use, when the second confidence score is greater than the first confidence score, an inference result obtained by the large model based on the second feature matrix as a target inference result;

[0035] The third processing module is further configured to determine the target inference result through the small model when the second confidence score is less than or equal to the first confidence score.

[0036] In a third aspect, the present application provides an industrial large and small model collaboration framework, the framework including a small model, a large model, and a self-reflective model;

[0037] The small model is configured to extract a first feature matrix of the time series data when receiving the time series data, and output an inference result when a first confidence score of the first feature matrix is greater than or equal to a preset threshold, and trigger the large model to receive the time series data when the first confidence score is less than the preset threshold;

[0038] The large model is used to output a second feature matrix to the self-reflective model when receiving time series data. The self-reflective model is used to output a second confidence score when receiving the second feature matrix, and when the second confidence score is greater than the first confidence score, the large model outputs an inference result. When the second confidence score is less than or equal to the first confidence score, the target inference result is determined by the small model.

[0039] In a fourth aspect, the present application provides an industrial-scale model collaborative training method for training the industrial-scale model collaborative framework described in the third aspect through the following process:

[0040] In a first training phase, the small model is independently trained, and the parameters of the large model and the self-reflective model are frozen, so that the small model learns to extract the first feature matrix and processes the first feature matrix whose first confidence score is greater than or equal to the preset threshold;

[0041] In the second training phase, the parameters of the small model are frozen, and only the input layer and output layer of the large model are fine-tuned. Based on the scenario where the first confidence score of the small model triggering intervention is less than the preset threshold, the large model learns to extract the second feature matrix and optimizes the inference error of the large model.

[0042] In the third training stage, the parameters of the large model and the small model are frozen, and the self-reflective model is trained so that the self-reflective model outputs a second confidence score with the second feature matrix of the large model as input, and a dynamic routing signal is generated according to the first confidence score and the second confidence score. By minimizing the overall error of collaborative reasoning, a confidence-driven model switching mechanism is constructed.

[0043] In a fifth aspect, the present application provides an industrial-scale model collaborative training device, the device comprising a first training module, a second training module, and a third training module;

[0044] The first training module is configured to independently train the small model in a first training phase, freeze the parameters of the large model and the self-reflective model, and enable the small model to learn to extract the first feature matrix and process the first feature matrix having the first confidence score greater than or equal to the preset threshold;

[0045] The second training module is configured to freeze the parameters of the small model in a second training phase, and only fine-tune the input layer and output layer of the large model; based on a scenario in which the first confidence score of the small model triggering intervention is less than the preset threshold, the large model learns to extract a second feature matrix, and optimize the inference error of the large model;

[0046] The third training module is used to freeze the parameters of the large model and the small model in the third training stage, train the self-reflective model, so that the self-reflective model outputs a second confidence score with the second feature matrix of the large model as input, generate a dynamic routing signal according to the first confidence score and the second confidence score, and construct a confidence-driven model switching mechanism by minimizing the overall error of collaborative reasoning.

[0047] In a sixth aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0048] The memory stores computer-executable instructions;

[0049] The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of the first aspects and / or the method according to the fourth aspect.

[0050] In the seventh aspect, the present application provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, they are used to implement the method as described in any one of the first aspect and / or the method as described in the fourth aspect.

[0051] In an eighth aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in any one of the first aspects and / or the method described in the fourth aspect.

[0052] The industrial large and small model collaborative reasoning method provided by the present application, when obtaining the time series data output by the industrial system, first calls the small model to perform feature extraction on the time series data, obtains the corresponding first feature matrix, and determines the first confidence score of the first feature matrix. Then, when the first confidence score is less than the preset threshold, calls the large model to perform feature extraction on the time series data, obtains the corresponding second feature matrix, and determines the second confidence score of the second feature matrix. Finally, when the second confidence score is greater than the first confidence score, the inference result obtained by the large model based on the second feature matrix is used as the target inference result, and when the second confidence score is less than or equal to the first confidence score, the target inference result is determined by the small model. In this process, since the uncertainty can be calculated according to the feature matrix of the time series data to determine whether to determine the inference result through the small model or the large model, the target inference result obtained in the end can be made more reliable under the premise of ensuring the efficiency of inference. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0054] Figure 1 A schematic diagram of an application scenario of an industrial-scale model collaborative reasoning method provided in an embodiment of the present application;

[0055] Figure 2 A schematic diagram of the process of collaborative reasoning method for industrial-scale models provided in an embodiment of the present application Figure 1 ;

[0056] Figure 3 A schematic diagram of the process of collaborative reasoning method for industrial-scale models provided in an embodiment of the present application Figure 2 ;

[0057] Figure 4 A schematic diagram of the structure of an industrial-scale model collaboration framework provided in an embodiment of the present application;

[0058] Figure 5 A schematic diagram of the training process of an industrial-scale model collaboration framework provided in an embodiment of the present application;

[0059] Figure 6 A schematic diagram of the structure of an industrial-scale model collaborative reasoning device provided in an embodiment of the present application;

[0060] Figure 7A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0061] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0062] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0063] In industrial systems, time series data (such as equipment sensor data, production process parameters, and energy consumption curves) is a core element reflecting the system's operating status. Its prediction and control capabilities directly impact production efficiency, equipment reliability, and energy utilization. With the advancement of industrial intelligence, data-driven models (such as machine learning and deep learning models) have become the mainstream solution for processing time series data, widely used in scenarios such as fault warning, production capacity forecasting, and process optimization. By exploiting the temporal dependencies and cyclical characteristics in time series, these models can predict the system's future state and enable real-time control, providing data support for industrial automation and decision-making.

[0064] Conventional machine learning models (such as traditional statistical models and deep learning models) or shallow neural networks are typically used to process industrial time series data. These methods extract time series features through manually designed feature engineering (such as sliding window statistics and frequency domain transformations), and then establish a mapping relationship between input features and output targets through model training. For example, in equipment failure prediction, a random forest model is used to determine the equipment's operating status by analyzing statistical features such as the root mean square value and crest factor of vibration signals. In energy consumption forecasting, a seasonal autoregressive integrated moving average (SARIMA) model is used to achieve short-term load forecasting based on the trend and seasonal terms of historical power data. These methods can achieve reasonable results in scenarios with low data noise and relatively regular time series patterns.

[0065] However, industrial time series data often exhibits strong dynamics, non-stationarity, and significant noise interference, which limits the feature expression and generalization capabilities of traditional models. On the one hand, artificially designed features struggle to capture the complex nonlinear relationships and multi-scale coupling characteristics in the data (such as the interaction between vibration components of different frequencies). On the other hand, shallow models have insufficient fitting capabilities, and when there is conceptual drift in the data (such as changes in performance degradation patterns due to equipment aging), the reliability of the inference results decreases significantly. In addition, traditional methods lack a quantitative assessment mechanism for model uncertainty and are unable to effectively distinguish between "model prediction errors" and "inherent data noise." This makes it difficult to trust model outputs in critical decision-making scenarios (such as high-risk equipment control and real-time process adjustments), limiting their in-depth application in core industrial links.

[0066] Therefore, known technologies propose collaborative reasoning methods for large and small models to solve the above problems. However, existing collaborative reasoning methods are mainly divided into task-level and sample-level methods. Task-level collaborative reasoning methods select models based on task types, such as the Chain of Experts Framework (CoE framework), but cannot be dynamically adjusted according to sample complexity. Sample-level collaborative reasoning methods such as the Frugal Generative Pre-trained Transformer (FrugalGPT) and the Routing Large Language Model (RouteLLM) can dynamically select large and small models, but are mainly targeted at natural language processing tasks. The time series data of industrial systems is continuous, high-dimensional, and highly noisy numerical data. Its complexity is not only reflected in single samples (such as outliers at a single time point), but also in temporal dependencies (such as multi-cycle coupling and gradual trends) and physical meaning associations (such as the coupling of vibration frequency domain and equipment speed). At this time, if the text similarity evaluation logic of natural language processing tasks is directly applied, it will not be able to capture the dynamic correlation and physical prior constraints of time series data, resulting in the model selection logic being out of touch with industrial scenario requirements, and thus still has the defect of low reliability.

[0067] The industrial large and small model collaborative reasoning method provided in this application is intended to solve the above technical problems of the prior art. Specifically, the small model is first called to perform feature extraction on the time series data to obtain the corresponding first feature matrix, and the first confidence score of the first feature matrix is determined. According to the first confidence score and the preset threshold, it is determined whether the small model can continue to be used for reasoning and the target reasoning result can be obtained. If the first confidence score is less than the preset threshold, the large model is further called to perform feature extraction on the time series data to obtain the corresponding second feature matrix, and the second confidence score of the second feature matrix is determined. Finally, according to the second confidence score and the first confidence score, it is determined whether to continue to obtain the target reasoning result through the small model or the large model.

[0068] From the above content, it can be seen that the collaborative reasoning method of this application addresses the defects of task-level collaborative reasoning (such as the CoE framework) that only statically allocates models according to task types and cannot dynamically adapt to sample complexity, as well as the lack of global task constraints and insufficient contextual analysis of industrial time series data in sample-level collaborative reasoning. It implements task-level global filtering through pre-set confidence thresholds and performs dynamic model selection in combination with real-time confidence evaluation of sample feature matrices. It ensures large-scale model verification for key tasks to improve reliability and reduces computational overhead by processing simple samples with small models. Ultimately, it achieves an optimized balance among reasoning accuracy, efficiency, and stability in the dynamic and high-reliability scenarios of industrial time series data.

[0069] It should be understood that the method of the present application can be widely applied to reasoning scenarios based on any type of time series data in industrial systems. For example, Figure 1 A schematic diagram of an application scenario of an industrial-scale model collaborative reasoning method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method of the present application can be used in scenarios where sensor data (such as temperature, pressure, and vibration frequency) from industrial processing steps is analyzed in real time to predict the probability of equipment failure. In this scenario, the method of the present application is specifically executed by a control device in the processing step of the industrial system. The control device is in communication with the sensor and is used to obtain sensor data within a preset time period as time series data.

[0070] When the control device obtains the time series data, it first calls the small model to extract features from it, and performs a confidence evaluation on the extracted features to obtain a corresponding first confidence score. Then, when the first confidence score is less than a preset threshold, it calls the large model to extract features from it, and performs a confidence evaluation on the extracted features to obtain a corresponding second confidence score. Based on the second confidence score and the first confidence score, it determines which model to use to obtain the final target inference result. Through the method of the present application, the control device can maintain the high efficiency of the small model in simple scenarios and utilize the deep reasoning capabilities of the large model in complex scenarios; identify and correct model errors through double-reset confidence evaluation and result comparison.

[0071] It is understood that control devices and sensors can transmit data via a variety of communication networks, including, for example, Ethernet, wireless networks (such as Wi-Fi and 5G), and industrial buses. Furthermore, the method of the present application can also be executed by a worker's computer, mobile phone, or other electronic device, as long as it can communicate with the corresponding sensor and obtain time series data. This embodiment is not limited to this.

[0072] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0073] The present application provides an industrial-scale model collaborative reasoning method, which is specifically executed by any electronic device in an industrial system. Figure 2 A schematic diagram of the process of collaborative reasoning method for industrial-scale models provided in an embodiment of the present application Figure 1 .like Figure 2 As shown, the method includes the following steps:

[0074] S201. When obtaining time series data to be processed output by the industrial system, extract features of the time series data using a small model to obtain a first feature matrix of the time series data, and determine a first confidence score of the first feature matrix.

[0075] In this embodiment, the time series data to be processed can specifically be a device or a dimensional indicator in an industrial system, or the time series data of two or more devices or two or more dimensional indicators within a preset time length, depending on the specific purpose of reasoning. For example, if the reasoning purpose is a scenario of a single device or a single dimensional indicator (such as speed fluctuation analysis of a motor), the time series data to be processed is specifically the time series data of the single device within a preset time length. If the reasoning purpose involves multivariate association analysis (such as coupled anomaly detection of equipment vibration, temperature, and current), the time series data of multiple sensors within a preset time length can be used as the time series data to be processed.

[0076] The preset duration may vary in different reasoning scenarios, i.e., the preset duration is determined specifically based on the corresponding reasoning scenario. For example, the preset duration is specifically related to the equipment inspection cycle, process response time, etc.

[0077] In this embodiment, the small model is a lightweight model. A fine-tuned Transformer encoder is used as the small model. Its feature extraction process includes sliding window segmentation, normalization, and multi-layer causal convolution operations on the time series data to capture local temporal dependencies and obtain the first feature matrix of the time series data. It is understood that in practical applications, the small model can also be any lightweight neural network, such as a modified Temporal Convolutional Network (TCN), and this embodiment is not limited to this.

[0078] Understandably, in industrial environments, time series data often contains significant uncertainty, making it difficult for deterministic deep learning methods to handle noise and strong dynamics. This issue is particularly pressing in predictive maintenance, as accurately assessing model confidence is crucial for subsequent decision-making. Therefore, in this embodiment, after the electronic device small model outputs the first feature matrix, a first confidence score for the first feature matrix is further determined.

[0079] In this embodiment, the electronic device further inputs the first feature matrix output by the small model into the pre-trained fuzzy neural network model to obtain a first confidence score. It should be understood that the fuzzy neural network model is trained to output a corresponding first confidence score when the first feature matrix is input.

[0080] In practical applications, electronic devices can also determine the first confidence score for the first feature matrix using any of a number of methods, such as statistical distance metrics, ensemble learning voting, and Bayesian inference. Specifically, when using statistical distance metrics, this can be achieved by quantifying the uncertainty of the data distribution. This uncertainty is measured by the degree of deviation between the first feature matrix and the training data distribution. A larger distance / error indicates a more anomalous sample and a higher uncertainty in the inference result. For example, a steel plant used nonparametric statistical methods to estimate the density distribution of blast furnace sensor features. When real-time data fell into low-density areas (e.g., an abnormal temperature-pressure combination), the confidence score was considered low. When using ensemble learning voting, this can be achieved by quantifying the prediction divergence between models. The uncertainty is measured by the inconsistency of the prediction results from multiple models. The greater the divergence (e.g., a large standard deviation of the predicted values), the less reliable the inference result. When using Bayesian inference, the uncertainty of the model parameters or prediction results is explicitly modeled using a probability distribution. The output confidence score is essentially a probabilistic representation of uncertainty.

[0081] S202: When the first confidence score is less than a preset threshold, extract features from the time series data using the large model to obtain a second feature matrix of the time series data, and determine a second confidence score of the second feature matrix.

[0082] In this embodiment, after obtaining the first confidence score output by the fuzzy neural network model, the electronic device compares the first confidence score with a preset threshold. In this embodiment, the confidence score is a value in the interval [0, 1], and the preset threshold is 0.9. The electronic device compares the first confidence score with 0.9. If the first confidence score is greater than or equal to 0.9, the electronic device continues to process the first feature matrix using the small model to obtain a corresponding inference result as the final target inference result.

[0083] If the first confidence score is less than 0.9, the electronic device calls the large model to extract features from the time series data, obtaining a second feature matrix for the time series data. The large model uses an improved Transformer model based on the pre-trained autoregressive language model GPT-2 architecture. To adapt to time series data, the model retains GPT-2's multi-layer self-attention mechanism and feedforward neural network model, but freezes its pre-trained parameters, fine-tuning only the input and output layers. The input layer uses a temporal patch embedding mechanism, which divides the time series data into multiple local patches (e.g., L = 16 time points) using a sliding window of length L. Each patch is mapped to an embedding vector of dimension d (e.g., d = 768) through a linear transformation, and positional encoding is added to preserve temporal information.

[0084] In this embodiment, the large model performs deep feature extraction on the input time series data, capturing long-range temporal dependencies through a self-attention mechanism to produce a second feature matrix containing richer temporal context information. It is understandable that compared to the small model, the large model is able to capture more complex nonlinear patterns and long-term dependencies (such as device performance degradation trends).

[0085] Furthermore, the electronic device inputs the second feature matrix into an independently trained confidence assessment module, which can have the same structure as the confidence assessment module of the small model (such as a fuzzy neural network), but with independent parameters. The confidence assessment module outputs a second confidence score, which reflects the reliability of the large model's reasoning for the current time series data. It is understandable that in actual applications, the second confidence score of the second feature matrix can also be determined by other methods. For details, please refer to the process of obtaining the first confidence score in the aforementioned embodiment, which will not be repeated here.

[0086] It is understood that in actual applications, the preset threshold value can also be any value in the range of 0.8-0.95, or an adaptive threshold value that is dynamically adjusted according to the risk level of the industrial scenario. This is not limited in this embodiment. The large model can also be an improved Transformer model based on other pre-trained language models (such as BERT, T5), or a specially designed deep time series model (such as Informer, Autoformer). It is understood that for scenarios with extremely high real-time requirements, lightweight large models (such as MiniLM, TinyBERT) can be used. For multimodal data (such as vibration signals + images), an improved version of a multimodal pre-trained model (such as CLIP) can be used. This is not limited in this embodiment.

[0087] As a possible design, in actual applications, a dynamically adjustable preset threshold can be configured. Electronic devices dynamically adjust the preset threshold based on the current scenario. For example, in high-risk scenarios (such as nuclear power plant monitoring), the threshold can be set to 0.95 to ensure high reliability. In common detection scenarios (such as ambient temperature monitoring), the threshold can be reduced to 0.8 to balance efficiency and accuracy. Electronic devices can also dynamically adjust the preset threshold based on historical false alarm rates, device health status, and other factors.

[0088] It is understandable that the small model can provide real-time response and its computational complexity is significantly lower than that of the large model. At this stage, filtering simple samples through the preliminary prediction of the small model can reduce the frequency of calling the large model.

[0089] S203: When the second confidence score is greater than the first confidence score, the inference result obtained by the large model based on the second feature matrix is used as the target inference result.

[0090] S204: When the second confidence score is less than or equal to the first confidence score, determine the target inference result using the small model.

[0091] In this embodiment, after obtaining the second confidence score, the electronic device compares the second confidence score with the first confidence score. If the second confidence score is greater than the first confidence score, it indicates that the large model has higher reasoning reliability for the current task than the small model, and that the large model provides a more reliable feature representation by capturing more complex temporal dependencies (such as device performance degradation trends and multivariate coupling characteristics). The electronic device then uses the inference result obtained by the large model based on the second feature matrix as the target inference result.

[0092] If the second confidence score is less than or equal to the first confidence score, the large model does not have higher inference reliability than the small model for the current task. In this case, although the large model has stronger theoretical expression capabilities, the first feature matrix extracted by the small model may more accurately reflect the key features of the current time series data. Or, the large model may have overfitted or experienced other anomalies when processing the data, resulting in its inference results being less reliable than those of the small model. In this case, the electronic device will use the small model to determine the target inference result.

[0093] Specifically, the electronic device can use the inference result obtained by the small model based on the first characteristic matrix as the target inference result; or, obtain the target inference result after averaging the inference result obtained by the small model based on the first characteristic matrix and the inference result obtained by the large model based on the second characteristic matrix; or, obtain the target inference result after weighted averaging the inference result obtained by the small model based on the first characteristic matrix and the inference result obtained by the large model based on the second characteristic matrix.

[0094] More specifically, when performing weighted averaging on the inference results obtained by the small model and the inference results obtained by the large model, the weight can be dynamically adjusted according to the ratio of the first confidence score to the second confidence score (for example, weight coefficient w1 = first confidence score / (first confidence score + second confidence score), weight coefficient w2 = 1 - w1) to ensure that the model with higher reliability contributes more to the final result.

[0095] By using the inference results from the small model as the target inference result, a more reliable inference result can be quickly fed back, effectively ensuring inference efficiency. By averaging or weighted averaging the inference results from the small model with those from the large model to obtain the target inference result, a comprehensive judgment that integrates the advantages of the large and small models can be fed back, further improving the robustness of inference results under complex working conditions (for example, mitigating the risk of overfitting or noise sensitivity of a single model).

[0096] For example, if a wind turbine's gearbox is equipped with a vibration sensor (sampling frequency 10kHz), it is necessary to analyze the vibration signal's characteristic matrix to provide real-time warnings of potential faults. Using the method of this embodiment, the electronic device first invokes a small model, which segments the vibration signal using a sliding window (window length 512 points, 50% overlap), extracts 12-dimensional features such as root mean square (RMS), kurtosis, and frequency entropy, and generates a first characteristic matrix. A fuzzy neural network then calculates a first confidence score to quantify the "normal mode matching" of the characteristic matrix. Assuming the vibration signal is stationary, the fuzzy neural network outputs a confidence score of 0.92 (higher than the threshold of 0.9), and the small model's inference result ("equipment normal") is directly used.

[0097] When the gearbox enters the abnormal wear stage, the vibration signal begins to exhibit non-stationary impact components. The kurtosis value in the feature matrix extracted by the small model increases significantly, and the fuzzy neural network outputs a confidence score of 0.85 (below the threshold of 0.9), triggering the call of the large model. The electronic device then uses the large model to embed a time series patch of the original vibration signal (a patch length of 1024 points, mapped to a 768-dimensional vector). Using a self-attention mechanism, it captures the coupling relationship between different frequency components and generates a second feature matrix containing time-frequency domain correlation features (such as the cross-correlation coefficient between the gear mesh frequency and the bearing fault frequency). Assuming that an independent confidence assessment module (using the Mahalanobis distance method) calculates the distance between the second feature matrix and the normal operating condition feature distribution and outputs a second confidence score of 0.94 (above the threshold of 0.9), since the second confidence score (0.94) is higher than the first confidence score (0.85), the system adopts the large model's inference result: "Early signs of abnormal gear wear have been detected, and maintenance is recommended within three days."

[0098] It can be seen that the collaborative reasoning method for industrial large and small models provided in this application achieves a balance between reasoning accuracy and computational efficiency through a confidence-driven dynamic model selection mechanism, a multi-dimensional uncertainty quantification method, and a feature extraction technology adapted to industrial scenarios, significantly improving the reliability and robustness of complex industrial time series data processing. It is especially suitable for scenarios such as predictive maintenance and process control that have extremely high requirements for real-time and security.

[0099] In addition, in this application, the large model is called to determine the setting of the inference result only when the second confidence score is greater than the first confidence score. The confidence-driven dynamic screening mechanism can effectively reduce the number of invalid calls to the large model, thereby saving the cost overhead caused by invalid calls to the large model.

[0100] As a preferred embodiment, in this embodiment, the electronic device is configured with a target threshold value in addition to the aforementioned preset threshold value, specifically a value within the range of 0-0.1. On this basis, only when the second confidence score is greater than the sum of the first confidence score and the target threshold value is the inference result obtained by the large model based on the second feature matrix used as the target inference result. When the second confidence score is less than or equal to the sum of the first confidence score and the target threshold value, the target inference result is determined using the small model.

[0101] Through the above settings, the second confidence score of the large model is allowed to still use the inference results of the small model when it does not exceed the error range. The call is triggered only when the second confidence score of the large model is significantly higher than that of the small model (outside the error range). This not only tolerates small fluctuations in confidence assessment and avoids frequent invalid calls of the large model due to errors, but also ensures that the large model intervenes when the reliability advantage is sufficient by limiting the error range, thus achieving a balance between error tolerance and inference accuracy.

[0102] In practical applications, the target threshold can also be flexibly adjusted according to the industrial scenario to adjust the strictness of large model calls, reduce invalid calculations while ensuring the reliability of reasoning, and achieve refined control of confidence decision boundaries and optimal allocation of computing resources.

[0103] As a further explanation, this embodiment further illustrates how to determine the first confidence score and the second confidence score based on the above embodiment. Figure 3 A schematic diagram of the process of collaborative reasoning method for industrial-scale models provided in an embodiment of the present application Figure 2 ,like Figure 3 As shown, the method of this embodiment includes:

[0104] S301 , when obtaining the time series data to be processed output by the industrial system, extracting features of the time series data through a small model to obtain a first feature matrix of the time series data.

[0105] S302: Decompose the first feature matrix into multiple dimensions to obtain multiple feature dimensions.

[0106] S303: Perform fuzzy processing on each feature dimension to obtain a fuzzy feature matrix corresponding to the first feature matrix.

[0107] Specifically, in this embodiment, for each feature dimension, a corresponding Gaussian membership function is used to map the feature dimension to a preset fuzzy membership range; wherein, the Gaussian membership function corresponding to each feature dimension is obtained through the following process: during the training process of the fuzzy neural network model, the fuzzy mean and fuzzy variance of the Gaussian membership function are set as learnable parameters; and the fuzzy membership of each feature dimension is merged into a fuzzy feature matrix.

[0108] In this embodiment, when obtaining the first feature matrix of the time series data, the first feature matrix is first decomposed to obtain multiple feature dimensions. Each feature dimension is then processed independently, and for each feature dimension d, a separate fuzzy membership function is defined to map the data of the feature dimension to the [0,1] fuzzy membership range. Specifically, this embodiment uses a Gaussian membership function to fuzzify each feature dimension, and the fuzzy membership function of each feature dimension in this embodiment is not manually determined by expert knowledge. Instead, this embodiment sets the fuzzy mean and fuzzy variance as learnable parameters in the fuzzy neural network so that an appropriate Gaussian membership function can be learned for each dimensional feature during the training process.

[0109] Next, the member values of all feature dimensions are combined into a fuzzy feature matrix. These fuzzy features removed by the fuzzy neural network effectively capture the uncertainty in time series data, thereby improving the model's predictive performance in complex time series forecasting tasks. Furthermore, they provide key information for subsequent confidence estimation, further enhancing overall inference accuracy.

[0110] It is understandable that in practical applications, other differentiable fuzzy membership functions such as triangular membership functions, trapezoidal membership functions, and S-shaped membership functions can also be used to replace Gaussian membership functions. For example, for scenarios with extremely high real-time requirements (such as high-speed production line monitoring), triangular membership functions / trapezoidal membership functions can be used. For scenarios that need to represent asymmetric distribution uncertainty (such as progressive performance degradation during equipment aging), S-shaped membership functions can be used. This is not limited in this embodiment. In this embodiment, the Gaussian membership function is used to determine the fuzzy feature matrix, which can adaptively fit the data distribution of different feature dimensions through learnable mean and variance parameters, more accurately capture the nonlinear uncertainty in industrial time series data, and maintain gradient continuity to support end-to-end model training.

[0111] S304: Input the fuzzy feature matrix into the trained fuzzy neural network model to obtain a first confidence score of the first feature matrix.

[0112] Among them, the fuzzy neural network model takes the first mapping relationship between learning the fuzzy feature matrix and the first confidence score as the training goal; the first confidence score is used to characterize the reliability of the small model reasoning result, and the first mapping relationship is determined based on the reasoning error of the small model.

[0113] More specifically, the first mapping relationship is obtained through the following process: using a preset loss function to calculate the first error between the inference result of the small model and the corresponding true label, and normalizing the first error to obtain the inference error of the small model; using a nonlinear activation function, by minimizing the loss function between the first inference confidence score and the first target confidence score, the inference error of the small model is mapped to the [0,1] interval to obtain the corresponding first confidence score, so as to obtain the first mapping relationship.

[0114] It should be understood that in fields such as natural language processing, confidence scores typically rely on metrics derived from soft-max outputs, such as the maximum output value or the entropy of the output distribution, to assess confidence. However, in regression tasks, the model output is no longer a probability distribution, but rather a continuous or multi-dimensional vector of actual values, making traditional confidence estimation methods based on maximum probability no longer applicable. Therefore, to determine whether to use a large model, this embodiment adopts a confidence-based decision strategy that can adaptively learn confidence scores for fuzzy neural network models.

[0115] As shown above, in this embodiment, according to the inference results of the small model and the true label , we can get the first target confidence score as follows: ,in, is a nonlinear activation function used to transform The value of is limited to the range [0,1]. is a parameter that controls the sensitivity of the function to inference errors, Used to represent the actual inference results of the small model, Used to represent the original inference result of the small model. The first inference confidence score The process can be expressed as: , where F is used to represent the fuzzy neural network model, Used to represent the first characteristic matrix, is a sigmoid function, W is a weight vector, and b is a bias term. In this embodiment, the first mapping relationship is finally obtained by minimizing the loss function between the first inference confidence score and the first target confidence score.

[0116] In this embodiment, a fuzzy feature matrix is obtained and input into a pre-trained fuzzy neural network model to obtain a first confidence score for the first feature matrix. In this embodiment, by integrating fuzzy learning methods into a deep learning framework and utilizing learnable fuzzy membership functions, fuzzy learning transforms uncertain and noisy time series features into fuzzy metrics that capture uncertainty, thereby providing strong support for confidence prediction. Compared to traditional deterministic neural networks, fuzzy neural network models are more robust to noise and heterogeneity in industrial time series data and can make more reliable decisions under uncertain conditions.

[0117] S305: When the first confidence score is less than a preset threshold, extract features from the time series data using the large model to obtain a second feature matrix of the time series data; and convert the second feature matrix into a one-dimensional feature vector along the time dimension.

[0118] S306 , inputting the one-dimensional feature vector into the self-reflective model to obtain a second confidence score of a second feature matrix.

[0119] Among them, the self-reflective model is implemented using a fully connected network, and the self-reflective model takes learning the second mapping relationship between the one-dimensional feature vector and the second confidence score as the training goal; the second confidence score is used to characterize the reliability of the large model reasoning result, and the second mapping relationship is determined based on the reasoning error of the large model.

[0120] More specifically, the second mapping relationship is obtained through the following process: using a preset loss function to calculate the second error between the inference result of the large model and the corresponding true label, and normalizing the second error to obtain the inference error of the large model; using a nonlinear activation function, by minimizing the loss function between the second inference confidence score and the second target confidence score, the inference error of the large model is mapped to the [0,1] interval to obtain the corresponding second confidence score, so as to obtain the second mapping relationship.

[0121] As described above, in this embodiment, the self-reflective model is implemented using a fully connected network, and its input is the second feature matrix extracted by the large model. In order to reduce the complexity of high-dimensional time series features, the input is first flattened along the time dimension and converted into a static one-dimensional feature vector. Then, a single-layer fully connected projection is designed to map the one-dimensional feature vector to the second inference confidence score. , the formula is as follows: , where R is used to represent the self-reflective model, It is used to represent the one-dimensional feature vector corresponding to the second feature matrix. The fully connected network is trained using a supervision signal based on the large model inference error. Similarly, the second target confidence score It can be defined as: ,in, is a nonlinear activation function used to transform The value of is limited to the range [0,1]. Used to represent the real inference results of large models, Used to represent the original inference results of large models, , is the residual scaling factor that controls the error sensitivity.

[0122] This embodiment obtains the second mapping relationship by minimizing the loss function between the second reasoning confidence score and the second target confidence score. Specifically, the loss function is expressed as: , where N is used to represent the batch size, and They are used to represent the second reasoning confidence score and the second target confidence score of the i-th sample, respectively. During the reasoning process, the self-reflective model quantifies the potential reasoning bias based on the second feature matrix extracted by the large model and generates a confidence score.

[0123] In this embodiment, a self-reflective model is constructed to perform metacognitive evaluation of deep time series features extracted from a large model. Combined with a hyperbolic tangent mapping function with adjustable error sensitivity, this allows for precise quantification of uncertainty in large model reasoning, providing an interpretable confidence reference for industrial decision-making. Simultaneously, time-dimensional feature compression reduces computational overhead, making real-time confidence assessment of complex time series models possible. For example, flattening a two-dimensional feature matrix (time step × feature dimension) into a one-dimensional vector can reduce the number of parameters (e.g., from 100 × 512 to 51,200), thereby improving computational efficiency.

[0124] It should be emphasized that the aforementioned S301-S304 process is specifically a process for determining a first confidence score, and the S305-S306 process is specifically a process for determining a second confidence score. In actual applications, the electronic device can call these two processes separately when needed, that is, in this embodiment, S301-S304 and S305-S306 are two independent processes, and there is no meaning that S305-S306 must be executed after executing S301-S304.

[0125] It should be understood that the method of the present application is implemented based on an industrial large-scale model collaborative framework. The present application also provides an industrial large-scale model collaborative framework. Specifically, the framework of the present application includes a small model, a large model, and a self-reflective model; wherein the small model is used to extract the first feature matrix of the time series data when receiving the time series data, and output the inference result when the first confidence score of the first feature matrix is greater than or equal to the preset threshold, and trigger the large model to receive the time series data when the first confidence score is less than the preset threshold; the large model is used to output the second feature matrix to the self-reflective model when receiving the time series data, and the self-reflective model is used to output the second confidence score when receiving the second feature matrix, and when the second confidence score is greater than the first confidence score, enable the large model to output the inference result, and when the second confidence score is less than or equal to the first confidence score, determine the target inference result through the small model.

[0126] More specifically, Figure 4 This is a schematic diagram of the structure of an industrial-scale model collaboration framework provided by an embodiment of the present application. Figure 4 As shown, the framework includes a small model, a large model, a fuzzy neural network model, and a self-reflective model, wherein the small model is used to output a first feature matrix to the fuzzy neural network model when receiving time series data, and the fuzzy neural network model outputs a first confidence score when receiving the first feature matrix; the small model is also used to output an inference result when the first confidence score is greater than or equal to a preset threshold, and trigger the large model to receive time series data when the first confidence score is less than the preset threshold; the large model is used to output a second feature matrix to the self-reflective model when receiving time series data, and the self-reflective model is used to output a second confidence score when receiving the second feature matrix, and when the second confidence score is greater than the first confidence score, the large model outputs an inference result, and when the second confidence score is less than or equal to the first confidence score, the target inference result is determined by the small model.

[0127] Next, based on the above content, combined with Figure 4The collaborative reasoning method of industrial large and small models of the present application is introduced in detail. In the present application, when processing time series data, a lightweight small model is first used to extract features of the input sample (i.e., time series data) to obtain a first feature matrix. Then, the fuzzy decision mechanism constructed based on the fuzzy neural network evaluates the first feature matrix extracted by the small model and generates a first confidence score. If the first confidence score is higher than the preset threshold, the reasoning result of the small model is directly output; otherwise, the large model is triggered to perform deep reasoning. Specifically, the large model is used to extract features to obtain a second feature matrix, and the self-reflective model is used to evaluate the second feature matrix of the large model to generate a second confidence score. If the second confidence score corresponding to the large model is lower than the first confidence score corresponding to the small model, the small model auxiliary reasoning mechanism is triggered to correct the reasoning result of the small model to generate the final target reasoning result.

[0128] Next, the training process of the aforementioned industrial-scale model collaboration framework is described in detail. Figure 5 A schematic diagram of the training process of an industrial-scale model collaboration framework provided in an embodiment of the present application. Figure 5 As shown, this application adopts a multi-stage training strategy and proposes a three-stage progressive training strategy to achieve modular collaborative reasoning and parameter decoupling, and adopts a hierarchical parameter freezing mechanism to reduce the risk of gradient interference.

[0129] Specifically, in the first training stage, the small model is trained independently, and the parameters of the large model and the self-reflective model are frozen, so that the small model learns to extract the first feature matrix and process the first feature matrix with a first confidence score greater than or equal to a preset threshold; in the second training stage, the small model parameters are frozen, and only the input layer and output layer of the large model are fine-tuned. Based on the scenario where the first confidence score of the small model triggers intervention less than the preset threshold, the large model extracts the second feature matrix and optimizes the reasoning error of the large model; in the third training stage, the parameters of the large and small models are frozen, and the self-reflective model is trained, so that the self-reflective model takes the second feature matrix of the large model as input and outputs the second confidence score, generates dynamic routing signals according to the first confidence score and the second confidence score, and constructs a confidence-driven model switching mechanism by minimizing the overall error of collaborative reasoning.

[0130] More specifically, during training phase 1, the small model (S) is trained independently. All other modules are frozen during training. S's input is time series data, and its output is the first feature matrix (used for confidence assessment). End-to-end optimization minimizes inference error, learns to establish a feature representation for the time series, and outputs the inference results. The training goal is to enable S to quickly extract effective features and process simple time series patterns.

[0131] Training Phase 2 - Fine-tuning the Large Model (L): The parameters of S are frozen during training to prevent optimization from being interfered with by S's backpropagation path. More specifically, L is initialized based on a pretrained language model (such as GPT-2). Only the input layer (temporal patch embedding) and output layer (inference head) are fine-tuned, while the parameters of the core layers (self-attention, FFN) remain frozen. L's input is high-uncertainty data (i.e., time series data where the small model's confidence is less than a preset threshold) that triggers the intervention of the small model. Its output is the second feature matrix (for confidence assessment) and the inference result (for error calculation).

[0132] Training Phase 3 - Training the Reflective Model (R) and the Decision Agent (F): With all other module parameters frozen, F is trained. F serves as a confidence-driven model switching controller. R and F learn the joint distribution of input features and predictions through a neural network, building a confidence-driven dynamic routing mechanism. R's input is the one-dimensional feature vector corresponding to the second feature matrix output by the large model, and its output is the second confidence score. F's input is the first confidence score corresponding to S and the second confidence score corresponding to L. Its output is the model selection signal. The training objective is to minimize the overall error of the collaborative reasoning.

[0133] Through the above three-stage training strategy, combined with the hierarchical parameter freezing mechanism, this solution significantly improves the training efficiency and generalization ability of the collaborative framework of industrial large and small models while ensuring the functional independence of each module. It is particularly suitable for industrial scenarios with complex data distribution and limited computing resources.

[0134] As can be seen from the foregoing, the method of this application has the following advantages: through a sample-level large and small model collaborative framework, the appropriate model is dynamically selected for inference based on sample characteristics, achieving flexible representation capabilities and improving inference reliability. Through fast inference using lightweight small models, combined with fuzzy decision-making mechanisms and self-reflection mechanisms, the frequency of calling large models is effectively reduced, thereby reducing computational costs. By adopting a phased training strategy and parameter decoupling technology, the problem of gradient imbalance during dynamic neural network training is solved, improving training efficiency and model performance.

[0135] Specifically, unlike static task-level collaboration, the method in this application dynamically selects the appropriate model at each sample level, assigning simple samples to the small model and only handing off complex samples to the large model. This significantly reduces the number of calls to the large model and lowers computational costs. Fuzzy neural networks are also used to quantify sample uncertainty and accurately assess sample complexity, enabling more efficient judgment on whether to call the large model and avoiding unnecessary resource waste.

[0136] In this application's method, when the large model's prediction confidence is low, the small model is called upon for correction, thereby reducing the risk of catastrophic errors in the large model in industrial scenarios. Simultaneously, sample-level collaboration is employed, with model selection and result fusion performed at each sample level. This ensures the accuracy and reliability of predictions and avoids the decision-making bias that can arise from static task-level collaboration.

[0137] This method achieves modular collaborative reasoning and parameter decoupling optimization through phased training, reducing the risk of gradient interference and improving the model's scalability and adaptability. Furthermore, this framework can be easily extended to more complex fusion strategies, such as adaptive soft fusion methods, to accommodate different application scenarios and requirements.

[0138] The above embodiments introduce an industrial-size model collaborative reasoning method and training method from the perspective of method flow. The following embodiments introduce an industrial-size model collaborative reasoning device and training device from the perspective of virtual modules or virtual units. Please see the following embodiments for details.

[0139] The present application embodiment provides an industrial-scale model collaborative reasoning device, Figure 6 A schematic diagram of the structure of an industrial-scale model collaborative reasoning device provided in an embodiment of the present application is shown as follows: Figure 6 As shown, the device includes:

[0140] The first processing module 61 is configured to, when acquiring time series data to be processed output by the industrial system, extract features of the time series data using the small model to obtain a first feature matrix of the time series data, and determine a first confidence score of the first feature matrix;

[0141] A second processing module 62 is configured to, when the first confidence score is less than a preset threshold, perform feature extraction on the time series data using the large model to obtain a second feature matrix of the time series data, and determine a second confidence score of the second feature matrix;

[0142] A third processing module 63 is configured to use the inference result obtained by the large model based on the second feature matrix as the target inference result when the second confidence score is greater than the first confidence score;

[0143] The third processing module 63 is further configured to determine the target inference result using the small model when the second confidence score is less than or equal to the first confidence score.

[0144] In another possible implementation of the embodiment of the present application, the first processing module 61 is specifically configured to:

[0145] Performing dimension decomposition on the first feature matrix to obtain multiple feature dimensions;

[0146] Perform fuzzy processing on each feature dimension to obtain a fuzzy feature matrix corresponding to the first feature matrix;

[0147] The fuzzy feature matrix is input into the trained fuzzy neural network model to obtain a first confidence score of the first feature matrix; the fuzzy neural network model takes learning the first mapping relationship between the fuzzy feature matrix and the first confidence score as the training goal; the first confidence score is used to characterize the reliability of the small model inference result, and the first mapping relationship is determined based on the inference error of the small model.

[0148] In another possible implementation of the embodiment of the present application, the first mapping relationship is obtained through the following process:

[0149] The first error between the inference result of the small model and the corresponding true label is calculated using a preset loss function, and the first error is normalized to obtain the inference error of the small model;

[0150] Using a nonlinear activation function, by minimizing the loss function between the first inference confidence score and the first target confidence score, the inference error of the small model is mapped to the interval [0,1] to obtain the corresponding first confidence score, so as to obtain a first mapping relationship.

[0151] In another possible implementation of the embodiment of the present application, the first processing module 61 is specifically configured to:

[0152] For each feature dimension, a corresponding Gaussian membership function is used to map the feature dimension to a preset fuzzy membership range; wherein the Gaussian membership function corresponding to each feature dimension is obtained by the following process: during the training process of the fuzzy neural network model, the fuzzy mean and fuzzy variance of the Gaussian membership function are set as learnable parameters;

[0153] The fuzzy membership of each feature dimension is combined into a fuzzy feature matrix.

[0154] In another possible implementation of the embodiment of the present application, the second processing module 62 is specifically configured to:

[0155] Convert the second feature matrix into a one-dimensional feature vector along the time dimension;

[0156] The one-dimensional feature vector is input into the self-reflective model to obtain a second confidence score of the second feature matrix; the self-reflective model is implemented using a fully connected network, and the self-reflective model takes learning the second mapping relationship between the one-dimensional feature vector and the second confidence score as the training goal; the second confidence score is used to characterize the reliability of the large model inference result, and the second mapping relationship is determined based on the inference error of the large model.

[0157] In another possible implementation of the embodiment of the present application, the second mapping relationship is obtained through the following process:

[0158] The preset loss function is used to calculate the second error between the inference result of the large model and the corresponding true label, and the second error is normalized to obtain the inference error of the large model;

[0159] By using a nonlinear activation function, the inference error of the large model is mapped to the interval [0,1] by minimizing the loss function between the second inference confidence score and the second target confidence score, and the corresponding second confidence score is obtained to obtain a second mapping relationship.

[0160] In another possible implementation of the embodiment of the present application, the second processing module 63 is specifically configured to:

[0161] The inference result obtained by the small model based on the first feature matrix is used as the target inference result;

[0162] Alternatively, the target inference result is obtained by averaging the inference result obtained by the small model based on the first feature matrix and the inference result obtained by the large model based on the second feature matrix;

[0163] Alternatively, the target inference result is obtained by performing weighted averaging processing on the inference result obtained by the small model based on the first feature matrix and the inference result obtained by the large model based on the second feature matrix.

[0164] An industrial-size model collaborative reasoning device provided in an embodiment of the present application is applicable to the above-mentioned industrial-size model collaborative reasoning method embodiment, which will not be described in detail here.

[0165] The embodiment of the present application provides an industrial-scale model collaborative training device, the device comprising a first training module, a second training module, and a third training module;

[0166] A first training module is configured to independently train the small model in a first training phase, freeze the parameters of the large model and the self-reflective model, and enable the small model to learn to extract the first feature matrix and process the first feature matrix having a first confidence score greater than or equal to a preset threshold;

[0167] The second training module is used to freeze the parameters of the small model in the second training phase and only fine-tune the input and output layers of the large model. Based on the scenario where the first confidence score of the small model triggering intervention is less than a preset threshold, the large model learns to extract the second feature matrix and optimize the large model inference error;

[0168] The third training module is used to freeze the parameters of the large model and the small model in the third training phase, train the self-reflective model, and enable the self-reflective model to output a second confidence score using the second feature matrix of the large model as input, generate dynamic routing signals based on the first confidence score and the second confidence score, and build a confidence-driven model switching mechanism by minimizing the overall error of collaborative reasoning.

[0169] An industrial-size model collaborative training device provided in an embodiment of the present application is applicable to the above-mentioned industrial-size model collaborative training method embodiment, and will not be described in detail here.

[0170] An electronic device is provided in an embodiment of the present application. Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, Figure 7 The electronic device shown includes: a processor 71 and a memory 72. The processor 71 and the memory 72 are connected, for example, via a bus 73. Optionally, the electronic device may further include a transceiver 74. It should be noted that in actual applications, the number of transceivers 74 is not limited to one, and the structure of the electronic device does not constitute a limitation on the embodiments of the present application.

[0171] Processor 71 may be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 71 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0172] The bus 73 may include a path for transmitting information between the above components. The bus 73 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus 73 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the figure, but it does not mean that there is only one bus 73 or one type of bus 73.

[0173] The memory 72 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.

[0174] The memory 72 is used to store application code for executing the solution of the present application, and the execution is controlled by the processor 71. The processor 71 is used to execute the application code stored in the memory 72 to implement the content shown in the above method embodiment.

[0175] The present application also provides a computer-readable storage medium, which may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program code. Specifically, the computer-readable storage medium stores program instructions, and the program instructions are used to implement the methods in the above embodiments.

[0176] A computer program product is also provided in an embodiment of the present application, including a computer program. When the computer program is executed by a processor, the technical solution of the above-mentioned method embodiment is implemented. Its implementation principle and technical effect are similar and will not be repeated here.

[0177] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered merely as exemplary, and the true scope and spirit of the present application are indicated by the claims.

[0178] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A collaborative reasoning method for industrial-scale models, characterized by: The method comprises: When obtaining time series data to be processed output by the industrial system, extracting features of the time series data using the small model to obtain a first feature matrix of the time series data, and determining a first confidence score of the first feature matrix; When the first confidence score is less than a preset threshold, extracting features from the time series data using a large model to obtain a second feature matrix of the time series data, and determining a second confidence score of the second feature matrix; When the second confidence score is greater than the first confidence score, using the inference result obtained by the large model based on the second feature matrix as the target inference result; When the second confidence score is less than or equal to the first confidence score, the target inference result is determined by the small model.

2. The method according to claim 1, characterized in that Determining a first confidence score of the first feature matrix includes: Performing dimensionality decomposition on the first feature matrix to obtain multiple feature dimensions; Performing fuzzy processing on each feature dimension to obtain a fuzzy feature matrix corresponding to the first feature matrix; The fuzzy feature matrix is input into a trained fuzzy neural network model to obtain a first confidence score of the first feature matrix; the fuzzy neural network model takes learning a first mapping relationship between the fuzzy feature matrix and the first confidence score as a training goal; the first confidence score is used to characterize the reliability of the small model inference result, and the first mapping relationship is determined based on the inference error of the small model.

3. The method according to claim 2, characterized in that The first mapping relationship is obtained through the following process: Calculating a first error between the inference result of the small model and the corresponding true label using a preset loss function, and normalizing the first error to obtain an inference error of the small model; Using a nonlinear activation function, by minimizing the loss function between the first inference confidence score and the first target confidence score, the inference error of the small model is mapped to the interval [0,1] to obtain the corresponding first confidence score, so as to obtain the first mapping relationship.

4. The method according to claim 2 or 3, characterized in that The fuzzy processing is performed on each feature dimension to obtain a fuzzy feature matrix corresponding to the first feature matrix, including: For each feature dimension, a corresponding Gaussian membership function is used to map the feature dimension to a preset fuzzy membership range; wherein the Gaussian membership function corresponding to each feature dimension is obtained by the following process: during the training process of the fuzzy neural network model, the fuzzy mean and fuzzy variance of the Gaussian membership function are set as learnable parameters; The fuzzy membership of each of the feature dimensions is combined into the fuzzy feature matrix.

5. The method according to any one of claims 1 to 3, characterized in that: Determining a second confidence score of the second feature matrix includes: Converting the second feature matrix into a one-dimensional feature vector along the time dimension; The one-dimensional feature vector is input into the self-reflective model to obtain a second confidence score of the second feature matrix; the self-reflective model is implemented using a fully connected network, and the self-reflective model takes learning a second mapping relationship between the one-dimensional feature vector and the second confidence score as a training goal; the second confidence score is used to characterize the reliability of the inference result of the large model, and the second mapping relationship is determined based on the inference error of the large model.

6. The method according to claim 5, characterized in that The second mapping relationship is obtained through the following process: Calculating a second error between the inference result of the large model and the corresponding true label using a preset loss function, and normalizing the second error to obtain an inference error of the large model; Using a nonlinear activation function, by minimizing the loss function between the second inference confidence score and the second target confidence score, the inference error of the large model is mapped to the interval [0,1] to obtain the corresponding second confidence score, so as to obtain the second mapping relationship.

7. The method according to any one of claims 1 to 3, characterized in that Determining the target inference result by using the small model includes: Using the inference result obtained by the small model based on the first feature matrix as the target inference result; Alternatively, the target inference result is obtained by averaging the inference result obtained by the small model based on the first feature matrix and the inference result obtained by the large model based on the second feature matrix; Alternatively, the target inference result is obtained by performing weighted averaging processing on the inference result obtained by the small model based on the first feature matrix and the inference result obtained by the large model based on the second feature matrix.

8. An industrial-scale model collaborative reasoning device, characterized in that: The device comprises: a first processing module configured to, when acquiring time series data to be processed output by the industrial system, extract features of the time series data using a small model to obtain a first feature matrix of the time series data, and determine a first confidence score of the first feature matrix; a second processing module, configured to, when the first confidence score is less than a preset threshold, perform feature extraction on the time series data using a large model to obtain a second feature matrix of the time series data, and determine a second confidence score of the second feature matrix; a third processing module, configured to use, when the second confidence score is greater than the first confidence score, an inference result obtained by the large model based on the second feature matrix as a target inference result; The third processing module is further configured to determine the target inference result through the small model when the second confidence score is less than or equal to the first confidence score.

9. An industrial-scale model collaboration framework, characterized in that The framework includes small models, large models, and self-reflective models; The small model is configured to extract a first feature matrix of the time series data when receiving the time series data, and output an inference result when a first confidence score of the first feature matrix is greater than or equal to a preset threshold, and trigger the large model to receive the time series data when the first confidence score is less than the preset threshold; The large model is used to output a second feature matrix to the self-reflective model when receiving time series data. The self-reflective model is used to output a second confidence score when receiving the second feature matrix, and when the second confidence score is greater than the first confidence score, the large model outputs an inference result. When the second confidence score is less than or equal to the first confidence score, the target inference result is determined by the small model.

10. A collaborative training method for industrial-scale models, characterized in that: The method is used to train the industrial-scale model collaborative framework according to claim 9 through the following process: In a first training phase, the small model is independently trained, and the parameters of the large model and the self-reflective model are frozen, so that the small model learns to extract the first feature matrix and processes the first feature matrix whose first confidence score is greater than or equal to the preset threshold; In the second training phase, the parameters of the small model are frozen, and only the input layer and output layer of the large model are fine-tuned. Based on the scenario where the first confidence score of the small model triggering intervention is less than the preset threshold, the large model learns to extract the second feature matrix and optimizes the inference error of the large model. In the third training stage, the parameters of the large model and the small model are frozen, and the self-reflective model is trained so that the self-reflective model outputs a second confidence score with the second feature matrix of the large model as input, and a dynamic routing signal is generated according to the first confidence score and the second confidence score. By minimizing the overall error of collaborative reasoning, a confidence-driven model switching mechanism is constructed.

11. An electronic device, characterized in that: include: a processor and a memory communicatively connected to the processor; The memory stores computer-executable instructions; When executing the computer-executable instructions, the processor is configured to implement the method according to any one of claims 1 to 7, and / or the method according to claim 10.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method according to any one of claims 1 to 7, and / or the method according to claim 10.

Citation Information

Patent Citations

  • Model reasoning method and device, electronic equipment, storage medium and program product

    CN118036751A

  • Elevator air pressure height self-correction method and device based on mechanical state

    CN119191010A

Cited By

  • Crane coupling time-varying working condition online fault diagnosis method based on large and small models

    CN121256302A