A distributed multi-center machine learning intelligent decision-making method
Through the combination of fuzzy rule model and multi-center graph matrix, the problems of poor generalization ability and insufficient interpretability in multi-center machine learning are solved, and an efficient and interpretable decision-making system is realized, which is suitable for medical and transportation fields.
Patent Information
- Application Number
- CN202510550322.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing multi-center machine learning methods have poor generalization capabilities in decision-making systems, which are difficult to meet interpretability needs, and have low decision accuracy and efficiency, so they cannot effectively utilize the collaboration and complementary relationship between multi-centers.
The fuzzy rule model and multi-center graph matrix are adopted, and through targeted knowledge transfer strategies, combining the optimization objective function of the fuzzy rule model and the multi-center graph matrix, the knowledge transfer process is adjusted, the structural characteristics and consistency of the multi-center data are maintained, and an independent decision-making system is established.
It improves the generalization ability and decision-making accuracy of the decision-making system, meets the needs of interpretability, and improves the applicability and efficiency in areas such as medical care and transportation.
Smart Images

Figure CN120069100B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to a distributed multi-center machine learning intelligent decision-making method. Background Art
[0002] As a distributed machine learning method, multi-center learning has the advantage that through the collaborative cooperation of multiple independent centers, integrating diverse data and knowledge can effectively improve the performance and generalization ability of the model. By integrating heterogeneous data and knowledge among different centers, the generalization ability, robustness, and flexibility of the model can be improved. It can be applied to various fields such as content / product recommendation, medical auxiliary diagnosis and drug research and development, financial risk assessment and prediction.
[0003] Especially in medical auxiliary diagnosis and drug research and development, in related technologies, the training of multi-center machine learning models generally uses transfer learning and federated learning. Although knowledge transfer and data integration are achieved to a certain extent, the problem of knowledge heterogeneity among multi-centers has not been completely solved, resulting in poor generalization ability of the final decision-making system model. Moreover, general machine learning models are mostly in the "black box mode" with complex internal mechanisms and are difficult to interpret, thus leading to poor applicability in the medical and transportation fields and being difficult to meet the requirements of interpretability.
[0004] Moreover, when existing multi-center learning methods transfer knowledge to a new target center in a multi-center framework, there are usually two aspects of limitations: on the one hand, the problems of consistency and complementarity in the knowledge transfer process have not been effectively solved. Existing models cannot fully maintain the structural characteristics of multi-center data during the transfer process, resulting in insufficient adaptability to the lower-level centers of non-top-level and sub-top-level.
[0005] On the other hand, existing methods rely too much on the multi-center framework during knowledge transfer and cannot establish independent models in the lower-level centers. This limitation leads to poor performance of existing models in the lower-level centers and difficulty in realizing localized knowledge application, and further leads to problems such as low decision accuracy and low efficiency in the decision-making systems finally obtained by the lower-level centers.
[0006] For the above reasons, the decision-making systems of each center finally trained have poor generalization ability, are difficult to meet the requirements of interpretability, and have problems of low decision accuracy and low efficiency.
[0007] For the problems of poor generalization ability, difficulty in meeting the requirements of interpretability, and low decision accuracy and low efficiency in the decision-making systems in related technologies, no effective solutions have been proposed yet. Summary of the Invention
[0008] A distributed multi - center machine learning intelligent decision - making method provided by an embodiment of the present invention solves at least the problems in the related - art decision - making system, such as poor generalization ability, difficulty in meeting the interpretability requirements, low decision - making accuracy, and low efficiency.
[0009] According to one aspect of the embodiments of the present invention, a distributed multi - center machine learning intelligent decision - making method is provided, which is applied to the lower - level centers other than the top - level center and the sub - top - level center in the distributed multi - center. The method includes: obtaining input data that needs to be decision - made; determining the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model; calculating the rule outputs of different fuzzy rules according to the input data and the consequent parameters corresponding to different antecedent parameters; performing weighted calculation according to the membership degrees and the rule outputs corresponding to different fuzzy rules to obtain the output decision data corresponding to the input data; wherein, the fuzzy rule model is calculated by the current lower - level center through knowledge transfer of the fuzzy rule model of the upper - level center in the distributed multi - center, using an optimization objective function, and using a multi - center graph matrix to determine a regularization term to adjust the optimization objective function.
[0010] As an optional solution, before determining the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model, the method further includes: determining the antecedent parameters and consequent parameters for transfer in the lower - level center through knowledge transfer for the remaining rules after partial forgetting of the fuzzy rules of the transferred upper - level center, where the fuzzy rules include corresponding antecedent parameters and consequent parameters; determining the cluster centers according to the remaining rules, calculating the distances between the data samples of the training data of the sub - lower - level center and the cluster centers, and performing clustering to obtain multiple clusters, and removing the empty clusters; calculating the corresponding antecedent parameters based on the cluster centers and the data samples within the clusters using the antecedent parameter optimization formula; updating the optimization objective function according to the optimization objective function of the fuzzy rule model of the upper - level center and using the multi - center graph matrix as a regularization term; updating the transferred consequent parameters according to the updated optimization objective function and the recalculated antecedent parameters.
[0011] As an alternative solution, after partially forgetting the fuzzy rules of the upper-level center for migration through knowledge transfer, the antecedent parameters and consequent parameters for the lower-level center to migrate are determined, including: obtaining the fuzzy rules of the upper-level center through knowledge transfer; determining the upper-level center with the fewest rules and the number of rules; forgetting some rules of other upper-level centers to the same number of rules as the upper-level center with the fewest rules to obtain the remaining rules; using the antecedent parameters and consequent parameters of the remaining rules as the antecedent parameters and consequent parameters after migration for the lower-level center.
[0012] As an alternative solution, according to the optimization objective function of the fuzzy rule model of the upper-level center and using the multi-center graph matrix as a regularization term, the optimization objective function is updated, including: calculating the improved maximum mean discrepancy according to the intermediate variable between the knowledge domain of the upper-level center and the knowledge domain of the current lower-level center; calculating the optimization objective function corresponding to the current lower-level center according to the improved maximum mean discrepancy and the optimization objective function of the fuzzy rule model of the upper-level center, where the optimization objective function includes a regularization term for solving the consequent parameters of the current lower-level center according to the consequent parameters of the upper-level center; correcting the regularization term of the optimization objective function according to the multi-center graph matrix to obtain the regularized optimization objective function.
[0013] According to another aspect of the embodiments of the present invention, a machine learning intelligent decision-making method for a distributed multi-center is further provided, which is applied to the sub-top-level center of the distributed multi-center, including: obtaining input data that needs to be decided; determining the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model; calculating the rule outputs of different fuzzy rules according to the input data and the consequent parameters corresponding to different antecedent parameters; performing weighted calculation according to the membership degrees and the rule outputs corresponding to different fuzzy rules to obtain the output decision data corresponding to the input data; where the fuzzy rule model is obtained by the current sub-top-level center through knowledge transfer calculation using the optimization objective function of the fuzzy rule model of the top-level center of the distributed multi-center.
[0014] As an alternative solution, before determining the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model, the method further includes: obtaining the fuzzy rules of the top-level center by means of knowledge transfer, where the fuzzy rules include antecedent parameters and consequent parameters; forgetting the fuzzy rules according to a preset probability to obtain the remaining rules, where the fuzzy rules include the antecedent parameters and consequent parameters of the top-level center; determining the cluster centers according to the remaining rules, calculating the distances between the data samples of the training data of the sub-top-level center and the cluster centers, and performing clustering to obtain a plurality of clusters, and removing the empty clusters; based on the cluster centers and the corresponding data samples within the clusters, using the antecedent parameter optimization formula to recalculate the corresponding antecedent parameters; calculating the improved maximum mean discrepancy according to the intermediate variables of the knowledge domain of the top-level center and the knowledge domain of the current sub-top-level center; calculating the optimization objective function corresponding to the current sub-top-level center according to the improved maximum mean discrepancy and the optimization objective function of the fuzzy rule model of the top-level center; calculating the consequent parameters of the current sub-top-level center according to the consequent parameters of the remaining rules and the optimization objective function.
[0015] According to another aspect of the embodiments of the present invention, there is also provided a distributed multi-center machine learning intelligent decision-making method, which is applied to the top-level center of the distributed multi-center, and includes: obtaining the input data that needs to be decided; determining the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model; calculating the rule outputs of different fuzzy rules according to the input data and the consequent parameters corresponding to different antecedent parameters; performing weighted calculation according to the membership degrees and the rule outputs corresponding to different fuzzy rules to obtain the output decision data corresponding to the input data; where the fuzzy rule model is trained by the current top-level center through the fuzzy c-means clustering algorithm FCM and a regression algorithm including a penalty term of an insensitive loss function.
[0016] As an alternative solution, before determining the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model, the method further includes: calculating the antecedent parameters of the fuzzy rule model of the top-level center by using the fuzzy c-means clustering algorithm FCM; training and optimizing the consequent parameters of the fuzzy rule model by a regression algorithm including a penalty term of an insensitive loss function; obtaining sample data according to the knowledge data of the top-level center, where the sample data includes training data and verification data; iteratively optimizing the antecedent parameters and the consequent parameters by using the training data and the verification data to obtain the final antecedent parameters and consequent parameters.
[0017] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including: a processor, and a memory storing a program, where the program includes instructions that, when executed by the processor, cause the processor to execute the above-mentioned distributed multi-center machine learning intelligent decision-making method.
[0018] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including a computer program / instructions that, when executed by a processor, implement the above-mentioned distributed multi-center machine learning intelligent decision-making method.
[0019] For the top-level center, sub-top-level center, and other-level centers in the distributed multi-center machine learning intelligent decision-making method provided by the embodiments of the present invention, different training methods are adopted. For non-top-level centers, a knowledge fusion strategy based on targeted knowledge transfer is adopted, which solves the problem of weak generalization ability of the decision-making system.
[0020] For each center, a fuzzy rule model is adopted as the decision-making system, which can effectively realize the interpretability of the decision-making and greatly improve the applicability in fields such as medical treatment and transportation.
[0021] Moreover, when training other-level centers, a multi-center graph matrix with the structural characteristics of multi-center data is combined, which effectively preserves the structural characteristics of multi-center data during the transfer learning process, and finally improves the decision-making accuracy and efficiency of the decision-making system of other-level centers.
[0022] Furthermore, it solves the problems of poor generalization ability, difficulty in meeting the interpretability requirements, low decision-making accuracy, and low efficiency in the decision-making system in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other embodiments according to these drawings without creative efforts.
[0024] Figure 1 It is a flowchart of a distributed multi-center machine learning intelligent decision-making method for a lower-level center in the embodiments of the present invention.
[0025] Figure 2 It is a flowchart of a distributed multi-center machine learning intelligent decision-making method for a sub-top-level center in the embodiments of the present invention.
[0026] Figure 3It is a flowchart of a distributed multi-center machine learning intelligent decision-making method for the top-level center of an embodiment of the present invention.
[0027] Figure 4 It is a schematic diagram of the training architecture of a distributed three-level center of an implementation manner of the present invention.
[0028] Figure 5 It is a schematic diagram of the structure of an electronic device of the present invention. Detailed implementation manners
[0029] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.
[0030] In related technologies, many existing machine learning methods and their variants, such as support vector machines, naive Bayes, TSK fuzzy systems, etc., are mostly based on a single-center architecture. These single-center learning methods concentrate all data and computing resources on a single central node. The performance of the model completely depends on the data quality and computing power of this center, lacking the supplement of external information, which inevitably limits the generalization ability of the model.
[0031] In addition, there are bottlenecks in resource utilization and system stability in the single-center architecture; once the central node fails, the entire system will face interruption, lacking flexibility and scalability. To address the challenges faced by single-center algorithms in practical applications, in recent years, researchers have gradually turned their attention to multi-center learning algorithms.
[0032] The advantages of multi-center learning not only lie in the fact that through the collaborative cooperation of multiple independent centers, integrating diverse data and knowledge, it can effectively improve the performance and generalization ability of the model. Moreover, it can integrate heterogeneous data and knowledge through cooperation between different centers to enhance the generalization ability, robustness, and flexibility of the model. At the same time, it makes full use of distributed computing resources, significantly enhancing the flexibility and scalability of the system. In addition, the multi-center architecture helps to protect data privacy and security because data can be stored on local nodes without the need to be centralized on a central server.
[0033] In the field of multi - center learning and distributed intelligent systems, the existing technologies mainly focus on single - center and early multi - center learning methods. A single - center model relies on a central node to process and store data, and the model performance mainly depends on the data quality and computing resources of this center. However, the limitations of this architecture are obvious: single - center learning cannot effectively process heterogeneous data from different sources, and its dependence on the central node results in insufficient system vulnerability and scalability, making it difficult to meet the diverse requirements in large - scale distributed application scenarios. In addition, in a single - center system, data privacy is also difficult to be effectively protected because all data needs to be centrally processed.
[0034] To address the above problems, researchers have proposed distributed methods such as multi - center learning and federated learning, aiming to collaborate among multiple independent centers to achieve broader data integration and model generalization. Multi - center learning enables the model to utilize dispersed data from multiple sources through cooperation among different centers, thus alleviating the limitations of the single - center mode to a certain extent. However, there are still several significant technical problems in existing multi - center learning methods, mainly including the following points.
[0035] (1) Insufficient integration of knowledge heterogeneity; in a multi - center system, the data of each center often has obvious heterogeneity, including differences in data quality, different feature spaces, inconsistent distributions, etc. Current multi - center methods, such as transfer learning and federated learning, although achieving knowledge transfer and data integration to a certain extent, have not completely solved the problem of knowledge heterogeneity among multi - centers.
[0036] Transfer learning usually relies on transferring knowledge from high - resource centers (such as large centers with high - quality data, the top - level and sub - top - level centers mentioned above) to low - resource centers (such as the lower - level centers mentioned above), but it mainly focuses on the knowledge transfer between a single source and the target, and fails to fully utilize the collaborative information among multiple source centers in the multi - center framework.
[0037] Federated learning allows multiple centers to jointly train a model without sharing the original data, and deals with data heterogeneity by aggregating the model updates of each center. Federated learning achieves collaborative training and data privacy protection among multiple centers. However, federated learning usually assumes that the data quality and importance of each center are comparable, and fails to fully consider the data quality differences and hierarchical relationships among centers. Therefore, the effect of federated learning in dealing with data heterogeneity is limited, and it cannot fully utilize the complementarity among multi - centers, resulting in poor model generalization ability.
[0038] In fact, heterogeneity is unique to individuals but complementary to the whole. Therefore, while maintaining the consistency of multi - center nodes to the greatest extent, their complementarity should be utilized to build a more complete and comprehensive model.
[0039] (2) Limited model interpretability and transparency; in fields such as medicine where extremely high requirements are placed on model transparency and interpretability, although existing deep learning models perform excellently in terms of performance, due to their complex internal mechanisms, they are usually regarded as "black boxes" and lack interpretability. Traditional interpretable models, such as decision trees and logistic regression, although having a certain degree of interpretability, are limited in performance when dealing with high-dimensional, non-linear, and heterogeneous data. When applied to key fields such as healthcare and transportation, the interpretability and transparency of the model are particularly important, especially in fields such as fuzzy systems that have strict requirements for high reliability and transparency. However, existing multi-center learning technologies have obvious deficiencies in this regard.
[0040] Although current methods such as deep learning have powerful pattern recognition capabilities, they are often regarded as "black box" models with complex internal mechanisms and are difficult to interpret. Although some traditional machine learning models (such as decision trees and logistic regression) have a certain degree of interpretability, they are weak in dealing with high-dimensional, non-linear, and heterogeneous data and are difficult to apply in multi-center systems. Therefore, it is difficult for existing technologies to provide sufficient interpretability while ensuring performance.
[0041] (3) Limited knowledge transfer effect; when existing multi-center learning models transfer knowledge to new lower-level centers in a multi-center framework, there are usually two aspects of limitations: on the one hand, the issues of consistency and complementarity in the knowledge transfer process have not been effectively resolved. Existing models cannot fully maintain the structural characteristics of multi-center data during the transfer process, resulting in insufficient adaptability to lower-level centers; on the other hand, existing methods rely too much on the multi-center framework during knowledge transfer and cannot establish independent models in lower-level centers. This limitation causes existing models to perform poorly in lower-level centers and makes it difficult to achieve localized knowledge applications.
[0042] In summary, the defects of the above existing technologies are mainly attributed to the complexity of the multi-center system and the problem of knowledge heterogeneity. Existing technologies mostly only limit to simple knowledge sharing or knowledge transfer from a single source when dealing with multi-center knowledge heterogeneity, and fail to deeply utilize the collaborative and complementary relationships among multi-centers. In addition, most models of multi-center learning adopt complex black box models such as deep neural networks, which although having a certain recognition ability, lack transparent interpretability for high-dimensional heterogeneous data. Finally, existing technologies do not consider the local characteristics and consistency in the multi-center structure during knowledge transfer, resulting in poor knowledge transfer effects and difficulty in meeting the localized needs of lower-level centers.
[0043] Figure 1 It is a flowchart of a distributed multi-center machine learning intelligent decision-making method for the lower-level center of an embodiment of the present invention, as Figure 1As shown, to solve the problems in the decision-making system in the related art, such as poor generalization ability, difficulty in meeting the interpretability requirements, low decision-making accuracy, and low efficiency, an embodiment of the present invention provides a distributed multi-center machine learning intelligent decision-making method applied to the lower-level center. The method includes the following steps.
[0044] Step S101: Obtain the input data that needs to be decision-making.
[0045] Step S102: Determine the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model. The fuzzy rule model is obtained by the current lower-level center through knowledge transfer using the fuzzy rule model of the upper-level center of the distributed multi-center and using the multi-center graph matrix to determine the regularization term to adjust the optimization objective function and then calculating.
[0046] Step S103: Calculate the rule outputs of different fuzzy rules according to the input data and the consequent parameters corresponding to different antecedent parameters.
[0047] Step S104: Perform weighted calculation according to the membership degrees and rule outputs corresponding to different fuzzy rules to obtain the output decision data corresponding to the input data.
[0048] Among them, for the distributed multi-center machine learning intelligent decision-making method provided by the embodiment of the present invention, for non-top-level centers, a knowledge fusion strategy based on targeted knowledge transfer is adopted to solve the problem of poor generalization ability of the decision-making system.
[0049] For each center, a fuzzy rule model is adopted as the decision-making system, which can effectively realize the interpretability of the decision-making and greatly improve the applicability in fields such as medical treatment and transportation.
[0050] Moreover, when training the centers of other levels, the multi-center graph matrix with the structural characteristics of multi-center data is combined, which effectively maintains the structural characteristics of multi-center data during the transfer learning process, and finally improves the decision-making accuracy and efficiency of the decision-making systems of the centers of other levels.
[0051] Furthermore, the problems in the decision-making system in the related art, such as poor generalization ability, difficulty in meeting the interpretability requirements, low decision-making accuracy, and low efficiency, are solved.
[0052] The execution subject of the above steps can be a server, a computer, etc. on which the decision-making system of the lower-level center runs. The above distributed multi-center includes at least three levels of centers, specifically the top-level center, the sub-top-level center, and the lower-level center. The top-level center can be the center in a first-tier city, the sub-top-level center can be the center in a second-tier city, and the lower-level center can be the center in other cities.
[0053] Generally, three levels can meet the requirements of the distributed multi - center architecture in common application scenarios. In some special scenarios, there can be more levels. The additional levels are preferably classified as lower - level centers. In cases where there are more lower - level centers or in artificially set situations, they can also be classified as sub - top - level centers.
[0054] The training methods of the centers at the above - mentioned three levels are different. When knowledge transfer is carried out among the centers at each level using the method of this embodiment during training, it can effectively improve the generalization ability of the decision - making systems corresponding to each center, especially the lower - level centers.
[0055] The above - mentioned decision - making system can be understood as the system for making decisions by the center in a specific application scenario. Its main core is the above - mentioned fuzzy rule model, combined with the configuration in the specific application scenario, to form a system for auxiliary intelligent decision - making in the specific application scenario.
[0056] The above - mentioned fuzzy rule model can be a Takagi - Sugeno - Kang (TSK) model. The TSK model combines expert knowledge and data - driven methods through fuzzy rules, and can clearly express the relationship between input and output. This model not only has good interpretability, but also can handle non - linear and complex data patterns, and is suitable for the scenario of multi - center data fusion. Therefore, using the TSK model helps to construct a model with both high performance and interpretability, meeting the dual requirements of the medical field for model performance and interpretability.
[0057] The TSK model consists of two key parts: rule antecedent parameters and rule consequent parameters. Rule antecedent parameters (Rule Antecedents) are the conditional part of the fuzzy rule. In the TSK model, it is usually a combination of fuzzy propositions composed of input variables and fuzzy sets. These fuzzy sets are defined by membership functions, such as Gaussian membership functions, triangular membership functions, etc., to determine the degree to which an input variable belongs to a certain fuzzy set.
[0058] Rule consequent parameters (Rule Consequents) are the conclusion part of the fuzzy rule. In the TSK model, it is usually a linear combination of input variables (there are also other forms). This linear form gives the TSK model certain advantages in dealing with complex non - linear systems and can effectively perform tasks such as function approximation.
[0059] The rule antecedent parameters may include the center and width of the membership function. The rule consequent parameters may include linear combination coefficients, etc. The determination of these parameters is the key to TSK model training. The training process usually uses optimization algorithms such as least squares method and gradient descent method to adjust parameters based on given input-output data pairs to achieve the optimization objective function, so that the output of the model is as close to the actual output value as possible, thereby improving the accuracy and generalization ability of the model.
[0060] The trained fuzzy rule model only needs to use the corresponding antecedent parameters and consequent parameters to perform corresponding calculations on the input data to obtain the corresponding output decision data. Specifically, the membership of the input data in different fuzzy rules is determined according to the antecedent parameters of different fuzzy rules in the fuzzy rule model; the rule outputs of different fuzzy rules are calculated according to the input data and the consequent parameters corresponding to different antecedent parameters; and the output decision data corresponding to the input data is obtained by performing weighted calculations based on the membership and rule outputs of different fuzzy rules.
[0061] The above fuzzy rule model is calculated by the current lower-level center through the fuzzy rule model of the upper-level center of the distributed multi-center, using the optimization objective function for knowledge transfer, and using the multi-center graph matrix as a regular term to adjust the optimization objective function. The final fuzzy rule model of the decision-making system improves the generalization ability and meets the requirements of explainability, as well as the problems of low decision accuracy and low efficiency.
[0062] This lower-level center adopts a knowledge fusion strategy based on targeted knowledge transfer to solve the problem of weak generalization ability of the decision-making system. For the current lower-level center, a fuzzy rule model is used as the decision-making system, which can effectively realize the interpretability of decisions and greatly improve the applicability in fields such as medical care and transportation. Moreover, when training the current lower-level center, a multi-center graph matrix with the structural characteristics of multi-center data is combined to effectively fully maintain the structural characteristics of multi-center data in the transfer learning process, and ultimately the decision-making system of centers at other levels improves the accuracy and efficiency of decisions.
[0063] As an alternative solution, according to the antecedent parameters of different fuzzy rules of the fuzzy rule model, the input data is determined. Before the membership degrees of different fuzzy rules, the method further includes: by means of knowledge transfer, for the remaining rules after partially forgetting the fuzzy rules of the transferred upper-level center, determining the antecedent parameters and consequent parameters for the lower-level center to transfer, where the fuzzy rules include corresponding antecedent parameters and consequent parameters; determining the clustering centers according to the remaining rules, calculating the distances between the data samples of the training data of the sub-lower-level center and the clustering centers, and performing clustering to obtain multiple clustering clusters, and removing the empty clusters; based on the clustering centers and the corresponding data samples within the clusters, using the antecedent parameter optimization formula to calculate the corresponding antecedent parameters; according to the optimization objective function of the fuzzy rule model of the upper-level center, and using the multi-center graph matrix as a regularization term, updating the optimization objective function; according to the updated optimization objective function, combined with the recalculated antecedent parameters, updating the transferred consequent parameters.
[0064] The above-mentioned way of knowledge transfer can be transductive transfer, which does not rely on the label information of the target domain, but realizes knowledge transfer by directly learning the data distribution of the target domain. And during knowledge transfer, by means of partial forgetting, irrelevant or outdated data can be filtered out to avoid interfering with the learning process, thereby improving the learning efficiency.
[0065] The above-mentioned upper-level center at least includes the sub-top-level center and the top-level center. During knowledge transfer, the fuzzy rules of the above two levels will be transferred simultaneously, so that the current lower-level center can simultaneously use the top-level center and the sub-top-level center to guide the establishment of the fuzzy rule model, making the fuzzy rule model of the lower-level center have better generalization ability. And it greatly reduces the training optimization time and improves the training efficiency.
[0066] Specifically, as an alternative embodiment, by means of knowledge transfer, after partially forgetting the fuzzy rules of the transferred upper-level center, determining the antecedent parameters and consequent parameters for the lower-level center to transfer includes: by means of knowledge transfer, obtaining the fuzzy rules of the upper-level center; determining the upper-level center with the least number of rules and the number of rules; forgetting some rules of other upper-level centers to the same number of rules as the upper-level center with the least number of rules to obtain the remaining rules; using the antecedent parameters and consequent parameters of the remaining rules as the antecedent parameters and consequent parameters after the lower-level center is transferred.
[0067] Performing random forgetting of fuzzy rules in the above way is for the convenience of calculation. Since the upper-level center corresponding to the lower-level center at least includes the top-level center and the sub-top-level center, in a distributed multi-center system, as the levels go from the top down, the number of centers at each level will increase.
[0068] If the proportion of random forgetting of the lower-level centers is set to be fixed, it will cause the number of fuzzy rules of each upper-level center to change with the number of fuzzy rules of the center itself, which is not convenient for subsequent calculations. Therefore, in this embodiment, the center with the least number of rules is first determined, and then some rules of the remaining centers will be forgotten until they have the same number of rules as the center with the least number of rules, which can greatly improve the training efficiency of the subsequent fuzzy rule model.
[0069] The fuzzy rules after forgetting are the remaining rules mentioned above. The antecedent parameters and consequent parameters of the remaining rules are used as the antecedent parameters and consequent parameters after migration of the lower-level centers, and thus become the basis for optimization. It avoids a series of time-consuming and inefficient steps such as creating and training the fuzzy rule models of a large number of lower-level centers.
[0070] As an optional embodiment, the corresponding antecedent parameters are calculated according to the migrated antecedent parameters and the antecedent parameter optimization formula, including: optimizing the antecedent parameters in the fuzzy rules by using the antecedent parameter optimization formula to obtain the corresponding antecedent parameters; the antecedent parameters include the center of the membership function of the fuzzy rule model and the kernel width ; the antecedent parameter optimization formula is as follows.
[0071]
[0072]
[0073] In the formula, k represents the number of clusters of the fuzzy rule model, x ji is the i-th training sample of the j-th cluster of the training data of the lower-level center, N k represents the number of samples of k clusters, and h is the scale parameter.
[0074] Through the above antecedent parameter optimization formula, the migrated antecedent parameters can be optimized by using the training data of the lower-level centers, so that the optimized antecedent parameters have better applicability and accuracy for the decision-making of this lower-level center.
[0075] As an alternative embodiment, according to the optimization objective function of the fuzzy rule model of the upper-level center and using the multi-center graph matrix as a regularization term, the optimization objective function is updated, including: calculating the improved maximum mean discrepancy according to the intermediate variable between the knowledge domain of the upper-level center and the knowledge domain of the current lower-level center; calculating the optimization objective function corresponding to the current lower-level center according to the improved maximum mean discrepancy and the optimization objective function of the fuzzy rule model of the upper-level center, where the optimization objective function includes a regularization term for solving the consequent parameters of the current lower-level center according to the consequent parameters of the upper-level center; and correcting the regularization term of the optimization objective function according to the multi-center graph matrix to obtain the regularized optimization objective function.
[0076] The above intermediate variable can be the intermediate variable of the above transductive transfer learning, and the intermediate variable can capture the shared information or structure between the source domain and the target domain. In this embodiment, the current lower-level center is the target domain and the upper-level center is the source domain.
[0077] The above improved maximum mean discrepancy is also PMMD. Here, it is necessary to clarify the concepts of the Maximum Mean Discrepancy (MMD) and its variant PMMD (Projected Maximum Mean Discrepancy).
[0078] MMD is a distance metric commonly used to measure the difference between two probability distributions and is often used in transfer learning to reduce the distribution difference between the source domain and the target domain data. PMMD is an improved form of MMD that calculates the distribution difference by projecting high-dimensional features into a low-dimensional space, thus effectively dealing with the problems of high data dimension and domain drift. MMD and PMMD transfer knowledge from the source domain D s to the target domain D t The theoretical formulas are as follows.
[0079]
[0080]
[0081] Among them, ϕ(·) and P represent the mapping function and the projection vector respectively.
[0082] The above calculates the improved maximum mean discrepancy through the intermediate variable, and then determines the optimization objective function through the maximum mean discrepancy, introducing the relationship between the source domain class information and the target domain class information of knowledge transfer into the optimization objective function, so that the optimization objective function can retain the knowledge related to the class information of the upper-level center in relevant scenarios to ensure the effectiveness of the transfer.
[0083] The above multi-center graph matrix is composed of an intra-center graph matrix and an inter-center graph matrix. By modifying the regularization term of the optimization objective function through the multi-center graph matrix, a regularized optimization objective function is obtained, which can make the optimization objective function maintain the consistency among multi-center data and mine the complementary information between different centers during the optimization process.
[0084] Specifically, the method for calculating the above improved maximum mean discrepancy and using the improved maximum mean discrepancy to optimize and determine the optimization objective function corresponding to the current lower-level center is as follows.
[0085] As an optional embodiment, the improved maximum mean discrepancy is calculated according to the intermediate variable between the knowledge domain of the upper-level center and the knowledge domain of the current lower-level center, including: the calculation formula of the intermediate variable according to the intermediate variable between the knowledge domain of the upper-level center and the knowledge domain of the current lower-level center is as follows.
[0086]
[0087] In the formula, D s is the source domain, including the knowledge domain of the upper-level center, D t is the target domain, including the knowledge domain of the current lower-level center, N s is the number of samples in the source domain, N t is the number of samples in the target domain, x g,s is the training data after the antecedent parameter conversion of the original data in the source domain, x g,t is the training data after the antecedent parameter conversion of the original data in the target domain, T is the matrix transpose; through the intermediate variable, the improved maximum mean discrepancy is calculated as follows.
[0088]
[0089] In the formula, PMMD represents the improved maximum mean discrepancy, D t(a) is the a-th target domain, is the consequent parameter of the j-th cluster of the current lower-level center corresponding to the a-th target domain.
[0090] As an optional embodiment, according to the improved maximum mean discrepancy and the optimization objective function of the fuzzy rule model of the upper-level center, the optimization objective function corresponding to the current lower-level center is calculated, including: according to the following formula, using the improved maximum mean discrepancy and the optimization objective function of knowledge transfer, the optimization objective function corresponding to the current lower-level center is obtained.
[0091]
[0092] In the formula, , , are the regularization parameters of the current lower-level center, c is the number of clusters, and j belongs to 1, 2, 3... c.
[0093] Specifically, the above-mentioned intra-center graph matrix represents the structural relationship within a single center, that is, when mapped to the low-dimensional label space, the data should maintain its neighborhood relationship in the original sample space. The inter-center graph, on the other hand, describes the structural relationship between different centers, indicating that when data of different modalities represent the same content or theme, they should be similar in the upper-level semantics. The specific calculation method is as follows.
[0094] As an alternative embodiment, before obtaining the regularized optimization objective function by modifying the regular term of the optimization objective function according to the multi-center graph matrix, the method further includes: calculating the intra-center graph matrix of the upper-level center through the following formula.
[0095]
[0096] Among them, is the connection node representing the intra-center graph matrix at the m-th row and n-th column and The weight of the edge between them; G uu represents the intra-center graph matrix of the u-th upper-level center; represents The K-nearest neighbor of, represents the Gaussian kernel bandwidth of the internal structure of the center; the intra-center structural relationship is maintained by minimizing the following formula.
[0097]
[0098] In the formula, represents the label of the m-th sample data in the u-th upper-level center; represents the label data label of the n-th sample data in the u-th upper-level center; the inter-center graph matrix of the upper-level center is calculated through the following formula.
[0099]
[0100] G uv represents the inter-center graph matrix between the u-th center and the v-th center, and its m-th row and n-th column represent the connection node of the inter-center graph matrix and The weight of the edge between them is ; the inter-center structural relationship is maintained by minimizing the following formula.
[0101]
[0102] According to the intra-center graph matrix and the inter-center graph matrix, the multi-center graph matrix is calculated through the following formula.
[0103]
[0104] Wherein, G is the multi - center graph matrix, , and β is the weight parameter.
[0105] As an alternative embodiment, the regularization term of the optimization objective function is corrected according to the multi - center graph matrix to obtain the regularized optimization objective function, including: being obtained by correcting the regularization term according to the multi - center graph matrix.
[0106]
[0107] Wherein, represents the u - th element in represents the u - th element on the diagonal in is the Laplacian matrix corresponding to the u - th mode and the v - th mode, B uv represents a diagonal matrix whose diagonal elements are ; the regularized optimization objective function is.
[0108]
[0109] Wherein, TC is the current lower - level center, , , , are all regularization parameters of the current lower - level center, c is the number of clusters, and j belongs to 1, 2, 3... c.
[0110] Figure 2 is the flowchart of a distributed multi - center machine - learning intelligent decision - making method for the sub - top - level center of the embodiments of the present invention. As Figure 2 shown, to solve the problems in the related - art decision - making system, such as poor generalization ability, difficulty in meeting the interpretability requirements, low decision - making accuracy, and low efficiency, the embodiments of the present invention provide a distributed multi - center machine - learning intelligent decision - making method applied to the sub - top - level center. The method includes the following steps.
[0111] Step S201, obtain the input data to be decided.
[0112] Step S202, determine the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model, wherein the fuzzy rule model is obtained by using the optimization objective function for knowledge transfer calculation of the fuzzy rule model of the distributed multi - center top - level center of the current sub - top - level center.
[0113] Step S203: Calculate the rule outputs of different fuzzy rules based on the input data and the consequent parameters corresponding to different antecedent parameters.
[0114] Step S204: Perform weighted calculation based on the membership degrees and rule outputs corresponding to different fuzzy rules to obtain the output decision data corresponding to the input data.
[0115] Among them, for the sub-top level center, the distributed multi-center machine learning intelligent decision-making method provided by the embodiments of the present invention adopts a knowledge fusion strategy based on targeted knowledge transfer, which solves the problem of weak generalization ability of the decision-making system.
[0116] For each center, a fuzzy rule model is adopted as the decision-making system, which can effectively realize the interpretability of the decision-making, and greatly improve the applicability in fields such as medical treatment and transportation.
[0117] Furthermore, it solves the problems of the decision-making system in the related technologies, such as poor generalization ability, difficulty in meeting the interpretability requirements, low decision-making accuracy, and low efficiency.
[0118] The execution subject of the above steps can be a server, a computer, etc. on which the decision-making system of the sub-top level center runs.
[0119] By integrating the knowledge of the top level center, the sub-top level center can efficiently utilize local data and quickly optimize a personalized model suitable for this region, thereby realizing the efficient migration of the model. Specifically as follows.
[0120] Before determining the input data according to the antecedent parameters of different fuzzy rules of the fuzzy rule model in the above step S201 and the membership degrees of different fuzzy rules, the method further includes.
[0121] Obtain the fuzzy rules of the top level center through the way of knowledge transfer, where the fuzzy rules include antecedent parameters and consequent parameters.
[0122] Forget the fuzzy rules according to a preset probability to obtain the remaining rules, where the fuzzy rules include the antecedent parameters and consequent parameters of the top level center.
[0123] Determine the cluster centers according to the remaining rules, calculate the distances between the data samples of the training data of the sub-top level center and the cluster centers, and perform clustering to obtain multiple clusters, and eliminate the empty clusters.
[0124] Based on the cluster centers and the corresponding data samples within the clusters, use the antecedent parameter optimization formula to recalculate the corresponding antecedent parameters.
[0125] Calculate the improved maximum mean discrepancy according to the intermediate variables between the knowledge domain of the top level center and the knowledge domain of the current sub-top level center.
[0126] Calculate the optimization objective function corresponding to the current sub-top level center according to the maximum average difference improvement and the optimization objective function of the fuzzy rule model of the top level center.
[0127] Calculate the consequent parameters of the current sub-top level center according to the consequent parameters of the remaining rules and the optimization objective function.
[0128] The training method of the sub-top level center is similar to that of the lower level center, mainly based on knowledge transfer, and adaptively optimized according to the situation. Compared with re-creating a model and using training data for training, it has higher efficiency and faster speed.
[0129] Specifically, before training, first obtain the fuzzy rules of the top level center through knowledge transfer and perform random forgetting. Different from the lower level center, the sub-top level center adopts a fixed frequency, such as 10%, for forgetting during forgetting.
[0130] This is because the number of top level centers is small and the amount of knowledge is large, which leads to a large data magnitude of the top level centers, and the data volume gap between different top level centers is not significant enough for the data volume of the top level centers themselves. Therefore, random forgetting by a fixed ratio is more efficient.
[0131] After random forgetting, use the remaining rules to determine the clustering centers, calculate the distances between the data samples of the training data of the sub-top level center and the clustering centers, and perform clustering to obtain multiple clustering clusters, and eliminate the empty clusters; based on the clustering centers and the corresponding data samples within the clusters, use the antecedent parameter optimization formula to recalculate the corresponding antecedent parameters. The calculation method of the antecedent parameters is the same as that of the antecedent parameters of the lower level center, and the corresponding formula can also be applied to the corresponding data.
[0132] Since there is only the top level center above the sub-top level center, the number of top level centers is very small, and the data heterogeneity between them is not high. Therefore, it is not very meaningful to use a multi-center graph matrix. Therefore, when improving the optimization objective function of the sub-top level center, only calculate the optimization objective function corresponding to the current sub-top level center according to the maximum average difference improvement and the optimization objective function of the fuzzy rule model of the top level center. Then calculate the consequent parameters of the current sub-top level center according to the consequent parameters of the remaining rules and the optimization objective function.
[0133] Figure 3 It is a flowchart of a distributed multi-center machine learning intelligent decision-making method for the top level center of an embodiment of the present invention, as Figure 3As shown in the figure, in order to solve the problems in the decision-making system in the related art, such as poor generalization ability, difficulty in meeting the interpretability requirements, low decision-making accuracy, and low efficiency, the embodiments of the present invention provide a distributed multi-center machine learning intelligent decision-making method applied to the top-level center. The method includes the following steps.
[0134] Step S301, obtain the input data that needs to be decision-making.
[0135] Step S302, determine the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model, where the fuzzy rule model is trained by the top-level center through the fuzzy c-means clustering algorithm FCM and the regression algorithm including the penalty term of the insensitive loss function.
[0136] Step S303, calculate the rule outputs of different fuzzy rules according to the input data and the consequent parameters corresponding to different antecedent parameters.
[0137] Step S304, perform weighted calculation according to the membership degrees and rule outputs corresponding to different fuzzy rules to obtain the output decision data corresponding to the input data.
[0138] Among them, for the top-level center, the decision-making method of the distributed multi-center machine learning decision-making system provided by the embodiments of the present invention adopts a fuzzy rule model as the decision-making system, which can effectively realize the interpretability of decision-making and greatly improve the applicability in fields such as medical treatment and transportation.
[0139] And use the regression algorithm with the penalty term of the insensitive loss function to train and optimize the antecedent parameters and consequent parameters, so that the fuzzy rule model has better robustness and generalization ability.
[0140] Furthermore, it solves the problems in the decision-making system in the related art, such as poor generalization ability, difficulty in meeting the interpretability requirements, low decision-making accuracy, and low efficiency.
[0141] The execution subject of the above steps can be a server, a computer, etc. on which the decision-making system of the top-level center runs.
[0142] The top-level center has no upper-level center, and the establishment of its fuzzy rule system mainly depends on its own training data. In this embodiment, the TSK fuzzy rule model of the top-level center is a method based on fuzzy logic, which consists of three key parts: rule antecedent, rule consequent, and model parameters. The rule antecedent describes the fuzzy conditions of the input variables, and the rule consequent defines the calculation method of the output variables. When appropriate rule antecedents and rule consequents are provided, the TSK model can effectively capture complex non-linear relationships.
[0143] The output of the TSK fuzzy system can be interpreted as a linear model in the implicit mapping feature space, thus meeting the requirement of output interpretability in some application scenarios, such as medical systems and transportation systems. Specifically, the training problem of the TSK fuzzy model can be transformed into a learning problem involving the parameters of the linear model. When solving it, various methods can be used to determine, including the global least squares method (GL), the local least squares method (LL), ridge regression (RR), L1 ε-insensitive loss (L1), and L2 ε-insensitive loss (L2), etc.
[0144] Preferably, the algorithm of ε-insensitive loss is adopted in this embodiment. Specifically, it can be L1 ε-insensitive loss (L1) or L2 ε-insensitive loss (L2). It has strong robustness, strong interpretability, better generalization ability, and is more accurate and in line with the actual situation.
[0145] As an alternative embodiment, according to the antecedent parameters of different fuzzy rules of the fuzzy rule model, to determine the input data, before the membership degrees of different fuzzy rules, the method further includes: calculating the antecedent parameters of the fuzzy rule model of the top-level center by using the fuzzy c-means clustering algorithm FCM; training and optimizing the consequent parameters of the fuzzy rule model by using a regression algorithm including a penalty term of the insensitive loss function; obtaining sample data according to the knowledge data of the top-level center, where the sample data includes training data and validation data; iteratively optimizing the antecedent parameters and consequent parameters by using the training data and validation data to obtain the final antecedent parameters and consequent parameters.
[0146] When training the TSK fuzzy rule model, the antecedent parameters of the fuzzy rule model of the top-level center can be calculated by using the fuzzy c-means clustering algorithm FCM; then, the consequent parameters of the fuzzy rule model can be trained and optimized by using a regression algorithm including a penalty term of the insensitive loss function, such as the regression algorithms of the above L1 ε-insensitive loss (L1) function and L2 ε-insensitive loss (L2) function. So that the finally trained fuzzy rule model has better generalization ability, higher robustness, and interpretability.
[0147] During specific training, first obtain sample data according to the knowledge data of the top-level center, where the sample data includes training data and validation data. The antecedent parameters are trained by using the above FCM fuzzy c-means clustering algorithm with the training data, and the consequent parameters are trained by using the regression algorithm of the insensitive loss function.
[0148] Then, use the verification data to verify the antecedent parameters and consequent parameters, and incorporate the incorrect samples into the error set. Conduct multiple iterations of optimization, and select the antecedent parameters and consequent parameters with the smallest error set at the iteration number with the smallest error set as the final antecedent parameters and consequent parameters. The smallest error set indicates the smallest error in this iteration of training, and the corresponding antecedent parameters and consequent parameters are the most accurate.
[0149] As an alternative embodiment, iteratively optimize the antecedent parameters and consequent parameters through training data to obtain the final antecedent parameters and consequent parameters, including: the antecedent parameters include the centers of the membership functions of the fuzzy rule model and the kernel widths ; optimize the antecedent parameters through the antecedent parameter optimization formula; optimize the consequent parameters through the following formula.
[0150]
[0151] In the formula, is the consequent parameter of the top-level center at the t-th iteration, the consequent parameter of the top-level center at the (t - 1)-th iteration at the t-th iteration; and are the regularization parameters of the top-level center, represents the consequent parameter corresponding to the j-th cluster at the t-th iteration , and respectively represent the data and labels of the i-th training sample regarding the j-th class; verify the optimized antecedent parameters and consequent parameters according to the verification data, take those that do not match the actual labels of the corresponding verification samples as incorrect samples, and incorporate the incorrect samples into the error set; increment the iteration number by one, continue the iteration until the iteration number reaches the preset number, and select the antecedent parameters and consequent parameters corresponding to the iteration round with the least error set as the final antecedent parameters and consequent parameters.
[0152] It should be noted that this embodiment also provides an alternative implementation manner, which will be described in detail below. This implementation manner belongs to the technical field of multi-center learning and distributed intelligent systems, and particularly focuses on the knowledge fusion and data consistency issues in fuzzy systems.
[0153] The multi - center TSK fuzzy system (KCF - LSG - MTSK) is mainly applied to address the challenges brought by knowledge heterogeneity in practical fields such as medicine. It realizes the knowledge integration of the base center and the auxiliary center through a knowledge calibration fusion strategy, and maintains the consistency and complementarity of multi - center data through the label - space graph regularization method, thereby enhancing the interpretability and adaptability of the system. Such a system has significant advantages in multi - center data integration, is suitable for complex distributed data analysis, and has broad application prospects especially in fields that require distributed processing and data privacy protection.
[0154] To address the challenges faced in practical applications such as medicine, this embodiment provides a specific multi - center architecture. Figure 4 It is a schematic diagram of the training architecture of the distributed three - level center of the embodiment of the present invention, as Figure 4 shown. This architecture consists of a top - level center (referred to as the base center) with extremely high data quality and reliability and a sub - top - level center (referred to as the auxiliary center) with relatively high data quality. The base center and the auxiliary center cooperate with each other to form a multi - center system, jointly guiding the establishment and model training of the new lower - level center (referred to as the target center).
[0155] In a practical scenario in the medical field, the base center usually consists of large hospitals located in first - tier cities, which have a doctor team with excellent professional qualities and rich clinical experience, especially having unique (representative) diagnostic capabilities when dealing with rare diseases.
[0156] The auxiliary centers are composed of relatively large hospitals in second - tier cities. They can combine the local actual situation, give full play to the advantages of adapting to local conditions, and provide effective diagnosis and treatment plans for these diseases.
[0157] The target center usually refers to small hospitals or clinics in small cities of the third and fourth tiers. The main challenge is the lack of sufficient understanding and diagnosis and treatment experience for certain diseases. By integrating the experience and knowledge of the base center and the auxiliary center, the target center can quickly establish a clear understanding and diagnostic ability for specific diseases, thereby improving the medical service level. This multi - center cooperation model helps the sharing and dissemination of medical knowledge, and promotes the optimization and improvement of the overall medical system.
[0158] To address the problem of "insufficient integration of knowledge heterogeneity" in existing multi - center learning methods, this embodiment proposes the "Knowledge Calibration Fusion Strategy", which makes full use of the hierarchical structure and data quality differences between the top - level center and the sub - top - level center, effectively fuses the heterogeneous data and knowledge of each center, and thus constructs a multi - center architecture to guide the establishment of the lower - level center.
[0159] To address the issue of "limited model interpretability and transparency" in existing multi - center learning methods, this embodiment adopts the Takagi - Sugeno - Kang (TSK) fuzzy rule model as the underlying model and optimizes it on this basis to improve the interpretability of the multi - center architecture. The TSK model combines expert knowledge and data - driven methods through fuzzy rules, and can clearly express the relationship between input and output. This model not only has good interpretability, but also can handle non - linear and complex data patterns, and is suitable for the scenario of multi - center data fusion. Therefore, adopting the TSK model helps to construct a model with both high performance and interpretability, meeting the dual requirements of model performance and interpretability in the medical field.
[0160] To address the issue of "limited knowledge transfer effect" in existing multi - center learning methods, previous work has tried to predict by inputting the samples of the target center into each center in the multi - center architecture, constructing a weight matrix reflecting the importance of each center, and using the prediction results and weight matrix of each center as the basis for determining the final samples. The advantage of this method is that it can quickly establish a model for the target center, but its drawback is that it fails to establish a local model and always relies on the multi - center architecture [MKTC - R0T] during the decision - making process. In addition, this method does not fully utilize the consistency and complementarity among multi - center data during the knowledge transfer process.
[0161] Based on this, this embodiment proposes a label - space graph regular term in the label space based on the theory that "if the goal is to maximize the separability between classes, then the label space can indeed be regarded as an optimal subspace". This regular term is composed of an intra - center graph and an inter - center graph, aiming to maintain the consistency among multi - center data and mine the complementary information between different centers. By introducing the label - space graph regularization term during the model training process, while maintaining the data structure within each center, it promotes knowledge sharing and fusion between different centers.
[0162] In summary, this embodiment proposes a new multi - center learning framework through "Knowledge Calibration Fusion (KCF)" and "Label - Space Graph (LSG) regularization" techniques, systematically addressing the deficiencies of existing methods in knowledge integration, model interpretability, and knowledge transfer.
[0163] First, the KCF strategy makes full use of the knowledge heterogeneity among centers. By implementing a random knowledge forgetting and knowledge transfer mechanism between the base center (BC) and the auxiliary center (AC), it realizes the complementary integration of data quality and hierarchical structure differences, thereby enhancing the generalization ability of the model.
[0164] Secondly, the present invention constructs a multi-center learning framework based on the TSK fuzzy system, expresses the input-output relationship based on rules, provides a clear decision-making process, and combines the data consistency graph structure constructed by LSG regularization, enabling the model to maintain control over local and global data consistency during the knowledge transfer process while ensuring the interpretability of decision-making, meeting the high interpretability requirements of medical diagnosis, intelligent transportation, etc.
[0165] Finally, in terms of knowledge transfer, the present invention transforms the knowledge transfer between multi-centers into a structure-preserving problem in the low-dimensional label space through the LSG regularization mechanism, ensures data consistency during the transfer process, avoids the problems of knowledge loss or insufficient transfer in traditional methods, and constructs a localized model on the target center to enhance the adaptability of the target center to local data. By improving the generalization ability, interpretability, and transfer efficiency of multi-center learning, the invention is particularly suitable for complex distributed systems with strong heterogeneity and high requirements for data privacy protection, and has significant practical value and application prospects.
[0166] Specifically, in this embodiment, KCF-LSG-MTSK incorporates the data-level differences between the base center and the auxiliary center into the model through a knowledge calibration fusion strategy, constructing a multi-center architecture that integrates heterogeneity and complementarity, thereby enhancing the collaborative effect between multi-centers. Secondly, using the interpretability of the TSK fuzzy system as the model basis, this paper realizes the transparent processing of high-dimensional and complex data through fuzzy rules, enabling the model to not only have good performance but also provide a clear reasoning mechanism. Finally, this embodiment introduces a graph regularization strategy in the label space. By constructing an intra-center graph and an inter-center graph, it further explores the data consistency and complementarity between centers, promoting the sharing and efficient transfer of knowledge.
[0167] Next, the training methods of the base center, auxiliary center, and target center in this embodiment will be described in detail.
[0168] I. Training of the base center.
[0169] The light green area in Figure 1 represents the training process of the base center. In a multi-center scenario, the base center has a large amount of high-quality data. Therefore, it is extremely crucial and necessary to build a model with high robustness and high generalization performance. The quality of its model directly affects the performance of the auxiliary center, the target center, and the entire multi-center framework. In previous work, in order to achieve rapid model deployment, various techniques were adopted to accelerate model training, such as using K-means to replace the FCM algorithm to calculate the antecedent parameters, and adopting a faster 0-TSK model. However, these strategies were achieved at the cost of sacrificing the accuracy and generalization ability of the model. In fact, as the cornerstone of the multi-center framework, the base center should prioritize developing a model with high accuracy and high generalization ability, even if it requires more computing resources and training time.
[0170] Based on this, this embodiment uses the most conventional FCM method to calculate the antecedent parameters of the fuzzy rules, and uses the rule consequent learning based on the -insensitive loss function (L2-TSK-FS) to find the optimal consequent parameters. In addition, inspired by the natural tendency of humans to continuously enhance and correct their understanding by using the "error set" (ES) during the learning process, an iterative method is designed to optimize the parameters of the fuzzy rules.
[0171] Specifically, the training data is divided into K folds, where K - 1 folds are used to train the model, and the remaining one fold is used to verify the model, and all samples with prediction errors are recorded, that is, the ES. For the samples within the error set in the th round of iteration, they are re-assigned to the nearest homogeneous center according to their labels. Then, the antecedent parameters of the fuzzy rules and are corrected using the following formula, and the antecedents and consequents of the rules corresponding to the possible empty clusters are removed.
[0172]
[0173]
[0174] Among them, the antecedent parameters include the center of the membership function and the kernel width , k represents the number of clusters of the fuzzy rule model, x ji is the i-th training sample of the j-th cluster of the training data, and N k represents the number of samples in the k-th cluster.
[0175] Similarly, the consequent parameters of the base center are not solved using the L2-TSK-FS method in each iteration process. The L2-TSK-FS method is to use L2 - A regression algorithm with an insensitive loss function, which only uses this method for the first time. In the t-th round of iteration, Optimize by mean square error combined with historical knowledge Its theoretical formula can be expressed as.
[0176]
[0177] Where, and Are the regularization parameters of the base center BC, Represents the consequent parameter corresponding to the j-th class in the t-th round of iteration , and Respectively represent the data and label of the i-th sample regarding the j-th class, c is the number of clusters, and j belongs to 1, 2, 3... c.
[0178] Taking the partial derivative of the above formula, setting it to 0 and simplifying, we can get.
[0179]
[0180] Where, Represents the identity matrix. The optimization of the above antecedent and consequent parameters is also called knowledge calibration. In addition, during the iteration process, two iteration termination conditions are set, namely the maximum number of iterations And / or the iteration termination threshold . When the maximum number of iterations condition terminates the iteration, the antecedent and consequent parameters of the round with the smallest error set are selected as the final antecedent and consequent parameters.
[0181] II. Training of the auxiliary center.
[0182] The auxiliary center plays a key role in the multi-center architecture. It not only makes up for the gap of the base center in the customized model, ensuring that the model can better fit the regional characteristics, but also improves the robustness of the entire multi-center system. By integrating the knowledge of the base center, the auxiliary center can efficiently utilize local data and quickly optimize a personalized model suitable for this region, thus realizing the efficient migration of the model. The processes of learning and forgetting are indeed complementary to each other, and they each play a crucial role in constructing and retaining useful representations of experience. Learning involves acquiring and integrating new information, which can lead to a more comprehensive understanding of the environment. On the contrary, forgetting is like a filter that eliminates irrelevant or outdated information that may interfere with the cognitive process. Learning and forgetting work together to ensure that the cognitive system is both adaptable and efficient. The processes of learning and forgetting are complementary to each other, and both contribute to constructing and retaining useful representations of experience.
[0183] Therefore, based on the above idea, before each auxiliary center borrows knowledge from the basic center, the knowledge of the basic center is randomly forgotten first. In this embodiment, this ratio is set to 10%. In addition, the basic center model trained in combination with the knowledge correction strategy has achieved good generalization performance. To reduce the overhead of the multi-center model, the auxiliary centers in the embodiment adopt a typical knowledge transfer method, rather than training with a model similar to the L2-TSK-FS model of the basic center.
[0184] The yellow area in Figure 1 depicts the training process of the auxiliary center. Further explanations will be provided for the antecedent parameters and consequent parameters of the auxiliary center.
[0185] As described above, before optimizing the antecedent and consequent parameters of the auxiliary center, the knowledge of the basic center is randomly forgotten first. The remaining antecedent parameters are used as the initial antecedent parameters of the auxiliary center, that is, the cluster centers for the auxiliary center clustering. Then, the principle of "the closer the distance, the higher the similarity" is applied. The auxiliary center samples are assigned to the nearest cluster according to their Euclidean distance to the clustering cluster. It should be noted that when all samples are assigned, the empty clusters without assigned samples will be discarded. Moreover, the clustering centers of the auxiliary center will be recalculated according to the samples of each cluster class through the above-mentioned formula for correcting the antecedent parameters and By avoiding using the typical fuzzy C-means (FCM) clustering technique, the training speed has been greatly improved.
[0186] When optimizing the consequent parameters of the auxiliary center , a = 1, ⋯, A, where A represents the number of auxiliary centers. To reduce the domain deviation between the basic center and the auxiliary centers, an improved maximum mean discrepancy PMMD is introduced in the conventional TSK transfer learning method. First, based on transductive transfer learning, for any intermediate variable Ω(D s , D t ), the following formula is set.
[0187]
[0188] In the formula, D s is the source domain, the knowledge domain of the knowledge transfer source side, D t is the target domain, the knowledge domain of the knowledge transfer target side, N s is the number of samples in the source domain, N t is the number of samples in the target domain, x g,s is the training data after the source domain's original data is transformed by the antecedent parameters, x g,t is the training data after the target domain's original data is transformed by the antecedent parameters, and T is the matrix transpose.
[0189] Give the basic center D BC and the auxiliary center D AC (a) The simplified equation of PMMD between them is as follows.
[0190]
[0191] In summary, the optimization objective function of the consequent parameters of the a-th auxiliary center is as follows.
[0192]
[0193] The above formula is equivalent to the following formula.
[0194]
[0195] Among them, represents the consequent parameters of the a-th auxiliary center, , . Taking the derivative of the above formula with respect to and setting it to 0 and simplifying, we can get:
[0196]
[0197] III. Training of the target center.
[0198] After the above-mentioned auxiliary center training is completed, the basic center and the auxiliary center together form a multi-center architecture to guide the establishment of the target center. The construction process of the target center is generally similar to that of the auxiliary center, but there are two significant differences: one is to use both the basic center and the auxiliary center to guide the establishment of the target center model, and the other is to use the label space graph regularization term to optimize the consequent parameters of the target center.
[0199] Figure 1 The light purple area in depicts the training process of the target center. Specifically, before using the antecedent parameters of the basic center to optimize the antecedent parameters of the target center, the knowledge of each center needs to be randomly forgotten. It should be noted that for the convenience of calculation, the proportion of random forgetting is not fixed at 10%, but first determine the center with the least number of rules, and then some rules of the remaining centers will be forgotten until they have the same number of rules as the center with the least number of rules.
[0200] In the process of integrating multi-center knowledge to guide the construction of the target center model, challenges such as data heterogeneity will inevitably be faced. Therefore, only by making full use of the consistency and complementarity between multi-center data can the knowledge in the multi-center architecture be effectively transferred to the new target center.
[0201] Based on the theory that "if the goal is to maximize the between-class separability, then the label space can be regarded as an optimal subspace", a label space graph regularizer is proposed. This regularizer consists of an intra-center graph and an inter-center graph, aiming to maintain the consistency among multi-center data and mine the complementary information between different centers.
[0202] Specifically, the intra-center graph represents the structural relationship within a single center, that is, when mapped to the low-dimensional label space, the data should maintain its neighborhood relationship in the original sample space; while the inter-center graph describes the structural relationship between different centers, meaning that when data of different modalities represent the same content or theme, they should have semantic similarity at the upper level. The following gives the detailed theoretical descriptions of the intra-center graph and the inter-center graph.
[0203] Intra-center graph. To preserve the local structural information of intra-center data in the original space in the low-dimensional label subspace, the intra-center graph is constructed based on the neighbor structure of samples in the original space. According to the Riemannian manifold structure, the intra-center graph matrix of the \(u\)-th center can be expressed as \(G\) uu , and its element in the \(i\)-th row and \(j\)-th column represents the weight of the edge connecting the node and in the graph, and its definition is as follows.
[0204]
[0205] where, where, is the element in the \(m\)-th row and \(n\)-th column representing the weight of the edge connecting the node and in the graph; \(G\) uu represents the intra-center graph of the \(u\)-th upper-level center; denotes 's \(K\) nearest neighbors, represents the Gaussian kernel bandwidth of the intra-center structure; \(\sigma\) represents the Gaussian kernel bandwidth of the intra-center structure.
[0206] Furthermore, the intra-center structural relationship can be maintained by minimizing the following formula.
[0207]
[0208] where, \(U\) represents the total number of all centers including the base center, auxiliary center and target center, represents the label of the \(m\)-th sample data in the \(u\)-th upper-level center; represents the label data label of the \(n\)-th sample data in the \(u\)-th upper-level center.
[0209] Inter - center graph. To preserve the high - level semantic information of the inter - center data in the original space in the low - dimensional label subspace, an inter - center graph is constructed in the original space according to the semantic relationship of the samples. According to the Riemannian manifold structure, the inter - center graph matrix between the \(u\) - th center and the \(v\) - th center can be expressed as \(G\) uv , and its element in the \(m\) - th row and \(n\) - th column represents the weight of the edge connecting the node and in the graph, and its definition is as follows.
[0210]
[0211] In the formula, \(G\) uv represents the inter - center graph matrix between the \(u\) - th center and the \(v\) - th center, and its element in the \(m\) - th row and \(n\) - th column represents the weight of the edge connecting the node and in the graph is .
[0212] Furthermore, the inter - center structural relationship can be preserved by minimizing the following formula.
[0213]
[0214] For convenience, the intra - center graph matrix \(G\) uu and the inter - center graph matrix \(G\) uv are integrated together to construct a block matrix \(G\), which is called the multi - center graph matrix, as follows.
[0215]
[0216] Among them, , and \(\beta\) is the weight parameter. The multi - center graph matrix constitutes the graph regularization term in the label space, which can be defined as.
[0217]
[0218] Among them, represents the \(u\) - th element in , represents the \(u\) - th diagonal element in , and \(L_{uv}\) is the Laplacian matrix corresponding to the \(u\) - th modality and the \(v\) - th modality, and \(B\) uv represents a diagonal matrix, and its diagonal element is .
[0219] It should be noted that, in order to construct the above Laplace equation and learn the idea of zero-padding alignment in deep learning, some empty samples are added to the target centers with fewer central samples. There is neither an intra-center structural relationship nor an inter-center structural relationship between these added samples and the original samples. They are just empty nodes added for alignment purposes.
[0220] Therefore, based on the above optimization objective function, the optimization objective function for the consequent parameters of the target center is as follows.
[0221]
[0222] In the formula, , , are all regularization parameters of the target center.
[0223] Taking the derivative of the above formula and setting it to 0 and simplifying, we can obtain.
[0224]
[0225] An embodiment of the present invention also provides a non-transitory machine-readable medium storing a computer program, wherein the above computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiment of the present invention.
[0226] An embodiment of the present invention also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiment of the present invention.
[0227] An embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The above memory stores a computer program capable of being executed by the at least one processor, and the above computer program, when executed by the at least one processor, is used to cause the electronic device to execute the method of the embodiment of the present invention.
[0228] Reference Figure 5, a block diagram of an electronic device of a server or a client that can be an embodiment of the present inventive concept will now be described. It is an example of a hardware device that can be applied to various aspects of the present inventive concept. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present inventive concept described and / or claimed herein.
[0229] As Figure 5 shown, the electronic device includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0230] Multiple components in the electronic device are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 can be any type of device that can input information into the electronic device. The input unit 506 can receive input digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device. The output unit 507 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 508 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 509 allows the electronic device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, and / or a wireless communication transceiver, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0231] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU, a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing units, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above. For example, in some embodiments, the method embodiments of the present invention can be implemented as a computer program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM 502 and / or the communication unit 509. In some embodiments, the computing unit 501 can be configured to execute the above methods in any other suitable way (e.g., by means of firmware).
[0232] The computer program for implementing the method of the embodiments of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0233] The user information (including, but not limited to, user device information, user personal information, etc.) and data (including, but not limited to, data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection.
[0234] The above-described embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but should not be construed as a limitation on the protection scope. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A distributed multi - center machine learning intelligent decision - making method, which is applied to the lower - level centers other than the top - level center and the sub - top - level center in the distributed multi - center. The top - level center is a large hospital in first - tier cities, the sub - top - level center is a relatively large hospital in second - tier cities, and the lower - level centers are small hospitals or clinics in third - and fourth - tier cities. It is characterized in that, Including: Obtain input data that requires decision-making; Determine the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model; Calculate the rule outputs of different fuzzy rules according to the input data and the consequent parameters corresponding to different antecedent parameters; Perform weighted calculation according to the membership degrees and the rule outputs corresponding to different fuzzy rules to obtain the output decision data corresponding to the input data; Among them, the fuzzy rule model is obtained by the lower-level center of the current layer through the fuzzy rule model of the upper-level center of the distributed multi-center, using the optimization objective function for knowledge transfer, and using the multi-center graph matrix to determine the regularization term to adjust the optimization objective function; by integrating the experience and knowledge of the top-level center and the sub-top-level center, the lower-level center can establish a clear understanding and diagnostic ability for specific diseases; Before determining the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model, the method further includes: Through the way of knowledge transfer, determine the antecedent parameters and consequent parameters for the lower-level center to migrate for the remaining rules after partially forgetting the fuzzy rules of the migrated upper-level center, where the fuzzy rules include corresponding antecedent parameters and consequent parameters; Determine the cluster centers according to the remaining rules, calculate the distances between the data samples of the training data of the sub-lower-level center and the cluster centers, and perform clustering to obtain multiple clusters, and eliminate the empty clusters; Based on the cluster centers and the corresponding data samples within the clusters, use the antecedent parameter optimization formula to calculate the corresponding antecedent parameters; Calculate the improved maximum mean difference according to the intermediate variable between the knowledge domain of the upper-level center and the knowledge domain of the current lower-level center; Calculate the optimization objective function corresponding to the current lower-level center according to the improved maximum mean difference and the optimization objective function of the fuzzy rule model of the upper-level center, where the optimization objective function includes a regularization term for solving the consequent parameters of the current lower-level center according to the consequent parameters of the upper-level center; Modify the regularization term of the optimization objective function according to the multi-center graph matrix to obtain the regularized optimization objective function; According to the updated optimization objective function, combine the recalculated antecedent parameters to update the migrated consequent parameters.
2. The method according to claim 1, wherein Determining the antecedent parameters and consequent parameters for the lower-level center to migrate after partially forgetting the fuzzy rules of the migrated upper-level center through the way of knowledge transfer includes: Obtain the fuzzy rules of the upper-level center through the way of knowledge transfer; Determine the upper-level center with the least number of rules and the number of rules; Forget the partial rules of other upper-level centers to the same number of rules as the upper-level center with the least number of rules to obtain the remaining rules; Use the antecedent parameters and consequent parameters of the remaining rules as the antecedent parameters and consequent parameters after the lower-level center migrates.
3. The method according to claim 1, wherein Calculating the corresponding antecedent parameters according to the migrated antecedent parameters and the antecedent parameter optimization formula includes: Optimize according to the antecedent parameters in the fuzzy rules using the antecedent parameter optimization formula to obtain the corresponding antecedent parameters; The antecedent parameters include the centers of the membership functions of the fuzzy rule model and the kernel widths ; The optimization formula for the antecedent parameters is as follows: ; ; where k represents the number of clusters of the fuzzy rule model, and x ji is the i-th training sample of the j-th cluster of the training data of the lower-level center, N k represents the number of samples of the k clusters, and h is the scale parameter.
4. The method according to claim 1, wherein Calculate the improved maximum average difference according to the intermediate variables between the knowledge domain of the upper-level center and the knowledge domain of the current lower-level center, including: According to the intermediate variables between the knowledge domain of the upper-level center and the knowledge domain of the current lower-level center, the calculation formula of the intermediate variables is as follows: ; Where D s is the source domain, including the knowledge domain of the upper-level center, D t is the target domain, including the knowledge domain of the current lower-level center, N s is the number of samples in the source domain, N t is the number of samples in the target domain, x g,s is the training data after the antecedent parameter conversion of the original data in the source domain, x g,t is the training data after the antecedent parameter conversion of the original data in the target domain, and T is the matrix transpose; Calculate the improved maximum average difference through the intermediate variables, as shown in the following formula: ; where PMMD represents the improved maximum mean difference, D t(a) is the a-th target domain, and is the consequent parameter of the j-th cluster of the current lower-level center corresponding to the a-th target domain.
5. The method according to claim 1, wherein Calculate the optimization objective function corresponding to the current lower-level center according to the improved maximum average difference and the optimization objective function of the fuzzy rule model of the upper-level center, including: Obtain the optimization objective function corresponding to the current lower-level center according to the following formula using the improved maximum average difference and the optimization objective function of knowledge transfer: ; In the formula, , , are all regularization parameters of the current lower-level center, c is the number of clusters, and j belongs to 1, 2, 3... c.
6. The method according to claim 5, characterized in that, Modify the regularization term of the optimization objective function according to the multi-center graph matrix to obtain the regularized optimization objective function, including: Modify the regularization term according to the multi-center graph matrix to obtain: ; In the formula, represents the u-th element in represents the u-th diagonal element in is the Laplacian matrix corresponding to the u-th mode and the v-th mode, B uv represents a diagonal matrix whose diagonal elements are ; The regularized optimization objective function is: ; Wherein, TC is the current lower-level center, , , , are all regularization parameters of the current lower-level center, c is the number of clusters, and j belongs to 1, 2, 3... c.
7. The method according to claim 1, wherein Before modifying the regularization term of the optimization objective function according to the multi-center graph matrix to obtain the regularized optimization objective function, the method further includes: Calculate the intra-center graph matrix of the upper-level center through the following formula: ; Among them, is the connection node of the center inner graph matrix represented by the m-th row and the n-th column and is the weight of the edge between; G uu represents the center inner graph matrix of the u-th upper-level center; represents the K-nearest neighbor of represents the Gaussian kernel bandwidth of the center internal structure; Minimize the following formula to maintain the intra-center structural relationship: ; where U represents the total number of all centers in the distributed multi-center, represents the label of the m-th sample data in the u-th upper-level center; represents the label data label of the n-th sample data in the u-th upper-level center; Calculate the inter-center graph matrix of the upper-level center through the following formula: ; Where G uv represents the center - to - center graph matrix between the u - th center and the v - th center, and the element in the m - th row and n - th column of it represents the connecting node of the center - to - center graph matrix and The weight of the edge between them is ; Minimize the following formula to maintain the inter-center structural relationship: ; Calculate the multi-center graph matrix through the following formula according to the intra-center graph matrix and the inter-center graph matrix: ; where G is the multi-center graph matrix, , and β is the weight parameter.
8. A distributed multi-center machine learning intelligent decision-making method, characterized in that, The distributed multi-center includes centers at three levels, specifically the top-level center, the sub-top-level center, and the lower-level center. The lower-level center executes the method according to any one of claims 1 to 7. The method executed by the sub-top-level center includes: Obtain the input data that needs to be decision-making; Determine the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model; Calculate the rule outputs of different fuzzy rules according to the input data and the consequent parameters corresponding to different antecedent parameters; Perform weighted calculation according to the membership degrees and the rule outputs corresponding to different fuzzy rules to obtain the output decision data corresponding to the input data; Among them, the fuzzy rule model is obtained by the current sub-top-level center through knowledge transfer calculation of the fuzzy rule model of the top-level center of the distributed multi-center using the optimization objective function.
9. The method according to claim 8, wherein Before determining the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model, the method further includes: Obtain the fuzzy rules of the top-level center through the way of knowledge transfer, where the fuzzy rules include antecedent parameters and consequent parameters; Forget the fuzzy rules according to a preset probability to obtain the remaining rules, where the fuzzy rules include the antecedent parameters and consequent parameters of the top-level center; Determine the clustering centers according to the remaining rules, calculate the distances between the data samples of the training data of the sub-top-level center and the clustering centers, and perform clustering to obtain multiple clustering clusters, and eliminate the empty clusters; Based on the clustering centers and the corresponding data samples within the clusters, use the antecedent parameter optimization formula to recalculate the corresponding antecedent parameters; Calculate the improved maximum mean difference according to the intermediate variables of the knowledge domain of the top-level center and the knowledge domain of the current sub-top-level center; Calculate the optimization objective function corresponding to the current sub-top-level center according to the improved maximum mean difference and the optimization objective function of the fuzzy rule model of the top-level center; Calculate the consequent parameters of the current sub-top-level center according to the consequent parameters of the remaining rules and the optimization objective function; 10. A distributed multi-center machine learning intelligent decision-making method, characterized in that, The distributed multi-center includes centers at three levels, specifically the top-level center, the sub-top-level center, and the lower-level center. The lower-level center executes the method described in any one of claims 1 to 7, the sub-top-level center executes the method described in any one of claims 8 to 9, and the method executed by the top-level center includes: Obtain the input data that needs to be decision-making; Determine the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model; Calculate the rule outputs of different fuzzy rules according to the input data and the consequent parameters corresponding to different antecedent parameters; Perform weighted calculation according to the membership degrees and the rule outputs corresponding to different fuzzy rules to obtain the output decision data corresponding to the input data; Wherein, the fuzzy rule model is obtained by the top-level center through the fuzzy c-means clustering algorithm FCM and a regression algorithm including a penalty term of an insensitive loss function; 11. The method according to claim 10, characterized in that Before determining the membership degrees of the input data in different fuzzy rules according to the antecedent parameters of different fuzzy rules of the fuzzy rule model, the method further includes: Calculate the antecedent parameters of the fuzzy rule model of the top-level center by using the fuzzy c-means clustering algorithm FCM; Train and optimize the consequent parameters of the fuzzy rule model by a regression algorithm including a penalty term of an insensitive loss function; Obtain sample data according to the knowledge data of the top-level center, wherein the sample data includes training data and verification data; Iteratively optimize the antecedent parameters and consequent parameters through the training data and the verification data to obtain the final antecedent parameters and consequent parameters; 12. The method according to claim 11, wherein Iteratively optimize the antecedent parameters and consequent parameters through the training data to obtain the final antecedent parameters and consequent parameters, including: The antecedent parameters include the centers of the membership functions of the fuzzy rule model and the kernel widths ; the antecedent parameters are optimized by the antecedent parameter optimization formula; Optimize the consequent parameters by the following formula: ; wherein, is the consequent parameter of the top-level center at the t-th iteration, is the consequent parameter of the top-level center at the (t - 1)-th iteration at the t-th iteration; and are the regularization parameters of the top-level center, represents the consequent parameter corresponding to the j-th cluster at the t-th iteration , and respectively represent the data and label of the i-th training sample with respect to the j-th class; Verify the optimized antecedent parameters and consequent parameters according to the verification data, regard those that do not match the actual labels of the corresponding verification samples as error samples, and incorporate the error samples into the error set; Increment the iteration count by one, continue the iteration until the iteration count reaches the preset number of times, and select the antecedent parameters and consequent parameters corresponding to the iteration round with the least corresponding error set as the final antecedent parameters and consequent parameters; 13. An electronic device, comprising: A processor and a memory storing a program, characterized in that the program includes instructions that, when executed by the processor, cause the processor to execute the method described in any one of claims 1 to 12.
14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Target prediction model construction method and prediction method in multi-center small sample scene
CN116596161A
Data classification method and system based on improved TSK fuzzy classification model
CN119691578A