A trapezoidal granularity data weighted fuzzy clustering method and system

By using the weighted fuzzy clustering method of trapezoidal granularity data and combining the FCM and PSO algorithms to optimize granularity data, the grouping accuracy and efficiency problems of traditional fuzzy clustering algorithms in medical pre-diagnosis are solved, and more efficient patient grouping and information retention are achieved.

CN116776182BActive Publication Date: 2025-10-10HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310804073.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-03
Publication Date
2025-10-10
Estimated Expiration
2043-07-03

AI Technical Summary

Technical Problem

Traditional fuzzy clustering algorithms have insufficient time performance and accuracy in pre-diagnosis of medical patients. The selection of granular data types, boundary optimization and weight setting are unreasonable, resulting in low grouping accuracy.

Method used

A weighted fuzzy clustering method based on trapezoidal granularity data was adopted, combined with the FCM algorithm and the PSO particle swarm optimization algorithm. By constructing trapezoidal granularity data and optimizing its coverage and specificity, the optimal membership and typicality matrices were calculated for cluster analysis.

Benefits of technology

It improves the accuracy and efficiency of medical patient grouping, avoids local optimality, enhances data interpretability and information retention, and is suitable for the representation of complex data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116776182B_ABST
    Figure CN116776182B_ABST
Patent Text Reader

Abstract

The application discloses a kind of weighted fuzzy clustering methods and systems for trapezoidal granularity data, comprising:1 selects the data representative of medical patient information dataset, and accordingly constructs granularity data;2 calculate the weight corresponding to granularity data;3 using weighted fuzzy clustering algorithm to cluster granularity data, and combine new distance measurement mode;4 calculate the reconstruction error corresponding to the result after clustering, according to the inflection point corresponding to the number of clustering center of reconstruction error, and combine membership matrix to carry out grouping in pre-diagnosis stage for medical patient.The PSO particle swarm optimization algorithm is introduced into the construction process of trapezoidal granularity data in the application, and combined with weighted fuzzy clustering algorithm, so as to improve the accuracy and reliability of medical patient grouping in pre-diagnosis stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of data mining, in particular to a weighted fuzzy clustering method and system for trapezoidal granularity data. Background Art

[0002] Cluster analysis, a key component of machine learning algorithms, integrates data with similar characteristics while minimizing the differences between clusters. Clustering can be categorized as hard clustering or soft clustering. Unlike hard clustering, which constrains data objects to belong to only one cluster, soft clustering allows data objects to belong to multiple clusters with varying degrees of membership. Fuzzy clustering is a type of soft clustering. Researchers have incorporated the concept of fuzzy sets into clustering, and fuzzy clustering algorithms have gradually gained widespread acceptance. Fuzzy clustering algorithms describe the degree to which data samples belong to clusters and better reflect real-world data. Consequently, fuzzy clustering algorithms have become a key area of ​​cluster analysis. However, when faced with pre-diagnosis questions for a large number of medical patients, traditional fuzzy clustering algorithms fall short of expectations in terms of time performance and accuracy.

[0003] Therefore, the concept of granular computing (GrC) was proposed, and people began to study the clustering problem of granular data and apply it to the pre-diagnosis stage of the medical system. Currently, there are three main problems with fuzzy clustering algorithms based on granular data:

[0004] 1) Granularity data type selection problem. Granularity data types such as interval information granularity and triangular fuzzy set information granularity lose information about medical patient data and are not very representative.

[0005] 2) There is no optimization process for the granular data boundary. After the granular data is constructed, the optimization process is also an important step. Directly determining the endpoint value without the optimization step may result in local optimality and insufficient accuracy of the endpoint value.

[0006] 3) Lack of granular data weights or unreasonable weight settings. After granular data construction and optimization, it is necessary to evaluate the indicators of granular data quality, usually from the two aspects of coverage and specificity. Constructing appropriate coverage and specificity functions for different granular data types will be effective. Summary of the Invention

[0007] In order to address the shortcomings of the above-mentioned existing technologies, the present invention proposes a weighted fuzzy clustering algorithm based on granular data, in order to construct granular data with better data interpretability and meaning, enhance the accuracy of grouping of medical patients in the pre-consultation stage, shorten the grouping time, and further improve the working efficiency of the medical system.

[0008] In order to achieve the above-mentioned object of the invention, the present invention adopts the following technical solutions:

[0009] The weighted fuzzy clustering method for trapezoidal granularity data of the present invention is characterized by comprising:

[0010] Step 1: Collect the corresponding medical patient information dataset x k represents the information of the kth medical patient, and N is the number of medical patients;

[0011] Step 2: Use the FCM algorithm to obtain the data set M data prototypes V={v1,v2,…,v s ,…,v M} and the cluster Ω={Ω1,Ω2,…,Ω s ,…,Ω M}; where v s Represents the sth data prototype; Ω s Indicates that the s-th data prototype v s The sth cluster with centered at ;

[0012] Step 3: Use PSO particle swarm optimization algorithm to construct the s-th data prototype v s and its cluster Ω s The sth optimal trapezoidal granularity data based on And calculate Particle size quality s∈[1,M];

[0013] Step 4: Use the weighted PFCM algorithm to calculate the optimal trapezoidal granularity data. Perform clustering to obtain the optimal membership matrix U * , the optimal typicality matrix T * and the optimal cluster center V * ;

[0014] Step 5: Use formula (1) to obtain the reconstructed s-th trapezoidal granularity data Thus, the reconstructed trapezoidal granularity data set is obtained

[0015]

[0016] In formula (1), m is the fuzzy coefficient, Represents the sth trapezoidal granularity data Belongs to the i-th optimal cluster center degree, and c represents the number of cluster centers;

[0017] Step 6: Calculate the reconstruction error E using formula (2):

[0018]

[0019] Step 7: According to the number of cluster centers corresponding to the inflection point of the reconstruction error E, combined with the optimal membership matrix U in the clustering result * Conduct group consultations for medical patients.

[0020] The weighted fuzzy clustering method for trapezoidal granularity data of the present invention is also characterized in that step 3 comprises:

[0021] Step 3.1, set the sth trapezoidal granularity data The size of the corresponding particle swarm is P; define the cognitive constant as c1, the social constant as c2, the inertia constant as ω, and ω∈[0,1];

[0022] Step 3.2, initialize s=1;

[0023] Step 3.3, initialize the current attribute j=1;

[0024] Step 3.4, define and initialize the current particle swarm iteration number iter=1, the maximum iteration number is iterNum, and the iteration stop threshold is ε;

[0025] Step 3.5, initialize the particle number p in the particle swarm of the current iter iteration to 1;

[0026] Step 3.6: Initialize the sth trapezoidal granularity data The trapezoidal granularity of the iter-th iteration in the j-th attribute The random vector of the pth particle in and Random initialization The position of the pth particle in and speed

[0027] Step 3.7, Initialization The trapezoidal particle size corresponding to the pth particle in, express The left endpoint of the lower base, express The right endpoint of the lower base, express The left endpoint of the upper base, express The right endpoint of the upper base; when iter = 1, let in, and They are The trapezoidal particle size corresponding to the pth particle The maximum and minimum values ​​of the data covered, is the sth data prototype v s The value of the jth attribute;

[0028] Step 3.8: Use formula (3) to get the sth trapezoidal particle size data The trapezoidal granularity of the iter+1th iteration in the jth attribute of The trapezoidal particle size corresponding to the pth particle

[0029]

[0030] In formula (3), express The left endpoint of the lower base, express The right endpoint of the lower base, express The left endpoint of the upper base, express The right endpoint of the upper base; Respectively The four endpoints of the p-th particle correspond to the position of the particle, j = 1, ... n; s = 1, ... M, n is the number of attributes;

[0031] Step 3.9: Use formula (4) to obtain the sth trapezoidal granularity data The trapezoidal granularity of the iter+1th iteration in the jth attribute of The position of the pth particle in

[0032]

[0033] In formula (4), in, Respectively The positions of the pth particle corresponding to the four endpoints; express The velocity of the pth particle in ;

[0034] Step 3.10: Using formula (5), we can get The velocity of the pth particle in

[0035]

[0036] In formula (5), and yes Two random vectors for the p-th particle in , and are uniformly distributed in the interval [0,1]; for The particle size quality, is the sth optimal trapezoidal granularity data The optimal trapezoidal granularity for the iter-th iteration of the j-th attribute Particle size quality;

[0037] Step 3.11: Calculate using formula (6) The trapezoidal particle size corresponding to the pth particle Coverage

[0038]

[0039] In formula (6), express the number of data points covered; yes The corresponding trapezoidal fuzzy membership function, x kj is the jth attribute of the kth patient information, and has:

[0040]

[0041] In formula (7), and They are The left and right endpoints of the lower base, and They are The left and right endpoints of the upper and lower edges of , when the upper boundary contains only one element, let

[0042] Step 3.12: Calculate using formula (8) Specificity

[0043]

[0044] In formula (8), express The maximum value of the data covered and minimum value The difference between

[0045] Step 3.13: Calculate using formula (9) Particle size quality

[0046]

[0047] Step 3.14, let p+1 be assigned to p, and determine whether p≤P is established. If so, return to step 3.6 to execute; otherwise, it means that the sth trapezoidal granularity data is obtained. The trapezoidal granularity set of the iter+1th iteration in the jth attribute of And execute step 3.15;

[0048] Step 3.15: Use formula (10) to optimally calculate the sth optimal trapezoidal granularity data The trapezoidal granularity of the iter+1th iteration in the jth attribute of Particle size quality

[0049]

[0050] Step 3.16: Judgment Or whether iter>iterNum is established. If so, it means that the sth optimal trapezoidal granularity data is obtained. The trapezoidal granularity of the j-th attribute And execute step 3.17; otherwise, assign iter+1 to iter and return to step 3.5.

[0051] Step 3.17, let j+1 be assigned to j, and determine whether j≤n is established. If so, return to step 3.4 to execute; otherwise, it means that the sth optimal trapezoidal granularity data is obtained. And execute step 3.18;

[0052] Step 3.18: Calculate using formula (11) Coverage

[0053]

[0054] In formula (11), express the number of data points covered; yes The corresponding trapezoidal fuzzy membership function, x kj is the jth attribute of the kth patient information, and has:

[0055]

[0056] In formula (12), and They are The left and right endpoints of the lower base, and They are The left and right endpoints of the upper and lower edges of , when the upper boundary contains only one element, let

[0057] Step 3.19: Calculate using formula (13) Specificity

[0058]

[0059] In formula (13), express The maximum value of the data covered and minimum value The difference between

[0060] Step 3.20: Calculate using formula (14) Particle size quality

[0061]

[0062] Step 3.21, let s+1 be assigned to s, and determine whether s≤M is true. If so, return to step 3.3 to execute; otherwise, it means that M optimal trapezoidal granularity data sets are obtained. and the corresponding particle size quality set

[0063] The step 4 comprises:

[0064] Step 4.1: Calculate the sth optimal trapezoidal granularity data Corresponding weight

[0065]

[0066] In formula (15), for The particle size quality, is the maximum value of the particle size quality of M optimal particle size data;

[0067] Step 4.2: Define and initialize the current number of iterations iter = 1, the iteration stop threshold is ε, and the maximum number of iterations is iterMax;

[0068] Step 4.3: Use formula (16) to establish the objective function of the iterth iteration

[0069]

[0070] In formula (16), c represents the number of clusters, M represents the number of granularities, m represents the fuzzy coefficient, η represents the typicality coefficient, a and b represent the relative importance parameters of membership and typicality, respectively. Represents the sth optimal trapezoidal granularity data The weight of represents the i-th trapezoidal granularity cluster center of the iter-th iteration, Represents the sth optimal trapezoidal granularity data Affiliated to the degree, Represents the sth optimal trapezoidal granularity data belong Typicality; is the penalty factor, and:

[0071]

[0072] In formula (17), K is a constant; yes and The distance between them is:

[0073]

[0074] In formula (18), Represents the sth optimal trapezoidal granularity data The trapezoidal granularity of the j-th attribute and the i-th trapezoidal granularity cluster center of the iter-th iteration The trapezoidal granularity of the jth attribute in The distance between them is:

[0075]

[0076] In formula (19), p1, p2, p3, and p4 represent four intermediate variables, and express The left endpoint of the lower base, express The left endpoint of the lower base, express The left endpoint of the upper base, express The left endpoint of the upper base, express The right endpoint of the upper base, express The right endpoint of the upper base, express The right endpoint of the lower base, express The right endpoint of the lower base;

[0077] Step 4.4: Calculate the sth optimal trapezoidal granularity data using formula (20) The i-th trapezoidal granularity cluster center belonging to the iter+1-th iteration degree Get the membership matrix of the iter+1th iteration

[0078]

[0079] In formula (20), represents the j-th trapezoidal granularity cluster center of the iter-th iteration;

[0080] Step 4.5: Calculate the sth optimal trapezoidal granularity data using formula (21) belong Typicality Get the typical degree matrix of the iter+1th iteration

[0081]

[0082] Step 4.6: Use equations (22)-(25) to calculate the trapezoidal granularity cluster center of the iter+1th iteration Thus, the trapezoidal granularity cluster center matrix of the iter+1th iteration is obtained

[0083]

[0084]

[0085]

[0086]

[0087] In formulas (22)-(25), Respectively represent the i-th trapezoidal granularity cluster center of the iter+1-th iteration The left and right endpoints of the upper and lower sides of the trapezoidal granularity of the j-th attribute; Respectively represent the i-th trapezoidal granularity cluster center of the iter+1-th iteration The left and right endpoints of the lower base of the trapezoidal granularity of the j-th attribute;

[0088] Step 4.7: If ||U iter+1 -U iter ||<ε or iter>iterMax, then the membership matrix U of the iter-th iteration iter , typicality matrix T iter and cluster center Viter As the optimal membership matrix U * , the optimal typicality matrix T * and the optimal cluster center V * Otherwise, assign iter+1 to iter and return to step 4.3.

[0089] The weighted fuzzy clustering system for trapezoidal granularity data of the present invention is characterized by comprising:

[0090] Data prototype extraction module: used to generate data prototypes and their clusters corresponding to the medical patient information dataset as the basis for granular data construction;

[0091] Granularity data construction module: used to construct trapezoidal granularity data and optimize its formation process;

[0092] Granular data clustering module: used to take granular data as input and perform cluster analysis;

[0093] Reconstruction error calculation module: used to calculate the reconstruction error and combine it with the optimal membership matrix to complete the grouping of medical patients before consultation;

[0094] Wherein, the granular data construction module includes:

[0095] PSO algorithm data initialization unit: determine the particle swarm size P, cognitive constant c1 and social constant c2, random vector r j and s j , inertia constant ω, and the maximum number of iterations iterNum of the particle swarm algorithm;

[0096] Granularity data initialization unit: used for initializing trapezoidal granularity parameters;

[0097] Trapezoidal granularity data optimization unit: based on the weighted fuzzy clustering method, the trapezoidal granularity data constructed by PSO algorithm is optimized;

[0098] The granular data clustering module includes:

[0099] Data initialization unit: determines the value of the number of clusters c of the medical patient information dataset, initializes the cluster center, and sets the threshold ε;

[0100] Data update unit: used to update cluster centers, membership and typicality functions;

[0101] Clustering condition determination unit: determines whether the clustering meets the condition || U new -U old ||<ε, if so, stop executing, otherwise continue clustering the granular data; where U new Represents the membership function of the current iteration, Uold represents the membership function of the previous iteration.

[0102] Compared with the prior art, the beneficial effects of the present invention are embodied in:

[0103] 1. The present invention proposes trapezoidal information granularity, which not only combines the characteristics of triangular fuzzy set information granularity and interval information granularity, but also approaches the original data membership and has a stronger ability to express complex data, making the clustering effect of trapezoidal granularity data more robust than the previous two. In addition, the trapezoidal information granularity retains the information of medical patient data sets intact, laying a solid foundation for the subsequent grouping process in the pre-diagnosis stage.

[0104] 2. The present invention introduces the PSO particle swarm optimization algorithm, which globally optimizes the granularity data, constructs the coverage function and the specificity function, and uses the product of the two as the optimization index of the granularity data parameters to adjust the core left and right endpoints and the left and right boundary positions of the trapezoidal granularity data to achieve the optimal trapezoidal information granularity, thereby avoiding falling into the local optimum and improving the reliability of the granularity data.

[0105] 3. This invention addresses the lack of guidance for granular weights or irrational weight settings during clustering. By considering both coverage and specificity, it optimizes the granularity quality evaluation formula. Furthermore, granular weights are introduced into the clustering process of medical patient information, and patients are grouped in the pre-consultation phase based on the clustering results, improving the accuracy of medical patient grouping. BRIEF DESCRIPTION OF THE DRAWINGS

[0106] Figure 1 This is a flow chart of the weighted fuzzy clustering method for medical patient information based on trapezoidal granularity data of the present invention;

[0107] Figure 2 Constructing a graph for the trapezoidal granularity data of the medical patient information dataset of the present invention;

[0108] Figure 3 This is a diagram of the medical patient information data set processing process of the present invention. DETAILED DESCRIPTION

[0109] In this embodiment, a weighted fuzzy clustering method for trapezoidal granularity data is used to cluster medical patient information and group patients for consultation based on the clustering results. Specifically, see Figure 1 , is done as follows:

[0110] Step 1: Collect information of N medical patients and use x k Represents the information of the kth medical patient, and x k ={x k1 ,x k2,…,x kj ,…,x kn}, x kj Represents the information x of the kth medical patient k The jth attribute of , j = 1, ... n, k = 1, ... N, n is the number of attributes of medical patient information. N attributes of medical patient information can form a data set

[0111] Step 2: Use the FCM algorithm to obtain the data set M data prototypes V={v1,v2,…,v s ,…,v M} and the cluster Ω={Ω1,Ω2,…,Ω s ,…,Ω M}; where v s Represents the sth data prototype; Ω s Indicates that the s-th data prototype v s Specifically, the information of the kth medical patient belongs to the sth cluster Ω s , if and only if u ks When the maximum value is taken, s=1,…M, and:

[0112]

[0113] Step 3: As an abstract data, the trapezoidal granularity data has multi-dimensional and high information content characteristics. The general granularity construction method is difficult to achieve the ideal effect. As a mature global optimization algorithm, the PSO particle swarm optimization algorithm has good stability and effectiveness. Therefore, this optimization algorithm is applied to the construction process of the trapezoidal granularity data. That is, construct the s-th data prototype v s and its cluster Ω s The sth optimal trapezoidal granularity data based on s∈[1,M];

[0114] Step 3.1, set the sth trapezoidal granularity data The size of the corresponding particle swarm is P; define the cognitive constant as c1, the social constant as c2, the inertia constant as ω, and ω∈[0,1];

[0115] Step 3.2, initialize s=1;

[0116] Step 3.3, initialize the current attribute j=1;

[0117] Step 3.4, define and initialize the current particle swarm iteration number iter=1, the maximum iteration number is iterNum, and the iteration stop threshold is ε;

[0118] Step 3.5, initialize the particle number p in the particle swarm of the current iter iteration to 1;

[0119] Step 3.6: Initialize the sth trapezoidal granularity data The trapezoidal granularity of the iter-th iteration in the j-th attribute The random vector of the pth particle in and Random initialization The position of the pth particle in and speed

[0120] Step 3.7, Initialization The trapezoidal particle size corresponding to the pth particle in, express The left endpoint of the lower base, express The right endpoint of the lower base, express The left endpoint of the upper base, express The right endpoint of the upper base; when iter = 1, let in, and They are The trapezoidal particle size corresponding to the pth particle The maximum and minimum values ​​of the data covered, is the sth data prototype v s The value of the jth attribute;

[0121] Step 3.8: Use formula (2) to get the sth trapezoidal granularity data The trapezoidal granularity of the iter+1th iteration in the jth attribute of The trapezoidal particle size corresponding to the pth particle

[0122]

[0123] In formula (2), express The left endpoint of the lower base, express The right endpoint of the lower base, express The left endpoint of the upper base, express The right endpoint of the upper base; Respectively The four endpoints of the p-th particle correspond to the position of the particle, j = 1, ... n; s = 1, ... M, n is the number of attributes;

[0124] Step 3.9: Use formula (3) to obtain the sth trapezoidal particle size data The trapezoidal granularity of the iter+1th iteration in the jth attribute of The position of the pth particle in

[0125]

[0126] In formula (3), Respectively The positions of the pth particle corresponding to the four endpoints; express The velocity of the pth particle in ;

[0127] Step 3.10: Using formula (4), we can get The velocity of the pth particle in

[0128]

[0129] In formula (4), and yes Two random vectors for the p-th particle in , and are uniformly distributed in the interval [0,1]; for The particle size quality, is the sth optimal trapezoidal granularity data The optimal trapezoidal granularity for the iter-th iteration of the j-th attribute Particle size quality;

[0130] Step 3.11: Calculate using formula (5) The trapezoidal particle size corresponding to the pth particle Coverage

[0131]

[0132] In formula (5), express the number of data points covered; yes The corresponding trapezoidal fuzzy membership function, x kj is the jth attribute of the kth patient information, and has:

[0133]

[0134] In formula (6), and They are The left and right endpoints of the lower base, and They are The left and right endpoints of the upper and lower edges of , when the upper boundary contains only one element, let

[0135] Step 3.12: Calculate using formula (7) Specificity

[0136]

[0137] In formula (7), express The maximum value of the data covered and minimum value The difference between

[0138] Step 3.13: Calculate using formula (8) Particle size quality

[0139]

[0140] Step 3.14, let p+1 be assigned to p, and determine whether p≤P is established. If so, return to step 3.6 to execute; otherwise, it means that the sth trapezoidal granularity data is obtained. The trapezoidal granularity set of the iter+1th iteration in the jth attribute of And execute step 3.15;

[0141] Step 3.15: Use formula (9) to optimally calculate the sth optimal trapezoidal granularity data The trapezoidal granularity of the iter+1th iteration in the jth attribute of Particle size quality

[0142]

[0143] Step 3.16: Judgment Or whether iter>iterNum is established. If so, it means that the sth optimal trapezoidal granularity data is obtained. The trapezoidal granularity of the j-th attribute And execute step 3.17; otherwise, assign iter+1 to iter and return to step 3.5.

[0144] Step 3.17, let j+1 be assigned to j, and determine whether j≤n is established. If so, return to step 3.4 to execute; otherwise, it means that the sth optimal trapezoidal granularity data is obtained. And execute step 3.18;

[0145] Step 3.18: Calculate using formula (10) Coverage

[0146]

[0147] In formula (10), express the number of data points covered; yes The corresponding trapezoidal fuzzy membership function, x kj is the jth attribute of the kth patient information, and has:

[0148]

[0149] In formula (11), and They are The left and right endpoints of the lower base, and They are The left and right endpoints of the upper and lower edges of , when the upper boundary contains only one element, let

[0150] Step 3.19: Calculate using formula (12) Specificity

[0151]

[0152] In formula (12), express The maximum value of the data covered and minimum value The difference between

[0153] Step 3.20: Calculate using formula (13) Particle size quality

[0154]

[0155] Step 3.21, let s+1 be assigned to s, and determine whether s≤M is true. If so, return to step 3.3 to execute; otherwise, it means that M optimal trapezoidal granularity data sets are obtained. and the corresponding particle size quality set

[0156] Step 4: Use the weighted PFCM algorithm to calculate the optimal trapezoidal granularity data Perform clustering to obtain the optimal membership matrix U * , the optimal typicality matrix T * and the optimal cluster center V * ;

[0157] Step 4.1: Calculate the sth optimal trapezoidal granularity data Corresponding weight

[0158]

[0159] In formula (14), for Particle size quality, max(Q * ) is the maximum value of the particle size quality of the M optimal particle size data;

[0160] Step 4.2: Define and initialize the current number of iterations iter = 1, the iteration stop threshold is ε, and the maximum number of iterations is iterMax;

[0161] Step 4.3: Use formula (15) to establish the objective function of the iterth iteration

[0162]

[0163] In formula (15), c represents the number of clusters, M represents the number of granularities, m represents the fuzzy coefficient, η represents the typicality coefficient, a and b represent the relative importance parameters of membership and typicality, respectively. Represents the sth optimal trapezoidal granularity data The weight of represents the i-th trapezoidal granularity cluster center of the iter-th iteration, Represents the sth optimal trapezoidal granularity data Affiliated to degree, Represents the sth optimal trapezoidal granularity data belong Typicality; is the penalty factor, and:

[0164]

[0165] In formula (16), K is a constant; yes and The distance between them is:

[0166]

[0167] In formula (17), Represents the sth optimal trapezoidal granularity data The trapezoidal granularity of the j-th attribute and the i-th trapezoidal granularity cluster center of the iter-th iteration The trapezoidal granularity of the j-th attribute in The distance between them is:

[0168]

[0169] In formula (18), p1, p2, p3, and p4 represent four intermediate variables, and express The left endpoint of the lower base, express The left endpoint of the lower base, express The left endpoint of the upper base, express The left endpoint of the upper base, express The right endpoint of the upper base, express The right endpoint of the upper base, express The right endpoint of the lower base, express The right endpoint of the lower base;

[0170] Step 4.4: Calculate the sth optimal trapezoidal granularity data using formula (19) The i-th trapezoidal granularity cluster center belonging to the iter+1-th iteration degree Get the membership matrix of the iter+1th iteration

[0171]

[0172] In formula (19), represents the j-th trapezoidal granularity cluster center of the iter-th iteration;

[0173] Step 4.5: Calculate the sth optimal trapezoidal granularity data using formula (20) belong Typicality Get the typical degree matrix of the iter+1th iteration

[0174]

[0175] Step 4.6: Calculate the trapezoidal granularity clustering centers of the (iter+1)th iteration by using formula (21)-(24) Thus, the trapezoidal granularity clustering center matrix of the (iter+1)th iteration is obtained

[0176]

[0177]

[0178]

[0179]

[0180] In formula (21)-(24), respectively represent the left and right endpoints of the upper base of the trapezoidal granularity of the jth attribute of the ith trapezoidal granularity clustering center of the (iter+1)th iteration; respectively represent the left and right endpoints of the lower base of the trapezoidal granularity of the jth attribute of the ith trapezoidal granularity clustering center of the (iter+1)th iteration;

[0181] Step 4.7: If ||U iter+1 -U iter ||<ε or iter>iterMax, then the membership matrix U iter , the typicality matrix T iter and the clustering centers V iter of the iter iteration are respectively taken as the optimal membership matrix U * , the optimal typicality matrix T * and the optimal clustering centers V * ; otherwise, iter+1 is assigned to iter, and then the step 4.3 is executed.

[0182] Step 5: Obtain the reconstructed s th trapezoidal granularity data by using formula (25) Thus, the reconstructed trapezoidal granularity data set is obtained

[0183]

[0184] In formula (25), m is a fuzzy coefficient, represents the degree to which the s th trapezoidal granularity data belongs to the ith optimal clustering center , and c represents the number of clustering centers.​​

[0185] Step 6: Calculate the reconstruction error E using formula (26):

[0186]

[0187] Step 7: According to the number of cluster centers corresponding to the inflection point of the reconstruction error E, combined with the optimal membership matrix U in the clustering result * Conduct group consultations for medical patients.

[0188] In this embodiment, a weighted fuzzy clustering system for trapezoidal granularity data includes:

[0189] Data prototype extraction module: used to generate data prototypes and their clusters corresponding to the medical patient information dataset as the basis for granular data construction;

[0190] Granularity data construction module: used to construct trapezoidal granularity data and optimize its formation process;

[0191] Granular data clustering module: used to take granular data as input and perform cluster analysis;

[0192] Reconstruction error calculation module: used to calculate the reconstruction error and, combined with the optimal membership matrix, to complete the grouping of medical patients before consultation;

[0193] Among them, the granular data building module includes:

[0194] PSO algorithm data initialization unit: determine the particle swarm size P, cognitive constant c1 and social constant c2, random vector r j and s j , inertia constant ω, and the maximum number of iterations iterNum of the particle swarm algorithm;

[0195] Granularity data initialization unit: used for initializing trapezoidal granularity parameters;

[0196] Trapezoidal granularity data optimization unit: Based on the above-mentioned weighted fuzzy clustering method, the trapezoidal granularity data is optimized and constructed through the PSO algorithm;

[0197] The granular data clustering module includes:

[0198] Data initialization unit: determines the value of the number of clusters c of the medical patient information dataset, initializes the cluster center, and sets the threshold ε;

[0199] Data update unit: used to update cluster centers, membership and typicality functions;

[0200] Clustering condition determination unit: determines whether the clustering meets the condition || U new -U old||<ε, if so, stop executing, otherwise continue clustering the granular data; where U new Represents the membership function of the current iteration, U old represents the membership function of the previous iteration.

[0201] This example focuses on the "medical consultation" phase of a medical system. The data to be clustered is patient information. Specifically, the dataset to be clustered consists of all patient information. Each piece of patient information includes numerical attributes such as name, age, gender, contact number, medical history, allergies, family medical history, and medical records. The clustering method used in this example is used to cluster all patient information. Based on the clustering results, pre-consultation services such as clinic diversion and general precautions are provided to different types of patients.

[0202] In this example, the clustering method was implemented on a PC running Windows 11 and MATLAB R2018a. The hardware requirements included an Intel Core 2.90 GHz CPU and 16 GB of RAM. The algorithm parameters m and η were both set to 2.0, a = b = 1, the particle swarm size P was set to 500, and the maximum number of iterations, iterNum, was set to 100.

[0203] The weighted fuzzy clustering method and system for trapezoidal granularity data of the present invention can be well applied to the "medical guidance and consultation" link in the medical system to correctly group medical patients and improve the efficiency of the medical system operation.

Claims

1. A weighted fuzzy clustering method for trapezoidal granularity data, characterized in that: include: Step 1: Collect the corresponding medical patient information dataset x k represents the information of the kth medical patient, and N is the number of medical patients; Step 2: Use the FCM algorithm to obtain the data set M data prototypes V={v1,v2,…,v s ,…,v M } and the cluster Ω={Ω1,Ω2,…,Ω s ,…,Ω M }; Among them, v s Represents the sth data prototype; Ω s Indicates that the s-th data prototype v s The sth cluster with centered at ; Step 3: Use PSO particle swarm optimization algorithm to construct the s-th data prototype v s and its cluster Ω s The sth optimal trapezoidal granularity data based on And calculate Particle size quality s∈[1,M]; Step 4: Use the weighted PFCM algorithm to calculate the optimal trapezoidal granularity data. Perform clustering to obtain the optimal membership matrix U * , the optimal typicality matrix T * and the optimal cluster center V * ; Step 5: Use formula (1) to obtain the reconstructed s-th trapezoidal granularity data Thus, the reconstructed trapezoidal granularity data set is obtained In formula (1), m is the fuzzy coefficient, Represents the sth trapezoidal granularity data Belongs to the i-th optimal cluster center degree, and c represents the number of cluster centers; Step 6: Calculate the reconstruction error E using formula (2): Step 7: According to the number of cluster centers corresponding to the inflection point of the reconstruction error E, combined with the optimal membership matrix U in the clustering result * Conduct group consultations for medical patients.

2. The weighted fuzzy clustering method for trapezoidal granularity data according to claim 1 is characterized in that: The step 3 includes: Step 3.1, set the sth trapezoidal granularity data The size of the corresponding particle swarm is P; define the cognitive constant as c1, the social constant as c2, the inertia constant as ω, and ω∈[0,1]; Step 3.2, initialize s=1; Step 3.3, initialize the current attribute j=1; Step 3.4, define and initialize the current particle swarm iteration number iter=1, the maximum iteration number is iterNum, and the iteration stop threshold is ε; Step 3.5, initialize the particle number p in the particle swarm of the current iter iteration to 1; Step 3.6: Initialize the sth trapezoidal granularity data The trapezoidal granularity of the iter-th iteration in the j-th attribute The random vector of the pth particle in and Random initialization The position of the pth particle in and speed Step 3.7, Initialization The trapezoidal particle size corresponding to the pth particle in, express The left endpoint of the lower base, express The right endpoint of the lower base, express The left endpoint of the upper base, express The right endpoint of the upper base; when iter = 1, let in, and They are The trapezoidal particle size corresponding to the pth particle The maximum and minimum values ​​of the data covered, is the sth data prototype v s The value of the jth attribute; Step 3.8: Use formula (3) to get the sth trapezoidal particle size data The trapezoidal granularity of the iter+1th iteration in the jth attribute of The trapezoidal particle size corresponding to the pth particle In formula (3), express The left endpoint of the lower base, express The right endpoint of the lower base, express The left endpoint of the upper base, express The right endpoint of the upper base; Respectively The four endpoints of the p-th particle correspond to the position of the particle, j = 1, ... n; s = 1, ... M, n is the number of attributes; Step 3.9: Use formula (4) to obtain the sth trapezoidal granularity data The trapezoidal granularity of the iter+1th iteration in the jth attribute of The position of the pth particle in In formula (4), in, Respectively The positions of the pth particle corresponding to the four endpoints; express The velocity of the pth particle in ; Step 3.10: Using formula (5), we can get The velocity of the pth particle in In formula (5), and yes Two random vectors for the p-th particle in , and are uniformly distributed in the interval [0,1]; for The particle size quality, is the sth optimal trapezoidal granularity data The optimal trapezoidal granularity for the iter-th iteration of the j-th attribute Particle size quality; Step 3.11: Calculate using formula (6) The trapezoidal particle size corresponding to the pth particle Coverage In formula (6), express the number of data points covered; yes The corresponding trapezoidal fuzzy membership function, x kj is the jth attribute of the kth patient information, and has: In formula (7), and They are The left and right endpoints of the lower base, and They are The left and right endpoints of the upper and lower edges of , when the upper boundary contains only one element, let Step 3.12: Calculate using formula (8) Specificity In formula (8), express The maximum value of the data covered and minimum value The difference between Step 3.13: Calculate using formula (9) Particle size quality Step 3.14, let p+1 be assigned to p, and determine whether p≤P is established. If so, return to step 3.6 to execute; otherwise, it means that the sth trapezoidal granularity data is obtained. The trapezoidal granularity set of the iter+1th iteration in the jth attribute of And execute step 3.15; Step 3.15: Use formula (10) to optimally calculate the sth optimal trapezoidal granularity data The trapezoidal granularity of the iter+1th iteration in the jth attribute of Particle size quality Step 3.16: Judgment Or whether iter>iterNum is established. If so, it means that the sth optimal trapezoidal granularity data is obtained. The trapezoidal granularity of the j-th attribute And execute step 3.17; otherwise, assign iter+1 to iter and return to step 3.5; Step 3.17, let j+1 be assigned to j, and determine whether j≤n is established. If so, return to step 3.4 to execute; otherwise, it means that the sth optimal trapezoidal granularity data is obtained. And execute step 3.18; Step 3.18: Calculate using formula (11) Coverage In formula (11), express the number of data points covered; yes The corresponding trapezoidal fuzzy membership function, x kj is the jth attribute of the kth patient information, and has: In formula (12), and They are The left and right endpoints of the lower base, and They are The left and right endpoints of the upper and lower edges of , when the upper boundary contains only one element, let Step 3.19: Calculate using formula (13) Specificity In formula (13), express The maximum value of the data covered and minimum value The difference between Step 3.20: Calculate using formula (14) Particle size quality Step 3.21, let s+1 be assigned to s, and determine whether s≤M is true. If so, return to step 3.3 to execute; otherwise, it means that M optimal trapezoidal granularity data sets are obtained. and the corresponding particle size quality set 3. The weighted fuzzy clustering method for trapezoidal granularity data according to claim 1, characterized in that: The step 4 comprises: Step 4.1: Calculate the sth optimal trapezoidal granularity data Corresponding weight In formula (15), for The particle size quality, is the maximum value of the particle size quality of M optimal particle size data; Step 4.2: Define and initialize the current number of iterations iter = 1, the iteration stop threshold is ε, and the maximum number of iterations is iterMax; Step 4.3: Use formula (16) to establish the objective function of the iterth iteration In formula (16), c represents the number of clusters, M represents the number of granularities, m represents the fuzzy coefficient, η represents the typicality coefficient, a and b represent the relative importance parameters of membership and typicality, respectively. Represents the sth optimal trapezoidal granularity data The weight of represents the i-th trapezoidal granularity cluster center of the iter-th iteration, Represents the sth optimal trapezoidal granularity data Affiliated to the degree, Represents the sth optimal trapezoidal granularity data belong Typicality; is the penalty factor, and: In formula (17), K is a constant; yes and The distance between them is: In formula (18), Represents the sth optimal trapezoidal granularity data The trapezoidal granularity of the j-th attribute and the i-th trapezoidal granularity cluster center of the iter-th iteration The trapezoidal granularity of the j-th attribute in The distance between them is: In formula (19), p1, p2, p3, and p4 represent four intermediate variables, and express The left endpoint of the lower base, express The left endpoint of the lower base, express The left endpoint of the upper base, express The left endpoint of the upper base, express The right endpoint of the upper base, express The right endpoint of the upper base, express The right endpoint of the lower base, express The right endpoint of the lower base; Step 4.4: Calculate the sth optimal trapezoidal granularity data using formula (20) The i-th trapezoidal granularity cluster center belonging to the iter+1-th iteration degree Get the membership matrix of the iter+1th iteration In formula (20), represents the j-th trapezoidal granularity cluster center of the iter-th iteration; Step 4.5: Calculate the sth optimal trapezoidal granularity data using formula (21) belong Typicality Get the typical degree matrix of the iter+1th iteration Step 4.6: Use equations (22)-(25) to calculate the trapezoidal granularity cluster center of the iter+1th iteration Thus, the trapezoidal granularity cluster center matrix of the iter+1th iteration is obtained In formulas (22)-(25), Respectively represent the i-th trapezoidal granularity cluster center of the iter+1-th iteration The left and right endpoints of the upper and lower sides of the trapezoidal granularity of the j-th attribute; Respectively represent the i-th trapezoidal granularity cluster center of the iter+1-th iteration The left and right endpoints of the lower base of the trapezoidal granularity of the j-th attribute; Step 4.7: If ||U iter+1 -U iter ||<ε or iter>iterMax, then the membership matrix U of the iter-th iteration iter , typicality matrix T iter and cluster center V iter As the optimal membership matrix U * , the optimal typicality matrix T * and the optimal cluster center V * Otherwise, assign iter+1 to iter and return to step 4.

3.

4. A weighted fuzzy clustering system for trapezoidal granularity data, characterized in that: include: Data prototype extraction module: used to generate data prototypes and their clusters corresponding to the medical patient information dataset as the basis for granular data construction; Granularity data construction module: used to construct trapezoidal granularity data and optimize its formation process; Granular data clustering module: used to take granular data as input and perform cluster analysis; Reconstruction error calculation module: used to calculate the reconstruction error and combine it with the optimal membership matrix to complete the grouping of medical patients before consultation; Wherein, the granular data construction module includes: PSO algorithm data initialization unit: determine the particle swarm size P, cognitive constant c1 and social constant c2, random vector r j and s j , inertia constant ω, and the maximum number of iterations iterNum of the particle swarm algorithm; Granularity data initialization unit: used for initializing trapezoidal granularity parameters; Trapezoidal granularity data optimization unit: trapezoidal granularity data constructed by optimizing the weighted fuzzy clustering method according to claim 2 through the PSO algorithm; The granular data clustering module includes: Data initialization unit: determines the value of the number of clusters c of the medical patient information dataset, initializes the cluster center, and sets the threshold ε; Data update unit: used to update cluster centers, membership and typicality functions; Clustering condition determination unit: determines whether the clustering meets the condition || U new -U old ||<ε, if so, stop executing, otherwise continue clustering the granular data; where U new Represents the membership function of the current iteration, U old represents the membership function of the previous iteration.

Citation Information

Patent Citations

  • Method of predicting driving range of all-electric passenger vehicles

    CN103745111A

  • Manufacturing process similarity measurement method

    CN111062574A