An intermittent process mode partitioning method of density-weighted and similar label assigned density peak clustering

By introducing weighting coefficients and relative distances to adjust local density, and combining modal evaluation indicators and data sample allocation strategies, the problem of density imbalance in the modal division of intermittent processes is solved, achieving more accurate selection of modal centers and reasonable modal division.

CN117235544BActive Publication Date: 2026-05-12BEIJING UNIV OF CHEM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING UNIV OF CHEM TECH
Filing Date
2023-10-29
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing intermittent process mode partitioning methods based on density peak clustering fail to effectively consider the imbalance in the density distribution of data samples, resulting in inaccurate selection of mode centers and incorrect allocation strategies for the remaining data samples, thus affecting the rationality of the mode partitioning results.

Method used

We introduce weighting coefficients to adjust the local density of data samples in low-density areas. Combining relative distance and local density, we determine the optimal number of modes through modality evaluation indicators and construct an allocation strategy for the remaining data samples to ensure accurate selection of mode centers and reasonable allocation of labels.

Benefits of technology

This improves the rationality of mode segmentation results for intermittent processes, avoids errors in mode center selection and label assignment, and achieves more accurate mode segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117235544B_ABST
    Figure CN117235544B_ABST
Patent Text Reader

Abstract

The application discloses a kind of density weighting and similar label allocation density peak clustering intermittent process modal division method, this method first considers the density distribution imbalance of intermittent process data sample, introduces weight coefficient to adjust the local density of low-density area data sample, obtains intermittent process data modal center;Then define modal evaluation index to determine the optimal modal number of intermittent process;Finally, the allocation strategy of the remaining data sample points of intermittent process is constructed, the modal division of intermittent process is realized.The application does not need the initial modal center and modal number of intermittent process data as input parameter, fully considers the influence of the density distribution imbalance of intermittent process data sample on modal center selection, avoids the allocation error of modal label by the allocation strategy of the remaining data sample points of intermittent process constructed, can realize the modal division of intermittent process, and improves the rationality of intermittent process modal division result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intermittent process monitoring technology, and in particular relates to an intermittent process mode segmentation method based on density-weighted and similarity label allocation density peaks clustering (WSDPC). Background Technology

[0002] Batch processes are an important production method in modern industry and have been widely used in fields such as chemical engineering, pharmaceuticals, and microelectronics. The frequent operational changes and complex production processes of batch production processes give them multimodal characteristics. Reasonable classification of the modes of batch processes can provide a basis for modal modeling and promote the improvement of the modeling accuracy of multimodal batch processes.

[0003] The intermittent process mode partitioning method based on density peak clustering selects mode centers by constructing a decision graph and allocates remaining data samples based on the local density and relative distance of the intermittent process data samples. However, this method does not consider the density imbalance of the intermittent process data samples, reducing the accuracy of mode center selection. Furthermore, its allocation strategy for remaining data samples may lead to the propagation of incorrect mode labels, affecting the rationality of the mode partitioning results. Therefore, to fully consider the density imbalance of the intermittent process data samples, this paper proposes a density-weighted and similarity label allocation density peak clustering intermittent process mode partitioning method. By introducing weight coefficients to adjust the local density of low-density region data samples, it accurately selects mode centers. It then combines relative distance, local density, and ε-nearest neighbors to construct an allocation strategy for the remaining data samples of the intermittent process, thereby improving the rationality of the intermittent process mode partitioning results. Summary of the Invention

[0004] This invention aims to improve the rationality of intermittent process mode segmentation results. It proposes a density-weighted and similar label-assigned density peak clustering method for intermittent process mode segmentation, comprising the following steps:

[0005] Step 1: Collect multiple batches of intermittent process data, standardize the process data, adjust the local density of low-density data samples in the intermittent process using the introduced weighting coefficients, and calculate the decision value for each candidate mode center.

[0006] Step 2: Determine the optimal number of modes for the batch process using the defined Mode Evaluation Index (MEI), and obtain the mode centers of the batch process data using the decision values;

[0007] Step 3: Construct an allocation strategy for the remaining data samples of the intermittent process, obtain the partitioning results under the optimal number of modes, and realize the mode partitioning of the intermittent process.

[0008] Step one specifically includes:

[0009] Collect batches of intermittent process data X(I×J×K), where I is the batch number, J is the number of variables, and K is the number of sampling points. Considering the data differences between different batches of the intermittent process, average the data along the batch direction. Further standardize each process variable by subtracting the mean and dividing by the standard deviation to obtain the intermittent process modal classification dataset.

[0010] right The local density ρ of each data sample in the intermittent process is calculated based on equations (1) and (2). j and relative distance δ j for

[0011]

[0012]

[0013] In the formula, e is the natural base, and d jh For intermittent process data sample points x j and x h The Euclidean distance between them; d c ρ is the cutoff distance parameter. j and ρ h These represent sample points x of the intermittent process data. j and x h The local density.

[0014] Based on the distance between intermittent process data samples, the density contribution r of intermittent process data samples at different distances to the current intermittent process data sample is defined. jh for

[0015]

[0016] In the formula, j and h represent the sampling point numbers of the intermittent process data. The local density mean of the intermittent process data is calculated. ρ j Greater than Intermittent process data samples and ρ j Less than The local density mean of the intermittent process data samples is divided as the upper limit of the weighting coefficient w. max Normalized density contribution r jh to 1 and w max The weighting coefficient w is obtained in the interim. jh for

[0017]

[0018] In the formula, r jh Data sample x representing an intermittent process h For intermittent process data sample x j The degree of density contribution; and The maximum and minimum values ​​of the contribution to density.

[0019] The local density ρ of the intermittent process data sample is recalculated using weighting coefficients. j 'for

[0020]

[0021] Considering the distribution differences between modal and non-modal centers in the decision graph of intermittent process data, calculate the local density mean of the intermittent process data samples. And the standard deviation of the relative distance σ(δ), select a local density greater than Alternatively, intermittent process data samples with a relative distance greater than σ(δ) are used as candidate mode center sets, and the relative distance δ of the data samples in the set is calculated. j 'for

[0022]

[0023] In the formula, This represents the distance between the two most distant intermittent process data samples in the set.

[0024] Considering the high-dimensional data characteristics of intermittent processes, the similarity NC(·) between two intermittent process data samples x and y with d dimensions is:

[0025]

[0026] In the formula, |x n -y n | represents the difference between samples x and y in dimension n of an intermittent process data set; span n The range of values ​​for the intermittent process data sample in dimension n.

[0027] To convert the similarity information between data samples of the intermittent process into a distance matrix, equation (7) is transformed to obtain the distance calculation function ND(·) for high-dimensional data samples of the intermittent process.

[0028] In the formula, a is a constant.

[0029] The distance matrix between the intermittent process data samples is calculated from equation (8). Substituting this matrix into equations (5) and (6) yields the local density and relative distance of the intermittent process data samples.

[0030] The decision value γ for each candidate mode center in the intermittent process is calculated using local density and relative distance. j 'for

[0031] γ j '=ρ j '·δ j ' (9)

[0032] Step two specifically includes:

[0033] To obtain the optimal number of modes for the intermittent process, the decision values ​​obtained in step one are sorted in descending order, and the top ten values ​​are used to define the Modal Evaluation Index (MEI).

[0034]

[0035]

[0036]

[0037] In the formula, γ descendf ' represents the f-th decision value after the decision values ​​are sorted in descending order; F f =f represents the number of modal divisions for the intermittent process; f = 1, 2, ..., 10 represents the sequence number of the decision value; and Representing γ descendf 'and F f The standardization results.

[0038] When MEI f F corresponding to the minimum f This represents the optimal number of modes. At this point, the data sample points corresponding to the first f decision values ​​are selected as the mode centers of the intermittent process.

[0039] Step three specifically includes:

[0040] To ensure high similarity among intermittent process data samples located in the same density distribution region, a similarity index S for intermittent process data samples is defined, taking into account both the distance and local density between the samples. jh for

[0041]

[0042] In the formula, e is the natural base; ND(x) j ,x h (x) represents the data sample point x for the intermittent process. j and x h Distance between; |ρ j '-ρ h '| represents the sample point x of the intermittent process data.j and x h The local density difference between them.

[0043] When allocating the remaining data samples of the intermittent process, the sample points around the mode center obtained in step two are first allocated to the corresponding mode. Then, the mode label is propagated outward with a small ε neighborhood. When the number of unallocated sample points remains unchanged, the similarity index is calculated and allocated to the mode to which the data sample with the highest similarity and existing mode label belongs.

[0044] To ensure the temporal order of the intermittent process mode division, the temporal constraint T for cross-modal sample point allocation is calculated as follows:

[0045]

[0046] In the formula, e is the natural base; x t For intermittent process data sample points assigned across modalities, the sampling time is t; c i The sampling time is i, representing the modal center before and after the sample point; |ti| is the difference between the sampling time of the cross-modal assignment point and the modal center point.

[0047] x t The modes are redistributed to modes with larger T values ​​to complete the mode division of the intermittent process, while ensuring the temporal order of the mode division results.

[0048] The advantages of this invention are as follows: Addressing the issues of intermittent process mode segmentation being easily affected by initial mode centers and the unbalanced density distribution of intermittent process data samples affecting the accuracy of mode center selection, a weighting coefficient is introduced to adjust the local density of data samples in low-density regions, accurately obtaining the mode centers of intermittent process data. A modal evaluation index is defined to determine the optimal number of modes for the intermittent process, and an allocation strategy for the remaining data samples of the intermittent process is constructed, avoiding errors in mode label allocation and enabling reasonable mode segmentation of the intermittent process. Attached Figure Description

[0049] Figure 1 This is a flowchart of an intermittent process mode partitioning method based on density-weighted and similar label-assigned density peak clustering, as described in this invention.

[0050] Figure 2 This is a local density variation map of the data sample from the intermittent process;

[0051] Figure 3 These are the decision values ​​of the top ten candidate mode centers in the intermittent process data;

[0052] Figure 4 This is the optimal mode number discriminant diagram for the penicillin fermentation process;

[0053] Figure 5This is a diagram showing the modal partitioning results of the method described in this invention. Detailed Implementation

[0054] The present invention will be further described below with reference to examples and accompanying drawings. It should be noted that the embodiments do not limit the scope of protection claimed by the present invention.

[0055] Example

[0056] The penicillin fermentation process is a typical multimodal batch process. Using the penicillin fermentation process simulation platform (Pensim V2.0), 25 batches of process data {Xi(400×17)} were generated under different initial conditions and Gaussian noise, where 1≤i≤25. The duration of each batch was 400 hours and the sampling interval was 1 hour. The penicillin fermentation process variables are shown in Table 1.

[0057] Table 1 Variables in the penicillin fermentation process

[0058]

[0059] The modality segmentation process of applying this invention to the penicillin fermentation process is as follows: Figure 1 As shown, the specific steps are as follows:

[0060] Step 1: Average and standardize the 25 batches of intermittent process data along the batch direction to obtain the modality partitioning dataset. First, the distance matrix between data samples is calculated using equation (8), and then the ρ of the sample is calculated using equation (1), where the local density changes are as follows: Figure 2 As shown. From Figure 2 It can be seen that the data sample density is relatively small in the early and middle stages of penicillin fermentation, while the data sample density is relatively large in the middle and later stages, indicating an imbalance in data sample density. Further, the distance matrix is ​​substituted into equations (5) and (6) to calculate the ρ' and δ' of the sample, and the decision value of the candidate mode center is calculated using equation (9).

[0061] Step 2: Use equations (10) to (12) to obtain the optimal number of modes for the intermittent process. The decision values ​​γ' of the top ten candidate mode centers and the values ​​of the mode division index are as follows: Figures 3-4 As shown, the MEI value is the smallest when the number of modes is 3. Therefore, the optimal number of modes for the penicillin fermentation process is 3, and the sample points corresponding to the first three decision values ​​are selected as mode centers.

[0062] Step 3: Assign the sample points around the modality center obtained in Step 2 to the corresponding modality, and propagate the modality label outward in a circular area with a radius of ε = 0.46; when the number of unassigned intermittent process data samples remains unchanged, calculate the similarity index S using equation (13), and assign it to the modality to which the data sample with the largest S and existing modality label belongs; further calculate the temporal constraint T of the cross-modality assignment point using equation (14), and reassign the misclassified modality point to the modality with a larger T value. Assign it to the modality with the largest similarity.

[0063] The modality division results of the penicillin fermentation process obtained through the above steps are as follows: Figure 5 As shown, the first mode ranges from 1 to 47 hours, the second mode ranges from 48 to 130 hours, and the third mode ranges from 131 to 400 hours.

Claims

1. A method for partitioning intermittent process modes using density-weighted and similar label-assigned density peak clustering, characterized in that: Includes the following steps: Step 1: Collect multiple batches of intermittent process data, standardize the process data, adjust the local density of low-density data samples in the intermittent process using the introduced weighting coefficients, and calculate the decision value for each candidate mode center. The penicillin fermentation process is a typical multimodal batch process. Using a penicillin fermentation process simulation platform, 25 batches of process data were generated under different initial conditions and Gaussian noise. , Each batch lasted for 400 hours, with a sampling interval of 1 hour. Variables in the penicillin fermentation process included aeration rate, stirring power, bottom stream acceleration rate, bottom stream temperature, dissolved oxygen concentration, biomass concentration, penicillin concentration, reactor volume, carbon dioxide concentration, pH, reactor temperature, heat production, acid addition flow rate, alkali addition flow rate, cooling water addition flow rate, and heating water flow rate. Step 2: Determine the optimal number of modes for the batch process using the defined Modal Evaluation Index (MEI), and obtain the mode centers of the batch process data using the decision values; Step 3: Construct an allocation strategy for the remaining data samples of the intermittent process, obtain the partitioning results under the optimal number of modes, and complete the mode partitioning of the intermittent process; Based on the distance between intermittent process data samples, the degree of density contribution of intermittent process data samples at different distances to the current intermittent process data sample is defined. for (3) In the formula, j and h represent the sampling point numbers of the intermittent process data; For intermittent process data sample points and Euclidean distance between them; calculate the local density mean of all data for the intermittent process. ,Will Greater than Intermittent process data samples and Less than The upper limit of the weighting coefficient is the sum of the local density mean of the data samples from the intermittent process, and the sum of the local density mean. Normalized density contribution level to 1 and Obtaining weight coefficients for (4); In the formula, Data sample representing intermittent process Data samples of intermittent processes The degree of density contribution; and The maximum and minimum values ​​representing the degree of contribution to density; Recalculate the local density of the intermittent process data samples using weighting coefficients. for (5); In the formula, The cutoff distance parameter is used to calculate the local density mean of the intermittent process data samples. and the standard deviation of relative distance Select a local density greater than or relative distance greater than The intermittent process data samples are used as the candidate mode center set.

2. The intermittent process mode partitioning method based on density-weighted and similar label-assigned density peak clustering according to claim 1, characterized in that: Step one specifically includes: Collect I batches of intermittent process data. Let I be the batch number, J be the number of variables, and K be the number of sampling points. The average of these values ​​along the batch direction is calculated. Then, each variable is standardized by subtracting the mean and dividing by the standard deviation, resulting in the batch process modal classification dataset. ; The local density of each data sample in the intermittent process is calculated based on equations (1) and (2). and relative distance for ; ; In the formula, The base is the natural number. For intermittent process data sample points and The Euclidean distance between them; This is the cutoff distance parameter; and These represent the data sample points of the intermittent process. and Local density; Calculate the relative distance between data samples in the set. for (6); In the formula, This represents the distance between the two most distant intermittent process data samples in the set. Define two intermittent process data samples with d dimensions. and Similarity between ; (7); In the formula, For intermittent process data samples and The difference in dimension n; Let n be the range of values ​​for the intermittent process data samples in dimension n; transform equation (7) to obtain the distance calculation function for the high-dimensional data samples of the intermittent process. for (8); In the formula, It is a constant; Substituting the distance matrix between the data samples of the intermittent process into equations (5) and (6), the local density and relative distance of the data samples of the intermittent process are obtained; Decision values ​​for each candidate mode center in the intermittent process are calculated using local density and relative distance. for (9)。 3. The intermittent process mode partitioning method based on density-weighted and similar label-assigned density peak clustering according to claim 1, characterized in that: Step two specifically includes: Sort the decision values ​​of each candidate mode center of the intermittent process obtained in step one in descending order, and take the top ten values ​​to define the mode evaluation index. for (10); (11); (12); In the formula, This represents the f-th decision value after the decision values ​​are sorted in descending order; The number of modal divisions representing an intermittent process; Indicates the sequence number of the decision value; and They represent and The standardization results; when The minimum value corresponding to This is the optimal number of modes; at this point, the first... The data sample points corresponding to each decision value are the modal centers of the intermittent process.

4. The intermittent process mode partitioning method based on density-weighted and similar label-assigned density peak clustering according to claim 1, characterized in that: Step three specifically includes: Based on the distance and local density between data samples of intermittent processes, a similarity index is defined between data samples of intermittent processes. for (13); In the formula, The base is the natural number; For intermittent process data sample points and The distance between; Represents data sample points for intermittent processes and The local density difference between them; When allocating the remaining data samples of the intermittent process, firstly, the sample points around the mode center obtained in step two are allocated to the corresponding modes, and then a smaller number of samples are allocated to each mode. The domain propagates modal labels outward. When the number of unassigned sample points remains unchanged, a similarity index is calculated and assigned to the modality to which the data sample with the highest similarity already belongs. To ensure the temporal order of mode partitioning, the temporal constraints for allocating sample points across modes are calculated. for (14); In the formula, The base is the natural number; For intermittent process data sample points allocated across modalities, the sampling time is ; This indicates the modal center before and after the sample point, with the sampling time being... ; The difference between the sampling times of the cross-modal assignment point and the modal center point; Will Reassigned to The mode with the larger value is selected; at this point, the mode division of the intermittent process is completed, and the temporal order of the mode division results is guaranteed.