Method, System and Clustering Method for Classifying Debris Detection Signals of Aerospace Hermetic Relays

CN119066473BActive Publication Date: 2025-07-29HARBIN INST OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411234858.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2025-07-29
Estimated Expiration
2044-09-04

AI Technical Summary

Technical Problem

[0005]本发明为了解决航天继电器多余物检测信号分类准确率有待于进一步提高的问题

Benefits of technology

[0054] The present invention can effectively improve the recognition accuracy of debris detection signals for different types of aerospace hermetic relays. For the mixed recognition of 6 types of signals, namely single - component signals, mixed signal I, debris signals, mixed signal II, multi - component signals, and over - sized component signals, the accuracy rate of single - component signals is 96.50%, the accuracy rate of mixed signal I is 95.00%, the accuracy rate of debris signals is 96.00%, the accuracy rate of mixed signal II is 93.00%, the accuracy rate of multi - component signals is 92.00%, and the accuracy rate of over - sized component signals is 98%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119066473B_ABST
    Figure CN119066473B_ABST
Patent Text Reader

Abstract

Method, system and clustering method for classifying debris detection signals of aerospace hermetic relays, belonging to the technical field of aerospace relay debris detection, to solve the problem that the classification accuracy of aerospace relay debris detection signals needs to be further improved. Pulse extraction and frame segmentation are performed on the PIND signals collected in the aerospace relay debris detection; then each segmented signal is used as a data point sample, and each segmented signal is converted into an m-dimensional vector and used as the coordinate of the data point. Based on the corresponding data point coordinates on all segmented segments, k-CAD is used for clustering, and at the same time, the abnormality degree of the outlier is evaluated based on the corresponding data point coordinates on all segmented segments, and then the classification of the detection signal is realized based on the clustering result and the abnormality degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of foreign object detection for aerospace relays, and particularly relates to a method for classifying foreign object detection signals of aerospace relays, a system, and a method for clustering foreign object detection signals. Background Art

[0002] As a switching component, an aerospace relay is an extremely important part of a spacecraft system. Before an aerospace relay becomes a qualified product, it needs to undergo multiple mandatory tests, and foreign object detection is an important part of the mandatory tests. Research shows that the foreign object problem is a serious problem faced by aerospace products. Based on existing relevant data estimates, foreign objects account for approximately 6% of the total number of faults in the aerospace system. A large number of cases show that the foreign object problem cannot be completely avoided, so foreign object detection is currently the only economical and effective solution. The Particle Impact Noise Detection (hereinafter referred to as: PIND) device is a special device for detecting foreign particles in sealed components, and its purpose is to effectively detect loose particles existing in components and semiconductor package cavities, thereby improving the reliability of components, semiconductors and other components. The Particle Impact Noise Detection method (hereinafter referred to as: PIND method) is a standard method in foreign object detection, and the accuracy of the detection results of the Particle Impact Noise Detection system is extremely important. Similarly, the accuracy of the foreign object detection results of aerospace relays based on the PIND method is extremely important for the quality and reliability of aerospace relays.

[0003] Currently, one of the difficulties in judging the results of foreign object detection of aerospace relays is the accurate identification of the foreign object collision signal inside the test piece and the component mechanical vibration signal (referred to as the component signal for short). It should be noted that in many cases, these two signals will appear simultaneously and be mixed together to form a mixed signal. In the mixed signal, how to effectively judge the foreign object signal has always been a difficult problem in this field. Since the foreign object signal is not easy to accurately identify, it is often considered to give priority to judging the component signal. The currently widely used method for identifying the component signal is the pulse occurrence time method, that is, it is considered that component pulses always appear at equal intervals. However, based on long-term aerospace relay foreign object detection results and analysis, the component signal is not always a single group of pulse sequences that appear continuously and at equal intervals. This characteristic leads to many misjudgment and missed judgment problems based on the pulse occurrence time method, and it is precisely because of this problem that there is still a large room for improvement in the discrimination method of the existing Particle Impact Noise Detection method.

[0004] In addition, as Figure 1 shown, generally the movement trajectory of foreign objects is random, and the movement trajectory of components is periodic. Figure 1 (a) in Figure 1(b) in it shows the pulses generated by the collision of debris particles and components at different times. In Figure 1 (a), t is the time axis, S is the displacement axis, H is the height of the cavity of the hermetic relay moving with equal amplitude in a single cycle, and t k (k = 1, …, ∞) is the collision time of the debris particles, and t m (m = 1, …, ∞) is the collision time of the components. The motion trajectories of the upper and lower walls of the cavity are represented by the yellow line and the green line respectively, which shows the motion state of the cavity. In Figure 1 (b), the reference period is used to illustrate the periodic characteristics of the vibration. And Figure 1 (a) The red trajectory describes the motion state of the debris in the periodic vibration system. The debris collides repeatedly between the upper and lower walls, and its path is irregular and difficult to predict. By Figure 1 it can better understand the manifestation forms of the debris signal and the component signal obtained by the PIND method detection. In fact, the debris detection signals of aerospace hermetic relays obtained based on the PIND method show various types. However, previous studies have a simple classification of the loose particle detection signals in the reliability assessment of aerospace hermetic relays, an incomplete analysis of the waveform manifestation forms, and often start from the static characteristics of the signals, making it difficult to accurately identify various types. Therefore, the accuracy rate of debris detection is relatively low. To solve these problems, a more reasonable and systematic signal processing method needs to be introduced, especially for the problem of the aliasing of non-stationary debris signals and stationary component signals in the debris detection signals. Therefore, there is still a relatively large room for improvement in the recognition effect of the existing aerospace hermetic relay debris detection signal recognition method for debris. To ensure the high safety and high stability requirements of spacecraft, for aerospace relays, the final recognition accuracy rate needs to be further improved. Summary of the Invention

[0005] The present invention aims to solve the problem that the classification accuracy rate of the debris detection signals of aerospace relays needs to be further improved.

[0006] A method for classifying debris detection signals of aerospace hermetic relays, comprising the following steps:

[0007] Perform pulse extraction on the PIND signals collected in the debris detection of aerospace relays, and frame the extracted pulse signals; then fold the framed signals;

[0008] Represent the pulse with the coordinates of the peak point of each pulse in the frame where the pulse is located. Display the peak points of the pulses on all framed segments on the x - y coordinate axes, with the x - axis being the length direction of each frame of the signal and the y - axis being the amplitude direction of each frame of the signal;

[0009] Based on the peak point coordinates of the pulses corresponding to all segmented frames on the x-y coordinate axes, multiple data points in the x-y coordinate system are obtained. For the multiple data points, k-CAD is used for clustering. At the same time, the degree of abnormality of the outliers is evaluated based on the peak point coordinates of the pulses corresponding to all segmented frames on the x-y coordinate axes, and then the classification of the detection signal is realized based on the clustering results and the degree of abnormality.

[0010] Further, in the process of frame segmentation for the extracted pulse signals, the vibration period length of the vibration table is used as the length of the frame segmentation window for frame segmentation.

[0011] Further, before frame segmentation of the pulse signals, adjacent pulses with a time interval less than the set threshold are merged into one pulse.

[0012] Further, the process of using k-CAD for clustering the multiple data points in the x-y coordinate system includes the following steps:

[0013] The set of data points corresponding to the peak points of the pulses corresponding to all segmented segments on the x-y coordinate axes is denoted as the data set C:

[0014] C = {x i ∣ i = 1, 2, …, n}, x i = [x i1 , x i2 , …, x im

[0015] where, x i represents the data point sample, [x i1 , x i2 , …, x im is the coordinate of x i , and n is the number of data point samples;

[0016] Based on the optimal number of clusters k, the k-means clustering algorithm is used to cluster the data set C, and the class label L = [l1, l2, …, l n of each point is obtained, where l i represents the class of the signal point x i , and l i ∈ {1, 2, …, k}; for each class j, j = 1, 2, …, k, all sample points belonging to this class are extracted and denoted as the set C j ; the calculation formula for the center μ j of class j is:

[0017]

[0018] where, |C j | is the number of sample points in class j;​

[0019] The calculation formula for category j based on the k-CAD algorithm is as follows:

[0020] Furthermore, the optimal number of clusters k is determined according to the following steps:

[0021] Use the k-means clustering algorithm to divide the dataset C into k categories. By calculating the SSE under different k values and plotting the relationship between the k value and the sum of squared errors within the clusters, the k value at the inflection point is taken as the finally determined optimal number of clusters.

[0022] Furthermore, during the process of evaluating the degree of abnormality of outliers based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, the AE-LOF algorithm is used for the evaluation of the degree of abnormality. The specific process includes:

[0023] For the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, first obtain the set of associated points as A i ; for each data point p i , the set of associated points A i is as follows:

[0024] A i ={p j ∈C\{p i}∣d(p i , p j )≤R i}

[0025] This set contains all other data points within the radius R of p i ; that is, d i (·) represents the k-nearest neighbor distance; k

[0026] For each point p in the dataset j , initialize a frequency counter F j =0; for each data point p i , check whether each other point p j is in A i . If it is, then F j =F j +1; after traversing all points, each F j will represent the total number of times the point p j appears within the radius of all other points as an associated point;

[0027] Then calculate AE-LOF(p j ):

[0028]

[0029] Among them, N k (·) represents the set of neighbor points, and |N k (P)| represents the set capacity; LRD(·) represents the reciprocal of the average reachability distance from a point to its surrounding neighbor points; Z is the internal adjustment coefficient, and w = Z×k is the external adjustment coefficient;

[0030] Using AE-LOF(p j ) to evaluate the degree of abnormality of the outliers of the peak points of the pulses corresponding to all the framed segments on the x-y coordinate axes.

[0031] Furthermore, the process of classifying the detection signal based on the clustering result and the degree of abnormality includes the following steps:

[0032] For multiple data points in the x-y coordinate system, calculate the pulse duty cycle D of the pulse corresponding to each point. When D≥the duty cycle threshold, it is judged as an oversized component signal; otherwise, continue the judgment;

[0033] Then calculate all points' Among them, |C| represents the capacity of the data set C corresponding to all points, and μ j represents the mean of the data set C; when CAD(xi)≤the clustering threshold and the AE-LOF algorithm detection result has no outliers, the judgment result is a single-component signal; when CAD(xi)≤the clustering threshold and the AE-LOF algorithm detection result has outliers, the judgment result is a mixture of a single-component signal and a foreign particle signal; otherwise, continue the judgment;

[0034] Record the calculation result of the k-CAD algorithm corresponding to the optimal clustering number k as CAD k (xi). If CAD k (xi)≤the clustering threshold and the AE-LOF algorithm detection result has no outliers, the judgment result is a multi-component signal; if CAD k (xi)≤the clustering threshold and the AE-LOF algorithm detection result has outliers, the judgment result is a mixture of a multi-component signal and a foreign particle signal; if CAD k (xi)>the clustering threshold and the AE-LOF algorithm detection result has outliers, the judgment result is a foreign particle signal.

[0035] A classification system for foreign particle detection signals of aerospace hermetic relays, including:

[0036] Pulse extraction module: Extract pulses from the PIND signals collected in the foreign particle detection of aerospace relays;

[0037] Pulse segmentation module: Frame the extracted pulse signals;

[0038] Signal folding and coordinate projection module: Fold the framed signals; then represent each pulse by the coordinates of the peak point of the pulse in the frame where the pulse is located, and display the peak points of the pulses on all framed segments on the x-y coordinate axes, with the x-axis being the length direction of each frame of the signal and the y-axis being the amplitude direction of each frame of the signal;

[0039] Detection signal classification module: Based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, obtain multiple data points in the x-y coordinate system. For the multiple data points, use k-CAD for clustering, and at the same time evaluate the degree of abnormality of the outliers based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, and then classify the detection signals based on the clustering results and the degree of abnormality.

[0040] Furthermore, the detection signal classification module includes a clustering result calculation unit, an outlier abnormality degree evaluation unit, and a classification unit; among them,

[0041] Clustering result calculation unit: Based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, obtain multiple data points in the x-y coordinate system. Calculate the CAD values of all points for the multiple data points, and at the same time, for the multiple data points, use k-CAD for clustering based on the optimal number of clusters k to obtain k-CAD values;

[0042] Outlier abnormality degree evaluation unit: Evaluate the degree of abnormality of the outliers based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes;

[0043] Classification and identification unit: Classify the detection signals according to the results of the clustering result calculation unit and the results of the outlier abnormality degree evaluation unit.

[0044] A debris detection signal clustering method for classifying debris detection signals in aerospace hermetic relays includes the following steps:

[0045] Extract pulses from the PIND signals collected in the debris detection of aerospace relays, and frame the extracted pulse signals; then take each framed signal as a data point sample x i , and convert each framed signal into an m-dimensional vector [x i1 ,x i2 ,…,x im , and use [x i1 ,x i2 ,…,x im as the coordinates of the data point x i ; Denote the set of all data points as the data set C:

[0046] C = {x i∣i = 1, 2, …, n}, x i = [x i1 , x i2 , …, x im

[0047] where x i represents the data point sample, and [x i1 , x i2 , …, x im are the coordinates of x i , and n is the number of data point samples;

[0048] Based on the optimal number of clusters k, the k - means clustering algorithm is used to cluster the data set C to obtain the class label L = [l1, l2, …, l n , where l i represents the class of the signal point x i , and l i ∈{1, 2, …, k}; for each class j, j = 1, 2, …, k, all the sample points belonging to this class are extracted and denoted as the set C j ; the formula for calculating the center μ j of class j is:

[0049]

[0050] where |C j | is the number of sample points in class j;

[0051] The formula for class j based on the k - CAD algorithm is:

[0052] k - CAD(C j ) is used for the classification of debris detection signals of aerospace hermetic relays.

[0053] Advantageous effects:

[0054] The present invention can effectively improve the recognition accuracy of debris detection signals for different types of aerospace hermetic relays. For the mixed recognition of 6 types of signals, namely single - component signals, mixed signal I, debris signals, mixed signal II, multi - component signals, and over - sized component signals, the accuracy rate of single - component signals is 96.50%, the accuracy rate of mixed signal I is 95.00%, the accuracy rate of debris signals is 96.00%, the accuracy rate of mixed signal II is 93.00%, the accuracy rate of multi - component signals is 92.00%, and the accuracy rate of over - sized component signals is 98%. Description of the Drawings

[0055] Figure 1 It is a schematic diagram of the movement trajectories and collision pulses of debris particles and components under periodic vibration. ​

[0056] Figure 2 It is a schematic diagram of the pulse extraction and waveform segmentation process;

[0057] Figure 3 It is a schematic diagram of the pulse framing and folding steps of the foreign matter signal;

[0058] Figure 4 It is a waveform diagram of the original component signals of the mixed signal;

[0059] Figure 5 It is a display diagram of the pulse points after processing the mixed signal;

[0060] Figure 6 It is a schematic diagram of outliers and clusters;

[0061] Figure 7 It is a schematic diagram of the local outlier detection algorithm;

[0062] Figure 8 It is a schematic diagram of the outlier anomaly degree based on the LOF algorithm;

[0063] Figure 9 It is a schematic diagram of the outlier anomaly degree based on the AE-LOF algorithm;

[0064] Figure 10 It is a comparison diagram of the LOF value and AE-LOF value of a certain mixed signal data;

[0065] Figure 11 It is a schematic diagram of the calculated values before and after based on the average distance method of the k clustering centers;

[0066] Figure 12 It is a schematic diagram of determining the number k of clusters after clustering based on the elbow method;

[0067] Figure 13 It is a flowchart of the classification system of the present invention;

[0068] Figure 14 It is a recognition effect diagram of a single-component signal;

[0069] Figure 15 It is a recognition effect diagram of the foreign matter particle signal;

[0070] Figure 16 It is a recognition effect diagram of the mixture of a single-component signal and a foreign matter particle signal - mixed signal I;

[0071] Figure 17 It is a recognition effect diagram of the mixture of a multi-component signal and a foreign matter particle signal - mixed signal II;

[0072] Figure 18 It is a recognition effect diagram of a multi-component signal;

[0073] Figure 19Line graph comparing the accuracy rates of the old and new discrimination methods for different signal types. Detailed implementation manners Detailed implementation manner 1:

[0075] This implementation manner is a method for classifying debris detection signals of aerospace hermetic relays, including the following steps:

[0076] Step 1. PIND signal preprocessing:

[0077] S100. Extract the pulses in the PIND signal. In this implementation manner, the three-threshold method is used to extract the effective pulses in the signal. It should be noted that the present invention includes, but is not limited to, using the three-threshold method to extract the effective pulses in the signal, and it can also be the two-threshold method, etc., as long as the effective pulses in the signal can be extracted.

[0078] The process of using the three-threshold method to extract the effective pulses in the signal is as follows:

[0079] (a) Determination of the reference threshold: First, take the overall collected signal as the object, calculate the average energy as the reference for the threshold, and determine the pulse main energy threshold, start energy threshold, and end energy threshold according to the average energy, as Figure 2 (a) shows. The threshold parameters need to be adjusted according to the actual situation and signal characteristics.

[0080] (b) Preliminary search: According to the principle of the search algorithm, process the single-pulse signal, as Figure 2 (b) shows. Determine whether there are pulses exceeding the threshold. When the average energy value or amplitude of a certain segment of signal data reaches the set threshold, it is determined that the pulse main body exists. Subsequently, continue to calculate the energy of each segment backward until the energy value is lower than the set main energy threshold. Count the peak points of the energy value, and obtain the time start point and position coordinates of the segment with the highest energy.

[0081] (c) Detailed search: On the basis of the preliminary search, take the pulse energy peak point as the starting point, divide it with the set time length as the step size, and search forward and backward for the exact starting point and exact ending point of the pulse occurrence moment. As Figure 2 (c) shows.

[0082] So far, the extraction work of the pulses in the signal is completed, ensuring the accuracy and integrity of the pulses.

[0083] S200. Pulse merging, splitting, and folding:

[0084] For the extracted pulse signal, use the vibration period time length of the vibration table as the length of the frame division window (i.e., the splitting step size) for frame division, that is, split according to the splitting step size, and then fold the split signal, as Figure 2As shown in (d) thereof. Before segmentation, considering the acoustic reflection characteristics and sensor sensitivity characteristics, adjacent pulses with a time interval less than a set value are merged into one pulse, that is, adjacent pulses with too close distances are merged.

[0085] (1) The segmentation is carried out with the period of the vibration cycle signal as the step size (or the length of the separable frame window).

[0086] (2) The merging of adjacent pulses is completed before the segmentation program is carried out. That is, many pulses are screened out by the three-threshold method. For example, pulses a1, a2, a3,..., an obtained from a certain signal. Among them, pulses a1 and a2 are very close. The distance between the end point of pulse a1 and the starting point of pulse a2 is less than the set value Δdmin = 500, then a1 and a2 are merged. That is, the end point of pulse a2 is assigned to the end point of pulse a1. Therefore, the actual range of the merged pulse a1-2 is: [the starting point of a1, the end point of a2].

[0087] (3) The folding program is carried out after the segmentation program. As Figure 3 shown, specifically, the respective sub-frame data obtained after the execution of the segmentation program are superimposed in a manner parallel to the x-y axis and perpendicular to the z axis in Figure 3 . Finally, the peak points of the pulses in each sub-frame are extracted as the representatives of the pulses and displayed to form Figure 3 the final two-dimensional scatter plot in

[0088] Actually, the x-axis is the length direction of each frame of signal, the y-axis is the amplitude direction of each frame of signal, and the z-axis is the superimposing direction of each frame of signal. In actual drawing, the range on the x-axis is the vibration cycle time multiplied by the sampling rate of the detection system (500k in this embodiment) and is used to represent the data position points.

[0089] S300. Pulse display: Represent the pulse with the coordinates of the peak point of each pulse in the sub-frame where the pulse is located, and display the peak points of the pulses on all the segmented segments on the x-y coordinate axes. As Figure 2 shown in (e) thereof are the coordinates of the pulse peak points after folding.

[0090] Step 2. Display of the pulse points after the preprocessing of the mixed signal:

[0091] Macroeconomic waveform of the mixed signal and image representation of the pulse peak point distribution after algorithm processing in the detection of foreign objects in aerospace relays:

[0092] (1) Affected by objective factors such as the shape, mass, material, and collision angle of the foreign object particles, the trajectory of the foreign object signal has the randomness similar to that of the free electron collision signal. Then, the image of the pulse peak points of the foreign object signal is an irregular scatter plot.

[0093] (2) The image of the pulse peak points of the component signals has the shrinkage of the cluster data set, similar to a narrow band-shaped cloud cluster.

[0094] (3) As Figure 4 shown in Figure 5 is the detection signal diagram of the mixed signal. It is difficult to see the pattern from this detection signal diagram. In fact, the pulse peak point distribution image of the mixed signal has two characteristics, which are: (1) The main part of the pulse peak points is distributed in clusters similar to narrow band-shaped cloud clusters, and most of them are distributed inside and around the cluster, as shown near the abscissa of 0.4×10 4 in

[0095] shown in Figure 5 . (2) A small part of the pulse peak points are randomly distributed in other areas, similar to scattered stars.

[0096] Actually, the pulse point distribution after the preprocessing of the mixed signal reflects both the periodicity of the component signal and the randomness of the foreign object signal. As can be seen from Figure 5 , the main problem of the present invention is to determine whether there is a foreign object signal similar to scattered stars outside the component signal similar to a narrow band-shaped cloud cluster. Further understood as, detecting whether there is a foreign object particle signal around the component signal with better periodicity.

[0096] It should be further noted that the component signal similar to a narrow band-shaped cloud cluster has a certain mobile redundancy range, and only the pulse peak points of the foreign object signals far from the "narrow band-shaped cloud cluster" range need to be clearly identified.

[0097] Step 3: For all points on the x-y coordinate axes, use the k-CAD algorithm to cluster the data points, evaluate the degree of abnormality based on the clustering results and in combination with the AE-LOF algorithm, and then use the classification system to classify the detection signals. The AE-LOF algorithm and the k-CAD algorithm are described separately below.

[0098] (A) For the peak point coordinates of the pulses corresponding to all segmented segments on the x-y coordinate axes, use the AE-LOF algorithm to evaluate the degree of abnormality.

[0099] Principle of the LOF algorithm (conventional existing algorithm):

[0100] In 2000, Markus M. Breunig et al. proposed the Local Outlier Factor (LOF) method. This method is used to identify abnormal data, and its core idea is to judge whether a point is abnormal by comparing the local density of a point with its neighboring points. Therefore, compared with the distance-based anomaly algorithm, density-based outlier detection can detect a type of abnormal data - local outliers. In the LOF algorithm, the degree of abnormality of the data is defined based on the concept of relative density, so it has a strong dependence on the surrounding data.

[0101] Figure 6 is a two-dimensional data set, which contains two clusters C1 = {c 11 , c 12 , …, c 1n} and C2 = {c 21 , c 22 , …, c 2n}, two types of outliers {o1, o3} and {o2, o4}. C1 is a dense-shaped cluster, and C2 is a sparse-shaped cluster. {o1, o3} are global outliers, and {o2, o4} are local outliers. Generally speaking, global outliers are easy to identify, but local outliers are difficult to detect. If we want to obtain local outliers, we need to adjust the distance parameter Dist(x). Suppose we make Dist(x) less than the minimum distance between C1 and o2, then some data points {c 2i , c 2j} in C2 may be mislabeled as outliers. Based on the above analysis, we need to obtain an algorithm with the ability to identify local outliers, that is, the LOF algorithm. In a given data set, if for any data point, the points within its local range are very dense, then this data point is considered likely to be a normal data point. Otherwise, it may be an outlier.

[0102] Let there be a data set sample C. Suppose there are a total of n samples and the data dimension is m. For

[0103]

[0104] any two points, the Euclidean distance can be expressed as:

[0105]

[0106] The k-th nearest point P to O is the k-th nearest distance point of O, that is, d k (O) is the k-th distance of point O, defined as follows:

[0107] d k (O) = d(O, P)

[0108] If the following conditions are met:

[0109] (1) There are at least k points P′ ∈ C\{O} in the set such that d(O, P′) ≤ d(O, P); where C\{O} represents the difference set of C with respect to {O}, and \ is the difference set symbol;

[0110] (2) There are at most k - 1 points P′ ∈ C\{O} in the set such that d(O, P′) < d(O, P).

[0111] At this time, point P is the k-th point closest to O.

[0112] For example, as Figure 7 shown, C is a set of points centered at O. Both P6 and P7 are the 6th closest points. There are 5 points inside the 6th distance, and a total of 2 points on the 6th distance.

[0113] RD(P, O) (Reachability Distance of P with respect to Object O) is a custom function that obtains the maximum distance. The reachability distance of point P with respect to data point O can be calculated according to the results of equations (3)-(4). For example, in Figure 5 , when the parameter k = 6 is set, the reachability distance of P5 is d k=6 (O), and the reachability distance of P 10 is dist(P 10 , O), where P 10 is outside the k-distance. In this case, the Euclidean distance dist(P 10 , O) needs to be used to represent it. Therefore, the reachability distance of point P with respect to point O is the maximum of the k-distance of point O and the Euclidean distance between P and O.

[0114] reach-dist k (P, O) = max(d k (O), dist(P, O)) (3)

[0115] That is:

[0116] RD(P, O) = max(d k (O), dist(P, O)) (4)

[0117] Among them, dist(·) represents the Euclidean distance.

[0118] Local reachability density (LRD) is the reciprocal of the average reachability distance from P to its surrounding neighbor points. Intuitively, according to the LRD formula, the larger the average reachability distance (that is, the farther the surrounding neighbor points are from point P), the smaller the density of the surrounding neighbor points of this point. Therefore, LRD represents the distance of a point from the closest point group. The lower the LRD value, the farther the closest point group is from this point.

[0119]

[0120] Among them, N k (P) represents the set of neighbor points of P, and |n k (P)| represents the set capacity.

[0121] The Local Outlier Factor (LOF) is the ratio of the average of the Local Reachability Density (LRD) of the k-nearest neighbors of P to the LRD of P:

[0122]

[0123] Furthermore, if point P is not an outlier, the ratio of the average LRD of its neighbors is approximately equal to the LRD of point P because the density of a point is roughly equal to the density of its neighboring points. In this case, LOF is almost equal to 1. On the other hand, if a point is an outlier, its LRD is less than the average LRD of its surrounding neighbors. Then the LOF value will be very high. Generally, if LOF > 1, it may be considered an outlier, but this is not always true. Suppose we know that there is only one outlier in the data, then the point corresponding to the maximum value among all LOF values will be considered an outlier. Obviously, the LOF value of a point directly represents the degree of its abnormality.

[0124] Improved LOF algorithm - Attention Enhancement AE-LOF algorithm:

[0125] Take "attention" as the core term to explain the AE-LOF algorithm. For example, in a social network, assume that A follows many people {B, C, D, E, F,...}. Just because A follows many people does not mean that A has a high degree of attention. Only when many people {B, C, D, E, F,...} also follow A can it be considered that A has a high degree of attention.

[0126] Therefore, the AE-LOF algorithm can be explained as follows: Suppose there is a dataset C consisting of several data points, and the data points in the dataset can be related to each other. Dataset C = {p1, p2,..., p n}, where C contains n data points. The attention radius of data point p i is R i . The Euclidean distance between data points p j and p j is d(p i , p j ).

[0127] The set of associated points is A i . For each data point p i , define the set of associated points A i as follows:

[0128] A i = {p j ∈ C\{p i} | d(p i , p j ) ≤ R i} (7)

[0129] This set contains all i Radius R i Other data points within .

[0130] Frequency statistics: For each point p in the data set j , initialize a frequency counter F j = 0. For each data point p i , check every other point p j Is it in A i If p j In A i In, F j Increase by 1. For each i and each j≠i, if p j ∈A i , then F j =F j +1. After traversing all points, each F j

[0131] Let point p be j As the total number of times the associated point appears within the radius of all other points. In summary, each point p j The number of occurrences of global correlation points F j This can be calculated by iterating over the dataset and updating the counter:

[0132]

[0133] Among them, 1 is the indicator function, when p j In A i The value is 1 when it is in, otherwise it is 0.

[0134] In order to adjust the LOF value of each point to a new value that takes into account its "attention", F is integrated on the basis of the original LOF algorithm. j Therefore, the present invention is implemented by adjusting the LRD calculation method of each point. Specifically, assuming F j The smaller the value, the more points p j The more likely it is to be isolated, the higher its abnormality should be. Then, by introducing F j The method of using the value to calculate the LOF value of these points is used to indirectly improve the abnormality of these points.

[0135] Adjusted AE-LOF(p j ) is defined as:

[0136]

[0137] Among them, Z is the internal adjustment coefficient, usually taking a positive integer, generally taking 1. w is the external adjustment coefficient, usually taking a positive integer, and generally w = Z×k can be taken.

[0138] When w = Z×k is taken, there are two main calculation results: (1) When the point p j is not an outlier, then F j is equal to or approximately equal to k. At this time can be simplified to At this time the calculated value of is approximately equal to 1. Therefore, in this case, the result of AE-LOF(p j ) is almost the same as the calculated result of the original LOF(p j ), ensuring the usability of the result of the original LOF(p j ). (2) When the point p j is an outlier, then F j will be less than k. At this time can be simplified to But at this time the calculated value of will become much larger. Therefore, in this case, the calculated result of AE-LOF(p j ) will be greater than the calculated result of the original LOF(p j ), ensuring that the calculated result of AE-LOF(p j ) can very effectively screen out outliers.

[0139] Through the above steps, the local density of the point p j is adjusted by the number of occurrences of its global associated points, so that those points that are less concerned are enhanced in the outlier degree evaluation. This not only strengthens the analysis of the outlier degree of each point, but also provides more flexibility and adaptability to handle different data and scenarios by considering the new value F j (the number of occurrences of associated points) and Z, w (internal and external adjustment coefficients).

[0140] (A1), the result of the original LOF algorithm

[0141] For mixed-signal data, the LOF value of each point to be detected can be obtained through the LOF algorithm. By setting the outlier detection threshold, outliers can be screened out. Generally speaking, after obtaining the outliers, it is also necessary to evaluate the outlier degree of each outlier, which is beneficial for visual inspection, display, and subsequent evaluation of relevant data. It should be emphasized that the evaluation of the outlier degree of the point to be detected is mainly completed by comparing the LOF values. Therefore, the LOF value is a key parameter. There are defects in the conventional outlier degree evaluation based on the LOF algorithm. Taking the data in Figure 8 as a typical object to illustrate the defects. As Figure 8The figure shows a schematic diagram of the outlier anomaly scores calculated based on the LOF algorithm. The outliers are mainly divided into three regions, and among them, the LOF value of No.1 in region C is the smallest. In fact, compared with the outlier set {No.75, No.40, No.80, No.57, No.60, No.62}, the outlier No.1 is about 4000 points away from the center line. Obviously, the anomaly degree of outlier No.1 is more obvious. Therefore, there are defects in the evaluation of the anomaly degree of outliers based on the LOF algorithm.

[0142] First, take the outlier No.40 in region A as an example to analyze the reason why the LOF value of this point is relatively large. The LOF algorithm is mainly a density-based outlier detection algorithm. To evaluate the anomaly degree of any data point in the dataset to be tested, the anomaly degrees of the k neighbor points of this data point usually need to be referred to. For the outlier No.40 in region A, its k neighbor points (k = 10) mainly include three parts, namely the outliers {No.75, No.80} in region A, the outliers {No.62, No.60, No.57} in region B, and the data points {No.81, No.70, No.74, No.69, No.73} in the central yellow area (because Figure 8 and Figure 9 the display space is limited, these points are not marked in Figure 8 and Figure 9 ). The calculation results show that the order of magnitude of the LRD values of the data points {No.81, No.70, No.74, No.69, No.73} in the central yellow area is all 1×10 -2 , while the order of magnitude of the LRD of the outliers {No.75, No.80} and {No.62, No.60, No.57} is all 1×10 -4 . Obviously, due to the large difference in the order of magnitude of the above data points, generally, the average LRD of the k neighbor points (k = 10) of the outlier No.40 in region A is significantly affected by the 5 data points in the central yellow area, resulting in a relatively high calculated LOF value of the outlier No.40 in region A.

[0143] Secondly, analyze the reason why the LOF value of the outlier No.1 in region C is relatively small. For the outlier No.1 in region C, its k neighbor points (k = 10) are all distributed in the central yellow area and concentrated in area C-1. It can be seen from area C-1 in the figure that the data points in area C-1 are relatively far from the surrounding points, and their LRD values are smaller than those of other data points in the central yellow area. Therefore, it will result in a relatively low calculated LOF value of the outlier No.1 in region C.

[0144] Through the above cause analysis, the LOF value of the outlier No.1 in region C is smaller than that of the outlier No.40 in region A. And so on, the LOF values of the outlier set {No.75, No.80, No.57, No.60, No.62} in regions A and B are larger than the LOF value of the outlier No.1 in region C.

[0145] Table 1 Schematic table of the calculation process of the LOF value of each outlier in the main region

[0146]

[0147] (A2), Results of the AE-LOF algorithm

[0148] For the mixed signal data, Table 2 shows the statistical situation of the association frequencies of the outliers in regions A, B, and C. Sorting according to the statistical values of the association frequencies: Region C < Region B < Region A. Based on the statistical values of the association frequencies, the LOF values of the outliers recalculated by the AE-LOF algorithm are plotted as Figure 9 shown. From Figure 9 it can be seen that, compared with the LOF value of the outliers in region A, the LOF value of the outlier No.1 in region C increases significantly, and the degree of abnormality increases significantly. At the same time, the LOF values of the outliers in region B are also significantly lower than those of the outliers in region A.

[0149] Table 2

[0150]

[0151] From Figure 10 the local enlarged detail diagram in it, it can be seen that for non-outliers, the calculated values of the LOF algorithm and the AE-LOF algorithm are approximately equal, and the overall change trends are the same. For the outlier No.1, the calculated value based on the AE-LOF algorithm is 80 times higher than that of the LOF algorithm. For the outliers {No.62, No.60, No.57}, it is increased by about 3 times. For the outliers {No.40, No.80, No.75}, it is increased by about 7 times. The detection value range of the AE-LOF algorithm is wider, and it can capture more significant outliers (such as the 1620.54 corresponding to the outlier No.1 in Table 2), which indicates that the AE-LOF algorithm is more sensitive in detecting abnormalities, and the AE-LOF algorithm can identify more obvious outliers in anomaly detection.

[0152] In outlier detection based on the LOF algorithm, the LOF result for any point P is influenced by its neighboring points. When some of point P's neighbors exhibit minor anomalies, the LOF value for point P may be somewhat undervalued from a global perspective, potentially causing some outlier anomalies to be overlooked or masked. The attention-enhanced density-based local outlier detection algorithm (AE-LOF algorithm) can, to a certain extent, mitigate this issue in the presence of outlier neighbors. Due to the confidentiality of aerospace testing, experimental data cannot be readily obtained or made public, and they also contain many complex anomalies found in redundant object detection. Due to these factors, the current ability to identify mixed signals for redundant object detection requires further improvement. Future field experiments are planned to increase the number of experimental data and further investigate outlier value detection in redundant object detection signals. Experimental results show that, in general, the AE-LOF calculated value for outliers is several times higher than the LOF calculated value. AE-LOF is more sensitive in detecting anomalies, has a wider detection range, and can capture more significant outliers. The AE-LOF algorithm is effective in detecting and identifying the abnormality of some outliers in the aerospace relay redundant object detection signal.

[0153] In fact, the present invention may also adopt any other improved LOF algorithm to implement evaluation.

[0154] (B) Clustering based on k-CAD is performed for the peak coordinates of the pulses corresponding to all segmented segments on the xy coordinate axis:

[0155] The data point set corresponding to the peak point of the pulse on all the segmented segments on the xy coordinate axis is recorded as data set C, where the data point is x i , assuming there are n samples in total, and the data dimension of each sample (actually the two-dimensional coordinates of the point) is m:

[0156] x i =[x i1 ,x i2 ,…,x im ]

[0157] The goal of k-means clustering is to divide the dataset C into k categories (clusters), and the center point of each category is denoted as μ j (j=1,2,…,k). The objective function of k-means clustering is to minimize the sum of squared errors (SSE) within the cluster:

[0158]

[0159] Among them, δ ij is the indicator function, when point xi When it belongs to category j, δ ij = 1, otherwise δ ij = 0. By calculating the SSE under different k values and plotting the relationship between SSE and k, select the "elbow" position (i.e., the inflection point position) as the optimal number of clusters. The specific steps are as follows: (1) Select a range of k values. Through the research and verification of the present invention, for the debris detection signals of aerospace hermetic relays, k generally does not exceed 5; (2) Perform k-means clustering for each k value and calculate the corresponding SSE; (3) Plot the relationship between k value and SSE; (4) Observe the "elbow" position and the number of data points contained in each cluster to determine the optimal k value.

[0160] Based on the optimal k value, use the k-means algorithm to cluster the dataset C to obtain the class label L = [l1, l2, …, l n , where l i represents the class of the signal point x i , and l i ∈{1, 2, …, k}. For each class j (j = 1, 2, …, k), extract all the sample points belonging to this class and denote it as the set C j . The calculation formula for the center μ j of class j is:

[0161]

[0162] where, |C j | is the number of sample points in class j.

[0163] The calculation formula for class j based on the k-CAD algorithm is:

[0164]

[0165] According to long-term actual test experience, if for all j, the condition k-CAD(C j ) < 1000 is satisfied, then it is considered that the clustering result of the debris detection signal is meaningful.

[0166] For the signal data of a certain component, such as Figure 11As shown in (a), before clustering multi-component signal data points by cluster, the calculated value of the average distance of the original center points usually exceeds 1000. In this case, the calculated value of the average distance of the original center points cannot accurately reflect the cluster characteristics of multi-component signal data, which may lead to misclassification. To solve this problem, the present invention combines the k-means algorithm and the Center Average Distance (CAD) algorithm to propose the k-means Center Average Distance (k-CAD) method. This algorithm can effectively measure the volatility of data points in the signal, thereby evaluating the complexity and composition of the signal.

[0167] Compared with the simple Standard Deviation (STD) algorithm, the Center Average Distance (CAD) algorithm is applicable to various data types and distributions, especially for data with non-normal distributions. In addition, the CAD algorithm is insensitive to outliers and will not be severely affected by extreme values. From Figure 11 (b), it can be seen that the calculation result based on the k-means Center Average Distance (k-CAD) method is 64.5, which is significantly reduced by about 40 times compared with the Figure 11 calculation result in (a). Therefore, it is effective to determine the number of clusters in multi-component signals based on the k-CAD method.

[0168] Using the k-means algorithm requires pre-determining the number of clusters. The present invention uses the Elbow Method to determine the optimal number of clusters. Although in many practical applications, the selection of the "elbow" is very ambiguous because it does not include an elbow with an obvious inflection point. However, in the foreign object detection signal, the attempt to capture the bend using the "slope" is very clear because the pre-processed component signals usually exhibit obvious cluster characteristics. In addition, in foreign object detection, when the number of data points in a certain cluster is less than 15, we consider this data cluster as a foreign object signal because the number of component signals is generally higher than 15. Subsequently, the k-CAD algorithm will skip this clustering to reduce the calculation time and improve the calculation efficiency. As Figure 12 shown, taking Figure 11Taking the data in as an example, when the number of clusters is 2, the maximum deviation multiple is close to 40, which is an obvious inflection point - the "elbow". As the number of clusters increases, the maximum deviation multiple tends to level off and is significantly lower than 40. It should be further noted that the maximum deviation multiple is the ratio used to calculate the difference between adjacent elements in a vector or matrix, and is a parameter used to measure the difference between two variables. In the present invention, according to the experience of foreign matter detection, if the multiple of the maximum deviation value between adjacent elements exceeds 10, it is considered to be an obvious inflection point - the "elbow", and the number of clusters at this time is usually the optimal number of clusters.

[0169] For all points on the x-y coordinate axis, the specific process of measuring signal classification is as Figure 13 shown. Combining the long-term experience of foreign matter detection signal recognition, the present invention proposes a new closed-loop classification system for analyzing and recognizing the types of foreign matter detection signals. This classification system takes the preprocessed signal as input, combines multiple algorithms and multi-level rules to achieve accurate recognition of the types of foreign matter detection signals. The main functions of this classification system include: 1) Input inspection of the preprocessed signal, which can make the input data meet the processing requirements. 2) Calculating the pulse duty cycle, which can quickly screen out signals of oversized components, reduce redundant analysis, and improve the recognition efficiency. 3) Combining the AE-LOF algorithm with the initial CAD algorithm can deeply analyze the composition of the signal, so as to accurately identify single-component signals and mixed signal I. 4) Combining the AE-LOF algorithm with the k-CAD algorithm can accurately identify foreign matter particle signals, mixed signal II, and multi-component signals. It should be noted again that the k-CAD algorithm determines the optimal number of clusters by the elbow method, and then calculates the CAD k (xi) value of each cluster to evaluate the volatility and composition complexity of the detection signal. It should be noted that in the process of finding the optimal number of clusters, the minimum number of data points in a certain cluster is set to not be less than 15, otherwise the cluster will not be recognized as a qualified cluster. According to the experience of foreign matter detection, the number of pulses of general component signals is not less than 15 within the vibration cycle of 5s. Setting this screening rule can quickly and effectively find the optimal k value. 5) For some rare and complex signals, the classification system will prompt that manual recognition is required, forming a closed loop of the recognition process to ensure the reliability of the final result. In summary, this classification system provides an efficient and accurate tool for recognizing foreign matter detection signals through the combination of multiple algorithms and multi-level rules.

[0170] The specific implementation steps are as follows:

[0171] (1) Judgment 1: Calculate the pulse duty cycle for preliminary identification:

[0172] Step1. For the data points in Step 2, ensure the validity of the input data;

[0173] Step 2. Calculate the pulse duty cycle D (Duty cycle). When D ≥ 65%, it is an excessive component signal; when D < 65%, continue to make the following identification. It should be noted that the excessive component signal is a special signal, usually generated by components with severe vibration. Since such signals will cover or interfere with the information of other types of signals, it is impossible to effectively identify the composition of the signals. Therefore, when D ≥ 65%, there is a high probability that this signal is an excessive component signal.

[0174] (2) Identification 2: Denote the result calculated by the CAD algorithm as CAD(xi) (it can be considered that in the k-CAD algorithm, k is first set to 1, and then the k-CAD algorithm corresponding to the optimal k value is used in the subsequent steps). Make an identification based on the results of CAD(xi) and the AE-LOF algorithm:

[0175] Step 3. When CAD(xi) ≤ 1000 and the detection result of the AE-LOF algorithm has no outliers, the identification result is: single-component signal;

[0176] Step 4. When CAD(xi) ≤ 1000 and the detection result of the AE-LOF algorithm has outliers, the identification result is: a mixture of single-component signal and foreign particle signal, that is, mixed signal I.

[0177] (3) Identification 3: Denote the result calculated by the k-CAD algorithm as CAD k (xi), and make an identification based on the results of CAD k (xi) and the AE-LOF algorithm:

[0178] Step 5. If CAD k (xi) ≤ 1000 and the detection result of the AE-LOF algorithm has no outliers, the identification result is: multi-component signal;

[0179] Step 6. If CAD k (xi) ≤ 1000 and the detection result of the AE-LOF algorithm has outliers, the identification result is: a mixture of multi-component signal and foreign particle signal, that is, mixed signal II;

[0180] Step 7. If CAD k (xi) > 1000 and the detection result of the AE-LOF algorithm has outliers, the identification result is: foreign particle signal;

[0181] Step 8. If CAD k (xi) > 1000 and the detection result of the AE-LOF algorithm has no outliers, the identification result is: manual identification is required.

[0182] To better illustrate the classification rules and logic of the classification system of the present invention, five typical types of foreign object detection signals will be analyzed and demonstrated one by one. Specifically, the five typical signals mainly include single-component signals, mixed signal type I, foreign object particle signals, multi-component signals, and mixed signal type II. The calculation results obtained based on the AE-LOF algorithm and the k-CAD algorithm are directly marked in the schematic diagram for convenient comparison and analysis. In addition, there are different specific problems in the identification of different signal types. To better illustrate these problems, different markings and explanations are made on different schematic diagrams. In summary, it can better help readers understand the classification rules, algorithm innovation, and the original intention of the early data preprocessing in the classification system of the present invention.

[0183] Figure 14 What is shown is the scatter plot of the single-component signal under two different abscissa scales. It should be noted that in Figure 14 (a), the range of the abscissa is expanded, and there are two main purposes: (1) To match the interface detected by the PIND instrument. (2) To match the number of samples in a complete sampling of the signal acquisition system. In addition, during the manual identification stage, Figure 14 the range of the abscissa in (a) is also the range used for the dynamic display of the detection results during manual identification. In Figure 14 (b), the range of the abscissa is narrowed, and the main purpose is to clearly show the positions of the center lines of different clustering regions. From Figure 14 (a), it can be seen that all data points are clustered according to k = 1 and 2 respectively, and the corresponding values based on the k-CAD algorithm are 95.40 and 74.14 respectively, and both are lower than 1000. From the comparison results, it is obvious that the change in the k-CAD value corresponding to different clustering numbers k is not obvious, and there are no outliers. In summary, this signal can be identified as a single-component signal.

[0184] Figure 15 What is shown is the scatter plot of the foreign object particle signal. From Figure 15 (a), it can be seen that since the initial CAD value is 4035.60, exceeding the threshold of 1000 in Identification 2, it is necessary to calculate and analyze it again using the k-CAD algorithm. In addition, from the locally enlarged image, it can be known that it is difficult to obtain some data points through manual observation due to image superposition. In fact, there is also the same problem of signal image coverage when manually identifying the detection waveform on the PIND instrument. This also once again reflects the necessity and scientificity of using the identification algorithm. From Figure 15 (b), it can be seen that when k = 2, the k-CAD value is 1844.10, exceeding the threshold of 1000 in Identification 3, and there are multiple outliers. In summary, this signal can be identified as a foreign object particle signal. Further, from Figure 15(b) It can be seen that the outliers obtained by the AE-LOF algorithm are P3 - P5. However, judging from the positions of each data point and the distances between them, points P1 and P2 are more likely to be outliers. To fully analyze this problem, we calculated the LOF values and AE-LOF values of points P1 - P5 respectively, and marked the corresponding values on the points. Through the comparative analysis of the calculated values, the main reasons are as follows: (1) The AE-LOF values of points P1 and P2 are highly dependent on the surrounding data points; (2) The distribution of the surrounding data points of P1 and P2 is loose; (3) Points P1 and P2 receive a high degree of attention from the surrounding points obtained; (4) The order of magnitude of the vertical coordinates of each point is much lower than that of the horizontal coordinates, which is also an important reason for the difference between the visually observed results and the calculation results based on the algorithm. Generally, the larger the outlier value of a certain point, the farther it is from the center line. However, the evaluation of actual outliers often needs to consider multiple factors such as the distribution, quantity, and attention of the surrounding points.

[0185] The initial CAD value is 678.74, which is lower than the threshold of 1000 in Judgment 2. From Figure 16 (b) It can be seen that multiple outliers are captured by the AE-LOF algorithm. Among them, point P1 has the largest outlier value, the calculated attention value is the lowest, and the degree of abnormality is the highest. This shows the obvious superiority of the AE-LOF algorithm in detecting outliers located at the edge of the cluster. At the same time, when k = 2, the k-CAD value is 355.44, which is lower than the threshold of 1000 in Judgment 3. Obviously, continuing to calculate the k-CAD value is redundant, which also reflects the scientificity of taking the initial CAD value as the priority judgment condition in the classification system of the present invention. In summary, this signal can be judged as a mixture of a single-component signal and a foreign particle signal - Mixed Signal I.

[0186] Figure 17 What is shown is a scatter plot of the mixture of multi-component signals and foreign particle signals. From Figure 17 (a) It can be seen that the initial CAD value is 1168.80, which is higher than the threshold of 1000 in Judgment 2, and it is necessary to use the k-CAD algorithm for in-depth calculation. From Figure 4(b) As can be seen, when k = 2, the k-CAD value is 522.25, which is lower than the threshold of 1000 in identification 3. Therefore, the AE-LOF algorithm can detect multiple outliers. In summary, this signal can be identified as a mixture of multi-component signals and excess particle signals—mixed signal II. It should be noted that achieving the optimal clustering result for multi-component signals may require multiple iterative calculations, as different k values yield different k-CAD values. If several outliers are located close together but are collectively far from the main data cluster of the component signals, the k-means algorithm may mistakenly identify these outliers as a single data cluster, and the k-CAD value may still be greater than 1000. Therefore, during the iterative calculation process, we set the rule that if the number of data points in a cluster is less than 15, it will not be considered a qualified cluster. According to the classification system process, this signal type may be identified as requiring manual identification. This situation is unique, and the identification result is acceptable.

[0187] Figure 18 Displayed is a scatter plot of a multi-component signal. Figure 18 (a) It can be seen that at that time, the initial CAD value was 1303.20, which was higher than the threshold of 1000 in the judgment 2. Figure 18 (b) As can be seen, when k = 2, the k-CAD value is 53.25. Clearly, for multi-component signal types, the k-CAD values corresponding to different k values vary significantly, often by more than a factor of 10. Generally speaking, the vibration frequency of component signals is relatively stable, with occasional shifts in the frequency center point, but the overall absolute value of the deviation is less than 1000. Otherwise, if the component vibrates too violently and the displacement is too flexible, it will be significantly detrimental to product quality and fail to meet product reliability design requirements.

[0188] In summary, we have explained the identification process of the five signal categories one by one, which will help readers understand the classification rules and design logic of the classification system. It is worth noting that the systematic strategy of this invention is also valuable for the identification of non-periodic signals and periodic signals containing abnormal signals.

[0189] To verify the accuracy of the classification system of the present invention, we randomly selected 800 sets of signal sample data from the research group's database as the identification targets according to the sampling rule. These signal sample data are as follows: (1) 200 sets of known single-component signal samples; (2) 200 sets of known mixed signal I samples; (3) 200 sets of known redundant signal samples; (4) 100 sets of known mixed signal II samples; (5) 50 sets of multi-component signal samples; (6) 50 sets of known oversized component signal samples. The identification statistics of these signal data are shown in Table 3.

[0190] Table 3 Recognition Results and Comparative Statistics Based on the Classification System

[0191]

[0192] Note:(1) * :The number of signal samples of mixed signal II is small because the actual occurrence probability is also small. (2)**:The number of samples of multi-component signals is small because the actual occurrence probability is also small. (3)***:The pulse duty cycle threshold D used in this test is 65%, from long-term data statistics. (4)△:The main parameters of data statistics are accuracy rate and improvement ratio. Since the random rule is used as the selection method for test sample data, the differences and contingencies brought by different signal categories and different numbers of test samples are not considered for the time being. (5)△△:Since the old algorithm cannot reclassify the component signal types, the recognition accuracy rate of the overall component signal is used instead.

[0193] It can be seen from the table that 193 groups of single-component signals are correctly recognized, with an accuracy rate of 96.50%; 190 groups of mixed signal I are correctly recognized, with an accuracy rate of 95.00%; 192 groups of foreign object signals are correctly recognized, with an accuracy rate of 96.00%; 86 groups of mixed signal II are correctly recognized, with an accuracy rate of 93.00%. 46 groups of multi-component signals are correctly recognized, with an accuracy rate of 92.00%; 49 groups of oversized component signals are correctly recognized, with an accuracy rate of 98%.

[0194] It should be additionally noted that this recognition accuracy rate is the result of the mixed recognition of signals of 6 types as the classification objects. Previous other studies were not based on the results of the mixed recognition of 6 types of signals. When the prior art is based on the mixed recognition of 6 types of signals, the disclosed accuracy rate will inevitably be greatly affected. The fact that the present invention obtains the above recognition results in the mixed recognition of 6 types of signals is sufficient to illustrate the effect of the present invention.

[0195] To reflect the continuity and improvement effect in the research of the classification system, a comparison is made with CN117009858B (the previous research method of the inventor's team). It is worth noting that the average accuracy rate of the strategy proposed in the present invention is 95.08%, while the average accuracy rate of the old method is only 89.10%, and the average accuracy rate is increased by about 6.71%. In addition, to more clearly show the comparison results between the two methods, we separately plot line charts of the strategy of the present invention and the previous research results according to signal categories. From Figure 19(As can be seen from the abscissa in the figure, where 1 is the single-component signal, 2 is the mixed signal I, 3 is the foreign object signal, 4 is the mixed signal II, 5 is the multi-component signal, and 6 is the oversized component signal), compared with the previous method, the strategy of the present invention has improved the recognition accuracy for all 6 types of signal sample data. Among them, the recognition accuracy for the mixed signal I has the highest improvement, reaching 13.42%. Specific Embodiment 2:

[0197] This embodiment is a classification system for foreign object detection signals of aerospace hermetic relays, which is actually a classification system composed of program modules corresponding to a classification method for foreign object detection signals of aerospace hermetic relays, including:

[0198] Pulse extraction module: Extract pulses from the PIND signals collected in the foreign object detection of aerospace relays;

[0199] Pulse segmentation module: Frame the extracted pulse signals;

[0200] Signal folding and coordinate projection module: Fold the framed signals; then use the coordinates of the peak points of each pulse in the frame where the pulse is located to represent the pulse, and display the peak points of the pulses on all framed segments on the x-y coordinate axes. The x-axis is the length direction of each frame of the signal, and the y-axis is the amplitude direction of each frame of the signal;

[0201] Classification module for detection signals: Based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, obtain multiple data points in the x-y coordinate system. For the multiple data points, use k-CAD for clustering, and at the same time evaluate the degree of abnormality of the outlier points based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes. Then, based on the clustering results and the degree of abnormality, evaluate, and then use the classification system to classify the detection signals;

[0202] The classification module for detection signals includes a clustering result calculation unit, an outlier abnormality degree evaluation unit, and a classification unit; among them,

[0203] Clustering result calculation unit: Based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, obtain multiple data points in the x-y coordinate system. Calculate the CAD values of all points for the multiple data points, and at the same time, for the multiple data points, use k-CAD for clustering based on the optimal number of clusters k, and calculate the k-CAD value;

[0204] Outlier abnormality degree evaluation unit: Evaluate the degree of abnormality of the outlier points based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes;

[0205] Classification and Identification Unit: Classify the detection signal according to the results of the Clustering Result Calculation Unit and the Outlier Anomaly Degree Evaluation Unit. Specific Embodiment 3:

[0207] This embodiment is a foreign object detection signal clustering method for classifying foreign object detection signals of aerospace hermetic relays, including the following steps:

[0208] Extract pulses from the PIND signals collected in the foreign object detection of aerospace relays, and frame the extracted pulse signals; then take each framed signal as a data point sample x i , convert each framed signal into an m-dimensional vector [x i1 , x i2 , …, x im , and use [x i1 , x i2 , …, x im as the coordinates of the data point x i ;

[0209] In this embodiment, each framed signal is taken as a data point sample x i , and when converting each framed signal into an m-dimensional vector [x i1 , x i2 , …, x im , the peak points of the pulses on all framed segments are displayed on the x-y coordinate axes, the x-axis is the length direction of each frame signal, and the y-axis is the amplitude direction of each frame signal; based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, multiple data points in the x-y coordinate system are obtained.

[0210] Denote the set of data points corresponding to the peak points of the pulses corresponding to all segmented segments on the x-y coordinate axes as the data set C:

[0211] C = {x i | i = 1, 2, …, n}, x i = [x i1 , x i2 , …, x im

[0212] Among them, x i represents the data point sample, and [x i1 , x i2 , …, x im is the coordinate of x i , and n is the number of data point samples;

[0213] Based on the optimal number of clusters k, the k-means clustering algorithm is used to cluster the dataset C, and the class label L = [l1, l2, …, l n is obtained, where l i represents the class of the signal point x i , and l i ∈{1, 2, …, k}; for each class j, j = 1, 2, …, k, all the sample points belonging to this class are extracted and denoted as the set C j ; the formula for calculating the center μ j of class j is:

[0214]

[0215] where |C j | is the number of sample points in class j;

[0216] The formula for class j based on the k-CAD algorithm is:

[0217] k-CAD(C j ) is used for the classification of debris detection signals of aerospace hermetic relays.

[0218] The above examples of the present invention are only for explaining in detail the calculation model and calculation process of the present invention, rather than limiting the implementation manner of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is impossible to list all the implementation manners here. Any obvious changes or modifications derived from the technical solutions of the present invention still fall within the protection scope of the present invention.

Claims

1. A method for classifying the signals of foreign matters detected in a space sealing relay, characterized in that It includes the following steps: Extract pulses from the PIND signals collected in the detection of debris in aerospace relays, and frame the extracted pulse signals; then fold the framed signals; Use the coordinates of the peak point of each pulse in the frame where the pulse is located to represent the pulse, and display the peak points of the pulses on all framed segments on the x-y coordinate axes, with the x-axis being the length direction of each frame of signal and the y-axis being the amplitude direction of each frame of signal; Based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, obtain multiple data points in the x-y coordinate system. For the multiple data points, use k-CAD for clustering, and at the same time evaluate the degree of abnormality of the outlier points based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, and then classify the detection signals based on the clustering results and the degree of abnormality; The process of using k-CAD for clustering for the multiple data points in the x-y coordinate system includes the following steps: Denote the set of data points corresponding to the peak points of the pulses corresponding to all segmented segments on the x-y coordinate axes as dataset C: C = {x i | i = 1, 2, …, n}, x i = [x i1 , x i2 , …, x im ​ Among them, x i represents a data point sample, [x i1 , x i2 , …, x im are the coordinates of x i , and n is the number of data point samples; Based on the optimal number of clusters k, the dataset C is clustered using the k-means clustering algorithm to obtain the class label L = [l1, l2, …, l n , where l i represents the class of the signal point x i , and l i ∈{1, 2, …, k}; for each class j, j = 1, 2, …, k, all the sample points belonging to this class are extracted and denoted as the set C j ; the formula for calculating the center μ j of class j is: where |C j | is the number of samples in class j; The calculation formula for category j based on the k-CAD algorithm is as follows:

2. A method for classifying redundant object detection signals of an aerospace hermetic relay according to claim 1, characterized in that, During the process of framing the extracted pulse signals, use the vibration period time length of the vibration table as the length of the framing window for framing.

3. A method for classifying the redundant object detection signals of a spaceborne sealed relay according to claim 1, characterized in that, Before framing the pulse signals, merge adjacent pulses with a time interval less than the set threshold into one pulse.

4. A method for classifying the signals of foreign matter detection of an aerospace sealed relay according to claim 1, characterized in that The optimal number of clusters k is determined according to the following steps: Use the k-means clustering algorithm to divide dataset C into k categories, calculate the SSE under different k values, and draw a relationship graph of the k value and the sum of squared errors within the clusters; take the k value at the inflection point position as the finally determined optimal number of clusters.

5. A method for classifying redundant object detection signals of an aerospace hermetic relay according to claim 4, characterized in that, During the process of evaluating the degree of abnormality of the outlier points based on the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, use the AE-LOF algorithm for the evaluation of the degree of abnormality. The specific process includes: For the peak point coordinates of the pulses corresponding to all framed segments on the x-y coordinate axes, first obtain the associated point set as A i ; For each data point p i , the associated point set A i is as follows: A i = {p j ∈ C \ {p i} | d(p i , p j ) ≤ R i} This set contains all other data points within the radius R of p i ; that is, d i where (·) represents the k-nearest neighbor distance; k (·) represents the k-nearest neighbor distance; For each point p in the dataset j , initialize a frequency counter F j = 0; For each data point p i , check each other point p j to see if it is in A i . If it is, then F j = F j + 1; After traversing all points, each F j will represent the total number of times point p j appears within the radius of all other points as an associated point; Then calculate AE-LOF(p j ): Among them, N k (·) represents the set of neighbor points, and |N k (P)| represents the set capacity; LRD(·) represents the reciprocal of the average reachable distance from a point to its surrounding neighbor points; Z is the internal adjustment coefficient, and w = Z × k is the external adjustment coefficient; Using AE-LOF(p j ) to evaluate the degree of abnormality of the outliers at the peak points of the pulses corresponding to all segmented segments on the x-y coordinate axes.

6. A method for classifying redundant object detection signals of an aerospace hermetic relay according to claim 5, characterized in that, The process of classifying the detection signals based on the clustering results and the degree of abnormality includes the following steps: For the multiple data points in the x-y coordinate system, calculate the pulse duty cycle D of each point corresponding to the pulse. When D≥the duty cycle threshold, it is judged as an over-sized component signal; otherwise, continue to judge; Then calculate for all points where |C| represents the capacity of the dataset C corresponding to all points, and μ j represents the mean of the dataset C; when CAD(xi) ≤ clustering threshold and the detection result of the AE-LOF algorithm has no outliers, the identification result is a single-component signal; when CAD(xi) ≤ clustering threshold and the detection result of the AE-LOF algorithm has outliers, the identification result is a mixture of a single-component signal and debris particle signals; otherwise, continue the identification; Record the calculation result of the k-CAD algorithm corresponding to the optimal number of clusters k as CAD k (xi). If CAD k (xi) ≤ the clustering threshold and the detection result of the AE-LOF algorithm has no outliers, the identification result is a multi-component signal; if CAD k (xi) ≤ the clustering threshold and the detection result of the AE-LOF algorithm has outliers, the identification result is a mixture of multi-component signal and foreign particle signal; if CAD k (xi) > the clustering threshold and the detection result of the AE-LOF algorithm has outliers, the identification result is a foreign particle signal.

7. A foreign object detection signal classification system for aerospace hermetic relays, characterized in that, It includes: Pulse extraction module: Extract pulses from the PIND signals collected in the detection of debris in aerospace relays; Pulse segmentation module: Frame the extracted pulse signals; Signal folding and coordinate projection module: Fold the framed signals; then use the coordinates of the peak point of each pulse in the frame where the pulse is located to represent the pulse, and display the peak points of the pulses on all framed segments on the x-y coordinate axes, with the x-axis being the length direction of each frame of signal and the y-axis being the amplitude direction of each frame of signal; Classification module for detection signals: Based on the peak point coordinates of the pulses corresponding to all segmented frames on the x-y coordinate axes, multiple data points in the x-y coordinate system are obtained. For the multiple data points, k-CAD is used for clustering. At the same time, the degree of abnormality of the outliers is evaluated based on the peak point coordinates of the pulses corresponding to all segmented frames on the x-y coordinate axes, and then the classification of the detection signals is carried out based on the clustering results and the evaluation of the degree of abnormality; The process of using k-CAD for clustering for multiple data points in the x-y coordinate system includes the following steps: Denote the set of data points corresponding to the peak points of the pulses corresponding to all segmented segments on the x-y coordinate axes as the data set C: C = {x i | i = 1, 2, …, n}, x i = [x i1 , x i2 , …, x im ​ Among them, x i represents a data point sample, [x i1 , x i2 , …, x im are the coordinates of x i , and n is the number of data point samples; Based on the optimal number of clusters k, the k-means clustering algorithm is used to cluster the dataset C, and the class label L = [l1, l2, …, l n is obtained, where l i represents the class of the signal point x i , and l i ∈{1, 2, …, k}; for each class j, j = 1, 2, …, k, all the sample points belonging to this class are extracted and denoted as the set C j ; the calculation formula for the center μ j of class j is: where, |C j | is the number of samples in class j; The calculation formula for category j based on the k-CAD algorithm is as follows:

8. A debris detection signal clustering method for classifying debris detection signals of aerospace hermetic relays, characterized in that, It includes the following steps: Extract pulses from the PIND signals collected in the detection of aerospace relay contaminants, and frame the extracted pulse signals; then take each framed signal as a data point sample x i , and convert each framed signal into an m-dimensional vector [x i1 , x i2 , …, x im . Use [x i1 , x i2 , …, x im as the coordinates of the data point x i . Denote the set of all data points as the data set C: C = {x i | i = 1, 2, …, n}, x i = [x i1 , x i2 , …, x im ​ Among them, x i represents a data point sample, [x i1 , x i2 , …, x im are the coordinates of x i , and n is the number of data point samples; Based on the optimal number of clusters k, the dataset C is clustered using the k-means clustering algorithm to obtain the class labels L = [l1, l2, …, l n , where l i represents the class of the signal point x i , and l i ∈{1, 2, …, k}; for each class j, j = 1, 2, …, k, all the sample points belonging to this class are extracted and denoted as the set C j ; the calculation formula for the center μ j of class j is: where, |C j | is the number of samples of class j; The calculation formula for category j based on the k-CAD algorithm is as follows: k-CAD (C j ) is used for the classification of foreign object detection signals of aerospace hermetic relays.

Citation Information

Patent Citations

  • A synchronous classification method for redundant detection signals of aerospace sealed electronic components

    CN117009858B

  • Based on improved fast density peak clustering and LOF outlier detection algorithm

    CN109102028A

  • Power equipment discharge signal separation and classification method based on kernel principal component analysis

    CN111444784A