A Fault Detection Method and System for LPCAMM2

By using the fault detection method of LPCAMM2 based on curve expansion, the construction of a generative adversarial network model, combined with the single-step random mapping gradient algorithm, the problem of low fault detection accuracy in the existing technology is solved, and more efficient and accurate fault detection is achieved.

CN120045378BActive Publication Date: 2025-06-27SHENZHEN COMOS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510519452.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-06-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing LPCAMM2 fault detection method is based on static thresholds, ignoring the coupling relationship between different state parameters, resulting in extremely low detection accuracy and insufficient practicality.

Method used

By obtaining multiple sets of normal state data of each state parameter of LPCAMM2, data augmentation based on curve expansion is carried out, a generative adversarial network model is constructed, combined with a single-step random mapping gradient algorithm, the optimal hidden variable is searched, the abnormal score of real-time state data is calculated, and the contribution value of each state parameter to the abnormal score is analyzed to complete fault detection.

Benefits of technology

It effectively expands the distribution range of normal state data, enhances the learning ability of the generative adversarial network model for complex failure modes, improves the speed and accuracy of fault detection, and can effectively deal with multi-parameter coupling exceptions in complex systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045378B_ABST
    Figure CN120045378B_ABST
Patent Text Reader

Abstract

The present invention provides a fault detection method and system for LPCAMM2, which relates to the technical field of memory testing. The method includes: obtaining multiple groups of normal state data of each state parameter of LPCAMM2; performing data augmentation based on curve expansion on the normal state data; constructing a generator and a discriminator of a generative adversarial network model; obtaining real-time state data; searching for the optimal latent variable; calculating the anomaly score of the real-time state data; calculating the contribution value of the real-time state data corresponding to each state parameter to the anomaly score, and determining the state parameter corresponding to the maximum contribution value as the abnormal state parameter; outputting the abnormal category based on the abnormal state parameter to complete the fault detection. While ensuring the fast response ability of fault detection, it effectively improves the speed and accuracy of fault detection, can effectively cope with multi-parameter coupling anomalies in complex systems, and has high practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of memory testing, and particularly to a method and system for fault detection of LPCAMM2. Background Art

[0002] LPCAMM2 (Low Power Compressed Add-on Memory Module) is a low-power add-on memory module based on LPDDR5X technology. It adopts a brand-new design architecture. Compared with traditional SODIMM memory modules, it can provide more efficient packaging and reduce space occupancy by 60%. LPCAMM2 supports higher data transfer rates and can significantly reduce power consumption, and has wide applications in high-performance computing, gaming workstations, data centers and other applications. Its detachable design enables manufacturers to more flexibly configure and upgrade memory modules in devices.

[0003] As a high-performance memory module, LPCAMM2 undertakes important computing tasks and data transfer functions. Any hardware failure will directly affect the stability and performance of the device. Fault detection can identify potential problems in advance, reduce the risks of system crashes, data loss or performance degradation, and ensure the efficient operation of the device and user experience. Timely detection and repair of LPCAMM2 faults are crucial for ensuring the stability of critical applications (such as high-performance computing, gaming, etc.).

[0004] However, the existing fault detection of LPCAMM2 often detects single faults based on static thresholds. Although the detection efficiency is high, due to ignoring the coupling relationship between different state parameters, the detection accuracy is extremely low and the practicality is insufficient. Summary of the Invention

[0005] In view of the above deficiencies of the prior art, the purpose of the embodiments of the present invention is to provide a method and system for fault detection of LPCAMM2, which can solve the technical problem that the existing fault detection of LPCAMM2 often detects single faults based on static thresholds. Although the detection efficiency is high, due to ignoring the coupling relationship between different state parameters, the detection accuracy is extremely low and the practicality is insufficient.

[0006] In the first aspect of the embodiments of the present invention, a method for fault detection of LPCAMM2 is proposed, including:

[0007] S1: Obtain multiple groups of normal state data of each state parameter of the LPCAMM2;

[0008] S2: Perform data augmentation based on curve extension on the normal state data with the calibration range of the state data of the LPCAMM2 as a constraint;

[0009] S3: Construct the generator and discriminator of the generative adversarial network model based on the enhanced normal state data;

[0010] S4: Obtain the real-time state data;

[0011] S5: Combine the single-step stochastic mapping gradient algorithm to search for the optimal latent variable in the input data of the generator, i.e., the latent space, that minimizes the recovery error of the real-time state data;

[0012] S6: Calculate the anomaly score of the real-time state data according to the optimal latent variable;

[0013] S7: In the case where the anomaly score is greater than the preset anomaly score, calculate the contribution value of the real-time state data corresponding to each state parameter to the anomaly score, and determine the state parameter corresponding to the maximum contribution value as the abnormal state parameter;

[0014] S8: Output the anomaly category based on the abnormal state parameter to complete the fault detection of the LPCAMM2.

[0015] In the second aspect of the embodiments of the present invention, a fault detection system for LPCAMM2 is proposed, including: a processor and a memory;

[0016] The memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, it implements the steps of the fault detection method for LPCAMM2 as described in the first aspect.

[0017] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention at least include:

[0018] In the embodiments of the present invention, through data enhancement based on curve expansion for normal state data including multiple state parameters, and then using the generative adversarial network model to learn the distribution pattern and coupling relationship of normal state data. Through the data enhancement method based on curve expansion, the distribution range of normal state data is effectively expanded, and the learning ability of the generative adversarial network model for complex fault patterns is enhanced. Then, combined with the single-step stochastic mapping gradient algorithm, the optimal latent variable that minimizes the recovery error of the real-time state data is quickly searched in the input data of the generator, i.e., the latent space, which can quickly infer the optimal latent variable and improve the fault detection speed. And based on this optimal dependent variable, the anomaly score of the real-time state data is calculated, and the contribution values of different categories of state parameters to the anomaly score are calculated respectively to complete the precise positioning of the anomaly category, which can fully learn and capture the coupling relationship between high-dimensional data, effectively improve the speed and accuracy of fault detection while ensuring the fast response ability of fault detection, and can effectively handle multi-parameter coupling anomalies in complex systems, with high practicality. Brief Description of the Drawings

[0019] The drawings are only for the purpose of illustrating specific embodiments and are not considered as limitations of the present invention. Throughout the drawings, the same reference numerals denote the same components. Obviously, the drawings in the following description are only some embodiments described in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 is a schematic flowchart of a fault detection method for LPCAMM2 provided by an embodiment of the present invention;

[0021] Figure 2 is a schematic structural diagram of a fault detection system for LPCAMM2 provided by an embodiment of the present invention. Detailed Embodiments

[0022] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. It should be understood that these descriptions are only exemplary and are not used to limit the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] The following will describe in detail the fault detection method for LPCAMM2 provided by the embodiments of the present invention with reference to the drawings through specific embodiments and their application scenarios.

[0024] Refer to the attached drawings of the specification Figure 1 , which shows a schematic flowchart of a fault detection method for LPCAMM2 provided by an embodiment of the present invention.

[0025] The embodiments of the present invention provide a fault detection method for LPCAMM2, which may include the following steps:

[0026] S1: Obtain multiple groups of normal state data of each state parameter of LPCAMM2.

[0027] Among them, the normal state data refers to the key state parameter data recorded by LPCAMM2 (Low Power Compression Additional Memory Module) under normal working conditions. These data represent the typical behaviors of LPCAMM2 without faults or abnormalities, such as power consumption, temperature, clock frequency, bit error rate, read rate, write rate, etc. The normal state data is used to establish a benchmark model to help the subsequent fault detection system determine whether the real-time data exceeds the normal range.

[0028] Specifically, the normal state data can be obtained by monitoring its key parameters in real time under different workloads during the operation of LPCAMM2. The specific steps include: First, ensure that LPCAMM2 runs in a stable and normal operating environment. Then, use hardware monitoring tools or embedded sensors to collect data such as power consumption, temperature, and frequency. Finally, store multiple sets of these data to ensure coverage of different operating conditions and environments as a benchmark for the subsequent fault detection system.

[0029] In a possible implementation, the state parameters include power consumption, temperature, clock frequency, clock jitter, bit error rate, read rate, write rate, and operating voltage.

[0030] S2: Constrained by the calibration range of the state data of LPCAMM2, perform data augmentation on the normal state data based on curve expansion.

[0031] Among them, the calibration range of the state data refers to the allowable value range of each state parameter (such as power consumption, temperature, clock frequency, etc.) when performing fault detection on LPCAMM2. This range is usually based on the design specifications of LPCAMM2 or the normal operating range obtained through experiments, and is used to limit and constrain the changes in the data. Through the calibration range, it can be ensured that the results of data augmentation do not exceed reasonable physical or operating limits. Data augmentation based on curve expansion is a mathematical method for expanding or changing the normal state data. Specifically, by expanding the normal state data based on a curve, the system can generate new data points, expand the range of the data set, and enhance the diversity of the data. This method can help the model better learn different change patterns of the state data, increase the diversity of the data set, and improve the robustness and generalization ability of the model.

[0032] It should be noted that by performing data augmentation based on curve expansion on the normal state data of LPCAMM2, more diverse samples can be generated within the calibration range of the state data to simulate different operating conditions. This method not only expands the scale of the data set but also enables the detection system to better learn various potential changes in the operating state, improving the accuracy and adaptability of fault detection. By controlling the curve expansion in the data augmentation process, it is ensured that the data changes conform to the actual situation, thereby avoiding the generation of unreasonable abnormal data.

[0033] In a possible implementation, the S2 specifically includes:

[0034] Fit the normal state data belonging to the same state parameter category in each group of normal state data to obtain a smooth function representing a smooth curve, where each smooth curve forms a high-dimensional space curve.

[0035] The generation formula of the high-dimensional space curve is specifically:

[0036] ;

[0037] ;

[0038] Among them, represents the path parameter, represents the smoothing function of the normal state data of the i-th type of state parameter under the path parameter ; m represents the number of categories of state parameters, represents the high-dimensional space curve related to including each group of normal state data, ; represents the j-th normal state data of the i-th type of state parameter given, ; n represents the number of groups of normal state data, represents B-spline interpolation.

[0039] In this process, first, the normal state data belonging to the same category of state parameters are fitted to obtain the smoothing function of each category. Using the B-spline interpolation method, these data can be smoothed to obtain a smooth curve. B-spline interpolation is an efficient smoothing method that can effectively reduce the noise in the data and accurately capture the trend of the normal state data. Doing so can not only avoid the influence of abnormal data on model training but also enable the model to better learn the internal laws and change trends of the data.

[0040] Calculate the derivative of each smoothing function with respect to the path parameter.

[0041] In this process, by taking the derivative of each smoothing function, the change rate of each state parameter under the path parameter is obtained. This step is to calculate the sensitivity of the state parameter to the change of the path. Calculating the derivative can help understand the rate of change of the state parameter with respect to the path parameter. This can capture the subtle changes between the data and contribute to more accurate modeling of the state data changes in the high-dimensional space in the subsequent process.

[0042] Calculate the coupled derivative between two derivatives.

[0043] The specific calculation formula of the coupled derivative is:

[0044] ;

[0045] Among them, represents the smoothing function of the normal state data of the v-th type of state parameter under the path parameter ; represents the coupled derivative between the i-th type of state parameter and the v-th type of state parameter.​​

[0046] In this process, for different state parameters (such as power consumption, temperature, clock frequency, etc.), the coupled derivatives between them are calculated. The coupled derivative represents the variation relationship between two state parameters. By comparing the derivatives of different state parameters, their mutual influence can be revealed. Furthermore, the correlation and interaction between multiple state parameters can be revealed. The multiple state parameters applicable to LPCAMM2 are in a mutually coupled working state. Understanding these coupling relationships helps improve the detection ability for complex fault modes.

[0047] Calculate the tangent vector of the high-dimensional space curve.

[0048] The specific calculation formula for the tangent vector is:

[0049] .

[0050] In this process, by calculating the tangent vector of the high-dimensional space curve, the change trend of each state parameter is understood. This step helps the model understand the multi-dimensional characteristics of the data. The tangent vector can describe the change direction of the data in the multi-dimensional space. Through this direction, the model can learn how to make trade-offs between various state parameters, thereby improving the understanding ability of complex data relationships.

[0051] Combine the state data calibration range and the coupled derivative to set the perturbation parameters for data augmentation.

[0052] The specific calculation formula for the perturbation parameters is:

[0053] ;

[0054] where, represents the perturbation parameter, and respectively represent the upper limit value and the lower limit value of the state data calibration range of the i-th type of state parameter.

[0055] In this process, the perturbation parameter is calculated based on the product of the calibration range of the state data (i.e., the allowable maximum and minimum values of each state parameter) and the coupled derivative. By combining the actual range of the state data, the perturbation parameter can ensure that the generated data does not exceed the physical or operational limits, avoiding the generation of unrealistic abnormal data. At the same time, the perturbation parameter can effectively expand the diversity of the data and improve the quality of the generated samples.

[0056] Combine the state data calibration range, use the tangent vector direction as the perturbation direction, and add the perturbation parameter on the high-dimensional space curve to obtain the enhanced normal state data.

[0057] The enhanced normal state data is specifically:

[0058] ;

[0059] ;

[0060] Among them, represents the smoothing function of the enhanced normal state data under the path parameter , and represents the state data of the i-th type of state parameter.

[0061] It should be noted that the calculated perturbation parameters are applied to the high-dimensional space curve, and perturbations are increased along the direction of the tangent vector to generate new data points. These data points represent the diverse situations that may occur under normal conditions. By generating enhanced data highly relevant to the actual normal working state, the model can better learn the potential laws of normal data. Since the enhanced data covers more possible situations, the fault detection system can make accurate judgments within a wider range, further improving the accuracy and efficiency of fault detection.

[0062] S3: Construct the generator and discriminator of the generative adversarial network model according to the enhanced normal state data.

[0063] Among them, the generative adversarial network (GAN) model is a deep learning model composed of two neural networks: the generator and the discriminator. The goal of the generator is to generate fake data as similar as possible to the real data, while the discriminator is responsible for distinguishing the generated data from the real data. The generator and the discriminator compete with each other during the training process. Through continuous optimization, the generator can finally generate high-quality samples, and the discriminator can more accurately distinguish between real data and generated data. Using normal data to train the GAN, the generator G learns the distribution law of the normal state data. In this way, the generator can inversely capture the coupling relationship of LPCAMM2 under normal conditions, providing an accurate benchmark for subsequent detection of abnormal data. This adversarial training method enables the model not only to generate similar normal data but also to better identify abnormal data with different distributions from these data.

[0064] In a possible implementation manner, the S3 specifically includes:

[0065] Input the enhanced normal state data as the training set into the generative adversarial network model.

[0066] Adopt the method of alternately fixing the discriminator and the generator, and train the unfixed discriminator and generator respectively to complete the construction of the generative adversarial network model.

[0067] It should be noted that during the training process of the Generative Adversarial Network (GAN), the common practice is to first fix the discriminator and then train the generator, and then fix the generator and train the discriminator. This is because the task of the discriminator is to distinguish between real and generated data, while the task of the generator is to generate as realistic data as possible. The training sequence is as follows: Fix the discriminator: Keep the discriminator unchanged and train the generator to generate more realistic data. Fix the generator: Keep the generator unchanged and train the discriminator to distinguish between real data and generated data. This alternating training method can help the generator continuously improve to generate more realistic samples, and at the same time enable the discriminator to more accurately distinguish between true and false data.

[0068] S4: Obtain real-time status data.

[0069] Among them, the real-time status data refers to various dynamic status information collected at the current moment during the actual operation of LPCAMM2.

[0070] S5: Combine the single-step stochastic mapping gradient algorithm to search for the optimal latent variable in the input data of the generator, i.e., the latent space, that minimizes the recovery error of the real-time status data.

[0071] Among them, the single-step stochastic mapping gradient algorithm is an optimization algorithm that combines stochastic mapping and gradient descent methods. By gradually adjusting the latent variables in the latent space, it seeks the optimal solution that can minimize a certain objective function (such as the recovery error). In each step of the optimization process, the algorithm randomly selects the initial latent variable and gradually approaches the optimal solution based on the gradient of the objective function, thereby improving the search efficiency. The recovery error refers to the difference between the samples generated from the latent space by the generator and the real-time status data. It reflects the quality of the generated samples. If the recovery error is large, it indicates a large deviation between the generated samples and the real data, indicating that the effect of the generator is not ideal. Minimizing the recovery error is a key objective in the optimization process. The optimal latent variable refers to the latent variable in the latent space that, after being optimized by the algorithm, can make the generated samples best match the actual status data. The latent variable is the input in the generative adversarial network, which determines the characteristics of the generated samples. By optimizing the optimal latent variable, the generator can more accurately generate samples that conform to the real data distribution, thereby improving the accuracy of fault detection.

[0072] It should be noted that by combining the single-step stochastic mapping gradient algorithm, the optimal latent variable can be efficiently searched in the latent space to minimize the recovery error of the real-time status data. In this way, the system can accelerate the fault detection process and reduce the number of iterations. The search for the optimal latent variable ensures that the difference between the generated data and the actual data is minimized, thereby improving the accuracy and efficiency of the model.

[0073] In a possible implementation, the recovery error includes the generator reconstruction error and the discriminator recognition error.

[0074] It should be noted that the recovery error consists of the generator reconstruction error and the discriminator recognition error. The generator reconstruction error represents the deviation between the data reconstructed by the generator when attempting to reconstruct real-time state data and the real data. The discriminator recognition error represents the error of the discriminator in distinguishing between generated data and real data. Combining these two errors can more comprehensively measure the quality of the generated data and optimize the GAN model to improve the accuracy and stability of fault detection.

[0075] In a possible implementation, the search process for the optimal latent variable specifically includes:

[0076] Randomly select a base latent variable from the latent space, and calculate the generated sample representing the state data through the generator.

[0077] Based on the base latent variable, calculate the discrimination probabilities of the generated sample and the real-time state data through the discriminator respectively:

[0078] Calculate the square of the L1 norm between the real-time state data and the generated sample to obtain the generator reconstruction error.

[0079] Calculate the square of the L1 norm between the discrimination probabilities of the generated sample and the real-time state data to obtain the discriminator recognition error.

[0080] Calculate the first weighted sum value between the generator reconstruction error and the discriminator recognition error.

[0081] In the latent space, optimize the base latent variable by combining the single-step random mapping gradient algorithm to obtain the optimized latent variable.

[0082] In a possible implementation, in the latent space, optimizing the base latent variable by combining the single-step random mapping gradient algorithm to obtain the optimized latent variable specifically includes:

[0083] Project the latent space onto a two-dimensional random subspace to obtain a low-dimensional latent variable.

[0084] The mapping process is specifically:

[0085] ;

[0086] where P represents a randomly generated orthogonal matrix, represents the square of the L1 norm, represents the real-time state data, represents the latent variable z when the function takes the minimum value, represents the real number field, d represents the dimension of the latent space, represents based on the latent variable The generated generated samples represent low-dimensional latent variables in a two-dimensional random subspace.

[0087] Among them, represents the two-dimensional real number space, represents a d×2 matrix. That is, P is a matrix of size d×2, which is used to map the original high-dimensional latent space data (dimension d) to a two-dimensional space.

[0088] Based on the local linear assumption of the generator, the low-dimensional latent variables are optimized once.

[0089] The specific formula for the first optimization is:[[]]

[0090] ;

[0091] Among them, represents the fixed step size, represents based on the low-dimensional latent variables The generated generated samples represents the Jacobian matrix of the generator G at the low-dimensional latent variables at that point, represents the discriminator D's gradient with respect to z of the generated sample represents the low-dimensional latent variables after the first optimization, represents the first weight factor.

[0092] Optionally, the fixed step size can be set to 0.1. The first weight factor can be set to 0.5.

[0093] The low-dimensional latent variables are optimized twice through the discriminator residual to obtain the optimized latent variables.

[0094] The specific formula for the second optimization is:[[]]

[0095] ;

[0096] Among them, represents the calibration coefficient, represents the discriminant probability of the real-time state data represents based on the latent variables after the first optimization The generated generated samples represents the optimized latent variables obtained by the second optimization, represents the discriminator's gradient with respect to z of

[0097] Optionally, the calibration coefficient can be set to 0.01. ​​​​

[0098] Specifically, first, the latent variables in the latent space are projected onto a two-dimensional random subspace. The purpose of this step is to reduce the dimension of the latent space for more efficient search of the optimal latent variables. Using a randomly generated orthogonal matrix P, the latent space is mapped to a two-dimensional subspace, and the base latent variable z0 is determined by minimizing the squared L1 norm between the generated samples and the real data. This ensures that the generated samples are as close as possible to the actual real-time state data. Based on the low-dimensional latent variable z0 in the low-dimensional subspace, the generator is optimized through the local linear assumption. By calculating the difference between the generated samples and the real-time state data, z0 is adjusted and updated to the latent variable z1. During one optimization process, the latent variable z1 is optimized by reducing the error between the generated samples and the real data. The optimization process involves the Jacobian matrix of the generator and the gradient of the discriminator, which can ensure that the generated samples are closer to the real data. The secondary optimization further reduces the difference between the generated samples and the real-time data by adjusting the output of the discriminator. By adjusting the gradient of the discriminator, the generated samples more precisely match the real data.

[0099] It can be understood that by reducing the dimension of the latent space to a two-dimensional subspace and combining with the single-step random mapping gradient algorithm, the dimension of the search space and the computational complexity are reduced. By gradually optimizing the latent variables, the number of iterations required is reduced, thus accelerating the determination speed of the optimal latent variables. This method is more efficient than traditional high-dimensional space optimization, enabling the generator to quickly generate samples that match the real-time data, thereby improving the speed of fault detection.

[0100] Update the base latent variable using the optimized latent variable, and return based on the base latent variable. Calculate the discrimination probabilities of the generated samples and the real-time state data through the discriminator respectively until the number of iterations is greater than the preset number of iterations.

[0101] It should be noted that those skilled in the art can set the size of the preset number of iterations according to actual needs, and the present invention does not limit this here.

[0102] Output the base latent variable that minimizes the first weighted sum value between the reconstruction error of the generator and the recognition error of the discriminator as the optimal latent variable.

[0103] Specifically, the search process for the optimal latent variable aims to find a latent variable that minimizes the error between the samples generated by the generator and the real-time state data. First, a base latent variable is randomly selected from the latent space, and the corresponding generated samples are calculated using the generator. Then, the discriminator is used to calculate the discrimination probability between the generated samples and the real-time state data, and the generator reconstruction error (the square of the L1 norm) and the discriminator recognition error (the square of the L1 norm between the discrimination probabilities of the generated samples and the real-time data) are calculated respectively. Next, these two errors are weighted and summed to measure the effect of the current latent variable. To find the optimal latent variable, the system optimizes the base latent variable in the latent space using the single-step stochastic mapping gradient algorithm to obtain the optimized latent variable, and continuously updates the base latent variable, repeating the above calculation process until the preset number of iterations is reached. Finally, the latent variable that minimizes the weighted error is selected as the optimal latent variable. Using the single-step stochastic mapping gradient algorithm can quickly optimize the latent variable, reduce the computational complexity, and speed up the fault detection speed. By combining the generator reconstruction error and the discriminator recognition error, it ensures the highest degree of matching between the generated samples and the real data, and improves the detection accuracy. By comprehensively using the generation ability and discrimination ability of the GAN, the fault detection not only depends on the data similarity, but also considers the recognition ability of the discriminator, thus reducing false alarms and missed alarms.

[0104] The calculation method of the optimal latent variable is specifically as follows:

[0105] ;

[0106] Wherein, represents the optimal latent variable, represents the base latent variable when the function takes the minimum value , represents the square of the L1 norm, represents the generated samples generated based on the base latent variable , represents the real-time state data of the discrimination probability, represents of the discrimination probability, represents the second weight factor.

[0107] Optionally, the second weight factor can be set to 0.5.

[0108] It should be noted that this method of calculating the optimal latent variable ensures the highest degree of matching between the generated samples and the real-time state data by minimizing the weighted sum of the generator reconstruction error and the discriminator recognition error, while considering the judgment ability of the discriminator. Using the square of the L1 norm to measure the error can more accurately measure the difference between the generated data and the real data, and further improve the quality of the generated samples. By introducing the second weight factor ( ), the influence weights of the generation error and the discrimination error can be flexibly adjusted to achieve better error balance, thereby improving the accuracy and reliability of the model.

[0109] S6: Calculate the anomaly score of the real-time status data according to the optimal latent variable.

[0110] Specifically, by inputting the optimal latent variable into the generative adversarial network (GAN) model, calculate the difference between the generated sample and the actual real-time status data, and evaluate whether the data is abnormal according to this difference (such as the reconstruction error). The level of the anomaly score indicates the deviation between the real-time status data and the normal data. The higher the score, the greater the degree of abnormality of the data, and vice versa, indicating that the data tends to be normal. The goal of this step is to provide a basis for fault detection by accurately calculating the anomaly score.

[0111] In a possible implementation manner, the S6 specifically includes:

[0112] Calculate the reconstruction error and the discriminator recognition error under the optimal latent variable respectively.

[0113] Calculate the second weighted sum value between the reconstruction error and the discriminator recognition error.

[0114] Output the second weighted sum value as the anomaly score.

[0115] The specific calculation formula of the anomaly score is:

[0116] ;

[0117] where, represents the anomaly score.

[0118] Among them, the reconstruction error represents the difference between the generated sample and the actual data, and the discriminator recognition error represents the discriminator's ability to distinguish between the generated sample and the real data. By calculating the weighted sum of these two errors, the final anomaly score is obtained. This score reflects the similarity between the generated sample and the real data. The higher the score, the more abnormal the data. By considering both the reconstruction error and the discriminator recognition error and comprehensively evaluating the quality of the generated data, the anomaly detection is made more accurate. The calculation of the weighted sum value can dynamically balance the influence of the generator and the discriminator, improving the robustness of the fault detection.

[0119] S7: When the anomaly score is greater than the preset anomaly score, calculate the contribution value of the real-time status data corresponding to each state parameter to the anomaly score, and determine the state parameter corresponding to the maximum contribution value as the abnormal state parameter.

[0120] Optionally, by analyzing the historical state data collected by LPCAMM2 during normal operation, the mean and standard deviation of each state parameter (such as power consumption, temperature, clock frequency, etc.) can be calculated. Then, according to statistical methods, the threshold for abnormal scoring is set. Usually, the mean plus a certain multiple (3 times or 5 times) of the standard deviation can be selected as the preset abnormal score. Or rely on the experience of engineers or the judgment of domain experts to set the threshold for abnormal scoring. According to the usage scenario of LPCAMM2 (such as high-performance computing, data center applications, etc.), experts will set appropriate scoring thresholds according to the performance requirements and tolerances of the system. The threshold for abnormal scoring can also be set according to the performance standards required by LPCAMM2. For some high-performance applications (such as gaming workstations or data centers), the threshold can be set to ensure that the system always operates in a high-performance state.

[0121] It should be noted that when the calculated abnormal score is greater than the preset abnormal score, the system will further calculate the contribution value of each state parameter (such as power consumption, temperature, clock frequency, etc.) to the abnormal score. This can help determine which specific parameters have the greatest impact on the abnormal score, thereby identifying the state parameters most likely to cause faults. In this way, the system can accurately locate the fault source and improve the accuracy and response speed of fault detection.

[0122] It should be noted that those skilled in the art can set the size of the preset abnormal score according to actual needs, and the present invention does not limit this here.

[0123] In a possible implementation manner, the S7 specifically includes:

[0124] Calculate the partial derivative of the abnormal score with respect to various state parameters in the real-time state data, and use the partial derivative as the contribution value.

[0125] The specific calculation method of the contribution value is:

[0126] ;

[0127] Wherein, represents the real-time state data corresponding to the i-th type of state parameter, represents taking the absolute value, represents taking the partial derivative, represents the partial derivative of the abnormal score with respect to the i-th type of state parameter in the real-time state data, that is, the contribution value.

[0128] Output the state parameter corresponding to the maximum contribution value as the abnormal state parameter.

[0129] The specific formula for determining the abnormal state parameter is:

[0130] ;

[0131] Among them, represents the state parameter index corresponding to the representative abnormal state parameter category of the maximum contribution value.

[0132] Specifically, in this process, by calculating the partial derivatives of the anomaly score with respect to each state parameter (such as power consumption, temperature, etc.) in the real-time state data, the contribution value of each state parameter to the anomaly score is obtained. The contribution value reflects the degree of influence of each state parameter on the anomaly score, and the calculation formula is the absolute value of the partial derivative. Finally, by selecting the state parameter with the largest contribution value, it is determined as the abnormal state parameter, indicating that this parameter has the greatest influence on the anomaly score and may be the main source of the fault. This method can accurately locate the source of the anomaly, quantify the contributions of various state parameters by calculating partial derivatives, and ensure that the system can accurately identify the parameter that has the greatest impact on the fault. Compared with the traditional threshold-based method, this fine-grained analysis improves the accuracy of fault detection and can effectively distinguish different abnormal causes in a complex working environment.

[0133] S8: Output the anomaly category based on the abnormal state parameter to complete the fault detection of LPCAMM2.

[0134] Specifically, the anomaly categories include power consumption anomaly, temperature anomaly, clock frequency anomaly, clock jitter anomaly, bit error rate anomaly, read rate anomaly, write rate anomaly, and working voltage anomaly.

[0135] In the actual application process, first, multiple groups of data of LPCAMM2 in the normal state are collected and the data is extended to simulate different working conditions. Then, an adversarial network is trained using the enhanced normal data, and the generator learns the distribution pattern of the normal data. Then, the current state data of LPCAMM2 is collected in real time, and the latent variables in the latent space are optimized by searching to minimize the reconstruction error. Then, the anomaly score is calculated according to the optimal latent variables, and the contributions of various state parameters to the anomaly score are further analyzed to finally accurately locate the fault source. The fault type corresponding to the abnormal state parameter is output. By intelligently learning the distribution of normal state data and the coupling relationship between multi-dimensional parameters, this system not only improves the accuracy of fault detection, but also significantly improves the speed of fault detection without increasing the computational burden, and has higher real-time performance and adaptability.

[0136] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:

[0137] In an embodiment of the present invention, data augmentation based on curve expansion is performed on normal state data including multiple state parameters, and then a generative adversarial network model is used to learn the distribution pattern and coupling relationship of the normal state data. Through the data augmentation method based on curve expansion, the distribution range of the normal state data is effectively expanded, and the learning ability of the generative adversarial network model for complex fault patterns is enhanced. Then, combined with the single-step random mapping gradient algorithm, the optimal latent variable that minimizes the recovery error of the real-time state data is quickly searched in the input data of the generator, i.e., the latent space, and the optimal latent variable can be quickly inferred, improving the fault detection speed. And an anomaly score of the real-time state data is calculated based on the optimal dependent variable, and the contribution value of different categories of state parameters to the anomaly score is calculated respectively, completing the precise positioning of the anomaly category. It can fully learn and capture the coupling relationship between high-dimensional data, effectively improving the speed and accuracy of fault detection while ensuring the fast response ability of fault detection, and can effectively cope with multi-parameter coupling anomalies in complex systems, with high practicability.

[0138] Refer to the appended drawings of the specification Figure 2 , which shows a schematic structural diagram of a fault detection system for LPCAMM2 provided by an embodiment of the present invention.

[0139] An embodiment of the present invention provides a fault detection system 20 for LPCAMM2, including: a processor 201 and a memory 202;

[0140] The memory 202 stores programs or instructions that can be run on the processor 201. When the programs or instructions are executed by the processor 201, the steps of the above-mentioned fault detection method for LPCAMM2 are implemented, and the same technical effects can be achieved. To avoid repetition, the present invention will not be described in detail.

[0141] It should be understood that the processor 201 in the embodiment of the present invention may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0142] It should also be understood that the memory 202 in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0143] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more sets of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0144] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0145] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0146] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0147] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0148] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0150] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0151] The embodiment of the present invention provides a readable storage medium including: a program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps of the above-mentioned fault detection method of LPCAMM2 are implemented, and the same technical effects can be achieved. To avoid repetition, the present invention will not be described in detail.

[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A fault detection method for LPCAMM2, characterized in that: include: S1: Acquire multiple groups of normal state data of various state parameters of the LPCAMM2; S2: taking the state data calibration range of the LPCAMM2 as a constraint, performing data enhancement based on curve extension on the normal state data; S3: Construct the generator and discriminator of the generative adversarial network model based on the enhanced normal state data; S4: Get real-time status data; S5: Combined with a single-step random mapping gradient algorithm, searching for an optimal latent variable that minimizes the recovery error of the real-time state data in the input data of the generator, i.e., the latent space; S6: Calculating an abnormality score of the real-time status data according to the optimal latent variable; S7: When the abnormality score is greater than a preset abnormality score, calculating the contribution value of the real-time state data corresponding to each of the state parameters to the abnormality score, and determining the state parameter corresponding to the maximum contribution value as the abnormal state parameter; S8: Outputting an abnormality category based on the abnormal state parameter to complete the fault detection of the LPCAMM2; Wherein, the recovery error includes the generator reconstruction error and the discriminator recognition error; The optimal latent variable search process specifically includes: Randomly select basic latent variables from the latent space, and calculate generated samples representing state data through the generator; Based on the basic latent variables, the discriminant probabilities of the generated samples and the real-time state data are respectively calculated by the discriminator: Calculating the square of the L1 norm between the real-time state data and the generated sample to obtain the generator reconstruction error; Calculating the square of the L1 norm between the generated sample and the discrimination probability of the real-time state data to obtain the discriminator recognition error; Calculating a first weighted value between the generator reconstruction error and the discriminator recognition error; In the latent space, the basic latent variables are optimized in combination with the single-step random mapping gradient algorithm to obtain optimized latent variables; Using the optimized latent variable to update the basic latent variable, returning the discriminant probability of the generated sample and the real-time state data respectively calculated by the discriminator based on the basic latent variable, until the number of iterations is greater than the preset number of iterations; Outputting the basic latent variable that minimizes the first weighted sum value between the generator reconstruction error and the discriminator recognition error as the optimal latent variable; In the latent space, the basic latent variables are optimized in combination with the single-step random mapping gradient algorithm to obtain optimized latent variables, specifically including: Projecting the latent space into a two-dimensional random subspace to obtain low-dimensional latent variables; Based on the local linearity assumption of the generator, optimizing the low-dimensional latent variables once; The low-dimensional latent variable is optimized twice through the discriminator residual to obtain the optimized latent variable.

2. The fault detection method of LPCAMM2 according to claim 1, characterized in that: The state parameters include power consumption, temperature, clock frequency, clock jitter, bit error rate, read rate, write rate and operating voltage.

3. The fault detection method of LPCAMM2 according to claim 1, characterized in that: The S2 specifically includes: Fitting the normal state data belonging to the same state parameter category in each group of the normal state data to obtain a smooth function representing a smooth curve, wherein each of the smooth curves constitutes a high-dimensional space curve; Calculate the derivative of each smoothing function with respect to the path parameter; Calculate the coupled derivatives between any two derivatives; Calculating a tangent vector of the high-dimensional space curve; Setting disturbance parameters for data enhancement in combination with the state data calibration range and the coupling derivative; In combination with the state data calibration range, the tangent vector direction is used as the disturbance direction, and the disturbance parameter is added to the high-dimensional space curve to obtain enhanced normal state data.

4. The fault detection method of LPCAMM2 according to claim 1, characterized in that: The S3 specifically includes: inputting the enhanced normal state data as a training set into the generative adversarial network model; The discriminator and the generator are alternately fixed, and the unfixed discriminator and generator are trained respectively to complete the construction of the generative adversarial network model.

5. The fault detection method of LPCAMM2 according to claim 1, characterized in that: The S6 specifically includes: Respectively calculating the reconstruction error and the discriminator recognition error under the optimal latent variable; Calculating a second weighted sum value between the reconstruction error and the discriminator recognition error; The second weighted sum value is output as the anomaly score.

6. The fault detection method of LPCAMM2 according to claim 5, characterized in that: The S7 specifically includes: Calculating partial derivatives of the anomaly score with respect to various state parameters in the real-time state data, and using the partial derivatives as contribution values; The state parameter corresponding to the maximum contribution value is output as the abnormal state parameter.

7. A fault detection system for LPCAMM2, characterized in that: include: Processor and memory; The memory stores a program or an instruction executable on the processor, and when the program or the instruction is executed by the processor, the steps of the fault detection method of the LPCAM M2 according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Equipment abnormality detection method and device, computer readable storage medium and chip

    CN113056750A

  • Abnormality detection system, abnormality detection method, abnormality detection program, and method for generating learned model

    US20180365089A1