LPCAMM2 fault detection method and system

Through the combination of curve expansion-based data augmentation and generative adversarial network model, combined with a single-step random mapping gradient algorithm, the problem of the existing LPCAMM2 fault detection method ignores state parameter coupling, and achieves more efficient and accurate fault detection.

CN120045378AActive Publication Date: 2025-05-27SHENZHEN COMOS TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510519452.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing LPCAMM2 fault detection method is based on static thresholds, ignoring the coupling relationship between different state parameters, resulting in extremely low detection accuracy and insufficient practicality.

Method used

Using a data enhancement method based on curve extension, multiple sets of normal state data are obtained, and a generative adversarial network model is constructed. Combined with a single-step random mapping gradient algorithm, the optimal hidden variable is searched, the abnormal score of real-time state data is calculated, and the abnormal state parameters are determined.

Benefits of technology

It effectively expands the distribution range of normal state data, enhances the learning ability of the generative adversarial network model for complex failure modes, improves the speed and accuracy of fault detection, and can effectively deal with multi-parameter coupling exceptions in complex systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045378A_ABST
    Figure CN120045378A_ABST
Patent Text Reader

Abstract

The invention provides a fault detection method and system for an LPCAMM2, and relates to the technical field of memory testing, and the method comprises the steps: obtaining a plurality of groups of normal state data of each state parameter of the LPCAMM2; performing data enhancement based on curve expansion on the normal state data; constructing a generator and a discriminator of the generative adversarial network model; acquiring real-time state data; searching an optimal hidden variable; calculating an abnormal score of the real-time state data; calculating the contribution value of the real-time state data corresponding to each state parameter to the abnormal score, and determining the state parameter corresponding to the maximum contribution value as an abnormal state parameter; and outputting an abnormal category based on the abnormal state parameters to complete fault detection. The fault detection speed and accuracy are effectively improved while the quick response capability of fault detection is ensured, multi-parameter coupling abnormity in a complex system can be effectively dealt with, and the practicability is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of memory testing, and particularly to a method and system for detecting faults in LPCAMM2. Background Art

[0002] LPCAMM2 (Low Power Compressed Add-on Memory Module) is a low-power add-on memory module based on LPDDR5X technology. It adopts a brand-new design architecture. Compared with traditional SODIMM memory modules, it can provide more efficient packaging and reduce the space occupation by 60%. LPCAMM2 supports higher data transfer rates and can significantly reduce power consumption, and has a wide range of applications in high-performance computing, gaming workstations, data centers and other applications. Its detachable design enables manufacturers to more flexibly configure and upgrade memory modules in devices.

[0003] As a high-performance memory module, LPCAMM2 undertakes important computing tasks and data transfer functions. Any hardware failure will directly affect the stability and performance of the device. Fault detection can identify potential problems in advance, reduce the risks of system crashes, data loss or performance degradation, and ensure the efficient operation of the device and the user experience. Timely detection and repair of LPCAMM2 faults are crucial for ensuring the stability of critical applications (such as high-performance computing, gaming, etc.).

[0004] However, the existing fault detection of LPCAMM2 often detects single faults based on static thresholds. Although the detection efficiency is high, due to ignoring the coupling relationship between different state parameters, the detection accuracy is extremely low and the practicability is insufficient. Summary of the Invention

[0005] In view of the above deficiencies in the prior art, the purpose of the embodiments of the present invention is to provide a method and system for detecting faults in LPCAMM2, which can solve the technical problem that the existing fault detection of LPCAMM2 often detects single faults based on static thresholds. Although the detection efficiency is high, due to ignoring the coupling relationship between different state parameters, the detection accuracy is extremely low and the practicability is insufficient.

[0006] In the first aspect of the embodiments of the present invention, a method for detecting faults in LPCAMM2 is proposed, including:

[0007] S1: Obtain multiple groups of normal state data of each state parameter of the LPCAMM2;

[0008] S2: Perform data enhancement based on curve extension on the normal state data with the calibration range of the state data of the LPCAMM2 as a constraint;

[0009] S3: Construct the generator and discriminator of the generative adversarial network model based on the enhanced normal state data;

[0010] S4: Obtain real-time state data;

[0011] S5: Combine the single-step random mapping gradient algorithm to search for the optimal latent variable in the input data of the generator, i.e., the latent space, that minimizes the recovery error of the real-time state data;

[0012] S6: Calculate the anomaly score of the real-time state data according to the optimal latent variable;

[0013] S7: When the anomaly score is greater than the preset anomaly score, calculate the contribution value of the real-time state data corresponding to each state parameter to the anomaly score, and determine the state parameter corresponding to the maximum contribution value as the abnormal state parameter;

[0014] S8: Output the anomaly category based on the abnormal state parameter to complete the fault detection of the LPCAMM2.

[0015] In the second aspect of the embodiments of the present invention, a fault detection system for LPCAMM2 is proposed, including: a processor and a memory;

[0016] The memory stores programs or instructions that can run on the processor, and when the programs or instructions are executed by the processor, the steps of the fault detection method for LPCAMM2 described in the first aspect are implemented.

[0017] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include: In the embodiments of the present invention, data enhancement based on curve expansion is performed on the normal state data including multiple state parameters, and then the generative adversarial network model is used to learn the distribution pattern and coupling relationship of the normal state data. Through the data enhancement method based on curve expansion, the distribution range of the normal state data is effectively expanded, and the learning ability of the generative adversarial network model for complex fault patterns is enhanced. Then, combined with the single-step random mapping gradient algorithm, the optimal latent variable that minimizes the recovery error of the real-time state data is quickly searched in the input data of the generator, i.e., the latent space, which can quickly infer the optimal latent variable and improve the fault detection speed. And based on this optimal dependent variable, the anomaly score of the real-time state data is calculated, and the contribution values of different categories of state parameters to the anomaly score are calculated respectively to complete the precise positioning of the anomaly category, which can fully learn and capture the coupling relationship between high-dimensional data, effectively improve the speed and accuracy of fault detection while ensuring the fast response ability of fault detection, and can effectively handle multi-parameter coupling anomalies in complex systems, with high practicality. Description of the Drawings

[0018] The accompanying drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference numerals represent the same components. Obviously, the drawings in the following description are only some embodiments described in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a schematic flowchart of a fault detection method for LPCAMM2 provided by an embodiment of the present invention;

[0020] Figure 2 is a schematic structural diagram of a fault detection system for LPCAMM2 provided by an embodiment of the present invention. Detailed Embodiments

[0021] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. It should be understood that these descriptions are only exemplary and are not used to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] The following will describe in detail the fault detection method for LPCAMM2 provided by the embodiments of the present invention through specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0023] Refer to the attached drawings of the specification Figure 1 , which shows a schematic flowchart of a fault detection method for LPCAMM2 provided by an embodiment of the present invention.

[0024] The embodiments of the present invention provide a fault detection method for LPCAMM2, which may include the following steps:

[0025] S1: Obtain multiple groups of normal state data of each state parameter of LPCAMM2.

[0026] Among them, the normal state data refers to the key state parameter data recorded by LPCAMM2 (Low Power Compression Additional Memory Module) under normal working conditions. These data represent the typical behaviors of LPCAMM2 without faults or abnormalities, such as power consumption, temperature, clock frequency, bit error rate, read rate, write rate, etc. The normal state data is used to establish a benchmark model to help the subsequent fault detection system determine whether the real-time data exceeds the normal range.

[0027] Specifically, the normal state data can be obtained by monitoring its key parameters in real time under different workloads during the operation of LPCAMM2. The specific steps include: First, ensure that LPCAMM2 runs in a stable and normal operating environment. Then, use hardware monitoring tools or embedded sensors to collect data such as power consumption, temperature, and frequency. Finally, store multiple sets of these data to ensure coverage of different operating conditions and environments as a benchmark for the subsequent fault detection system.

[0028] In a possible implementation, the state parameters include power consumption, temperature, clock frequency, clock jitter, bit error rate, read rate, write rate, and operating voltage.

[0029] S2: Constrained by the calibration range of the state data of LPCAMM2, perform data augmentation based on curve expansion on the normal state data.

[0030] Among them, the calibration range of the state data refers to the allowable value range of each state parameter (such as power consumption, temperature, clock frequency, etc.) when performing fault detection on LPCAMM2. This range is usually based on the design specifications of LPCAMM2 or the normal operating range obtained through experiments, and is used to limit and constrain the variation of the data. By the calibration range, it can be ensured that the results of data augmentation do not exceed reasonable physical or operating limits. The data augmentation based on curve expansion is a mathematical method for expanding or varying the normal state data. Specifically, by expanding the normal state data based on a curve, the system can generate new data points, expand the range of the data set, and enhance the diversity of the data. This method can help the model better learn different variation patterns of the state data, increase the diversity of the data set, and improve the robustness and generalization ability of the model.

[0031] It should be noted that by performing data augmentation based on curve expansion on the normal state data of LPCAMM2, more diverse samples can be generated within the calibration range of the state data to simulate different operating conditions. This method not only expands the scale of the data set but also enables the detection system to better learn various potential changes in the operating state, improving the accuracy and adaptability of fault detection. By controlling the curve expansion during the data augmentation process, it is ensured that the data changes conform to the actual situation, thus avoiding the generation of unreasonable abnormal data.

[0032] In a possible implementation, the specific content of S2 includes:

[0033] Fit the normal state data belonging to the same state parameter category in each group of normal state data to obtain a smooth function representing a smooth curve, where each smooth curve forms a high-dimensional space curve.

[0034] The generation formula of the high-dimensional space curve is specifically: ; ; Among them, represents a path parameter, represents the normal state data of the i-th type of state parameter under the path parameter of the smoothing function, , m represents the number of state parameter categories, represents the high-dimensional space curve related to including each group of normal state data and , represents the j-th normal state data of the given i-th type of state parameter, , n represents the number of groups of normal state data, represents B-spline interpolation.

[0035] In this process, first, the normal state data belonging to the same state parameter category are fitted to obtain the smoothing function of each category. Using the B-spline interpolation method, these data can be smoothed to obtain a smooth curve. B-spline interpolation is an efficient smoothing method that can effectively reduce the noise in the data and accurately capture the trend of the normal state data. Doing so can not only avoid the influence of abnormal data on model training but also enable the model to better learn the internal laws and change trends of the data.

[0036] Calculate the derivative of each smoothing function with respect to the path parameter.

[0037] In this process, by taking the derivative of each smoothing function, the change rate of each state parameter under the path parameter is obtained. This step is to calculate the sensitivity of the state parameter to the path change. Calculating the derivative can help understand the rate of change of the state parameter with respect to the path parameter. This can capture the subtle changes between the data and contribute to more accurate modeling of the state data change in the high-dimensional space in the future.

[0038] Calculate the coupled derivative between two derivatives.

[0039] The specific calculation formula for the coupled derivative is: ; Among them, represents the smoothing function of the normal state data of the v-th type of state parameter under the path parameter , represents the coupled derivative between the i-th type of state parameter and the v-th type of state parameter.

[0040] In this process, for different state parameters (such as power consumption, temperature, clock frequency, etc.), the coupled derivatives between them are calculated. The coupled derivative represents the change relationship between two state parameters. By comparing the derivatives of different state parameters, the mutual influence between them can be revealed. Furthermore, the correlation and interaction between multiple state parameters can be revealed. The multiple state parameters applicable to LPCAMM2 are in a mutually coupled working state. Understanding these coupling relationships helps improve the detection ability for complex fault modes.

[0041] Calculate the tangent vector of the high-dimensional space curve.

[0042] The specific calculation formula for the tangent vector is: .

[0043] In this process, by calculating the tangent vector of the high-dimensional space curve, the change trend of each state parameter is understood. This step helps the model understand the multi-dimensional characteristics of the data. The tangent vector can describe the change direction of the data in the multi-dimensional space. Through this direction, the model can learn how to make trade-offs between various state parameters, thereby improving the understanding ability of complex data relationships.

[0044] Combine the state data calibration range and the coupled derivative to set the perturbation parameter for data augmentation.

[0045] The specific calculation formula for the perturbation parameter is: ; where represents the perturbation parameter, and respectively represent the upper limit value and the lower limit value of the state data calibration range of the i-th type of state parameter.

[0046] In this process, the perturbation parameter is calculated based on the product of the calibration range of the state data (i.e., the maximum and minimum allowed values of each state parameter) and the coupled derivative. By combining the actual range of the state data, the perturbation parameter can ensure that the generated data does not exceed the physical or operational limits, avoiding the generation of unrealistic abnormal data. At the same time, the perturbation parameter can effectively expand the diversity of the data and improve the quality of the generated samples.

[0047] Combine the state data calibration range, use the tangent vector direction as the perturbation direction, and add the perturbation parameter on the high-dimensional space curve to obtain the enhanced normal state data.

[0048] The enhanced normal state data is specifically: ; ; where Indicates the smoothing function of the enhanced normal state data under the path parameter and represents the state data of the i-th type of state parameter.

[0049] It should be noted that the calculated perturbation parameters are applied to the high-dimensional space curve, and perturbations are increased along the direction of the tangent vector to generate new data points. These data points represent the diverse situations that may occur under normal conditions. By generating enhanced data highly correlated with the actual normal working state, the model can better learn the potential laws of normal data. Since the enhanced data covers more possible situations, the fault detection system can make accurate judgments within a wider range, further improving the accuracy and efficiency of fault detection.

[0050] S3: Construct the generator and discriminator of the generative adversarial network model based on the enhanced normal state data.

[0051] Among them, the generative adversarial network (GAN) model is a deep learning model composed of two neural networks: the generator and the discriminator. The goal of the generator is to generate fake data as similar as possible to the real data, while the discriminator is responsible for distinguishing the generated data from the real data. The generator and the discriminator compete with each other during the training process. Through continuous optimization, the generator can finally generate high-quality samples, and the discriminator can more accurately distinguish real data from generated data. Training the GAN with normal data, the generator G learns the distribution law of the normal state data. In this way, the generator can inversely capture the coupling relationship of LPCAMM2 under normal conditions, providing an accurate benchmark for subsequent detection of abnormal data. This adversarial training method enables the model not only to generate similar normal data but also to better identify abnormal data with different distributions from these data.

[0052] In a possible implementation manner, the S3 specifically includes: Input the enhanced normal state data as the training set into the generative adversarial network model.

[0053] Adopt the method of alternately fixing the discriminator and the generator, and train the unfixed discriminator and generator respectively to complete the construction of the generative adversarial network model.

[0054] It should be noted that during the training process of the Generative Adversarial Network (GAN), the common practice is to first fix the discriminator and then train the generator, and then fix the generator and train the discriminator. This is because the task of the discriminator is to distinguish between real and generated data, while the task of the generator is to generate as realistic data as possible. The training sequence is as follows: Fix the discriminator: Keep the discriminator unchanged and train the generator to generate more realistic data. Fix the generator: Keep the generator unchanged and train the discriminator to distinguish between real data and generated data. This alternating training method can help the generator continuously improve to generate more real samples, and at the same time enable the discriminator to more accurately distinguish between true and false data.

[0055] S4: Obtain real-time status data.

[0056] Among them, the real-time status data refers to various dynamic status information collected at the current moment during the actual operation of LPCAMM2.

[0057] S5: Combine the single-step stochastic mapping gradient algorithm to search for the optimal latent variable in the input data of the generator, i.e., the latent space, that minimizes the recovery error of the real-time status data.

[0058] Among them, the single-step stochastic mapping gradient algorithm is an optimization algorithm that combines stochastic mapping and gradient descent methods. By gradually adjusting the latent variables in the latent space, it seeks the optimal solution that can minimize a certain objective function (such as the recovery error). In each step of the optimization process, the algorithm randomly selects the initial latent variable and gradually approaches the optimal solution based on the gradient of the objective function, thereby improving the search efficiency. The recovery error refers to the difference between the samples generated from the latent space by the generator and the real-time status data. It reflects the quality of the generated samples. If the recovery error is large, it indicates a large deviation between the generated samples and the real data, indicating that the effect of the generator is not ideal. Minimizing the recovery error is a key objective in the optimization process. The optimal latent variable refers to the latent variable in the latent space that, after being optimized by the algorithm, can make the generated samples best match the actual status data. The latent variable is the input in the generative adversarial network, and it determines the characteristics of the generated samples. By optimizing the optimal latent variable, the generator can more accurately generate samples that conform to the real data distribution, thereby improving the accuracy of fault detection.

[0059] It should be noted that by combining the single-step stochastic mapping gradient algorithm, the optimal latent variable can be efficiently searched in the latent space to minimize the recovery error of the real-time status data. In this way, the system can accelerate the fault detection process and reduce the number of iterations. The search for the optimal latent variable ensures that the difference between the generated data and the actual data is minimized, thereby improving the accuracy and efficiency of the model.

[0060] In a possible implementation, the recovery error includes the generator reconstruction error and the discriminator recognition error.

[0061] It should be noted that the recovery error is composed of the generator reconstruction error and the discriminator recognition error. The generator reconstruction error represents the deviation between the generator and the real data when attempting to reconstruct the real-time state data. The discriminator recognition error represents the error of the discriminator in distinguishing between the generated data and the real data. Combining these two errors can more comprehensively measure the quality of the generated data and optimize the GAN model to improve the accuracy and stability of fault detection.

[0062] In a possible implementation, the search process for the optimal latent variable specifically includes: Randomly select a base latent variable from the latent space and calculate the generated sample representing the state data through the generator.

[0063] Based on the base latent variable, calculate the discrimination probabilities of the generated sample and the real-time state data through the discriminator respectively: Calculate the square of the L1 norm between the real-time state data and the generated sample to obtain the generator reconstruction error.

[0064] Calculate the square of the L1 norm between the discrimination probabilities of the generated sample and the real-time state data to obtain the discriminator recognition error.

[0065] Calculate the first weighted sum value between the generator reconstruction error and the discriminator recognition error.

[0066] In the latent space, optimize the base latent variable by combining the single-step random mapping gradient algorithm to obtain the optimized latent variable.

[0067] In a possible implementation, in the latent space, optimize the base latent variable by combining the single-step random mapping gradient algorithm to obtain the optimized latent variable, which specifically includes: Project the latent space onto a two-dimensional random subspace to obtain a low-dimensional latent variable.

[0068] The mapping process is specifically: ; where P represents a randomly generated orthogonal matrix, represents the square of the L1 norm, represents the real-time state data, represents the latent variable z that makes the function take the minimum value, represents the real number field, d represents the dimension of the latent space, represents based on the latent variable generated generated sample, represents the low-dimensional latent variable in the two-dimensional random subspace.

[0069] Among them, represents a two-dimensional real number space, represents a d×2 matrix. That is, P is a matrix of size d×2, which is used to map the original high-dimensional latent space data (dimension d) to a two-dimensional space.

[0070] Based on the local linear assumption of the generator, the low-dimensional latent variable is optimized once.

[0071] The specific formula for the first optimization is: ; Among them, represents a fixed step size, represents the generated sample based on the low-dimensional latent variable represents the Jacobian matrix of the generator G at the low-dimensional latent variable represents the gradient of the discriminator D with respect to z for the generated sample represents the low-dimensional latent variable after the first optimization, represents the first weight factor.

[0072] Optionally, the fixed step size can be set to 0.1. The first weight factor can be set to 0.5.

[0073] The low-dimensional latent variable is optimized twice through the discriminator residual to obtain the optimized latent variable.

[0074] The specific formula for the second optimization is: ; Among them, represents the calibration coefficient, represents the discriminant probability of the real-time state data represents the discriminant probability of the generated sample generated based on the latent variable after the first optimization, represents the optimized latent variable obtained by the second optimization, represents the gradient of the discriminator with respect to z for

[0075] Optionally, the calibration coefficient can be set to 0.01.

[0076] ​​​​​Specifically, first, the latent variables in the latent space are projected onto a two-dimensional random subspace. The purpose of this step is to reduce the dimension of the latent space for more efficient search of the optimal latent variables. Using a randomly generated orthogonal matrix P, the latent space is mapped to a two-dimensional subspace, and the base latent variable z is determined by minimizing the squared L1 norm between the generated samples and the real data. 0 . This ensures that the generated samples are as close as possible to the actual real-time state data. Based on the low-dimensional latent variable z in the low-dimensional subspace 0 , the generator is optimized through the local linear assumption. By calculating the difference between the generated samples and the real-time state data, z 0 is adjusted and updated to the latent variable z 1 . During one optimization process, the latent variable z 1 is optimized by reducing the error between the generated samples and the real data. The optimization process involves the Jacobian matrix of the generator and the gradient of the discriminator, which can ensure that the generated samples are closer to the real data. The secondary optimization further reduces the difference between the generated samples and the real-time data by adjusting the output of the discriminator. By adjusting the gradient of the discriminator, the generated samples more precisely match the real data.

[0077] It can be understood that by reducing the latent space to a two-dimensional subspace and combining with the single-step random mapping gradient algorithm, the dimension of the search space and the computational complexity are reduced. By gradually optimizing the latent variables, the number of iterations required is reduced, thus accelerating the determination speed of the optimal latent variables. This method is more efficient than traditional high-dimensional space optimization, enabling the generator to quickly generate samples that match the real-time data, thereby improving the speed of fault detection.

[0078] Update the base latent variable using the optimized latent variable, and return based on the base latent variable. Calculate the discrimination probabilities of the generated samples and the real-time state data through the discriminator respectively until the number of iterations is greater than the preset number of iterations.

[0079] It should be noted that those skilled in the art can set the size of the preset number of iterations according to actual needs, and the present invention does not make any limitations in this regard.

[0080] Output the base latent variable that minimizes the first weighted sum value between the reconstruction error of the generator and the recognition error of the discriminator as the optimal latent variable.

[0081] Specifically, the search process for the optimal latent variable aims to find a latent variable that minimizes the error between the samples generated by the generator and the real-time state data. First, a basic latent variable is randomly selected from the latent space, and the corresponding generated samples are calculated using the generator. Then, the discriminator is used to calculate the discrimination probability between the generated samples and the real-time state data, and the generator reconstruction error (the square of the L1 norm) and the discriminator recognition error (the square of the L1 norm between the discrimination probabilities of the generated samples and the real-time data) are calculated respectively. Next, these two errors are weighted and summed to measure the effect of the current latent variable. To find the optimal latent variable, the system optimizes the basic latent variable in the latent space using the single-step random mapping gradient algorithm to obtain the optimized latent variable, and continuously updates the basic latent variable, repeating the above calculation process until the preset number of iterations is reached. Finally, the latent variable that minimizes the weighted error is selected as the optimal latent variable. Using the single-step random mapping gradient algorithm can quickly optimize the latent variable, reduce the computational complexity, and accelerate the fault detection speed. By combining the generator reconstruction error and the discriminator recognition error, it ensures the highest degree of matching between the generated samples and the real data, and improves the detection accuracy. By comprehensively using the generation ability and discrimination ability of the GAN, the fault detection not only depends on the data similarity, but also considers the recognition ability of the discriminator, thereby reducing false alarms and missed detections.

[0082] The calculation method of the optimal latent variable is specifically as follows: ; Wherein, represents the optimal latent variable, represents the basic latent variable when the function takes the minimum value , represents the square of the L1 norm, represents the generated samples based on the basic latent variable , represents the real-time state data of the discrimination probability, represents of the discrimination probability, represents the second weight factor.

[0083] Optionally, the second weight factor can be set to 0.5.

[0084] It should be noted that this method of calculating the optimal latent variable ensures the highest degree of matching between the generated samples and the real-time state data by minimizing the weighted sum of the generator reconstruction error and the discriminator recognition error, while considering the judgment ability of the discriminator. Using the square of the L1 norm to measure the error can more accurately measure the difference between the generated data and the real data, and further improve the quality of the generated samples. By introducing the second weight factor ( ), it can flexibly adjust the influence weights of the generation error and the discrimination error, achieve better error balance, and thus improve the accuracy and reliability of the model.

[0085] S6: Calculate the anomaly score of the real-time status data according to the optimal latent variable.

[0086] Specifically, by inputting the optimal latent variable into the generative adversarial network (GAN) model, calculate the difference between the generated sample and the actual real-time status data, and evaluate whether the data is abnormal according to this difference (such as the reconstruction error). The level of the anomaly score indicates the deviation between the real-time status data and the normal data. The higher the score, the greater the degree of data abnormality, and vice versa, indicating that the data tends to be normal. The goal of this step is to provide a basis for fault detection by accurately calculating the anomaly score.

[0087] In a possible implementation manner, the S6 specifically includes:

[0088] Calculate the reconstruction error and the discriminator recognition error under the optimal latent variable respectively.

[0089] Calculate the second weighted sum value between the reconstruction error and the discriminator recognition error.

[0090] Output the second weighted sum value as the anomaly score.

[0091] The specific calculation formula of the anomaly score is: ; where represents the anomaly score.

[0092] Among them, the reconstruction error represents the difference between the generated sample and the actual data, and the discriminator recognition error represents the discriminator's ability to distinguish between the generated sample and the real data. By calculating the weighted sum of these two errors, the final anomaly score is obtained. This score reflects the similarity between the generated sample and the real data. The higher the score, the more abnormal the data. By considering both the reconstruction error and the discriminator recognition error, the quality of the generated data is comprehensively evaluated, making the anomaly detection more accurate. The calculation of the weighted sum value can dynamically balance the influence of the generator and the discriminator, and improve the robustness of fault detection.

[0093] S7: When the anomaly score is greater than the preset anomaly score, calculate the contribution value of the real-time status data corresponding to each state parameter to the anomaly score, and determine the state parameter corresponding to the maximum contribution value as the abnormal state parameter.

[0094] Optionally, by analyzing the historical status data collected by LPCAMM2 during normal operation, the mean and standard deviation of each status parameter (such as power consumption, temperature, clock frequency, etc.) can be calculated. Then, according to statistical methods, the threshold of the anomaly score is set. Usually, the mean plus a certain multiple (3 times or 5 times) of the standard deviation can be selected as the preset anomaly score. Or rely on the experience of engineers or the judgment of domain experts to set the threshold of the anomaly score. According to the usage scenario of LPCAMM2 (such as high-performance computing, data center applications, etc.), experts will set appropriate scoring thresholds according to the performance requirements and tolerances of the system. The threshold of the anomaly score can also be set according to the performance standards required by LPCAMM2. For some high-performance applications (such as gaming workstations or data centers), the threshold can be set to ensure that the system always operates in a high-performance state.

[0095] It should be noted that when the calculated anomaly score is greater than the preset anomaly score, the system will further calculate the contribution value of each status parameter (such as power consumption, temperature, clock frequency, etc.) to the anomaly score. This can help determine which specific parameters have the greatest impact on the anomaly score, thereby identifying the status parameters most likely to cause failures. In this way, the system can accurately locate the fault source and improve the accuracy and response speed of fault detection.

[0096] It should be noted that those skilled in the art can set the size of the preset anomaly score according to actual needs, and the present invention does not make any limitations here.

[0097] In a possible implementation manner, the S7 specifically includes:

[0098] Calculate the partial derivative of the anomaly score with respect to various status parameters in the real-time status data, and use the partial derivative as the contribution value.

[0099] The specific calculation method of the contribution value is: ; Wherein, represents the real-time status data corresponding to the i-th type of status parameter, represents taking the absolute value, represents taking the partial derivative, represents the partial derivative of the anomaly score with respect to the i-th type of status parameter in the real-time status data, that is, the contribution value.

[0100] Output the status parameter corresponding to the maximum contribution value as the abnormal status parameter.

[0101] The specific formula for determining the abnormal status parameter is: ; Wherein, represents the status parameter index corresponding to the abnormal status parameter category corresponding to the maximum contribution value.

[0102] Specifically, in this process, by calculating the partial derivatives of the anomaly score with respect to each state parameter (such as power consumption, temperature, etc.) in the real-time state data, the contribution value of each state parameter to the anomaly score is obtained. The contribution value reflects the degree of influence of each state parameter on the anomaly score, and the calculation formula is the absolute value of the partial derivative. Finally, by selecting the state parameter with the largest contribution value, it is determined as the abnormal state parameter, indicating that this parameter has the greatest influence on the anomaly score and may be the main source of the fault. This method can accurately locate the source of the anomaly, quantify the contributions of various state parameters by calculating partial derivatives, and ensure that the system can accurately identify the parameter that has the greatest impact on the fault. Compared with the traditional threshold-based method, this fine-grained analysis improves the accuracy of fault detection and can effectively distinguish different anomaly causes in a complex working environment.

[0103] S8: Output the anomaly category based on the abnormal state parameter to complete the fault detection of LPCAMM2.

[0104] Specifically, the anomaly categories include power consumption anomaly, temperature anomaly, clock frequency anomaly, clock jitter anomaly, bit error rate anomaly, read rate anomaly, write rate anomaly, and working voltage anomaly.

[0105] In the actual application process, first, multiple groups of data of LPCAMM2 in the normal state are collected and the data is extended to simulate different working conditions. Then, the enhanced normal data is used to train a generative adversarial network, and the generator learns the distribution pattern of the normal data. Then, the current state data of LPCAMM2 is collected in real time, and the latent variables in the latent space are optimized by searching to minimize the reconstruction error. Then, the anomaly score is calculated according to the optimal latent variables, and the contributions of various state parameters to the anomaly score are further analyzed to finally accurately locate the fault source. The fault type corresponding to the abnormal state parameter is output. By intelligently learning the distribution of the normal state data and the coupling relationship between multi-dimensional parameters, this system not only improves the accuracy of fault detection, but also significantly improves the speed of fault detection without increasing the computational burden, and has higher real-time performance and adaptability.

[0106] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include: In the embodiment of the present invention, data augmentation based on curve expansion is performed on the normal state data including multiple state parameters, and then the generative adversarial network model is used to learn the distribution pattern and coupling relationship of the normal state data. Through the data augmentation method based on curve expansion, the distribution range of the normal state data is effectively expanded, and the learning ability of the generative adversarial network model for complex fault patterns is enhanced. Then, combined with the single-step stochastic mapping gradient algorithm, the optimal latent variable that minimizes the recovery error of the real-time state data is quickly searched in the input data of the generator, that is, the latent space, and the optimal latent variable can be quickly inferred, improving the fault detection speed. And based on this optimal dependent variable, the anomaly score of the real-time state data is calculated, and the contribution value of different categories of state parameters to the anomaly score is calculated respectively, completing the precise positioning of the anomaly category. It can fully learn and capture the coupling relationship between high-dimensional data, effectively improve the speed and accuracy of fault detection while ensuring the fast response ability of fault detection, and can effectively handle the multi-parameter coupling anomaly in a complex system, with high practicability.

[0107] Refer to the attached drawings of the specification Figure 2 , which shows a schematic structural diagram of a fault detection system for LPCAMM2 provided by an embodiment of the present invention.

[0108] An embodiment of the present invention provides a fault detection system 20 for LPCAMM2, including: a processor 201 and a memory 202; The memory 202 stores programs or instructions that can be run on the processor 201. When the programs or instructions are executed by the processor 201, the steps of the above-mentioned fault detection method for LPCAMM2 are implemented, and the same technical effects can be achieved. To avoid repetition, the present invention will not be described in detail.

[0109] It should be understood that the processor 201 in the embodiment of the present invention may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0110] It should also be understood that the memory 202 in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0111] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0112] It should be understood that in various embodiments of the present invention, the order numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0113] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0114] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated herein.

[0115] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.

[0116] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0117] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0118] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0119] The embodiment of the present invention provides a readable storage medium, including: a program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, it implements the steps of the above-mentioned fault detection method of LPCAMM2 and can achieve the same technical effect. To avoid repetition, the present invention will not be elaborated herein.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily conceived by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A fault detection method for LPCAMM2, characterized in that: include: S1: Acquire multiple groups of normal state data of various state parameters of the LPCAMM2; S2: taking the state data calibration range of the LPCAMM2 as a constraint, performing data enhancement based on curve extension on the normal state data; S3: Construct the generator and discriminator of the generative adversarial network model based on the enhanced normal state data; S4: Get real-time status data; S5: Combined with a single-step random mapping gradient algorithm, searching for an optimal latent variable that minimizes the recovery error of the real-time state data in the input data of the generator, i.e., the latent space; S6: Calculating an abnormality score of the real-time status data according to the optimal latent variable; S7: When the abnormality score is greater than a preset abnormality score, calculating the contribution value of the real-time state data corresponding to each of the state parameters to the abnormality score, and determining the state parameter corresponding to the maximum contribution value as the abnormal state parameter; S8: Outputting an abnormality category based on the abnormal state parameter to complete the fault detection of the LPCAMM2.

2. The fault detection method of LPCAMM2 according to claim 1, characterized in that: The state parameters include power consumption, temperature, clock frequency, clock jitter, bit error rate, read rate, write rate and operating voltage.

3. The fault detection method of LPCAMM2 according to claim 1, characterized in that: The S2 specifically includes: Fitting the normal state data belonging to the same state parameter category in each group of the normal state data to obtain a smooth function representing a smooth curve, wherein each of the smooth curves constitutes a high-dimensional space curve; Calculate the derivative of each smoothing function with respect to the path parameter; Calculate the coupled derivatives between any two derivatives; Calculating a tangent vector of the high-dimensional space curve; Setting disturbance parameters for data enhancement in combination with the state data calibration range and the coupling derivative; In combination with the state data calibration range, the tangent vector direction is used as the disturbance direction, and the disturbance parameter is added to the high-dimensional space curve to obtain enhanced normal state data.

4. The fault detection method of LPCAMM2 according to claim 1, characterized in that: The S3 specifically includes: Inputting the enhanced normal state data into the generative adversarial network model as a training set; The discriminator and the generator are alternately fixed, and the unfixed discriminator and generator are trained respectively to complete the construction of the generative adversarial network model.

5. The fault detection method of LPCAMM2 according to claim 1, characterized in that: The recovery error includes a generator reconstruction error and a discriminator recognition error.

6. The fault detection method of LPCAMM2 according to claim 5, characterized in that: The optimal latent variable search process specifically includes: Randomly select basic latent variables from the latent space, and calculate generated samples representing state data through the generator; Based on the basic latent variables, the discriminant probabilities of the generated samples and the real-time state data are respectively calculated by the discriminator: Calculating the square of the L1 norm between the real-time state data and the generated sample to obtain the generator reconstruction error; Calculating the square of the L1 norm between the generated sample and the discrimination probability of the real-time state data to obtain the discriminator recognition error; Calculating a first weighted value between the generator reconstruction error and the discriminator recognition error; In the latent space, the basic latent variables are optimized in combination with the single-step random mapping gradient algorithm to obtain optimized latent variables; Using the optimized latent variable to update the basic latent variable, returning the discriminant probability of the generated sample and the real-time state data respectively calculated by the discriminator based on the basic latent variable, until the number of iterations is greater than the preset number of iterations; The basic latent variable that minimizes the first weighted sum value between the generator reconstruction error and the discriminator recognition error is output as the optimal latent variable.

7. The fault detection method of LPCAMM2 according to claim 6, characterized in that: In the latent space, the basic latent variables are optimized in combination with the single-step random mapping gradient algorithm to obtain optimized latent variables, specifically including: Projecting the latent space into a two-dimensional random subspace to obtain low-dimensional latent variables; Based on the local linearity assumption of the generator, optimizing the low-dimensional latent variables once; The low-dimensional latent variable is optimized twice through the discriminator residual to obtain the optimized latent variable.

8. The fault detection method of LPCAMM2 according to claim 1, characterized in that: The S6 specifically includes: Respectively calculating the reconstruction error and the discriminator recognition error under the optimal latent variable; Calculating a second weighted sum value between the reconstruction error and the discriminator recognition error; The second weighted sum value is output as the anomaly score.

9. The fault detection method of LPCAMM2 according to claim 8, characterized in that: The S7 specifically includes: Calculating partial derivatives of the anomaly score with respect to various state parameters in the real-time state data, and using the partial derivatives as contribution values; The state parameter corresponding to the maximum contribution value is output as the abnormal state parameter.

10. A fault detection system for LPCAMM2, characterized in that: include: Processor and memory; The memory stores a program or an instruction executable on the processor, and when the program or the instruction is executed by the processor, the steps of the fault detection method of the LPCAM M2 according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Equipment abnormality detection method and device, computer readable storage medium and chip

    CN113056750A

  • Abnormality detection system, abnormality detection method, abnormality detection program, and method for generating learned model

    US20180365089A1

  • Online fault detection in reram-based ai / ml

    US20220066888A1

  • Machine learning apparatus, abnormality detection apparatus, and abnormality detection method

    US20230022566A1

  • Systems, methods, kits, and apparatuses for specialized chips for robotic intelligence layers

    WO2025050067A1