Failure mode data analysis method and system based on generative language model
Through the method based on the generative language model, the text data and risk dimensions of the failure mode are analyzed, and the risk assessment problem caused by relying on expert experience in the existing technology is solved, and more accurate and objective failure mode sorting is achieved, providing better management decision guidance.
Patent Information
- Application Number
- CN202510213219.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The existing failure mode analysis methods rely on expert experience. Subjective evaluation and simple product sorting lead to one-sided risk assessment. The evaluation results are not instructive to management practices and cannot accurately guide management decisions.
Using a generative language model method, by obtaining multiple failure modes of the product, the initial scores of each failure mode on multiple risk dimensions, and the text data of the product, the text data are split, classified and vectorized, the similarity between the failure modes is analyzed, the scores are predicted, and the failure modes are corrected according to the correction weights, and the failure modes are finally sorted according to the priority.
Reduce dependence on expert experience, improve the objectivity and accuracy of analysis, and can sort the failure mode more refinedly, providing better guidance on product research and development, management and testing.
Smart Images

Figure CN120146031A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and more particularly, to a method and system for analyzing failure mode data based on a generative language model. Background Art
[0002] Failure mode and effects analysis (FMEA) is a risk management method to ensure the safety and reliability of product development. By identifying potential failure modes, evaluating risk factors (occurrence probability O, severity S, detectability D), and calculating the risk priority number (RPN = O × S × D), it is widely used in high-risk fields such as aerospace and medical.
[0003] Traditional FMEA methods rely on experts to evaluate and score risk factors, and then multiply the scores of multiple risk factors to obtain the priority of risk factors.
[0004] The above traditional FMEA methods rely on expert experience, the evaluation is relatively subjective, lacking objective data support, and the calculation method of risk factor priority is single. Different combinations of scores of risk factors may get the same product, making it difficult to distinguish the differentiated response requirements of different risk scenarios, resulting in insufficient guidance of the evaluation results for management practice. Due to subjective evaluation and simple product ranking, the risk assessment is one-sided and cannot accurately guide management decisions. Summary of the Invention
[0005] The problem to be solved by the present invention is that the existing failure mode analysis methods rely on expert experience, resulting in one-sided risk assessment due to subjective evaluation and simple product ranking, insufficient guidance of the evaluation results for management practice, and inability to accurately guide management decisions.
[0006] To solve the above problems, in the first aspect, the present invention provides a method for analyzing failure mode data based on a generative language model, including:
[0007] Obtaining multiple failure modes of a product, initial scores of each failure mode in multiple risk dimensions, and text data of the product;
[0008] Splitting the text data into sentences according to punctuation marks, and classifying the multiple sentences based on the content of each sentence and the interpretations of each failure mode in multiple risk dimensions to obtain the text information corresponding to each failure mode in multiple risk dimensions;
[0009] Using a generative language model to vectorize the text information corresponding to each risk dimension to obtain vectors of each failure mode in multiple risk dimensions;
[0010] Analyze the similarity between multiple failure modes based on the vectors of each failure mode in multiple risk dimensions;
[0011] Based on the similarity and the initial scores of multiple failure modes in multiple risk dimensions, obtain the predicted scores of multiple failure modes in multiple risk dimensions;
[0012] Based on the predicted scores of each failure mode in multiple risk dimensions and the corrected weights of multiple risk dimensions, obtain the corrected scores of multiple failure modes in multiple risk dimensions;
[0013] Based on the corrected scores of multiple failure modes in multiple risk dimensions, obtain the priority of each failure mode;
[0014] Sort multiple failure modes according to the priority to guide the management decision-making of the product.
[0015] Optionally, classify multiple sentences using a generative language model.
[0016] Optionally, the analyzing the similarity between multiple failure modes based on the vectors of each failure mode in multiple risk dimensions includes:
[0017] Arbitrarily select two failure modes, and based on the vector of one failure mode in one risk dimension and the vector of the other failure mode in the same risk dimension, use the cosine similarity algorithm to obtain the similarity between the two failure modes in the same risk dimension;
[0018] Loop through all failure modes for analysis to obtain the similarity between any two failure modes in multiple risk dimensions.
[0019] Optionally, the obtaining the predicted scores of multiple failure modes in multiple risk dimensions based on the similarity and the initial scores of multiple failure modes in multiple risk dimensions includes:
[0020] Select a failure mode, compare the similarities corresponding to the selected failure mode in the current risk dimension, and determine the maximum similarity corresponding to the selected failure mode in the current risk dimension, where the similarity corresponding to the selected failure mode is the similarity between the selected failure mode and other failure modes;
[0021] Based on the similarity corresponding to the selected failure mode in the current risk dimension and the maximum similarity corresponding to the selected failure mode in the current risk dimension, obtain the absolute differences of multiple similarities corresponding to the selected failure mode in the current risk dimension, where the absolute differences of similarities corresponding to all failure modes in the current risk dimension form a set of absolute differences of similarities;
[0022] Taking each absolute difference in the set of absolute differences of similarities corresponding to the current risk dimension as a candidate threshold, and determining multiple screening thresholds based on the multiple candidate thresholds and the maximum similarity corresponding to the failure mode to be predicted, where the multiple candidate thresholds and the multiple screening thresholds are in one-to-one correspondence;
[0023] Based on the similarity between the failure mode to be predicted and other failure modes in the current risk dimension, each screening threshold of the current risk dimension, and the initial scores of the multiple failure modes in the current risk dimension, obtaining the predicted value of the failure mode to be predicted corresponding to each candidate threshold;
[0024] Analyzing in a loop to determine the predicted values of all failure modes corresponding to each candidate threshold in the current risk dimension;
[0025] Based on the predicted values of all failure modes corresponding to each candidate threshold in the current risk dimension and the initial scores of all failure modes in the current risk dimension, obtaining the mean absolute error of all failure modes corresponding to each candidate threshold;
[0026] Comparing the mean absolute errors of all failure modes corresponding to all candidate thresholds, and taking the predicted value corresponding to the minimum mean absolute error as the predicted score of the multiple failure modes in the current risk dimension;
[0027] Analyzing in a loop to determine the predicted scores of the multiple failure modes in each risk dimension, where the risk dimensions include occurrence probability, occurrence severity, detectability, and management cost.
[0028] Optionally, the obtaining the predicted value of the failure mode to be predicted corresponding to each candidate threshold based on the similarity between the failure mode to be predicted and other failure modes in the current risk dimension, each screening threshold of the current risk dimension, and the initial scores of the multiple failure modes in the current risk dimension includes:
[0029] Based on the similarity between the failure mode to be predicted and other failure modes in the current risk dimension and each screening threshold of the current risk dimension, determining the similar failure modes corresponding to the failure mode to be predicted under each screening threshold, where the other failure modes are the other failure modes except the current failure mode among the multiple failure modes;
[0030] Determining the similarity weight corresponding to each similar failure mode according to the similarity of the similar failure modes corresponding to each screening threshold;
[0031] Based on the similarity weight and the initial scores of each similar failure mode in the current risk dimension, obtaining the predicted value of the failure mode to be predicted corresponding to each candidate threshold.
[0032] Optionally, obtaining the corrected scores of multiple failure modes on multiple risk dimensions according to the predicted scores of each failure mode on multiple risk dimensions and the corrected weights of multiple risk dimensions includes:
[0033] Normalize the predicted scores of multiple failure modes on multiple risk dimensions to obtain standardized scores;
[0034] Determine the variability of each risk dimension according to the standardized scores on each risk dimension;
[0035] Determine the corrected weight of each risk dimension according to the variability of multiple risk dimensions;
[0036] Obtain the corrected scores of multiple failure modes on multiple risk dimensions according to the corrected weights and standardized scores.
[0037] Optionally, the variability is:
[0038]
[0039] where G j represents the variability of the jth risk dimension, h nj represents the standardized score of the nth failure mode on the jth risk dimension, and q represents the total number of failure modes;
[0040] The corrected weight is:
[0041]
[0042] where t j represents the corrected weight of the jth risk dimension, and G j represents the variability of the jth risk dimension;
[0043] The corrected score is:
[0044] U nj =t j h nj ,
[0045] where U nj represents the corrected score of the nth failure mode on the jth risk dimension, and h nj represents the standardized score of the nth failure mode on the jth risk dimension.
[0046] Optionally, obtaining the priority of each failure mode according to the corrected scores of multiple failure modes on multiple risk dimensions includes:
[0047] Determine the maximum corrected score and the minimum corrected score on each risk dimension according to the corrected scores of multiple failure modes on multiple risk dimensions;
[0048] Determine the priority of each failure mode according to the maximum correction score and the minimum correction score on each risk dimension and the correction scores of each failure mode on multiple risk dimensions.
[0049] Optionally, the priority is:
[0050]
[0051] where l n represents the priority of the nth failure mode, a represents the category of the maximum and minimum values of the correction scores on the risk dimension, and U nj represents the correction score of the nth failure mode on the jth risk dimension, represents the maximum correction score on the jth risk dimension, represents the minimum correction score on the jth risk dimension.
[0052] In a second aspect, the present invention also provides a failure mode data analysis system based on a generative language model, including:
[0053] A data acquisition module for acquiring multiple failure modes of a product, the initial scores of each failure mode on multiple risk dimensions, and the text data of the product;
[0054] An information classification module for splitting the text data into sentences according to punctuation marks, and classifying the multiple sentences according to the content of each sentence and the paraphrases of each failure mode on multiple risk dimensions, so as to obtain the text information corresponding to each failure mode on multiple risk dimensions;
[0055] An information vectorization module for vectorizing the text information corresponding to each risk dimension by using a generative language model to obtain vectors of each failure mode on multiple risk dimensions;
[0056] A similarity analysis module for analyzing the similarity between multiple failure modes according to the vectors of each failure mode on multiple risk dimensions;
[0057] A score prediction module for obtaining the predicted scores of multiple failure modes on multiple risk dimensions according to the similarity and the initial scores of multiple failure modes on multiple risk dimensions;
[0058] A score correction module for obtaining the correction scores of multiple failure modes on multiple risk dimensions according to the predicted scores of each failure mode on multiple risk dimensions and the correction weights of multiple risk dimensions;
[0059] A priority analysis module for obtaining the priority of each failure mode according to the correction scores of multiple failure modes on multiple risk dimensions;
[0060] A failure mode sorting module for sorting multiple failure modes according to priority to guide the management decision-making of products.
[0061] The present invention provides a method and system for analyzing failure mode data based on a generative language model. Compared with the prior art, it has the following beneficial effects:
[0062] The text data of the product obtained can be split by using a generative language model, etc., and combined with the interpretations of each failure mode in multiple risk dimensions to classify the sentences in the text data, and the text information corresponding to each failure mode in multiple risk dimensions can be obtained, avoiding the dependence on expert experience in the failure mode recognition stage; then use the generative language model to vectorize the text information corresponding to each risk dimension, analyze the similarity between multiple failure modes according to the vectors of each failure mode in multiple risk dimensions, and similar failure modes to the failure mode to be predicted can be screened out according to the similarity, and then use the initial scores of multiple failure modes in multiple risk dimensions to obtain the predicted scores of multiple failure modes in multiple risk dimensions one by one; further consider the reliability and importance degrees of multiple risk dimensions, and use the corrected weights of multiple risk dimensions to correct the predicted scores to obtain the corrected scores of multiple failure modes in multiple risk dimensions; based on the corrected scores, obtain the priority of each failure mode, and sort multiple failure modes according to the priority. The sorting result weakens the dependence on expert experience and considers the reliability of multiple risk dimensions. Even if the final priorities are the same, the sorting can be further refined according to the reliability of the risk dimension and the score of the corresponding dimension, and the failure modes can be sorted in a refined manner, which has good guiding significance for the research and development, management, and testing of products. Description of the Drawings
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0064] Figure 1 It is a schematic flowchart of a method for analyzing failure mode data based on a generative language model provided by an embodiment of the present invention;
[0065] Figure 2 It is a simplified execution flowchart of a method for analyzing failure mode data based on a generative language model provided by an embodiment of the present invention;
[0066] Figure 3Schematic diagram of a failure mode data analysis system based on a generative language model provided by an embodiment of the present invention. Detailed implementation manners
[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0068] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0069] Generative language model: It is a language model whose core function is to generate new and coherent text content based on the input text content. It learns the patterns, structures, and semantics of the language from the given text data and can generate natural language text according to the given prompts or contexts. Through training or adjustment, it can be used for text classification tasks; for example, a generative model using the Transformer architecture can be used for text classification by adding a classifier to the output layer of the model. It can also extract the vector representation of the text through its internal embedding layer to vectorize the text.
[0070] As Figure 1 shown, a failure mode data analysis method based on a generative language model provided by an embodiment of the present application includes:
[0071] S1: Obtain multiple failure modes of the product, the initial scores of each failure mode on multiple risk dimensions, and the text data of the product.
[0072] Specifically, data is mainly collected from various sources such as expert opinions, historical records, and technical documents related to the development of the product to be analyzed. Then, this data is processed and organized into a structured format to form a list of typical potential failure modes in the development process of a certain product. This failure mode list includes descriptions of each failure mode (i.e., the paraphrases of the failure mode on multiple risk dimensions), which is compiled by the relevant project departments according to historical experience and management priorities, and initial expert scoring is performed on each failure mode (abbreviated as FM) as the initial score to distinguish it from the subsequent predicted scores. The text data includes text data related to the development process of the product (similar products, existing batches, etc.), such as various expert opinions, meeting records, technical documents, etc. The risk dimensions of each failure mode include occurrence probability, occurrence severity, detectability, and management cost.
[0073] S2: Split the text data into sentences according to punctuation marks. Based on the content of each sentence and the interpretations of each failure mode in multiple risk dimensions, classify multiple sentences to obtain the text information corresponding to each failure mode in multiple risk dimensions. A generative language model can be used to split and classify the text data.
[0074] S3: Use a generative language model to vectorize the text information corresponding to each risk dimension to obtain the vectors of each failure mode in multiple risk dimensions.
[0075] S4: Analyze the similarity between multiple failure modes based on the vectors of each failure mode in multiple risk dimensions. The similarity between two text vectors can be analyzed by calculating the cosine similarity, Euclidean distance, Manhattan distance, etc. between the two vectors.
[0076] S5: Obtain the predicted scores of multiple failure modes in multiple risk dimensions based on the similarity and the initial scores of multiple failure modes in multiple risk dimensions.
[0077] Specifically, by analyzing the text similarity of multiple failure modes in multiple risk dimensions, it is possible to know the failure modes with similar text descriptions. If two failure modes have similar or close descriptions in a certain risk dimension, it indicates that the two failure modes are related and even have mutual triggering properties. Therefore, the evaluation scores of the two failure modes in this risk dimension should be similar or close. According to this characteristic, it is possible to first analyze the text similarity of multiple failure modes in multiple risk dimensions, screen out the failure modes similar to the failure mode to be predicted through similarity, and then use the initial scores of these similar failure modes to predict the score of the failure mode to be predicted, rather than directly using the initial score of the failure mode to be predicted for evaluation and ranking.
[0078] S6: Obtain the corrected scores of multiple failure modes in multiple risk dimensions based on the predicted scores of each failure mode in multiple risk dimensions and the corrected weights of multiple risk dimensions.
[0079] Specifically, although the predicted scores of multiple failure modes have been calculated, there is no correlation between the scores of these failure modes in each risk dimension, that is, there is no distinction between the importance and urgency of each risk dimension. The corrected weights corresponding to multiple risk dimensions can be assigned according to the importance of each risk dimension or the reliability of each risk dimension, so as to correct the predicted scores, make the scores of failure modes in multiple risk dimensions have a bias, and the obtained corrected scores are more instructive.
[0080] S7: Obtain the priority of each failure mode based on the corrected scores of multiple failure modes in multiple risk dimensions.
[0081] S8: Sort multiple failure modes according to their priorities to guide the management decision-making of the product.
[0082] In this embodiment, a generative language model or the like can be used to split the obtained text data of the product (such as text data in the product development process, data in the product production process, data in the test process, etc.), and in combination with the interpretations of each failure mode in multiple risk dimensions, classify the sentences in the text data to obtain the text information corresponding to each failure mode in multiple risk dimensions, avoiding the dependence on expert experience in the failure mode identification stage; then use the generative language model to vectorize the text information corresponding to each risk dimension, analyze the similarity between multiple failure modes according to the vectors of each failure mode in multiple risk dimensions, and according to the similarity, filter out the failure modes similar to the to-be-predicted failure mode, and then use the initial scores of multiple failure modes in multiple risk dimensions to obtain the predicted scores of multiple failure modes in multiple risk dimensions one by one; further consider the reliability and importance levels of multiple risk dimensions, use the corrected weights of multiple risk dimensions to correct the predicted scores, and obtain the corrected scores of multiple failure modes in multiple risk dimensions; based on the corrected scores, obtain the priority of each failure mode, and sort multiple failure modes according to the priority. The sorting result weakens the dependence on expert experience and considers the reliability of multiple risk dimensions. Even if the final priorities are the same, the sorting can be further refined according to the reliability of the risk dimension and the score of the corresponding dimension, enabling a refined sorting of failure modes, which has good guiding significance for product R & D, management, and testing, etc.
[0083] Combined with Figure 2 , each step will be described in detail below.
[0084] S1: Obtain multiple failure modes of the product, the initial scores of each failure mode in multiple risk dimensions, and the text data of the product.
[0085] Specifically, there are 4 types of risk dimensions, namely occurrence probability (O), occurrence severity (S), detectability (D), and management cost (M). Specifically: O represents the frequency of occurrence of problems related to the failure mode; S represents the degree of impact that may be brought once the failure mode occurs; D represents the possibility of being observed and detected in actual management of the failure mode; M represents the cost and feasibility of managing the failure mode once it occurs.
[0086] S2: Split the text data into sentences according to punctuation marks, and classify the multiple sentences according to the content of each sentence and the interpretations of each failure mode in multiple risk dimensions to obtain the text information corresponding to each failure mode in multiple risk dimensions.
[0087] Specifically, first, split all the collected failure mode text information into logically complete sentences according to punctuation marks. Second, pre-train the generative language model. In the input of the generative language model, first give several examples as the execution reference of the generative language model. Specifically, we need to input three types of information into the generative language model: A. The task objective of executing text information (that is, classify the text information of failure modes according to the four categories of O, S, D, and M, and at the same time give the specific connotations of O, S, D, and M). B. All logically complete sentences. C. Given text information classification examples. Finally, the generative language model will output the text information of n failure modes in different risk dimensions according to the task setting. The generative language model is trained using the few-shot learning method, which helps to provide better classification results, has strong reasoning ability, and is more objective and accurate than manual classification operations. Finally, use the generative language model to map the logically complete sentences to different risk dimensions (O, S, D, M) of the failure mode to classify the risk information of the failure mode. Using the generative language model to split and classify text data can avoid relying on experts' recognition experience and classification experience, making the recognition and classification more objective and improving the classification efficiency.
[0088] S3: Use the generative language model to vectorize the text information corresponding to each risk dimension to obtain the vectors of each failure mode in multiple risk dimensions.
[0089] Specifically, use the generative language model to vectorize the text information, and vectorize the written descriptions of the occurrence probability (O), occurrence severity (S), detectability (D), and management cost (M) of the n failure modes FM. Considering the operability of data processing, we average all the word vectors (token vectors) in a sentence or text segment to obtain a vector representing the entire text. Finally, output the vectorization results of n failure modes FM in 4 risk dimensions.
[0090] S4: Analyze the similarity between multiple failure modes according to the vectors of each failure mode in multiple risk dimensions. Specifically include:
[0091] S410: Arbitrarily select two failure modes, and use the cosine similarity algorithm to obtain the similarity between the two failure modes in the same risk dimension according to the vector of one failure mode in one risk dimension and the vector of the other failure mode in the same risk dimension.
[0092] S420: Analyze all failure modes in a loop to obtain the similarity between any two failure modes in multiple risk dimensions.
[0093] For two failure modes with vectors on the same risk dimension as (x 11 , x 12 , x 13 ,..., x 1k ) and (x 21 , x 22 , x 23 ,..., x 2k ), the similarity calculation formula is as follows:
[0094]
[0095] where x 1p represents the p-th element of the vector of a failure mode on a risk dimension, and x 2p represents the p-th element of the vector of another failure mode on the same risk dimension, and k represents the total number of elements of the vector corresponding to the risk dimension.
[0096] For example, if it is necessary to predict the occurrence severity (S) of failure mode FM1, we first vectorize the descriptions of the occurrence severity (S) of failure mode FM1 and the other n - 1 failure modes, and then calculate the similarity between the vector of the occurrence severity (S) of failure mode FM1 and the vector of the occurrence severity (S) of each of the other n - 1 failure modes.
[0097] S5: Obtain the predicted scores of multiple failure modes on multiple risk dimensions based on the similarity and the initial scores of multiple failure modes on multiple risk dimensions. Specifically, it includes:
[0098] S510: Select a failure mode, compare the similarities corresponding to the selected failure mode on the current risk dimension, and determine the maximum similarity corresponding to the selected failure mode on the current risk dimension. Among them, the similarity corresponding to the selected failure mode is the similarity between the selected failure mode and other failure modes, and other failure modes are each of the other failure modes except the selected failure mode among all failure modes. For example, the similarities between FM1 and other failure modes (FM2, FM3, and FM4) are 0.6, 0.7, and 0.4 in sequence.
[0099] S520: Obtain multiple absolute differences of similarities corresponding to the selected failure mode on the current risk dimension based on the similarity corresponding to the selected failure mode on the current risk dimension and the maximum similarity corresponding to the selected failure mode on the current risk dimension. Among them, the set of absolute differences of similarities corresponding to all failure modes on the current risk dimension constitutes the set of absolute differences of similarities.
[0100] Suppose there are four failure modes FM1, FM2, FM3 and FM4, and the severity of occurrence (S) is selected as the risk dimension for the current analysis. First, select FM1. Suppose the similarities between FM1 and the other three failure modes are 0.6, 0.7 and 0.4 in sequence. Then the maximum similarity corresponding to FM1 is 0.7. Subtract each similarity from the maximum similarity to obtain the absolute difference of similarity corresponding to FM1, which are 0.1, 0 and 0.3 respectively. Then select FM2. Suppose the similarities between FM2 and the other three failure modes are 0.6, 0.5 and 0.8 in sequence. It can be seen that the maximum similarity corresponding to FM2 is 0.8. Then the absolute differences of similarity corresponding to FM2 are 0.2, 0.3 and 0 respectively. Similarly, suppose the similarities between FM3 and the other three failure modes are 0.7, 0.5 and 0.9 in sequence. It can be seen that the maximum similarity corresponding to FM3 is 0.9. Then the absolute differences of similarity corresponding to FM3 are 0.2, 0.4 and 0 respectively. Suppose the similarities between FM4 and the other three failure modes are 0.4, 0.8 and 0.9 in sequence. It can be seen that the maximum similarity corresponding to FM4 is 0.9. Then the absolute differences of similarity corresponding to FM4 are 0.5, 0.1 and 0 respectively. As a result, 12 absolute difference results are obtained. The repeated absolute differences can be merged and omitted in the subsequent calculation process. All the obtained absolute differences of similarity form a set of absolute differences of similarity, such as (0.1, 0, 0.3, 0.2, 0.4, 0.5). In addition, it can be seen from the above that the maximum similarity corresponding to each failure mode may be different.
[0101] S530: Take each absolute difference in the set of absolute differences of similarity corresponding to the current risk dimension as a candidate threshold, and determine multiple screening thresholds according to the multiple candidate thresholds and the maximum similarity corresponding to the failure mode to be predicted, where the multiple candidate thresholds and the multiple screening thresholds correspond one by one.
[0102] Specifically, subtract the candidate threshold from the maximum similarity corresponding to the failure mode to be predicted to obtain the screening threshold, and each candidate threshold corresponds to a screening threshold. For example, to predict the severity of occurrence S of the failure mode FM1, the failure mode to be predicted is FM1, the candidate threshold is 0.2, and the maximum similarity corresponding to FM1 is 0.7. Then the screening threshold is 0.5. Similarly, each candidate threshold can be calculated to obtain multiple screening thresholds.
[0103] S540: Obtain the predicted value of the failure mode to be predicted corresponding to each candidate threshold according to the similarity between the failure mode to be predicted and other failure modes in the current risk dimension, each screening threshold of the current risk dimension, and the initial scores of multiple failure modes in the current risk dimension.
[0104] Specifically, for each screening threshold, failure modes similar to the failure mode to be predicted can be screened out, and the average value of the initial scores of all similar failure modes is calculated to obtain the predicted value of the failure mode to be predicted. Similarly, the predicted values of other failure modes are obtained. Finally, the predicted value of each failure mode is subtracted from the initial score (true value) to obtain the total error of all failure modes, and the predicted value corresponding to the minimum total error is used as the final predicted score.
[0105] In another embodiment, S540 specifically includes:
[0106] S541: Determine the similar failure modes corresponding to the failure mode to be predicted under each screening threshold according to the similarity between the failure mode to be predicted and other failure modes in the current risk dimension and each screening threshold of the current risk dimension, where the other failure modes are the other failure modes except the current failure mode among multiple failure modes.
[0107] For example, among the similarities between failure mode FM1 and other failure modes, the failure modes corresponding to the similarities with a similarity greater than or equal to the screening threshold are screened out. These failure modes are the similar failure modes of failure mode FM1. According to the above example, when the candidate threshold is 0.2 and the screening threshold is 0.5, FM2 and FM3 are screened as the similar failure modes of FM1.
[0108] S542: Determine the similarity weight corresponding to each similar failure mode according to the similarity of the similar failure modes corresponding to each screening threshold. The similarity weight of a similar failure mode is the similarity of the similar failure mode divided by the sum of the similarities of all failure modes.
[0109] S543: Obtain the predicted value of the failure mode to be predicted corresponding to each screening threshold according to the similarity weight and the initial score of each similar failure mode in the current risk dimension. Since the screening threshold corresponds one-to-one with the candidate threshold, the predicted value of the failure mode to be predicted corresponding to each candidate threshold is obtained.
[0110] The occurrence severity (S) of failure mode FM1 is the weighted sum of the initial scores of the occurrence severities (S) of the similar failure modes of FM1, where the similarity weight is proportional to their similarity scores. Assume that the initial scores of the occurrence severities corresponding to FM1, FM2, FM3, and FM4 are 4, 3, 6, and 8. Then, when the candidate threshold is 0.2, the predicted value of FM1 in the occurrence severity dimension is 0.6×3 / (0.6 + 0.7) + 0.7×6 / (0.6 + 0.7) = 4.62.
[0111] S550: The loop analysis determines the predicted values of all failure modes corresponding to each candidate threshold on the current risk dimension. Similarly, similar failure modes of the remaining failure modes can be screened, and then the initial scores of the similar failure modes are weighted and summed to obtain the predicted values of each failure mode on each risk dimension. This method can weight and sum the initial scores evaluated by experts without directly using the initial scores, reducing the errors caused by the subjectivity of expert scoring and avoiding misleading the final ranking due to personal subjectivity.
[0112] S560: Based on the predicted values of all failure modes corresponding to each candidate threshold on the current risk dimension and the initial scores of all failure modes on the current risk dimension, the mean absolute error of all failure modes corresponding to each candidate threshold is obtained.
[0113] For example, when the candidate threshold is selected as 0.2, the absolute error of FM1 in the occurrence severity dimension is |4.62 - 4| = 0.62. Similarly, when the candidate threshold is selected as 0.2, the absolute errors of FM2, FM3, and FM4 are calculated, assumed to be 0.2, 0.1, and 0.5. Then the mean absolute error of all failure modes corresponding to the candidate threshold 0.2 is (0.62 + 0.2 + 0.1 + 0.5) / 4 = 0.355.
[0114] S570: Compare the mean absolute errors of all failure modes corresponding to all candidate thresholds, and take the predicted value corresponding to the minimum mean absolute error as the predicted score of multiple failure modes on the current risk dimension.
[0115] S580: The loop analysis determines the predicted scores of multiple failure modes on each risk dimension, where the risk dimensions include occurrence probability, occurrence severity, detectability, and management cost.
[0116] Specifically, expert scoring has certain reference value, but the scoring of individual experts may also involve too much personal subjectivity. Therefore, instead of directly using expert scoring for analysis and ranking, expert scoring is weighted and summed to reduce the influence of the personal subjectivity of expert scoring. And when considering risk dimensions, management cost is added to the evaluation system, which can provide more operable guidance for risk management activities in the complex product development process and is conducive to comprehensively managing risk activities and cost allocation in the product development process.
[0117] S6: Based on the predicted scores of each failure mode on multiple risk dimensions and the corrected weights of multiple risk dimensions, the corrected scores of multiple failure modes on multiple risk dimensions are obtained. Specifically, it includes the following steps:
[0118] S610: Normalize the prediction scores of multiple failure modes across multiple risk dimensions to obtain standardized scores, forming the following standardized score matrix H. Specifically, divide the corrected score of a failure mode in a risk dimension by the square root of the sum of the squares of the corrected scores of all failure modes in the same risk dimension to obtain the normalized result (i.e., the standardized score) of this failure mode in this risk dimension.
[0119]
[0120] Among them, h nj represents the standardized score of the nth failure mode in the jth risk dimension, and q represents the total number of failure modes. In this embodiment, the matrix H has four dimensions, namely occurrence probability (O), occurrence severity (S), detectability (D), and management cost (M).
[0121] S620: Determine the variability of each risk dimension based on the standardized scores of each risk dimension.
[0122] The variability is:
[0123]
[0124] Among them, G j represents the variability of the jth risk dimension. A smaller variability means that the jth risk dimension can provide more certain and reliable information. Therefore, a larger weight should be assigned to the jth risk dimension in the weight allocation; h nj represents the standardized score of the nth failure mode in the jth risk dimension, and q represents the total number of failure modes.
[0125] S630: Determine the corrected weight of each risk dimension based on the variability of multiple risk dimensions.
[0126] The corrected weight is:
[0127]
[0128] Among them, t j represents the corrected weight of the jth risk dimension, and G j represents the variability of the jth risk dimension.
[0129] The correction weight is inversely proportional to the variability, that is, the greater the variability, the smaller the correction weight; conversely, the greater the correction weight. A greater correction weight is assigned to the risk dimension with more reliable information. Since the information of some risk dimensions may not be accurate, if the scores obtained from accurate information and inaccurate information are regarded as the same, the scores obtained from inaccurate information will mislead the results. To reduce the interference of uncertain information, a smaller weight can be assigned to the scores of these risk dimensions during correction, reducing their proportion in the final calculation of the priority. Thus, the accuracy of the final ranking is ensured, and the value of the final ranking is improved.
[0130] S640: Obtain the corrected scores of multiple failure modes on multiple risk dimensions according to the correction weight and the standardized scores.
[0131] The corrected score is:
[0132] U nj =t j h nj ,
[0133] where u nj represents the corrected score of the nth failure mode on the jth risk dimension, and h nj represents the standardized score of the nth failure mode on the jth risk dimension.
[0134] S7: Obtain the priority of each failure mode according to the corrected scores of multiple failure modes on multiple risk dimensions. Specifically, it includes:
[0135] Determine the maximum corrected score and the minimum corrected score on each risk dimension according to the corrected scores of multiple failure modes on multiple risk dimensions.
[0136] Determine the priority of each failure mode according to the maximum corrected score and the minimum corrected score on each risk dimension and the corrected scores of each failure mode on multiple risk dimensions.
[0137] The priority is:
[0138]
[0139] where l n represents the priority of the nth failure mode, a represents the category of the maximum value of the corrected scores on the risk dimension, represents the maximum corrected score on the jth risk dimension, represents the minimum corrected score on the jth risk dimension, and U nj represents the corrected score of the nth failure mode on the jth risk dimension.
[0140] S8: Sort multiple failure modes according to their priorities to guide the management decisions of the product.
[0141] Specifically, sort the priorities in descending order. If the priorities of two failure modes are equal, compare the prediction scores corresponding to the risk dimensions with the largest revised weights, and place the one with the larger prediction score in the front. If they are still equal, compare the prediction scores on the risk dimensions corresponding to the revised weights in descending order of the revised weights one by one until the size relationship between the two is determined.
[0142] In summary, compared with the prior art, the following beneficial effects are achieved:
[0143] 1. In this application, first use the language model to split and classify the text data of the product. In the failure mode recognition stage, it does not rely on the experience of experts, reduces the introduction of expert personal experience, and improves the classification accuracy, objectivity and classification efficiency. Then analyze the similarity of the text information of each failure mode in each risk dimension, and then screen the similar failure modes according to the similarity. Determine the optimal prediction value as the prediction score based on the size of the mean absolute error between the prediction value and the true value. Instead of directly using the initial score evaluated by experts, the prediction score is obtained based on the text vector similarity and combined with the initial scores of multiple similar failure modes, avoiding the adverse effects caused by misevaluation by experts or insufficient expert experience on a certain failure mode. Then calculate the variability of each failure mode in each risk dimension to obtain the revised weight, and assign different revised weights to different risk dimensions, which can show the primary and secondary relationships of multiple risk dimensions, improve the accuracy and reference of the final priority, and is also conducive to refining the final ranking. The method of this application can be applied to the data analysis in the product development stage, and can also be applied to the data analysis in stages such as product testing, product improvement, product production and qualification inspection. This method is not only applicable to the analysis of the data of simple products, but can also be applied to the data analysis of complex products, which can better reflect the advantages of this method, can analyze the priorities of a large number of failure modes of complex products, is conducive to refining the development focus of complex products, and has great guiding and promoting significance for the development process.
[0144] 2. Using the generative language model for the classification and vectorization of text data can more accurately understand the text content and semantics, helps to improve the use effect of objective data, provides better classification and vectorization results, and can quickly set instructions and tasks, and can avoid the disadvantages of over-reliance on a large amount of training data and complex model operation of the deep learning algorithm.
[0145] 3. By increasing the measurement and evaluation of the failure mode management cost, it can provide more operable guidance for risk management activities in the complex product development process, facilitate the comprehensive management of risk activities and cost allocation in the product development process, and also give full play to the advantages of the new generation of information technology.
[0146] As Figure 3 shown, a failure mode data analysis system based on a generative language model provided by an embodiment of the present application includes:
[0147] A data acquisition module 100, configured to acquire multiple failure modes of a product, initial scores of each failure mode in multiple risk dimensions, and text data of the product.
[0148] An information classification module 200, configured to split the text data into sentences according to punctuation marks, and classify the multiple sentences according to the content of each sentence and the paraphrases of each failure mode in multiple risk dimensions, so as to obtain the text information corresponding to each failure mode in multiple risk dimensions.
[0149] An information vectorization module 300, configured to vectorize the text information corresponding to each risk dimension by using a generative language model, so as to obtain vectors of each failure mode in multiple risk dimensions.
[0150] A similarity analysis module 400, configured to analyze the similarity between multiple failure modes according to the vectors of each failure mode in multiple risk dimensions.
[0151] A score prediction module 500, configured to obtain predicted scores of multiple failure modes in multiple risk dimensions according to the similarity and the initial scores of multiple failure modes in multiple risk dimensions.
[0152] A score correction module 600, configured to obtain corrected scores of multiple failure modes in multiple risk dimensions according to the predicted scores of each failure mode in multiple risk dimensions and the correction weights of multiple risk dimensions.
[0153] A priority analysis module 700, configured to obtain the priority of each failure mode according to the corrected scores of multiple failure modes in multiple risk dimensions.
[0154] A failure mode sorting module 800, configured to sort multiple failure modes according to the priority to guide the management decision-making of the product.
[0155] In this embodiment, the beneficial effects of a failure mode data analysis system based on a generative language model are similar to those of the above-mentioned failure mode data analysis method based on a generative language model, and will not be elaborated here.
[0156] An electronic device provided by an embodiment of the present application includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement the failure mode data analysis method based on the generative language model as described above when executing the computer program.
[0157] A computer-readable storage medium provided by an embodiment of the present application has a computer program stored thereon, and when the computer program is executed by a processor, the failure mode data analysis method based on the generative language model as described above is implemented.
[0158] In this embodiment, the beneficial effects of the electronic device and the computer-readable storage medium are similar to those of the failure mode data analysis method based on the generative language model as described above, and will not be elaborated here.
[0159] Now, an electronic device that can be used as a server or a client of the present application will be described. It is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described herein and / or claimed.
[0160] The electronic device includes a computing unit, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) or the computer program loaded from the storage unit into the random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The computing unit, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.
[0161] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc. In this application, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. One can select some or all of the units according to actual needs to achieve the purpose of the solution of the embodiments of this application. In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0162] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0163] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A failure mode data analysis method based on a generative language model, characterized in that: include: Obtain multiple failure modes of the product, initial scores of each failure mode in multiple risk dimensions, and text data of the product; Split the text data into sentences according to punctuation marks, classify the sentences according to the content of each sentence and the interpretation of each failure mode in multiple risk dimensions, and obtain the text information corresponding to each failure mode in multiple risk dimensions; The generative language model is used to vectorize the text information corresponding to each risk dimension to obtain the vector of each failure mode in multiple risk dimensions. Analyze the similarity between multiple failure modes based on the vectors of each failure mode in multiple risk dimensions; According to the similarity and the initial scores of the multiple failure modes in the multiple risk dimensions, the prediction scores of the multiple failure modes in the multiple risk dimensions are obtained; According to the prediction score of each failure mode in multiple risk dimensions and the correction weights of the multiple risk dimensions, the correction scores of the multiple failure modes in multiple risk dimensions are obtained; According to the modified scores of multiple failure modes in multiple risk dimensions, the priority of each failure mode is obtained; Rank multiple failure modes according to priority to guide product management decisions.
2. The failure mode data analysis method based on a generative language model according to claim 1, characterized in that: Classify multiple sentences using a generative language model.
3. The failure mode data analysis method based on a generative language model according to claim 1, characterized in that: Analyzing the similarity between multiple failure modes according to the vectors of each failure mode on multiple risk dimensions includes: Select any two failure modes, and use the cosine similarity algorithm to obtain the similarity of the two failure modes on the same risk dimension based on the vector of one failure mode on a risk dimension and the vector of another failure mode on the same risk dimension; All failure modes are analyzed cyclically to obtain the similarity between any two failure modes in multiple risk dimensions.
4. The failure mode data analysis method based on a generative language model according to claim 1, characterized in that: The step of obtaining predicted scores of the multiple failure modes in the multiple risk dimensions according to the similarity and the initial scores of the multiple failure modes in the multiple risk dimensions includes: A failure mode is selected, and similarities corresponding to the selected failure mode on the current risk dimension are compared to determine the maximum similarity corresponding to the selected failure mode on the current risk dimension, wherein the similarity corresponding to the selected failure mode is the similarity between the selected failure mode and other failure modes; According to the similarity corresponding to the failure mode selected on the current risk dimension and the maximum similarity corresponding to the failure mode selected on the current risk dimension, a plurality of similarity absolute differences corresponding to the failure mode selected on the current risk dimension are obtained, wherein the similarity absolute differences corresponding to all failure modes on the current risk dimension constitute a similarity absolute difference set; Taking each absolute difference in the set of absolute differences of similarity corresponding to the current risk dimension as a candidate threshold, and determining multiple screening thresholds according to the maximum similarities corresponding to the multiple candidate thresholds and the failure mode to be predicted, wherein the multiple candidate thresholds correspond to the multiple screening thresholds one by one; According to the similarity between the failure mode to be predicted and other failure modes in the current risk dimension, each screening threshold of the current risk dimension and the initial scores of multiple failure modes in the current risk dimension, the predicted value of the failure mode to be predicted corresponding to each candidate threshold is obtained; The predicted values of all failure modes corresponding to each candidate threshold in the current risk dimension are determined by cyclic analysis; According to the predicted values of all failure modes corresponding to each candidate threshold in the current risk dimension and the initial scores of all failure modes in the current risk dimension, the mean absolute errors of all failure modes corresponding to each candidate threshold are obtained; Compare the mean absolute errors of all failure modes corresponding to all candidate thresholds, and use the predicted value corresponding to the minimum mean absolute error as the predicted score of multiple failure modes in the current risk dimension; The cyclic analysis determines a prediction score for a plurality of failure modes on each risk dimension, wherein the risk dimensions include probability of occurrence, severity of occurrence, detectability, and management cost.
5. The failure mode data analysis method based on a generative language model according to claim 4, characterized in that: The method of obtaining the predicted value of the failure mode to be predicted corresponding to each candidate threshold value according to the similarity between the failure mode to be predicted and other failure modes in the current risk dimension, each screening threshold of the current risk dimension and the initial scores of multiple failure modes in the current risk dimension includes: According to the similarity between the failure mode to be predicted and other failure modes in the current risk dimension and each screening threshold of the current risk dimension, determine the similar failure modes corresponding to the failure mode to be predicted under each screening threshold, wherein the other failure modes are other failure modes among the multiple failure modes except the current failure mode; According to the similarity of similar failure modes corresponding to each screening threshold, a similarity weight corresponding to each similar failure mode is determined; According to the similarity weight and the initial score of each similar failure mode in the current risk dimension, the predicted value of the failure mode to be predicted corresponding to each candidate threshold is obtained.
6. The failure mode data analysis method based on a generative language model according to claim 1, characterized in that: The step of obtaining the correction scores of the multiple failure modes in the multiple risk dimensions according to the prediction scores of each failure mode in the multiple risk dimensions and the correction weights of the multiple risk dimensions includes: Normalize the prediction scores of multiple failure modes in multiple risk dimensions to obtain standardized scores; Based on the standardized scores on each risk dimension, the variability of each risk dimension was determined; Determine the correction weight of each risk dimension based on the variability of multiple risk dimensions; According to the modified weights and standardized scores, modified scores of multiple failure modes in multiple risk dimensions are obtained.
7. The failure mode data analysis method based on a generative language model according to claim 6, characterized in that: The variability is: Among them, G j represents the variability of the j-th risk dimension, h nj represents the standardized score of the nth failure mode on the jth risk dimension, and q represents the total number of failure modes; The correction weight is: Among them, t j represents the modified weight of the j-th risk dimension, G j represents the variability of the j-th risk dimension; The modified score is: U nj =t j h nj , Among them, U nj represents the correction score of the nth failure mode in the jth risk dimension, h nj Represents the standardized score of the nth failure mode on the jth risk dimension.
8. The failure mode data analysis method based on a generative language model according to claim 1, characterized in that: The step of obtaining the priority of each failure mode according to the correction scores of the multiple failure modes in the multiple risk dimensions includes: According to the correction scores of the multiple failure modes in the multiple risk dimensions, a maximum correction score and a minimum correction score in each risk dimension are determined; The priority of each failure mode is determined based on the maximum and minimum correction scores on each risk dimension and the correction scores of each failure mode on multiple risk dimensions.
9. The failure mode data analysis method based on a generative language model according to claim 8, characterized in that: The priorities are: Among them, l n represents the priority of the nth failure mode, a represents the maximum value category of the correction score on the risk dimension, and U nj represents the correction score of the nth failure mode in the jth risk dimension, represents the maximum correction score on the j-th risk dimension, Represents the minimum correction score on the j-th risk dimension.
10. A failure mode data analysis system based on a generative language model, characterized in that: include: A data acquisition module, used to acquire multiple failure modes of a product, initial scores of each failure mode in multiple risk dimensions, and text data of the product; An information classification module is used to split the text data into sentences according to punctuation marks, classify multiple sentences according to the content of each sentence and the interpretation of each failure mode in multiple risk dimensions, and obtain the text information corresponding to each failure mode in multiple risk dimensions; An information vectorization module is used to vectorize the text information corresponding to each risk dimension using a generative language model to obtain the vector of each failure mode in multiple risk dimensions; A similarity analysis module is used to analyze the similarity between multiple failure modes based on the vectors of each failure mode in multiple risk dimensions; A score prediction module, used for obtaining prediction scores of multiple failure modes in multiple risk dimensions according to similarities and initial scores of multiple failure modes in multiple risk dimensions; A score correction module, used to obtain correction scores of multiple failure modes in multiple risk dimensions according to the predicted scores of each failure mode in multiple risk dimensions and the correction weights of the multiple risk dimensions; A priority analysis module is used to obtain the priority of each failure mode according to the correction scores of multiple failure modes in multiple risk dimensions; The failure mode ranking module is used to rank multiple failure modes according to priority to guide product management decisions.
Citation Information
Patent Citations
Risk assessment method and device for failure mode of power battery system of new energy automobile
CN114154252A
Automobile oil can manufacturing process state evaluation method based on G-TOPSIS-FMEA
CN115689279A
Soil environment parameter information monitoring system based on intelligent sensor network
CN116894166A
Vehicle defect misjudgment identification method and system based on model rule
CN118965120A
FMEA creation assist system and method
JP2018045548A