Product quality problem correlation analysis method based on bipartite graph subgraph mining
By constructing a two-part graph model and using sub-graph mining algorithms to deeply analyze the relationship between product components and quality data, the problem of difficulty in revealing deep correlations and inefficiency of large-scale data analysis is solved, and effective prediction and identification of potential quality problems are achieved.
Patent Information
- Application Number
- CN202411909317.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to deeply reveal the deep correlation between product components and quality issues, and there are problems of inefficiency and insufficient prediction capabilities when dealing with unstructured and large-scale data.
By constructing a two-part graph model, it represents the relationship between product components and quality data, and uses a sub-graph mining algorithm to deeply analyze these relationships to identify potential quality problems and key influencing factors. Specific steps include data preprocessing, label extraction, two-part graph modeling and cohesive subgraph query index construction.
It has realized the disclosure of the complex relationship between products and quality issues, improved the efficiency of large-scale data analysis, enhanced the ability to predict potential quality issues, and can take preventive measures in advance to improve product quality.
Smart Images

Figure CN120047019A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and in particular relates to a product quality problem association analysis method based on bipartite graph subgraph mining. Background Art
[0002] In the wave of digital transformation, data has become the core asset of enterprises and plays a vital role in enterprise innovation and technological innovation. With the digitalization of industry, intelligent information management systems have provided new tools for enterprise management, especially in the field of product quality management. Enterprises have deployed a variety of independently operated information management systems, such as production process management systems, product structure management systems, and finished product quality management systems, to collect and manage data from different sources. However, the data they collect are interrelated in the process of product quality formation. Product quality problems can often be identified by tracing the process data of the processing and manufacturing process, thereby discovering the key process links that affect quality.
[0003] Current technologies have defects in processing industrial data, especially unstructured data. These data include product process text, quality description text and fault cause description, which are crucial for identifying key factors affecting product quality. At present, there are relatively few studies targeting the industrial sector, especially the lack of systematic analysis and knowledge construction of process and quality data. In order to solve these problems, the latest research progress has provided some innovative methods. For example, [1] proposed a quality association analysis method that does not rely on data distribution assumptions. Through data collection, preprocessing, modeling, training and partitioning, this method can be applied to situations where parameter distribution is unknown. [2] By obtaining data related to quality problems, extracting quality labels, and performing product quality problem analysis and fault process association analysis based on quality labels, data resources are integrated and association analysis is performed on these data, which is conducive to tracing the source of quality problems.
[0004] However, existing data analysis methods often have difficulty revealing the deep correlation between product components and quality issues. Existing product quality problem correlation analysis methods have defects and limitations in deep correlation analysis, making it difficult to fully reveal the complex relationship between products and quality issues. In addition, this method mainly relies on the distribution transformation model constructed by reversible neural networks, which may not be able to effectively extract and utilize key information when processing unstructured data such as process documents and quality failure descriptions. In addition, the ability to predict and identify potential quality problems may be insufficient, making it difficult to take preventive measures in advance during the production process. At the same time, this method may face challenges when processing large-scale data.
[0005] On the other hand, the existing technology for mining cohesive subgraphs in bipartite graphs also has defects. This is because with the rapid growth of data scale, existing algorithms usually face the problem of low efficiency or difficult maintenance when dealing with subgraph search tasks. Furthermore, the currently proposed index BiCore-Index, although it has good space efficiency and maintenance efficiency, does not store the connectivity information between vertices, and needs to perform breadth traversal on the bipartite graph to find the cohesive subgraph where the query point is located, which is time-consuming. The established index structure skyline also faces the above-mentioned time-consuming problem. Summary of the invention
[0006] In view of the above problems, a product quality problem association analysis method based on bipartite graph subgraph mining has emerged. This method constructs a bipartite graph model to represent the relationship between product components and quality data, and uses a subgraph mining algorithm to deeply analyze these relationships to identify potential quality problems and key influencing factors. For example, in the field of aviation product manufacturing, a bipartite graph model can be used to analyze the association between various components and data labels in an aircraft engine, helping engineers identify parts prone to failure and predict and identify parts that may have quality problems.
[0007] Specifically, the present invention aims to disclose a method and system for product quality problem association analysis based on bipartite graph subgraph mining, the method comprising: obtaining product and quality-related data. Quality-related data include processing and manufacturing process data, product structure data, maintenance and repair history data, and batch and traceability data; formatting and cleaning the data; extracting quality-related data for label extraction; establishing a bipartite graph model based on the association relationship between products and quality labels; and analyzing product quality problem associations based on a method for bipartite graph subgraph mining. The present invention integrates data resources and optimizes the bipartite graph subgraph mining algorithm to make the system suitable for large-scale data analysis, facilitate experts to perform association analysis and optimize manufacturing processes, and facilitate product quality problem tracing and mining products with potential quality problems. In addition, the method also helps to predict and identify products that may have quality problems and improve product quality.
[0008] The present invention specifically provides a product quality problem association analysis method based on bipartite graph subgraph mining, comprising the following steps:
[0009] Establish a product information database that will centrally store and manage all product quality-related data;
[0010] Preprocess the collected data;
[0011] The deep learning framework of BERT combined with BiLSTM and CRF is used to extract labels from the preprocessed data;
[0012] According to the different dimensional information of the product, multiple bipartite graphs are modeled. Since these bipartite graphs share the same set of product nodes, in order to reduce storage overhead, multiple bipartite graphs are further fused into a heterogeneous graph containing multiple entities and relationships.
[0013] Based on heterogeneous graphs, a cohesive subgraph query index is constructed for bipartite graphs of different dimensions, and the correlation analysis of product quality issues is performed based on the index-based bipartite graph cohesive subgraph mining algorithm.
[0014] Furthermore, the data includes product structure data, manufacturing process data, maintenance and repair history data, and batch and traceability data.
[0015] Furthermore, data preprocessing includes data cleaning, specifically, processing missing values in product structure data, detecting outliers in manufacturing process data, converting data types of batch and traceability data, deleting duplicate records, and normalizing or standardizing data.
[0016] Furthermore, the data is labelled and the specific steps are as follows:
[0017] 1) Corpus preprocessing: perform word segmentation, stop word removal and normalization on process text and quality description data to make the input data format uniform;
[0018] 2) BERT embedding generation: Use the BERT pre-trained model to generate high-dimensional semantic embeddings for each word, capturing the contextual information of the word and its deep semantic features;
[0019] 3) BiLSTM sequence modeling: The word embedding generated by BERT is fed into a bidirectional LSTM (BiLSTM) network to further model the dependencies between word sequences and extract deep features.
[0020] 4) CRF tag prediction: The conditional random field (CRF) layer combines the word sequence characteristics to output the tag category corresponding to each word to achieve keyword extraction.
[0021] Furthermore, based on the cohesive model (α, β)-core of the bipartite graph, the concept of (α, β)-shell is proposed; by studying the association between the connected components of different (α, β)-shells, an index structure ASG is established, specifically,
[0022] First, in descending order of α and β values, α and β refer to the constraints that the node degree must satisfy, i.e., the degree of the upper node is not less than α, and the degree of the lower node is not less than β. Access the corresponding (α, β)-core and identify the connected components in each subgraph. Merge vertices with the same number of dual cores and located in the same connected component into an aggregate supernode; following the above steps, a summary graph SG consisting of these aggregate supernodes can be constructed;
[0023] Secondly, traverse the edge set E of the bipartite graph. If (u, v) ∈ E and u and v belong to two different aggregation supernodes, project (u, v) into the summary graph SG.
[0024] Next, the redundant projection edges are deleted and the current summary graph SG is constructed as an acyclic graph ASG. A top-down strategy is adopted to reduce the summary graph SG to an acyclic graph ASG by constructing a local minimum spanning tree of the summary graph.
[0025] Finally, a community search algorithm is designed: first, the query point is located and the supernode where it is located is determined. Starting from the supernode, a breadth-first traversal is performed in the area defined by specific α and β values. During the traversal, the algorithm will identify and record all vertices connected to the query point, thereby determining the complete scope of the community. The algorithm will analyze the current search path and predict whether it may contain the connected components of the query point, thereby deciding whether to continue searching along the path or terminate early.
[0026] Further, the correlation analysis of product quality issues includes:
[0027] (1) Correlation analysis of quality issues of a single defective product: Input a defective product p∈U, and its corresponding cohesive subgraph Cp can be obtained through the index-based cohesive subgraph mining algorithm. The subgraph consists of vertex sets Up and Vp that meet the cohesive conditions, that is:
[0028]
[0029] For the input defective product p and its associated product pi, the similarity of the corresponding cohesive subgraphs Cp and Cpi can be defined by the following formula:
[0030]
[0031] Among them, |Cp∩Cpi| represents the intersection size of two subgraphs (such as common vertex pairs or edges), and |Cp∪Cpi| represents the union size; when S≥δ (the set similarity threshold), p and pi are considered to have potential association and can be marked as products that may have similar quality problems.
[0032] (2) Correlation analysis of quality issues of multiple defective products: Input a set of defective products P = {p1, p2, p3, ..., pk}, and analyze the common features by finding the intersection method. Define the common cohesive subgraph CP = (UP, VP, EP) of these products as:
[0033] CP=Cp1∩Cp2∩...∩Cpk
[0034] If CP satisfies: |UP|≥τ1 (minimum threshold of the number of upper vertices); |VP|≥τ2 (minimum threshold of the number of lower vertices), then it can be considered that these products have a common defect mode, and the key quality problem influencing factors can be further extracted.
[0035] (3) Based on the topological structure of defective products found in historical queries, the correlation analysis of product quality issues is further expanded: First, for each bipartite graph, its cohesive subgraph is independently mined. For each bipartite graph, a series of subgraphs that meet the cohesive conditions are mined. Each cohesive subgraph contains a product set (U) and a related quality label set (V). These subgraphs represent the close correlation between the product and the quality issue in this dimension;
[0036] Based on the defective product set {p1, p2, p3, ..., pk} in the historical query, extract its corresponding cohesive subgraph set {CP1, CP2, ..., CPk}; each cohesive subgraph represents the defective product CPi and its associated quality label. For the target product, calculate the comparison between its cohesive subgraph Cp and the cohesive subgraph set of defective products in the historical query, and evaluate its similarity with each CPi; if its similarity with any defective product in the historical query reaches the threshold δ, mark it as a potential defective product;
[0037] Combined with the expert verification results, the analysis threshold δ is adjusted, the cohesion conditions are optimized, the key quality label Vp is further extracted, and the common defect patterns are identified.
[0038] The advantages of the present invention are:
[0039] (i) Revealing the complex relationship between products and quality issues: A bipartite graph model is constructed to deeply analyze the complex relationship between products and parameter labels, providing support for quality issue correlation analysis.
[0040] (ii) Efficient processing of large-scale data: In response to the growth of industrial data scale, the present invention improves the efficiency of cohesive subgraph mining by optimizing the algorithm to meet the needs of large-scale data analysis.
[0041] (iii) Predicting and identifying potential quality issues: The present invention enhances the ability to predict potential quality issues through collaborative analysis of multi-dimensional bipartite graphs so that preventive measures can be taken in advance during the production process. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the process of the present invention;
[0043] Figure 2 is a sample bipartite graph;
[0044] Figure 3 It is the summary graph SG and the acyclic graph ASG;
[0045] Figure 4 is the search sample graph;
[0046] Figure 5 This is an experimental effect diagram. DETAILED DESCRIPTION
[0047] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0048] refer to Figure 1 The present invention aims to provide a product quality problem association analysis method based on bipartite graph subgraph mining, the main contents are as follows Figure 1 The technical solution provided by the present invention has the following steps:
[0049] Step 1: Data collection, product information database.
[0050] In order to implement the present invention, it is necessary to first establish a product information database, which will centrally store and manage all data related to product quality. These data include product structure data, manufacturing process data, maintenance and repair history data, and batch and traceability data. This information is crucial for subsequent quality data label extraction, bipartite graph modeling, and quality problem association analysis.
[0051] The establishment of the database will ensure the consistency and integrity of the data and provide the necessary input for the bipartite graph subgraph mining algorithm, so that the system can effectively identify and analyze product quality issues.
[0052] Step 2: Data cleaning.
[0053] After establishing the product information database, data cleaning is a crucial next step. Data cleaning is the process of preprocessing the collected raw data, aiming to eliminate inconsistencies, errors and redundant information in the data, and ensure the accuracy and reliability of the data used in subsequent analysis. Specifically, data cleaning includes the following aspects.
[0054] Handling missing values in product structure data: First, it is necessary to identify missing values in product structure data and take appropriate treatment methods, such as deletion, filling or interpolation to remedy them. For example, in the aircraft structure data, the weight information of a certain aircraft wing component is missing. To fill this missing value, you can interpolate and estimate the weight of the wing component of a similar model aircraft, or fill it in according to the weight of other components in the same design scheme.
[0055] Detecting outliers in manufacturing process data: Detecting and handling outliers in manufacturing process data is another important task. This process usually relies on statistical analysis methods or domain knowledge to determine which data are outliers and decide how to deal with them. For example, in the manufacturing process of aircraft engines, the temperature record of a certain casting process is obviously abnormal and far exceeds the normal range. This may be caused by a temperature sensor failure or an operating error. By performing statistical analysis on the process data, it is found that the value is obviously beyond the temperature range of the historical record, so it can be marked as an outlier and further verified whether the data needs to be recollected.
[0056] Convert batch and traceability data types: Convert non-numeric data in batch numbers and traceability data to numeric data for subsequent mathematical operations and statistical analysis. For example, the production batch number of an aircraft engine is initially recorded in text format, but when performing statistical analysis, it needs to be converted to numeric data. This makes it easier to track, compare, and analyze batches, such as comparing the performance of engines from different batches.
[0057] Deleting duplicate records: Identify and delete duplicate records in the database to avoid affecting the accuracy of subsequent analysis due to duplicate data. For example, in the aircraft maintenance record, a maintenance record was entered repeatedly due to a system error. In order to avoid this duplicate data affecting the analysis results, duplicate records should be deleted during the data cleaning process to ensure that each maintenance record appears only once in the database to ensure the uniqueness and accuracy of the data.
[0058] Data normalization or standardization: In order to make data with different characteristics comparable on the same scale, data normalization or standardization may be required. For example, when analyzing the performance of different aircraft models, multiple indicators (such as fuel consumption, flight speed, load capacity, etc.) may be involved, and these indicators have different dimensions. In order to make them comparable on the same scale, these indicators can be normalized, such as converting fuel consumption into units of per kilometer per ton, which makes it easier to compare the performance of different aircraft.
[0059] Step 3: Extract data labels.
[0060] Product process descriptions, quality problem descriptions, etc. are mostly unstructured texts, in which there are few words directly related to quality problems. Through label extraction technology, process texts can be converted into process keywords, quality descriptions can be converted into problem labels, high-credibility keywords can be screened, and redundant information can be filtered. This invention uses the BERT+BiLSTM+CRF deep learning framework for label extraction for unstructured texts combined with corpus context information. The specific steps are as follows:
[0061] Corpus preprocessing: Segment the process text and quality description, remove stop words, and normalize them to ensure that the input data format is unified.
[0062] BERT embedding generation: Use the BERT pre-trained model to generate high-dimensional semantic embeddings for each word, capturing the contextual information of the word and its deep semantic features.
[0063] BiLSTM sequence modeling: The word embedding generated by BERT is fed into the bidirectional LSTM (BiLSTM) network to further model the dependencies between word sequences and extract deep features.
[0064] CRF tag prediction: The conditional random field (CRF) layer combines the word sequence characteristics to output the tag category corresponding to each word to achieve accurate keyword extraction.
[0065] Compared with traditional methods, this framework can more accurately extract keywords related to quality issues from unstructured text, providing more reliable data support for quality analysis and problem tracing.
[0066] Step 4: Model the products and labels as a bipartite graph.
[0067] The present invention models the relationship between products and quality issues based on a bipartite graph. A bipartite graph is a special graph structure in which nodes are divided into two non-overlapping sets, and edges only connect nodes from different sets. In the present invention, one set represents the product or its components, and the other set represents quality-related labels (such as defects, failures, etc.), which are used to intuitively represent the relationship between products and quality issues. However, due to the large scale and complex structure of industrial data, a single bipartite graph is difficult to meet storage and analysis requirements.
[0068] To this end, the present invention proposes a modeling method based on heterogeneous graphs, which characterizes industrial data as a network of multi-type nodes and multi-relationship edges. Heterogeneous graphs can be regarded as a collection of multiple bipartite graphs, each of which focuses on a specific entity type and its relationship, which can not only support distributed storage and management, but also significantly improve the efficiency of data processing. For example, a product-structure bipartite graph: describes the assembly relationship between a product and its components. The nodes represent product instances and components respectively.
[0069] Step 5: Build a cohesive subgraph query index.
[0070] The existing cohesive subgraph mining algorithm is inefficient. To this end, the present invention studies the cohesive model characteristics, combines the actual application requirements, designs a more reasonable and effective index structure, and designs an accurate and effective search algorithm based on the index structure to improve the search efficiency.
[0071] Based on the cohesive model (α, β)-core of the bipartite graph, the present invention proposes the concept of (α, β)-shell. By studying the association relationship between the connected components of different (α, β)-shells, an index structure ASG is established, which is essentially an undirected acyclic graph composed of the connected components of the (α, β)-shell. Then, the present invention designs accurate query algorithms and maintenance algorithms based on the characteristics of the ASG structure, thereby improving the effectiveness and efficiency of the product quality problem association analysis method based on bipartite graph subgraph mining.
[0072] First, in descending order of α and β values, the corresponding (α, β)-cores are accessed and the connected components in each subgraph are identified. Vertices with the same number of dual cores and located in the same connected component are merged into an aggregate supernode. Following the above steps, a summary graph SG consisting of these aggregate supernodes can be constructed. Figure 3 As shown, Figure 3 The supernodes in Figure 2 There are 7 aggregated supernodes in total, and their coordinates are their corresponding α and β values.
[0073] Secondly, traverse the edge set E of the bipartite graph, if (u, v) ∈ E and u, v belong to two different aggregation supernodes, then project (u, v) into the summary graph SG.
[0074] Next, the redundant projection edges are deleted and the current summary graph SG is constructed as an acyclic graph ASG. The summary graph SG constitutes the core of the index, which provides a solid foundation for efficient subgraph search and discovery. However, in the initially constructed summary graph, there may be some redundant edges. These edges do not increase the amount of information in the index, but may reduce the search efficiency. In order to improve the performance of the index, the summary graph needs to be further optimized. In order to achieve the goal, a top-down strategy is adopted to reduce the summary graph SG to an acyclic graph ASG by constructing a local minimum spanning tree of the summary graph, such as Figure 3As shown in the figure. The minimum spanning tree is an effective data structure that connects all vertices in the graph while ensuring that the total number of edges is minimized, which helps reduce redundant information in the index and improves the index retrieval efficiency. By constructing a local minimum spanning tree, not only can redundancy be reduced, but the structure of the index can also be optimized to make it more compact and efficient. Such an index structure will facilitate fast retrieval and traversal, thereby providing faster response speed in practical applications.
[0075] Finally, the community search algorithm is designed. First, the query point is located and the supernode where it is located is determined. Starting from the supernode, a breadth-first traversal is performed in the area defined by the specific α and β values. During the traversal, the algorithm will identify and record all vertices connected to the query point to determine the complete scope of the community. The algorithm will analyze the current search path and predict whether it may contain the connected components of the query point, so as to decide whether to continue searching along the path or terminate early. The general process of the search is as follows: Figure 4 shown.
[0076] Step 6: Product quality correlation analysis.
[0077] (1) Single defective product quality problem association analysis: Input a defective product p∈U, and its corresponding cohesive subgraph C can be obtained through the index-based cohesive subgraph mining algorithm p , the subgraph consists of a set of vertices U that satisfy the cohesion condition p and V p Composition, namely:
[0078]
[0079] For the input defective product p and its associated product p i Its corresponding cohesive subgraph C p and C pi The similarity can be defined by the following formula:
[0080]
[0081] Among them, |C p ∩C pi ∣ represents the size of the intersection of two subgraphs (e.g. common vertex pairs or edges), ∣C p ∪C pi ∣ represents the size of the union. When S ≥ δ (the set similarity threshold), it is considered that p and p i There is a potential connection and products can be flagged as potentially having similar quality issues.
[0082] (2) Correlation analysis of multiple defective product quality issues: Input a set of defective products P = {p 1 ,p 2 ,p 3,...,p k}, the common features can be analyzed by finding the intersection method. Define the common cohesive subgraph C of these products P =(U P ,V P ,E P )for:
[0083] C P =C p1 ∩C p2 ∩...∩C pk
[0084] If C P Satisfaction: |U P ∣≥τ 1 (minimum threshold of the number of upper-level vertices); |V P ∣≥τ 2 (The minimum threshold of the number of lower-level vertices). It can be considered that these products have common defect patterns, and the key quality problem influencing factors can be further extracted.
[0085] (4) Based on the topological structure of defective products found in historical queries, the correlation analysis of product quality issues is further expanded: First, for each bipartite graph, its cohesive subgraph is independently mined. For each bipartite graph, a series of subgraphs that meet the cohesive conditions are mined. Each cohesive subgraph contains a product set (U) and a related quality label set (V). These subgraphs represent the close correlation between the product and the quality issue in this dimension.
[0086] Take the defective product set in the historical query {p 1 ,p 2 ,p 3 ,...,p k}, extract the corresponding cohesive subgraph set {C P1 ,C P2 ,…,C Pk}. Each cohesive subgraph represents a defective product C Pi The quality label associated with it. For the target product, calculate its cohesive subgraph C p Compare with the cohesive subgraph set of defective products in historical queries to evaluate its compatibility with each C Pi If its similarity with any historical query defective product reaches a threshold δ, it is marked as a potential defective product.
[0087] Combined with the expert verification results, the analysis threshold δ is adjusted, the cohesion conditions are optimized, the key quality label Vp is further extracted, and the common defect patterns are identified.
[0088] Technical effects of the present invention:
[0089] (1) The present invention proposes to construct a bipartite graph model and a heterogeneous graph model, which overcomes the problems of low level of informatization in the industrial manufacturing process, scattered original data structure, and difficulty in intuitively discovering rules due to the lack of correlation between data.
[0090] (2) The present invention optimizes the cohesive subgraph mining algorithm of bipartite graphs and conducts comparative experiments on four real open source bipartite graph datasets. The dataset parameters are shown in Table 1 below.
[0091] Table 1
[0092]
[0093] To enhance the credibility of the results, the present invention randomly selects 500 vertices from each data set as query vertices and averages the results. Figure 5 As shown, the time unit is milliseconds (ms). Experimental results show that the search efficiency of the present invention is improved by 1-2 orders of magnitude compared with the existing advanced methods.
[0094] (3) In response to the needs of product quality association analysis, the present invention proposes an analysis method based on cohesive subgraph similarity and common cohesive subgraph, which realizes the rapid identification of potential quality problem associations and in-depth mining of key influencing factors, and enhances the prediction ability of potential quality problems.
[0095] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A product quality problem association analysis method based on bipartite graph subgraph mining, characterized in that: The following steps are included: Establish a product information database that will centrally store and manage all product quality-related data; Preprocess the collected data; The deep learning framework of BERT combined with BiLSTM and CRF is used to extract labels from the preprocessed data; Model multiple bipartite graphs based on different dimensional information of products, where the multiple bipartite graphs share the same set of product nodes; merge the multiple bipartite graphs into a heterogeneous graph containing multiple entities and relationships; Based on heterogeneous graphs, a cohesive subgraph query index is constructed for bipartite graphs of different dimensions, and the correlation analysis of product quality issues is performed based on the index-based bipartite graph cohesive subgraph mining algorithm.
2. The product quality problem association analysis method based on bipartite graph subgraph mining according to claim 1 is characterized in that: The data includes product structure data, manufacturing process data, maintenance and repair history data, and batch and traceability data.
3. The product quality problem association analysis method based on bipartite graph subgraph mining according to claim 1 is characterized in that: Data preprocessing includes data cleaning, specifically, processing missing values in product structure data, detecting outliers in manufacturing process data, converting data types of batch and traceability data, deleting duplicate records, and normalizing or standardizing data.
4. The product quality problem association analysis method based on bipartite graph subgraph mining according to claim 1 is characterized in that: Extract labels from data. The specific steps are as follows: 1) Corpus preprocessing: perform word segmentation, stop word removal and normalization on process text and quality description data to make the input data format uniform; 2) BERT embedding generation: Use the BERT pre-trained model to generate high-dimensional semantic embeddings for each word, capturing the contextual information of the word and its deep semantic features; 3) BiLSTM sequence modeling: The word embedding generated by BERT is fed into a bidirectional LSTM (BiLSTM) network to further model the dependencies between word sequences and extract deep features. 4) CRF tag prediction: The conditional random field (CRF) layer combines the word sequence characteristics to output the tag category corresponding to each word to achieve keyword extraction.
5. The product quality problem association analysis method based on bipartite graph subgraph mining according to claim 1 is characterized in that: Based on the cohesive model (α, β)-core of the bipartite graph, the concept of (α, β)-shell is proposed. By studying the association between the connected components of different (α, β)-shells, the index structure ASG is established, which is as follows: First, in descending order of α and β values, the corresponding (α, β)-cores are accessed and the connected components in each subgraph are identified. The vertices with the same number of dual cores and located in the same connected component are merged into an aggregate supernode; Following the above steps, a summary graph SG consisting of these aggregated super nodes can be constructed; Secondly, traverse the edge set E of the bipartite graph. If (u, v) ∈ E and u and v belong to two different aggregation supernodes, project (u, v) into the summary graph SG. Next, the redundant projection edges are deleted and the current summary graph SG is constructed as an acyclic graph ASG. A top-down strategy is adopted to reduce the summary graph SG to an acyclic graph ASG by constructing a local minimum spanning tree of the summary graph. Finally, a community search algorithm is designed: first, the query point is located and the supernode where it is located is determined. Starting from the supernode, a breadth-first traversal is performed in the area defined by specific α and β values. During the traversal, the algorithm will identify and record all vertices connected to the query point, thereby determining the complete scope of the community. The algorithm will analyze the current search path and predict whether it may contain the connected components of the query point, thereby deciding whether to continue searching along the path or terminate early.
6. The product quality problem association analysis method based on bipartite graph subgraph mining according to claim 1 is characterized in that: The correlation analysis of product quality issues includes: (1) Correlation analysis of quality issues of a single defective product: Input a defective product p∈U, and its corresponding cohesive subgraph Cp can be obtained through the index-based cohesive subgraph mining algorithm. The subgraph consists of vertex sets Up and Vp that meet the cohesive conditions, that is: For the input defective product p and its associated product pi, the similarity of the corresponding cohesive subgraphs Cp and Cpi can be defined by the following formula: Among them, |Cp∩Cpi| represents the intersection size of two subgraphs (such as common vertex pairs or edges), and |Cp∪Cpi| represents the union size; when S≥δ (the set similarity threshold), p and pi are considered to have potential association and can be marked as products that may have similar quality problems. (2) Correlation analysis of quality issues of multiple defective products: Input a set of defective products P = {p1, p2, p3, ..., pk}, and analyze the common features by finding the intersection method. Define the common cohesive subgraph CP = (UP, VP, EP) of these products as: CP=Cp1∩Cp2∩...∩Cpk If CP satisfies: |UP|≥τ1 (minimum threshold of the number of upper vertices); |VP|≥τ2 (minimum threshold of the number of lower vertices), then it can be considered that these products have a common defect mode, and the key quality problem influencing factors can be further extracted. Based on the topological structure of defective products found in historical queries, the correlation analysis of product quality issues is further expanded: First, for each bipartite graph, its cohesive subgraph is mined independently. For each bipartite graph, a series of subgraphs that meet the cohesive conditions are mined. Each cohesive subgraph contains a product set (U) and a related quality label set (V). These subgraphs represent the close correlation between the product and the quality issue in this dimension; Based on the defective product set {p1, p2, p3, ..., pk} in the historical query, extract its corresponding cohesive subgraph set {CP1, CP2, ..., CPk}; each cohesive subgraph represents the defective product CPi and its associated quality label. For the target product, calculate the comparison between its cohesive subgraph Cp and the cohesive subgraph set of defective products in the historical query, and evaluate its similarity with each CPi; if its similarity with any defective product in the historical query reaches the threshold δ, mark it as a potential defective product; Combined with the expert verification results, the analysis threshold δ is adjusted, the cohesion conditions are optimized, the key quality label Vp is further extracted, and the common defect patterns are identified.
Citation Information
Cited By
FPC flexible circuit board intelligent production system based on image recognition
CN120317636A
An intelligent production system for FPC flexible circuit boards based on image recognition
CN120317636B
Large aircraft complex design and assembly modeling method based on unified networked representation
CN120724586A