Aircraft general assembly information association rule mining method and equipment based on fine tuning large language model, and medium
Through the method of mining the aircraft assembly information association rules based on the fine-tuning large language model, the problem that traditional methods are difficult to accurately analyze the association relationship of production factors is solved, and high-precision association relationship recognition and optimization of the assembly process are achieved.
Patent Information
- Application Number
- CN202510685297.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-27
AI Technical Summary
It is difficult for traditional aircraft assembly information management and analysis methods to comprehensively and accurately extract and analyze the relationship between production factors, resulting in challenges in the optimization of assembly process and efficiency improvement.
The aircraft assembly information correlation rule mining method based on the fine-tuning large language model is adopted. By collecting multi-source production factor data, performing domain adaptation, extracting structured factor information, building semantic optimization project sets, mining association rules and optimizing the output of high confidence correlation rule sets.
It significantly enhances the identification and information extraction capabilities of the model in the professional terms of aviation field, improves the identification accuracy of the correlation between production factors, and provides reliable data support and decision-making basis for the optimization of the assembly process.
Smart Images

Figure CN120197688A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital processing of electronic data, and particularly to information retrieval. Specifically, it is a method, device, and medium for mining aircraft final assembly information association rules based on fine-tuning large language models. Background Art
[0002] During the aircraft final assembly manufacturing process, it involves a vast amount of production factor data, including complex information such as process information, material requirements, tooling equipment codes, and workstation positions. These data are usually scattered in structured (such as AO process cards in the MES system), semi-structured (such as BOM lists in the PDM system), and unstructured (such as natural language descriptions in process guides) forms. These factors are coupled with each other in the final assembly link, having a significant impact on the assembly efficiency.
[0003] Traditional aircraft final assembly information management and analysis methods mainly rely on manual experience and simple data analysis, making it difficult to comprehensively and accurately extract and analyze the association relationships between these production factors, resulting in many challenges in optimizing the assembly process and improving efficiency. These methods often have technical problems such as inaccurate information extraction and incomplete association analysis when dealing with complex production factor data.
[0004] Therefore, a method that can efficiently and accurately mine the association relationships between production factors in aircraft final assembly process information is urgently needed. Summary of the Invention
[0005] The present invention aims to solve the problem of inaccurate extraction and association analysis of production factor information in the aircraft final assembly process, overcomes the deficiencies of the prior art, and provides a method, device, and medium for mining aircraft final assembly information association rules based on fine-tuning large language models.
[0006] To achieve the above object, the technical solution adopted by the present invention is: A method for mining aircraft final assembly information association rules based on fine-tuning large language models, including the following steps: S1. Collect multi-source production factor data in aircraft final assembly process documents and perform preprocessing; S2. Perform domain adaptation based on the large language model fine-tuning method of low-rank adaptation to enhance the ability to recognize professional terms; S3. Extract structured element information through the fine-tuned large language model to construct a semantically optimized item set; S4. Mine the association rules between production factors and the final assembly link based on the association rule mining algorithm; S5. Combine word vector clustering and semantic verification to optimize and output a high-confidence association rule set.
[0007] In a preferred embodiment of the present invention, in the step of S1, the multi-source production factor data includes structured and unstructured data, and comprises the following sub-steps: S11. Extract the AO number, station number, and position number from the structured data, and perform standardized coding; S12. Parse the material code and tooling equipment code from the semi-structured data; S13. Extract professional terms, material codes, and tooling equipment codes from the unstructured text through natural language processing technology, and perform data cleaning based on regular expressions, domain dictionaries, and edit distance.
[0008] In a preferred embodiment of the present invention, in the step of S13, the extraction of the unstructured text is performed using a large language model from the process steps described in natural language, and the input format is defined as: ; Wherein, is the th process; is the AO number of process ; is the station number where process is located; is the position number where process is located; is the material required for process , is the set of material numbers; is the tooling equipment required for process , is the set of numbers; is the set of key process information of process .
[0009] In a preferred embodiment of the present invention, in the step of S2, it comprises the following sub-steps: S21. Insert a low-rank adapter module into the Transformer layer of the pre-trained large language model, and adjust the model parameters through a low-rank matrix; S22. Freeze the original model parameters, and only train the low-rank adapter module, and perform fine-tuning using the cross-entropy loss function and the dynamic learning rate strategy; S23. Construct a fine-tuning data set based on the domain-adapted prompt template, and enhance the information extraction accuracy by combining role definition, task statement, and output example.
[0010] In a preferred embodiment of the present invention, in the step of S21, the insertion of the low-rank matrix is specifically: insert a matrix with a rank of beside the Query and Value matrices of each Transformer layer of the model The low-rank adapter module has a weight update formula as follows: ( , ); Among them, is the rank number; and are the row and column dimensions of the original model weight matrix; is the set of real numbers; Transformer is the model architecture. The Transformer layer includes a self-attention mechanism and a feed-forward neural network, which are used to capture the context dependencies in the text; Query is the query matrix in the self-attention mechanism, which is used to calculate the correlation weights between different positions in the input sequence; Value is the value matrix in the self-attention mechanism, which is used to apply the weights to the original input to generate the final context-aware representation.
[0011] In a preferred embodiment of the present invention, in the step of S3, the following sub-steps are included: S31. Use a multi-agent mechanism to perform format verification and domain term matching on the extraction results, and merge synonym classes and eliminate meaningless words through a clustering algorithm; The multi-agent mechanism includes: the first agent module: verifying the format compliance of the extraction results; the second agent module: verifying the compliance of terms and codes based on the aviation manufacturing domain term library; S32. Encode the clustering results into standard labels to generate a semantically consistent structured item set.
[0012] In a preferred embodiment of the present invention, in the step of S4, the association rule between the production factors and the general assembly link is based on the FP-Growth algorithm, specifically including: frequent item set mining and generating association rules based on the frequent item sets.
[0013] In a preferred embodiment of the present invention, for the frequent item set mining: count the occurrence frequency of each item set in the transaction database. For the item set ( is the set of all items, and all items are all different single basic items in the dataset), its support degree is calculated as: , ; Among them, is the total number of transactions; is the indicator function, which is 1 when and 0 otherwise; is the item set in the entire transaction database the number of times it appears; is the transaction database a single transaction, i.e., a set of data items; Itemsets with support ≥ minimum threshold The set of frequent itemsets is defined as: ; Build an index for each frequent item, record its position and global frequency in the FP-tree, and the header table is a dictionary with the frequent item as the key and the value is: ; where is the support count of item ; is a linked list pointing to all nodes in the FP-tree that contain ; Compress the transaction database into a prefix tree structure, and each node contains the following attributes: ; where is the item represented by the current node; is the number of occurrences of the path from the root to this node; is the parent node pointer; is the child node dictionary; Construction rule: For each transaction , insert it into the tree after sorting in descending order of the support of frequent items, and merge and count the paths with the same prefix; Conditional pattern base generation: For item , its conditional pattern base is the set of all paths in the FP-tree that end with , and each path is in the form of: ; Recursive formula: For each item , recursively construct the conditional FP-tree and mine its frequent itemsets: ; where represents the set of new frequent itemsets generated from the conditional pattern base of the current item during the recursive process of the FP-Growth algorithm; is to re-run the FP-Growth algorithm on the conditional pattern base of item using the same minimum support threshold to mine the frequent itemsets; Finally, merge all Obtain global frequent itemsets; The association rule calculation: Generate rules based on the frequent itemsets, and screen for high-value rules through confidence and lift; The confidence represents the proportion of transactions that contain the consequent in the transactions that contain the antecedent , that is, the accuracy of the rule: ; Among them, is the number of transactions that contain and ; is the number of transactions that contain ; The lift represents the association strength between and . The higher the lift, the stronger the association between and . If the lift value is greater than 1, the association rule is considered valuable: ; Among them, is to measure the association strength of the rule , reflecting whether the co-occurrence of and has statistical significance; is the confidence of the rule ; is the support of the itemset .
[0014] The present invention provides an electronic device, which includes: at least one processor, and a memory communicatively connected to at least one of the processors; the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, the processor can execute the method for mining association rules of aircraft final assembly information based on a fine-tuned large language model as described in any one of the above.
[0015] The present invention provides a computer-readable storage medium, which stores computer instructions for causing a processor to implement the method for mining association rules of aircraft final assembly information based on a fine-tuned large language model as described in any one of the above when executed.
[0016] The present invention solves the defects in the background technology, and the present invention has the following beneficial effects: (1) The present invention provides a method, device, and medium for mining aircraft final assembly information association rules based on fine-tuning large language models. By combining the fine-tuning of large language models with association rule mining, the model's recognition and information extraction capabilities for aviation domain-specific terms are significantly enhanced. The introduction of low-rank matrices enables the model fine-tuning process to only adjust a small number of parameters, which not only preserves the general knowledge of the pre-trained model but also improves the accuracy of aviation professional tasks, effectively reducing the training complexity and resource consumption. The in-depth combination of mining algorithms effectively improves the accuracy of information extraction and the depth of production factor association analysis, thereby greatly improving the recognition accuracy of the association relationships between production factors and providing reliable data support and decision-making basis for the optimization of the final assembly process.
[0017] (2) In the model fine-tuning stage of the present invention, the low-rank adapter technology is adopted. By inserting low-rank matrices into the Transformer layer of the pre-trained model and freezing the original parameters, efficient domain adaptation of the large language model is achieved. The lightweight nature of the low-rank matrices significantly reduces the scale of parameter adjustment, enabling the model to converge quickly in the aviation domain-specific term recognition task. At the same time, it avoids the excessive dependence on computing resources for full-parameter fine-tuning. The combination of the dynamic learning rate strategy and the cross-entropy loss function further improves the stability and convergence speed of the fine-tuning process, significantly enhancing the model's adaptability in the aviation domain while greatly reducing the training cost.
[0018] (3) In the present invention, association rule mining adopts the FP-Growth algorithm combined with a multi-index screening mechanism. Through frequent itemset mining and lift screening, it not only ensures the high-frequency co-occurrence of rules but also quantifies the actual association strength between the antecedent and the consequent, avoiding false associations caused by the high-confidence trap in traditional methods. The combination of multi-dimensional evaluation rules ensures the actual business association strength of the rules, enhances the interpretability of the rules, thereby improving the mining accuracy of association rules and providing accurate basis for workstation resource allocation and material scheduling. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings; Figure 1 It is a flowchart of a method for mining aircraft final assembly information association rules based on fine-tuning large language models according to a preferred embodiment of the present invention; Figure 2 It is a data schematic diagram of the ChatGLM model fine-tuning process according to a preferred embodiment of the present invention; Figure 3It is a schematic diagram of the extraction structure of the preferred embodiment of the present invention; Figure 4 It is a schematic diagram of the association rule of the preferred embodiment of the present invention. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0022] As Figure 1 shown, a method for mining the association rules of aircraft final assembly information based on fine-tuning a large language model includes the following steps: S1. Collect multi-source production element data in aircraft final assembly process documents and perform preprocessing; S2. Perform domain adaptation based on the large language model fine-tuning method of low-rank adaptation to enhance the professional term recognition ability; S3. Extract structured element information through the fine-tuned large language model and construct a project set with optimized semantics; S4. Mine the association rules between production elements and final assembly links based on the association rule mining algorithm; S5. Combine word vector clustering and semantic verification to optimize and output a high-confidence association rule set.
[0023] In this embodiment, in the step of S1, the multi-source production element data includes structured and unstructured data, and includes the following sub-steps: S11. Extract the AO number, station number and position number from the structured data and perform standardized coding; S12. Parse the material code and tooling equipment code from the semi-structured data; S13. Extract professional terms, material codes and tooling equipment codes from unstructured text through natural language processing technology, and perform data cleaning based on regular expressions, domain dictionaries and edit distances.
[0024] In a specific implementation manner, in the step of S11, the structured data source comes from the AO process card exported from the MES system. Among them, the standardized coding rule is specifically: AO Number: The format is "AO-XXXXX" (e.g., AO-5720C41073), and it needs to be unified as the "AO" prefix + 10-digit alphanumeric combination; Workstation Number (WS): The format is "number + letter" (e.g., 200B), and it needs to be unified as 3 characters. If it is less than 3 characters, fill in zeros (e.g., 200B → 200B, no zero filling); Station Number (ST): The format is "XXX" (e.g., 101), and it needs to be unified as 3 digits. If it is less than 3 digits, fill in zeros (e.g., 1 → 001).
[0025] In a specific embodiment, in the step of S12, the semi-structured data source comes from the BOM list of the PDM system, and the parsing rules are specifically as follows: Material Code: Follow the format of "MAT-XXXXX" (e.g., MAT-3205A); Tooling Equipment Code: Follow the format of "TOOL-XXXX" (e.g., TOOL-8892).
[0026] In a specific embodiment, in the step of S13, the extraction of unstructured text uses a large language model (ChatGLM3-6B) to extract from the process steps described in natural language, and the input format is defined as: ; Among them, is the rd process; is the AO number of process ; is the workstation number where process is located; is the station number where process is located; is the material required for process , is the set of material codes; is the tooling equipment required for process , is the set of numbers; is the set of key process information of process .
[0027] Furthermore, for the text preprocessing process, a multi-stage cleaning strategy is adopted: Regular expression filtering: Extract standard codes in the form of "XXX-XXXX"; Domain dictionary matching: Load domain-specific vocabulary for exact matching; Correct spelling mistakes based on the edit distance.
[0028] In this embodiment, in the step of S2, the following sub-steps are included: S21. Insert a low-rank adapter module in the Transformer layer of the pre-trained large language model, and adjust the model parameters through a low-rank matrix; S22. Freeze the original model parameters, only train the low-rank adapter module, and perform fine-tuning using the cross-entropy loss function and the dynamic learning rate strategy; S23. Construct a fine-tuning dataset based on the prompt template for domain adaptation, and enhance the information extraction accuracy by combining role definitions, task statements, and output examples.
[0029] In a specific implementation, in the step of S21, the insertion of the low-rank matrix is specifically: beside the Query and Value matrices of each Transformer layer of the model (ChatGLM3-6B), insert a low-rank adapter module with a rank of , and the weight update formula is: ( , ) where, is the rank number ( ); and are the row and column dimensions of the original model weight matrix; is the set of real numbers; Transformer is the model architecture, and the Transformer layer contains a self-attention mechanism and a feed-forward neural network, which are used to capture the context dependencies in the text; Query is the query matrix in the self-attention mechanism, which is used to calculate the correlation weights between different positions in the input sequence; Value is the value matrix in the self-attention mechanism, which is used to apply the weights to the original input to generate the final context-aware representation.
[0030] Example: Assume that the dimension of the Query matrix of a certain Transformer layer is 768×768, then the dimension of the low-rank matrix is 768×8, is 8×768, and the dimension of the adjusted is still 768×768.
[0031] In a specific implementation, in the step of S22, as Figure 2 shown, the input hidden state representation is , where is the feature dimension, and freezing the original model parameters specifically means: during the fine-tuning process, only update the parameters of the low-rank matrices and . Among them, the parameters of are initialized from the normal distribution , is the variance, which represents the degree of dispersion of the data, Initializes to 0, the original model parameters Remains fixed; For the input sequence and the label , the cross-entropy loss function is: ; Wherein, is the number of classification categories (such as aviation terms, material codes, etc.); is the probability of the th class predicted by the model; is the one-hot encoding of the true label; Example: If the model needs to extract the term "wing-body docking" and the code "MAT-3205A" from the text "Wing-body docking requires MAT-3205A", then the label is the corresponding one-hot encoded vector, and the loss function measures the difference between the predicted distribution and the true label; Dynamically adjust the learning rate according to the gradient variance of the low-rank adapter module to prevent gradient explosion or oscillation: After each training epoch, calculate the gradient variance of all low-rank adapter parameters: ; Wherein, is the total number of low-rank adapter parameters; is a certain parameter in the low-rank adapter module (i.e., the inserted low-rank matrix and ); is the gradient mean; Learning rate adjustment rule: If , then the learning rate decays to 50% of the original value: ; Adopt the Adam optimizer, set the learning rate to , the batch size to 32, and the number of training epochs to 10 until the loss function converges.
[0032] In a specific implementation, in the step of S23, design a prompt template based on the CRISPE framework, and construct a fine-tuning dataset in combination with Few-shot examples. The example template is as follows: "Technical term": ["Wing-body docking", "Fuel pipeline installation"], "Material code": ["MAT-3205A", "MAT-1876B"], "Tooling equipment code": ["TOOL-8892"].
[0033] In this implementation, in the step of S3, the following sub-steps are included: S31, using a multi-agent mechanism to verify the format of the extraction results and match them with domain terms, and using a clustering algorithm to merge synonym classes and remove meaningless words; S32. Encode the clustering results into standard labels to generate a semantically consistent structured item set.
[0034] In a specific implementation, in step S31, the extraction result is as follows: Figure 3 As shown; the multi-agent mechanism includes: The first agent module: verify the format standardization of the extraction results; The second agent module: Verify the compliance of terms and codes based on the terminology database in the aviation manufacturing field.
[0035] Further, the input example of the first agent module is: “Professional terms”: [“wing-body docking”, “fuel line installation”, “name”], “Material Code”: [“MAT-3205A”, “MAT-32O5A”, “MAT-1876B”], “Tooling equipment code”: [“TOOL-8892”, “TOOL-8892”].
[0036] First-agent processing: Eliminate invalid items: "name" (non-term), "MAT-32O5A" (including the letter O), "TOOL-8892" (including Chinese characters).
[0037] Output: “Professional terms”: [“wing-body docking”, “fuel line installation”], “Material Code”: [“MAT-3205A”, “MAT-1876B”], “Tooling equipment code”: [“TOOL-8892”].
[0038] Verification term similarity matching rules of the second agent module: Exact match: directly compare whether the term / code exists in the domain library; Fuzzy matching: For terms that are not exactly matched, calculate the cosine similarity with the terms in the domain library, and retain the terms with similarity ≥ 0.8; Cosine Similarity: ; in, and is the word vector of the two words (such as the 300-dimensional vector generated by Word2Vec); is the vector dimension (300).
[0039] Furthermore, the K-means clustering algorithm is used to group the word vectors. By presetting the number of clusters (e.g., K = 10), the algorithm automatically divides words with similar semantics into the same cluster. For example, "connect" and "interconnect" may be grouped into the same cluster, while "fasten" forms a separate cluster. During the clustering process, the system continuously optimizes the distribution of words within the clusters until the cluster centers no longer change significantly.
[0040] In a specific implementation, in step S32, based on a professional dictionary or expert knowledge in the field of aviation manufacturing, semantic labels are assigned to each clustering cluster for the clustering result. For example, the cluster containing "connect" and "interconnect" is defined as the "connection category", and the cluster containing "fasten" and "bolt" is defined as the "fastener category". Automated processing is carried out through the following rules: Domain dictionary matching: Count the attribution categories of the words in the cluster in the domain dictionary, and select the category with the highest frequency as the label.
[0041] Label encoding rule: Assign a unique code (e.g., C1 - C10) to each semantic category, and establish a mapping table, as shown in Table 1: Table 1: Semantic category mapping table
[0042] In this implementation, in step S4, the association rules between production factors and the general assembly link are based on the FP-Growth algorithm, specifically including: frequent item set mining and generating association rules based on frequent item sets.
[0043] Frequent item set mining: Count the occurrence frequency of each item set in the transaction database. For the item set ( is the set of all items, and all items are all different single basic items in the dataset, such as workstations, materials, terms), its support is calculated as: , ; Among them, is the total number of transactions; is the indicator function, which is 1 when and 0 otherwise; is the item set in the entire transaction database ; is the single transaction in the transaction database , that is, a set of data items; Retain the item sets with support ≥ the minimum threshold , and the set of frequent item sets is defined as: ; Build an index for each frequent item, record its position in the FP-tree and its global frequency, and the header table is a dictionary with the frequent item as the key , and the value is: ; Among them, is the support count of the item ; is a linked list pointing to all nodes in the FP-tree that contain ; Compress the transaction database into a prefix tree structure, and each node contains the following attributes: ; Among them, is the item represented by the current node; is the number of occurrences of the path from the root to this node; is the parent node pointer; is a dictionary of child nodes (with items as keys); Construction rule: For each transaction , insert it into the tree after sorting in descending order of the support of frequent items, and merge and count the paths with the same prefix; Conditional pattern base generation: For the item , its conditional pattern base is the set of all paths in the FP-tree that end with , and each path is in the form of: ; Recursive formula: For each item , recursively construct the conditional FP-tree , and mine its frequent item sets: ; Among them, represents the set of new frequent item sets generated by the conditional pattern base passing through the current item during the recursive process of the FP-Growth algorithm; is to re-run the FP-Growth algorithm on the conditional pattern base of the item using the same minimum support threshold to mine frequent item sets; Finally, merge all to obtain the global frequent item sets. Association rule calculation: Generate rules based on frequent item sets, and screen high-value rules through confidence and lift;
[0044] Confidence indicates that in the case of containing the antecedent (combination of production factors) The transactions also include the consequent (final assembly station or key process). The proportion of transactions, which is the accuracy of the rule: ; Among them, is the number of transactions that include and ; is the number of transactions that include ; The lift represents and The strength of the association between: ; Among them, is to measure the association strength of the rule , reflecting and Whether the co-occurrence has statistical significance; is the confidence of the rule ; is the support of the item set ; Specifically, if the lift value is greater than 1, the association rule is considered valuable: Lift = 1: is independent of and has no association; Lift > 1: is positively correlated with , The occurrence of will increase The probability of occurrence; Lift < 1: is negatively correlated with , The occurrence of will reduce The probability of occurrence.
[0045] It should be noted that in the aircraft final assembly scenario, if the lift of a certain rule (such as {workstation 200B} ⇒ {material MS21043 - 3}) > 1, it means that the probability of using this material at this workstation is significantly higher than the global distribution probability of this material, indicating a strong business association between the two, rather than a random coincidence.
[0046] It can be understood that the proposed algorithm can avoid the high-confidence trap and the problem of ignoring independence testing existing in traditional association rule mining methods (such as Apriori, FP - Growth), which usually rely only on support and confidence to screen rules: High-confidence trap: If the consequent ( ) itself has a very high support (such as a certain material is widely used at all workstations), even if Irrelevant to rules may also have a high confidence level, leading to false associations; Ignoring independence tests: not quantifying the actual association strength, may misjudge accidental co-occurrences as valid rules.
[0047] Example: Suppose in the process of a certain station 200B, in 90% of cases, a general-purpose screw is required (support = 30%). Then the confidence level of the rule {Station 200B} ⇒ {General-purpose screw} = 90%, seemingly a strong association. However, if the lift = 1 (i.e., the average usage rate of general-purpose screws at all stations is also 30%), it indicates that this association has no practical significance, only because the general-purpose screw itself is frequently used.
[0048] In this embodiment, in the step of S5, the following sub-steps are included: S51. Perform secondary clustering on the optimized structured item set to eliminate noise data; S52. Align the structured item set with the association rule set to generate a rule set with unified expression, and output the final high-confidence rule set.
[0049] In a specific embodiment, in the step of S51, first use the Word2Vec model to generate 300-dimensional word vectors for the optimized Chinese word segmentation results in step S32 to capture the semantic similarity between terms, and then apply the same K-means clustering algorithm as in step S31 to perform unsupervised clustering on the word vectors, set the number of clusters K = 10, and iterate until the within-cluster distance converges.
[0050] In a specific embodiment, in the step of S52, align the structured item set with the association rule set, and generate a semantically unified and concise rule set through mapping, merging, and recalculation; Update label encoding through rule item mapping: Based on the secondary clustering results of step S51, record the attribution relationship from the old label to the new label, replace the original items (such as "C1", "C2") in the association rules with the new labels after secondary clustering (such as "C1-2"), and establish a label mapping table; traverse the antecedent ( ) and consequent ( ) of each rule, and replace the old labels according to the mapping table; Eliminate redundant rules through rule merging: If the antecedents and consequents of multiple rules are exactly the same, merge them into one rule, accumulate the support count, and recalculate the support, confidence, and lift metrics; Delete rules whose metrics do not meet the standards due to label merging or data update, and its filtering conditions are: Support: lower than the threshold (0.03); Confidence level: lower than the threshold (0.85); Lift: ≤ 1 (no practical relevance); Therefore, rules with a confidence level ≥ 0.85 and a lift > 1 are output to guide the production scheduling and resource optimization of final assembly production.
[0051] Example: As Figure 4 shown, the confidence level of the merged rule {C1-2} ⇒ {Station 200B} is 1.0, and the lift is 1.0 / 0.02 = 50 (assuming the global support of Station 200B is 0.02), and this rule is retained; if the lift of a certain rule = 0.8 (< 1), it is excluded.
[0052] The present invention provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor; the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, the processor can execute the method for mining aircraft final assembly information association rules based on fine-tuning the large language model as described above.
[0053] The present invention provides a computer-readable storage medium, which stores computer instructions for causing a processor to implement the method for mining aircraft final assembly information association rules based on fine-tuning the large language model as described above when executed.
[0054] In summary, through the fine-tuning of the large language model combined with association rule mining, the present invention significantly enhances the model's recognition and information extraction capabilities for aviation domain-specific terms. The introduction of the low-rank matrix enables the model fine-tuning process to only adjust a small number of parameters, which not only maintains the general knowledge of the pre-trained model but also improves the accuracy of aviation professional tasks, effectively reducing the training complexity and resource consumption. The in-depth combination of the mining algorithm effectively improves the accuracy of information extraction and the depth of production factor association analysis, thereby greatly improving the recognition accuracy of the association relationship between production factors and providing reliable data support and decision-making basis for the optimization of the final assembly process.
[0055] Based on the ideal embodiments of the present invention as inspiration, through the above description, for those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to include all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.
[0056] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for mining aircraft final assembly information association rules based on fine-tuning large language models, characterized in that, It includes the following steps: S1. Collect multi-source production factor data from aircraft final assembly process documents and perform preprocessing; S2. Perform domain adaptation based on the low-rank adaptation large language model fine-tuning method to enhance the professional term recognition ability; S3. Extract structured element information through the fine-tuned large language model and construct a semantically optimized project set; S4. Mine the association rules between production factors and final assembly links based on the association rule mining algorithm; S5. Combine word vector clustering and semantic verification to optimize and output an association rule set with a confidence level ≥ 0.
85.
2. The aircraft final assembly information association rule mining method based on fine-tuning a large language model according to claim 1, wherein: In the step of S1, the multi-source production factor data includes structured and unstructured data, and includes the following sub-steps: S11. Extract the AO number, station number, and position number from the structured data and perform standardized coding; S12. Parse the material code and tooling equipment code from the semi-structured data; S13. Extract professional terms, material codes, and tooling equipment codes from the unstructured text through natural language processing technology, and perform data cleaning based on regular expressions, domain dictionaries, and edit distances.
3. The aircraft final assembly information association rule mining method based on fine-tuning the large language model according to claim 2, characterized in that: In the step of S13, the extraction of the unstructured text uses a large language model to extract from the process steps described in natural language, and the input format is defined as: ; Among them, is the th process; is the AO number of process ; is the station number where process is located; is the position number where process is located; is the material required for process , is the set of material numbers; is the tooling equipment required for process ; is the set of numbers; is the set of key process information of process .
4. A method for mining aircraft final assembly information association rules based on fine-tuning large language models according to claim 1, characterized in that: In the step of S2, it includes the following sub-steps: S21. Insert a low-rank adapter module into the Transformer layer of the pre-trained large language model, and adjust the model parameters through a low-rank matrix; The Transformer is a model architecture, and the Transformer layer includes a self-attention mechanism and a feed-forward neural network, which are used to capture the context dependencies in the text; S22. Freeze the original model parameters, only train the low-rank adapter module, and perform fine-tuning using the cross-entropy loss function and the dynamic learning rate strategy; S23. Construct a fine-tuning data set based on the domain adaptation prompt template, and enhance the information extraction accuracy by combining role definitions, task statements, and output examples.
5. A method for mining aircraft final assembly information association rules based on fine-tuning large language models according to claim 4, characterized in that: In the step of S21, the specific insertion of the low-rank matrix is as follows: Insert a low-rank adapter module with a rank of beside the Query and Value matrices of each Transformer layer in the model, and the weight update formula is: ( , ); Among them, is the rank; and are the row and column dimensions of the original model weight matrix; is the set of real numbers; Transformer is the model architecture, and the Transformer layer contains the self-attention mechanism and the feed-forward neural network, which are used to capture the context dependencies in the text; Query is the query matrix in the self-attention mechanism, which is used to calculate the correlation weights between different positions in the input sequence; Value is the value matrix in the self-attention mechanism, which is used to apply the weights to the original input to generate the final context-aware representation.
6. A method for mining aircraft final assembly information association rules based on fine-tuning large language models according to claim 1, characterized in that: In the step of S3, it includes the following sub-steps: S31. Use a multi-agent mechanism to verify the format of the extraction results and match domain terms, and merge synonym classes and eliminate meaningless words through a clustering algorithm; The multi-agent mechanism includes: the first agent module: verify the format compliance of the extraction results; the second agent module: verify the compliance of terms and codes based on the aviation manufacturing domain term library; S32. Encode the clustering results into standard labels to generate a structured project set with consistent semantics.
7. A method for mining aircraft final assembly information association rules based on fine-tuning large language models according to claim 1, characterized in that: In the step of S4, the association rules between production factors and final assembly links are based on the FP-Growth algorithm, specifically including: frequent item set mining and generating association rules based on frequent item sets.
8. A method for mining aircraft final assembly information association rules based on fine-tuning large language models according to claim 7, characterized in that: The frequent itemset mining: count the occurrence frequency of each itemset in the transaction database. For the itemset , which is the set of all items, and all items are all different single basic items in the dataset, its support is calculated as: , ; Among them, is the total number of transactions; is an indicator function that is 1 when and 0 otherwise; is an itemset in the entire transaction database the number of occurrences; is a single transaction in the transaction database i.e., a set of data items; Keep the item sets with support degree ≥ the minimum threshold The set of frequent item sets is defined as: ; Build an index for each frequent item, recording its position and global frequency in the FP-tree, and the header table is a dictionary with the frequent item as the key , and the value is: ; Among them, is the support count of item ; is the linked list of nodes that point to all nodes in the FP-tree that contain ; Compress the transaction database into a prefix tree structure, where each node contains the following attributes: ; Among them, is the item represented by the current node; is the number of occurrences of the path from the root to this node; is the pointer to the parent node; is the dictionary of child nodes; Construction rule: For each transaction , insert it into the tree after sorting in descending order of the support of frequent items, and merge and count the paths sharing the same prefix; Conditional Pattern Base Generation: For item , its conditional pattern base is the set of all paths in the FP-tree that end with . Each path has the form: ; Recursive formula: For each item , recursively construct the conditional FP-tree , and mine its frequent item sets: ; Among them, represents a set of new frequent item sets generated from the conditional pattern bases of the current item during the recursive process of the FP-Growth algorithm; is to re-run the FP-Growth algorithm on the conditional pattern base of the item using the same minimum support threshold to mine frequent item sets; Finally, merge all to obtain the global frequent item sets; The association rule calculation: generate rules based on frequent item sets, and screen high-value rules through confidence and lift; The confidence represents the proportion of transactions that contain the antecedent and also contain the consequent in the transactions, that is, the accuracy of the rule: ; Among them, is the number of transactions that contain and ; is the number of transactions that contain ; Lift representation and The strength of the association between them. The higher the lift, the stronger the association between and is. If the lift value is greater than 1, the association rule is considered valuable: ; Among them, is the measurement rule of the association strength, reflecting and whether the co-occurrence of is the confidence of the rule ; is the support of the item set .
9. An electronic device, characterized in that: The electronic device includes: at least one processor, and a memory communicatively connected to at least one of the processors; the memory stores a computer program executed by the processor, and the computer program is executed by the processor so that the processor can execute the method for mining the aircraft final assembly information association rules based on the fine-tuned large language model according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the method for mining the aircraft final assembly information association rules based on the fine-tuned large language model according to any one of claims 1-8 when executed by a processor.
Citation Information
Patent Citations
Manufacturing process instruction standardization method based on process composition elements
CN111461912A
Gearbox front-middle shell assembly process and operation data structured management method and system
CN116468384A
PVB product quality association rule analysis method and system
CN116882822A
Model information interaction and integrated analysis method based on digital principal line
CN117252359A
Distributed FP-growth with node table for large-scale association rule mining
US20180107695A1
Cited By
Task processing method and device based on fine-tuning large language model, equipment and medium
CN120851126A
Multi-source heterogeneous data element disassembling and comparing method and related device
CN121168623A