Industrial MBSE model generation and evaluation method based on large language model
By constructing industrial MBSE model generation and evaluation methods, and using large language models to automatically generate MBSE models, the problems of manual intervention dependence and cross-view consistency in the existing technology are solved, the intelligence and consistency verification of MBSE models are realized, and the digital design of complex systems is supported.
Patent Information
- Application Number
- CN202510435198.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-18
AI Technical Summary
The existing MBSE modeling methods rely too much on manual intervention, and the level of intelligence is limited, making it difficult to meet the needs of multi-dimensional consistency verification, and there are technical bottlenecks in cross-view model collaboration.
Through industrial MBSE model generation and evaluation methods based on large language models, pneumatic models and component design rules are collected and organized, a dedicated enhanced knowledge base is built, user needs are converted into standard code using prompt word engineering technology, and real-time optimization is combined with function calls and simulation algorithms, and design consistency is ensured through anchor entity retrieval.
It realizes the automated generation and evaluation of MBSE models, improves the intelligence level of design, ensures the global consistency and logical consistency of the model, and supports the digital design and verification of complex systems.
Smart Images

Figure CN120337918A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of neural networks, specifically an industrial MBSE model generation and evaluation method based on large language models. Background Art
[0002] Model-based systems engineering (MBSE) is a method for efficiently managing the product design and manufacturing process by uniformly digitally modeling various dimensions of the product's entire life cycle. Existing MBSE model modeling methods rely too much on manual intervention and have limited intelligence. At the same time, the ability of large language models to generate and parse domain-specific modeling languages is not yet perfect, and there are technical bottlenecks in cross-view model collaboration, making it difficult to meet the requirements of multi-dimensional consistency verification and facing major technical challenges. Summary of the Invention
[0003] In view of the key and difficult problems in the existing industrial system design, such as relying on development experience, too long design cycle, and lack of digital assets, the present invention proposes an industrial MBSE model generation and evaluation method based on large language models. By specifying the design output requirements and target parameters, the large model automatically calls optimization algorithms and simulation tools through function call technology to generate the MBSE model, minimizing the dependence on the experience of designers. During the tuning process, the large model analyzes the generated MBSE model in the application scenario and test data through an automated analysis tool, and confirms the consistency between the design output and the actual requirements. Finally, a consistency check is performed on the output through anchor point entity retrieval technology to ensure the standardization of the design output.
[0004] The present invention is implemented through the following technical solutions:
[0005] The present invention relates to an industrial MBSE model generation and evaluation method based on large language models. It collects and organizes aerodynamic models, component design rules, and relevant basic design parameters, and formats the rules applied in MBSE to form a dedicated enhanced knowledge base in the industrial field. Through prompt engineering technology, the requirements described in the user's natural language are converted into standard MBSE model codes based on the large language model. By function call technology, optimization algorithms and simulation algorithms are selected, and the model is optimized in real time according to the design goal. The generated MBSE model is applied to the design and verification test of the propulsion device to determine usability and model consistency. Through automated configuration based on requirements and components, the reconstruction and acceleration of digital design are realized, specifically including:
[0006] Step 1. Knowledge base construction, including: data cleaning, tokenizer preprocessing, inverted index annotation, text embedding encoding, and knowledge graph construction. Industrial data knowledge described in text is subjected to part-of-speech tagging and stop word filtering methods through a language model-based tokenization method in the module, and is decomposed into a set of several terms W = {w1, w2, …, w m}, where: each w i is a word. For each word w i , the stop word filtering rule f(w i ) and part-of-speech tagging p(w i ) are used to determine whether to retain the word, and the calculation process is The construction of the knowledge graph organizes entities, attributes, and relationships extracted from the text into a graph structure. Entities E = {e1, e2, …, e k} mentioned in the text, relationships R = {r1, r2, …, r l}, and attributes A = {a1, a2, …, a m}. The constructed graph structure is G = (V, E, R), where: V is the set of nodes, E is the set of edges, and they are stored in the graph database. Text data is transformed into semantic vectors through a pre-trained private industrial encoder, and the labeled industrial data knowledge is stored in the vector database, jointly forming a knowledge base with the graph database.
[0007] Step 2. Prompt construction, including: semantic analysis, requirement extraction, and domain knowledge supplementation to generate high-quality instruction text. Through a deep learning-based classification model f class , text T can be mapped to different requirement categories C = {c1, c2, …, c m}, and classification is completed using the maximum class probability , where: P(c i |T) = f class (T, c i ) is the conditional probability that text T belongs to category c i , and c i are experimental requirements such as design optimization, performance requirements, and material selection. After the original input requirements are transformed into standardized instruction text based on the extracted target intent, the instructions are formatted according to predefined templates and rules to avoid ambiguity and redundant information.
[0008] Step 3. MBSE model generation, including: MBSE view generation, global consistency check, and model view integration. Three steps are finally used to output a Figure 1 consistent MBSE model, specifically including:
[0009] 3.1. MBSE view generation: Receive the standardized instructions generated by the prompt word building module and the supplementary information provided by the knowledge base building module, and extract the core goals, constraints and priority information in the requirements. Since MBSE modeling has many different grammatical rules, and the general large language model lacks deep domain knowledge, the module provides clear grammatical rules for the large language model based on the received prompt words to support the generation of MBSE models under a unified grammatical structure: Use key-value pairs to describe the general grammatical elements in SysML, where: the key is the grammatical element in SysML, including diagram types such as behavior diagrams, requirement diagrams, and structure diagrams, and model elements such as executors and value types, that is, the set S = {s1, s2, ..., s k}; The value is a brief introduction and rule description of the key. i =Rule(s i ),fors i ∈S. We further introduce the few-sample learning technique to train a small number of similar cases D samples ={(T1,M1),)(T2,M2),…,(T n ,M n )} modeling requirements T i The modeling results M i As input, guide the large language model to have stronger emergence ability: use multimodal entity vectors to search for similar entities, and use the domain document vector library D doc Middle domain document block d i ∈D doc With entity , find the entities involved in the domain knowledge, and use the entity as an index to retrieve its model representation and question-answering history, where: each entity e j vector Where: d is the dimension of the vector space: using Computational Entity j and entity e k The similarity between them: v(e j ) and v(e k ) is entity e j and e k The semantic vector representation of ||v(e j )|| and||v(e k )|| is the L2 norm of the two vectors. When the current modeling object has other view contexts, the cosine similarity is calculated between the model representation containing the existing view and the model representation of the corresponding view of the candidate similar entity. The question and answer history can be analogous. When there is no view context, the entity obtained from the domain knowledge is regarded as a similar entity. K cases are used as the prompt word case for few-shot learning and spliced after the above prompt word, formalized
[0010] 3.2. Use the method of anchor entity matching and tree traversal for consistency checking, that is, establish a detection mechanism based on bidirectional mapping of anchor entities, and achieve global logical verification through tree structure traversal and feature matching across iterative views, specifically including:
[0011] 3.2.1 Convert the view set generated in the historical iteration stage and the current iteration view G i (V i ) into a tree - like hierarchical structure v prev and v i . The end - point nodes v ∈ v of the view correspond to an element entity of the view. For each element v ∈ v of the current view, for the feature vector f of each entity v ∈ V, calculate its similarity value S(f, f′) with the historical view element set V j .
[0012] 3.2.2 Denote τ as the defined similarity threshold. When S(f, f′) ≥ τ, it is considered that the entity v of the current view is an anchor entity. According to this definition, form a set of anchor entity pairs
[0013] 3.2.3 When that is, no anchor entity is found, it indicates that there is a consistency problem with the MBSE view generated in step 2, and then the conflict information is passed to the large - model for iterative optimization. When that is, after the anchor entity is found, use the bidirectional tree - structure traversal algorithm, and perform breadth - first expansion detection with the anchor pair as the root node.
[0014] 3.2.4 During the synchronous traversal process, the system compares the feature vectors of the corresponding nodes of the two structure trees in real - time. When it is found that the entities corresponding to the nodes are different during the traversal process, it is considered that there is a conflict between the two nodes, that is where: the node information label structure conflict node pair (n, n′), and the conflict result are passed to the large - model as input parameters for iterative generation together, otherwise continue traversing until there are no conflicts in all global nodes.
[0015] 3.3. Model view integration. After ensuring the consistency between views, find the anchor entities A = {a1, a2,..., a m} shared among multiple views in the view set {V1, V2,..., V k}. Each anchor entity a i ∈ A exists in at least two views simultaneously. Denote the anchor entity a i ∈ A in view V1 as the reference entity, and then it is necessary to find the associated entity a i corresponding to the reference entity a i' ∈ E2, forming a bidirectional association mapping Each anchor entity a i ∈ A can have a corresponding entity in view V2 in view V1, map(a i ) = a' i . The corresponding relationship of anchor entities in the view set satisfies map(a i ) = a' i . After the corresponding relationship is established, continue to align other nodes and edges in the view. Specifically, for each pair of associated entity pairs (a i , a' i ), after that, check and align by calculating the similarity of other nodes and edges in view V1 and view V2. The aligned anchor entities ensure that similar entities in each view can be correctly mapped during the integration process, ensuring the global consistency of the generated MBSE model.
[0016] Step 4. Result simulation evaluation: Based on the function call method, through the full-process closed-loop of requirement dependency parsing, parameter performance simulation, simulation result evaluation, and optimization suggestion generation, to support the continuous optimization of the system and improve the intelligent design ability, specifically including:
[0017] 4.1. Requirement dependency parsing. The core goal of requirement dependency parsing is to decompose the performance requirements that the digital prototype design needs to meet into several subtasks, and map the requirement indicators corresponding to each subtask into the objective function and constraint conditions represented by operational research mathematics. After the design requirements are jointly composed of the general specifications of the private knowledge base and the requirement view of the MBSE model, based on the semantic graph neural network, parse the MBSE requirement graph and parameter graph, extract the performance indicators, constraint boundaries, and association relationships in the node attributes, and use the private knowledge base to align and provide semantic completion to construct a quadruple (Q, Ω, Φ, Ψ), where: Q = {q i} represents the atomized requirement elements extracted from the requirement graph; Ω is the domain constraint rule base provided by the knowledge base (such as the material strength-weight relationship formula); Φ: Q × Ω → F generates the objective function set; Ψ: Q × Ω → C generates the constraint conditions. Each objective function F i = f(D i ) ∈ F in the objective function set corresponds to the constraint condition C i : g i (D i ) ≤ 0 ∈ C: Construct g0 as a hard constraint, and other requirements as soft constraints. Each subtask corresponds to a specific simulation module. Since the simulation process of some subtasks may depend on the simulation results of other subtasks, the module uses a tree-type dependency parsing method that combines topological sorting and hierarchical traversal to construct a requirement dependency tree to represent the dependency relationship between subtasks, and the performance simulation part starts running from the root node of the dependency tree.
[0018] 4.2. Simulation Environment Setup. For different simulation tasks, different simulation environments will be designed according to the design parameters. First, extract the corresponding material properties M = (ρ, E, ν, k, σ yield ) from the knowledge base, and define the external environment and operating conditions of the simulation model. The simulation environment boundaries include but are not limited to mechanical boundaries, thermal boundaries, and fluid boundaries. Suppose we conduct a thermal analysis, and the boundary conditions set may be to set the temperature of the structure surface as T0; apply an external heat flux at a certain position The analysis results of all material properties and boundary conditions will be integrated into the simulation model to form a complete physical model.
[0019] 4.3. Solver Invocation: Based on the simulation environment, first discretize the continuous physical domain into discrete grid cells to realize the modeling of complex geometries. After a three-dimensional object space domain Ω, this domain will be divided into multiple small, finite-sized grid cells ω, and ω are arranged and connected in different ways: where: ω i is the i-th grid cell, and N is the total number of grid cells. After the refined grid may have problems of being too distorted or non-uniform, Laplacian smoothing is used to improve the shape factor of the grid. When denoting the grid node as x i , and its set of adjacent nodes is then the new node position is updated as When a node x i moves, all grid cells connected to this node will be affected. Since the grid cells are triangular cells, the shape factor is where: A is the area of the triangle, expanded as When the node x i undergoes Laplacian smoothing, the new position x′ i will change the triangle area and side lengths, thus affecting Q. To optimize the overall shape factor of the grid, define the objective function represents the average value of the shape factors of all grid cells, where: N T is the total number of cells in the grid. The goal of Laplacian smoothing is to maximize Q avg . To achieve this goal, based on the gradient, update the node positions. For the node x i , its update formula is where: α is the learning rate, is the gradient of the objective function, expanded as: Repeatedly update the node positions until Q avgThe change is less than the threshold value or reaches the upper limit of the number of iterations. After the mesh generation is completed, the physical field is solved numerically on the discretized network by discretizing the continuous differential equations into algebraic equations for solution. Taking the elastic problem as an example, the equilibrium equation The weak form of this problem on the spatial domain Ω is ∫ Ω B T DBudΩ = ∫ Ω B T fdΩ + ∫ Γ N T tdΓ, where: u is the displacement vector, B is the strain-displacement matrix, D is the stiffness matrix of the material, f is the body force, t is the external force on the boundary, and Γ is the boundary. Denote ∫ Ω B T DBdΩ as the global stiffness matrix K, and ∫ Ω B T fdΩ + ∫ Γ N T tdΓ as the force vector F, then the original formula can be converted into the matrix form Ku = F. In the finite element method, the stiffness matrix and the force matrix of each element are derived through a similar weak form. Based on the connection relationship between nodes, by combining the stiffness matrices and force matrices of all elements, the global stiffness matrix and the force vector After the discretized equation is a system of linear equations, depending on the complexity, the Gaussian elimination method or the conjugate gradient method is used to solve it. The solution is passed to the next simulation module that depends on this subtask.
[0020] 4.4. Evaluation of simulation results. After obtaining the global results of the simulation solver, based on the large model, a set of performance indicators P = f(σ, u, T, v, etc.) are defined according to the design objectives. Each design objective will have a clear performance requirement. The model compares the simulation results and the performance indicators to give an evaluation report on the digital prototype. When the digital prototype meets the performance indicators, it is judged that the design scheme meets this requirement; otherwise, optimization is required.
[0021] 4.5. Generation of model optimization suggestions. When the evaluation of the simulation results fails, optimization suggestions will be generated as a reference to assist in guiding the next iterative optimization direction of the digital prototype. For example, in the module definition view, adjust the geometric design, add support structures, change the material thickness, optimize the shape, etc.; in the internal module view, adjust the load distribution, modify the heat flow or fluid pressure, etc.; in the parameter view, put forward improvement suggestions for the variables in the manufacturing process. Finally, a result report containing the simulation results and evaluation conclusions is generated. Technical effects
[0022] The present invention realizes a cross-iteration visual Figure 1 consistency checking method through anchor entity matching and bidirectional tree structure traversal. First, the historical iteration view set and the current view are respectively converted into a tree-like hierarchical structure. The similarity between the current view entity and the historical entity is calculated through the feature vector space, and the anchor entity pairs with an inheritance relationship are identified according to the preset threshold τ to form a cross-version semantic association benchmark. When there are anchor entities, bidirectional breadth-first traversal is implemented with the anchor pairs as the root nodes, and the feature vectors and structural relationships of the corresponding nodes of the two trees are compared in real time. When node feature deviation or topological relationship conflict is detected, a structured conflict parameter is generated and fed back as an optimization input to the large model for iterative correction; if there are no anchor entities, a global consistency warning is directly triggered.
[0023] The present invention provides a method for constructing a requirements - operations research representation through semantic-driven model collaborative parsing. This method integrates the domain specifications of the MBSE requirements view and the private knowledge base, jointly parses the requirements diagram and the parameter diagram, extracts performance indicators, constraint boundaries, and their topological relationships, and forms an atomic requirements element set Q = {q i}; Subsequently, combined with the engineering constraint rules Ω in the knowledge base, the objective function set F = {∪ i f(D i )} and the constraint condition set C = {∪ i g i (D i ) ≤ 0} of the operations research model are constructed through the objective function generator Φ: Q×Ω→F and the constraint condition generator Ψ: Q×Ω→C respectively, where the key constraints are defined as hard constraints g0, and the remaining requirements are used as soft constraints.
[0024] The present invention realizes the dynamic programming of the simulation execution order by reversely constructing a dependency tree. First, each subtask is abstracted into a tree node, and a directed graph is formed by analyzing the mapping relationship of the simulation input and output parameters to establish parent-child dependency connections. Subsequently, the initial terminal nodes are identified based on depth-first search, and recursive hierarchical division is performed starting from these nodes. When all the parent nodes of a certain node are included in the current level, the node is classified into the next processing level, and finally a hierarchical tree structure arranged according to the dependency depth is generated. The nodes in the same level represent that the simulation tasks can be executed in parallel in this simulation stage.
[0025] The present invention realizes multi-granularity consistency verification from local to global, ensuring logical self-consistency and cross-view semantic coherence during the evolution of model versions. Through the dual mechanisms of anchor entity matching and cross-view topological traversal, the global consistency guarantee ability during the iterative process of system engineering models is significantly improved. Establishing a dynamically evolving semantic association network can automatically identify the logical relationships between different views and locate cross-view conflicts caused by possible hallucination problems during the generation of large models. Through the synergistic effect of similarity calculation in the feature vector space and bidirectional tree structure traversal, entity-level attribute deviations are captured, and at the same time, structural topological relationship contradictions are discovered, forming structured conflict description parameters to drive model iterative optimization. Meanwhile, the global warning mechanism without anchor entities effectively prevents the risk of model evolution interruption, realizing automated conflict detection and closed-loop correction under multi-dimensional specification constraints.
[0026] The present invention realizes the intelligent conversion of system engineering requirements into a mathematically model that can be formally expressed. Semantic extraction is performed on the MBSE requirement view and parameter diagram, and a quadruple generation mechanism (Q, Ω, Φ, ψ) guided by knowledge rules is established to integrate engineering constraints and domain knowledge into an operational research model with dynamically divided hard and soft constraints. Breakthroughs have been achieved in the accuracy of demand mathematical representation, the integration ability of multi-source constraints, and the parsing efficiency of task dependencies, providing an automated and traceable demand parsing solution for complex system design and a modeling system with self-consistent verification ability for complex system design.
[0027] The present invention upgrades the traditional sequential execution mode to a hierarchical parallel architecture. First, a dependency directed graph is reversely constructed through parameter mapping relationships, and the depth-first search is used to accurately locate the terminal nodes as the initial level. The recursive hierarchical division algorithm is adopted to classify the child nodes into the next processing level only when all the parent nodes are ready, generating a tree-like hierarchical structure with strict dependency relationships. The execution order is dynamically adjusted according to the dependency depth to improve the utilization rate of simulation resources, providing an extensible and adaptive task scheduling solution for the dynamic collaborative simulation of complex systems. Brief Description of the Drawings
[0028] Figure 1 is the method framework diagram of the present invention;
[0029] Figure 2 is the system structure diagram of the embodiment of the present invention. Detailed Embodiments
[0030] Such as Figure 2As shown in the figure, this embodiment relates to an implementation architecture of an industrial MBSE model generation and evaluation system, including: a Web application layer, a business processing layer, and a data storage layer, where: the Web application layer serves as an interaction interface between the system and users, receives the data input by users, and transfers the data to the subsequent knowledge processing layer and modeling and reasoning layer; the front-end application receives two types of data, structured data and unstructured data, and the Web application layer transfers the modeling requirements input by users as instructions for large model reasoning to the business processing layer.
[0031] The structured data includes process parameter data, equipment operation data, and quality inspection and maintenance data, and the Web uploads the data through Restful Api.
[0032] The unstructured data includes product maintenance manuals, equipment operation instructions, and production design reports, and is processed through methods such as document upload, text parsing, and data annotation.
[0033] As Figure 1 shown in the figure, the business processing layer includes: a knowledge base construction module, a prompt word construction module, an MBSE model generation module, and an MBSE model evaluation module, where: the knowledge base construction module performs unit standardization and data normalization processing on structured data such as process parameters and equipment operation data based on pandas; uses natural language processing technology based on BERT to convert text data into vector form; the prompt word construction module extracts key information from the requirements and instructions input by users and converts it into standardized task instructions to provide accurate input for subsequent large model reasoning; the MBSE model generation module converts the processed task instructions into actual reasoning and modeling processes. For more complex or professional requirements.
[0034] When processing the original text, the knowledge base construction module preprocesses the text using the SpaCy tokenizer to split the text into words or sub-words; uses the inverted index annotation technology of the ElasticSearch engine to establish the index relationship between keywords and documents; for more complex text data, the knowledge base construction module identifies important entities in the text through a natural language processing model and identifies the relationships between entities through semantic analysis to form a knowledge graph.
[0035] During the implementation process, the prompt word construction module analyzes the user's industrial domain background, model format requirements, and personal interest preferences based on the LangChain framework, judges the context and requirements of the task, and extracts domain background knowledge related to the task from the knowledge base through a hybrid retrieval of ElasticSearch, vector knowledge, and knowledge graph.
[0036] The described MBSE model generation module transforms requirements by decomposition into actionable tasks with clear goals and constraints, and adopts a hybrid cloud-edge inference deployment method during implementation. For core sensitive design requirements, they are analyzed by a local inference engine. After the inference results are generated, a post-processing unit combines them into a final design solution.
[0037] The actionable tasks include:
[0038] a) The MBSE model evaluation module transforms the inference results into a physical scenario simulation and evaluates the feasibility of the design based on the simulation results. During implementation, while constructing a simulation dependency tree, two sets of indexes, namely a hierarchical sequence index table and an adjacency matrix, are constructed. The hierarchical sequence index table records the processing level where the node is located, and the adjacency matrix stores cross-level dependency relationships. When the simulation is executed, a breadth-first strategy is adopted to start the simulation modules layer by layer in ascending order of the hierarchical sequence numbers. Within the same level, the execution priority is dynamically adjusted based on the weight values in the adjacency matrix. When it is detected that a pre-task is completed, the successor task that depends on the result is automatically triggered to execute.
[0039] b) The MBSE model evaluation module calls the API simulation interfaces provided by external simulation toolchains such as SimuLink according to the design solution obtained by inference to perform simulations of various physical scenarios such as fluid simulation, thermodynamics simulation, and rigid body simulation. Based on the quantitative comparison between the simulation results and the preset performance indicators, the MBSE model evaluation module matches the performance thresholds corresponding to each design goal, and uses a weighted scoring model to calculate the goal achievement degree. When the pass rate of all key indicators is greater than the set threshold, the solution is determined to be feasible; otherwise, an optimization process is triggered. In the optimization process, for the unqualified indicators, trace back to a specific view layer of the MBSE model: recommend a geometric topology optimization solution in the defined view of the MBSE model evaluation module; propose a physical parameter iteration strategy in the internal module view; guide the improvement of manufacturing processes in the parameter view. After that, integrate the simulation data curves, index deviation degrees, and optimization paths to generate a structured result report, which includes a performance comparison matrix and a confidence evaluation value for each design iteration version.
[0040] The described data storage layer stores and manages various data resources in the project, including similar model cases, domain background knowledge, and component entity relationships. Among them: domain background knowledge is the most basic data resource in the project, including process parameter data, equipment operation data, quality inspection and maintenance data, and unstructured data input by users. During implementation, the data storage layer uses MongoDB for data storage. In scenarios with a large amount of corpus and complex queries, ElasticSearch is used as a search engine. After the similar model case library is semantically encoded, Qdrant is selected as the vector storage library. For component entity relationship data, the present invention uses Neo4j to store and manage the domain knowledge graph formed after data extraction and processing.
[0041] Table 1 Comparison of Technical Characteristics
[0042] The present invention shows significant advantages in terms of reliability, effectiveness, and adaptability. In terms of reliability, the MBSE model evaluation module of the present invention ensures the reliability of the MBSE model through closed-loop simulation and iterative optimization methods. Requirement dependency parsing decomposes performance requirements into mathematical objective functions and constraint conditions, and constructs a quadruple (Q, Ω, φ, ψ) by combining knowledge base semantic completion to achieve a reliable mapping from requirements to mathematical models. Through dynamic grid optimization and accurate numerical solution, the calculation accuracy of the physical field is guaranteed. Finally, based on performance index quantitative evaluation and closed-loop feedback mechanism, multi-view optimization suggestions are generated for unqualified results, and the design parameters and structures are iteratively corrected to avoid error accumulation and ensure the reliability and physical authenticity of the model under multi-field constraints such as mechanics and thermotics.
[0043] In terms of effectiveness, the MBSE model generation module of the present invention ensures the high accuracy and stability of the model through automated global consistency checking and intelligent iterative optimization mechanisms. When facing complex cross-view constraints, through combining the generation process context of multiple views and global consistency checking, automated management from requirement analysis to model integration is achieved. Whenever model conflicts or inconsistencies are found, the system can automatically feedback the conflict information to the large model for iterative correction to ensure that each generated model version meets the design requirements and avoid human errors.
[0044] In terms of adaptability, the knowledge base construction module of the present invention achieves adaptability through dynamic data processing and multi-modal knowledge fusion methods. Industrial texts are dynamically cleaned using configurable stop word filtering rules and part-of-speech tagging, and an extensible set of terms W is retained. Heterogeneous data is transformed into semantic vectors through pre-trained industrial encoders to support flexible embedding of terms in different fields. The dual-modal storage mechanism of the knowledge graph G=(V, E, R) and the vector database retains the structured information and implicit semantics of entity relationships, is compatible with the diverse requirements of industrial scenarios, and adapts to complex and changing industrial requirements and emerging terms.
[0045] The above specific embodiments can be locally adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific embodiments, and all implementation solutions within its scope are subject to the present invention.
Claims
1. An industrial MBSE model generation and evaluation method based on large language models, characterized in that, By collecting and organizing aerodynamic models, component design rules, and relevant basic design parameters, and formatting the rules applied in MBSE to form a dedicated enhanced knowledge base for the industrial field; through prompt engineering techniques, converting the requirements described in natural language by users into standard MBSE model codes based on large language models, selecting optimization algorithms and simulation algorithms through function call techniques, and performing real-time optimization on the model according to the design objectives, and applying the generated MBSE model to the design and verification test of the propulsion device to determine usability and model consistency.
2. The method for generating and evaluating an industrial MBSE model based on a large language model according to claim 1, characterized in that, The described dedicated enhanced knowledge base is constructed through the following steps: data cleaning, preprocessing by a tokenizer, inverted index annotation, text embedding encoding, and knowledge graph construction. The industrial data knowledge described in text is subjected to part-of-speech tagging and stop-word filtering methods through a language model-based tokenization method in the module, and decomposed into a set of several terms W = {w1, w2, …, w m}, where each w i is a word. For each word w i , the stop-word filtering rule f(w i ) and part-of-speech tagging p(w i ) are used to determine whether to retain the word. The calculation process can The construction of the knowledge graph is achieved by extracting entities, attributes, and relationships in the text and organizing them into a graph structure. The entities mentioned in the text are E = {e1, e2, …, e k}, the relationships are R = {r1, r2, …, r l}, and the attributes are A = {a1, a2, …, a m}. The constructed graph structure is G = (V, E, R), where: V is the set of nodes, E is the set of edges, and they are stored in a graph database. The text data is passed through a pre-trained private industrial encoder, and the annotated industrial data knowledge is transformed into semantic vectors and stored in a vector database, jointly forming a knowledge base with the graph database.
3. The method for generating and evaluating an industrial MBSE model based on a large language model according to claim 1, characterized in that, The described prompt engineering techniques include: semantic analysis, requirement extraction, and domain knowledge supplementation to generate high-quality instruction texts, through a deep learning-based classification model f class , the text T is mapped to different requirement categories C = {c1, c2, …, c m}, and classification is completed using the maximum class probability , where: P(c i |T) = f class (T, c i ) is the conditional probability that the text T belongs to the category c i , c i is design optimization, performance requirements, material selection. After the original input requirements are transformed into standardized instruction texts based on the extracted target intent, the instructions will be formatted according to predefined templates and rules to avoid ambiguity and redundant information.
4. The method for generating and evaluating an industrial MBSE model based on a large language model according to claim 1, characterized in that, For the MBSE model described above, through generating MBSE views, global consistency checking, and model view integration, an MBSE model with consistent views is finally output.
5. The method for generating and evaluating an industrial MBSE model based on a large language model according to claim 1 or 4, characterized in that, The MBSE model described above is generated through the following methods, specifically including: 3.
1. MBSE view generation: Receive the standardized instructions generated by the prompt word building module and the supplementary information provided by the knowledge base building module, extract the core goals, constraints and priority information in the requirements. Since MBSE modeling has many different grammatical rules and the general large language model lacks deep domain knowledge, the module provides clear grammatical rules for the large language model based on the received prompt words to support the generation of MBSE models under a unified grammatical structure: Use key-value pairs to describe the general grammatical elements in SysML, where: the key is the grammatical element in SysML, including diagram types and model elements, that is, the set S = {s1, s2, ..., s k }; The value refers to a brief introduction and rule description of the key v i =Rule(s i ),fors i ∈S, we further introduce the few-sample learning technique to samples ={(T1,M1),)(T2,M2),…,(T n ,M n )} modeling requirements T i The modeling results M i As input, guide the large language model to have stronger emergence ability: use multimodal entity vectors to search for similar entities, and use the domain document vector library D doc Middle domain document block d i ∈D doc With entity , find the entities involved in the domain knowledge, and use the entity as an index to retrieve its model representation and question-answering history, where: each entity e j vector Where: d is the dimension of the vector space: using Computational Entity j and entity e k The similarity between them, where: v(e j ) and v(e k ) is entity e j and e k The semantic vector representation of ||v(e j )|| and||v(e k )|| is the L2 norm of the two vectors. When the current modeling object has other view contexts, the cosine similarity calculation is performed between the model representation containing the existing view and the model representation of the corresponding view of the candidate similar entity. The question-answering history can be analogous. When there is no view context, the entity obtained from the domain knowledge is used as a similar entity, and the K cases are used as the prompt words case for few-shot learning, which are spliced after the above prompt words. 3.
2. Global consistency check. The present invention proposes a method of using anchor entity matching and tree traversal for consistency checking. Specifically, first, the view set generated in the historical iteration stage and the current iteration view G i (V i ) are transformed into a tree-like hierarchical structure and the end nodes of the view correspond to an element entity of the view. For each element of the current view For the feature vector f of each entity v ∈ V, the system calculates its similarity value S(f, f′) with the historical view element set V j in the feature vector space. Let τ be the defined similarity threshold. When S(f, f′) ≥ τ, the entity v of the current view is considered an anchor entity. According to this definition, a set of anchor entity pairs is formed. The anchor entity set A is defined as When that is, no anchor entity is found, it indicates that there is a consistency problem with the MBSE view generated in step 2. Then, the conflict information is passed to the large model for iterative optimization. When that is, after an anchor entity is found, the bidirectional tree structure traversal algorithm is used, and the anchor pair is used as the root node to perform breadth-first expansion detection. During the synchronous traversal process, the system compares the feature vectors of the corresponding nodes of the two structure trees in real time. When it is found that the entities corresponding to the nodes are different during the traversal, it is considered that there is a conflict between the two nodes, that is where: the node information label structure conflict node pair (n, n′), and the conflict result are passed to the large model as input parameters for iterative generation together, otherwise continue traversing until there are no conflicts in all global nodes; 3.
3. Model view integration. After ensuring the consistency among views, anchor entities A = {a1, a2, …, a m} that are shared among multiple views will be found in the view set {V1, V2, …, V k}. Each anchor entity a i ∈ A exists in at least two views simultaneously. After designating the anchor entity a i ∈ A in view V1 as the reference entity, it is necessary to find the associated entity a i ′ ∈ E2 in view V2 that corresponds to the reference entity a i to form a bidirectional association mapping Each anchor entity a i ∈ A has a corresponding entity map(a i ) = a′ i in view V2. The corresponding relationships of the anchor entities in the view set satisfy map(a i ) = a′ i . Based on the established corresponding relationships, continue to align other nodes and edges in the views. Specifically, for each pair of associated entity pairs (a i , a i ′), verify and align them by calculating the similarity of other nodes and edges in view V1 and view V2. The aligned anchor entities ensure that similar entities in each view can be correctly mapped during the integration process, guaranteeing the global consistency of the generated MBSE model.
6. The method for generating and evaluating an industrial MBSE model based on a large language model according to claim 1, characterized in that, The design and verification test of the propulsion device includes: a full-process closed-loop generated based on the function call method through requirement dependency parsing, parameter performance simulation, simulation result evaluation, and optimization suggestion generation to support continuous system optimization and improve the intelligent design ability.
7. The method for generating and evaluating an industrial MBSE model based on a large language model according to claim 1 or 6, characterized in that, The design and verification test of the propulsion device specifically includes: 4.
1. Requirement Dependency Analysis: Decompose the performance requirements that need to be met in the digital prototype design into several subtasks, and map the requirement indicators corresponding to each subtask into the objective function and constraint conditions represented by the mathematical characterization of operations research. After the design requirements are jointly composed of the general specifications of the private knowledge base and the requirement view of the MBSE model, parse the MBSE requirement graph and parameter graph based on the semantic graph neural network, extract the performance indicators, constraint boundaries and association relationships in the node attributes, and use the private knowledge base alignment to provide semantic completion to construct a quadruple (Q, Ω, Φ, Ψ), where: Q = {q i} represents the atomized requirement elements extracted from the requirement graph; Ω is the domain constraint rule base provided by the knowledge base; Φ: Q × Ω → F generates the objective function set; Ψ: Q × Ω → C generates the constraint conditions. Each objective function F i = f(D i ) ∈ F in the objective function set corresponds to the constraint condition C i : g i (D i ) ≤ 0 ∈ C: Construct g0 as a hard constraint and other requirements as soft constraints. Each subtask corresponds to a specific simulation module. Since the simulation process of some subtasks may depend on the simulation results of other subtasks, the module will generate a requirement dependency tree to represent the dependency relationship between subtasks. The performance simulation part starts running from the root node of the dependency tree; 4.
2. Simulation Environment Setup: For different simulation tasks, different simulation environments will be designed according to the design parameters. First, extract the corresponding material properties M = (ρ, E, v, k, σ yield ) from the knowledge base, and define the external environment and operating conditions of the simulation model. The simulation environment boundaries include but are not limited to mechanical boundaries, thermal boundaries, and fluid boundaries. Suppose we conduct a thermal analysis, and the boundary conditions set may be to set the temperature of the structure surface as T0; apply an external heat flux at a certain position The analysis results of all material properties and boundary conditions will be integrated into the simulation model to form a complete physical model; 4.3 Solver Call: Based on the simulation environment, first, the continuous physical domain is discretized into discrete grid cells to model complex geometries. After a three-dimensional object space domain Ω, this domain is divided into multiple small, finite-sized grid cells ω, which are arranged and connected in different ways: where: ω i is the i-th grid cell, N is the total number of grid cells. After the refined grid may have problems of excessive distortion or non-uniformity, Laplacian smoothing is used to improve the shape factor of the grid. When the grid nodes are denoted as x i , and its set of adjacent nodes is then the updated position of the new node is When a node x i moves, all the grid cells connected to this node will be affected. Since the grid cells are triangular cells, the shape factor is where: A is the area of the triangle, expanded as When the node x i after Laplacian smoothing, the new position x′ i will change the triangle area and side length, thus affecting Q. To optimize the overall shape factor of the grid, the objective function represents the average value of the shape factors of all grid cells, where: N T is the total number of cells in the grid. The goal of Laplacian smoothing is to maximize Q avg . To achieve this goal, based on the gradient, the node positions are updated. For the node x i , its update formula is where: α is the learning rate, is the gradient of the objective function, expanded as: The node positions are repeatedly updated until the change in Q avg is less than the threshold or the iteration count limit is reached. After completing the grid division, the physical field is solved numerically on the discretized network by discretizing the continuous differential equations into algebraic equations for solution; 4.
4. Simulation result evaluation: Based on the large model, a set of performance indicators P = f(σ, u, T, v, etc.) are defined according to the design objectives. Each design objective will have a clear performance requirement. The model compares the simulation results with the performance indicators and gives an evaluation report on the digital prototype. When the digital prototype meets the performance indicators, it is determined that the design scheme meets this requirement; otherwise, optimization is required. 4.
5. Generation of model optimization suggestions: When the simulation result evaluation fails, optimization suggestions will be generated as a reference to assist in guiding the next iterative optimization direction of the digital prototype and generating a result report containing the simulation results and evaluation conclusions.
8. An industrial MBSE model generation and evaluation system based on a large language model for implementing the method according to any one of claims 1-7, characterized in that, Including: A knowledge base construction module, a prompt construction module, an MBSE model generation module, and an MBSE model evaluation module, where: the knowledge base construction module uses natural language processing technology to perform word segmentation preprocessing on text data, and generates graph-structured data using the word segmentation results and domain rules; The prompt construction module formats the instructions according to predefined templates and rules after extracting key requirements and target intentions through semantic analysis of the design task description; the MBSE model generation module receives the standardized instructions generated by the prompt construction module and the supplementary information provided by the knowledge base construction module, and constructs an end-to-end inference process framework; the MBSE model evaluation module clarifies the corresponding simulation objectives and constraint conditions according to the input design requirements and performance goals, calls the simulation solver to complete the calculation, and comprehensively evaluates the simulation results based on the large language model.
9. The industrial MBSE model generation and evaluation system based on a large language model according to claim 8, characterized in that, The knowledge base construction module performs unit standardization and data normalization processing on process parameters and equipment operation data based on pandas to ensure that data from different sources and dimensions can be uniformly measured. For text data, through natural language processing technology based on BERT, the text data is converted into vector form. When processing the original text, the SpaCy tokenizer is used to preprocess the text, splitting the text into words or subwords for further analysis. The inverted index annotation technology of the ElasticSearch engine is utilized to establish the index relationship between keywords and documents, facilitating efficient retrieval and information extraction. For relatively complex text data, important entities in the text are identified through a natural language processing model, and the relationships between entities are recognized through semantic analysis to form a knowledge graph; The described prompt construction module extracts key information from the user's input requirements and instructions and converts it into standardized task instructions, providing accurate input for subsequent large model reasoning. The system analyzes the user's industrial field background, model format requirements, and personal interest preferences based on the LangChain framework, determines the context and requirements of the task, ensuring that the prompts can be customized according to different fields and user needs. Through ElasticSearch retrieval, vector knowledge retrieval, and knowledge graph retrieval, domain background knowledge related to the task is extracted from the knowledge base. After obtaining the background knowledge and the original requirements, the generated preliminary instructions are fine-tuned, including adjusting the language style of the instructions according to the requirements of specific roles, injecting background information, and detailed description of the task instructions; The described MBSE model generation module converts the processed task instructions into an actual reasoning and modeling process. For relatively complex or professional requirements, the system decomposes the requirements into actionable tasks, ensuring that each task has clear goals and constraints. For core sensitive design requirements, they are analyzed through a local reasoning engine. After the reasoning results are generated, the post-processing unit merges them into the final design solution. Throughout the reasoning process, the system continuously checks the consistency of the reasoning state to ensure the traceability of the process and the credibility of the generated results; The described result simulation evaluation module converts the reasoning results into physical scenario simulations and evaluates the feasibility of the design based on the simulation results. According to the design solution obtained from the reasoning, the API simulation interfaces provided by the SimuLink toolchain are called to perform simulations of various physical scenarios such as fluid simulation, thermodynamics simulation, and rigid body simulation. After obtaining the simulation results by calling the external toolchain, qualified determination is made through parameter analysis based on the large model to judge whether the design meets the performance requirements. When the design solution meets the requirements, the final design model, design description, and simulation report are generated; when the simulation results fail to pass the evaluation, modification suggestions will be generated according to the simulation analysis results and enter the next iteration to gradually optimize the design until the performance requirements are met.
Citation Information
Cited By
System and method for automatically generating SysML model based on mixed AI and domain knowledge
CN120911452A
A SysML Model Automatic Generation System and Method Based on Hybrid AI and Domain Knowledge
CN120911452B
Method and system for automatically adjusting parameters of power distribution network data quality improvement algorithm
CN121031999A
User behavior data enhancement method of large language model based on RFLP driving
CN121350606A
Test evaluation method and system based on large language model
CN121414079A