Spacecraft overall design large model training method based on knowledge graph
By adopting the knowledge graph-based overall design big model training method of spacecraft in the aerospace field, the problems of insufficient data and heterogeneity are solved, and efficient design decision support and computational cost reduction are achieved.
Patent Information
- Application Number
- CN202510303905.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-05-02
AI Technical Summary
The application of large-scale models in the aerospace field faces the problems of insufficient effective data, multidisciplinary collaborative design data heterogeneity, and lack of interpretability in traditional expert system reasoning.
The spacecraft overall design big model training method based on knowledge graph is adopted. By obtaining the spacecraft overall design knowledge graph, a deep entity disambiguation model based on BERT is constructed and trained through the SFT algorithm.
Effectively train the model in the case of scarcity of data, improve inference ability and design decision support, integrate heterogeneous data, generate highly interpretable design decisions, and reduce computing costs.
Smart Images

Figure CN119918187A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aerospace, and in particular to a method for training a large model of spacecraft overall design based on a knowledge graph. Background Art
[0002] In recent years, big model technology has become an important paradigm to promote the development of general artificial intelligence by combining massive data pre-training with downstream task fine-tuning. Open source and field-specific big models represented by LLaMA and Wenxin Yiyan have achieved remarkable results in finance, medical care, education and other fields with their hundreds of billions of parameters and multimodal data processing capabilities. For example, in the financial field, complex text analysis and intelligent decision-making have been achieved through models such as "Xuanyuan" and "Tianjing"; in the medical field, an intelligent system covering the entire diagnosis and treatment process has been built with the help of models such as MedGPT and "Jingyi Qianxun"; in the education field, personalized knowledge service capabilities have been strengthened by relying on the "Zi Yue" big model. These achievements highlight the generalization potential of big models in vertical fields, and their core lies in the feature extraction and logical reasoning capabilities driven by high-quality data.
[0003] However, the application of large models in the aerospace field faces unique challenges. First, spacecraft design involves complex scenarios such as multi-physical field coupling and extreme environment simulation. The relevant data is extremely scarce due to high experimental costs and strong confidentiality, resulting in insufficient effective data in model training, which restricts feature learning and generalization performance. Secondly, the data generated by multidisciplinary collaborative design has heterogeneity (such as misalignment between dynamics, thermodynamics, and materials parameters), noise, and default value problems, which can easily cause the model to learn incorrect associations and reduce robustness. More importantly, there is a fundamental contradiction between the "black box" characteristics of large models and the high safety requirements of aerospace missions-traditional expert systems rely on rule bases and preset interpreters, and it is difficult to deal with cross-system coupling problems; and although large models can break through rule restrictions, their reasoning process lacks explainability and cannot meet the strict requirements of aerospace design for decision transparency. These technical bottlenecks urgently need to be broken through through interdisciplinary methods to achieve the reliable implementation of intelligent technologies in the aerospace field. Summary of the invention
[0004] Aiming at the aerospace field, the present invention proposes a large-scale model training method for overall spacecraft design based on knowledge graph, in which there is insufficient effective data in model training, the data generated by multidisciplinary collaborative design is heterogeneous, and the traditional expert system is limited by rule-based reasoning and preset interpreters, which makes it difficult to answer design problems that go beyond preset rules and cross subsystems.
[0005] The method comprises:
[0006] Obtain the knowledge graph of spacecraft overall design;
[0007] Build a deep entity disambiguation model based on BERT;
[0008] The BERT-based deep entity disambiguation model is trained through the SFT algorithm.
[0009] Furthermore, a preferred method is proposed, whereby the node labels of the spacecraft overall design knowledge graph include: attitude control module, actuator, fault, sensor, measurement and control station, spacecraft, and orbit.
[0010] Furthermore, a preferred method is proposed, wherein the BERT-based deep entity disambiguation model includes:
[0011]
[0012] Among them, CF is the dimensional feature of the candidate entity.
[0013] Furthermore, a preferred embodiment is proposed, wherein the dimension features of the candidate entity include:
[0014]
[0015] Among them, F m,ct,e is the entity disambiguation module, m is the entity referent, ct is the context of m, and e is the entity.
[0016] Furthermore, a preferred embodiment is proposed, wherein the entity disambiguation module comprises:
[0017]
[0018] Among them, BERT [CLS] is the BERT model parameter of the features of the entire input sequence, BERT dot BERT model parameter for the similarity between query and key, prior e is the prior probability of entity e, prior m,e is the prior probability of the entity referent m being linked to the entity e in the candidate list, max_prior m,e is the maximum prior probability value of the entity referent m linking to the entity e in the candidate list, is the candidate text of entity referent m, e_qual_m means entity e matches entity referent m, e_contain_m means entity e contains entity referent m, e_start_m is the starting position of entity e and its referent m, and e_end_m is the ending position of entity e and the corresponding entity referent m.
[0019] Furthermore, a preferred method is proposed, wherein the deep entity disambiguation model based on BERT is trained by the SFT algorithm, including:
[0020] Build a pre-trained model;
[0021] Check whether the characteristics of the dataset meet the requirements of the entity disambiguation task;
[0022] Fine-tune the pre-trained model using a dataset that meets the requirements of the entity disambiguation task and optimize it using cross entropy loss;
[0023] Fine-tune the ERNIE Speed model and adjust the hyperparameters until the loss function is stable.
[0024] Furthermore, a preferred method is proposed, wherein the pre-training model is:
[0025]
[0026] Among them, x i is the i-th word in the sentence, x <i are all the words before the i-th word, θ is the model parameter, and N is the total number of candidate entities.
[0027] Furthermore, a preferred method is proposed, wherein the optimization using cross entropy loss includes:
[0028]
[0029] Among them, y ij is the true label of sample i in category j, P(y i |x i ; θ) is the probability that the model predicts sample i as category j, and M is the total number of categories.
[0030] Based on the same inventive concept, the present invention also proposes a computer device, including a memory and a processor, wherein the memory stores a computer program. When the processor runs the computer program stored in the memory, the processor executes a large model training method for spacecraft overall design based on a knowledge graph as described in any one of the above items.
[0031] Based on the same inventive concept, the present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a method for training a large model of spacecraft overall design based on a knowledge graph as described above are executed.
[0032] The present invention is beneficial in that:
[0033] The present invention proposes a method for training a large model of spacecraft overall design based on knowledge graph, which uses small sample high-quality labeled data for training. This has significant advantages over the problem of large-scale high-quality data being difficult to obtain in the aerospace field. In this way, even when data is scarce, the model can still be effectively trained to improve reasoning ability and design decision support.
[0034] Spacecraft design involves multiple subsystems and disciplines, and data is usually heterogeneous. Through the construction of a knowledge graph, the method proposed in this invention can integrate and associate design data from different sources and in different formats, so that information between different disciplines can be seamlessly connected, effectively overcoming the challenges brought by heterogeneous data.
[0035] The present invention adopts a deep entity disambiguation model based on BERT and combines it with the SFT algorithm for training. It can generate accurate and explanatory design decisions through the reasoning ability of the model. When faced with complex cross-subsystem and multidisciplinary cross-coupling problems, it can provide reasonable design solutions under the overall optimal principle, avoiding the shortcomings of traditional expert systems that are limited to rule reasoning.
[0036] Compared with traditional large-scale training methods, the present invention can effectively perform training in a low-computing power environment by fine-tuning large domain models, so that it can still achieve efficient design decision support under limited resources and reduce computing costs.
[0037] In spacecraft design, the coupling and complexity between subsystems make cross-subsystem design problems extremely difficult. By training a large model and utilizing the reasoning ability of the large model, the present invention can break the dilemma of traditional methods in design limitations, avoid the design deviation caused by the optimization of a single subsystem, and achieve the optimal design decision across subsystems. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A flow chart of a method for training a large model of spacecraft overall design based on a knowledge graph as described in implementation mode 1;
[0039] Figure 2 This is a schematic diagram of the knowledge graph of the overall design of a spacecraft according to the eleventh embodiment;
[0040] Figure 3 This is a schematic diagram of model training according to the eleventh embodiment;
[0041] Figure 4 This is a schematic diagram of model evaluation according to the eleventh embodiment;
[0042] Figure 5 This is a flowchart of the algorithm of the BERT-based deep entity disambiguation model described in Implementation Example 11;
[0043] Figure 6 This is a flow chart of the entity disambiguation algorithm described in implementation mode eleven. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.
[0045] Implementation method 1, see Figure 1 The present embodiment describes a method for training a large model of spacecraft overall design based on a knowledge graph, the method comprising:
[0046] Obtain the knowledge graph of spacecraft overall design;
[0047] Build a deep entity disambiguation model based on BERT;
[0048] The BERT-based deep entity disambiguation model is trained through the SFT algorithm.
[0049] The method proposed in this implementation uses small sample high-quality labeled data for training, which has significant advantages in the aerospace field where large-scale high-quality data is difficult to obtain. In this way, even when data is scarce, the model can still be effectively trained to improve reasoning capabilities and design decision support.
[0050] Spacecraft design involves multiple subsystems and disciplines, and data is usually heterogeneous. Through the construction of a knowledge graph, the present invention can integrate and associate design data from different sources and in different formats, so that information between different disciplines can be seamlessly connected, effectively overcoming the challenges brought by heterogeneous data.
[0051] In this implementation, a deep entity disambiguation model based on BERT is used in combination with the SFT algorithm for training, which can generate accurate and explanatory design decisions through the model's reasoning ability. When faced with complex cross-subsystem and multidisciplinary cross-coupling problems, it can provide a reasonable design solution under the overall optimal principle, avoiding the shortcomings of traditional expert systems that are limited to rule reasoning.
[0052] Compared with traditional large-scale training methods, this implementation can effectively perform training in a low-computing power environment by fine-tuning the large domain model, so that it can still achieve efficient design decision support under limited resources and reduce computing costs.
[0053] In spacecraft design, the coupling and complexity between subsystems make cross-subsystem design problems extremely difficult. The method proposed in this implementation mode, by training a large model and utilizing the reasoning ability of the large model, can break the dilemma of traditional methods in design limitations, avoid the design deviation caused by the optimization of a single subsystem, and achieve the optimal design decision across subsystems.
[0054] Implementation method 2. This implementation method is a further limitation of the large model training method for spacecraft overall design based on knowledge graph described in implementation method 1. The node labels of the spacecraft overall design knowledge graph include: attitude control module, actuator, fault, sensor, measurement and control station, spacecraft, and orbit.
[0055] Implementation method 3: This implementation method further limits the spacecraft overall design large model training method based on knowledge graph described in implementation method 1. The BERT-based deep entity disambiguation model includes:
[0056]
[0057] Among them, CF is the dimensional feature of the candidate entity.
[0058] Implementation method 4: This implementation method is a further limitation of the method for training a large model of spacecraft overall design based on a knowledge graph described in implementation method 3. The dimensional features of the candidate entities include:
[0059]
[0060] Among them, F m,ct,e is the entity disambiguation module, m is the entity referent, ct is the context of m, and e is the entity.
[0061] By further limiting the dimensional features of candidate entities, this embodiment can more accurately identify and eliminate ambiguity. For example, in spacecraft design, a large number of terms and equipment names are involved, and the same term may have different meanings in different contexts. Through precise disambiguation processing, the large model can better understand the exact meaning of different entities in the design, thereby improving the accuracy of reasoning and decision-making. By incorporating dimensional features into the disambiguation process and considering more entity attributes (such as function, structure, performance, etc.), similar entities can be effectively distinguished, and incorrect disambiguation caused by insufficient features can be avoided, thereby improving the decision-making ability of the model. By combining contextual information, the model can identify the specific meaning of candidate entities in the current design task. For example, some design elements may have different functions and attributes in different design stages or subsystems. Through contextual information, the model can more intelligently infer the true reference of the entity and avoid making incorrect matches based only on surface information. By deeply mining contextual information, the correlation between entities can be strengthened. For example, in the design process of a spacecraft, the selection of a certain component may be directly related to the selection of another component. Through accurate understanding of the context, the coupling between designs can be better understood, helping designers make better decisions.
[0062] As new design requirements or data are added, the knowledge graph can be continuously expanded and updated. This implementation further limits the dimensional features of candidate entities so that the model can more flexibly adapt to the introduction of new entities or design problems. When faced with new types of design tasks or system components, the model can quickly adapt and reason based on known entity features.
[0063] Implementation mode 5: This implementation mode further limits the spacecraft overall design large model training method based on knowledge graph described in implementation mode 4, and the entity disambiguation module includes:
[0064]
[0065] Among them, BERT [CLS] is the BERT model parameter of the features of the entire input sequence, BERT dot BERT model parameter for the similarity between query and key, prior e is the prior probability of entity e, prior m,e is the prior probability of the entity referent m being linked to the entity e in the candidate list, max_prior m,e is the maximum prior probability value of the entity referent m linking to the entity e in the candidate list, is the candidate text of entity referent m, e_qual_m means entity e matches entity referent m, e_contain_m means entity e contains entity referent m, e_start_m is the starting position of entity e and its referent m, and e_end_m is the ending position of entity e and the corresponding entity referent m.
[0066] In this embodiment, by using BERT model parameters to process the characteristics of the input sequence and the similarity between the query and the key, the meaning of the words in different contexts can be better captured, thereby improving the system's understanding and matching capabilities for complex texts. Entities are disambiguated by using the prior probability of the entity, the maximum prior probability in the candidate list, and the candidate text. The relevance evaluation of these prior probabilities and candidate entities helps the system accurately identify the correct match between the referent and the entity, reduces ambiguity, and ensures that each entity in the knowledge graph can be clearly identified, thereby improving the accuracy of the model. By considering the starting and ending positions of the entity e, as well as the matching degree between the entity and its referent, the entity in the text can be accurately located. This approach can further enhance the model's understanding of entity relationships, especially in more complex design documents or texts, where entities are often cited multiple times, and accurately matching their positions is crucial for subsequent knowledge reasoning and task execution. By linking the entity referent m to the maximum prior probability value in the candidate entity list, the model can effectively select the most likely entity from multiple candidate entities. This method of maximizing matching can improve the matching accuracy of the entity referent m, reduce the risk of mismatching, and thus improve the quality and reliability of the knowledge graph.
[0067] By combining multiple matching indicators (such as entity prior probability, position matching, etc.), this method can build a richer entity relationship graph and provide stronger support for subsequent reasoning and decision-making. In the process of overall spacecraft design, there are many entities involved. Accurate entity disambiguation and relationship reasoning can help the model generate more forward-looking and accurate design solutions.
[0068] Implementation 6. This implementation is a further limitation of the method for training a large model of spacecraft overall design based on a knowledge graph described in Implementation 1. The method of training a deep entity disambiguation model based on BERT by using the SFT algorithm includes:
[0069] Build a pre-trained model;
[0070] Check whether the characteristics of the dataset meet the requirements of the entity disambiguation task;
[0071] Fine-tune the pre-trained model using a dataset that meets the requirements of the entity disambiguation task and optimize it using cross entropy loss;
[0072] Fine-tune the ERNIE Speed model and adjust the hyperparameters until the loss function is stable.
[0073] In this embodiment, it is necessary to first ensure that the characteristics of the data set meet the requirements of the entity disambiguation task. This data quality control can ensure that the input data in the training process is valid and meets the standards, avoiding the adverse effects of low-quality data on model training. Ensuring the quality of the data set is one of the key steps to optimize the effect of the deep entity disambiguation model. If it is found that the data set does not meet the requirements, further optimization can be performed. By screening and adjusting the data set, it is ensured that it contains sufficient diversity and high-quality information, and ultimately improves the generalization ability of the model. The ERNIE Speed model is used to further improve the model's ability to understand complex tasks by integrating external knowledge sources, and to make up for the limitations of the traditional BERT model by combining external knowledge such as knowledge graphs, so that the model performs better in specific tasks. By fine-tuning based on the ERNIE Speed model and adjusting the hyperparameters, the loss function can be stabilized during the training process. This refined adjustment can effectively improve the performance of the model, avoid overfitting or underfitting, so as to better adapt to actual tasks and ensure the stability and accuracy of the entity disambiguation effect.
[0074] Requirements for spacecraft overall design tasks: Spacecraft overall design involves multiple disciplines and complex engineering requirements. The model needs to be able to understand complex entity relationships, eliminate ambiguity, and provide accurate design support. The BERT-based deep entity disambiguation model, combined with the SFT algorithm and the fine-tuning of the ERNIE Speed model, can effectively deal with entity ambiguity problems that may arise in spacecraft design and improve the reliability and accuracy of the model during the design process.
[0075] Implementation method 7: This implementation method is a further limitation of the method for training a large model of spacecraft overall design based on a knowledge graph described in implementation method 6. The pre-trained model is:
[0076]
[0077] Among them, x i is the i-th word in the sentence, x <i are all the words before the i-th word, θ is the model parameter, and N is the total number of candidate entities.
[0078] By using the knowledge graph, the model is able to extract entities, attributes, and relationships related to spacecraft design from a structured knowledge base. This approach not only captures the vocabulary information in the text, but also matches it with the existing knowledge graph, thereby enhancing the model's semantic understanding ability and being able to handle more complex spacecraft design scenarios.
[0079] By setting the total number of candidate entities N, the model can select more possible entities to process during reasoning. This approach increases the diversity of entities during training, allowing the model to better understand and respond to different design situations.
[0080] Implementation 8. This implementation is a further limitation of the spacecraft overall design large model training method based on knowledge graph described in Implementation 6. The optimization using cross entropy loss includes:
[0081]
[0082] Among them, y ij is the true label of sample i in category j, P(y i |x i ; θ) is the probability that the model predicts sample i as category j, and M is the total number of categories.
[0083] Embodiment 9. A computer device described in this embodiment includes a memory and a processor, wherein the memory stores a computer program. When the processor runs the computer program stored in the memory, the processor executes a knowledge graph-based large model training method for spacecraft overall design described in any one of embodiments 1 to 7.
[0084] Embodiment 10. A computer-readable storage medium described in this embodiment stores a computer program, and when the computer program is executed by a processor, the steps of a method for training a large model of spacecraft overall design based on a knowledge graph are executed as described in any one of embodiments 1 to 7.
[0085] Implementation Method 11: See Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 6 This embodiment provides a specific example of a method for training a large model of spacecraft overall design based on a knowledge graph as described in the first embodiment, and is also used to explain the second to eighth embodiments, specifically:
[0086] (1) Dataset preparation: The data comes from the knowledge graph of spacecraft overall design and is structured data in JSON format. The data has the characteristics of high quality, high accuracy, strong reliability, multi-domain entity alignment and de-fragmentation. The knowledge distillation experiment has been completed to ensure that duplicate data is removed and missing values and outliers are handled.
[0087] (2) Figure 5 and Figure 6As shown in the algorithm, a deep entity disambiguation model based on BERT is constructed;
[0088] A multi-layer perceptron (MLP) will be used to solve the entity disambiguation problem. For each candidate entity in the candidate list, the following features are obtained from the fine-tuned BERT module as the input of the MLP, which is called BASE features in this model, including: the vector of the candidate entity; the vector of the entity referent and its context.
[0089] In the following experiments, we added prior features (denoted as prior) and string similarity features (denoted as stringsimi) to compare with the basic feature BASE. Specifically, the prior features include: the prior probability of entity e o e It is expressed as the number of occurrences of e, o all is the total number of occurrences of all entities; the prior probability That is, the prior probability of entity referent m linking to entity e in the candidate list; the total number of candidate entities of m in the corpus with the maximum prior probability that each entity referent in the context points to entity e.
[0090] The string similarity features include whether the title of e is exactly equal to or contains the string of m, and whether the title of e starts or ends with the string of m.
[0091] The characteristics of the entity disambiguation module are as follows:
[0092]
[0093] All of these constitute a 1546-dimensional feature. m,ct,e In , the dimension of the first two features is 768, and the dimension of the other features is 1. Then the 1546-dimensional features of each entity are embedded into a fixed dimension dim, and then the features of each candidate entity are combined to obtain dim*30-dimensional features:
[0094]
[0095] On top of the features, a hidden layer using the dropout strategy is superimposed, and this hidden layer is activated by the Relu function. In this implementation, a hidden layer and an output layer are further added to use softmax to predict the entity most relevant to the entity referent among the candidate entities. This process can be formalized as:
[0096]
[0097] In addition, this module uses the classification cross entropy cost function as the loss function:
[0098]
[0099] in, represents the predicted probability, y is a Boolean vector indicating whether the candidate entity is mentioned, and N is the total number of candidate entities.
[0100] (3) Select the pre-trained model and model fine-tuning algorithm:
[0101] In view of the data set size and platform computing power limitations, the SFT (Supervised Fine-tuning) algorithm was selected. The core idea of the SFT algorithm is to fine-tune the model in a supervised manner based on the pre-trained model. Specifically, it uses labeled data to train the model to minimize the loss function on the labeled data. During the training process, the model parameters are updated according to the gradient of the loss function to gradually adapt to the data distribution of the new task.
[0102] The SFT process is divided into the following steps:
[0103] Pre-trained models:
[0104]
[0105] Among them, x i is the i-th word in the sentence, x <i are all the words before it, and θ is the model parameter.
[0106] Check the feature values of the dataset;
[0107] Supervised Fine-tuning:
[0108] Fine-tune the pre-trained model using a task-specific dataset;
[0109] During fine-tuning, the model’s parameters are updated based on task-specific data to optimize the model’s performance on that task;
[0110] The objective function used in the fine-tuning phase is the loss function of the supervised learning task, and the cross entropy loss is selected:
[0111]
[0112] Among them, y ij is the true label of sample i in category j, P(y i |x i ; θ) is the probability that the model predicts sample i as category j.
[0113] (4) Model training: Based on the knowledge graph, ERNIE Speed is selected as the base model for fine-tuning, and model parameters are adjusted, including hyperparameters such as learning rate, number of training rounds (Epochs) and batch size (Batch Sizes). When the loss function is stable, the model training is completed.
[0114] (5) Model evaluation: Select the model evaluation method, and select the GPT-J-6B and Llama-2-7B models to evaluate the language understanding and reasoning ability of the generated model. The evaluation indicators include accuracy, recall rate, F1 value, etc. When the F1 value is greater than 75%, it indicates that the model has semantic understanding ability.
[0115] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0116] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure rather than to limit its protection scope. Although the present disclosure has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that after reading the present disclosure, those skilled in the art can still make various changes, modifications or equivalent substitutions to the specific implementation methods of the invention, but these changes, modifications or equivalent substitutions are all within the protection scope of the disclosed claims to be approved.
Claims
1. A large model training method for spacecraft overall design based on knowledge graph, characterized in that: The method comprises: Obtain the knowledge graph of spacecraft overall design; Build a deep entity disambiguation model based on BERT; The BERT-based deep entity disambiguation model is trained through the SFT algorithm.
2. According to claim 1, a method for training a large model of spacecraft overall design based on knowledge graph is characterized in that: The node labels of the spacecraft overall design knowledge graph include: attitude control module, actuator, fault, sensor, measurement and control station, spacecraft, and orbit.
3. According to claim 1, a method for training a large model of spacecraft overall design based on knowledge graph is characterized in that: The BERT-based deep entity disambiguation model includes: Among them, CF is the dimensional feature of the candidate entity.
4. The method for training a large model of spacecraft overall design based on knowledge graph according to claim 3 is characterized in that: The dimension features of the candidate entity include: Among them, F m,ct,e is the entity disambiguation module, m is the entity referent, ct is the context of m, and e is the entity.
5. The method for training a large model of spacecraft overall design based on knowledge graph according to claim 4 is characterized in that: The entity disambiguation module includes: Among them, BERT [CLS] is the BERT model parameter of the features of the entire input sequence, BERT dot BERT model parameter for the similarity between query and key, prior e is the prior probability of entity e, prior m,e is the prior probability of the entity referent m being linked to the entity e in the candidate list, max_prior m,e is the maximum prior probability value of the entity referent m linking to the entity e in the candidate list, is the candidate text of entity referent m, e_qual_m means entity e matches entity referent m, e_contain_m means entity e contains entity referent m, e_start_m is the starting position of entity e and its referent m, and e_end_m is the ending position of entity e and the corresponding entity referent m.
6. The method for training a large model of spacecraft overall design based on knowledge graph according to claim 1 is characterized in that: The deep entity disambiguation model based on BERT is trained by the SFT algorithm, including: Build a pre-trained model; Check whether the characteristics of the dataset meet the requirements of the entity disambiguation task; Fine-tune the pre-trained model using a dataset that meets the requirements of the entity disambiguation task and optimize it using cross entropy loss; Fine-tune the ERNIE Speed model and adjust the hyperparameters until the loss function is stable.
7. The method for training a large model of spacecraft overall design based on knowledge graph according to claim 6 is characterized in that: The pre-trained model is: Among them, x i is the i-th word in the sentence, x <i are all the words before the i-th word, θ is the model parameter, and N is the total number of candidate entities.
8. The method for training a large model of spacecraft overall design based on knowledge graph according to claim 6 is characterized in that: The optimization using cross entropy loss includes: Among them, y ij is the true label of sample i in category j, P(y i |x i ; θ) is the probability that the model predicts sample i as category j, and M is the total number of categories.
9. A computer device, characterized in that: It includes a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes a large model training method for overall design of a spacecraft based on a knowledge graph according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of a method for training a large model of spacecraft overall design based on a knowledge graph as described in any one of claims 1-7.