Method and device for improving robustness of satellite-borne large model and storage medium

By optimizing the structure and training process of the satellite-borne big model, using task and equipment attribute information, the robustness of the model is improved, and the problem of low efficiency in the existing technology is solved, thereby achieving higher prediction accuracy and training efficiency.

CN120067673APending Publication Date: 2025-05-30XINGHAN SPACE TIME (SHENZHEN) AEROSPACE INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411984424.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Star-borne large models face the problem of low efficiency in complex spatial environments and data changes, and are prone to underfitting or overfitting, resulting in low prediction accuracy.

Method used

By responding to the robustness improvement signal of the initial satellite-borne big model, the sample data set, task attribute information, device attribute information and expected robustness performance indicators are obtained, the adapted expected model structure and external network to be integrated, the structure optimization and adjustment are performed, and the sample data set is used for training to improve the robustness of the model.

Benefits of technology

It improves the robustness and efficiency of the large-scale satellite models, avoids overfitting and underfitting, and enhances the model's adaptability to environmental changes and prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067673A_ABST
    Figure CN120067673A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for improving robustness of a satellite-borne large model, and a storage medium, and the method comprises the steps: responding to a robustness improvement signal of an initial satellite-borne large model; obtaining an initial sample data set, task attribute information of a to-be-predicted task corresponding to the initial satellite-borne large model, equipment attribute information of satellite-borne computing equipment deployed by the initial satellite-borne large model, and an expected robustness performance index of the initial satellite-borne large model; determining an expected model structure matched with the to-be-predicted task based on the task attribute information and the equipment attribute information; based on the task attribute information and the expected robustness performance index, determining a to-be-integrated external network matched with the initial satellite-borne large model; based on the expected model structure and a to-be-integrated external network, carrying out structure optimization adjustment on the initial satellite-borne large model to obtain a to-be-trained satellite-borne large model; and training a to-be-trained satellite-borne large model by using the sample data set to obtain a satellite-borne large model meeting a robustness improvement condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model optimization, and in particular, to a method, device, and storage medium for improving the robustness of on-board large models. Background Art

[0002] With the continuous development of space technology, on-board large models are undertaking increasingly important tasks, such as spacecraft navigation and control, satellite applications, and space exploration and data analysis. However, on-board large models face complex space environments and data changes, which pose higher requirements for their robustness.

[0003] Currently, during the process of improving model robustness, the original model is usually directly trained using the original data obtained from the network. However, this direct training method for the original model may, due to the complexity of the model, lead to the risk of underfitting or overfitting during the training process, resulting in poor efficiency in improving the robustness of the large model, that is, low accuracy in improving robustness. At the same time, during the overfitting process of the model, there will also be time-consuming and laborious training, that is, low efficiency in improving the robustness of the model, and further resulting in the inability of the trained model to make accurate predictions. Summary of the Invention

[0004] The present invention provides a method, device, and storage medium for improving the robustness of on-board large models, mainly capable of improving the improvement effect and efficiency of the robustness of on-board large models.

[0005] According to the first aspect of the present invention, there is provided a method for improving the robustness of an on-board large model, including:

[0006] In response to a signal for improving the robustness of the initial on-board large model, obtaining a sample data set, task attribute information of the prediction task corresponding to the initial on-board large model, device attribute information of the on-board computing device on which the initial on-board large model is deployed, and the expected robustness performance index of the initial on-board large model, where the sample data set includes sample data matching the prediction task and its corresponding annotation information;

[0007] Based on the task attribute information and the device attribute information, determining an expected model structure adapted to the prediction task;

[0008] Based on the task attribute information and the expected robustness performance index, determining an external network to be integrated adapted to the initial on-board large model;

[0009] Based on the expected model structure and the external network to be integrated, performing structural optimization and adjustment on the initial on-board large model to obtain a to-be-trained on-board large model;

[0010] The sample data set is used to train the onboard large model to be trained, so as to obtain the onboard large model that meets the robustness improvement condition.

[0011] Optionally, before using the sample data set to train the onboard large model to be trained to obtain the onboard large model that meets the robustness improvement condition, the method further includes:

[0012] Determine abnormal data in the sample data, and remove the abnormal data in the sample data to obtain abnormal sample data;

[0013] Smoothing the removed-odds sample data to obtain smoothed sample data;

[0014] Determine whether there are missing values ​​in the smoothed sample data, and if so, supplement the missing values ​​in the smoothed sample data to obtain supplemented sample data;

[0015] Data enhancement is performed on the supplementary sample data to obtain the sample data set after adaptive adjustment for improving robustness.

[0016] Optionally, the determining abnormal data in the sample data includes:

[0017] Determine the data density in a preset neighborhood corresponding to each of the sample data, and determine core data points and non-core data points in each of the sample data based on the data density;

[0018] Taking any core data point among the core data points as a target core data point, taking the target core data point as the cluster center, clustering the remaining core data points to obtain core data points under different clustering categories;

[0019] Taking any non-core data point among the non-core data points as a target non-core data point, determining a reference core data point that is closest to the non-core data point among the core data points under each clustering category, and judging whether the distance between the non-core data point and the reference core data point is less than a preset distance threshold;

[0020] If the distance between the non-core data point and the reference core data point is less than the preset distance threshold, the non-core data point is classified into the cluster category to which the reference core data point belongs; otherwise, the non-core data point is determined as abnormal data.

[0021] Optionally, the supplementary sample data includes at least one of image data and text data;

[0022] The step of performing data enhancement on the supplementary sample data to obtain the sample data set after adaptive adjustment for improving robustness includes:

[0023] When the supplementary sample data is image data, perform at least one of rotation, translation, scaling, cropping, and color transformation on the image data to generate enhanced image data, and the sample data set after adaptation adjustment is composed of the enhanced image data and the image data;

[0024] When the supplementary sample data is text data, perform at least one of synonym replacement, sentence restructuring, and text back-translation on the text data to generate enhanced text data, and the sample data set after adaptation adjustment is composed of the enhanced text data and the text data.

[0025] Optionally, the to-be-trained spaceborne large model is an information prediction model for predicting the carbon dioxide column concentration in the target area. The sample data set includes sample data and its corresponding labeled carbon dioxide column concentration. The sample data includes: meteorological data, environmental data, terrain data, historical carbon dioxide column concentration data, production activity data, ecosystem data, and prediction time of multiple regions;

[0026] Training the to-be-trained spaceborne large model using the sample data set to obtain a spaceborne large model that meets the robustness improvement condition includes:

[0027] Determine the current model parameters of the to-be-trained spaceborne large model;

[0028] Based on the current model parameters, determine the overfitting constraint term of the to-be-trained spaceborne large model;

[0029] According to the cross-entropy between the labeled carbon dioxide column concentration corresponding to the sample data in the sample data set and the predicted carbon dioxide column concentration corresponding to the sample data predicted by the to-be-trained spaceborne large model, determine the initial loss function of the to-be-trained spaceborne large model;

[0030] Add the overfitting constraint term to the initial loss function to obtain a target loss function, and use the target loss function to train the to-be-trained spaceborne large model to obtain a spaceborne large model that meets the robustness improvement condition.

[0031] Optionally, training the to-be-trained spaceborne large model using the sample data set to obtain a spaceborne large model that meets the robustness improvement condition includes:

[0032] Determine the number of categories of the labeled information corresponding to the sample data in the sample data set, and determine the information vector corresponding to the labeled information;

[0033] Based on the number of categories and the information vector, determine the fuzzy labeled information corresponding to the sample data;

[0034] Determine the loss function of the to-be-trained on-board large model based on the cross-entropy between the fuzzy annotation information corresponding to the sample data and the predicted annotation information corresponding to the sample data predicted by the to-be-trained on-board large model;

[0035] Use the loss function to train the to-be-trained on-board large model to obtain an on-board large model that meets the conditions for robustness improvement.

[0036] Optionally, the expected model structure includes the expected number of network layers and the expected activation function;

[0037] The structural optimization and adjustment of the initial on-board large model based on the expected model structure and the to-be-integrated external network to obtain the to-be-trained on-board large model includes:

[0038] Gradually adjust the number of network layers in the initial on-board large model according to the expected number of network layers to obtain the adjusted initial on-board large model;

[0039] Use the expected activation function to replace the original activation function in the adjusted initial on-board large model to obtain the replaced initial on-board large model;

[0040] Integrate the to-be-integrated external network into the replaced initial on-board large model to obtain the to-be-trained on-board large model, where the method of integrating the to-be-integrated external network into the replaced initial on-board large model includes:

[0041] Package the to-be-integrated external network as a to-be-called function;

[0042] Add a function reference statement to the code of the replaced initial on-board large model and determine the integration position of the to-be-integrated external network in the to-be-trained initial on-board large model;

[0043] Based on the function reference statement, reference the to-be-called function to the integration position.

[0044] According to the second aspect of the present invention, there is provided a device for improving the robustness of an on-board large model, including:

[0045] An acquisition unit, configured to acquire a sample data set, task attribute information of a to-be-predicted task corresponding to the initial on-board large model, device attribute information of an on-board computing device where the initial on-board large model is deployed, and an expected robustness performance index of the initial on-board large model in response to a robustness improvement signal of the initial on-board large model, where the sample data set includes sample data matching the to-be-predicted task and its corresponding annotation information;

[0046] A structure determination unit, configured to determine an expected model structure adapted to the to-be-predicted task based on the task attribute information and the device attribute information;

[0047] A network determination unit, configured to determine an external network to be integrated and adapted to the initial on-board large model based on the task attribute information and the expected robustness performance index;

[0048] A structure adjustment unit, configured to perform structure optimization and adjustment on the initial on-board large model based on the expected model structure and the external network to be integrated, so as to obtain a to-be-trained on-board large model;

[0049] A training unit, configured to train the to-be-trained on-board large model by using the sample data set to obtain an on-board large model that meets the conditions for robustness improvement.

[0050] According to a third aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above method for improving the robustness of the on-board large model is implemented.

[0051] According to a fourth aspect of the present invention, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the above method for improving the robustness of the on-board large model is implemented.

[0052] A method, device, and storage medium for improving the robustness of an on-board large model provided by the present invention, compared with the current method of directly training an original model using the original data obtained from a network, the present invention obtains a sample data set, task attribute information of a to-be-predicted task corresponding to the original on-board large model, device attribute information of an on-board computing device on which the original on-board large model is deployed, and an expected robustness performance index in response to a robustness improvement signal of the original on-board large model. Among them, the sample data set includes sample data matching the to-be-predicted task and its corresponding annotation information; and based on the task attribute information and the device attribute information, an expected model structure adapted to the to-be-predicted task is determined; at the same time, based on the task attribute information and the expected robustness performance index, an external network to be integrated with the original on-board large model is determined; then, based on the expected model structure and the external network to be integrated, the structure of the original on-board large model is optimized and adjusted to obtain a to-be-trained on-board large model; finally, the to-be-trained on-board large model is trained using the sample data set to obtain an on-board large model that meets the robustness improvement conditions. Thus, before training the original on-board large model using the sample data, first, according to the task attribute information and the device attribute information, an expected model structure capable of improving robustness is determined, and according to the task attribute information and the expected robustness performance index, an external network to be integrated capable of improving robustness is determined. Finally, the expected model structure and the external network to be integrated are used to optimize and adjust the structure of the original on-board large model, and the optimized model is trained using the sample data. The appropriate on-board large model structure can fully capture the complex features in the data and will not exhibit underfitting or overfitting during the training process, thereby improving the training effect of the on-board large model, that is, improving the robustness improvement effect of the on-board large model. At the same time, training the appropriate on-board large model structure will avoid the time and resources consumed by training a model with an overly complex structure. Therefore, the present invention can improve the efficiency of improving the robustness of the on-board large model and save the resources for improving the robustness, and further improve the prediction accuracy of the trained on-board large model. Description of the Drawings

[0053] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0054] Figure 1 Shows a flowchart of a method for improving the robustness of an on-board large model provided by an embodiment of the present invention;

[0055] Figure 2 Shows a flowchart of another method for improving the robustness of an on-board large model provided by an embodiment of the present invention;

[0056] Figure 3 shows a schematic structural diagram of a device for improving the robustness of an on-board large model provided by an embodiment of the present invention;

[0057] Figure 4 shows a schematic structural diagram of another device for improving the robustness of an on-board large model provided by an embodiment of the present invention;

[0058] Figure 5 shows a schematic physical structure diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0059] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0060] Currently, the method of directly training an original model using the original data obtained from the network to improve the model robustness will result in overfitting and underfitting situations, thereby leading to a poor improvement effect on the model robustness and reducing the improvement efficiency of the model robustness.

[0061] To solve the above problems, an embodiment of the present invention provides a method for improving the robustness of an on-board large model, as Figure 1 shown, the method includes:

[0062] 101. In response to a signal for improving the robustness of the initial on-board large model, obtain a sample data set, task attribute information of the prediction task corresponding to the initial on-board large model, device attribute information of the on-board computing device on which the initial on-board large model is deployed, and an expected robustness performance index of the initial on-board large model, where the sample data set includes sample data matching the prediction task and its corresponding annotation information.

[0063] Among them, the initial on-board large model is the model mainly for improving robustness; the prediction task may be: CO 2 column concentration prediction task, ship recognition task in the sea area and other target recognition tasks, environmental change monitoring task, predicting future environmental conditions based on historical data, etc.; if the prediction task is CO 2 column concentration prediction task. Then the sample data set is meteorological data (wind speed, wind direction, etc.), environmental data (temperature, humidity, etc.), terrain data (altitude, terrain undulation, landform features, surface type, vegetation coverage, soil type, etc.) of multiple regions with CO 2 column concentration label information, historical O 2Column concentration, human activities (such as industrial emissions, traffic emissions, agricultural activities, etc.); if the task to be predicted is a target recognition task such as vehicles on the road, the corresponding sample data set is road image data with vehicle annotation frames; the task attribute information includes: task type, task data volume and diversity, task complexity, etc.; the device attribute information includes the computing power, storage capacity, architecture, processing power, etc. of the spaceborne computing device; the expected robustness performance indicators can be: stability, reliability, accuracy, response time, resource occupancy, etc.

[0064] For the embodiments of the present invention, in order to improve the robustness of the initial spaceborne large model, it is first necessary to determine the task to be predicted corresponding to the initial spaceborne large model, then obtain a sample data set matching the task to be predicted, and determine the device attribute information of the spaceborne computing device to which the initial spaceborne large model belongs, the task attribute information of the task to be predicted, and the expected robustness performance indicators of the initial spaceborne large model. Then, according to the above information, determine the expected model structure and external network for the initial spaceborne large model to achieve the expected improvement in robustness. Finally, based on the expected model structure and external network, optimize and adjust the structure of the initial spaceborne large model, and use the training set to train the initial spaceborne large model after the structure optimization and adjustment, so as to train a spaceborne large model with satisfactory robustness. In the process of improving the robustness of the spaceborne large model in the embodiments of the present invention, by optimizing and adjusting the structure of the spaceborne large model, it is possible to avoid the situation of overfitting caused by an overly complex structure of the spaceborne large model, and also avoid the risk of underfitting due to an overly simple structure of the spaceborne large model. Therefore, the embodiments of the present invention can improve the effect of improving the robustness of the spaceborne large model.

[0065] 102. Determine the expected model structure adapted to the task to be predicted based on the task attribute information and the device attribute information.

[0066] Among them, the expected model structure includes the number of network layers of the large model, activation functions, etc. For the embodiments of the present invention, first, analyze the initial network structure of the initial spaceborne large model, understand its basic architecture and performance characteristics, evaluate whether the initial network structure is suitable for the task to be predicted, and whether adjustment is needed, considering the task type and data scale of the task to be predicted: According to the type of the task to be predicted (such as tasks like image recognition, speech recognition, natural language processing, etc.), determine the required network depth and complexity, and consider the size of the data scale. Larger data sets usually require deeper networks to capture complex features in the data. Evaluate the processing capacity of the spaceborne computing device: Understand the hardware configuration (such as CPU, GPU, etc.) and processing capacity of the deployed spaceborne computing device, and determine the upper limit of the number of network layers according to the processing capacity of the device to avoid exceeding the computing capacity of the device. Comprehensive trade-off and selection: On the basis of considering the initial network structure, task type, data scale, and the processing capacity of the spaceborne computing device, comprehensively trade off and select the most suitable number of network layers. It can be adjusted continuously through experiments and verification until the optimal configuration is found. Further, the activation function is an important non-linear element in the neural network, which can increase the expressive ability of the network. The activation function determines the output characteristics of the neuron. Different activation functions have different properties and will also affect the performance and robustness of the model. For example, the ReLU activation function is widely used in deep learning and has advantages such as simple calculation and fast convergence speed. However, the ReLU activation function will have the problem of gradient disappearance when the input is negative, which affects the training effect of the model. Specifically, the activation function can be selected according to the task type: For example, for the image recognition task, the ReLU activation function is usually a good choice because it can accelerate training and reduce the problem of gradient disappearance. For the binary classification task, the Sigmoid activation function may be more suitable because it can limit the output between 0 and 1. For other types of tasks, the appropriate activation function can be selected according to specific needs. When selecting the activation function, its computational efficiency and stability also need to be considered. Some activation functions may cause the problem of neuron death during the training process, so careful selection is required. Different activation functions can also be tried and their performance can be evaluated through experiments and verification. In this way, the appropriate number of network layers and activation function can be determined. From the appropriate number of network layers, activation function, and the initial model structure of the spaceborne large model, the expected model structure of the spaceborne large model with enhanced robustness can be obtained. In the process of enhancing the robustness of the large model in the embodiments of the present invention, by optimizing the number of network layers and activation function of the large model, the appropriate number of network layers can improve the learning ability of the spaceborne large model, enabling it to capture deep features in the data, thereby improving the performance of the spaceborne large model. At the same time, the appropriate activation function can enable the spaceborne large model to learn the non-linear relationship in the data, thereby enhancing the expressive ability of the spaceborne large model.

[0067] In another embodiment of the present invention, the expected model structure can be determined in the following manner: determining the initial network structure of the initial spaceborne large model; respectively determining the task feature vector corresponding to the task attribute information, the device feature vector corresponding to the device attribute information, and the structure feature vector corresponding to the initial network structure; performing cross-processing on the task feature vector, the device feature vector, and the structure feature vector to obtain a cross feature vector; inputting the cross feature vector into a preset network structure prediction model for structure prediction to obtain the expected number of network layers and the expected activation function. Among them, the preset network structure prediction model is trained based on a sample data set with label information of the expected number of network layers and the expected activation function. The cross-processing method includes: performing feature-level cross-processing on the task feature vector, the device feature vector, and the structure feature vector to obtain a feature cross vector; performing element-level cross-processing on the task feature vector, the device feature vector, and the structure feature vector to obtain an element cross vector; performing low-order cross-processing on the task feature vector, the device feature vector, and the structure feature vector to obtain a low-order cross vector; using a preset transformation function to perform transformation processing on the feature cross vector, the element cross vector, and the low-order cross vector to obtain a cross feature vector. Among them, the preset transformation function is set according to actual requirements. Through the above cross-processing of the features in the embodiment of the present invention, more implicit information can be extracted, making the prediction of the expected structure of the model more accurate.

[0068] 103. Based on the task attribute information and the expected robustness performance index, determine the external network to be integrated that is adapted to the initial spaceborne large model.

[0069] For the embodiments of the present invention, task type: determine whether the task to be predicted is a classification task, a regression task, an association analysis, a clustering analysis, or other types. Different types of tasks have different requirements for the external network. For example, a classification task may require a more refined feature extraction network, while a regression task may pay more attention to the smoothness and generalization ability of the network. Data scale: evaluate the size of the data volume of the task to be predicted, including the scales of training data, validation data, and test data. The size of the data scale will directly affect the selection of the external network. For example, large-scale data may require a more complex network structure to process. Data characteristics: analyze the characteristics of the data, such as distribution, features, noise level, etc. Select a suitable external network structure according to the data characteristics. For example, for high-noise data, it may be necessary to introduce a noise suppression mechanism or a more robust network structure. Determine the expected robustness performance indicators, stability: evaluate the ability of the model to maintain stable output when facing changes in input data. Stability is one of the important indicators of robustness, especially important for spaceborne large models because the spaceborne environment may have various uncertain factors. Reliability: measure the ability of the spaceborne large model to complete the specified functions within the specified time and under the specified conditions. Reliability is a key indicator for evaluating the performance of spaceborne large models, especially important for spaceborne large models that need to run stably for a long time. Fault tolerance: evaluate the ability of the spaceborne large model to still maintain certain performance and functions when faults or errors occur. In summary, according to the expected task attribute information and expected robustness performance indicators, select and determine the external network to be integrated that is adapted to the initial spaceborne large model from multiple external networks, and at the same time, it is also possible to determine the integration position of the network to be integrated in the initial spaceborne large model. For example, for a classification task that requires high stability, structures such as deep convolutional neural networks or deep residual networks can be selected. For tasks that require high reliability, redundant design or backup mechanisms can be considered to improve the fault tolerance of the network. The embodiments of the present invention improve the robustness of the spaceborne large model by introducing an external network. Since the external network can introduce additional robustness mechanisms, such as noise suppression, anomaly detection, etc., the model can maintain a stable output when facing changes in input data, thereby effectively improving the robustness of the spaceborne large model.

[0070] 104. Based on the expected model structure and the external network to be integrated, perform structural optimization and adjustment on the initial spaceborne large model to obtain the spaceborne large model to be trained.

[0071] For the embodiments of the present invention, first, determine the initial network structure of the initial spaceborne large model. For example, the initial network structure includes an input layer, a hidden layer, and an output layer, and there is an activation function 1 before the output layer. The expected model structure includes adding a new hidden layer after the existing hidden layer and adjusting the activation function before the output layer to activation function 2. Then, according to the expected network structure, add a new hidden layer after the hidden layer of the initial spaceborne large model and replace activation function 1 with activation function 2. Further, according to the external network to be integrated and the integration position, integrate the external network to be integrated at the integration position of the initial spaceborne large model. In another embodiment of the present invention, the integration method further includes: determining the function, input and output formats of the network to be integrated and its compatibility with other networks in the initial spaceborne large model; selecting a suitable integration method according to the characteristics of the network to be integrated and the requirements of the initial spaceborne large model; making corresponding adjustments to the network structure of the initial spaceborne large model according to the selected integration method so as to connect or fuse with the network to be integrated; and integrating the network to be integrated into the adjusted initial spaceborne large model by using the integration method to obtain the spaceborne large model to be trained. Among them, the integration methods include: direct connection, that is, directly connecting the output of the network to be integrated to the input or a certain intermediate layer of the initial spaceborne large model; feature fusion, that is, fusing the features extracted by the network to be integrated with the features of the initial spaceborne large model to improve the performance of the initial spaceborne model; model integration, that is, regarding the network to be integrated as a separate model and making decisions or predictions together with the initial spaceborne large model.

[0072] 105. Use the sample data set to train the spaceborne large model to be trained, and obtain a spaceborne large model that meets the conditions for improving robustness.

[0073] For the embodiments of the present invention, after optimizing and adjusting the structure of the initial spaceborne large model to obtain the spaceborne large model to be trained, in order to further improve the robustness of the initial spaceborne large model, it is necessary to use the sample data set to train the spaceborne large model to be trained. Based on this, the method includes: The specific training method includes: dividing the sample data set into training data and test data; using the training data to train the spaceborne large model to be trained to obtain the trained spaceborne large model; using the test data to test the trained spaceborne large model, and determining the trained spaceborne large model that meets the test conditions as the spaceborne large model that meets the conditions for improving robustness. Since the spaceborne large model will face various environmental interferences during operation, such as cosmic radiation, temperature changes, etc., the embodiments of the present invention optimize and adjust the structure of the spaceborne large model and train it, so that the model can better adapt to environmental changes, maintain stable performance output, and thus effectively improve the robustness of the spaceborne large model.

[0074] A method for improving the robustness of an on-board large model provided by the present invention. Compared with the current method of directly training the original model using the raw data obtained from the network, the present invention obtains a sample data set, task attribute information of the to-be-predicted task corresponding to the initial on-board large model, device attribute information of the on-board computing device where the initial on-board large model is deployed, and the expected robustness performance index in response to the robustness improvement signal of the initial on-board large model. Among them, the sample data set includes sample data matching the to-be-predicted task and its corresponding annotation information; and based on the task attribute information and the device attribute information, an expected model structure adapted to the to-be-predicted task is determined; at the same time, based on the task attribute information and the expected robustness performance index, an external network to be integrated with the initial on-board large model is determined; then, based on the expected model structure and the external network to be integrated, the structure of the initial on-board large model is optimized and adjusted to obtain a to-be-trained on-board large model; finally, the to-be-trained on-board large model is trained using the sample data set to obtain an on-board large model that meets the robustness improvement conditions. Thus, before training the initial on-board large model using the sample data, first, according to the task attribute information and the device attribute information, an expected model structure that can improve the robustness is determined, and according to the task attribute information and the expected robustness performance index, an external network to be integrated that can improve the robustness is determined. Finally, the expected model structure and the external network to be integrated are used to optimize and adjust the structure of the initial on-board large model, and the optimized model structure is trained using the sample data. The appropriate on-board large model structure can fully capture the complex features in the data and will not suffer from underfitting or overfitting during the training process, thereby improving the training effect of the on-board large model, that is, improving the robustness improvement effect of the on-board large model. At the same time, training the appropriate on-board large model structure will avoid the time and resources consumed by training a model with an overly complex structure. Therefore, the present invention can improve the efficiency of improving the robustness of the on-board large model and save the resources for improving the robustness, and further improve the prediction accuracy of the trained on-board large model.

[0075] Further, to better illustrate the process of improving the robustness of the on-board large model above, as a refinement and extension of the above embodiment, the embodiment of the present invention provides another method for improving the robustness of the on-board large model, as Figure 2 shown, the method includes:

[0076] 201. In response to a robustness enhancement signal of an initial satellite-borne large model, a sample data set, task attribute information of a task to be predicted corresponding to the initial satellite-borne large model, device attribute information of a satellite-borne computing device deployed by the initial satellite-borne large model, and expected robustness performance indicators of the initial satellite-borne large model are obtained, wherein the sample data set includes sample data matching the task to be predicted and its corresponding annotation information.

[0077] Specifically, the expected robustness performance index is set according to the prediction requirements of the onboard large model. A sample data set corresponding to the task to be predicted can be obtained in the network, and task attribute information of the task to be predicted, device attribute information of the onboard computing device, and other information can be obtained in the database.

[0078] 202. Determine abnormal data in the sample data, and remove the abnormal data in the sample data to obtain abnormal sample data.

[0079] For the embodiment of the present invention, in order to improve the quality of the sample data in the sample data set, it is first necessary to perform anomaly detection on the sample data in the sample data set. Based on this, step 202 specifically includes: determining the data density within the preset neighborhood corresponding to each of the sample data, and determining the core data points and non-core data points in each of the sample data based on the data density; taking any core data point in the core data points as a target core data point, clustering the remaining core data points with the target core data point as the cluster center, and obtaining core data points under different cluster categories; taking any non-core data point in the non-core data points as a target non-core data point, determining the reference core data point closest to the non-core data point in the core data points under each of the cluster categories, and judging whether the distance between the non-core data point and the reference core data point is less than a preset distance threshold; if the distance between the non-core data point and the reference core data point is less than the preset distance threshold, the non-core data point is classified into the cluster category to which the reference core data point belongs, otherwise, the non-core data point is determined as abnormal data.

[0080] Among them, the preset neighborhood is set according to actual needs, and the preset distance threshold is set according to actual needs. The data density can specifically be the data volume. Specifically, if the data density within the preset neighborhood corresponding to a certain sample data is greater than the preset density threshold, then this sample data is determined as a core data point. On the contrary, if the data density within the preset neighborhood corresponding to a certain sample data is less than or equal to the preset density threshold, then this sample data is determined as a non-core data point. Then, each core data point is used as a clustering center respectively to cluster the remaining core data points to obtain different clustering categories. Then, it is determined whether each non-core data point can be clustered into the above-mentioned clustering categories. Finally, the non-core points that cannot be clustered into all the above-mentioned clustering categories are determined as abnormal data. Then, the abnormal data in the sample data is deleted to obtain the sample data without anomalies. In another embodiment of the present invention, if the spaceborne large model is a CO2 column concentration prediction model, the sample data includes: L1-level data (spectral data), L2-level data (CO2 component content data with geographical information), meteorological data, pollutant data, vegetation data, terrain data, etc. It is possible to perform operations such as screening, spatio-temporal matching, derivative transformation, and correlation analysis on the L1-level data and L2-level data. According to the correlation analysis results, redundant data is removed, and spectral feature screening, longitude and latitude identification, date, observation parameter identification, etc. are performed on the data after removing the redundant data. For meteorological data, pollutant data, vegetation data, terrain data, etc., operations such as format conversion and variable extraction, data cleaning, band superposition, resampling, spatio-temporal matching, etc. can be performed. Finally, operations such as data integration, data cleaning, and variable correlation analysis are sequentially performed on the L1-level data, L2-level data, meteorological data, pollutant data, vegetation data, terrain data, etc. after the above operations. By preprocessing the sample data in the embodiments of the present invention, the data quality can be improved, the model performance can be enhanced, and the calculation cost can be reduced.

[0081] 203. Smooth the sample data without anomalies to obtain smooth sample data.

[0082] For the embodiments of the present invention, the specific smoothing method includes: arranging the outlier-removed sample data in the order of sampling time to obtain an outlier-removed sample data sequence; determining the moving average window value k; starting from the first data point of the outlier-removed sample data sequence, selecting a data segment with a window size of k, and calculating the average value of the data segment as the first smoothed data point; moving the window one data point to the right, and repeating the above steps until all smoothed data points are calculated; arranging all the calculated smoothed data points in the order of window movement to form a smoothed data sequence, and determining the data in the smoothed data sequence as the smoothed sample data, wherein the length of the smoothed data sequence is the same as that of the original outlier-removed sample data sequence. In another embodiment of the present invention, noise data in the sample data can also be determined and removed by means of filtering. By smoothing the sample data, noise data in the sample data can be removed, the data quality can be improved, and thus the training effect of the spaceborne large model can be improved.

[0083] 204. Determine whether there are missing values in the smoothed sample data. If so, supplement the missing values in the smoothed sample data to obtain supplemented sample data.

[0084] Further, if there are missing values in the smoothed sample data, interpolation, filling and other methods are used to supplement the missing values to obtain supplemented sample data. For example, interpolation is performed based on the linear relationship or polynomial relationship between adjacent data points to estimate the missing values. For time series data, the missing values can be filled with the previous or next known value. Thus, by filling the missing values, the integrity of the sample data set can be ensured, and the improvement effect of the robustness of the spaceborne large model can be enhanced.

[0085] 205. Perform data augmentation on the supplemented sample data to obtain a sample data set adaptively adjusted for robustness improvement.

[0086] Among them, the supplemented sample data can be at least one of image data, audio data, text data, spectral data, etc. For the embodiments of the present invention, in order to increase the amount of data in the sample data set and thus improve the training effect of the spaceborne large model, data augmentation needs to be performed on the sample data in the sample data set. Based on this, step 205 specifically includes: when the supplemented sample data is image data, performing at least one of rotation, translation, scaling, cropping, and color transformation on the image data to generate enhanced image data, and the enhanced image data and the image data constitute the sample data set adaptively adjusted; when the supplemented sample data is text data, performing at least one of synonym replacement, sentence restructuring, and text backtranslation on the text data to generate enhanced text data, and the enhanced text data and the text data constitute the sample data set adaptively adjusted.

[0087] Specifically, if the task to be predicted is an image classification task, the sample data is image data. At this time, the image data can be rotated at different angles around a certain center point to obtain multiple pieces of image data. The image can also be transformed such as translated and scaled in size to obtain more image data, thereby realizing the enhancement of the sample data, that is, increasing the data volume of the sample data. If the task to be predicted is a natural language processing task, the sample data can be text data. At this time, more training texts can be generated by performing operations such as random replacement, deletion, insertion, synonym replacement, sentence recombination, and text back-translation on the text. At the same time, data enhancement can also be performed through methods such as logarithmic transformation, differential transformation, reciprocal transformation, and power operation. By increasing the sample data volume in the embodiments of the present invention, the diversity of the data can be increased, enabling the model to learn more robust features, thereby improving the adaptability of the model to input changes and the understanding ability of the model to different expression forms.

[0088] In another embodiment of the present invention, in order to increase the data volume in the sample data set, derivative transformation can also be performed on the sample data in the sample data set. Based on this, the method includes: determining the correlation between the sample data in the sample data set by using the correlation coefficient analysis method; determining the highly correlated sample data with a correlation greater than a preset threshold; performing derivative transformation on the highly correlated sample data to obtain derivative data; adding the derivative data to the sample data set to obtain an expanded sample data set, and finally using the expanded sample data set to train the on-orbit large model to be trained.

[0089] In another embodiment of the present invention, new training samples and labels can also be constructed in a linear interpolation manner. That is, two randomly selected training samples and their corresponding labels are mixed to generate new training samples and labels. Specifically, given two samples (x1, y1) and (x2, y2), new samples (λx1+(1 - λ)x2, λy1+(1 - λ)y2) are generated through the parameter λ (λ ∈ [0, 1], usually following the Beta distribution). This method introduces prior knowledge to the model: linear interpolation in the feature space corresponds to linear interpolation in the label space, thereby enhancing the generalization ability of the model.

[0090] 206. Based on the task attribute information and device attribute information, determine the expected model structure suitable for the task to be predicted.

[0091] Specifically, according to information such as the task type, task complexity, and data volume of the task to be predicted, as well as information such as the computing resources and processing capabilities of the on-orbit computing device where the on-orbit large model is deployed, determine the number of network layers and activation functions of the on-orbit large model with robustness meeting the requirements.

[0092] 207. Determine the external network to be integrated that is adapted to the initial on-board large model based on the task attribute information and the expected robustness performance metrics.

[0093] Among them, the external network to be integrated includes convolutional neural network, recurrent neural network, generative adversarial network, etc. Specifically, according to information such as the task type, task complexity, and data volume of the task to be predicted, as well as the expected robustness performance metrics (such as expected stability, etc.) preset for the on-board large model, determine whether an external network is needed to improve the robustness of the on-board large model. In the case where an external network is needed, specifically what kind of external network structure is required.

[0094] 208. Based on the expected model structure and the external network to be integrated, perform structural optimization and adjustment on the initial on-board large model to obtain the on-board large model to be trained.

[0095] For the embodiments of the present invention, after determining the expected model structure and activation function that meet the robustness improvement conditions, it is necessary to perform structural optimization and adjustment on the initial on-board large model based on the expected model structure and activation function. Based on this, step 208 specifically includes: gradually adjusting the number of network layers in the initial on-board large model according to the expected number of network layers to obtain the adjusted initial on-board large model; using the expected activation function to replace the original activation function in the adjusted initial on-board large model to obtain the replaced initial on-board large model; integrating the external network to be integrated into the replaced initial on-board large model to obtain the on-board large model to be trained. Among them, the method of integrating the external network to be integrated into the replaced initial on-board large model includes: encapsulating the external network to be integrated as a function to be called; adding a function reference statement in the code of the replaced initial on-board large model and determining the integration position of the external network to be integrated in the on-board large model to be trained; based on the function reference statement, referring the function to be called to the integration position.

[0096] Specifically, determine which specific network layers to increase or decrease, and then gradually adjust the number of network layers in the initial on-board large model. At the same time, use activation functions to replace the original activation functions in the initial on-board large model to obtain the replaced initial on-board large model. Experiments and parameter tuning can be carried out when replacing the activation functions to find the activation function most suitable for specific tasks and data. Further, encapsulate the external network to be integrated as a function to be called, and add a reference statement to the external network function at an appropriate position in the code of the replaced initial on-board large model. Then, based on the function of the external network to be integrated, determine its integration position in the on-board large model, and select a suitable point in the forward propagation process of the on-board large model to insert a call to the external network function using the reference statement, that is, use the function reference statement to reference the function to be called to the integration position. This point should be able to receive the necessary input data and combine the output of the external network with other parts of the model.

[0097] In another embodiment of the present invention, the external network to be integrated can also be integrated into the on-board large model in the following manner. First, it is necessary to clarify the specific functions and objectives of the external network to be integrated and their relationship with the initial on-board large model. When integrating the external network, it is necessary to define the interfaces and communication protocols between them to ensure that they can communicate and interact with each other. Interface standardization can ensure that each module can be correctly called and communicate with each other. For example, if the external network to be integrated is a new model, fusion can be considered at the output level. If both the external network to be integrated and the on-board large model provide probability outputs, fusion can be carried out at the probability level. Or graft part of the structure and weights of the external network to be integrated onto the on-board large model and undergo a certain continued pre-training process to enable its model parameters to adapt to the new model.

[0098] In another embodiment of the present invention, a standardized external module can also be referenced in the on-board large model, that is, the input standardization of the activation layer. The standardized external module standardizes the input of the activation layer so that the standardized input can fall within the non-saturated region of the activation function. This helps to improve the stability and convergence speed of the on-board large model, and at the same time can also improve the adaptability of the model to different data distributions, thereby enhancing the robustness.

[0099] 209. Train the on-board large model to be trained using the adaptively adjusted sample data set to obtain an on-board large model that meets the conditions for improving robustness.

[0100] For the embodiments of the present invention, the on-orbit large model to be trained can be an information prediction model for predicting the carbon dioxide column concentration in a target area. At this time, the sample data set includes sample data and the label information of the corresponding carbon dioxide column concentration. The sample data includes: meteorological data, environmental data, terrain data, historical carbon dioxide column concentration data, production activity data, ecosystem data, prediction time, etc. of multiple regions. To train the information prediction model, step 209 specifically includes: determining the current model parameters of the on-orbit large model to be trained; based on the current model parameters, determining the overfitting constraint term of the on-orbit large model to be trained; according to the cross-entropy between the labeled carbon dioxide column concentration corresponding to the sample data in the sample data set and the predicted carbon dioxide column concentration corresponding to the sample data predicted by the on-orbit large model to be trained, determining the initial loss function of the on-orbit large model to be trained; adding the overfitting constraint term to the initial loss function to obtain the target loss function, and using the target loss function to train the on-orbit large model to be trained to obtain an on-orbit large model that meets the robustness improvement condition.

[0101] Among them, the model parameters can be the weights of each network structure in the on-orbit large model to be trained. Specifically, first, the sample data (such as meteorological data, environmental data, terrain data, historical carbon dioxide column concentration data, production activity data, ecosystem data, prediction time) in the sample data set is input into the on-orbit large model to be trained for CO2 column concentration prediction to obtain the predicted CO2 column concentration. The difference between the predicted CO2 column concentration and the labeled carbon dioxide column concentration in the same area is used to determine the initial loss function L 0 , and then the target loss function L is determined according to the following formula:

[0102]

[0103] Among them, λ is the overfitting constraint intensity parameter, and ω i is the i-th model parameter. Finally, the on-orbit large model to be trained is trained using the target loss function and the sample data set until an on-orbit large model that meets the robustness improvement condition is obtained. By constraining the loss function with the overfitting constraint term in the embodiments of the present invention, the large model can minimize the original loss function during training and maintain the simplicity of the large model. This helps to prevent the large model from overfitting to the noise and details in the training data, thereby improving the generalization ability of the model on unseen data. That is, by constraining the complexity of the large model parameters, the large model becomes more robust and is not easily affected by the random fluctuations in the training data, thereby improving the robustness improvement effect of the on-orbit large model.

[0104] In another embodiment of the present invention, the spaceborne large model can also be trained in the following manner: determining the number of categories of the annotation information corresponding to the sample data in the sample dataset, and determining the information vector corresponding to the annotation information; based on the number of categories and the information vector, determining the fuzzy annotation information corresponding to the sample data; determining the loss function of the spaceborne large model to be trained according to the cross entropy between the fuzzy annotation information corresponding to the sample data and the predicted annotation information corresponding to the sample data predicted by the spaceborne large model to be trained; using the loss function to train the spaceborne large model to be trained to obtain a spaceborne large model that meets the condition of improved robustness.

[0105] Specifically, for example, the annotation information of a sample is represented as a vector with a length equal to the number of categories, where the position of the correct category is 1 and the other positions are 0. For example, if the number of categories is 3 and the sample belongs to the second category, the information vector is [0, 1, 0]. Select a small positive number ε (set according to actual needs). For the probability of the correct category, reduce it from 1 to 1 - ε. For the probabilities of all other categories, evenly distribute the remaining ε among these categories. If ε = 0.1, the smoothed label distribution is [0.05, 0.9, 0.05]. Here, the probability of the correct category (the second category) is reduced from 1 to 0.9, and the remaining probability of 0.1 is evenly distributed to the other two categories. In this way, the fuzzy annotation information of the sample data can be obtained. Then, the sample data is input into the spaceborne large model to be trained for label prediction. By calculating the cross entropy between the fuzzy label distribution and the probability distribution predicted by the model, the loss function is obtained. Finally, the loss function and the sample dataset are used to train the spaceborne large model to be trained until a spaceborne large model that meets the condition of improved robustness is obtained. In the embodiment of the present invention, by performing fuzzy processing on the original label information of the sample data, the label distribution becomes smoother, thereby avoiding the over-reliance of the model on the training data, improving its generalization ability, and further enhancing the robustness improvement effect of the spaceborne large model.

[0106] Furthermore, during the model training process, the model can also be prevented from overfitting by randomly discarding (i.e., setting the output to 0) a part of the neurons in the neural network. During training, each random discard will result in a new sub-network. During prediction, all neurons will function, which can be regarded as the average of multiple sub-networks. The parameters are shared among multiple sub-networks, and at the same time, the neurons are randomly discarded. By the above method, the cooperative relationship between neurons can be effectively alleviated, and the robustness of the model can be improved.

[0107] Another method for improving the robustness of on-board large models provided by the present invention. Compared with the current method of directly training the original model using the raw data obtained from the network, the present invention obtains a sample data set, task attribute information of the to-be-predicted task corresponding to the initial on-board large model, device attribute information of the on-board computing device on which the initial on-board large model is deployed, and the expected robustness performance index in response to the robustness improvement signal of the initial on-board large model. Among them, the sample data set includes sample data matching the to-be-predicted task and its corresponding annotation information; and based on the task attribute information and the device attribute information, an expected model structure adapted to the to-be-predicted task is determined; at the same time, based on the task attribute information and the expected robustness performance index, an external network to be integrated with the initial on-board large model is determined; then, based on the expected model structure and the external network to be integrated, the structure of the initial on-board large model is optimized and adjusted to obtain a to-be-trained on-board large model; finally, the to-be-trained on-board large model is trained using the sample data set to obtain an on-board large model that meets the robustness improvement conditions. Thus, before training the initial on-board large model using the sample data, first, according to the task attribute information and the device attribute information, an expected model structure that can improve the robustness is determined, and according to the task attribute information and the expected robustness performance index, an external network to be integrated that can improve the robustness is determined. Finally, the expected model structure and the external network to be integrated are used to optimize and adjust the structure of the initial on-board large model, and the optimized model is trained using the sample data. The appropriate on-board large model structure can fully capture the complex features in the data and will not underfit or overfit during the training process, thereby improving the training effect of the on-board large model, that is, improving the robustness improvement effect of the on-board large model. At the same time, training the appropriate on-board large model structure will avoid the time and resources consumed by training a model with an overly complex structure. Therefore, the present invention can improve the efficiency of improving the robustness of the on-board large model and save the resources for improving the robustness, and further improve the prediction accuracy of the trained on-board large model.

[0108] Further, as Figure 1 a specific implementation of Figure 3 shown, the embodiment of the present invention provides a device for improving the robustness of an on-board large model, as

[0109] The obtaining unit 31 can be used to obtain a sample data set, task attribute information of a to-be-predicted task corresponding to the initial on-board large model, device attribute information of an on-board computing device on which the initial on-board large model is deployed, and an expected robustness performance index of the initial on-board large model in response to a robustness improvement signal of the initial on-board large model. Wherein, the sample data set includes sample data matching the to-be-predicted task and corresponding annotation information.

[0110] The structure determination unit 32 can be used to determine an expected model structure adapted to the to-be-predicted task based on the task attribute information and the device attribute information.

[0111] The network determination unit 33 can be used to determine an external network to be integrated adapted to the initial on-board large model based on the task attribute information and the expected robustness performance index.

[0112] The structure adjustment unit 34 can be used to perform structure optimization and adjustment on the initial on-board large model based on the expected model structure and the external network to be integrated to obtain a to-be-trained on-board large model.

[0113] The training unit 35 can be used to train the to-be-trained on-board large model by using the sample data set to obtain an on-board large model that meets the robustness improvement conditions.

[0114] In a specific application scenario, in order to perform adaptation adjustment on the sample data in the sample data set, as Figure 4 shown, the device further includes an adaptation adjustment unit 36.

[0115] The adaptation adjustment unit 36 can be used to determine abnormal data in the sample data, remove the abnormal data in the sample data to obtain de-abnormalized sample data; perform smoothing processing on the de-abnormalized sample data to obtain smoothed sample data; determine whether there are missing values in the smoothed sample data, and if so, supplement the missing values in the smoothed sample data to obtain supplemented sample data; perform data augmentation on the supplemented sample data to obtain the sample data set adapted and adjusted for robustness improvement.

[0116] In a specific application scenario, in order to determine abnormal data in the sample data, the adaptation adjustment unit 36 includes a first determination module 361, a clustering module 362, a judgment module 363, and a division module 364.

[0117] The first determination module 361 can be used to determine the data density within a preset neighborhood corresponding to each sample data, and determine core data points and non-core data points in each sample data based on the data density.

[0118] The clustering module 362 can be used to take any of the core data points in the core data points as a target core data point respectively, and use the target core data point as the clustering center to cluster the remaining core data points to obtain the core data points under different clustering categories.

[0119] The judgment module 363 can be used to take any of the non-core data points in the non-core data points as a target non-core data point respectively, determine the reference core data point closest to the non-core data point among the core data points under each clustering category, and judge whether the distance between the non-core data point and the reference core data point is less than a preset distance threshold.

[0120] The partitioning module 364 can be used to, if the distance between the non-core data point and the reference core data point is less than the preset distance threshold, partition the non-core data point into the clustering category to which the reference core data point belongs; otherwise, determine the non-core data point as abnormal data.

[0121] In a specific application scenario, in order to enhance the sample data, the adaptation and adjustment unit 36 further includes: an enhancement module 36.

[0122] The enhancement module 36 can be used to, when the supplementary sample data is image data, perform at least one of operations such as rotation, translation, scaling, cropping, and color transformation on the image data to generate enhanced image data, and the enhanced image data and the image data constitute the sample data set after adaptation and adjustment; when the supplementary sample data is text data, perform at least one of operations such as synonym replacement, sentence restructuring, and text back-translation on the text data to generate enhanced text data, and the enhanced text data and the text data constitute the sample data set after adaptation and adjustment.

[0123] In a specific application scenario, the on-orbit large model to be trained is an information prediction model for predicting the carbon dioxide column concentration in a target area. The sample data set includes sample data and its corresponding labeled carbon dioxide column concentration. The sample data includes: meteorological data, environmental data, terrain data, historical carbon dioxide column concentration data, production activity data, ecosystem data, and prediction time of multiple regions. In order to train the on-orbit large model to be trained, the training unit 35 includes a second determination module 351 and a training module 352.

[0124] The second determination module 351 can be used to determine the current model parameters of the on-orbit large model to be trained.

[0125] The second determination module 351 may specifically be configured to determine an overfitting constraint term of the to-be-trained spaceborne large model based on the current model parameters.

[0126] The second determination module 351 may also be configured to determine an initial loss function of the to-be-trained spaceborne large model according to the cross entropy between the labeled carbon dioxide column concentration corresponding to the sample data in the sample dataset and the predicted carbon dioxide column concentration corresponding to the sample data predicted by the to-be-trained spaceborne large model.

[0127] The training module 352 may be configured to add the overfitting constraint term to the initial loss function to obtain a target loss function, and use the target loss function to train the to-be-trained spaceborne large model to obtain a spaceborne large model that meets the robustness improvement condition.

[0128] In a specific application scenario, in order to train the to-be-trained spaceborne large model, the second determination module 351 may also be configured to determine the number of categories of the labeled information corresponding to the sample data in the sample dataset, and determine the information vector corresponding to the labeled information.

[0129] The second determination module 351 may also be configured to determine the fuzzy labeled information corresponding to the sample data based on the number of categories and the information vector.

[0130] The second determination module 351 may also be configured to determine the loss function of the to-be-trained spaceborne large model according to the cross entropy between the fuzzy labeled information corresponding to the sample data and the predicted labeled information corresponding to the sample data predicted by the to-be-trained spaceborne large model.

[0131] The training module 352 may also be configured to use the loss function to train the to-be-trained spaceborne large model to obtain a spaceborne large model that meets the robustness improvement condition.

[0132] In a specific application scenario, the expected model structure includes an expected number of network layers and an expected activation function. In order to optimize the structure of the initial spaceborne large model, the structure adjustment unit 34 includes a layer adjustment module 341, a function replacement module 342, and an integration module 343.

[0133] The layer adjustment module 341 may be configured to gradually adjust the number of network layers in the initial spaceborne large model according to the expected number of network layers to obtain the adjusted initial spaceborne large model.

[0134] The function replacement module 342 may be configured to use the expected activation function to replace the original activation function in the adjusted initial spaceborne large model to obtain the replaced initial spaceborne large model.

[0135] The integration module 343 can be used to integrate the external network to be integrated into the replaced initial on-board large model to obtain the on-board large model to be trained.

[0136] In a specific application scenario, in order to integrate the external network to be integrated into the replaced initial on-board large model, the integration module 343 can specifically be used to encapsulate the external network to be integrated into a function to be called; add a function reference statement to the code of the replaced initial on-board large model, and determine the integration position of the external network to be integrated in the on-board large model to be trained; based on the function reference statement, reference the function to be called to the integration position.

[0137] It should be noted that for other corresponding descriptions of each functional module involved in the device for improving the robustness of the on-board large model provided in the embodiments of the present invention, reference can be made to Figure 1 the corresponding description of the method shown, which will not be elaborated here.

[0138] Based on the above method as Figure 1 shown, correspondingly, the embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the following steps are implemented: in response to a signal for improving the robustness of the initial on-board large model, obtain a sample data set, task attribute information of the task to be predicted corresponding to the initial on-board large model, device attribute information of the on-board computing device on which the initial on-board large model is deployed, and the expected robustness performance index of the initial on-board large model, where the sample data set includes sample data matching the task to be predicted and its corresponding annotation information; based on the task attribute information and the device attribute information, determine an expected model structure adapted to the task to be predicted; based on the task attribute information and the expected robustness performance index, determine an external network to be integrated adapted to the initial on-board large model; based on the expected model structure and the external network to be integrated, perform structural optimization and adjustment on the initial on-board large model to obtain an on-board large model to be trained; use the sample data set to train the on-board large model to be trained to obtain an on-board large model that meets the conditions for improving robustness.

[0139] Based on the above method as Figure 1 shown and the embodiments of the device as Figure 3 shown, the embodiments of the present invention further provide an entity structure diagram of a computer device, as Figure 5As shown, the computer device includes: a processor 41, a memory 42, and a computer program stored on the memory 42 and executable on the processor. Both the memory 42 and the processor 41 are provided on a bus 43. When the processor 41 executes the program, the following steps are implemented: in response to a robustness improvement signal of the initial on-board large model, obtain a sample data set, task attribute information of the prediction task corresponding to the initial on-board large model, device attribute information of the on-board computing device on which the initial on-board large model is deployed, and the expected robustness performance index of the initial on-board large model. Among them, the sample data set includes sample data matching the prediction task and its corresponding annotation information; based on the task attribute information and the device attribute information, determine an expected model structure adapted to the prediction task; based on the task attribute information and the expected robustness performance index, determine an external network to be integrated adapted to the initial on-board large model; based on the expected model structure and the external network to be integrated, perform structural optimization and adjustment on the initial on-board large model to obtain a large on-board model to be trained; use the sample data set to train the large on-board model to be trained to obtain an on-board large model that meets the robustness improvement conditions.

[0140] Through the technical solution of the present invention, the present invention obtains a sample data set, task attribute information of a to-be-predicted task corresponding to the initial on-board large model, device attribute information of an on-board computing device on which the initial on-board large model is deployed, and an expected robustness performance index of the initial on-board large model in response to a robustness improvement signal of the initial on-board large model. Wherein, the sample data set includes sample data matching the to-be-predicted task and corresponding annotation information thereof; and based on the task attribute information and the device attribute information, an expected model structure adapted to the to-be-predicted task is determined; at the same time, based on the task attribute information and the expected robustness performance index, an external network to be integrated with the initial on-board large model is determined; then, based on the expected model structure and the external network to be integrated, the structure of the initial on-board large model is optimized and adjusted to obtain a to-be-trained on-board large model; finally, the to-be-trained on-board large model is trained using the sample data set to obtain an on-board large model that meets the robustness improvement condition. Thus, before training the initial on-board large model using sample data, first, according to the task attribute information and the device attribute information, an expected model structure capable of improving robustness is determined, and according to the task attribute information and the expected robustness performance index, an external network to be integrated capable of improving robustness is determined. Finally, the expected model structure and the external network to be integrated are used to optimize and adjust the structure of the initial on-board large model, and the optimized model structure is trained using sample data. A suitable on-board large model structure can fully capture complex features in the data and will not exhibit underfitting or overfitting during the training process, thereby being able to improve the training effect of the on-board large model, that is, improve the robustness improvement effect of the on-board large model. At the same time, training a suitable on-board large model structure will avoid the time and resources consumed by training a model with an overly complex structure. Thus, the present invention can improve the efficiency of improving the robustness of the on-board large model and save the resources for improving the robustness, and further improve the prediction accuracy of the trained on-board large model.

[0141] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to be implemented. In this way, the present invention is not limited to any specific combination of hardware and software.

[0142] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for improving the robustness of a large satellite model, characterized in that: include: In response to the robustness improvement signal of the initial onboard large model, a sample data set, task attribute information of the task to be predicted corresponding to the initial onboard large model, device attribute information of the onboard computing device deployed by the initial onboard large model, and an expected robustness performance index of the initial onboard large model are acquired, wherein the sample data set includes sample data matching the task to be predicted and its corresponding annotation information; Determining an expected model structure that matches the task to be predicted based on the task attribute information and the device attribute information; Based on the mission attribute information and the expected robustness performance index, determining an external network to be integrated that is compatible with the initial onboard large model; Based on the expected model structure and the external network to be integrated, the initial onboard large model is structurally optimized and adjusted to obtain the onboard large model to be trained; The sample data set is used to train the onboard large model to be trained, so as to obtain the onboard large model that meets the robustness improvement condition.

2. The method according to claim 1, characterized in that Before using the sample data set to train the onboard large model to be trained to obtain the onboard large model that meets the robustness improvement condition, the method further includes: Determine abnormal data in the sample data, and remove the abnormal data in the sample data to obtain abnormal sample data; Smoothing the removed-odds sample data to obtain smoothed sample data; Determine whether there are missing values ​​in the smoothed sample data, and if so, supplement the missing values ​​in the smoothed sample data to obtain supplemented sample data; Data enhancement is performed on the supplementary sample data to obtain the sample data set after adaptive adjustment for improving robustness.

3. The method according to claim 2, characterized in that The determining of abnormal data in the sample data includes: Determine the data density in a preset neighborhood corresponding to each of the sample data, and determine core data points and non-core data points in each of the sample data based on the data density; Taking any core data point among the core data points as a target core data point, taking the target core data point as the cluster center, clustering the remaining core data points to obtain core data points under different clustering categories; Taking any non-core data point among the non-core data points as a target non-core data point, determining a reference core data point that is closest to the non-core data point among the core data points under each clustering category, and judging whether the distance between the non-core data point and the reference core data point is less than a preset distance threshold; If the distance between the non-core data point and the reference core data point is less than the preset distance threshold, the non-core data point is classified into the cluster category to which the reference core data point belongs; otherwise, the non-core data point is determined as abnormal data.

4. The method according to claim 2, characterized in that: The supplementary sample data includes at least one of image data and text data; The step of performing data enhancement on the supplementary sample data to obtain the sample data set after adaptive adjustment for improving robustness includes: In the case where the supplementary sample data is image data, performing at least one operation of rotation, translation, scaling, cropping, and color conversion on the image data to generate enhanced image data, wherein the enhanced image data and the image data constitute the sample data set after adaptive adjustment; In the case where the supplementary sample data is text data, at least one operation of synonym replacement, sentence reorganization, and text back translation is performed on the text data to generate enhanced text data, and the enhanced text data and the text data constitute the sample data set after adaptation and adjustment.

5. The method according to claim 1, characterized in that The large satellite-borne model to be trained is an information prediction model, which is used to predict the carbon dioxide column concentration in the target area. The sample data set includes sample data and its corresponding annotated carbon dioxide column concentration. The sample data includes: meteorological data, environmental data, terrain data, historical carbon dioxide column concentration data, production activity data, ecosystem data, and prediction time of multiple regions; The step of training the onboard large model to be trained by using the sample data set to obtain the onboard large model that meets the robustness improvement condition includes: Determining current model parameters of the large onboard model to be trained; Based on the current model parameters, determining an overfitting constraint item of the large onboard model to be trained; Determining an initial loss function of the onboard large model to be trained according to a cross entropy between the annotated carbon dioxide column concentration corresponding to the sample data in the sample data set and the predicted carbon dioxide column concentration corresponding to the sample data predicted by the onboard large model to be trained; The overfitting constraint term is added to the initial loss function to obtain a target loss function, and the target loss function is used to train the large satellite-borne model to be trained to obtain a large satellite-borne model that meets the robustness improvement conditions.

6. The method according to claim 1, characterized in that The step of training the onboard large model to be trained by using the sample data set to obtain the onboard large model that meets the robustness improvement condition includes: Determine the number of categories of the annotation information corresponding to the sample data in the sample data set, and determine the information vector corresponding to the annotation information; Determining fuzzy labeling information corresponding to the sample data based on the number of categories and the information vector; Determining a loss function of the onboard large model to be trained according to a cross entropy between the fuzzy annotation information corresponding to the sample data and the predicted annotation information corresponding to the sample data predicted by the onboard large model to be trained; The loss function is used to train the large satellite-borne model to be trained, so as to obtain a large satellite-borne model that meets the robustness improvement condition.

7. The method according to claim 1, characterized in that The expected model structure includes an expected number of network layers and an expected activation function; The step of optimizing and adjusting the structure of the initial onboard large model based on the expected model structure and the external network to be integrated to obtain the onboard large model to be trained includes: Stepwise adjusting the number of network layers in the initial onboard large model according to the expected number of network layers to obtain the adjusted initial onboard large model; Replacing the original activation function in the adjusted initial onboard large model with the expected activation function to obtain the replaced initial onboard large model; Integrating the external network to be integrated into the replaced initial onboard large model to obtain the onboard large model to be trained, wherein the method of integrating the external network to be integrated into the replaced initial onboard large model comprises: Encapsulating the external network to be integrated as a function to be called; Adding a function reference statement in the code of the replaced initial onboard large model, and determining the integration position of the external network to be integrated in the initial onboard large model to be trained; Based on the function reference statement, the to-be-called function is referenced to the integration location.

8. A device for improving the robustness of a large satellite model, characterized in that: include: an acquisition unit, configured to acquire, in response to a robustness enhancement signal of an initial onboard large model, a sample data set, task attribute information of a task to be predicted corresponding to the initial onboard large model, device attribute information of an onboard computing device deployed by the initial onboard large model, and an expected robustness performance index of the initial onboard large model, wherein the sample data set includes sample data matching the task to be predicted and its corresponding annotation information; A structure determination unit, configured to determine an expected model structure that matches the task to be predicted based on the task attribute information and the device attribute information; A network determination unit, configured to determine an external network to be integrated that is compatible with the initial onboard large model based on the mission attribute information and the expected robustness performance index; A structure adjustment unit, configured to optimize and adjust the structure of the initial onboard large model based on the expected model structure and the external network to be integrated, so as to obtain the onboard large model to be trained; A training unit is used to train the onboard large model to be trained using the sample data set to obtain a onboard large model that meets the robustness improvement condition.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Method and device for improving performance of remote sensing large model based on composite visual coding

    CN120997529A

  • A remote sensing large model performance improvement method and device based on composite visual coding

    CN120997529B