Method and apparatus for training of large model based solution generation and optimization model

By using a large-scale model-based scheme generation and optimization model training method, the problem of low efficiency and poor reliability of manual scheme writing is solved, realizing the automated design and optimization of special equipment or systems, and improving the efficiency and accuracy of scheme generation.

CN119721940BActive Publication Date: 2025-11-18INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411478926.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-11-18
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

In existing technologies, the generation of experimental schemes for specialized equipment or systems relies on manual writing, which suffers from low efficiency, poor reliability, and strong subjectivity, making it difficult to meet the design requirements of complex tasks.

Method used

A training method based on a large model for scheme generation and optimization is adopted. The language model is iteratively trained using an autoregressive loss function and an improved opinion loss function to construct a scheme generation model, a feedback model, and an optimization model, thereby achieving automated design and optimization.

Benefits of technology

It improves the efficiency and accuracy of experimental scheme generation, realizes the automated design and optimization of special equipment or system schemes, and meets the design requirements of complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721940B_ABST
    Figure CN119721940B_ABST
Patent Text Reader

Abstract

The application provides a large model-based scheme generation and optimization model training method and device, the large model-based scheme generation and optimization model training method comprises the following steps: obtaining a sample experiment scheme and sample requirements thereof; taking the sample experiment scheme and the sample requirements as training samples, and training a scheme generation model by taking a first autoregressive loss as a loss function; taking a historical optimization trajectory and a target scheme as training samples, and training a feedback model by taking an improved opinion loss as a loss function; taking the sample requirements, the historical optimization trajectory, and sample improved opinions output by the feedback model in a historical training process as training samples, and training an optimization model by taking a second autoregressive loss as a loss function; and obtaining a scheme generation and optimization model based on the scheme generation model, the feedback model, and the optimization model; the method realizes automatic design and optimization of a scheme for a special device or system, and improves experiment scheme generation efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and in particular to a method and apparatus for training scheme generation and optimization models based on large models. Background Technology

[0002] With the rapid development of technology, the design complexity of specialized equipment or systems used for tasks such as rescue, disaster relief, or training has increased significantly. Such equipment or systems require precise experimental plans for performance verification and optimization. Test reports generated through these plans can effectively assess whether the equipment meets various performance requirements in specific mission scenarios, thereby ensuring its reliability and effectiveness in practical applications.

[0003] In related technologies, the generation of traditional experimental schemes mainly relies on manual writing, manual error correction and evaluation. This process requires designers to have rich professional knowledge and experience. Moreover, manually designed schemes have problems such as strong subjectivity and inconsistent standards, making it difficult to guarantee their scientificity and repeatability. In addition, as the complexity of the aforementioned special equipment or systems increases and the scheme generation efficiency is low, it cannot meet the design requirements of equipment or systems when performing complex tasks. Summary of the Invention

[0004] This invention provides a method and apparatus for generating and optimizing schemes based on large models, which solves the problem that existing technologies rely on human experience to specify special equipment or system design schemes, resulting in low scheme generation efficiency and low reliability, and improves the efficiency and accuracy of special equipment or system scheme generation.

[0005] This invention provides a method for generating and optimizing schemes based on large models, including:

[0006] Obtain the experimental design and sample requirements;

[0007] Using the sample experiment scheme and the sample requirements as training samples, the first large language model is iteratively trained with the first autoregressive loss as the loss function to obtain the scheme generation model; wherein, the first autoregressive loss is determined based on the cross-entropy between the sample experiment scheme and the sample requirements;

[0008] Using historical optimization trajectories and target schemes as training samples, and using improved opinion loss as the loss function, the second language processing is iteratively trained to obtain a feedback model; wherein, the improved opinion loss is determined based on the feedback value between the sample optimization scheme and the sample experimental scheme; the target scheme includes the sample optimization scheme or the sample preliminary scheme output by the scheme generation model during the historical training process;

[0009] Using the sample requirements, the historical optimization trajectory, and the sample improvement suggestions output by the feedback model during historical training as training samples, the third language processing is iteratively trained using the second autoregressive loss as the loss function to obtain an optimization model; a scheme generation and optimization model is obtained based on the scheme generation model, the feedback model, and the optimization model; wherein, the second autoregressive loss is determined based on the cross-entropy between the sample optimization scheme and the sample experimental scheme.

[0010] According to the present invention, a training method for a scheme generation and optimization model based on a large model is provided, wherein the feedback model and the optimization model are iteratively trained through cyclic optimization;

[0011] The training termination conditions for the feedback model and the optimization model include: determining, based on the improvement suggestions output by the feedback model, that the sample optimization scheme output by the optimization model meets the target requirements.

[0012] According to the present invention, a training method for a scheme generation and optimization model based on a large model is provided. The scheme generation model, the feedback model, and the optimization model are respectively trained through pre-trained language processing, and all of them use an autoregressive mechanism to encode and decode sample requirements. Different LoRA weights are used to fine-tune different pre-trained models.

[0013] According to the present invention, a training method for a scheme generation and optimization model based on a large model is provided, wherein the scheme generation model adopts a first LoRA weight; the first LoRA weight is obtained by autoregressive training of the pre-trained model through sample requirements and sample experimental schemes;

[0014] The feedback model employs a second LoRA weight; the second LoRA weight is obtained by autoregressive training of the pre-trained model using a PPO method optimized by combining sample improvement opinions with a proximal strategy.

[0015] The optimized model uses a third LoRA weight; the third LoRA weight is obtained by autoregressive training of the pre-trained model based on the improvement suggestions data corresponding to the sample requirements, sample experimental schemes, and sample preliminary schemes.

[0016] According to the training method for scheme generation and optimization models based on a large model provided by the present invention, the first autoregressive loss is expressed by the following formula:

[0017] ;

[0018] in, This is the first autoregressive loss value. For the sample experiment plan targeting sample requirements, the sentence in each text is the first... The probability of each word;

[0019] The loss from the proposed improvements is expressed by the following formula:

[0020] ;

[0021] in, To improve the loss value of the suggestion, For experimental feedback; For large language models that require parameter updates; This serves as the initial large language model; As an adaptive penalty factor;

[0022] The second autoregressive loss is expressed by the following formula:

[0023] ;

[0024] in, This is the second autoregressive loss value. The sentence representing the sample experiment plan, sample improvement suggestions, and historical optimization trajectory in each text related to sample requirements is the first... The probability of each word.

[0025] This invention also provides a method for generating and optimizing solutions based on a large model, comprising:

[0026] Obtain the target requirements of the experimental design to be developed;

[0027] The target requirements are processed based on the scheme generation and optimization model to obtain the target scheme; wherein the scheme generation and optimization model is trained by the training method of the scheme generation and optimization model based on the large model.

[0028] The present invention also provides a training device for scheme generation and optimization models based on large models, comprising:

[0029] The first acquisition module is used to acquire the sample experimental plan and its sample requirements;

[0030] The first training module is used to iteratively train the first large language model using the sample experimental scheme and the sample requirements as training samples and the first autoregressive loss as the loss function to obtain the scheme generation model; wherein, the first autoregressive loss is determined based on the cross-entropy between the sample experimental scheme and the sample requirements.

[0031] The second training module is used to iteratively train the second language processing model using historical optimization trajectories and target schemes as training samples and improved opinion loss as the loss function to obtain a feedback model; wherein, the improved opinion loss is determined based on the feedback value between the sample optimization scheme and the sample experimental scheme; the target scheme includes the sample optimization scheme or the preliminary sample scheme output by the scheme generation model during the historical training process;

[0032] The third training module is used to iteratively train the third language processing module using the sample requirements, the historical optimization trajectory, and the sample improvement suggestions output by the feedback model during the historical training process as training samples, and using the second autoregressive loss as the loss function to obtain an optimization model; and to obtain a scheme generation and optimization model based on the scheme generation model, the feedback model, and the optimization model; wherein, the second autoregressive loss is determined based on the cross-entropy between the sample optimization scheme and the sample experimental scheme.

[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method as described above.

[0034] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described above.

[0035] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the method as described above.

[0036] The present invention provides a training method and apparatus for a scheme generation and optimization model based on a large model. The method trains a scheme generation model using sample experimental schemes and sample requirements as training samples and a first autoregressive loss as the loss function. It then trains a feedback model using historical optimization trajectories and target schemes as training samples and an improvement suggestion loss as the loss function. Finally, it trains an optimization model using sample requirements, historical optimization trajectories, and sample improvement suggestions output by the feedback model during historical training as training samples and a second autoregressive loss as the loss function, through third language processing. This process determines the scheme generation and optimization model, enabling automated design and optimization of schemes for specialized equipment or systems, and improving the efficiency and accuracy of experimental scheme generation. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0038] Figure 1 This is a flowchart illustrating the training method for scheme generation and optimization models based on large models provided by the present invention.

[0039] Figure 2 This is one of the flowcharts illustrating the scheme generation and optimization method based on a large model provided by the present invention.

[0040] Figure 3 This is the second flowchart illustrating the scheme generation and optimization method based on a large model provided by the present invention.

[0041] Figure 4 This is a schematic diagram of the training device for scheme generation and optimization model based on a large model provided by the present invention.

[0042] Figure 5 This is a schematic diagram of the structure of the scheme generation and optimization device based on a large model provided by the present invention.

[0043] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0045] The following is combined Figures 1-5 This invention describes a method and apparatus for generating and optimizing schemes based on large models.

[0046] Figure 1 This is a flowchart illustrating the training method for scheme generation and optimization models based on large models provided by this invention, as shown below. Figure 1 As shown, the training method for scheme generation and optimization models based on large models includes the following steps:

[0047] Step 110: Obtain the sample experimental plan and its sample requirements;

[0048] Step 120: Using the sample experimental plan and sample requirements as training samples, and using the first autoregressive loss as the loss function, iteratively train the first large language model to obtain the plan generation model; wherein, the first autoregressive loss is determined based on the cross-entropy between the sample experimental plan and the sample requirements.

[0049] Step 130: Using historical optimization trajectories and target schemes as training samples, and using improved opinion loss as the loss function, iteratively train the second language processing model to obtain a feedback model; wherein, the improved opinion loss is determined based on the feedback value between the sample optimization scheme and the sample experimental scheme; the target scheme includes the sample optimization scheme or the sample preliminary scheme output by the scheme generation model during the historical training process;

[0050] Step 140: Using sample requirements, historical optimization trajectories, and sample improvement suggestions output by the feedback model during historical training as training samples, iteratively train the third language processing using the second autoregressive loss as the loss function to obtain the optimization model; obtain the scheme generation and optimization model based on the scheme generation model, feedback model, and optimization model; wherein, the second autoregressive loss is determined based on the cross-entropy between the sample optimization scheme and the sample experimental scheme.

[0051] In step 110, the sample experiment scheme can be an existing design scheme for special equipment or systems used for training in rescue, disaster relief or specific tasks. This design scheme can be determined based on expert experience or obtained through natural language processing technology.

[0052] In this embodiment, the requirements corresponding to the experimental scheme include relevant information about the device under test and the required indicators. The relevant information about the device under test may include, but is not limited to, the application scenario of the device under test, the function of the device under test, and the operating frequency range of the device under test. The required indicators mainly include the subjects to be examined and the specific indicators and requirements of the relevant subjects (such as the specific functional boundary range).

[0053] In this embodiment, the sample requirements for dedicated wearable devices used for frontline disaster relief include lightweight, pressure resistance, and reliable communication.

[0054] In this embodiment, the sample requirements can be pre-written manually, based on the requirements of the desired generated solution, and the front-end can be written accordingly.

[0055] Taking a radar with a working frequency in the GHz range as an example, the device under test is used for communication between aircraft, navigation, and communication with the base.

[0056] In this embodiment, the evaluation criteria for the radar are as follows: ① Examine the operational stability and signal response of the equipment in different frequency bands to ensure that the radar has no significant distortion or interference throughout the entire frequency range; ② Evaluate the communication performance of the radar at different distances to ensure that sufficient signal strength and clarity are maintained within the specified flight distance range; ③ Operational stability and signal response in different frequency bands; ④ Test the radar's ability to resist interference in complex electromagnetic environments, etc.

[0057] In this embodiment, the radar design can be used to test the radar's operational stability in a specified GHz frequency band, ensuring that its signal response is free from significant distortion or interference; and the following method is employed:

[0058] A spectrum analyzer is used to perform a comprehensive spectrum scan test on the radar equipment. Within the entire specified frequency band, the frequency is gradually adjusted and the signal strength, stability, and noise level at each frequency are recorded. The response of each frequency band is then compared and analyzed to detect any distortion, power drop, or other anomalies. Detailed spectrum records are made for each signal segment to ensure coverage of the entire frequency range.

[0059] Specifically, a high-precision spectrum analyzer is used to conduct tests in a shielded laboratory environment to minimize external electromagnetic interference. The frequency is gradually covered to cover the entire specified GHz band, and the test duration should exceed the equipment's predetermined operating time limit to examine its stability under long-term operation. The experimental scheme must meet the following target requirements: a reference signal-to-noise ratio (SNR) is set, and when the signal SNR is greater than a specified dB and there is no obvious spectral fluctuation, the signal response is considered good. If the radar signal does not show obvious distortion, noise, or interference throughout the entire frequency band, and the signal strength is stable, it is considered to have met the standards for operational stability and signal response.

[0060] In this embodiment, the radar design can also be used to evaluate the signal strength and clarity of the radar at different communication distances, ensuring that the signal can be transmitted clearly and stably within a specified flight range.

[0061] The specific method includes: first, simulating an aircraft flight environment, gradually increasing the distance between the radar and the remote communication equipment (from 0 to 100 kilometers), testing the radar's signal strength and communication clarity, and recording the signal attenuation at different distances; starting at 0 kilometers, gradually increasing to 100 kilometers, with each 10-kilometer distance test point; using a remote communication terminal and an aircraft simulator as the receiving end, recording indicators such as signal strength, data loss rate, and communication delay at each distance. Multiple tests are conducted at each distance point to eliminate random errors, and environmental conditions (such as weather conditions, interference sources, etc.) are recorded. The experimental scheme must meet the following objectives: conduct the test in open airspace or a simulated flight environment, ensuring no additional obstructions or interference; equip signal analysis equipment to monitor data transmission in real time, ensuring the reliability and repeatability of the test data; if the signal remains clear and stable within 90 kilometers, and the packet loss rate is less than 1%, it is considered to have passed the test standard; if signal attenuation occurs at 100 kilometers, but basic communication quality can still be maintained, further optimization of the communication distance is required.

[0062] In this embodiment, the radar design can also be used to evaluate the radar's anti-interference capability in complex electromagnetic environments, ensuring that it can maintain normal operation of communication and navigation functions when facing external electromagnetic interference, and avoiding signal interruption or failure due to interference.

[0063] Specifically, in a simulated electromagnetic interference environment, the radar's anti-interference capability is tested by gradually increasing the interference signal. First, in an interference-free environment, the normal signal output of the radar is confirmed, and baseline data is recorded. Then, using electromagnetic interference simulation equipment, different types of interference signals (such as continuous wave, pulse interference, and frequency hopping interference) are injected into the radar's operating frequency band. The intensity of the interference signal is gradually increased, and parameters such as radar signal quality, communication packet loss rate, navigation error, and communication delay are recorded. The experimental scheme must meet the following objectives: gradually increasing the interference intensity from low (10 dBm) to high (50 dBm) and observing the radar's response capability; simulating various types of interference, including frequency overlap interference and signal blocking interference, to comprehensively evaluate the radar's anti-interference performance; if the radar signal packet loss rate remains below 5% and signal quality can be quickly restored after the interference intensity increases to a certain threshold, it is considered to have passed the anti-interference test.

[0064] In the above steps, the first, second, and third language models can be models with the same architecture, or they can be different.

[0065] In the above steps, the historical optimization trajectory is used to represent the log data or label data generated during the process of the model generating a preliminary experimental plan according to the requirements. By analyzing this type of data, we can guide the generation of an experimental plan suitable for the specific task based on the requirements of that task.

[0066] In the above steps, the sample improvement suggestions are determined based on the differences between the simulated scheme output by the model and the actual scheme. For example, the preliminary experimental scheme output by the model does not include the indicator of "ability to resist interference in complex electromagnetic environments", while in the actual scheme, this indicator is a necessary setting indicator. Therefore, a sample modification suggestion based on self-reflection could be "there is no coverage of the indicator of ability to resist interference in complex electromagnetic environments".

[0067] In the above steps, the feedback model can predict the error between the actual value and the simulated value of various equipment attributes simulated by the large language model. The optimization model adaptively adjusts the design parameters of the experimental scheme based on the reference actual value and the corresponding error, which is conducive to the optimization model outputting a more reliable experimental scheme.

[0068] In this embodiment, the optimization model, combining the preliminary experimental plan and improvement suggestions, outputs the following optimization plan:

[0069] (1) Experimental objective: To evaluate the anti-interference capability of radar equipment in complex electromagnetic environments and ensure that it can still work stably in strong interference and complex signal environments, and maintain communication and navigation functions; this test is particularly critical to ensure the reliability of the equipment when facing various electromagnetic interferences in actual missions.

[0070] (2) Implementation method: A specialized electromagnetic interference simulator and signal generator were used to simulate various complex electromagnetic interference scenarios. By gradually increasing the interference intensity and changing the interference type, the performance of the radar equipment under strong interference and complex signal environments was tested; the interference types included:

[0071] Continuous wave interference: Simulates long-term, sustained interference at a specific frequency;

[0072] Pulse interference: Simulates intermittent strong interference signals;

[0073] Frequency jump interference: simulates interference signals with rapidly changing frequencies;

[0074] Multi-source interference: Simulates multiple interference sources applying interference simultaneously;

[0075] Interference intensity: The interference signal intensity gradually increases from low to high, with an initial intensity of 10 dBm and a maximum intensity of 50 dBm.

[0076] The experimental design must meet the following objectives: conduct the experiment in an electromagnetically shielded laboratory to avoid the influence of external interference on the test results; use multiple interference signal generators to simultaneously apply interference of different types and intensities to simulate a real complex electromagnetic environment; and install the equipment according to the actual task deployment method to ensure that the test environment is as close as possible to the actual application conditions.

[0077] In some embodiments, another self-reflective sample modification suggestion is that "no specific experimental equipment is given in the experimental conditions for the ability to resist interference in complex electromagnetic environments." Correspondingly, the optimization model, combining the preliminary experimental plan and the improvement suggestions, outputs the following optimized solutions:

[0078] The experimental objective and implementation method in this embodiment are the same as "(1) experimental objective and (2) implementation method" in the previous embodiment, and will not be repeated in this embodiment.

[0079] It should be noted that the experimental scheme implemented in this instance needs to meet the following objectives: The test should be conducted in an electromagnetically shielded laboratory to avoid the influence of external interference on the test results; multiple interference signal generators should be used to simultaneously apply interference of different types and intensities to simulate a real, complex electromagnetic environment; the equipment should be installed according to the actual task deployment method to ensure that the test environment is as close as possible to the actual application conditions. A spectrum analyzer is used to monitor the radar's spectrum response in real time under electromagnetic interference, recording signal strength, spectrum changes, signal loss, etc.; an interference power amplifier is used to transmit and receive both radar and interference signals; in addition, an antenna array is required to support the transmission and reception of specified GHz-level frequency bands, ensuring that the test radar signal and the interference signal cover the same frequency band.

[0080] In this embodiment, during the training phase of the above-mentioned major language models, the sample experiment scheme and sample requirements are first obtained, and the two constitute the training dataset. The training dataset can be divided into a training set and a test set. The training set is used for model training, and the test set is used for model testing.

[0081] In this embodiment, the scheme generation model, feedback model, and optimization model are trained through pre-trained language processing, and all of them use an autoregressive mechanism to encode and decode sample requirements. Different LoRA weights are used to fine-tune different pre-trained models.

[0082] In this embodiment, the scheme generation model is built based on the first large language model. The structure of the first large language model is similar to that of the Transformer-Decoder model. The first large language model adopts the attention mechanism, which can determine the vector representation of the preceding sentence based on the vector representation of the following sentence.

[0083] In this embodiment, the structures of the second and third language models can also be similar to those of the Transformer-Decoder model, and these two models can also use the Attention mechanism to focus on the corresponding text features.

[0084] In this embodiment, the three large language models are fine-tuned with different LoRA (Low-Rank Adaptation) pre-trained model weights, and the corresponding LoRA weights are obtained through autoregressive training.

[0085] In this embodiment, the first autoregressive loss is expressed by the following formula:

[0086] ;

[0087] in, This is the first autoregressive loss value. For the sample experiment plan targeting sample requirements, the sentence in each text is the first... The probability of each word.

[0088] In this embodiment, the loss of improvement suggestions is represented by the following formula:

[0089] ;

[0090] in, To improve the loss value of the suggestion, For experimental feedback; For large language models that require parameter updates; This serves as the initial large language model; This is an adaptive penalty factor.

[0091] In this embodiment, experimental feedback can be written by humans or obtained based on EM matching; for example, experimental feedback can come from experts or researchers scoring the experimental design and improvement suggestions generated by the feedback model; EM matching includes the character-level matching degree between the experimental design and the historical optimization trajectory.

[0092] In this embodiment, the second autoregressive loss is expressed by the following formula:

[0093] ;

[0094] in, This is the second autoregressive loss value. The sentence representing the sample experiment plan, sample improvement suggestions, and historical optimization trajectory in each text related to sample requirements is the first... The probability of each word.

[0095] In this step, the joint loss of the entire scheme generation model can be constructed using the first autoregressive loss value, the improvement suggestion loss, and the second autoregressive loss value. This joint loss is expressed as:

[0096] .

[0097] In this embodiment, the parameters of the initial experimental scheme automatic generation and optimization model are iterated by minimizing the joint loss. The resulting scheme generation model can be used to realize the automatic generation and optimization of experimental schemes.

[0098] In this embodiment, the above-mentioned scheme generation and optimization model is jointly constructed by the three large language models. After obtaining the experimental scheme requirements, the trained scheme generation model is first used to analyze the requirements and output a preliminary experimental scheme report. Then, the trained feedback model is used to generate improvement suggestions based on self-reflection. Finally, the trained optimization model is used to optimize and adjust the experimental scheme and output the optimized experimental scheme report.

[0099] The training method for a scheme generation and optimization model based on a large model provided in this invention obtains a scheme generation model by training sample experimental schemes and sample requirements as training samples and using a first autoregressive loss as the loss function; a feedback model is obtained by training historical optimization trajectories and target schemes as training samples and using improvement suggestion loss as the loss function; and an optimization model is obtained by training sample requirements, historical optimization trajectories, and sample improvement suggestions output by the feedback model during historical training, using a second autoregressive loss as the loss function and third language processing, thereby determining the scheme generation and optimization model. This achieves automated design and optimization of schemes for dedicated equipment or systems, improving the efficiency and accuracy of experimental scheme generation.

[0100] In some embodiments, the scheme generation model uses a first LoRA weight; the first LoRA weight is obtained by autoregressive training of the pre-trained model using sample requirements and sample experimental schemes; the feedback model uses a second LoRA weight; the second LoRA weight is obtained by autoregressive training of the pre-trained model using sample improvement opinions combined with the proximal strategy optimization PPO method; the optimization model uses a third LoRA weight; the third LoRA weight is obtained by autoregressive training of the pre-trained model using improvement opinion data corresponding to sample requirements, sample experimental schemes, and sample preliminary schemes.

[0101] In this embodiment, the scheme generation model is constructed based on the base large language model word segmenter, the base large language model original weights, and the first fine-tuned LoRA weights.

[0102] For example, the experimental scheme requirements are input into the base large language model word segmenter of the scheme generation model to obtain the word vector encoding of the experimental scheme requirements; then the word vector encoding is input into the original weight of the base large language model that integrates the first LoRA weight, and the probability of each word output is predicted step by step until an end symbol is generated or the set maximum length is reached. Finally, the text is converted back to readable text according to the probability to obtain the preliminary experimental scheme report (i.e., the preliminary sample scheme).

[0103] In this embodiment, the feedback model is constructed based on the base large language model word segmenter, the base large language model original weights, and the second LoRA weights.

[0104] In this embodiment, the sample optimization scheme and historical optimization trajectory can be input into the base large language model word segmenter of the feedback model to obtain the word vector encoding of the sample optimization scheme and historical optimization trajectory; then the word vector encoding is input into the original weight of the base large language model that integrates the second LoRA weight, and the probability of each output word is predicted step by step until the end symbol is generated or the set maximum length is reached. Finally, the sample improvement suggestions are obtained by converting the probability back into readable text.

[0105] In this embodiment, the preliminary sample scheme and historical optimization trajectory can also be input into the base large language model word segmenter of the feedback model to obtain the word vector encoding of the preliminary sample scheme and historical optimization trajectory, and then determine the corresponding sample improvement suggestions.

[0106] In this embodiment, the feedback model is constructed based on the base large language model word segmenter, the base large language model original weights, and the third LoRA weights.

[0107] Specifically, the sample requirements, improvement suggestions, and historical optimization trajectory are input into the base large language model word segmenter to obtain the word vector encoding of the experimental scheme requirements, the self-reflection-based improvement suggestions, and the historical optimization trajectory; then the word vector encoding is input into the original weights of the base large language model that incorporate the third LoRA weights, and the probability of each output word is predicted step by step until an end symbol is generated or the set maximum length is reached. Finally, the text is converted back into readable text based on the probability to obtain the sample optimization scheme.

[0108] The training method for scheme generation and optimization models based on large models provided in this invention fine-tunes the models by using different LoRA weights during the training of the scheme generation model, feedback model, and optimization model, and obtains the corresponding language processing model through autoregressive training. This enables automated processing of experimental scheme design, obtaining improvement opinions, and optimizing schemes based on improvement opinions, thereby further improving the efficiency of scheme generation.

[0109] In some embodiments, the feedback model and the optimization model are iteratively trained through cyclic optimization; the training termination conditions for the feedback model and the optimization model include: determining that the sample optimization scheme output by the optimization model meets the target requirements based on the improvement suggestions data output by the feedback model.

[0110] In this embodiment, the objective requirement includes whether the sample optimization scheme covers all the indicators required by the scheme. If all are covered, the sample optimization scheme currently output by the optimization model is determined to be effective; or if the sample optimization scheme covers the specified core scheme requirements, the current sample optimization scheme can also be considered effective.

[0111] Specifically, the optimization process is cyclical, and the termination condition is: when the feedback model determines, based on the improvement suggestions, that the sample optimization scheme output by the optimization model has met all the preset requirements and achieved the expected optimization effect, the cyclic optimization ends, and the current output of the optimization model is recognized as the final experimental scheme report. This cyclic optimization mechanism ensures the continuous improvement of the experimental scheme, gradually approaching the optimal scheme, and can achieve comprehensive optimization of the experimental scheme through self-reflection and gradual adjustment.

[0112] The method for generating and optimizing schemes based on large models provided in this invention improves the reliability of scheme generation by setting up a feedback model and an optimization model for iterative training through cyclic optimization. This allows the optimization model to improve its performance through self-reflection and gradual adjustment.

[0113] The following describes the scheme generation and optimization method based on a large model provided by the present invention. The scheme generation and optimization method based on a large model described below can be referred to in correspondence with the training method of the scheme generation and optimization model based on a large model described above.

[0114] Figure 2 This is one of the flowcharts illustrating the scheme generation and optimization method based on a large model provided by the present invention, such as... Figure 2 As shown, the scheme generation and optimization method based on the large model includes the following steps:

[0115] Step 210: Obtain the target requirements of the experimental design.

[0116] In this step, in the same application field as the above sample experimental scheme, the experimental scheme to be designed can also be an existing design scheme for special equipment or systems for performing rescue, disaster relief or specific mission training. This design scheme can be determined based on expert experience or obtained through natural language processing technology.

[0117] In this embodiment, the target requirements for dedicated wearable devices used for frontline disaster relief include lightweight, pressure resistance, and reliable communication.

[0118] In this embodiment, the target requirements can be pre-written manually or the front-end can be written according to the solution requirements.

[0119] Step 220: Process the target requirements based on the scheme generation and optimization model to obtain the target scheme; wherein, the scheme generation and optimization model is trained by the training method of the scheme generation and optimization model based on the large model.

[0120] In this step, the solution generation and optimization model includes a solution generation model, a feedback model, and an optimization model.

[0121] The scheme generation model is obtained by iteratively training the first large language model using sample experimental schemes and sample requirements as training samples and the first autoregressive loss as the loss function; the first autoregressive loss is determined based on the cross-entropy between the sample experimental schemes and sample requirements.

[0122] The feedback model is obtained by iteratively training second language processing using historical optimization trajectories and target schemes as training samples and improved opinion loss as the loss function. The improved opinion loss is determined based on the feedback value between the sample optimization scheme and the sample experimental scheme. The target scheme includes the sample optimization scheme or the preliminary sample scheme output by the scheme generation model during historical training.

[0123] Among them, the optimization model is obtained by iteratively training the third language processing with sample requirements, historical optimization trajectories, and sample improvement opinions output by the feedback model during historical training as training samples, and using the second autoregressive loss as the loss function; the scheme generation and optimization model is obtained based on the scheme generation model, the feedback model, and the optimization model.

[0124] In this embodiment, the specific training process of the scheme generation model, feedback model and optimization model can be referred to the embodiment corresponding to steps 110 to 140, and will not be repeated in this embodiment.

[0125] In this embodiment, the above-mentioned scheme generation and optimization model is jointly constructed by the three large language models. After obtaining the experimental scheme requirements, the trained scheme generation model is first used to analyze the requirements and output a preliminary experimental scheme report. Then, the trained feedback model is used to generate improvement suggestions based on self-reflection. Finally, the trained optimization model is used to optimize and adjust the experimental scheme and output the optimized experimental scheme report.

[0126] The scheme generation and optimization method based on a large model provided in this invention analyzes and processes the target requirements of the experimental scheme to be designed by constructing a scheme generation and optimization model using a scheme generation model, a feedback model, and an optimization model, and obtains the target scheme. This achieves automated design and optimization of schemes for special equipment or systems, and improves the efficiency and accuracy of experimental scheme generation.

[0127] Figure 3This is the second flowchart illustrating the scheme generation and optimization method based on a large model provided by the present invention. Figure 3 In the illustrated embodiment, a scheme generation and optimization model includes an automatic experimental scheme generation module and an automatic experimental scheme optimization module. The automatic experimental scheme generation module includes a scheme generation and optimization model (corresponding to the scheme generation model), and the automatic experimental scheme optimization module includes a planner model (corresponding to the feedback model) and an actor model (corresponding to the optimization model). Each model is trained based on a large language model. During model training, an experimental requirement document (including multiple experimental requirements) is input into the scheme generation and optimization model to generate a preliminary experimental scheme document. Then, the experimental requirement document, the preliminary experimental scheme document, and the optimized experimental scheme document are all input into the planner model. If it is determined that the preliminary experimental scheme document needs further improvement, the self-reflection improvement suggestions output by the planner model are input into the actor model to obtain a new optimized experimental scheme document. The new optimized experimental scheme document is then input into the planner model, and the current experimental scheme document is determined based on the sample requirements and the preliminary experimental scheme document. If no improvement is needed, the final experimental scheme document is input.

[0128] The training apparatus for scheme generation and optimization model based on a large model provided by the present invention will be described below. The training apparatus for scheme generation and optimization model based on a large model described below can be referred to in correspondence with the training method for scheme design model described above.

[0129] Figure 4 This is a schematic diagram of the training device for scheme generation and optimization models based on large models provided by the present invention, as shown below. Figure 4 As shown, the training device for the scheme generation and optimization model based on the large model includes: a first acquisition module 410, a first training module 420, a second training module 430, and a third training module 440.

[0130] The first acquisition module 410 is used to acquire the sample experimental plan and its sample requirements;

[0131] The first training module 420 is used to iteratively train the first large language model using sample experimental schemes and sample requirements as training samples and the first autoregressive loss as the loss function to obtain the scheme generation model; wherein, the first autoregressive loss is determined based on the cross-entropy between sample experimental schemes and sample requirements.

[0132] The second training module 430 is used to iteratively train the second language processing model using historical optimization trajectories and target schemes as training samples and improved opinion loss as the loss function to obtain a feedback model. The improved opinion loss is determined based on the feedback value between the sample optimization scheme and the sample experimental scheme. The target scheme includes the sample optimization scheme or the preliminary sample scheme output by the scheme generation model during the historical training process.

[0133] The third training module 440 is used to iteratively train the third language processing module using sample requirements, historical optimization trajectories, and sample improvement suggestions output by the feedback model during historical training as training samples, and using the second autoregressive loss as the loss function to obtain an optimized model; and to obtain a scheme generation and optimization model based on the scheme generation model, the feedback model, and the optimization model; wherein, the second autoregressive loss is determined based on the cross-entropy between the sample optimization scheme and the sample experimental scheme.

[0134] The training device for a scheme generation and optimization model based on a large model provided in this invention obtains a scheme generation model by training sample experimental schemes and sample requirements as training samples and using a first autoregressive loss as the loss function; a feedback model is obtained by training historical optimization trajectories and target schemes as training samples and using improvement suggestion loss as the loss function; and an optimization model is obtained by training sample requirements, historical optimization trajectories, and sample improvement suggestions output by the feedback model during historical training with a second autoregressive loss as the loss function and using third language processing. This determines the scheme generation and optimization model, achieving automated design and optimization of schemes for dedicated equipment or systems, and improving the efficiency and accuracy of experimental scheme generation.

[0135] Figure 5 This is a schematic diagram of the structure of the scheme generation and optimization device based on a large model provided by the present invention, as shown below. Figure 5 As shown, the scheme generation and optimization device based on the large model includes: a second acquisition module 510 and a scheme generation module 520.

[0136] The second acquisition module 510 is used to acquire the target requirements of the experimental scheme to be designed.

[0137] The solution generation module 520 is used to process the target requirements based on the solution generation and optimization model to obtain the target solution; wherein the solution generation and optimization model is trained by the training method of the solution generation and optimization model based on the large model.

[0138] The large-model-based scheme generation and optimization device provided in this invention analyzes and processes the target requirements of the experimental scheme to be designed by constructing a scheme generation and optimization model using a scheme generation model, a feedback model, and an optimization model, and obtains the target scheme. This realizes the automated design and optimization of schemes for special equipment or systems, and improves the efficiency and accuracy of experimental scheme generation.

[0139] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 6 As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communications bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other through the communications bus 640. The processor 610 can call logic instructions in the memory 630 to execute a training method for a scheme generation and optimization model based on a large model. This method includes: acquiring sample experimental schemes and their sample requirements; using the sample experimental schemes and sample requirements as training samples, and using a first autoregressive loss as the loss function, iteratively training a first large language model to obtain a scheme generation model; wherein the first autoregressive loss is determined based on the cross-entropy between the sample experimental schemes and sample requirements; using historical optimization trajectories and target schemes as training samples, and using improvement suggestion loss as the loss function, iteratively training a second language processing model to obtain a feedback model; wherein the improvement suggestion loss is determined based on the feedback value between the sample optimized scheme and the sample experimental scheme; the target scheme includes the sample optimized scheme or the sample preliminary scheme output by the scheme generation model during historical training; using sample requirements, historical optimization trajectories, and sample improvement suggestions output by the feedback model during historical training as training samples, and using a second autoregressive loss as the loss function, iteratively training a third language processing model to obtain an optimization model; and obtaining a scheme generation and optimization model based on the scheme generation model, the feedback model, and the optimization model; wherein the second autoregressive loss is determined based on the cross-entropy between the sample optimized scheme and the sample experimental scheme.

[0140] Alternatively, a scheme generation and optimization method based on a large model can be implemented, which includes: obtaining the target requirements of the experimental scheme to be designed; processing the target requirements based on the scheme generation and optimization model to obtain the target scheme; wherein the scheme generation and optimization model is trained by a training method based on a large model for scheme generation and optimization.

[0141] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0142] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the training method for a scheme generation and optimization model based on a large model provided by the above methods. This method includes: obtaining sample experimental schemes and their sample requirements; using the sample experimental schemes and sample requirements as training samples, and using a first autoregressive loss as a loss function to iteratively train a first large language model to obtain a scheme generation model; wherein, the first autoregressive loss is determined based on the cross-entropy between the sample experimental schemes and the sample requirements; using historical optimization trajectories and target schemes as training samples, and using... The suggestion improvement loss is used as the loss function to iteratively train the second language processing to obtain the feedback model. The suggestion improvement loss is determined based on the feedback value between the sample optimization scheme and the sample experimental scheme. The target scheme includes the sample optimization scheme or the preliminary sample scheme output by the scheme generation model during historical training. Using sample requirements, historical optimization trajectories, and the suggestion improvement output by the feedback model during historical training as training samples, the third language processing is iteratively trained using the second autoregressive loss as the loss function to obtain the optimization model. Based on the scheme generation model, the feedback model, and the optimization model, a scheme generation and optimization model is obtained. The second autoregressive loss is determined based on the cross-entropy between the sample optimization scheme and the sample experimental scheme.

[0143] Alternatively, a scheme generation and optimization method based on a large model can be implemented, which includes: obtaining the target requirements of the experimental scheme to be designed; processing the target requirements based on the scheme generation and optimization model to obtain the target scheme; wherein the scheme generation and optimization model is trained by a training method based on a large model for scheme generation and optimization.

[0144] Furthermore, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a training method for a scheme generation and optimization model based on a large model, as provided by the methods described above. This method includes: acquiring sample experimental schemes and their sample requirements; using the sample experimental schemes and sample requirements as training samples, iteratively training a first large language model using a first autoregressive loss as a loss function to obtain a scheme generation model; wherein the first autoregressive loss is determined based on the cross-entropy between the sample experimental schemes and the sample requirements; using historical optimization trajectories and target schemes as training samples, iteratively training a second large language model using an improved suggestion loss as a loss function. The language processing is iteratively trained to obtain a feedback model; the improvement suggestion loss is determined based on the feedback value between the sample optimization scheme and the sample experimental scheme; the target scheme includes the sample optimization scheme or the sample preliminary scheme output by the scheme generation model during historical training; using sample requirements, historical optimization trajectory, and sample improvement suggestions output by the feedback model during historical training as training samples, the third language processing is iteratively trained with the second autoregressive loss as the loss function to obtain an optimization model; based on the scheme generation model, the feedback model, and the optimization model, a scheme generation and optimization model is obtained; the second autoregressive loss is determined based on the cross-entropy between the sample optimization scheme and the sample experimental scheme.

[0145] Alternatively, a scheme generation and optimization method based on a large model can be implemented, which includes: obtaining the target requirements of the experimental scheme to be designed; processing the target requirements based on the scheme generation and optimization model to obtain the target scheme; wherein the scheme generation and optimization model is trained by a training method based on a large model for scheme generation and optimization.

[0146] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0147] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for training a scheme generation and optimization model based on a large model, characterized in that, include: Obtain the experimental design and sample requirements; Using the sample experiment scheme and the sample requirements as training samples, the first large language model is iteratively trained with the first autoregressive loss as the loss function to obtain the scheme generation model; wherein, the first autoregressive loss is determined based on the cross-entropy between the sample experiment scheme and the sample requirements. Using historical optimization trajectories and target schemes as training samples, and using improved opinion loss as the loss function, the second language processing is iteratively trained to obtain a feedback model; wherein, the improved opinion loss is determined based on the feedback value between the sample optimization scheme and the sample experimental scheme; the target scheme includes the sample optimization scheme or the sample preliminary scheme output by the scheme generation model during the historical training process; Using the sample requirements, the historical optimization trajectory, and the sample improvement suggestions output by the feedback model during historical training as training samples, the third language processing is iteratively trained using the second autoregressive loss as the loss function to obtain an optimization model; a scheme generation and optimization model is obtained based on the scheme generation model, the feedback model, and the optimization model; wherein, the second autoregressive loss is determined based on the cross-entropy between the sample optimization scheme and the sample experimental scheme; The first autoregressive loss is expressed by the following formula: Among them, L reg_1 Let y be the first autoregressive loss value. j This represents the probability of the j-th word in a sentence of each text in the sample experiment scheme designed to meet the sample requirements. The loss from the proposed improvements is expressed by the following formula: Among them, L PPO To improve the opinion loss value, r θ (x,y r (This is for experimental feedback;) For large language models that require parameter updates; This represents the initial large language model; β is the adaptive penalty factor. The second autoregressive loss is expressed by the following formula: Among them, L reg_2 is the second autoregressive loss value, where w represents the probability of the k-th word in the sentence of each text in the sample experiment plan, sample improvement suggestions, and historical optimization trajectory, based on the sample requirements.

2. The training method for scheme generation and optimization models based on large models according to claim 1, characterized in that, The feedback model and the optimization model are iteratively trained through cyclic optimization. The training termination conditions for the feedback model and the optimization model include: determining, based on the improvement suggestions output by the feedback model, that the sample optimization scheme output by the optimization model meets the target requirements.

3. The training method for scheme generation and optimization models based on large models according to claim 1, characterized in that, The scheme generation model, the feedback model, and the optimization model are all trained through pre-trained language processing, and all employ an autoregressive mechanism to encode and decode sample requirements. Different LoRA weights are used to fine-tune different pre-trained models.

4. The training method for scheme generation and optimization models based on large models according to any one of claims 3, characterized in that, The scheme generates a model using a first LoRA weight; the first LoRA weight is obtained by autoregressive training of the pre-trained model based on sample requirements and sample experiment scheme. The feedback model employs a second LoRA weight; the second LoRA weight is obtained by autoregressive training of the pre-trained model using a PPO method optimized by combining sample improvement opinions with a proximal strategy. The optimized model uses a third LoRA weight; the third LoRA weight is obtained by autoregressive training of the pre-trained model based on the improvement suggestions data corresponding to the sample requirements, sample experimental schemes, and sample preliminary schemes.

5. A method for generating and optimizing solutions based on a large model, characterized in that, include: Obtain the target requirements of the experimental design to be developed; The target requirements are processed based on the scheme generation and optimization model to obtain the target scheme; wherein the scheme generation and optimization model is trained by the training method of the scheme generation and optimization model based on the large model as described in any one of claims 1-4.

6. A training apparatus for a scheme generation and optimization model based on a large model, employing the training method for a scheme generation and optimization model based on a large model as described in claim 1, characterized in that, include: The first acquisition module is used to acquire the sample experimental plan and its sample requirements; The first training module is used to iteratively train the first large language model using the sample experimental scheme and the sample requirements as training samples and the first autoregressive loss as the loss function to obtain the scheme generation model; wherein, the first autoregressive loss is determined based on the cross-entropy between the sample experimental scheme and the sample requirements. The second training module is used to iteratively train the second language processing model using historical optimization trajectories and target schemes as training samples and improved opinion loss as the loss function to obtain a feedback model; wherein, the improved opinion loss is determined based on the feedback value between the sample optimization scheme and the sample experimental scheme; the target scheme includes the sample optimization scheme or the preliminary sample scheme output by the scheme generation model during the historical training process; The third training module is used to iteratively train the third language processing module using the sample requirements, the historical optimization trajectory, and the sample improvement suggestions output by the feedback model during the historical training process as training samples, and using the second autoregressive loss as the loss function to obtain an optimization model; and to obtain a scheme generation and optimization model based on the scheme generation model, the feedback model, and the optimization model; wherein, the second autoregressive loss is determined based on the cross-entropy between the sample optimization scheme and the sample experimental scheme.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Deep Learning Training Method for Computing Device and Apparatus

    US20230206069A1

  • Artificial intelligence-based sample evaluation method, apparatus, device, and storage medium

    WO2021121128A1