Self-adaptive system capability evolution method based on human feedback learning

By designing adaptive conditions, constructing human preference data pairs, and optimizing the loss function in the urban drone detection system, and utilizing human feedback learning to optimize the parameters of the large model, the problem of limited adaptive capability was solved, and the system's autonomous evolution and improved environmental perception accuracy were achieved.

CN121436764APending Publication Date: 2026-01-30THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511545703.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

The adaptive capabilities of existing urban drone detection systems are limited by the diversity of predetermined rules and human experience, making it difficult to adapt to unexpected changes in mission environments and unknown threats, and lacking autonomous evolution capabilities.

Method used

By designing adaptive conditions, establishing a comprehensive evaluation index system, constructing human preference data pairs, and designing a human preference optimization loss function, the large model parameters of the adaptive system are optimized using human feedback learning, thus skipping the training process of traditional reward models and reinforcement learning.

Benefits of technology

The system achieves autonomous evolution, improves the accuracy of environmental change perception and the superiority or inferiority of strategies, reduces training complexity, and enhances the system's adaptive capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121436764A_ABST
    Figure CN121436764A_ABST
Patent Text Reader

Abstract

The invention provides an adaptive system capability evolution method based on human feedback learning, and the method comprises the following steps: 1, designing adaptive conditions, employing different strategies to carry out the evolution of a large model in an adaptive system according to different adaptive conditions, and outputting an execution effect; performing comprehensive evaluation on the execution effect; 3, constructing a human preference data pair according to the comprehensive evaluation score; 4, designing a human preference optimization loss function; and step 5, training by using a human preference optimization method, and performing optimization fine tuning on the large model parameters of the adaptive system according to the human preference data pair and the human preference optimization loss function to realize adaptive system capability evolution based on human feedback learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an adaptive system capability evolution method, and more particularly to an adaptive system capability evolution method based on human feedback learning. Background Technology

[0002] This section provides only background information relevant to this disclosure and is not necessarily prior art.

[0003] Currently, the adaptive adjustment of urban drone detection systems is mainly based on fixed preset rules and relies on the experience and knowledge of system operation and maintenance personnel. It can adapt to a certain extent to predictable external mission environments, resource damage, and changes in its own operating status. However, its adaptive capability is limited by the diversity and completeness of the preset rules and human experience and knowledge. It lacks the ability to evolve autonomously during the system's adaptive adjustment process, and cannot achieve continuous growth in the system's adaptive capability. Moreover, the system lacks the generalization of adaptive adjustment and is difficult to adapt to unexpected changes in mission environments and unknown threat environments.

[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] Purpose of the invention: The technical problem to be solved by the present invention is to provide an adaptive system capability evolution method based on human feedback learning, which addresses the shortcomings of the existing technology.

[0006] To address the aforementioned technical problems, this invention discloses an adaptive system capability evolution method based on human feedback learning, comprising the following steps:

[0007] Step 1: Design adaptive conditions. The large model in the adaptive system evolves according to different adaptive conditions using different strategies and outputs the execution results.

[0008] Step 2: Establish a comprehensive evaluation index system to comprehensively evaluate the implementation effect;

[0009] Step 3: Construct human preference data pairs based on the comprehensive evaluation scores;

[0010] Step 4: Design a loss function to optimize human preferences;

[0011] Step 5: Train the system using a human preference optimization method. Based on human preference data pairs and the human preference optimization loss function, fine-tune the large model parameters of the adaptive system to achieve the capability evolution of the adaptive system based on human feedback learning.

[0012] Furthermore, the adaptive conditions described in step 1 include:

[0013] The four categories are: resource damage, resource inefficiency, electronic interference, and sudden targets.

[0014] Among them, resource damage-related adaptive conditions include: damage to sensing resources, damage to decision-making resources, damage to platform resources, and damage to communication resources;

[0015] Resource-inefficient adaptive conditions include: resource overload attacks and resource denial-of-service attacks;

[0016] Adaptive conditions to electronic interference include: radar interference, communication interference, and navigation interference;

[0017] Adaptive conditions for sudden target types include: target type, quantity, orientation, altitude, and speed.

[0018] Furthermore, the comprehensive evaluation index system described in step 2 includes the following four indicators:

[0019] Resource cost is used to measure the type and quantity of resources required by the generation strategy.

[0020] Strategy execution time is used to measure the time interval required from the start of resource acceptance strategy execution to the completion of strategy execution;

[0021] The gain-loss ratio measures the proportion of gains gained to resources lost after a strategy is executed.

[0022] Capability recovery is used to measure the degree to which an adaptive system recovers its capabilities after it autonomously adjusts its strategy.

[0023] Furthermore, the comprehensive evaluation of the execution effect described in step 2 includes:

[0024] Step 2-1: Construct a resource cost calculation model and calculate the resource cost. The details are as follows:

[0025]

[0026] in, Assign the first in the strategy The number of class resources, Assign the first in the strategy The cost of such resources, For the number of resource categories;

[0027] Step 2-2: Construct a strategy execution time calculation model and calculate the strategy execution time. The details are as follows:

[0028]

[0029] in, For resources The time interval from receiving the policy to the completion of policy execution. For resources Complete the strategy execution time. For resources The execution strategy completes at the specified time.

[0030] Steps 2-3: Construct a revenue-loss ratio calculation model, and calculate the revenue-loss ratio by obtaining the ratio of revenue value to the number of damaged resources. The details are as follows:

[0031]

[0032] in, For profit value, The quantity of damaged resources;

[0033] Steps 2-4: Construct a capacity recovery rate calculation model and calculate the capacity recovery rate. The details are as follows:

[0034]

[0035] in, The initial system capability value, The restored system capability value;

[0036] Steps 2-5: Construct a comprehensive evaluation model for execution effectiveness and calculate the comprehensive evaluation score. The details are as follows:

[0037]

[0038] in, For the first large model Each output result , , and These represent the maximum values ​​of resource cost, strategy execution time, profit-loss ratio, and capability recovery rate for different strategies. and The weighting coefficient for the indicator;

[0039] Steps 2-6 use the Sigmoid function mapping method to map the comprehensive evaluation score to the interval [0,100].

[0040] Furthermore, the mapping method for the Sigmoid function described in steps 2-6 is as follows:

[0041]

[0042] in, This is the overall evaluation score after mapping.

[0043] Furthermore, the construction of human preference data pairs described in step 3 is as follows:

[0044] Step 3-1, design the human preference data pairs in the following form:

[0045]

[0046] in, These are prompts for the large model; The results are either the preferred responses from a large model or adjusted responses based on human feedback. The results of rejection responses in a large model;

[0047] Step 3-2: Based on the comprehensive evaluation score of the execution effect output by the large model, determine the preferred answer results. and refusing to answer results Outputs with a comprehensive evaluation score higher than the threshold are considered preferred answers, while outputs with a evaluation score lower than the threshold are considered rejected answers.

[0048] Step 3-3 involves manually fine-tuning the performance evaluation scores to obtain the final preferred and rejected responses.

[0049] Furthermore, step 4 involves designing a human preference optimization loss function to maximize the preference response result. and refusing to answer results The objective is the difference between the logarithmic probability values.

[0050] Furthermore, the human preference optimization loss function described in step 4 is expressed as follows:

[0051]

[0052] in, Optimize the loss function to suit human preferences; This represents the large model, i.e., the policy model, in an adaptive system. For the given prompt words Generate preference response results The probability of; Representational Strategy Model For the given prompt words Generate a rejection response. The probability of; Represents the logarithmic difference, where and For length standardization rewards, Scaling hyperparameters to reward differences The target reward margin hyperparameter.

[0053] Furthermore, the training using human preference optimization methods described in step 5 includes:

[0054] Preparation phase:

[0055] Step 5-1: Obtain the pre-trained model and select a basic large model, LLM;

[0056] Step 5-2, supervised fine-tuning of SFT: fine-tuning the basic large model LLM using pre-set question and answer data to obtain the fine-tuned pre-trained SFT model.

[0057] Step 5-3: Prepare preference data. Collect and record all human preference data using a large-sample parallel simulation experiment method. Based on this, the training dataset D for the human preference algorithm is constructed;

[0058] Training phase:

[0059] Step 5-4: Initialize the model and load the SFT model as a policy model. That is, the model that needs to be trained and optimized;

[0060] Step 5-5: Extract a portion of batch human preference data pairs from the human preference dataset D. Create a training dataset;

[0061] Steps 5-6: Forward propagation; for each data item in the batch, calculate the strategy model. Generate preference response results and refusing to answer results log probability and ;

[0062] Steps 5-7: Optimize the loss function based on human preferences, calculate the average damage value of the batch data, perform backpropagation, and update the policy model. Parameters;

[0063] Steps 5-8, repeat steps 5-5 to 5-7 until the preset conditions are met.

[0064] Furthermore, the preset condition mentioned in steps 5-8 is that the policy model converges or reaches a preset number of training steps.

[0065] Beneficial effects:

[0066] 1. This invention can continuously optimize and fine-tune the output of large models, improve the accuracy of environmental change perception and the quality of adaptation strategies given by the large model of the system, and provide new system adaptation strategies to realize the autonomous evolution of the system's adaptive capabilities.

[0067] 2. In this invention, human preference optimization training is highly efficient and the training convergence process is controllable, reducing the complexity of model tuning and training. Compared with traditional human feedback learning methods, this invention directly uses human preference data to optimize and fine-tune the large system model, skipping the design and training of reward models and policy reinforcement learning models. The training process is similar to conventional fine-tuning.

[0068] 3. This invention avoids the reliance on the calculation of preference probabilities of the reference model in existing large model fine-tuning algorithms (DPO, OPO, etc.). This method directly uses the difference in log probabilities of human preference data to optimize the target loss function, making learning and training simpler, more efficient, and more effective. Attached Figure Description

[0069] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.

[0070] Figure 1 This is a schematic diagram of the overall process of the present invention.

[0071] Figure 2 This is a schematic diagram of the system's adaptability design.

[0072] Figure 3 This is a schematic diagram illustrating the comprehensive evaluation of the execution effect of the large model output.

[0073] Figure 4 This is a schematic diagram of human preference data. Detailed Implementation

[0074] This invention proposes an adaptive system capability evolution method based on human feedback learning for adaptive systems such as urban UAV detection systems. The core idea is to construct the preferred and rejected answers of a large model through human feedback, and use human preference data to directly optimize the large model. This fine-tunes the parameters of the large model of the adaptive system, skipping the complex processes of reward model training and reinforcement learning training in traditional human feedback learning. This improves the accuracy, timeliness, and strategy superiority of the environmental change perception of the urban UAV detection system, thereby iteratively optimizing the system's adaptive capability to adapt to external disturbances and internal faults.

[0075] This invention provides an adaptive system capability evolution method based on human feedback learning for urban unmanned aerial vehicle (UAV) detection systems, comprising the following five steps: Figure 1 As shown.

[0076] Step 1: Design System Adaptability Conditions. A large-sample parallel simulation experiment method is used to design four types of adaptability conditions for the system: resource damage, resource degradation, electronic interference, and sudden targets. Resource damage adaptability conditions include: perception resource damage, decision-making resource damage, platform resource damage, and communication resource damage; resource degradation includes: resource overload attacks and resource denial-of-service attacks; electronic interference adaptability conditions include: radar interference, communication interference, and navigation interference; and sudden target adaptability conditions include: target type, quantity, azimuth, altitude, and speed.

[0077] Step 2: Establish a comprehensive evaluation index system and model for the execution effect of the large model output. The execution effect of the large model output is comprehensively evaluated from four aspects: resource cost, strategy execution time, benefit-loss ratio, and capability recovery rate.

[0078] Among them, resource cost is used to measure the types and quantities of resources required in generating the autonomous adjustment strategy; strategy execution time is used to measure the time interval from the start of resource acceptance strategy execution to the completion of strategy execution; benefit-loss ratio is used to measure the ratio of the benefit value obtained after strategy execution to the number of damaged resources; and capability recovery rate is used to measure the degree of capability recovery of the adaptive system after autonomous adjustment strategy.

[0079] By combining the above four categories of indicators (resource cost, strategy execution time, benefit-loss ratio, and capability recovery rate), a weighted summation method is used to comprehensively evaluate the performance of the large model output.

[0080] Step 3: Construct human preference data pairs. Human preference data pairs are used to directly optimize large models, fine-tuning the parameters of large models in adaptive systems, skipping the complex processes of reward model training and reinforcement learning training in traditional human feedback learning. The constructed human preference data pairs consist of... It consists of three parts:

[0081] · These are prompt words;

[0082] · The results could be based on the model's preferred responses or adjusted responses based on human feedback.

[0083] · This indicates a refusal to answer the question.

[0084] Preference Answer Refuse to answer The determination of the preferred answer is based on the evaluation score of the execution effect output by the large model. A higher evaluation score is considered a preferred answer, while a lower score is considered a rejected answer. Simultaneously, the evaluation score can be manually fine-tuned to adjust the preferred and rejected answers.

[0085] Step 4: Design the human preference optimization loss function. The human preference optimization loss function designed in this invention... As shown in the following formula, its goal is to maximize the preferred answer. The logarithmic probability value and the refusal to answer The difference between the logarithmic probability values.

[0086]

[0087] In the formula:

[0088] : Representation of the strategy model For a given prompt word x, generate a preferred answer. The probability of;

[0089] : Representation of the strategy model Given a prompt word x, generate a rejection response. The probability of;

[0090] : Represents the logarithmic probability difference, measuring Generate preferred answers Instead How much has the probability increased? The larger the expected value, the better. If the difference is greater than 1 and the logarithm is greater than 0, it indicates... More inclined to generate ;

[0091] The length-normalized reward is used to control the length of the output sequence and avoid generating longer but lower-quality output sequences.

[0092] The reward difference scaling hyperparameter controls the degree of reward difference scaling, and is usually set between 2 and 2.5, with the expectation that this log probability value will be larger.

[0093] The target reward margin hyperparameter is used to control the difference between the probability of a preferred answer and the probability of a rejection in the output of a large model, preventing excessive deviation. The value should be between 0.5 and 1.5, increasing it... Values ​​can improve the model's generalization ability.

[0094] Step 5: Learn and train human preference optimization algorithms.

[0095] (1) Initialize the model: Load the pre-trained large model into the policy model. That is, the model that needs to be trained and optimized;

[0096] (2) Training loop. A portion of the batch dataset < is extracted from the human preference dataset D. >;

[0097] (3) Forward propagation. For each data point in the batch, the strategy model is calculated. generate Log probability: , value;

[0098] (4) Calculate the target damage value based on the above target loss function formula, combined with Hyperparameters are used to calculate the average damage value of a batch of data.

[0099] (5) Backpropagation: Calculate the loss function Compared to the strategy model Find the gradient of the parameters and their derivatives.

[0100] (6) Parameter update. The optimizer updates the policy model based on the gradient. Parameters;

[0101] (7) Repeat the training cycle. Repeat the above process until the model converges or reaches the preset number of training steps.

[0102] Example:

[0103] This invention proposes an adaptive system capability evolution method based on human feedback learning. The specific implementation process, using a UAV detection system as an example, includes five steps: designing system adaptation conditions, establishing a comprehensive evaluation index system and model for large-scale model output execution performance, constructing human preference data pairs, designing a human preference optimization loss function, and learning and training a human preference optimization algorithm. The overall process is as follows: Figure 1 As shown.

[0104] 1. Design system adaptability conditions

[0105] System adaptability simulation is mainly used to simulate the excitation information that triggers the dynamic adjustment of the adaptive system. By analyzing the environmental changes and its own operating state changes during the operation of the adaptive system, a large-sample parallel simulation test method is used to design four types of adaptability conditions faced by the system: resource damage, resource degradation, electronic interference, and sudden targets. Among them, resource damage adaptability conditions include: perception resource damage, decision resource damage, platform resource damage, and communication resource damage; resource degradation adaptability conditions include: resource overload attack and resource denial-of-service attack; electronic interference adaptability conditions include: radar interference, communication interference, and navigation interference; sudden target adaptability conditions include: target type, quantity, azimuth, altitude, and speed.

[0106] To achieve intelligent generation of diverse, boundary-limit-oriented system adaptation condition sets, this paper employs experimental design theory, defining the adaptation conditions that trigger adaptive adjustments in the system as experimental factors, and the specific values ​​of these adaptation conditions as experimental factor levels. A system adaptation condition design method is proposed, such as... Figure 2 As shown, it includes 3 sub-steps:

[0107] (1) Design of experimental factors

[0108] This invention designs four types of test factors: resource damage, resource inefficiency, electronic interference, and sudden target.

[0109] 1) Resource damage factors: Experimental factors are designed from the perspective of resource category and damage scale, including: number of perceived resource damages, perceived resource damage ratio, number of decision-making resource damages, decision-making resource damage ratio, number of platform resource damages, platform resource damage ratio, number of communication resource damages, and communication resource damage ratio.

[0110] 2) Resource Degradation Factors: Experimental factors are designed from the perspective of resource function service degradation that cannot meet the normal user service requests, including: traffic replay intensity, CPU overload intensity, memory overload intensity, IO overload length and overload duration.

[0111] 3) Electronic interference factors: Test factors are designed from the perspective of interference patterns and interference operating parameters, including: radar interference frequency band, radar interference intensity, communication interference frequency band, communication interference intensity, navigation interference frequency band, and navigation interference intensity.

[0112] 4) Sudden Target Factors: Design experimental factors from the perspective of sudden target scale and region, including: target type, target quantity, target orientation, target speed, target altitude, and target latitude and longitude.

[0113] (2) Design of experimental factor levels

[0114] Designing experimental factor levels refers to recommending the number and value of levels for the screened experimental factors, i.e., setting the value space for the experimental factors. Based on the continuity of the experimental factor values, they can be divided into two categories: discrete and continuous.

[0115] For discrete experimental factor levels, an enumeration method is used to design the experimental factor level space. For example, the resource damage quantity can be 2, 4, 6, 8, 10, or 15, and the resource damage rate can be 5%, 10%, 15%, 20%, or 30%.

[0116] For continuous test factor levels, including interference frequency band, interference intensity, flow playback intensity, azimuth, and velocity, the test factor values ​​are used as benchmark values ​​to determine the spatial range of test factor values, and the test factor level values ​​and number of levels are designed. Two methods are specifically used for design: uniform value selection and random value selection.

[0117] • Uniform value selection can generate a list of experimental factor values ​​that conforms to an arithmetic sequence pattern based on the required number of experimental factor levels, between the minimum and maximum values;

[0118] • Random value selection: The experimental factor level is randomly selected between the minimum and maximum values ​​according to a set distribution.

[0119] The adaptive conditions of the designed adaptive system are shown in the table below.

[0120] Table 1 System Adaptability Design

[0121]

[0122] (3) Generation of system adaptation condition set

[0123] The generation of the system adaptation condition set is mainly based on the designed experimental factors, the number of experimental factor levels, and the interaction relationships of experimental factors. Orthogonal design, analytical design, and Latin square experimental design methods are selected to complete the experimental factor header, experimental factor level combination scheme, determination of the number of experiments, etc., and generate the experimental plan table.

[0124] Table 2 System Adaptation Condition Set

[0125] 2. Establish a comprehensive evaluation index system and model for the output execution effect of the large model.

[0126] To objectively evaluate the output results of a large-scale model in an urban UAV detection system, a human feedback model is employed. Bonus points are awarded for the model's environmental change perception and dynamic reconstruction strategy execution results. To improve the rationality and accuracy of the evaluation of the large-scale model's output results and reduce human feedback bias, a combination of parallel simulation experiments and human feedback is used to comprehensively evaluate the performance of the large-scale model's output results. Figure 3 As shown.

[0127] This invention comprehensively evaluates the execution performance of large model outputs from four aspects: resource cost, strategy execution time, benefit-loss ratio, and capability recovery rate.

[0128] (1) Resource cost calculation model

[0129] Resource cost is used to measure the types and quantities of resources required in a generation strategy, and its calculation model is... As shown in the following formula:

[0130]

[0131] In the formula, Allocate the quantity of the i-th type of resource in the strategy. The cost of allocating resource type i in the strategy. Resource types allocated in urban drone detection systems include: drone catchers, radar jammers, communication jammers, navigation decoys, etc.

[0132] (2) Strategy execution time calculation model

[0133] Policy execution time measures the time interval required from the start of resource acceptance policy execution to the completion of policy execution. Its calculation model... As shown in the following formula:

[0134]

[0135] In the formula, For resources The time interval from receiving the policy to the completion of policy execution. For resources Complete the strategy execution time. For resources The execution of the strategy is completed at the specified time.

[0136] (3) Profit-loss ratio

[0137] The gain-loss ratio is used to measure the proportion of gains gained to resources lost after a strategy is executed. It is calculated by the ratio of the gains gained to the amount of resources lost, and its calculation model is shown in the following formula.

[0138]

[0139] In the formula, For profit value, The quantity of damaged resources;

[0140] (4) Calculation model for capacity recovery rate

[0141] Capability recovery rate is used to measure the degree of capability recovery of an adaptive system after it adopts an autonomous adjustment strategy. The calculation model is as follows:

[0142]

[0143] In the formula, The initial system capability value, This represents the system's restored capability value.

[0144] (5) Comprehensive evaluation model of implementation effect

[0145] By combining the above four categories of indicators (resource cost, strategy execution time, benefit-loss ratio, and capability recovery rate), a weighted summation method is used to comprehensively evaluate the performance of the large model output. Its calculation model is shown in the following formula:

[0146]

[0147] In the formula, This is the i-th output result of the large model. , , , These are the maximum resource cost, longest strategy execution time, maximum benefit-loss ratio, and maximum capability recovery rate among multiple strategies, aiming to normalize these four indicators. The weighting coefficients are the weighting coefficients for the four indicators.

[0148] Considering that the overall evaluation value of the large model output is [0,1], to facilitate querying the evaluation of the merits of the adjustment strategy, it is mapped to the interval [0,100]. This mapping is achieved using the Sigmoid function, as shown in the following equation.

[0149]

[0150] 3. Construct human preference data pairs, such as Figure 4 As shown, the details are as follows:

[0151] Human preference data pairs are used to directly optimize large models and fine-tune the parameters of large models in adaptive systems. They consist of three parts: prompt words, preferred answers, and rejected answers. >, is a prompt is the preferred answer result of the large model, or it can also be the answer after human feedback adjustment; is the result of refusing to answer.

[0152] The present invention is based on the evaluation score of the execution effect output by the large model. The output with a high evaluation score is used as the preferred answer, and the output with a low evaluation score is used as the refusal answer. At the same time, the execution effect evaluation score can be manually fine-tuned to adjust the preferred answer and the refusal answer.

[0153] Taking the urban drone detection system as an example, when facing sudden targets, the large model in the command system will give various adaptable strategy outputs such as "radar jamming, deception jamming, communication jamming, navigation deception", etc. Conduct large-sample parallel simulation experiments on each adaptable strategy, collect experimental records, and comprehensively evaluate the execution effect scores of each strategy, including task completion rate, resource cost, strategy execution time, benefit loss ratio, ability recovery rate and other index evaluation results. By comparing the evaluation results, it can be obtained that:

[0154] · The execution effect score of the unmanned capture strategy is the lowest;

[0155] · The execution effect score of the combined strategy of communication jamming and navigation jamming is the highest.

[0156] Then, the constructed pair of human preference data is:

[0157] 《<prompt x: There are 4 sudden targets and the penetration responsibility area>, <preferred answer yw: Combined strategy of communication jamming and navigation jamming>, <refusal answer yl: Unmanned capture strategy>》

[0158] 4. Design the human preference optimization loss function

[0159] The core of the human preference optimization algorithm is to directly train and optimize the large model using the pair of preference data < >, which skips the training of the reward model RM and the reinforcement learning RL model in human feedback reinforcement learning.

[0160] Suppose there is an initial SFT large model, and it is hoped to obtain an optimized large model through the training of the human preference data set , and this new model can better reflect human preferences. The goal of human preference optimization is: adjust such that for each piece of preference data < >, the probability of generating is greater than the probability of generating , should increase the probability of generating and at the same time reduce the probability of generating , while reducing the probability of generating probability Human preferences ( > The strength of ) can be determined by right , It is expressed as the difference between the logarithmic probability estimates, i.e. .

[0161] In summary, the human preference optimization loss function designed in this invention... As shown in the following formula, its goal is to maximize the preferred answer. The logarithmic probability value and the refusal to answer The difference between the logarithmic probability values.

[0162]

[0163] In the formula:

[0164] : Representation of the strategy model For a given prompt word x, generate a preferred answer. The probability of;

[0165] : Representation of the strategy model Given a prompt word x, generate a rejection response. The probability of;

[0166] : Represents the logarithmic probability difference, measuring Generate preferred answers Instead How much has the probability increased? The larger the expected value, the better. If the difference is greater than 1 and the logarithm is greater than 0, it indicates... More inclined to generate ;

[0167] The length-normalized reward is used to control the length of the output sequence and avoid generating longer but lower-quality output sequences.

[0168] The sigmoid function is defined as follows: Normalize to the [0,1] interval;

[0169] For batch data size;

[0170] The reward difference scaling hyperparameter controls the degree of reward difference scaling, and is usually set between 2 and 2.5, with the expectation that this log probability value will be larger.

[0171] The target reward margin hyperparameter is used to control the difference between the probability of a preferred answer and the probability of a rejection in the output of a large model, preventing excessive deviation. The value should be between 0.5 and 1.5, increasing it... Values ​​can improve the model's generalization ability.

[0172] 5. Learn and train human preference optimization algorithms

[0173] The learning and training process of the human preference optimization algorithm is divided into two stages: the preparation stage and the training stage.

[0174] (1) Preparation stage

[0175] Step 1: Obtain a pre-trained model: Choose a basic large model, LLM;

[0176] Step 2: Supervised fine-tuning of SFT: Fine-tuning the base LLM using high-quality quality and response data to obtain the fine-tuned pre-trained SFT model.

[0177] Step 3: Prepare preference data: Collect and record data using large-sample parallel simulation experiments. Given pairs of human preference data, construct a training dataset D for the human preference algorithm.

[0178] The human preference dataset constructed for urban drone detection systems is shown in the table below:

[0179] Table 3 Human Preference Dataset

[0180]

[0181] (2) Training phase

[0182] Step 1: Initialize the model: Load the SFT model as a policy model That is, the model that needs to be trained and optimized;

[0183] Step 2: Training loop. Extract a partial batch of data from the human preference dataset D. ;

[0184] Step 3: Forward Propagation. For each data point in the batch, calculate the strategy model. generate Log probability: , value;

[0185] Step 4: Calculate the target damage value based on the target loss function formula mentioned above, combined with... Hyperparameters are used to calculate the average damage value of a batch of data.

[0186]

[0187] In the formula, Batch size;

[0188] Step 6: Backpropagation: Calculate the loss function Compared to the strategy model Find the gradient of the parameters and their derivatives.

[0189] Step 7: Parameter Update. Update the policy model using the optimizer based on the gradient. Parameters;

[0190] Step 8: Repeat the training loop. Continuously repeat the above process until the model converges or reaches the preset number of training steps.

[0191] In its specific implementation, this application provides a computer storage medium and a corresponding data processing unit. The computer storage medium is capable of storing a computer program, which, when executed by the data processing unit, can run the invention's content regarding an adaptive system capability evolution method based on human feedback learning, as well as some or all of the steps in various embodiments. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0192] Those skilled in the art will clearly understand that the technical solutions in the embodiments of the present invention can be implemented using computer programs and their corresponding general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of computer programs, i.e., software products. These computer program software products can be stored in a storage medium and include several instructions to cause a device containing a data processing unit (which may be a personal computer, server, microcontroller, MCU, or network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.

[0193] This invention provides an idea and method for the evolution of adaptive system capabilities based on human feedback learning. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A method for learning-based adaptive system capability evolution based on human feedback, characterized in that, The method comprises the following steps: Step 1: design adaptive conditions, and the large model in the adaptive system evolves according to different adaptive conditions and adopts different strategies to output execution effects; Step 2: establish a comprehensive evaluation index system to comprehensively evaluate the execution effects; Step 3: construct a human preference data pair according to the comprehensive evaluation score; Step 4: design a human preference optimization loss function; Step 5: use a human preference optimization method for training, optimize and fine-tune the large model parameters of the adaptive system according to the human preference data pair and the human preference optimization loss function, and realize the adaptive system capability evolution based on human feedback learning.

2. The method of claim 1, wherein, The adaptive conditions in step 1 include: Four types of resource damage, resource degradation, electronic interference and sudden target; Among them, the resource damage type adaptive condition includes: perception resource damage, decision resource damage, platform resource damage and communication resource damage; The resource degradation type adaptive condition includes: resource overload attack and resource denial of service attack; The electronic interference type adaptive condition includes: radar interference, communication interference and navigation interference; The sudden target type adaptive condition includes: target type, number, direction, height and speed.

3. The method of claim 2, wherein, The comprehensive evaluation index system in step 2 includes the following four indexes: Resource cost, used to measure the types and quantities of resources required for generating strategies; Strategy execution time, used to measure the time interval from the start of resource accepting strategy to the completion of strategy execution; Profit loss ratio, used to measure the proportion of income and loss resources after the completion of strategy execution; Ability recovery, used to measure the degree of system capability recovery of the adaptive system based on autonomous adjustment strategy.

4. The method of claim 3, wherein, The comprehensive evaluation of the execution effects in step 2 includes: Step 2-1, constructing resource cost calculation model, calculating resource cost as follows: wherein, a number of resource classes allocated in the policy, a number of resource classes allocated in the policy, a cost of resource class allocated in the policy, a cost of resource class allocated in the policy, a number of resource classes; Step 2-2, build a strategy execution time calculation model to calculate the strategy execution time as follows: wherein, is a resource a time interval from receiving the policy to completion of the policy execution, is a resource a time of completion of the policy execution, is a resource a time of completion of the policy execution; Step 2-3, build the yield loss ratio calculation model, calculate the yield loss ratio by obtaining the ratio of yield value to the number of damaged resources , as follows: wherein, is a benefit value, is a number of damaged resources; Step 2-4, build the ability recovery rate calculation model, calculate the ability recovery rate as follows: wherein, is an initial system capability value, is a restored system capability value; Step 2-5, build the execution effect comprehensive evaluation model, calculate the comprehensive evaluation score Specifically as follows: wherein, is the first output result of the large model, , , , and are the maximum values of resource cost, strategy execution time, revenue loss ratio and capability recovery rate in different strategies, respectively, and are index weighting coefficients. Step 2-6: use the mapping method of Sigmoid function to map the comprehensive evaluation score to the interval [0, 100].

5. The method of claim 4, wherein, The mapping method of Sigmoid function in step 2-6 is as follows: wherein, is the mapped overall evaluation score.

6. The method of claim 5, wherein, The construction of human preference data pair in step 3 is as follows: Step 3-1: design the human preference data pair in the following form: wherein, is a prompt word for a large model; is a preferred answer result of a large model or an answer adjusted by human feedback; is a rejected answer result of a large model; Step 3-2, determine the preferred answer result based on the comprehensive evaluation score of the execution effect output by the large model and the rejection answer result , output with a comprehensive evaluation score higher than the threshold as the preferred answer, and output with a comprehensive evaluation score lower than the threshold as the rejection answer; Step 3-3: fine-tune the execution effect evaluation score by artificial fine-tuning to obtain the final preference answer and rejection answer.

7. The method of claim 6, wherein, The design human preference optimization loss function described in Step 4 to maximize the difference between the log probability values of the preferred answer result and the rejected answer result .

8. The method of claim 7, wherein, The human preference optimization loss function in step 4 is expressed as follows: wherein, is a human preference optimization loss function; denotes a large model in an adaptive system, i.e., a policy model for a given prompt word generates a probability of a preference answer result ; denotes a policy model for a given prompt word generates a probability of a rejection answer result ; denotes a log probability difference, wherein and is a length-normalized reward, is a reward difference scaling hyperparameter, is a target reward margin hyperparameter.

9. The method of claim 8, wherein, The training using the human preference optimization method in step 5 includes: Preparation stage: Step 5-1: obtain a pre-trained model, and select a basic large model LLM; Step 5-2: supervise fine-tuning SFT, and fine-tune the basic large model LLM using the pre-set question and answer data pair to obtain a fine-tuned pre-trained SFT model; Step 5-3, prepare preference data, collect and record all human preference data pairs by large sample parallel simulation test method and construct the human preference algorithm training data set D accordingly; Training stage: Step 5 - 4, initialize model, load SFT model as policy model i.e., the model that needs to be trained and optimized; Step 5-5, extracting a partial batch of human preference data pairs from the human preference dataset D forming a training dataset; Step 5-6, forward pass, for each data in batch, compute policy model Generate preference answer result And reject answer result Log probability of And ; Step 5-7, optimize the loss function based on human preference, calculate the average loss value of the batch data, and perform back propagation, and update the parameters of the policy model ; Step 5-8: repeat steps 5-5 to 5-7 until the pre-set condition is met.

10. The method of claim 9, wherein, The pre-set condition in step 5-8 is that the strategy model converges or reaches the pre-set training step number.