Model generalization ability improvement method, device and equipment and readable storage medium

By constructing a reward model and optimizing the basic model using reinforcement learning algorithms, and adjusting parameters using user feedback, the problem of insufficient generalization ability of general models in optical communication networks was solved, and the performance of the model was improved and continuously optimized in specific environments.

CN121998024APending Publication Date: 2026-05-08FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD
Filing Date
2026-01-26
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing pre-trained general models lack generalization ability in optical communication network application scenarios, leading to performance degradation. Furthermore, traditional fine-tuning methods are costly, rely on manual intervention, and may impair the model's generalization ability.

Method used

By collecting user preference feedback, a reward model is built and the basic model is optimized by combining reinforcement learning algorithms. User feedback is used to adjust model parameters to improve generalization ability.

Benefits of technology

It achieves model performance optimization and generalization capability improvement in specific environments, reduces data threshold and privacy risks, and supports personalized deployment and continuous optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998024A_ABST
    Figure CN121998024A_ABST
Patent Text Reader

Abstract

The invention discloses a model generalization ability improvement method and device, equipment and a readable storage medium. The method comprises the steps that by collecting user preference feedback, model optimization can directly utilize local knowledge and intention of domain users (users), sensitive original scene data or a large number of standard labels do not need to be obtained, and the data threshold and the privacy risk are reduced; a reward model is constructed according to all preference feedbacks, scattered and subjective user preferences are summarized and extracted into a stable and computable scoring function, and a quantitative standard is provided for automatic optimization; feedback from a reward model (representing user preferences) is continuously received in the reinforcement learning process, and internal parameters of the reward model are gradually adjusted to generate output better meeting user expectations. According to the method and the device, the obtained optimized artificial intelligence model shows remarkably improved task conformity and scene adaptability in a specific application environment, namely, the model generalization ability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to a method, apparatus, device, and computer-readable storage medium for improving model generalization ability. Background Technology

[0002] With the increasing application of artificial intelligence technology in complex fields such as optical communication network planning, optimization, and operation and maintenance, pre-trained general-purpose models face severe generalization challenges. These general-purpose models are typically trained on broad benchmark datasets, and while they perform well on general tasks, their performance often degrades significantly when directly deployed to specific, diverse production networks or customer scenarios. This is because the data distribution, business objectives, and constraints of the target application scenario differ from those of the training data, leading to a decline in model performance.

[0003] In existing technologies, this problem is typically addressed by retraining or fine-tuning the model based on target scene data. However, this approach has significant drawbacks: First, it requires collecting a large amount of labeled data from the target scene, which is costly and raises data privacy and security concerns. Second, the process usually requires manual intervention from algorithm experts and cannot be completed independently and quickly by domain users (such as network engineers), making it difficult to support large-scale, personalized model deployment needs. Finally, traditional fine-tuning methods focus on fitting the model to limited scene data, which may impair the model's inherent generalization ability to handle unseen situations within the scene, leading to overfitting. Summary of the Invention

[0004] This application provides a method, apparatus, device, and computer-readable storage medium for improving the generalization capability of a model, which can solve the technical problem of insufficient generalization capability of generalized models in optical communication network application scenarios in the prior art.

[0005] In a first aspect, embodiments of this application provide a method for improving model generalization ability, the method comprising: Obtain user feedback on the basic artificial intelligence model based on input Output prediction results Preference feedback The preference feedback is used to instruct the user on preferences. Expected model output Superior , The value can range from 1 to n, where n is a preset value; A reward model is constructed based on all preference feedback, wherein the reward model is used to score the output of the strategy model; An optimized artificial intelligence model is obtained by using a basic artificial intelligence model as the initial model for the policy model, combined with the reward model, and then training the policy model using a reinforcement learning algorithm.

[0006] In conjunction with the first aspect, in one implementation, constructing the reward model based on all preference feedback includes: Based on preference feedback To obtain training data ;in, include: , as well as ; The reward model is trained using all the training data, wherein the training objective of the reward model is to provide... To make it The rating is higher than that of The rating.

[0007] In conjunction with the first aspect, in one implementation, the reward model is trained using the cross-entropy loss function.

[0008] In conjunction with the first aspect, in one implementation, the reinforcement learning algorithm is a proximal policy optimization algorithm.

[0009] In conjunction with the first aspect, in one implementation, the basic artificial intelligence model is deployed in a target application scenario, which is a network planning, network optimization, or network operation and maintenance scenario in an optical communication network.

[0010] In conjunction with the first aspect, in one implementation, the acquisition of user feedback on the basic artificial intelligence model based on input... Output prediction results Preference feedback Previously, it also included: The initial model is pre-trained to obtain a pre-trained model; The pre-trained model is then subjected to supervised fine-tuning using labeled data to obtain a basic artificial intelligence model.

[0011] In conjunction with the first aspect, in one implementation, after obtaining the optimized artificial intelligence model, the method further includes: The reward model is iteratively updated based on newly acquired user preference feedback. The optimized artificial intelligence model is iteratively trained using a reinforcement learning algorithm, in conjunction with the updated reward model. The model after iterative training was evaluated using an independent test set to verify its generalization ability. The evaluated model is deployed to real-world application scenarios, and its performance is continuously monitored, along with new user feedback, for subsequent optimization.

[0012] Secondly, embodiments of this application provide a model generalization capability enhancement device, the model generalization capability enhancement device comprising: The acquisition module is used to acquire user feedback on the basic artificial intelligence model based on input. Output prediction results Preference feedback The preference feedback is used to instruct the user on preferences. Expected model output Superior , The value can range from 1 to n, where n is a preset value; A building module is used to construct a reward model based on all preference feedback, wherein the reward model is used to score the output of the strategy model; The optimization module is used to initialize the model with the basic artificial intelligence model as the policy model, and then train the policy model using a reinforcement learning algorithm in conjunction with the reward model to obtain an optimized artificial intelligence model.

[0013] Thirdly, embodiments of this application provide a model generalization capability enhancement device, the model generalization capability enhancement device including a processor, a memory, and a model generalization capability enhancement program stored in the memory and executable by the processor, wherein when the model generalization capability enhancement program is executed by the processor, it implements the steps of the model generalization capability enhancement method as described in the first aspect.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a model generalization capability enhancement program, wherein when the model generalization capability enhancement program is executed by a processor, it implements the steps of the model generalization capability enhancement method as described in the first aspect.

[0015] The beneficial effects of the technical solutions provided in this application include: In this embodiment, by collecting user preference feedback, model optimization can directly utilize the local knowledge and intent of domain users (users) without needing to obtain sensitive raw scene data or a large amount of standard annotations, thus reducing the data threshold and privacy risks. A reward model is constructed based on all preference feedback, summarizing and refining scattered and subjective user preferences into a stable and computable scoring function, providing a quantitative standard for automated optimization. During reinforcement learning, feedback from the reward model (representing user preferences) is continuously received, and its internal parameters are gradually adjusted to produce outputs that better meet user expectations. Through this embodiment, the resulting optimized artificial intelligence model exhibits significantly improved task compliance and scene adaptability in specific application environments, i.e., improved model generalization ability. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating an embodiment of the method for improving the generalization ability of the model in this application; Figure 2 This is a schematic diagram of the functional modules of an embodiment of the model generalization capability enhancement device of this application; Figure 3 This is a schematic diagram of the hardware structure of the model generalization capability enhancement device involved in the embodiments of this application. Detailed Implementation

[0017] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0018] First, the core concept of this application will be explained to enable those skilled in the art to understand this application.

[0019] In this embodiment, user preference feedback for the output of the base model is first collected; then, a reward model that can quantify user preferences is constructed using this feedback; finally, the reward model is used as an automated guiding signal to fine-tune the base model through reinforcement learning, thereby driving the model's behavior to align with the user's expectations, achieving performance optimization and generalization improvement in specific environments.

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0021] In a first aspect, embodiments of this application provide a method for improving model generalization ability.

[0022] In one embodiment, reference is made to Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the method for improving the generalization ability of the model in this application. Figure 1 As shown, methods to improve model generalization ability include: Step S10: Obtain user feedback on the basic artificial intelligence model based on input. Output prediction results Preference feedback The preference feedback is used to instruct the user on preferences. Expected model output Superior , The value can range from 1 to n, where n is a preset value; In this embodiment, before optimizing the basic artificial intelligence model, it is necessary to prepare the basic artificial intelligence model in advance and deploy it to a specific application scenario.

[0023] Furthermore, in one embodiment, before step S10, the method further includes: The initial model is pre-trained to obtain a pre-trained model; The pre-trained model is then subjected to supervised fine-tuning using labeled data to obtain a basic artificial intelligence model.

[0024] Furthermore, in one embodiment, the basic artificial intelligence model is deployed in a target application scenario, which is a network planning, network optimization, or network operation and maintenance scenario in an optical communication network.

[0025] In this embodiment, the root cause analysis of optical communication network faults is used as an example for explanation as follows: First, large-scale network device logs and performance indicator sequences are collected as pre-training data; this data can be unlabeled. A self-supervised learning algorithm (such as masked language modeling) is used to pre-train an initial neural network model (such as the Transformer architecture) to obtain a pre-trained model capable of understanding the basic patterns and semantics of network data. Next, a subset of network fault case data with root cause annotations is collected as supervised fine-tuning data. This labeled data is used to perform supervised fine-tuning of the pre-trained model, training it to output the most probable root cause category based on the input network alarm and performance indicator sequences. The final model, denoted as M_base, serves as the base AI model for deployment. Then, M_base is deployed to a carrier's optical network operation and maintenance system.

[0026] Once deployed, when a network failure occurs, the system will use the current alarm set, performance metric snapshots, etc., as input. Input M_base; after processing, M_base outputs what it considers the most likely root cause, i.e., the prediction result. (For example (This indicates that the optical module is aging).

[0027] Operations engineers (users) are viewing the prediction results. Subsequently, based on their experience, they might conclude that... Inaccurate or incomplete. At this point, the engineer provides a more accurate or reasonable root cause analysis result through the system interface, that is, the result the user has identified regarding the root cause. Expected model output (For example This indicates a loose fiber optic connector, accompanied by performance degradation of the optical module. The system records this selection behavior of the engineer, forming a preference feedback. Its semantic meaning is to indicate Superior .

[0028] Repeat this process to collect n feedback records (n is a preset positive integer, such as 1000), forming a feedback set. This collection is automated through the human-computer interaction logs of the operations and maintenance system.

[0029] Step S20: Construct a reward model based on all preference feedback, wherein the reward model is used to score the output of the strategy model; In this embodiment, a reward model is constructed based on the feedback set. Its function is to score any (input, output) pair. The higher the score, the more the output matches the user's preferences given the input.

[0030] Further, in one embodiment, step S20 includes: Based on preference feedback To obtain training data ;in, include: , as well as ; The reward model is trained using all the training data, wherein the training objective of the reward model is to provide... To make it The rating is higher than that of The rating.

[0031] In this embodiment, for each preference feedback Parse out what it contains , as well as This constitutes a training dataset. That is, a triple ( , , ).all This constitutes the training dataset for the reward model.

[0032] The reward model itself can be a neural network (such as a multilayer perceptron), whose input is a concatenated or encoded "input-output" pair, and whose output is a scalar score.

[0033] During training, for each triplet ( , , The training objective is to make the reward model effective against ( , The score for ) was higher than that for ( , (Rating)

[0034] To achieve this goal, the reward model is trained using the cross-entropy loss function. Specifically, the loss function is defined as follows:

[0035] in, For reward model pairs ( , The rating of () and the () , The difference in scores; This is the Sigmoid function, used to map the difference to the interval (0, 1).

[0036] Step S30: Using the basic artificial intelligence model as the initialization model of the policy model, and combining it with the reward model, the policy model is trained using a reinforcement learning algorithm to obtain an optimized artificial intelligence model.

[0037] In this embodiment, M_base is used as the initialization model of the policy model, and the reward model is used as the guide to train a policy model through reinforcement learning, ultimately obtaining an optimized artificial intelligence model.

[0038] Furthermore, in one embodiment, the reinforcement learning algorithm is a proximal policy optimization algorithm.

[0039] In this embodiment, the steps for training the policy model are as follows: The initialization model uses the basic artificial intelligence model as the policy model M_policy; The M_policy is trained using a proximal policy optimization algorithm, specifically: Generate a large number of network states using historical fault data or simulation environments as input; The policy model M_policy generates a root cause analysis output A based on the input I; Reward function: It is provided by a trained reward model, which outputs a scalar reward value r based on (I, A). This reward value represents the degree to which the policy model outputs A based on output I, which is consistent with the user's preference. The Proximal Policy Optimization (PPO) algorithm calculates the policy gradient by sampling a large amount of (I, A, r) data and updates the parameters of M_policy according to its unique objective function (which includes reward terms and policy change constraints). After multiple iterations, M_policy is optimized to tend to generate outputs that achieve higher scores, i.e., root cause analysis results that better align with user preferences.

[0040] Once training is complete, M_policy represents the optimized AI model, which can provide judgments that are closer to the experience of operations engineers when dealing with scenarios similar to the feedback data.

[0041] In this embodiment, by collecting user preference feedback, model optimization can directly utilize the local knowledge and intent of domain users (users) without needing to obtain sensitive raw scene data or a large amount of standard annotations, thus reducing the data threshold and privacy risks. A reward model is constructed based on all preference feedback, summarizing and refining scattered and subjective user preferences into a stable and computable scoring function, providing a quantitative standard for automated optimization. During reinforcement learning, feedback from the reward model (representing user preferences) is continuously received, and its internal parameters are gradually adjusted to produce outputs that better meet user expectations. Through this embodiment, the resulting optimized artificial intelligence model exhibits significantly improved task compliance and scene adaptability in specific application environments, i.e., improved model generalization ability.

[0042] Furthermore, in one embodiment, after step S30, the method further includes: The reward model is iteratively updated based on newly acquired user preference feedback. The optimized artificial intelligence model is iteratively trained using a reinforcement learning algorithm, in conjunction with the updated reward model. The model after iterative training was evaluated using an independent test set to verify its generalization ability. The evaluated model is deployed to real-world application scenarios, and its performance is continuously monitored, along with new user feedback, for subsequent optimization.

[0043] In this embodiment, model optimization is not a one-time event. After optimizing the AI ​​model deployment, new feedback from operations engineers continues to be collected. The reward model is iteratively updated periodically (e.g., monthly) using the newly added feedback data. Then, using the optimized AI model as the initialization model of the strategy model, combined with the new reward model, step S30 is repeated to obtain a new version of the optimized AI model.

[0044] Before each release of a new version of the optimized AI model, a batch of historical failure cases that have never been used in training are used as an independent test set to evaluate the new version of the optimized AI model and calculate its root cause localization accuracy and other indicators to verify its generalization ability.

[0045] Finally, the evaluated model is deployed to real-world application scenarios, replacing the old model. The system continuously monitors the performance of the new model, such as prediction confidence and user adoption rate, and continues to collect new user feedback, thus forming a continuous improvement loop of deployment-feedback-optimization-redeployment for subsequent optimization.

[0046] In this embodiment, by continuously incorporating new user feedback to update the reward model and optimize the artificial intelligence model, the system can dynamically adapt to changes in the application environment and the evolution of user needs. Combined with evaluation on an independent test set, the generalization ability of the model is continuously verified and guaranteed during the iteration process. Ultimately, this method supports long-term deployment and performance maintenance of the model, forming a complete operation and maintenance closed loop.

[0047] Secondly, embodiments of this application also provide a model generalization capability enhancement device.

[0048] In one embodiment, reference is made to Figure 2 , Figure 2 This is a schematic diagram of the functional modules of an embodiment of the model generalization capability enhancement device of this application. Figure 2 As shown, the model generalization capability enhancement device includes: Module 10 is used to acquire user feedback on the basic artificial intelligence model based on input. Output prediction results Preference feedback The preference feedback is used to instruct the user on preferences. Expected model output Superior , The value can range from 1 to n, where n is a preset value; Module 20 is used to construct a reward model based on all preference feedback, wherein the reward model is used to score the output of the strategy model; The optimization module 30 is used to initialize the model with the basic artificial intelligence model as the policy model, and combine it with the reward model to train the policy model using a reinforcement learning algorithm to obtain an optimized artificial intelligence model.

[0049] Furthermore, in one embodiment, the construction module 20 is used for: Based on preference feedback To obtain training data ;in, include: , as well as ; The reward model is trained using all the training data, wherein the training objective of the reward model is to provide... To make it The rating is higher than that of The rating.

[0050] Furthermore, in one embodiment, the loss function used to train the reward model is the cross-entropy loss function.

[0051] Furthermore, in one embodiment, the reinforcement learning algorithm is a proximal policy optimization algorithm.

[0052] Furthermore, in one embodiment, the basic artificial intelligence model is deployed in a target application scenario, which is a network planning, network optimization, or network operation and maintenance scenario in an optical communication network.

[0053] Furthermore, in one embodiment, the model generalization capability enhancement device further includes a pre-training module, used for: The initial model is pre-trained to obtain a pre-trained model; the pre-trained model is then fine-tuned in a supervised manner using labeled data to obtain a basic artificial intelligence model.

[0054] Furthermore, in one embodiment, the model generalization capability enhancement device further includes a continuous optimization module, used for: The reward model is iteratively updated based on newly acquired user preference feedback. The optimized artificial intelligence model is iteratively trained using a reinforcement learning algorithm, in conjunction with the updated reward model. The model after iterative training was evaluated using an independent test set to verify its generalization ability. The evaluated model is deployed to real-world application scenarios, and its performance is continuously monitored, along with new user feedback, for subsequent optimization.

[0055] The functions of each module in the above-mentioned model generalization ability enhancement device correspond to the steps in the above-mentioned model generalization ability enhancement method embodiment, and their functions and implementation processes will not be described in detail here.

[0056] Thirdly, embodiments of this application provide a model generalization capability enhancement device, which can be a personal computer (PC), laptop computer, server, or other device with data processing capabilities.

[0057] Reference Figure 3 , Figure 3 This is a schematic diagram of the hardware structure of the model generalization capability enhancement device involved in the embodiments of this application. In the embodiments of this application, the model generalization capability enhancement device may include a processor, a memory, a communication interface, and a communication bus.

[0058] The communication bus can be of any type and is used to interconnect the processor, memory, and communication interface.

[0059] Communication interfaces include input / output (I / O) interfaces, physical interfaces, and logical interfaces used for interconnecting devices within the model generalization enhancement device, as well as interfaces used for interconnecting the model generalization enhancement device with other devices (such as other computing devices or user equipment). Physical interfaces can be Ethernet interfaces, fiber optic interfaces, ATM interfaces, etc.; user equipment can be displays, keyboards, etc.

[0060] Memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0061] The processor can be a general-purpose processor, which can call the model generalization capability enhancement program stored in memory and execute the model generalization capability enhancement method provided in the embodiments of this application. For example, the general-purpose processor can be a central processing unit (CPU). The method executed when the model generalization capability enhancement program is called can be referred to in the various embodiments of the model generalization capability enhancement method of this application, and will not be repeated here.

[0062] Those skilled in the art will understand that Figure 3 The hardware structure shown does not constitute a limitation of this application and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0063] Fourthly, embodiments of this application also provide a computer-readable storage medium.

[0064] The present application has a computer-readable storage medium storing a model generalization capability improvement program, wherein when the model generalization capability improvement program is executed by a processor, it implements the steps of the model generalization capability improvement method as described above.

[0065] The method implemented when the model generalization capability enhancement procedure is executed can be referred to in various embodiments of the model generalization capability enhancement method of this application, and will not be repeated here.

[0066] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0067] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. The terms "first," "second," and "third," etc., are used to distinguish different objects, etc., and do not indicate a sequence, nor do they limit "first," "second," and "third" to different types.

[0068] In the description of the embodiments of this application, terms such as "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a concrete manner.

[0069] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.

[0070] In some processes described in the embodiments of this application, multiple operations or steps are included in a specific order. However, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or they may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.

[0071] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.

[0072] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for improving the generalization ability of a model, characterized in that, The methods for improving the model's generalization ability include: Obtain user feedback on the basic artificial intelligence model based on input Output prediction results Preference feedback The preference feedback is used to instruct the user on preferences. Expected model output Superior , The value can range from 1 to n, where n is a preset value; A reward model is constructed based on all preference feedback, wherein the reward model is used to score the output of the strategy model; An optimized artificial intelligence model is obtained by using a basic artificial intelligence model as the initial model for the policy model, combined with the reward model, and then training the policy model using a reinforcement learning algorithm.

2. The method for improving model generalization ability as described in claim 1, characterized in that, The construction of the reward model based on all preference feedback includes: Based on preference feedback To obtain training data ;in, include: , as well as ; The reward model is trained using all the training data, wherein the training objective of the reward model is to provide... To make it The rating is higher than that of The rating.

3. The method for improving model generalization ability as described in claim 2, characterized in that, The reward model is trained using the cross-entropy loss function.

4. The method for improving model generalization ability as described in claim 1, characterized in that, The reinforcement learning algorithm is a near-end policy optimization algorithm.

5. The method for improving model generalization ability as described in claim 1, characterized in that, The basic artificial intelligence model is deployed in the target application scenario, which is a network planning, network optimization, or network operation and maintenance scenario in an optical communication network.

6. The method for improving model generalization ability as described in claim 1, characterized in that, The acquisition of user feedback on the basic artificial intelligence model based on input Output prediction results Preference feedback Previously, it also included: The initial model is pre-trained to obtain a pre-trained model; The pre-trained model is then subjected to supervised fine-tuning using labeled data to obtain a basic artificial intelligence model.

7. The method for improving model generalization ability as described in claim 1, characterized in that, After obtaining the optimized artificial intelligence model, the following is also included: The reward model is iteratively updated based on newly acquired user preference feedback. The optimized artificial intelligence model is iteratively trained using a reinforcement learning algorithm, in conjunction with the updated reward model. The model after iterative training was evaluated using an independent test set to verify its generalization ability. The evaluated model is deployed to real-world application scenarios, and its performance is continuously monitored, along with new user feedback, for subsequent optimization.

8. A device for enhancing model generalization ability, characterized in that, The model generalization capability enhancement device includes: The acquisition module is used to acquire user feedback on the basic artificial intelligence model based on input. Output prediction results Preference feedback The preference feedback is used to instruct the user on preferences. Expected model output Superior , The value can range from 1 to n, where n is a preset value; A building module is used to construct a reward model based on all preference feedback, wherein the reward model is used to score the output of the strategy model; The optimization module is used to initialize the model with the basic artificial intelligence model as the policy model, and then train the policy model using a reinforcement learning algorithm in conjunction with the reward model to obtain an optimized artificial intelligence model.

9. A device for improving model generalization ability, characterized in that, The model generalization capability enhancement device includes a processor, a memory, and a model generalization capability enhancement program stored in the memory and executable by the processor, wherein when the model generalization capability enhancement program is executed by the processor, it implements the steps of the model generalization capability enhancement method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a model generalization capability improvement program, wherein when the model generalization capability improvement program is executed by a processor, it implements the steps of the model generalization capability improvement method as described in any one of claims 1 to 7.