Method, system and equipment for optimizing partner training strategy engine and medium

By constructing a digital twin environment and fusing multimodal features, training data is dynamically generated and training strategies are optimized, solving the problem of poor model training results in existing technologies and achieving efficient and personalized model training results.

CN121503731APending Publication Date: 2026-02-10E-SURFING DIGITAL LIFE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511666923.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies cannot dynamically cover the ever-increasing new anomalies and unknown attack patterns in big data generation environments during model training. This leads to performance degradation and drift of models when facing real-world challenges. Furthermore, the low simulation environment results in evaluation distortion, and the training strategy is rigid and lacks adaptability, resulting in low training efficiency.

Method used

Construct a digital twin environment based on business training data, dynamically generate training data through multimodal feature fusion and defect root cause analysis, conduct simulated training in the digital twin environment, and optimize the training strategy engine based on multidimensional model indicators.

Benefits of technology

It improves the model's performance in the face of complex and unknown challenges, enhances the simulation accuracy and scene coverage, reduces training costs and risks, and improves training efficiency and personalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121503731A_ABST
    Figure CN121503731A_ABST
Patent Text Reader

Abstract

The invention provides a partner training strategy engine optimization method, system and device and a medium, and the method comprises the steps: obtaining a first partner training strategy engine and a to-be-partner training model in a current partner training optimization process, and business training data; the first partner training strategy engine is an original partner training strategy engine or a second partner training strategy engine obtained in a previous partner training optimization process; constructing a digital twin environment according to the service training data; according to the digital twin environment and the first partner training strategy engine, performing simulated environment partner training on the to-be-partner training model to obtain multi-dimensional model index data of the to-be-partner training model; and according to the multi-dimensional model index data, performing reward optimization on the first partner training strategy engine to obtain a second partner training strategy engine. According to the method, a partner training strategy engine can be provided, the partner training effect of the model can be effectively improved, and then the performance of the partner training model facing complex unknown challenges can be improved. The invention relates to the technical field of artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an optimization method, system, device and medium for a coaching strategy engine. Background Technology

[0002] With the rapid development of artificial intelligence technology, AI models have been widely applied in big data analysis and decision-making across various fields such as finance, healthcare, the internet, and industry. However, the performance and reliability of these models heavily depend on the quality and coverage of their training data, as well as the effectiveness of the training methods. To improve the generalization ability and robustness of AI models in complex real-world environments, "AI coaching" technology has emerged.

[0003] Currently, related technologies typically rely on historically accumulated test case libraries or predefined rules to generate test data and behaviors for model training. For example, this involves randomly injecting noise, simulating several preset failure modes, or using publicly available adversarial sample libraries (such as static samples generated by FGSM or PGD) to attack and train the model. However, this approach fails to dynamically cover the ever-increasing number of novel anomalies, unknown attack patterns (0-day attacks), and rare long-tail cases in the big data generation environment when facing complex and ever-changing big data scenarios. This results in models trained through this method performing poorly when facing dynamic and unknown challenges in the real world, exhibiting severe performance degradation and model drift.

[0004] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention

[0005] The purpose of this invention is to at least partially solve one of the technical problems existing in the related art.

[0006] The main objective of this application is to propose an optimization method, system, device, and medium for a training strategy engine. This method can provide a training strategy engine that can effectively improve the training effect of the model, thereby improving the performance of the trained model in the face of complex and unknown challenges.

[0007] To achieve the above objectives, one aspect of this application proposes an optimization method for a coaching strategy engine, comprising: Obtain the first training strategy engine and the training model to be trained in the current training optimization process, as well as the business training data. The first training strategy engine is either the original training strategy engine or the second training strategy engine obtained in the previous training optimization process. Based on the aforementioned business training data, construct a digital twin environment; Based on the digital twin environment and the first training strategy engine, the model to be trained is trained in a simulated environment to obtain multi-dimensional model index data of the model to be trained. Based on the multidimensional model index data, the first coaching strategy engine is optimized for rewards to obtain the second coaching strategy engine.

[0008] In some embodiments, obtaining business training data includes: Acquire multi-source business data from the business system; Multimodal feature extraction is performed on the multi-source business data to obtain multimodal feature data; The multimodal feature data is processed to generate synthetic data, thereby obtaining the business training data.

[0009] In some embodiments, the process of synthesizing the multimodal feature data to obtain the business training data includes: Obtain the root cause information of the defects of the model to be trained in the current training optimization process. The root cause information is used to characterize the reasons for the model performance defects of the model to be trained in the previous training optimization process. Based on the defect root cause information, the multimodal feature data is filtered to obtain few-sample feature data and adversarial example feature data; The few-sample feature data are used to synthesize samples to obtain synthesized sample data; Adversarial perturbations are added to the adversarial sample feature data to obtain perturbation sample data; Based on the synthetic sample data and the perturbation sample data, the multimodal feature data is fused to obtain the business training data.

[0010] In some embodiments, constructing a digital twin environment based on the business training data includes: Obtain logical topology data from the business system and construct a hardware resource model for the business system; The business training data is input into the hardware resource model, and the hardware resource model is used to simulate the environment of the business training data to obtain the digital twin environment of the business system.

[0011] In some embodiments, the step of conducting simulated environment training on the model to be trained based on the digital twin environment and the first training strategy engine to obtain multi-dimensional model indicator data of the model to be trained includes: Obtain the decision strategy output by the model to be trained during the current training optimization process; The decision-making strategy is executed through the digital twin environment to determine the target environmental state of the digital twin environment; The first training strategy engine performs index analysis on the decision strategy and the target environment state to obtain multi-dimensional model index data of the training model.

[0012] In some embodiments, obtaining the decision strategy output by the training model in the current training optimization process includes: The training strategy engine generates training action instructions. Based on the training action instructions, the state of the digital twin environment is updated to obtain an intermediate environment state; The intermediate environment state is input into the training model to make model decisions and obtain the decision strategy.

[0013] In some embodiments, the method further includes: Obtain the decision-making strategy of the model to be trained; The decision-making strategy is subjected to feature importance analysis to obtain first analytical information; The decision-making strategy is subjected to interpretability analysis to obtain second analytical information; Based on the first and second analysis information, root cause localization processing is performed to obtain the defect root cause information of the model to be trained.

[0014] To achieve the above objectives, another aspect of this application proposes an optimization system for a coaching strategy engine, comprising: The first processing unit is used to obtain the first training strategy engine and the training model to be trained in the current training optimization process, as well as business training data. The first training strategy engine is the original training strategy engine or the second training strategy engine obtained in the previous training optimization process. The second processing unit is used to construct a digital twin environment based on the business training data; The third processing unit is used to conduct simulated environment training on the model to be trained based on the digital twin environment and the first training strategy engine, and obtain multi-dimensional model index data of the model to be trained. The fourth processing unit is used to optimize the rewards of the first coaching strategy engine based on the multidimensional model index data to obtain the second coaching strategy engine.

[0015] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0016] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.

[0017] On the other hand, embodiments of this application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the method described above.

[0018] The embodiments of this application include at least the following beneficial effects: This application provides a method, system, device, and medium for optimizing a training strategy engine. The method involves acquiring a first training strategy engine and a model to be trained during the current training optimization process, as well as business training data. The first training strategy engine is either the original training strategy engine or a second training strategy engine obtained in a previous training optimization process. Based on the business training data, a digital twin environment is constructed. Using the digital twin environment and the first training strategy engine, the model to be trained is subjected to simulated environment training to obtain multi-dimensional model indicator data. Based on the multi-dimensional model indicator data, the first training strategy engine is optimized for rewards to obtain a second training strategy engine. This method constructs a digital twin environment based on business training data and uses the training strategy engine to conduct simulated environment training on the model to be trained within the digital twin environment. It allows for attack training of the model within this simulated environment, effectively improving the model training effect and enhancing the performance of the trained model when facing dynamic and unknown challenges in real-time. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating an optimization method for a coaching strategy engine provided in an embodiment of this application; Figure 2 This is a schematic diagram of a process for obtaining business training data provided in an embodiment of this application; Figure 3 This is a detailed flowchart of step S230 provided in an embodiment of this application; Figure 4 This is a detailed flowchart of step S120 provided in an embodiment of this application; Figure 5 This is a detailed flowchart of step S130 provided in an embodiment of this application; Figure 6 This is a logical diagram of step S130 provided in an embodiment of this application; Figure 7This is a detailed flowchart of step S510 provided in an embodiment of this application; Figure 8 This is a schematic diagram of one optional process of an optimization method for a coaching strategy engine provided in an embodiment of this application; Figure 9 This is a schematic diagram of the framework of an optimization system for a coaching strategy engine provided in an embodiment of this application; Figure 10 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses / devices and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0021] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0022] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0024] The following is an explanation of the terms used in the embodiments of this application: AI training companion refers to an intelligent simulation training system that provides dynamic, adaptive, and continuous training challenges to the target AI model by simulating various scenarios, faults, and attacks in a real big data environment, in order to improve its performance and robustness.

[0025] A digital twin environment is a virtual simulation environment that models a real business system with high precision. It can simulate data flow, user behavior, system status, and network conditions, providing a safe, controllable, and resettable test sandbox for AI tutoring.

[0026] Related technologies typically rely on historically accumulated test case libraries or predefined rules to generate test data and behaviors for model training; for example, attacking and training the model by randomly injecting noise, simulating several preset fault modes, or using publicly available adversarial sample libraries (such as static samples generated by FGSM and PGD); this approach faces the following challenges when dealing with complex and ever-changing big data scenarios: 1. The contradiction between static data and insufficient scenario coverage: Related technologies heavily rely on static, pre-built datasets and rule bases. The scenarios they generate are limited and fixed, unable to dynamically cover the ever-emerging new anomalies, unknown attack patterns (0-day attacks), and rare long-tail cases in big data production environments. This leads to poor performance of trained models when facing dynamic and unknown challenges in the real world, resulting in severe performance degradation and model drift, turning them into "lab champions." Essentially, the breadth and diversity of training data cannot match the complexity of the production environment.

[0027] 2. Risk of Evaluation Distortion Due to Low Environmental Simulation: Related technologies are typically practiced in highly simplified or abstract simulation environments, which cannot faithfully reproduce the complex topology, microservice dependencies, network latency fluctuations, and hardware resource contention of real big data platforms. Excellent practice results achieved in low-simulation environments often cannot be effectively transferred to production environments. Optimization directions may deviate from real-world requirements, and the model's performance under real-world extreme failures cannot be verified, significantly reducing the effectiveness of model practice and resulting in high trial-and-error costs and substantial risks.

[0028] 3. Rigid and inflexible training strategies lack adaptability: Related techniques typically train models using fixed strategies. They cannot dynamically adjust based on the current capabilities and weaknesses of the model being trained. Whether it's random fault injection or using a fixed adversarial sample library, it's a "crude" training method, resulting in low training efficiency and significant waste of computational resources. The training optimization process cannot be "tailored to individual needs," overtraining the model on its already mastered knowledge points while undertraining its true "knowledge blind spots," failing to achieve personalized and efficient model capability improvement.

[0029] It should be noted that the aforementioned related technologies are only used to assist in understanding the technical solutions of this application and do not mean that they belong to the publicly disclosed prior art.

[0030] In view of this, this application provides an optimization method, system, device, and medium for a training strategy engine. The method constructs a digital twin environment based on business training data. Specifically, it builds a hardware resource model based on logical topology data from the business system and uses the business training data as the driving data source to determine the digital twin environment of the business system. This is combined with subsequent simulated environment training of the model to be trained within the digital twin environment using a first training strategy engine (specifically, in any training optimization process, training action instructions are executed through the digital twin environment to update the state, and the updated intermediate environment state is input into the model to be trained for model decision-making). This method can faithfully reproduce the complex topology, microservice dependencies, network latency fluctuations, and hardware resource competition of a real big data platform. It effectively improves the environmental simulation degree during subsequent model training, effectively improves the model training effect, and reduces training costs and risks.

[0031] Furthermore, this method constructs a digital twin environment based on business training data, and uses a first training strategy engine to simulate the training environment of the model to be trained within the digital twin environment. Specifically, it performs multimodal feature fusion on multi-source business data, and processes the obtained multimodal feature data by sample synthesis and adversarial perturbation based on defect root cause information. This allows the obtained business training data to be dynamically updated as the training optimization process progresses, thereby enabling the subsequently generated data twin environment to change dynamically as well. This effectively improves the scenario coverage and helps to dynamically cover the emerging new anomalies, unknown attack patterns (0-day attacks), and rare long-tail cases in the big data production environment, thereby effectively improving the performance of the training model when facing dynamic and unknown challenges in the real world.

[0032] Furthermore, in any training optimization process, this method optimizes the first training strategy engine by using multi-dimensional model index data of the model to be trained. This allows the training strategy engine to dynamically adjust the training strategy based on the actual ability level and weaknesses of the "model to be trained", effectively improving the efficiency of model training, reducing the waste of computing resources, and increasing the personalization of model training.

[0033] This application provides an optimization method, system, device, and medium for a coaching strategy engine, which can be applied to artificial intelligence application scenarios. In these scenarios, AI service providers can use the method provided in this application to construct a digital twin and use a coaching decision model to train the model to be coached, thereby improving the model's performance in complex and ever-changing big data scenarios.

[0034] The method provided in this application can be applied to a terminal, a server, or software running on a terminal or server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited thereto; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the method, but is not limited to the above forms.

[0035] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0036] Reference Figure 1 , Figure 1This is an optional flowchart illustrating an optimization method for a coaching strategy engine provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S110 to S140.

[0037] Step S110: Obtain the first training strategy engine and the training model to be trained in the current training optimization process, as well as the business training data. The first training strategy engine is the original training strategy engine or the second training strategy engine obtained in the previous training optimization process. In this application embodiment, the model to be trained can be an AI model that needs AI training. There are various specific model types, including Large Language Models (LLMs) undergoing fine-tuning training, machine learning decision models, deep learning decision models, reinforcement learning decision models, and generative decision models. Specifically, the machine learning decision model can be any one of the following: Decision Tree model, Random Forest model, Support Vector Machine (SVM) model, Logistic Regression model, etc.; the deep learning decision model can be a Feedforward Neural Network model, etc.; the reinforcement learning decision model can be a Q-Learning model, Deep Reinforcement Learning (DRL) model, etc.; and the generative decision model can be a Generative Adversarial Network (GAN), etc. The examples in this application are for illustrative purposes only and do not limit the models to be trained in this application.

[0038] Understandably, a training strategy engine is a rule-based and algorithm-driven system used to formulate different training decisions and evaluate their effectiveness. In a digital twin environment, these decisions are typically represented as training action instructions. Specifically, in the first training optimization process, the first training strategy engine can be a pre-set engine (i.e., the original engine), while the training data can be the existing business data of the business system. If the business system is a communication network system, its corresponding business data could be data such as channel delay, channel bandwidth, signal strength, and data transmission volume of each node in the communication network. Alternatively, if the business system is a cloud computing system, its corresponding business data could be data such as the tasks being processed by each computing node in the cloud computing network, the load resource data of each computing node (e.g., CPU utilization, GPU utilization), and remaining available storage space.

[0039] In the second training optimization process, the first training strategy engine can be the second training strategy engine obtained after reward optimization in the previous training optimization process; and the business training data can be the original business data of the business system, or the business training data obtained by optimizing the multi-dimensional model indicator data based on the previous training optimization process.

[0040] Reference Figure 2 In some embodiments, acquiring business training data includes: Step S210: Obtain multi-source business data from the business system; Step S220: Extract multimodal features from the multi-source business data to obtain multimodal feature data; In this application embodiment, multi-source business data can be a collection of data from different data sources / types in the business system. Specifically, it can be collected / extracted from multi-source heterogeneous data such as distributed logs, real-time data streams, and historical databases in the business system. There are already various ways to obtain this multi-source business data, which will not be described in detail here.

[0041] Understandably, in one optional implementation, after obtaining multi-source business data from the business system, the multi-source business data can typically be preprocessed. Specific preprocessing operations may include missing value imputation, outlier handling, and format standardization. Multimodal feature fusion can involve extracting and aligning features from different modalities of multi-source business data, such as structured data, text data, and time-series data, to form a unified feature representation. All determined unified feature representations are then integrated into multimodal feature data. Various methods for multimodal feature extraction already exist, which will not be elaborated upon here. For example, structured data can have its corresponding feature vector representation extracted based on preset rules, text data can have its corresponding feature vector representation extracted based on the BERT model, and time-series data can have its temporal feature vector extracted.

[0042] Step S230: Perform synthetic data generation processing on the multimodal feature data to obtain the business training data.

[0043] Reference Figure 3 Furthermore, step S230, which involves performing synthetic data generation processing on the multimodal feature data to obtain the business training data, includes: Step S310: Obtain the root cause information of the defects of the model to be trained in the current training optimization process. The root cause information is used to characterize the reasons for the model performance defects of the model to be trained in the previous training optimization process. Step S320: Based on the defect root cause information, perform data filtering on the multimodal feature data to obtain few-sample feature data and adversarial example feature data; Step S330: Perform sample synthesis on the few sample feature data to obtain synthesized sample data; Step S340: Add adversarial perturbations to the adversarial sample feature data to obtain perturbation sample data; Step S350: Based on the synthesized sample data and the perturbation sample data, perform data fusion on the multimodal feature data to obtain the service training data.

[0044] In the embodiments of this application, for any training optimization process, step S310 may be to obtain the root cause of the performance defect of the training model in the previous training optimization process, which is recorded as the defect root cause information. One way to represent the defect root cause information may be "the training model depends on irrelevant features", "the training model is too sensitive to certain disturbances", "the ability to understand rare cases is poor", etc.

[0045] Understandably, data filtering can be based on the root cause information of defects, selecting corresponding feature data from multimodal feature data. Specifically, when the root cause information of defects includes "the model to be trained is too sensitive to certain perturbations" and "has poor understanding of rare cases", data filtering can be to select a few feature data with a small total sample size from the multimodal feature data, denoted as few-sample feature data; and to randomly select a few feature data from the multimodal feature data to add random perturbations, denoted as adversarial example feature data.

[0046] It should be noted that sample synthesis can be based on a small number of feature data and specified conditions such as "anomaly type" and "fault severity", and a conditional generative adversarial network (C-GAN) can be used to generate synthetic sample data corresponding to the small number of feature data; while adversarial perturbation can be added by using adversarial attack algorithms such as item gradient descent (FGSM / PGD) to add small perturbations to the original data, generating adversarial samples to improve the robustness of the model, which are denoted as perturbation sample data.

[0047] Data fusion can be performed by adding and updating multimodal feature data based on synthetic sample data and perturbation sample data. Specifically, it can be done by concatenating the feature vectors of synthetic sample data and perturbation sample data and adding them to the multimodal feature data to obtain the business training data in the current training optimization process.

[0048] Step S120: Construct a digital twin environment based on the business training data; In this application embodiment, a digital twin environment for the business system can be constructed based on business training data using digital twin technology.

[0049] Reference Figure 4 In some embodiments, step S120, constructing a digital twin environment based on the business training data, includes: Step S410: Obtain logical topology data from the business system and construct a hardware resource model for the business system; Step S420: Input the business training data into the hardware resource model, and simulate the environment of the business training data through the hardware resource model to obtain the digital twin environment of the business system.

[0050] In this embodiment, the logical topology data can be a data set of business logic and system topology of a business system. Specifically, the system topology includes relevant data of each node in the business system. For example, in a communication network business system, the system topology includes business-related data such as CPU, memory, and network bandwidth of each node in the business system.

[0051] Understandably, after constructing the hardware resource model in the business system, the business training data can be identified as the input source for environment simulation. By inputting the business training data into the hardware resource model, a digital twin environment of the business system can be obtained, which is used to simulate the complexity and uncertainty of the real world.

[0052] Step S130: Based on the digital twin environment and the first training strategy engine, conduct simulated environment training on the model to be trained to obtain multi-dimensional model index data of the model to be trained. In this embodiment of the application, for any training optimization process, the first training strategy engine can train the training model based on the digital twin environment to obtain multi-dimensional model index data of the training model.

[0053] Reference Figure 5 and Figure 6 In some embodiments, step S130, which involves providing simulated environment training to the model to be trained based on the digital twin environment and the first training strategy engine, to obtain multi-dimensional model indicator data of the model to be trained, includes: Step S510: Obtain the decision strategy output by the training model in the current training optimization process; Reference Figure 7 Furthermore, step S510, obtaining the decision strategy output by the training model in the current training optimization process, includes: Step S710: Generate training action instructions through the first training strategy engine; Step S720: Update the state of the digital twin environment according to the training action instructions to obtain the intermediate environment state; Step S730: Input the intermediate environment state into the training model to make model decisions and obtain the decision strategy.

[0054] In this embodiment, the training strategy engine (i.e., the first training strategy engine) in the current training optimization process can dynamically generate the next training decision (i.e., training action instruction) based on meta-learning or reinforcement learning (RL) algorithms. The training decision can typically include: task type (such as injecting noise, simulating faults), task difficulty, task duration, etc. Then, the training action instruction is executed through the digital twin environment, thereby realizing the internal state update of the digital twin environment and obtaining the intermediate environment state. Specifically, the digital twin environment can calculate the new environment state (i.e., intermediate environment state) of the digital twin environment based on the original environment state according to physical rules and logic. The calculation of this environment state has many implementation methods in digital twin technology, which will not be elaborated here.

[0055] Understandably, model decision-making involves inputting the obtained intermediate environment state into the model to be trained, and then using the model to make decisions based on this intermediate environment state to obtain the decision strategy output by the model. This decision strategy characterizes the action required by the digital twin environment, as indicated by the model, and the expected state of the digital twin environment corresponding to that action. The specific content of this action can be flexibly set according to the actual situation. For example, for a digital twin environment of a communication network service system, if the intermediate environment state indicates a failure of a communication node in the digital twin environment, the corresponding action could be to instruct that communication node in the digital twin environment to perform a restart operation; or, if the intermediate environment state indicates a high packet loss rate between two adjacent communication nodes in the digital twin environment, the corresponding action could be to instruct that communication node in the digital twin environment to perform a data forwarding route switching operation.

[0056] Step S520: Execute the decision-making strategy through the digital twin environment to determine the target environmental state of the digital twin environment; Step S530: The decision-making strategy and the target environment state are analyzed by the first training strategy engine to obtain multi-dimensional model index data of the training model.

[0057] In this embodiment, after receiving the decision strategy output by the model to be trained, the digital twin environment can execute the decision strategy and calculate the new environmental state of the digital twin environment, i.e., the target environmental state. Then, through the first training strategy engine combined with the indicator data retained in the previous training optimization process, the target environmental state is analyzed based on the expected state of the digital twin environment in the decision strategy, thereby obtaining multi-dimensional model indicator data of the model to be trained. The multi-dimensional model indicator data includes at least one of accuracy, latency, robustness, fairness, and accuracy change trend.

[0058] Step S140: Based on the multidimensional model index data, optimize the rewards of the first coaching strategy engine to obtain the second coaching strategy engine.

[0059] In this embodiment, reward optimization can be based on optimizing the training strategy engine using multidimensional model indicator data. Specifically, taking accuracy and robustness as examples of multidimensional model indicator data, reward optimization can involve assigning weights to accuracy and robustness respectively, and determining a reward optimization value by synthesizing all weights. The specific synthesis method can be summation, averaging, or other operations. Then, the reward optimization value is fed back to the training strategy engine so that the training strategy engine can achieve optimization based on reinforcement learning technology and the reward optimization value.

[0060] Understandably, in one optional implementation, the accuracy metric and its weight can have a negative correlation; that is, the lower the accuracy of the model to be trained, the greater the weight of the accuracy metric. Specifically, if the accuracy of the model to be trained is lower (i.e., the model's performance is poorer), it indicates that the model is more likely to face dynamic and unknown challenges in the digital twin environment, or that the current environment faced by the model is its weak point; in this case, a larger weight can be determined through the negative correlation. The robustness metric and its weight can also have a negative correlation, and its content is similar to that of the aforementioned accuracy metric, which can be easily deduced by analogy.

[0061] Based on the negative correlation between robustness and accuracy metrics and their weights, if the accuracy and robustness of the decision strategy output by the model to be trained in the current training optimization process are low, the training strategy engine can determine a larger reward optimization value. Combined with reinforcement learning techniques, this reward optimization value can encourage the training strategy engine to generate training action instructions corresponding to the current training optimization process in subsequent training optimization processes. This allows the model to be trained to target weak links or dynamic unknown challenge environments, which is beneficial to increasing the training volume of the model in the "knowledge blind spot" and thus effectively improving the model training effect and personalization.

[0062] It should be noted that, in practical applications, the second training measurement engine obtained in step S140 can optimize the training model and process its input data to obtain data processing results. Specifically, in a communication network service system, the training model can be a communication network fault detection model. This model is used to detect existing faults and / or potential faults in the future. The input data can be data such as channel delay, channel bandwidth, signal strength, and data transmission volume of various nodes in the communication network. The communication network fault detection model performs fault detection on the input data to obtain the fault detection results of the communication network.

[0063] In cloud computing business systems, the training model can be a resource scheduling model, which is used to make decisions on the tasks that each node in the cloud computing cluster needs to execute. Its corresponding input data can be the tasks being processed by each computing node in the cloud computing cluster, the load resource data of each computing node (such as CPU utilization, GPU utilization), and the remaining available storage space. After making scheduling decisions on the input data, the data processing result obtained by the resource model can be the resource scheduling result of each computing node in the cloud computing cluster.

[0064] Reference Figure 8 In some embodiments, the method further includes: Step S810: Obtain the decision strategy of the model to be trained; Step S820: Perform feature importance analysis on the decision-making strategy to obtain first analysis information; Step S830: Perform interpretability analysis on the decision-making strategy to obtain second analysis information; Step S840: Based on the first analysis information and the second analysis information, perform root cause localization processing to obtain the defect root cause information of the model to be trained.

[0065] In the embodiments of this application, feature importance analysis can be based on SHAP (SHapley Additive ex Planations) to perform feature importance analysis on the decision strategy to obtain first analysis information; while interpretability analysis can be based on LIME (Local Interpretable Model-agnostic Explanations) technology to perform interpretability analysis on the decision strategy to obtain second analysis information.

[0066] It is understandable that root cause localization processing can be based on the Root Cause Analysis (RCA) method to locate the root cause of the first analysis information and the second analysis information, thereby obtaining the root cause information of the defects of the model to be trained.

[0067] Please see Figure 9 This application also provides an optimization system for a coaching strategy engine, comprising: The first processing unit 901 is used to obtain the first training strategy engine and the training model to be trained in the current training optimization process, as well as business training data. The first training strategy engine is the original training strategy engine or the second training strategy engine obtained in the previous training optimization process. The second processing unit 902 is used to construct a digital twin environment based on the business training data; The third processing unit 903 is used to perform simulated environment training on the model to be trained based on the digital twin environment and the first training strategy engine, and obtain multi-dimensional model index data of the model to be trained. The fourth processing unit 904 is used to optimize the first coaching strategy engine for rewards based on the multidimensional model index data to obtain the second coaching strategy engine.

[0068] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0069] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0070] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0071] Please see Figure 10 , Figure 10 This illustration depicts the hardware structure of an electronic device according to one embodiment. The electronic device includes: The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1002 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001. Input / output interface 1003 is used to implement information input and output; The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004); The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.

[0072] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0073] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0074] This application also discloses a computer program product or computer program, which includes computer instructions stored in the aforementioned computer-readable storage medium; the processor of the aforementioned electronic device can read the computer instructions from the aforementioned computer-readable storage medium, and the processor executes the computer instructions, causing the electronic device to perform the aforementioned method embodiment.

[0075] It is understood that the content of the above method embodiments is applicable to this computer program product or computer program embodiment. The specific functions implemented by this computer program product or computer program embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0076] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0077] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0078] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0079] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0080] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0081] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0082] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0083] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0084] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0085] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0086] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0087] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. An optimization method for a coaching strategy engine, characterized in that, include: Obtain the first training strategy engine and the training model to be trained in the current training optimization process, as well as the business training data. The first training strategy engine is either the original training strategy engine or the second training strategy engine obtained in the previous training optimization process. Based on the aforementioned business training data, construct a digital twin environment; Based on the digital twin environment and the first training strategy engine, the model to be trained is trained in a simulated environment to obtain multi-dimensional model index data of the model to be trained. Based on the multidimensional model index data, the first coaching strategy engine is optimized for rewards to obtain the second coaching strategy engine.

2. The method according to claim 1, characterized in that, Obtain business training data, including: Acquire multi-source business data from the business system; Multimodal feature extraction is performed on the multi-source business data to obtain multimodal feature data; The multimodal feature data is processed to generate synthetic data, thereby obtaining the business training data.

3. The method according to claim 2, characterized in that, The process of generating synthetic data from the multimodal feature data to obtain the business training data includes: Obtain the root cause information of the defects of the model to be trained in the current training optimization process. The root cause information is used to characterize the reasons for the model performance defects of the model to be trained in the previous training optimization process. Based on the defect root cause information, the multimodal feature data is filtered to obtain few-sample feature data and adversarial example feature data; The few-sample feature data are used to synthesize samples to obtain synthesized sample data; Adversarial perturbations are added to the adversarial sample feature data to obtain perturbation sample data; Based on the synthetic sample data and the perturbation sample data, the multimodal feature data is fused to obtain the business training data.

4. The method according to claim 1, characterized in that, The step of constructing a digital twin environment based on the business training data includes: Obtain logical topology data from the business system and construct a hardware resource model for the business system; The business training data is input into the hardware resource model, and the hardware resource model is used to simulate the environment of the business training data to obtain the digital twin environment of the business system.

5. The method according to claim 1, characterized in that, The step of conducting simulated environment training on the model to be trained based on the digital twin environment and the first training strategy engine to obtain multi-dimensional model index data of the model to be trained includes: Obtain the decision strategy output by the model to be trained during the current training optimization process; The decision-making strategy is executed through the digital twin environment to determine the target environmental state of the digital twin environment; The first training strategy engine performs index analysis on the decision strategy and the target environment state to obtain multi-dimensional model index data of the training model.

6. The method according to claim 5, characterized in that, The step of obtaining the decision strategy output by the model to be trained during the current training optimization process includes: The training strategy engine generates training action instructions. Based on the training action instructions, the state of the digital twin environment is updated to obtain an intermediate environment state; The intermediate environment state is input into the training model to make model decisions and obtain the decision strategy.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Obtain the decision-making strategy of the model to be trained; The decision-making strategy is subjected to feature importance analysis to obtain first analytical information; The decision-making strategy is subjected to interpretability analysis to obtain second analytical information; Based on the first and second analysis information, root cause localization processing is performed to obtain the defect root cause information of the model to be trained.

8. An optimization system for a coaching strategy engine, characterized in that, include: The first processing unit is used to obtain the first training strategy engine and the training model to be trained in the current training optimization process, as well as business training data. The first training strategy engine is the original training strategy engine or the second training strategy engine obtained in the previous training optimization process. The second processing unit is used to construct a digital twin environment based on the business training data; The third processing unit is used to conduct simulated environment training on the model to be trained based on the digital twin environment and the first training strategy engine, and obtain multi-dimensional model index data of the model to be trained. The fourth processing unit is used to optimize the rewards of the first coaching strategy engine based on the multidimensional model index data to obtain the second coaching strategy engine.

9. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method as described in any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.