A multi-model meta-evolution search method and device for expert-oriented load balancing

CN122816894APending Publication Date: 2026-09-25NINGXIA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611061149.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0010]有鉴于此,本申请提供一种面向专家负载均衡的多模型元进化搜索方法及装置,以解决背景技术中静态专家放置、基于专家热度的贪心复制、基于设备负载的均衡分配以及人工阈值局部调整方法难以协同利用专家激活频率矩阵、token路由计数数据、设备负载向量、设备拓扑关系和跨设备通信开销矩阵的问题,并解决现有自动程序优化方法中固定搜索策略容易停滞、单一生成模型稳定性不足、候选专家映射程序结构错误率较高以及真实推理评估成本较大的问题

Benefits of technology

[0021]与现有技术相比,本申请的有益效果在于:本申请在搜索方法优化过程中采用双层优化结构,内层优化候选程序;外层优化搜索策略。同时维护候选程序数据库、模型权重表和搜索策略数据库,使候选程序生成、模型选择和搜索策略选择形成闭环反馈。通过根据候选程序的实际得分和静态预检结果动态调整大语言模型选择权重,可以提高多模型调用的稳定性和有效性。通过在停滞时基于种群状态描述符更新搜索策略,可以避免固定搜索策略在优化后期陷入局部停滞。通过在真实评估之前执行面向EPLB的静态预检,可以提前拦截索引越界、张量维度错误、副本计数非法和映射不一致等候选程序,减少无效评估次数。通过将错误类型转换为下一轮提示约束,可以持续降低同类错误重复出现的概率,从而提高自动发现过程中的搜索效率和最终优化性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816894A_ABST
    Figure CN122816894A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence model deployment and inference acceleration, in particular to a multi-model meta-evolution search method and device for expert load balancing. The method comprises the following steps: obtaining a to-be-optimized task, an initial program, an evaluation standard, a model library and a strategy set; establishing a candidate program database, a model weight table and a search strategy database; calculating a population state descriptor according to the candidate program database, and selecting a target large language model and a target search strategy; constructing prompt information from the target search strategy, calling the target large language model to generate a new candidate program; performing static pre-check before task evaluation; when the pre-check passes, calling a task evaluator and updating the related database and weight table, and when the pre-check fails, generating constraint feedback and reducing the selection weight of the corresponding model; updating the search strategy when the window promotion amount is lower than a stagnation threshold value, and outputting a final candidate program until the final candidate program is output. The application can reduce the proportion of invalid candidate programs, improve the search efficiency and optimization stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence model deployment and inference acceleration technology, specifically involving expert parallel load balancing technology for Mixture of Experts (MoE) in multiple graphics processors, neural network accelerators or heterogeneous computing clusters, and particularly involving a multi-model meta-evolutionary search method and apparatus for expert load balancing. Background Technology

[0002] In large-scale hybrid expert model inference services, input text, images, or multimodal data, after passing through routers, will form activation distributions across multiple logical experts at different network layers of the model. This activation distribution can typically be represented as a layer-wise statistical matrix of expert activation frequencies, an expert heat vector, or token-to-expert routing counts. The deployment environment usually includes multiple graphics processors, neural network accelerators, or heterogeneous computing nodes, and corresponds to device load vectors, device topology relationships, cross-card communication overhead matrices, and the number of available physical expert slots at each layer. The goal of the expert parallel load balancing problem is to determine, based on the aforementioned domain data, the number of copies each logical expert should replicate and the placement of each copy in the physical expert slots, thereby forming the expert replica count matrix logcnt, the logical expert-to-physical expert mapping tensor log2phy, and the physical expert-to-logical expert mapping tensor phy2log.

[0003] In existing hybrid expert model systems, common practices for parallel load balancing of experts mainly include static expert placement, greedy replication based on expert popularity, balanced allocation based on device load, and local adjustment based on manual thresholds. Static expert placement typically predetermines the correspondence between experts and devices based on historical sample statistics before model deployment; the greedy replication method based on expert popularity usually assigns more physical expert copies to logical experts with higher activation frequencies; the balanced allocation method based on device load typically distributes expert copies as widely as possible across different devices based on the current computing load of each graphics processor or accelerator card; and the local adjustment method based on manual thresholds manually adds copies or adjusts physical slots when the load of certain experts or devices exceeds a threshold.

[0004] With the rapid development of Mixture of Experts (MoE) models and large-scale deep learning systems, how to place experts, replicate experts, and load balance among multiple layers, multiple experts, multiple graphics processors, or other heterogeneous computing resources has become a crucial factor affecting model inference throughput and service stability. In Expert Parallelism Load Balancing (EPLB) tasks, the system typically needs to generate mapping relationships from logical experts to physical experts and vice versa, based on the activation frequency of logical experts at each layer, device load, communication overhead, and the number of physical expert slots.

[0005] Traditional expert placement methods often employ manually designed heuristic rules, such as sorting experts by popularity, greedily allocating them based on equipment load, replicating high-frequency experts at a fixed ratio, or making local adjustments based on preset thresholds. These methods can be effective when the problem size is small or the load distribution is relatively stable. However, when faced with scenarios involving dynamically changing expert activation distribution, significant load differences between different layers, limited physical expert slots, and interdependent multi-objective evaluation metrics, problems such as load skew, excessive concentration of hot experts, or mismatched expert replication numbers can easily arise.

[0006] In recent years, large language model-driven program optimization methods have been used to automatically generate heuristic algorithms, optimize code, and develop search strategies. These methods typically maintain a set of evaluated candidate programs. A search strategy selects parent programs, constructs hints, and calls a large language model to generate new candidate programs. A task evaluator then calculates the performance scores of these candidate programs. However, existing methods often employ fixed search strategies. The parent selection rules, exploration utilization ratios, mutation operators, and heuristic sample selection methods within the search strategy remain unchanged throughout the optimization process. When the search space changes or the candidate program population gradually converges, fixed strategies are prone to stagnation, requiring manual parameter retuning to achieve further effective improvements.

[0007] Furthermore, programs generated by large language models are inherently uncertain. For EPLB-type programs, candidate programs must not only achieve high load balancing scores but also meet strict structural validity requirements. For example, the sum of the number of replicas in each layer of the expert replica count matrix must match the number of physical experts; the mapping tensors from logical experts to physical experts and from physical experts to logical experts must be consistent; physical expert number slices must not exceed boundaries; and tensor index dimensions and broadcast relationships must be valid. If candidate programs are not effectively checked before entering the real evaluator, they are prone to evaluation failures due to syntax errors, dimension errors, index out-of-bounds errors, and inconsistent mappings, wasting significant generation and evaluation costs.

[0008] Existing automated program optimization systems generally rely on a single large language model or a fixed model invocation order. Different large language models vary in their capabilities in code generation, local repair, structural rewriting, constraint compliance, and handling long contexts. A single model is prone to instability in performance for specific error types or at specific optimization stages. If the model selection probability cannot be dynamically adjusted based on the actual evaluation performance of candidate programs and static pre-detection results, it is difficult to fully utilize the complementary capabilities of multiple models.

[0009] Therefore, an automatic program optimization scheme is needed that can simultaneously adaptively select a large language model, dynamically update the search strategy, and perform static pre-examination for the EPLB task before evaluation, in order to reduce the proportion of invalid candidate programs and improve the stability, search efficiency, and quality of the final candidate programs in the optimization process. Summary of the Invention

[0010] In view of this, this application provides a multi-model meta-evolutionary search method and apparatus for expert load balancing, to solve the problems in the background art of static expert placement, greedy replication based on expert popularity, balanced distribution based on device load, and the difficulty in synergistically utilizing expert activation frequency matrix, token routing count data, device load vector, device topology relationship, and cross-device communication overhead matrix. It also solves the problems in existing automatic program optimization methods such as the tendency of fixed search strategies to stagnate, insufficient stability of single generative models, high error rate of candidate expert mapping program structure, and high cost of real inference evaluation.

[0011] A multi-model meta-evolutionary search method for expert load balancing includes: Obtain the expert parallel load balancing task to be optimized, the task evaluator, the static pre-detection rules, multiple candidate large language models, and the search strategy set; A candidate program database is established, which is used to record candidate programs, their evaluation scores, execution logs, parent sources, evolution methods, and static pre-detection results. Establish a model weight table and a search strategy database. The model weight table is used to record the selection weights of each candidate large language model, and the search strategy database is used to record search strategies and their historical performance. The population state descriptor is calculated based on the candidate program database. The population state descriptor includes the optimal score, score distribution, recent window boost, parent selection frequency, and static pre-detection failure type distribution. The target large language model is selected from multiple candidate large language models according to the model weight table, and the target search strategy is selected from the search strategy set according to the search strategy database; The target search strategy is used to select parent programs, evolutionary methods, and heuristic sample sets from the candidate program database, and to construct prompt information for code generation. New candidate programs are generated based on the prompt information using the target large language model, and a static pre-check oriented towards expert parallel load balancing is performed on the new candidate programs before the task evaluation is executed. When the static pre-detection passes, the task evaluator is invoked to obtain the evaluation score of the new candidate program, and the candidate program database, model weight table, and search strategy database are updated based on the evaluation score; when the static pre-detection fails, constraint feedback is generated according to the failure type and the selection weight of the corresponding large language model or search strategy is reduced. When the recent increase in the window is lower than the preset stagnation threshold, the search strategy is generated or updated based on the search strategy database and the current population state descriptor until the preset termination condition is met and the final candidate program is output.

[0012] Preferably, the selection of the target large language model from multiple candidate large language models includes: sampling according to the selection weight of each candidate large language model using a roulette wheel selection method, and updating the selection weight of the corresponding candidate large language model according to the score change of the candidate program relative to the parent program or the current best baseline after the candidate program has completed static pre-detection and task evaluation.

[0013] Preferably, updating the selection weights of the corresponding candidate large language model includes: applying a positive reward to the candidate large language model that generated the new candidate program when the evaluation score of the new candidate program is higher than the score of the parent program or the current best baseline score; and applying a negative penalty to the candidate large language model that generated the new candidate program when the new candidate program has a syntax error, a runtime error, a static preflight failure, or an evaluation score lower than the score of the parent program.

[0014] Preferably, the search strategy set includes genetic algorithm strategy, differential evolution strategy, and particle swarm optimization strategy; the mutation operator includes random strategy, global refinement strategy, and local repair strategy.

[0015] Preferably, the historical performance in the search strategy database is calculated based on the optimal score improvement within a monitoring window, the optimal score at the beginning of the window, and the window length, and is used to bias the selection of search strategies.

[0016] Preferably, generating or updating a search strategy based on the search strategy database and the current population state descriptor includes: selecting a parent strategy from a set of search strategies whose historical performance meets preset conditions; mutating at least one of the parent selection rule, heuristic sample construction rule, and search utilization ratio of the parent strategy according to the current population state descriptor to obtain a candidate search strategy; performing a validity verification on the candidate search strategy; if the verification passes, using it as a new target search strategy; if the verification fails, retaining the original target search strategy or reselecting a search strategy from the set of search strategies.

[0017] Preferably, the static pre-check for expert parallel load balancing includes: checking whether the number of expert replicas at each layer is consistent with the number of physical experts; checking whether the physical expert ID slice is non-empty and not out of bounds; checking whether the number of physical expert IDs is consistent with the number of corresponding logical expert replicas; checking whether the mapping tensor from logical experts to physical experts and the mapping tensor from physical experts to logical experts have the same source; and checking whether the tensor index dimension and broadcast relationship are valid; when it is detected that a two-dimensional index is directly used for one-dimensional layer index update, a two-dimensional replica count tensor is directly passed to the repeated expansion function, the physical expert ID slice length is insufficient, or the number of mapping tensors written is inconsistent with the number of expert replicas, the new candidate program is rejected from entering the task evaluation, and the corresponding error type is added to the constraint feedback.

[0018] Preferably, the constraint feedback is written into the prompt information of the next round of code generation, including: adding the error type, error location, corresponding repair constraints, and error patterns that should not be repeated obtained from the static pre-detection to the prompt information of the next round of code generation, so as to constrain the new candidate program generated in the next round to meet the preset structural legality requirements in the process of constructing the expert replica count matrix, selecting the expert index, slicing the physical expert number, and writing the mapping tensor; wherein, the new candidate program is generated using at least one of the following methods: layer-by-layer explicit loop, one-dimensional expert index, boundary assertion, or mapping consistency assertion.

[0019] Preferably, the method model framework is a two-layer evolutionary structure model: outer layer model: optimizes the search strategy; inner layer model: optimizes the candidate program.

[0020] A multi-model meta-evolutionary search device for expert load balancing includes: A two-layer evolutionary structure, consisting of an inner layer of evolutionary candidate programs and an outer layer of evolutionary search strategies; The initialization module is used to obtain the expert load balancing task to be optimized, the task evaluator, static pre-detection rules, multiple candidate large language models, and a set of search strategies. The candidate program database module is used to record candidate programs, their evaluation scores, execution logs, parent sources, mutation operators, and static pre-detection results. The model selection module is used to select the target large language model from multiple candidate large language models according to the model weight table, and update the model weights according to the static pre-detection results and evaluation scores of the candidate programs. The search strategy management module is used to establish a search strategy database and select, generate, or update search strategies based on population state descriptors. The hint construction module is used to select the parent program, evolutionary method and heuristic sample set according to the target search strategy, and construct hint information for code generation; The program generation module is used to generate new candidate programs based on the target large language model; The static pre-inspection module is used to perform checks on the new candidate program before task evaluation, including consistency of expert replica count, physical expert number boundaries, consistency of mapping tensors, and validity of tensor index dimensions. The evaluation and update module is used to call the task evaluator to obtain the evaluation score and update the candidate program database, model weight table and search strategy database when the static pre-detection passes. When the static pre-detection fails, it generates constraint feedback and reduces the selection weight of the corresponding large language model or search strategy.

[0021] Compared with existing technologies, the advantages of this application are as follows: This application employs a two-layer optimization structure in the search method optimization process: the inner layer optimizes candidate programs, and the outer layer optimizes the search strategy. Simultaneously, it maintains a candidate program database, a model weight table, and a search strategy database, creating a closed-loop feedback between candidate program generation, model selection, and search strategy selection. By dynamically adjusting the selection weights of large language models based on the actual scores of candidate programs and static pre-detection results, the stability and effectiveness of multi-model calls can be improved. By updating the search strategy based on the population state descriptor when stagnation occurs, the fixed search strategy can avoid local stagnation in the later stages of optimization. By performing static pre-detection for EPLB before the actual evaluation, candidate programs with issues such as index out-of-bounds errors, tensor dimension errors, illegal replica counts, and inconsistent mappings can be intercepted in advance, reducing the number of invalid evaluations. By converting error types into hint constraints for the next round, the probability of recurring similar errors can be continuously reduced, thereby improving the search efficiency and final optimization performance during the automatic discovery process. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating a multi-model meta-evolutionary search method for expert load balancing provided in the first embodiment of this application.

[0023] Figure 2 This is a schematic diagram (inner layer) of a multi-model roulette wheel selection and weight update process provided in the first embodiment of this application.

[0024] Figure 3This is a schematic diagram (outer layer) of a search strategy database and strategy update process provided in the first embodiment of this application.

[0025] Figure 4 This is a schematic diagram of a static pre-inspection process for expert-oriented parallel load balancing provided in the first embodiment of this application.

[0026] Figure 5 This is a schematic diagram illustrating the feedback relationship between a candidate program database, a model weight table, and a search strategy database, as provided in the first embodiment of this application.

[0027] Figure 6 This is a schematic diagram of the structure of a multi-model meta-evolutionary search method for expert-oriented parallel load balancing provided in the second embodiment of this application.

[0028] Figure 7 The second embodiment of this application provides a schematic diagram of the structure of a multi-model meta-evolutionary search device for expert load balancing. Detailed Implementation

[0029] To more clearly illustrate the technical solutions of the embodiments of this application, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the scope of protection of this application.

[0030] In this application, the candidate program can be executable code, function, script, configuration generation logic, or prompt template for solving the target optimization task; the task evaluator is used to run the candidate program and return the evaluation result; the large language model is used to generate, repair, or rewrite the candidate program based on the prompt information; the search strategy is used to determine the parent program selection method, the heuristic sample selection method, the mutation operator type, and the exploration utilization ratio; meta-evolution refers to not only evolving the candidate program itself, but also evolving the search strategy used to generate the candidate program.

[0031] In the EPLB task, logical experts can be understood as expert units to be scheduled in the model structure, and physical experts can be understood as expert replicas deployed on computing devices or physical slots. Candidate programs typically need to output three types of data structures: an expert replica count matrix `logcnt`, a logical expert-to-physical expert mapping tensor `log2phy`, and a physical expert-to-logical expert mapping tensor `phy2log`. Here, `logcnt` represents the number of replicas of each logical expert in each layer, `log2phy` represents the set of physical expert IDs corresponding to each logical expert, and `phy2log` represents the logical expert ID corresponding to each physical expert slot.

[0032] Please refer to Figure 1 The first embodiment of this application provides a multi-model meta-evolutionary search method for expert load balancing, including the following steps.

[0033] S101: Obtain the expert parallel load balancing task to be optimized, initial candidate programs, task evaluator, static pre-detection rules, multiple candidate large language models, and a set of search strategies.

[0034] Specifically, the expert parallel load balancing task can include the number of model layers, the number of logical experts per layer, the number of physical experts, the number of devices, expert activation frequency, device load, communication constraints, and evaluation metrics. Initial candidate programs can be manually written baseline algorithms or simple heuristic algorithms generated from a large language model, such as uniform replication, replication based on expert popularity, or greedy placement based on device load. The task evaluator executes the candidate programs and returns a comprehensive score, inference time, load balancing metrics, error logs, and other auxiliary information.

[0035] S102: Establish a candidate program database, a model weight table, and a search strategy database.

[0036] The candidate program database stores generated or evaluated candidate programs. Each record may include the candidate program text, program identifier, the large language model used to generate the candidate program, parent program identifier, search strategy identifier, mutation operator, heuristic sample set, static pre-detection results, runtime log, evaluation score, whether it has entered the current best set, and generation time. The candidate program database provides a source of candidates for subsequent parent selection and also provides a data foundation for population state descriptor calculation.

[0037] The model weight table records the selection weights of multiple candidate large language models. Initially, the selection weights of each candidate large language model can be equal, or they can be pre-set based on historical task performance. The search strategy database records individual search strategies and their historical performance. Each individual search strategy can include parent selection rules, heuristic sample construction rules, mutation operator preferences, exploration utilization ratio, error repair preferences, and strategy validity constraints.

[0038] S103: Calculate the population state descriptor based on the candidate program database, select the target large language model based on the model weight table, and select the target search strategy based on the search strategy database.

[0039] The population state descriptor is used to summarize the current stage of the automatic optimization process. The population state descriptor may include the current best score, mean score, median score, score variance, score changes in the most recent few generated iterations, improvement within the most recent window, iterations since the last significant improvement, frequency of parent program selection, frequency of mutation operator usage, distribution of static preflight failure types, proportion of syntax errors, proportion of runtime errors, and similarity between high-scoring candidate programs.

[0040] Please refer to Figure 2 A roulette wheel selection method is used to select the target large language model. Let w be the selection weight of the i-th candidate large language model. i Then the probability of this candidate large language model being selected is w. i Divide by the sum of the weights of all candidate large language models. This approach preserves the preference for high-performance models while giving other models some opportunity to explore, thus preventing the system from becoming overly reliant on a single model too early.

[0041] S104: Select parent programs, mutation operators, and heuristic sample sets from the candidate program database using a target search strategy, and construct hints for code generation.

[0042] The parent program can be selected from the current best candidate program, high-scoring candidate programs, candidate programs with significant recent improvements, or candidate programs with a large structural difference from the current best program. Mutation operators can include local refinement, structural mutation, free mutation, error correction, parameter tuning, and code refactoring. The heuristic sample set can include several high-scoring candidate programs, complementary candidate programs, recent failure cases and their reasons for failure. The prompt information can also include task descriptions, interface signatures, input / output constraints, static preflight rules, historical error types, and the expected mutation direction.

[0043] S105: Generate new candidate programs based on prompt information using the target large language model.

[0044] The target large language model generates new candidate programs based on the prompts. Generation can be a complete rewrite, or it can involve partial replacement of the parent program, repair of specific functions, rewriting of search strategy parameters, or restructuring of mapping logic. To ensure that candidate programs can be automatically extracted, the prompts can require the candidate programs to be output in a preset code block format, including specified function names, parameter names, and return values.

[0045] S106: Perform a static pre-check of the new candidate program for expert-oriented parallel load balancing before performing the task evaluation.

[0046] Static preflight checks are used to determine whether a candidate program has obvious structural errors before actual execution, based on the code text, abstract syntax tree, simple symbolic execution results, or small-scale sample execution results. Compared to directly entering task evaluation, static preflight checks can discover issues such as illegal copy counts, dimension errors, index out-of-bounds errors, and mapping inconsistencies in candidate programs at a lower cost.

[0047] S107: When the static pre-detection passes, call the task evaluator to obtain the evaluation score of the new candidate program, and update the candidate program database, model weight table and search strategy database.

[0048] After passing the static pre-detection, the task evaluator executes the candidate program and returns a comprehensive score. The comprehensive score can be obtained by combining at least one of the following metrics: GPU-level load balancing score, expert-level load balancing score, average inference time, operational stability, and resource utilization. The system determines the reward value based on the improvement of the new candidate program relative to the parent program or the current best baseline, and updates the weights of the large language model that generated the candidate program and the historical performance of the corresponding search strategy accordingly.

[0049] S108: When the static pre-detection fails, generate constraint feedback based on the failure type and reduce the corresponding large language model.

[0050] When the static preflight check fails, the system does not invoke the high-cost task evaluator. Instead, it categorizes the failure reasons into types such as syntax error, function signature error, illegal replica count, physical number out of bounds, insufficient mapping tensor capacity, tensor broadcast error, mapping inconsistency, or runtime interface error. The corresponding error type is written to the candidate program database and can be used as constraint feedback in the next round of prompts.

[0051] S109: When the recent increase in the population size falls below a preset stagnation threshold, generate or update the search strategy based on the search strategy database and the current population state descriptor. Please refer to [link / reference]. Figure 3 .

[0052] The system monitors search progress using a preset window length W. If the optimal score at the start of the window is sstart and the optimal score at the end of the window is send, then the improvement within the most recent window is the difference between send and sstart. When the improvement is less than a preset stagnation threshold, the system triggers a search strategy update. During the update, a parent strategy can be selected from strategies with high historical performance, and combined with the current population state descriptor, the parent selection rule, heuristic sample construction rule, mutation operator preference, and exploration utilization ratio are mutated to obtain candidate search strategies.

[0053] S110: Repeat the above process until the preset termination condition is met and the final candidate program is output.

[0054] Preset termination conditions may include reaching the maximum number of iterations, reaching the maximum generation cost, failing to improve within multiple consecutive windows, reaching the target score threshold, or the user voluntarily stopping. The final candidate program can be the program with the highest evaluation score in the candidate program database, or the program with the best overall stability and performance selected after ranking by multiple indicators.

[0055] The static pre-detection rules in S106 are explained in further detail below. Please refer to [link / reference]. Figure 4 .

[0056] S201: Check the validity of the expert replica count matrix logcnt. logcnt should be a two-dimensional integer matrix whose shape matches the number of layers and the number of logical experts. For each layer, the replica count of each logical expert should be greater than or equal to one, and the sum of the replica counts of all logical experts in that layer should equal the number of available physical experts in that layer. If it is detected that the candidate program updates logcnt by subtraction, directly overwrites logcnt with the rounded result, uses an allocation method that does not guarantee that each logical expert has at least one replica, or that the sum of the replica counts in each layer is not equal to the number of physical experts, then the candidate program is rejected from entering the task evaluation.

[0057] S202: Check the capacity of the logic expert-to-physical expert mapping tensor log2phy. Since different logic experts may have different numbers of replicas, the third-dimensional capacity of log2phy should not be less than the maximum number of replicas in logcnt. In one implementation, the maximum number of replicas maxlogcnt is first calculated based on logcnt, and then maxlogcnt is used as the capacity of the third dimension of log2phy. If the candidate program has a fixed capacity hardcoded, and this capacity may be less than the actual number of replicas for some logic experts, then the static preflight check determines that the capacity is insufficient.

[0058] S203: Check the boundaries and number of physics expert number slices. For each layer and each logical expert, let the number of copies of that logical expert be count, the start position of the physics expert number slice be pos, and the end position be end. end should be equal to the sum of pos and count. The preflight check should confirm that count is greater than zero, end does not exceed the number of physics experts, and the number of physics expert numbers obtained from the slice is equal to count. If the candidate program forcibly truncates end through the min function, resulting in an insufficient number of slices but still writing to the mapping tensor, it is considered illegal.

[0059] S204: Check the consistency between phy2log and log2phy. For the same logical expert, the set of physical expert IDs written to phy2log should be the same as the set of physical expert IDs written to log2phy. If the candidate program uses ID sets from different sources in the two mapping tensors, or if the order of writing causes the mapping relationship to be non-reverse, then the mapping is considered inconsistent.

[0060] S205: Check that each layer of physics expert slots is completely and uniquely covered. For the same layer, all physics expert numbers should be assigned exactly once. Pre-checking can be done by checking whether the allocation cursor ultimately equals the number of physics experts, whether there are unassigned slots in phy2log, and whether there are duplicate covers. If there are missed allocations, over-allocations, or duplicate allocations, the candidate program is rejected.

[0061] S206: Check the validity of tensor index dimensions and broadcast relationships. If the candidate program directly uses the two-dimensional topk index together with the one-dimensional layer index to update the two-dimensional logcnt, or directly uses the two-dimensional logcnt as the repeats parameter of the repeat expansion function, it is prone to broadcast errors or dimension mismatches. For scenarios where only one expert is selected per layer, the candidate program should use a one-dimensional expert index; for scenarios where multiple experts are selected per layer, the layer index and expert index should be explicitly expanded.

[0062] S207: Check the interface, syntax, and runtime environment constraints of the candidate program. The candidate program should contain predefined function signatures, and the number and type of return values ​​should be consistent with the task evaluator requirements. Violations of import statements, function definitions, indentation structures, and return statements are prohibited. For scenarios where only specific libraries or device tensors are allowed, the pre-check can also check whether the candidate program calls unauthorized libraries, creates tensors on the wrong device, and contains obviously inefficient or non-terminating loops.

[0063] The following section provides further explanation of the multi-model weight updates in S104 and S108. Please refer to [link / reference needed]. Figure 5 .

[0064] In one implementation, each candidate large language model in the model weight table has equal initial weights. If the new candidate program generated by the i-th candidate large language model passes the static pre-detection and obtains an evaluation score snew, the parent program score is spar, and the current best baseline score is sbest, then the reward value can be determined based on the difference between snew and max(spar, sbest). If snew is higher than the parent program or the current best baseline, the corresponding model receives a positive reward; if the candidate program fails the static pre-detection or fails to run, a negative penalty is given according to the severity of the failure.

[0065] In this application, the reward value refers to a numerical feedback quantity calculated based on the static pre-detection results, running results, and changes in evaluation scores of the new candidate program. This reward value is used to adjust the selection weight of the large language model that generated the new candidate program in subsequent roulette wheel selections. This reward value is not limited to environmental rewards in the field of reinforcement learning, but rather serves as a weight update feedback signal for multi-model adaptive selection. Specifically, the reward value can be positive, zero, or negative. When the new candidate program improves upon its parent program or the current best baseline, the reward value is positive, used to increase the subsequent selection probability of the corresponding large language model. When the new candidate program fails the static pre-detection, runs unsuccessfully, or its evaluation score drops significantly, the reward value is negative, used to decrease the subsequent selection probability of the corresponding large language model.

[0066] The current best baseline score refers to the highest evaluation score of the candidate program in the candidate program database that has passed the static pre-test and completed the task evaluation before the current round of candidate program generation and evaluation. This current best baseline score is dynamically updated during the iteration process. Specifically, if a new candidate program passes the static pre-test and runs successfully, and its evaluation score snew is higher than the current best baseline score sbest, then the current best baseline score is updated to snew, and the new candidate program is marked as the current best candidate program; if snew does not exceed sbest, the current best baseline score remains unchanged; if the new candidate program fails the static pre-test or fails to run, the current best baseline score is not updated.

[0067] In one specific implementation, let the parent program score be `spar`, the current best baseline score be `sbest`, and the new candidate program score be `snew`. Then, let the comparison benchmark be `b = max(spar, sbest)`, and determine the reward value based on the difference `Δ = snew - b`. When `Δ` is greater than zero, it indicates that the new candidate program is superior to both the parent program and the current best baseline, corresponding to a strong positive reward for the large language model. When `snew` is higher than `spar` but not higher than `sbest`, it indicates that the new candidate program has local improvements over the parent program but has not yet refreshed the global optimum, corresponding to a weak positive reward or neutral reward for the large language model. When `snew` is lower than `spar`, the large language model receives a negative penalty. For cases such as failing static pre-detection, syntax errors, function signature errors, tensor dimension errors, out-of-bounds physical expert numbering, or inconsistent mapping relationships, different negative penalty coefficients are set according to the severity of the errors.

[0068] For example, a first penalty coefficient can be applied to candidate programs that cannot enter the evaluator due to syntax errors, corrupted function signatures, or obvious out-of-bounds errors; a second penalty coefficient can be applied to candidate programs that can run but score significantly lower than their parent programs; and a positive reward can be applied to candidate programs whose evaluation scores exceed the current best baseline, increasing the probability of the corresponding model in subsequent roulette wheel selections. After the update, all model weights are normalized so that the sum of the selection probabilities of each model is one.

[0069] By employing the above methods, this application can automatically identify large language models that are better at generating effective candidate programs in the current task and current search phase, and reduce the probability of models that frequently produce similar errors being selected. Meanwhile, low-weight models can still retain the lowest exploration probability, allowing for complementary capabilities in subsequent search phases.

[0070] The following section provides further details on the search strategy database and strategy updates.

[0071] Search strategies can be represented as executable strategy templates, sets of strategy parameters, or sets of hint construction rules. Genetic algorithm strategies can prioritize multiple high-scoring candidate programs as parents and generate hints through code snippet combinations, local repair, and parameter perturbations. Differential evolution strategies can construct mutation instructions that "preserve the advantages of parents and introduce differential fragments" based on structural or performance differences among multiple candidate programs. Particle swarm optimization strategies can maintain a historical best state and a global best state for each strategy, adjusting the parent selection ratio, exploration intensity, and mutation operator probability based on the historical best strategy and the global best strategy.

[0072] The historical performance of a search strategy can be calculated based on the optimal score improvement within a monitoring window. Let the window length be W, the optimal score at the start of the window be sstart, and the optimal score at the end of the window be send. The strategy performance can then be positively correlated with send minus sstart, and normalized by combining the initial window score and the window length. This rewards strategies that still bring improvement at high score stages and avoids incomparable strategy performance due to different window lengths.

[0073] When stagnation is detected, the system selects a parent strategy with good historical performance from the search strategy database, prioritizing strategies that perform well in similar population states as heuristic samples. Subsequently, the system can generate new search strategies using the target large language model, or modify the parent selection rules, the number of heuristic samples, the local refinement ratio, the structural mutation ratio, and the error correction ratio using preset strategy mutation operators. Candidate search strategies are deployed to subsequent windows after passing validity verification; if verification fails, the original strategy is retained or a new evolutionary approach is selected.

[0074] During strategy switching, the candidate program database is not cleared; historical candidate programs, evaluation scores, and error feedback are all retained. This avoids losing high-quality candidate programs that have already been discovered due to strategy switching, while allowing the new search strategy to continue searching based on the existing population.

[0075] The following is a specific implementation example using the EPLB task.

[0076] Suppose the objective task is to optimize the expert parallel load balancing function of a hybrid expert model. This function receives activation statistics, device information, and the number of physical experts at each layer, and outputs three results: phy2log, log2phy, and logcnt. The system first adds a uniformly distributed baseline function as an initial candidate program to the candidate program database and obtains its baseline score through a task evaluator. The system also initializes the weights of multiple candidate large language models, using a set of search strategies including genetic algorithm, differential evolution, and particle swarm optimization. Mutation operators include randomization, greedy refinement, and local repair strategies.

[0077] In the initial iterations, the system can employ random or high-diversity strategies to encourage the large language model to explore different expert replication and physical slot ranking schemes. Once high-scoring candidate programs appear in the candidate program database, the system can gradually increase the selection probability of greedy refinement or local repair strategies, fine-tuning the expert replica count, physical slot ranking, and load estimation formulas for high-scoring candidate programs. If the score no longer improves within the most recent window, and the population state descriptor indicates high structural similarity among candidate programs, the system triggers a strategy update, increasing the proportion of structural mutation or differential evolution strategies to attempt new expert replication structures or load modeling methods.

[0078] After each candidate program is generated, the static preflight module first checks whether the candidate program contains the correct function signature and whether it returns phy2log, log2phy, and logcnt. Subsequently, the static preflight module uses a combination of code text rules, abstract syntax tree checks, and small-scale sample execution to check whether logcnt is constructed starting from a one-matrix matrix and only adding extra copies, whether the sum of logcnt at each level equals the number of physics experts, whether the capacity of log2phy is determined by the maximum value of logcnt, whether the physics number slice has boundary assertions, and whether phy2log and log2phy are written using the same set of numbers.

[0079] If a candidate program is found to have written the physics expert number of a certain layer into an index that exceeds the legal range, for example, if the number of physics experts is R but the number is R or larger, the static pre-detection module will mark the error type as "physics expert number out of bounds", reject the candidate program from entering the task evaluation, and add the constraint "end must be less than or equal to the number of physics experts, and the number of ids must be verified to be equal to count before writing" to the prompt message in the next round.

[0080] If a candidate program is found to be using the two-dimensional topk index directly with the one-dimensional layer index for logcnt updates, the static preflight module will mark the error type as "illegal broadcast of two-dimensional index", reject the candidate program from entering the task evaluation, and add the constraint "when only one expert is selected for each layer, the one-dimensional argmax index should be used; when multiple experts are selected for each layer, layer_ids and expert_ids should be explicitly expanded" to the next round of prompts.

[0081] If a candidate program passes the static pre-test, the task evaluator runs the candidate program and calculates its overall load balancing score. If the candidate program improves the GPU-level load balancing score without significantly increasing inference time, it is added to the current preferred set, and the corresponding large language model and search strategy receive a positive reward. If a candidate program passes the static pre-test but its overall score is lower than its parent program, the system retains its record for diversity analysis but reduces the short-term selection weight of the corresponding model or strategy.

[0082] In one embodiment, the overall score of the parent program is spar, the overall score of the new candidate program is snew, the current best baseline score is sbest, the change in GPU-level load balancing score is Δgpu, and the change rate of average inference time is Δtime. When snew is greater than max(spar, sbest) and Δtime is not greater than a preset time threshold τt, the system applies a first positive reward to the corresponding large language model and search strategy. When snew is greater than spar but not greater than sbest and Δtime is not greater than the preset time threshold τt, the system applies a second positive reward, which is less than the first positive reward. When a candidate program passes the static pre-detection but snew is less than spar, or Δtime is greater than the preset time threshold τt, the system retains the candidate program record for diversity analysis but reduces the short-term selection weight of the corresponding large language model or search strategy.

[0083] For example, the current weight of the i-th large language model in the model weight table is wi, and the current weight of the j-th search strategy in the search strategy database is vj. If the new candidate program generated by the i-th large language model under the j-th search strategy passes the static pre-detection, and the comprehensive score increases from 0.760 of the parent program to 0.785, while the average inference time only increases by 1.5%, which is lower than the preset time threshold of 3%, then the model reward value ri = +0.10, the strategy reward value rj = +0.08, and the weights are updated as follows: wi′=max(wmin,wi·exp(η·ri)); vj′=max(vmin,vj·exp(ηs·rj)).

[0084] Where wmin represents the minimum exploration weight of the large language model, vmin represents the minimum exploration weight of the search strategy, and η and ηs represent the update step size of the model weights and strategy weights, respectively. After the update, all large language model weights and all search strategy weights are normalized to make the sum of the roulette wheel selection probabilities equal to one.

[0085] For example, if a candidate program passes the static preflight check but its overall score decreases from 0.760 (parent program) to 0.720, the system will not consider it the current optimal candidate program, nor will it increase the selection probability of its corresponding model or strategy. Instead, the system will save the candidate program, along with its generating model, search strategy, parent source, sub-scores, and runtime logs, to the candidate program database for subsequent diversity analysis or failure mode analysis. Simultaneously, the model reward value ri = -0.05 and the strategy reward value rj = -0.04 can be set, and the short-term selection weights of the corresponding model and search strategy can be reduced according to the aforementioned exponential update method. If a candidate program fails the static preflight check, for example, due to function signature errors, tensor dimension errors, out-of-bounds physics expert IDs, or inconsistent mapping relationships, a larger negative reward value can be set according to the severity of the error, such as ri = -0.20 or ri = -0.30, thereby more significantly reducing the selection probability of the corresponding large language model in subsequent rounds.

[0086] Positive rewards are not abstract evaluations, but rather direct effects on the selection weights in the model weight table and search strategy database. Reducing short-term selection weights does not mean deleting the corresponding model or search strategy, but rather reducing its probability of being selected in subsequent roulette wheel selections while retaining the lowest exploration probability. This increases the probability of calling large language models and search strategies that have performed well recently, while preserving the opportunity for low-weight models and strategies to play a role again in subsequent search phases.

[0087] Through the above process, the system can maintain a high level of exploration capability in the early stage of the search, perform strategy mutation based on historical effective strategies in the middle stage of the search, and reduce invalid candidate programs through local refinement and error feedback in the later stage of the search, ultimately outputting EPLB candidate programs with high performance and valid structure.

[0088] Please refer to Figure 6 The second embodiment of this application also provides an apparatus 300 for a multi-model meta-evolutionary search method oriented towards expert parallel load balancing. The apparatus may include an initialization module, a candidate program database module, a model selection module, a search strategy management module, a suggestion construction module, a program generation module, a static pre-detection module, and an evaluation and update module.

[0089] The initialization module 301 is used to obtain the EPLB task to be optimized, initial candidate programs, task evaluator, static pre-detection rules, multiple candidate large language models, and a set of search strategies. The initialization module can also determine the maximum number of iterations, monitoring window length, stagnation threshold, minimum exploration probability of the model, static pre-detection failure penalty coefficient, and strategy update frequency according to the task configuration.

[0090] The candidate program database module 302 stores candidate programs and their metadata. The metadata includes candidate program scores, static pre-detection results, execution logs, parent source, generative model, search strategy, mutation operator, heuristic sample set, and error type. The candidate program database module can also provide functions for sorting by score, filtering by error type, searching by structural similarity, and tracing back by parent relationship.

[0091] The model selection module 303 is used to maintain the model weight table and select the target large language model according to the roulette wheel selection method. The model selection module updates the model weights after the candidate program passes the static pre-detection and completes the evaluation. If the candidate program fails the static pre-detection, the corresponding model weight is reduced according to the failure type.

[0092] The search strategy management module 304 is used to maintain the search strategy database, calculate the historical performance of each strategy, and generate or update search strategies when stagnation is detected. The search strategy management module is also used to verify the legality of candidate search strategies, preventing strategies from generating illegal prompts, selecting non-existent parent programs, or violating task interface constraints.

[0093] The hint construction module 305 is used to select the parent program, mutation operator, and heuristic sample set according to the target search strategy, and to construct hint information for code generation. The hint construction module can combine the task description, parent program fragment, high-scoring candidate program summary, failure case summary, static pre-detection rules, and the mutation target of this round into structured hints.

[0094] The program generation module 306 is used to call the target large language model to generate new candidate programs and extract executable code from the model output. The program generation module can also trigger a format correction prompt when extraction fails, or mark the output as having a format error and provide feedback to the model selection module.

[0095] The static preflight module 307 is used to perform EPLB-specific structural checks before task evaluation. The static preflight module may include a syntax checking unit, an interface checking unit, a replica count checking unit, a physical number boundary checking unit, a mapping consistency checking unit, a tensor dimension checking unit, and an error feedback generation unit.

[0096] The evaluation and update module 308 is used to invoke the task evaluator when the static pre-detection is passed, and to update the candidate program database, model weight table, and search strategy database according to the evaluation results. The evaluation and update module is also used to record the overall score, sub-scores, running time, and resource consumption of the candidate programs, and to determine whether to end the optimization based on preset termination conditions.

[0097] like Figure 7 The diagram shows the structure of the dual-layer optimization device. (The following three figures are illustrated in conjunction with the attached figures, or the diagram is as follows...) Figure 7 (These three words should be inserted into the preceding content related to the double-layer structure) Figure 7 is a schematic diagram of a two-layer optimization device for a multi-model meta-evolutionary search method for expert parallel load balancing provided in the second embodiment of this application.

[0098] As shown in Figure 7, the two-layer optimization device of this application includes an inner-layer candidate program evolution structure and an outer-layer search strategy evolution structure. The inner-layer candidate program evolution structure uses expert parallel load balancing candidate programs as the optimization object. It selects a target large language model through a large language model pool, and the large language model generator generates initial programs or new candidate programs. The generated candidate programs enter the evaluator, which evaluates their program validity, load balancing effect, runtime, and overall score. The evaluation results, along with candidate program samples, program execution logs, and detailed program information, are fed back to the search strategy module. The search strategy module selects parent programs, constructs heuristic samples, and determines the next round of candidate program generation, thus forming an inner-layer closed loop of candidate program generation, evaluation, and feedback updates. This structure corresponds to the technical solution of "inner-layer evolutionary candidate program" in the specification, which achieves automatic optimization of candidate programs through the collaborative efforts of a candidate program database, model selection module, prompt construction module, program generation module, static pre-detection module, and evaluation update module.

[0099] The outer search strategy evolution structure optimizes the search strategy itself. When the stagnation mechanism detects that the score improvement of a candidate program within the most recent window is lower than a preset stagnation threshold, it triggers the outer strategy optimization process. In this outer process, the evaluator assesses the current search strategy based on its best score improvement in the historical window, the distribution of static preflight failure types, candidate program diversity, and strategy execution effectiveness. Subsequently, the search strategy module selects a parent strategy based on sample strategies or strategy details, and the large language model generator generates new candidate search strategies. After the candidate search strategies are validated, they are written into the search strategy database and, through the feedback path of "updating search strategy" in the diagram, influence the inner candidate program evolution process, enabling subsequent candidate program generation to adopt updated parent selection rules, heuristic sample construction rules, mutation operator preferences, and exploration utilization ratios.

[0100] Through the two-layer structure shown in Figure 7, this application does not perform a single-layer search on candidate programs. Instead, it continuously optimizes the expert replica count matrix, the logical expert-to-physical expert mapping tensor, and the physical expert-to-logical expert mapping tensor in the inner layer, while dynamically optimizing the search strategy in the outer layer based on the search stagnation state and historical feedback. Therefore, the candidate program optimization results can inversely influence the selection weights of the large language model and the performance evaluation of the search strategy. The updated search strategy results can further guide the generation of subsequent candidate programs, thus forming a closed-loop meta-evolutionary mechanism of "candidate program evolution—strategy evaluation—strategy generation—strategy feedback," improving the search efficiency, stability, and final optimization effect in the expert parallel load balancing task.

[0101] The domain data processed in this application includes activation frequency matrices of logical experts at each layer of the hybrid expert model, expert weight or popularity statistics, device load data, graphics processor or accelerator card topology data, communication overhead data, and the number of physical expert slots at each layer. The output domain results include an expert replica count matrix (logcnt), a logical expert-to-physical expert mapping tensor (log2phy), a physical expert-to-logical expert mapping tensor (phy2log), a model selection weight table, a search strategy database, and a load balancing score. The above data and results are used to automatically optimize expert replica allocation, physical slot placement, and cross-device load balancing in large-scale hybrid expert model inference services.

[0102] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A multi-model meta-evolutionary search method for expert load balancing, characterized in that, include: Obtain the expert parallel load balancing task to be optimized, the task evaluator, the static pre-detection rules, multiple candidate large language models, and the search strategy set; A candidate program database is established, which is used to record candidate programs, their evaluation scores, execution logs, parent sources, evolution methods, and static pre-detection results. Establish a model weight table and a search strategy database. The model weight table is used to record the selection weights of each candidate large language model, and the search strategy database is used to record search strategies and their historical performance. The population state descriptor is calculated based on the candidate program database. The population state descriptor includes the optimal score, score distribution, recent window boost, parent selection frequency, and static pre-detection failure type distribution. The target large language model is selected from multiple candidate large language models according to the model weight table, and the target search strategy is selected from the search strategy set according to the search strategy database; The target search strategy is used to select parent programs, evolutionary methods, and heuristic sample sets from the candidate program database, and to construct prompt information for code generation. New candidate programs are generated based on the prompt information using the target large language model, and a static pre-check oriented towards expert parallel load balancing is performed on the new candidate programs before the task evaluation is executed. When the static pre-detection passes, the task evaluator is invoked to obtain the evaluation score of the new candidate program, and the candidate program database, model weight table, and search strategy database are updated based on the evaluation score; when the static pre-detection fails, constraint feedback is generated according to the failure type and the selection weight of the corresponding large language model or search strategy is reduced. When the recent increase in the window is lower than the preset stagnation threshold, the search strategy is generated or updated based on the search strategy database and the current population state descriptor until the preset termination condition is met and the final candidate program is output.

2. The multi-model meta-evolutionary search method for expert load balancing according to claim 1, characterized in that, Selecting a target large language model from multiple candidate large language models includes: The selection method is adopted by roulette wheel selection to sample according to the selection weight of each candidate large language model. After the candidate program completes static pre-detection and task evaluation, the selection weight of the corresponding candidate large language model is updated according to the score change of the candidate program relative to the parent program or the current best baseline.

3. The multi-model meta-evolutionary search method for expert load balancing according to claim 2, characterized in that, The updated selection weights for the corresponding candidate large language models include: When the evaluation score of the new candidate program is higher than the score of the parent program or the current best baseline score, a positive reward is applied to the candidate large language model that generated the new candidate program; when the new candidate program has a syntax error, a runtime error, a static preflight failure, or an evaluation score lower than the score of the parent program, a negative penalty is applied to the candidate large language model that generated the new candidate program.

4. The multi-model meta-evolutionary search method for expert load balancing according to claim 1, characterized in that, The set of search strategies includes genetic algorithm strategy, differential evolution strategy, and particle swarm strategy; the mutation operator includes random strategy, global refinement strategy, and local repair strategy.

5. The multi-model meta-evolutionary search method for expert load balancing according to claim 1, characterized in that, The historical performance in the search strategy database is calculated based on the optimal score improvement within a monitoring window, the optimal score at the beginning of the window, and the window length, and is used to bias the selection of search strategies.

6. The multi-model meta-evolutionary search method for expert load balancing according to claim 1 or 4, characterized in that, Generating or updating a search strategy based on the search strategy database and the current population state descriptor includes: From the set of search strategies whose historical performance meets the preset conditions, a parent strategy is selected, and at least one of the parent selection rule, heuristic sample construction rule and search utilization ratio of the parent strategy is mutated according to the current population state descriptor to obtain candidate search strategies. The candidate search strategy is validated. If the validation passes, it is adopted as the new target search strategy. If the validation fails, the original target search strategy is retained or a new search strategy is selected from the set of search strategies.

7. The multi-model meta-evolutionary search method for expert load balancing according to claim 1, characterized in that, The static pre-check for expert load balancing includes: Check if the number of expert replicas at each level matches the number of physical experts; check if the physical expert ID slice is not empty and does not exceed the bounds; check if the number of physical expert IDs matches the number of corresponding logical expert replicas; check if the mapping tensor from logical experts to physical experts and the mapping tensor from physical experts to logical experts have the same source; and check if the tensor index dimension and broadcast relationship are valid. When it is detected that a two-dimensional index is directly used for updating a one-dimensional layer index, a two-dimensional replica count tensor is directly passed to a repeated expansion function, the length of the physical expert number slice is insufficient, or the number of mapped tensors written is inconsistent with the number of expert replicas, the new candidate program is rejected from entering the task evaluation, and the corresponding error type is added to the constraint feedback.

8. The multi-model meta-evolutionary search method for expert load balancing according to claim 7, characterized in that, The constraint feedback is written into the prompt information of the next round of code generation, including: adding the error type, error location, corresponding repair constraints, and error patterns that should not be repeated obtained from the static pre-detection to the prompt information of the next round of code generation, so as to constrain the new candidate program generated in the next round to meet the preset structural legality requirements in the process of constructing the expert replica count matrix, selecting the expert index, slicing the physical expert number, and writing the mapping tensor; wherein, the new candidate program is generated using at least one of the following methods: layer-by-layer explicit loop, one-dimensional expert index, boundary assertion, or mapping consistency assertion.

9. The multi-model meta-evolutionary search method for expert load balancing according to claim 1, characterized in that, The method model framework is a two-layer evolutionary structure model: Outer model: Optimize search strategy; Inner model: Optimize candidate programs.

10. A multi-model meta-evolutionary search device for expert load balancing, used to implement the method according to any one of claims 1-9, characterized in that, include: The initialization module is used to obtain the expert load balancing task to be optimized, the task evaluator, static pre-detection rules, multiple candidate large language models, and a set of search strategies. The candidate program database module is used to record candidate programs, their evaluation scores, execution logs, parent sources, evolution methods, and static pre-detection results. The model selection module is used to select the target large language model from multiple candidate large language models according to the model weight table, and update the model weights according to the static pre-detection results and evaluation scores of the candidate programs. The search strategy management module is used to establish a search strategy database and select, generate, or update search strategies based on population state descriptors. The hint construction module is used to select the parent program, evolutionary method and heuristic sample set according to the target search strategy, and construct hint information for code generation; The program generation module is used to generate new candidate programs based on the target large language model; The static pre-inspection module is used to perform checks on the new candidate program before task evaluation, including consistency of expert replica count, physical expert number boundaries, consistency of mapping tensors, and validity of tensor index dimensions. The evaluation and update module is used to call the task evaluator to obtain the evaluation score and update the candidate program database, model weight table and search strategy database when the static pre-detection passes. When the static pre-detection fails, it generates constraint feedback and reduces the selection weight of the corresponding large language model or search strategy.