Parallel strategy generation method, model training method and equipment

By utilizing hardware configuration information to adjust the parallel parameter set, generating fitness evaluation results and adapting them, the problems of low efficiency and resource waste in hybrid parallel training strategies are solved, and efficient parallel training strategy generation is achieved.

CN121599162APending Publication Date: 2026-03-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610123919.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-29
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

In existing technologies, the search method of hybrid parallel training strategies is difficult to quickly determine a strategy that is suitable for the model to be trained, resulting in low training efficiency and high resource consumption.

Method used

By obtaining the configuration information of the hardware device as a constraint, the candidate parallel parameter set is adjusted to generate fitness evaluation results. After the training conditions are met, the model parameters and training data are adapted to the hardware device to generate a parallel training strategy.

Benefits of technology

While improving training efficiency, it reduces resource consumption, and the generated parallel training strategy is more adaptable and efficient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121599162A_ABST
    Figure CN121599162A_ABST
Patent Text Reader

Abstract

The invention provides a parallel strategy generation method, a model training method and equipment, which can be applied to the technical field of artificial intelligence. The parallel strategy generation method comprises the following steps: acquiring equipment configuration information of hardware equipment for training a to-be-trained model; the device configuration information serves as a constraint condition, parallel parameters in the candidate parallel parameter set are adjusted, a to-be-evaluated parallel parameter set with updated parallel parameters is obtained, and the parallel parameters are used for indicating the parallel mode of at least one of model parameters, training data and hardware devices in the to-be-trained model in the training process; based on the equipment configuration information and the model parameters, evaluating the to-be-evaluated parallel parameter set to obtain a fitness evaluation result of the to-be-evaluated parallel parameter set; and under the condition that the fitness assessment result represents that the to-be-assessed parallel parameter set meets the training condition, based on the to-be-assessed parallel parameter set, matching the model parameters and the training data with hardware equipment, and generating a parallel training strategy of the to-be-trained model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and more specifically to a parallel policy generation method, a model training method, and an apparatus. Background Technology

[0002] With the development of artificial intelligence technology, the size of models has also increased significantly. In order to improve training efficiency, large-scale models are usually distributed across multiple computing nodes, and a hybrid parallel training strategy is used for model training.

[0003] In related technologies, grid search or random search are commonly used to search for parallel training strategies to determine a hybrid parallel training strategy for training the model. However, these methods are difficult to quickly determine a hybrid parallel strategy that is suitable for the model to be trained, resulting in low training efficiency and high resource consumption. Summary of the Invention

[0004] In view of the above problems, this application provides a parallel policy generation method, a model training method, and an apparatus.

[0005] According to a first aspect of this application, a parallel strategy generation method is provided, comprising: acquiring device configuration information of a hardware device used to train a model to be trained; adjusting each parallel parameter in a candidate parallel parameter set using the device configuration information as a constraint to obtain an updated set of parallel parameters to be evaluated, wherein the parallel parameters are used to indicate the parallel mode of at least one of the model parameters, training data, and hardware device in the model to be trained during training; evaluating the set of parallel parameters to be evaluated based on the device configuration information and the model parameters to obtain a fitness evaluation result of the set of parallel parameters to be evaluated; and, if the fitness evaluation result indicates that the set of parallel parameters to be evaluated meets the training conditions, adapting the model parameters and the training data to the hardware device based on the set of parallel parameters to be evaluated to generate a parallel training strategy for the hardware device to train the model to be trained using the training data.

[0006] The second aspect of this application provides a model training method, comprising: training a model to be trained on a hardware device using a parallel training strategy corresponding to a target parallel parameter set, to obtain a trained model; wherein the target parallel parameter set is obtained using the parallel training strategy generation method described above.

[0007] A third aspect of this application provides a parallel training strategy generation apparatus, comprising: an acquisition module for acquiring device configuration information of a hardware device used to train a model to be trained; an adjustment module for adjusting each parallel parameter in a candidate parallel parameter set using the device configuration information as a constraint to obtain an updated set of parallel parameters to be evaluated, wherein the parallel parameters are used to indicate the parallelism of at least one of the model parameters, training data, and hardware device in the model to be trained during training; an evaluation module for evaluating the set of parallel parameters to be evaluated based on the device configuration information and the model parameters to obtain a fitness evaluation result of the set of parallel parameters to be evaluated; and a generation module for adapting the model parameters and the training data to the hardware device based on the set of parallel parameters to be evaluated, provided that the fitness evaluation result indicates that the set of parallel parameters to be evaluated meets the training conditions, to generate a parallel training strategy for the hardware device to train the model to be trained using the training data.

[0008] A fourth aspect of this application provides a model training apparatus, comprising: a model training module, used to train a model to be trained on a hardware device using a parallel training strategy corresponding to a target parallel parameter set, to obtain a trained model; wherein the target parallel parameter set is obtained using the parallel training strategy generation method described above.

[0009] A fifth aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0010] A sixth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0011] A seventh aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.

[0012] According to embodiments of this application, by using device configuration information as constraints and adjusting the candidate parallel parameter set based on fitness evaluation results, a set of parallel parameters to be evaluated that is adapted to the hardware device and whose fitness evaluation results meet the training conditions is obtained. Furthermore, model parameters and training data are adapted to the hardware device to generate a parallel training strategy for the hardware device to train the model to be trained using the training data, thereby improving training efficiency while reducing resource consumption. Attached Figure Description

[0013] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments of this application with reference to the accompanying drawings.

[0014] Figure 1 The diagram illustrates an application scenario of the parallel policy generation method, model training method, and device according to embodiments of this application.

[0015] Figure 2 A flowchart of a parallel training strategy generation method according to an embodiment of this application is shown.

[0016] Figure 3 A schematic diagram illustrating the determination of spatial positioning parameters for the next round according to an embodiment of this application is shown.

[0017] Figure 4 A flowchart of a model training method according to an embodiment of this application is shown.

[0018] Figure 5 A structural block diagram of a parallel training strategy generation apparatus according to an embodiment of this application is shown.

[0019] Figure 6 A structural block diagram of a model training apparatus according to an embodiment of this application is shown.

[0020] Figure 7 A block diagram of an electronic device suitable for implementing a parallel training policy generation method according to an embodiment of this application is shown. Detailed Implementation

[0021] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0022] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0024] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0025] The mainstream parallel strategies in related technologies include data parallelism, tensor parallelism, pipelined parallelism, and sequence parallelism. In practical applications, in order to effectively distribute a single model across multiple computing nodes, it is usually necessary to combine the above parallel strategies to form a hybrid parallel scheme.

[0026] While hybrid parallelism offers the possibility of training ultra-large models, the increased configuration complexity it introduces also presents significant performance tuning challenges. How to automatically and efficiently select the optimal combination of hybrid parallelism strategies for specific model architectures and cluster hardware environments has become a core bottleneck restricting the training efficiency and development progress of large models.

[0027] Currently, configuration processes typically rely on manual debugging and intuition from distributed system experts. This approach is not only time-consuming and labor-intensive, but also, due to the vast search space, experts often find locally optimal solutions, making it difficult to guarantee optimal global performance. Furthermore, using grid search or random search to train and evaluate the performance of each configuration, even with only a few iterations, results in extremely high cumulative computational costs, leading to significant time overhead and making it impractical in practice. While some reinforcement learning techniques theoretically offer ideal solutions to these problems, they require execution in a real-world environment for reward feedback in each policy iteration, resulting in an exceptionally lengthy training process and negating any efficiency advantages.

[0028] Embodiments of this application provide a method for generating a parallel training strategy, comprising: obtaining device configuration information of a hardware device used to train a model to be trained; adjusting each parallel parameter in a candidate parallel parameter set using the device configuration information as a constraint to obtain an updated set of parallel parameters to be evaluated, wherein the parallel parameters are used to indicate the parallel mode of at least one of the model parameters, training data, and hardware device in the model to be trained during the training process; evaluating the set of parallel parameters to be evaluated based on the device configuration information and the model parameters to obtain a fitness evaluation result of the set of parallel parameters to be evaluated; and, if the fitness evaluation result indicates that the set of parallel parameters to be evaluated meets the training conditions, adapting the model parameters and training data to the hardware device based on the set of parallel parameters to be evaluated to generate a parallel training strategy for the hardware device to train the model to be trained using the training data.

[0029] The embodiments of this application adjust the candidate parallel parameter set by using device configuration information as a constraint, thereby obtaining a parallel training strategy adapted to the hardware device, thereby improving training efficiency while reducing resource consumption.

[0030] Figure 1 The diagram illustrates an application scenario of the parallel policy generation method, model training method, and device according to embodiments of this application.

[0031] like Figure 1 As shown, the application scenario according to this embodiment may include a terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, a server 105, and a distributed cluster 106. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server cluster 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc. The distributed cluster 106 may include multiple computing nodes, such as a first computing node 106_1 and a second computing node 106_2.

[0032] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0033] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0034] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0035] The distributed cluster 106 can deploy a model to be trained. For example, each computing node in the distributed cluster 106 can deploy a portion of the parameters of the model to be trained.

[0036] It should be noted that the parallel training strategy generation method provided in this application embodiment can generally be executed by server 105. Correspondingly, the parallel training strategy generation device provided in this application embodiment can generally be located in server 105. The parallel training strategy generation method provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the parallel training strategy generation device provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0037] It should be understood that Figure 1 The number of terminal devices, networks, servers, distributed clusters, and computing nodes shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, servers, distributed clusters, and computing nodes can be included.

[0038] The following will be based on Figure 1 The described scene, through Figures 2-3 The parallel training strategy generation method of the application embodiments is described in detail.

[0039] Figure 2 A flowchart of a parallel training strategy generation method according to an embodiment of this application is shown.

[0040] like Figure 2 As shown, the parallel training strategy generation method in this embodiment includes operations S210 to S240.

[0041] In operation S210, the device configuration information of the hardware device used to train the model to be trained is obtained.

[0042] In operation S220, the device configuration information is used as a constraint to adjust each parallel parameter in the candidate parallel parameter set, so as to obtain the parallel parameter set to be evaluated with updated parallel parameters.

[0043] In operation S230, based on the device configuration information and model parameters, the parallel parameter set to be evaluated is evaluated to obtain the fitness evaluation result of the parallel parameter set to be evaluated.

[0044] In operation S240, if the fitness evaluation result indicates that the parallel parameter set to be evaluated meets the training conditions, the model parameters and training data are adapted to the hardware device based on the parallel parameter set to be evaluated, and a parallel training strategy for the hardware device to train the model to be trained using the training data is generated.

[0045] In the embodiments of this application, the model to be trained can be distributed across multiple computing nodes in a distributed cluster. The hardware configuration information can include key indicators such as the computing power, memory size, and network bandwidth of the computing nodes, accurately reflecting the performance of the hardware and providing basic data for adjusting parallel parameters. Model parameters can include the number of layers, hidden layer dimensions, and number of attention heads in the model to be trained, which may affect the feasibility of parallel parameters.

[0046] After obtaining the device configuration information, it can be used as an important constraint to adjust each parallel parameter in the candidate parallel parameter set, ensuring that the parallel parameters can fully adapt to the actual performance of the hardware device.

[0047] In embodiments of this application, parallel parameters are used to indicate the parallelism of at least one of the model parameters, training data, and hardware devices in the model to be trained during the training process. It is understood that the candidate parallel parameter set may include multiple parallel parameters capable of characterizing hybrid parallelism. The parallel parameters can be the degree of parallelism of any of the following parallelism methods: data parallelism, model parallelism, pipeline parallelism, and sequence parallelism, as well as the splitting point for dividing the model to be trained under model parallelism.

[0048] After adjusting the parallel parameters, the set of parallel parameters to be evaluated is obtained. In one example, the parallel parameters can be adjusted according to the fitness evaluation results while satisfying the constraints using an optimization model, until the fitness evaluation results indicate that the set of parallel parameters to be evaluated meets the training conditions.

[0049] When evaluating the parallel parameters to be evaluated, the training efficiency and resource consumption of the model to be trained based on the parallel parameters to be evaluated can be evaluated. The corresponding training conditions can be that the training efficiency is greater than a preset efficiency value or the resource consumption is less than a preset resource value.

[0050] After determining the set of parallel parameters to be evaluated, the model parameters and training data can be precisely adapted to the hardware devices based on this set to generate a targeted parallel training strategy. In some embodiments, training conditions may include a fitness evaluation result greater than a preset fitness evaluation result threshold.

[0051] According to embodiments of this application, by using device configuration information as constraints and adjusting the candidate parallel parameter set based on fitness evaluation results, a set of parallel parameters to be evaluated that is adapted to the hardware device and whose fitness evaluation results meet the training conditions is obtained. Furthermore, model parameters and training data are adapted to the hardware device to generate a parallel training strategy for the hardware device to train the model to be trained using the training data, thereby improving training efficiency while reducing resource consumption.

[0052] According to an embodiment of this application, the parallel training strategy generation method further includes iteratively adjusting the parallel parameter set to be evaluated; wherein performing one adjustment includes: determining spatial positioning parameters for the search space used to collect the parallel parameter set of the next round based on multiple candidate parallel parameter sets in the current round; determining the search space of the next round based on the spatial positioning parameters; and determining multiple parallel parameter sets to be evaluated in the next round from the search space of the next round.

[0053] The candidate parallel parameter sets for the current round can be selected from the multiple parallel parameter sets to be evaluated in the current round. Specifically, the candidate parallel parameter sets can be selected based on their respective fitness evaluation scores, with the highest fitness evaluation scores being chosen as the candidate parallel parameter sets. When adjusting the parallel parameter sets to be evaluated, an iterative adjustment method can also be used. In the embodiments of this application, the search space can be iteratively adjusted by adjusting the spatial positioning parameters, and multiple parallel parameter sets to be evaluated for the next round can be resampled from the search space of the next round until the fitness evaluation results indicate that the parallel parameter sets to be evaluated meet the training conditions.

[0054] For example, candidate parallel parameter sets It can be expressed by the following formulas (1) to (2):

[0055] (1);

[0056] (2);

[0057] in, It can represent the degree of data parallelism. It can represent the degree of parallelism in a pipeline. It can represent the parallelism of tensors. It can represent the degree of expert parallelism. It can represent the degree of parallelism of a sequence. This represents the model splitting weight parameter, guiding the selection of splitting points for pipeline parallelism and model parallelism. k represents the number of splitting points, each... This corresponds to a potential split point between layers in the model. This represents the number of layers in the model to be trained.

[0058] Spatial positioning parameters can be used to define the search space. By iteratively adjusting the search space, the search range can be gradually narrowed to find a better combination of parallel parameters. Spatial positioning parameters may include, for example, the search center and the search range.

[0059] In some embodiments, in each iteration, the location of the search center can be dynamically determined based on the fitness evaluation results of multiple parallel parameters to be evaluated in the current round. This search center represents the region currently considered most likely to contain the optimal solution. Simultaneously, the size of the search range is determined according to preset rules or algorithms, which determines the breadth of the search in the next round. By continuously iterating and adjusting the search center and search range, the focus can be gradually shifted to the region where the optimal solution is located, thereby efficiently finding a set of parallel parameters that meets the training conditions and has high training efficiency.

[0060] According to embodiments of this application, by iteratively adjusting the search space of the parallel parameter set to be evaluated, the optimal parallel parameter set to be evaluated can be gradually locked from a global perspective, avoiding local optima, improving the accuracy of the generated parallel training strategy, and thus improving training efficiency.

[0061] According to an embodiment of this application, based on multiple candidate parallel parameter sets in the current round, spatial positioning parameters for collecting the parallel parameter sets in the next round are determined, including: weighted summation of multiple candidate parallel parameter sets in the current round based on their respective weights to obtain the search center of the search space in the next round; and spatial positioning parameters representing the search space in the next round are determined based on the search center of the current round, the search center of the next round, and multiple candidate parallel parameter sets in the current round.

[0062] In the embodiments of this application, the search center for the next round can be determined first based on multiple candidate parallel parameter sets, and then the spatial positioning parameters for the next round can be further determined based on the determined search center for the next round. For example, the spatial positioning parameters can be adjusted using a covariance matrix adaptive evolution strategy (CMA-ES).

[0063] When determining the search center, a weighted sum can be applied to multiple candidate parallel parameter sets in the current round. For example, the weights of each candidate parallel parameter set can be determined based on their fitness evaluation results, causing the search space to shrink towards candidate parallel parameter sets with higher fitness. Alternatively, the weights of multiple candidate parallel parameter sets can be determined based on their distances, giving greater weight to closer candidate parallel parameter sets, thus favoring the exploration of these nearby regions during the search.

[0064] According to an embodiment of this application, by weighted summation of multiple candidate parallel parameter sets in the current round, the search center of the search space in the next round is obtained. This makes the search center tend to be closer to the candidate parallel parameter sets that perform better in the current round, thereby making it more likely to discover the globally optimal parallel parameter set in the subsequent search process, and improving the generation efficiency and accuracy of the parallel training strategy.

[0065] According to an embodiment of this application, the parallel training strategy generation method further includes: sorting the multiple candidate parallel parameter sets in the current round based on their respective fitness evaluation results to obtain a sorting result; and determining the weights of the multiple candidate parallel parameter sets in the current round according to the sorting result and a predetermined weight function.

[0066] When determining the weights of each candidate parallel parameter set, the multiple candidate parallel parameter sets can first be sorted according to the fitness evaluation results to determine the performance of each candidate parallel parameter set in the current round, providing an orderly reference for subsequent weight determination.

[0067] After obtaining the ranking results, the weights of each candidate parallel parameter set can be determined according to the pre-set weight function and the ranking results. The pre-set weight function can be flexibly designed according to actual needs and scenarios. For example, it can be a function that is linearly related to the ranking, or it can be other complex functions that are more in line with actual needs. The pre-set weight function can be shown in the following formula (3):

[0068] , (3);

[0069] in, This represents the weight of the i-th candidate parallel parameter set sorted by fitness. Indicates the number of candidate parallel parameter sets. This represents the i-th candidate set of parallel parameters sorted by fitness.

[0070] According to embodiments of this application, by using fitness evaluation scores to assign weights to candidate parallel parameter sets, the importance of each candidate parallel parameter set in the current round can be more accurately reflected. This not only makes candidate parallel parameter sets with higher fitness evaluation scores have greater influence, but also avoids dependence on a single candidate parallel parameter set, improves the efficiency of determining the search center, and thus improves search efficiency.

[0071] Figure 3 A schematic diagram illustrating the determination of spatial positioning parameters for the next round according to an embodiment of this application is shown.

[0072] like Figure 3As shown, fitness evaluation is performed on candidate parallel parameter set 1, candidate parallel parameter set 2, ..., candidate parallel parameter set i respectively to obtain fitness evaluation result 1, fitness evaluation result 2, ..., fitness evaluation result i. Based on the ranking result obtained by sorting fitness evaluation result 1, fitness evaluation result 2, ..., fitness evaluation result i, the weights of candidate parallel parameter set 1, candidate parallel parameter set 2, ..., candidate parallel parameter set i are determined, including weight 1, weight 2, ..., weight i.

[0073] like Figure 3 As shown, after obtaining weights 1, 2, ..., i, we can use weights 1, 2, ..., i to perform a weighted summation on candidate parallel parameter sets 1, 2, ..., i to obtain the search center for the next round. Based on candidate parallel parameter sets 1, 2, ..., i, the search center for the next round, and the search center for the current round, we can determine the spatial positioning parameters for the next round.

[0074] According to embodiments of this application, spatial positioning parameters characterizing the search space of the next round are determined based on the search center of the current round, the search center of the next round, and multiple candidate parallel parameter sets in the current round. This includes: updating the evolutionary path parameters based on the search center of the current round and the search center of the next round to obtain updated evolutionary path parameters; updating the step size based on the updated evolutionary path parameters to obtain updated step size; updating the covariance matrix based on the updated evolutionary path parameters and multiple candidate parallel parameter sets in the current round to obtain updated covariance matrix; and obtaining spatial positioning parameters based on the updated evolutionary path parameters, updated step size, and updated covariance matrix.

[0075] In embodiments of this application, spatial positioning parameters may include evolutionary path parameters, step size, and covariance matrix. When updating spatial positioning parameters, the evolutionary path parameters, step size, and covariance matrix all need to be updated.

[0076] When updating the evolutionary path parameters, updates can be made based on the search center of the current round and the search center of the next round. The method for determining the updated evolutionary path parameters can be shown in the following formula (4):

[0077] (4);

[0078] in, This indicates the updated evolutionary path parameters. It is the learning rate, usually taken as 4 / (n + 4), where n is the dimension of the search space. These are the evolution path parameters, initially set as a vector of zeros. = 1 / Σ ² represents the number of independent individuals equivalent in a weighted reorganization, quantifying the number of individuals that actually contribute information during the weighted reorganization process. Indicates the search center for the next round. Indicates the search center for the current round. Indicates the step size.

[0079] When updating the step size, it can be updated based on the updated evolutionary path parameters. The method for determining the updated step size is as shown in the following formula (5):

[0080] (5);

[0081] in, Indicates the updated step size. The learning rate for the step-size path is typically set to... , The damping coefficient is usually taken as... , Let be the expectation of an n-dimensional standard Gaussian vector.

[0082] When updating the covariance matrix, it can be updated based on the updated evolutionary path parameters and the multiple candidate parallel parameter sets of the current round. The updated covariance matrix can be determined by the following formulas (6) to (10):

[0083] × (6);

[0084] (7);

[0085] (8);

[0086] (9);

[0087] (10);

[0088] in, For the updated covariance matrix, Let covariance matrix be the variance matrix. The learning rate is rank 1, approximately 2 / (n² + 6). Let μ be the rank learning rate. Update the matrix for rank 1. Update the matrix for rank μ. This indicates the offset of the i-th candidate parallel parameter set, sorted by fitness, relative to the search center in the current round. The step size.

[0089] According to embodiments of this application, by updating the evolution path parameters, step size, and covariance matrix, the search space can be dynamically adjusted, the search range can be gradually narrowed, and the region where the optimal solution is located can be focused on. This allows for the efficient finding of a set of parallel parameters to be evaluated that meets the training conditions and has high training efficiency, providing strong support for generating high-quality parallel training strategies.

[0090] According to an embodiment of this application, determining multiple sets of parallel parameters to be evaluated in the next round from the search space of the next round includes: determining a predetermined number of initial sets of parallel parameters to be evaluated in the next round from the search space using a random sampling method; and modifying the parameters of the multiple initial sets of parallel parameters to be evaluated based on device configuration information to obtain multiple sets of parallel parameters to be evaluated.

[0091] In the embodiments of this application, the search space can be a multivariate Gaussian distribution, and a predetermined number of initial parallel parameter sets to be evaluated for the next round can be sampled randomly. The predetermined number can be determined by the following formula (11):

[0092] (11);

[0093] in, Indicates the number of orders.

[0094] After collecting multiple initial sets of parallel parameters to be evaluated, parameter corrections, such as feasibility corrections, can be performed on each of these initial sets of parallel parameters to ensure that all parallel parameters to be evaluated can be decoded into valid parallel configurations.

[0095] According to embodiments of this application, by modifying the parameters of the initial set of parallel parameters to be evaluated, all sets of parallel parameters to be evaluated can better adapt to the actual hardware environment, improve the accuracy of the determined set of parallel parameters to be evaluated, and thus ensure the feasibility and effectiveness of the parallel training strategy in practical applications.

[0096] According to an embodiment of this application, based on device configuration information and model parameters of the model to be trained, an evaluation of the parallel parameter set to be evaluated is performed to obtain a fitness evaluation result of the parallel parameter set to be evaluated. This includes: decoding the parallel parameters to be decoded in the parallel parameter set to be evaluated to obtain decoded parallel parameters; extracting features from the device configuration information, model parameters, and decoded parallel parameters to obtain fused features; performing an initial evaluation based on the fused features to obtain a first evaluation result and a second evaluation result; and performing an fitness evaluation based on the first evaluation result and the second evaluation result to obtain a fitness evaluation result.

[0097] Since the set of parallel parameters to be evaluated is generated by an evolutionary algorithm such as CMA-ES, the parallel parameters in the set of parallel parameters to be evaluated are usually in encoded form so that they can be directly processed by the algorithm. Therefore, when evaluating the parallel parameters to be evaluated, it is necessary to decode the parallel parameters to be decoded in the set of parallel parameters to be evaluated so as to facilitate subsequent evaluation.

[0098] Decoding parallelism parameters can include decoding data parallelism, decoding pipeline parallelism, decoding tensor parallelism, decoding expert parallelism, and decoding sequence parallelism. The decoding process can be represented by the following formula (12):

[0099] DP = clamp(round( ), 1, DP_max) (12;

[0100] Here, DP represents the parallelism of the decoding data, round() is the rounding function, DP_max represents the maximum value of the data parallelism, and the clamp function ensures that the result is within a valid range. The decoding method for other parallel parameters to be decoded is similar to that for data parallelism, and will not be described in detail here.

[0101] During evaluation, features can be extracted from device configuration information, model parameters, and decoding parallel parameters to obtain fused features. These fused features can more comprehensively reflect the complex relationships between the device, model, and parallel parameters. These fused features not only include the performance characteristics of the device itself but also cover the model training requirements and the impact of parallel parameters on the training process. By fusing this multi-dimensional information, the fitness of the set of parallel parameters to be evaluated can be assessed more accurately.

[0102] After obtaining the fusion features, different aspects of the fusion features can be evaluated to obtain multiple evaluation results. In the embodiments of this application, the evaluation results include a first evaluation result and a second evaluation result. The first evaluation result represents the training performance of the parallel training strategy corresponding to the parallel parameter set to be evaluated, and the second evaluation result represents the hardware adaptation performance of the parallel training strategy corresponding to the parallel parameter set to be evaluated.

[0103] After obtaining the first and second assessment results, the adaptive assessment result can be obtained by fusing the first and second assessment results. The process of determining the adaptive assessment result can be shown in the following formula (13):

[0104] = - × (13);

[0105] in, This represents the fitness evaluation result of the i-th parallel parameter set to be evaluated. Represents the first evaluation result of the i-th set of parallel parameters to be evaluated, memory_penalty i This represents the second evaluation result for the i-th set of parallel parameters to be evaluated. The weight of the second evaluation result.

[0106] In one specific embodiment, the second evaluation result can be a key penalty term, used to quantify and penalize parallel strategies that exceed hardware configuration limits or have low memory efficiency. Its core function is to ensure that the strategies selected by the evolutionary algorithm do not fail due to insufficient memory during actual deployment, while guiding the search towards memory-efficient configurations. Other penalty terms and corresponding influence weights can also be added according to actual needs.

[0107] According to embodiments of this application, by evaluating the parallel parameter set to be evaluated from two aspects—training performance and hardware adaptation performance—the evaluation results are more comprehensive and accurate, and can comprehensively consider various key factors in the practical application of parallel training strategies. This evaluation method not only focuses on performance indicators such as training speed, but also fully considers the limitations and utilization efficiency of hardware resources, ensuring that the generated parallel training strategy is both efficient and feasible.

[0108] According to an embodiment of this application, an evaluation model is used to evaluate the parallel parameter set to be evaluated based on device configuration information, model parameters, and the set of parallel parameters to be evaluated, to obtain the fitness evaluation result of the set of parallel parameters to be evaluated; wherein, the evaluation model is trained in the following manner: obtaining a training parallel parameter set, a first sample evaluation label, and a second sample evaluation label; inputting the training parallel parameter set into an initial evaluation model to obtain a first sample evaluation result and a second sample evaluation result; and using a first loss function value determined based on the first sample evaluation label and the first sample evaluation result, and a second loss function value determined based on the second sample evaluation label and the second sample evaluation result, the parameters of the initial evaluation model are adjusted to obtain the evaluation model.

[0109] In embodiments of this application, fitness can be evaluated using an evaluation model. The evaluation model may employ activation functions in the middle of fully connected layers to balance nonlinearity and gradient flow, and may introduce regularization techniques to prevent overfitting. In one example, the evaluation model may also incorporate a graph neural network module to capture the topological properties of the model's computational graph.

[0110] Before applying the evaluation model to fitness evaluation, the evaluation model can be trained. Specifically, the parameters of the initial evaluation model can be adjusted using a training parallel parameter set, a first sample evaluation label, and a second sample evaluation label to obtain the evaluation model. The first sample evaluation label characterizes the training performance of the parallel training strategy corresponding to the training parallel parameter set, and the second sample evaluation label characterizes the hardware adaptation performance of the parallel training strategy corresponding to the training parallel parameter set.

[0111] When training the evaluation model, a first loss function value can be determined based on the evaluation label and evaluation result of the first sample, and a second loss function value can be determined based on the evaluation label and evaluation result of the second sample. The process of determining the first loss function value can be shown in the following formula (14):

[0112] (14);

[0113] in, Here, N represents the first loss function value, and N is the number of parallel training parameters. As the first sample evaluation label, The results of the first sample evaluation. This is the logarithmic offset, which can be a very small positive number to prevent mathematical calculation errors when the first sample evaluation label or the first sample evaluation result is zero.

[0114] The process of determining the value of the second loss function can be shown in the following formula (15):

[0115] (15);

[0116] Where memory_loss is the value of the second loss function, memory_true is the evaluation label of the second sample, and memory_pred is the evaluation result of the second sample.

[0117] Furthermore, the value of the third loss function can be determined using the evaluation results of the first samples with different training parallel parameters. The method for determining the value of the third loss function is as follows (16):

[0118] ranking_loss = × Σ max(0, - ( - ) × sign( - (16);

[0119] Where ranking_loss is the value of the third loss function, M is the number of parallel parameter pairs, and i and j represent the first and second parallel parameters in the parallel parameter pair, respectively. This represents the evaluation result of the first sample for the first training parallel parameters. This represents the evaluation result of the second sample for the second parallel training parameter. This represents the evaluation label of the first sample for the first parallel training parameter. This represents the first sample evaluation label for the first training parallel parameter.

[0120] In one example, the first loss function value, the second loss function value, and the third loss function value can be weighted and summed to obtain the total loss function value. The parameters of the initial evaluation model can then be adjusted based on the total loss function value to obtain the evaluation model. The process of determining the total loss function value can be shown in the following formula (17):

[0121] (17);

[0122] in, This is the total loss function value. , , These are the weights corresponding to the first, second, and third loss function values, respectively.

[0123] According to embodiments of this application, by using a trained evaluation model to evaluate the parallel parameter set to be evaluated, the fitness evaluation results of the parallel parameter set to be evaluated can be obtained more efficiently and accurately. The evaluation model is carefully trained, comprehensively considering multiple key factors such as training performance and hardware adaptation performance, making the evaluation results more comprehensive and reliable. In practical applications, the evaluation model can be used to quickly screen out parallel parameter sets with high fitness, further improving the efficiency and effectiveness of parallel training.

[0124] According to embodiments of this application, based on a set of parallel parameters to be evaluated, model parameters, training data, and hardware devices are adapted to generate a parallel training strategy for training a model on a hardware device using training data. This includes: decoding the parameters to be decoded in the set of parallel parameters to be evaluated to obtain decoded parallel parameters; using device configuration information to perform hardware resource constraint verification on the decoded parallel parameters to obtain a verification result; if the verification result indicates that the parallel mode indicated by the decoded parallel parameters meets the hardware constraints indicated by the device configuration information, the training parameters are split based on the set of parallel parameters to be evaluated to obtain a set of training sub-parameters for parallel training; and based on the device configuration information and the set of training sub-parameters, the set of training sub-parameters is adapted to the hardware device to generate a parallel training strategy.

[0125] If the fitness evaluation results meet the training conditions, a parallel training strategy can be generated based on the set of parallel parameters to be evaluated. First, the parameters to be decoded can be decoded to obtain the decoded parallel parameters. The specific decoding process is the same as described above and will not be repeated here. After obtaining the decoded parallel parameters, hardware resource constraints can be verified. Specifically, the number of computing nodes in the device configuration information can be used to verify the product of the parallelism represented by each of the multiple decoded parallel parameters. If the number of computing nodes matches the product, the verification result indicates that the parallelism indicated by the decoded parallel parameters meets the hardware constraints indicated by the device configuration information. Otherwise, subsequent parallel training strategies can be generated based on other sets of parallel parameters to be evaluated.

[0126] Training parameters include the training data and the model parameters of the model to be trained. Parallel training typically requires splitting the training data and model parameters to utilize the resulting sets of training sub-parameters for parallel training. For example, if the parallel parameter set to be evaluated indicates a parallelism of 3, the training data can be divided into 3 equal parts. As another example, if the parallel parameter set to be evaluated indicates a parallelism of 2, each layer of the model to be trained can be divided into 2 equal parts.

[0127] Subsequently, based on device configuration information, such as the computing power and storage capacity of each computing node in the hardware device, and the obtained training sub-parameter set, the training sub-parameter set is rationally allocated to each hardware device, enabling each hardware device to undertake the corresponding training task, thereby generating a complete parallel training strategy. This parallel training strategy can fully utilize the resources of the hardware devices and improve the efficiency and performance of model training.

[0128] After determining the parallel training strategy, the parallel training strategy can be further generated into a strategy deployment configuration file, which contains complete parallel parameters, so as to train the model to be trained by running the strategy deployment configuration file.

[0129] According to embodiments of this application, by decoding and verifying the parallel parameter set to be evaluated, the generated parallel training strategy can be adapted to the actual limitations of the hardware device, ensuring that it will not fail due to insufficient resources or misconfiguration during actual deployment. Simultaneously, by splitting and rationally allocating the training parameters, this strategy can fully utilize the computing power and storage capacity of the hardware device to achieve efficient parallel training.

[0130] According to an embodiment of this application, the parallel training strategy generation method further includes: initializing the initial parallel parameters; if the type of the parallel parameters is determined to be discrete, performing data conversion on the assigned initial parallel parameters to obtain continuous type initial parallel parameters; obtaining an initial parallel parameter set based on multiple initial parallel parameters; and performing at least one iterative adjustment on the initial parallel parameter set using device configuration information as a constraint to obtain a candidate parallel parameter set.

[0131] When generating candidate parallel parameter sets, in order for the parallel parameters to be directly processed by the evolutionary algorithm, the complex hybrid parallel configuration and related information need to be encoded into a fixed-length candidate parallel parameter set. The candidate parallel parameter set needs to maintain the continuity of the encoding space to facilitate the exploration of the evolutionary algorithm.

[0132] First, the initial parallel parameters can be initialized. When initializing, the values ​​can be determined based on the number of processors, the storage capacity of the storage device, and the network bandwidth of the communication device. For example, the initial data parallelism parameter in the initial parallelism parameters can be determined as shown in formula (18):

[0133] = min( , , (18);

[0134] in, The initial data parallelism parameter, For the number of processors, For storage capacity, This refers to network bandwidth.

[0135] In some embodiments, Specifically, this involves the total number of available GPUs in the cluster, ensuring that resources are not over-segmented. Specifically, regarding the available memory capacity of a single GPU, you can refer to information such as model parameter memory, activation value memory, and optimizer state memory for configuration. Specifically, this involves the communication overhead of gradient synchronization, to avoid communication bottlenecks caused by excessive data parallelism. The other initialization assignments follow the same logic and will not be elaborated upon here.

[0136] When the parallel parameter type is discrete, it can be converted into a continuous type of initial parallel parameter through data transformation, such as logarithmic transformation. The data transformation process can be shown in the following formula (19):

[0137] (19);

[0138] in, This represents the degree of data parallelism in the initial parallel parameters. The logarithmic transformation linearizes the exponentially growing search space, improving search efficiency; it is also more in line with the actual situation of hardware resource allocation (usually allocated in powers of 2).

[0139] In some embodiments, the candidate parallel parameter set should also ensure feasibility. Therefore, after obtaining the initial parallel parameter set, the device configuration information can be used as a constraint to perform at least one iterative adjustment on the initial parallel parameter set to obtain the candidate parallel parameter set so that it meets the hardware constraints.

[0140] According to embodiments of this application, by assigning values ​​to the initial parallel parameters, converting data, and iteratively adjusting them, the resulting candidate parallel parameter set can not only better adapt to the actual conditions of the hardware device, but also facilitate the processing of evolutionary algorithms and improve search efficiency.

[0141] According to an embodiment of this application, the assigned initial parallel parameters are converted into data to obtain continuous type initial parallel parameters, including: based on device configuration information, the assigned initial parallel parameters are compared for adaptability to obtain a comparison result; if the comparison result indicates that the assigned initial parallel parameters meet the hardware constraints indicated by the device configuration information, the assigned initial parallel parameters are converted into data to obtain initial parallel parameters.

[0142] In embodiments of this application, random values ​​can also be assigned to the initial parallel parameters. Therefore, it is necessary to perform an adaptability comparison of the initial parallel parameters based on the device configuration information to ensure that they meet the hardware constraints. The device configuration information includes at least one of the following: the number of processors, the storage capacity of the storage device, and the network bandwidth of the communication device.

[0143] If the initial parallel parameters after assignment meet the hardware constraints indicated by the device configuration information, the assigned initial parallel parameters are converted using the same data conversion method as described above, and will not be repeated here.

[0144] According to embodiments of this disclosure, by comparing the adaptability of the initial parallel parameters assigned based on device configuration information, the obtained initial parallel parameters can better adapt to the actual conditions of the hardware device, avoiding resource waste or training failure due to unreasonable parameter settings, thereby improving search efficiency and search accuracy.

[0145] Figure 4 A flowchart of a model training method according to an embodiment of this application is shown.

[0146] like Figure 4 As shown, the model training method includes operation S410.

[0147] In operation S410, the training model is trained on the hardware device using the parallel training strategy corresponding to the target parallel parameter set to obtain the trained model; wherein, the target parallel parameter set is obtained using the parallel training strategy generation method.

[0148] After determining the parallel training strategy for the model to be trained using the parallel training strategy generation method in this application, the model to be trained can be trained using the parallel training strategy to obtain the trained model.

[0149] During training, it is essential to first ensure that the hardware meets the preset configuration requirements. These requirements include, but are not limited to, the number of processors, the storage capacity of the storage devices, and the network bandwidth of the communication devices, to ensure the stability and efficiency of the training process. Then, the model to be trained is trained according to a defined parallel training strategy until the preset training objective is reached or the stopping condition is met, thus obtaining the trained model.

[0150] According to embodiments of this application, by utilizing the parallel training strategy generation method of this application, a parallel training strategy adapted to the actual configuration of the hardware device can be automatically and accurately generated, thereby maximizing the utilization of hardware resources and improving training speed.

[0151] According to an embodiment of this application, the model training method further includes: searching in a parallel training strategy library based on device configuration information and model parameters to obtain search results; and training the model to be trained using the target parallel training strategy if the search results indicate that the parallel training strategy library includes a target parallel training strategy that matches the device configuration information and model parameters.

[0152] In the embodiments of this application, after generating a parallel training strategy using the parallel training strategy generation method of this application, the parallel training strategy can be stored in the parallel training strategy library in association with the device configuration information and model parameters, so that the matching target parallel training strategy can be directly retrieved from the parallel training strategy library during subsequent model training, thereby improving training efficiency.

[0153] Based on the above-described parallel training strategy generation method, this application also provides a parallel training strategy generation apparatus. The following will combine... Figure 5 The device is described in detail.

[0154] Figure 5 A structural block diagram of a parallel training strategy generation apparatus according to an embodiment of this application is shown.

[0155] like Figure 5 As shown, the parallel training strategy generation device 500 of this embodiment includes an acquisition module 510, an adjustment module 520, an evaluation module 530, and a generation module 540.

[0156] The acquisition module 510 is used to acquire device configuration information of the hardware device used to train the model to be trained. In one embodiment, the acquisition module 510 can be used to perform the operation S210 described above, which will not be repeated here.

[0157] The adjustment module 520 is used to adjust each parallel parameter in the candidate parallel parameter set using device configuration information as constraints, to obtain an updated set of parallel parameters to be evaluated. The parallel parameters indicate the parallelism of at least one of the model parameters, training data, and hardware devices in the model to be trained during training. In one embodiment, the adjustment module 520 can be used to perform the operation S220 described above, which will not be repeated here.

[0158] The evaluation module 530 is used to evaluate the parallel parameter set to be evaluated based on the device configuration information and model parameters, and obtain the fitness evaluation result of the parallel parameter set to be evaluated. In one embodiment, the evaluation module 530 can be used to perform the operation S230 described above, which will not be repeated here.

[0159] The generation module 540 is used to, when the fitness evaluation result indicates that the parallel parameter set to be evaluated meets the training conditions, adapt the model parameters and training data to the hardware device based on the parallel parameter set to be evaluated, and generate a parallel training strategy for the hardware device to train the model to be trained using the training data. In one embodiment, the generation module 540 can be used to perform the operation S240 described above, which will not be repeated here.

[0160] According to an embodiment of this application, the parallel training strategy generation device 500 further includes an iterative adjustment module.

[0161] The iterative tuning module is used to adjust the set of parallel parameters to be evaluated in an iterative manner. The iterative tuning module also includes a location determination submodule, a space determination submodule, and a parameter determination submodule.

[0162] The location determination submodule is used to determine the spatial location parameters for the search space used to collect the parallel parameter set in the next round, based on multiple candidate parallel parameter sets in the current round.

[0163] The spatial determination submodule is used to determine the search space for the next round based on spatial positioning parameters.

[0164] The parameter determination submodule is used to determine the set of multiple parallel parameters to be evaluated in the next round from the search space of the next round.

[0165] According to an embodiment of this application, the spatial determination submodule includes a center determination unit and a spatial determination unit.

[0166] The center determination unit is used to perform a weighted summation of multiple candidate parallel parameter sets in the current round based on their respective weights, to obtain the search center of the search space in the next round.

[0167] The spatial determination unit is used to determine the spatial positioning parameters that characterize the search space of the next round based on the search center of the current round, the search center of the next round, and multiple candidate parallel parameter sets in the current round.

[0168] According to an embodiment of this application, the spatial determination submodule further includes a sorting determination unit and a weight determination unit.

[0169] The sorting determination unit is used to sort multiple candidate parallel parameter sets in the current round based on their respective fitness evaluation results, and obtain the sorting result.

[0170] The weight determination unit is used to determine the weights of multiple candidate parallel parameter sets in the current round according to the sorting results and the predetermined weight function.

[0171] According to embodiments of this application, the positioning determination submodule includes a path update submodule, a step size update submodule, a matrix update submodule, and a positioning update submodule.

[0172] The path update submodule is used to update the evolutionary path parameters based on the search center of the current round and the search center of the next round, so as to obtain the updated evolutionary path parameters.

[0173] The step size update submodule is used to update the step size based on the updated evolution path parameters to obtain the updated step size.

[0174] The matrix update submodule is used to update the covariance matrix based on the updated evolution path parameters and multiple candidate parallel parameter sets for the current round, thus obtaining the updated covariance matrix.

[0175] The positioning update submodule is used to obtain spatial positioning parameters based on the updated evolutionary path parameters, the updated step size, and the updated covariance matrix.

[0176] According to an embodiment of this application, the parameter determination submodule includes a parameter sampling unit and a parameter correction unit.

[0177] The parameter sampling unit is used to determine a predetermined number of initial parallel parameter sets to be evaluated for the next round from the search space using a random sampling method.

[0178] The parameter correction unit is used to correct the parameters of multiple initial parallel parameter sets to be evaluated based on the device configuration information, thereby obtaining multiple parallel parameter sets to be evaluated.

[0179] According to an embodiment of this application, the evaluation module 530 includes a decoding submodule, an extraction submodule, an initial evaluation submodule, and an adaptive evaluation submodule.

[0180] The decoding submodule is used to decode the parallel parameters to be decoded in the set of parallel parameters to be evaluated, and obtain the decoded parallel parameters.

[0181] The extraction submodule is used to extract features from device configuration information, model parameters, and decoding parallel parameters to obtain fused features.

[0182] The initial evaluation submodule is used to perform an initial evaluation based on the fusion features to obtain a first evaluation result and a second evaluation result. The first evaluation result represents the training performance of the parallel training strategy corresponding to the parallel parameter set to be evaluated, and the second evaluation result represents the hardware adaptation performance of the parallel training strategy corresponding to the parallel parameter set to be evaluated.

[0183] The adaptive assessment submodule is used to perform an adaptive assessment based on the first and second assessment results to obtain the fitness assessment results.

[0184] According to an embodiment of this application, the parallel training strategy generation device 500 further includes a model evaluation module.

[0185] The model evaluation module is used to evaluate the device configuration information, model parameters, and the set of parallel parameters to be evaluated using an evaluation model, thereby obtaining the fitness evaluation result of the set of parallel parameters to be evaluated. The evaluation model is trained as follows: A training parallel parameter set, a first sample evaluation label, and a second sample evaluation label are obtained. The first sample evaluation label represents the training performance of the parallel training strategy corresponding to the training parallel parameter set, and the second sample evaluation label represents the hardware adaptation performance of the parallel training strategy corresponding to the training parallel parameter set. The training parallel parameter set is input into the initial evaluation model to obtain the first sample evaluation result and the second sample evaluation result. Finally, the parameters of the initial evaluation model are adjusted using the first loss function value determined based on the first sample evaluation label and the first sample evaluation result, and the second loss function value determined based on the second sample evaluation label and the second sample evaluation result, to obtain the evaluation model.

[0186] According to an embodiment of this application, the generation module 540 includes a decoding submodule, a verification submodule, a splitting submodule, and a generation submodule.

[0187] The decoding submodule is used to decode the parameters to be decoded in the set of parallel parameters to be evaluated, and obtain the decoded parallel parameters.

[0188] The verification submodule is used to perform hardware resource constraint verification on the decoding parallel parameters using device configuration information, and obtain the verification result.

[0189] The splitting submodule is used to split the training parameters based on the set of parallel parameters to be evaluated, under the condition that the parallel mode indicated by the verification result characterizes the parallel parameters and meets the hardware constraints indicated by the device configuration information, to obtain a set of training sub-parameters for parallel training. The training parameters include training data and model parameters of the model to be trained.

[0190] The generation submodule is used to adapt the training sub-parameter set to the hardware device based on the device configuration information and the training sub-parameter set, and generate a parallel training strategy.

[0191] According to embodiments of this application, the parallel training strategy generation device 500 further includes an initialization module, a conversion module, a set generation module, and a candidate generation module.

[0192] The initialization module is used to initialize and assign values ​​to the initial parallel parameters.

[0193] The conversion module is used to convert the assigned initial parallel parameters into continuous initial parallel parameters when the type of the parallel parameters is determined to be discrete.

[0194] The set generation module is used to obtain an initialization parallel parameter set based on multiple initialization parallel parameters.

[0195] The candidate generation module is used to perform at least one iteration adjustment on the initial parallel parameter set using device configuration information as constraints to obtain a candidate parallel parameter set.

[0196] According to embodiments of this application, the device configuration information includes at least one of the following: the number of processors, the storage capacity of the storage device, and the network bandwidth of the communication device; the conversion module includes a comparison submodule and a conversion submodule.

[0197] The comparison submodule is used to compare the adaptability of the assigned initial parallel parameters based on the device configuration information and obtain the comparison results.

[0198] The conversion submodule is used to perform data conversion on the assigned initial parallel parameters to obtain the initial parallel parameters, provided that the initial parallel parameters after the comparison result characterization meets the hardware constraints indicated by the device configuration information.

[0199] According to embodiments of this application, any multiple modules among the acquisition module 510, adjustment module 520, evaluation module 530, and generation module 540 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the acquisition module 510, adjustment module 520, evaluation module 530, and generation module 540 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the acquisition module 510, adjustment module 520, evaluation module 530, and generation module 540 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0200] Figure 6 A structural block diagram of a model training apparatus according to an embodiment of this application is shown.

[0201] like Figure 6 As shown, the model training device 600 includes a model training module 610.

[0202] The model training module 610 is used to train the model to be trained on the hardware device using the parallel training strategy corresponding to the target parallel parameter set, so as to obtain the trained model; wherein, the target parallel parameter set is obtained using the parallel training strategy generation method.

[0203] Figure 7 A block diagram of an electronic device suitable for implementing a parallel training policy generation method according to an embodiment of this application is shown.

[0204] like Figure 7As shown, an electronic device 700 according to an embodiment of this application includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0205] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 702 and / or RAM 703. It should be noted that programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in one or more memories.

[0206] According to embodiments of this application, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.

[0207] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0208] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703 described above.

[0209] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the parallel training strategy generation method provided in the embodiments of this application.

[0210] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0211] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0212] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0213] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0214] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0215] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

[0216] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.

Claims

1. A method for generating parallel training strategies, characterized in that, The method includes: Obtain the device configuration information of the hardware device used to train the model to be trained; Using the device configuration information as a constraint, each parallel parameter in the candidate parallel parameter set is adjusted to obtain an updated set of parallel parameters to be evaluated. The parallel parameters are used to indicate the parallelism of at least one of the model parameters, training data, and hardware device in the model to be trained during the training process. Based on the device configuration information and the model parameters, the parallel parameter set to be evaluated is evaluated to obtain the fitness evaluation result of the parallel parameter set to be evaluated. If the fitness evaluation result indicates that the set of parallel parameters to be evaluated meets the training conditions, the model parameters and the training data are adapted to the hardware device based on the set of parallel parameters to be evaluated, and a parallel training strategy for the hardware device to train the model to be trained using the training data is generated.

2. The method according to claim 1, characterized in that, This also includes adjusting the set of parallel parameters to be evaluated in an iterative manner; One adjustment includes: Based on multiple candidate parallel parameter sets in the current round, determine the spatial positioning parameters of the search space used to collect the parallel parameter set in the next round; Based on the spatial positioning parameters, the search space for the next round is determined; From the search space of the next round, determine multiple sets of parallel parameters to be evaluated in the next round.

3. The method according to claim 2, characterized in that, The spatial positioning parameters for determining the search space for collecting the parallel parameter set in the next round, based on multiple candidate parallel parameter sets in the current round, include: Based on the weights of the multiple candidate parallel parameter sets in the current round, the multiple candidate parallel parameter sets in the current round are weighted and summed to obtain the search center of the search space in the next round. Based on the search center of the current round, the search center of the next round, and multiple candidate parallel parameter sets in the current round, spatial positioning parameters characterizing the search space of the next round are determined.

4. The method according to claim 3, characterized in that, The method further includes: Based on the fitness evaluation results of the multiple candidate parallel parameter sets in the current round, the multiple candidate parallel parameter sets in the current round are sorted to obtain a sorting result; and Based on the sorting results and the predetermined weight function, the weights of the multiple candidate parallel parameter sets in the current round are determined.

5. The method according to claim 3, characterized in that, The determination of spatial positioning parameters characterizing the search space of the next round, based on the search center of the current round, the search center of the next round, and multiple candidate parallel parameter sets in the current round, includes: Based on the search center of the current round and the search center of the next round, the evolutionary path parameters are updated to obtain the updated evolutionary path parameters; Based on the updated evolution path parameters, the step size is updated to obtain the updated step size; Based on the updated evolutionary path parameters and the multiple candidate parallel parameter sets for the current round, the covariance matrix is ​​updated to obtain the updated covariance matrix; and The spatial positioning parameters are obtained based on the updated evolutionary path parameters, the updated step size, and the updated covariance matrix.

6. The method according to claim 2, characterized in that, The step of determining the set of multiple parallel parameters to be evaluated in the next round from the search space of the next round includes: Using random sampling, a predetermined number of initial sets of parallel parameters to be evaluated for the next round are determined from the search space; and Based on the device configuration information, the parameters of the multiple initial parallel parameter sets to be evaluated are corrected to obtain the multiple parallel parameter sets to be evaluated.

7. The method according to claim 1 or 2, characterized in that, Based on the device configuration information and the model parameters of the model to be trained, the parallel parameter set to be evaluated is evaluated to obtain the fitness evaluation result of the parallel parameter set to be evaluated, including: The parallel parameters to be decoded in the set of parallel parameters to be evaluated are decoded to obtain the decoded parallel parameters; Feature extraction is performed on the device configuration information, the model parameters, and the decoding parallel parameters to obtain fused features; Based on the fusion features, an initial evaluation is performed to obtain a first evaluation result and a second evaluation result. The first evaluation result represents the training performance of the parallel training strategy corresponding to the set of parallel parameters to be evaluated, and the second evaluation result represents the hardware adaptation performance of the parallel training strategy corresponding to the set of parallel parameters to be evaluated. Based on the first evaluation result and the second evaluation result, an adaptation evaluation is performed to obtain the fitness evaluation result.

8. The method according to claim 1 or 2, characterized in that, The device configuration information, the model parameters, and the set of parallel parameters to be evaluated are evaluated using an evaluation model to obtain the fitness evaluation result of the set of parallel parameters to be evaluated. The evaluation model is trained in the following manner: Obtain a training parallel parameter set, a first sample evaluation label, and a second sample evaluation label, wherein the first sample evaluation label characterizes the training performance of the parallel training strategy corresponding to the training parallel parameter set, and the second sample evaluation label characterizes the hardware adaptation performance of the parallel training strategy corresponding to the training parallel parameter set. The training parallel parameter set is input into the initial evaluation model to obtain the first sample evaluation result and the second sample evaluation result; and The initial evaluation model is adjusted by using the first loss function value determined based on the first sample evaluation label and the first sample evaluation result, and the second loss function value determined based on the second sample evaluation label and the second sample evaluation result, to obtain the evaluation model.

9. The method according to claim 1, characterized in that, The step of adapting the model parameters and the training data to the hardware device based on the set of parallel parameters to be evaluated, and generating a parallel training strategy for the hardware device to train the model to be trained using the training data, includes: The parameters to be decoded in the set of parallel parameters to be evaluated are decoded to obtain the decoding parallel parameters; Using the device configuration information, the hardware resource constraints of the decoding parallel parameters are verified to obtain the verification result; If the verification result indicates that the parallel mode indicated by the decoding parallel parameters meets the hardware constraints indicated by the device configuration information, the training parameters are split based on the set of parallel parameters to be evaluated to obtain a set of training sub-parameters for parallel training. The training parameters include training data and model parameters of the model to be trained. Based on the device configuration information and the training sub-parameter set, the training sub-parameter set is adapted to the hardware device to generate the parallel training strategy.

10. The method according to claim 1, characterized in that, The method further includes: Initialize the initial parallel parameters; If the type of parallel parameters is determined to be discrete, the initial parallel parameters after assignment are converted to obtain the initial parallel parameters of continuous type. Based on the multiple initialization parallel parameters, an initialization parallel parameter set is obtained; and Using the device configuration information as a constraint, the initialization parallel parameter set is adjusted at least once to obtain the candidate parallel parameter set.

11. The method according to claim 10, characterized in that, The device configuration information includes at least one of the following: number of processors, storage capacity of storage devices, and network bandwidth of communication devices; The process of converting the assigned initial parallel parameters to obtain continuous-type initial parallel parameters includes: Based on the device configuration information, the adaptability of the assigned initial parallel parameters is compared to obtain the comparison results. If the comparison result indicates that the assigned initial parallel parameters meet the hardware constraints indicated by the device configuration information, the assigned initial parallel parameters are converted to obtain the initial parallel parameters.

12. A model training method, characterized in that, The method includes: Using the parallel training strategy corresponding to the target parallel parameter set, the model to be trained is trained on the hardware device to obtain the trained model; The target parallel parameter set is obtained using the parallel training strategy generation method as described in any one of claims 1 to 11.

13. An electronic device, comprising: One or more processors; Memory, used to store one or more programs. The feature is that, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 12.

Citation Information

Patent Citations

  • Hyper-parameter determination method and device, computer equipment and storage medium

    CN112529211A

  • Model distributed parallel acceleration training strategy generation method, device and equipment

    CN118797332A

  • Large model parallel training strategy generation method oriented to domestic supercomputing system

    CN119127477A