Configuration methods, devices, electronic equipment, and storage media for machine learning engineering

By unifying and managing the computing power of multiple computing resources, the problem of low efficiency in machine learning engineering caused by insufficient computing resources is solved, and efficient scheduling of computing resources and maximum utilization of computing resources are achieved.

CN115292028BActive Publication Date: 2026-06-30NEW H3C TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NEW H3C TECH CO LTD
Filing Date
2022-06-17
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Insufficient computing power leads to low processing efficiency in machine learning engineering.

Method used

By unifying and managing the computing power of multiple computing resources, a computing power resource selection area is provided to determine the target computing power resources and configure the target operator cores, thereby achieving full lifecycle support for the operator cores and dynamically adjusting computing resources to meet the computing power requirements of machine learning projects.

Benefits of technology

It improves the processing efficiency of machine learning engineering, ensures the maximum utilization of computing resources, supports diverse and unified scheduling of computing resources, and simplifies user operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115292028B_ABST
    Figure CN115292028B_ABST
Patent Text Reader

Abstract

This invention relates to the field of machine learning technology, specifically to a configuration method, apparatus, electronic device, and storage medium for machine learning projects. The method includes, in response to a configuration request for a target machine learning project, displaying a configuration interface for the machine learning project. The configuration interface includes a computing resource selection area for selecting computing resources, which are obtained by managing the computing power of multiple computing resources. In response to a selection operation of computing resources in the computing resource selection area, determining the target computing resources corresponding to the target machine learning project. In response to a configuration operation of an operator core for the target machine learning project, determining the target operator core, and determining the target machine learning project based on the correspondence between the target computing resources and the target operator core. The operator core includes data input, functional operators, and data output. This method improves the processing efficiency of the target machine learning project.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine learning technology, and more specifically to configuration methods, apparatus, electronic devices, and storage media for machine learning projects. Background Technology

[0002] As societal technology advances and the needs and practical applications of artificial intelligence increase, data collection, cleaning, and computation become more frequent. Mature algorithms, data processing techniques, and result analysis are gradually forming general-purpose processing modules. Among these, some commonly used methods and approaches in machine learning are integrated into standardized functional modules, creating modular, convenient operation. After aggregating commonly used functions and establishing standards, these specific functional modules are used to construct the operators for computation—machine learning operators, or simply operators.

[0003] Furthermore, as the complexity of machine learning projects increases, the performance requirements for computing resources responsible for running operators become increasingly demanding. If computing resources are insufficient to provide the computational power for the machine learning project, the processing efficiency will be low. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a configuration method, apparatus, electronic device, and storage medium for machine learning engineering, in order to solve the problem of low processing efficiency of machine learning engineering due to low computing power.

[0005] According to a first aspect, embodiments of the present invention provide a configuration method for a machine learning project, comprising:

[0006] In response to a configuration request for a target machine learning project, a configuration interface for the machine learning project is displayed. The configuration interface includes a computing resource selection area, which is used to provide a selection of computing resources. The computing resources are obtained by managing the computing power of multiple computing resources.

[0007] In response to the selection operation of computing resources in the computing resource selection area for the target machine learning project, a target operator core is determined, and the target computing resources corresponding to the target machine learning project are determined based on the correspondence between the target computing resources and the target operator core;

[0008] In response to the configuration operation of the operator core for the target machine learning project, a target operator core is determined, and the target machine learning project is determined based on the correspondence between the target computing resources and the target operator core. The operator core includes data input, functional operators, and data output.

[0009] The machine learning project configuration method provided in this embodiment of the invention obtains computing power resources by managing the computing power of multiple computing resources. Resource scheduling is achieved through unified management of computing resources. When creating a machine learning project, computing power resources are selected first, that is, the target computing power resources used to provide computing power to the target machine learning project are determined. In other words, the target computing power resources include multiple computing resources, so that the target machine learning project can run with the support of diverse computing power resources, thereby improving the processing efficiency of the target machine learning project.

[0010] In conjunction with the first aspect, in the first embodiment of the first aspect, the method for determining the computing power resources is as follows:

[0011] Obtain preset configuration resources, wherein the preset configuration resources include user-level resource permissions of the computing resources or host resources of the computing resources;

[0012] Configure the computing resources based on the preset configuration resources to determine the available computing resources;

[0013] In response to the computing power management operation of the available computing resources, the computing power resources are determined.

[0014] The machine learning engineering configuration method provided in this embodiment of the invention uses preset configuration resources to select and restrict computing resources, so that the computing power management is processed under conditional constraints, thus ensuring the reliability of computing power resources.

[0015] In conjunction with the first embodiment of the first aspect, in the second embodiment of the first aspect, when the computing resource is a computing cluster resource, determining the computing resource in response to the computing power management operation of the available computing resources includes:

[0016] Displays the interface for managing computing cluster resources;

[0017] In response to the selection operation of available computing clusters in the computing cluster resource management interface, the computing power resources are determined.

[0018] In conjunction with the first embodiment of the first aspect, in the third embodiment of the first aspect, when the computing resource is a container cluster resource, determining the computing resource in response to the computing power management operation of the available computing resource includes:

[0019] Displays the container cluster resource management interface;

[0020] In response to the partition setting operation of the available container clusters in the container cluster resource management interface, the computing resources are determined.

[0021] In conjunction with the first embodiment of the first aspect, in the fourth embodiment of the first aspect, the method further includes:

[0022] When the target computing resources do not meet the computing resource requirements of the target machine learning project, a request for target resource configuration is sent.

[0023] When the application is approved, the preset configuration resources of the target computing power resources are updated based on the target resource configuration to update the target computing power resources.

[0024] The machine learning project configuration method provided in this embodiment of the invention triggers a target resource configuration application when the preset configuration of the target computing power resources does not meet the requirements, so that the updated target computing power resources can meet the resource requirements of the target machine learning project.

[0025] In conjunction with the first aspect, in the fifth embodiment of the first aspect, the step of determining the target operator core in response to a configuration operation for the operator core used in the target machine learning project, and based on the correspondence between the target computing resources and the target operator core, includes:

[0026] In response to the selection operation of the operator kernel for the target machine learning project, the target operator kernel is determined;

[0027] In response to the parameter configuration operation of the target operator kernel, the target parameters of the target operator kernel are determined;

[0028] Based on the correspondence between the target computing resources and the target operator kernel, the target machine learning project is determined.

[0029] The machine learning engineering configuration method provided in this embodiment of the invention supports the entire lifecycle of operator cores based on the configuration of computing resources, and realizes unified configuration and operation logic. Users are not aware of the computing power forms supported by the operator cores, realizing the diversity of computing power support, the diversity of operator cores and the unified orchestration and scheduling, making it more convenient to use and expand.

[0030] In conjunction with the first aspect, in the sixth embodiment of the first aspect, the method further includes:

[0031] In response to the start operation of the target machine learning project, the execution of the target machine learning project is triggered;

[0032] Obtain the operating status of the computing resources in the target computing power resources;

[0033] Based on the operating status, determine the computing resources in the target computing resources that are actually used to provide computing power to the target machine learning project.

[0034] The configuration method for a machine learning project provided in this embodiment of the invention dynamically adjusts the computing resources actually used to provide computing power to the target machine learning project by utilizing the running status of each computing resource during the operation of the target machine learning project. This ensures the maximum utilization of computing resources, enabling the target machine learning project to run with maximum computing power and improving the operating efficiency of the target machine learning project.

[0035] According to a second aspect, embodiments of the present invention also provide a configuration apparatus for a machine learning project, comprising:

[0036] The display module is used to respond to a configuration request for a target machine learning project and display the configuration interface of the machine learning project. The configuration interface includes a computing resource selection area, which is used to provide a selection of computing resources. The computing resources are obtained by managing the computing power of multiple computing resources.

[0037] The first response module is used to respond to the selection operation of computing resources in the computing resource selection area and determine the target computing resources corresponding to the target machine learning project;

[0038] The second response module is used to determine the target operator core in response to the configuration operation of the operator core for the target machine learning project, and to determine the target machine learning project based on the correspondence between the target computing power resources and the target operator core. The operator core includes data input, functional operators and data output.

[0039] According to a third aspect, embodiments of the present invention provide an electronic device, including: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the configuration method of the machine learning project described in the first aspect or any embodiment of the first aspect.

[0040] According to a fourth aspect, embodiments of the present invention provide a computer-readable storage medium storing computer instructions for causing the computer to execute the configuration method of the machine learning engineering described in the first aspect or any embodiment of the first aspect.

[0041] It should be noted that the corresponding beneficial effects of the machine learning engineering configuration device, electronic device and computer-readable storage medium provided in the embodiments of the present invention can be found in the description of the corresponding beneficial effects of the machine learning engineering configuration method above, and will not be repeated here. Attached Figure Description

[0042] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0043] Figure 1 This is a schematic diagram of the operator kernel according to an embodiment of the present invention;

[0044] Figure 2 This is a schematic diagram of the algorithm repository according to an embodiment of the present invention;

[0045] Figure 3 This is a schematic diagram of the operator kernel arrangement according to an embodiment of the present invention;

[0046] Figure 4 This is a flowchart of a machine learning project configuration method according to an embodiment of the present invention;

[0047] Figure 5 This is a schematic diagram of the configuration interface of a machine learning project according to an embodiment of the present invention;

[0048] Figure 6 This is a flowchart of a machine learning project configuration method according to an embodiment of the present invention;

[0049] Figure 7 This is a schematic diagram of preset configuration resources according to an embodiment of the present invention;

[0050] Figure 8 This is a schematic diagram of resource application according to an embodiment of the present invention;

[0051] Figure 9a This is a schematic diagram of the cluster resource management interface according to an embodiment of the present invention;

[0052] Figure 9b This is a schematic diagram of the container resource management interface according to an embodiment of the present invention;

[0053] Figure 10 This is a flowchart of a machine learning project configuration method according to an embodiment of the present invention;

[0054] Figure 11 This is a schematic diagram illustrating the configuration of a machine learning model according to an embodiment of the present invention;

[0055] Figure 12 This is a structural block diagram of a configuration device for a machine learning project according to an embodiment of the present invention;

[0056] Figure 13 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] The configuration method for a machine learning project provided in this embodiment of the invention first manages the computing power of computing resources by centralizing and unifying the computing power of computing resources before creating a new machine learning project. This ensures that the corresponding computing resources can be dynamically called for computing processing during the operation of the machine learning project, rather than being limited to a certain computing node.

[0059] Furthermore, the operator core provided in this embodiment integrates data input, functional operators, and data output, achieving support for the entire lifecycle of the operator core. The operator core is adaptable to various computing resources. Therefore, during the automatic allocation of computing resources and the construction of machine learning projects, users do not need to concern themselves with the specific implementation of the operator core, achieving diverse computing power support, diverse and unified orchestration and scheduling of operator cores, making it more convenient to use and expand.

[0060] For ease of description below, the terms used in the embodiments of the present invention are explained as follows:

[0061] Operator kernel, as described above, such as Figure 1 As shown, this is an independent module containing data input (input), functional operators, and data output (out). The operator core adopts a planned computational pattern, which is conducive to forming a complete and unified computational pattern and provides a foundation for subsequent expansion and arrangement of operator cores. Among them, the expansion of the operator core is used to provide users with custom operator cores, allowing customization of data input, functional operators, and data output according to actual needs.

[0062] An operator, as mentioned above, refers to a functional operator. An operator represents an implementation module containing a complete algorithm or standard function, which can be used for specific algorithms and customized functions. Here, "operator" is a general term for a specific function and does not have a fixed algorithmic structure.

[0063] Data input and data output, i.e. Figure 1 The input and output components are composed of standard interfaces. Within the structure, it can support read and write operations from multiple data sources. This functional module, in conjunction with operators, constructs a complete computational entity, namely the operator core.

[0064] Scalability is enhanced because the input and output of the operator cores, as well as the operators themselves, all have corresponding interface standards. The types and number of operator cores can be customized and pre-built, fully meeting expansion needs. Furthermore, while reusing the input and output modules, operators can be custom-developed independently.

[0065] Operator cores are diverse, and the internal implementation of operator cores varies depending on the computing power supported. For example, Hadoop clusters use Spark tasks by default, and data sources include other big data components such as Hive and HDFS; container clusters use custom containers by default and are configured with limited resources.

[0066] The core computing warehouse is constructed from a series of operator cores and the variable connections between them formed through standard interfaces. The core computing warehouse is a logical visualization used for logical analysis, and also provides orchestration operations, such as visualizing content. Figure 2 As shown, after selecting computing resources, data is transmitted to the core computing warehouse through the standard data source module. All operator cores within the core computing warehouse complete specific connections to meet corresponding functions or requirements. Finally, the entire processing is completed through the underlying resource scheduling module, which allocates resources and performs computations. Figure 2 Taking two types of computing resources as examples, one is Hadoop cluster resources, and the other is container cluster resources. When machine learning projects run within the core computing warehouse, computing resources are dynamically scheduled to maximize computing power support for machine learning projects.

[0067] Specifically, in terms of computing power support, this mainly includes underlying support and resource scheduling. Underlying support provides independent computing power for different operator core orchestrations, ensuring a complete lifecycle and consistent user operation methods; that is, users are unaware of differences in operator cores and resource scheduling methods. Underlying resource scheduling methods differ, and the granularity of invocation varies. Since Hadoop ecosystem computing power is managed in a cluster-based manner by default, it supports authentication for common big data components and user-level resource permission restrictions. Furthermore, based on the characteristics of the Hadoop ecosystem itself, it can achieve specified target space resource size limits. For container cluster computing power, resource configuration supports different levels of differentiation, including host-level hardware resources, host-level resources, and cluster resources.

[0068] The operator kernel is formed by encapsulating data input, functional modules, and data output. Its structural parameters include the following parts:

[0069] User information identification code: A unique identifier for users, used to identify associated resources, such as stored data, computing resource usage, etc.

[0070] Operator core parameter list: Configuration parameters for the corresponding operator core;

[0071] Operator kernel type identifier: used to indicate the types of operators it supports, such as Figure 3 As shown, in Figure 3 The left side represents the computing power operators supported by the Hadoop cluster, and the right side represents the computing power operators of the container cluster. Comparative analysis shows that when displayed on the visualization interface, the different operator core types do not affect the arrangement of operator cores; that is, the user is unaware of the operator type.

[0072] Data source type identification code: A unique path (library / table) identifier used to support data source read operations, corresponding to the data input in the operator core, and supporting the expansion of data source types;

[0073] Operator implementation class: A specific functional module used to implement the functional operators in the operator core;

[0074] Data flow type identification code: A unique path (library / table) identifier used to support storage write operations for newly generated data, support data storage type expansion, and correspond to data output in the operator core.

[0075] The unique paths for the data source type identification code and the data flow type identification code are determined by convention. For example, they can be generated by combining items such as project, task, user, and unique ID.

[0076] like Figure 3 As shown, during the operator core orchestration process, all operator cores are used in a unified manner and with unified operating logic. For operator cores with the same name, the usage and specific parameters show minimal difference to the user under different computing power support. For the constructed machine learning project, computing power resources can be selected, supporting the selection of the underlying computing power type and configuring the same data source. Among these, Figure 3 The dashed outline on the left represents the computing power support of the computing cluster, while the solid outline on the right represents the computing power support of the container cluster.

[0077] According to an embodiment of the present invention, a configuration method for machine learning engineering is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0078] This embodiment provides a configuration method for machine learning engineering, which can be used on electronic devices such as computers and tablets. Figure 4 This is a flowchart of a machine learning project configuration method according to an embodiment of the present invention, such as... Figure 4 As shown, the process includes the following steps:

[0079] S11, in response to a configuration request for the target machine learning project, displays the configuration interface for the machine learning project.

[0080] The configuration interface includes a computing resource selection area, which is used to provide a selection of computing resources. The computing resources are obtained by managing the computing power of multiple computing resources.

[0081] When a user needs to configure a target machine learning project, they interact with an electronic device. The device responds to the configuration request by displaying the machine learning project's configuration interface. Computing resources are obtained by uniformly managing the computing power of multiple allocatable computing resources before creating a new machine learning project. These multiple computing resources include big data cluster computing power or container computing power, etc. Big data cluster computing power can be configured with components such as Spark, Hive, and HDFS, and supports decoupling. Decoupling means that the aforementioned computing resources can be used independently, without relying on the cluster.

[0082] Electronic devices can connect to multiple computing resources, treating them as readily available resources. Alternatively, they can connect to multiple computing resources through third-party devices; for example, the electronic device can connect to a data center, allowing the data center to adjust the frequency of multiple computing resources. No restrictions are placed on the method of allocating multiple computing resources here.

[0083] The aforementioned computing power management involves the unified management of multiple computing resources, resulting in individual computing power resources. Within these resources, dynamic resource allocation is supported. For example... Figure 2 As shown, the Hadoop cluster resources and container cluster resources are displayed respectively. Resource scheduling is implemented within the Hadoop cluster. For example, the computing tasks of hadoop1 are allocated to hadoop2.

[0084] Once the computing resources are configured, subsequent machine learning projects can directly select various computing resources without needing to configure them repeatedly. Of course, if the configured computing resources cannot meet the requirements, they can be reconfigured.

[0085] The specific process of configuring computing resources will be described in detail below.

[0086] When creating a new machine learning model, a configuration interface for the machine learning project is displayed on the electronic device's screen. This configuration interface includes a computing resource selection area for users to choose computing resources. To establish a mapping between computing resources and the newly created machine learning project, the configuration interface also displays an identifier for the machine learning project, such as its name.

[0087] S12, in response to the selection operation of computing resources in the computing resource selection area, determines the target computing resources corresponding to the target machine learning project.

[0088] like Figure 5 As shown, the configuration interface for this machine learning project includes the name of each machine learning project, its corresponding computing resources, and creation time, etc. The computing resource selection area provides a choice of computing resources; users select based on their needs to obtain the computing resources corresponding to each machine learning project. In this embodiment, taking the target machine learning project as an example, through interaction with the computing resource selection area, the electronic device obtains the target computing resources corresponding to the target machine learning project.

[0089] in, Figure 5 The computing cluster resources mentioned above correspond to the computing clusters in the embodiments below, and the platform cluster resources correspond to the container clusters in the embodiments below.

[0090] S13, in response to the configuration operation of the operator kernel for the target machine learning project, the target operator kernel is determined, and the target machine learning project is determined based on the correspondence between the target computing resources and the target operator kernel.

[0091] The operator kernel includes data input, functional operators, and data output.

[0092] Computing resources provide computational support for the operation of the machine learning project, upon which the target machine learning project is constructed. As mentioned above, the operator core includes encapsulation modules for data input, functional operators, and data output. Therefore, the construction of the target machine learning project is achieved by configuring the operator core used for the target machine learning project. The configuration operation of the operator core includes, but is not limited to, assembling each operator core according to the implementation logic of the target machine learning project to obtain the architecture of the target machine learning project. Simultaneously, since the target computing resources for the target machine learning project were configured in S12 above, the target operator cores for the target machine learning project are configured here. Therefore, both the target computing resources and the target operator cores correspond to the target machine learning project. The target computing resources provide computational support to the target machine learning project, while the target operator cores provide specific implementation logic to the target machine learning project. The implementation of this logic depends on the computing support; these two are complementary and both belong to the target machine learning project.

[0093] The specifics of this step will be described in detail below.

[0094] Once the target machine learning project is determined, it can be visualized on the interface of an electronic device.

[0095] The configuration process for the target machine learning project described above mainly includes configuring computing resources and operator cores. After the project is fully built, the electronic device also supports online debugging of the target machine learning project. In case of errors, error alerts are displayed on the interface so that the target machine learning project can be adjusted in a timely manner.

[0096] In some specific implementations, the target machine learning project can be a machine learning model used to achieve a specific function, such as face recognition, object detection, etc. The operator kernel can be a convolutional module, a pooling module, etc. Specific application scenarios are constructed according to actual needs, and no limitations are imposed here.

[0097] The machine learning project configuration method provided in this embodiment obtains computing power resources by managing the computing power of multiple computing resources. Resource scheduling is achieved through unified management of computing resources. When creating a machine learning project, computing power resources are selected first, that is, the target computing power resources used to provide computing power to the target machine learning project are determined. In other words, the target computing power resources include multiple computing resources, so that the target machine learning project can run with the support of diverse computing power resources, thereby improving the processing efficiency of the target machine learning project.

[0098] This embodiment provides a configuration method for machine learning engineering, which can be used on electronic devices such as computers and tablets. Figure 6 This is a flowchart of a machine learning project configuration method according to an embodiment of the present invention, such as... Figure 6 As shown, the process includes the following steps:

[0099] S21, Obtain preset configuration resources.

[0100] The preset configuration resources include user-level resource permissions of the computing resources or host resources of the computing resources.

[0101] The user-level resource permissions for computing resources in the preset configuration resources correspond to the big data cluster and are used to manage user permissions. Specifically, it is used to determine whether the current user has permission to call the corresponding computing resources when configuring computing resources.

[0102] The aforementioned preset configuration resources are pre-set and used to limit the configuration of computing resources. These preset configuration resources can be a default, uniform resource size configuration, or a user-defined resource size configuration. For example, ... Figure 7 As shown, resource size configurations are customized for different users, and the customization results are sent to the users.

[0103] In some alternative implementations, the method further includes:

[0104] (1) When the target computing resources do not meet the computing resource requirements of the target machine learning project, send an application for target resource configuration.

[0105] (2) When the application is approved, the preset configuration resources of the target computing power resources are updated based on the target resource configuration to update the target computing power resources.

[0106] The target computing power resources are used to provide computing power support for the target machine learning project. The determination of whether the target computing power resources meet the computing power requirements of the target machine learning project can be achieved through several methods. First, the user estimates the computational workload of the target machine learning project and determines whether the target computing power resources can provide that workload. If not, the user interacts with the electronic device to fill out an online application form for target resource configuration and submits the application. Second, the target machine learning project can be run based on the target computing power resources, and the efficiency of the running project can be used to determine whether the target computing power resources meet the requirements. The running efficiency is compared with a threshold; if it is lower than the threshold, it indicates a failure to meet the requirements, and a reminder can be displayed on the electronic device's interface to inform the user that a resource application is needed. Alternatively, the electronic device can automatically send a resource application form to the server, and so on. Of course, the methods for determining whether the target computing power resources meet the requirements of the target machine learning project after running it are not limited to the above-mentioned efficiency assessment; other methods can also be used, and no limitations are imposed here.

[0107] When the target computing resources cannot meet the computing resource requirements of the target machine learning project, the user fills out an online application form for target resource configuration and triggers an electronic device to send the application form to the server. The server then provides interactive operations to review the application form. Figure 8 The diagram illustrates the various work order formats. Upon approval, the target resource configuration from the work order is sent to the electronic device. Accordingly, the electronic device updates the preset configuration of the target computing power resources based on the target resource configuration, and then updates the target computing power resources based on the updated preset configuration, thereby providing computing power support for the target machine learning project using the updated target computing power resources. The process of updating the target computing power resources using the updated preset configuration resources is the same as the process of generating computing power resources based on the preset configuration resources described in this embodiment.

[0108] In some specific implementations, the configuration method of this machine learning project relies on an application running on the electronic device. The user interacts with this application, triggering the electronic device to execute the machine learning project configuration method. Each application has a corresponding application server. The application server receives work orders, distributes them to the appropriate personnel for approval, and then sends the approval results back to the electronic device that submitted the work order.

[0109] S22, Configure computing resources based on preset configuration resources and determine available computing resources.

[0110] After determining the preset resource configuration, the electronic device can configure the computing resources to identify the available computing resources that can be used for subsequent computing power management. Alternatively, this can be understood as imposing computing power management restrictions on each computing resource after configuration. For example, computing resource A may have 2GB of memory, but the preset resource configuration limits it to 1GB. Therefore, when managing computing power for computing resource A, only 1GB of memory can be used. The reason for setting preset resource configurations is to avoid using all the computing power of the resources for running machine learning models while neglecting other unprocessed tasks on the computing resources.

[0111] S23, in response to the computing power management operation of available computing resources, determines the computing power resources.

[0112] Different computing power management interfaces are provided for different forms of computing power management. This embodiment describes computing clusters and container clusters as examples. Specifically,

[0113] When the computing resources are computing cluster resources, the above S23 includes:

[0114] (1) Display the interface for managing computing cluster resources.

[0115] (2) In response to the selection operation of available computing clusters in the computing cluster resource management interface, determine the computing resources.

[0116] like Figure 9a As shown, specific Hadoop computing cluster resources are managed in... Figure 9a The interface shown allows users to add or perform other operations on the computing cluster to obtain various computing resources. These computing resources are presented as big data cluster computing power.

[0117] When the computing resource is a container cluster, S23 above includes:

[0118] (1) Display the container cluster resource management interface.

[0119] (2) In response to the partition setting operation of the available container clusters in the container cluster resource management interface, determine the computing resources.

[0120] like Figure 9b As shown, cluster resources on the Kubernetes platform are managed, and available container cluster resources can be partitioned for resource allocation.

[0121] S24, in response to a configuration request for the target machine learning project, displays the configuration interface for the machine learning project.

[0122] The configuration interface includes a computing resource selection area, which is used to provide a selection of computing resources. The computing resources are obtained by managing the computing power of multiple computing resources.

[0123] Please see details Figure 4 S11 of the illustrated embodiment will not be described again here.

[0124] S25, in response to the selection operation of computing resources in the computing resource selection area, determines the target computing resources corresponding to the target machine learning project.

[0125] Please see details Figure 4 S12 of the illustrated embodiment will not be described again here.

[0126] S26, in response to the configuration operation of the operator core for the target machine learning project, the target operator core is determined, and the target machine learning project is determined based on the correspondence between the target computing resources and the target operator core.

[0127] The operator kernel includes data input, functional operators, and data output.

[0128] Please see details Figure 4 S13 of the illustrated embodiment will not be described again here.

[0129] The machine learning engineering configuration method provided in this embodiment uses preset configuration resources to select and restrict computing resources, so that the management of computing power is processed under conditional constraints, thus ensuring the reliability of computing power resources.

[0130] This embodiment provides a configuration method for machine learning engineering, which can be used on electronic devices such as computers and tablets. Figure 10 This is a flowchart of a machine learning project configuration method according to an embodiment of the present invention, such as... Figure 10 As shown, the process includes the following steps:

[0131] S31, in response to a configuration request for the target machine learning project, displays the configuration interface for the machine learning project.

[0132] The configuration interface includes a computing resource selection area, which is used to provide a selection of computing resources. The computing resources are obtained by managing the computing power of multiple computing resources.

[0133] For the process of configuring computing resources, please refer to [link / reference]. Figure 6 The corresponding descriptions of S21-S23 in the illustrated embodiment will not be repeated here.

[0134] For further details, please see Figure 4 S11 of the illustrated embodiment will not be described again here.

[0135] S32, in response to the selection operation of computing resources in the computing resource selection area, determines the target computing resources corresponding to the target machine learning project.

[0136] For further details, please see Figure 4 S12 of the illustrated embodiment will not be described again here.

[0137] S33, in response to the configuration operation on the operator kernel, determines the target machine learning project.

[0138] The operator kernel includes data input, functional operators, and data output.

[0139] Specifically, S33 includes:

[0140] S331, in response to the selection operation of the operator kernel for the target machine learning engineering, determines the target operator kernel.

[0141] Users select operator kernels by interacting with electronic devices. For example... Figure 11 As shown, the target operator kernel is determined by dragging or copying. If only one operator kernel is needed to achieve the target machine learning project, and since the operator kernel is integrated through data input, functional operators, and data output, this operator kernel can be directly selected as the target operator kernel. If multiple operator kernels are needed to achieve the target machine learning project, then the required operator kernels must first be selected, and then the operator kernels with related relationships must be connected according to the processing logic to obtain the connection relationship between the various operator kernels and determine the target operator kernel.

[0142] S332, in response to the parameter configuration operation of the target operator kernel, determines the target parameters of the target operator kernel.

[0143] like Figure 11 As shown, after selecting the target operator kernel, the corresponding parameter configuration page is displayed on the interface. By configuring the parameters on this page, the target machine learning project is determined. In one specific implementation, as... Figure 11 As shown, this description uses cluster resources as an example. The left side represents the core compute warehouse, and the right side represents the parameter configuration. Within the core compute warehouse, operator cores are selected. For example, the core compute warehouse categorizes operator cores based on their characteristics, allowing for direct drag-and-drop usage.

[0144] S333, based on the correspondence between target computing resources and target operator kernels, determines the target machine learning project.

[0145] As mentioned above, the target computing power resources and the target operator core are complementary, and both correspond to the target machine learning project. Therefore, once these two are determined, the target machine learning project can be determined by using their correspondence with the target machine learning project.

[0146] S34, in response to the start operation of the target machine learning project, triggers the execution of the target machine learning project.

[0147] After the operator core orchestration is completed and the corresponding parameter configuration information is correct, the target machine learning project is obtained. The user triggers the execution of the target machine learning project through an interactive operation with the electronic device. The electronic device obtains the operator core orchestration corresponding to the target machine learning project, schedules the corresponding computing power support, enters the running state, and obtains the returned results.

[0148] S35: Obtain the running status of computing resources in the target computing power resources.

[0149] During the operation of the target machine learning project, electronic devices obtain the operating status of computing resources in the target computing power resources in order to dynamically allocate computing resources.

[0150] S36, based on the operating status, determine the computing resources in the target computing resources that are actually used to provide computing power to the target machine learning project.

[0151] Electronic devices utilize the operating status of various computing resources to determine which computing resources are in good operating condition or are idle, and then provide computing power to the target machine learning project.

[0152] The configuration method for the machine learning project provided in this embodiment dynamically adjusts the computing resources actually used to provide computing power to the target machine learning project by utilizing the running status of each computing resource during the operation of the target machine learning project. This ensures the maximum utilization of computing resources, enabling the target machine learning project to run with maximum computing power and improving the operating efficiency of the target machine learning project.

[0153] The machine learning project configuration method provided in this invention determines the corresponding computing resources for each machine learning project through diverse orchestration and scheduling supported by computing power; it also supports the diversity of operator cores and unified orchestration and scheduling. Specifically, this method supports two different forms of computing power: containerized resources and big data computing cluster computing power based on Hadoop clusters. Operator cores are used to support different computing powers, meaning that operator cores with the same function have two types available for orchestration. The selection of computing power has a unified entry point, and the computing warehouses supported by both types of computing power have a unified operation page and operation logic. For the combination and orchestration of operator cores, after selecting a specific operator core on the page, the backend automatically assembles the corresponding type of operator core and starts the corresponding computing process.

[0154] Furthermore, for big data computing clusters, they can exist decoupled, do not strongly depend on the cluster, and support a variety of cluster combinations that meet the requirements; container clusters support the allocation of hardware resources such as partitions, nodes, and specific CPUs and GPUs, as well as personalized resource requests.

[0155] This embodiment also provides a configuration device for machine learning engineering, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0156] This embodiment provides a configuration device for machine learning engineering, such as... Figure 12 As shown, it includes:

[0157] Display module 41 is used to display the configuration interface of the machine learning project in response to the configuration request of the target machine learning project. The configuration interface includes a computing resource selection area, which is used to provide the selection of computing resources. The computing resources are obtained by managing the computing power of multiple computing resources.

[0158] The first response module 42 is used to respond to the selection operation of computing resources in the computing resource selection area and determine the target computing resources corresponding to the target machine learning project;

[0159] The second response module 43 is used to determine the target operator core in response to the configuration operation of the operator core for the target machine learning project, and to determine the target machine learning project based on the correspondence between the target computing power resources and the target operator core. The operator core includes data input, functional operators and data output.

[0160] In some implementations, the method for determining the computing resources includes:

[0161] The first acquisition module is used to acquire preset configuration resources, wherein the preset configuration resources include user-level resource permissions of the computing resources or host resources of the computing resources;

[0162] The configuration module is used to configure the computing resources based on the preset configuration resources and determine the available computing resources;

[0163] The first determining module is used to determine the computing resources in response to the computing power management operation of the available computing resources.

[0164] In some implementations, when the computing resources are computing cluster resources, the first determining module includes:

[0165] The first display unit is used to display the computing cluster resource management interface;

[0166] The first response unit is used to determine the computing resources in response to the selection operation of available computing clusters in the computing cluster resource management interface.

[0167] In some implementations, when the computing resources are container cluster resources, the first determining module includes:

[0168] The second display unit is used to display the container cluster resource management interface;

[0169] The second response module is used to determine the computing resources in response to the partition setting operation of the available container clusters in the container cluster resource management interface.

[0170] In some embodiments, the apparatus further includes:

[0171] The sending module is used to send a request for target resource configuration when the target computing resources do not meet the computing resource requirements of the target machine learning project.

[0172] An update module is used to update the preset configuration resources of the target computing power resources based on the target resource configuration when the application is approved, so as to update the target computing power resources.

[0173] In some implementations, the second response module 43 includes:

[0174] The third response unit is used to determine the target operator kernel in response to the selection operation of the operator kernel for the target machine learning project;

[0175] The fourth response unit is used to determine the target parameters of the target operator core in response to the parameter configuration operation of the target operator core;

[0176] The determining unit is used to determine the target machine learning project based on the correspondence between the target computing power resources and the target operator kernel.

[0177] In some embodiments, the apparatus further includes:

[0178] A triggering module is used to trigger the execution of the target machine learning project in response to a startup operation on the target machine learning project;

[0179] The second acquisition module is used to acquire the operating status of the computing resources in the target computing power resources;

[0180] The second determining module is used to determine, based on the operating state, the computing resources in the target computing resources that are actually used to provide computing power to the target machine learning project.

[0181] In this embodiment, the configuration device for the machine learning project is presented in the form of a functional unit. Here, a unit refers to an ASIC circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above-mentioned functions.

[0182] Further functional descriptions of the above modules are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0183] This invention also provides an electronic device having the above-described features. Figure 12 The configuration device for the machine learning project is shown.

[0184] Please see Figure 13 , Figure 13 This is a schematic diagram of the structure of an electronic device provided in an optional embodiment of the present invention, such as... Figure 13 As shown, the electronic device may include: at least one processor 51, such as a CPU (Central Processing Unit), at least one communication interface 53, memory 54, and at least one communication bus 52. The communication bus 52 is used to enable communication between these components. The communication interface 53 may include a display screen or a keyboard; optionally, the communication interface 53 may also include a standard wired interface or a wireless interface. The memory 54 may be high-speed RAM (Random Access Memory) or non-volatile memory, such as at least one disk storage device. Optionally, the memory 54 may also be at least one storage device located remotely from the aforementioned processor 51. The processor 51 may be combined with... Figure 12 The described apparatus has an application program stored in memory 54, and the processor 51 calls the program code stored in memory 54 to perform any of the above method steps.

[0185] The communication bus 52 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus 52 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 13The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0186] The memory 54 may include volatile memory, such as random-access memory (RAM); the memory may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory 54 may also include a combination of the above types of memory.

[0187] The processor 51 can be a central processing unit (CPU), a network processor (NP), or a combination of CPU and NP.

[0188] The processor 51 may further include a hardware chip. This hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0189] Optionally, memory 54 is also used to store program instructions. Processor 51 can invoke program instructions to implement the configuration method of a machine learning project as shown in any embodiment of this application.

[0190] This invention also provides a non-transitory computer storage medium storing computer-executable instructions that can execute the configuration method of the machine learning project in any of the above method embodiments. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium may also include combinations of the above types of memory.

[0191] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A configuration method of machine learning engineering, characterized by, include: In response to a configuration request for a target machine learning project, a configuration interface for the machine learning project is displayed. The configuration interface includes a computing resource selection area, which is used to provide a selection of computing resources. The computing resources are obtained by managing the computing power of multiple computing resources. In response to the selection operation of computing resources in the computing resource selection area, the target computing resources corresponding to the target machine learning project are determined; In response to the configuration operation of the operator core for the target machine learning project, a target operator core is determined, and the target machine learning project is determined based on the correspondence between the target computing power resources and the target operator core. The operator core includes data input, functional operators, and data output. The method for determining the computing resources is as follows: Obtain preset configuration resources, wherein the preset configuration resources include user-level resource permissions of the computing resources or host resources of the computing resources; Configure the computing resources based on the preset configuration resources to determine the available computing resources; In response to the computing power management operation of the available computing resources, the computing power resources are determined; The step of determining the target operator core in response to a configuration operation for the operator core used in the target machine learning project, and determining the target machine learning project based on the correspondence between the target computing resources and the target operator core, includes: In response to the selection operation of the operator kernel for the target machine learning project, a target operator kernel is determined; In response to the parameter configuration operation of the target operator kernel, the target parameters of the target operator kernel are determined; Based on the correspondence between the target computing resources and the target operator kernel, the target machine learning project is determined.

2. The method according to claim 1, characterized in that, When the computing resource is a computing cluster resource, determining the computing resource in response to the computing power management operation of the available computing resources includes: Displays the interface for managing computing cluster resources; In response to the selection operation of available computing clusters in the computing cluster resource management interface, the computing power resources are determined.

3. The method according to claim 1, characterized in that, When the computing resource is a container cluster resource, determining the computing resource in response to the computing power management operation of the available computing resource includes: Displays the container cluster resource management interface; In response to the partition setting operation of the available container clusters in the container cluster resource management interface, the computing resources are determined.

4. The method according to claim 1, characterized in that, The method further includes: When the target computing resources do not meet the computing resource requirements of the target machine learning project, a request for target resource configuration is sent. When the application is approved, the preset configuration resources of the target computing power resources are updated based on the target resource configuration to update the target computing power resources.

5. The method according to claim 1, characterized in that, The method further includes: In response to the start operation of the target machine learning project, the execution of the target machine learning project is triggered; Obtain the operating status of the computing resources in the target computing power resources; Based on the operating status, determine the computing resources in the target computing resources that are actually used to provide computing power to the target machine learning project.

6. A configuration device for machine learning engineering, characterized in that, include: The display module is used to respond to a configuration request for a target machine learning project and display the configuration interface of the machine learning project. The configuration interface includes a computing resource selection area, which is used to provide a selection of computing resources. The computing resources are obtained by managing the computing power of multiple computing resources. The first response module is used to respond to the selection operation of computing resources in the computing resource selection area and determine the target computing resources corresponding to the target machine learning project; The second response module is used to determine the target operator core in response to the configuration operation of the operator core for the target machine learning project, and to determine the target machine learning project based on the correspondence between the target computing power resources and the target operator core. The operator core includes data input, functional operators and data output. The method for determining the computing resources is as follows: Obtain preset configuration resources, wherein the preset configuration resources include user-level resource permissions of the computing resources or host resources of the computing resources; Configure the computing resources based on the preset configuration resources to determine the available computing resources; In response to the computing power management operation of the available computing resources, the computing power resources are determined; The step of determining the target operator core in response to a configuration operation for the operator core used in the target machine learning project, and determining the target machine learning project based on the correspondence between the target computing resources and the target operator core, includes: In response to the selection operation of the operator kernel for the target machine learning project, a target operator kernel is determined; In response to the parameter configuration operation of the target operator kernel, the target parameters of the target operator kernel are determined; Based on the correspondence between the target computing resources and the target operator kernel, the target machine learning project is determined.

7. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the configuration method of the machine learning project according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to execute the configuration method of the machine learning project according to any one of claims 1-5.

Citation Information

Patent Citations

  • Computing power resource management method and device, storage medium and electronic device

    CN112817751A

  • Computing power resource allocation method and device, electronic equipment and storage medium

    CN114416352A