A machine learning system

Through the distributed resource management and algorithm framework of the machine learning system, the high development and operation and maintenance costs of multi-scenario applications have been solved, and efficient utilization of hardware resources and cost reduction have been achieved.

CN111353609BActive Publication Date: 2025-09-12PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010127495.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-02-28
Publication Date
2025-09-12
Estimated Expiration
2040-02-28

AI Technical Summary

Technical Problem

Building machine learning applications for various scenarios has the problem of high development and operation costs.

Method used

A machine learning system is provided, including a computing resource module, a machine learning algorithm and framework module, a resource management module, an operation module and a data module. It avoids repeated deployment of cluster environments through the integration of distributed CPU and GPU resources, multiple machine learning algorithms and frameworks, drag-and-drop components for model building, resource scheduling and big data platforms.

Benefits of technology

Effectively avoid waste of hardware resources, reduce development and operation and maintenance costs, and improve computing resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111353609B_ABST
    Figure CN111353609B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of artificial intelligence technology and provides a machine learning system, including: a computing resource module, a machine learning algorithm and framework module, a resource management module, an operation module and a data module; the computing resource module is used to provide computing resources for the machine learning system; the machine learning algorithm and framework module is used to provide machine learning algorithms and frameworks for building machine learning models; the resource management module is used to schedule computing resources; the operation module is used to provide an operation platform for building machine learning models; the data module is used to provide sample data for machine learning models. Based on the big data platform, the machine learning algorithm and framework module is used to provide a variety of machine learning algorithms and frameworks, the operation module is used to build a machine learning model, and the computing resources are scheduled based on the resource management module to train the built machine learning model. There is no need to repeatedly deploy a cluster environment, which effectively avoids waste of hardware resources and reduces development and operation and maintenance costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and in particular to a machine learning system. Background Art

[0002] With the evolution of big data applications, more and more businesses are requiring the construction of machine learning applications for various scenarios based on big data platforms. For example, building a machine learning environment on the Hadoop platform is necessary. However, building machine learning applications for various scenarios requires deploying a separate cluster environment for each type of machine learning application. This results in significant waste of hardware resources and high development and maintenance costs.

[0003] In summary, the current construction of machine learning applications in various scenarios has the problem of high development and operation and maintenance costs. Summary of the Invention

[0004] The embodiments of the present application provide a machine learning system that can solve the problem of high development and operation and maintenance costs in building machine learning applications in various scenarios.

[0005] The embodiment of the present application provides a machine learning system, including a computing resource module, a machine learning algorithm and framework module, a resource management module, an operation module, and a data module;

[0006] The computing resource module is used to provide computing resources for the machine learning system;

[0007] The machine learning algorithm and framework module is used to provide a machine learning algorithm and framework for building a machine learning model;

[0008] The resource management module is used to schedule the computing resources;

[0009] The operation module is used to provide an operation platform for building a machine learning model;

[0010] The data module is used to provide sample data for the machine learning model.

[0011] In a possible implementation, the computing resource module is distributed CPU resources and / or GPU resources.

[0012] It should be understood that based on the above-mentioned computing resource module, it is possible to provide computing power for the machine learning model built by the machine learning system, thereby realizing the training and model evaluation of the machine learning model.

[0013] In one possible implementation, the machine learning algorithm and framework module are encapsulated in the machine learning system through a computing engine.

[0014] For example, commonly used machine learning algorithms and frameworks are encapsulated in the machine learning system using the TensorFlow on YARN (TonY) computing engine, the Spark computing engine, and the like. These machine learning frameworks include, but are not limited to, deep learning frameworks such as TensorFlow, PyTorch, MXNet, and Caffe, distributed machine learning frameworks such as Spark, and lightweight single-cluster learning frameworks such as Python. These machine learning algorithms include, but are not limited to, linear regression, regression tree, logistic regression, support vector machine, decision tree, affine propagation, and clustering algorithms.

[0015] In a possible implementation, the operation module includes an interactive operation unit, a batch operation unit, and an interface operation unit.

[0016] It should be understood that the operation module is an operating platform for the machine learning system to build a machine learning model based on the machine learning task. It can provide draggable components, and build a machine learning module through draggable components. The machine learning algorithm and framework are encapsulated in the above-mentioned draggable components, and the machine learning model is built by dragging the components. For example, a machine learning model can be built by dragging a component encapsulated with the tensorflow framework and a component encapsulated with the decision tree algorithm.

[0017] Furthermore, the interactive operation unit constructs a front-end operating system of the machine learning system based on visualization technology.

[0018] Furthermore, the batch operation unit constructs a batch processing service framework based on a scheduling system.

[0019] Furthermore, the interface operation unit constructs an interface service framework based on Knox technology and Livy technology, and the interface service framework interacts with external systems based on the Hypertext Transfer Protocol.

[0020] In a possible implementation, the resource management module uses a scheduling management mode to uniformly schedule computing resources of the distributed system.

[0021] Furthermore, the resource management module adopts a master-slave mode to implement the scheduling of computing resources of a single cluster.

[0022] In one possible implementation, the resource management module is specifically used to allocate resources based on the resource requirements of the machine learning task and the computing resources of each computing node, and to schedule the machine learning task to the corresponding computing node for execution based on the resource allocation results.

[0023] The beneficial effects of the embodiments of the present application compared with the prior art are: the above-mentioned machine learning system is based on the big data provided by the data module, uses the machine learning algorithm and framework module to provide a variety of machine learning algorithms and frameworks, uses the operation module to build a machine learning model, and schedules computing resources based on the resource management module to train the constructed machine learning model, without the need for repeated deployment of the cluster environment, effectively avoiding waste of hardware resources and reducing development and operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] Figure 1 This is a schematic diagram of the structure of a machine learning system provided in one embodiment of the present application;

[0026] Figure 2 It is a structural diagram of a machine learning system provided in another embodiment of the present application. DETAILED DESCRIPTION

[0027] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0028] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0029] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0030] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0031] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0032] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0033] The machine learning system provided in this embodiment can specifically be a cloud server or a computer cluster system. Individual computers in a cluster system are typically called nodes and are connected via a communication connection such as a local area network, thereby providing faster computing speed, stronger computing power, and system reliability. The computer cluster system can be either a homogeneous or heterogeneous computer cluster system.

[0034] See also Figure 1 ,like Figure 1 A machine learning system is shown, which includes a computing resource module 11, a machine learning algorithm and framework module 12, a resource management module 13, an operation module 14 and a data module 15.

[0035] Specifically, the computing resource module 11 is used to provide computing resources for the machine learning system.

[0036] Specifically, the above-mentioned computing resource module is distributed CPU resources and / or GPU resources, which can provide basic computing capabilities for the machine learning system, and can provide the computing capabilities required for machine learning model training and model evaluation based on distributed CPU resources and / or GPU resources.

[0037] Specifically, the distributed CPU resources and GPU resources may be CPU resources and GPU resources provided by a hybrid cluster system of CPUs and GPUs.

[0038] Specifically, the machine learning algorithm and framework module 12 is used to provide machine learning algorithms and frameworks for building machine learning models. It can provide multiple algorithms and frameworks to support the needs of different application scenarios.

[0039] In this embodiment, the above-mentioned machine learning algorithm and framework module 12 are encapsulated in the machine learning system through a computing engine.

[0040] Specifically, commonly used machine learning algorithms and frameworks are encapsulated within the aforementioned machine learning system through the TensorFlow on YARN (TonY) computing engine and the Spark computing engine. These machine learning frameworks include, but are not limited to, deep learning frameworks such as TensorFlow, pyTorch, MXNet, and Caffe, the distributed machine learning framework SparkMLlib, and the lightweight single-cluster learning framework Python. These machine learning algorithms include, but are not limited to, linear regression, regression tree, logistic regression, support vector machine, decision tree, affine propagation, and clustering algorithms.

[0041] Specifically, the resource management module 13 is used to schedule the computing resources.

[0042] In this embodiment, the resource management module 13 is specifically used to allocate resources according to the resource requirements of the machine learning task and the computing resources of each computing node, and schedule the machine learning task to the corresponding computing node for execution according to the resource allocation results.

[0043] Specifically, the resource management module 13 initializes the underlying computing nodes (devices that provide computing resources) of the machine learning system and then obtains the CPU and GPU resources available for each computing node. The obtained CPU and GPU resources available for each computing node are then fed back to the management node. The management node then allocates computing tasks to each computing node based on the resource requirements of the current machine learning task and schedules each allocated computing task to the corresponding computing node for execution.

[0044] It should be noted that the resource management module 13 divides the nodes into a management node and several computing nodes in advance, and the management node manages and uniformly schedules the computing resources.

[0045] Specifically, the resource management module 13 manages machine learning tasks through a queue management mechanism. After receiving a machine learning task from a user, the resource management module submits the task to a queue according to the task priority. The resource management module then allocates computing resources in the order of the queue. The priority of the machine learning task can be determined based on the permissions of the user who submitted the task, or based on the execution time of the machine learning task, without limitation here.

[0046] It is understandable that for machine learning tasks that require large computing resources, the machine learning tasks can be decomposed, and then computing resources can be allocated according to the subtasks obtained by the decomposition, and each subtask can be scheduled to the corresponding computing node, and the corresponding computing node executes the subtask. It should be noted that in order not to occupy the computing resources of other machine learning tasks, a set amount of resources can be allocated to each subtask. If the computing resources used by a subtask exceed the set amount of resources, the subtask is forced to exit from the computing node to achieve the purpose of not occupying the computing resources of other tasks. It should be noted that the decomposition rules of machine learning tasks can be set according to actual conditions and are not limited here. For example, data preprocessing and model training can be decomposed into two subtasks.

[0047] Specifically, the operation module 14 is used to provide an operation platform for building a machine learning model.

[0048] Specifically, the operation module 14 is an operating platform for the machine learning system to build a machine learning model based on the machine learning task. It can provide draggable components, and build a machine learning module through the draggable components. The machine learning algorithm and framework are encapsulated in the above-mentioned draggable components, and the machine learning model is built by dragging the components. For example, a machine learning model can be built by dragging a component encapsulated with a tensorflow framework and a component encapsulated with a decision tree algorithm.

[0049] It is understood that the operation module 14 is a docking window between the user and the machine learning system. To facilitate the user's use of the machine learning system to build machine learning models, train machine learning modules, and evaluate machine learning models, a variety of docking windows can be provided for the user to use, such as interactive docking windows, batch docking windows, and interface-based docking windows.

[0050] Specifically, the data module 15 is used to provide sample data for the machine learning model.

[0051] In this embodiment, the above-mentioned data module 15 relies on the Hadoop big data platform, and the Hadoop big data platform provides the sample data required for machine learning model training and model evaluation. The above-mentioned sample data includes but is not limited to audio data, video data, image data, text data, etc.

[0052] Specifically, the above data module can be stored in the Hadoop Distributed File System (HDFS), in a data warehouse (HIVE), in an open source database (HBASE), or in a network attached storage (NAS), without limitation herein.

[0053] This embodiment provides a machine learning system that is based on the big data provided by the data module, uses the machine learning algorithm and framework module to provide a variety of machine learning algorithms and frameworks, uses the operation module to build a machine learning model, and schedules computing resources based on the resource management module to train the constructed machine learning model. There is no need to repeatedly deploy a cluster environment, which effectively avoids waste of hardware resources and reduces development and operation and maintenance costs.

[0054] See also Figure 2 , Figure 2 A structural diagram of a machine learning system provided by another embodiment of the present application is shown. The difference between this embodiment and the previous embodiment is that the operation module 14 includes an interactive operation unit 141, a batch operation unit 142 and an interface operation unit 143.

[0055] Specifically, the interactive operation unit 141 constructs a front-end operating system of the machine learning system based on visualization technology.

[0056] Specifically, Apache Zeppelin is used to build a front-end operating system for the machine learning system. The front-end operating system implements operations such as modeling, training, evaluation, and data preprocessing of the machine learning model. The draggable components are set through the interactive front-end system's drag-and-drop component module, and the machine learning algorithm and machine learning framework are encapsulated in the above-mentioned draggable components. The machine learning model required by the user is built by dragging the above-mentioned draggable components, and machine learning tasks are set based on the built machine learning model. The machine learning tasks are added to the queue of the resource management module 13. The resource management module 13 allocates computing resources (underlying CPU and GPU resources) according to the machine learning tasks set by the front-end operating system (including but not limited to model training, model evaluation, data preprocessing, and other computing tasks). It should be noted that the Zeppelin component of Apache Zeppelin can be used as the front-end operating system of the above-mentioned machine learning system. The above-mentioned Zeppelin component can be a B / S architecture system that can support all machine learning frameworks under the TensorFlow on YARN (TonY) computing engine and the spark computing engine, such as sparkMLlib, scikit-learn, TensorFlow, pyTorch, etc.

[0057] Specifically, the batch operation unit 142 constructs a batch processing service framework based on the scheduling system.

[0058] Specifically, the scheduling system is a scheduler system, which packages and uploads the machine learning model code and the execution environment (i.e., the machine learning framework) that the machine learning model relies on to the machine learning system. The scheduler system then triggers scheduling periodically to initiate batch machine learning model training tasks. It should be noted that the scheduler system can also be a B / S architecture system, using the browser window as the front-end operating platform to input batch machine learning tasks, and then the server as the batch processing service framework to respond to the learning tasks.

[0059] Specifically, the interface operation unit 143 constructs an interface service framework based on the Knox technology and the Livy technology, and the interface service framework interacts with the external system based on the Hypertext Transfer Protocol.

[0060] Specifically, the interface-based application approach can use Knox technology to implement file synchronization for machine learning models, supporting the integration of machine learning models trained by the machine learning system with other application systems. The interface to the machine learning system, based on Livy technology, facilitates the deployment of machine learning applications developed by other application systems into the machine learning system of this embodiment, facilitating the utilization of the computing power of the machine learning system provided by this embodiment. Interaction with external systems is achieved through the Hypertext Transfer Protocol.

[0061] In this embodiment, the resource management module uses a scheduling management mode to uniformly schedule the computing resources of the distributed system.

[0062] Specifically, the scheduling of computing resources of a distributed system can be achieved through a scheduling management mode. The above-mentioned scheduling management mode is the YARN mode. After the operation module submits the machine learning task, the resource manager will select a computing node and control the computing node to start the container. The computing node is set as the management node, and the management node requests the computing resources required for calculation from the resource manager. After the resource manager agrees to the request, it allocates computing resources to the management node. The management node then sends the machine learning task to the computing node corresponding to the allocated computing resources for execution based on the allocated computing resources, and obtains the execution results. When the machine learning task is completed, the management node will release the computing resources.

[0063] In this embodiment, the resource management module uses a master-slave mode to implement the scheduling of computing resources of a single cluster.

[0064] Specifically, the standalone mode can be used to schedule small-scale computing resources. Each computing node has a heartbeat mechanism that maintains communication with the resource manager. When a machine learning task is received, the SparkContext object will request computing resources from the resource manager. The resource manager will allocate computing resources based on the heartbeat signals of each computing node and start the scheduling process of the computing node. The SparkContext object will then parse the program code of the machine learning task into a DAG structure and submit it to the DagScheduler. The DAG will be broken down into many steps in the DagScheduler, each step containing multiple tasks. The steps are then submitted to the TaskScheduler, which will assign tasks to computing nodes and submit the assignment status to the scheduling process. The scheduling process will create a thread pool to execute tasks and report the execution status until all tasks are completed and the computing resources are released.

[0065] This embodiment provides multiple operating platforms based on interactive, batch, and interface-based operating units, making it easy for users to use the machine learning system to perform operations such as model building, model training, model evaluation, and data preprocessing. It also enables seamless interoperability between the machine learning system and external systems, thereby fully utilizing the computing resources of the machine learning system. Furthermore, system resource management is performed based on the YANR and Standalone modes, enabling the scheduling of heterogeneous computing resources and standalone computing resources. This allows for full utilization of computing resources, avoids waste of hardware resources, and reduces development and operation costs.

[0066] It should be noted that all or part of the above embodiments can be implemented by instructing the relevant hardware through a computer program. The above computer program can be stored in a computer-readable medium. When the program is executed, it can implement the functions of the above machine learning system. The computer-readable medium can at least include: any entity or device capable of carrying computer program code to the camera / access control system, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, mobile hard disk, magnetic disk, or optical disk.

[0067] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0068] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0069] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0070] Those skilled in the art will appreciate that the units of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0071] In the embodiments provided in this application, it should be understood that the disclosed machine learning system can be implemented in other ways. For example, the machine learning system embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0072] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0073] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A machine learning system, characterized in that include: Computing resource module, machine learning algorithm and framework module, resource management module, operation module and data module; The computing resource module is used to provide computing resources for the machine learning system; The machine learning algorithm and framework module is used to provide machine learning algorithms and frameworks for building machine learning models. The machine learning frameworks include deep learning frameworks such as tensorflow, pytorch, MXNet, and caffe, distributed machine learning frameworks such as spark, and lightweight single-cluster learning frameworks such as python. The machine learning algorithms include linear regression algorithms, regression tree algorithms, logistic regression algorithms, support vector machine algorithms, decision tree algorithms, affine propagation algorithms, and clustering algorithms. The resource management module is used to schedule the computing resources; The operation module is used to provide an operation platform for building a machine learning model; The data module is used to provide sample data for the machine learning model; The operation module includes an interactive operation unit, a batch operation unit and an interface operation unit. The interactive operation unit constructs the front-end operating system of the machine learning system based on visualization technology; a draggable component is set through the drag component module of the interactive front-end system, and the machine learning algorithm and the machine learning framework are encapsulated in the draggable component. The machine learning model required by the user is built by dragging the above-mentioned draggable component, and a machine learning task is set based on the built machine learning model. The machine learning task is added to the queue of the resource management module, and computing resources are allocated according to the machine learning task set by the front-end operating system through the resource management module. The machine learning task includes model training, model evaluation, and data preprocessing; The resource management module first initializes the underlying computing nodes of the machine learning system, obtains the CPU resources and GPU resources available for each computing node, and feeds back the obtained CPU resources and GPU resources available for each computing node to the management node. The management node allocates computing tasks to each computing node according to the resource requirements of the current machine learning task, and schedules each allocated computing task to the corresponding computing node to perform calculations.

2. The machine learning system according to claim 1, wherein: The computing resource module is distributed CPU resources and / or GPU resources.

3. The machine learning system according to claim 1, wherein: The machine learning algorithm and framework module are encapsulated in the machine learning system through a computing engine.

4. The machine learning system according to claim 1, wherein: The batch operation unit constructs a batch processing service framework based on the scheduling system.

5. The machine learning system according to claim 1, wherein: The interface operation unit constructs an interface service framework based on Knox technology and Livy technology, and the interface service framework interacts with external systems based on the Hypertext Transfer Protocol.

6. The machine learning system of claim 1, wherein: The resource management module uses a scheduling management mode to uniformly schedule the computing resources of the distributed system.

7. The machine learning system of claim 1, wherein: The resource management module uses a master-slave mode to implement the scheduling of computing resources of a single cluster.

8. The machine learning system according to any one of claims 1 to 7, wherein: The resource management module is specifically used to allocate resources according to the resource requirements of the machine learning task and the computing resources of each computing node, and to schedule the machine learning task to the corresponding computing node for execution according to the resource allocation results.

Citation Information

Patent Citations

  • Artificial intelligent platform system based on deep learning

    CN108881446A