Model inference method and cloud management platform
Patent Information
- Application Number
- PCT/CN2026/084489
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-24
- Filing Date
- 2026-03-19
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026084489_01102026_PF_FP_ABST
Abstract
Description
A method for model inference and a cloud management platform
[0001] This application claims priority to Chinese Patent Application No. 202510354032.X, filed with the China National Intellectual Property Administration on March 24, 2025, entitled "A Method for Model Reasoning and a Cloud Management Platform", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of machine learning technology, and more specifically, to a method for model inference and a cloud management platform. Background Technology
[0003] "Model inference" refers to the process of predicting input data using a trained model. During inference, the input data undergoes a series of calculations and transformations to ultimately arrive at the model's prediction. The model calculates based on the features of the input data and its parameters to arrive at the prediction. In practical applications, both the speed and accuracy of model inference are crucial factors. To improve the efficiency of model inference, optimization techniques such as model pruning, quantization, and model reduction are commonly employed.
[0004] Typically, "single-model inference service" has a faster inference speed, but low robustness, poor generalization ability, and is prone to overfitting; while "model ensemble inference service" requires multiple models to perform inference each time, has better accuracy of inference results, higher computational resource requirements, slower inference speed, and can only support a lower number of concurrent connections when computational resources are limited.
[0005] Therefore, how to reasonably perform model inference on users' inference tasks and provide users with a good business experience is a technical problem that needs to be solved. Summary of the Invention
[0006] This application provides a model reasoning method that can dynamically adjust the number of reasoning models provided for the user's reasoning tasks based on user information, thereby meeting the model reasoning needs of different reasoning scenarios and providing users with a good business experience.
[0007] Firstly, a method for model inference is provided, applied to a cloud management platform for managing infrastructure providing cloud services. This infrastructure includes at least one computing node, which deploys N inference models, where N is an integer greater than or equal to 2. The method includes: the cloud management platform receiving a first message from a user, the first message carrying information about an inference task and user information, wherein the first message instructs the inference task to be inferred, and the user information includes performance requirements for the inference task; the cloud management platform sending a second message carrying inference results from model integration, the inference results being derived from M inference results obtained by inferring the inference task using M inference models respectively, the M inference models being determined based on the user information, and the M inference models belonging to N inference models, where M is a positive integer greater than or equal to 1 and less than or equal to N.
[0008] In conjunction with the first aspect, in one possible implementation, the performance requirements information for model inference includes the accuracy requirements for model inference and / or the speed requirements for model inference.
[0009] In conjunction with the first aspect, in one possible implementation, M is positively correlated with the accuracy requirements of the inference task.
[0010] In this embodiment, when the accuracy requirement of the inference task is high, the number of inference models corresponding to the inference task increases (i.e., M increases); when the accuracy requirement of the inference task is low, the number of inference models corresponding to the inference task decreases (i.e., M decreases). For example, if the user has a high accuracy requirement, it is determined that an inference service with more models will be provided; for example, if the user has a low accuracy requirement, it is determined that an inference service with fewer models will be provided.
[0011] In conjunction with the first aspect, in one possible implementation, M is negatively correlated with the speed requirements of the inference task.
[0012] In this embodiment, when the speed requirement of the inference task is high, the number of inference models corresponding to the inference task is reduced (i.e., M is reduced); when the speed requirement of the inference task is low, the number of inference models corresponding to the inference task is increased (i.e., M is increased). For example, if the user has high speed requirements, it is determined that an inference service with fewer models or a single-model inference service will be provided. For example, if the user has low accuracy requirements, it is determined that an inference service with more models will be provided.
[0013] In this embodiment of the application, the number of inference models can also be understood as M.
[0014] For example, if a user has high requirements for both accuracy and speed, the number of inference models can be reasonably determined to ensure both the accuracy and speed of the inference results. For example, if a user has low requirements for both accuracy and speed, inference models can be allocated to other users' inference tasks first, and then the number of inference models allocated to that user can be reasonably determined based on the current resource usage.
[0015] In conjunction with the first aspect, in one possible implementation, the method further includes: the cloud management platform determining the information of the inference model corresponding to the inference task based on the user's information, the information of the inference model indicating M inference models, wherein the M inference models are used to perform model inference on the inference task to obtain M inference results; the cloud management platform determining the inference result of model integration based on the M inference results.
[0016] In conjunction with the first aspect, in one possible implementation, the information of the inference model includes the model identifiers corresponding to each of the M inference models.
[0017] In conjunction with the first aspect, in one possible implementation, the information of the inference model includes the number of inference models M and the model identifier corresponding to each of the M inference models.
[0018] Based on the above technical solution, this application embodiment can not only dynamically adjust the number of inference models providing model inference services to each user, but also determine which specific inference models each model is (i.e., determine the model identifier). Typically, each inference model has its own area of expertise or characteristics. Therefore, based on user information, it is possible to determine which inference models can provide inference services to that user, thereby improving the user's business experience.
[0019] In conjunction with the first aspect, in one possible implementation, the M inference models are determined based on user information and resource control information of at least one computing node, which is used to indicate the usage of computing resources of at least one computing node.
[0020] In conjunction with the first aspect, in one possible implementation, the method further includes: a cloud management platform acquiring resource control information of at least one computing node, the resource control information being used to indicate the usage of computing resources of at least one computing node; and the cloud management platform determining information of the inference model corresponding to the inference task based on user information, including: the cloud management platform determining information of the inference model corresponding to the inference task based on user information and the resource control information of at least one computing node.
[0021] Furthermore, in this embodiment, the number of inference models providing inference services to users can be dynamically controlled in conjunction with resource control information, thereby maximizing the utilization of system resources.
[0022] For example, when computing resources are sufficient and speed requirements are met, more inference models can be allocated to users for model inference, thereby improving the accuracy of model inference and ensuring service quality and user business experience.
[0023] For example, when computing resources are scarce but accuracy requirements are met, the number of inference models can be reduced, or the model integration service mode can be switched to a single model service mode, thereby ensuring both model inference accuracy and inference speed, as well as service quality and user business experience.
[0024] For example, when the load is high, a single-model inference can be used. For instance, each user's inference task is assigned only one inference model, prioritizing that the service can meet more concurrent requests.
[0025] In conjunction with the first aspect, in one possible implementation, the method further includes: a cloud management platform receiving M inference response messages from M inference models, each of the M inference response messages carrying the inference result of the inference model for the inference task; and the cloud management platform determining the model ensemble inference result of the model inference based on the M inference response messages.
[0026] In conjunction with the first aspect, in one possible implementation, the M inference models are determined based on user information and user instruction information, which is used to instruct the M inference models.
[0027] In conjunction with the first aspect, in one possible implementation, the M inference models are selected by the user from Q inference models, which are determined based on the user's information. The Q inference models belong to N inference models, where Q is a positive integer greater than or equal to 1 and less than or equal to N, and M is less than or equal to Q.
[0028] In conjunction with the first aspect, in one possible implementation, the method further includes: the cloud management platform sending a third message, the third message indicating Q inference models and their prices, wherein the Q inference models are determined based on user information, the Q inference models belong to N inference models, and Q is a positive integer greater than or equal to 1 and less than or equal to N; the cloud management platform receiving user instruction information, the instruction information indicating M inference models, the M inference models being selected by the user from the Q inference models, where M is less than or equal to Q.
[0029] Based on the above technical solution, in this embodiment of the application, for example, the price corresponding to different numbers of inference models can be displayed to the user. For example, as the number of inference models increases, the cost of model inference will also increase. At this time, users can choose the inference models they need according to their own needs and budget, and then perform inference on the user's inference task based on the M inference models selected by the user, thereby improving the user experience.
[0030] In conjunction with the first aspect, in one possible implementation, the price of the Q inference models is determined based on resource control information of at least one computing node, which indicates the usage of computing resources of at least one computing node.
[0031] Based on the above technical solution, in this embodiment, the price of the Q inference models can be determined based on the usage of computing resources. For example, the price of the Q inference models can vary in different scenarios. For instance, in scenarios where computing resources are scarce, the price of the Q inference models may be higher than the price of inference models under normal conditions. In this case, it can be understood that the price of the inference models is related to the current load and computing power.
[0032] In conjunction with the first aspect, in one possible implementation, the user's information also includes the user's attribute information.
[0033] In this embodiment of the application, a corresponding inference model can also be customized and assigned to the user based on the user's information.
[0034] Secondly, this application proposes a cloud management platform for executing the method described in the first aspect. Specifically, the cloud management platform may include units and / or modules for executing the method proposed in this application, such as a transceiver module and a processing module.
[0035] For example, the cloud management platform can be a server, a server cluster, or the infrastructure that provides cloud services in a cloud service system.
[0036] For example, the cloud management platform can be a virtual instance, such as a virtual machine, a container bare metal server, etc.
[0037] Thirdly, this application provides a cloud management platform. The device includes: at least one processor for executing computer programs or instructions stored in a memory to perform the method described in the first aspect. Optionally, the device further includes a memory for storing the computer programs or instructions. Optionally, the device further includes a communication interface through which the processor reads the computer programs or instructions stored in the memory.
[0038] In one implementation, the cloud management platform is a device used to implement the functions of the above-described methods in the chip.
[0039] In another implementation, the cloud management platform is a chip, chip system, or circuit used to implement the functions described above in the chip.
[0040] Fourthly, this application provides a processor, including: an input circuit, an output circuit, and a processing circuit. The processing circuit is used to receive signals through the input circuit and to transmit signals through the output circuit, causing the processor to execute the method described in the first aspect.
[0041] In specific implementation, the processor can be one or more chips, the input circuit can be input pins, the output circuit can be output pins, and the processing circuit can be transistors, gate circuits, flip-flops, and various logic circuits. The input signal received by the input circuit can be received and input by, for example, but not limited to, a transceiver, and the signal output by the output circuit can be, for example, but not limited to, output to and transmitted by a transmitter. Furthermore, the input circuit and the output circuit can be the same circuit, which is used as both the input circuit and the output circuit at different times. This application does not limit the specific implementation of the processor and various circuits.
[0042] Unless otherwise specified, or if it does not contradict its actual function or internal logic in the relevant description, the transmission and acquisition / reception operations involved in the processor can be understood as processor output and reception, input and other operations, or as transmission and reception operations performed by radio frequency circuits and antennas. This application does not limit them in this regard.
[0043] Fifthly, a processing apparatus is provided, including a processor and a memory. The processor is used to read instructions stored in the memory and to receive signals via a transceiver and transmit signals via a transmitter to execute the method described in the first aspect.
[0044] Optionally, the processor may be one or more, and the memory may be one or more.
[0045] Optionally, the memory may be integrated with the processor, or the memory may be separated from the processor.
[0046] In specific implementation, the memory can be a non-transitory memory, such as read-only memory (ROM), which can be integrated with the processor on the same chip or set on different chips. The embodiments of this application do not limit the type of memory or the way the memory and processor are set.
[0047] It should be understood that the relevant data interaction process, such as sending the first information, can be the process of the processor outputting the first information, and the receiving capability information can be the process of the processor receiving input capability information. Specifically, the data output by the processor can be sent to the transmitter, and the input data received by the processor can come from the transceiver. Here, the transmitter and the transceiver can be collectively referred to as the transceiver.
[0048] The processing device mentioned in the fifth aspect above can be one or more chips. The processor in the processing device can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc.; when implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. This memory can be integrated into the processor or located outside the processor and exist independently.
[0049] In a sixth aspect, a computing cluster is provided, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, such that the computing device cluster performs the method described in any possible implementation of the first aspect.
[0050] Optionally, the processor can be a general-purpose processor, which can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc.; when implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. This memory can be integrated into the processor or located outside the processor and exist independently.
[0051] In a seventh aspect, a computer-readable storage medium is provided, the computer-readable medium storing program code for execution by a computing device, the program code including the method described in the first aspect.
[0052] Eighthly, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform the method described in the first aspect.
[0053] A ninth aspect provides a chip system including a processor for calling and running a computer program from a memory, causing a device equipped with the chip system to perform the method of the first aspect described above. Attached Figure Description
[0054] Figure 1 is a schematic diagram of model training and inference provided in this application.
[0055] Figure 2 is a schematic diagram of a cloud service system architecture applicable to an embodiment of this application.
[0056] Figure 3 is a schematic flowchart of a model reasoning method 300 provided in this application.
[0057] Figure 4 is a schematic diagram of an applicable system architecture.
[0058] Figure 5 is a schematic block diagram of the cloud management platform 500 provided in an embodiment of this application.
[0059] Figure 6 is a schematic block diagram of the cloud management platform 600 provided in an embodiment of this application.
[0060] Figure 7 is a schematic diagram of the architecture of a computing device cluster provided in an embodiment of this application.
[0061] Figure 8 is a schematic diagram of the connection between computing devices 700A and 700B via a network provided in an embodiment of this application. Detailed Implementation
[0062] The technical solutions in this application will now be described with reference to the accompanying drawings.
[0063] To facilitate understanding of the technical solutions provided in the embodiments of this application, the technical terms involved in this application are briefly introduced below. It should be noted that the introduction of technical terms in this application is only for the purpose of helping to understand the technical solutions and should not be construed as limiting the application.
[0064] 1. Ensemble learning
[0065] Ensemble learning is a machine learning technique that improves overall prediction or classification performance by combining multiple models. In other words, it fuses the predictions from multiple trained models to enhance overall predictive ability. The basic idea behind ensemble learning is that the combined performance of multiple models often outperforms a single model because they complement each other's weaknesses. Ensemble learning typically provides better generalization ability than a single model and reduces the risk of overfitting. Its advantages include relatively fast inference speed and lower computational resource consumption. It is suitable for high-concurrency scenarios and situations with relatively limited computing resources.
[0066] The following introduces several typical ensemble learning algorithms:
[0067] (1) Bagging:
[0068] The principle is as follows: Multiple bootstrap samplings with replacement are performed on the original dataset to generate multiple different training datasets. Then, a base model is trained on each training dataset, and the prediction results of each base model are combined using methods such as voting or averaging. For example, a representative algorithm is random forest.
[0069] (2) Boosting:
[0070] The principle is as follows: base learners are trained sequentially, with each base learner adjusted based on the performance of the previous one. Specifically, the boosting algorithm increases the weight of samples misclassified by the previous base learner, causing subsequent base learners to pay more attention to these samples. Finally, the final prediction result is obtained by weightedly combining the predictions of all base learners. A representative algorithm for this is the gradient boosting tree.
[0071] (3) Stacking method:
[0072] The principle is as follows: train multiple different base learners, and then use the prediction results of these base learners as new features to input into a meta-learner for training, so as to obtain the final prediction result.
[0073] The combination strategy of ensemble learning can be understood as how to combine the prediction results of base learners. Common strategies include: (1) Voting method: This refers to voting on the prediction results of multiple models to obtain the final prediction result. Voting can be divided into "hard voting" and "soft voting". "Hard voting" means that when more than half of the votes are for the same prediction result among multiple models, the result is selected as the final prediction result. "Soft voting" means that the prediction results of multiple models are converted into probability values, and then the probability values are weighted and averaged to obtain the final prediction result. Its advantages are high robustness, strong generalization ability and strong resistance to overfitting, which is suitable for tasks with high accuracy requirements. For example, it can be used for classification problems. (2) Averaging method: This refers to taking the average of the prediction results of multiple models to obtain the final prediction result. This method is suitable for situations where the prediction results of multiple models are not significantly different. For example, it can be used for regression problems.
[0074] 2. Model training and inference
[0075] Before any artificial intelligence (AI) model can be used to solve a specific technical problem, it needs to be trained. AI model training refers to using a specified initial model to compute on training data, and then adjusting the parameters of the initial model based on the computation results, allowing the model to gradually learn certain patterns and acquire specific functions. Once trained and possessing stable functionality, the AI model can be used for inference. AI model inference is the process of using the trained AI model to compute on input data and obtain predicted inference results. The most common approach is supervised training of AI models. For example, most deep learning models are trained using supervised training methods. The following section, with reference to Figure 1, illustrates one such model training and inference process.
[0076] As shown in Figure 1, during the training phase, a training set for the deep learning model needs to be constructed based on the objective. The training set includes multiple training data points, each labeled. The label of a training data point represents the correct answer to a specific question, and the label can indicate the objective of training the deep learning model using the training data. For example, to train a deep learning model that can identify different animals, the training set can include images of multiple different animals (i.e., training data). Each image can have a label identifying the type of animal it contains, such as cat or dog. In this example, the type of animal corresponding to each image is the label of that training data.
[0077] When training a deep learning model, training data can be input into the model in batches after parameter initialization. The deep learning model performs calculations (i.e., inference) on the training data to obtain prediction results. The prediction results obtained through inference, along with the corresponding labels of the training data, are used to calculate the loss based on the loss function. The loss function is used during the model training phase to calculate the difference (i.e., the loss value) between the model's prediction results on the training data and the labels of that training data. Loss functions can be implemented using different mathematical functions; commonly used expressions for loss functions include: mean squared error loss function, logarithmic loss function, least squares method, etc.
[0078] The loss value calculated based on the loss function can be used to update the parameters of the deep learning model. The gradient descent method is commonly used for parameter updating. Model training is a repetitive iterative process. Each iteration performs inference on different training data and calculates the loss value. The goal of multiple iterations is to continuously update the parameters of the deep learning model and find the parameter configuration that minimizes or stabilizes the loss value of the loss function.
[0079] During the training phase, to improve training efficiency and post-training model performance, it's necessary to set appropriate hyperparameters. Hyperparameters in deep learning models refer to parameters that cannot be obtained through learning from training data or that cannot be changed by training data; they are a concept relative to the parameters in the model. Hyperparameters of deep learning models are typically set manually based on experience or experiments. These hyperparameters include: learning rate, batch size, and network structure hyperparameters (e.g., number of layers (also called depth), interaction methods between layers, number and size of convolutional kernels, activation functions, etc.). Among these, the learning rate, as a hyperparameter, controls the magnitude of parameter weight updates during training, significantly impacting training speed and accuracy.
[0080] Once trained, a deep learning model can be used to infer from the input data. In the inference phase, data from real-world application scenarios is typically used as input. The trained deep learning model then infers the results. The inference phase is the practical application of the trained deep learning model, allowing for the rapid use of AI capabilities to solve specific technical problems. Today, AI has numerous applications, and the inference capabilities of deep learning models can be used in various scenarios, such as personnel identification in access control and security systems, video content detection (including pornography and violence detection), and express delivery tracking number detection and recognition.
[0081] Generally, "single-model inference service" can be understood as using a single trained model to infer or predict input data, with the output determined solely by the model's predictive ability. "Model ensemble inference service," on the other hand, combines multiple different models to achieve better prediction or decision-making results. However, while "single-model inference service" offers faster inference speed, it suffers from low robustness, poor generalization ability, and is prone to overfitting. "Model ensemble inference service," requiring multiple models for inference, offers better accuracy but demands significant computational resources and is slower, limiting its concurrency to a lower limit when computational resources are limited. Therefore, how to perform model inference effectively and provide a good user experience is a key technical challenge that needs to be addressed.
[0082] In view of this, this application proposes a model inference method that, based on user information, can determine M inference models assigned to a user's inference task. That is, this application can dynamically adjust the number of inference models providing inference services for a user's specified inference task according to the performance requirements of that task, satisfying the model inference needs of different inference scenarios and thus providing users with a better business experience.
[0083] Figure 2 is a schematic diagram of a cloud scenario to which this application applies. As shown in Figure 2, this cloud scenario may include: a cloud management platform 210, the Internet 220, and a client 230. As shown in Figure 2, the cloud management platform 210 is used to manage the infrastructure providing multiple cloud services. The infrastructure includes multiple cloud data centers, each cloud data center including at least one computing node (e.g., a server), and each computing node including cloud service resources to provide corresponding cloud services to tenants. For example, the cloud service resources may be cloud databases.
[0084] The cloud management platform 210 can be located in a cloud data center and can provide access interfaces (such as user interfaces or application program interfaces, APIs). Tenants can use client 230 to remotely access the access interface to register a cloud account and password on the cloud management platform 210 and log in. After successful authentication of the cloud account and password on the cloud management platform 210, the tenant can further select and purchase virtual machines with specific specifications (processor, memory, disk) on the cloud management platform 210. After successful purchase, the cloud management platform 210 provides the remote login account and password for the purchased virtual machine, and client 230 can remotely log in to the virtual machine to install and run the tenant's applications. Therefore, tenants can create, manage, log in to, and operate virtual machines in the cloud data center through the cloud management platform 210.
[0085] The cloud management platform 210 includes, but is not limited to, a tenant console, compute management services, network management services, storage management services, authentication services, and image management services. The tenant console provides an interface or API for interaction with tenants. The compute management services manage servers running virtual machines and containers, as well as bare metal servers. The network management services manage network services (such as gateways and firewalls). The storage management services manage storage services (such as data bucket services). The authentication services manage tenant account passwords. The image management services manage virtual machine images. Tenants use client 230 and can log in to the cloud management platform 210 via the internet 220 to manage their rented cloud services.
[0086] In this embodiment of the application, for example, the at least one computing node deploys N inference models.
[0087] In this embodiment of the application, for example, the steps of the following method 300 can be performed by a cloud management platform in a cloud service system. For example, the infrastructure in the cloud service system deploys N inference models.
[0088] Figure 3 is a schematic flowchart of a model inference method 300 provided in this application. As shown in Figure 3, for example, this method can be executed by the cloud management platform shown in Figure 2. In this embodiment of the application, the input of each of the N inference models is the inference task, and the output of each inference model is the inference result. For example, the N inference models can be pre-trained and deployed on at least one computing node. The method 300 includes:
[0089] 310. The cloud management platform receives the user's first message, which carries information about the inference task and the user's information.
[0090] The first message is used to request model inference, and the user's information includes performance requirements for model inference.
[0091] This can also be understood as the inference task being the content information that the user needs to perform model inference on, and the first message being used to request model inference for this inference task. For example, the inference task could be a text analysis task, an image analysis task, a video analysis task, an audio analysis task, etc. In this embodiment, the specific type of inference task for the user is not limited.
[0092] In this embodiment of the application, the performance requirements information of the model inference includes the accuracy requirements information and / or the speed requirements information of the model inference.
[0093] In this embodiment, "accuracy of model inference" can be understood as the degree of accuracy during inference tasks. Common accuracy evaluation metrics, such as precision, recall, accuracy, and F1 score, are typically obtained by calculating the proportion of correctly predicted samples, which falls within the range of 0 to 1. For example, "precision" can be understood as the ratio of correctly predicted samples to the total number of samples; "precision" can be understood as the ratio of correctly predicted positive samples to all samples predicted as positive; "recall" can be understood as the ratio of correctly predicted positive samples to all samples that are actually positive; and "F1 score" can be understood as the harmonic mean of precision and recall.
[0094] For example, performance requirements for model inference may indicate that the user needs high-precision model inference. For instance, the performance requirements may indicate that the user needs a model inference accuracy of 0.8. Alternatively, performance requirements may indicate that the user needs low-precision model inference. For instance, the performance requirements may indicate that the user needs a model inference accuracy of 0.4.
[0095] In this embodiment of the application, "model inference speed" can be understood as the time required for the model to complete one inference. For example, model inference speed can be understood as the time spent by the model in processing input data and generating prediction results, or it can be understood as the time interval between input and output.
[0096] For example, the model inference requirement information indicates that the user needs fast model inference. For instance, the user can select "Fast Mode" on the interface to indicate the need for fast model inference. Alternatively, the model inference requirement information indicates that the user needs normal-speed model inference. For instance, the user can select "General Mode" on the interface to indicate the need for normal-speed model inference. Or, the user can specify a time range for model inference to indicate the need for fast, normal, or slow model inference, and so on.
[0097] For example, the model inference requirement information indicates to the user that they need high-precision, high-speed model inference. For instance, the user could indicate an accuracy requirement of 0.9 and specify that they need model inference in "fast mode".
[0098] It should be noted that there are many ways for a specific user to indicate the "accuracy requirements of model inference" and the "speed requirements of model inference". The implementation methods listed above are just some examples and do not constitute a limitation.
[0099] In one possible implementation, the user information also includes user attribute information. In this embodiment, the user attribute information is used to describe the user's characteristics. For example, the attribute information includes one or more of the following: (1) basic user information: such as username, registration duration, geographical location, etc.; (2) user behavior information: such as historical information of model inference, performance requirements of model inference (e.g., accuracy requirements of model inference or speed requirements of model inference), etc.; (3) user relationship information: such as membership level, ordinary user, large customer, small customer, new user, old user, active user, etc.
[0100] 320. The cloud management platform sends a second message, which carries the inference results of the model integration.
[0101] The inference results integrated by this model are obtained from M inference results. The M inference results are obtained by M inference models performing model inference on the inference task respectively. The M inference models are determined based on user information. The M inference models belong to N inference models.
[0102] In this embodiment of the application, M is a positive integer greater than or equal to 1 and less than or equal to N.
[0103] In this embodiment of the application, the N inference models may include inference models for functions such as text analysis, image analysis, video analysis, and audio analysis. Therefore, a suitable inference model can always be found for the user to complete the model inference for the user's inference task.
[0104] In one possible implementation, the method 300 further includes: the cloud management platform determining the information of the inference model corresponding to the inference task based on the user's information, the information of the inference model indicating M inference models, wherein the M inference models are used to perform model inference on the inference task message to obtain M inference results; the cloud management platform determining the inference result of model integration based on the M inference results.
[0105] In one possible implementation, M is positively correlated with the accuracy requirements of the inference task. In this embodiment, "positively correlated" can be understood as M increasing as the user's accuracy requirements for the inference task increase, and M decreasing as the user's accuracy requirements for the inference task decrease.
[0106] In this embodiment, the number of inference models can be dynamically controlled based on the user's information. For example, when the user has high accuracy requirements for the inference task, the number of inference models corresponding to the inference task increases; when the user has low accuracy requirements for the inference task, the number of inference models corresponding to the inference task decreases.
[0107] For example, if a user has high accuracy requirements, then it is determined to provide inference services with more models, i.e., increase M; for example, if a user has low accuracy requirements, then it is determined to provide inference services with fewer models or single-model inference services, i.e., decrease M.
[0108] In one possible implementation, M is negatively correlated with the speed requirement of the inference task. In this embodiment, "negative correlation" can be understood as M decreasing as the user's speed requirement for the inference task increases, and M increasing as the user's speed requirement for the inference task decreases.
[0109] In this embodiment, the number of inference models can be dynamically controlled based on the user's information. For example, when the user has a high speed requirement for the inference task, the number of inference models corresponding to the inference task is reduced; when the user has a low speed requirement for the inference task, the number of inference models corresponding to the inference task is increased.
[0110] For example, if a user has high speed requirements, it is determined that an inference service with fewer models or a single model will be provided, i.e., M will be decreased. For example, if a user has low speed requirements, it is determined that an inference service with more models will be provided, i.e., M will be increased.
[0111] In this embodiment of the application, the number of inference models can also be understood as M.
[0112] For example, if a user has high requirements for both accuracy and speed, the number of inference models can be reasonably determined to ensure both the accuracy and speed of the inference results. For example, if a user has low requirements for both accuracy and speed, inference models can be allocated to other users' inference tasks first, and then the number of inference models allocated to that user can be reasonably determined based on the current resource usage.
[0113] Furthermore, in this embodiment, the inference model can also be customized for users based on their attribute information. For example, more inference models can be provided for VIP customers and active users. For instance, the number of inference models can be increased based on user information, i.e., M can be increased. For example, if a user has higher speed requirements and is identified as a VIP user, the accuracy of model inference can be improved as much as possible while maintaining speed. For example, as many inference models as possible can be allocated to this user while ensuring inference speed. For instance, the number of inference models M can be reasonably determined based on user information, ensuring both inference speed and accuracy.
[0114] In this embodiment, determining the model-integrated inference result based on M inference results can be understood as integrating the inference results obtained from each inference model to form the model-integrated inference result, and then sending a second message to the user carrying the integrated model-integrated inference result. For example, a "voting method" or an "averaging method" can be used to integrate the inference results of each inference model to obtain the model-integrated inference result. Other methods can also be used, which are not limited in this application.
[0115] In one possible implementation, the inference model information includes model identifiers corresponding to M inference models. For example, based on the model identifiers of the M inference models in the inference model information, the user's inference task can be assigned to the corresponding M inference models. For example, assuming a total of 10 inference models are deployed on the computing node, and assuming the model identifiers in the determined inference model information are model #1, model #4, and model #5, model inference information can be sent to model #1, model #4, and model #5 respectively.
[0116] For example, considering the inference tasks on each inference model, the inference model serving the current inference task can be flexibly assigned based on the current occupancy of each inference model. For instance, assuming 10 models are deployed on a computing node, and models #1 to #5 are currently providing model inference services for user #1, then for the current inference task of user #2, models #6 to #9, which are in an idle state, can be assigned to user #2. Alternatively, inference models can be randomly assigned to users. For example, assuming models #1 to #10 can all complete user #3's inference task and are all in an idle state, and assuming that based on user #3's information it is determined that 3 inference models are needed, then 3 inference models can be randomly selected to provide model inference services for user #3.
[0117] In one possible implementation, the information of the inference model includes the number M of inference models and the model identifiers corresponding to each of the M inference models. For example, based on the number M of inference models and the model identifiers of the M models in the inference model information, the user's inference task can be assigned to the corresponding M inference models. Assume that the determined inference model information contains 4 inference models, and these 4 inference models are identified as: Inference Model #1, Inference Model #2, Inference Model #4, and Inference Model #6. Assume that a total of 10 inference models are deployed on the computing node, namely: Inference Model #0 to Inference Model #9. In this case, the user's inference task can be sent to Inference Model #1, Inference Model #2, Inference Model #4, and Inference Model #6, respectively, for model inference services.
[0118] This can also be understood as follows: in this embodiment, not only can the number of inference models providing inference services to each user be dynamically adjusted, but the specific inference models themselves can also be determined (i.e., model identifiers are determined). Typically, each inference model has its own strengths or characteristics in inference. Therefore, based on the inference content information carried in each user's inference task information, it is possible to determine which inference models can provide inference services, thereby improving the user's business experience.
[0119] In one possible implementation, the M inference models are determined based on user information and resource control information of at least one computing node, which is used to indicate the usage of computing resources of at least one computing node.
[0120] In one possible implementation, method 300 further includes: a cloud management platform acquiring resource control information of at least one computing node, the resource control information of the at least one computing node being used to indicate the usage of computing resources of the at least one computing node; and the cloud management platform determining information about the inference model corresponding to the inference task based on user information, including: the cloud management platform determining information about the inference model corresponding to the inference task based on user information and the resource control information of the at least one computing node. For example, the resource control information of the at least one computing node can be obtained through real-time monitoring tools and / or system logs.
[0121] In this embodiment of the application, "resource control information" can indicate the usage status of computing resources. For example, the resource control information may include information on computing power resources and current load information.
[0122] For example, "information on computing resources" may include one or more of the following: hardware resources, software resources, and resource availability.
[0123] For example, "hardware resources" may include one or more of the following: (1) Central processing unit (CPU) information: number of cores, clock speed, model, architecture, cache size, etc.; (2) Graphics processing unit (GPU) information: number of cores, video memory size, model, architecture, computing power; (3) Memory information: capacity, type, speed; (4) Storage information: hard disk type, capacity, read / write speed; (5) Network information: bandwidth, latency, network interface type (such as Ethernet).
[0124] For example, "software resources" may include one or more of the following: (1) Operating system information: type (e.g., Linux, Windows), version; (2) Virtualization environment information: number of virtual machines, configuration, resource allocation; (3) Container environment information: number of containers, image information, resource limits.
[0125] For example, "resource availability" can be understood as: which resources are idle, or which resources have been allocated or occupied, or the resource allocation strategy and scheduling rules, etc.
[0126] For example, "current load information" can be understood as real-time data on the workload and resource usage that the system is currently experiencing. This includes, but is not limited to: (1) CPU utilization: the percentage of utilization of each core or the entire CPU. (2) GPU utilization: the percentage of utilization of GPU cores and video memory. (3) Memory utilization: the percentage of memory used relative to total memory. (4) Disk input / output: disk read / write speed and request queue length. (5) Network traffic: the amount of data sent and received by the network interface and bandwidth utilization. (6) Process / thread information: the number of running processes and threads, and their respective resource consumption. (7) Task queue: the number and type of tasks waiting to be executed. (8) System load average: for example, the 1-minute, 5-minute, and 15-minute load averages in Linux systems, reflecting the overall busyness of the system. (9) Latency: the response time of various operations, such as network latency and database query latency. (10) Temperature and power consumption: temperature and power consumption data of hardware components, used to monitor the health status of the system.
[0127] Furthermore, in this embodiment, the number of inference models provided to users for inference services can be dynamically controlled in conjunction with resource control information, thereby maximizing system resource utilization. For example, when computing resources are sufficient and speed requirements are met, more inference models can be allocated to users for model inference, thereby improving the accuracy of model inference and ensuring service quality and user experience. For example, when computing resources are scarce but accuracy requirements are met, the number of inference models can be reduced, or the model integration service mode can be switched to a single model.
[0128] For example, when the load is high, a single-model inference can be used. For instance, only one inference model is assigned to each user's inference task, prioritizing that the service can meet more concurrent requests.
[0129] In one possible implementation, the M inference models are determined based on user information and user instructions, which instruct the M inference models. For example, the instructions may include model identifiers for the M inference models.
[0130] In this embodiment of the application, the M inference models are selected by the user from the Q inference models. The Q inference models are determined based on the user's information. The Q inference models belong to the N inference models. Q is a positive integer greater than or equal to 1 and less than or equal to N. M is less than or equal to Q.
[0131] In this implementation, for example, Q inference models can be initially determined based on user information, and then the user can determine which of the Q inference models should ultimately be used.
[0132] In this embodiment, the number of inference models can be dynamically controlled based on the user's information. For example, if the user has high accuracy requirements, it is determined to provide inference services with more models. Exemplarily, this can be achieved by increasing the number of inference models, i.e., increasing Q, based on the user's information. Conversely, if the user has lower accuracy requirements but higher speed requirements, it is determined to provide inference services with fewer models or single-model inference services. Exemplarily, this can be achieved by decreasing the number of inference models, i.e., decreasing Q, based on the user's information.
[0133] Furthermore, in this embodiment, the inference model can also be customized for users based on their attribute information. For example, more inference models can be provided for VIP customers and active users. For instance, the number of inference models can be increased based on user information, i.e., Q can be increased. For example, if a user has higher speed requirements and is identified as a VIP user, the accuracy of model inference can be improved as much as possible while maintaining speed. For example, as many inference models as possible can be allocated to this user while ensuring inference speed. For instance, the number of inference models Q can be reasonably determined based on user information, ensuring both inference speed and accuracy.
[0134] In one possible implementation, the method 300 further includes: the cloud management platform sending a third message, the third message indicating Q inference models and the prices of the Q inference models, wherein the Q inference models are determined based on user information and belong to N inference models; the cloud management platform receiving user instruction information, the instruction information indicating M inference models, the M inference models being selected by the user from the Q inference models.
[0135] In this embodiment, "the price of Q inference models" can be understood as the total price of Q inference models, or it can be understood as the price of each of the Q inference models.
[0136] In one possible implementation, the price of the Q inference models is determined based on resource control information of at least one computing node, which indicates the usage of computing resources of the at least one computing node.
[0137] In this embodiment, the price of the Q inference models can be determined based on the usage of computing resources. For example, the price of the Q inference models can vary in different scenarios. For instance, in a scenario where computing resources are scarce, the price of the Q inference models may be higher than the price of inference models under normal conditions. In this case, it can be understood that the price of the inference models is related to the current load and computing power.
[0138] Based on the above technical solution, in this embodiment of the application, the Q determined inference models and their corresponding prices can be displayed to the user (for example, as the number of inference models increases, the cost of model inference will also increase). At this time, the user can choose the required M inference models according to the needs and budget of model inference, and then perform inference on the user's inference task based on the M inference models selected by the user, thereby improving the user experience.
[0139] In one possible implementation, method 300 further includes: a cloud management platform receiving M inference response messages from M inference models, each of the M inference response messages carrying the inference result of the inference model for the inference task; and the cloud management platform determining the inference result of model ensemble based on the M inference response messages. For example, an ensemble learning method (e.g., voting) can be used to obtain the final inference result of model ensemble.
[0140] For example, Figure 4 is a schematic diagram of a system architecture applicable to this application. As shown in Figure 4, the system architecture includes a client 410 and a cloud system 420, wherein the cloud system 420 includes an integrated control module 421, an integrated distribution module 422, a model inference module 423, and an integrated decision module 424.
[0141] In conjunction with method 300 described above, in one example, the integration control module 421 can receive a first message from the client 410 and determine the information of the inference model corresponding to the inference task based on the user information in the received first message. The inference model information is used to indicate M inference models. Then, the inference model information and the inference task are sent to the integration distribution module 422. The integration distribution module 422 can send the user's inference task (e.g., the content information the user needs to infer) to the corresponding inference model based on the received inference model information. After each inference model completes its inference, it can send the inference result to the integration decision module 424. The integration decision module 424 obtains the model integration inference result based on the received inference results of each inference model and sends a second message to the client 410, carrying the model integration inference result.
[0142] In conjunction with method 300 above, in another example, the integration control module 421 can receive a first message from client 410 and determine the information of the inference model corresponding to the inference task based on the user information in the received first message. The inference model information is used to indicate Q inference models. Then, the integration control module 421 or other modules determine the price corresponding to the Q inference models and send a third message to client 410, carrying the Q inference models and their corresponding prices. Then, the integration control module 421 can receive instruction information from client 410, which indicates P inference models. The integration control module 421 can send this instruction information to the integration distribution module 422, which can then send the user's inference task (e.g., the content information the user needs to infer) to the corresponding inference model based on this instruction information. After each inference model completes its inference, it can send its inference result to the integration decision module 424. The integration decision module 424 obtains the model integration inference result based on the received inference results of each inference model and sends a second message to client 410, carrying the model integration inference result.
[0143] For example, how the integrated decision module 424 integrates the reasoning results of various reasoning models to form the final model reasoning result can be referred to the "voting method" and "averaging method" described above, and will not be repeated here. For example, the reasoning results of various reasoning models can also be integrated based on other strategies in existing schemes, which is not limited in this application.
[0144] In one possible implementation, the information of the inference model includes pattern identifiers corresponding to each of the M inference models. In another possible implementation, the information of the inference model includes the number of inference models M and the pattern identifiers corresponding to each of the M inference models.
[0145] In one possible implementation, the integrated control module 421 can also acquire resource control information, and then the integrated control module 421 determines the information of the inference model corresponding to the inference task based on the user information and the resource control information.
[0146] It should be noted that the method 300 in Figure 3 and the steps in Figure 4 provided above in this application can be combined, as long as the logic is reasonable. For example, the modules shown in Figure 4 can execute the steps provided in the method 300 above.
[0147] It should be noted that the system architecture shown in Figure 4 is merely illustrative. The names of the modules included in Figure 4 can also be changed, and the modules in Figure 4 can also be integrated together. For example, the integrated control module 421 and the integrated distribution module 422 can be integrated into one module and named differently. It can also be understood that the specific names of the modules or the specific system framework are not limited in the embodiments of this application. As long as the system architecture and modules that can achieve the above method 300 are within the scope of protection claimed in this application.
[0148] It is understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0149] It should also be understood that the ordinal numbers such as "first" and "second" mentioned in the embodiments of this application are used to distinguish multiple objects, and are not used to limit the size, content, order, timing, priority or importance of multiple objects.
[0150] It should also be understood that, in this application, "at least one" means one or more, and "more than one" means two or more. "At least one item" or similar expressions mean one or more items, that is, any combination of these items, including any combination of single items or multiple items. For example, at least one of a, b, or c means: a, b, c, a and b, a and c, b and c, or a and b and c.
[0151] It should also be understood that, in the various embodiments of this application, determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information.
[0152] It should be understood that the various implementations in this application can be combined with each other according to their internal implementation logic.
[0153] Those skilled in the art will recognize that, based on the units and algorithm steps described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is implemented in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0154] This application embodiment can divide the computing device into functional modules according to the above method example. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods. The following description uses the division of functional modules according to each function as an example.
[0155] Figure 5 is a schematic block diagram of a cloud management platform 500 provided in an embodiment of this application. As shown in the figure, the cloud management platform 500 may include a transceiver module 510; optionally, it may also include a processing module 520.
[0156] The modules described above are used to execute the respective steps of the methods mentioned above, which will not be elaborated here.
[0157] It should also be understood that the cloud management platform 500 here is represented in the form of functional units. The term "unit" here can refer to application-specific integrated circuits (ASICs), electronic circuits, processors (e.g., shared processors, proprietary processors, or group processors) and memory for executing one or more software or firmware programs, integrated logic circuits, and / or other suitable components that support the described functions.
[0158] The cloud management platform 500 of each of the above solutions has the function of implementing the corresponding steps of the method 300. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions; for example, a processing module can be replaced by a processor to execute the send / receive operations and related processing operations in each method embodiment. Furthermore, the processing module can also be a processing circuit.
[0159] It should be noted that the cloud management platform in Figure 5 can be the computing device in the aforementioned method embodiments, or it can be the chip or chip system corresponding to the computing device, such as a system-on-a-chip (SoC). The processing module is the processor, microprocessor, or integrated circuit integrated on the chip. No limitation is made here.
[0160] Figure 6 is a schematic block diagram of another cloud management platform 600 provided in an embodiment of this application. As shown, the device 600 includes at least one processor 620. The processor 620 is coupled to a memory and is used to execute instructions stored in the memory to send and / or receive signals. Optionally, the device 600 also includes a memory 630 for storing instructions. Optionally, the device 600 also includes a transceiver 610, and the processor 620 controls the transceiver 610 to send and / or receive signals.
[0161] It should be understood that the processor 620 and memory 630 described above can be combined into a single processing device, with the processor 620 executing the program code stored in the memory 630 to achieve the aforementioned functions. In specific implementations, the memory 630 can be integrated into the processor 620 or independent of the processor 620.
[0162] It should also be understood that transceiver 610 may include a transceiver (or receiver) and a transmitter (or transmitter). The transceiver may further include an antenna, and the number of antennas may be one or more. Transceiver 610 may have a communication interface or interface circuitry.
[0163] Bus 640 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 6, but this does not imply that there is only one bus or one type of bus. Bus 640 can include pathways for transmitting information between various components of computing device 600 (e.g., memory 630, processor 620, transceiver 610).
[0164] The memory 630 stores executable program code, and the processor 620 executes this executable program code to implement the functions of the aforementioned transceiver module and processing module, thereby implementing the method in this embodiment. That is, the memory 630 stores instructions for executing the above-described method 300. For example, the processor 620 executes the computer program or instructions stored in the memory 630 to implement the various steps in the above-described method 300.
[0165] Figure 7 is a schematic diagram of the architecture of a computing device cluster provided in an embodiment of this application. The computing device cluster includes at least one computing device. This computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a desktop computer, a laptop computer, or a smartphone, or other terminal device. As shown in Figure 7, the computing device cluster includes at least one computing device 700. The memory 730 in one or more computing devices 700 in the computing device cluster can store the same instructions for performing the actions executed in the above embodiment 300.
[0166] In some possible implementations, the memory 730 of one or more computing devices 700 in the computing device cluster may also store partial instructions for performing the actions executed in the method 300 described in the above embodiments. In other words, a combination of one or more computing devices 700 can jointly execute instructions for performing the actions executed by the method 300 described in the above embodiments.
[0167] It should be noted that the memory 730 in different computing devices 700 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the computing device 700. That is, the instructions stored in the memory 730 of different computing devices 700 can implement the functions of one or more of the aforementioned transceiver module and processing module.
[0168] Alternatively, the memories 730 in different computing devices 700 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the devices corresponding to the aforementioned devices 500-600. That is, the instructions stored in the memories 730 of different computing devices 700 can implement the functions of one or more modules, such as the transceiver module and the processing module.
[0169] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 8 illustrates one possible implementation, where two computing devices 700A and 700B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device.
[0170] It should be understood that the functions of computing device 700A shown in Figure 8 can also be performed by multiple computing devices 700. Similarly, the functions of computing device 700B can also be performed by multiple computing devices 700.
[0171] The connection method between the computing device clusters shown in Figure 8 can be considered as follows: taking into account that the method provided in this application needs to integrate the inference results of various inference models, the function implemented by the processing module is handed over to the computing device 700B for execution.
[0172] In this embodiment, a computer program product containing instructions is also provided. The computer program product may be software or program products containing instructions capable of running on a computing device cluster or stored on any available medium. When run by the computing device cluster, it causes the computing device cluster to perform the methods provided above, or causes the computing device cluster to implement the functions of the computing devices provided above.
[0173] In this embodiment, a computer-readable storage medium is also provided. This computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that, when executed on a computing device, cause the computing device to perform the method described above.
[0174] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0175] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0176] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0177] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0178] In addition, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0179] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0180] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for model reasoning, characterized in that, The method is applied to a cloud management platform, which manages the infrastructure providing cloud services. The infrastructure includes at least one computing node, and each computing node deploys N inference models, where N is an integer greater than or equal to 2. The method includes: The cloud management platform receives a first message from the user, the first message carrying information about the inference task and the user's information, wherein the first message is used to instruct the inference task to be inferred, and the user's information includes the performance requirements information of the inference task; The cloud management platform sends a second message carrying the inference results of the model integration. The inference results of the model integration are obtained based on M inference results. The M inference results are obtained by M inference models performing inference on the inference task respectively. The M inference models are determined based on the user's information. The M inference models belong to the N inference models. M is a positive integer greater than or equal to 1 and less than or equal to N.
2. The method according to claim 1, characterized in that, The performance requirements of the inference task include the accuracy requirements and / or the speed requirements of the inference task.
3. The method according to claim 2, characterized in that, The value of M is positively correlated with the accuracy requirements of the inference task.
4. The method according to claim 2, characterized in that, The value of M is negatively correlated with the speed requirement of the reasoning task.
5. The method according to any one of claims 1 to 4, characterized in that, The M inference models are determined based on the user's information and the resource control information of the at least one computing node, wherein the resource control information is used to indicate the usage of computing resources of the at least one computing node.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: The cloud management platform determines the information of the inference model corresponding to the inference task based on the user's information. The information of the inference model is used to instruct the M inference models, wherein the M inference models are used to infer the inference task and obtain the M inference results. The cloud management platform determines the inference result of the model integration based on the M inference results.
7. The method according to claim 6, characterized in that, The information of the inference model includes the model identifier corresponding to each of the M inference models.
8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: The cloud management platform sends a third message, which indicates Q inference models and their prices. The Q inference models are determined based on the user's information. The Q inference models belong to the N inference models, and Q is a positive integer greater than or equal to 1 and less than or equal to N. The cloud management platform receives instruction information from the user, which is used to indicate the M inference models. The M inference models are selected by the user from the Q inference models, and M is less than or equal to Q.
9. The method according to claim 8, characterized in that, The price of the Q inference models is determined based on the resource control information of the at least one computing node, which is used to indicate the usage of computing resources of the at least one computing node.
10. The method according to any one of claims 1 to 9, characterized in that, The user's information also includes the user's attribute information.
11. A cloud management platform, characterized in that, The cloud management platform is used to manage the infrastructure providing cloud services. The infrastructure includes at least one computing node, and each computing node deploys N inference models, where N is an integer greater than or equal to 2. The cloud management platform includes a transceiver module. The transceiver module is used to receive a first message from the user, the first message carrying information about the inference task and information about the user, wherein the first message is used to instruct the inference task to be inferred, and the user information includes performance requirement information of the inference task; The transceiver module is used to send a second message to the user. The second message carries the inference results of model integration. The inference results of model integration are obtained based on M inference results. The M inference results are obtained by M inference models performing inference on the inference task respectively. The M inference models are determined based on the user's information. The M inference models belong to the N inference models. M is a positive integer greater than or equal to 1 and less than or equal to N.
12. The cloud management platform according to claim 11, characterized in that, The performance requirements of the inference task include the accuracy requirements and / or the speed requirements of the inference task.
13. The cloud management platform according to claim 12, characterized in that, The value of M is positively correlated with the accuracy requirements of the inference task.
14. The cloud management platform according to claim 12, characterized in that, The value of M is negatively correlated with the speed requirement of the reasoning task.
15. The cloud management platform according to any one of claims 11 to 14, characterized in that, The M inference models are determined based on the user's information and the resource control information of the at least one computing node, which is used to indicate the usage of computing resources of the at least one computing node.
16. The cloud management platform according to any one of claims 11 to 15, characterized in that, The computing device further includes: a processing module, The processing module is used to determine the information of the reasoning model corresponding to the reasoning task based on the user's information. The information of the reasoning model indicates the M reasoning models, wherein the M reasoning models are used to reason about the reasoning task and obtain the M reasoning results. The processing module is used to determine the inference result of the model integration based on the M inference results.
17. The cloud management platform according to claim 16, characterized in that, The information of the inference model includes the model identifier corresponding to each of the M inference models.
18. The cloud management platform according to any one of claims 11 to 17, characterized in that, The transceiver module is also used to send a third message, which is used to indicate Q inference models and the prices of the Q inference models, wherein the Q inference models are determined based on the user's information, the Q inference models belong to the N inference models, and Q is a positive integer greater than or equal to 1 and less than or equal to N; The transceiver module is further configured to receive instruction information from the user, the instruction information being used to indicate the M inference models, the M inference models being selected by the user from the Q inference models, wherein M is less than or equal to Q.
19. The cloud management platform according to claim 18, characterized in that, The price of the Q inference models is determined based on the resource control information of the at least one computing node, which is used to indicate the usage of computing resources of the at least one computing node.
20. The cloud management platform according to any one of claims 11 to 19, characterized in that, The user's information also includes the user's attribute information.
21. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute a computer program or instructions stored in the memory of the at least one computing device, so that the cluster of computing devices performs the method as described in any one of claims 1 to 10.
22. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1 to 10.
23. A computer-readable storage medium, characterized in that, Includes a computer program or instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 10.