Vector processing method based on cloud management platform, and cloud management platform

Through the cloud management platform, the state analysis and load balancing selection of computing nodes is realized, resource sharing of different types of computing nodes is solved, the problems of resource waste and high costs in cloud service systems are solved, and resource utilization efficiency is improved.

WO2025153068A1PCT designated stage expired Publication Date: 2025-07-24HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/073036
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-30
Filing Date
2025-01-17
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

In a cloud service system with a separate architecture of storage computing, multiple computing nodes are divided into different types, resulting in a specific type of computing node that can only complete a specific type of vector processing task and cannot share resources, resulting in waste of resources and high costs.

Method used

The status information of multiple computing nodes is analyzed through the cloud management platform, and the load balancing algorithm is used to select the computing node with the best state to execute vector processing requests to realize global sharing of computing node resources.

Benefits of technology

Effectively utilize computing node resources, reduce the cost of cloud service systems, and improve resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025073036_24072025_PF_FP_ABST
    Figure CN2025073036_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a vector processing method based on a cloud management platform, and a cloud management platform, which can realize global sharing of resources of a plurality of computing nodes, thereby more effectively utilizing the resources of the computing nodes, and thus reducing the cost to a certain extent. The method in the present application comprises: after a cloud management platform receives a vector processing request triggered by a tenant, the cloud management platform acquiring, on the basis of the vector processing request, state information of a plurality of computing nodes serving the tenant; after obtaining the state information of the plurality of computing nodes, the cloud management platform analyzing the states of the plurality of computing nodes by using the states of the plurality of computing nodes, so as to select, from among the plurality of computing nodes, a computing node with an optimal state as a first computing node, and sending the vector processing request to the first computing node; and after obtaining the vector processing request, the first computing node executing the vector processing request, so as to complete vector processing.
Need to check novelty before this filing date? Find Prior Art

Description

A vector processing method based on cloud management platform and cloud management platform

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 18, 2024, with application number 202410077190.0, entitled “A method, device and other equipment for task scheduling”, and claims priority to the Chinese patent application filed with the State Intellectual Property Office on April 30, 2024, with application number 202410543817.7, entitled “A vector processing method and cloud management platform based on cloud management platform”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of cloud technology, and in particular to a vector processing method based on a cloud management platform and a cloud management platform. Background Art

[0003] With the rapid development of cloud technology, more and more tenants choose to store their unstructured data in the form of vector representation in the cloud service systems provided by cloud vendors, so that tenants can access and query a certain vector at any time, thereby meeting the tenants' vector processing needs.

[0004] In related technologies, a cloud service system built with a storage-computing separation architecture may include multiple compute nodes and multiple storage nodes. These compute nodes are divided into different types based on their functions. For example, a compute node with vector generation capabilities can convert a tenant's unstructured data into vectors and store them in a storage node. Alternatively, a compute node with vector query capabilities can retrieve the corresponding vector for a tenant from a storage node based on a tenant's request and return it to the tenant.

[0005] In the above system, since multiple computing nodes are divided into different types of computing nodes, a specific type of computing node can only complete a specific type of vector processing task, resulting in the inability to share resources of multiple computing nodes, causing a certain amount of resource waste and causing the overall cost of the system to be too high. Summary of the Invention

[0006] The embodiments of the present application provide a vector processing method and a cloud management platform based on a cloud management platform, which can enable the resources of multiple computing nodes to be shared globally, more effectively utilize the resources of the computing nodes, and thus reduce costs to a certain extent.

[0007] A first aspect of an embodiment of the present application provides a vector processing method based on a cloud management platform. The cloud management platform used to implement the method can manage an infrastructure that provides cloud services to tenants. The infrastructure includes multiple computing nodes and multiple storage nodes that serve the tenants. The multiple storage nodes are used to store multiple data of the tenants. The method includes:

[0008] When the cloud management platform receives a vector processing request triggered by a tenant, the cloud management platform may obtain status information of multiple computing nodes serving the tenant based on the vector processing request.

[0009] After obtaining the status information of multiple computing nodes, the cloud management platform can use a load balancing algorithm to analyze the status information of multiple computing nodes to select the computing node with the best status (i.e., the lowest load) from the multiple computing nodes as the first computing node. It is worth noting that for any computing node among the multiple computing nodes, the computing node is equipped with the ability to process vector processing requests, that is, even if the vector processing request is any one of a vector generation request, a vector query request, a vector addition request, and a vector deletion request, the computing node can execute the vector processing request.

[0010] After determining the first computing node, the cloud management platform may send a vector processing request to the first computing node, so that the first computing node may execute the vector processing request, thereby completing the vector processing.

[0011] For example, when the vector processing request is a vector generation request for multiple data from a tenant, the first computing node may obtain the multiple data from multiple storage nodes to generate multiple vectors corresponding to the multiple data, and store the multiple vectors in the multiple storage nodes, thereby meeting the tenant's vector generation requirements.

[0012] For another example, when a vector processing request is a vector query request from a tenant for a first vector, the first compute node can obtain a third vector that matches the first vector from its own stored second vector or from multiple vectors stored by multiple storage nodes. The multiple vectors stored by the multiple storage nodes include the second vector stored by the first compute node, thereby satisfying the tenant's vector query requirement.

[0013] For another example, when the vector processing request is a vector addition request for first data, the first compute node may obtain the first data from multiple storage nodes to generate a fourth vector corresponding to the first data, and store the fourth vector in the multiple storage nodes. The first data is new data written by the tenant to the multiple storage nodes.

[0014] For another example, when the vector processing request is a vector delete request for second data, the first computing node may delete a fifth vector corresponding to the second data from the multiple vectors stored by the multiple storage nodes, where the second data is data deleted by the tenant from the multiple data stored by the multiple storage nodes.

[0015] The above method demonstrates that, for vector processing requests triggered by tenants, the cloud management platform can obtain status information for each compute node serving the tenant and, based on this information, select the compute node with the best status from among these multiple compute nodes to execute the vector processing request. Regardless of whether the vector processing request is a vector generation request, a vector query request, a vector addition request, or a vector deletion request, from the cloud management platform's perspective, the functions of the multiple compute nodes are identical; each compute node can execute a vector processing request. Therefore, the cloud management platform can select the compute node with the best status from among the multiple compute nodes and call it to execute the vector processing request. This demonstrates that, because the cloud management platform possesses task scheduling capabilities for multiple compute nodes, each compute node can execute various types of vector processing requests, thereby enabling global resource sharing across multiple compute nodes and more efficiently utilizing their resources, thereby reducing the cost of the cloud service system to a certain extent.

[0016] In one possible implementation, a vector increase request or a vector deletion request comes from multiple storage nodes. In the aforementioned implementation, after a tenant writes first data to multiple storage nodes, the multiple storage nodes can generate vector increase requests for the first data by themselves, and send them to the cloud management platform, so that the cloud management platform calls the first computing node to execute the vector increase request, thereby completing the vector increase demand by themselves. Similarly, after a tenant deletes first data from multiple storage nodes, the multiple storage nodes can generate vector deletion requests for the second data by themselves, and send them to the cloud management platform, so that the cloud management platform calls the first computing node to execute the vector deletion request, thereby completing the vector deletion demand by themselves. It can be seen from this that vector increase requests and vector deletion requests can be generated and processed by the cloud management platform, computing nodes, and storage nodes by themselves, without the need for tenants to send requests to the cloud management platform, which can reduce the tenant's operation workload and improve the tenant experience.

[0017] In one possible implementation, a vector generation request is used to instruct: the first computing node to divide multiple vectors into multiple vector groups, each vector group containing a number of vectors equal to a preset value; and the first computing node to write the multiple vector groups into multiple storage nodes, using the vector group as a unit. In the aforementioned implementation, since the tenant has multiple data, the first computing node can generate a vector corresponding to the first data among the multiple data and store the vector, then generate a vector corresponding to the second data and store the vector, and so on. When the number of vectors stored by the first computing node is equal to the preset value, the first computing node can divide the stored vectors into a vector group and store the vector group in a storage node in the object storage pool. By repeating the aforementioned operation, the first computing node can convert multiple data into corresponding vectors and divide the multiple vectors into multiple vector groups and store them in the object storage pool, thereby meeting the tenant's vector generation needs.

[0018] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group. The vector generation request is further used to indicate that, after obtaining multiple vector groups, the first computing node may select a first vector group from the multiple vector groups and store the first vector group; and the first computing node may notify the cloud management platform that the first computing node has stored the first vector group. In the aforementioned implementation, in the process of obtaining multiple vector groups, the first computing node may independently select the vector group to be stored, namely, the first vector group, and store the first vector group. The first computing node may report to the cloud management platform that it has stored the first vector group, so that the cloud management platform records the vector storage status of the first computing node, so that it can subsequently obtain status information of the first computing node and perform an accurate status assessment of the first computing node.

[0019] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group, and the vector query request is used to instruct: the first computing node to detect whether there is a third vector matching the first vector in the first vector group; if the third vector exists in the first vector group, the first computing node to obtain the third vector from the first vector group; or, if the third vector does not exist in the first vector group, the first computing node to obtain a second vector group other than the first vector group from multiple storage nodes, and obtain the third vector from the second vector group, wherein the multiple vector groups are obtained based on the division of multiple vectors. In the aforementioned implementation, when a tenant needs to query the first vector, the first computing node can first detect whether there is a third vector matching the first vector in the first vector group stored by itself. If so, the first computing node can obtain the third vector from the first vector group stored by itself and return it to the tenant. If not, the first computing node can obtain a second vector group other than the first vector group from multiple storage nodes, obtain the third vector from the second vector group, and return it to the tenant, thereby meeting the tenant's vector query requirements.

[0020] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group, and the vector query request is further used to instruct: the first computing node to determine, from the second vector group, a third vector group in which the third vector is located, and to update the first vector group based on the third vector group to obtain an updated first vector group, wherein the updated first vector group includes the third vector group; and the first computing node to notify the cloud management platform that the first computing node stores the updated first vector group. In the aforementioned implementation, since the first computing node does not itself store the third vector group in which the third vector is located, when the first computing node determines that it needs to store the third vector group, the first computing node may replace a vector group in the first vector group with the third vector group, thereby obtaining an updated first vector group, and notify the cloud management platform that it stores the updated first vector group, so that the cloud management platform can update the first computing node's latest vector storage status, thereby subsequently obtaining status information of the first computing node and performing a more accurate status assessment of the first computing node.

[0021] In one possible implementation, a vector increase request is used to instruct: a first computing node to generate a fourth vector corresponding to the first data and store the fourth vector; if the number of vectors stored by the first computing node that are not divided into vector groups is equal to a preset value, the first computing node divides the vectors that are not divided into vector groups into a fourth vector group and writes the fourth vector group to multiple storage nodes, where the vectors that are not divided into vector groups include the fourth vector. In the aforementioned implementation, when it is necessary to increase the fourth vector corresponding to the first data, the first computing node may first obtain the first data from multiple storage nodes, generate the fourth vector corresponding to the first data, and then store the fourth vector. Then, the first computing node may determine whether the number of vectors stored by itself that are not divided into vector groups (including the fourth vector) is equal to the preset value. If so, the first computing node divides these vectors into a fourth vector group and writes the fourth vector group to multiple storage nodes, thereby completing the vector increase by itself.

[0022] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group, and the vector increase request is further used to indicate: the first computing node updates the first vector group based on the fourth vector group to obtain an updated first vector group, and the updated first vector group includes the fourth vector group; the first computing node notifies the cloud management platform that the first computing node stores the updated first vector group. In the aforementioned implementation, since the first computing node itself does not store the newly generated fourth vector group, when the first computing node determines that it needs to store the fourth vector group, the first computing node can replace a vector group in the first vector group with the fourth vector group, thereby obtaining an updated first vector group, and notifying the cloud management platform that it stores the updated first vector group, so that the cloud management platform updates the latest vector storage status of the first computing node, thereby obtaining the status information of the first computing node in the future, and further performing a more accurate status assessment of the first computing node.

[0023] In one possible implementation, multiple vectors are stored in the form of multiple vector groups in multiple storage nodes, and a vector deletion request is used to instruct: a first computing node to obtain a fifth vector group containing a fifth vector corresponding to the second data from the multiple vector groups stored by the multiple storage nodes, and to delete the fifth vector group from the multiple vector groups stored by the multiple storage nodes; the first computing node to generate a sixth vector group based on a sixth vector other than the fifth vector in the fifth vector group, and to store the sixth vector group in the multiple storage nodes. In the aforementioned implementation, when it is necessary to delete the fifth vector corresponding to the second data, the first computing node may determine the fifth vector group containing the fifth vector from the multiple vector groups stored by the multiple storage nodes, and delete the fifth vector group. In addition, the first computing node may also generate a sixth vector group based on a sixth vector other than the fifth vector in the fifth vector group, and write the sixth vector group to the multiple storage nodes, thereby completing vector deletion and recycling on its own.

[0024] In one possible implementation, the vector deletion request is further used to instruct: the first computing node to update the first vector group based on the sixth vector group to obtain an updated first vector group, wherein the updated first vector group includes the sixth vector group; and the first computing node to notify the cloud management platform that the first computing node stores the updated first vector group. In the aforementioned implementation, since the first computing node itself does not store the newly generated sixth vector group, when the first computing node determines that it needs to store the sixth vector group, the first computing node may replace a vector group in the first vector group with the sixth vector group, thereby obtaining an updated first vector group, and notifying the cloud management platform that it stores the updated first vector group, so that the cloud management platform updates the first computing node's latest vector storage status, thereby subsequently obtaining the first computing node's status information, and thereby performing a more accurate status assessment of the first computing node.

[0025] In one possible implementation, the cloud management platform determines the first computing node with the lowest load from multiple computing nodes according to a load balancing algorithm based on a vector processing request and status information of multiple computing nodes, including: the cloud management platform calculates the status information of multiple computing nodes according to a load balancing algorithm based on the vector processing request to obtain evaluation values ​​of the multiple computing nodes, and the evaluation values ​​of the multiple computing nodes are used to indicate the loads of the multiple computing nodes; the cloud management platform determines the computing node with the smallest evaluation value as the first computing node from the multiple computing nodes; wherein, for any one of the multiple computing nodes, the status information of the computing node includes at least one of the following: the time required for the computing node to load a vector stored in itself, the time required for the computing node to load an index of a vector stored in itself, the time required for the computing node to load a vector stored in multiple storage nodes, the time required for the computing node to load an index of a vector stored in multiple storage nodes, and the time required for the computing node to execute the vector processing request. In the aforementioned implementation, for any one of the multiple computing nodes, the cloud management platform can calculate the state information of the computing node, such as the time required for the computing node to load the vector stored by itself, the time required for the computing node to load the index of the vector stored by itself, the time required for the computing node to load the vectors stored by multiple storage nodes, the time required for the computing node to load the index of the vectors stored by multiple storage nodes, and the time required for the computing node to execute a vector processing request, through a load balancing algorithm, to obtain an evaluation value of the computing node. The same is true for the remaining computing nodes, so the cloud management platform can obtain evaluation values ​​of multiple computing nodes. Since the smaller the evaluation value, the better the state of the computing node (i.e., the lower the load of the computing node), the cloud management platform can select the computing node with the smallest evaluation value as the first computing node to efficiently execute the vector processing request.

[0026] In one possible implementation, the method further includes at least one of the following: the cloud management platform analyzes the vectors stored by the computing node and the vectors stored by multiple storage nodes to obtain the time required for the computing node to load the vectors stored by itself and the time required for the computing node to load the vectors stored by multiple storage nodes; the cloud management platform analyzes the indexes of the vectors stored by the computing node and the indexes of the vectors stored by multiple storage nodes to obtain the time required for the computing node to load the indexes stored by itself and the time required for the computing node to load the indexes stored by multiple storage nodes; the cloud management platform analyzes the resource usage of the computing node to obtain the time required for the computing node to execute the vector processing request. In the aforementioned implementation, since the cloud management platform has a vector status table, a vector index status table, and a computing node status table, the vector status table records the vectors stored by each computing node, the vector index status table records the indexes of the vectors stored by each computing node, and the computing node status table records the resource usage of each computing node, the cloud management platform can use these three tables for analysis to accurately obtain the status information of each computing node, and then accurately evaluate each computing node.

[0027] A second aspect of an embodiment of the present application provides a cloud management platform, which is used to manage an infrastructure that provides cloud services. The infrastructure includes multiple computing nodes and multiple storage nodes, and the multiple storage nodes are used to store multiple data of tenants. The cloud management platform includes: a receiving module, which is used to receive a vector generation request for multiple data from a tenant; a processing module, which is used to determine a first computing node with the lowest load from the multiple computing nodes according to a load balancing algorithm based on the vector generation request and status information of the multiple computing nodes, and the multiple computing nodes are all equipped with the ability to process vector generation requests; the processing module is also used to send a vector generation request to the first computing node, and the vector generation request is used to instruct the first computing node to generate multiple vectors corresponding to the multiple data and store the multiple vectors in multiple storage nodes.

[0028] In one possible implementation, the receiving module is further configured to receive a vector query request for a first vector from a tenant; the processing module is further configured to determine, based on the vector query request and status information of multiple computing nodes, a first computing node with the lowest load from among the multiple computing nodes according to a load balancing algorithm, wherein the multiple computing nodes are all configured with the ability to process vector query requests; and the processing module is further configured to send a vector query request to the first computing node, wherein the vector query request is configured to instruct the first computing node to obtain a third vector that matches the first vector from a second vector stored in the first computing node or from multiple vectors stored by multiple storage nodes, where the multiple vectors include the second vector.

[0029] In one possible implementation, the receiving module is further used to receive a vector increase request for first data, where the first data is new data written by a tenant to multiple storage nodes; the processing module is further used to determine, based on the vector increase request and status information of the multiple computing nodes, a first computing node with the lowest load from the multiple computing nodes according to a load balancing algorithm, where the multiple computing nodes are all capable of processing the vector increase request; the processing module is further used to send a vector increase request to the first computing node, where the vector increase request is used to instruct the first computing node to generate a fourth vector corresponding to the first data and store the fourth vector in the multiple storage nodes.

[0030] In one possible implementation, the receiving module is further configured to receive a vector deletion request for second data, where the second data is data deleted by the tenant from multiple data stored by multiple storage nodes. The processing module is further configured to determine, based on the vector deletion request and status information of the multiple computing nodes, a first computing node with the lowest load from the multiple computing nodes according to a load balancing algorithm, where all the multiple computing nodes are capable of processing the vector deletion request. The processing module is further configured to send a vector deletion request to the first computing node, where the vector deletion request is configured to instruct the first computing node to delete a fifth vector corresponding to the second data from the multiple vectors stored by the multiple storage nodes.

[0031] In a possible implementation, the vector addition request or the vector deletion request comes from multiple storage nodes.

[0032] In one possible implementation, the vector generation request is used to instruct: a first computing node to divide a plurality of vectors into a plurality of vector groups, where the number of vectors included in each vector group is equal to a preset value; and the first computing node to write the plurality of vector groups into the plurality of storage nodes respectively in units of vector groups.

[0033] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group. The vector generation request is further used to indicate that: after obtaining multiple vector groups, the first computing node may select a first vector group from the multiple vector groups and store the first vector group; and the first computing node notifies the cloud management platform that the first computing node has stored the first vector group.

[0034] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group, and the vector query request is used to instruct: the first computing node to detect whether there is a third vector matching the first vector in the first vector group; if the third vector exists in the first vector group, the first computing node to obtain the third vector from the first vector group; or, if the third vector does not exist in the first vector group, the first computing node to obtain a second vector group other than the first vector group from multiple vector groups from multiple storage nodes, and obtain the third vector from the second vector group, wherein the multiple vector groups are obtained based on the division of the multiple vectors.

[0035] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group. The vector query request is further used to instruct: the first computing node to determine, from the second vector group, a third vector group in which the third vector is located, and to update the first vector group based on the third vector group to obtain an updated first vector group, where the updated first vector group includes the third vector group; and the first computing node to notify the cloud management platform that the updated first vector group is stored in the first computing node.

[0036] In one possible implementation, the vector add request is used to instruct: a first computing node to generate a fourth vector corresponding to the first data and store the fourth vector; if the number of vectors not divided into vector groups stored by the first computing node is equal to a preset value, the first computing node to divide the vectors not divided into vector groups into a fourth vector group, and write the fourth vector group to multiple storage nodes, where the vectors not divided into vector groups include the fourth vector.

[0037] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group, and the vector increase request is further used to indicate: the first computing node updates the first vector group based on the fourth vector group to obtain an updated first vector group, where the updated first vector group includes the fourth vector group; and the first computing node notifies the cloud management platform that the updated first vector group is stored in the first computing node.

[0038] In one possible implementation, multiple vectors are stored in the form of multiple vector groups in multiple storage nodes. The vector deletion request is used to instruct: a first computing node to obtain a fifth vector group containing a fifth vector corresponding to the second data from the multiple vector groups stored in the multiple storage nodes, and delete the fifth vector group from the multiple vector groups stored in the multiple storage nodes; and the first computing node to generate a sixth vector group based on a sixth vector excluding the fifth vector in the fifth vector group, and store the sixth vector group in the multiple storage nodes.

[0039] In one possible implementation, the vector deletion request is further used to instruct: the first computing node to update the first vector group based on the sixth vector group to obtain an updated first vector group, where the updated first vector group includes the sixth vector group; and the first computing node to notify the cloud management platform that the first computing node stores the updated first vector group.

[0040] In one possible implementation, a processing module is configured to: generate a request based on a vector, calculate status information of multiple computing nodes according to a load balancing algorithm, and obtain evaluation values ​​of the multiple computing nodes, where the evaluation values ​​of the multiple computing nodes are used to indicate the loads of the multiple computing nodes; and determine, from the multiple computing nodes, a computing node with a smallest evaluation value as a first computing node; wherein, for any one of the multiple computing nodes, the status information of the computing node includes at least one of the following: the time required for the computing node to load a vector stored in itself, the time required for the computing node to load an index of a vector stored in itself, the time required for the computing node to load vectors stored in multiple storage nodes, the time required for the computing node to load indexes of vectors stored in multiple storage nodes, and the time required for the computing node to execute a vector processing request.

[0041] In one possible implementation, the cloud management platform further includes: an analysis module, the analysis module being configured to perform at least one of the following: analyzing vectors stored by a computing node and vectors stored by multiple storage nodes to obtain the time required for the computing node to load the vectors stored by itself and the time required for the computing node to load the vectors stored by multiple storage nodes; analyzing the indexes of the vectors stored by the computing node and the indexes of the vectors stored by multiple storage nodes to obtain the time required for the computing node to load the indexes stored by itself and the time required for the computing node to load the indexes stored by multiple storage nodes; and analyzing the resource usage of the computing node to obtain the time required for the computing node to execute a vector processing request.

[0042] A third aspect of an embodiment of the present application provides a computing device cluster, which includes at least one computing device, each computing device including a processor and a memory: the memory is used to store instructions; the processor is used to enable the computing device cluster to execute the method described in the first aspect or any possible implementation method of the first aspect according to the instructions.

[0043] A fourth aspect of an embodiment of the present application provides a computer storage medium storing one or more instructions, which, when executed by one or more computers, enables the one or more computers to implement the method described in the first aspect or any possible implementation method of the first aspect.

[0044] A fifth aspect of the embodiments of the present application provides a computer program product, which stores instructions. When the instructions are executed by a computer, the computer implements the method described in the first aspect or any possible implementation method of the first aspect.

[0045] In an embodiment of the present application, for a vector processing request triggered by a tenant, the cloud management platform can obtain the status information of each computing node serving the tenant, and based on this, select the computing node with the best status from these multiple computing nodes to execute the vector processing request. Regardless of whether the vector processing request is a vector generation request, a vector query request, a vector addition request, or a vector deletion request, from the perspective of the cloud management platform, the functions of the multiple computing nodes are the same, and each computing node can execute the vector processing request. Therefore, the cloud management platform can select a computing node with the best status from multiple computing nodes to call the computing node to execute the vector processing request. It can be seen that since the cloud management platform has a task scheduling function for multiple computing nodes, each computing node can execute various types of vector processing requests, thereby enabling the resources of multiple computing nodes to be shared globally, more effectively utilizing the resources of the computing nodes, and thus reducing the cost of the cloud service system to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] FIG1 is a schematic diagram of the structure of a cloud service system provided in an embodiment of the present application;

[0047] FIG2 is a flow chart of a vector processing method based on a cloud management platform according to an embodiment of the present application;

[0048] FIG3 is another schematic diagram of the structure of the cloud service system provided in an embodiment of the present application;

[0049] FIG4 is another schematic flow chart of a vector processing method based on a cloud management platform according to an embodiment of the present application;

[0050] FIG5 is another schematic diagram of the structure of the cloud service system provided in an embodiment of the present application;

[0051] FIG6 is another flow chart of a vector processing method based on a cloud management platform according to an embodiment of the present application;

[0052] FIG7 is another schematic diagram of the structure of the cloud service system provided in an embodiment of the present application;

[0053] FIG8 is another schematic flow chart of a vector processing method based on a cloud management platform according to an embodiment of the present application;

[0054] FIG9 is another schematic diagram of the structure of the cloud service system provided in an embodiment of the present application;

[0055] FIG10 is a schematic diagram of the structure of a cloud management platform provided in an embodiment of the present application;

[0056] FIG11 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0057] FIG12 is a schematic diagram of a structure of a computing device cluster provided in an embodiment of the present application;

[0058] FIG13 is a schematic diagram of computer devices in a computer cluster provided by an embodiment of the present application being connected via a network. DETAILED DESCRIPTION

[0059] The embodiments of the present application provide a vector processing method and a cloud management platform based on a cloud management platform, which can enable the resources of multiple computing nodes to be shared globally, more effectively utilize the resources of the computing nodes, and thus reduce costs to a certain extent.

[0060] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0061] With the rapid development of cloud technology, more and more tenants choose to store their unstructured data (such as images, text, and voice, etc.) in the form of vector representation in the cloud service systems provided by cloud vendors, so that tenants can access and query a certain vector at any time, thereby meeting the tenants' vector processing needs.

[0062] In related technologies, a cloud service system built with a storage-computing separation architecture may include multiple computing nodes and multiple storage nodes. These multiple computing nodes are divided into different types of computing nodes based on their functions. For example, a computing node with a vector generation function can convert a tenant's unstructured data into vectors and store them in a storage node. Another example is a computing node with an index generation function that can generate an index for a vector and store the index in a storage node. Another example is a computing node with a vector query function that can search for the corresponding vector for a tenant from a storage node based on a tenant's request and return it to the tenant, and so on.

[0063] In the aforementioned system, because multiple compute nodes are divided into different types, certain types of compute nodes can only complete certain types of vector processing tasks and are unable to complete other types. For example, a compute node with vector generation capabilities cannot complete vector query tasks. This prevents the global sharing of resources among multiple compute nodes. When a burst of vector processing tasks occurs in the system, some compute nodes will be idle, resulting in a certain amount of resource waste and excessively high overall system costs.

[0064] To address the above issues, the present invention provides a vector processing method based on a cloud management platform. This method can be implemented through a cloud service system. Figure 1 is a schematic diagram of the structure of the cloud service system provided by the present invention. As shown in Figure 1, the system includes an infrastructure that can provide cloud services and a cloud management platform that manages the infrastructure. The following describes the cloud management platform and infrastructure separately:

[0065] The cloud management platform can coordinate the management of the infrastructure in the entire cloud service system (for example, in the infrastructure, according to the instructions of the tenant, multiple computing nodes and multiple storage nodes are created for the tenant to serve the tenant. These computing nodes and storage nodes can provide vector generation services, vector query services, vector addition services and vector deletion services, etc. for the tenant). It can also be open to tenants outside the cloud service system and respond to their requests. For example, the cloud management platform can provide various interfaces such as login interfaces and processing interfaces for access by tenants' clients (for example, the terminal device used by the tenant or the browser on the terminal device, etc.). Among them, the cloud management platform can authenticate the tenant's client through the login interface, and after successful authentication, the tenant's client can be allowed to log in to the cloud management platform. For another example, the cloud management platform can also allow the tenant's client to send a vector generation request for the tenant's multiple (unstructured) data (these multiple data have been stored in advance in multiple storage nodes serving the tenant) to the cloud management platform through a processing interface. The cloud management platform can select a computing node from the multiple computing nodes serving the tenant so that the computing node processes the vector generation request, generates multiple vectors corresponding to the tenant's multiple data, and stores these multiple vectors in multiple storage nodes. For another example, the cloud management platform can also allow the tenant's client to send a vector query request for a target vector to the cloud management platform through a processing interface. The cloud management platform can select a computing node from the multiple computing nodes so that the computing node processes the vector query request, finds and returns another output vector that matches the target vector from several vectors stored in itself or from multiple vectors stored in multiple storage nodes to the tenant. For another example, the cloud management platform can also allow tenants to write new data to multiple storage nodes through their clients through a processing interface, so that multiple storage nodes send vector addition requests for the new data to the cloud management platform. The cloud management platform can select a computing node from the multiple computing nodes to have the computing node process the vector addition request, generate a new vector corresponding to the new data, and store the new vector in the multiple storage nodes. For another example, the cloud management platform can also allow tenants to delete certain data from multiple data stored by multiple storage nodes through their clients through a processing interface, so that multiple storage nodes send vector deletion requests for the deleted data to the cloud management platform. The cloud management platform can select a computing node from the multiple computing nodes to have the computing node process the vector deletion request and delete the vector to be deleted corresponding to the deleted data from the multiple vectors stored by the multiple storage nodes, etc.

[0066] The infrastructure includes multiple compute nodes and multiple storage nodes serving tenants. For any of the multiple compute nodes, the compute node may include a vector cache and a vector index cache. The vector cache is used to store several vectors from a plurality of vectors corresponding to the tenant's data. These vectors are typically stored in the form of several vector groups (also called vector segments), each vector group containing at least one vector. The vector index cache is used to store indexes of these vector groups, allowing the compute node to access these vector groups. For the multiple storage nodes, the vectors corresponding to the tenant's data may be stored in the form of several vector groups, each vector group containing at least one vector from the plurality of vectors. Furthermore, the indexes of these vector groups are also stored in the multiple storage nodes. It should be noted that each storage node may store at least one vector group from the plurality of vector groups and its corresponding index (of course, each storage node may also store these vector groups and their indexes).

[0067] It's worth noting that the multiple storage nodes serving tenants can be divided into two parts: one part constitutes the memory pool, and the other part constitutes the object storage pool. The storage nodes in the memory pool can store a portion of a tenant's multiple vector groups and their corresponding indexes. It should be noted that the number of vector groups stored in the memory pool is typically greater than the number of vector groups stored in any one compute node. The storage nodes in the object storage pool can store not only multiple tenant data but also multiple vector groups corresponding to these data and their corresponding indexes. In other words, the object storage pool stores all of a tenant's vector groups, while the compute nodes and memory pool store a portion of the tenant's vector groups.

[0068] It is also worth noting that the cloud management platform has a task scheduling function. Based on this function, the cloud management platform can record the following data:

[0069] (1) Vector status table: The vector status table can record several vector groups stored by each computing node and multiple vector groups stored by multiple storage nodes. It can be understood that for any computing node, the vector status table records several vector groups stored by the computing node. Since the several vector groups stored by the computing node are part of the multiple vector groups stored by multiple storage nodes, the cloud management platform can perform a comprehensive analysis on the several vector groups stored by the computing node and the remaining vector groups stored by multiple storage nodes other than these several vector groups (for example, the computing node stores a certain part of the multiple vector groups but does not store another part of the multiple vector groups), thereby calculating the time required for the computing node to load (obtain) the several vector groups stored by itself (i.e., the time required for the computing node to load the vectors stored by itself), and the time required for the computing node to load the remaining vector groups other than these several vector groups from multiple storage nodes (i.e., the time required for the computing node to load the vectors stored by multiple storage nodes). This information can be used as the status information of the computing node.

[0070] For example, a vector status table is shown in Table 1. Assume that the compute nodes serving the tenant include compute node 1, compute node 2, and so on, and that there are multiple storage nodes serving the tenant. Assume that, according to the tenant's requirements, vector groups 1, 2, 3, 4, 5, and 6 are generated for the tenant and stored in multiple remote storage nodes. Furthermore, in the subsequent process, compute node 1 caches vector groups 1, 2, and 5, and compute node 2 caches vector groups 2, 3, and 5. Therefore, the cloud management platform can generate the vector group status table shown in Table 1 to record the current vector storage status of each compute node and storage node. (Vector storage status is typically reported to the cloud management platform by each compute node and storage node. It should be noted that in this example, the multiple storage nodes are considered as a whole; vector groups 1 through 6 can be deployed on the same or different storage nodes.)

[0071] Table 1

[0072] Based on the vector status table, the cloud management platform can determine the time required for compute node 1 to load vector groups 1, 2, and 5 stored on itself, as well as the time required for compute node 1 to load vector groups 3, 4, and 6 stored on multiple storage nodes. Similarly, the cloud management platform also determines these times for compute node 2 and other compute nodes, which will not be further detailed here.

[0073] (2) Vector index status table: The vector index status table can record the indexes of several vector groups stored by each computing node, as well as the indexes of multiple vector groups stored by multiple storage nodes. It is understandable that, for any computing node, the vector index status table records the indexes of several vector groups stored by the computing node. Since the indexes of the several vector groups stored by the computing node are part of the indexes of the multiple vector groups stored by multiple storage nodes, the cloud management platform can perform a comprehensive analysis based on the indexes of the several vector groups stored by the computing node and the indexes of the remaining vector groups stored by the multiple storage nodes other than the indexes of the several vector groups (for example, the computing node stores the indexes of some vector groups among the multiple vector groups but does not store the indexes of another part of the multiple vector groups). The platform calculates the time required for the computing node to load the indexes of the several vector groups stored by itself (i.e., the time required for the computing node to load the indexes stored by itself), and the time required for the computing node to load the indexes of the remaining vector groups other than the indexes of the several vector groups from the multiple storage nodes (i.e., the time required for the computing node to load the indexes stored by the multiple storage nodes). This information can also serve as the status information of the computing node.

[0074] Still using the above example, the vector index status table is shown in Table 2. The cloud management platform can also generate a vector index status table as shown in Table 2 to record the current vector index storage status of each computing node and storage node (the vector index storage status is usually reported to the cloud management platform by each computing node and storage node).

[0075] Table 2

[0076] Based on the vector index status table, the cloud management platform can determine the time required for compute node 1 to load the indexes of vector group 1, vector group 2, and vector group 5 stored on the compute node. It can also determine the time required for compute node 1 to load the indexes of vector group 3, vector group 4, and vector group 6 stored on multiple storage nodes. Similarly, the cloud management platform also determines these times for compute node 2 and other compute nodes, which will not be further described here.

[0077] (3) Computing node status table: The computing node status table can record the number of pending tasks (pending vector processing requests) of each computing node, the used amount of the central processing unit (CPU), the used amount of memory, etc. It can be understood that for any computing node, the computing node status table records the number of pending tasks of the computing node, the used amount of the CPU of the computing node, the used amount of memory, and other data (used to indicate the resource usage of the computing node). Therefore, the cloud management platform can analyze this data and calculate the time required for the computing node to execute (process) a certain task (a certain vector processing request). This information can also be used as the status information of the computing node.

[0078] Still as in the above example, the computing node status table is shown in Table 3. The cloud management platform can also generate a computing node status table as shown in Table 3 to record the current resource usage of each computing node (resource usage is usually reported to the cloud management platform by each computing node).

[0079] Table 3

[0080] Then, for a pending task, the cloud management platform can determine the time required for compute node 1 to execute the pending task based on data such as the number of pending tasks, CPU usage, and memory usage for compute node 1 in the vector state table. Similarly, the cloud management platform also determines the time required for compute node 2 and other compute nodes, which will not be further described here.

[0081] Based on this, for a vector processing request triggered by a tenant, the cloud management platform can determine the status information of each computing node serving the tenant based on the vector status table, the vector index status table, and the computing node status table, so as to select the computing node with the best status (i.e., the lowest load) from the multiple computing nodes serving the tenant to execute the vector processing request. In other words, regardless of whether the vector processing request is a vector generation request, a vector query request, a vector addition request, or a vector deletion request, from the perspective of the cloud management platform, the functions of the multiple computing nodes are the same, and each computing node has the ability to execute the vector processing request. Therefore, the cloud management platform can select a computing node with the best status from the multiple computing nodes to call the computing node to execute the vector processing request. It can be seen from this that because the cloud management platform has a task scheduling function for multiple computing nodes, each computing node can execute various types of vector processing requests, thereby enabling the resources of multiple computing nodes to be shared globally, more effectively utilizing the resources of the computing nodes, and thus reducing the cost of the cloud service system to a certain extent. To further understand the workflow of the cloud service system, the following describes the workflow in conjunction with multiple embodiments. First, the first embodiment is described. The first embodiment is shown in FIG2 (FIG. 2 is a flow chart of a vector processing method based on a cloud management platform provided in an embodiment of the present application). The first embodiment can be implemented using the cloud service system shown in FIG1 . The cloud service system includes an infrastructure that provides cloud services to tenants and a cloud management platform that manages the infrastructure. The infrastructure includes multiple computing nodes and multiple storage nodes that serve tenants. The multiple storage nodes are used to store multiple data of the tenants. The first embodiment includes:

[0082] 201. The cloud management platform receives a vector generation request for multiple data from a tenant.

[0083] In this embodiment, when a tenant needs to convert multiple data stored in multiple storage nodes into multiple vectors, the tenant can send a vector generation request for the multiple data to the cloud management platform through its client. After receiving the vector generation request for the multiple data, the cloud management platform can determine that the tenant's multiple data need to be converted into multiple vectors.

[0084] For example, as shown in Figure 3 (Figure 3 is another schematic diagram of the structure of the cloud service system provided by an embodiment of the present application), assume that a tenant has previously stored multiple (unstructured) data in an object storage pool through the cloud management platform. When the tenant needs to store these multiple data in the form of vectors, the tenant can send a vector generation request for these multiple data to the cloud management platform.

[0085] 202. The cloud management platform determines a first computing node with the lowest load from the multiple computing nodes based on the vector generation request and status information of the multiple computing nodes according to a load balancing algorithm, and sends the vector generation request to the first computing node. The multiple computing nodes are all configured with the ability to process the vector generation request.

[0086] After receiving a vector generation request for multiple data points, the cloud management platform can, based on the vector generation request, obtain the status information of multiple computing nodes serving the tenant using the vector status table, vector index status table, and computing node status table. Based on this status information of multiple computing nodes, the cloud management platform selects the computing node with the best status (i.e., the lowest load) from these multiple computing nodes as the first computing node according to the load balancing algorithm, and sends the vector generation request to the first computing node. It should be noted that for any of the multiple computing nodes, the cloud management platform has already configured the computing node with the ability to process vector generation requests when it was created by the cloud management platform.

[0087] Specifically, the cloud management platform may determine the first computing node from the multiple computing nodes in the following manner:

[0088] After receiving the vector generation request, for any one of the multiple computing nodes serving the tenant, the cloud management platform can obtain the time required for the computing node to load the vector stored by itself (since the computing node does not store any vector group at this time, the time is zero), the time required for the computing node to load the vectors stored by multiple storage nodes (since the multiple storage nodes do not store any vector group at this time, the time is zero), the time required for the computing node to load the index stored by itself (since the computing node does not store any vector group at this time, the time is zero), and the time required for the computing node to load the index stored by itself (since the computing node does not store any vector group at this time). The cloud management platform calculates the status information of the computing node, including the time required for the computing node to load the indexes stored by multiple storage nodes (since the multiple storage nodes do not store any indexes of the vector groups at this time, this time is zero), and the time required for the computing node to execute the vector generation request. Then, the cloud management platform can calculate the status information of the computing node through a load balancing algorithm (there can be multiple load balancing algorithms, which are not limited here), thereby obtaining an evaluation value for the computing node. The evaluation value of the computing node is used to indicate the quality of the computing node's status, that is, the load of the computing node. The larger the evaluation value, the worse the computing node's status, that is, the higher the load of the computing node. The smaller the evaluation value, the better the computing node's status, that is, the lower the load of the computing node.

[0089] Similarly, the cloud management platform can perform the same operations on the remaining compute nodes other than the specified compute node as it did on the specified compute node. Ultimately, the cloud management platform can obtain evaluation values ​​for the multiple compute nodes. Then, among the multiple compute nodes, the cloud management platform can determine the compute node with the lowest evaluation value as the first compute node to execute the vector generation request.

[0090] Continuing with the above example, after receiving a vector generation request, the cloud management platform can treat the vector generation request as a vector generation task. Using the data recorded in the vector status table, vector index status table, and compute node status table, it can determine the time required for compute node 1 to load its own stored vector group (which is zero), the time required for compute node 1 to load vector groups stored by multiple storage nodes (which is zero), the time required for compute node 1 to load the index of its own stored vector group (which is zero), the time required for compute node 1 to load the index of vector groups stored by multiple storage nodes (which is zero), and the time required for compute node 1 to execute the vector generation task. The cloud management platform can then sum these five times (i.e., the aforementioned load balancing algorithm) to obtain the evaluation value of compute node 1. Similarly, the cloud management platform can perform similar operations on the remaining compute nodes, including compute node 2, ultimately obtaining the evaluation values ​​of all compute nodes, including compute node 1 and compute node 2.

[0091] Among these computing nodes, since computing node 1 has the smallest evaluation value (indicating that computing node 1 has the lowest load and is in the best state), the cloud management platform can determine computing node 1 as the computing node to execute the vector generation task and schedule the vector generation task to be executed by computing node 1.

[0092] 203. The first computing node generates multiple vectors corresponding to the multiple data based on the vector generation request, and stores the multiple vectors in multiple storage nodes.

[0093] Upon receiving a vector generation request for multiple data items, the first computing node can retrieve the multiple data items from multiple storage nodes based on the vector generation request and embed the multiple data items to obtain multiple vectors corresponding to the multiple data items, where each data item corresponds to at least one vector. The first computing node can then write the multiple vectors to the multiple storage nodes (e.g., writing to an object storage pool via a memory pool or directly writing to the object storage pool), so that the multiple storage nodes store the multiple vectors, thereby meeting the tenant's vector generation requirements.

[0094] Specifically, the first computing node may store the multiple vectors in the multiple storage nodes in the following manner:

[0095] After receiving a vector generation request for multiple data items, the first computing node can generate a vector corresponding to the first data item and store it. It can then generate a vector corresponding to the second data item and store it. This process continues. When the number of vectors stored by the first computing node reaches a preset value (this value can be set based on actual needs and is not limited here), the first computing node can divide the stored vectors into a vector group and store the vector group on a storage node in the object storage pool. By repeating the above operations, the first computing node can convert multiple data items into corresponding vectors and divide the multiple vectors into multiple vector groups for storage in the object storage pool.

[0096] Furthermore, after the first computing node writes a vector group to a storage node in the object storage pool, it can also generate an index of the vector group and write the index of the vector group to the storage node in the object storage pool. In this way, the first computing node can write the indexes of multiple vector groups to the object storage pool.

[0097] It is worth noting that each time the first computing node obtains a vector group, it can choose whether to store the vector group in itself according to a preset strategy. Regardless of whether the first computing node chooses to store the vector group or not, after the first computing node writes the vector group to the object storage pool, it can report the several vector groups it has stored to the cloud management platform so that the cloud management platform can update the vector status table. It should be noted that the first computing node does not store the multiple vector groups it generates, but only selectively stores some of the multiple vector groups (i.e., the first vector group in the subsequent embodiments, and the vectors contained in this part of the vector group are also the second vectors in the subsequent embodiments). In addition, since the process of writing the vector group to the object storage pool will pass through the memory pool, the storage nodes in the memory pool can also store some of the multiple vector groups.

[0098] It is also worth noting that each time the first computing node obtains the index of a vector group, it can choose whether to store the index of the vector group in its own memory based on whether the vector group is already stored. Regardless of whether the first computing node chooses to store the index of the vector group or not, after writing the index of the vector group to the object storage pool, the first computing node can report the indexes of several vector groups it has stored to the cloud management platform, so that the cloud management platform can update the vector index status table. It should be noted that the first computing node only selectively stores the indexes of some vector groups among the multiple vector groups. In addition, because the process of writing the vector group index to the object storage pool passes through the memory pool, the storage nodes in the memory pool may also store the indexes of some vector groups among the multiple vector groups.

[0099] Continuing with the above example, after receiving a vector generation task, compute node 1 can, based on the vector generation task, read multiple tenant data items from the object storage pool, including data 1 through data 36. Compute node 1 then uses the mapping model to process data 1, obtaining vector 1 corresponding to data 1. It then processes data 2, obtaining vector 2 corresponding to data 2, and so on. Since six vectors form a vector group, after obtaining vector 6, compute node 1 can treat vectors 1 through 6 as vector group 1 and write vector group 1 to the object storage pool via the memory pool. Compute node 1 can then choose to cache vector group 1 and report the currently stored vector group, i.e., vector group 1, to the cloud management platform. Compute node 1 can then generate an index for vector group 1 and write it to the object storage pool via the memory pool. Compute node 1 can then choose to cache the index for vector group 1 and report the index for the currently stored vector group, i.e., vector group 1, to the cloud management platform.

[0100] By repeating the above operations, compute node 1 can continue to generate vectors 7 to 36, that is, vector group 2 to vector group 5, and the index of vector group 2 to the index of vector group 5, and write vector group 2 to vector group 5 and the index of vector group 2 to vector group 5 to the object storage pool. This will not be repeated here. It is worth noting that compute node 1 is designed to cache only three vector groups. Based on the preset policy, compute node 1 selects to cache vector group 1, vector group 2, vector group 5, and the indexes of these three vector groups, and reports them to the cloud management platform in real time. Therefore, the cloud management platform can generate the vector storage status of compute node 1 in Table 1 and the vector index storage status of compute node 2 in Table 2.

[0101] The above is an introduction to the first embodiment. The second embodiment will be introduced below. Figure 4 is another flow chart of the vector processing method based on the cloud management platform provided in the embodiment of the present application. As shown in Figure 4, the second embodiment can be implemented through the cloud service system shown in Figure 1. The cloud service system includes an infrastructure that provides cloud services to tenants and a cloud management platform that manages this infrastructure. This infrastructure includes multiple computing nodes and multiple storage nodes that serve tenants. The multiple storage nodes are used to store multiple data of tenants. The second embodiment includes:

[0102] 401. The cloud management platform receives a vector query request for a first vector from a tenant.

[0103] In this embodiment, when a tenant needs to query the first vector (i.e., the target vector), the tenant can send a vector query request for the first vector to the cloud management platform through its client. After receiving the vector generation request for the first vector, the cloud management platform can determine that the tenant needs to be queried and return the first vector.

[0104] For example, as shown in FIG5 ( FIG5 is another structural diagram of the cloud service system provided in an embodiment of the present application), assuming that a tenant needs to query a target vector, the tenant can send a vector query request for the target vector to the cloud management platform.

[0105] 402. The cloud management platform determines a first computing node with the lowest load from among the multiple computing nodes based on the vector query request and status information of the multiple computing nodes according to a load balancing algorithm, and sends the vector query request to the first computing node. The multiple computing nodes are all configured to process the vector query request.

[0106] After receiving a vector query request for the first vector, the cloud management platform can, based on the vector query request, obtain the status information of multiple computing nodes serving the tenant through the vector status table, the vector index status table, and the computing node status table. Based on the status information of these multiple computing nodes, the cloud management platform selects the computing node with the best status (i.e., the lowest load) from these multiple computing nodes according to the load balancing algorithm. For example, if this computing node is still the first computing node, the cloud management platform can send the vector query request to the first computing node. It should be noted that for any of the multiple computing nodes, the cloud management platform has already configured the computing node with the ability to process vector query requests when it was created by the cloud management platform.

[0107] Specifically, the cloud management platform may determine the first computing node from the multiple computing nodes in the following manner:

[0108] After receiving a vector generation request, the cloud management platform can obtain, for any one of the multiple compute nodes serving a tenant, information about the compute node, including the time required for the compute node to load its own stored vectors (since the compute node typically already stores some vector groups from multiple vector groups at this time, this time is non-zero), the time required for the compute node to load vectors stored by multiple storage nodes (since multiple storage nodes already store multiple vector groups at this time, this time is non-zero), the time required for the compute node to load its own stored indexes (since the compute node already stores indexes of some vector groups from multiple vector groups at this time, this time is non-zero), the time required for the compute node to load indexes stored by multiple storage nodes (since multiple storage nodes already store indexes of multiple vector groups at this time, this time is non-zero), and the time required for the compute node to execute a vector query request. The cloud management platform can then calculate the compute node's status information using a load balancing algorithm to obtain an evaluation value for the compute node.

[0109] Similarly, the cloud management platform can perform the same operations on the remaining compute nodes as it did on the first compute node. Ultimately, the cloud management platform can obtain evaluation values ​​for the multiple compute nodes. Then, the cloud management platform can select the compute node with the lowest evaluation value among the multiple compute nodes. For example, this compute node remains the first compute node and can execute the vector query request.

[0110] Continuing with the above example, after receiving a vector query request, the cloud management platform can treat the vector query request as a vector query task. Using the data recorded in the vector status table, vector index status table, and compute node status table, it can determine the time required for compute node 1 to load its own stored vector group (non-zero), the time required for compute node 1 to load vector groups stored by multiple storage nodes (non-zero), the time required for compute node 1 to load the index of its own stored vector group (non-zero), the time required for compute node 1 to load the index of vector groups stored by multiple storage nodes (non-zero), and the time required for compute node 1 to execute the vector query task. The cloud management platform can then add these five times to obtain the evaluation value of compute node 1. Similarly, the cloud management platform can perform similar operations on the remaining compute nodes, including compute node 2, ultimately obtaining the evaluation values ​​of all compute nodes, including compute node 1 and compute node 2.

[0111] Among these computing nodes, since computing node 1 has the smallest evaluation value, the cloud management platform can determine computing node 1 as the computing node that executes the vector query task and schedule the vector query task to be executed by computing node 1.

[0112] 403. Based on the vector query request, the first computing node obtains a third vector matching the first vector from the second vector stored in the first computing node or multiple vectors stored in multiple storage nodes, where the multiple vectors include the second vector.

[0113] After receiving a vector query request for a first vector, the first computing node can, based on the vector query request, obtain a third vector (i.e., an output vector, which can be identical or similar to the target vector) that matches the first vector from its own stored second vectors (i.e., several vectors stored by the first computing node itself, which are some of the multiple vectors stored by multiple storage nodes) or multiple vectors stored by multiple storage nodes, and return it to the tenant's client for the tenant's use, thereby satisfying the tenant's vector query requirement.

[0114] Specifically, the first computing node may obtain the third vector in the following manner:

[0115] Based on the aforementioned embodiment, it can be seen that the multiple vectors stored by the first computing node can be stored in the form of multiple vector groups. That is, the second vector stored by the first computing node can be stored in the form of a first vector group, and the first vector group is also a portion of the multiple vector groups stored in the object storage pool. Therefore, after receiving a vector query request, the first computing node can retrieve the first vector group it has stored based on the index of the first vector group it has stored, and then detect whether there is a third vector matching the first vector from the first vector group. If so, it indicates that the first computing node has stored the third vector, so the first computing node can retrieve the third vector from the first vector group and return it to the tenant. If not, it indicates that the first computing node does not store the third vector, so the first computing node can retrieve the second vector group (i.e., the remaining vector groups in the multiple vector groups excluding the several vector groups stored by the first computing node) and the index of the second vector group from the memory pool or object storage pool, retrieve the second vector group based on the index of the second vector group, then retrieve the third vector from the second vector group, and then return the third vector to the tenant.

[0116] It is worth noting that when the first computing node itself does not store the third vector, the first computing node can determine the third vector group in which the third vector is located from the second vector group according to a preset strategy, and further determine whether to store the third vector group. If the first computing node chooses to store the third vector group, the first computing node can, according to the preset strategy, make the third vector group replace a certain vector group among the several vector groups (i.e., the first vector group) that it has stored (of course, it is also possible not to replace it, but to make the third vector group directly added to the first vector group), and report the several latest vector groups stored by itself (i.e., the updated first vector group) to the cloud management platform in real time, so that the cloud management platform updates the vector status table.

[0117] Correspondingly, if the first computing node chooses to store the vector group where the third vector is located, the first computing node can also replace the index of a vector group among the several vector groups it has stored with the index of the third vector group where the third vector is located (of course, it can also not be replaced), and report the indexes of the several vector groups most recently stored by itself (that is, the index of the updated first vector group) to the cloud management platform in real time, so that the cloud management platform updates the vector index status table.

[0118] Based on this, for the remaining computing nodes among the multiple computing nodes except the first computing node, once the remaining computing nodes also execute other vector query requests, the remaining computing nodes will also store several vector groups and corresponding indexes, which will not be repeated here.

[0119] As in the previous example, after receiving a vector query task, compute node 1 caches vector groups 1, 2, and 5. If vector 8 matches the tenant's target vector, which is in vector group 2, compute node 1 can read vector 8 from its cached vector group 2 and return it to the tenant. If vector 12 matches the tenant's target vector, which is in vector group 3, compute node 1 can read vector groups 3, 4, and 6, along with their corresponding indexes, from the memory pool or object storage pool to obtain vector 12 from vector group 3 and return it to the tenant.

[0120] Furthermore, when compute node 1 does not cache vector group 3, where vector 12 resides, compute node 1 can choose whether to cache vector group 3 according to a preset policy. If caching is chosen, compute node 1 can select vector group 3 to replace its already stored vector group 5 according to the preset policy, and replace the index of vector group 3 with the index of its already stored vector group 5. At this point, compute node 1 has cached vector group 1, vector group 2, vector group 3, and their corresponding indexes. Therefore, compute node 1 can report the latest vector storage status and index storage status to the cloud management platform in real time, allowing the cloud management platform to update the vector status table and vector index status table.

[0121] The above is an introduction to the second embodiment. The third embodiment will be introduced below. Figure 6 is another flow chart of the vector processing method based on the cloud management platform provided in the embodiment of the present application. As shown in Figure 6, the third embodiment can be implemented through the cloud service system shown in Figure 1. The cloud service system includes an infrastructure that provides cloud services to tenants and a cloud management platform that manages this infrastructure. This infrastructure includes multiple computing nodes and multiple storage nodes that serve tenants. The multiple storage nodes are used to store multiple data of tenants. The third embodiment includes:

[0122] 601. A cloud management platform receives a vector addition request for first data, where the first data is new data written by a tenant into multiple storage nodes.

[0123] In this embodiment, when a tenant writes new data, namely, first data, to multiple storage nodes, this may trigger the multiple storage nodes to send a vector increment request for the first data to the cloud management platform. Upon receiving the vector increment request for the first data, the cloud management platform may determine that the first data needs to be converted into a fourth vector (i.e., a new vector).

[0124] For example, as shown in Figure 7 (Figure 7 is another structural diagram of the cloud service system provided by an embodiment of the present application), assume that the tenant writes new data to the object storage pool, triggering the object storage pool to send a vector increase request for the new data to the cloud management platform.

[0125] 602. The cloud management platform determines a first computing node with the lowest load from among the multiple computing nodes based on the vector increase request and status information of the multiple computing nodes according to a load balancing algorithm, and sends a vector increase request to the first computing node. The multiple computing nodes are all configured with the ability to process the vector increase request.

[0126] After receiving the vector increase request for the first data, the cloud management platform can obtain the status information of multiple computing nodes serving the tenant through the vector status table, the vector index status table, and the computing node status table based on the vector increase request. Based on the status information of these multiple computing nodes, the cloud management platform selects the computing node with the best status (i.e., the lowest load) from these multiple computing nodes according to the load balancing algorithm. For example, if this computing node is still the first computing node, the cloud management platform can send the vector increase request to the first computing node. It should be noted that for any one of the multiple computing nodes, when the computing node was created by the cloud management platform, the cloud management platform has been configured with the ability to process vector increase requests.

[0127] For an introduction to step 602 , please refer to the relevant description of step 402 in the embodiment shown in FIG. 4 , which will not be repeated here.

[0128] 603. The first computing node generates a fourth vector corresponding to the first data based on the vector addition request, and stores the fourth vector in multiple storage nodes.

[0129] After receiving a vector increase request for the first data, the first computing node can retrieve the first data from multiple storage nodes based on the vector increase request and map the first data to obtain a fourth vector corresponding to the first data. The first computing node can then store the fourth vector in the multiple storage nodes, automatically satisfying the tenant's vector increase request.

[0130] Specifically, the first computing node may write the fourth vector into the multiple storage nodes in the following manner:

[0131] After receiving the vector addition request, the first computing node may read the first data from the object storage pool, map the first data to obtain a fourth vector corresponding to the first data, and store the fourth vector. At this point, the first computing node determines whether the number of vectors (including the fourth vector) stored in the node that have not been divided into vector groups equals the preset number. If so, the first computing node may divide the vectors into the fourth vector group (a new vector group) and write the fourth vector group to the object storage pool.

[0132] Furthermore, the first computing node may also generate an index of the fourth vector group, and write the index of the fourth vector group into the object storage pool.

[0133] It is worth noting that the first computing node can determine whether to store the fourth vector group based on a preset policy. If the first computing node chooses to store the fourth vector group, the first computing node can, according to the preset policy, cause the fourth vector group to replace a certain vector group among the several vector groups (i.e., the first vector group) stored by itself (of course, it is also possible not to replace it), and report the several newly stored vector groups (i.e., the updated first vector group) to the cloud management platform in real time, so that the cloud management platform updates the vector status table.

[0134] Correspondingly, the first computing node may also replace the index of a certain vector group among the several vector groups it has stored with the index of the fourth vector group (of course, it may not be replaced), and report the indexes of the several vector groups it has most recently stored (i.e., the updated index of the first vector group) to the cloud management platform in real time, so that the cloud management platform updates the vector index status table.

[0135] Continuing with the previous example, after receiving the vector addition task, Compute Node 1 reads the new data from the object storage pool, maps the new data, and obtains a new vector, which it then caches. Since Compute Node 1 already has some cached vectors (including the new vector) that haven't yet been grouped together, when the number of these vectors reaches a preset value, Compute Node 1 groups them together into a new vector group, Vector Group 7, and writes Vector Group 7 to the object storage pool. Compute Node 1 then generates an index for Vector Group 7 and writes it to the object storage pool.

[0136] Furthermore, compute node 1 can choose whether to cache vector group 7 according to a preset policy. If caching is selected, compute node 1 can select vector group 7 to replace its own stored vector group 5 according to the preset policy, and replace the index of vector group 7 with the index of vector group 5. At this point, compute node 1 has cached vector group 1, vector group 2, vector group 7, and their corresponding indexes. Therefore, compute node 1 can report this to the cloud management platform in real time, allowing the cloud management platform to update the vector status table and the vector index status table.

[0137] The above is an introduction to the third embodiment. The fourth embodiment will be introduced below. Figure 8 is another flow chart of the vector processing method based on the cloud management platform provided in the embodiment of the present application. As shown in Figure 8, the fourth embodiment can be implemented by the cloud service system shown in Figure 1. The cloud service system includes an infrastructure that provides cloud services to tenants and a cloud management platform that manages this infrastructure. This infrastructure includes multiple computing nodes and multiple storage nodes that serve tenants. The multiple storage nodes are used to store multiple data of tenants. The fourth embodiment includes:

[0138] 801. A cloud management platform receives a vector deletion request for a second vector, where the second vector is data deleted by a tenant from multiple data stored in multiple storage nodes.

[0139] In this embodiment, when a tenant deletes a piece of data, namely, the second data, from among the multiple data stored by multiple storage nodes, this may trigger the multiple storage nodes to send a vector deletion request for the second data to the cloud management platform. Upon receiving the vector deletion request for the second data, the cloud management platform may determine that the fifth vector corresponding to the second data (i.e., the vector to be deleted) needs to be deleted.

[0140] For example, as shown in Figure 9 (Figure 9 is another structural diagram of the cloud service system provided by an embodiment of the present application), assume that a tenant deletes certain data from the object storage pool, triggering the object storage pool to send a vector deletion request for the deleted data to the cloud management platform.

[0141] 802. The cloud management platform determines a first computing node with the lowest load from among the multiple computing nodes based on the vector deletion request and status information of the multiple computing nodes according to a load balancing algorithm, and sends the vector deletion request to the first computing node. The multiple computing nodes are all configured with the ability to process the vector deletion request.

[0142] After receiving the vector delete request for the second data, the cloud management platform can obtain the status information of multiple computing nodes serving the tenant based on the vector delete request through the vector status table, vector index status table, and computing node status table. Based on the status information of these multiple computing nodes, the cloud management platform selects the computing node with the best status (i.e., the lowest load) from these multiple computing nodes according to the load balancing algorithm. For example, if this computing node is still the first computing node, the cloud management platform can send the vector delete request to the first computing node. It should be noted that for any of the multiple computing nodes, the cloud management platform has already configured the computing node with the ability to process vector delete requests when it was created by the cloud management platform.

[0143] For an introduction to step 802 , please refer to the relevant description of step 402 in the embodiment shown in FIG4 , which will not be repeated here.

[0144] 803. The first computing node deletes a fifth vector corresponding to the second data from the multiple vectors stored in the multiple storage nodes based on the vector deletion request.

[0145] After receiving the vector deletion request for the second data, the first computing node may determine, based on the vector deletion request, a fifth vector corresponding to the second data from the multiple vectors stored in the multiple storage nodes, and delete the fifth vector.

[0146] Specifically, the first computing node may delete the fifth vector in the following manner:

[0147] After receiving the vector add request, the first computing node may determine, from among the multiple vector groups stored in the object storage pool, the fifth vector group containing the fifth vector corresponding to the second data (i.e., the vector group to be deleted), and delete the fifth vector group. Before deleting the fifth vector group, the first computing node may read the sixth vector in the fifth vector group (i.e., the remaining vectors in the vector group to be deleted, excluding the vector to be deleted). Thereafter, the first computing node may generate a sixth vector group (a new vector group) based on the sixth vector. For example, the first computing node may merge the sixth vector with the remaining vectors (these vectors may be vectors stored by the first computing node that are not divided into vector groups) to obtain the sixth vector group. The first computing node may then write the sixth vector group to the object storage pool.

[0148] Furthermore, the first computing node may also generate an index of the sixth vector group, and write the index of the sixth vector group into the object storage pool.

[0149] It is worth noting that the first computing node can determine whether to store the sixth vector group based on a preset strategy. If the first computing node chooses to store the sixth vector group, the first computing node can, according to the preset strategy, make the sixth vector group replace a certain vector group among the several vector groups (i.e., the first vector group) stored by itself (of course, it can also not replace it), and report the several newly stored vector groups (i.e., the updated first vector group) to the cloud management platform in real time, so that the cloud management platform updates the vector status table.

[0150] Correspondingly, the first computing node may also replace the index of a certain vector group among the several vector groups it has stored with the index of the sixth vector group (of course, it may not be replaced), and report the indexes of the several vector groups it has most recently stored (i.e., the updated index of the first vector group) to the cloud management platform in real time, so that the cloud management platform updates the vector index status table.

[0151] Continuing with the above example, after receiving a vector deletion task, compute node 1 can determine from vector groups 1 to 6 in the object storage pool that vector 31 corresponding to the deleted data is in vector group 6 (the vector group to be deleted). Compute node 1 can then delete vector group 6 from the object storage pool and cache vectors 32 to 36. Since compute node 1 also has other cached vectors that are not grouped, it can select one of these vectors and merge it with vectors 32 to 36 to form a new vector group, namely, vector group 8. Compute node 1 can then generate an index for vector group 8. Compute node 1 can then store vector group 8 and its index in the object storage pool.

[0152] Furthermore, compute node 1 can choose whether to cache vector group 8 according to a preset policy. If caching is selected, compute node 1 can select vector group 8 to replace its own stored vector group 5, and replace the index of vector group 8 with the index of vector group 5. At this point, compute node 1 has cached vector group 1, vector group 2, vector group 8, and their corresponding indexes. Therefore, compute node 1 can report this to the cloud management platform in real time, allowing the cloud management platform to update the vector status table and the vector index status table.

[0153] In an embodiment of the present application, for a vector processing request triggered by a tenant, the cloud management platform can obtain the status information of each computing node serving the tenant, and based on this, select the computing node with the best status from these multiple computing nodes to execute the vector processing request. Regardless of whether the vector processing request is a vector generation request, a vector query request, a vector addition request, or a vector deletion request, from the perspective of the cloud management platform, the functions of the multiple computing nodes are the same, and each computing node can execute the vector processing request. Therefore, the cloud management platform can select a computing node with the best status from multiple computing nodes to call the computing node to execute the vector processing request. It can be seen that since the cloud management platform has a task scheduling function for multiple computing nodes, each computing node can execute various types of vector processing requests, thereby enabling the resources of multiple computing nodes to be shared globally, more effectively utilizing the resources of the computing nodes, and thus reducing the cost of the cloud service system to a certain extent.

[0154] The above is a detailed description of the vector processing method based on the cloud management platform provided in the embodiment of the present application. The cloud management platform provided in the embodiment of the present application will be introduced below. Figure 10 is a structural schematic diagram of the cloud management platform provided in the embodiment of the present application. As shown in Figure 10, the cloud management platform is used to manage the infrastructure for providing cloud services. The infrastructure includes multiple computing nodes and multiple storage nodes. The multiple storage nodes are used to store multiple data of tenants. The cloud management platform includes:

[0155] Receiving module 1001, configured to receive a vector generation request for multiple data from a tenant;

[0156] The processing module 1002 is configured to determine a first computing node with the lowest load from among the multiple computing nodes according to a load balancing algorithm based on the vector generation request and status information of the multiple computing nodes, wherein the multiple computing nodes are all configured with the ability to process the vector generation request.

[0157] The processing module 1002 is further configured to send a vector generation request to the first computing node, where the vector generation request is configured to instruct the first computing node to generate multiple vectors corresponding to the multiple data and store the multiple vectors in multiple storage nodes.

[0158] In one possible implementation, receiving module 1001 is further configured to receive a vector query request for a first vector from a tenant; processing module 1002 is further configured to determine, based on the vector query request and status information of multiple computing nodes, a first computing node with the lowest load from among the multiple computing nodes according to a load balancing algorithm, where the multiple computing nodes are all capable of processing vector query requests; and processing module 1002 is further configured to send a vector query request to the first computing node, where the vector query request is configured to instruct the first computing node to obtain a third vector that matches the first vector from a second vector stored in the first computing node or from multiple vectors stored by multiple storage nodes, where the multiple vectors include the second vector.

[0159] In one possible implementation, the receiving module 1001 is further used to receive a vector increase request for first data, where the first data is new data written by a tenant to multiple storage nodes; the processing module 1002 is further used to determine, based on the vector increase request and status information of the multiple computing nodes, a first computing node with the lowest load from the multiple computing nodes according to a load balancing algorithm, where the multiple computing nodes are all capable of processing the vector increase request; the processing module 1002 is further used to send a vector increase request to the first computing node, where the vector increase request is used to instruct the first computing node to generate a fourth vector corresponding to the first data and store the fourth vector in the multiple storage nodes.

[0160] In one possible implementation, receiving module 1001 is further configured to receive a vector deletion request for second data, where the second data is data deleted by a tenant from multiple data stored by multiple storage nodes. Processing module 1002 is further configured to determine, based on the vector deletion request and status information of multiple computing nodes, a first computing node with the lowest load from the multiple computing nodes according to a load balancing algorithm, where all of the multiple computing nodes are capable of processing vector deletion requests. Processing module 1002 is further configured to send a vector deletion request to the first computing node, where the vector deletion request is used to instruct the first computing node to delete a fifth vector corresponding to the second data from the multiple vectors stored by the multiple storage nodes.

[0161] In a possible implementation, the vector addition request or the vector deletion request comes from multiple storage nodes.

[0162] In one possible implementation, the vector generation request is used to instruct: a first computing node to divide a plurality of vectors into a plurality of vector groups, where the number of vectors included in each vector group is equal to a preset value; and the first computing node to write the plurality of vector groups into the plurality of storage nodes respectively in units of vector groups.

[0163] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group. The vector generation request is further used to indicate that: after obtaining multiple vector groups, the first computing node may select a first vector group from the multiple vector groups and store the first vector group; and the first computing node notifies the cloud management platform that the first computing node has stored the first vector group.

[0164] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group, and the vector query request is used to instruct: the first computing node to detect whether there is a third vector matching the first vector in the first vector group; if the third vector exists in the first vector group, the first computing node to obtain the third vector from the first vector group; or, if the third vector does not exist in the first vector group, the first computing node to obtain a second vector group other than the first vector group from multiple vector groups from multiple storage nodes, and obtain the third vector from the second vector group, wherein the multiple vector groups are obtained based on the division of the multiple vectors.

[0165] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group. The vector query request is further used to instruct: the first computing node to determine, from the second vector group, a third vector group in which the third vector is located, and to update the first vector group based on the third vector group to obtain an updated first vector group, where the updated first vector group includes the third vector group; and the first computing node to notify the cloud management platform that the updated first vector group is stored in the first computing node.

[0166] In one possible implementation, the vector add request is used to instruct: a first computing node to generate a fourth vector corresponding to the first data and store the fourth vector; if the number of vectors not divided into vector groups stored by the first computing node is equal to a preset value, the first computing node to divide the vectors not divided into vector groups into a fourth vector group, and write the fourth vector group to multiple storage nodes, where the vectors not divided into vector groups include the fourth vector.

[0167] In one possible implementation, the second vector is stored in the first computing node in the form of a first vector group, and the vector increase request is further used to indicate: the first computing node updates the first vector group based on the fourth vector group to obtain an updated first vector group, where the updated first vector group includes the fourth vector group; and the first computing node notifies the cloud management platform that the updated first vector group is stored in the first computing node.

[0168] In one possible implementation, multiple vectors are stored in the form of multiple vector groups in multiple storage nodes. The vector deletion request is used to instruct: a first computing node to obtain a fifth vector group containing a fifth vector corresponding to the second data from the multiple vector groups stored in the multiple storage nodes, and delete the fifth vector group from the multiple vector groups stored in the multiple storage nodes; and the first computing node to generate a sixth vector group based on a sixth vector excluding the fifth vector in the fifth vector group, and store the sixth vector group in the multiple storage nodes.

[0169] In one possible implementation, the vector deletion request is further used to instruct: the first computing node to update the first vector group based on the sixth vector group to obtain an updated first vector group, where the updated first vector group includes the sixth vector group; and the first computing node to notify the cloud management platform that the first computing node stores the updated first vector group.

[0170] In one possible implementation, a processing module is configured to: generate a request based on a vector, calculate status information of multiple computing nodes according to a load balancing algorithm, and obtain evaluation values ​​of the multiple computing nodes, where the evaluation values ​​of the multiple computing nodes are used to indicate the loads of the multiple computing nodes; and determine, from the multiple computing nodes, a computing node with a smallest evaluation value as a first computing node; wherein, for any one of the multiple computing nodes, the status information of the computing node includes at least one of the following: the time required for the computing node to load a vector stored in itself, the time required for the computing node to load an index of a vector stored in itself, the time required for the computing node to load vectors stored in multiple storage nodes, the time required for the computing node to load indexes of vectors stored in multiple storage nodes, and the time required for the computing node to execute a vector processing request.

[0171] In one possible implementation, the cloud management platform further includes: an analysis module, the analysis module being configured to perform at least one of the following: analyzing vectors stored by a computing node and vectors stored by multiple storage nodes to obtain the time required for the computing node to load the vectors stored by itself and the time required for the computing node to load the vectors stored by multiple storage nodes; analyzing the indexes of the vectors stored by the computing node and the indexes of the vectors stored by multiple storage nodes to obtain the time required for the computing node to load the indexes stored by itself and the time required for the computing node to load the indexes stored by multiple storage nodes; and analyzing the resource usage of the computing node to obtain the time required for the computing node to execute a vector processing request.

[0172] It should be noted that the information interaction, implementation process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the embodiment of the present application, and no further details will be given here.

[0173] Please refer to Figure 11, which is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. As shown in Figure 11, the computing device 1100 (which can be used to present the aforementioned cloud management platform) includes: a processor 1101, a memory 1102, a communication interface 1103, and a bus 1104. The processor 1101, the memory 1102, and the communication interface 1103 are coupled via a bus (not labeled in the figure). The memory 1102 stores instructions. When the execution instructions in the memory 1102 are executed, the computing device 1100 executes the method executed by the cloud management platform in the above method embodiment.

[0174] The computing device 1100 may be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASICs), one or more digital signal processors (DSPs), one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms. For example, when a unit in the device can be implemented in the form of a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call a program. For example, these units can be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0175] The processor 1101 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0176] Memory 1102 may be volatile memory or nonvolatile memory, or may include both volatile and nonvolatile memory. Nonvolatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0177] Memory 1102 stores executable program code, and processor 1101 executes the executable program code to implement the functions of the aforementioned receiving module, processing module, and other modules, thereby implementing the aforementioned cloud management platform-based vector processing method. In other words, memory 1102 stores instructions for executing the aforementioned cloud management platform-based vector processing method.

[0178] The communication interface 1103 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.

[0179] In addition to the data bus, bus 1104 may also include a power bus, a control bus, and a status signal bus. The bus may be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a unified bus (Ubus or UB), a Compute Express Link (CXL), or a Cache Coherent Interconnect for Accelerators (CCIX). Buses can be categorized as address buses, data buses, and control buses.

[0180] Please refer to Figure 12 , which is a schematic diagram of a computing device cluster provided in an embodiment of the present application. As shown in Figure 12 , the computing device cluster 1200 includes at least one computing device 1100 .

[0181] As shown in Figure 12, the computing device cluster 1200 includes at least one computing device 1100. The memory 1102 in one or more computing devices 1100 in the computing device cluster 1200 may store the same instructions for executing the above-mentioned vector processing method based on the cloud management platform.

[0182] In some possible implementations, the memory 1102 of one or more computing devices 1100 in the computing device cluster 1200 may also store partial instructions for executing the aforementioned cloud management platform-based vector processing method. In other words, the combination of one or more computing devices 1100 can jointly execute the aforementioned cloud management platform-based vector processing method.

[0183] It should be noted that the memory 1102 in different computing devices 1100 in the computing device cluster 1200 may store different instructions, each for executing a portion of the functions of the aforementioned cloud management platform. In other words, the instructions stored in the memory 1102 in different computing devices 1100 may implement the functions of one or more modules such as the receiving module and the processing module.

[0184] In some possible implementations, one or more computing devices 1100 in the computing device cluster 1200 may be connected via a network, which may be a wide area network or a local area network.

[0185] Please refer to Figure 13, which is a schematic diagram of computer devices in a computer cluster provided by an embodiment of the present application being connected via a network. As shown in Figure 13, two computing devices 1100A and 1100B are connected via a network. Specifically, each computing device is connected to the network via a communication interface.

[0186] In a possible implementation, the memory of the computing device 1100A stores instructions for executing functions of a receiving module and the like. Meanwhile, the memory of the computing device 1100B stores instructions for executing functions of a processing module and the like.

[0187] It should be understood that the functions of the computing device 1100A shown in Figure 13 may also be completed by multiple computing devices. Similarly, the functions of the computing device 1100B may also be completed by multiple computing devices.

[0188] An embodiment of the present application also relates to a computer storage medium, which stores a program for signal processing. When the computer storage medium runs on a computer, it enables the computer to execute the steps executed by the cloud management platform in the embodiment shown in Figures 2, 4, 6 or 8.

[0189] An embodiment of the present application also relates to a computer program product, which stores instructions that, when executed by a computer, enable the computer to execute the steps performed by the cloud management platform in the embodiment shown in Figures 2, 4, 6 or 8.

[0190] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0191] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0192] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of the solution of this embodiment according to actual needs.

[0193] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0194] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A vector processing method based on a cloud management platform, characterized in that, The cloud management platform is used to manage the infrastructure that provides cloud services. The infrastructure includes multiple computing nodes and multiple storage nodes. The multiple storage nodes are used to store multiple data of tenants. The method includes: The cloud management platform receives a vector generation request from the tenant for the multiple data; Based on the vector generation request and the status information of the multiple computing nodes, the cloud management platform determines a first computing node with the lowest load from the multiple computing nodes according to the load balancing algorithm. The multiple computing nodes are all configured with the ability to process the vector generation request; The cloud management platform sends the vector generation request to the first computing node. The vector generation request is used to instruct the first computing node to generate multiple vectors corresponding to the multiple data and store the multiple vectors in the multiple storage nodes.

2. The method according to claim 1, wherein The method further includes: The cloud management platform receives a vector query request from the tenant for a first vector; Based on the vector query request and the status information of the multiple computing nodes, the cloud management platform determines the first computing node with the lowest load from the multiple computing nodes according to the load balancing algorithm. The multiple computing nodes are all configured with the ability to process the vector query request; The cloud management platform sends the vector query request to the first computing node. The vector query request is used to instruct the first computing node to obtain a third vector that matches the first vector from the second vector stored in itself or the multiple vectors stored in the multiple storage nodes. The multiple vectors include the second vector.

3. The method according to claim 1 or 2, characterized in that, The method further includes: The cloud management platform receives a vector addition request for first data. The first data is new data written by the tenant into the multiple storage nodes; Based on the vector addition request and the status information of the multiple computing nodes, the cloud management platform determines the first computing node with the lowest load from the multiple computing nodes according to the load balancing algorithm. The multiple computing nodes are all configured with the ability to process the vector addition request; The cloud management platform sends the vector addition request to the first computing node. The vector addition request is used to instruct the first computing node to generate a fourth vector corresponding to the first data and store the fourth vector in the multiple storage nodes.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The cloud management platform receives a vector deletion request for second data. The second data is data deleted by the tenant from the multiple data stored in the multiple storage nodes; Based on the vector deletion request and the status information of the multiple computing nodes, the cloud management platform determines the first computing node with the lowest load from the multiple computing nodes according to the load balancing algorithm. The multiple computing nodes are all configured with the ability to process the vector deletion request; The cloud management platform sends the vector deletion request to the first computing node, where the vector deletion request is used to instruct the first computing node to delete a fifth vector corresponding to the second data from the multiple vectors stored in the multiple storage nodes.

5. The method according to claim 3 or 4, characterized in that, The vector addition request or the vector deletion request comes from the multiple storage nodes.

6. The method according to any one of claims 1 to 5, characterized in that, The vector generation request is used to instruct: The first computing node divides the multiple vectors into multiple vector groups, and the number of vectors included in each vector group is equal to a preset value; The first computing node writes the multiple vector groups into the multiple storage nodes respectively in units of vector groups.

7. The method according to claim 6, wherein The second vector is stored in the first computing node in the form of a first vector group, and the vector generation request is further used to instruct: After obtaining the multiple vector groups, the first computing node can select the first vector group from the multiple vector groups and store the first vector group; The first computing node notifies the cloud management platform that the first computing node stores the first vector group.

8. The method according to claim 2, wherein The second vector is stored in the first computing node in the form of a first vector group, and the vector query request is used to instruct: The first computing node detects whether there is a third vector matching the first vector in the first vector group; If the third vector exists in the first vector group, the first computing node obtains the third vector from the first vector group; or, If the third vector does not exist in the first vector group, the first computing node obtains a second vector group other than the first vector group from the multiple vector groups at the multiple storage nodes and obtains the third vector from the second vector group, where the multiple vector groups are obtained based on the division of the multiple vectors.

9. The method according to claim 8, wherein The second vector is stored in the first computing node in the form of a first vector group, and the vector query request is further used to instruct: The first computing node determines the third vector group where the third vector is located from the second vector group, and updates the first vector group based on the third vector group to obtain an updated first vector group, where the updated first vector group includes the third vector group; The first computing node notifies the cloud management platform that the first computing node stores the updated first vector group.

10. The method according to claim 3, characterized in that The vector addition request is used to instruct: The first computing node generates a fourth vector corresponding to the first data and stores the fourth vector; If the number of vectors not divided into vector groups stored in the first computing node is equal to the preset value, the first computing node divides the vectors not divided into vector groups into a fourth vector group and writes the fourth vector group into the multiple storage nodes, where the vectors not divided into vector groups include the fourth vector.

11. The method according to claim 10, wherein The second vector is stored in the first computing node in the form of a first vector group, and the vector addition request is further used to instruct: The first computing node updates the first vector group based on the fourth vector group to obtain an updated first vector group, and the updated first vector group includes the fourth vector group; The first computing node notifies the cloud management platform that the first computing node stores the updated first vector group.

12. The method according to claim 4, characterized in that, The multiple vectors are stored in the multiple storage nodes in the form of multiple vector groups, and the vector deletion request is used to indicate: The first computing node obtains the fifth vector group where the fifth vector corresponding to the second data is located from the multiple vector groups stored in the multiple storage nodes, and deletes the fifth vector group from the multiple vector groups stored in the multiple storage nodes; The first computing node generates a sixth vector group based on the sixth vectors other than the fifth vector in the fifth vector group, and stores the sixth vector group in the multiple storage nodes.

13. The method according to claim 12, characterized in that, The vector deletion request is further used to indicate: The first computing node updates the first vector group based on the sixth vector group to obtain an updated first vector group, and the updated first vector group includes the sixth vector group; The first computing node notifies the cloud management platform that the first computing node stores the updated first vector group.

14. The method according to any one of claims 1 to 13, characterized in that, The cloud management platform determines the first computing node with the lowest load from the multiple computing nodes according to the load balancing algorithm based on the vector generation request and the status information of the multiple computing nodes, including: The cloud management platform calculates the status information of the multiple computing nodes according to the load balancing algorithm based on the vector generation request to obtain evaluation values of the multiple computing nodes, and the evaluation values of the multiple computing nodes are used to indicate the loads of the multiple computing nodes; The cloud management platform determines the computing node with the smallest evaluation value as the first computing node from the multiple computing nodes; Among them, for any one of the multiple computing nodes, the status information of the computing node includes at least one of the following: the time required for the computing node to load the vectors stored by itself, the time required for the computing node to load the indexes of the vectors stored by itself, the time required for the computing node to load the vectors stored in the multiple storage nodes, the time required for the computing node to load the indexes of the vectors stored in the multiple storage nodes, and the time required for the computing node to execute the vector processing request.

15. The method according to claim 14, characterized in that, The method further includes at least one of the following: The cloud management platform analyzes the vectors stored by the computing node and the vectors stored in the multiple storage nodes to obtain the time required for the computing node to load the vectors stored by itself and the time required for the computing node to load the vectors stored in the multiple storage nodes; The cloud management platform analyzes the indexes of the vectors stored by the computing node and the indexes of the vectors stored in the multiple storage nodes to obtain the time required for the computing node to load the indexes stored by itself and the time required for the computing node to load the indexes stored in the multiple storage nodes; The cloud management platform analyzes the resource usage of the computing nodes to obtain the time required for the computing nodes to execute the vector processing request.

16. A cloud management platform, characterized in that, The cloud management platform is used to manage the infrastructure that provides cloud services. The infrastructure includes multiple computing nodes and multiple storage nodes. The multiple storage nodes are used to store multiple data of tenants. The cloud management platform includes: a receiving module, configured to receive a vector generation request for the multiple data from the tenant; a processing module, configured to determine, based on the vector generation request and the status information of the multiple computing nodes, a first computing node with the lowest load from the multiple computing nodes according to a load balancing algorithm. The multiple computing nodes are all configured with the ability to process the vector generation request; The processing module is further configured to send the vector generation request to the first computing node. The vector generation request is used to instruct the first computing node to generate multiple vectors corresponding to the multiple data and store the multiple vectors in the multiple storage nodes.

17. The cloud management platform according to claim 16, characterized in that, The receiving module is further configured to receive a vector query request for a first vector from the tenant; The processing module is further configured to determine, based on the vector query request and the status information of the multiple computing nodes, the first computing node with the lowest load from the multiple computing nodes according to the load balancing algorithm. The multiple computing nodes are all configured with the ability to process the vector query request; The processing module is further configured to send the vector query request to the first computing node. The vector query request is used to instruct the first computing node to obtain a third vector matching the first vector from the second vector stored in itself or the multiple vectors stored in the multiple storage nodes. The multiple vectors include the second vector.

18. The cloud management platform according to claim 16 or 17, characterized in that The receiving module is further configured to receive a vector addition request for first data. The first data is new data written by the tenant into the multiple storage nodes; The processing module is further configured to determine, based on the vector addition request and the status information of the multiple computing nodes, the first computing node with the lowest load from the multiple computing nodes according to the load balancing algorithm. The multiple computing nodes are all configured with the ability to process the vector addition request; The processing module is further configured to send the vector addition request to the first computing node. The vector addition request is used to instruct the first computing node to generate a fourth vector corresponding to the first data and store the fourth vector in the multiple storage nodes.

19. The cloud management platform according to any one of claims 16 to 18, characterized in that, The receiving module is further configured to receive a vector deletion request for second data. The second data is data deleted by the tenant from the multiple data stored in the multiple storage nodes; The processing module is further configured to determine, based on the vector deletion request and the status information of the multiple computing nodes, the first computing node with the lowest load from the multiple computing nodes according to the load balancing algorithm. The multiple computing nodes are all configured with the ability to process the vector deletion request; The processing module is further configured to send the vector deletion request to the first computing node, where the vector deletion request is used to instruct the first computing node to delete the fifth vector corresponding to the second data from the multiple vectors stored in the multiple storage nodes.

20. The cloud management platform according to claim 18 or 19, characterized in that, The vector addition request or the vector deletion request comes from the multiple storage nodes.

21. The cloud management platform according to any one of claims 16 to 20, characterized in that, The processing module is configured to: Based on the vector generation request, calculate the status information of the multiple computing nodes according to the load balancing algorithm to obtain the evaluation values of the multiple computing nodes, where the evaluation values of the multiple computing nodes are used to indicate the loads of the multiple computing nodes; Determine the computing node with the smallest evaluation value as the first computing node from the multiple computing nodes; Wherein, for any one of the multiple computing nodes, the status information of the computing node includes at least one of the following: the time required for the computing node to load the vectors stored by itself, the time required for the computing node to load the indexes of the vectors stored by itself, the time required for the computing node to load the vectors stored in the multiple storage nodes, the time required for the computing node to load the indexes of the vectors stored in the multiple storage nodes, and the time required for the computing node to execute the vector processing request.

22. The cloud management platform according to claim 21, characterized in that, The cloud management platform further includes an analysis module, and the analysis module is configured to perform at least one of the following: Analyze the vectors stored in the computing node and the vectors stored in the multiple storage nodes to obtain the time required for the computing node to load the vectors stored by itself and the time required for the computing node to load the vectors stored in the multiple storage nodes; Analyze the indexes of the vectors stored in the computing node and the indexes of the vectors stored in the multiple storage nodes to obtain the time required for the computing node to load the indexes stored by itself and the time required for the computing node to load the indexes stored in the multiple storage nodes; Analyze the resource usage of the computing node to obtain the time required for the computing node to execute the vector processing request.

23. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device, and each computing device includes a processor and a memory: The memory is used to store instructions; The processor is configured to, according to the instructions, enable the computing device cluster to execute the method according to any one of claims 1 to 15.

24. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions, and when the instructions are executed by one or more computers, the one or more computers are enabled to implement the method according to any one of claims 1 to 15.

25. A computer program product, characterized in that, The computer program product stores instructions, and when the instructions are executed by a computer, the computer is enabled to implement the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Vector processing method based on cloud management platform and cloud management platform

    CN120343021A

  • Task scheduling method based on edge cloud, electronic equipment and storage medium

    CN115242798A

  • Dynamic feedback weighted cloud storage resource scheduling method, device and equipment

    CN115714817A

  • Service execution method, storage medium, equipment and distributed system

    CN116708583A

  • Cross-cluster load balancing method and device, equipment and storage medium

    CN117149445A