Algorithm service deployment method and apparatus, electronic device, and readable storage medium
By acquiring and processing the information of the target algorithm service on the first server side, determining the service access type and obtaining the corresponding access method, the problem that the existing cloud platform cannot support multiple algorithm service access types at the same time is solved, and a wider and more efficient service deployment is achieved.
Patent Information
- Application Number
- PCT/CN2024/099828
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-06-18
- Publication Date
- 2025-05-22
AI Technical Summary
Existing cloud platforms cannot support the deployment of algorithm service for multiple access methods at the same time, such as the access to application program interfaces (APIs), image files and model files.
By obtaining the service operation data and service information of the target algorithm service on the first server side, determining the service access type, and obtaining the corresponding access method according to the type, the algorithm service deployment for various ways of access is realized.
The service deployment scope and performance of the first server side are improved, allowing it to support the deployment of a variety of different algorithm service access types.
Smart Images

Figure CN2024099828_22052025_PF_FP_ABST
Abstract
Description
Deployment method, device, electronic device and readable storage medium for algorithm services
[0001] This application claims priority to a Chinese patent application with application date of November 13, 2023, application number 202311501500.9 and invention name “Deployment method, device, electronic device and readable storage medium for algorithm services”. Technical Field
[0002] The present disclosure relates to the field of Internet technology, particularly to artificial intelligence technologies such as cloud services, big data, and large language models. A method, device, electronic device, and readable storage medium for deploying an algorithm service are provided. Background Art
[0003] In addition to the ability to provide algorithm services to the outside world, cloud platforms in the existing technology also need to have the ability to deploy the external algorithm services that are connected. Usually, the external algorithm services that are connected include access in the form of application program interfaces, access in the form of image files, or access in the form of model files. The cloud platforms in the existing technology only support the deployment of algorithm services that are accessed in a certain way, and cannot realize the deployment of algorithm services that are accessed in multiple different ways.
[0004] Summary of the Invention
[0005] According to a first aspect of the present disclosure, a method for deploying an algorithm service is provided, including: a first server end obtains service operation data and service information of a target algorithm service, the service operation data including one of an application program interface, a mirror file, and a model file; a service access type is determined according to the service nature of the target algorithm service, and a target access method corresponding to the service access type is obtained, the service access type including one of application program interface access, mirror access, and model access; and the target algorithm service is deployed on the first server end according to the target access method, the service operation data, and the service information.
[0006] According to a second aspect of the present disclosure, a deployment device for an algorithm service is provided, comprising: an acquisition unit for acquiring service operation data and service information of a target algorithm service, wherein the service operation data includes one of an application program interface, a mirror file, and a model file; a processing unit for determining a service access type according to the service nature of the target algorithm service, and acquiring a target access method corresponding to the service access type, wherein the service access type includes one of application program interface access, mirror access, and model access; and a deployment unit for deploying the target algorithm service on the first server end according to the target access method, the service operation data, and the service information.
[0007] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0008] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method as described above.
[0009] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method described above when executed by a processor.
[0010] It can be seen from the above technical solutions that the first server in the present disclosure supports the deployment of algorithm services corresponding to different service access types, which can improve the service deployment scope of the first server and enable the first server to have stronger service deployment performance.
[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0013] FIG1 is a schematic diagram of a first embodiment of the present disclosure;
[0014] FIG2 is a schematic diagram of a second embodiment of the present disclosure;
[0015] FIG3 is a schematic diagram of a third embodiment of the present disclosure;
[0016] FIG4 is a schematic diagram of a fourth embodiment according to the present disclosure;
[0017] FIG5 is a block diagram of an electronic device for implementing the algorithm service deployment method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0018] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, and various details of the embodiments of the present disclosure are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and mechanisms are omitted in the following description.
[0019] FIG1 is a schematic diagram of a first embodiment of the present disclosure. As shown in FIG1 , the algorithm service deployment method of this embodiment specifically includes the following steps:
[0020] S101: A first server obtains service operation data and service information of a target algorithm service, wherein the service operation data includes one of an application program interface, an image file, and a model file;
[0021] S102: Determine a service access type based on the service nature of the target algorithm service, and obtain a target access method corresponding to the service access type, where the service access type includes one of application program interface access, mirror access, and model access.
[0022] S103: Deploy the target algorithm service on the first server according to the target access method, the service operation data, and the service information.
[0023] The execution subject of the algorithm service deployment method of this embodiment is the first server end. After obtaining the service operation data and service information of the target algorithm service, the first server end obtains the target access method according to the service access type corresponding to the target algorithm service, and then deploys the target algorithm service on the first server end according to the target access method, service operation data and service information, so that external users can access and use the target algorithm service deployed on the first server end. The first server end in this embodiment supports the deployment of algorithm services corresponding to different service access types, which can enhance the service deployment scope of the first server end and enable the first server end to have stronger service deployment performance.
[0024] In this embodiment, the first server can be a virtual machine, an independent server, or a node located in a Kubernetes (K8s) cluster; the target algorithm service to be deployed in the first server is a service based on a neural network model or a large language model for natural language processing (text recognition, text generation, etc.), computer vision (image recognition, image detection) and other algorithms.
[0025] When executing S101 in this embodiment, the first server end can first obtain the identification information of the algorithm service based on the input or selection of the input end. The identification information can be the name of the algorithm service, the application field of the algorithm service, etc., and then use the algorithm service corresponding to the obtained identification information as the target algorithm service.
[0026] When executing S101 in this embodiment, the first server can obtain the service operation data and service information of the target algorithm service from the second server by accessing the second server, that is, the second server pre-stores the service operation data and service information corresponding to different algorithm services.
[0027] In addition, when executing S101 in this embodiment, the first server can also obtain the service operation data and service information of the target algorithm service from the first server by accessing the local area, that is, the first server pre-stores the service operation data and service information corresponding to different algorithm services.
[0028] In this embodiment, the service operation data of the target algorithm service is the relevant data used to run the neural network model or large language model corresponding to the target algorithm service, including one of an API (Application Programming Interface), a mirror file (running the mirror file can provide the algorithm service corresponding to the mirror file to the outside world), and a model file (the model file corresponds to the model providing the algorithm service, for example, a neural network model or a large language model); the service information of the target algorithm service is the information required to access the target algorithm service deployed on the first server end, including information such as the external request address, interface call parameters, request method, and authentication method.
[0029] After executing S101 to obtain the service operation data and service information of the target algorithm service, the first server of this embodiment executes S102 to determine the service access type according to the service nature of the target algorithm service, and obtains the target access method corresponding to the service access type; wherein, the service access type in this embodiment includes one of API access, mirror access and model access.
[0030] When executing S102 in this embodiment, the service nature of the target algorithm service is one of providing services based on API, providing services based on mirror files, and providing services based on model files; among which, the access type corresponding to providing services based on API is API access, the access type corresponding to providing services based on mirror files is mirror access, and the access type corresponding to providing services based on model files is model access.
[0031] The service nature of the target algorithm service can be determined based on the identification information of the target algorithm service. For example, the identification information clearly indicates the content based on which the target algorithm service provides the service. It can also be determined based on the data type of the service operation data.
[0032] For example, if the service operation data obtained is an API, the first server determines that the service nature is to provide services based on the API; if the service operation data obtained is an image file, the first server determines that the service nature is to provide services based on the image file; if the service operation data obtained is a model file, the first server determines that the service nature is to provide services based on the model file.
[0033] In this embodiment, when executing S102 , after the first server determines the service access type, a target access method is acquired according to the determined service access type.
[0034] It can be understood that the first server end of this embodiment can pre-set access methods corresponding to different service access types. When executing S102, the first server end will use the access method corresponding to the determined service access type as the target access method; wherein different service access types correspond to different access methods.
[0035] That is to say, the first server end of this embodiment pre-sets access methods corresponding to different service access types, so that when the first server end faces algorithm services corresponding to different service access types, it can adopt an access method that matches the target algorithm service to be deployed for deployment, ensuring that algorithm services of different access types can be deployed, thereby improving the deployment accuracy and efficiency of the target algorithm service.
[0036] In this embodiment, after executing S102 to obtain the target access method corresponding to the service access type, the first server executes S103 to deploy the target algorithm service on the first server according to the target access method, service operation data, and service information.
[0037] In this embodiment, when executing S103, the target algorithm service is deployed on the first server according to the target access method, service operation data and service information. The implementation method that can be adopted is: when the service access type is API access, the API access information is configured in the first server according to the service information (such as interface call parameters, external request address, authentication method, request method, etc.) to complete the deployment of the target algorithm service in the first server.
[0038] That is to say, when the first server end of this embodiment determines that the target algorithm service is accessed in the form of an API, since the API can directly provide algorithm services, this embodiment only needs to configure the access information of the API according to the service information in the first server end to complete the deployment of the target algorithm service in the first server end.
[0039] In this embodiment, when executing S103, the target algorithm service is deployed on the first server according to the target access method, service operation data and service information. The implementation method that can be adopted is: when the service access type is mirror access, the target cluster (for example, K8s cluster) is opened, and the mirror file is run in the opened target cluster; according to the service information (for example, external request address, request method, authentication method), the access information of the mirror file is configured in the first server to complete the deployment of the target algorithm service in the first server.
[0040] That is to say, when the first server end of this embodiment determines that the target algorithm service is accessed in the form of a mirror file, since the mirror file cannot directly provide the algorithm service, the first server end of this embodiment needs to enable the target cluster to run the mirror file, and then configure the access information of the mirror file according to the service information, thereby completing the deployment of the target algorithm service in the first server end.
[0041] In this embodiment, when executing S103, the target algorithm service is deployed on the first server according to the target access method, service operation data and service information. The implementation method that can be adopted is: when the service access type is model access, the basic image file is obtained; according to the obtained basic image file, the model file is converted into a first target image file; the target cluster (for example, a K8s cluster) is opened, and the first target image file is run in the opened target cluster; according to the service information (for example, the external request address, the request method, and the authentication method), the access information of the first target image file is configured in the first server to complete the deployment of the target algorithm service in the first server.
[0042] That is to say, when the first server of this embodiment determines that the target algorithm service is accessed in the form of a model file, since the model file cannot directly provide the algorithm service, the first server of this embodiment needs to convert the model file into a mirror file, and then start the target cluster to run the converted mirror file, thereby completing the deployment of the target algorithm service in the first server.
[0043] It can be understood that this embodiment configures the access information of the API, mirror file and the first target mirror file in the first server according to the service information when executing S103, that is, the external request address, request method, authentication method, interface call parameters and other information required when accessing the target algorithm service are synchronized to the first server, so that external users can access the algorithm service deployed in the first server according to the configured access information.
[0044] When executing S103 to configure the access information of the API, mirror file and the first target mirror file in the first server according to the service information, this embodiment can first determine the gateway corresponding to the first server, and then synchronize the service information corresponding to the target algorithm service to the gateway, so that external users can access the algorithm service deployed in the first server through the gateway.
[0045] When executing S103, the first server of this embodiment may send the acquired image file to the image warehouse for storage; it may also send the acquired model file to the model warehouse for storage; it may also send the first target image file converted according to the model file to the image warehouse for storage; wherein, the image warehouse and the model warehouse may be located on the first server or on other servers independent of the first server.
[0046] After executing S103, the first server of this embodiment may further include the following contents: after receiving the update message sent by the second server, obtaining the model file to be updated corresponding to the algorithm service to be updated from the second server, the update message including the identification information of the algorithm service to be updated, different identification information is used to obtain different model files to be updated, and the model corresponding to the model file to be updated may be a large language model; according to the basic image file, converting the model file to be updated into a second target image file, the second target image file being the image file after the update; stopping the operation of the first target image file corresponding to the algorithm service to be updated in the target cluster, that is, temporarily shutting down the algorithm service to be updated; after replacing the first target image file with the second target image file, running the second target image file in the target cluster, that is, reopening the algorithm service to be updated, thereby completing the update of the algorithm service to be updated.
[0047] That is to say, the first server of this embodiment can also use the model file to be updated to update the running algorithm service according to the update information sent by the second server, so that the updated algorithm service has better performance, such as better text recognition performance, text generation performance, image recognition performance, image detection performance, etc., thereby improving the user experience of external users.
[0048] Figure 2 is a schematic diagram of the second embodiment of the present disclosure. Figure 2 shows a structural diagram of the first server of this embodiment when deploying an algorithm service. The first server includes a gateway, an API management module, a model management module, an image management module, an image warehouse, a model warehouse and a storage module, etc.; wherein, the gateway is used to synchronize the service information of the target algorithm service so that the client can access the target algorithm service contained in the service management module; the service management module is used to manage the deployed target algorithm service; the API management module is used to complete the deployment of the target algorithm service accessed by API; the model management module is used to complete the deployment of the target algorithm service accessed by model file, and the model management module can also include a model conversion module for converting the model file into an image file; the image management module is used to complete the deployment of the target algorithm service accessed by image file; the image warehouse is used to store image files; the model warehouse is used to store model files; the storage module is used to store the service information of the target algorithm service, and then synchronize the stored service information to the gateway.
[0049] Figure 3 is a schematic diagram of a third embodiment of the present disclosure. Figure 3 shows a structural diagram of the second server sending an update message to the first server in this embodiment: the second server includes a scheduling module, a training module, a testing module, an update module, a storage module, and a model repository.
[0050] Among them, the scheduling module is used to execute the following contents: 1) sending training instructions to the training module, so that the training module trains the model corresponding to the algorithm service to be updated (the model corresponding to the algorithm service in this embodiment can be a large language model) to obtain the model file to be updated; 2) sending test instructions to the testing module, so that the testing module compares the old model file with the model corresponding to the model file to be updated; 3) sending update instructions to the update module, so that the update module sends an update message to the first server.
[0051] The training module is used to perform the following: 1) After receiving the training instruction sent by the scheduling module, it obtains the old model file from the model warehouse and the training data from the storage module; 2) Use the training data to train the model corresponding to the old model file (the model corresponding to the model file can be a large language model) to obtain the model file to be updated, and store the model file to be updated in the model warehouse; 3) Send feedback information to the scheduling module that the model has completed training.
[0052] The testing module is used to perform the following: 1) After receiving the test instructions sent by the scheduling module, obtain the old model file and the model file to be updated from the model warehouse; 2) obtain test data from the storage module; 3) use the test data to obtain the reasoning performance of the models corresponding to the old model file and the model file to be updated respectively (the model corresponding to the model file can be a large language model); 4) when it is determined that the reasoning performance of the two models meets the preset requirements, send feedback information to the scheduling module that the model file to be updated has passed the test.
[0053] The update module is used to execute the following: 1) after receiving the update instruction sent by the scheduling module, replace the old model file in the model warehouse with the model file to be updated; 2) send an update message to the first server.
[0054] FIG4 is a schematic diagram of a fourth embodiment of the present disclosure. As shown in FIG4 , the algorithm service deployment device 400 of this embodiment is located at the first service end and includes:
[0055] An acquisition unit 401 is used to acquire service operation data and service information of a target algorithm service, wherein the service operation data includes one of an application program interface, an image file, and a model file;
[0056] Processing unit 402 is configured to determine a service access type according to a service property of the target algorithm service, and obtain a target access method corresponding to the service access type, wherein the service access type includes one of application program interface access, mirror access, and model access;
[0057] The deployment unit 403 is configured to deploy the target algorithm service on the first server according to the target access method, the service operation data and the service information.
[0058] The acquisition unit 401 may first acquire identification information of the algorithm service according to the input or selection of the input terminal. The identification information may be the name of the algorithm service, the application field of the algorithm service, etc., and then use the algorithm service corresponding to the acquired identification information as the target algorithm service.
[0059] The acquisition unit 401 can acquire the service operation data and service information of the target algorithm service from the second server by accessing the second server. That is, the second server pre-stores the service operation data and service information corresponding to different algorithm services.
[0060] In addition, the acquisition unit 401 can also obtain the service operation data and service information of the target algorithm service from the first server by accessing the local server, that is, the first server pre-stores the service operation data and service information corresponding to different algorithm services.
[0061] After the first server end of this embodiment obtains the service operation data and service information of the target algorithm service by the acquisition unit 401, the processing unit 402 determines the service access type according to the service nature of the target algorithm service, and obtains the target access method corresponding to the service access type; wherein, the service access type in this embodiment includes one of API access, mirror access and model access.
[0062] The service nature of the target algorithm service is one of API-based service, mirror file-based service and model file-based service; among them, the access type corresponding to API-based service is API access, the access type corresponding to mirror file-based service is mirror access, and the access type corresponding to model file-based service is model access.
[0063] The processing unit 402 can determine the service nature based on the identification information of the target algorithm service, for example, the identification information clearly indicates the content based on which the target algorithm service provides the service; the processing unit 402 can also determine the service nature based on the data type of the service operation data.
[0064] After the first server determines the service access type, the processing unit 402 acquires a target access method according to the determined service access type.
[0065] It can be understood that the first server end of this embodiment can pre-set access methods corresponding to different service access types, and the processing unit 402 will use the access method corresponding to the determined service access type as the target access method; wherein different service access types correspond to different access methods.
[0066] That is to say, the processing unit 402 pre-sets access methods corresponding to different service access types so that when the first server faces algorithm services corresponding to different service access types, it can adopt an access method that matches the target algorithm service to be deployed for deployment, ensuring that algorithm services of different access types can be deployed, thereby improving the deployment accuracy and efficiency of the target algorithm service.
[0067] In this embodiment, after the processing unit 402 obtains the target access method corresponding to the service access type, the deployment unit 403 deploys the target algorithm service on the first server according to the target access method, service operation data and service information.
[0068] When the deployment unit 403 deploys the target algorithm service on the first server side according to the target access method, service operation data and service information, the implementation method that can be adopted is: when the service access type is API access, the API access information is configured in the first server side according to the service information (such as interface call parameters, external request address, authentication method, request method, etc.) to complete the deployment of the target algorithm service in the first server side.
[0069] That is to say, when the deployment unit 403 determines that the target algorithm service is accessed in the form of an API, since the API can directly provide algorithm services to the outside world, this embodiment only needs to configure the access information of the API corresponding to the target algorithm service obtained according to the service information in the first server end to complete the deployment of the target algorithm service in the first server end.
[0070] When the deployment unit 403 deploys the target algorithm service on the first server end according to the target access method, service operation data and service information, the implementation method that can be adopted is: when the service access type is mirror access, open the target cluster (for example, K8s cluster) and run the mirror file in the opened target cluster; configure the access information of the mirror file in the first server end according to the service information (for example, external request address, request method, authentication method) to complete the deployment of the target algorithm service in the first server end.
[0071] That is to say, when the deployment unit 403 determines that the target algorithm service is accessed in the form of a mirror file, since the mirror file cannot directly provide the algorithm service, the first server end of this embodiment needs to enable the target cluster to run the mirror file, and then configure the access information of the mirror file corresponding to the target algorithm service according to the service information, thereby completing the deployment of the target algorithm service in the first server end.
[0072] When the deployment unit 403 deploys the target algorithm service on the first server side according to the target access method, service operation data and service information, the implementation method that can be adopted is: when the service access type is model access, obtain the basic image file; according to the obtained basic image file, convert the model file into a first target image file; open the target cluster (for example, a K8s cluster), and run the first target image file in the opened target cluster; configure the access information of the first target image file in the first server side according to the service information (for example, the external request address, the request method, and the authentication method), and complete the deployment of the target algorithm service in the first server side.
[0073] That is to say, when the deployment unit 403 determines that the target algorithm service is accessed in the form of a model file, since the model file cannot directly provide the algorithm service, the first server end of this embodiment needs to convert the model file corresponding to the target algorithm service into a mirror file, and then start the target cluster to run the converted mirror file, thereby completing the deployment of the target algorithm service in the first server end.
[0074] It can be understood that the deployment unit 403 configures the access information of the API, mirror file and the first target mirror file in the first server according to the service information, that is, the external request address, request method, authentication method, interface call parameters and other information required when accessing the target algorithm service are synchronized to the first server, so that external users can access the algorithm service deployed in the first server according to the configured access information.
[0075] When the deployment unit 403 configures the access information of the API, image file and the first target image file in the first server according to the service information, it can first determine the gateway corresponding to the first server, and then synchronize the service information corresponding to the target algorithm service to the gateway. External users can access the algorithm service deployed in the first server through the gateway.
[0076] The deployment unit 403 can send the obtained image file to the image warehouse for storage; it can also send the obtained model file to the model warehouse for storage, and it can also send the first target image file converted according to the model file to the image warehouse for storage; wherein, the image warehouse and the model warehouse can be located on the first server end, or on other server ends independent of the first server end.
[0077] The algorithm service deployment device 400 of this embodiment may further include an update unit 404, which is used to perform the following: after receiving an update message sent by the second server, obtain the model file to be updated corresponding to the algorithm service to be updated from the second server, and the update message includes identification information of the algorithm service to be updated; according to the basic image file, convert the model file to be updated into a second target image file; stop the operation of the first target image file corresponding to the algorithm service to be updated in the target cluster, that is, temporarily shut down the algorithm service to be updated; after replacing the first target image file with the second target image file, run the second target image file in the target cluster.
[0078] That is to say, the update unit 404 can also use the model file to be updated (i.e., the model file corresponding to the retrained model) to update the running algorithm service according to the update information sent by the second server, so that the updated algorithm service has better performance, such as better text recognition performance, text generation performance, image recognition performance, image detection performance, etc., thereby improving the user experience of external users.
[0079] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0080] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0081] As shown in Figure 5, it is a block diagram of an electronic device according to the deployment method of the algorithm service of an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0082] As shown in Figure 5, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0083] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0084] The computing unit 501 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as the deployment method of the algorithm service. For example, in some embodiments, the deployment method of the algorithm service can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508.
[0085] In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the algorithm service deployment method described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to execute the algorithm service deployment method in any other appropriate manner (e.g., by means of firmware).
[0086] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0087] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. Such program code may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other device that deploys a programmable algorithm service, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0088] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0089] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0090] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0091] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service system that addresses the management difficulties and poor business scalability of traditional physical hosts and VPS services ("Virtual Private Servers," or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0092] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0093] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for deploying an algorithm service, comprising: The first server obtains service operation data and service information of the target algorithm service, wherein the service operation data includes one of an application program interface, an image file, and a model file; Determine a service access type according to the service property of the target algorithm service, and obtain a target access method corresponding to the service access type, wherein the service access type includes one of application program interface access, mirror access, and model access; The target algorithm service is deployed on the first server according to the target access method, the service operation data and the service information.
2. The method according to claim 1, wherein: The deploying the target algorithm service on the first server according to the target access method, the service operation data and the service information includes: In the case where the service access type is application program interface access, the access information of the application program interface is configured in the first server according to the service information to complete the deployment of the target algorithm service in the first server.
3. The method according to claim 1, wherein: The deploying the target algorithm service on the first server according to the target access method, the service operation data and the service information includes: When the service access type is mirror access, starting the target cluster and running the mirror file in the target cluster; The access information of the image file is configured in the first server according to the service information, and the deployment of the target algorithm service in the first server is completed.
4. The method according to claim 1, wherein: The deploying the target algorithm service on the first server according to the target access method, the service operation data and the service information includes: When the service access type is model access, obtaining a basic image file; Converting the model file into a first target image file according to the base image file; Starting the target cluster and running the first target image file in the target cluster; The access information of the first target image file is configured in the first server according to the service information, and the deployment of the target algorithm service in the first server is completed.
5. The method according to any one of claims 2 to 4, wherein: The configuring the access information of the application program interface, the image file or the first target image file in the first server according to the service information includes: Determining a gateway corresponding to the first server; The service information is synchronized to the gateway to complete configuration of the application program interface, the image file or the first target image file access information in the first server.
6. The method according to claim 4, further comprising: After receiving the update message sent by the second server, obtaining the to-be-updated model file corresponding to the to-be-updated algorithm service from the second server; Converting the to-be-updated model file into a second target image file according to the base image file; Stopping the operation of the first target image file corresponding to the algorithm service to be updated in the target cluster; After replacing the first target image file with the second target image file, the second target image file is run in the target cluster.
7. The method according to claim 1, wherein: The model corresponding to the model file is a large language model.
8. A deployment device for an algorithm service, located at a first service end, comprising: An acquisition unit, used to acquire service operation data and service information of a target algorithm service, wherein the service operation data includes one of an application program interface, an image file, and a model file; A processing unit, configured to determine a service access type according to a service property of the target algorithm service, and obtain a target access method corresponding to the service access type, wherein the service access type includes one of application program interface access, mirror access, and model access; A deployment unit is used to deploy the target algorithm service on the first server according to the target access method, the service operation data and the service information.
9. The device according to claim 8, wherein: When the deployment unit deploys the target algorithm service on the first server according to the target access method, the service operation data and the service information, the deployment unit specifically performs: In the case where the service access type is application program interface access, the access information of the application program interface is configured in the first server according to the service information to complete the deployment of the target algorithm service in the first server.
10. The device according to claim 8, wherein: When the deployment unit deploys the target algorithm service on the first server according to the target access method, the service operation data and the service information, the deployment unit specifically performs: When the service access type is mirror access, starting the target cluster and running the mirror file in the target cluster; The access information of the image file is configured in the first server according to the service information, and the deployment of the target algorithm service in the first server is completed.
11. The device according to claim 8, wherein: When the deployment unit deploys the target algorithm service on the first server according to the target access method, the service operation data and the service information, the deployment unit specifically performs: When the service access type is model access, obtaining a basic image file; Converting the model file into a first target image file according to the base image file; Starting the target cluster and running the first target image file in the target cluster; The access information of the first target image file is configured in the first server according to the service information, and the deployment of the target algorithm service in the first server is completed.
12. The device according to any one of claims 9 to 11, wherein: When the deployment unit configures the access information of the application program interface, the image file, or the first target image file in the first server according to the service information, specifically: Determining a gateway corresponding to the first server; The service information is synchronized to the gateway to complete configuration of the access information of the application program interface, the image file or the first target image file in the first server.
13. The apparatus according to claim 8, further comprising an updating unit, configured to execute: After receiving the update message sent by the second server, obtaining the to-be-updated model file corresponding to the to-be-updated algorithm service from the second server; Converting the to-be-updated model file into a second target image file according to the base image file; Stopping the operation of the first target image file corresponding to the algorithm service to be updated in the target cluster; After replacing the first target image file with the second target image file, the second target image file is run in the target cluster.
14. The device according to claim 8, wherein: The model corresponding to the model file acquired by the acquisition unit is a large language model.
15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
17. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Project deployment method and device
CN114416112A
Private deployment method and device suitable for multiple clouds
CN114679458A
Service deployment method of hybrid cloud platform, management platform, equipment and medium
CN116684421A
Algorithm service deployment method and device, electronic equipment and readable storage medium
CN117591275A
Determining root causes of anomalies in services
WO2023154051A1