Neural network model task processing method and device, storage medium and program
By adopting a shared storage mechanism in the FaaS system, the neural network model is preloaded into the shared storage space of the cloud server, which solves the problems of long model loading time and bandwidth bottleneck in high-concurrency scenarios, and achieves faster task processing response and improved resource utilization.
Patent Information
- Application Number
- CN202410621623.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-11-18
AI Technical Summary
In high-concurrency scenarios, existing serverless architecture systems experience performance bottlenecks when loading large neural network models, including long model loading times and a surge in bandwidth requirements, which affect the response time of task processing requests and user experience.
By utilizing a shared storage mechanism on the cloud server in the FaaS system, neural network models can be mapped to local shared storage space, allowing different task processing instances to share the same model, thereby reducing loading time and bandwidth consumption.
Local shared storage reduces the loading time of neural network models, improves the system's concurrent processing capabilities and user experience, while ensuring data security and isolation.
Smart Images

Figure CN120973428A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, device, storage medium, and program for processing neural network model tasks. Background Technology
[0002] With the rapid development of artificial intelligence (AI) technology, more and more applications need to integrate AI models (usually neural network models) to provide functions such as data analysis, image processing, and natural language processing. This integration is mainly manifested in the following ways: the application code is deployed on a computing platform, and the neural network model may have a large amount of data, reaching hundreds of gigabytes. Therefore, the neural network model is often stored in remote storage space, such as object storage or network storage. When the application code is triggered by the user, the neural network model is loaded from the remote storage space on demand.
[0003] This approach can lead to performance bottlenecks in high-concurrency scenarios. For example, in a high-concurrency scenario where many users simultaneously trigger task processing requests to the application, loading a neural network model from remote storage for each task processing request can be time-consuming due to the large size of the neural network model, which requires significant network bandwidth and results in a long loading time. This can hinder timely responses to task processing requests and negatively impact user experience. Summary of the Invention
[0004] This invention provides a method, device, storage medium, and program for processing neural network model tasks, in order to reduce the loading time of neural network models.
[0005] In a first aspect, embodiments of the present invention provide a neural network model task processing method, applied to a function computing service system in a serverless architecture system, the method comprising:
[0006] Obtain the target application code and target application configuration file input by the user, wherein the target application configuration file includes the storage address of the neural network model called in the target application code in the remote storage space;
[0007] Identify the first cloud server in the cloud server cluster used to deploy the application;
[0008] Based on the storage address, the neural network model is loaded from the remote storage space into a shared storage space created in the first cloud server. The shared storage space can be accessed by all task processing instances running in the first cloud server.
[0009] In response to a task processing request for the target application code, a task processing instance corresponding to the task processing request is created in the first cloud server, so that the task processing instance loads the neural network model from the shared storage space and performs task processing based on the neural network model.
[0010] Secondly, embodiments of the present invention provide a neural network model task processing device, applied to a function computing service system in a serverless architecture system, the device comprising:
[0011] The acquisition module is used to acquire the target application code and the target application configuration file input by the user. The target application configuration file includes the storage address of the neural network model called in the target application code in the remote storage space.
[0012] A loading module is used to determine a first cloud server in a cloud server cluster used for deploying applications, and load the neural network model from the remote storage space into a shared storage space created in the first cloud server according to the storage address. The shared storage space can be accessed by all task processing instances running in the first cloud server.
[0013] The processing module is configured to, in response to a task processing request for the target application code, create a task processing instance in the first cloud server corresponding to the task processing request, so that the task processing instance loads the neural network model from the shared storage space and performs task processing based on the neural network model.
[0014] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a communication interface; wherein, the memory stores executable code, and when the executable code is executed by the processor, the processor can at least implement the neural network model task processing method as described in the first aspect.
[0015] Fourthly, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, wherein when the executable code is executed by a processor of an electronic device, the processor is able to at least implement the neural network model task processing method as described in the first aspect.
[0016] Fifthly, embodiments of the present invention provide a computer program product comprising a computer program that, when executed by a processor of an electronic device, enables the processor to at least implement the neural network model task processing method as described in the first aspect.
[0017] The neural network model task processing method provided in this invention is applied to a Function as a Service (FaaS) system within a serverless computing system. This FaaS system includes a cloud server cluster for deploying applications. Deploying a target application in the FaaS system primarily involves inputting the target application code and a target application configuration file. The target application consists of two parts: code and a configuration file. The configuration file includes the storage address of the neural network model called by the target application code in a remote storage space. Based on the user's input, the FaaS system can determine the first cloud server within its cloud server cluster based on the load of each cloud server, and load the neural network model from the remote storage space into a shared storage space created on the first cloud server. Upon receiving a task processing request for the target application code, a task processing instance corresponding to the request can be created on the first cloud server. This task processing instance loads the target application code neural network model from the shared storage space and performs task processing based on the neural network model.
[0018] By creating a shared storage space on the first cloud server, the neural network model is pre-loaded into this space for use by all task processing instances on the first cloud server that require it. This allows task processing instances to load the neural network model from the local storage space on the first cloud server. Compared to multiple task processing instances loading the model separately from remote storage spaces, this significantly reduces network bandwidth and loading time, enabling more timely responses to task processing requests. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the composition of the FaaS system provided in an embodiment of the present invention;
[0021] Figure 2 A flowchart of a neural network model task processing method provided in an embodiment of the present invention;
[0022] Figure 3 A flowchart of an instance creation process provided in an embodiment of the present invention;
[0023] Figure 4 A flowchart of another neural network model task processing method provided in an embodiment of the present invention;
[0024] Figure 5 A flowchart of a verification process provided in an embodiment of the present invention;
[0025] Figure 6 A flowchart of an example preheating process provided in an embodiment of the present invention;
[0026] Figure 7 This is a schematic diagram of the structure of a neural network model task processing device provided in an embodiment of the present invention;
[0027] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0030] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Where there is no conflict between the embodiments, the following embodiments and features can be combined with each other. Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.
[0031] Serverless computing systems consist of a Function as a Service (FaaS) system and a Backend as a Service (BaaS) system. Through the BaaS system, cloud service providers offer users (developers) various cloud backend services, such as storage services, database services, messaging services, network services, and security services. Through the FaaS system, cloud service providers offer a platform that allows users to focus solely on application development and execution without managing the underlying cloud servers involved in application development and startup. For example, it eliminates the need to build and maintain the runtime environment on the cloud servers where the application resides, thereby improving application development efficiency and the elasticity of server resource utilization.
[0032] Therefore, an increasing number of users are choosing to use FaaS systems for application development and deployment. These applications may integrate AI models (usually neural network models) to provide functions such as data analysis, image processing, and natural language processing.
[0033] Currently, FaaS systems focus solely on "function computing" capabilities, while services such as large-scale data storage are provided by BaaS systems. Therefore, FaaS systems do not have large storage capacity, especially the cloud server clusters used to deploy user application code. Since application code is typically small (usually in the KB range), it does not require a large amount of disk storage. Neural network models, on the other hand, can be stored in BaaS systems or other remote storage environments.
[0034] Therefore, despite the numerous advantages offered by serverless computing systems, several challenges remain when running large neural network models. For example, long model loading times: Neural network models are often enormous; many current large language models, for instance, require significant time to load from remote storage into the application's execution environment within the FaaS system. This leads to cold start issues, impacting task processing request response times and user experience. Another challenge is bandwidth bottlenecks: In high-concurrency scenarios, receiving a large number of task processing requests simultaneously and attempting to load the required neural network models from remote storage can cause a surge in bandwidth demands, potentially encountering network bottlenecks and limiting the ability to process task requests.
[0035] In view of this, embodiments of the present invention provide a model sharing scheme, which maps neural network models to local shared storage space (such as a fixed directory) on a cloud server in a FaaS system using a shared storage mechanism. This allows different task processing instances to share the same neural network model without repeatedly downloading it from remote storage. This not only significantly reduces the loading time of neural network models and solves the cold start problem, but also effectively reduces bandwidth consumption, improves the system's concurrent processing capabilities, and ensures data security and isolation.
[0036] The following describes a model sharing scheme provided by an embodiment of the present invention. This scheme can be applied to various applications developed based on a FaaS system within a serverless computing (Server-as-a-Service) system, and the model sharing method can be executed by the FaaS system. Firstly, in conjunction with... Figure 1 This section provides a brief overview of the components of a FaaS system.
[0037] From a macro perspective, based on the services provided, a FaaS system can include various service clusters, such as Figure 1 The diagram illustrates scheduling service clusters, storage service clusters, user application runtime service clusters, etc. Each service cluster can contain multiple cloud servers, and the computing and storage resources provided by each type of service cluster can differ. Specifically, the user application runtime service cluster is the cloud server cluster used to deploy applications, while the user's application code and configuration files can be stored in the storage service cluster. The scheduling service cluster can be used for deploying applications and allocating resources within the user application runtime service cluster. Therefore, the model sharing method provided in this embodiment can be executed by this scheduling service cluster. However, this is not a limitation; when some scheduling services are provided by the user application runtime service cluster, they can also be executed by that user application runtime service cluster.
[0038] In fact, taking a user application running service cluster as an example, the cloud servers in this service cluster can be divided into availability zones. That is, the service cluster can include multiple availability zones, each of which corresponds to a different location region, such as different availability zones like Hangzhou and Shanghai. Each availability zone can contain multiple cloud servers.
[0039] Figure 2 A flowchart of a neural network model task processing method provided in an embodiment of the present invention is shown below. Figure 2 As shown, the method may include the following steps:
[0040] 201. Obtain the target application code and target application configuration file input by the user. The target application configuration file includes the storage address of the neural network model called in the target application code in the remote storage space.
[0041] 202. Identify the first cloud server in the cloud server cluster used to deploy the application.
[0042] 203. Based on the storage address of the neural network model in the remote storage space, load the neural network model from the remote storage space into the shared storage space created in the first cloud server.
[0043] 204. In response to a task processing request for the target application code, a task processing instance corresponding to the task processing request is created in the first cloud server, so that the task processing instance loads the neural network model from the shared storage space and performs task processing based on the neural network model.
[0044] The FaaS system can provide an application configuration interface where users (developers) can configure their applications. In this embodiment, any user-configured application is referred to as the target application. The configuration interface allows users to input the target application's code and configuration file: the target application code and the target application configuration file. In other words, the target application code calls a neural network model; that is, the target application code contains code statements that call a neural network model.
[0045] The target application code describes the task processing logic. For example, if the target application is a program that provides image classification functionality, then the target application code describes the logic of how to classify the input images to obtain the classification results.
[0046] The neural network model invoked by the target application code is initially stored in remote storage, which is not located within the FaaS system. The target application configuration file includes the storage address of the neural network model invoked by the target application code in the remote storage.
[0047] The target application configuration file includes relevant configuration information for the target application. In addition to the storage address of the neural network model in the remote storage space, it may also include, but is not limited to: user identification (i.e., developer identification) information, language information used in the code, required runtime environment information, required computing resource information (such as memory, CPU, etc.), etc.
[0048] After obtaining the target application code and target application configuration file entered by the user in the above configuration interface, the FaaS system can store them in the system.
[0049] Subsequently, the FaaS system can identify a cloud server as the first cloud server from the cloud server cluster used to deploy applications (such as the user application runtime service cluster mentioned above). Based on the storage address of the aforementioned neural network model in the remote storage space, the system loads the neural network model into the shared storage space created on the first cloud server. The first cloud server can be randomly selected from the cloud server cluster, or a server with a lower load can be chosen based on the load of each cloud server.
[0050] If the shared storage space has not yet been created on the first cloud server, then the shared storage space will be created; if it has already been created, then the neural network model can be directly loaded into it.
[0051] Optionally, a shared storage space can be pre-configured on each cloud server in the aforementioned cloud server cluster. This shared storage space can be a specific directory on the local hard drive of the cloud server, specifically used to store neural network models, i.e., to store model data, such as weights and metadata information, including but not limited to the model data version number, update time, etc.
[0052] The aforementioned shared storage space is used to store one or more neural network models. That is, the shared storage space can store different neural network models corresponding to different applications. Therefore, the storage capacity of the shared storage space needs to be large. Thus, based on the existing FaaS system architecture, it is necessary to configure large-capacity storage media for each cloud server in the aforementioned cloud server cluster.
[0053] In practical applications, taking the first cloud server as an example, assuming that the first cloud server includes at least one large-capacity storage medium, such as multiple disks, then a disk can be selected to create a shared storage space, which can be a file system directory.
[0054] Furthermore, in practical applications, the FaaS system can complete the aforementioned operation of loading the neural network model from the remote storage space to the shared storage space during periods of low task processing load, thus avoiding impact on running task processing instances. In other words, the model loading process is not performed in real-time after obtaining the target application code and configuration file input by a user.
[0055] It should be noted that at this point, the aforementioned neural network model has only been "pre-loaded" from the remote storage space to the local shared storage space of the first cloud server, and the target application code and target application configuration files have been stored in a storage location within the FaaS system, such as the storage service cluster mentioned above. The target application has not yet been deployed or is not yet running. "Pre-loading" here means that the loading process is completed before the target application is deployed.
[0056] The target application configured by the user (developer) is essentially a service that can be used by different users. For example, an application developed by a developer that provides certain image processing functions can be used by a wide range of users to complete their respective image processing tasks.
[0057] Therefore, different users may simultaneously trigger task processing requests targeting the aforementioned target application code (i.e., the target application). These task processing requests may include user identification information corresponding to the target application. Upon receiving the task processing request, the FaaS system creates a corresponding task processing instance to process the request.
[0058] In practical applications, users can trigger the above task processing requests through various event channels, such as HTTP requests, message queues, database events, and so on.
[0059] In this scenario, if N task processing requests for the target application code are received simultaneously (N>1), and if there are currently no idle task processing instances corresponding to the target application code, then the FaaS system needs to create N task processing instances to respond to these N task processing requests in parallel.
[0060] Furthermore, assuming the neural network model corresponding to the aforementioned target application code has already been loaded into the shared storage space of the first cloud server, the N task processing instances corresponding to the aforementioned N task processing requests can all reside on the first cloud server. Of course, if the load on the first cloud server exceeds a set threshold, task processing instances corresponding to some task processing requests can also be created on other cloud servers, and the neural network model can be stored in the shared storage space of those other cloud servers, as will be explained in subsequent embodiments.
[0061] For ease of description, taking any task processing request triggered by the target application code as an example, after creating a task processing instance corresponding to the task processing request in the first cloud server, the task processing instance can load the neural network model corresponding to the target application code from the shared storage space in the first cloud server and perform task processing based on the neural network model.
[0062] like Figure 3 As shown, the process of creating a task processing instance may include the following steps:
[0063] 301. In response to a task processing request for the target application code, determine the target application configuration file and the target application code based on the user identification information in the task processing request.
[0064] 302. Create a container on the first cloud server according to the target application configuration file and allocate computing resources. Deploy the target application code into the container to form a task processing instance corresponding to the task processing request.
[0065] 303. Configure an access path for the task processing instance to access the neural network model in the shared storage space, so that the task processing instance can load the neural network model from the shared storage space according to the access path.
[0066] Since different applications have different user identification information, parsing the user identification information from the task processing request allows us to determine the corresponding target application, namely the target application code and the target application configuration file. Then, based on the configuration information described in the target application configuration file, a task processing instance is created: first, a corresponding container is created based on the code execution environment requirements in the configuration file, and corresponding computing resources are allocated (e.g., according to the computing resources required in the configuration file) to deploy the target application code into this container, thus forming a task processing instance. To enable this task processing instance to access the neural network model corresponding to the target application code in the shared storage space of the first cloud server, a corresponding access path needs to be configured for the task processing instance, allowing it to access the shared storage space and load the neural network model based on this access path. In practical applications, this access path can be set by adding an environment variable or specified in the target application configuration file.
[0067] In this embodiment of the invention, since the FaaS system is a function computing service system, the above-mentioned task processing instance can also be referred to as a function instance.
[0068] After creating a task processing instance corresponding to the task processing request through the above process, the task processing instance can load the neural network model corresponding to the target application code from the shared storage space in the first cloud server, and perform task processing based on the neural network model.
[0069] The task processing request carries data to be processed (i.e., data that needs to be processed by the neural network model), such as images, text, audio, and video. The task processing instance can parse the data from this data and perform some preprocessing based on the description in the target application code, such as format conversion and normalization, to meet the input requirements of the neural network model. Then, the neural network model loaded from shared storage space processes (i.e., performs inference) on the data to be processed, yielding the processing result.
[0070] In practical applications, task processing instances typically need to load the neural network model from shared storage into a suitable framework or library for model inference, such as TensorFlow or PyTorch. This instantiates the neural network model (including weights and metadata) into a model object that can be used to perform inference, and then uses this model object to process the data to be processed. The task processing instance then feeds back the processing results to the triggerer of the task processing request, which can be achieved through channels such as HTTP responses, message queues, and database updates.
[0071] During inference, the task processing instance may also access other services or databases to obtain additional data or store inference results. After obtaining the processing result from the neural network model, the task processing instance can also perform further processing on the result based on the description in the target application code, such as probability threshold filtering, result aggregation, format encapsulation, etc.
[0072] In an optional embodiment, after a task processing instance completes its task processing, if a set destruction condition is met, the computing resources corresponding to the task processing instance are released, and the task processing instance is destroyed to achieve resource recycling and reuse, thereby improving resource utilization. The destruction condition, for example, is that the task processing instance remains idle for a set duration. Therefore, in this embodiment of the invention, the state information of the task processing instance can be maintained, and its state can include busy, idle, destroyed, etc. Thus, when a task processing request for the target application code is received, if it is found that a task processing instance running the target application code is in an idle state, the task processing request can be assigned to the idle task processing instance without creating a new task processing instance.
[0073] In addition, when the task processing instance completes its task processing and is destroyed, the temporary data generated can also be cleared to free up storage space.
[0074] In summary, by creating a public shared storage space on the cloud server and pre-loading the neural network model into this space, it can be shared by various task processing instances on the cloud server that require the model. This allows task processing instances to load the neural network model from the local shared storage space on the cloud server. Compared to multiple task processing instances loading the model separately from remote storage, this significantly reduces network bandwidth and loading time, enabling more timely responses to task processing requests. Furthermore, since the model only needs to be loaded from remote storage once, multiple task processing instances can share it, improving network bandwidth utilization.
[0075] Figure 4 A flowchart of another neural network model task processing method provided in an embodiment of the present invention is shown below. Figure 4 As shown, the method may include the following steps:
[0076] 401. Obtain the target application code and target application configuration file input by the user. The target application configuration file includes the storage address of the neural network model called in the target application code in the remote storage space.
[0077] 402. Identify the first cloud server in the cloud server cluster used to deploy the application.
[0078] 403. Based on the storage address of the neural network model in the remote storage space, load the neural network model from the remote storage space into the shared storage space created in the first cloud server.
[0079] 404. Based on the user identification information contained in the target application configuration file, set the access permissions corresponding to the neural network model in the shared storage space.
[0080] 405. In response to a task processing request for the target application code, based on the user identification information in the task processing request and the above-mentioned access permission settings, determine that the task processing request has access permission to the above-mentioned neural network model in the shared storage space.
[0081] 406. Create a task processing instance corresponding to the task processing request in the first cloud server, so that the task processing instance loads the neural network model from the shared storage space and performs task processing based on the neural network model.
[0082] In this embodiment of the invention, in order to ensure the security of the neural network model data corresponding to different applications, data isolation settings are required, which can be achieved through permission configuration.
[0083] Specifically, access permissions for the neural network model in the shared storage space are set based on the user identification information contained in the target application's configuration file. A whitelist of access permissions for the neural network model in the shared storage space can be established, containing the user identification information. This means that each task processing request accessing the target application code corresponding to the user identification information has this access permission, and that is, the task processing instances corresponding to these task processing requests can all load the neural network model.
[0084] In practical applications, access permissions for the neural network model in the shared storage space can also be set by adjusting the Access Control List (ACL) of the first cloud server. A mapping between the neural network model and the user's identification information can be established within this ACL.
[0085] Based on this, upon receiving a task processing request targeting the application code, the system determines, based on the user identification information in the task processing request and the aforementioned access permission settings, that the task processing request has access to the aforementioned neural network model in the shared storage space. Then, a corresponding task processing instance is created, and an access path is configured for this instance to point to the neural network model in the shared storage space. Conversely, if the user identification information indicates that the task processing request does not have the aforementioned access permission, an error message is displayed.
[0086] By setting the access permissions as described above, we can ensure that task processing instances with permissions can access the corresponding neural network model, while ensuring that task processing instances without permissions cannot access the corresponding neural network model, thereby achieving secure isolation of different neural network models corresponding to different applications.
[0087] In practical applications, a neural network model may be continuously updated, for example, by continuously optimizing and training to update some of its weights. Therefore, in an optional embodiment, the model sharing method may further include: periodically querying the remote storage space based on a set timed task to determine the update information of the neural network model, and loading the updated neural network model into the shared storage space in the first cloud server according to the update information.
[0088] To ensure the neural network model in the shared storage space remains up-to-date, a periodic synchronization mechanism is needed. This can be automated using a scheduled task. The version number or update timestamp of the neural network model in the remote storage space can be queried periodically. If the version number or update timestamp differs from the version number or update timestamp of the neural network model already loaded into the shared storage space, it indicates that the neural network model has been updated. Updates can be implemented using methods such as full updates or incremental updates, downloading the updated neural network model to the shared storage space to replace or update the old model. Incremental updates synchronize only the model data (such as weights) that has changed since the last update, reducing data transfer volume and synchronization time.
[0089] In addition, the synchronous update process should ensure data consistency to prevent interference with ongoing task processing requests. Data consistency means ensuring that the new neural network model downloaded to the shared storage space is identical to the newly updated neural network model in the remote storage space, without any data loss, errors, or other anomalies occurring during transmission. Optionally, before loading, the MD5 (Message-Digest Algorithm 5) value of the new neural network model in the remote storage space can be calculated. After loading into the shared storage space, the MD5 value of the new neural network model in the shared storage space can be calculated again. The two MD5 values are compared; if they match, data consistency is satisfied; otherwise, an error message can be output, or the download can be restarted.
[0090] In this embodiment of the invention, whether the neural network model corresponding to the target application code is initially loaded into the shared storage space of the first cloud server, or the neural network model is periodically updated to the shared storage space, some testing and verification processes can be performed after each loading into the shared storage space. This ensures that the neural network model is correctly loaded into the shared storage space, and that the task processing instance can normally access and use the updated neural network model. It also verifies the improvement of some performance indicators after using the shared storage space. The initial loading can also be regarded as an update operation from scratch.
[0091] The neural network model is correctly loaded into the shared storage space, for example, by using the MD5 value mentioned above, to verify whether any anomalies such as data loss occur during the transmission of the neural network model from the remote storage space to the shared storage space.
[0092] Alternatively, it can be adopted Figure 5 The method shown is used for verification, and the method includes the following steps:
[0093] 501. Generate test tasks corresponding to the target application code.
[0094] 502. Create a test task processing instance corresponding to the test task in the first cloud server, so that the test task processing instance loads the neural network model corresponding to the target application code from the shared storage space, and processes the test task based on the neural network model.
[0095] 503. Determine the performance evaluation metrics corresponding to the test tasks.
[0096] For example, if the target application code is based on the neural network model to classify images input by the user, then the above test task would be a task that carries the images to be classified and the user identification information corresponding to the target application code, and would be generated automatically by the FaaS system.
[0097] Similar to the process of handling user-triggered task requests, a task processing instance corresponding to the test task is created in the first cloud server, called the test task processing instance. Then, the test task processing instance loads the neural network model corresponding to the target application code from the shared storage space and processes the test task based on this neural network model. The creation process of the test task processing instance is implemented as described in the aforementioned embodiments and will not be repeated here.
[0098] After completing the test task processing and obtaining the processing results, the corresponding execution evaluation metrics for the test task can be determined. Optionally, these execution evaluation metrics may include the loading time of the neural network model and / or the response time of the test task.
[0099] Specifically, log files generated in the FaaS system can be collected, and the loading time of the neural network model and the response time of the test task can be determined based on these log files. Loading time refers to the time consumed from the start of the loading action of the test task processing instance to the completion of loading the neural network model. The response time of the test task refers to the time consumed from the generation of the test task to the acquisition of the task processing result.
[0100] Based on the above performance evaluation metrics, they can be compared with the corresponding metrics of the method that does not use shared storage space but loads the neural network model from remote storage space to determine whether there is an improvement in latency performance.
[0101] For example, if the loading time and response time based on the shared storage space are both lower than the loading time and response time corresponding to the remote storage space, then the latency is determined to meet the requirements. The loading time and response time corresponding to the remote storage space refer to the loading time of the neural network model from the remote storage space to process the test task and the response time of the test task.
[0102] In an optional embodiment, a first access path can be configured for the test task processing instance to access the neural network model A corresponding to the target application code in the shared storage space, and a second access path can be configured to access another neural network model B in the shared storage space, to verify whether the data security meets the requirements. Specifically, based on the access permission settings of the neural network model A and neural network model B in the shared storage space as described in the foregoing embodiments, if it is determined that the test task processing instance can load the neural network model A from the shared storage space based on the first access path, and cannot load the other neural network model B from the shared storage space based on the second access path, then it is determined that the data security meets the requirements, wherein the user identification information corresponding to the test task satisfies the access permission of the neural network model A.
[0103] In other words, assuming a user ID (a1) has been set to have access to neural network model A, and a user ID (b1) has access to neural network model B, and assuming the user ID corresponding to the test task is a1, then normally the first access path of the test task processing instance should point to neural network model A. However, for data security verification, the second access path of the test task processing instance is configured to point to neural network model B. When the test task processing instance loads the neural network model from the shared storage space, permission verification is performed based on the set access permissions: if the first access path successfully loads neural network model A, but the second access path fails to load neural network model B, the access permission settings are effective, i.e., they meet the data security requirements. Conversely, if both the first and second access paths successfully load neural network model A and neural network model B, the access permission settings are ineffective, i.e., they do not meet the data security requirements.
[0104] To further reduce cold start latency—that is, the excessively long wait for a response after a user triggers a task processing request—this embodiment of the invention provides an instance preheating mechanism based on the aforementioned shared storage space. This mechanism preheats some task processing instances. Preheating means creating task processing instances in advance, before receiving a task processing request, so that they can be directly allocated and used once the request arrives. Combined with… Figure 6 The illustrated embodiment will be used for illustration.
[0105] Figure 6 A flowchart of an example preheating process provided for an embodiment of the present invention, such as... Figure 6 As shown, it may include the following steps:
[0106] 601. Determine the remaining available resources of the first cloud server.
[0107] 602. If the remaining available resources of the first cloud server are lower than the set threshold, the neural network model corresponding to the target application code is loaded from the shared storage space of the first cloud server to the shared storage space of the second cloud server, and a task processing instance corresponding to the target application code is created in the second cloud server. The second cloud server is a cloud server in the cloud server cluster whose remaining available resources are higher than the set threshold.
[0108] In this embodiment, when the remaining resources (such as CPU, memory, disk, etc.) in the first cloud server are below a set threshold due to the creation of a large number of task processing instances, all or part of the neural network models stored in the shared storage space of the first cloud server can be loaded into the shared storage space of the second cloud server with more remaining available resources. Furthermore, task processing instances corresponding to all or part of these transferred neural network models can be pre-established in the second cloud server. Pre-establishing task processing instances means creating task processing instances before receiving new task processing requests, and waiting for the arrival of task processing requests.
[0109] For example, the shared storage space of the first cloud server already contains neural network model A, neural network model B, and neural network model C. After loading neural network model A and neural network model B into the shared storage space of the second cloud server, a task processing instance corresponding to neural network model A can be created, or a task processing instance corresponding to neural network model A and a task processing instance corresponding to neural network model B can be created.
[0110] Loading neural network model A and neural network model B into the shared storage space of the second cloud server does not mean that they are deleted from the shared storage space of the first cloud server; they are still retained.
[0111] In practical applications, the number of neural network models to be loaded from the shared storage space of the first cloud server to the shared storage space of the second cloud server can be determined based on the remaining available resources of the first cloud server and the second cloud server. Alternatively, a set number of the above task processing instances can be created by default on the second cloud server.
[0112] However, alternatively, it can also be handled based on the following load mechanism:
[0113] The load information of at least one application code corresponding to multiple task processing instances already created in the first cloud server is determined, and the priority of the at least one application code is determined based on the load information. Based on the priority of the at least one application code, the neural network model corresponding to the target application code is loaded from the shared storage space in the first cloud server to the shared storage space in the second cloud server, and based on the load information of the target application code, a target number of task processing instances corresponding to the target application code are created in the second cloud server.
[0114] The shared storage space of the first cloud server can store multiple neural network models, each corresponding to different application code. Some of these application codes may be triggered by a large number of task processing requests (e.g., high concurrency), resulting in multiple busy task processing instances for a single application code on the first cloud server. Optionally, the number of task processing instances corresponding to an application code on the first cloud server can be used as its load information. Alternatively, based on the historical number and timing of task processing requests for an application code within a defined time period, a pre-defined predictive model can be used to predict the load of that application code within a future defined time period, and the prediction result can be used as the load information for that application code.
[0115] After obtaining the load information of each application code in the first cloud server, the priority of each application code is determined based on the load information: the higher the load, the higher the priority.
[0116] Therefore, based on the priority ranking results and the remaining available resources of the second cloud server, the neural network models corresponding to certain applications in the first cloud server are loaded from the shared storage space of the first cloud server to the shared storage space of the second cloud server. Assuming this includes the neural network model corresponding to the target application code, then, based on the load information of the target application code, the target number of task processing instances corresponding to the target application code can be created in the second cloud server. A pre-established correspondence between different load ranges and the number of pre-warmed task processing instances can be established to determine the target number.
[0117] In summary, the aforementioned preheating mechanism is actually a whole-machine-level preheating mechanism. It takes into account the overall available resources of the cloud server and the load of each application in the cloud server to carry out preheating processing across cloud servers, rather than preheating processing only for a certain application code. This is conducive to improving the overall resource utilization of the cloud server cluster and providing a better user experience for users of high-load application code.
[0118] The neural network model task processing apparatus of one or more embodiments of the present invention will be described in detail below. Those skilled in the art will understand that these apparatuses can be configured using commercially available hardware components through the steps taught in this solution.
[0119] Figure 7 This is a schematic diagram of the structure of a neural network model task processing device provided in an embodiment of the present invention, as shown below. Figure 7 As shown, the device includes: an acquisition module 11, a loading module 12, and a processing module 13.
[0120] The acquisition module 11 is used to acquire the target application code and target application configuration file input by the user. The target application configuration file includes the storage address of the neural network model called in the target application code in the remote storage space.
[0121] The loading module 12 is used to determine a first cloud server in the cloud server cluster used to deploy the application, and load the neural network model from the remote storage space into a shared storage space created in the first cloud server according to the storage address. The shared storage space can be accessed by all task processing instances running in the first cloud server.
[0122] The processing module 13 is configured to, in response to a task processing request for the target application code, create a task processing instance in the first cloud server corresponding to the task processing request, so that the task processing instance loads the neural network model from the shared storage space and performs task processing based on the neural network model.
[0123] Figure 7 The device shown can perform the steps in the foregoing embodiments. For detailed execution process and technical effects, please refer to the description in the foregoing embodiments, which will not be repeated here.
[0124] In one possible design, the above Figure 7 The structure of the device shown can be implemented as an electronic device. For example... Figure 8 As shown, the electronic device may include: a processor 21, a memory 22, and a communication interface 23. The memory 22 stores executable code, which, when executed by the processor 21, enables the processor 21 to at least implement the neural network model task processing method provided in the foregoing embodiments.
[0125] In addition, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, which, when executed by a processor of an electronic device, enables the processor to at least implement the neural network model task processing method provided in the foregoing embodiments.
[0126] The device embodiments described above are merely illustrative. The network elements described as separate components may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0127] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of a necessary general-purpose hardware platform, or by a combination of hardware and software. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing neural network model tasks, characterized in that, Function computing service systems applied to serverless computing systems include: Obtain the target application code and target application configuration file input by the user, wherein the target application configuration file includes the storage address of the neural network model called in the target application code in the remote storage space; Identify the first cloud server in the cloud server cluster used to deploy the application; Based on the storage address, the neural network model is loaded from the remote storage space into a shared storage space created in the first cloud server. The shared storage space can be accessed by all task processing instances running in the first cloud server. In response to a task processing request for the target application code, a task processing instance corresponding to the task processing request is created in the first cloud server, so that the task processing instance loads the neural network model from the shared storage space and performs task processing based on the neural network model.
2. The method according to claim 1, characterized in that, After loading the neural network model from the remote storage space into the shared storage space created in the first cloud server, the method further includes: Based on the user identification information contained in the target application configuration file, set the access permissions corresponding to the neural network model in the shared storage space; In response to a task processing request for the target application code, based on the user identification information in the task processing request and the access permission setting result, it is determined that the task processing request has access permission to the neural network model in the shared storage space.
3. The method according to claim 1, characterized in that, In response to a task processing request for the target application code, creating a task processing instance corresponding to the task processing request in the first cloud server, so that the task processing instance loads the neural network model from the shared storage space, includes: In response to a task processing request for the target application code, the target application configuration file and the target application code are determined based on the user identification information in the task processing request; Create a container and allocate computing resources in the first cloud server according to the target application configuration file; The target application code is deployed into the container to form a task processing instance corresponding to the task processing request; Configure an access path for the task processing instance to access the neural network model in the shared storage space, so that the task processing instance loads the neural network model from the shared storage space according to the access path.
4. The method according to claim 3, characterized in that, The method further includes: In response to the task processing instance completing the task processing, if the set destruction conditions are met, the computing resources corresponding to the task processing instance are released and the task processing instance is destroyed.
5. The method according to claim 1, characterized in that, The method further includes: Based on a pre-defined scheduled task, the remote storage space is periodically queried to determine the update information of the neural network model; Based on the updated information, the updated neural network model is loaded into the shared storage space.
6. The method according to claim 2, characterized in that, After loading the neural network model from the remote storage space into the shared storage space created in the first cloud server, the method further includes: Generate the test task corresponding to the target application code; A test task processing instance corresponding to the test task is created in the first cloud server, so that the test task processing instance loads the neural network model from the shared storage space and processes the test task based on the neural network model; Determine the execution evaluation metrics corresponding to the test task.
7. The method according to claim 6, characterized in that, The test task processing instances are respectively configured to access the first access path of the neural network model in the shared storage space and the second access path of another neural network model in the shared storage space; The determination of the execution evaluation metrics corresponding to the test task includes: Based on the access permission settings of the neural network model and the other neural network model, if it is determined that the test task processing instance can load the neural network model from the shared storage space based on the first access path, but cannot load the other neural network model from the shared storage space based on the second access path, then the data security is determined to meet the requirements, wherein the user identification information corresponding to the test task satisfies the access permission of the neural network model; Determine the loading time of the test task processing instance when it loads the neural network model from the shared storage space based on the first access path, and the response time of the test task; If both the loading time and the response time are lower than the loading time and response time corresponding to the remote storage space, then the latency is determined to meet the requirements. The loading time and response time corresponding to the remote storage space refer to the loading time of the neural network model when the test task processing instance loads the neural network model from the remote storage space to process the test task, and the response time of the test task.
8. The method according to any one of claims 1-7, characterized in that, The method further includes: Determine the remaining available resources of the first cloud server; If the remaining available resources are lower than a set threshold, the neural network model corresponding to the target application code is loaded from the shared storage space of the first cloud server to the shared storage space of the second cloud server, and a task processing instance corresponding to the target application code is created in the second cloud server, wherein the second cloud server is a cloud server in the cloud server cluster whose remaining available resources are higher than the set threshold.
9. The method according to claim 8, characterized in that, The step of loading the neural network model corresponding to the target application code from the shared storage space of the first cloud server to the shared storage space of the second cloud server, and creating a task processing instance corresponding to the application code in the second cloud server, includes: Determine the load information of at least one application code corresponding to the multiple task processing instances created in the first cloud server; The priority of the at least one application code is determined based on the load information; Based on the priority of the at least one application code, the neural network model corresponding to the target application code is loaded from the shared storage space in the first cloud server to the shared storage space in the second cloud server; Based on the load information of the target application code, a target number of task processing instances corresponding to the target application code are created in the second cloud server.
10. An electronic device, characterized in that, include: The system includes a memory, a processor, and a communication interface; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor performs the neural network model task processing method as described in any one of claims 1 to 9.
11. A non-transitory machine-readable storage medium, characterized in that, The non-transitory machine-readable storage medium stores executable code that, when executed by a processor of an electronic device, causes the processor to perform the neural network model task processing method as described in any one of claims 1 to 9.
12. A computer program product, characterized in that, include: A computer program, when executed by a processor of an electronic device, causes the processor to perform the neural network model task processing method as described in any one of claims 1 to 9.