Resource calculation method for LLM application development platform based on domestic CPU and OS
By performing appropriate hardware selection, platform deployment, distributed container deployment, load balancing configuration, knowledge base acquisition, monitoring and alarming, and resource recovery on the LLM application development platform of domestic CPU and OS, the problem of low resource utilization of large models on domestic hardware is solved, and efficient resource utilization and response performance improvement are achieved.
Patent Information
- Application Number
- CN202510926470.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-07
AI Technical Summary
When using the LLM application development platform of domestic CPU and OS, how to efficiently utilize large model computing resources, improve the utilization rate of large model resource allocation, and solve the problems of slow response and lag, especially on hardware devices with limited performance.
By selecting suitable domestic CPU and OS hardware configurations, deploying the LLM application development platform, deploying large model distributed containers, configuring load balancing, setting up knowledge base acquisition units, monitoring platforms and alarm units, and resource recovery units, we can optimize the resource utilization of large models.
It has achieved efficient use of large model computing resources on domestic hardware, improved the response performance and resource utilization of large models, reduced hardware costs, and enhanced the market competitiveness of large model service solutions.
Smart Images

Figure CN120429126B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large model analysis, and in particular to a resource calculation method for an LLM application development platform based on a domestically produced CPU and OS. Background Art
[0002] In recent years, domestic CPU technology has made significant progress. Through independent research and development and technological innovation, domestic manufacturers have mastered core technologies, freed themselves from dependence on foreign technology, and ensured national information security. They have also continuously launched new generations of products with continuously improved performance. Domestic systems have not only achieved breakthroughs in the desktop sector but have also been widely adopted in servers, embedded systems, and other fields. As domestic CPU and OS technology continues to mature, their costs are also gradually decreasing. Compared with similar foreign products, domestic CPUs and OSes are also more competitive in price, helping to reduce users' procurement costs.
[0003] As the application of large models becomes more widespread in the current era, the impact and challenges they bring are becoming increasingly prominent. As digital transformation deepens across various industries, enterprises and organizations are increasingly demanding intelligent and automated solutions. Large model knowledge Q&A, as a core technology for applications such as intelligent customer service and intelligent assistants, can significantly improve service efficiency and quality and reduce labor costs. However, as the industry researches the use of large models, users often face problems such as slow response and lag when using large models. This is especially true for hardware devices with limited performance. How to maximize hardware performance and reduce hardware costs has become an industry challenge. Improving the performance utilization of large models is a technical issue currently facing domestic CPUs. Summary of the Invention
[0004] The technical task of this invention is to provide a resource calculation method for an LLM application development platform based on domestic CPU and OS, which can efficiently utilize large model computing resources and improve the allocation and utilization of large model resources under the application of domestic CPU and OS.
[0005] The technical solution adopted by the present invention to solve its technical problem is:
[0006] The resource calculation method of the LLM application development platform based on domestic CPU and OS includes the following implementation:
[0007] Selection of domestic CPU and OS: Choose hardware configuration suitable for the specific model scale and user performance requirements, and be able to smoothly run the required scale model while controlling costs;
[0008] LLM application development platform deployment: The LLM application development platform includes a front-end operation platform, a request distribution platform, and a plug-in calling platform, which are used to conveniently and quickly integrate and call large models;
[0009] Large model distributed container deployment: Initialize multiple large model containers to improve system resource utilization while receiving a large number of concurrent requests;
[0010] Load balancing configuration: Configure multiple large model containers in Nginx to evenly distribute the load of each large model container and reduce the pressure on a single model. The load balancing configuration is used to repost and distribute input requests, receive a single user request, send the request to the large model to call the API, and return the result to the requester. Pay attention to the configuration and use of ports to avoid conflicts between multiple reposting ports.
[0011] The knowledge base acquisition unit customizes the knowledge base content to match the call of the large model, greatly improving the accuracy of the large model in business scenarios and avoiding invalid calculations that occupy system resources;
[0012] Platform monitoring and alarm unit, used to monitor machine performance and interface survival status to ensure smooth system operation while improving resource utilization of large models;
[0013] The large model resource recovery unit is used to timely recover the computing resources occupied by the large model to avoid the large model from continuing to calculate in the background, thus reserving resource space for subsequent large model calculations.
[0014] This method utilizes the LLM application development platform to distribute HTTP requests, including installation and deployment of the LLM application development platform, deployment of large models and plugins in distributed containers, and forwarding of load request flows. As an AI-powered large-scale model platform currently used in education, legal, medical, and archival services, the LLM application platform's use of large models is becoming a leading solution for industry-specific knowledge question-and-answer services. Efficient allocation of large-scale model computing resources fully utilizes machine performance, improving the computational efficiency and response rate of large models, significantly increasing the market competitiveness of large-scale model services.
[0015] Furthermore, regarding the choice of domestic CPU and OS, choose the combination of Hygon CPU with Kylin operating system and Ascend GPU. If the performance requirements for large models are average, then choose Hygon mid-range CPU, which can meet the general computing needs of 8B large model training and reasoning, while taking into account cost-effectiveness, and choose Ascend mid-range GPU with high computing power and low power consumption, which can significantly accelerate the training and reasoning process of 8B large models. If the performance requirements for large models are high, then choose Hygon high-end CPU, which has multi-core, high main frequency, and strong computing power, which can efficiently handle complex computing tasks in large model training. Choose Ascend high-performance GPU, which has ultra-high computing power and low power consumption, which can significantly accelerate the training and reasoning process of 32B large models. Therefore, flexible adjustment and optimization according to the specific model scale and performance requirements can give full play to the performance of the model and reduce hardware procurement costs.
[0016] Furthermore, the deployment of the LLM application development platform includes determining the use of platform tools, installing and deploying the LLM application development platform on the machine using containers, including three major container modules: the front-end page platform, the request forwarding platform, and the plug-in calling platform, and operating the platform to start and terminate conversation requests;
[0017] The platform embeds a large model for language analysis and information feedback. Depending on the type of large model in the system, corresponding large model plug-ins are installed so that they can be integrated into the application platform for use and call. This improves user response while quickly requesting and calling the large model, and streams the returned results while ensuring that the large model continues to occupy resources.
[0018] The LLM application development platform deployment is based on the Python development platform, the request delivery platform is based on the Go development platform, and the large model plug-in platform uses the Python development platform, which is more convenient for the development and use of the entire process.
[0019] Furthermore, the distributed container deployment of large models is a unique deployment method that controls multiple large models. When the performance of the board is limited, a single container deployment is used, or multiple devices are configured. It adapts to one or more application platforms to improve performance. Based on system performance control, the start and stop control and parameter control of large models are performed. The specific implementation is as follows:
[0020] Each large model is distributed on a single container, and the initialized large model is cached in a queue for use in controlling the large model. The number of models is opened according to the current model level and the maximum performance of the board, so that the video memory occupied by multiple initialized models does not exceed the upper limit of the board's video memory. In this way, the operation and calculation of the large model can be managed in a multi-model, multi-process, and batch manner based on the current hardware conditions. The large model is called through the distribution of requests, which greatly improves the response performance of the large model. At the same time, it also maximizes the use of domestic board resources and avoids the waste of hardware resource costs.
[0021] Furthermore, the load balancing configuration,
[0022] Repost the access requests to the API interface opened by the LLM application development platform, deploy nginx service for load balancing, modify the corresponding configuration file of nginx deployment for the multi-container distributed deployment of large models, add multiple repost requests in the nginx configuration file according to the access domain names provided by different containers, and reasonably distribute request resources through nginx load balancing. You can choose to round-robin request access interface, or choose to evenly distribute request interface according to the current request load to complete the round-robin call of the large model. The balanced load of the large model is achieved through reposting to avoid the situation where a single process is insufficiently resourced due to excessive pressure on a certain model.
[0023] Furthermore, the knowledge base acquisition unit supplements the basic data, enriches its own knowledge retrieval warehouse, implements valid data, adds valid tags to the data to be retrieved, and classifies the corresponding business; the specific implementation includes:
[0024] Acquisition of knowledge base. Knowledge base data can generally come from data sets provided by specific projects to provide basic knowledge retrieval, or use crawlers to retrieve the latest content from business-related official websites to obtain the latest authoritative information and update the knowledge base.
[0025] Classify the knowledge base according to the different contents or usage methods of the knowledge base, and classify the knowledge base for use in conjunction with the large model to achieve the highest accuracy retrieval; design an automatic knowledge base classification tool, combine it with the knowledge base acquisition tool to land the data according to the classification conditions, manually test the accuracy of the knowledge base after the data is entered into the database to maintain the update and iteration of the knowledge base, and dynamically bind the use of the large model and the knowledge base to improve the retrieval accuracy, meet user needs while avoiding repeated searches by users to waste machine performance resources.
[0026] Furthermore, the platform monitoring and alarm unit implements data monitoring and surveillance, intuitively displays the user interface in the form of charts, adds an alarm mechanism to directly notify developers of exposed problems, and dynamically analyzes factors affecting system operation in real time to ensure the smooth operation of the entire system. The specific implementation is as follows:
[0027] The platform monitoring and alerting unit coordinates risk management between machines and the platform. It uses the Grafana front-end monitoring platform and Prometheus to monitor system performance, generating machine performance charts for display on the front-end page. It also sets an alarm mechanism to trigger an alarm when a machine's load continuously exceeds a set threshold.
[0028] Email alerts are sent to the responsible person's mailbox. Based on the email content, precautions and adjustments are made to the machine performance to avoid risks in project operation and resolve them in a timely manner.
[0029] Use the Loki log management tool to generate logs on the Grafana interface to manage the logs of your own project's backend code. Categorize and manage logs according to log levels, add alarm prompts and log category retrieval to quickly resolve and troubleshoot project problems. Alarm monitoring can dynamically detect existing system resources in real time, ensuring maximum utilization of machine resources when large models are running.
[0030] Furthermore, the large model resource recovery unit takes out the large model that needs to terminate calculation for resource recovery according to the current number of large models in the queue, and tilts system resources to the calculation consumption of other large models in the queue when reloading the large model.
[0031] The large model resource recovery unit terminates the response and recycles the continuous calculation of the large model. According to the operating mechanism of the LLM application calling platform and the large model's own plug-in, the computing resources of the large model can recycle resources normally after the calculation is completed. However, the upper layer of the platform often terminates the current large model to continue responding. It is necessary to forcibly terminate the use and call of the large model based on the LLM application development platform, and then transfer the next request to the next initialized large model for calling. Relying on the multi-container deployment mechanism of the large model, when the user requests to terminate the response, the large model process is terminated, and the large model is reloaded and initialized. In this way, the recycling of the large model computing resources can be completed without affecting other large model processes and user requests.
[0032] The present invention also claims protection for an LLM application development platform resource computing device based on a domestically produced CPU and OS, comprising: at least one memory and at least one processor;
[0033] The at least one memory is configured to store a machine-readable program;
[0034] The at least one processor is configured to call the machine-readable program to implement the above method.
[0035] The present invention also claims protection for a computer-readable medium having computer instructions stored thereon, which, when executed by a processor, can implement the above method.
[0036] Compared with the prior art, the resource calculation method of the LLM application development platform based on domestic CPU and OS of the present invention has the following beneficial effects:
[0037] The selection of domestic CPUs and OS can maximize the cost of deploying large models of applicable scale; the distributed deployment of large models with multiple containers and the reprinting of requests targets the performance bottleneck of large model calls and is more suitable for scenarios with large concurrent users; the customization of the knowledge base and the deployment of the subsequent alarm platform, including resource recycling, improve the accuracy of the large model and enable the entire device and platform to have a more reasonable allocation and better control of resources. The present invention provides a text work knowledge question and answer artificial intelligence platform based on domestic hardware resources that is more suitable for a large number of requests. While ensuring the accuracy of the large model, it can fully call the resources of the machine, analyze and feedback the problem at a high quality and high speed, strictly control the process and system working status, ensure the smooth operation of the large model, and cooperate with the application platform of the large model to make the platform use faster and smarter. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1This is a flowchart of a resource calculation method for an LLM application development platform based on a domestic CPU and OS provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The present invention will be further described below with reference to specific embodiments.
[0040] The embodiment of the present invention provides a method for calculating resources of an LLM application development platform based on a domestically produced CPU and OS. The implementation of the method includes:
[0041] Selection of domestic CPU and OS: Choose hardware configuration suitable for the specific model scale and user performance requirements, and be able to smoothly run the required scale model while controlling costs;
[0042] LLM application development platform deployment: The LLM application development platform includes a front-end operation platform, a request distribution platform, and a plug-in calling platform, which are used to conveniently and quickly integrate and call large models;
[0043] Large model distributed container deployment: Initialize multiple large model containers to improve system resource utilization while receiving a large number of concurrent requests;
[0044] Load balancing configuration: Multiple large model containers are configured in Nginx to evenly distribute the load of each large model container, reducing the pressure on a single model.
[0045] The knowledge base acquisition unit customizes the knowledge base content to match the call of the large model, greatly improving the accuracy of the large model in business scenarios and avoiding invalid calculations that occupy system resources;
[0046] The platform monitoring and alarm unit monitors machine performance and interface survival status, ensuring smooth system operation while increasing the resource utilization limit of large models;
[0047] The large model resource recovery unit recovers the computing resources occupied by the large model in a timely manner, avoiding the continued calculation of the large model in the background and reserving resource space for subsequent large model calculations.
[0048] Among them, the selection of domestic CPU and OS should reduce costs, select hardware that is suitable for large model scale, avoid insufficient hardware cost-effectiveness or excessive hardware performance overflow, and maximize the use of computer resources at the lowest cost.
[0049] The LLM application development platform deployment is based on the Python development platform, the request delivery platform is based on the Go development platform, and the large model plug-in platform uses the Python development platform, which is more convenient for the development and use of the entire process.
[0050] Distributed container deployment of large models is a unique deployment method for controlling multiple large models. When the performance of the board is limited, a single container deployment may be required. Multiple devices can also be configured to adapt to multiple application platforms or one application platform to improve performance. The start and stop control and parameter control of large models need to be based on system performance control.
[0051] The load balancing configuration is used to repost and distribute input requests, obtain a single user request, send the request to the large model to call the API, and return the result to the requester. Pay attention to the configuration and use of ports to avoid conflicts between multiple reposting ports.
[0052] The knowledge base acquisition unit is used to supplement the basic data, enrich its own knowledge retrieval warehouse, implement effective data, add effective tags to the data to be retrieved, and classify the corresponding business.
[0053] The platform monitoring and alarm unit is used for data listening and monitoring. It displays the user interface intuitively in the form of charts, adds an alarm mechanism to directly notify developers of exposed problems, and analyzes factors affecting system operation in real time and dynamically to ensure the smooth operation of the entire system.
[0054] The large model resource recovery unit is used to recycle data computing resources. According to the current number of large model queues, the large model that needs to terminate the calculation is taken out for resource recovery. When the large model is reloaded, the system resources are tilted to the computing consumption of other large models in the queue.
[0055] LLMs (Large Language Models), with their powerful natural language processing capabilities, demonstrate immense potential in areas such as text generation, dialogue systems, and content creation, becoming a crucial tool for enterprise digital transformation. The rapid development of large model technology, particularly the maturity of pre-training and fine-tuning techniques, has enabled models to handle more complex natural language tasks, including knowledge question answering. Furthermore, increased computing power and reduced costs are making the deployment and application of large model knowledge question answering more economical and feasible. The Large Language Model (LLM) platform integrates high-performance computing, large-scale data, and engineering capabilities to provide an efficient and flexible AI development and application environment for both enterprises and individual users. Designing a large model call architecture and rationally allocating large model computing resources are crucial to meeting market expectations and reducing enterprise development costs.
[0056] The specific implementation of this method is as follows:
[0057] 1. Selection of domestic CPU and OS.
[0058] The selection of domestic CPUs and OSs is primarily based on performance and cost. Based on product development and user needs, and when running large models with concurrency requirements of 8 or 32 Bytes, consider performance, stability, cost-effectiveness, and market feedback. A Hygon CPU paired with the Kirin OS and Ascend GPUs is ideal for running large models with multiple concurrent workloads, ensuring compatibility with the Hygon CPUs and Ascend GPUs, providing excellent system performance and security. For those with moderate performance requirements for large models, a Hygon mid-range CPU can meet the general computing needs of training and inference for 8 Byte models while maintaining cost-effectiveness. A mid-range Ascend GPU offers high computing power and low power consumption, significantly accelerating the training and inference process for 8 Byte models. For those with higher performance requirements for large models, a Hygon high-end CPU with multiple cores, high clock speeds, and high computing power can efficiently handle the complex computational tasks of large model training. Choosing Ascend high-performance GPUs, with their ultra-high computing power and low power consumption, can significantly accelerate the training and inference of models as large as 32 Bytes. Therefore, flexible adjustments and optimizations based on the specific model size and performance requirements can fully maximize the model's performance and reduce hardware procurement costs.
[0059] 2. Deployment of LLM application development platform.
[0060] The deployment of the LLM application development platform includes determining the use of platform tools, using Docker to install and deploy the LLM application development platform on the machine, including three major container modules: the front-end page platform, the request forwarding platform, and the plug-in call platform. The platform operates to start and end conversation requests. Large models can be embedded in the platform for language analysis and information feedback. Corresponding large model plug-ins are installed according to the different types of large models in the system so that they can be integrated into the application platform for use and call, improving user response while quickly requesting to call the large model, and streaming the returned results while ensuring that the large model continues to occupy resources.
[0061] 3. Large model distributed container deployment.
[0062] Large-model distributed container deployment is used to distribute large models. There are many ways to deploy large models. Direct deployment to a physical machine is the most common and takes up the least resources, but it also has obvious disadvantages. When a large model runs on a single process, it may occupy resources between multiple threads when there are many concurrent threads, resulting in slow model response and resource grabbing. Therefore, large-model distributed container deployment can better solve this problem. The main feature is that each large model is distributed on a single container, and the initialized large model is cached in a queue for large-model control. The number of models is opened according to the current model level and the maximum performance of the board, so that the video memory occupied by multiple initialized models does not exceed the upper limit of the board's video memory. In this way, the operation and calculation of large models can be managed in a multi-model, multi-process, and batch manner based on the current hardware conditions. The large model is called by distributing requests, which greatly improves the response performance of the large model. At the same time, it also maximizes the use of domestic board resources and avoids the waste of hardware resource costs.
[0063] 4. Load balancing configuration.
[0064] The load balancing configuration is used to reprint the access requests to the API interface opened by the LLM application development platform. It is necessary to deploy the nginx service for load balancing. The deployment of nginx requires modifying the corresponding configuration file for the multi-container distributed deployment of large models. According to the access domain names provided by different containers, multiple reprint requests are added to the nginx configuration file to allow the nginx load balancing to reasonably allocate request resources. You can choose to round-robin the request access interface, or you can choose to evenly distribute the request interface according to the current request load to complete the round-robin call of the large model. The balanced load of the large model is achieved through reprinting to avoid the situation where a single process lacks resources due to excessive pressure on a certain model.
[0065] 5. Knowledge base acquisition unit.
[0066] The knowledge base acquisition unit refers to the knowledge base data that the LLM application development platform can rely on before performing large-scale model retrieval. This data can be customized according to user needs to improve the accuracy of the large-scale model. The first step is to acquire the knowledge base. The knowledge base data generally comes from the data set provided by a specific project, providing basic knowledge retrieval. You can also use crawlers to retrieve the latest content of business-related official websites to obtain the latest authoritative news and update the knowledge base; then the knowledge base is classified. According to the different content of the knowledge base or the distinction in usage, the knowledge base needs to be classified and used in conjunction with the large model for the most accurate retrieval. Design a knowledge base automatic classification tool, combine the knowledge base acquisition tool to implement data according to the classification conditions, manually test the knowledge base usage accuracy after the data is entered into the database to maintain the update and iteration of the knowledge base, and dynamically bind the use of large models and knowledge bases to improve retrieval accuracy, meet user needs, and avoid users from repeatedly searching to waste machine performance resources.
[0067] 6. Platform monitoring and alarm unit.
[0068] The platform monitoring and alarm unit refers to the risk control coordinated with the machine and platform. It uses the grafana front-end page monitoring platform to cooperate with prometheus to detect system performance, generates machine performance charts and displays them on the front-end page, and sets an alarm mechanism. When a certain load of the machine exceeds a certain threshold continuously, an alarm will be triggered and sent to the mailbox of the person in charge in the form of an email alarm. According to the content of the email, precautions and adjustments are made to the machine performance to avoid risks in project operation and resolve them in time. The loki log management tool is used to generate logs on the grafana interface, manage the logs of the background code of the project itself, classify and manage according to the log level, add alarm prompts and log classification retrieval to quickly solve and troubleshoot problems in the project. Alarm monitoring can dynamically detect existing system resources in real time to ensure the maximum utilization of machine resources when large models are running.
[0069] 7. Large model resource recovery unit.
[0070] The large model resource recovery unit pointer terminates the response and recycles the continuous calculation of the large model. According to the operating mechanism of the LLM application calling platform and the large model's own plug-in, the computing resources of the large model can recycle resources normally after the calculation is completed. However, the upper layer of the platform often terminates the current large model to continue responding. It is necessary to forcibly terminate the use and call of the large model based on the LLM application development platform, and then transfer the next request to the next initialized large model for calling. Relying on the multi-container deployment mechanism of the large model, when the user requests to terminate the response, the large model process is terminated, and the large model is reloaded and initialized. In this way, the recycling of the large model computing resources can be completed without affecting other large model processes and user requests.
[0071] This method is applicable and optimized for general large-model analysis application systems. The processes of general large-model analysis application systems are generally simple. However, this method proposes a method for efficiently utilizing large-model computing resources based on the LLM application development platform of domestic CPUs and OS. Its specific implementation includes multiple steps including the selection of domestic CPUs and OSs, LLM application development platform deployment, large-model distributed container deployment, load balancing configuration, knowledge base acquisition unit, platform monitoring and alarm unit, and large-model resource recovery unit. Specific application cases are as follows:
[0072] like Figure 1 The following are the implementation steps of this method. According to the process steps, when the entire project system is built and the machine, project platform and large model are deployed, the entire process is as follows: the user enters the question in the dialog box, internally requests the interface provided by the LLM application development platform, and the interface calls the corresponding large model to start the large language analysis work. After a series of processing flows, the analysis of the large model begins. The large model obtains the knowledge base as its own backup retrieval library to search and intelligently organize the language to return the answer. After processing and screening, the returned answer is displayed on the user platform for streaming output of the answer. After the user terminates the output, the internal platform automatically calls the termination of the large model to reclaim the occupied resources. In this way, the user can access the large model multiple times and quickly and get information feedback.
[0073] The first step is hardware selection. Taking the deployment of a 32B large model as an example, the CPU is the Hygon 7385, which is configured with more than 32 cores to provide sufficient parallel computing capabilities to meet the computing power requirements of the 32B large model. The main frequency is not less than 2.5GHz to ensure high performance when processing large-scale data. The GPU is the Ascend 910, with a video memory capacity of not less than 64GB to ensure that it can load and process the massive parameters and data of the 32B large model. In conjunction with the Galaxy Kirin V10 SP3 system, officially certified drivers and toolkits are installed to optimize hardware performance and ensure stable system operation.
[0074] Next is the deployment of the application. The LLM application platform is installed and deployed on the machine through Docker. The platform integrates and embeds the large model for language analysis and information feedback, installs the large model call plug-in, classifies the large model business tags, writes the large model business process, installs the large model usage tools, and installs the model supplier tools.
[0075] Then comes the deployment of large models. Prioritize deploying efficient and reliable large models on containers, specify the container network segment and domain name, and specify that the large model calculations will be run on domestic boards to improve the resource utilization and computing speed of the large model. Initialize the large model and start a single process to occupy the video memory of the board. Establish the number of containers based on the current model level and the maximum performance of the board. Each container initializes a large model process for use.
[0076] It is also necessary to deploy nginx for load balancing. Install and deploy nginx through docker, modify the nginx configuration file, add the large model container network segment and domain name as the request load, add automatic request allocation configuration based on the current minimum number of requests, and modify the LLM application platform to call the large model's http request to the current nginx request port.
[0077] Finally, deploy a monitoring platform to monitor system resources, deploy the Grafana graphical page monitoring platform, deploy Prometheus and node_exporter to monitor system performance, install Influxdb for monitoring data persistence, import the dashboard configuration in Grafana and add the Prometheus data source, add alarm configuration in Grafana, configure alarm rules and establish a new contact point for alarms.
[0078] After the application and tools are deployed, the design of the software solution generally starts with the project architecture, using a flexible, lightweight and quickly integrated web platform for technical adaptation to facilitate the integration of various third-party large models. The project is divided according to scenarios, modules and functions. Different scenarios use different branch modules, different functions are divided into different request categories, and sub-branches of different scenario modules are established under the same function. Finally, the large model call interface document is designed to provide interface R&D instructions for the project in the current scenario.
[0079] When designing a knowledge base architecture, the first thing is to design a knowledge base data acquisition platform. The main functions are to collect knowledge base data corresponding to the business according to business classification. One is to create a knowledge base upload function for manual review and offline acquisition. The other is to implement data search through crawler collection and other methods, and further data screening and review of the searched data. The second is to persist reliable data in the database and label the knowledge base according to business types. Dynamic binding of large models and the use of knowledge bases can improve retrieval accuracy.
[0080] Design a large model resource recovery plan, modify the LLM application platform stop response interface, obtain the unique identifier of the current request distribution, restart the corresponding large model container based on the unique identifier, and at the same time directly stop the continued calculation of the large model and reinitialize the large model. Subsequent requests are transferred to the initialized large model container for calling.
[0081] The embodiment of the present invention also provides a resource calculation device for an LLM application development platform based on a domestically produced CPU and OS, comprising: at least one memory and at least one processor;
[0082] The at least one memory is configured to store a machine-readable program;
[0083] The at least one processor is used to call the machine-readable program to implement the resource calculation method of the LLM application development platform based on the domestic CPU and OS described in the above embodiment.
[0084] An embodiment of the present invention further provides a computer-readable medium storing computer instructions. When executed by a processor, these instructions cause the processor to execute the resource calculation method for an LLM application development platform based on a domestically produced CPU and OS, as described in the aforementioned embodiments. Specifically, a system or device equipped with a storage medium storing software program code implementing the functions of any of the aforementioned embodiments can be provided, and a computer (or CPU or MPU) of the system or device can be configured to read and execute the program code stored in the storage medium.
[0085] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute part of the present invention.
[0086] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (e.g., CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RAMs, DVD-RWs, and DVD+RWs), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer via a communications network.
[0087] In addition, it should be clear that the functions of any of the above embodiments can be achieved not only by executing the program code read by the computer, but also by enabling the operating system operating on the computer to complete part or all of the actual operations based on the instructions of the program code.
[0088] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU installed on the expansion board or expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above embodiments.
[0089] The present invention has been shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the scope of protection of the present invention.
Claims
1. The resource calculation method of LLM application development platform based on domestic CPU and OS is characterized by: The implementation of this method includes: Selection of domestic CPU and OS: Choose hardware configuration suitable for the specific model scale and user performance requirements, and be able to smoothly run the required scale model while controlling costs; LLM application development platform deployment: The LLM application development platform includes a front-end operation platform, a request distribution platform, and a plug-in calling platform, which are used to quickly integrate and call large models; Large model distributed container deployment: Initialize multiple large model containers to receive concurrent requests; Load balancing configuration: Multiple large model containers are configured in Nginx to evenly distribute the load of each large model container. The load balancing configuration is used to reprint and distribute input requests, obtain a single user request, send the request to the large model call API, and return the result to the requester. The knowledge base acquisition unit customizes the knowledge base content to match the call of the large model to avoid invalid calculations that occupy system resources; Platform monitoring and alarm unit, used to monitor machine performance and interface survival status to ensure smooth system operation while improving resource utilization of large models; Large model resource recovery unit, used to promptly recover large model computing resources to avoid continued calculations in the background of the large model; The LLM application development platform deployment includes determining the use of platform tools, installing and deploying the LLM application development platform on a machine using containers, including three major container modules: the front-end page platform, the request forwarding platform, and the plug-in call platform, and operating the platform to initiate and terminate conversation requests. The platform embeds a large model for language analysis and information feedback, installs corresponding large model plug-ins based on different types of large models in the system, and streams the returned results while ensuring that the large model continues to occupy resources. The distributed container deployment of large models is a unique deployment method that controls multiple large models. When the performance of the board is limited, a single container deployment is used, or multiple devices are configured. It adapts to one or more application platforms to improve performance. Based on system performance control, the start and stop control and parameter control of large models are performed. The specific implementation is as follows: Each large model is distributed on a single container. The initialized large model is cached in a queue for large model control. The number of models is opened based on the current model level and the maximum performance of the board, so that the video memory occupied by multiple initialized models does not exceed the upper limit of the board's video memory. Based on the current hardware conditions, the operation and calculation of the large model are managed in a multi-model, multi-process, and batch manner, and the large model is called by request distribution. The load balancing configuration, Repost access requests to the API interface opened by the LLM application development platform, deploy nginx service for load balancing, modify the corresponding configuration file of nginx deployment for the multi-container distributed deployment of large models, add multiple repost requests to the nginx configuration file based on the access domain names provided by different containers, and reasonably allocate request resources through nginx load balancing. You can choose to round-robin request access interfaces or evenly distribute request interfaces based on the current request load to complete round-robin calls of large models; The knowledge base acquisition unit supplements the basic data, enriches the knowledge retrieval warehouse, implements valid data, adds valid tags to the data to be retrieved, and classifies the corresponding business. The specific implementation includes: Acquisition of knowledge base. Knowledge base data can come from data sets provided by specific projects to provide basic knowledge retrieval, or use crawlers to retrieve the latest content from business-related official websites to obtain the latest authoritative information and update the knowledge base. Classify the knowledge base according to the different contents or usage methods of the knowledge base, and classify the knowledge base for use in conjunction with the large model to achieve the highest accuracy retrieval; design an automatic knowledge base classification tool, combine it with the knowledge base acquisition tool to land the data according to the classification conditions, manually test the accuracy of the knowledge base after the data is entered into the database to maintain the update and iteration of the knowledge base, and dynamically bind the use of the large model and the knowledge base to improve the retrieval accuracy, meet user needs while avoiding repeated searches by users to waste machine performance resources.
2. The resource calculation method for the LLM application development platform based on domestic CPU and OS according to claim 1 is characterized in that: The choice of domestic CPU and OS is to choose Hygon CPU with Kirin operating system and Ascend GPU.
3. The resource calculation method for the LLM application development platform based on domestic CPU and OS according to claim 1 is characterized in that: The LLM application development platform deployment is based on the python development platform, the request delivery platform is based on the go development platform, and the large model plug-in platform uses the python development platform.
4. The resource calculation method for the LLM application development platform based on domestic CPU and OS according to claim 1 is characterized in that: The platform monitoring and alarm unit implements data monitoring and display, displays the user interface in the form of charts and graphs, adds an alarm mechanism to directly notify developers of exposed problems, and dynamically analyzes factors affecting system operation in real time to ensure the smooth operation of the entire system. The specific implementation is as follows: The platform monitoring and alerting unit coordinates risk management between machines and the platform. It uses the Grafana front-end monitoring platform and Prometheus to monitor system performance, generating machine performance charts for display on the front-end page. It also sets an alarm mechanism to trigger an alarm when a machine's load continuously exceeds a set threshold. Email alerts are sent to the responsible person's mailbox. Based on the email content, precautions and adjustments are made to the machine performance to avoid risks in project operation and resolve them in a timely manner. Use the Loki log management tool to generate logs on the Grafana interface to manage the logs of your own project's background code, classify and manage them according to the log level, add alarm prompts and log classification retrieval to quickly solve and troubleshoot project problems.
5. The resource calculation method for the LLM application development platform based on domestic CPU and OS according to claim 1 is characterized in that: The large model resource recovery unit takes out the large model that needs to terminate calculation for resource recovery according to the current number of large model queues, and tilts system resources to the calculation consumption of other large models in the queue when reloading the large model.
6. LLM application development platform resource computing device based on domestic CPU and OS, characterized by: include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to implement the method according to any one of claims 1 to 5.
7. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, can implement the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Large model cloud computing dynamic updating method based on elastic adaptive incremental learning
CN119323265A
Efficient deployment method and system for privatized large language model ecological service
CN119440849A