Local large language model operation framework transplantation and adaptation method based on domestic DCU environment

By performing hardware environment detection, deep learning framework adaptation and model quantitative compression on the domestic DCU platform, the transplantation and optimization problems of large language models on the heterogeneous computing platform are solved, efficient domestic hardware adaptation is achieved, and computing efficiency and compatibility are improved.

CN120234040APending Publication Date: 2025-07-01NORTHEASTERN UNIV CHINA
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510304829.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

There are insufficient porting and optimization of existing large language model operating frameworks on heterogeneous computing platforms, especially insufficient support for domestic GPU architectures, resulting in low computing efficiency and poor hardware compatibility.

Method used

Using model quantization and compression technology, the mainstream large language model operation framework is adapted to the domestic DCU platform. Through hardware environment detection, deep learning framework adaptation, model quantization and compression, multi-model management and parallel adaptation, local operation framework porting and adaptation in the domestic DCU environment can be realized.

Benefits of technology

It greatly increases the hardware selection range of mainstream large language model operation frameworks, reduces the resource consumption of local deployment, and improves the competitiveness of domestic hardware in large language model applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234040A_ABST
    Figure CN120234040A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence and localization basic platforms, and discloses a local large language model operation framework transplantation and adaptation method based on a domestic DCU environment. According to the method provided by the invention, DCU-based local operation framework adaptation is realized, and the hardware selection range of a mainstream large language model operation framework is greatly expanded; through the model quantification and compression technology, the hardware of the local deployment large language model can be quickly applied to the field of natural language processing, and the competitiveness of the domestic hardware in the large language model application is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and domestic basic platform, and particularly to a method for transplanting and adapting the operation framework of a local large language model based on a domestic DCU environment. Background Art

[0002] Large Language Models (LLMs) are natural language processing models based on deep learning. Through training with a large amount of text data, they can understand and generate language texts. With the progress of technology, people are becoming increasingly interested in LLMs. Researchers, developers, and enterprises are all trying to use LLMs to assist in work and learning and develop new applications for them. Large language models can handle various natural language tasks, such as text generation, dialogue systems, information retrieval and question answering, and translation and summarization, etc.; they can be widely applied to industries such as customer service, e-commerce, finance, education, medical care, law, tourism and hotels. However, the operation of large language models usually requires a large amount of computing resources. Especially when deployed in a local environment, problems such as low computing efficiency and poor hardware compatibility are faced.

[0003] Although existing operation frameworks (such as Ollama) provide lightweight local deployment solutions, there are still deficiencies in transplantation and optimization on heterogeneous computing platforms (such as CPUs, GPUs, etc.), and there is insufficient support for emerging domestic GPU architectures. Therefore, there is an urgent need for a technical solution that can efficiently transplant and adapt the operation framework of large language models on heterogeneous computing platforms to improve the model operation efficiency and reduce resource consumption. Summary of the Invention

[0004] The purpose of the present invention is to use model quantization and compression to adapt the mainstream large language model operation framework to the domestic DCU platform, and provide a method for transplanting and adapting the operation framework of a local large language model based on a domestic DCU environment, so as to obtain a local model operation framework applied to the domestic DCU environment. The present invention can fully evaluate the adaptation performance and reliability of the domestic hardware platform to actual business requirements, and ensure that the domestic hardware platform can meet project requirements. Combining potential development and optimization requirements, comprehensively examine whether the various capabilities of the hardware device can form good support.

[0005] The technical solution of the present invention is as follows: A method for transplanting and adapting the operation framework of a local large language model based on a domestic DCU environment, including the following steps:

[0006] S1: Detection of the hardware device environment and adaptation of the deep learning framework;

[0007] S2: Adaptation of the large language model operation framework;

[0008] S3: Quantization and compression of the large language model;

[0009] S4: Parallelism adaptation of multi-model management and operation framework.

[0010] The specific steps of step S1 include:

[0011] S1.1: Install DCU hardware firmware and driver programs;

[0012] S1.1.a: Install DCU hardware firmware;

[0013] S1.1.b: Detect and install driver program dependency component libraries;

[0014] S1.1.c: Obtain the driver program installation package and compile the driver program;

[0015] S1.1.d: Repeat S1.1.a - S1.1.c until the DCU hardware firmware and driver program effectiveness tests are qualified;

[0016] S1.2: Deep learning framework adaptation;

[0017] S1.2.a: Install and configure the environment manager;

[0018] S1.2.b: Obtain and install the deep learning framework package;

[0019] S1.2.c: Repeat S1.2.a - S1.2.b until the deep learning environment framework effectiveness verification is qualified.

[0020] The driver program dependency component libraries include compilers, cross-compilation tools, and dependent dynamic link libraries.

[0021] The specific steps of step S2 include:

[0022] S2.1: Obtain the source code of the large language operation framework;

[0023] S2.2: Detect and install the large language operation framework dependency libraries;

[0024] S2.3: Modify and optimize the compilation configuration and hardware detection logic;

[0025] S2.4: Repeat S2.1 - S2.3 until the large language operation framework adaptation effectiveness verification is qualified.

[0026] The modification and optimization of the compilation configuration and hardware detection logic are specifically:

[0027] S2.3.1: Change the software build build file path detection code: Add a new environment variable to directly point to the software build build file path;

[0028] S2.3.2. Modify the GPU environment detection code: Change DriverVersionFile to the DCU path;

[0029] S2.3.3. Modify the DCU parallel running code: Modify the DCU parallel running code to an explicit 1024-core function;

[0030] S2.3.4. Modify the compilation file: Delete the git_module_setup module to disable the llama.cpp version update, and change HIP_PATH, ROCM_PATH, and LIBRARY_PATH to the default configuration to adapt to DCU compilation;

[0031] S2.3.5 Hardware pre-detection and isolation optimization: First, identify the software parameter OLLAMA_DCU_ONLY_UUID. If the value of the OLLAMA_DCU_ONLY_UUID parameter is OFF, then use the default hardware pre-detection; if the value of the OLLAMA_DCU_ONLY_UUID parameter is ON, then use the hardware detection under software control;

[0032] After the hardware pre-detection is completed, output the available hardware information. If there is no available GPU node, then set the available hardware to the CPU and use the deep learning framework to call the system hardware.

[0033] The process of the default hardware pre-detection is as follows: First, identify the HIP_VISIBLE_DEVICES system parameter to record the allocated hardware serial number. Secondly, sequentially scan the hardware nodes under the / sys / class / kfd / kfd / topology / nodes / file to obtain the system hardware information through the driver file, including the system architecture and the hardware serial number; According to the obtained hardware information, judge whether the hardware architecture of the current node is GFX000 or a hardware architecture not supported by the software; If the current node is valid, then judge whether the allocated hardware exists according to the hardware serial number. If it does not exist, skip the current node and perform the next hardware node detection.

[0034] The process of the hardware detection under software control is as follows: First, record the hardware UUID applied by the user through the ROCR_VISIBLE_DEVICES system parameter. Secondly, scan the hardware nodes under the / sys / class / kfd / kfd / topology / nodes / file to obtain the system hardware architecture and the hardware independent id. Judge whether the current node is an abnormal architecture according to the hardware architecture. If the current node is a valid node, then judge whether the hardware UUID applied by the user exists according to the hardware independent id. If abnormal information is detected, skip the current node detection; If there is no abnormality in the hardware detection of the current node, then add the current node to the available GPU node queue and perform the next hardware node detection.

[0035] The specific steps of step S3 include:

[0036] S3.1: Model selection and format conversion;

[0037] S3.2: Model quantization and compression.

[0038] The specific steps of step S4 include:

[0039] S4.1: Import of multiple models and model switching: Import multiple different models into the framework and run the corresponding models individually in batches, monitoring the model resource usage;

[0040] S4.2: Adaptation of the parallelism of the running framework: Run multiple models simultaneously, monitoring the resource usage of each model to verify the parallel usage of the running framework.

[0041] Advantages of the present invention: The method provided by the present invention realizes the adaptation of the local running framework based on DCU, greatly increasing the hardware selection range of the running frameworks of mainstream large language models; Through model quantization and compression technologies, the hardware requirements for local deployment of large language models are reduced, enabling it to be quickly applied to the field of natural language processing and enhancing the competitiveness of domestic hardware in the application of large language models. Description of the Drawings

[0042] Figure 1 It is the schematic diagram of the overall architecture of the method;

[0043] Figure 2 It is the flowchart of the adaptation of the local large language model running framework;

[0044] Figure 3 It is the flowchart of the import adaptation of the local large language model;

[0045] Figure 4 It is the specific flowchart of the modification and optimization of the compilation configuration and hardware detection logic. Detailed Implementation Manner

[0046] Figure 1 It is the schematic diagram of a method for transplanting and adapting the local large language model running framework based on the domestic DCU environment of the present invention, and the specific steps include:

[0047] S1: Detection of the hardware device environment and adaptation of the deep learning framework;

[0048] S2: Adaptation of the large language model running framework;

[0049] S3: Quantization and compression of the large language model;

[0050] S4: Multi-model management and adaptation of the parallelism of the running framework.

[0051] In the example, the test environment is a domestic DCU intelligent acceleration card, model: K100L-DCU, the deep learning framework is the DAS (DCU AI Software Stack) framework self-developed by Dawning, and the mainstream operating framework it adapts to is OLLAMA.

[0052] Step S1: Detection of the hardware device environment and adaptation of the deep learning framework:

[0053] S1.1: Install the DCU hardware firmware and driver program. The specific method includes the following steps:

[0054] S1.1.a: Install the DCU hardware firmware;

[0055] S1.1.b: Detect and install the driver program dependency component library. The specific steps are as follows:

[0056] Step 1: Install the compiler: In the example, install the compilation tools GCC and G++ (GNU Compiler Collection) that support C and C++ languages. Step 2: Install the cross-compilation tool: In the example, install the cross-compilation tool Cmake that supports multiple target architectures. Step 3: Install dependencies such as dynamic link libraries: In the example, install libelf-dev libdrm-amdgpu1 libtinfo5 pciutils libdrm-dev, etc.

[0057] S1.1.c: Obtain the driver program installation package and compile the driver program: Download the adapted driver version through the DCU developer information platform, and install it after changing the driver permissions.

[0058] S1.1.d: Repeat S1.1.a - S1.1.c until the effectiveness detection of the DCU hardware firmware and driver program is qualified; after the installation is completed, use the terminal command (lsmod|grep hydcu) to confirm the effective installation of the firmware and driver. If the firmware or driver version is too low during the installation process, download and install the high-version driver program; if there are problems such as "card dropout" during the driver installation, reinstall the driver program.

[0059] S1.2: Adapt the deep learning framework. The specific steps are as follows:

[0060] S1.2.a: Install and configure the environment manager;

[0061] S1.2.b: Obtain and install the deep learning framework package. The specific steps are as follows:

[0062] Step 1: Download the PyTorch deep learning environment framework package supported by DSA through the DCU developer information platform; Step 2: Use Anaconda to install the PyTorch deep learning environment framework package and the remaining components for model quantization and compression; Step 3: Use the terminal command (conda list) to confirm the successful installation of the deep learning framework.

[0063] S1.2.c: Repeat S1.2.a - S1.2.b until the deep learning environment framework is successfully verified: The effectiveness of the installation can be confirmed through the Demo program provided by DSA or a program written by the user himself. In the example, it is verified by running the model conversion program.

[0064] The specific steps of the large language model running framework adaptation in step S2 include:

[0065] S2.1: Obtain the source code of the large language running framework: Find the source code of the running framework on the official website of the project or in the GitHub repository.

[0066] S2.2: Detect and install the dependent libraries of the running framework: Install the corresponding dependencies according to the project documentation requirements. In the example, GCC, GO, and CMAKE are installed to meet the OLLAMA requirements.

[0067] S2.3: Compile the configuration and modify and optimize the hardware detection logic. The specific steps are as follows:

[0068] Step 1: Modify the system path detection code: Since OLLAMA uses PWD to read the current system path for CPU and GPU build file detection, for the software startup method using the IDE, the original detection code of OLLAMA will become invalid. Therefore, by adding a new environment variable, the value of the environment variable is the build file path of ollama. At this time, the CPU and GPU build files can be recognized by directly reading the environment variable.

[0069] Step 2: Modify the GPU environment detection code: Since the overall system architecture of DCU and AMD is different, OLLAMA will not be able to detect the corresponding DCU driver file information. Therefore, modify the configuration DriverVersionFile = " / sys / module / hydcu / version" and ROCmLibGlobs = []string{"libhipblas.so.0.*","rocblas"} RocmStandardLocations = []string{" / opt / dtk / lib"," / usr / lib64"} for adaptation.

[0070] Step 3: Modify the DCU parallel running code: OLLAMA writes deep learning running code using CUDA, and its default code conflicts with the DCU system architecture. Therefore, modify the parallel running code to an explicit 1024-core function for DCU and running framework adaptation.

[0071] Step 4: Modify the compilation file: OLLAMA uses llama.cpp as the backend for running. llama.cpp has changed its software structure multiple times, resulting in runtime errors in some versions of OLLAMA. Therefore, delete the git_module_setup module to disable llama.cpp version updates and instead use a fixed version for adaptation; when compiling with OLLAMA, it will default to the default configuration of AMD, resulting in the inability to find some required dependencies. Adapt to DCU by changing default configurations such as HIP_PATH= / opt / dtk-24.04.2 / hip, ROCM_PATH= / opt / dtk-24.04.2, and LIBRARY_PATH= / opt / dtk-24.04.2 / llvm / lib / clang / 15.0.0 / lib / linux / :$LIBRARY_PATH.

[0072] Step 5: Optimize the hardware detection logic: The flowchart for hardware pre-detection and isolation optimization is as Figure 4As shown, for hardware pre-detection and isolation optimization, first, by identifying the software parameter OLLAMA_DCU_ONLY_UUID, if the parameter value is OFF, the default hardware pre-detection is adopted. The default hardware pre-detection process is as follows: First, identify the HIP_VISIBLE_DEVICES system parameter to record the allocated hardware serial numbers. Secondly, sequentially scan the hardware nodes under the / sys / class / kfd / kfd / topology / nodes / file to obtain information such as the system hardware architecture and hardware serial numbers through the driver files, and determine whether the hardware architecture of the current node is GFX000 or a hardware architecture not supported by the software based on the obtained hardware information. If the current node is valid, determine whether the allocated hardware exists based on the hardware serial number. If it does not exist, skip the current node and proceed to the next hardware node detection; if the parameter value is ON, the hardware detection under software control is adopted. The process is as follows: First, record the hardware UUID applied by the user through the ROCR_VISIBLE_DEVICES system parameter. Secondly, scan the hardware nodes under the / sys / class / kfd / kfd / topology / nodes / file to obtain information such as the system hardware architecture and hardware unique id, and determine whether the current node is an abnormal architecture based on the hardware architecture. If the current node is a valid node, determine whether the hardware UUID applied by the user exists based on the hardware unique id. If abnormal information such as using a hardware serial number or an incorrect UUID is detected, skip the current node detection; if there is no abnormality in the hardware detection of the current node, add the current node to the available GPU node queue and proceed to the next hardware node detection. After the hardware detection is completed, output the available hardware information. If there are no available GPU nodes, set the available hardware to CPU and call the system hardware using the deep learning framework.

[0073] In Example 1, 2 domestic intelligent acceleration cards of DCU with the architecture of GFX916 and 2 CPUs are used for verification. First, default hardware pre-detection is performed. If the OLLAMA_DCU_ONLY_UUID parameter is not set, the default hardware pre-detection is used at this time. Set the system parameter HIP_VISIBLE_DEVICES=0,1. At this time, the string is read through visibleDevices and the available hardware serial numbers are recorded. Secondly, by scanning 4 hardware information files under / sys / class / kfd / kfd / topology / nodes / , the hardware architecture information is read and recorded into the major, minor, and path variables. If it is determined that the current hardware is a CPU, it is directly skipped. If it is a GPU, it is detected whether it is GFX000. If so, it means that the hardware information file is damaged, and the current hardware node is skipped. After the hardware architecture detection is completed, it is judged whether the current hardware serial number exists in visibleDevices. If a match is detected, the detection of the next hardware node is performed. After the hardware detection is completed, the hardware information is output:

[0074] time=2025-03-10T09:11:47.293Z level=INFO source=types.go:107 msg="inference compute" library=rocm variant="" compute=gfx916 driver=6.2 name=1d94:55b7 total="32.0GiB" available="32.0GiB"

[0075] time=2025-03-10T09:11:47.293Z level=INFO source=types.go:107 msg="inference compute" library=rocm variant="" compute=gfx916 driver=6.2 name=1d94:55b7 total="32.0GiB" available="32.0GiB"

[0076] At this time, if there are available GPU nodes, the deep learning framework is used to call the GPU. If the above system parameter HIP_VISIBLE_DEVICES=-1 is set, no available GPU hardware is detected at this time, the system hardware is set to CPU, and the deep learning framework is used to call the CPU. The output hardware information is as follows:

[0077] time = 2025-03-10T09:16:38.976Z level = INFO source = types.go:107 msg = "inference compute" id = 0 library = cpu variant = avx2 compute = "" driver = 0.0 name = "" total = "110.0GiB" available = "106.8GiB"

[0078] In Example 2, 2 domestic intelligent acceleration cards of GFX928 and 2 CPUs are applied through the computing platform for verification. Set OLLAMA_DCU_ONLY_UUID = ON. At this time, hardware detection for software permission control is performed. Set ROCR_VISIBLE_DEVICES = GPU-6f0cb05fea483a21,GPU-6f09642fc1ca30e1. At this time, the hardware UUID applied by the user is recorded through visibleDevices. Secondly, by scanning 4 hardware information files under / sys / class / kfd / kfd / topology / nodes / , the hardware architecture information is read and recorded into the major, minor, and path variables. If it is determined that the current hardware is a CPU, it is directly skipped. If it is a GPU, it is detected whether it is GFX000. If so, it means that the hardware information file is damaged, and the current hardware node is skipped. After the hardware architecture detection is completed, the unique id is read according to the current hardware node information file, and it is judged whether the current unique id exists in visibleDevices. If a match is detected, the detection of the next hardware node is performed.

[0079] Step S3: Model quantization and compression:

[0080] S3.1 Basic model quantization and compression:

[0081] In the first step, download the mainstream open-source model or import the privately trained model. In the example, the DEEPSEEK-R1-7B model is used. The steps for quantizing the model include:

[0082] 1) Model acquisition: Download the model from the modelscape or huggingface repository. In the example, download the model example through the modelscape storage platform, and download it through the modelscope download --model deepseek-ai / DeepSeek-R1-Distill-Qwen-7B --local_dir deepseek-7b command.

[0083] 2) Model conversion: Use the script code provided by llama.cpp to convert the model to adapt to the DCU and model management software. In the example, use Python to run the convert_hf_to_gguf.py code for model conversion. The specific command is: python convert_hf_to_gguf.py --outfile / home / model / trans / deepseek-7b.gguf / home / model / deepseek-7b /

[0084] 3) Model quantization: The model can be quantized through ollama commands or by running the executable file using llama.cpp. In the example, use llama.cpp to run. / llama-quantize / home / model / deepseek-7b / / home / model / deepseek-7b_q4 / Q4 command to perform Q4 quantization on the DEEPSEEK-R1-7B model.

[0085] Step S4 mentioned above: Multi-model management and application verification.

[0086] S4.1: Multi-model import and model switching: In the example, import models such as qwen-0.5b-f16, deepseek-7b-f16, qwen-7b-f16, etc. Use ollama list to detect the available model list as follows:

[0087] NAME ID SIZE test_model_qwen0.5b:latest 372109218d7d994 994MB deepseek-7b:latest d1a3c615d4cc 15GB qwen-7b:latest 63bef4791d43 15GB

[0088] By sequentially executing ollama run model name, start the local model for simple execution, and monitor the usage of hardware VARM and DCU core resources through HY-SMI.

[0089] S4.2: Adaptation of running framework parallelism: In the example, import the qwen-70b-q4 model. By executing the ollama run qwen70b_test command, start the qwen70b model after Q4 quantization locally, and monitor the usage of hardware VARM and DCU core resources through HY-SMI to conduct the test of single-model multi-card parallelism; by executing the ollama run deepseek-7b and ollama run deepseek-7b commands, start the deepseek-7B and qwen-7b models locally to conduct the multi-model multi-card parallelism detection, and monitor the usage of hardware VARM and DCU core resources through HY-SMI.

Claims

1. A method for transplanting and adapting a local large language model operating framework based on a domestic DCU environment, characterized in that: The following steps are involved: S1: Hardware equipment environment detection and deep learning framework adaptation; S2: Large language model operation framework adaptation; S3: Quantization and compression of large language models; S4: Multi-model management and parallelism adaptation of the running framework.

2. The method for transplanting and adapting a local large language model operating framework based on a domestic DCU environment according to claim 1 is characterized in that: The specific steps of step S1 include: S1.1: Install DCU hardware firmware and drivers; S1.1.a: Install DCU hardware firmware; S1.1.b: Detect and install driver dependency libraries; S1.1.c: Get the driver installation package and driver compilation; S1.1.d: Repeat S1.1.a-S1.1.c until the DCU hardware firmware and driver validity test passes; S1.2: Deep learning framework adaptation; S1.2.a: Install and configure the environment manager. S1.2.b: Obtain and install deep learning framework packages; S1.2.c: Repeat S1.2.a-S1.2.b until the effectiveness of the deep learning environment framework is verified.

3. The method for transplanting and adapting a local large language model operating framework based on a domestic DCU environment according to claim 2 is characterized in that: The driver program dependent component library includes a compiler, a cross-compilation tool and a dependent dynamic link library.

4. The method for transplanting and adapting a local large language model operating framework based on a domestic DCU environment according to claim 1 is characterized in that: The specific steps of step S2 include: S2.1: Get the source code of the big language runtime framework; S2.2: Detect and install the dependency library of the large language runtime framework; S2.3: Compile configuration and hardware detection logic modification and optimization; S2.4: Repeat S2.1-S2.3 until the large language runtime framework adaptation validity verification is qualified.

5. The method for transplanting and adapting a local large language model operating framework based on a domestic DCU environment according to claim 4 is characterized in that: The compilation configuration and hardware detection logic modification and optimization are specifically as follows: S2.3.

1. Change the software build file path detection code: add a new environment variable to directly point to the software build file path; S2.3.

2. Change GPU environment detection code: change DriverVersionFile to DCU path; S2.3.

3. Modify the DCU parallel running code: Modify the DCU parallel running code to an explicit 1024 kernel function; S2.3.4, modify the compilation file: delete the git_module_setup module to disable the llama.cpp version update, change HIP_PATH, ROCM_PATH and LIBRARY_PATH to the default configuration to adapt to DCU compilation; S2.3.5, Hardware pre-detection and isolation optimization: First, identify the software parameter OLLAMA_DCU_ONLY_UUID. If the OLLAMA_DCU_ONLY_UUID parameter value is OFF, the default hardware pre-detection is used; if the OLLAMA_DCU_ONLY_UUID parameter value is ON, the hardware detection under software control is used; After the hardware pre-detection is completed, the available hardware information is output. If there is no available GPU node, the available hardware is set to CPU, and the deep learning framework is used to call the system hardware.

6. The method for transplanting and adapting a local large language model operating framework based on a domestic DCU environment according to claim 5 is characterized in that: The default hardware pre-detection process is as follows: first identify the hardware serial number assigned to the HIP_VISIBLE_DEVICES system parameter record, then scan the hardware nodes under the / sys / class / kfd / kfd / topology / nodes / file in turn, and obtain the system hardware information, including the system architecture and hardware serial number, through the driver file; determine whether the hardware architecture of the current node is GFX000 or a hardware architecture not supported by the software based on the obtained hardware information; If the current node is valid, determine whether the allocated hardware exists based on the hardware serial number. If not, skip the current node and proceed to the next hardware node detection.

7. The method for transplanting and adapting a local large language model operating framework based on a domestic DCU environment according to claim 5 is characterized in that: The hardware detection process under the software control is as follows: first, the hardware UUID applied for by the user is recorded through the ROCR_VISIBLE_DEVICES system parameter, and then the hardware nodes under the / sys / class / kfd / kfd / topology / nodes / file are scanned to obtain the system hardware architecture and hardware independent id, and whether the current node is an abnormal architecture is determined according to the hardware architecture. If the current node is a valid node, whether the hardware UUID applied for by the user exists is determined according to the hardware independent id. If abnormal information is detected, the current node detection is skipped; if there is no abnormality in the current node hardware detection, the current node is added to the available GPU node queue and the next hardware node detection is performed.

8. The method for transplanting and adapting a local large language model operating framework based on a domestic DCU environment according to claim 1 is characterized in that: The specific steps of step S3 include: S3.1: Model selection and format conversion; S3.2: Model quantization and compression.

9. The method for transplanting and adapting a local large language model operating framework based on a domestic DCU environment according to claim 1 is characterized in that: The specific steps of step S4 include: S4.1: Multi-model import and model switching: import multiple different models into the framework and run the corresponding models individually in batches to monitor model resource usage; S4.2: Run framework parallelism adaptation: Run multiple models simultaneously and monitor the resource usage of each model to verify the parallel usage of the run framework.

Citation Information

Cited By

  • Method and system for adapting credential host to deep learning framework based on edge calculation

    CN121031726A

  • A method and system for adapting a deep learning framework based on edge computing of a signal creation host

    CN121031726B

  • Ollama reasoning optimization method for domestic platform

    CN121706995A