A large-scale model hardware and software integrated machine tuning method and device for rapid deployment

Obtain hardware information and business needs through artificial intelligence models, optimize resource allocation and deployment of large-scale software and hardware all-in-one machines, solve the problems of long deployment time and low resource utilization, and achieve fast and efficient model deployment and resource management.

CN120255906BActive Publication Date: 2025-08-15ZHEJIANG HUATIE EMERGENCY EQUIP SCI & TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510757041.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-08-15
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The deployment process of large-model hardware and software all-in-one machines takes a long time, has low deployment efficiency, extensive resource management, serious fragmentation of video and memory, and low hardware resource utilization rate.

Method used

Hardware information and general model requirements are obtained through artificial intelligence models, initial deployment plans are generated, and hardware resource allocation is optimized according to business needs, and model deployment is dynamically adjusted based on visual configuration panels and industry knowledge graphs.

Benefits of technology

It realizes reasonable resource allocation and flexible optimization of large-model software and hardware all-in-one machines without relying on specific model information, improves deployment efficiency, reduces deployment time, and adapts to the needs of different industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255906B_ABST
    Figure CN120255906B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a large-scale model hardware and software integrated machine tuning method and device for rapid deployment, which can improve the overall deployment efficiency of the large-scale model hardware and software integrated machine. The method includes: in response to a user powering on the large-scale model hardware and software integrated machine, obtaining hardware information of each hardware module, and generating a first deployment plan based on the hardware information of each hardware module and preset general model requirements through an artificial intelligence model; in response to a user configuring the business requirements of the large-scale model hardware and software integrated machine in a visual interface, determining a target large model from multiple pre-trained large models built into the large-scale model hardware and software integrated machine through an artificial intelligence model according to the business requirements, and adjusting the first deployment plan according to the hardware conditions required by the target large model to obtain a second deployment plan for the large-scale model hardware and software integrated machine; and deploying the target large model in the large-scale model hardware and software integrated machine based on the second deployment plan.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of large model technology, and in particular to a large model hardware and software integrated machine tuning method and device for rapid deployment. Background Art

[0002] A large-scale model hardware and software all-in-one device refers to a product that integrates the hardware and software related to large-scale model applications into a single device or system. The hardware includes high-performance processors, large-capacity memory, high-speed storage devices, and chips specifically designed to accelerate large-scale model operations, providing the computing power and data storage capacity required for large-scale model operation. The software encompasses the large-scale model itself, as well as the associated operating system, drivers, model call interfaces, management, and optimization tools, ensuring stable and efficient operation of the large model and facilitating user interaction and use.

[0003] In related technologies, large model deployment requires manual configuration of drivers, model loading, and network environment, which takes a long time, for example, more than 48 hours, and has low deployment efficiency. Summary of the Invention

[0004] In order to overcome the problems existing in the related art, the embodiments of the present disclosure provide a large-model hardware and software integrated machine tuning method and device for rapid deployment, which are used to solve the defects in the related art.

[0005] According to a first aspect of an embodiment of the present disclosure, a method for optimizing a large-scale hardware-software integrated machine for rapid deployment is provided, comprising:

[0006] In response to a user powering on a large-model hardware and software all-in-one machine, hardware information of each hardware module in the large-model hardware and software all-in-one machine is obtained, and a first deployment plan for the large-model hardware and software all-in-one machine is generated through an artificial intelligence model based on the hardware information of each hardware module and preset general model requirements, wherein the hardware information includes the hardware model of the hardware module and transmission information between each of the hardware modules, and the first deployment plan includes a model quantization plan determined according to the hardware model and a model data transmission plan determined according to the transmission information;

[0007] In response to the user configuring business requirements for the large model hardware and software all-in-one machine in the visual interface, the artificial intelligence model determines a target large model from a plurality of pre-trained large models built into the large model hardware and software all-in-one machine according to the business requirements, and adjusts the first deployment plan according to the hardware conditions required by the target large model to obtain a second deployment plan for the large model hardware and software all-in-one machine;

[0008] Based on the second deployment solution, the target large model is deployed in the large model hardware and software integrated machine.

[0009] In one embodiment, the generating of a first deployment plan for the large-model hardware-software integrated machine using an artificial intelligence model based on the hardware information of each hardware module and preset general model requirements includes:

[0010] Generate a hardware topology diagram of the large-scale hardware and software integrated machine using the artificial intelligence model based on the hardware information of each hardware module, wherein a node in the hardware topology diagram represents a hardware module in the large-scale hardware and software integrated machine, and the node is labeled with the hardware model of the corresponding hardware module; the connection relationship between the nodes in the hardware topology diagram is used to represent the physical connection link between the corresponding hardware modules, and the edge is labeled with the transmission information of the corresponding physical connection link;

[0011] The artificial intelligence model is used to generate a first deployment plan for the large-model hardware and software integrated machine based on the hardware topology diagram and preset general model requirements.

[0012] In one embodiment, deploying the target large model in the large model hardware and software integrated machine based on the second deployment solution includes:

[0013] Displaying the second deployment plan on the large model hardware and software integrated machine;

[0014] In response to the user's adjustment operation on the second deployment plan, the artificial intelligence model determines, based on the hardware information of the hardware modules, whether the hardware conditions of the large-model hardware and software all-in-one machine can achieve a target adjustment plan for the second deployment plan corresponding to the adjustment operation; and if the hardware conditions of the large-model hardware and software all-in-one machine cannot achieve the target adjustment plan, the artificial intelligence model generates a recommended adjustment plan for the second deployment plan based on the adjustment operation and the hardware information of the hardware modules;

[0015] In response to the user's confirmation operation on the recommended adjustment plan, the target large model is deployed in the large model hardware and software integrated machine based on the recommended adjustment plan.

[0016] In one embodiment, the large model hardware and software all-in-one machine has built-in preset industry knowledge graphs corresponding to different industries, and also includes:

[0017] Displaying a visual configuration panel for a target agent, wherein the target agent is an agent associated with the target large model;

[0018] In response to a natural language input operation in the visualization configuration panel, obtaining a business requirement for the target agent corresponding to the natural language operation;

[0019] The artificial intelligence model is used to obtain the target industry knowledge graph corresponding to the business needs in the preset industry knowledge graph, and based on the business needs, the optimal work node combination that can achieve the business needs is recommended in the target industry knowledge graph. Based on the optimal work node combination, the agent workflow corresponding to the target agent is generated, wherein the target industry knowledge graph includes the agent work nodes required by the industry corresponding to the business needs and the execution order between each of the agent work nodes.

[0020] In one embodiment, obtaining the target industry knowledge graph corresponding to the business requirement from the preset industry knowledge graph through the artificial intelligence model includes:

[0021] Identify industry keywords based on the business needs through the artificial intelligence model, and search in the preset industry knowledge graph based on the industry keywords;

[0022] If the preset industry knowledge graph corresponding to the business demand is not found, the industry characteristics of the business demand are understood through the artificial intelligence model to obtain the demand industry characteristics corresponding to the business demand and, based on the demand industry characteristics, similar industries are matched among all industries corresponding to the preset industry knowledge graph, and the preset industry knowledge graph corresponding to the similar industry is determined as the target industry knowledge graph corresponding to the business demand.

[0023] In one embodiment, it further includes:

[0024] When it is monitored that the video memory usage of the large model hardware and software all-in-one machine exceeds a preset threshold, the artificial intelligence model is used to predict the time point when the video memory resources of the large model hardware and software all-in-one machine will be exhausted based on the video memory usage trend of the large model hardware and software all-in-one machine and the inference task characteristics of the target large model;

[0025] When the time distance from the time point is less than a first preset time distance, the reasoning task of the target large model is analyzed by the artificial intelligence model to obtain non-critical weight parameters of the target large model, and the non-critical weight parameters are quantized and compressed, and / or, the computational graph structure of the target large model is analyzed by the artificial intelligence model to obtain target computing nodes to be merged, and the target computing nodes are merged.

[0026] In one embodiment, it further includes:

[0027] Acquire data access information in the historical reasoning task of the target large model, wherein the data access information includes data access time and data access frequency;

[0028] Analyzing the data access information through the artificial intelligence model to capture the periodicity and correlation of the target large model data access, predicting frequent data that may be frequently accessed within a second preset time period in the future based on the periodicity and correlation of the target large model data access, and loading the frequent data into a cache, wherein the frequent access indicates that the access frequency exceeds a preset threshold;

[0029] The real-time reasoning tasks of the target large model are analyzed by the artificial intelligence model to obtain the predicted access frequency of the reasoning tasks of the target large model to the data in the cache within a third preset time period in the future, and the data in the cache is updated based on the predicted access frequency.

[0030] According to a second aspect of an embodiment of the present disclosure, a large-scale model hardware and software integrated machine tuning device for rapid deployment is provided, comprising:

[0031] a first tuning module, configured to obtain hardware information of each hardware module in the large-model hardware and software integrated machine in response to a user power-on operation on the large-model hardware and software integrated machine, and generate a first deployment plan for the large-model hardware and software integrated machine based on the hardware information of each hardware module and preset general model requirements through an artificial intelligence model, wherein the hardware information includes the hardware model of the hardware module and transmission information between the hardware modules, and the first deployment plan includes a model quantization plan determined according to the hardware model and a model data transmission plan determined according to the transmission information;

[0032] A second tuning module is configured to respond to the user's business requirement configuration for the large model hardware and software appliance in the visual interface, determine a target large model from multiple pre-trained large models built into the large model hardware and software appliance through the artificial intelligence model according to the business requirement configuration, and adjust the first deployment plan according to the hardware conditions required by the target large model to obtain a second deployment plan for the large model hardware and software appliance;

[0033] A deployment module is used to deploy the target large model in the large model hardware and software integrated machine based on the second deployment solution.

[0034] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a memory and a processor, wherein the memory is used to store computer instructions that can be executed on the processor, and the processor is used to implement any one of the methods described in the first aspect when executing the computer instructions.

[0035] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method described in any one of the first aspects is implemented.

[0036] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which implements any one of the methods described in the first aspect when executed by a processor.

[0037] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:

[0038] The large-model hardware and software integration machine optimization method for rapid deployment provided by the embodiment of the present disclosure can automatically perform hardware detection after the large-model hardware and software integration machine is powered on to obtain the hardware information of each hardware module in the large-model hardware and software integration machine, and then generate a first deployment plan based on the hardware information and general model requirements through the artificial intelligence model. Afterwards, the first deployment plan can be adjusted by the artificial intelligence model according to the hardware conditions required by the target large model in the actual business to obtain a second deployment plan, and finally the model is deployed based on the second deployment plan. Thus, the large-model hardware and software integration machine can be enabled to first perform reasonable resource allocation without relying on specific model information through a dynamic adjustment mechanism based on hardware information, and then perform flexible optimization according to the needs of the actual model, thereby improving the overall deployment efficiency of the large-model hardware and software integration machine. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0040] Figure 1 This is a flow chart of a method for optimizing a large-scale hardware-software integrated machine for rapid deployment, as shown in an exemplary embodiment of the present disclosure;

[0041] Figure 2 This is a block diagram of a large-scale model hardware and software integrated machine optimization device for rapid deployment, shown in an exemplary embodiment of the present disclosure;

[0042] Figure 3 FIG. 4 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0043] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0044] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0045] It should be understood that although the terms "first," "second," and "third" may be used in this disclosure to describe various types of information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information.

[0046] A large-scale model hardware and software all-in-one device refers to a product that integrates the hardware and software related to large-scale model applications into a single device or system. The hardware includes high-performance processors, large-capacity memory, high-speed storage devices, and chips specifically designed to accelerate large-scale model operations, providing the computing power and data storage capacity required for large-scale model operation. The software encompasses the large-scale model itself, as well as the associated operating system, drivers, model call interfaces, management, and optimization tools, ensuring stable and efficient operation of the large model and facilitating user interaction and use.

[0047] In related technologies, deploying large models requires manual configuration of drivers, model loading, and network environments, which takes a long time, potentially exceeding 48 hours, and results in low deployment efficiency. Furthermore, resource management is extensive, primarily manifested in severe graphics memory fragmentation and low hardware resource utilization.

[0048] Based on this, at least one embodiment of the present disclosure provides a large-scale model hardware and software integration optimization method for rapid deployment. Figure 1 , please refer to Figure 1 , which shows the process of the method, including steps S101 to S103.

[0049] In step S101, in response to a user powering on a large-scale hardware and software integrated machine, hardware information of each hardware module in the large-scale hardware and software integrated machine is obtained, and a first deployment plan for the large-scale hardware and software integrated machine is generated using an artificial intelligence model based on the hardware information of each hardware module and preset general model requirements. The hardware information includes the hardware models of the hardware modules and transmission information between the hardware modules, and the first deployment plan includes a model quantization plan determined based on the hardware models and a model data transmission plan determined based on the transmission information.

[0050] It should be understood that after powering on, the large-model hardware and software all-in-one machine can perform preliminary resource allocation based on common model operation requirements. This is because most large models have some common requirements during operation, such as GPU (Graphics Processing Unit) computing resources, storage read and write bandwidth, and network transmission bandwidth. The large-model hardware and software all-in-one machine can use the artificial intelligence model to allocate reasonable data transmission paths and bandwidth resources for these common requirements based on the hardware information of each hardware module to meet the basic operating conditions of most models. This preliminary allocation can ensure that the large-model hardware and software all-in-one machine has a certain degree of versatility and flexibility, providing a good basic environment for the specific deployment and operation of subsequent models, without the need for special customization for specific models.

[0051] For example, the model quantization scheme can be determined based on the hardware model, such as automatically matching the preset quantization strategy after identifying the Kunlun Core P800 GPU.

[0052] In step S102, in response to the user's business requirement configuration of the large model hardware and software integrated machine in the visual interface, the artificial intelligence model is configured according to the business requirements, and the target large model is determined from the multiple pre-trained large models built into the large model hardware and software integrated machine. According to the hardware conditions required by the target large model, the first deployment plan is adjusted to obtain the second deployment plan of the large model hardware and software integrated machine.

[0053] For example, when the AI model is configured based on the business requirements and determines a target large model from multiple pre-trained large models built into the large model hardware and software appliance, the AI model can consider the hardware resources available for model execution in the large model hardware and software appliance, such as the number, performance, and storage capacity of GPUs. Specifically, the AI model can use this hardware resource information to select target large models that match the hardware resources. For example, if the hardware is equipped with multiple high-performance GPUs, the AI model can recommend complex large models that fully utilize the GPU's parallel computing capabilities and meet the business requirements. If the GPU performance is average, the AI model can recommend relatively lightweight large models with lower hardware requirements that still meet the business requirements.

[0054] Once the model to be deployed is determined, that is, the target large model is determined, the first deployment plan can be further dynamically adjusted based on the characteristics of the target large model. The large model hardware and software integration machine will determine the hardware conditions required for the target large model based on information such as the size, computational complexity, and data flow pattern of the target large model, and then optimize and fine-tune the previously allocated data transmission paths and bandwidth resources. For example, if the target large model is an ultra-large-scale pre-trained language model that requires a large amount of video memory to store model parameters and generates frequent video memory access and data exchange during the inference process, the allocation ratio of GPU video memory can be appropriately increased, and the data transmission path between the GPU and video memory can be optimized to improve the performance of the model operation.

[0055] In step S103, based on the second deployment solution, the target large model is deployed in the large model hardware and software integrated machine.

[0056] Therefore, a dynamic adjustment mechanism based on hardware information can be used to enable large-model hardware and software integrated machines to first make reasonable resource allocations without relying on specific model information, and then flexibly optimize according to the actual model requirements, thereby improving the overall deployment efficiency of large-model hardware and software integrated machines.

[0057] For ease of understanding, the above steps are described below with examples in detail.

[0058] In one embodiment, in step S101, a first deployment plan for the large-model hardware and software integrated machine is generated by an artificial intelligence model based on the hardware information of the hardware modules and the preset general model requirements, including: generating a hardware topology diagram of the large-model hardware and software integrated machine based on the hardware information of the hardware modules by the artificial intelligence model, wherein a node in the hardware topology diagram represents a hardware module in the large-model hardware and software integrated machine, and the node is marked with the hardware model of the corresponding hardware module, and the connection relationship between the nodes in the hardware topology diagram is used to represent the physical connection link between the corresponding hardware modules, and the edge is marked with the transmission information of the corresponding physical connection link; generating the first deployment plan for the large-model hardware and software integrated machine by an artificial intelligence model based on the hardware topology diagram and the preset general model requirements.

[0059] It should be understood that the hardware of a large-scale hardware and software all-in-one machine usually includes several key parts:

[0060] Computing cores: such as GPU clusters, used to parallelize the model's large number of computing tasks, such as matrix operations;

[0061] Storage module: includes cache, SSD (Solid State Drive), etc., used to store model parameters, intermediate calculation results and other data;

[0062] Network interface: For example, a high-speed Ethernet card or InfiniBand network card is responsible for data transmission between different computing nodes or with external systems;

[0063] Mainboard and backplane: used to connect various hardware components and build physical communication links between hardware.

[0064] After powering on, the large-scale hardware and software appliance can activate its intelligent hardware topology awareness function. Dedicated sensing chips or firmware embedded in the hardware's underlying infrastructure can send identification signals to the appliance's topology awareness module, reporting its model, performance parameters, and other information. After receiving these signals, the topology awareness module automatically draws a hardware topology map using a pre-set topology recognition algorithm. For example, it can identify the system's four high-performance GPUs, each connected to a specific motherboard slot via a PCIe (Peripheral Component Interconnect Express) bus. These GPUs are interconnected via NVLink high-speed interconnects. Furthermore, multiple SSDs in the storage module are connected to the motherboard via SATA (Serial ATA) or NVMe interfaces, and the network interface is a dual-channel 100Gbps Ethernet card connected to the motherboard's network slot. Based on this information, the hardware topology map is then automatically drawn.

[0065] For example, a hardware topology diagram can be a logical diagram that can be used to clearly show the connection relationship and data transmission path between various hardware components. In this diagram, each node can represent a hardware component, such as a GPU, SSD hard disk, network interface, etc., and the corresponding hardware model can be marked next to the node. The lines between the nodes can represent the physical connection links between them, and key parameters such as the bus type and bandwidth of the connection can be marked next to the lines. For example, GPU nodes are connected through NVLink links with a marked bandwidth of 50GB / s; SSD hard disk nodes are connected to the motherboard through NVMe interfaces with a bandwidth of 4GB / s; network interface nodes are connected to the external network through 100Gbps Ethernet links.

[0066] As a result, the large-model hardware and software all-in-one machine can have intelligent hardware topology perception capabilities. After power-on, it can automatically draw a hardware topology map and identify the connection relationship and data transmission path between hardware. Based on the hardware topology map and general model requirements, it can dynamically adjust the data transmission path and resource allocation, such as automatically allocating more bandwidth resources to high-load modules, thereby improving the deployment efficiency and overall performance of the large-model hardware and software all-in-one machine.

[0067] In one embodiment, in step S103, based on the second deployment plan, the target large model is deployed in the large model hardware and software all-in-one machine, including: displaying the second deployment plan on the large model hardware and software all-in-one machine; in response to the user's adjustment operation on the second deployment plan, judging by the artificial intelligence model whether the hardware conditions of the large model hardware and software all-in-one machine can achieve the target adjustment plan for the second deployment plan corresponding to the adjustment operation based on the hardware information of the hardware modules, and when the hardware conditions of the large model hardware and software all-in-one machine cannot achieve the target adjustment plan, generating a recommended adjustment plan for the second deployment plan based on the adjustment operation and the hardware information of the hardware modules by the artificial intelligence model; in response to the user's confirmation operation on the recommended adjustment plan, deploying the target large model in the large model hardware and software all-in-one machine based on the recommended adjustment plan.

[0068] For example, the user can adjust the second deployment plan based on the visual operation interface. The visual operation interface can display the deployment parameters corresponding to the second deployment plan, and the user can adjust the second deployment plan by re-entering the deployment parameters. Afterwards, the artificial intelligence model can automatically identify whether the user-adjusted plan matches the hardware conditions of the large-model hardware and software all-in-one machine. If it matches, the user's adjustment operation can be directly applied. If it does not match, the artificial intelligence model can generate a recommended adjustment plan that matches the hardware conditions of the large-model hardware and software all-in-one machine. Thus, the user can first make a visual adjustment to the second deployment plan, and then the artificial intelligence model can automatically adjust and optimize the second deployment plan, thereby meeting user needs and improving the deployment efficiency and applicability of the large-model hardware and software all-in-one machine.

[0069] In one embodiment, the large model hardware and software integrated machine has built-in preset industry knowledge graphs corresponding to different industries, and can also display a visual configuration panel for the target intelligent agent, wherein the target intelligent agent is the intelligent agent associated with the target large model; in response to the natural language input operation in the visual configuration panel, the business requirements for the target intelligent agent corresponding to the natural language operation are obtained; the target industry knowledge graph corresponding to the business requirements is obtained in the preset industry knowledge graph through the artificial intelligence model, and based on the business requirements, the optimal work node combination that can realize the business requirements is recommended in the target industry knowledge graph, and based on the optimal work node combination, the intelligent agent workflow corresponding to the target intelligent agent is generated, wherein the target industry knowledge graph includes the intelligent agent work nodes required by the industry corresponding to the business requirements and the execution order between each of the intelligent agent work nodes.

[0070] It should be understood that in related technologies, pre-trained models and tool chains are not optimized for specific scenarios, resulting in high secondary development costs. Furthermore, the lack of a visual tool chain makes it difficult for non-technical personnel to fine-tune models and orchestrate processes. However, the present disclosure provides a visual configuration panel where users can input natural language input. This panel then automatically optimizes the AI model for specific scenarios, allowing even non-technical personnel to quickly fine-tune models and orchestrate processes for specific scenarios.

[0071] For example, taking the medical field as an example, the preset industry knowledge graph can include the following core information:

[0072] Disease entity: disease name (e.g., diabetes), symptoms (e.g., polydipsia, polyuria), complications (e.g., retinopathy);

[0073] Diagnosis and treatment process entities: examination items (such as blood routine, imaging examination), diagnostic methods (such as imaging diagnosis, expert consultation), and treatment methods (such as drugs, surgery).

[0074] Rules and regulatory entities: medical compliance reviews, clinical pathway guidelines (such as standard procedures for diabetes management).

[0075] Technical tool entities: available models and tools (such as multimodal diagnosis models, natural language understanding models, and medical record structured parsing engines).

[0076] Causal relationship: symptom → disease association, treatment → efficacy relationship (such as insulin injection → blood sugar control).

[0077] For example, a user can use natural language input in the visual configuration panel to enter the following: "Build a medical consultation agent that automatically analyzes patient medical records, provides diagnostic recommendations, and ensures compliance with medical regulations." This content is then semantically parsed and intent extracted. For example, an NLP (Natural Language Processing) model analyzes the keywords "medical record analysis," "diagnostic recommendations," and "medical regulations." Then, through intent recognition, it determines that the user's goal is to build a compliant automated diagnostic process.

[0078] After that, knowledge graph retrieval and matching can be performed. First, entity matching is performed to retrieve technical tool entities related to medical record analysis, such as "medical record structured parsing"; matching entities related to diagnostic recommendations, such as "multimodal diagnostic model"; and association rules and regulatory entities, such as "medical compliance review." Then, based on the causal relationship in the knowledge graph, the following optimal work node combination is determined: medical record structured parsing → medical compliance review → multimodal diagnosis → medical compliance review → result output. Finally, based on this optimal work node combination, the following intelligent agent workflow can be obtained: perform medical record structured parsing on the input content → perform medical compliance review on the parsed content → perform multimodal diagnosis on the input content after passing the medical compliance review → perform medical compliance review on the diagnosis results → generate a diagnosis report → output a diagnosis report.

[0079] This can avoid users from manually configuring complex processes, which is especially suitable for practitioners who lack technical background. It can compress process design that originally took several days to minutes, and can automatically iterate as the knowledge graph is updated, thereby improving the deployment efficiency of large-model hardware and software integrated machines.

[0080] In one embodiment, the target industry knowledge graph corresponding to the business demand is obtained from the preset industry knowledge graph by the artificial intelligence model, including: identifying industry keywords based on the business demand by the artificial intelligence model, and searching in the preset industry knowledge graph based on the industry keywords; if the preset industry knowledge graph corresponding to the business demand is not found, the industry characteristics of the business demand are understood by the artificial intelligence model to obtain the demand industry characteristics corresponding to the business demand, and based on the demand industry characteristics, similar industries are matched in all industries corresponding to the preset industry knowledge graph, and the preset industry knowledge graph corresponding to the similar industry is determined as the target industry knowledge graph corresponding to the business demand.

[0081] For example, continuing with the above example, for example, if the business requirement input by the user is: "Build a medical consultation intelligent agent that can automatically analyze patient medical records, give diagnostic suggestions, and ensure compliance with medical regulations", the industry keyword obtained can be "medical". Then, based on the industry keyword, search in the preset industry knowledge graph. It should be understood that each preset industry knowledge graph is pre-associated with a corresponding industry keyword. Therefore, when the corresponding preset industry knowledge graph is found based on the industry keyword, the subsequent process can be executed directly based on the preset industry knowledge graph. In the case that the corresponding preset industry knowledge graph is not found, the industry characteristics of the business demand can be understood through the artificial intelligence model, so as to obtain similar industries corresponding to the business demand, and then the subsequent process can be executed based on the similar industry.

[0082] As a result, the deployment efficiency of large-model hardware and software integrated machines in different industries can be improved, and through intelligent identification of industry commonalities, it can dynamically adapt to industry differences and provide technical support for rapid implementation in multiple fields.

[0083] In one embodiment, when it is monitored that the video memory occupancy of the large model hardware and software integrated machine exceeds a preset threshold, the artificial intelligence model can be used to predict the time point when the video memory resources of the large model hardware and software integrated machine are exhausted based on the video memory usage trend of the large model hardware and software integrated machine and the reasoning task characteristics of the target large model; when the time distance from the time point is less than a first preset time length, the artificial intelligence model is used to analyze the reasoning task of the target large model to obtain non-critical weight parameters of the target large model, and quantize and compress the non-critical weight parameters, or, the artificial intelligence model is used to analyze the computational graph structure of the target large model to obtain the target computing nodes to be merged, and merge the target computing nodes.

[0084] It should be understood that when using large-scale all-in-one hardware and software appliances to train natural language processing models (such as those based on the Transformer architecture), the model parameters are enormous, and video memory usage fluctuates constantly and is difficult to predict during training. As training progresses, video memory usage gradually increases. If not adjusted promptly, insufficient video memory may cause training to stall or even be interrupted. Alternatively, when using large-scale all-in-one hardware and software appliances for inference tasks on image recognition models (such as ResNet-50), video memory requirements can rapidly increase when multiple high-resolution image requests need to be processed simultaneously. This is especially true when performing image recognition on real-time video streams, where low latency and high throughput are essential. Efficient use of video memory is crucial.

[0085] Taking natural language processing as an example, a large-scale model hardware and software appliance can utilize its built-in hardware performance monitoring module to collect key performance indicators (KPIs) such as video memory usage, usage trends, and GPU compute task queue length at regular intervals (e.g., 1 second). This data is then fed into an AI model for prediction. This AI model is trained based on historical performance data from extensive training processes, learning patterns in video memory usage and their relationship to training task characteristics (e.g., number of model layers, batch size, and optimization algorithm). When video memory usage approaches a threshold (e.g., 90%), the AI model can predict the impending exhaustion of video memory based on current training task characteristics and video memory usage trends. The AI model can then employ a dynamic quantization strategy to quantize non-critical weight parameters of the target large model from 32-bit floating point (FP32) to 16-bit floating point (FP16), thereby reducing video memory usage. Furthermore, the large-scale model appliance can analyze the computational graph structure of the target large model, identifying compute nodes that can be merged or optimized, and adopting a more efficient computational graph structure to further free up video memory.

[0086] This enables AI-based hardware performance prediction and tuning. AI models are used to monitor and predict hardware performance in real time, identifying performance bottlenecks in advance. For example, if GPU memory is predicted to be exhausted, the model's GPU memory allocation strategy is automatically adjusted. This frees up GPU memory by compressing model parameters and / or adopting a more efficient computational graph structure, thus avoiding system lags.

[0087] In one embodiment, data access information in historical reasoning tasks of the target large model can also be obtained, and the data access information includes data access time and data access frequency; the data access information is analyzed by the artificial intelligence model to capture the periodicity and correlation of data access of the target large model, and based on the periodicity and correlation of data access of the target large model, frequent data that may be frequently accessed within a second preset time period in the future is predicted, and the frequent data is loaded into the cache, wherein the frequent access represents that the access frequency exceeds a preset threshold; the real-time reasoning tasks of the target large model are analyzed by the artificial intelligence model to obtain the predicted access frequency of the target large model's reasoning tasks to the data in the cache within a third preset time period in the future, and the data in the cache is updated based on the predicted access frequency.

[0088] It should be understood that, for example, when using a large-model integrated hardware and software machine for text generation tasks, the model needs to frequently access data such as the vocabulary and word vectors. Information such as the access time and frequency of access to these data during the inference process can be collected through artificial intelligence models. Artificial intelligence models, such as Transformer-based time series analysis models, are then used to analyze and learn data characteristics to capture the periodicity and correlation of data access. For example, it is found that certain common words are frequently accessed in specific types of text generation tasks. Then, based on the periodicity and correlation of data access, the artificial intelligence model predicts vocabulary data and word vectors that may be frequently accessed and loads them into the cache. For example, for news headline generation tasks, it is predicted that the access frequency of words related to news hotspots will increase. The large-model integrated hardware and software machine loads the data corresponding to these words into the cache, reducing data loading delays during inference and improving inference speed.

[0089] It should also be understood that traditional cache eviction strategies may suffer from low cache hit rates when dealing with complex access patterns. However, the disclosed embodiments can dynamically adjust cache eviction strategies based on predicted access frequencies. If a word vector is predicted to be accessed less frequently in subsequent tasks, it will be prematurely evicted from the cache even if it has been recently accessed, making room for new data that may be accessed frequently.

[0090] Therefore, AI models analyze and predict cache data access patterns, pre-loading data that is likely to be frequently accessed into the cache. At the same time, cache eviction policies are dynamically adjusted based on the prediction results to ensure efficient use of cache space and improve data access efficiency.

[0091] Through any of the above embodiments, the present disclosure can provide a large-scale model hardware and software integrated machine system that is ready to use out of the box. Through software and hardware collaborative tuning technology, automated deployment process and scenario-based tool chain, enterprise users can complete deployment within 30 minutes, zero-code operation and efficient resource utilization.

[0092] According to the second aspect of the embodiment of the present disclosure, a large-scale model hardware and software integrated machine optimization device for rapid deployment is provided. Figure 2 The large-scale model hardware and software integrated machine optimization device 200 for rapid deployment includes:

[0093] The first tuning module 201 is used to obtain hardware information of each hardware module in the large model hardware and software integrated machine in response to the user's power-on operation on the large model hardware and software integrated machine, and generate a first deployment plan for the large model hardware and software integrated machine based on the hardware information of each hardware module and preset general model requirements through an artificial intelligence model, wherein the hardware information includes the hardware model of the hardware module and the transmission information between each hardware module, and the first deployment plan includes a model quantization plan determined according to the hardware model and a model data transmission plan determined according to the transmission information;

[0094] A second tuning module 202 is configured to respond to the user's business requirement configuration for the large model hardware and software appliance in the visual interface, determine a target large model from multiple pre-trained large models built into the large model hardware and software appliance based on the business requirement configuration using the artificial intelligence model, and adjust the first deployment plan based on the hardware conditions required by the target large model to obtain a second deployment plan for the large model hardware and software appliance;

[0095] The deployment module 203 is used to deploy the target large model in the large model hardware and software integrated machine based on the second deployment solution.

[0096] In one embodiment, the first tuning module 201 is configured to:

[0097] Generate a hardware topology diagram of the large-scale hardware and software integrated machine using the artificial intelligence model based on the hardware information of each hardware module, wherein a node in the hardware topology diagram represents a hardware module in the large-scale hardware and software integrated machine, and the node is labeled with the hardware model of the corresponding hardware module; the connection relationship between the nodes in the hardware topology diagram is used to represent the physical connection link between the corresponding hardware modules, and the edge is labeled with the transmission information of the corresponding physical connection link;

[0098] The artificial intelligence model is used to generate a first deployment plan for the large-model hardware and software integrated machine based on the hardware topology diagram and preset general model requirements.

[0099] In one embodiment, the deployment module 203 is used to:

[0100] Displaying the second deployment plan on the large model hardware and software integrated machine;

[0101] In response to the user's adjustment operation on the second deployment plan, the artificial intelligence model determines, based on the hardware information of the hardware modules, whether the hardware conditions of the large-model hardware and software all-in-one machine can achieve a target adjustment plan for the second deployment plan corresponding to the adjustment operation; and if the hardware conditions of the large-model hardware and software all-in-one machine cannot achieve the target adjustment plan, the artificial intelligence model generates a recommended adjustment plan for the second deployment plan based on the adjustment operation and the hardware information of the hardware modules;

[0102] In response to the user's confirmation operation on the recommended adjustment plan, the target large model is deployed in the large model hardware and software integrated machine based on the recommended adjustment plan.

[0103] In one embodiment, the large-scale model hardware and software integrated machine has built-in preset industry knowledge graphs corresponding to different industries. The large-scale model hardware and software integrated machine tuning device 200 for rapid deployment further includes:

[0104] A display module, configured to display a visual configuration panel for a target agent, wherein the target agent is an agent associated with the target macro model;

[0105] A first acquisition module is configured to, in response to a natural language input operation in the visual configuration panel, acquire a business requirement for the target agent corresponding to the natural language operation;

[0106] A workflow module is used to obtain the target industry knowledge graph corresponding to the business needs in the preset industry knowledge graph through the artificial intelligence model, and based on the business needs, recommend the optimal work node combination that can achieve the business needs in the target industry knowledge graph, and generate the agent workflow corresponding to the target agent based on the optimal work node combination, wherein the target industry knowledge graph includes the agent work nodes required by the industry corresponding to the business needs and the execution order between each of the agent work nodes.

[0107] In one embodiment, the workflow module is used to:

[0108] Identify industry keywords based on the business needs through the artificial intelligence model, and search in the preset industry knowledge graph based on the industry keywords;

[0109] If the preset industry knowledge graph corresponding to the business demand is not found, the industry characteristics of the business demand are understood through the artificial intelligence model to obtain the demand industry characteristics corresponding to the business demand and, based on the demand industry characteristics, similar industries are matched among all industries corresponding to the preset industry knowledge graph, and the preset industry knowledge graph corresponding to the similar industry is determined as the target industry knowledge graph corresponding to the business demand.

[0110] In one embodiment, the large-scale model hardware and software integrated machine optimization device 200 for rapid deployment further includes:

[0111] The first analysis module is configured to, when monitoring that the video memory usage of the large model hardware and software integrated machine exceeds a preset threshold, predict, by using the artificial intelligence model, a time point when the video memory resources of the large model hardware and software integrated machine will be exhausted based on the video memory usage trend of the large model hardware and software integrated machine and the inference task characteristics of the target large model;

[0112] The second analysis module is used to analyze the inference task of the target large model through the artificial intelligence model when the time distance from the time point is less than the first preset time length, obtain the non-critical weight parameters of the target large model, and quantize and compress the non-critical weight parameters, and / or analyze the computational graph structure of the target large model through the artificial intelligence model to obtain the target computing nodes to be merged, and merge the target computing nodes.

[0113] In one embodiment, the large-scale model hardware and software integrated machine optimization device 200 for rapid deployment further includes:

[0114] A second acquisition module is used to obtain data access information in the historical reasoning task of the target large model, wherein the data access information includes data access time and data access frequency;

[0115] a third analysis module, configured to analyze the data access information using the artificial intelligence model to capture the data periodicity and correlation of the target large model data accesses, predict frequent data that may be frequently accessed within a second preset time period in the future based on the periodicity and correlation of the target large model data accesses, and load the frequent data into a cache, wherein the frequent access indicates that the access frequency exceeds a preset threshold;

[0116] The fourth analysis module is used to analyze the real-time reasoning tasks of the target large model through the artificial intelligence model, obtain the predicted access frequency of the reasoning tasks of the target large model to the data in the cache within a third preset time period in the future, and update the data in the cache based on the predicted access frequency.

[0117] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the first aspect related to the method, and will not be elaborated here.

[0118] According to the third aspect of the embodiment of the present disclosure, please refer to Figure 3 , which exemplarily shows a block diagram of an electronic device, the electronic device 700 may include: a processor 701, a memory 702. The electronic device 700 may also include one or more of a multimedia component 703, an input / output (I / O) interface 704, and a communication component 705.

[0119] The processor 701 is used to control the overall operation of the electronic device 700 to complete all or part of the steps in any of the above methods. The memory 702 is used to store various types of data to support the operation of the electronic device 700. This data may include, for example, instructions for any application or method operating on the electronic device 700, as well as application-related data such as contact information, sent and received messages, images, audio, video, etc. The memory 702 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 703 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 702 or sent through the communication component 705. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 704 provides an interface between the processor 701 and other interface modules. The above-mentioned other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 705 is used for wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them, so the corresponding communication component 705 may include: a Wi-Fi module, a Bluetooth module, an NFC module.

[0120] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform any of the above methods.

[0121] In one embodiment, a computer-readable storage medium including program instructions is further provided, wherein the program instructions, when executed by a processor, implement the steps of any of the above methods. For example, the computer-readable storage medium may be the memory 702 including the program instructions, and the program instructions may be executed by the processor 701 of the electronic device 700 to perform any of the above methods.

[0122] In one embodiment, a computer program product is further provided. The computer program product includes a computer program that can be executed by a processor. When the computer program is executed by the processor, the steps of any of the above methods are implemented.

[0123] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details of the above embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the scope of protection of the present disclosure.

[0124] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. In order to avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.

[0125] In addition, the various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.

Claims

1. A large-scale model hardware and software integration optimization method for rapid deployment, characterized by: include: In response to a user powering on a large-model hardware and software all-in-one machine, hardware information of each hardware module in the large-model hardware and software all-in-one machine is obtained, and a first deployment plan for the large-model hardware and software all-in-one machine is generated through an artificial intelligence model based on the hardware information of each hardware module and preset general model requirements, wherein the hardware information includes the hardware model of the hardware module and transmission information between each of the hardware modules, and the first deployment plan includes a model quantization plan determined according to the hardware model and a model data transmission plan determined according to the transmission information; In response to the user configuring business requirements for the large-model hardware and software appliance in the visual interface, the artificial intelligence model determines a target large model from a plurality of pre-trained large models built into the large-model hardware and software appliance according to the business requirements, and adjusts the first deployment plan according to the hardware conditions required by the target large model to obtain a second deployment plan for the large-model hardware and software appliance; Based on the second deployment solution, the target large model is deployed in the large model hardware and software integrated machine.

2. The large-scale model hardware and software integrated machine optimization method for rapid deployment according to claim 1 is characterized in that: The generating, by the artificial intelligence model, of a first deployment plan for the large-model hardware-software integrated machine according to the hardware information of each hardware module and the preset general model requirements includes: Generate a hardware topology diagram of the large-model hardware and software all-in-one machine based on the hardware information of each hardware module through the artificial intelligence model, wherein a node in the hardware topology diagram represents a hardware module in the large-model hardware and software all-in-one machine, and the node is marked with the hardware model of the corresponding hardware module; the connection relationship between the nodes in the hardware topology diagram is used to represent the physical connection link between the corresponding hardware modules, and the edges between the nodes are marked with the transmission information of the corresponding physical connection link; The artificial intelligence model is used to generate a first deployment plan for the large-model hardware and software integrated machine based on the hardware topology diagram and preset general model requirements.

3. The large-scale model hardware and software integrated machine optimization method for rapid deployment according to claim 1 is characterized in that: The step of deploying the target large model in the large model hardware and software integrated machine based on the second deployment solution includes: Displaying the second deployment plan on the large model hardware and software integrated machine; In response to the user's adjustment operation on the second deployment plan, the artificial intelligence model determines, based on the hardware information of the hardware modules, whether the hardware conditions of the large-model hardware and software all-in-one machine can achieve a target adjustment plan for the second deployment plan corresponding to the adjustment operation; and if the hardware conditions of the large-model hardware and software all-in-one machine cannot achieve the target adjustment plan, the artificial intelligence model generates a recommended adjustment plan for the second deployment plan based on the adjustment operation and the hardware information of the hardware modules; In response to the user's confirmation operation on the recommended adjustment plan, the target large model is deployed in the large model hardware and software integrated machine based on the recommended adjustment plan.

4. The large-scale model hardware and software integrated machine optimization method for rapid deployment according to any one of claims 1 to 3, characterized in that: The large-scale hardware and software all-in-one machine has built-in preset industry knowledge graphs corresponding to different industries, and also includes: Displaying a visual configuration panel for a target agent, wherein the target agent is an agent associated with the target large model; In response to a natural language input operation in the visualization configuration panel, obtaining a business requirement for the target agent corresponding to the natural language operation; The artificial intelligence model is used to obtain the target industry knowledge graph corresponding to the business needs in the preset industry knowledge graph, and based on the business needs, the optimal work node combination that can achieve the business needs is recommended in the target industry knowledge graph. Based on the optimal work node combination, the agent workflow corresponding to the target agent is generated, wherein the target industry knowledge graph includes the agent work nodes required by the industry corresponding to the business needs and the execution order between each of the agent work nodes.

5. The large-scale model hardware and software integrated machine optimization method for rapid deployment according to claim 4 is characterized in that: The obtaining of the target industry knowledge graph corresponding to the business requirement from the preset industry knowledge graph by the artificial intelligence model includes: Identify industry keywords based on the business needs through the artificial intelligence model, and search in the preset industry knowledge graph based on the industry keywords; If the preset industry knowledge graph corresponding to the business demand is not found, the industry characteristics of the business demand are understood through the artificial intelligence model to obtain the demand industry characteristics corresponding to the business demand and, based on the demand industry characteristics, similar industries are matched among all industries corresponding to the preset industry knowledge graph, and the preset industry knowledge graph corresponding to the similar industry is determined as the target industry knowledge graph corresponding to the business demand.

6. The large-scale model hardware and software integrated machine optimization method for rapid deployment according to any one of claims 1 to 3, characterized in that: Also includes: When it is monitored that the video memory usage of the large model hardware and software all-in-one machine exceeds a preset threshold, the artificial intelligence model is used to predict the time point when the video memory resources of the large model hardware and software all-in-one machine will be exhausted based on the video memory usage trend of the large model hardware and software all-in-one machine and the inference task characteristics of the target large model; When the time distance from the time point is less than a first preset time distance, the reasoning task of the target large model is analyzed by the artificial intelligence model to obtain non-critical weight parameters of the target large model, and the non-critical weight parameters are quantized and compressed, and / or, the computational graph structure of the target large model is analyzed by the artificial intelligence model to obtain target computing nodes to be merged, and the target computing nodes are merged.

7. The large-scale model hardware and software integrated machine optimization method for rapid deployment according to any one of claims 1 to 3, characterized in that: Also includes: Acquire data access information in the historical reasoning task of the target large model, wherein the data access information includes data access time and data access frequency; Analyzing the data access information through the artificial intelligence model to capture the periodicity and correlation of the target large model data access, predicting frequent data that may be frequently accessed within a second preset time period in the future based on the periodicity and correlation of the target large model data access, and loading the frequent data into a cache, wherein the frequent access indicates that the access frequency exceeds a preset threshold; The real-time reasoning tasks of the target large model are analyzed by the artificial intelligence model to obtain the predicted access frequency of the reasoning tasks of the target large model to the data in the cache within a third preset time period in the future, and the data in the cache is updated based on the predicted access frequency.

8. A large-scale hardware and software integrated machine tuning device for rapid deployment, characterized in that: include: a first tuning module, configured to obtain hardware information of each hardware module in the large-model hardware and software integrated machine in response to a user power-on operation on the large-model hardware and software integrated machine, and generate a first deployment plan for the large-model hardware and software integrated machine based on the hardware information of each hardware module and preset general model requirements through an artificial intelligence model, wherein the hardware information includes the hardware model of the hardware module and transmission information between the hardware modules, and the first deployment plan includes a model quantization plan determined according to the hardware model and a model data transmission plan determined according to the transmission information; A second tuning module is configured to respond to the user's business requirement configuration for the large-model hardware and software appliance in the visual interface, determine a target large model from a plurality of pre-trained large models built into the large-model hardware and software appliance based on the business requirement configuration using the artificial intelligence model, and adjust the first deployment plan based on the hardware conditions required by the target large model to obtain a second deployment plan for the large-model hardware and software appliance; A deployment module is used to deploy the target large model in the large model hardware and software integrated machine based on the second deployment solution.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, wherein the memory is used to store computer instructions that can be executed on the processor, and the processor is used to implement the steps of the method according to any one of claims 1 to 7 when executing the computer instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Model deployment method and device, storage medium and electronic equipment

    CN119883295A

  • Task scheduling method for deep learning service, and related apparatus

    WO2023050712A1