Large-model soft and hard all-in-one machine adjusting and optimizing method and device oriented to rapid deployment

Through artificial intelligence models generation and adjustment of the deployment solution of large-model software and hardware all-in-one machines, the problems of long deployment and low resource utilization are solved, rapid deployment and efficient resource management are achieved, adapting to the needs of multiple industries, and overall performance and efficiency are improved.

CN120255906AActive Publication Date: 2025-07-04ZHEJIANG HUATIE EMERGENCY EQUIP SCI & TECH CO LTD +1

Patent Information

Application Number
CN202510757041.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-04
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

In the prior art, the deployment process of large-model hardware and software all-in-one machines takes a long time, has low deployment efficiency, extensive resource management, serious fragmentation of video and memory, and low hardware resource utilization rate.

Method used

Hardware information and business requirements are obtained through artificial intelligence models, initial deployment plans are generated, and adjusted according to the hardware conditions of the target big model. Finally, the target big model is deployed in the big model software and hardware all-in-one machine, including hardware topology map generation, resource optimization and the use of visual configuration panels.

Benefits of technology

It realizes reasonable resource allocation and flexible optimization of large-model software and hardware all-in-one machines without relying on specific model information, improves deployment efficiency, shortens deployment time, adapts to the needs of different industries, and improves overall performance and resource utilization through AI monitoring and optimization of hardware performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255906A_ABST
    Figure CN120255906A_ABST
Patent Text Reader

Abstract

The invention relates to a large model soft and hard all-in-one machine adjusting and optimizing method and device oriented to rapid deployment. The overall deployment efficiency of a large model soft and hard all-in-one machine can be improved. The method comprises the steps of obtaining hardware information of each hardware module in response to a power-on operation of a user on the large-model software and hardware all-in-one machine, and generating a first deployment scheme through an artificial intelligence model according to the hardware information of each hardware module and a preset general model requirement; in response to service demand configuration of a user on the large-model software and hardware all-in-one machine in the visual interface, determining a target large model in a plurality of pre-trained large models built in the large-model software and hardware all-in-one machine through the artificial intelligence model according to the service demand configuration, and adjusting the first deployment scheme according to hardware conditions required by the target large model, obtaining a second deployment scheme of the large-model soft and hard all-in-one machine; and based on the second deployment scheme, deploying the target large model in the large model soft and hard all-in-one machine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of large models, and specifically relates to a tuning method and device for a large model software and hardware integrated machine for rapid deployment. Background Art

[0002] A large model software and hardware integrated machine refers to a product that integrates the hardware and software related to large model applications in a device or system. In terms of hardware, it includes high-performance processors, large-capacity memories, high-speed storage devices, and chips specifically used to accelerate large model operations, etc., to provide the computing power and data storage capabilities required for large model operation. At the software level, it covers the large model itself and related operating systems, drivers, model call interfaces, management and optimization tools, etc., to ensure that the large model can run stably and efficiently, and facilitate user interaction and use.

[0003] In related technologies, the deployment of large models requires manual configuration of drivers, model loading, and network environments, which takes a long time, such as possibly exceeding 48 hours, and the deployment efficiency is low. Summary of the Invention

[0004] To overcome the problems existing in related technologies, embodiments of the present disclosure provide a tuning method and device for a large model software and hardware integrated machine for rapid deployment to solve the defects in related technologies.

[0005] According to the first aspect of the embodiments of the present disclosure, a tuning method for a large model software and hardware integrated machine for rapid deployment is provided, including: In response to the user's power-on operation of the large model software and hardware integrated machine, obtain the hardware information of each hardware module in the large model software and hardware integrated machine, and generate a first deployment plan for the large model software and hardware integrated machine through an artificial intelligence model according to the hardware information of each hardware module and the preset general model requirements, where the hardware information includes the hardware models of the hardware modules and the transmission information between the hardware modules, and the first deployment plan includes a model quantization plan determined according to the hardware models and a model data transmission plan determined according to the transmission information; In response to the user's configuration of the business requirements of the large model software and hardware integrated machine in the visualization interface, determine a target large model among multiple pre-trained large models built in the large model software and hardware integrated machine through the artificial intelligence model according to the business requirement configuration, and adjust the first deployment plan according to the hardware conditions required by the target large model to obtain a second deployment plan for the large model software and hardware integrated machine; Based on the second deployment plan, deploy the target large model in the large model software and hardware integrated machine.

[0006] In one embodiment, the artificial intelligence model generates the first deployment plan of the large model software and hardware integrated machine according to the hardware information of each hardware module and the preset general model requirements, including: The artificial intelligence model generates a hardware topology diagram of the large model software and hardware integrated machine according to the hardware information of each hardware module. Wherein, a node in the hardware topology diagram represents a hardware module in the large model software and hardware integrated machine, and the node is marked with the hardware model of the corresponding hardware module. The connection relationship between the nodes in the hardware topology diagram is used to represent the physical connection link between the corresponding hardware modules, and the edge is marked with the transmission information of the corresponding physical connection link; The artificial intelligence model generates the first deployment plan of the large model software and hardware integrated machine according to the hardware topology diagram and the preset general model requirements.

[0007] In one embodiment, deploying the target large model in the large model software and hardware integrated machine based on the second deployment plan includes: Display the second deployment plan on the large model software and hardware integrated machine; In response to the user's adjustment operation on the second deployment plan, the artificial intelligence model determines whether the hardware conditions of the large model software and hardware integrated machine can implement the target adjustment plan corresponding to the adjustment operation for the second deployment plan according to the hardware information of each hardware module. And when the hardware conditions of the large model software and hardware integrated machine cannot implement the target adjustment plan, the artificial intelligence model generates a recommended adjustment plan for the second deployment plan according to the adjustment operation and the hardware information of each hardware module; In response to the user's confirmation operation on the recommended adjustment plan, the target large model is deployed in the large model software and hardware integrated machine based on the recommended adjustment plan.

[0008] In one embodiment, the large model software and hardware integrated machine is built in with preset industry knowledge graphs corresponding to different industries, and further includes: Display a visual configuration panel for the target intelligent agent, where the target intelligent agent is an intelligent agent associated with the target large model; In response to a natural language input operation in the visual configuration panel, obtain the business requirements for the target intelligent agent corresponding to the natural language operation; Obtain the target industry knowledge graph corresponding to the business requirement in the preset industry knowledge graph through the artificial intelligence model, and based on the business requirement, recommend the optimal combination of working nodes that can achieve the business requirement in the target industry knowledge graph. Based on the optimal combination of working nodes, generate the agent workflow corresponding to the target agent, where the target industry knowledge graph includes the agent working nodes required by the industry corresponding to the business requirement and the execution order between the agent working nodes.

[0009] In one embodiment, the obtaining the target industry knowledge graph corresponding to the business requirement in the preset industry knowledge graph through the artificial intelligence model includes: Identify industry keywords based on the business requirement through the artificial intelligence model, and search in the preset industry knowledge graph based on the industry keywords; In the case where the preset industry knowledge graph corresponding to the business requirement is not found, understand the industry characteristics of the business requirement through the artificial intelligence model to obtain the demand industry characteristics corresponding to the business requirement, and based on the demand industry characteristics, match similar industries in all industries corresponding to the preset industry knowledge graph, and determine the preset industry knowledge graph corresponding to the similar industry as the target industry knowledge graph corresponding to the business requirement.

[0010] In one embodiment, it further includes: When it is monitored that the video memory occupancy of the large model software and hardware integrated machine exceeds the preset threshold, predict the time point when the video memory resources of the large model software and hardware integrated machine are exhausted through the artificial intelligence model according to the video memory usage trend of the large model software and hardware integrated machine and the inference task characteristics of the target large model; When the duration from the time point is less than the first preset duration, analyze the inference task of the target large model through the artificial intelligence model to obtain the non-critical weight parameters of the target large model, and perform quantization compression on the non-critical weight parameters, and / or analyze the computational graph structure of the target large model through the artificial intelligence model to obtain the target computational nodes to be merged, and merge the target computational nodes.

[0011] In one embodiment, it further includes: Obtain the data access information in the historical inference tasks of the target large model, where the data access information includes the data access time and the data access frequency; Analyze the data access information through the artificial intelligence model to capture the periodicity and correlation of the target large model's data access. According to the periodicity and correlation of the target large model's data access, predict the frequently accessed data that may be frequently accessed within a second preset duration in the future, and load the frequently accessed data into the cache, where the frequent access indicates that the access frequency exceeds a preset threshold; Analyze the real-time inference task of the target large model through the artificial intelligence model to obtain the predicted access frequency of the inference task of the target large model to the data in the cache within a third preset duration in the future, and update the data in the cache based on the predicted access frequency.

[0012] According to a second aspect of the embodiments of the present disclosure, there is provided an optimization device for a large model software and hardware integrated machine for rapid deployment, including: A first optimization module, configured to, in response to a user's power-on operation of the large model software and hardware integrated machine, obtain the hardware information of each hardware module in the large model software and hardware integrated machine, and generate a first deployment plan for the large model software and hardware integrated machine through an artificial intelligence model according to the hardware information of each hardware module and preset general model requirements, where the hardware information includes the hardware model of the hardware module and the transmission information between the hardware modules, and the first deployment plan includes a model quantization plan determined according to the hardware model and a model data transmission plan determined according to the transmission information; A second optimization module, configured to, in response to the user's configuration of the service requirements of the large model software and hardware integrated machine in the visualization interface, determine a target large model among multiple pre-trained large models built in the large model software and hardware integrated machine through the artificial intelligence model according to the service requirement configuration, and adjust the first deployment plan according to the hardware conditions required by the target large model to obtain a second deployment plan for the large model software and hardware integrated machine; A deployment module, configured to deploy the target large model in the large model software and hardware integrated machine based on the second deployment plan.

[0013] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, where the electronic device includes a memory and a processor, the memory is used to store computer instructions that can be run on the processor, and the processor is used to implement the method according to any one of the first aspects when executing the computer instructions.

[0014] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and the program implements the method according to any one of the first aspects when executed by a processor.

[0015] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method described in any one of the first aspect.

[0016] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: For the method for optimizing the large model software and hardware integrated machine for rapid deployment provided by the embodiments of the present disclosure, after the large model software and hardware integrated machine is powered on, hardware detection can be automatically performed to obtain the hardware information of each hardware module in the large model software and hardware integrated machine. Then, an artificial intelligence model generates a first deployment plan according to the hardware information and the general model requirements. After that, the artificial intelligence model can adjust the first deployment plan according to the hardware conditions required by the target large model in the actual business to obtain a second deployment plan. Finally, model deployment is performed based on the second deployment plan. Thus, through the dynamic adjustment mechanism based on hardware information, the large model software and hardware integrated machine can first perform reasonable resource allocation without relying on specific model information, and then perform flexible optimization according to the requirements of the actual model, thereby improving the overall deployment efficiency of the large model software and hardware integrated machine. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.

[0018] Figure 1 is a flowchart of a method for optimizing a large model software and hardware integrated machine for rapid deployment shown in an exemplary embodiment of the present disclosure; Figure 2 is a block diagram of a device for optimizing a large model software and hardware integrated machine for rapid deployment shown in an exemplary embodiment of the present disclosure; Figure 3 is a block diagram of an electronic device shown in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0020] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms "a", "the", and "said" used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.

[0021] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information.

[0022] The large model software and hardware integrated machine refers to a product that integrates the hardware and software related to large model applications in a device or system. In terms of hardware, it includes high-performance processors, large-capacity memories, high-speed storage devices, and chips dedicated to accelerating large model operations, etc., to provide the computing power and data storage capacity required for the operation of large models. At the software level, it covers the large model itself and related operating systems, drivers, model call interfaces, management and optimization tools, etc., to ensure that the large model can run stably and efficiently and facilitate user interaction and use.

[0023] In the related art, the deployment of large models requires manual configuration of drivers, model loading, and network environments, which takes a long time, such as possibly exceeding 48 hours, and the deployment efficiency is low. In addition, resource management is extensive, mainly reflected in serious fragmentation of video memory and low utilization rate of hardware resources.

[0024] Based on this, at least one embodiment of this disclosure provides a tuning method for a large model software and hardware integrated machine for rapid deployment. Please refer to Figure 1 Please refer to Figure 1 which shows the flow of the method, including steps S101 to S103.

[0025] In step S101, in response to the user's power-on operation of the large model software and hardware integrated machine, obtain the hardware information of each hardware module in the large model software and hardware integrated machine, and generate a first deployment plan for the large model software and hardware integrated machine through an artificial intelligence model according to the hardware information of each hardware module and the preset general model requirements. Among them, the hardware information includes the hardware models of the hardware modules and the transmission information between the hardware modules, and the first deployment plan includes a model quantization plan determined according to the hardware models and a model data transmission plan determined according to the transmission information.

[0026] It should be understood that after power-on, the large model software and hardware integrated machine can perform preliminary resource allocation based on common model operation requirements. Since most large models have some common requirements during operation, such as the requirements for GPU (Graphics Processing Unit) computing resources, storage read / write bandwidth, and network transmission bandwidth. The large model software and hardware integrated machine can use an artificial intelligence model to allocate reasonable data transmission paths and bandwidth resources for these general requirements based on the hardware information of each hardware module to meet the basic operation conditions of most models. This preliminary allocation can ensure that the large model software and hardware integrated machine has a certain degree of generality and flexibility, providing a good basic environment for the subsequent specific deployment and operation of the model, without the need for special customization for specific models.

[0027] Exemplarily, the model quantization scheme can be determined based on the hardware model. For example, after identifying the Kunlun Core P800 GPU, the preset quantization strategy is automatically matched.

[0028] In step S102, in response to the user's configuration of the business requirements for the large model software and hardware integrated machine in the visualization interface, the artificial intelligence model determines the target large model among multiple pre-trained large models built in the large model software and hardware integrated machine according to the business requirement configuration, and adjusts the first deployment plan according to the hardware conditions required by the target large model to obtain the second deployment plan of the large model software and hardware integrated machine.

[0029] Exemplarily, in the process of the artificial intelligence model determining the target large model among multiple pre-trained large models built in the large model software and hardware integrated machine according to the business requirement configuration, the artificial intelligence model can combine the hardware resources available for model operation in the large model software and hardware integrated machine, such as the number, performance, and storage capacity of GPUs. Specifically, the artificial intelligence model can screen out the target large model that matches the hardware resources according to this hardware resource information. For example, if there are multiple high-performance GPUs in the hardware, complex large models that can make full use of the parallel computing power of GPUs and meet the business requirements can be recommended; if the GPU performance is average, relatively lightweight large models with lower hardware requirements but still able to meet the business requirements are recommended.

[0030] After the model to be deployed is determined, i.e., the target large model is determined, the first deployment plan can be further dynamically adjusted according to the characteristics of the target large model. The large model hardware-software integrated machine will determine the hardware conditions required by the target large model based on information such as the size, computational complexity, and data flow pattern of the target large model, and then optimize and fine-tune the previously allocated data transmission paths and bandwidth resources, etc. For example, if the target large model is an ultra-large-scale pre-trained language model that requires a large amount of video memory to store model parameters and will generate frequent video memory accesses and data exchanges during the inference process, then the allocation ratio of GPU video memory can be appropriately increased, and the data transmission path between the GPU and the video memory can be optimized to improve the performance of model operation.

[0031] In step S103, based on the second deployment plan, deploy the target large model in the large model hardware-software integrated machine.

[0032] Thus, through the dynamic adjustment mechanism based on hardware information, the large model hardware-software integrated machine can first perform reasonable resource allocation without relying on specific model information, and then flexibly optimize according to the actual model requirements, thereby improving the overall deployment efficiency of the large model hardware-software integrated machine.

[0033] For ease of understanding, the above steps will be described in detail with examples below.

[0034] In one embodiment, in step S101, the artificial intelligence model generates the first deployment plan of the large model hardware-software integrated machine based on the hardware information of each hardware module and the preset general model requirements, including: generating the hardware topology diagram of the large model hardware-software integrated machine through the artificial intelligence model according to the hardware information of each hardware module, where a node in the hardware topology diagram represents a hardware module in the large model hardware-software integrated machine, and the node is marked with the hardware model of the corresponding hardware module, the connection relationship between the nodes in the hardware topology diagram is used to represent the physical connection link between the corresponding hardware modules, and the edge is marked with the transmission information of the corresponding physical connection link; generating the first deployment plan of the large model hardware-software integrated machine through the artificial intelligence model according to the hardware topology diagram and the preset general model requirements.

[0035] It should be understood that the hardware of the large model hardware-software integrated machine usually includes multiple key parts: Computing core: such as a GPU cluster, which is used to parallel process a large number of computing tasks of the model, such as matrix operations, etc.; Storage module: includes cache, SSD (Solid State Drive), etc., which is used to store data such as model parameters and intermediate calculation results; Network interface: Taking a high-speed Ethernet network card or an InfiniBand network card as an example, it is responsible for data transmission between different computing nodes or with external systems; Motherboard and backplane: Used to connect various hardware components and build a physical communication link between the hardware.

[0036] After power-on, the large model software and hardware integrated machine can start the intelligent hardware topology perception function. Components such as dedicated perception chips or firmware embedded at the hardware bottom layer can send recognition signals to the topology perception module of the large model software and hardware integrated machine, reporting information such as their own model numbers and performance parameters. After receiving these signals, the topology perception module uses a preset topology recognition algorithm to automatically draw a hardware topology diagram. For example, it is recognized that there are 4 high-performance GPUs in the system, which are respectively connected to specific slots on the motherboard through the PCIe (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard) bus, and these GPUs are connected to each other through NVLink high-speed interconnect links; at the same time, multiple SSD hard disks in the storage module are connected to the motherboard through the SATA (Serial ATA, Serial Advanced Technology Attachment) bus or NVMe interface, and the network interface is a dual-channel 100Gbps Ethernet network card, which is connected to the network slot on the motherboard. After that, based on this information, a hardware topology diagram is automatically drawn.

[0037] Exemplarily, the hardware topology diagram can be a logical schematic diagram, which can be used to clearly show the connection relationship and data transmission path between each hardware component. In this diagram, each node can represent a hardware component, such as a GPU, an SSD hard disk, a network interface, etc., and the corresponding hardware model can be marked beside the node. The connection lines between the nodes can represent the physical connection links between them, and key parameters such as the connection bus type and bandwidth can be marked beside the connection lines. For example, the GPU nodes are connected to each other through NVLink links, and the marked bandwidth is 50GB / s; the SSD hard disk nodes are connected to the motherboard through the NVMe interface, and the bandwidth is 4GB / s; the network interface nodes are connected to the external network through 100Gbps Ethernet links.

[0038] Thus, the large model software and hardware integrated machine can have the intelligent hardware topology perception function, automatically draw a hardware topology diagram after power-on, identify the connection relationship and data transmission path between the hardware, and thus, based on the hardware topology diagram and the general model requirements, dynamically adjust the data transmission path and resource allocation, such as automatically allocating more bandwidth resources to high-load modules, thereby improving the deployment efficiency and overall performance of the large model software and hardware integrated machine.

[0039] In one embodiment, in step S103, based on the second deployment plan, deploying the target large model in the large model software and hardware integrated machine includes: displaying the second deployment plan on the large model software and hardware integrated machine; in response to the user's adjustment operation on the second deployment plan, the artificial intelligence model determines whether the hardware conditions of the large model software and hardware integrated machine can implement the target adjustment plan corresponding to the adjustment operation for the second deployment plan according to the hardware information of each hardware module, and in the case where the hardware conditions of the large model software and hardware integrated machine cannot implement the target adjustment plan, the artificial intelligence model generates a recommended adjustment plan for the second deployment plan according to the adjustment operation and the hardware information of each hardware module; in response to the user's confirmation operation on the recommended adjustment plan, deploying the target large model in the large model software and hardware integrated machine based on the recommended adjustment plan.

[0040] Exemplarily, the user can adjust the second deployment plan based on the visual operation interface. The visual operation interface can display the deployment parameters corresponding to the second deployment plan, and the user can adjust the second deployment plan by re-entering the deployment parameters. Then, the artificial intelligence model can automatically identify whether the adjusted plan by the user matches the hardware conditions of the large model software and hardware integrated machine. If it matches, the user's adjustment operation can be directly applied. If it does not match, the artificial intelligence model can generate a recommended adjustment plan that matches the hardware conditions of the large model software and hardware integrated machine. Thus, the user can first visually adjust the second deployment plan, and then the artificial intelligence model can automatically adjust and optimize the second deployment plan, which can not only meet the user's needs but also improve the deployment efficiency and applicability of the large model software and hardware integrated machine.

[0041] In one embodiment, the large model software and hardware integrated machine is built-in with preset industry knowledge graphs corresponding to different industries, and can also display a visual configuration panel for the target intelligent agent, where the target intelligent agent is the intelligent agent associated with the target large model; in response to the natural language input operation in the visual configuration panel, obtaining the business requirements for the target intelligent agent corresponding to the natural language operation; the artificial intelligence model obtains the target industry knowledge graph corresponding to the business requirements in the preset industry knowledge graph and recommends the optimal combination of working nodes that can implement the business requirements in the target industry knowledge graph based on the business requirements, and generates the intelligent agent workflow corresponding to the target intelligent agent based on the optimal combination of working nodes, where the target industry knowledge graph includes the intelligent agent working nodes required by the industry corresponding to the business requirements and the execution order between the intelligent agent working nodes.

[0042] It should be understood that in the related art, the pre-trained model and the tool chain are not optimized for vertical scenarios, and the cost of secondary development is high. Moreover, there is a lack of a visualization tool chain, making it difficult for non-technical personnel to complete model fine-tuning and process orchestration. However, the present disclosure can provide a visualization configuration panel where users can perform natural language input operations. Then, through an artificial intelligence model, automatic optimization for vertical scenarios can be carried out, enabling even non-technical personnel to quickly perform model fine-tuning and process orchestration in vertical scenarios.

[0043] Exemplarily, taking the medical field as an example, the preset industry knowledge graph may include the following core information: Disease entities: disease names (such as diabetes), symptoms (such as polydipsia, polyuria), complications (such as retinopathy); Diagnosis and treatment process entities: examination items (such as blood routine, imaging examination), diagnosis methods (such as imaging diagnosis, expert consultation), treatment means (such as drugs, surgery).

[0044] Rules and regulations entities: medical compliance review, clinical pathway guidelines (such as the standard process for diabetes management).

[0045] Technical tool entities: available models and tools (such as multi-modal diagnosis models, natural language understanding models, medical record structured parsing engines).

[0046] Causal relationships: symptom → disease association, treatment → efficacy relationship (such as insulin injection → blood glucose control).

[0047] For example, the user inputs the following content through natural language input operations in the visualization configuration panel: "Build a medical consultation intelligent agent that can automatically analyze patient medical records, give diagnosis suggestions, and ensure compliance with medical regulations." Then, semantic parsing and intention extraction are performed on this content. For example, through an NLP (Natural Language Processing) model, keywords such as "medical record analysis", "diagnosis suggestions", and "medical regulations" are parsed, and then through intention recognition, it is determined that the user's goal is to build a compliant automated diagnosis process.

[0048] After that, knowledge graph retrieval and matching can be performed. First, entity matching is carried out to retrieve technical tool entities related to medical record analysis, such as "medical record structured parsing"; entities related to diagnostic suggestions are matched, such as "multi-modal diagnostic model"; association rules and specification entities are associated, such as "medical compliance review". Then, according to the causal relationship in the knowledge graph, the following optimal working node combination is determined: medical record structured parsing → medical compliance review → multi-modal diagnosis → medical compliance review → result output. Finally, based on this optimal working node combination, the following intelligent agent workflow can be obtained: perform medical record structured parsing on the input content → perform medical compliance review on the parsed content → perform multi-modal diagnosis on the input content after passing the medical compliance review → perform medical compliance review on the diagnosis result → generate a diagnostic report → output the diagnostic report.

[0049] Thus, it is possible to avoid users from manually configuring complex processes, which is especially suitable for practitioners without a technical background, compress the originally time-consuming process design from several days to minutes, and can be automatically iterated with the update of the knowledge graph, thereby improving the deployment efficiency of the large model software and hardware integrated machine.

[0050] In one embodiment, obtaining the target industry knowledge graph corresponding to the business requirement in the preset industry knowledge graph by the artificial intelligence model includes: identifying industry keywords based on the business requirement by the artificial intelligence model, and searching in the preset industry knowledge graph based on the industry keywords; in the case where the preset industry knowledge graph corresponding to the business requirement is not found, understanding the industry characteristics of the business requirement by the artificial intelligence model to obtain the demand industry characteristics corresponding to the business requirement and based on the demand industry characteristics, matching similar industries in all industries corresponding to the preset industry knowledge graph, and determining the preset industry knowledge graph corresponding to the similar industry as the target industry knowledge graph corresponding to the business requirement.

[0051] Exemplarily, continuing with the above example, for instance, the business requirement input by the user is: "Build a medical consultation intelligent agent that can automatically analyze patient medical records, give diagnostic suggestions, and ensure compliance with medical regulations", the industry keyword can be obtained as "medical". Then, search in the preset industry knowledge graph based on this industry keyword. It should be understood that each preset industry knowledge graph is pre-associated with a corresponding industry keyword. Thus, in the case where the corresponding preset industry knowledge graph is found based on the industry keyword, the subsequent process can be directly executed based on this preset industry knowledge graph. In the case where the corresponding preset industry knowledge graph is not found, the artificial intelligence model can be used to understand the industry characteristics of the business requirement, thereby obtaining a similar industry corresponding to the business requirement, and then executing the subsequent process based on this similar industry.

[0052] Thus, the deployment efficiency of the large model software and hardware integrated machine in different industries can be improved, and through intelligent identification of industry commonalities, it can dynamically adapt to industry differences, providing technical support for rapid implementation in multiple fields.

[0053] In one embodiment, when it is detected that the video memory occupancy of the large model software and hardware integrated machine exceeds a preset threshold, the artificial intelligence model can predict the time point when the video memory resources of the large model software and hardware integrated machine are exhausted according to the video memory usage trend of the large model software and hardware integrated machine and the inference task characteristics of the target large model; when the duration from the time point is less than a first preset duration, the artificial intelligence model analyzes the inference task of the target large model to obtain the non-critical weight parameters of the target large model, and quantizes and compresses the non-critical weight parameters, or analyzes the computational graph structure of the target large model through the artificial intelligence model to obtain the target computational nodes to be merged, and merges the target computational nodes.

[0054] It should be understood that when using a large model software and hardware integrated machine to train a natural language processing model (such as a model based on the Transformer architecture), the number of model parameters is huge, and the video memory occupancy during training changes continuously and is difficult to predict. As training progresses, the video memory occupancy gradually increases. If not adjusted in time, there may be a situation where the video memory is insufficient, resulting in training stuttering or even interruption. Or, when using a large model software and hardware integrated machine to perform inference tasks for an image recognition model (such as ResNet-50), when there are multiple high-resolution image requests to be processed simultaneously, the video memory demand will increase rapidly. Especially when performing real-time video stream image recognition, ensuring low latency and high throughput, the efficient utilization of video memory is crucial.

[0055] Taking natural language processing as an example, the large model software and hardware all-in-one machine can use the built-in hardware performance monitoring module to collect key performance indicators such as video memory occupancy, video memory usage trend, and GPU computing task queue length every fixed time (such as 1 second), and input this data into an artificial intelligence model for prediction. Among them, the artificial intelligence model can be trained based on historical performance data during a large number of previous training processes, and can learn the change rules of video memory occupancy and the relationship with training task characteristics (such as the number of model layers, batch size, optimization algorithm, etc.). When it is monitored that the video memory occupancy is close to the threshold (such as 90%), the artificial intelligence model can predict the time point when the video memory is about to run out according to the current training task characteristics and video memory usage trend. After that, the artificial intelligence model can adopt the dynamic quantization compression (Dynamic Quantization) strategy to quantize the non-critical weight parameters of the target large model from 32-bit floating-point numbers (FP32) to 16-bit floating-point numbers (FP16), thereby reducing the video memory occupancy. At the same time, the large model software and hardware all-in-one machine can analyze the computational graph structure of the target large model, find out the computational nodes that can be merged or optimized, and adopt a more efficient computational graph structure to further release the video memory.

[0056] Thus, it is possible to achieve AI (Artificial Intelligence)-based hardware performance prediction and optimization, use the AI model to monitor and predict the hardware performance in real time, and identify performance bottlenecks in advance. When it is predicted that the video memory is about to run out, the video memory allocation strategy of the model is automatically adjusted, and then the video memory is released by compressing the model parameters and / or adopting a more efficient computational graph structure to avoid system jamming.

[0057] In one embodiment, it is also possible to obtain the data access information in the historical inference tasks of the target large model, where the data access information includes data access time and data access frequency; analyze the data access information through the artificial intelligence model to capture the periodicity and correlation of the data access of the target large model, and predict the frequently accessed data that may be frequently accessed within a second preset duration in the future according to the periodicity and correlation of the data access of the target large model, and load the frequently accessed data into the cache, where the frequent access means that the access frequency exceeds a preset threshold; analyze the real-time inference task of the target large model through the artificial intelligence model to obtain the predicted access frequency of the inference task of the target large model to the data in the cache within a third preset duration in the future, and update the data in the cache based on the predicted access frequency.

[0058] It should be understood that, for example, when using a large model software and hardware all-in-one machine for text generation tasks, the model needs to frequently access data such as vocabulary tables and word vectors. Information such as the access time and access frequency of these data during the inference process can be collected through an artificial intelligence model. Then, an artificial intelligence model, such as a time series analysis model based on Transformer, is used to analyze and learn the data characteristics, capturing the periodicity and correlation of data access. For example, it is found that certain commonly used words are frequently accessed in specific types of text generation tasks. Then, based on the periodicity and correlation of data access, the artificial intelligence model predicts the vocabulary table data and word vectors that may be frequently accessed and loads them into the cache. For example, for the news headline generation task, it is predicted that the access frequency of words related to news hotspots will increase, and the large model software and hardware all-in-one machine loads the data corresponding to these words into the cache, reducing the data loading latency during the inference process and improving the inference speed.

[0059] It should also be understood that traditional cache eviction policies may have problems with low cache hit rates when dealing with complex access patterns. However, the embodiments of the present disclosure can dynamically adjust the cache eviction policy according to the predicted access frequency. If it is predicted that the access frequency of a certain word vector will decrease in subsequent tasks, even if it has been recently accessed, it will be evicted from the cache in advance to make room for new data that may be frequently accessed.

[0060] Thus, by analyzing and predicting the access patterns of cache data through an AI model, data that may be frequently accessed is loaded into the cache in advance. At the same time, the cache eviction policy is dynamically adjusted according to the prediction results to ensure the effective utilization of the cache space and improve the data access efficiency.

[0061] Through any of the above embodiments, the present disclosure can provide an out-of-the-box large model software and hardware all-in-one machine system. Through software and hardware collaborative optimization technology, an automated deployment process, and a scenario-based tool chain, enterprise users can complete deployment within 30 minutes, perform zero-code operations, and achieve efficient resource utilization.

[0062] According to the second aspect of the embodiments of the present disclosure, there is provided a tuning device for a large model software and hardware all-in-one machine for rapid deployment. Please refer to Figure 2 , the tuning device 200 for a large model software and hardware all-in-one machine for rapid deployment includes: The first tuning module 201 is configured to, in response to the user's power-on operation of the large model software and hardware integrated machine, obtain the hardware information of each hardware module in the large model software and hardware integrated machine, and generate a first deployment plan for the large model software and hardware integrated machine through an artificial intelligence model based on the hardware information of each hardware module and a preset general model requirement. Wherein, the hardware information includes the hardware model of the hardware module and the transmission information between the hardware modules, and the first deployment plan includes a model quantization plan determined according to the hardware model and a model data transmission plan determined according to the transmission information; The second tuning module 202 is configured to, in response to the user's configuration of the service requirements of the large model software and hardware integrated machine in the visualization interface, determine a target large model from multiple pre-trained large models built in the large model software and hardware integrated machine through the artificial intelligence model according to the service requirement configuration, and adjust the first deployment plan according to the hardware conditions required by the target large model to obtain a second deployment plan for the large model software and hardware integrated machine; The deployment module 203 is configured to deploy the target large model in the large model software and hardware integrated machine based on the second deployment plan.

[0063] In one embodiment, the first tuning module 201 is configured to: Generate a hardware topology diagram of the large model software and hardware integrated machine through the artificial intelligence model according to the hardware information of each hardware module. Wherein, a node in the hardware topology diagram represents a hardware module in the large model software and hardware integrated machine, and the node is marked with the hardware model of the corresponding hardware module. The connection relationship between the nodes in the hardware topology diagram is used to represent the physical connection link between the corresponding hardware modules, and the edge is marked with the transmission information of the corresponding physical connection link; Generate a first deployment plan for the large model software and hardware integrated machine through the artificial intelligence model according to the hardware topology diagram and a preset general model requirement.

[0064] In one embodiment, the deployment module 203 is configured to: Display the second deployment plan on the large model software and hardware integrated machine; In response to the user's adjustment operation on the second deployment plan, determine whether the hardware conditions of the large model software and hardware integrated machine can implement the target adjustment plan corresponding to the adjustment operation for the second deployment plan through the artificial intelligence model according to the hardware information of each hardware module, and in the case where the hardware conditions of the large model software and hardware integrated machine cannot implement the target adjustment plan, generate a recommended adjustment plan for the second deployment plan through the artificial intelligence model according to the adjustment operation and the hardware information of each hardware module; In response to the user's confirmation operation on the recommended adjustment plan, based on the recommended adjustment plan, deploy the target large model in the large model software and hardware integrated machine.

[0065] In one embodiment, the large model software and hardware integrated machine is built with preset industry knowledge graphs corresponding to different industries. The optimization device 200 for the large model software and hardware integrated machine for fast deployment further includes: A display module, configured to display a visual configuration panel for a target intelligent agent, where the target intelligent agent is an intelligent agent associated with the target large model; A first acquisition module, configured to, in response to a natural language input operation in the visual configuration panel, acquire the business requirements for the target intelligent agent corresponding to the natural language operation; A workflow module, configured to, through the artificial intelligence model, acquire the target industry knowledge graph corresponding to the business requirements in the preset industry knowledge graph, and based on the business requirements, recommend an optimal combination of working nodes that can implement the business requirements in the target industry knowledge graph, and generate an intelligent agent workflow corresponding to the target intelligent agent based on the optimal combination of working nodes, where the target industry knowledge graph includes the intelligent agent working nodes required by the industry corresponding to the business requirements and the execution order between the intelligent agent working nodes.

[0066] In one embodiment, the workflow module is configured to: Identify industry keywords based on the business requirements through the artificial intelligence model, and search in the preset industry knowledge graph based on the industry keywords; In the case where the preset industry knowledge graph corresponding to the business requirements is not found, understand the industry characteristics of the business requirements through the artificial intelligence model to obtain the demand industry characteristics corresponding to the business requirements, and based on the demand industry characteristics, match similar industries in all industries corresponding to the preset industry knowledge graph, and determine the preset industry knowledge graph corresponding to the similar industries as the target industry knowledge graph corresponding to the business requirements.

[0067] In one embodiment, the optimization device 200 for the large model software and hardware integrated machine for fast deployment further includes: A first analysis module, configured to, when it is monitored that the video memory occupancy of the large model software and hardware integrated machine exceeds a preset threshold, predict the time point when the video memory resources of the large model software and hardware integrated machine are exhausted through the artificial intelligence model according to the video memory usage trend of the large model software and hardware integrated machine and the inference task characteristics of the target large model; A second analysis module, configured to, when a duration from the time point is less than a first preset duration, analyze an inference task of the target large model through the artificial intelligence model to obtain non-critical weight parameters of the target large model, and perform quantization compression on the non-critical weight parameters, and / or analyze a computational graph structure of the target large model through the artificial intelligence model to obtain target computational nodes to be merged, and merge the target computational nodes.

[0068] In one embodiment, the large model software and hardware integrated machine tuning device 200 for fast deployment further includes: A second acquisition module, configured to acquire data access information in a historical inference task of the target large model, where the data access information includes a data access time and a data access frequency; A third analysis module, configured to analyze the data access information through the artificial intelligence model to capture data periodicity and correlation of data access of the target large model, and predict frequently accessed data that may be frequently accessed within a second preset duration in the future according to the periodicity and correlation of data access of the target large model, and load the frequently accessed data into a cache, where the frequent access means that the access frequency exceeds a preset threshold; A fourth analysis module, configured to analyze a real-time inference task of the target large model through the artificial intelligence model to obtain a predicted access frequency of the inference task of the target large model to data in the cache within a third preset duration in the future, and update the data in the cache based on the predicted access frequency.

[0069] Regarding the device in the above embodiment, the specific manners in which each module performs operations have been described in detail in the embodiment of the method in the first aspect, and will not be elaborated herein.

[0070] According to a third aspect of the embodiments of the present disclosure, please refer to Figure 3 , which exemplarily shows a block diagram of an electronic device. The electronic device 700 may include: a processor 701, a memory 702. The electronic device 700 may further include one or more of a multimedia component 703, an input / output (I / O) interface 704, and a communication component 705.

[0071] Among them, the processor 701 is used to control the overall operation of the electronic device 700 to complete all or part of the steps in any of the above methods. The memory 702 is used to store various types of data to support the operation of the electronic device 700. These data may include, for example, instructions for any application or method operating on the electronic device 700, as well as application-related data, such as contact data, messages sent and received, pictures, audio, video, and so on. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The multimedia component 703 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 702 or sent through the communication component 705. The audio component further includes at least one speaker for outputting audio signals. The I / O interface 704 provides an interface between the processor 701 and other interface modules, and the other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 705 is used for wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Accordingly, the communication component 705 may include: a Wi-Fi module, a Bluetooth module, and an NFC module.

[0072] In an exemplary embodiment, the electronic device 700 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, and is used to execute any of the above methods.

[0073] In one embodiment, there is also provided a computer-readable storage medium including program instructions. When the program instructions are executed by a processor, the steps of any of the above methods are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 702 including program instructions, and the above program instructions can be executed by the processor 701 of the electronic device 700 to complete any of the above methods.

[0074] In one embodiment, there is also provided a computer program product. The computer program product includes a computer program that can be executed by a processor. When the computer program is executed by the processor, the steps of any of the above methods are implemented.

[0075] The preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0076] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present disclosure does not separately describe various possible combination methods.

[0077] In addition, any combination can be made between various different embodiments of the present disclosure as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.

Claims

1. A tuning method for a large model software and hardware integrated machine oriented to rapid deployment, characterized in that, Including: In response to the user's power-on operation of the large model software and hardware integrated machine, obtain the hardware information of each hardware module in the large model software and hardware integrated machine, and generate a first deployment plan for the large model software and hardware integrated machine through an artificial intelligence model according to the hardware information of each hardware module and the preset general model requirements. Wherein, the hardware information includes the hardware models of the hardware modules and the transmission information between the hardware modules, and the first deployment plan includes a model quantization plan determined according to the hardware models and a model data transmission plan determined according to the transmission information; In response to the user's configuration of the business requirements of the large model software and hardware integrated machine in the visualization interface, determine a target large model among multiple pre-trained large models built in the large model software and hardware integrated machine through the artificial intelligence model according to the business requirement configuration, and adjust the first deployment plan according to the hardware conditions required by the target large model to obtain a second deployment plan for the large model software and hardware integrated machine; Based on the second deployment plan, deploy the target large model in the large model software and hardware integrated machine.

2. The optimization method for the large model software and hardware integrated machine for rapid deployment according to claim 1, wherein The step of generating the first deployment plan for the large model software and hardware integrated machine through an artificial intelligence model according to the hardware information of each hardware module and the preset general model requirements includes: Generate a hardware topology diagram of the large model software and hardware integrated machine through the artificial intelligence model according to the hardware information of each hardware module. Wherein, a node in the hardware topology diagram represents a hardware module in the large model software and hardware integrated machine, and the node is marked with the hardware model of the corresponding hardware module. The connection relationship between the nodes in the hardware topology diagram is used to represent the physical connection link between the corresponding hardware modules, and the edge is marked with the transmission information of the corresponding physical connection link; Generate the first deployment plan for the large model software and hardware integrated machine through an artificial intelligence model according to the hardware topology diagram and the preset general model requirements.

3. The method for optimizing the software and hardware integrated machine of large models for rapid deployment according to claim 1, wherein The step of deploying the target large model in the large model software and hardware integrated machine based on the second deployment plan includes: Display the second deployment plan on the large model software and hardware integrated machine; In response to the user's adjustment operation on the second deployment plan, determine whether the hardware conditions of the large model software and hardware integrated machine can implement the target adjustment plan corresponding to the adjustment operation for the second deployment plan through the artificial intelligence model according to the hardware information of each hardware module. And in the case where the hardware conditions of the large model software and hardware integrated machine cannot implement the target adjustment plan, generate a recommended adjustment plan for the second deployment plan through the artificial intelligence model according to the adjustment operation and the hardware information of each hardware module; In response to the user's confirmation operation on the recommended adjustment plan, deploy the target large model in the large model software and hardware integrated machine based on the recommended adjustment plan.

4. The optimization method for the large model software and hardware integrated machine for rapid deployment according to any one of claims 1-3, characterized in that The large model software and hardware integrated machine is built in with preset industry knowledge graphs corresponding to different industries, and further includes: Display a visual configuration panel for the target agent, where the target agent is the agent associated with the target large model; In response to a natural language input operation in the visual configuration panel, obtain the business requirements for the target agent corresponding to the natural language operation; Through the artificial intelligence model, obtain the target industry knowledge graph corresponding to the business requirements in the preset industry knowledge graph, and based on the business requirements, recommend the optimal combination of working nodes that can achieve the business requirements in the target industry knowledge graph. Based on the optimal combination of working nodes, generate the agent workflow corresponding to the target agent, where the target industry knowledge graph includes the working nodes required by the industry corresponding to the business requirements and the execution order between the working nodes.

5. The optimization method for the large model software and hardware integrated machine for rapid deployment according to claim 4, wherein The obtaining of the target industry knowledge graph corresponding to the business requirements in the preset industry knowledge graph through the artificial intelligence model includes: Through the artificial intelligence model, identify industry keywords based on the business requirements, and search in the preset industry knowledge graph based on the industry keywords; In the case where the preset industry knowledge graph corresponding to the business requirements is not found, through the artificial intelligence model, understand the industry characteristics of the business requirements, obtain the demand industry characteristics corresponding to the business requirements, and based on the demand industry characteristics, match similar industries in all industries corresponding to the preset industry knowledge graph, and determine the preset industry knowledge graph corresponding to the similar industries as the target industry knowledge graph corresponding to the business requirements.

6. The optimization method for the large model software and hardware integrated machine for rapid deployment according to any one of claims 1-3, characterized in that, It further includes: In the case where it is monitored that the video memory occupancy of the large model software and hardware integrated machine exceeds a preset threshold, through the artificial intelligence model, predict the time point when the video memory resources of the large model software and hardware integrated machine are exhausted according to the video memory usage trend of the large model software and hardware integrated machine and the inference task characteristics of the target large model; In the case where the duration from the time point is less than a first preset duration, through the artificial intelligence model, analyze the inference task of the target large model to obtain the non-critical weight parameters of the target large model, and perform quantization compression on the non-critical weight parameters, and / or, through the artificial intelligence model, analyze the computational graph structure of the target large model to obtain the target computational nodes to be merged, and merge the target computational nodes.

7. The method for optimizing the software and hardware integrated machine of the large model for rapid deployment according to any one of claims 1-3, characterized in that It further includes: Obtain the data access information in the historical inference tasks of the target large model, where the data access information includes the data access time and the data access frequency; Through the artificial intelligence model, analyze the data access information to capture the periodicity and correlation of the data access of the target large model. According to the periodicity and correlation of the data access of the target large model, predict the frequently accessed data that may be frequently accessed within a future second preset duration, and load the frequently accessed data into the cache, where the frequent access means that the access frequency exceeds a preset threshold; Analyze the real-time inference task of the target large model through the artificial intelligence model to obtain the predicted access frequency of the data in the cache by the inference task of the target large model within the third preset duration in the future, and update the data in the cache based on the predicted access frequency.

8. An optimization device for a large model software and hardware integrated machine for rapid deployment, characterized in that, It includes: A first tuning module, configured to, in response to the user's power-on operation of the large model software and hardware integrated machine, obtain the hardware information of each hardware module in the large model software and hardware integrated machine, and generate a first deployment plan for the large model software and hardware integrated machine through the artificial intelligence model according to the hardware information of each hardware module and the preset general model requirements. Among them, the hardware information includes the hardware models of the hardware modules and the transmission information between the hardware modules, and the first deployment plan includes a model quantization plan determined according to the hardware models and a model data transmission plan determined according to the transmission information; A second tuning module, configured to, in response to the user's configuration of the business requirements of the large model software and hardware integrated machine in the visualization interface, determine a target large model from multiple pre-trained large models built in the large model software and hardware integrated machine through the artificial intelligence model according to the business requirements configuration, and adjust the first deployment plan according to the hardware conditions required by the target large model to obtain a second deployment plan for the large model software and hardware integrated machine; A deployment module, configured to deploy the target large model in the large model software and hardware integrated machine based on the second deployment plan.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory is used to store computer instructions that can run on the processor, and the processor is used to implement the steps of the method according to any one of claims 1-7 when executing the computer instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Self-adaptive AI model deployment method

    CN113050955A

  • Method and apparatus for hardware aware machine learning model training

    CN114139714A

  • Model deployment method and device, electronic equipment and storage medium

    CN114594963A

  • Model deployment method and device, model operation method and device, medium and electronic equipment

    CN115658083A

  • Soft and hard all-in-one machine integrating large model training and reasoning and large model training method

    CN118153649A

Cited By

  • Deployment method, device and equipment of large model agent and medium

    CN121116649A