Computing network large model adaptation method, device, equipment and storage medium

By scheduling resources and compressing models on the computing network, and using professional domain corpora to fine-tune large models in certain fields, we address the insufficient end-side service requirements for training general large models on the cloud, achieve efficient deployment of large models on the edge and end, and improve the industry-specificity and security of the models.

CN118214729BActive Publication Date: 2025-09-26INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410209357.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2025-09-26
Estimated Expiration
2044-02-26

AI Technical Summary

Technical Problem

The existing technology model of training general large models on the cloud to support end-side service needs has problems such as insufficient industry specificity and accuracy, large network latency, low security, and high deployment costs.

Method used

By performing business and resource perception on the computing network, obtaining computing resources and business demand indicators, and performing resource scheduling and allocation, we can obtain a target large model that is adapted to the demand intent, use the corpus of professional fields to perform domain fine-tuning and model compression, and realize the edge deployment of the large model.

Benefits of technology

It improves the industry targeting and accuracy of large models, reduces network latency and deployment costs, improves security, and enables flexible deployment of large models at the edge and end sides.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118214729B_ABST
    Figure CN118214729B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, equipment and storage medium for adapting a computing network large model, and relates to the field of artificial intelligence technology. The method obtains the computing network resource indicators and business demand indicators of the computing network by perceiving the business and resources of the computing network, thereby performing resource scheduling and allocation results on the computing power side, network side and application side of the computing network, and determining the demand intention of the application side for the computing network. Based on the resource scheduling and allocation results of the computing network, a target large model adapted to the demand intention is obtained for deployment. The general large model is fine-tuned in the field through the corpus of the professional field, and the professional capability of the large model is adapted. While improving the industry pertinence and accuracy of the large model, the size of the large model is reduced through model compression, and the edge deployment application of the large model is realized. The reduction in the size of the large model is conducive to reducing network latency, improving the security of the large model, and reducing the deployment cost of the large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and storage medium for adapting a large computing network model. Background Art

[0002] The rapid development of big models is enabling new intelligent, multimodal content production models, along with human-like thinking and interaction, to rapidly penetrate various industries. However, the development and application of big models still present a series of challenges. For example, due to the enormous computing power consumed by big models, effective coordination between efficient training and downstream task service delivery is needed, which places higher demands on the upgrade of computing infrastructure. Currently, existing big models are becoming more generalized, incorporating multimodal technology and content diversity, providing high-precision, high-quality content and robust content productivity, and supporting cross-task, cross-modal, multi-functional, and multi-scenario applications. The gradual shift from a single domain to a multimodal domain significantly improves production intelligence while also placing higher demands on the upgrade of computing infrastructure. In the context of the "computing network + big model" converged service model, current big model services primarily rely on the traditional model of training general-purpose big models in the cloud to support on-device service needs. This model has limitations in adapting to specialized scenarios, industry applications, and performance in specific, refined and personalized tasks. On the one hand, mainstream large language models are mainly aimed at general fields, with insufficient knowledge accumulation in various industry scenarios and specific professional fields, weak comprehension capabilities, and insufficient industry-specificity and accuracy of the models. On the other hand, the current large model parameters are large in scale and volume. For example, the GLM-130B model has 260GB of parameters, requiring at least 260GB of GPU memory for inference, making it difficult to deploy on the edge and end side. The traditional model of training large models on the cloud to support end-side service needs has problems such as high network latency, low security, and high deployment costs. Summary of the Invention

[0003] The present invention provides a computing network large model adaptation method, device, equipment and storage medium to solve the existing technology mode of training general large models on the cloud to support terminal-side service needs, resulting in insufficient industry specificity and accuracy of the large model, and the defects of large network latency, low security and high deployment cost.

[0004] The present invention provides a method for adapting a large computing network model, comprising:

[0005] Performing business and resource awareness on the computing network to obtain computing network resource indicators of the computing network's computing resources and business demand indicators on the computing network's user side; the computing network resource indicators include resource load, and the business demand indicators include latency, data transmission volume, and upload traffic;

[0006] According to the computing network resource indicators and the business demand indicators, the computing network is scheduled and allocated resources on the computing power side, the network side, and the application side, and the demand intention of the application side for the computing power network is determined;

[0007] Based on the resource scheduling and allocation results of the computing power network, a target large model adapted to the demand intention is obtained for deployment; the target large model is obtained by fine-tuning the domain and compressing the preset general large model based on the corpus of the professional field corresponding to the demand intention.

[0008] According to the computing network large model adaptation method provided by the present invention, the resource scheduling and allocation on the computing power side, the network side, and the application side of the computing power network is performed based on the computing network resource indicators and the business demand indicators, and the demand intention of the application side for the computing power network is determined, including:

[0009] Based on the computing network resource indicators and the business demand indicators, mapping the business demand and resource capacity of the computing network is performed rule-based, so as to perform business rule analysis on the computing network and generate a scheduling strategy for the computing network;

[0010] Perform resource scheduling and allocation on the computing power side, network side, and application side of the computing power network according to the scheduling strategy;

[0011] The business intention of the application side is translated according to the business demand indicator to determine the demand intention of the application side for the computing power network.

[0012] According to the computing network large model adaptation method provided by the present invention, after performing resource scheduling and allocation on the computing power side, the network side, and the application side of the computing network according to the computing network resource indicators and the business demand indicators, the method further includes:

[0013] Return and execute the step of performing business and resource awareness on the computing power network to dynamically adjust the scheduling strategy.

[0014] According to the computing network large model adaptation method provided by the present invention, the step of acquiring and deploying a target large model adapted to the demand intent includes:

[0015] Obtaining a general large model that is adapted to the demand intention, and integrating and packaging the general large model to obtain a large model base;

[0016] Obtain a corpus of professional fields in the business scenario corresponding to the demand intent, and use the corpus to fine-tune the large model base in the field;

[0017] The knowledge distillation method is used to compress the large model base after domain fine-tuning to obtain the target large model that is adapted to the demand intention and deploy it.

[0018] According to the computing network large model adaptation method provided by the present invention, the knowledge distillation method is used to compress the large model base after domain fine-tuning to obtain a target large model adapted to the demand intent and deploy it, including:

[0019] Based on the preset small model, the knowledge distillation method is used to perform behavioral learning on the large model base after domain fine-tuning. The small model is fine-tuned and trained in a back-propagation manner to transfer the knowledge of the large model base after domain fine-tuning to the small model;

[0020] The small model is trained in a data parallel manner to perform model compression on the large model base after domain fine-tuning, so as to obtain the target large model adapted to the demand intention and deploy it.

[0021] According to the method for adapting a large computing network model provided by the present invention, the domain fine-tuning of the large model base using the corpus includes:

[0022] Inputting the corpus into the large model base, and generating key information and summary information of the corpus;

[0023] Using the key information and summary information as indexes, projecting the corpus into the semantic space of the large model base to generate a high-dimensional vector;

[0024] The large model base is domain-fine-tuned based on the high-dimensional vector.

[0025] According to the computing network large model adaptation method provided by the present invention, obtaining a corpus of professional fields in the business scenario corresponding to the demand intention includes:

[0026] Obtaining a knowledge base under the business scenario corresponding to the demand intention; the knowledge base includes multiple text libraries;

[0027] Performing parsing analysis and standardization on the knowledge base to obtain text data corresponding to the knowledge base;

[0028] The text data is preprocessed to obtain a corpus of professional fields in the business scenario corresponding to the demand intention; the preprocessing operation includes text format cleaning, special format processing, resampling, deduplication, low-information text elimination, semantic enhancement, keyword extraction and word segmentation processing.

[0029] The present invention also provides a computing network large model adaptation device, comprising:

[0030] A computing network perception module is used to perceive the services and resources of the computing network, obtain computing network resource indicators of the computing resources of the computing network, and business demand indicators on the user side of the computing network; the computing network resource indicators include resource load, and the business demand indicators include latency, transmission data volume, and upload traffic;

[0031] A computing network scheduling module is used to schedule and allocate resources on the computing power side, network side, and application side of the computing network based on the computing network resource indicators and the business demand indicators, and determine the demand intention of the application side for the computing power network;

[0032] The model adaptation module is used to obtain and deploy a target large model that is adapted to the demand intention based on the resource scheduling and allocation results of the computing power network; the target large model is obtained by fine-tuning and compressing the preset general large model based on the corpus of the professional field corresponding to the demand intention.

[0033] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the large-scale computing network model adaptation method as described above are implemented.

[0034] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described methods for adapting a large computing network model.

[0035] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of any of the above-described methods for adapting a large computing network model.

[0036] The present invention provides a method, device, equipment and storage medium for adapting a computing network large model. By perceiving the business and resources of the computing network, the computing network resource indicators and business demand indicators of the computing network are obtained, thereby performing resource scheduling and allocation results on the computing power side, network side and application side of the computing network, and determining the application side's demand intention for the computing network. Based on the resource scheduling and allocation results of the computing network, a target large model adapted to the demand intention is obtained for deployment. The general large model is fine-tuned in the field through the corpus of professional fields to adapt the professional capabilities of the large model. While improving the industry pertinence and accuracy of the large model, the size of the large model is reduced through model compression to realize the edge deployment application of the large model. The reduction in the size of the large model is conducive to reducing network latency, improving the security of the large model, and reducing the deployment cost of the large model. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0038] Figure 1 It is a flowchart of the method for adapting a large computing network model provided by the present invention;

[0039] Figure 2 This is a schematic diagram of the network large model adaptation process provided by the present invention;

[0040] Figure 3 This is a schematic diagram of the large model capability adaptation and deployment application process provided by the present invention;

[0041] Figure 4 It is a structural diagram of the computing network large model adaptation device provided by the present invention;

[0042] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0043] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0044] It should be noted that, in the description of the present invention, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, the phrase "comprises a..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus comprising the elements. Terms such as "upper" and "lower" indicate positions or relationships based on those shown in the accompanying drawings and are intended solely to facilitate the description of the present invention and simplify the description. They are not intended to indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation, and are therefore not to be construed as limitations on the present invention. Unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be broadly construed, for example, to mean fixed, removable, or integral; mechanical or electrical; direct or indirect through an intermediary; or internal communication between two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0045] The terms "first," "second," and so forth, used herein are used to distinguish similar objects, not to describe a specific order or precedence. It should be understood that such terms are interchangeable where appropriate, allowing embodiments of the present invention to be implemented in an order other than that illustrated or described herein. Furthermore, the terms "first," "second," and so forth generally distinguish objects of a single type, and do not limit the number of objects. For example, the first object may be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the connected objects.

[0046] An embodiment of the present invention provides a method for adapting a large model of a computing network. By combining the large model algorithm under the computing network with application scenarios such as daily work, business analysis, content creation, and video generation, it integrates professional field knowledge on top of the basic capabilities of the general model, effectively improving the scene perception and understanding capabilities of the large model, flexibly supporting "image-text-audio" multimodal applications, and better meeting various business demand scenarios. On the one hand, based on the basic capabilities of general models, industry-specific and scenario-based professional knowledge is integrated to significantly enhance the scene perception and understanding capabilities of large models, and to improve the model performance of large models for specific industries and application scenarios, as well as for refined and personalized specific tasks, to ensure the practicality, usability and reliability of large models, and to flexibly support multimodal applications such as "image-text-audio"; on the other hand, the intelligent scheduling capabilities of the computing network are utilized to achieve the coordinated integration of cloud services and edge reasoning, give full play to the best effects of cross-layer and cross-domain computing resources, realize multi-factor fusion, cross-capability integration, and elastic on-demand AI services, give full play to the computing power and algorithm advantages of the central intelligent computing cloud, provide diversified and high-precision model services, as well as the network and latency advantages of the edge cloud, to meet the needs of specific industries.

[0047] Specifically, refer to Figure 1 , Figure 1 A flow chart of the method for adapting a large computing network model according to an embodiment of the present invention is provided. Figure 1 The embodiment of the present invention provides a method for adapting a large computing network model, including:

[0048] Step 100: Perform business and resource awareness on the computing network to obtain computing network resource indicators of the computing network's computing resources and business demand indicators on the computing network's user side; the computing network resource indicators include resource load, and the business demand indicators include latency, data transmission volume, and upload traffic;

[0049] First, the computing network is perceived for its services and resources to obtain the computing network resource indicators of the computing network's computing resources and the business demand indicators on the user side of the computing network. The computing network resource indicators of the computing network include resource load and performance, and the business demand indicators on the user side include latency, transmission data volume, and upload traffic.

[0050] The computing power network is the computing network. It performs business and resource perception on the computing power network, that is, business perception and resource perception on the computing network. Through computing network resource perception, the computing network resource indicators corresponding to the underlying resources of the computing network are obtained. Through computing network business perception, the business demand indicators on the computing network user side are obtained.

[0051] In one embodiment, the business and resource perception of the computing network is implemented based on a preset computing network perception component, which is used to collect and perceive the load and performance data of the computing network resources such as the underlying computing power, network and storage, as well as the business requirements of the computing network user side for latency, transmission data volume and upload traffic.

[0052] The computing network perception component further includes a computing network resource perception engine and a computing network service perception engine. The computing network resource perception engine connects to the resource layer and perceives network resource metrics, such as the load and performance of underlying computing resources, including computing power, network resources, and storage resources. It also assesses the quality of computing network services based on these perceived network resource metrics. By accessing information such as available computing power and overall network load, the computing network resource perception engine can flexibly match and process various protocol interfaces, parse, match, convert, and standardize field data collected from underlying resources, dynamically manage various access tasks, monitor and track their execution in real time, and retry or repeat abnormal and erroneous tasks according to scheduling rules.

[0053] The computing network service perception engine is used to perceive user-side computing network latency, data transmission volume, and upload traffic requirements. By connecting to service portals and user service systems, the computing network service perception engine collects user-related business demand data, including business scenarios, resource requirements, and SLA (Service-Level Agreement) requirements. It then maps user business requirements into standardized requirements for underlying computing power, network, and service capabilities. It also connects to computing network service quality data, including service success rate, resource utilization, and user SLAs, to enable a perceptual assessment of computing network service quality.

[0054] Step 200: Based on the computing network resource indicators and the business demand indicators, resource scheduling and allocation are performed on the computing power network on the computing power side, the network side, and the application side, and the demand intention of the application side for the computing power network is determined;

[0055] Based on the perceived computing network resource indicators and business demand indicators, resources are scheduled and allocated on the computing power, network, and application sides of the computing network. Application-side demand intentions for the computing network are determined, corresponding to the network's business scenarios. Based on the perceived computing network resource indicators and business demand indicators, business rules analysis is performed on the computing network, including traffic forecasting and business demand analysis. The corresponding perception analysis results are then generated, enabling the coordinated management of computing network tasks such as cloud-edge-end computing power and network bandwidth allocation, enabling the scheduling and allocation of computing network resources.

[0056] In one embodiment, resource scheduling and allocation for the computing network, as well as the perception of application-side demand for the computing network, are implemented based on a pre-set computing network scheduling component. This computing network scheduling component automatically maps business requirements with resource capabilities based on the computing network resource indicators and business demand indicators perceived by the computing network perception component. Through holistic modeling, business rule analysis, and optimization strategy solution, it automatically allocates and coordinates resources across the computing network, network, and application sides, meeting users' personalized needs for cost and latency.

[0057] Step 300: Based on the resource scheduling and allocation results of the computing power network, a target large model adapted to the demand intention is obtained for deployment; the target large model is obtained by fine-tuning and compressing the preset general large model based on the corpus of the professional field corresponding to the demand intention.

[0058] Based on the results of scheduling and allocating computing network resources, a target big model that matches the intended demand is obtained and deployed. This target big model is obtained by fine-tuning and compressing a pre-set general big model based on the corresponding professional domain corpus. Domain fine-tuning adapts the big model's professional capabilities to improve its performance in professional scenarios. Model compression adjusts the model's size to enable edge deployment of the big model and support its application in industry scenarios.

[0059] In one embodiment, the acquisition and edge deployment of the target big model is implemented based on a pre-set big model capability adaptation component. Based on the scheduling capabilities of the computing network scheduling component, this big model capability adaptation component uses a corpus of specialized domains corresponding to demand intent to fine-tune and compress the general big model. This improves the big model's performance in specialized scenarios while enabling edge deployment of the big model and supporting its application in various industry scenarios.

[0060] In this embodiment, by perceiving the business and resources of the computing network, the computing network resource indicators and business demand indicators of the computing network are obtained, thereby performing resource scheduling and allocation results on the computing power side, network side, and application side of the computing network, and determining the application side's demand intention for the computing network. Based on the resource scheduling and allocation results of the computing network, a target large model adapted to the demand intention is obtained for deployment. The general large model is fine-tuned in the field through the corpus of professional fields to adapt the professional capabilities of the large model. While improving the industry-specificity and accuracy of the large model, the size of the large model is reduced through model compression, realizing the edge deployment application of the large model. The reduction in the size of the large model is conducive to reducing network latency and deployment costs, and improving security.

[0061] In one embodiment, step 200 performs resource scheduling and allocation on the computing power network, the network side, and the application side based on the perceived computing network resource indicators and business demand indicators, and determines the application side's demand intention for the computing power network. This may also include:

[0062] Step 201: Based on the computing network resource indicators and the business demand indicators, rule mapping is performed between the business demand and resource capacity of the computing network to perform business rule analysis on the computing network and generate a scheduling strategy for the computing network.

[0063] Step 202: Scheduling and allocating resources on the computing power side, network side, and application side of the computing power network according to the scheduling strategy;

[0064] Step 203: Translate the business intention of the application side according to the business demand indicator to determine the demand intention of the application side for the computing power network.

[0065] First, based on perceived computing network resource indicators and business demand indicators, a rule-based mapping is performed between the computing network's business requirements and resource capabilities. This allows for business rule analysis of the computing network and generates a scheduling strategy for the computing network. Then, based on the generated scheduling strategy, resource scheduling and allocation are performed on the computing power, network, and application sides of the computing network. Finally, the business intent of the computing network's application side is translated based on the business demand indicators to determine the application's intended demand for the computing network.

[0066] In one embodiment, the computing network scheduling component may further include a computing network intent parsing and translation engine, a computing network scheduling strategy optimization engine, and a computing network intelligent scheduling execution engine. The computing network intent parsing and translation engine can translate the business requirements of the application side based on business demand indicators to determine the application side's demand intentions for the computing power network; the computing network scheduling strategy optimization engine can analyze the business rules of the computing power network based on computing network resource indicators and business demand indicators to generate a scheduling strategy for the computing power network; and the computing network intelligent scheduling execution engine can perform resource scheduling and allocation on the computing power side, network side, and application side of the computing power network based on the scheduling strategy generated by the computing network scheduling strategy optimization engine.

[0067] Specifically, the computing network intent analysis and translation engine supports a variety of intent input methods such as text, voice, and pictures. Based on the accurate perception and capture of user demand intentions, it automatically and accurately translates the user's business intentions into demand for computing network resources according to rule strategies and combined with corresponding artificial intelligence / machine learning algorithms, and generates the application side's evaluation results for computing network resources such as computing power, network, and storage, thereby determining the application side's demand intentions for the computing network.

[0068] The computing network scheduling strategy optimization engine can automatically map business needs and resource capabilities through rules based on factors such as the SLA requirements of different businesses, the overall network load, and available computing resources through modeling, target definition, algorithm optimization, and policy generation. It can dynamically generate scheduling optimization strategies for the overall management of tasks such as cloud-edge-end computing power and network bandwidth allocation, and adopt dynamic and adaptive transmission and processing strategies in different time periods and / or different spatial points to meet business requirements while ensuring real-time processing.

[0069] The intelligent scheduling execution engine of the computing network dispatches qualified process templates to execute according to the generated scheduling optimization strategy and the needs of specific business scenarios to meet the corresponding needs, completes the automatic allocation and scheduling of resources on the computing power side, network side and application side, and realizes the automatic activation of end-to-end services.

[0070] In one embodiment, based on the computing network perception component and the computing network scheduling component, closed-loop optimization of the computing network resource scheduling and allocation can be achieved. In step 200, after the computing network is scheduled and allocated with computing power, network, and application resources according to the computing network resource indicators and business demand indicators, the following steps may also be included:

[0071] Step 210, return and execute the step of performing business and resource perception on the computing power network to dynamically adjust the scheduling strategy.

[0072] After completing the resource scheduling and allocation of the computing power network, the business and resource perception of the computing power network is re-performed to obtain the computing network resource indicators and business demand indicators after resource scheduling and allocation, so as to dynamically adjust the scheduling strategy.

[0073] In one embodiment, after the computing network scheduling component completes the scheduling execution of computing network resources, the computing network perception component will feed back the perception evaluation results of computing network resource indicators and business demand indicators to the computing network scheduling component, evaluate in real time whether the current scheduling strategy can meet the user's business needs, and dynamically adjust the scheduling strategy according to the evaluation results, thereby realizing a feedback and optimization closed loop for the scheduling allocation of computing network resources.

[0074] In step 300, a target large model adapted to the application-side demand intent is obtained and deployed, specifically including:

[0075] Step 310: Obtain a general large model that is adapted to the demand intent, and integrate and package the general large model to obtain a large model base;

[0076] Step 320: Obtain a corpus of professional fields in the business scenario corresponding to the demand intent, and use the corpus to fine-tune the large model base in the field;

[0077] In step 330, the knowledge distillation method is used to compress the large model base after domain fine-tuning to obtain a target large model that is adapted to the demand intent and deploy it.

[0078] First, a universal big model that matches the application's needs and intent is obtained and integrated into a package to create the corresponding big model base. Different needs and intents can correspond to the same or different universal big models, each with different parameter sizes. To meet different business scenarios, different big model bases can be flexibly and dynamically switched.

[0079] Then, a corpus of professional domains in business scenarios corresponding to the application's demand intent is obtained, and this corpus is used to fine-tune the constructed large model base in the domain, adapting the large model to its professional capabilities. Finally, the knowledge distillation method is used to compress the large model base after domain fine-tuning, obtaining the target large model that is adapted to the application's demand intent and deploying it. Among them, using a corpus of professional domains to fine-tune the large model base in the domain, that is, fine-tuning the large model in the domain, can increase the large model's knowledge accumulation in the professional domain, enhance the large model's knowledge comprehension ability in the professional domain, and improve the model's industry-specific targeting and accuracy.

[0080] In one embodiment, the adaptation and deployment of large models primarily involves several phases: large model integration and packaging, corpus collection and preprocessing, domain fine-tuning, and model compression. This integration and packaging of large models allows for integration and packaging testing of large models with varying parameter sizes, establishing a large model base. Furthermore, based on online testing and verification results, different base large models can be flexibly and dynamically switched to suit different usage scenarios.

[0081] In the corpus collection and preprocessing stage, the corpus of professional fields under various business scenarios is mainly collected as the basis for domain fine-tuning of the large model. In step 320, the corpus of professional fields under the business scenarios corresponding to the application-side demand intent is obtained, specifically including:

[0082] Step 321: Obtain a knowledge base in a business scenario corresponding to the demand intention; the knowledge base includes multiple text bases;

[0083] Step 322: performing parsing analysis and standardization on the knowledge base to obtain text data corresponding to the knowledge base;

[0084] Step 323, preprocessing operations are performed on the text data to obtain a corpus of professional fields under the business scenario corresponding to the demand intention; the preprocessing operations include text format cleaning, special format processing, resampling, deduplication, low-information text elimination, semantic enhancement, keyword extraction and word segmentation processing.

[0085] First, obtain the knowledge base for the business scenario corresponding to the application's demand intent, including open source communities, blogs, industry-related online public data, electronic document materials, and open data sets. This knowledge base includes multiple text libraries such as PDF, HTML, docx, and TXT. Through standardized parsing programs, etc., parse, analyze, and standardize each text library in the acquired knowledge base to obtain the corresponding text data. Preprocess the text data through preprocessing programs, etc., to obtain a corpus that can ultimately be used for domain fine-tuning of the large model. Preprocessing operations include but are not limited to text format cleaning, special format processing, resampling, deduplication, low-information text removal, semantic enhancement, keyword extraction, and word segmentation.

[0086] Furthermore, in step 320, the domain-specific corpus obtained is used to fine-tune the large model base, specifically including:

[0087] Step 324: input the corpus into the large model base and generate key information and summary information of the corpus;

[0088] Step 325 , using the key information and summary information as indexes, projecting the corpus into the semantic space of the large model base to generate a high-dimensional vector;

[0089] Step 326: Perform domain fine-tuning on the large model base based on the high-dimensional vector.

[0090] The preprocessed corpus is fed into the large model base, generating key information and summary information for the corpus. Using this key information and summary information as indexes, the corpus is projected into the semantic space of the large model base, generating corresponding high-dimensional vectors. Based on these high-dimensional vectors, the large model base is fine-tuned for its specific domain. Furthermore, by continuously improving the size and quality of the corpus, the large model can be continuously iterated and optimized.

[0091] Domain fine-tuning of the large model base can be achieved through model fine-tuning technology. Specifically, the large model is further trained through downstream task datasets, the performance of the fine-tuned model is evaluated using a validation set, and the large model is optimized and adjusted to achieve the construction of a personalized scenario large model, making the large model more adaptable and accurate when facing specific downstream tasks.

[0092] In the actual application of large models, Prompt technology can also be used. By developing standardized interfaces, for question-answering scenarios, after the user enters a question, the user question is first mapped into a high-dimensional vector, and similarity matching is performed with the high-dimensional vector in the semantic space of the large model using methods such as Euclidean distance. Knowledge prompts are then provided based on the matched knowledge base that has a high similarity with the user question, thereby targeting user questions and improving the accuracy and professionalism of the large model's responses in professional fields.

[0093] Furthermore, in step 330, the knowledge distillation method is used to compress the large model base after domain fine-tuning to obtain a target large model that is adapted to the application's requirements and intent and then deployed. Specifically, the following steps are performed:

[0094] Step 331: Based on the preset small model, the knowledge distillation method is used to perform behavioral learning on the domain-fine-tuned large model base, and the small model is fine-tuned and trained in a back-propagation manner to transfer the knowledge of the domain-fine-tuned large model base to the small model;

[0095] Step 332: The small model is trained in a data parallel manner to perform model compression on the large model base after domain fine-tuning, and a target large model adapted to the demand intent is obtained and deployed.

[0096] During the model compression stage, knowledge distillation is used to compress the fine-tuned large model, weakening the irrelevant general capabilities of the large model while retaining its professional capabilities. The computing power and storage resources required for compressed model reasoning are greatly reduced, enabling the service to be further sunk to the edge side, providing users with lower-latency, higher-bandwidth AI services.

[0097] Specifically, first, based on a pre-set small model, knowledge distillation is used to learn the behavior of the large model. Backpropagation is then used to fine-tune the small model, transferring the knowledge of the large model to the small model, which is a smaller compressed language model. Then, data parallelism is used to train the small model, compressing the size of the large model while retaining the thought chain and emergence capabilities of the large model. This improves model performance and generalization, facilitates edge deployment of the large model, and reduces its deployment cost.

[0098] In one embodiment, referring to Figure 2In the large-scale computing network model adaptation process shown, the computing network perception component performs computing network business and resource perception. Based on the computing network perception component's business and resource perception results, the computing network scheduling component performs computing network intent interpretation and translation, optimizes computing network scheduling policies, and executes intelligent computing network scheduling. Based on the scheduling capabilities of the computing network perception component, the large-scale model capability adaptation component adapts the large model to specialized domain capabilities, compresses the model, and deploys and applies it. After the large-scale model is deployed and applied, the computing network scheduling component optimizes and adjusts the computing network scheduling policy based on the computing network perception component's business and resource perception results, thus achieving a closed-loop feedback optimization for computing network resource scheduling.

[0099] Further, refer to Figure 3 The large-scale model capability adaptation and deployment application process shown in the figure integrates and encapsulates common large models to form a large-scale model base. The large-scale model base is then fine-tuned and compressed using a pre-processed corpus. The target large-scale model is then deployed on the edge, resulting in an edge model. The fine-tuned large-scale model base is deployed to the cloud, forming a cloud-side model. The corpus is stored to form the large-scale model's knowledge base. Based on user input in the cloud-side and edge models, as well as their output in response to user input, the large-scale model base can be flexibly and dynamically switched for different business scenarios. Furthermore, in different business scenarios, the edge model can be tested and verified online using user input and model output, continuously iteratively optimizing the cloud-side model. The cloud-side model, in turn, continuously iteratively optimizes the large-scale model knowledge base based on user input and model output. This increases the large-scale model's domain-specific knowledge accumulation and enhances its understanding capabilities.

[0100] In this embodiment, based on the perception and optimized scheduling of the computing network, the underlying ubiquitous heterogeneous computing network resources are fully utilized, and the industry-specific and scenario-specific data sets and knowledge bases are combined. Through technologies such as model fine-tuning and model compression, the understanding ability of large models in professional fields is improved, as well as the model performance and controllability in refined and personalized specific tasks. The model volume is effectively reduced through model compression, and the large model service is deployed to the edge and end devices to achieve further coordination between cloud training and edge inference, providing users with lower thresholds, low latency and on-demand flexible AI service capabilities, giving full play to the advantages of the computing network, providing diversified and high-precision large model services, meeting specific industry needs, improving the practicality, ease of use and reliability of large models, and achieving rapid business response.

[0101] Under the integrated service background of "computing network + industry big model", business efficiency, automation and intelligence levels are improved. Through model compression, the consumption of computing power and storage resources by big model applications is alleviated, and the energy consumption and cost of big model applications are reduced. Through dynamic perception of computing network resources, optimal scheduling and reasonable allocation of resources, the overall resource utilization of the computing network is improved.

[0102] The following describes the large-scale computing network model adaptation device provided by the present invention. The large-scale computing network model adaptation device described below and the large-scale computing network model adaptation method described above can be referenced to each other.

[0103] Reference Figure 4 The embodiment of the present invention provides a large-scale computing network model adaptation device, including:

[0104] The computing network perception module 10 is used to perceive the services and resources of the computing network, obtain computing network resource indicators of the computing resources of the computing network, and business demand indicators of the computing network user side; the computing network resource indicators include resource load, and the business demand indicators include latency, transmission data volume, and upload traffic;

[0105] The computing network scheduling module 20 is used to schedule and allocate resources on the computing power side, network side, and application side of the computing network according to the computing network resource indicators and the business demand indicators, and determine the demand intention of the application side for the computing power network;

[0106] The model adaptation module 30 is used to obtain and deploy a target large model that is adapted to the demand intention based on the resource scheduling and allocation results of the computing power network; the target large model is obtained by fine-tuning and compressing the preset general large model based on the corpus of the professional field corresponding to the demand intention.

[0107] In one embodiment, the computing network scheduling module 20 is further configured to:

[0108] Based on the computing network resource indicators and the business demand indicators, mapping the business demand and resource capacity of the computing network is performed rule-based, so as to perform business rule analysis on the computing network and generate a scheduling strategy for the computing network;

[0109] Perform resource scheduling and allocation on the computing power side, network side, and application side of the computing power network according to the scheduling strategy;

[0110] The business intention of the application side is translated according to the business demand indicator to determine the demand intention of the application side for the computing power network.

[0111] In one embodiment, the computing network scheduling module 20 is further configured to:

[0112] Return and execute the step of performing business and resource awareness on the computing power network to dynamically adjust the scheduling strategy.

[0113] In one embodiment, the model adaptation module 30 is further configured to:

[0114] Obtaining a general large model that is adapted to the demand intention, and integrating and packaging the general large model to obtain a large model base;

[0115] Obtain a corpus of professional fields in the business scenario corresponding to the demand intent, and use the corpus to fine-tune the large model base in the field;

[0116] The knowledge distillation method is used to compress the large model base after domain fine-tuning to obtain the target large model that is adapted to the demand intention and deploy it.

[0117] In one embodiment, the model adaptation module 30 is further configured to:

[0118] Based on the preset small model, the knowledge distillation method is used to perform behavioral learning on the large model base after domain fine-tuning. The small model is fine-tuned and trained in a back-propagation manner to transfer the knowledge of the large model base after domain fine-tuning to the small model;

[0119] The small model is trained in a data parallel manner to perform model compression on the large model base after domain fine-tuning, so as to obtain the target large model adapted to the demand intention and deploy it.

[0120] In one embodiment, the model adaptation module 30 is further configured to:

[0121] Inputting the corpus into the large model base, and generating key information and summary information of the corpus;

[0122] Using the key information and summary information as indexes, projecting the corpus into the semantic space of the large model base to generate a high-dimensional vector;

[0123] The large model base is domain-fine-tuned based on the high-dimensional vector.

[0124] In one embodiment, the model adaptation module 30 is further configured to:

[0125] Obtaining a knowledge base under the business scenario corresponding to the demand intention; the knowledge base includes multiple text libraries;

[0126] Performing parsing analysis and standardization on the knowledge base to obtain text data corresponding to the knowledge base;

[0127] The text data is preprocessed to obtain a corpus of professional fields in the business scenario corresponding to the demand intention; the preprocessing operation includes text format cleaning, special format processing, resampling, deduplication, low-information text elimination, semantic enhancement, keyword extraction and word segmentation processing.

[0128] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the large-scale computing network model adaptation method, which includes:

[0129] Performing business and resource awareness on the computing network to obtain computing network resource indicators of the computing network's computing resources and business demand indicators on the computing network's user side; the computing network resource indicators include resource load, and the business demand indicators include latency, data transmission volume, and upload traffic;

[0130] According to the computing network resource indicators and the business demand indicators, the computing network is scheduled and allocated resources on the computing power side, the network side, and the application side, and the demand intention of the application side for the computing power network is determined;

[0131] Based on the resource scheduling and allocation results of the computing power network, a target large model adapted to the demand intention is obtained for deployment; the target large model is obtained by fine-tuning the domain and compressing the preset general large model based on the corpus of the professional field corresponding to the demand intention.

[0132] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0133] On the other hand, the present invention further provides a computer program product, comprising a computer program, which may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the large-scale computing network model adaptation method provided by the above methods, which includes:

[0134] Performing business and resource awareness on the computing network to obtain computing network resource indicators of the computing network's computing resources and business demand indicators on the computing network's user side; the computing network resource indicators include resource load, and the business demand indicators include latency, data transmission volume, and upload traffic;

[0135] According to the computing network resource indicators and the business demand indicators, the computing network is scheduled and allocated resources on the computing power side, the network side, and the application side, and the demand intention of the application side for the computing power network is determined;

[0136] Based on the resource scheduling and allocation results of the computing power network, a target large model adapted to the demand intention is obtained for deployment; the target large model is obtained by fine-tuning the domain and compressing the preset general large model based on the corpus of the professional field corresponding to the demand intention.

[0137] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the large-scale computing network model adaptation method provided by the above methods, the method comprising:

[0138] Performing business and resource awareness on the computing network to obtain computing network resource indicators of the computing network's computing resources and business demand indicators on the computing network's user side; the computing network resource indicators include resource load, and the business demand indicators include latency, data transmission volume, and upload traffic;

[0139] According to the computing network resource indicators and the business demand indicators, the computing network is scheduled and allocated resources on the computing power side, the network side, and the application side, and the demand intention of the application side for the computing power network is determined;

[0140] Based on the resource scheduling and allocation results of the computing power network, a target large model adapted to the demand intention is obtained for deployment; the target large model is obtained by fine-tuning the domain and compressing the preset general large model based on the corpus of the professional field corresponding to the demand intention.

[0141] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0142] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for adapting a large computing network model, characterized in that: include: Performing business and resource awareness on the computing network to obtain computing network resource indicators of the computing network's computing resources and business demand indicators on the computing network's user side; the computing network resource indicators include resource load, and the business demand indicators include latency, data transmission volume, and upload traffic; According to the computing network resource indicators and the business demand indicators, the computing network is scheduled and allocated resources on the computing power side, the network side, and the application side, and the demand intention of the application side for the computing power network is determined; Based on the resource scheduling and allocation results of the computing power network, a target large model adapted to the demand intention is obtained and deployed; The target large model is obtained by fine-tuning and compressing the preset general large model based on the corpus of the professional field corresponding to the demand intention; The step of scheduling and allocating resources on the computing power network, the network side, and the application side based on the computing network resource indicators and the business demand indicators, and determining the application side's demand intention for the computing power network, includes: Based on the computing network resource indicators and the business demand indicators, mapping the business demand and resource capacity of the computing network is performed rule-based, so as to perform business rule analysis on the computing network and generate a scheduling strategy for the computing network; Perform resource scheduling and allocation on the computing power side, network side, and application side of the computing power network according to the scheduling strategy; The business intention of the application side is translated according to the business demand indicator to determine the demand intention of the application side for the computing power network.

2. The method for adapting a large computing network model according to claim 1, characterized in that: After scheduling and allocating resources on the computing power network, the network side, and the application side according to the computing network resource indicators and the business demand indicators, the method further includes: Return and execute the step of performing business and resource awareness on the computing power network to dynamically adjust the scheduling strategy.

3. The method for adapting a large computing network model according to claim 1, characterized in that: The acquiring and deploying a target large model adapted to the demand intention includes: Obtaining a general large model that is adapted to the demand intention, and integrating and packaging the general large model to obtain a large model base; Obtain a corpus of professional fields in the business scenario corresponding to the demand intent, and use the corpus to fine-tune the large model base in the field; The knowledge distillation method is used to compress the large model base after domain fine-tuning to obtain the target large model that is adapted to the demand intention and deploy it.

4. The method for adapting a large computing network model according to claim 3, characterized in that: The method of using the knowledge distillation method to compress the large model base after domain fine-tuning to obtain a target large model that is adapted to the demand intent and deploy it includes: Based on the preset small model, the knowledge distillation method is used to perform behavioral learning on the large model base after domain fine-tuning. The small model is fine-tuned and trained in a back-propagation manner to transfer the knowledge of the large model base after domain fine-tuning to the small model; The small model is trained in a data parallel manner to perform model compression on the large model base after domain fine-tuning, so as to obtain the target large model adapted to the demand intention and deploy it.

5. The method for adapting a large computing network model according to claim 3, wherein: The domain fine-tuning of the large model base using the corpus includes: Inputting the corpus into the large model base, and generating key information and summary information of the corpus; Using the key information and summary information as indexes, projecting the corpus into the semantic space of the large model base to generate a high-dimensional vector; The large model base is domain-fine-tuned based on the high-dimensional vector.

6. The method for adapting a large computing network model according to claim 3, characterized in that: The obtaining of a corpus of professional fields in a business scenario corresponding to the demand intention includes: Obtaining a knowledge base under the business scenario corresponding to the demand intention; the knowledge base includes multiple text libraries; Performing parsing analysis and standardization on the knowledge base to obtain text data corresponding to the knowledge base; The text data is preprocessed to obtain a corpus of professional fields in the business scenario corresponding to the demand intention; the preprocessing operation includes text format cleaning, special format processing, resampling, deduplication, low-information text elimination, semantic enhancement, keyword extraction and word segmentation processing.

7. A large-scale model adaptation device for a computing network, characterized in that: include: A computing network perception module is used to perceive the services and resources of the computing network, obtain computing network resource indicators of the computing resources of the computing network, and business demand indicators on the user side of the computing network; the computing network resource indicators include resource load, and the business demand indicators include latency, transmission data volume, and upload traffic; A computing network scheduling module is used to schedule and allocate resources on the computing power side, network side, and application side of the computing network based on the computing network resource indicators and the business demand indicators, and determine the demand intention of the application side for the computing power network; A model adaptation module is used to obtain and deploy a target large model adapted to the demand intention based on the resource scheduling and allocation results of the computing power network; The target large model is obtained by fine-tuning and compressing the preset general large model based on the corpus of the professional field corresponding to the demand intention; The computing network scheduling module is further used to: Based on the computing network resource indicators and the business demand indicators, mapping the business demand and resource capacity of the computing network is performed rule-based, so as to perform business rule analysis on the computing network and generate a scheduling strategy for the computing network; Perform resource scheduling and allocation on the computing power side, network side, and application side of the computing power network according to the scheduling strategy; The business intention of the application side is translated according to the business demand indicator to determine the demand intention of the application side for the computing power network.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the large-scale computing network model adaptation method as described in any one of claims 1 to 6 are implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for adapting a large computing network model as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method and system for scheduling computing power network resources

    CN115914392A

  • Large model parallel training method and system and readable storage medium

    CN117311975A