Method and device for finely adjusting image segmentation large model based on computing power of intelligent computing center
Through distributed parallel fine-tuning and evaluation of the computing power resources of the intelligent computing center, the defects of traditional image segmentation large models in terms of computing performance, scalability and accuracy are solved, and efficient and accurate image segmentation effect is achieved, meeting the needs of modern industrial quality inspection.
Patent Information
- Application Number
- CN202510212995.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional image segmentation large models have obvious shortcomings in computing performance, scalability and model accuracy, and it is difficult to meet the needs of modern industrial quality inspection.
Through the computing power resources of the intelligent computing center, distributed parallel fine-tuning and evaluation methods are adopted to fine-tune the image segmentation model, and the rich hyperparameter space and parallel fine-tuning strategies are used to improve the computing performance and accuracy of the model.
It significantly improves the computing performance and accuracy of image segmentation large models, can effectively meet the needs of modern industrial quality inspection, and improves quality inspection accuracy and efficiency.
Smart Images

Figure CN120147335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a method and device for fine-tuning an image segmentation large model based on the computing power of an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, to mainly provide the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). An intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] An "intelligent computing center" includes, but is not limited to, an "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", which is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0007] In the field of image segmentation, SAM (Segment Anything Model, a large image segmentation model) is a powerful image segmentation model. Although its basic performance is excellent, in specific fields (such as industrial inspection), the accuracy often cannot meet the business requirements. It is necessary to use the annotation data in this field to fine-tune the model to achieve the ideal effect.
[0008] In the industrial production and manufacturing industry, real-time quality inspection is an important link to ensure product quality and a key measure for enterprises to ensure that products meet standards and improve competitiveness. The large image segmentation model can perform per-pixel classification on the input quality inspection images, distinguish different instances of the same category, and accurately calculate quality inspection indicators through business post-processing logic. However, traditional solutions have the following defects under the requirements of high precision, high performance, and wide application scenarios:
[0009] 1. Low computing performance: The computing power and memory capacity of a single server are limited. The dynamic memory (such as activation values, gradients, etc.) and static memory (model parameters) generated during the training process will quickly exhaust the available video memory resources, and a single machine is limited by the internal communication bandwidth and latency characteristics of the node.
[0010] 2. Scalability: It is difficult for a single server to flexibly adjust resource allocation according to business requirements, unable to adapt to changing workloads, and difficult to minimize communication overhead.
[0011] 3. Low model accuracy: The basic large image segmentation model has poor adaptability to unknown categories in specific fields and insufficient generalization performance, making it difficult to meet the project requirements for high accuracy and wide applicability.
[0012] In summary, the traditional large image segmentation model has obvious defects in computing performance, scalability, and model accuracy, and it is difficult to meet the needs of modern industrial quality inspection. Summary of the Invention
[0013] The present invention provides a method and device for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center to solve the technical problem that the traditional large image segmentation model has obvious defects in computing performance, scalability, and model accuracy and is difficult to meet the needs of modern industrial quality inspection.
[0014] To solve the above technical problems, the present invention is implemented as follows:
[0015] In the first aspect, the present invention provides a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center, and the method includes:
[0016] Step S1: Obtain a task to be processed, and determine corresponding training data based on the task to be processed;
[0017] Step S2: Determine a hyperparameter space, where the hyperparameter space includes multiple hyperparameter groups;
[0018] Step S3: Execute the fine-tuning task of the large image segmentation model. The fine-tuning task of the large image segmentation model includes: based on the hyperparameter space, the training data, and the computing power resources of the intelligent computing center, perform distributed parallel fine-tuning on the large image segmentation model to obtain multiple fine-tuned large image segmentation models, where each fine-tuned large image segmentation model corresponds to one hyperparameter group in the hyperparameter space;
[0019] Step S4: Perform the evaluation task of the large image segmentation model. The evaluation task of the large image segmentation model includes: based on the preset evaluation metrics and the computing power resources of the intelligent computing center, perform distributed parallel evaluation on each fine-tuned large image segmentation model to obtain the evaluation results corresponding to each fine-tuned large image segmentation model, where the evaluation result is the value of the evaluation metric;
[0020] Step S5: Perform the task of determining the large image segmentation model. The task of determining the large image segmentation model includes: determine the best evaluation result among the evaluation results, and determine the fine-tuned large image segmentation model corresponding to the best evaluation result as the target large image segmentation model, and record the hyperparameter group corresponding to the target large image segmentation model; where the best evaluation result is the evaluation result corresponding to the best value of the evaluation metric.
[0021] Optionally, step S3 includes:
[0022] Step S31: Determine the fine-tuning strategy of the large image segmentation model, and perform the fine-tuning task of the large image segmentation model based on the fine-tuning strategy of the large image segmentation model;
[0023] Among them, the fine-tuning strategy of the large image segmentation model includes at least one of the following:
[0024] Perform staged fine-tuning on the large image segmentation model. Among them, the staged fine-tuning includes: fixing all the current layers of the large image segmentation model; training the newly added layers; after the training of the newly added layers is completed, select the target layer from all the layers according to the user's instruction, and perform joint training on the target layer and the newly added layers;
[0025] Based on the image enhancement technology, represent the same sample image in the training data at multiple different resolutions;
[0026] Use the AdamW optimizer to decay the learning rate in the later stage of the fine-tuning of the large image segmentation model;
[0027] Construct a mask loss function and a prompt loss function; perform fine-tuning on the large image segmentation model based on the mask loss function and the prompt loss function;
[0028] Use Dropout and / or Batch Normalization technology to reduce overfitting during the fine-tuning process of the large image segmentation model;
[0029] Use L2 regularization technology to improve the generalization of the large image segmentation model during the fine-tuning process of the large image segmentation model.
[0030] Optionally, step S1 includes:
[0031] Step S11: Obtain the task to be processed, and determine the image to be annotated based on the task to be processed;
[0032] Step S12: Input the image to be annotated into the large image segmentation model for pre-annotation, and output the pre-annotated image;
[0033] Step S13: Verify the pre-annotated image based on the Labelme tool to obtain the verified image;
[0034] Step S14: Determine the verified image and the image to be annotated as the training data.
[0035] Optionally, after step S5, the method further includes:
[0036] Step S6: Perform distributed parallel inference based on the target large image segmentation model and the computing power resources of the intelligent computing center; wherein, the target large image segmentation model uses the ONNX inference framework.
[0037] Optionally, the hyperparameters include at least one of the following: learning rate, batch size, number of iterations, optimizer, regularization parameter, non-maximum suppression parameter.
[0038] Optionally, the computing power resources of the intelligent computing center are distributed computing power resources, and the distributed computing power resources include: graphics processing units (GPUs) located on different nodes, and at least one GPU is set on each node; on one GPU, only one fine-tuning task or evaluation task of the large image segmentation model is executed at a time.
[0039] In a second aspect, the present invention provides a device for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center, and the device includes:
[0040] An acquisition module, configured to acquire the task to be processed, and determine the corresponding training data based on the task to be processed;
[0041] An execution module, configured to determine a hyperparameter space, where the hyperparameter space includes multiple hyperparameter groups;
[0042] Execute the fine-tuning task of the large image segmentation model, where the fine-tuning task of the large image segmentation model includes: performing distributed parallel fine-tuning on the large image segmentation model based on the hyperparameter space, the training data, and the computing power resources of the intelligent computing center to obtain multiple fine-tuned large image segmentation models, where each fine-tuned large image segmentation model corresponds to one hyperparameter group in the hyperparameter space;
[0043] Execute the evaluation task of the large image segmentation model. The evaluation task of the large image segmentation model includes: based on the preset evaluation metrics and the computing power resources of the intelligent computing center, perform distributed parallel evaluation on each fine-tuned large image segmentation model to obtain the evaluation results corresponding to each fine-tuned large image segmentation model, where the evaluation result is the value of the evaluation metric;
[0044] Execute the determination task of the large image segmentation model. The determination task of the large image segmentation model includes: determine the best evaluation result among the evaluation results, determine the fine-tuned large image segmentation model corresponding to the best evaluation result as the target image segmentation model, and record the hyperparameter group corresponding to the target image segmentation model; where the best evaluation result is the evaluation result corresponding to the best value of the evaluation metric.
[0045] Optionally, the execution module is further configured to determine the fine-tuning strategy of the large image segmentation model, and based on the fine-tuning strategy of the large image segmentation model, execute the fine-tuning task of the large image segmentation model;
[0046] Among them, the fine-tuning strategy of the large image segmentation model includes at least one of the following:
[0047] Perform staged fine-tuning on the large image segmentation model. Among them, the staged fine-tuning includes: fixing all the current layers of the large image segmentation model; training the newly added layers; after the training of the newly added layers is completed, select the target layers from all the layers according to the user's instructions, and perform joint training on the target layers and the newly added layers;
[0048] Based on the image enhancement technology, represent the same sample image in the training data at multiple different resolutions;
[0049] Use the AdamW optimizer to decay the learning rate in the later stage of the fine-tuning of the large image segmentation model;
[0050] Construct a mask loss function and a prompt loss function; fine-tune the large image segmentation model based on the mask loss function and the prompt loss function;
[0051] Use Dropout and / or Batch Normalization technology to reduce overfitting during the fine-tuning process of the large image segmentation model;
[0052] Use L2 regularization technology to improve the generalization of the large image segmentation model during the fine-tuning process of the large image segmentation model.
[0053] Optionally, the acquisition module is further configured to acquire the task to be processed, and determine the image to be annotated based on the task to be processed;
[0054] Input the image to be labeled into the large image segmentation model for pre-labeling, and output the pre-labeled image;
[0055] Based on the Labelme tool, verify the pre-labeled image to obtain the verified image;
[0056] Determine the verified image and the image to be labeled as the training data.
[0057] Optionally, the execution module is further configured to perform distributed parallel inference based on the target large image segmentation model and the computing power resources of the intelligent computing center after the large image segmentation model determination task is executed; wherein, the target large image segmentation model uses the ONNX inference framework.
[0058] Optionally, the hyperparameters include at least one of the following: learning rate, batch size, number of iterations, optimizer, regularization parameter, non-maximum suppression parameter.
[0059] Optionally, the computing power resources of the intelligent computing center are distributed computing power resources, and the distributed computing power resources include: GPUs located on different nodes, and at least one GPU is set on each node; on one GPU, only one fine-tuning task or evaluation task of the large image segmentation model is executed at a time.
[0060] In a third aspect, the present invention provides an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor, and when the program is executed by the processor, the steps of a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center as described in the first aspect are implemented.
[0061] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor, the steps of a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center as described in the first aspect are implemented.
[0062] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and when the computer instructions are executed by the processor, the steps of a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center as described in the first aspect are implemented.
[0063] Compared with the existing fine-tuning methods for large image segmentation models, the present invention utilizes the computing power resources of an intelligent computing center, which can significantly improve computing performance and avoid resource exhaustion problems when a single server processes dynamic and static memory. Moreover, through a distributed parallel fine-tuning and evaluation process, the training of the large image segmentation model can be carried out simultaneously on multiple nodes, thus greatly improving computing efficiency.
[0064] In terms of scalability, the intelligent computing center allows for flexible adjustment of the allocation of computing power resources to adapt to changing workloads. Therefore, computing resources can be dynamically expanded or reduced according to the business requirements during the fine-tuning of the large image segmentation model, avoiding waste of computing power resources.
[0065] In terms of the accuracy of the large image segmentation model, through a rich hyperparameter space and a parallel fine-tuning strategy, it is ensured that the best combination of hyperparameters can be found, thereby improving the accuracy of the large image segmentation model.
[0066] In summary, through the computing power resources of the intelligent computing center, distributed parallel fine-tuning and evaluation of the large image segmentation model are realized, significantly improving the computing performance and accuracy of the large image segmentation model, thus effectively meeting the needs of modern industrial quality inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0068] Figure 1 is a flowchart of a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center provided by the present invention;
[0069] Figure 2 is a flowchart of a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center provided by the present invention;
[0070] Figure 3 is a flowchart of a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center provided by the present invention;
[0071] Figure 4 is a structural block diagram of a device for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center provided by the present invention;
[0072] Figure 5 is a schematic structural diagram of an electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0074] The following briefly explains the technical terms related to the present invention.
[0075] The "computing power" as described in the present invention refers to: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to process information data and achieve the output of the target result, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0076] The "computational power" (Computational Power, CP) as described in the present invention refers to: the ability of a data center server to process data and achieve the output of the result, a comprehensive indicator for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP_general + CP_intelligent + CP_super
[0077] The "carrying capacity" (Network Power, NP) as described in the present invention refers to: the performance of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability.
[0078] The "Storage Power" (SP) described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, which is a comprehensive indicator for measuring the data storage capacity of a data center and includes external storage devices such as storage arrays and built-in storage devices of servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read / write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0079] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage power, which can realize the centralized computing, storage, transmission, and application of information, presenting characteristics such as multi-element ubiquitous, intelligent and agile, secure and reliable, and green and low-carbon.
[0080] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing. With the emergence and popularization of new general technologies, the form of the new type of information infrastructure will be more diverse.
[0081] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.
[0082] The "general computing power" described in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0083] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform based on the large-scale deployment of dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.
[0084] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0085] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0086] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".
[0087] The "intelligent computing center" described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0088] The "computing power center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0089] The "supercomputing center" described in the present invention refers to: namely the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0090] The "computing power resources" described in the present invention refers to: technologies and facilities required for the development of the digital society with information computing, transmission, storage, and application capabilities, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0091] The "large image segmentation model" described in the present invention refers to: a deep learning model in the field of computer vision that can complete image segmentation tasks and has a large number of parameters.
[0092] Figure 1 There is shown a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center, asFigure 1 As shown, the method includes:
[0093] Step S1: Obtain the task to be processed, and determine the corresponding training data based on the task to be processed;
[0094] Step S2: Determine the hyperparameter space, where the hyperparameter space includes multiple hyperparameter groups;
[0095] Step S3: Perform the fine-tuning task of the large image segmentation model. The fine-tuning task of the large image segmentation model includes: based on the hyperparameter space, training data, and computing power resources of the intelligent computing center, perform distributed parallel fine-tuning on the large image segmentation model to obtain multiple fine-tuned large image segmentation models;
[0096] Among them, each fine-tuned large image segmentation model corresponds to a hyperparameter group in the hyperparameter space;
[0097] Step S4: Perform the evaluation task of the large image segmentation model. The evaluation task of the large image segmentation model includes: based on the preset evaluation metrics and the computing power resources of the intelligent computing center, perform distributed parallel evaluation on each fine-tuned large image segmentation model to obtain the corresponding evaluation results of each fine-tuned large image segmentation model;
[0098] Among them, the evaluation result is the value of the evaluation metric;
[0099] Step S5: Perform the determination task of the large image segmentation model. The determination task of the large image segmentation model includes: determine the best evaluation result among the evaluation results, determine the fine-tuned large image segmentation model corresponding to the best evaluation result as the target large image segmentation model, and record the hyperparameter group corresponding to the target large image segmentation model;
[0100] Among them, the best evaluation result is the evaluation result corresponding to the best value of the evaluation metric.
[0101] In step S1, the training set refers to the training data used to fine-tune the large image segmentation model, usually including a series of well-annotated images and corresponding segmentation labels. The sources of these images are diverse. For example, they can be from real-time acquisitions on industrial production lines, public data sets, synthetic data, or historical record data. The content of the training set should cover a variety of scenarios and conditions to ensure that the large image segmentation model can learn rich features and patterns, thereby improving its generalization ability in practical applications.
[0102] In one possible implementation, as Figure 2 shown, step S1 includes:
[0103] Step S11: Obtain the task to be processed, and determine the images to be annotated based on the task to be processed;
[0104] Step S12: Input the image to be labeled into the large image segmentation model for pre-labeling, and output the pre-labeled image;
[0105] Step S13: Verify the pre-labeled image based on the Labelme tool to obtain the verified image;
[0106] Step S14: Determine the verified image and the image to be labeled as the training data.
[0107] It should be noted that in the industrial quality inspection scenario, the system will obtain the task to be processed according to the current industrial quality inspection requirements. The task may involve the quality inspection of specific products or processes; according to the task requirements, the system will screen out the images that need to be labeled from the existing dataset. The images to be labeled will be input into the trained large image segmentation model for pre-labeling. The large image segmentation model will automatically identify the objects in the image and generate a preliminary segmentation result based on the features it has learned. This process can greatly reduce the workload of manual labeling and improve the labeling efficiency (in this step, manual labeling can also be combined); after the pre-labeling is completed, tools such as Labelme can be used to perform manual verification on the pre-labeled images. The verification process includes checking and making necessary modifications to the segmentation results output by the model to ensure the accuracy of the labeling, correct the possible errors of the large image segmentation model, and improve the labeling precision; finally, the verified images and the original images to be labeled will be integrated to form the final training dataset. Thus, high-quality training data can be constructed to ensure the performance improvement of the large image segmentation model.
[0108] In step S2, a hyperparameter space needs to be defined, which contains multiple hyperparameter groups. Each hyperparameter group contains different types of hyperparameters, and each type of hyperparameter has only one value in each group. There are different hyperparameter combinations between different hyperparameter groups, so that the performance of the large image segmentation model under different hyperparameter settings can be explored. And by systematically defining the hyperparameter space, the impact of different hyperparameter combinations on the performance of the large image segmentation model can be comprehensively evaluated, providing diversity and flexibility for the subsequent fine-tuning of the large image segmentation model.
[0109] In a possible implementation, the hyperparameters include at least one of the following: learning rate, batch size, number of epochs, optimizer, regularization parameter, non-maximum suppression (NMS) parameter.
[0110] In step S3, using the defined hyperparameter space and training set, with the computing power resources of the intelligent computing center, the large image segmentation model can be fine-tuned in a distributed parallel manner. Each fine-tuned large image segmentation model corresponds to a set of hyperparameters. Eventually, multiple fine-tuned large image segmentation models can be obtained. Among them, the computing power resources of the intelligent computing center are distributed computing power resources, and the distributed computing power resources include GPUs located on different nodes, with at least one GPU set on each node. On one GPU, only one fine-tuning task of the large image segmentation model is executed at a time. It should be noted that using the computing power resources of the intelligent computing center for distributed parallel fine-tuning can significantly shorten the training time of the large image segmentation model. At the same time, through the combination of different hyperparameters, diverse fine-tuned large image segmentation models can be generated, which is convenient for subsequent evaluation and selection of the best large image segmentation model.
[0111] In a specific application scenario, assuming there are 30 sets of hyperparameters, then there are 30 fine-tuning tasks for the large image segmentation model. Each task requires 1 GPU. The computing power resources of the intelligent computing center are as follows: currently, there are 10 idle and schedulable GPUs. Then 10 tasks can be executed in parallel. Using K8S, 10 Pods are started simultaneously. Each Pod is bound to 1 GPU, and 10 tasks are run simultaneously. And for each completed task, 1 GPU can be released for the next task. The total time consumption is approximately 3 rounds (10, 10, 10), and the total time consumption is 3 times that of a single task. Thus, the overall fine-tuning time is reduced, and the fine-tuning efficiency of the large image segmentation model is improved.
[0112] In step S4, evaluation metrics (such as accuracy, etc.) need to be used to evaluate each fine-tuned large image segmentation model. And with the computing power resources of the intelligent computing center, the evaluation of multiple large image segmentation models can be processed in a distributed parallel manner. Eventually, the evaluation results of each large image segmentation model are obtained. Similarly, the computing power resources of the intelligent computing center are distributed computing power resources, and the distributed computing power resources include GPUs located on different nodes, with at least one GPU set on each node. On one GPU, only one evaluation task of the large image segmentation model is executed at a time. Through the distributed parallel evaluation of multiple large image segmentation models, the performance metrics of each large image segmentation model can be quickly obtained, providing a basis for the subsequent selection of the large image segmentation model. And the distributed parallel evaluation improves the evaluation efficiency of the large image segmentation model, reduces the overall evaluation time, and enhances the optimization efficiency of the large image segmentation model.
[0113] In step S5, it is necessary to analyze the evaluation results, determine the best evaluation result, and determine the corresponding fine-tuned large image segmentation model as the target large image segmentation model. At the same time, record the information of the hyperparameter group corresponding to the target large image segmentation model. Thus, the fine-tuning evaluation task of the large image segmentation model is completed, and the large image segmentation model with the optimal performance (i.e., the target large image segmentation model) can be obtained in this fine-tuning evaluation.
[0114] Compared with the existing fine-tuning methods for large image segmentation models, the present invention adds the computing power resources of the intelligent computing center, which can significantly improve the computing performance and avoid the problem of resource exhaustion of a single server when processing dynamic and static memory. Moreover, through the distributed parallel fine-tuning and evaluation process, the training of the large image segmentation model can be carried out simultaneously on multiple nodes, thus greatly improving the computing efficiency.
[0115] In terms of scalability, the intelligent computing center allows flexible adjustment of the allocation of computing power resources to adapt to changing workloads. Therefore, the computing resources can be dynamically expanded or reduced according to the business requirements during the fine-tuning of the large image segmentation model, avoiding the waste of computing power resources.
[0116] In terms of the accuracy of the large image segmentation model, through the rich hyperparameter space and parallel fine-tuning strategy, it can be ensured that the best hyperparameter combination can be found, thereby improving the accuracy of the large image segmentation model.
[0117] In summary, through the computing power resources of the intelligent computing center, the distributed parallel fine-tuning and evaluation of the large image segmentation model are realized, significantly improving the computing performance and accuracy of the large image segmentation model, thus effectively meeting the needs of modern industrial quality inspection.
[0118] In a possible implementation, as Figure 3 shown, step S3 includes: step S31: Determine the fine-tuning strategy of the large image segmentation model, and based on the fine-tuning strategy of the large image segmentation model, execute the fine-tuning task of the large image segmentation model.
[0119] Among them, the fine-tuning strategy of the large image segmentation model includes at least one of the following:
[0120] Perform staged fine-tuning on the large image segmentation model, where the staged fine-tuning includes: fixing all current layers of the large image segmentation model; training the newly added layers; after the training of the newly added layers is completed, select the target layer from all layers according to the user's instruction, and jointly train the target layer and the newly added layers.
[0121] It should be noted that the staged fine-tuning is an effective strategy for optimizing large image segmentation models to meet the requirements of specific tasks. This process first fixes all the current layers of the large image segmentation model to ensure that the existing knowledge and feature extraction capabilities are not affected. Then, new layers are added to the large image segmentation model, and only these new layers are trained to enable them to learn features related to specific tasks. After completing the training of the new layers, the user can select the target layers to be fine-tuned according to the needs, and then jointly train these target layers and the newly added layers to achieve a better collaborative effect. Through this staged fine-tuning method, the large image segmentation model can effectively improve its performance on specific tasks while retaining its original performance, and at the same time reduce the training time and resource consumption.
[0122] Based on the image enhancement technology, the same sample image in the training data is represented at multiple different resolutions.
[0123] It should be noted that the fine-tuning strategy of the large image segmentation model can enhance the boundary clarity of the sample images in the training data by applying image enhancement technology. This method includes operations such as rotating, translating, scaling, and adjusting the contrast of the images, thereby generating diverse training samples. This not only increases the richness of the dataset but also strengthens the large image segmentation model's ability to learn boundary details, improving its accuracy and robustness in the segmentation task. By enhancing the boundary clarity of the images, the large image segmentation model can better identify and segment different objects, ultimately improving the overall performance. And it is especially applicable to the tasks to be processed that require high-precision boundaries.
[0124] Moreover, the fine-tuning strategy of the large image segmentation model can enhance the generalization ability of the large image segmentation model by representing the same sample image in the training data at multiple different resolutions. This method generates high and low different resolution versions of the image, enabling the large image segmentation model to learn features at multiple scales, thus better adapting to various possible input conditions. Images with different resolutions help the large image segmentation model capture different levels of details and context information, improving its ability to recognize object boundaries and details. Ultimately, this multi-scale training strategy helps improve the performance and robustness of the large image segmentation model in practical applications. And it is especially applicable to processing images with complex structures.
[0125] Use the AdamW optimizer to decay the learning rate in the later stage of the fine-tuning of the large image segmentation model.
[0126] It should be noted that in the fine-tuning strategy of the large image segmentation model, the AdamW optimizer is used, combined with weight decay to prevent overfitting, and learning rate decay is performed in the later stage of training, which can help the large image segmentation model to converge. The AdamW optimizer combines the advantages of adaptive learning rate and weight decay, and can automatically adjust the learning rate of each parameter during training, thereby improving the convergence speed and stability. In the later stage of fine-tuning, gradually decaying the learning rate helps the large image segmentation model to make more delicate adjustments when approaching the optimal solution, avoiding performance fluctuations caused by overly large updates. This strategy can improve the final performance of the large image segmentation model, enabling it to achieve higher accuracy and robustness on specific tasks.
[0127] Construct a mask loss function and a prompt loss function; fine-tune the large image segmentation model based on the mask loss function and the prompt loss function.
[0128] It should be noted that when fine-tuning the large image segmentation model, a mask loss function and a prompt loss function can be constructed to better optimize the performance of the large image segmentation model. The mask loss function is used to measure the difference between the segmentation mask generated by the large image segmentation model and the true label, ensuring that the large image segmentation model accurately identifies and segments the target area. The prompt loss function, on the other hand, is used to guide the large image segmentation model to learn under specific contexts or conditions. By optimizing the response to the input prompt, the adaptability of the large image segmentation model to specific tasks is enhanced. Combining these two loss functions can simultaneously focus on segmentation accuracy and task relevance during the fine-tuning process, thereby improving the overall performance of the large image segmentation model.
[0129] Use Dropout and / or Batch Normalization techniques to reduce overfitting during the fine-tuning process of the large image segmentation model.
[0130] It should be noted that in the fine-tuning process of the large image segmentation model, using Dropout (random dropout or stochastic inactivation) and Batch Normalization (batch normalization) techniques can effectively reduce the overfitting phenomenon. Dropout randomly discards a certain proportion of neurons during training, enabling the large image segmentation model to learn different combinations of features in each iteration, thereby enhancing the generalization ability of the large image segmentation model and reducing its dependence on specific features. On the other hand, Batch Normalization standardizes the input of each layer, reduces internal covariate shift, accelerates the training process, and improves the stability of the large image segmentation model. Combining these two techniques can effectively control the model complexity during the fine-tuning stage, improve its performance on unseen data, and reduce the risk of overfitting.
[0131] Use the L2 regularization technique to improve the generalization of the large image segmentation model during the fine-tuning process.
[0132] It should be noted that during the fine-tuning process of the large image segmentation model, using the L2 regularization technique can effectively constrain the hyperparameters and prevent the model from overfitting. L2 regularization promotes the large image segmentation model to learn smaller weight values by adding a penalty term proportional to the sum of the squares of the weights of the large image segmentation model to the loss function, thereby reducing the complexity of the large image segmentation model. This constraint mechanism makes the large image segmentation model more inclined to choose a smoother decision boundary during training, avoiding being overly sensitive to noise and outliers in the training data. By introducing L2 regularization, the fine-tuned large image segmentation model can not only perform well on the training set but also maintain good generalization ability on unseen data, improving its reliability in practical applications.
[0133] In a possible implementation, after step S5, the method further includes: step S6: performing distributed parallel inference based on the target large image segmentation model and the computing power resources of the intelligent computing center; wherein, the target large image segmentation model uses the ONNX inference framework.
[0134] It should be noted that the ONNX (Open Neural Network Exchange) format is an open neural network exchange format used for model conversion and communication between different deep learning frameworks. Its design goal is to provide an intermediate representation format that allows users to seamlessly convert models between different deep learning training and inference frameworks. By using the ONNX inference framework and distributed parallel inference, the target large image segmentation model can achieve efficient, fast, and accurate image segmentation tasks with the support of the high-performance computing power of the intelligent computing center, not only improving the inference speed and resource utilization rate but also maintaining the accuracy and compatibility of the large image segmentation model, which is suitable for large-scale and high-complexity image processing scenarios.
[0135] Now, from the perspective of specific application scenarios, a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center shown in the present invention is introduced as a whole.
[0136] The large image segmentation model is deployed in VKS (Elastic Container Cluster).
[0137] It should be noted that VKS is an elastic container cluster that optimizes resource utilization, reduces hardware and maintenance costs through the cluster, and can dynamically adjust resource allocation according to the actual load to ensure that the service always remains in the best state.
[0138] 1. Pre - deployment preparation: Apply for Vcluster resources and obtain the config configuration file; Apply for an image repository instance to host the image segmentation large - model image; Apply for S3 storage for synchronizing and storing the image segmentation large - model files.
[0139] It should be noted that Vcluster is short for virtual cluster, which is a fully functional, lightweight, and well - isolated Kubernetes cluster running on top of a regular Kubernetes cluster.
[0140] 2. Environment setup and image preparation: Prepare the development environment, including: Python version, Python package version, CUDA version. Pay attention to the compatibility between CUDA, Python, and PyTorch versions. Then, package the environment into a Docker image and push the built Docker image to the image repository.
[0141] 3. Training data preparation: Prepare sufficient labeled data in this industrial quality inspection field as training data, and perform pre - annotation using a pre - trained image segmentation large - model. Then, experts in the quality inspection field can use the labelme annotation tool for annotation verification to improve the annotation efficiency of training images, and store the training samples in a distributed database.
[0142] 4. Model fine - tuning:
[0143] (1) The techniques for fine - tuning the image segmentation large - model include the following points: 1. Perform staged fine - tuning. First, fix all pre - trained layers and only train the newly added layers, and then release some specific layers for joint training; 2. For tasks that require high - precision boundaries, special attention should be paid to how to retain and improve boundary details; 3. When processing images with complex structures, explore improving the performance of the image segmentation large - model by providing input images with different resolutions; 4. Use the AdamW optimizer, combined with weight decay to prevent overfitting, and perform learning rate decay in the later stage of training to help the image segmentation large - model converge; 5. Combine mask loss and prompt loss to improve the overall performance of the image segmentation large - model; 6. Other hyperparameter optimization strategies, such as using techniques like Dropout and Batch Normalization to reduce overfitting and promote a more stable training process; 7. Appropriately apply L2 regularization (weight decay) to constrain the parameters of the image segmentation large - model and avoid overfitting the training data.
[0144] (2) Build the hyperparameter optimization space for the image segmentation large - model: For example: The key hyperparameters are: learning rate, batch size, number of iterations, optimizer, regularization parameter, non - maximum suppression parameter to build the hyperparameter optimization space.
[0145] (3)Perform hyperparameter tuning: Configure deployment parameters through the VKS configuration file, such as the computing resources, images, memory, and mounted directories required for the fine-tuning task. Based on the fine-tuning technology in (1) and the hyperparameter optimization space in (2), allocate the hyperparameter optimization task to the computing power resources of the intelligent computing center for parallel fine-tuning. After the fine-tuning is completed, perform an evaluation, select the optimal large image segmentation model according to the evaluation metrics, save the corresponding hyperparameter combination, and store the optimal large image segmentation model in S3.
[0146] 5. Deployment of the large image segmentation model: Preparing the activation and running environment for the large image segmentation model inference service includes the following steps: First, load the specified model weights in the model repository; Second, perform pruning and distillation operations on the model to optimize the model structure and reduce the computational complexity. These operations help to achieve efficient distributed deployment, improve the model inference performance, and enhance the generalization ability of the model.
[0147] 6. Model inference: Start the model inference service, use the efficient ONNX inference framework, and perform distributed parallel inference. The inference speed can meet the real-time requirements of quality inspection, and perform post-processing on the inference results to meet the performance required for quality inspection and output the indicators concerned by quality inspection.
[0148] In summary, the method for fine-tuning the large image segmentation model based on the computing power of the intelligent computing center has the following advantages:
[0149] 1. High fine-tuning performance: The VCluster of the intelligent computing center deeply integrates the new generation of cloud native technologies, provides a high-performance Kubernetes container cluster management service with containers as the core, provides a stable, reliable and high-performance development environment for the large image segmentation model, and realizes the strong combination of the use and operation of the large image segmentation model; High availability and load balancing for autonomous container orchestration, realizing resource sharing and load balancing, ensuring the high availability and stability of applications; In the multi-machine distributed architecture, through the network for large-scale parallel computing, the iteration efficiency of the large image segmentation model is improved.
[0150] 2. High scalability: The distributed solution based on the cluster or multi-cloud environment of the intelligent computing center VCluster can more flexibly adapt to changing workloads, effectively divide the workload, maintain gradient consistency, and minimize communication overhead, improve the fine-tuning and inference performance of the large image segmentation model, dynamically adjust resources according to actual needs, and avoid over-allocation and waste of resources.
[0151] 3. High model accuracy: Through a carefully designed fine-tuning strategy, and based on the labeled data in a specific domain, fine-tuning the large image segmentation model can effectively improve the accuracy of the large image segmentation model and better apply it to specific application scenarios.
[0152] In summary, based on the computing power resources of the intelligent computing center, rapid and accurate fine-tuning of the large image segmentation model can be achieved, which helps the full-process industrial quality inspection application practice from data annotation to the deployment of the large image segmentation model. Using the large image segmentation model to collect the production status and conduct quality inspection, compared with manual or traditional quality inspection methods, it can significantly improve the quality inspection accuracy and efficiency.
[0153] Figure 4 There is shown a device for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center, as Figure 4 shown. The device 40 includes:
[0154] An acquisition module 401, configured to acquire a task to be processed and determine corresponding training data based on the task to be processed;
[0155] An execution module 402, configured to determine a hyperparameter space, where the hyperparameter space includes multiple hyperparameter groups;
[0156] Execute the fine-tuning task of the large image segmentation model. The fine-tuning task of the large image segmentation model includes: based on the hyperparameter space, training data, and the computing power resources of the intelligent computing center, performing distributed parallel fine-tuning on the large image segmentation model to obtain multiple fine-tuned large image segmentation models, where each fine-tuned large image segmentation model corresponds to one hyperparameter group in the hyperparameter space;
[0157] Execute the evaluation task of the large image segmentation model. The evaluation task of the large image segmentation model includes: based on a preset evaluation metric and the computing power resources of the intelligent computing center, performing distributed parallel evaluation on each fine-tuned large image segmentation model to obtain the evaluation result corresponding to each fine-tuned large image segmentation model, and the evaluation result is the value of the evaluation metric;
[0158] Execute the determination task of the large image segmentation model. The determination task of the large image segmentation model includes: determining the best evaluation result in the evaluation results, determining the fine-tuned large image segmentation model corresponding to the best evaluation result as the target large image segmentation model, and recording the hyperparameter group corresponding to the target large image segmentation model; where the best evaluation result is the evaluation result corresponding to the best value of the evaluation metric.
[0159] In a possible implementation manner, the execution module 402 is further configured to determine a fine-tuning strategy for the large image segmentation model and execute the fine-tuning task of the large image segmentation model based on the fine-tuning strategy for the large image segmentation model;
[0160] where the fine-tuning strategy for the large image segmentation model includes at least one of the following:
[0161] Perform staged fine-tuning on the large image segmentation model, where the staged fine-tuning includes: fixing all current layers of the large image segmentation model; training the newly added layers; after the training of the newly added layers is completed, select target layers from all layers according to the user's instructions, and jointly train the target layers and the newly added layers;
[0162] Based on image enhancement technology, represent the same sample image in the training data at multiple different resolutions;
[0163] Use the AdamW optimizer to decay the learning rate in the later stage of fine-tuning the large image segmentation model;
[0164] Construct a mask loss function and a prompt loss function; fine-tune the large image segmentation model based on the mask loss function and the prompt loss function;
[0165] Use Dropout and / or Batch Normalization techniques to reduce overfitting during the fine-tuning process of the large image segmentation model;
[0166] Use L2 regularization technology to improve the generalization of the large image segmentation model during the fine-tuning process.
[0167] In a possible implementation, the acquisition module 401 is further configured to acquire a task to be processed, and determine an image to be annotated based on the task to be processed;
[0168] Input the image to be annotated into the large image segmentation model for pre-annotation, and output the pre-annotated image;
[0169] Verify the pre-annotated image based on the Labelme tool to obtain the verified image;
[0170] Determine the verified image and the image to be annotated as training data.
[0171] In a possible implementation, the execution module 402 is further configured to perform distributed parallel inference based on the target image segmentation model and the computing power resources of the intelligent computing center after the task determined by the image segmentation model is executed; where the target image segmentation model uses the ONNX inference framework.
[0172] In a possible implementation, the hyperparameters include at least one of the following: learning rate, batch size, number of iterations, optimizer, regularization parameter, non-maximum suppression parameter.
[0173] In a possible implementation, the computing power resources of the intelligent computing center are distributed computing power resources, and the distributed computing power resources include: graphics processing units (GPUs) located on different nodes, and at least one GPU is set on each node; on one GPU, only one fine-tuning task or evaluation task of the large image segmentation model is executed at a time.
[0174] In summary, compared with the existing fine-tuning methods of large image segmentation models, the present invention adds the computing power resources of the intelligent computing center, which can significantly improve the computing performance and avoid the problem of resource exhaustion of a single server when processing dynamic and static memory; and through the distributed parallel fine-tuning and evaluation process, the training of the large image segmentation model can be carried out simultaneously on multiple nodes, thus greatly improving the computing efficiency.
[0175] In terms of scalability, the intelligent computing center allows flexible adjustment of the allocation of computing power resources to adapt to changing workloads. Therefore, the computing resources can be dynamically expanded or reduced according to the business requirements during the fine-tuning of the large image segmentation model, avoiding waste of computing power resources.
[0176] In terms of the accuracy of the large image segmentation model, through the rich hyperparameter space and parallel fine-tuning strategy, it can be ensured that the best combination of hyperparameters can be found, thereby improving the accuracy of the large image segmentation model.
[0177] In summary, through the computing power resources of the intelligent computing center, the distributed parallel fine-tuning and evaluation of the large image segmentation model are realized, significantly improving the computing performance and accuracy of the large image segmentation model, thus effectively meeting the needs of modern industrial quality inspection.
[0178] Please refer to Figure 5 , the present invention also provides an electronic device 50, including a processor 501, a memory 502, and a computer program stored on the memory 502 and executable on the processor 501. When the computer program is executed by the processor 501, it implements the steps of the above method for fine-tuning a large image segmentation model based on the computing power of the intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0179] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above method for fine-tuning a large image segmentation model based on the computing power of the intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0180] The present invention also provides a computer program product, including computer instructions which, when executed by a processor, implement the steps of the above method for fine-tuning an image segmentation large model based on the computing power of an intelligent computing center, and can achieve the same technical effects. To avoid repetition, details are not described herein again.
[0181] It should be noted that in this document, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0182] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the method described in the present invention.
[0183] The present invention has been described above with reference to the accompanying drawings, but the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all of them fall within the protection scope of the present invention.
Claims
1. A method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center, characterized in that: The method comprises: Step S1: Obtain a task to be processed, and determine corresponding training data based on the task to be processed; Step S2: determining a hyperparameter space, wherein the hyperparameter space includes a plurality of hyperparameter groups; Step S3: executing a large image segmentation model fine-tuning task, wherein the large image segmentation model fine-tuning task includes: based on the hyperparameter space, the training data and the computing power resources of the intelligent computing center, performing distributed parallel fine-tuning on the large image segmentation model to obtain a plurality of fine-tuned large image segmentation models, wherein each fine-tuned large image segmentation model corresponds to a hyperparameter group in the hyperparameter space; Step S4: executing an image segmentation large model evaluation task, wherein the image segmentation large model evaluation task includes: based on a preset evaluation index and the computing power resources of the intelligent computing center, performing a distributed parallel evaluation on each fine-tuned image segmentation large model, and obtaining an evaluation result corresponding to each fine-tuned image segmentation large model, wherein the evaluation result is the value of the evaluation index; Step S5: Execute the image segmentation large model determination task, the image segmentation large model determination task includes: determining the best evaluation result in the evaluation results, determining the fine-tuned image segmentation large model corresponding to the best evaluation result as the target image segmentation large model, and recording the hyperparameter group corresponding to the target image segmentation large model; wherein the best evaluation result is the evaluation result corresponding to the best value of the evaluation indicator.
2. The method according to claim 1, characterized in that Step S3 includes: Step S31: determining a fine-tuning strategy for a large image segmentation model, and executing the fine-tuning task for the large image segmentation model based on the fine-tuning strategy for the large image segmentation model; The image segmentation large model fine-tuning strategy includes at least one of the following: The image segmentation model is fine-tuned in stages, wherein the fine-tuning in stages includes: fixing all current layers of the image segmentation model; training the newly added layers; after the training of the newly added layers is completed, selecting a target layer from all the layers according to a user's instruction, and jointly training the target layer and the newly added layers; Based on the image enhancement technology, the same sample image in the training data is represented with multiple different resolutions; Use the AdamW optimizer to decay the learning rate in the later stages of fine-tuning the large image segmentation model; Constructing a mask loss function and a prompt word loss function; fine-tuning the image segmentation model based on the mask loss function and the prompt word loss function; Using Dropout and / or Batch Normalization techniques to reduce overfitting during fine-tuning of the large image segmentation model; L2 regularization technology is used to improve the generalization of the image segmentation model during the fine-tuning process of the image segmentation model.
3. The method according to claim 1, characterized in that Step S1 includes: Step S11: obtaining a task to be processed, and determining an image to be annotated based on the task to be processed; Step S12: inputting the image to be annotated into the image segmentation model for pre-annotation, and outputting the pre-annotated image; Step S13: verifying the pre-labeled image based on the Labelme tool to obtain a verified image; Step S14: determining the verified image and the image to be labeled as the training data.
4. The method according to claim 1, characterized in that After step S5, the method further comprises: Step S6: Perform distributed parallel reasoning based on the target image segmentation large model and the computing resources of the intelligent computing center; wherein the target image segmentation large model uses the ONNX reasoning framework.
5. The method according to claim 1, characterized in that The hyperparameters include at least one of the following: learning rate, batch size, number of iterations, optimizer, regularization parameter, and non-maximum suppression parameter.
6. The method according to any one of claims 1 to 5, characterized in that The computing power resources of the intelligent computing center are distributed computing power resources, which include: graphics processing units (GPUs) located on different nodes, with at least one GPU set on each node; on one GPU, only one image segmentation large model fine-tuning task or one image segmentation large model evaluation task is executed at a time.
7. A device for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center, characterized in that: The device comprises: An acquisition module, used for acquiring a task to be processed and determining corresponding training data based on the task to be processed; An execution module, configured to determine a hyperparameter space, wherein the hyperparameter space includes a plurality of hyperparameter groups; Execute a large image segmentation model fine-tuning task, the large image segmentation model fine-tuning task comprising: based on the hyperparameter space, the training data and the computing power resources of the intelligent computing center, performing distributed parallel fine-tuning on the large image segmentation model to obtain a plurality of fine-tuned large image segmentation models, wherein each fine-tuned large image segmentation model corresponds to a hyperparameter group in the hyperparameter space; Executing an image segmentation large model evaluation task, the image segmentation large model evaluation task comprising: based on a preset evaluation index and the computing power resources of the intelligent computing center, performing a distributed parallel evaluation on each fine-tuned image segmentation large model, and obtaining an evaluation result corresponding to each fine-tuned image segmentation large model, wherein the evaluation result is the value of the evaluation index; Execute an image segmentation large model determination task, the image segmentation large model determination task comprising: determining the best evaluation result in the evaluation results, determining the fine-tuned image segmentation large model corresponding to the best evaluation result as the target image segmentation large model, and recording the hyperparameter group corresponding to the target image segmentation large model; wherein the best evaluation result is the evaluation result corresponding to the best value of the evaluation indicator.
8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by the processor, implements the steps of a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center as described in any one of claims 1 to 6.
10. A computer program product, characterized in that It includes computer instructions, which, when executed by the processor, implement the steps of a method for fine-tuning a large image segmentation model based on the computing power of an intelligent computing center as described in any one of claims 1 to 6.
Citation Information
Cited By
Automatic image classification and packaging system and method based on container technology
CN120563942A