Large model reasoning acceleration system for industrial quality inspection

Through the large-model inference acceleration system, combined with data pruning and quantization technology, the industrial quality inspection process is optimized, solving the problems of low efficiency and incomplete detection in traditional testing, and realizing efficient and accurate detection of complex products, which is suitable for edge hardware computing devices.

CN120806153APending Publication Date: 2025-10-17JIANGXI INST OF FASHION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510931966.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional industrial quality inspection methods are inefficient and lack comprehensive testing capabilities for complex products and diverse defects.

Method used

A large-model inference acceleration system is used, including industrial data acquisition, processing, fusion, update and output modules, combined with data pruning, quantization and hardware acceleration technologies to optimize the use of neurons and computing resources of large models.

Benefits of technology

It achieves efficient and accurate quality inspection of complex products, reduces computational complexity and storage requirements, improves inspection efficiency and accuracy, is suitable for edge hardware computing devices, and reduces hardware costs and energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806153A_ABST
    Figure CN120806153A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial detection, and discloses an industrial quality inspection-oriented large model reasoning acceleration system, which combines an industrial data acquisition module, an industrial data processing module, an industrial data fusion module, an industrial data updating module, a model reasoning acceleration module and a model reasoning acceleration module. In a target industrial quality inspection scene, quality detection can be efficiently and accurately carried out on various industrial data, and comprehensive detection can be carried out on the quality problem of a complex product.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of industrial detection, and particularly relates to a large model inference acceleration system for industrial quality inspection. BACKGROUND

[0002] Industrial quality inspection is to detect the appearance defects, production process, appearance image, etc. of products in the industrial production process to ensure that the products meet the design standards.

[0003] In the related art, the quality of products is mainly detected by manual detection or automatic detection. The manual detection has high flexibility, but has low detection efficiency and strong subjectivity, and the manual fatigue may cause misjudgment. The automatic detection mainly detects the quality of products by automatic detection equipment in a specific single manner, and lacks comprehensive detection capability for complex products and diversified defects. SUMMARY

[0004] Therefore, the present application provides a large model inference acceleration system for industrial quality inspection to solve the problems of low detection efficiency and incomplete detection of the traditional detection method.

[0005] In a first aspect, the present application provides a large model inference acceleration system for industrial quality inspection, which comprises:

[0006] An industrial data acquisition module is configured to acquire a plurality of industrial data in a target industrial quality inspection scene.

[0007] An industrial data processing module is configured to process the plurality of industrial data to obtain processing results of the plurality of industrial data.

[0008] An industrial data fusion module is configured to extract and fuse features of the processing results of the plurality of industrial data to obtain fusion results of the plurality of industrial data.

[0009] An industrial data updating module is configured to update weight parameters of a large model applied to the plurality of industrial data, so that the inference accuracy of the large model meets the requirements of industrial quality inspection.

[0010] A model inference acceleration module is configured to perform data pruning, data quantization and data acceleration on the large model applied to the plurality of industrial data according to the fusion results of the plurality of industrial data, to obtain inference acceleration results of the plurality of industrial data.

[0011] An industrial data output module is configured to output the inference acceleration results of the plurality of industrial data to a target application interface for display.

[0012] The embodiment can realize efficient and accurate quality detection of various industrial data in the target industrial quality inspection scene, and can also realize comprehensive detection of quality problems of complex products.

[0013] In some optional embodiments, the industrial data acquisition module comprises:

[0014] The industrial camera is configured to acquire video image data in the target industrial quality inspection scene.

[0015] The industrial camera can acquire video image data.

[0016] In some optional embodiments, the industrial data acquisition module comprises:

[0017] The temperature sensor is configured to acquire environmental temperature data in the target industrial quality inspection scene.

[0018] The vibration sensor is configured to acquire vibration data of the industrial equipment in the target industrial quality inspection scene.

[0019] The pressure sensor is configured to acquire pressure data of the industrial equipment in the target industrial quality inspection scene.

[0020] The above-mentioned various types of sensors can acquire environmental temperature data, vibration data of the industrial equipment, and pressure data of the industrial equipment in the target industrial quality inspection scene.

[0021] In some optional embodiments, the industrial data processing module comprises:

[0022] The industrial data processing unit is configured to clean, normalize, and align the various industrial data.

[0023] The industrial data processing unit can process various business data to ensure data quality.

[0024] In some optional embodiments, the industrial data fusion module comprises:

[0025] The feature extraction unit is configured to extract target features corresponding to each type of industrial data using a convolutional neural network model.

[0026] The feature fusion unit is configured to fuse the various industrial data according to the target features corresponding to each type of industrial data using a cross-modal attention module to obtain a fusion result of the various industrial data.

[0027] The embodiment can obtain the correlation between various industrial data, and further accurately infer the various industrial data.

[0028] In some optional embodiments, the model inference acceleration module comprises:

[0029] The data pruning unit is configured to perform a pruning action on neurons of the large model applied to the various industrial data according to a preset weight, monitor a weight change trend of the large model in real time, calculate an importance score of a task of the large model according to the weight change trend of the large model, and perform the pruning action on the neurons of the large model according to the importance score of the task of the large model, wherein the preset weight is determined according to a task type and a task priority of the task of the large model.

[0030] The embodiment can reduce the data storage amount by pruning the neurons of the large model through the data pruning unit, and improve the data processing capability of the various industrial data.

[0031] In some optional embodiments, the data pruning unit comprises:

[0032] The first pruning subunit is configured to perform a pruning action on a neuron of the large model when a weight parameter of the neuron of the large model is less than a preset weight during processing of a task of the large model by the large model.

[0033] The data calculation subunit is configured to calculate an importance score of the task of the large model by using the following importance score function formula.

[0034]

[0035] wherein I is the importance score of the task of the large model, a is a hyperparameter balance amplitude of the large model, β is a weight change trend of the large model, and γ is a curvature. i is a weight importance of the large model, is a weight parameter generated at a t time of an i training of the large model, is a weight parameter generated at a t-1 time of the i training, and ω i is a maximum weight parameter generated at the i training of the large model, and T is a training period.

[0036] The second pruning subunit is configured to perform a pruning action on the neuron of the large model when the importance score of the task of the large model is less than or equal to a preset score.

[0037] ​The first pruning unit prunes the neurons whose weight parameters are less than the preset weight, and the pruning action is performed on the neurons of the large model when the importance score of the large model task is less than or equal to the preset score, thereby enhancing the data processing capability of the large model task.

[0038] In some optional embodiments, the model inference acceleration module comprises a data quantization unit, which comprises:

[0039] The level determination subunit determines the current computing level of the large model task according to the task type of the large model task.

[0040] The first quantization subunit quantizes the pruned large model task according to the quantization precision of 16-bit floating point numbers when the current computing level of the large model task is the first level.

[0041] The second quantization subunit quantizes the pruned large model task according to the quantization precision of 8-bit floating point numbers when the current computing level of the large model task is the second level.

[0042] The third quantization subunit quantizes the pruned large model task according to the quantization precision of 4-bit floating point numbers when the current computing level of the large model task is the third level.

[0043] The data compensation subunit monitors the current error of the large model in real time during the quantization of the large model task, and performs a compensation action on the large model when the current error of the large model is greater than a preset error.

[0044] In the above manner, the data quantization unit uses an elastic quantization method to dynamically select the number of quantization bits, thereby enhancing the data processing capability of the large model task.

[0045] In some optional embodiments, the model inference acceleration module comprises:

[0046] The hardware acceleration subunit uses application specific integrated circuit chips or field programmable gate array chips to accelerate the inference acceleration results of various industrial data in parallel.

[0047] The hardware acceleration subunit is advantageous in improving the data processing speed.

[0048] In some optional embodiments, the industrial data updating module comprises:

[0049] The data updating unit updates the large model applied to various industrial data according to the sample data of the various industrial data at a preset period.

[0050] The embodiment updates the large model applied to various industrial data in the above manner, and is beneficial to enhancing the accuracy of the large model. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings required to be used in the specific embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0052] Figure 1 is a large model reasoning acceleration system for industrial quality inspection according to an embodiment of the present application;

[0053] Figure 2 is another large model reasoning acceleration system for industrial quality inspection according to an embodiment of the present application. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical scheme and advantages of the embodiments of the present application more clear, the technical scheme in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0055] According to the present embodiment, a large model reasoning acceleration system for industrial quality inspection is provided, as shown in Figure 1 including: an industrial data acquisition module 11, an industrial data processing module 12, an industrial data fusion module 13, an industrial data updating module 14, a model reasoning acceleration module 15 and an industrial data output module 16.

[0056] The industrial data acquisition module is configured to acquire various industrial data in a target industrial quality inspection scene.

[0057] In some optional embodiments, the industrial data acquisition module includes:

[0058] The industrial camera is configured to acquire video image data in the target industrial quality inspection scene.

[0059] The industrial camera in the present embodiment adopts an industrial camera with a resolution ≥12MP and a frame rate ≥60fps, thereby realizing high-speed acquisition of video images. The number of industrial cameras can be multiple. In the present embodiment, the number of industrial cameras can be multiple. For example, the present embodiment acquires video images of the surface of a part through the industrial camera.

[0060] In some optional embodiments, in Figure 2 the industrial data acquisition module 11 further comprises:

[0061] a temperature sensor 111 configured to acquire environmental temperature data in the target industrial quality inspection scene.

[0062] a vibration sensor 112 configured to acquire vibration data of the industrial equipment in the target industrial quality inspection scene.

[0063] a pressure sensor 113 configured to acquire pressure data of the industrial equipment in the target industrial quality inspection scene.

[0064] Therefore, the industrial quality inspection-oriented large model inference acceleration system in the embodiment can acquire various industrial data such as environmental temperature data, vibration data of the industrial equipment, and pressure data of the industrial equipment in the target industrial quality inspection scene through the industrial data acquisition module.

[0065] the industrial data processing module is configured to process the various industrial data to obtain processing results of the various industrial data.

[0066] In a specific example, as shown in Figure 2 the industrial data processing module 12 comprises:

[0067] an industrial data processing unit 121 configured to perform cleaning, normalization, and alignment processing on the various industrial data.

[0068] For example, the industrial data processing unit is configured to remove abnormal data such as noise and missing values from the various industrial data, normalize the various industrial data to a uniform scale, and then align the various industrial data after the normalization processing by using timestamps or spatial coordinates.

[0069] The embodiment performs cleaning, normalization, and alignment processing on the various industrial data, thereby ensuring data quality and providing reliable input for subsequent industrial data fusion and inference.

[0070] the industrial data fusion module is configured to perform feature extraction and feature fusion on the processing results of the various industrial data to obtain fusion results of the various industrial data.

[0071] In a specific example, as shown in Figure 2 the industrial data fusion module 13 comprises:

[0072] a feature extraction unit 131 configured to extract target features corresponding to each type of industrial data by using a convolutional neural network model.

[0073] The feature fusion unit 132 is configured to fuse the multiple industrial data by using the cross-modal attention module according to the target features corresponding to each kind of industrial data, to obtain a fusion result of the multiple industrial data.

[0074] For example, when the industrial data is video image data, the target features corresponding to the video image data are extracted by using a convolutional neural network (CNN), and the target features can be image features.

[0075] For example, when the industrial data is environmental temperature data, vibration data and pressure data, the time sequence features of the data are extracted by using a long short-term memory (LSTM).

[0076] Further, the cross-modal attention module is used to assign weight parameters to the target features corresponding to each kind of industrial data, to obtain the fusion result of the multiple industrial data.

[0077] Since a large model usually contains a large number of parameters, the inference calculation of the model needs strong computing resources. If the important levels of different neurons of the large model are not considered, the unimportant neuron data is compressed together, which leads to a large resource occupation and affects the data calculation efficiency.

[0078] The industrial data updating module is configured to update the weight parameters of the large model applied to the multiple industrial data, so that the inference accuracy of the large model meets the requirements of industrial quality inspection.

[0079] The industrial data updating module is configured to update the weight parameters of the large model applied to the multiple industrial data, so that the inference accuracy of the large model meets the requirements of industrial quality inspection.

[0080] Therefore, in Figure 2 , the model inference acceleration module is further improved, and the model inference module 15 includes a data pruning unit 151 and a data quantization unit 152.

[0081] In some specific embodiments, in Figure 2 , the model inference acceleration module 15 includes:

[0082] The data pruning unit 151 is configured to perform a pruning action on the neurons of the large model applied to the multiple industrial data according to a preset weight, and to monitor the weight change trend of the large model in real time, and to calculate the importance score of the large model task according to the weight change trend of the large model, and to perform a pruning action on the neurons of the large model according to the importance score of the large model task, wherein the preset weight is determined according to the task type and the task priority of the large model task.

[0083] In a specific example, inFigure 2 In the embodiment, the data pruning unit 151 includes:

[0084] The weight setting subunit 1511 is used to determine the importance of the large model task according to the task type and task priority of the large model task, and set the preset weight according to the importance of the large model task.

[0085] The first pruning subunit 1512 is configured to prune the neurons of the large model when the weight parameter of the neurons of the large model is less than a preset weight during the process of the large model processing the large model task;

[0086] The data calculation subunit 1513 is used to calculate the importance score of the large model task using the following importance score function formula;

[0087]

[0088] Among them, I i is the importance score of the large model task, α is the hyperparameter balance amplitude of the large model, β is the weight change trend of the large model, and γ is the curvature. is the weight importance of the large model, is the weight parameter generated during the i-th training of the large model at time t, is the weight parameter generated during the i-th training at time t-1, ω i is the final weight parameter generated during the i-th training of the large model, and T is the training cycle;

[0089] The second pruning subunit 1514 is used to perform pruning on the neurons of the large model when the importance score of the large model task is less than or equal to a preset score.

[0090] This embodiment uses the first and second pruning subunits in the model inference acceleration module, and utilizes the importance scoring function formula to determine the importance of neurons in the large model by considering not only the hyperparameter balance range of the large model but also the weight change trend of the large model. For example, during training, neurons with smaller weight changes and lower hyperparameter balance ranges are prioritized for pruning. During the pruning process, the second pruning subunit also needs to monitor the accuracy changes of the model in real time and immediately stop the pruning operation if the accuracy drops beyond the allowable range.

[0091] In a specific example, the data quantization precision includes 16-bit floating point number quantization precision, 8-bit floating point number quantization precision, and 4-bit floating point number quantization precision.

[0092] In some specific embodiments, Figure 2 In the embodiment, the model inference acceleration module 15 includes a data quantization unit 152, which includes:

[0093] The level determination subunit 1521 is configured to determine the current computing level of the large model task according to the task type of the large model task.

[0094] The first quantization subunit 1522 is configured to quantize the pruned large model task according to the quantization precision of 16-bit floating point numbers when the current computing level of the large model task is the first level.

[0095] The second quantization subunit 1523 is configured to quantize the pruned large model task according to the quantization precision of 8-bit floating point numbers when the current computing level of the large model task is the second level.

[0096] The third quantization subunit 1524 is configured to quantize the pruned large model task according to the quantization precision of 4-bit floating point numbers when the current computing level of the large model task is the third level.

[0097] The data compensation subunit 1525 is configured to monitor the current error of the large model in real time during the quantization of the large model task, and perform a compensation action on the large model when the current error of the large model is greater than a preset error.

[0098] In this embodiment, the data quantization unit adopts an elastic quantization method. For a large model task with strong computing capability and high accuracy requirement, a higher bit quantization is adopted. For example, when the current computing level of the large model task is the first level, the large model task is quantized according to the quantization precision of 16-bit floating point numbers. For a task with limited computing capability and relatively low accuracy requirement, a lower bit quantization is adopted. For example, when the current computing level of the large model task is the fourth level, a lower bit quantization is adopted, and the large model task is quantized according to the quantization precision of 4-bit floating point numbers.

[0099] This embodiment is applicable to an edge computing system of a large model. The data pruning unit and the data quantization unit in the model inference acceleration module prune, adaptively quantize and adaptively compensate the neurons of the large model, which is beneficial to improving the intelligent processing capability of various industrial data. Moreover, the data pruning unit and the data quantization unit prune and adaptively quantize and adaptively compensate the neurons of the large model, which can realize efficient deployment of the large model on a resource-limited edge hardware computing device, significantly improve the intelligent level of the edge hardware computing device, and solve the contradiction between model accuracy and resource occupation, network transmission delay and the like in the traditional way, thereby enhancing the practicability of the computing development of the edge hardware computing device.

[0100] The embodiment monitors the current error of the large model in real time in the process of quantizing the large model task, and performs a compensation action on the large model when the current error of the large model is greater than a preset error. For example, in a speech recognition task, the word error rate may be as high as 15% when quantized with a traditional 8-bit floating-point quantization precision. However, by using the compensation mechanism, the word error rate can be reduced to within 10% under the same quantization bit, thereby improving the accuracy of the quantized model.

[0101] In some specific embodiments, in Figure 2 The model inference acceleration module 15 comprises:

[0102] The hardware acceleration unit 153 uses an application-specific integrated circuit (ASIC) chip or a field-programmable gate array (FPGA) chip to accelerate the inference acceleration results of multiple industrial data in parallel.

[0103] For example, the embodiment uses an application-specific integrated circuit (ASIC) chip or a field-programmable gate array (FPGA) chip to accelerate the inference acceleration results of multiple industrial data in parallel, greatly shortening the inference time. Taking product quality inspection on a high-speed pipeline as an example, it used to take several seconds to complete one multiple industrial data inference detection. After using the embodiment, the inference time can be shortened to milliseconds, meeting the real-time requirement of industrial production scenarios, greatly improving the product detection efficiency, and avoiding production stagnation caused by detection delay.

[0104] The industrial data output module is configured to output the inference acceleration results of the multiple industrial data to a target application interface for display.

[0105] The target application interface in the embodiment can be a human-computer interface, which displays the inference acceleration results of the multiple industrial data in the target industrial quality inspection scenario in real time.

[0106] The large model inference acceleration system for industrial quality inspection in the embodiment reduces the computational complexity and storage requirements of the multiple industrial data by performing pruning actions on the neurons of the large model applied to the multiple industrial data through the data pruning unit and performing data quantization on the large model task after pruning according to different quantization precision. The optimized large model has a significantly reduced computational load when running in the edge inference engine, which significantly improves the inference speed. Experimental data shows that the large model before optimization can process dozens of samples per second on the same edge hardware computing device, and the processing speed is improved by several times or even dozens of times after optimization. The optimized large model can quickly process a large amount of industrial quality inspection data and meet the detection needs brought by the expanding industrial production scale.

[0107] The data pruning unit and the data quantization unit in the model inference acceleration module optimize the large model task of various business data, reducing the requirements of the model on hardware computing capacity and storage capacity. The large model originally requiring high-performance and high-cost edge hardware computing devices to run can be smoothly run on relatively inexpensive industrial-grade edge hardware computing devices after optimization. This not only reduces the hardware procurement cost of the industrial quality inspection system, but also reduces the energy consumption and maintenance cost of the equipment, enabling enterprises to realize the application of various industrial data in industrial quality inspection without increasing too much hardware investment, thereby improving the economic benefits of enterprises.

[0108] The embodiment can fully mine the value of multi-modal data, and the industrial data fusion module extracts and fuses features from the processing results of various industrial data to obtain a fusion result of the various industrial data. The model inference acceleration module performs data pruning, data quantization and data acceleration on the large model applied to the various industrial data according to the fusion result of the various industrial data to obtain an inference acceleration result of the various industrial data, and finally realizes the full mining of the correlation between video image data, environmental temperature data, vibration data of industrial equipment and pressure data audio of industrial equipment and other various data. For example, in electronic product quality inspection, the product appearance image and the vibration data of the industrial equipment can be combined to comprehensively and accurately judge the product quality, which is more efficient and accurate than the traditional manual quality inspection method and the single automatic quality inspection method. Therefore, the embodiment can detect more subtle defects and potential quality problems in the target industrial quality inspection scene, thereby improving the quality inspection accuracy and reducing the false detection and missed detection rates.

[0109] Although the embodiments of the present application are described in conjunction with the drawings, various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application, and such modifications and changes fall within the scope defined by the appended claims.

Claims

1. A large-model inference acceleration system for industrial quality inspection, characterized by: The system comprises: Industrial data acquisition module, used to collect various industrial data in target industrial quality inspection scenarios; An industrial data processing module, configured to process the various industrial data and obtain processing results of the various industrial data; An industrial data fusion module, configured to perform feature extraction and feature fusion on the processing results of the plurality of industrial data to obtain a fusion result of the plurality of industrial data; An industrial data updating module, configured to update the weight parameters of the large model for the various industrial data applications so that the inference accuracy of the large model meets the requirements of industrial quality inspection; a model reasoning acceleration module, configured to perform data pruning, data quantization, and data acceleration on a large model applied to the multiple industrial data based on the fusion results of the multiple industrial data, to obtain reasoning acceleration results for the multiple industrial data; The industrial data output module is used to output the inference acceleration results of the various industrial data to the target application interface for display.

2. The large-model inference acceleration system for industrial quality inspection according to claim 1 is characterized in that: The industrial data acquisition module includes: An industrial camera is used to collect video image data in the target industrial quality inspection scenario.

3. The large-model inference acceleration system for industrial quality inspection according to claim 1 is characterized in that: The industrial data acquisition module includes: A temperature sensor is used to collect ambient temperature data in the target industrial quality inspection scenario; A vibration sensor is used to collect vibration data of industrial equipment in the target industrial quality inspection scenario; The pressure sensor is used to collect pressure data of industrial equipment in the target industrial quality inspection scenario.

4. The large-model inference acceleration system for industrial quality inspection according to claim 1 is characterized in that: The industrial data processing module includes: Industrial data processing unit: used for cleaning, normalizing and aligning the various industrial data.

5. The large-model inference acceleration system for industrial quality inspection according to claim 1 is characterized in that: The industrial data fusion module includes: A feature extraction unit is used to extract target features corresponding to each type of industrial data using a convolutional neural network model; The feature fusion unit is used to fuse the multiple industrial data using a cross-modal attention module according to the target features corresponding to each type of industrial data to obtain a fusion result of the multiple industrial data.

6. The large-model inference acceleration system for industrial quality inspection according to claim 1 is characterized in that: The model reasoning acceleration module includes: A data pruning unit is used to perform pruning actions on the neurons of the large model of the various industrial data applications according to preset weights, and monitor the weight change trend of the large model in real time, and calculate the importance score of the large model task according to the weight change trend of the large model, and perform pruning actions on the neurons of the large model according to the importance score of the large model task, wherein the preset weights are determined according to the task type and task priority of the large model task.

7. The large-model inference acceleration system for industrial quality inspection according to claim 6 is characterized in that: The data pruning unit includes: A first pruning subunit is configured to, during the process of the large model processing the large model task, perform a pruning action on the neurons of the large model when the weight parameter of the neurons of the large model is less than the preset weight; The data calculation subunit is used to calculate the importance score of the large model task using the following importance score function formula; Among them, I i is the importance score of the large model task, α is the hyperparameter balance amplitude of the large model, β is the weight change trend of the large model, γ is the curvature, H i -1 is the weight importance of the large model, ω i t is the weight parameter generated during the training of the large model at time t for the i-th time, ω i t-1 is the weight parameter generated during the i-th training at time t-1, ω i is the final weight parameter generated during the i-th training of the large model, and T is the training cycle; The second pruning subunit is used to perform pruning actions on the neurons of the large model when the importance score of the large model task is less than or equal to a preset score.

8. The large-model inference acceleration system for industrial quality inspection according to claim 1 is characterized in that: The model reasoning acceleration module includes: a data quantization unit, Data quantification unit, including: The level determination subunit is used to determine the current calculation level of the large model task according to the task type of the large model task; A first quantization subunit is configured to quantize the pruned large model task according to a quantization accuracy of 16-bit floating point numbers when the current calculation level of the large model task is the first level; A second quantization subunit is configured to quantize the pruned large model task according to a quantization accuracy of 8-bit floating point numbers when the current calculation level of the large model task is the second level; A third quantization subunit is configured to quantize the pruned large model task according to a quantization accuracy of a 4-bit floating point number when the current calculation level of the large model task is the third level; The data compensation subunit is used to monitor the current error of the large model in real time during the process of quantifying the large model task, and perform compensation action on the large model when the current error of the large model is greater than a preset error.

9. The large-model inference acceleration system for industrial quality inspection according to claim 1 is characterized in that: The model reasoning acceleration module includes: The hardware acceleration subunit utilizes a dedicated integrated circuit chip or a field programmable logic gate array chip to accelerate the inference acceleration results of the multiple industrial data in parallel.

10. The large-model inference acceleration system for industrial quality inspection according to claim 1 is characterized in that: The industrial data update module includes: The data updating unit is used to update the large model of the multiple industrial data applications according to a preset period using the sample data of the multiple industrial data.