An industrial defect detection method based on cloud-edge collaboration

CN122820591APending Publication Date: 2026-09-25INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610956933.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]综上所述,现有工业视觉检测技术存在以下显著缺陷:语义增强能力有限,难以输出缺陷严重程度、成因等深层信息;边云协同缺乏存算协同性,静态调度无法适应资源动态变化且忽视存储与缓存重用;思维链机制僵化,无法根据场景需求灵活调节推理深度,从而在实时性、稳定性与可扩展性方面严重制约了系统在复杂工业环境下的高效应用

Benefits of technology

[0034]与现有技术相比,本发明的优点在于:(1)通过对多边缘端与云端之间进行实时监测计算资源占用率、存储资源占用率,网络带宽占用率以及任务队列长度这四类核心指标,对待检测图像进行跨边缘端动态分配,从而在网络波动与任务激增的情况下仍能保持高吞吐率、低延迟与高资源利用率;(2)将轻量化目标检测模型与多模态大模型通过串联式结构深度融合,显著提升了缺陷检测的准确率与可解释性;(3)通过融合粗检测结果、任务复杂度评估以及云边端系统资源状态,实现思维链结构的动态选择与深度自适应控制,使得云边端系统在不同任务复杂度与资源环境下均能保持推理速度与检测精度的动态平衡。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820591A_ABST
    Figure CN122820591A_ABST
Patent Text Reader

Abstract

The application provides an industrial defect detection method based on cloud-edge-terminal cooperation, which comprises the following steps: configuring each client to continuously collect product images on an industrial production line to obtain to-be-detected images, and uploading the to-be-detected images to the assigned edge terminal according to the distribution decision issued by the cloud terminal; configuring each edge terminal to periodically collect the resource state of itself and upload it to the cloud terminal, continuously receive and store the to-be-detected images, and use a target detection model to perform coarse detection on each to-be-detected image to obtain a defect region image set and a coarse detection information tuple of each to-be-detected image and upload them to the cloud terminal; and configuring the cloud terminal to use a multi-modal large model to perform fine detection on each defect region image set to obtain the comprehensive defect category, the comprehensive detection confidence and the defect semantic description of each defect region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, specifically to industrial visual inspection technology in the field of image processing, and more specifically, to a cloud-edge-device collaborative industrial defect detection method. Background Technology

[0002] With the deepening of industrial intelligence, computer vision technology has been widely applied in the field of defect detection. Lightweight object detection models, represented by the YOLO series, have become one of the mainstream solutions for industrial defect detection due to their compact structure and fast inference speed. In recent years, the deployment of deep learning models, especially lightweight object detection models, in industrial scenarios has grown rapidly. These models account for a high proportion in actual production and can meet basic real-time localization and classification needs. However, the detection output of these models is usually limited to detection boxes and category numbers, lacking the mining and interpretation of higher-level semantic information such as defect nature, severity, and potential causes. This makes it difficult for the detection results to directly support subsequent intelligent decision-making and complex scene analysis.

[0003] To address the aforementioned issues, the development of Multimodal Large Models (MLLMs) has offered new possibilities for semantic enhancement detection. MLLMs can establish deep associations between images and text, generating structured or natural language interpretations, effectively supplementing the shortcomings of traditional detection models in semantic description. However, MLLMs suffer from high computational complexity and large inference latency. Direct application in industrial inspection would lead to significant resource pressure and efficiency problems, making it difficult to meet the stringent real-time requirements of production lines. To alleviate this contradiction, a collaborative model between edge computing and cloud computing has gradually emerged. The edge relies on local computing power for real-time response, while the cloud undertakes complex inference tasks. However, current edge-cloud collaboration mostly adopts static task allocation or simple rule scheduling, lacking adaptability to dynamic resource states and task complexity. System performance significantly degrades under network fluctuations or uneven computational load. With the increasing heterogeneity of edge devices and the growing demand for multi-task parallelism, scheduling strategies based solely on computing resources are no longer suitable for real-world industrial environments. There is an urgent need to establish a collaborative allocation mechanism between computing, storage, and network resources to achieve dynamic task migration and cache reuse. However, existing research in this area is insufficient, lacking a unified optimization framework.

[0004] Furthermore, even in the application of multimodal large models, balancing inference accuracy and speed remains a significant challenge. While multimodal large models possess powerful semantic understanding and complex reasoning capabilities, completing defect identification, property analysis, and semantic interpretation often requires lengthy thought chain reasoning processes. Although this improves accuracy, it significantly increases time and resource consumption. In scenarios with simple tasks or limited resources, using excessively long thought chains leads to unnecessary delays and energy waste. Current technologies have not yet proposed a mechanism that can utilize structured information in visual inspection results and flexibly adjust the length of the thought chain according to task complexity and resource status, making it difficult to dynamically balance inference depth and efficiency.

[0005] In summary, existing industrial vision inspection technologies have the following significant drawbacks: limited semantic enhancement capabilities, making it difficult to output in-depth information such as the severity and cause of defects; lack of storage-computing synergy in edge-cloud collaboration, with static scheduling unable to adapt to dynamic resource changes and neglecting storage and cache reuse; and a rigid thinking chain mechanism that cannot flexibly adjust the inference depth according to scenario requirements, thus severely restricting the efficient application of the system in complex industrial environments in terms of real-time performance, stability, and scalability.

[0006] It should be noted that the background information presented here is only for illustrating relevant information about the present invention to aid in understanding the technical solution of the present invention, and does not imply that the relevant information is necessarily prior art. The relevant information was submitted and disclosed together with the present invention, and should not be considered prior art unless there is evidence that the relevant information was disclosed before the filing date of the present invention. Summary of the Invention

[0007] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a cloud-edge-device collaborative industrial defect detection method.

[0008] The objective of this invention is achieved through the following technical solution:

[0009] According to a first aspect of the present invention, a cloud-edge-device collaborative industrial defect detection method is provided, which is used to collaboratively implement industrial defect detection by multiple clients, multiple edge devices, and the cloud. Each edge device is configured with a target detection model, and the cloud is configured with a multimodal large model. The method includes configuring each client, each edge device, and the cloud to perform industrial defect detection in the following manner: each client continuously collects product images on the industrial production line to obtain images to be detected, and uploads the images to be detected to the assigned edge device according to the allocation decision issued by the cloud, wherein the allocation decision instructs the image to be detected to be assigned to the designated edge device for processing; each edge device periodically collects its own resource status and uploads it to the cloud, and continuously receives and stores the images to be detected, and uses the target detection model to perform coarse detection on each stored image to obtain a set of defect region images and coarse detection for each image to be detected. Information tuples are uploaded to the cloud. The defect region image set includes multiple defect regions cropped from the image to be detected. The coarse detection information tuple includes the location information, basic category, and basic detection confidence of each defect region. The cloud receives the resource status collected periodically by each edge terminal and, based on the resource status of all edge terminals in the previous period, generates and distributes allocation decisions for each image to be detected collected by each client using a preset storage-computing collaborative allocation strategy. The cloud also receives the defect region image set and coarse detection information tuple of each image to be detected uploaded by each edge terminal and uses a multimodal large model to perform fine detection on each defect region image set to obtain the comprehensive defect category, comprehensive detection confidence, and defect semantic description of each defect region. The defect semantic description includes the description of the location of the defect region, the rating of the defect severity, and the inference of the defect cause.

[0010] According to some embodiments of the present invention, a scheduling control module is configured on the cloud to receive resource status data periodically collected from each edge terminal. The method further includes: configuring the scheduling control module on the cloud to calculate the periodically changing load score of each edge terminal based on the received resource status data periodically collected from each edge terminal, and calculating the periodically changing average load and load variance to assess whether the load is balanced among all edge terminals; if the current load variance is less than a preset load balancing threshold, it indicates that the load is balanced among all edge terminals; conversely, if the current load variance is greater than or equal to the preset load balancing threshold, it indicates that the load is unbalanced among all edge terminals. In this case, the scheduling control module performs load balancing on some edge terminals according to a preset dynamic reallocation scheduling mechanism. The redistribution process, wherein the preset dynamic redistribution scheduling mechanism is as follows: each edge with a current load score greater than or equal to the current average load and exceeding a preset load deviation threshold is designated as a migrating-out edge, and each edge with a current load score less than the current average load is designated as a migrating-in edge. Multiple images to be detected that have not yet undergone coarse detection on each migrating-out edge are migrated to one or more migrating-in edges for coarse detection. During the migration process, the real-time load score of all edge edges is calculated, and the real-time average load and real-time load variance are calculated based on the real-time load scores of all edge edges, until the real-time load variance is less than a preset load balancing threshold, at which point the migration stops. The scheduling control module is configured to periodically calculate the load score of each edge edge in the following manner:

[0011]

[0012] in, Indicates the first The current load score of each edge device. Indicates the first The weight of the calculation index at each edge end, Indicates the first Current computing resource utilization at each edge device Indicates the first The weight of storage metrics at each edge. Indicates the first Current storage resource utilization at each edge device Indicates the first The weights of network metrics at the edge. Indicates the first Current network bandwidth utilization at each edge. Indicates the first Weighting of task metrics at the edge Indicates the first The number of currently unprocessed images to be detected at each edge. This represents the maximum number of images currently awaiting processing across all edge endpoints.

[0013] According to some embodiments of the present invention, the method further includes: configuring a scheduling control module on the cloud to periodically calculate the average network latency; if the average network latency is greater than or equal to a preset latency threshold in multiple consecutive periods indicated by a first preset value, it indicates that the network link between the cloud and all edge terminals is congested; the cloud triggers a local asynchronous mitigation mechanism to cause some edge terminals to suspend uploading the defect region image set and coarse detection information tuple of the image to be detected; wherein the local asynchronous mitigation mechanism is: obtaining the current round-trip latency from each edge terminal to the cloud, and causing all edge terminals with a current round-trip latency greater than or equal to the preset latency threshold to suspend uploading the defect region image set and coarse detection information tuple until the average network latency in multiple consecutive periods indicated by the first preset value is less than the preset latency threshold, and then resuming all edge terminals from uploading the defect region image set and coarse detection information tuple; wherein the scheduling control module is configured to periodically calculate the average network latency in the following manner:

[0014]

[0015] in, This represents the average network latency for the current period. Represents the set consisting of all edge ends. This indicates the number of all edge ends. Indicates the number of times within the current period The average round-trip latency from each edge node to the cloud.

[0016] According to some embodiments of the present invention, the method further includes: configuring the cloud to generate an allocation decision for each image to be detected in the following manner: based on the resource status of all edge terminals in the scheduling control module in the previous cycle, an allocation decision is generated for each image to be detected collected by each client using a preset storage-computing collaborative allocation strategy, wherein the preset storage-computing collaborative allocation strategy is:

[0017]

[0018]

[0019] in, This represents the allocation decision, which indicates the allocation of the image to be detected. Assigned to the One edge end, Indicates the first The computational latency weights of each edge end Indicates the first Process the image to be detected at each edge. Required computation latency, Indicates the first Storage access weights at each edge. Indicates the image to be detected In the Storage overhead generated at each edge. Indicates the first Network transmission weights at each edge end Indicates the image to be detected In the The network latency required at each edge is specified. The goal of the preset memory-computation co-allocation strategy is to allocate an edge for each image to be detected, so as to minimize the weighted sum of the computation latency, storage overhead and network transmission of all edge edges. The preset memory-computation co-allocation strategy is solved using the particle swarm optimization algorithm.

[0020] According to some embodiments of the present invention, the method further includes: configuring a first adaptive optimization strategy on the cloud to update the computational indicator weight, storage indicator weight, network indicator weight, and task indicator weight corresponding to each edge when the scheduling control module calculates the load score, and updating the computational latency weight, storage access weight, and network transmission weight corresponding to each edge when the scheduling control module generates the allocation decision. The first adaptive optimization strategy is: periodically recording the average coarse detection completion time of each edge for coarse detection of the image to be detected, the throughput rate corresponding to each edge, and the cache hit rate; if, within multiple consecutive periods indicated by a second preset value, the average coarse detection completion time of any edge continuously increases and the throughput rate corresponding to that edge continuously decreases, then the computational indicator weight, network indicator weight, or task indicator weight corresponding to that edge is increased, and the computational latency weight, storage access weight, and network transmission weight corresponding to that edge are increased. The calculation latency weight or network transmission weight is adjusted, where the increase is determined based on the change in the average coarse detection completion time and throughput of the edge terminal. If, within multiple consecutive periods indicated by the second preset value, the average coarse detection completion time and the throughput of the corresponding edge terminal remain within a specified range, and the cache hit rate of the corresponding edge terminal continues to decrease, then the storage metric weight and storage access weight of the corresponding edge terminal are increased, where the increase is determined based on the change in the cache hit rate of the corresponding edge terminal. If, within multiple consecutive periods indicated by the second preset value, the average coarse detection completion time of the corresponding edge terminal continues to shorten, the throughput continues to rise, and the cache hit rate continues to increase, then the calculation metric weight, storage metric weight, network metric weight, task metric weight, calculation latency weight, storage access weight, and network transmission weight of the corresponding edge terminal are not adjusted.

[0021] According to some embodiments of the present invention, the multimodal large model is configured with a chainless mode, a short-chain mode, and a long-chain mode to perform fine detection on each defect region image set, wherein the method further includes configuring the cloud to select a chainless mode, a short-chain mode, or a long-chain mode for each defect region image set in the following manner:

[0022]

[0023] in,

[0024]

[0025] in, This represents the joint decision value for task resources, when... When selecting the no-chain mode, When selecting short link mode, When selecting long chain mode, Indicates the first boundary threshold. Indicates the second boundary threshold. Indicates the task complexity weight. Representing the task complexity of a set of defect region images. Indicates the defect size weight. This represents the average defect size corresponding to all defect regions in the defect region image set. Indicates the ambiguity weight. This represents the image blur score corresponding to the image to be detected within the set of defect region images. Indicates the weight of the number of defects. This indicates the number of defective regions in the defective region image set. This indicates that the GPU utilizes weights. Indicates the real-time utilization of cloud GPUs. Indicates CPU usage weight. This represents the real-time average CPU utilization across all edge devices. Indicates network bandwidth weight. This represents the real-time average network bandwidth utilization corresponding to all edge devices. Among them, the no-chain mode means that the multimodal large model does not execute the thinking chain when performing fine detection on the defect region image set; the short-chain mode means that the multimodal large model executes a three-step reasoning thinking chain when performing fine detection on the defect region image set; and the long-chain mode means that the multimodal large model executes a five-step reasoning thinking chain when performing fine detection on the defect region image set.

[0026] According to some embodiments of the present invention, the method further includes: configuring a second adaptive optimization strategy on the cloud to update task complexity weight, GPU utilization weight, CPU usage weight, network bandwidth weight, a first boundary threshold, and a second boundary threshold, wherein the second adaptive optimization strategy is: periodically recording the average precision detection completion time and average inference accuracy of the cloud for each defect region; if the average inference accuracy of the current period is lower than the average inference accuracy of the previous period, and the average precision detection completion time is still within a specified range, then increasing the task complexity weight and decreasing the first boundary threshold or the second boundary threshold, wherein the increase is determined according to the degree of change in the average inference accuracy; if the average precision detection completion time of the current period is longer than the average precision detection completion time of the previous period, and the average inference accuracy is still within a specified range, then increasing the GPU utilization weight, CPU usage weight, or network bandwidth weight, and increasing the first boundary threshold or the second boundary threshold, wherein the increase is determined according to the degree of change in the GPU utilization rate of the cloud, the average CPU usage rate of all edge terminals, or the average network bandwidth usage rate of all edge terminals, and the increase in the threshold is determined according to the degree of change in the average precision detection completion time.

[0027] According to some embodiments of the present invention, an industrial defect knowledge base is configured on the cloud, which includes a variety of known defect categories. The method further includes: configuring the cloud to process each defect region image set in the following manner: performing fine detection on each defect region image set using a self-configured multimodal large model to obtain the fine category, fine category probability, and defect semantic description of each defect region; if the fine category of the defect region is a subclass of the basic category, then the fine category is taken as the comprehensive defect category of the defect region, and the comprehensive detection confidence is calculated according to a preset first fusion mechanism, wherein the fine category being a subclass of the basic category indicates that the fine category is a defect category that refines the basic category; if the fine category of the defect region is not a subclass of the basic category, then multiple compatibility candidate categories are constructed based on the basic category, the fine category, and the industrial defect knowledge base, and the score of each compatibility candidate category is calculated according to a preset second fusion mechanism, and based on the score of each compatibility candidate category, the compatibility candidate category with the highest score is selected as the comprehensive defect category, and the score of the compatibility candidate category is taken as the comprehensive detection confidence, wherein the compatibility candidate category represents a defect category that extends the basic category and the fine category to the same classification level.

[0028] Preferably, the preset first fusion mechanism is:

[0029]

[0030] in, Indicates the overall detection confidence level. This indicates the third adjustment weight. Represents the fine category probability, This indicates the confidence level of the basic test.

[0031] Preferably, the preset second fusion mechanism is:

[0032]

[0033] in, Indicates rating, This indicates the compatibility candidate category output by the target detection model at the corresponding edge. The probability, Indicates candidate categories for compatibility of multimodal large model output. The probability, Indicates category compatibility score, This indicates the compatibility weight.

[0034] Compared with the prior art, the advantages of the present invention are: (1) By monitoring the four core indicators of computing resource utilization, storage resource utilization, network bandwidth utilization and task queue length between multiple edge terminals and the cloud in real time, the images to be detected are dynamically allocated across edge terminals, so that high throughput, low latency and high resource utilization can still be maintained under network fluctuations and task surges; (2) The lightweight target detection model and the multimodal large model are deeply integrated through a serial structure, which significantly improves the accuracy and interpretability of defect detection; (3) By integrating coarse detection results, task complexity assessment and cloud edge system resource status, the dynamic selection and deep adaptive control of the thinking chain structure are realized, so that the cloud edge system can maintain a dynamic balance between inference speed and detection accuracy under different task complexity and resource environments. Attached Figure Description

[0035] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:

[0036] Figure 1 This is a schematic diagram of the cloud-edge-device collaborative industrial defect detection method according to an embodiment of the present invention. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the invention.

[0038] As mentioned in the background technology section, existing industrial vision inspection technologies have the following significant drawbacks: limited semantic enhancement capabilities, making it difficult to output in-depth information such as the severity and cause of defects; lack of storage-computing synergy in edge-cloud collaboration, static scheduling cannot adapt to dynamic resource changes and ignores storage and cache reuse; and rigid thinking chain mechanism, which cannot flexibly adjust the inference depth according to scenario requirements, thus severely restricting the efficient application of the system in complex industrial environments in terms of real-time performance, stability, and scalability.

[0039] To address the aforementioned issues, the inventors analyzed existing technologies and discovered significant shortcomings in three areas: semantic understanding, resource coordination, and reasoning efficiency.

[0040] First, the lack of semantic enhancement capabilities makes it difficult to meet the detection needs of complex scenarios. While existing lightweight detection models excel in inference speed and target localization accuracy, their output is typically limited to bounding boxes and category numbers, lacking a deeper semantic description of the detected object. For example, in the inspection of industrial parts or circuits, the model can mark the location of defects such as open circuits or bubbles, but cannot further explain the severity of the defects, their possible causes, or their impact on the overall circuit performance. This limitation makes it difficult for the detection results to directly support subsequent intelligent decision-making. The reason for this problem lies in the compact structure of lightweight models themselves, their limited feature extraction capabilities, and their lack of joint modeling and semantic generation capabilities for multimodal information, making it difficult to cope with complex and ambiguous industrial scenarios.

[0041] Secondly, edge-cloud collaboration suffers from insufficient resource utilization and a lack of dynamic scheduling mechanisms. Most current edge-cloud collaboration solutions employ a static division of labor model, typically pre-setting lightweight detection at the edge and deep inference in the cloud. However, in actual operation, devices often experience fluctuations in computing power, storage usage, and network status. When there is a surge in tasks, network congestion, or overload of some edge devices, static partitioning cannot perform task migration and load balancing based on real-time resource status, leading to edge performance bottlenecks, cloud queuing backlogs, and increased system latency. More importantly, most current methods only focus on computing resource scheduling, neglecting to include storage resources and caching mechanisms in their optimization objectives. The lack of a three-dimensional coordinated control mechanism integrating "computing-storage-network" results in high task migration costs, low cache hit rates, and overall low resource utilization.

[0042] Thirdly, it is difficult to balance inference accuracy and speed, and there is a lack of an adjustable thought chain mechanism. Multimodal large models possess strong semantic understanding and complex reasoning capabilities, but when handling highly complex tasks such as industrial inspection, they often require a long thought chain reasoning process to gradually complete defect identification, property analysis, and semantic interpretation. While this improves inference accuracy, it also significantly increases inference time and computational resource consumption. In scenarios with simple tasks and limited resources, continuing to use complex long-chain inference will cause unnecessary delays and energy waste. Current technologies have not yet proposed a mechanism that can utilize the structured information in visual inspection results and flexibly select the thought chain length according to task complexity and resource status to achieve a dynamic balance between inference depth and efficiency. Therefore, existing methods either sacrifice real-time performance for high accuracy or can only maintain a fast response in simple modes, failing to meet the multiple needs of different scenarios.

[0043] Based on the above analysis, the inventors propose a multimodal large-scale cloud-edge-device collaborative industrial inspection method based on dynamic thinking chain enhancement. The core idea of ​​this method is to achieve high-precision, high-real-time, and scalable industrial defect detection on a cloud-edge-device system (which includes multiple clients, multiple edge devices, and the cloud) through the collaborative optimization of semantic enhancement, dynamic scheduling, and adjustable thinking chains. In short, to address the issue of insufficient semantics, this method employs a "detection-semantic concatenation" mechanism. This involves configuring edge devices to quickly locate defects and uploading the cropped feature regions to a multimodal large model in the cloud for semantic generation and natural language interpretation, thereby enhancing the semantics and interpretability of the localization results. To address the issues of static scheduling and lack of storage-computing collaboration, this method uses a dynamic scheduling mechanism based on resource and storage-computing collaboration awareness. This involves configuring the cloud to periodically record the resource status of each edge device and, based on the periodically changing resource status of each edge device, migrating unprocessed tasks from high-load edge devices to low-load edge devices, achieving dynamic scheduling to improve the stability and throughput of the cloud-edge-device system in fluctuating scenarios. To address the issue of rigid thought chains, this method employs an adaptive reasoning optimization mechanism based on visually guided dynamic thought chains. This involves configuring the cloud to dynamically select the thought chain used when the multimodal large model performs semantic generation based on the difficulty of the semantic generation task performed by the cropped feature regions and the resource status of the cloud-edge-device system, enabling the cloud-edge-device system to maintain a dynamic balance between speed and accuracy in complex scenarios.

[0044] In summary, such as Figure 1As shown, this invention proposes a cloud-edge-device collaborative industrial defect detection method for coordinating multiple clients, multiple edge devices, and the cloud to achieve industrial defect detection. Each edge device is configured with a target detection model, and the cloud is configured with a multimodal large model. The method includes configuring each client, each edge device, and the cloud to perform industrial defect detection in the following manner: Each client continuously collects product images on the industrial production line to obtain images to be detected, and uploads the images to be detected to the assigned edge device according to the allocation decision issued by the cloud, wherein the allocation decision instructs the image to be detected to be processed by the designated edge device; Each edge device periodically collects its own resource status and uploads it to the cloud, and continuously receives and stores the images to be detected, and uses the target detection model to perform coarse detection on each stored image to obtain a set of defect region images and coarse detection information for each image to be detected. The data is processed and uploaded to the cloud. The defect region image set includes multiple defect regions cropped from the image to be detected. The coarse detection information tuple includes the location information, basic category, and basic detection confidence of each defect region. The cloud receives the resource status collected periodically by each edge terminal and, based on the resource status of all edge terminals in the previous period, generates and distributes allocation decisions for each image to be detected collected by each client using a preset storage-computing collaborative allocation strategy. The cloud also receives the defect region image set and coarse detection information tuple of each image to be detected uploaded by each edge terminal and performs fine detection on each defect region image set using a multimodal large model to obtain the comprehensive defect category, comprehensive detection confidence, and defect semantic description of each defect region. The defect semantic description includes the description of the location of the defect region, the rating of the defect severity, and the inference of the defect cause.

[0045] To better understand the cloud-edge-device collaborative industrial defect detection method proposed in this invention, the following describes how to configure the cloud-edge-device system to achieve industrial defect detection, with specific embodiments as an example.

[0046] I. Client

[0047] In the industrial defect detection method proposed in this invention, each client is configured to continuously collect product images on the industrial production line to obtain images to be detected, and upload the images to be detected to the assigned edge end according to the allocation decision issued by the cloud. The allocation decision instructs the image to be detected to be assigned to the designated edge end for processing, and the edge end only performs coarse detection, that is, the edge end is configured to perform defect localization and coarse classification on the image to be detected in order to locate all defect areas in the image to be detected and the defect category corresponding to each defect area.

[0048] II. Edge End

[0049] In the industrial defect detection method proposed in this invention, each edge terminal is configured to periodically collect its own resource status and upload it to the cloud, and continuously receive and store images to be detected. A target detection model is used to perform coarse detection on each stored image to be detected to obtain a set of defect region images and a coarse detection information tuple for each image to be detected, which are then uploaded to the cloud. The set of defect region images includes multiple defect regions cropped from the image to be detected, and the coarse detection information tuple includes the location information, basic category, and basic detection confidence of each defect region.

[0050] The object detection model configured on the edge devices is set based on actual needs, and this invention does not impose any special restrictions on it. For example, all edge devices can be configured with the same object detection model, or different edge devices can be configured with different object detection models independently. The object detection model can be selected from existing mainstream deep learning detection frameworks such as the YOLO series, SSD series, Faster R-CNN series, or EfficientDet series. For example, all edge devices can be uniformly configured with the YOLOv12l model as the object detection model; or, some edge devices can be configured with the YOLOv12l model, and the remaining edge devices can be configured with the YOLOv12m model, so as to achieve differentiated deployment in terms of accuracy and speed.

[0051] Specifically, the target detection model configured on each edge performs coarse detection on each image to be detected in the following manner.

[0052] The image to be detected is input into the target detection model, which then outputs a set of detection boxes through convolutional feature extraction and a multi-scale detection head. , ,in, Indicates the first detection box in the set. One detection box, Represents the detection box The coordinates of the center point, Represents the detection box width, Represents the detection box of high, Represents the detection box The basic category corresponding to the selected defect area. Represents the detection box The original confidence level.

[0053] Furthermore, to enhance the target detection model's ability to detect small and ambiguous defects, an adaptive confidence recalibration mechanism is introduced to weight and correct the original confidence of each detection box in the detection box set:

[0054]

[0055] in, Indicates the first The weighted confidence score of each detection box is adjusted. This adjusted confidence score can effectively reduce false alarms in low-quality defect areas while maintaining the recall rate of the detection boxes. This represents the first adjustment coefficient. This represents the second adjustment coefficient. This indicates the current computing load at the corresponding edge. Indicates the first The normalized local texture sharpness index corresponds to each detection box. The larger the index value, the more obvious the edge changes and the clearer the local texture of the defect area selected by the corresponding detection box. The smaller the index value, the more likely the defect area selected by the corresponding detection box may have problems such as blurriness, low contrast, or weak texture. This represents the maximum value of the normalized local texture sharpness index corresponding to all detection boxes in the detection box set. It should be noted that the values ​​of the first and second adjustment coefficients are determined based on actual needs, and this invention does not impose any special limitations. For example, setting the first adjustment coefficient... Set a second adjustment coefficient for values ​​such as 0.3, 0.5, and 0.6. The values ​​are 0.2, 0.7, 0.5, etc.

[0056] The normalized local texture sharpness index for each detection box is calculated as follows:

[0057]

[0058] in,

[0059]

[0060] in,

[0061]

[0062]

[0063] in,

[0064]

[0065] in, Indicates the first The original local texture sharpness index corresponding to each detection box This represents the maximum value of the original local texture sharpness index corresponding to all detection boxes in the detection box set. To prevent extremely small constants with a denominator of 0, Indicates the first The number of pixels within the defect area corresponding to each detection box. Represents pixel coordinates, Indicates the first Each detection box corresponds to a set of pixel coordinates within the defect area. Represents pixel coordinates gradient magnitude at that point Indicates the first The average gradient magnitude of pixels within the defect region corresponding to each detection box. Indicates the first The grayscale image of the defect area corresponding to each detection box. Indicates grayscale image The gradient map obtained after performing horizontal convolution with the Sobel operator (reflecting the intensity of grayscale changes in the image in the horizontal direction). Indicates grayscale image The gradient map obtained after performing vertical convolution with the Sobel operator (reflecting the intensity of grayscale changes in the image in the vertical direction). Indicates the first grayscale image of the defect area corresponding to each detection box medium pixel The square of the horizontal gradient value, Indicates the first grayscale image of the defect area corresponding to each detection box medium pixel The square of the vertical gradient value.

[0066] Furthermore, the image to be detected is cropped based on each detection box in the detection box set to obtain a set of defect region images corresponding to the image to be detected. And by using the weighted and corrected confidence score of each detection box as the base detection confidence score for that detection box, a coarse detection information tuple corresponding to the defect region image set is obtained. ,in, This indicates the first defect region in the image to be detected. This indicates the second defect region in the image to be detected. Indicates the corresponding image to be detected. One defective area.

[0067] In the industrial defect detection method proposed in this invention, in addition to configuring each edge terminal to use its own target detection model to perform coarse detection on each stored image to be detected, each edge terminal is also configured to periodically collect its own resource status and upload it to the cloud, so that the cloud can perceive the overall resource distribution of the cloud-edge system in real time.

[0068] Specifically, each edge device is configured to periodically collect four core metrics: computing resource utilization, storage resource utilization, network bandwidth utilization, and task queue length, in order to determine the resource status of each edge device.

[0069] Specifically, the computing resource utilization rate of each edge device is collected periodically in the following manner:

[0070]

[0071] in, Indicates the first The utilization rate of computing resources at each edge. Indicates the first CPU utilization at the edge Indicates the first GPU utilization at the edge Indicates the first Total computing resources at each edge.

[0072] Specifically, the storage resource utilization rate of each edge device is collected periodically in the following manner:

[0073]

[0074] in, Indicates the first Storage resource utilization at each edge end Indicates the first The memory already used at the edge Indicates the first Total memory capacity at each edge.

[0075] The network bandwidth utilization rate of each edge terminal is collected in the following manner:

[0076]

[0077] in, Indicates the first Network bandwidth utilization at the edge Indicates the first Bandwidth consumption at each edge end Indicates the first The maximum available bandwidth at each edge.

[0078] Specifically, the task queue length of each edge endpoint is collected periodically in the following manner. : Count the number of unprocessed images to be detected at the edge.

[0079] The resource status of each edge node is periodically collected and uploaded to the cloud via a message queue to form a node status matrix. ,in, This indicates the number of edge nodes. The node status matrix is ​​used to reflect the overall resource distribution of all edge nodes in real time.

[0080] III. Cloud

[0081] In the industrial defect detection method proposed in this invention, the cloud is configured to receive the resource status periodically collected by each edge terminal, and based on the resource status of all edge terminals in the previous period, a preset storage-computing collaborative allocation strategy is adopted to generate and distribute allocation decisions for each image to be detected collected by each client. The cloud also receives the defect region image set and coarse detection information tuple of each image to be detected uploaded by each edge terminal, and uses a multimodal large model to perform fine detection on each defect region image set to obtain the comprehensive defect category, comprehensive detection confidence and defect semantic description of each defect region. The defect semantic description includes the description of the location of the defect region, the rating of the defect severity and the inference of the defect cause.

[0082] The cloud also includes a scheduling and control module to receive the resource status collected periodically from each edge device and dynamically reallocate resources based on the periodically changing resource status of all edge devices, thereby improving the stability and throughput of the cloud-edge-device system in fluctuating scenarios.

[0083] Based on this, it can be seen that the cloud is configured to perform three types of tasks: task allocation, dynamic rescheduling, and fine-grained defect detection. To better understand this invention, the following detailed explanation, using specific embodiments, illustrates how to configure the cloud to perform task allocation, dynamic rescheduling, and fine-grained defect detection.

[0084] 3.1 Task Allocation

[0085] In the industrial defect detection method proposed in this invention, the cloud is configured to receive the resource status collected periodically by each edge terminal, and based on the resource status of all edge terminals in the previous period, a preset storage-computing collaborative allocation strategy is adopted to generate and distribute allocation decisions for each image to be detected collected by each client.

[0086] According to an embodiment of the present invention, the method further includes: configuring the cloud to generate an allocation decision for each image to be detected in the following manner: based on the resource status of all edge terminals in the previous cycle in the scheduling control module, an allocation decision is generated for each image to be detected collected by each client using a preset storage-computing collaborative allocation strategy, wherein the preset storage-computing collaborative allocation strategy is:

[0087]

[0088]

[0089] in, This represents the allocation decision, which indicates the allocation of the image to be detected. Assigned to the One edge end, Indicates the first The computational latency weights of each edge end Indicates the first Process the image to be detected at each edge. Required computation latency, Indicates the first Storage access weights at each edge. Indicates the image to be detected In the Storage overhead generated at each edge. Indicates the first Network transmission weights at each edge end Indicates the image to be detected In the The network latency required at each edge is determined by the preset memory-computing co-allocation strategy, which aims to allocate an edge to each image to be detected so as to minimize the weighted sum of computation latency, storage overhead, and network transmission of all edge edges. The preset memory-computing co-allocation strategy is solved using the particle swarm optimization algorithm.

[0090] As can be seen from the foregoing, the present invention configures the cloud with the goal of minimizing the weighted sum of computation latency, storage overhead, and network transmission of all edge devices, so as to allocate edge devices to perform coarse detection for each image to be detected, thereby effectively utilizing the computing resources, storage resources, and network transmission resources of the cloud-edge-device system and avoiding the waste of resources in the cloud-edge-device system.

[0091] 3.2 Dynamic Rescheduling

[0092] In the industrial defect detection method proposed in this invention, a scheduling and control module on the cloud is configured to receive the resource status collected periodically by each edge terminal, and dynamically reallocate the resources according to the periodically changing resource status of all edge terminals, so as to improve the stability and throughput of the cloud-edge-end system under fluctuating scenarios.

[0093] The reason for configuring dynamic reallocation in the cloud is that after the cloud-edge-device system starts running, task traffic and network status are constantly changing, making it difficult to maintain long-term stable low latency and high throughput with only single-round scheduling. Therefore, this invention configures a scheduling control module in the cloud to receive resource status data collected periodically from all edge devices, and analyzes whether the load is balanced among all edge devices through the periodically changing resource status. If it is unbalanced, images to be detected that have not undergone coarse detection on high-load edge devices are preferentially migrated to low-load edge devices for coarse detection. For tasks that have completed coarse detection but have not yet performed fine detection, their defect area image set and coarse detection information tuple are uploaded to the cloud for fine detection. This avoids directly using raw images lacking coarse detection results as fine detection input, and achieves load balancing among edge devices, ensuring the long-term stability of the cloud-edge-device system.

[0094] According to an embodiment of the present invention, the method further includes: configuring a scheduling control module on the cloud to calculate the periodically changing load score of each edge terminal based on the resource status periodically collected by each edge terminal, and to calculate the periodically changing average load and load variance to assess whether the load is balanced among all edge terminals; if the current load variance is less than a preset load balancing threshold, it indicates that the load is balanced among all edge terminals; conversely, if the current load variance is greater than or equal to the preset load balancing threshold, it indicates that the load is unbalanced among all edge terminals. In this case, the scheduling control module performs a reallocation process on some edge terminals according to a preset dynamic reallocation scheduling mechanism.

[0095] The scheduling control module calculates the periodically changing average load and load variance as follows:

[0096]

[0097]

[0098] in, This indicates the current average load. This indicates the total number of edge ends. Represents the set consisting of all edge ends. Indicates the first The current load score of each edge device. Indicates the current load variance, if This indicates that the load is currently unbalanced among all edge devices, requiring dynamic reallocation. This indicates that the load is currently balanced across all edge devices. This indicates the preset load balancing threshold, which is set based on actual needs.

[0099] The preset dynamic redistribution scheduling mechanism is as follows: each edge with a current load score greater than or equal to the current average load and exceeding a preset load deviation threshold (set based on actual needs) is designated as a migration-out edge, and each edge with a current load score less than the current average load is designated as a migration-in edge. Multiple images to be detected that have not yet undergone coarse detection on each migration-out edge are migrated to one or more migration-in edges for coarse detection. During the migration process, the real-time load score of all edge edges is calculated, and the real-time average load and real-time load variance are calculated based on the real-time load scores of all edge edges. The migration stops when the real-time load variance is less than the preset load balancing threshold.

[0100] The configuration scheduling control module periodically calculates the load score for each edge terminal in the following manner:

[0101]

[0102] in, Indicates the first The current load score of each edge device. Indicates the first The weight of the calculation index at each edge end, Indicates the first Current computing resource utilization at each edge device Indicates the first The weight of storage metrics at each edge. Indicates the first Current storage resource utilization at each edge device Indicates the first The weights of network metrics at the edge. Indicates the first Current network bandwidth utilization at each edge. Indicates the first Weighting of task metrics at the edge Indicates the first The number of currently unprocessed images to be detected at each edge. This represents the maximum number of images currently awaiting processing across all edge endpoints.

[0103] According to one embodiment of the present invention, the method further includes: configuring a scheduling control module on the cloud to periodically calculate the average network latency; if the average network latency is greater than or equal to a preset latency threshold in multiple consecutive periods indicated by a first preset value, it indicates that the network link between the cloud and all edge terminals is congested, and the cloud triggers a local asynchronous mitigation mechanism to cause some edge terminals to suspend uploading the defect region image set and coarse detection information tuple of the image to be detected. It should be noted that when the local asynchronous mitigation mechanism is triggered, it does not suspend all computing tasks of all edge terminals, but rather suspends or delays non-urgent, cloud-dependent subsequent tasks. That is, the edge terminals continue to perform coarse detection, but do not upload the coarse detection results. After the network is restored, the coarse detection results are uploaded to the cloud for fine detection.

[0104] The local asynchronous mitigation mechanism is as follows: obtain the current round-trip latency from each edge terminal to the cloud, and suspend all edge terminals with current round-trip latency greater than or equal to a preset latency threshold from uploading the defect area image set and coarse detection information tuple until the average network latency in multiple consecutive periods indicated by the first preset value is less than the preset latency threshold, and then resume all edge terminals from uploading the defect area image set and coarse detection information tuple.

[0105] The configuration scheduling control module periodically calculates the average network latency in the following manner:

[0106]

[0107] in, This represents the average network latency for the current period. Represents the set consisting of all edge ends. This indicates the number of all edge ends. Indicates the number of times within the current period The average round-trip latency from each edge node to the cloud.

[0108] According to an embodiment of the present invention, the method further includes: configuring a first adaptive optimization strategy on the cloud to update the computational indicator weight, storage indicator weight, network indicator weight, and task indicator weight corresponding to each edge when the scheduling control module calculates the load score, and updating the computational latency weight, storage access weight, and network transmission weight corresponding to each edge when the scheduling control module generates the allocation decision. The first adaptive optimization strategy is as follows: periodically recording the average coarse detection completion time of each edge for coarse detection of the image to be detected, the throughput rate corresponding to each edge, and the cache hit rate; if, within multiple consecutive periods indicated by a second preset value, the average coarse detection completion time of any edge continuously increases and the throughput rate corresponding to that edge continuously decreases, then increasing the computational indicator weight, network indicator weight, or task indicator weight corresponding to that edge, and increasing the computational latency corresponding to that edge. The weights, or network transmission weights, are adjusted based on the changes in the average coarse detection completion time and throughput of the edge endpoint. If, within multiple consecutive periods indicated by the second preset value, the average coarse detection completion time and the throughput of the corresponding edge endpoint remain within a specified range (determined based on actual needs), and the cache hit rate of the corresponding edge endpoint continues to decrease, then the storage metric weights and storage access weights of the corresponding edge endpoint are increased, with the increase determined based on the changes in the cache hit rate of the corresponding edge endpoint. If, within multiple consecutive periods indicated by the second preset value, the average coarse detection completion time of the corresponding edge endpoint continues to shorten, the throughput continues to rise, and the cache hit rate continues to increase, then the computation metric weights, storage metric weights, network metric weights, task metric weights, computation latency weights, storage access weights, and network transmission weights of the corresponding edge endpoint are not adjusted.

[0109] Based on the foregoing, this invention proposes a resource-aware dynamic rescheduling mechanism. This mechanism can monitor four core indicators in real time between multiple edge devices and the cloud: computing resource utilization, storage resource utilization, network bandwidth utilization, and task queue length. This allows for dynamic reallocation of images to be detected across edge devices, ensuring the cloud-edge system maintains high throughput, low latency, and high resource utilization even under network fluctuations and task surges. Furthermore, this invention proposes a local asynchronous mitigation mechanism. This mechanism monitors whether the network links between the cloud and all edge devices are congested, controlling the upload of defect region image sets and coarse detection information tuples by the edge devices.

[0110] 3.3 Fine-grained defect detection

[0111] In the industrial defect detection method proposed in this invention, the cloud is configured to receive a set of defect region images and a coarse detection information tuple for each image to be detected uploaded by each edge device. A multimodal large model is then used to perform fine detection on each set of defect region images to obtain the comprehensive defect category, comprehensive detection confidence, and defect semantic description for each defect region. The defect semantic description includes an explanation of the location of the defect region, a rating of the defect severity, and an inference of the defect cause. The multimodal large model is set based on actual needs, and this invention does not impose special limitations. For example, models supporting multimodal input, such as the GPT series vision models, Qwen-VL series models, InternVL series models, or Llama Vision series, can be selected. In addition, an industrial defect knowledge base is configured on the cloud, which includes various known defect categories and the judgment rules for each defect category.

[0112] The multimodal large model is configured with three modes: no-chain, short-chain, and long-chain, to perform fine detection on each defect region image set. No-chain mode means the multimodal large model does not execute a thought chain when performing fine detection on the defect region image set. Short-chain mode means the multimodal large model executes a three-step reasoning thought chain when performing fine detection on the defect region image set. Long-chain mode means the multimodal large model executes a five-step reasoning thought chain when performing fine detection on the defect region image set. It should be noted that the adjustable multi-level thought chain template is constructed to dynamically adjust the thought chain length based on task complexity and cloud-edge-device system resource status, thereby achieving adaptive optimization of cloud-based inference depth and balancing real-time performance and accuracy.

[0113] To better understand this invention, the chainless mode, short chain mode, and long chain mode are briefly described below.

[0114] In chainless mode, each defect region in the defect region image set, its corresponding location information, basic category, and basic detection confidence score are combined into a single multimodal input. The multimodal large model does not generate intermediate analysis steps and directly outputs the location, severity, defect cause judgment, fine category, and corresponding fine category probability of the defect region. Specifically, the multimodal large model outputs a fine defect category probability distribution and selects the defect category with the highest probability as the fine category output. For example, if the defect region is a scratch image selected from the center of a clearly illuminated steel plate image, with a basic category of "scratch" and a basic detection confidence score of 0.94, the multimodal large model can directly combine the location information and basic category of the defect region to output "Fine category: surface scratch, scratch located slightly to the right of the center of the steel plate, scratch is a continuous linear mechanical damage caused by relative sliding friction, defect severity is mild, fine category probability is 0.96". Chainless mode is suitable for scenarios with a single target, clear boundaries, and high basic detection confidence scores.

[0115] In short-chain mode, each defect region in the defect region image set, its corresponding location information, basic category, and basic detection confidence are still combined into a one-time multimodal input. At this time, the multimodal large model performs three intermediate inference steps: First, it reads the location, size, texture, and color differences of the defect region to form a visual description of the defect. Second, it matches the visual description of the defect with the defect features and discrimination rules in the industrial defect knowledge base to exclude similar categories. Third, it outputs the location, severity, defect cause judgment, fine category, and corresponding fine category probability of the defect region. For example, if the defect area is a dark, nearly circular area selected from an image of an aluminum alloy surface, first read the location information of the defect area (located in the upper right corner of the image), size (approximately 2.1 mm in diameter), texture (irregular edges, rough surface), and color difference (low brightness in the center, obvious depression around the perimeter) to form a visual description of the defect: a dark, nearly circular depression area with an uneven boundary and a shallow depression depth. Then, compare the visual features with oil stains (dark color but blurred edges, no depression), indentations (depressions, clear edges or mechanical deformation marks), and corrosion spots (irregular, accompanied by rust or peeling) in the industrial defect knowledge base, exclude oil stains and corrosion spots, and determine that it is more consistent with the "indentation" feature. Finally, output "Fine category: local indentation, location: upper right corner, severity: moderate, indentation is caused by excessive pressure from the conveyor fixture or hard particles remaining on the workpiece surface under impact, fine category probability: 0.90, and it is recommended to re-inspect the dimensions and check the pressure of the conveyor fixture."

[0116] In the long-chain mode, each defect region in the defect region image set, its corresponding location information, basic category, and basic detection confidence are still combined into a one-time multimodal input. At this time, the multimodal large model performs five intermediate inference steps: First, it evaluates the image quality, product model, and inspection station of the defect region to determine usable evidence; second, it locates, groups, and describes the morphology of the defect region; third, it fuses multimodal information such as visible light, infrared, or historical process parameters, and retrieves defect categories from the industrial defect knowledge base to obtain multiple candidate defect categories; fourth, it compares the retrieved multiple candidate defect categories to analyze the cause and severity of the defect; fifth, it outputs the location, severity, defect cause judgment, fine category, and corresponding fine category probability of the defect region.

[0117] It should be noted that the above descriptions of the chainless mode, short-chain mode, and long-chain mode are merely illustrative. In actual implementation, the thought chain mode on the multimodal large model is set based on actual needs, and this invention does not impose any special limitations. It should also be noted that in multi-edge-cloud collaboration, the edge can perform pre-processing steps related to thought chain inference, such as candidate region cropping, basic category recognition, visual feature extraction, coarse detection information structuring, and local caching; the cloud performs short-chain or long-chain inference of the multimodal large model based on the above pre-processing results. This reduces the amount of data received and processed by the cloud, while avoiding describing the edge as directly executing the thought chain inference of the multimodal large model.

[0118] According to one embodiment of the present invention, the method further includes configuring the cloud to select a chainless mode, a short chain mode, or a long chain mode for each set of defective region images in the following manner:

[0119]

[0120] in,

[0121]

[0122] in, Represents the joint decision value of task resources, when When selecting the no-chain mode, When selecting short link mode, When selecting long chain mode, Indicates the first boundary threshold. Indicates the second boundary threshold. Indicates the task complexity weight. The task complexity of representing a set of defect region images. Indicates the defect size weight. This represents the average defect size corresponding to all defect regions in the defect region image set. Indicates the ambiguity weight. This represents the image blur score corresponding to the image to be detected within the set of defect region images. Indicates the weight of the number of defects. This indicates the number of defective regions in the defective region image set. This indicates that the GPU utilizes weights. Indicates the real-time utilization of cloud GPUs. Indicates CPU usage weight. This represents the real-time average CPU utilization across all edge devices. Indicates network bandwidth weight. This represents the real-time average network bandwidth utilization across all edge devices. , and The value is determined based on actual needs, and this invention does not impose any special restrictions.

[0123] The image blur score is calculated as follows:

[0124]

[0125] in,

[0126]

[0127]

[0128] in, Indicates the first Image blur score of the image to be detected. Indicates the first The variance of the Laplacian response corresponding to each image to be detected This represents the minimum variance of the Laplacian response corresponding to the image to be detected, obtained from historical statistics. This represents the maximum variance of the Laplacian response corresponding to the image to be detected, obtained from historical statistics. This represents a very small constant to prevent the denominator from being zero. Indicates the first The second-order gradient response of the image to be detected. Indicates the first The grayscale image of the image to be detected. Indicates the first The total number of pixels contained in the image to be detected Indicates the first A set consisting of all pixels in an image to be detected. Represents pixel coordinates The two-stage gradient response at the location, Indicates the first The mean of the two-stage gradient response corresponding to each image to be detected.

[0129] According to one embodiment of the present invention, the method further includes: configuring a second adaptive optimization strategy in the cloud to update task complexity weight, GPU utilization weight, CPU usage weight, network bandwidth weight, a first boundary threshold, and a second boundary threshold, wherein the second adaptive optimization strategy is: periodically recording the average precision detection completion time and average inference accuracy of each defect region performed by the cloud; if the average inference accuracy of the current period is lower than the average inference accuracy of the previous period, but the average precision detection completion time is still within a specified range, then increasing the task complexity weight and decreasing the first boundary threshold or the second boundary threshold, so that more defect region image sets enter short chain mode or long chain mode. The magnitude of the increase is determined based on the degree of change in the average inference accuracy. If the average precision detection completion time of the current period is longer than that of the previous period, and the average inference accuracy is still within the specified range (determined based on actual needs), then the GPU utilization weight, CPU usage weight, or network bandwidth weight is increased, and the first boundary threshold or the second boundary threshold is raised to enable more defective region image sets to adopt the chainless mode or short chain mode. The magnitude of the increase is determined based on the degree of change in the GPU utilization rate in the cloud, the average CPU usage rate of all edge terminals, or the average network bandwidth usage rate of all edge terminals, and the magnitude of the threshold increase is determined based on the degree of change in the average precision detection completion time.

[0130] The second adaptive optimization strategy should also satisfy the constraints of the feedback objective function, that is, to maximize the feedback objective function, dynamically adjust the task complexity weight, GPU utilization weight, CPU usage weight, network bandwidth weight, first boundary threshold, and second boundary threshold. The feedback objective function is: , Indicates the feedback value. Indicates the reasoning speed weight. This indicates the average time to complete the precision testing. Indicates the inference precision weight. This represents the average inference accuracy. The inference speed weight is included. Weights related to inference accuracy All are determined based on actual needs; for example, they can be... and Setting it to 0.5, for example, can... Set to 0.3, The value is set to 0.7. Therefore, the second adaptive optimization strategy is used to increase the triggering opportunities of short-chain or long-chain modes when inference accuracy decreases, and to reduce the triggering opportunities of long-chain modes when the precision detection completion time becomes longer.

[0131] According to an embodiment of the present invention, the method further includes: configuring the cloud to process each defect region image set in the following manner: performing fine detection on each defect region image set using a self-configured multimodal large model to obtain the fine category, fine category probability, and defect semantic description of each defect region; if the fine category of the defect region is a subclass of the basic category, then the fine category is taken as the comprehensive defect category of the defect region, and the comprehensive detection confidence is calculated according to a preset first fusion mechanism, wherein the fine category being a subclass of the basic category indicates that the fine category is a defect category that refines the basic category; if the fine category of the defect region is not a subclass of the basic category, then multiple compatibility candidate categories are constructed based on the basic category, the fine category, and the industrial defect knowledge base, and the score of each compatibility candidate category is calculated according to a preset second fusion mechanism, and based on the score of each compatibility candidate category, the compatibility candidate category with the highest score is selected as the comprehensive defect category, and the score of the compatibility candidate category is taken as the comprehensive detection confidence, wherein the compatibility candidate category represents a defect category that extends the basic category and the fine category to the same classification level.

[0132] According to an embodiment of the present invention, the preset first fusion mechanism is as follows:

[0133]

[0134] in, Indicates the overall detection confidence level. This represents the third adjustment weight, which is determined based on actual needs. Represents the fine category probability, This indicates the confidence level of the basic test.

[0135] According to one embodiment of the present invention, the preset second fusion mechanism is as follows:

[0136]

[0137] in, Indicates rating, This indicates the compatibility candidate category output by the target detection model at the corresponding edge. The probability, Indicates candidate categories for compatibility of multimodal large model output. The probability, Indicates category compatibility score, This represents the compatibility weight, which is determined based on actual needs.

[0138] To better understand the process of performing detailed defect detection in the cloud, the following detailed explanation is provided in conjunction with the aforementioned embodiments.

[0139] The first step is to configure the defect region image set and coarse detection information tuple of the image to be detected uploaded by the cloud receiving edge terminal, and select the chainless mode, short chain mode or long chain mode for the defect region image set. The specific selection method has been described in the previous embodiments and will not be repeated here.

[0140] The second step, after determining the depth of the thought chain, is to configure the cloud to construct a visual feature vector for each defect region in the defect region image set in order to make the reasoning results make fuller use of the coarse detection results information. ( Indicates the first Location information of each defect area Indicates the first Scale characteristics of each defect region Indicates the first Local texture gradient features of each defect region Indicates the first The contextual information of each defect region (i.e., the surrounding visual environment, structural relationships, and scene background) is used to map the visual feature vectors through a visual semantic mapping function. Matching with an industrial defect knowledge base to extract visual feature vectors Convert to structured reasoning hints And organize the structured reasoning hints into text hints. and text prompts With visual feature vectors Input a large multimodal model to perform fine detection to obtain the fine category, fine category probability and defect semantic description for each defect region.

[0141] Among them, text prompts The constraints include task role description, input image description, coarse edge detection results, candidate category range, domain knowledge constraints, and output format constraints. The task role description requires the multimodal large model to act as an industrial defect detection expert, classifying, interpreting, and assessing the risk of the input defect image. The input image description indicates that the input image is a defect region image cropped from the edges. The edge detection results include coarse detection results such as basic category labels, bounding box positions, confidence scores, defect sizes, and local texture clarity. The candidate category range constraint restricts the model to select only from a given basic category and its fine-grained defect subcategories to reduce false positives. The domain knowledge constraint requires selecting typical defect appearances, causes, and severity judgment rules from an industrial defect knowledge base. The output format constraint requires outputting structured fields, such as defect category, defect category probability, location explanation, severity, possible causes, and processing suggestions.

[0142] To improve the professionalism and consistency of the detection results, this invention also introduces an adaptive prompt optimization mechanism, which dynamically adjusts the priority, appearance position, candidate category order, weight of relevant knowledge fragments, and example matching degree of domain words in the text prompt during the prompt generation stage.

[0143] The specific optimization process of the adaptive suggestion optimization mechanism is as follows:

[0144] First, establish a domain-specific vocabulary set:

[0145]

[0146] in, The terminology used in the field of industrial defects includes terms such as "microcracks," "oxidative contamination," "poor weld joints," and "edge damage," with each term assigned a corresponding weight. .

[0147] Then, generate a task vector based on the current task characteristics. The task vector It consists of basic category labels, visual features of defective regions, device type, historical misjudgment records, etc., and calculates the relevance of domain vocabulary to the current task:

[0148]

[0149] in, This indicates the correlation between task vectors and the domain vocabulary set. Represents text or multimodal embedding functions. Cosine similarity can be used.

[0150] Then, update the domain term weights based on the calculated relevance:

[0151]

[0152] in, Update the relevance coefficients for the current task. Historical feedback coefficient, This indicates the accuracy, false positive rate, or expert feedback score of the word in historical testing.

[0153] Finally, according to the updated The text prompts are adjusted to prioritize high-weight words in the candidate category list, inference rules, and knowledge fragments, while reducing the frequency of low-weight words or placing them in the alternative categories. This allows the multimodal large model to focus more on the most relevant defect concepts in the current industrial scenario.

[0154] The third step involves combining the basic category, basic detection confidence, fine category, fine category probability, and defect semantic description of each defect region to determine the comprehensive defect category, comprehensive detection confidence, and defect semantic description. The specific analysis process has been given in the preceding embodiments and will not be repeated here.

[0155] As described in the foregoing embodiments, this invention proposes a cross-model semantic enhancement detection mechanism that integrates a lightweight target detection model with a multimodal large model through a cascaded structure, thereby achieving industrial defect detection. Specifically, the lightweight target detection model is first configured at the edge to quickly output defect locations and basic categories. Then, the cropped defect region is fed into the multimodal large model in the cloud, which generates more refined defect classifications and natural language descriptions, including location interpretation, severity, and cause inference. Through this mechanism, the efficient localization of the lightweight target detection model and the semantic generation capabilities of the large model complement each other, ultimately outputting a comprehensive detection report that integrates structured detection boxes and natural language text, thereby significantly improving the accuracy and interpretability of defect detection.

[0156] In addition, this invention proposes a dynamic thinking chain generation and adaptive reasoning optimization mechanism. This mechanism integrates coarse detection results, task complexity assessment, and cloud-edge-device system resource status to achieve dynamic selection and deep adaptive control of the thinking chain structure, enabling the cloud-edge-device system to maintain a dynamic balance between reasoning speed and detection accuracy under different task complexities and resource environments.

[0157] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) By monitoring the four core indicators of computing resource occupancy, storage resource occupancy, network bandwidth occupancy and task queue length in real time between multiple edge terminals and the cloud, the images to be detected are dynamically allocated across edge terminals, so that high throughput, low latency and high resource utilization can still be maintained under network fluctuations and task surges; (2) The lightweight target detection model and the multimodal large model are deeply integrated through a serial structure, which significantly improves the accuracy and interpretability of defect detection; (3) By integrating coarse detection results, task complexity assessment and cloud-edge system resource status, the dynamic selection and deep adaptive control of the thinking chain structure are realized, so that the cloud-edge system can maintain a dynamic balance between inference speed and detection accuracy under different task complexity and resource environments.

[0158] It should be noted that the above text involves multiple formulas, each of which includes multiple parameters, and some parameters inevitably use the same letter representation. Therefore, the meaning of each parameter shall be based on its corresponding interpretation.

[0159] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.

[0160] This invention can be a system, method, electronic device, computing device, computer-readable medium, and / or computer program product. A computer program product primarily refers to a software product that implements this solution through a computer program.

[0161] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A cloud-edge-device collaborative industrial defect detection method, used to collaboratively detect industrial defects by multiple clients, multiple edge devices, and the cloud, characterized in that, Each edge device is configured with a target detection model, and the cloud device is configured with a multimodal large model. The method includes configuring each client, each edge device, and the cloud device to perform industrial defect detection in the following manner: Each client continuously collects product images on the industrial production line to obtain images to be inspected, and uploads the images to be inspected to the assigned edge end according to the allocation decision issued by the cloud. The allocation decision instructs the images to be inspected to be assigned to the designated edge end for processing. Each edge device periodically collects its own resource status and uploads it to the cloud, and continuously receives and stores images to be detected. It then uses a target detection model to perform coarse detection on each stored image to obtain a set of defect region images and a coarse detection information tuple for each image to be detected, and uploads them to the cloud. The set of defect region images includes multiple defect regions cropped from the image to be detected, and the coarse detection information tuple includes the location information, basic category, and basic detection confidence of each defect region. The cloud receives the resource status collected periodically from each edge terminal, and based on the resource status of all edge terminals in the previous period, it generates and distributes allocation decisions for each image to be detected collected by each client using a preset storage-computing collaborative allocation strategy. It also receives the defect region image set and coarse detection information tuple of each image to be detected uploaded by each edge terminal, and uses a multimodal large model to perform fine detection on each defect region image set to obtain the comprehensive defect category, comprehensive detection confidence, and defect semantic description of each defect region. The defect semantic description includes the description of the location of the defect region, the rating of the defect severity, and the inference of the defect cause.

2. The method according to claim 1, characterized in that, The cloud-based scheduling and control module receives resource status data periodically collected from each edge device. The method also includes: The cloud-based scheduling and control module is configured to calculate the periodically changing load score of each edge based on the resource status collected periodically from each edge, and to calculate the periodically changing average load and load variance to assess whether the load is balanced among all edge devices. If the current load variance is less than the preset load balancing threshold, it indicates that the load is currently balanced among all edge devices. Conversely, if the current load variance is greater than or equal to the preset load balancing threshold, it indicates that the load is unbalanced among all edge endpoints. In this case, the scheduling control module performs a redistribution process on some edge endpoints according to a preset dynamic redistribution scheduling mechanism. The preset dynamic redistribution scheduling mechanism is as follows: Each edge with a current load score greater than or equal to the current average load and exceeding a preset load deviation threshold is designated as a migration-out edge, and each edge with a current load score less than the current average load is designated as a migration-in edge. Multiple images to be detected that have not yet undergone coarse detection on each migration-out edge are migrated to one or more migration-in edges for coarse detection. During the migration process, the real-time load score of all edge edges is calculated, and the real-time average load and real-time load variance are calculated based on the real-time load scores of all edge edges. The migration stops when the real-time load variance is less than a preset load balancing threshold. The configuration scheduling control module periodically calculates the load score for each edge terminal in the following manner: in, Indicates the first The current load score of each edge device. Indicates the first The weight of the calculation index at each edge end, Indicates the first Current computing resource utilization at each edge device Indicates the first The weight of storage metrics at each edge. Indicates the first Current storage resource utilization at each edge device Indicates the first The weights of network metrics at the edge. Indicates the first Current network bandwidth utilization at each edge. Indicates the first Weights of task metrics at the edge. Indicates the first The number of currently unprocessed images to be detected at each edge. This represents the maximum number of images currently awaiting processing across all edge endpoints.

3. The method according to claim 2, characterized in that, The method further includes: The scheduling and control module on the cloud periodically calculates the average network latency. If the average network latency is greater than or equal to a preset latency threshold for multiple consecutive periods indicated by a first preset value, it indicates that the network link between the cloud and all edge devices is congested. The cloud triggers a local asynchronous mitigation mechanism, causing some edge devices to suspend uploading the defect region image set and coarse detection information tuple of the image to be detected. The local asynchronous mitigation mechanism is as follows: Get the current round-trip latency from each edge to the cloud, and pause all edge devices whose current round-trip latency is greater than or equal to a preset latency threshold from uploading the defect area image set and coarse detection information tuple until the average network latency in multiple consecutive periods indicated by the first preset value is less than the preset latency threshold, and then resume all edge devices from uploading the defect area image set and coarse detection information tuple. The configuration scheduling control module periodically calculates the average network latency in the following manner: in, This represents the average network latency for the current period. Represents the set consisting of all edge ends. This indicates the number of all edge ends. Indicates the number of times within the current period The average round-trip latency from each edge node to the cloud.

4. The method according to claim 2, characterized in that, The method further includes configuring the cloud to generate and assign decisions for each image to be detected in the following manner: Based on the resource status of all edge terminals in the previous cycle in the scheduling control module, a preset in-memory computing collaborative allocation strategy is used to generate an allocation decision for each image to be detected acquired by each client. The preset in-memory computing collaborative allocation strategy is as follows: in, This represents the allocation decision, which indicates the allocation of the image to be detected. Assigned to the One edge end, Indicates the first The computational latency weights of each edge end Indicates the first Process the image to be detected at each edge. Required computation latency, Indicates the first Storage access weights at each edge. Indicates the image to be detected In the Storage overhead generated at each edge. Indicates the first Network transmission weights at each edge end Indicates the image to be detected In the The required network latency at each edge; The goal of the preset memory-computation collaborative allocation strategy is to allocate an edge to each image to be detected so as to minimize the weighted sum of computation latency, storage overhead and network transmission of all edge edges. The preset memory-computation collaborative allocation strategy is solved using the particle swarm optimization algorithm.

5. The method according to claim 4, characterized in that, The method further includes: configuring a first adaptive optimization strategy in the cloud to update the computation metric weights, storage metric weights, network metric weights, and task metric weights corresponding to each edge when the scheduling control module calculates the load score, and to update the computation latency weights, storage access weights, and network transmission weights corresponding to each edge when the scheduling control module generates allocation decisions, wherein the first adaptive optimization strategy is: The average coarse detection completion time, throughput, and cache hit rate of each edge are periodically recorded for each edge. If, within multiple consecutive periods indicated by the second preset value, the average coarse detection completion time of any edge continuously increases and the throughput corresponding to that edge continuously decreases, then the weight of the calculation indicator, network indicator, or task indicator corresponding to that edge is increased, as well as the weight of the calculation delay or network transmission corresponding to that edge is increased. The magnitude of the increase is determined based on the degree of change in the average coarse detection completion time and throughput of the edge. If, within multiple consecutive periods indicated by the second preset value, the average coarse detection completion time and the throughput corresponding to the edge end remain within the specified range, and the cache hit rate corresponding to the edge end continues to decrease, then the storage metric weight and storage access weight corresponding to the edge end are increased. The increase is determined based on the degree of change in the cache hit rate corresponding to the edge end. If, within multiple consecutive periods indicated by the second preset value, the average coarse detection completion time of any edge continuously shortens, the throughput continuously increases, and the cache hit rate continuously improves, then the weights of the corresponding computation metric, storage metric, network metric, task metric, computation latency, storage access, and network transmission for that edge will not be adjusted.

6. The method according to claim 1, characterized in that, The multimodal large model is configured with chainless mode, short chain mode, and long chain mode to perform fine detection on each defect region image set. The method further includes configuring the cloud to select chainless mode, short chain mode, or long chain mode for each defect region image set in the following manner: in, in, Represents the joint decision value of task resources, when When selecting the no-chain mode, When selecting short link mode, When selecting long chain mode, Indicates the first boundary threshold. Indicates the second boundary threshold. Indicates the task complexity weight. The task complexity of representing a set of defect region images. Indicates the defect size weight. This represents the average defect size corresponding to all defect regions in the defect region image set. Indicates the ambiguity weight. This represents the image blur score corresponding to the image to be detected within the set of defect region images. Indicates the weight of the number of defects. This indicates the number of defective regions in the defective region image set. This indicates that the GPU utilizes weights. Indicates the real-time utilization of cloud GPUs. Indicates CPU usage weight. This represents the real-time average CPU utilization across all edge devices. Indicates network bandwidth weight. This represents the real-time average network bandwidth utilization for all edge devices. Among them, the no-chain mode means that the multimodal large model does not execute the thought chain when performing fine detection on the defect region image set; the short-chain mode means that the multimodal large model executes a three-step reasoning thought chain when performing fine detection on the defect region image set; and the long-chain mode means that the multimodal large model executes a five-step reasoning thought chain when performing fine detection on the defect region image set.

7. The method according to claim 6, characterized in that, The method further includes: configuring a second adaptive optimization strategy in the cloud to update the task complexity weight, GPU utilization weight, CPU usage weight, network bandwidth weight, first boundary threshold, and second boundary threshold, wherein the second adaptive optimization strategy is: The average completion time and average inference accuracy of fine inspection for each defect area are periodically recorded in the cloud. If the average inference accuracy of the current period is lower than that of the previous period, and the average precision detection completion time is still within the specified range, then the task complexity weight is increased and the first boundary threshold or the second boundary threshold is decreased. The increase is determined according to the degree of change in the average inference accuracy. If the average precision detection completion time in the current period is longer than the average precision detection completion time in the previous period, and the average inference accuracy is still within the specified range, then the GPU utilization weight, CPU usage weight, or network bandwidth weight will be increased, and the first boundary threshold or the second boundary threshold will be raised. The magnitude of the increase will be determined based on the degree of change in the GPU utilization rate in the cloud, the average CPU usage rate of all edge devices, or the average network bandwidth usage rate of all edge devices. The magnitude of the threshold increase will be determined based on the degree of change in the average precision detection completion time.

8. The method according to claim 1, characterized in that, An industrial defect knowledge base is configured on the cloud, which includes various known defect categories. The method further includes configuring the cloud to process each defect region image set in the following manner: The system employs a self-configured multimodal large model to perform fine detection on each defect region image set, in order to obtain the fine category, fine category probability, and defect semantic description of each defect region. If the fine category of the defect region is a subclass of the basic category, then the fine category is used as the comprehensive defect category of the defect region, and the comprehensive detection confidence is calculated according to the preset first fusion mechanism. Here, the fine category is a subclass of the basic category, which means that the fine category is a defect category that is a refinement of the basic category. If the fine category of the defect region is not a subclass of the basic category, then multiple compatibility candidate categories are constructed based on the basic category, the fine category, and the industrial defect knowledge base. The score of each compatibility candidate category is calculated according to the preset second fusion mechanism. Based on the score of each compatibility candidate category, the compatibility candidate category with the highest score is selected as the comprehensive defect category. The score of the compatibility candidate category is used as the comprehensive detection confidence. Here, the compatibility candidate category represents the defect category that extends the basic category and the fine category to the same classification level.

9. The method according to claim 8, characterized in that, The preset first fusion mechanism is as follows: in, Indicates the overall detection confidence level. This indicates the third adjustment weight. Represents the fine category probability, This indicates the confidence level of the basic test.

10. The method according to claim 8, characterized in that, The preset second fusion mechanism is as follows: in, Indicates rating, This indicates the compatibility candidate category output by the target detection model at the corresponding edge. The probability, Indicates candidate categories for compatibility of multimodal large model output. The probability, Indicates category compatibility score, This indicates the compatibility weight.