An industrial image detection teaching method and system based on multi-agent cooperation

CN122821172APending Publication Date: 2026-09-25UNIV OF ELECTRONICS SCI & TECH OF CHINA ZHONGSHAN INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611033369.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

学生难以即时了解检测结果是否达到预期、各环节是否存在改进空间,从而难以及时调整学习策略

Benefits of technology

[0026]与现有技术相比,本发明的有益效果是:通过教学智能体集群(老师提问智能体、算法架构智能体、质量评估智能体、教育智能体)与工具智能体集群(工具智能体、代码参数智能体、沙箱运行环境)的协作架构,将全流程自动化,学生仅需上传图像和文字描述即可获得完整的检测结果与教学报告,无需掌握OpenCV编程细节,大幅降低了工业视觉检测的学习门槛。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821172A_ABST
    Figure CN122821172A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image detection teaching, and particularly discloses an industrial image detection teaching method and system based on multi-agent cooperation, which comprises the following steps: acquiring a user-input image to be detected and description information, performing demand analysis on the description information to generate a structured detection strategy; generating corresponding executable image processing code based on the detection steps in the structured detection strategy, and performing parameter iterative optimization on the executable image processing code; quantitatively scoring the execution result after parameter optimization, and collecting process data of each agent to generate a structured teaching report after meeting preset qualified conditions. Through the cooperation architecture of a teaching agent cluster and a tool agent cluster, the whole process is automated, and students can obtain complete detection results and teaching reports by only uploading images and text descriptions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image detection teaching technology, specifically to an industrial image detection teaching method and system based on multi-agent collaboration. Background Technology

[0002] With the rapid development of Industry 4.0 and intelligent manufacturing, machine vision inspection technology is playing an increasingly important role in industrial production processes such as product quality control, defect identification, and dimensional measurement. At the same time, the market demand for professionals with industrial vision inspection skills continues to grow, but traditional talent training models are struggling to meet industry needs in terms of efficiency and quality. In recent years, the rise of artificial intelligence technologies such as Large Language Modeling (LLM) and Retrieval Augmentation (RAG) has provided new technological possibilities for building intelligent and personalized teaching support systems.

[0003] Currently, industrial image detection teaching mainly adopts the following methods:

[0004] Traditional image processing instruction based on OpenCV. This approach teaches students the principles and methods of using various image preprocessing, feature extraction, and detection algorithms through programming practice. Typically, the instructor demonstrates a standard workflow, and students manually write code and adjust parameters to complete a specified detection task.

[0005] Teaching object detection based on deep learning. This approach mainly focuses on deep learning models such as YOLO and Faster R-CNN, teaching topics such as dataset annotation, model training, and inference deployment, emphasizing the application of neural networks in visual detection.

[0006] Training courses for commercial machine vision software. Professional software such as Halcon and VisionPro provide graphical environments for setting up inspection workflows and rich algorithm libraries. Enterprises often use these software for employee training, but they are costly and have relatively closed software ecosystems.

[0007] The existing teaching methods generally adopt a linear transmission model of "one-way teacher instruction + student hands-on practice," which has the following shortcomings in practical application:

[0008] Lack of intelligent teaching aids: The teaching process relies on teachers to provide targeted guidance based on individual student differences, making it difficult to achieve personalized teaching on a large scale. When students encounter problems, they often need to wait for teacher intervention, and learning efficiency is limited by the availability of teachers.

[0009] Parameter tuning relies heavily on human experience: whether it's traditional OpenCV algorithms or deep learning hyperparameters, parameter tuning requires extensive practical experience and trial and error. Beginners often find themselves overwhelmed by dozens of parameter combinations, resulting in a steep learning curve and a lengthy teaching period.

[0010] Lack of standardized detection process references: The algorithm combinations and parameter configurations required for different detection tasks (such as surface scratch detection, crooked cover detection, color anomaly detection, etc.) vary significantly. The existing teaching environment lacks a validated standardized process library, and students need to explore the algorithm links for each task on their own, resulting in high trial and error costs.

[0011] Delayed feedback on learning outcomes: Most existing tools only output detection result images or numerical indicators, lacking a systematic explanation of the detection process, parameter adjustment strategies, and result quality. Students find it difficult to understand in a timely manner whether the detection results meet expectations or whether there is room for improvement in each step, thus hindering their ability to adjust their learning strategies promptly.

[0012] To address the aforementioned issues, some studies have attempted to introduce agents or expert systems to assist in the design of teaching and testing processes. However, existing solutions typically only address the automation of a single step (e.g., focusing solely on parameter tuning or code generation), lacking a fully intelligent closed loop from requirements analysis to result feedback. Furthermore, the scalability and security of the system architecture (especially code execution security in teaching environments) still have significant room for improvement. Summary of the Invention

[0013] The purpose of this invention is to provide a teaching method and system for industrial image detection based on multi-agent collaboration, so as to solve the problems mentioned in the background art.

[0014] To achieve the above objectives, the present invention provides the following technical solution:

[0015] A teaching method for industrial image detection based on multi-agent collaboration, the method comprising:

[0016] Obtain the image to be detected and its description information input by the user, and perform requirement parsing on the description information to determine at least one detection task;

[0017] A structured detection strategy is generated based on the detection task according to the preset golden process mapping table.

[0018] Based on the detection steps in the structured detection strategy, corresponding executable image processing code is generated, and the executable image processing code is iteratively optimized for parameters.

[0019] The execution results after parameter optimization are quantitatively scored, and the preset qualification conditions are determined based on the scoring results.

[0020] After meeting the preset qualification conditions, the process data of each agent is summarized to generate a structured teaching report.

[0021] This invention also provides an industrial image detection teaching system based on multi-agent cooperation, the system comprising:

[0022] The teaching intelligent agent cluster is used to acquire the image to be detected and the description information input by the user, perform requirement parsing on the description information to determine at least one detection task, and generate a structured detection strategy based on the detection task according to the preset golden flow mapping table.

[0023] A cluster of intelligent tools is used to generate corresponding executable image processing code based on the detection steps in the structured detection strategy, and to perform parameter iterative optimization on the executable image processing code.

[0024] The teaching intelligent agent cluster also includes a quality assessment intelligent agent and an educational intelligent agent. The quality assessment intelligent agent is used to quantify and score the execution results after parameter optimization, and determine whether the preset qualification conditions are met based on the scoring results.

[0025] After the preset qualification conditions are met, the educational agent summarizes the process data of each agent and generates a structured teaching report.

[0026] Compared with existing technologies, the beneficial effects of this invention are: through the collaborative architecture of teaching intelligent agent cluster (teacher questioning intelligent agent, algorithm architecture intelligent agent, quality assessment intelligent agent, education intelligent agent) and tool intelligent agent cluster (tool intelligent agent, code parameter intelligent agent, sandbox running environment), the entire process is automated. Students only need to upload images and text descriptions to obtain complete detection results and teaching reports, without needing to master the details of OpenCV programming, which greatly reduces the learning threshold of industrial vision inspection. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention.

[0028] Figure 1 This is an overall architecture diagram provided for an embodiment of the present invention.

[0029] Figure 2 The flowchart illustrates the data structure and matching mechanism of the gold flow mapping table provided in this embodiment of the invention.

[0030] Figure 3 A flowchart illustrating the workflow of a teacher-asked intelligent agent provided in an embodiment of the present invention.

[0031] Figure 4 A flowchart illustrating the security mechanism for the parameter tuning execution environment provided in this embodiment of the invention.

[0032] Figure 5 The RAG knowledge retrieval and injection method provided in the embodiments of the present invention.

[0033] Figure 6 A flowchart for generating teaching agent reports is provided for embodiments of the present invention. Detailed Implementation

[0034] To make the technical problems to be solved, the technical solutions, and the beneficial effects of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention.

[0035] In this embodiment of the invention, an industrial image detection teaching method based on multi-agent cooperation is provided, the method comprising:

[0036] Obtain the image to be detected and its description information input by the user, and perform requirement parsing on the description information to determine at least one detection task;

[0037] A structured detection strategy is generated based on the detection task according to the preset golden process mapping table.

[0038] Based on the detection steps in the structured detection strategy, corresponding executable image processing code is generated, and the executable image processing code is iteratively optimized for parameters.

[0039] The execution results after parameter optimization are quantitatively scored, and the preset qualification conditions are determined based on the scoring results.

[0040] After meeting the preset qualification conditions, the process data of each agent is summarized to generate a structured teaching report.

[0041] like Figure 1 As shown, in this embodiment, the present invention adopts a multi-agent collaborative architecture, consisting of two major clusters: a teaching agent cluster and a tool agent cluster. The teaching agent cluster includes a teacher questioning agent, an algorithm architecture agent, a quality assessment agent, and an educational agent; the tool agent cluster includes tool agents, code parameter agents, and a sandbox runtime environment. The two clusters communicate and collaborate through structured strategy data objects DetectionStrategy and DetectionResult.

[0042] The overall system workflow is as follows: The front end receives images and text descriptions of industrial products to be inspected uploaded by users, packages the input information into a structured request, and sends it to the backend API interface; After receiving the request, the teacher-asked agent in the teaching agent cluster calls the image feature extraction module to analyze the multi-dimensional attributes of the image, and performs requirement analysis and detection task generation by combining keyword matching and large-scale language model inference; Based on whether the detection task exists in the golden process mapping table, the algorithm process is directly provided by the golden process mapping table to generate the algorithm process; After receiving the algorithm process, the tool agent cluster generates corresponding OpenCV tool code for each detection step in the process and performs sandbox security verification; The code parameter agent assigns optimization strategies to the verified code framework according to the algorithm type and performs parameter iteration and tuning; The sandbox running environment securely executes the detection code and produces result images and quantitative data in each iteration; The quality assessment agent performs a five-dimensional weighted score judgment on the execution results; If the comprehensive score is lower than 60 points, it is returned to the code parameter agent for re-tuning (maximum 20 iterations); If the comprehensive score reaches 60 points or above, it is handed over to the education agent to summarize the data of the whole process and generate a structured teaching report according to the six-segment report template, and finally returned to the front end for display.

[0043] The Golden Process Mapping Table uses a multi-level nested dictionary as its data storage structure. The top-level key is the string representing the detection task name, and the values ​​are dictionary objects containing four substructures. The specific organization of this data structure is as follows:

[0044] The algorithm pipeline substructure is identified by the key name "pipeline", and its value is an ordered list (list type). Each element in the list is a dictionary object (dict type) containing an "algorithm" field. The steps in the algorithm pipeline are arranged in the order of "preprocessing-feature extraction-quantization analysis".

[0045] The parameter configuration substructure is identified by the key name "default_params". Its value is a dictionary object, where the key is the parameter name string and the value is the default value of the parameter (which can be a number, list or string type). The parameter configuration content is different for different detection tasks. Most detection tasks (such as skewed cover detection, surface scratch detection, etc.) use the default OpenCV parameters of each algorithm, and default_params is an empty dictionary.

[0046] The validation metrics substructure is identified by the key "validation_metrics", whose values ​​are dictionary objects. Each entry contains a metric name and an expected range, using a nested dictionary format: {"<metric name>":{"expected":"<expected range>","description":"<metric description>"}}. For example, the crooked cap detection task includes two validation metrics: contour_count (expected range "1-5", indicating "number of detected contours") and aspect_ratio (expected range "0.8-1.2", indicating "aspect ratio, normal bottle caps are close to 1.0"). The bubble impurity detection task includes contour_count (expected range "1-100") and circularity (expected range "0.5-1.0", indicating "roundness, 1.0 is a perfect circle").

[0047] The verified accuracy field is identified by the key name "verified_accuracy" and its value is a floating-point number (float type), representing the percentage accuracy of the inspection process on the standard test set. Accuracy values ​​are derived from actual test verification, with accuracy ranging from 82% to 95% for different inspection tasks. Specifically, the distribution is as follows: crooked cap inspection 95% (highest), shape and contour inspection 93%, color inspection 92%, bubble and impurity inspection 91%, surface scratch inspection and high cap inspection 90%, tamper-evident ring inspection and comprehensive cap inspection 88%, surface stain inspection and thread inspection 85%, and surface defect inspection 82% (lowest).

[0048] In a preferred embodiment of the present invention, the step of performing requirement parsing on the description information to determine at least one detection task includes:

[0049] The pre-configured multi-level matching mechanism determines the standard algorithm combination corresponding to the detection task. The multi-level matching mechanism attempts the following matching levels in order of priority: exact matching based on the detection task name, knowledge base matching based on product description keywords, mapping matching based on material type and defect type, mapping matching based on educational scenario, and default degradation strategy.

[0050] In this embodiment, the system first performs hierarchical matching of user text using a keyword mapping table to accurately identify material and defect types. If "comprehensive detection" is matched, the system automatically aggregates all algorithm processes corresponding to the product. Subsequently, the system constructs structured prompts (Prompts) from image quality diagnostic results, relevant domain knowledge from RAG retrieval, and standard algorithm combination sequences, which are then input into a large language model. These structured prompts constitute the detection task. Based on this, the large model generates a detection strategy in JSON format containing complete steps. After rigorous verification of the number of steps and compliance review of auxiliary functions, a standardized algorithm tree structure is finally output to guide subsequent execution.

[0051] like Figure 2 As shown, the golden flow mapping table is configured with a multi-level matching mechanism, and algorithm matching is performed according to the following priority order to ensure that a suitable algorithm combination can be located from any perspective:

[0052] Exact matching is the first priority. The matching rule is: perform a complete string match (case sensitive) between the detection task name string and the top-level key of the GOLDEN_PIPELINE_MAPPING table. If the match is successful, the corresponding algorithm flow (pipeline) and parameter configuration will be returned directly.

[0053] Keyword matching in the knowledge base is the second priority. The matching rule is as follows: when an exact match fails, the `get_detection_tasks_from_description()` function is called to perform keyword substring matching (case-insensitive) between the user-input product description and the `PRODUCT_DETECTION_MAPPING` table in the knowledge base. After identifying the product type, the corresponding detection task list is obtained. Then, the first entry matching `GOLDEN_PIPELINE_MAPPING` is found in the detection task list, and the algorithm flow is returned. This mapping table covers multiple product categories, including bottle caps, glass bottles, metal parts, textiles, ceramic products, and food packaging.

[0054] Material-defect matching is the third priority. The matching rule is as follows: based on the identified material type (material_type) and defect type (defect_type), a two-level search is performed in the MATERIAL_ALGORITHM_MAPPING mapping table. First, the defect mapping sub-dictionary of the material is located using the material type as the key. Then, the corresponding algorithm combination is searched in the sub-dictionary using the defect type as the key.

[0055] Educational scenario matching is the fourth priority, and the matching rule is as follows: when the material-defect matching still fails, substring matching is performed in the educational scenario mapping table EDUCATION_TASK_MAPPING based on the semantics of the detection task description (case-insensitive). This table contains 8 independent educational scenarios and their corresponding algorithm combinations: shape recognition, color learning, texture analysis, size measurement, edge detection, contour extraction, image segmentation, and feature matching.

[0056] The degradation strategy is the fifth priority, and the matching rule is: if all the above levels fail to match, the default algorithm combination [GaussianBlur, Canny, findContours] is used as the degradation strategy. This algorithm combination covers the most common "preprocessing-edge detection-contour extraction" process and can give usable preliminary results in most scenarios.

[0057] like Figure 3 As shown, after the teacher asks a question, the intelligent agent receives user input and performs the following technical steps to complete requirement parsing and detection strategy generation:

[0058] The `extract_all_features()` function of the image feature extraction module is called to perform 8-dimensional feature analysis on the input image. The multi-dimensional feature results returned by `extract_all_features()` have three subsequent processing paths: injection into the RAG vector retrieval engine to recall relevant domain knowledge, input into `recommend_algorithm()` to generate preprocessing parameter recommendations and drive automatic image quality correction field by field, and serialization into natural language text and incorporating it into the LLM Prompt as the image context for code generation. This function first determines the number of channels in the input image (1 channel for grayscale images, 1 channel for color images). Figure 1 Typically 3 channels), record image size information, and then perform the following feature calculations sequentially:

[0059] Brightness calculation: After converting the input image to grayscale, calculate the arithmetic mean of the grayscale values ​​of all pixels. The calculation formula is: ,in B represents the grayscale value of the i-th pixel in the grayscale image (range 0-255), and N is the total number of pixels. A brightness value B < 60 indicates the image is too dark, and B > 220 indicates the image is overexposed.

[0060] Contrast calculation: Calculate the standard deviation of all pixel values ​​in a grayscale image. The formula is: ,in This represents the grayscale mean. A higher contrast value indicates richer image detail; C < 30 is considered low contrast, and C > 100 is considered high contrast.

[0061] Noise level estimation: Based on the median filtering residual method. First, a 3×3 kernel median filter is applied to the grayscale image to obtain a denoised image. Then, the standard deviation of the pixel difference between the original image and the denoised image is calculated. The calculation formula is: ,in The grayscale image matrix This is the standard deviation operator (i.e., calculating the standard deviation of all elements in the difference matrix). Noise level. >20 is judged as high noise 10< ≤20 is considered medium noise. ≤10 is considered low noise.

[0062] Edge density calculation: Apply the Canny edge detection algorithm (default low threshold 50, high threshold 150) to the grayscale image to obtain a binary edge image E, and calculate the proportion of edge pixels (pixels with a value greater than 0) to the total number of pixels. The calculation formula is: Edge density values ​​are between 0 and 1, with 0.01 to 0.15 being a moderate range.

[0063] Texture complexity calculation: Construct a Gabor filter bank with four directions (0°, 45°, 90°, 135°), a Gabor kernel size of 31×31, a sine wave standard deviation σ=5.0, a wavelength λ=10.0, and a spatial aspect ratio of [missing information]. Gabor filtering is applied to the grayscale image, and the arithmetic mean of the response energies (the mean of the absolute values ​​of the filtering results) in the four directions is calculated as the texture complexity index. A texture complexity greater than 20 indicates that the image contains rich texture details.

[0064] Color distribution analysis: The BGR image was converted to the HSV color space, and the mean and standard deviation of the H (hue), S (saturation), and V (lightness) channels were calculated. The percentage of pixels with a saturation greater than 30 in the S channel (colorful_ratio) was counted; when the percentage exceeded 20%, the image was considered a color image (is_colorful=True). The number of pixels in each hue interval was counted using an 18-interval histogram (each interval is 10 degrees) of the H channel. The hues corresponding to the top 3 intervals with the highest pixel percentages were selected as the dominant colors. The hue values ​​were then converted to color names using a hue threshold mapping table (0-15° corresponds to red, 15-30° to orange, 30-45° to yellow, 45-75° to green, 75-105° to cyan, 105-135° to blue, 135-165° to purple, and 165-180° to red). The color complexity calculation formula is as follows: .

[0065] Illumination uniformity detection: The grayscale image is uniformly divided into 8×8=64 grid regions. The average brightness value of each grid region is calculated, and then the average deviation of the brightness of each region from the global brightness is calculated. The formula for calculating the illumination uniformity value is: ,in This represents the mean of the brightness deviation in each region (the mean of |region_mean - global_mean|). This represents the global average brightness. The value range is from 0 to 1. A value less than 0.7 indicates uneven lighting.

[0066] Defect scale estimation: First, adaptive threshold binarization (adaptiveThreshold, Gaussian method, 11×11 neighborhood, constant 2) is applied to the grayscale image. Then, morphological closing operation with a 3×3 kernel is used for denoising. Finally, findContours is used to extract contours. For all valid contours with an area greater than 10 pixels, area statistics (minimum, maximum, and average) are calculated. The pixel size is converted to the physical size (micrometers) using the equivalent diameter (sqrt(area)) and the pixel / micrometer conversion ratio (default 5 pixels / micrometer). Defects are classified according to their maximum size: less than 100 micrometers are small-scale defects, 100-1000 micrometers are medium-scale defects, and greater than or equal to 1000 micrometers are large-scale defects (millimeters are also calculated).

[0067] If none of the above levels are matched, "Unknown" is returned. This "Unknown" flag is passed to the subsequent S304 algorithm matching process and will be used as the lookup key for the material dimension in the "Material-Defect 2D Lookup Table" step. Since this lookup table requires that both the material (material_type) and defect (defect_type) dimensions are not "unknown" to match a valid algorithm combination, the missing material dimension will cause this step to be skipped, and the system will directly trigger the final degradation strategy: use the hard-coded fallback algorithm combination ["GaussianBlur", "Canny", "findContours"] as the detection step, write it into the detection scheme (DetectionScheme) and return it to the front end for display.

[0068] The PRODUCT_DETECTION_MAPPING data structure is a dictionary with product type names as keys. When a user enters a product description, the system iterates through the keyword list of all product type entries, performing substring matching (case-insensitive) on each keyword against the user description. The first successful match determines the product type, and then the system retrieves the complete list of testing tasks corresponding to that product type. For example, if the user description contains any of the keywords "bottle cap," "cap," "mineral water bottle cap," or "beverage bottle cap," the system automatically matches the "bottle cap" product type and generates six testing items: color detection, crooked cap detection, high cap detection, anti-theft ring detection, surface scratch detection, and surface stain detection.

[0069] The DEFECT_KEYWORDS defect keyword mapping table contains 16 defect types and their respective keyword lists. The system matches the user-inputted detection task description in the following priority order:

[0070] Level 1 (Special Handling for Comprehensive Testing): The system first checks whether the testing task description contains the keyword "comprehensive testing". If it does, it directly returns the "comprehensive testing" flag.

[0071] Level 2: If comprehensive detection is not triggered, check if the detection task name is a standard detection task name in GOLDEN_PIPELINE_MAPPING (such as "color detection", "crooked cover detection", "surface scratch detection", etc.), and perform substring matching (case-insensitive). If the match is successful, return the corresponding standard detection task name.

[0072] Level 3: Check if the task name contains any of the five task type names in TASK_ALGORITHM_MAPPING (surface scratches, surface cracks, air bubbles, shape deformation, and color abnormalities), and perform substring matching (case-insensitive). If a match is successful, return the corresponding task type name.

[0073] Level 4: Iterate through the 16 defect types in DEFECT_KEYWORDS (scratches, cracks, rust, dents, burrs, deformation, bubbles, color difference, holes, stains, decay, delamination, breakage, glaze defects, thickness, and threads), and perform substring matching (case-insensitive) between the keywords of each defect type and the detection task description. If a match is successful, return the corresponding defect type name.

[0074] If none of the above levels are matched, "Unknown" is returned. This "Unknown" flag, after being passed to the algorithm matching process, will serve as the lookup key for the defect dimension in the "Material-Defect 2D Lookup Table" step. Since this lookup table requires both the material (material_type) and defect (defect_type) dimensions to be non-"Unknown" for a valid algorithm combination to be matched, the missing defect dimension will cause this step to be skipped, and the system will directly trigger the final degradation strategy: using a hard-coded fallback algorithm combination ["GaussianBlur", "Canny", "findContours"] as the detection step, writing it into the detection scheme (DetectionScheme), and returning it to the front end for display. This fallback combination is a general edge detection pipeline that can ensure the process is not interrupted, but its accuracy and specificity are not as good as the dedicated algorithm combination when a match is successful.

[0075] The system call to the `generate_detection_strategy()` function first determines the standard algorithm combination `standard_tools` through the following 5-level priority decision. This combination serves as an immutable constraint for subsequent LLM strategy generation:

[0076] The first priority for the golden pipeline is exact matching: check if the user's detection task name `detection_task` is an exact key name in `GOLDEN_PIPELINE_MAPPING` (such as "color detection", "crooked cover detection", "surface scratch detection", etc.). If a match is found, directly extract the `pipeline` array from the golden pipeline configuration, extract the `algorithm` field of each step in sequence as the `standard_tools` list, and set the matching flag `smart_match_used` to `True`.

[0077] Task algorithm mapping matching is the second priority: check if `detection_task` is a key in `TASK_ALGORITHM_MAPPING`. This mapping table contains 5 task types (surface scratches, surface cracks, bubble impurities, shape deformation, color anomalies) and their corresponding algorithm combinations. If a match is found, the corresponding algorithm list is extracted and `smart_match_used=True` is set.

[0078] Material-defect joint identification and matching is the third priority: material type identification and defect type identification are performed sequentially to obtain two labels, material_type and defect_type. The system first determines the value of defect_type:

[0079] If `defect_type` is "Comprehensive Inspection": Trigger the comprehensive inspection expansion branch, call the `get_detection_tasks_from_description(product_desc)` function to match the product description to the product inspection mapping table (`PRODUCT_DETECTION_MAPPING`), and obtain the list of all sub-inspection tasks corresponding to this product type. Iterate through each sub-task; if the sub-task name matches exactly in `GOLDEN_PIPELINE_MAPPING`, extract the list of algorithms in its pipeline; if the sub-task name matches in `TASK_ALGORITHM_MAPPING`, extract the corresponding list of algorithms. All collected algorithms are deduplicated and merged (keeping the order of first appearance), directly assigned to `standard_tools`, `smart_match_used=True`, and the subsequent material-defect joint lookup step is skipped.

[0080] If `defect_type` is neither "Comprehensive Inspection" nor "Unknown", and `material_type` is also not "Unknown", and the previous steps have all failed (`smart_match_used` is still False): Perform a two-dimensional lookup in the material-defect algorithm mapping table `MATERIAL_ALGORITHM_MAPPING` (this table is an 8-row × N-column matrix, with rows corresponding to metal / plastic / glass / ceramics / textiles / wood / composite materials / food, and columns corresponding to 16 defect types such as scratches / cracks / rust / dents / burrs / deformation / bubbles / color differences / thickness / glaze defects / holes / rot / delamination / damage / color abnormalities / surface stains, etc.). If a match is found, extract the corresponding algorithm list and set `smart_match_used` to `True`.

[0081] The education scenario mapping is downgraded to the fourth priority: If none of the above paths produce valid standard_tools (smart_match_used is still False), traverse the education task mapping table EDUCATION_TASK_MAPPING (which includes 8 education scenario tasks: shape recognition, color learning, texture analysis, size measurement, edge detection, contour extraction, image segmentation, and feature matching), and perform substring matching (case-insensitive) between the task type name and detection_task. If a match is found, extract the corresponding algorithm list.

[0082] Final default fallback: If all matches fail, use the final default algorithm combination ["GaussianBlur", "Canny", "findContours"] as a fallback.

[0083] After determining standard_tools, the system constructs a structured prompt containing the following four parts of information and sends it to the large language model:

[0084] Image quality diagnostic results text. Generated by the analyze_image_quality_before_strategy() function, it contains a list of detected quality issues (such as "image too dark", "severely blurry", "spectral reflection", etc.) and corresponding preprocessing suggestions (such as "it is recommended to add histogram equalization in step 1", "it is recommended to use sharpening filter or reshoot", "it is recommended to add GaussianBlur to smooth highlight areas in step 1", etc.).

[0085] The RAG knowledge retrieval module injects relevant domain knowledge text. If the RAG module is available, it extracts multi-dimensional image features by calling the extract_all_features() function of the image feature extraction module, retrieves the top_k relevant knowledge documents from the ChromaDB vector database using the feature vectors, and formats them into structured paragraphs (including the recommendation algorithm name, description, applicable conditions, typical use cases, and relevance percentage). If the RAG module is unavailable or no valid results are found, it uses static knowledge base text as a fallback (formatting the purpose description, applicable scenarios, and recommendation parameters of the specified algorithm combination into text paragraphs using the build_knowledge_prompt() function).

[0086] Standard algorithm combination sequence and algorithm description text. The standard_tools list determined in the aforementioned decision steps is used as a binding standard combination, along with description text of the purpose of each algorithm (ALGORITHM_PURPOSES mapping table, including GaussianBlur: "Gaussian filtering for denoising and smoothing images", Canny: "Canny edge detection, extracting edge features", threshold: "threshold binarization, segmenting target regions", morphologyEx: "morphological operations, optimizing detection results", findContours: "contour extraction, quantizing target analysis", cvtColor: "color space conversion, used for color analysis", medianBlur: "median filtering for denoising and removing salt-and-pepper noise", etc., totaling 12 algorithms) and description text of expected output (ALGORITHM_OUTPUTS mapping table).

[0087] The core rules and constraints include the following mandatory instructions: Strictly use the specified standard algorithm combination; do not add, remove, or replace algorithms; the number of steps must be exactly equal to the number of algorithms in the combination; the execution order of the algorithms must follow the specified order in the combination; It is prohibited to use the seven auxiliary functions—contourArea, arcLength, boundingRect, minEnclosingCircle, fitEllipse, convexHull, and approxPolyDP—as independent detection steps (these functions can only be used as internal calls within main functions such as findContours).

[0088] The large language model generates a detection strategy in JSON format based on the above prompts, which includes six fields: product_description (product description string), detection_task (detection task name string), image_paths (image path dictionary), detection_steps (detection step list, each step includes step_number, tool_name, purpose, and expected_output), and strategy_reasoning (strategy reasoning description string).

[0089] The system then performs two compliance checks on the generated policy:

[0090] Step count verification: If the length of the detection_steps array returned by LLM is inconsistent with the number of elements in standard_tools, the algorithm metadata query interface of AlgorithmRegistry and the ALGORITHM_PURPOSES / ALGORITHM_OUTPUTS mapping table are automatically called to reconstruct the standard step sequence in the order of standard_tools and replace the original result.

[0091] Auxiliary function interception: Iterate through the tool_name field of all steps and compare it one by one with the blacklist of auxiliary functions (contourArea, arcLength, boundingRect, minEnclosingCircle, fitEllipse, convexHull, approxPolyDP). If a blacklist function is detected as an independent step, the process is terminated immediately and an error response is returned, along with corresponding alternative suggestions (such as "It is recommended to use the findContours tool, which has integrated contour area and perimeter calculation functions").

[0092] After both verifications pass, the `_build_algorithm_tree_for_strategy()` function is called to build a complete algorithm tree structure (algorithm_tree) based on `standard_tools` and `AlgorithmRegistry`. The name, type, parameters, and parent-child relationships of each algorithm node are organized in a hierarchical tree data format starting from the root node, and the algorithm tree is attached to the strategy data for use by the subsequent tool agent during execution.

[0093] In a preferred embodiment of the present invention, the step of generating corresponding executable image processing code according to the detection steps in the structured detection strategy includes:

[0094] Obtain the name of the tool function in the detection step, and search for a code template that matches the name of the tool function in the preset security template library;

[0095] If the search is successful, the executable image processing code is generated based on the code template; if the search fails, the current detection process is terminated and a failure message is returned.

[0096] In this embodiment, after receiving the DetectionStrategy structured object from the golden process mapping table, the tool agent performs the following technical steps:

[0097] Strategy parsing and verification: The static method `DetectionStrategy.from_dict()` is called to deserialize the JSON-formatted strategy data into a type-safe data object. This method performs integrity verification: it checks whether the `step_number` of each detection step starts from 1 and increments consecutively (skipping or repeating numbers is not allowed), whether the `tool_name` utility function name is a non-empty string, and whether the `purpose` description is a non-empty string. If the verification fails, a `ValueError` exception containing the specific error field names is thrown.

[0098] Auxiliary Function Interception: Before code generation, all detection steps of the detection strategy are traversed. The `tool_name` field of each step is converted to lowercase and compared with the auxiliary function set. The auxiliary function set contains 7 function names: `contourArea` (contour area calculation), `arcLength` (contour perimeter calculation), `boundingRect` (boundary rectangle calculation), `minEnclosingCircle` (minimum enclosing circle calculation), `fitEllipse` (ellipse fitting), `convexHull` (convex hull calculation), and `approxPolyDP` (polygon approximation). These functions are all auxiliary functions in OpenCV used for secondary calculations on existing contours. They must be called internally by main functions such as `findContours` and cannot be used independently as a step in the detection process. If an auxiliary function is detected as an independent step, the process is immediately terminated and an error response containing the following information is returned: `success=False`, the `error_message` field contains the violation step number and function name, and a suggestion to use the `findContours` tool (which integrates contour area and perimeter calculation functions).

[0099] Tool code generation: For each detection step in the detection strategy, the generate_cv_toolLangChain tool function is called sequentially to generate the corresponding tool code. This function accepts two parameters: function_name (OpenCV function name string) and image_context (image context string, in the format "{product description}-{step purpose description}").

[0100] Code generation employs a template library strategy: it searches the pre-built template library TARGET_FUNCTIONS for a code template with the function name. The template library includes pre-built LangChain tool code templates for 15 commonly used OpenCV functions.

[0101] If the function name exists in the template library, code is generated based on the template, and the image_context parameter is injected into the docstring and tool description of the code to make the generated code scene-aware.

[0102] If the function name is not in the template library, `generate_cv_tool` immediately returns a failure response (success=False). The `error_message` field indicates the missing function name and the reason for the error, and suggests adding the function to the template library and retrying. Upon receiving this failure response, the tool agent terminates the current detection process, passing the failure information to the upper-level caller as is, without executing any degradation or replacement logic. This mechanism ensures that all tool code generated by the system originates from security templates that have been manually reviewed, fundamentally eliminating syntax errors, logical defects, and security risks introduced by LLM-generated code.

[0103] Sandbox security verification: Functions not in the template library have already been intercepted and returned as failures during the code generation stage. This step only processes code that is found in the template library and generated successfully. Since the template code is pre-built and manually reviewed static code, the possibility of it containing syntax errors or dangerous calls is zero. The focus of sandbox verification shifts from "security review" to "functional correctness verification," and it is divided into the following two layers:

[0104] Syntax compatibility check: The code string is compiled using Python's built-in `compile()` function. If `compile()` throws a `SyntaxError` exception, the verification fails, and a failure response containing the line number and description of the syntax error is returned. This layer serves as a defensive check to prevent template files from being accidentally corrupted during maintenance.

[0105] Synthetic Image Runtime Testing: The generated tool code is executed in a restricted sandbox environment. The sandbox loads the code via `exec()`, exposing only cv2 (OpenCV), numpy (np), PIL.Image, and a whitelist of secure built-in functions. A 128×128 pixel randomly noisy synthetic image is used as input, and the generated `apply_<function name lowercase>()` entry function is called to execute the test. Verification passes if the function throws no exceptions and its return value includes `success=True`. If any exception occurs during execution or `success=False` is returned, verification fails, and a failure response containing the exception type and error description is returned. This layer ensures the practical usability of the template code in the current runtime environment.

[0106] After successful verification, the code is encapsulated into a GeneratedTool data object, which records three fields: tool_name (tool function name string), tool_code (complete code string), and function_signature (function signature, in the format "apply_<function name lowercase>()").

[0107] Code delivery: After the tool agent completes code generation and sandbox verification for all detection steps, it encapsulates all verified GeneratedTool objects (generated_tools field) along with a technical pipeline description list (technical_pipeline field, formatted as ["Step 1:<function name>(<purpose>)","Step 2:<function name>(<purpose>)",...]) and the original policy data into a DetectionResult data object, which is then delivered to the code parameter agent via a cross-cluster communication protocol. At this point, the parameters in the code are all default values ​​(e.g., Canny threshold is (50,150), filter kernel size is 3, binarization threshold is 127, etc.), and are marked as pending optimization.

[0108] In a preferred embodiment of the present invention, the step of performing parameter iterative optimization on the executable image processing code includes:

[0109] The corresponding parameter optimization strategy is invoked based on the algorithm type of the current detection step;

[0110] During the parameter iterative optimization process, the historical memory instances of parameters are queried to obtain historically successful parameters that match the current image features as initial values. A convergence monitoring mechanism is used to determine whether the parameter adjustment has stalled. If it stalls, the iterative optimization of the current step is terminated.

[0111] In this embodiment, after receiving the code framework and DetectionResult object passed by the tool agent, the code parameter agent performs adaptive parameter tuning.

[0112] Before parameter optimization begins, the system first calls the `_load_images_for_detection()` function to load the images to be detected and execute the quality preprocessing pipeline. Image loading supports three input methods: file path (supports relative and absolute paths, first checking if it is Base64 data format before concatenating the current working directory), and Base64 encoded string (formatted as "data:image / ). <format>;base64,<base64_data> The loaded image is converted to OpenCV format using cv2.imdecode after parsing the MIME header and Base64 decoding, and then used as a PILImage object (used directly as global state). The loaded image undergoes the following six quality diagnostics and automatic repair pipelines:

[0113] Underexposure correction: When the brightness is less than 60, convert the BGR image to the YCrCb color space, perform cv2.equalizeHist() histogram equalization on the Y (brightness) channel, and then convert it back to the BGR color space. This operation enhances shadow details by stretching the brightness histogram.

[0114] Overexposure repair: When the brightness is greater than 220, apply gamma correction. Perform a transformation on each channel value p (range 0-255) for each pixel: ,in =0.7. This nonlinear transformation compresses highlight areas and stretches shadow areas.

[0115] Minor blur restoration: When the Laplacian variance is between 100 and 300, a 3×3 sharpening convolution kernel is applied for filtering. The convolution kernel matrix is: [[-1,-1,-1],[-1,9,-1],[-1,-1,-1]]. This high-pass filter kernel enhances the difference between pixels and their neighbors, improving image sharpness.

[0116] Specular reflection restoration: When oversaturated pixels (grayscale value > 250) account for more than 5%, apply a 5×5 Gaussian smoothing filter (cv2.GaussianBlur). This operation replaces the highlight pixels with the neighborhood mean, reducing the contrast of the specular reflection area.

[0117] High noise restoration: When the noise level is >15, apply a medium filter (cv2.medianBlur). The filter kernel size is adaptively determined by the algorithm recommendation module recommend_algorithm() based on the noise level: the kernel size is 7 when the noise is greater than 20, and the kernel size is 5 when the noise is less than or equal to 20.

[0118] Lighting unevenness repair: When the lighting uniformity is less than 0.7 and recommend_algorithm() suggests using equalizeHist, perform histogram equalization on the Y channel of the YCrCb color space.

[0119] After preprocessing, the preprocessed image (PILImage format) is stored in the global image state variable _current_image of the smart_metatools module for parameter tuning in subsequent detection steps. Regardless of whether preprocessing is performed, a processing log is output to record the execution status of each repair operation.

[0120] The system iterates through each detection step (DetectionStep object) in the detection strategy and selects the corresponding parameter optimization function based on the step's tool_name field (converted to lowercase). The dispatch logic consists of four branches:

[0121] Branch 1: The step where tool_name contains the substring "canny" (case-insensitive) calls the _optimize_canny() function to optimize the Canny edge detection parameters. This function manages two sets of parameters: edge detection parameters (integers low_threshold and high_threshold) and preprocessing strategy parameters (preprocessing string enumeration values: "none" / "strong_blur" / "adaptive_threshold", and an integer blur_kernel).

[0122] Branch 2: For steps where `tool_name` contains the substring "blur" or "filter" (case-insensitive), the `_optimize_blur()` function is called to optimize the filtering and denoising parameters. This function manages a single parameter, `kernel_size` (an integer representing the kernel size, forced to be odd). During iteration, the specific filtering type is selected based on the function name: `cv2.medianBlur()` is used if "median" is included, `cv2.GaussianBlur()` is used if "gaussian" is included, and `cv2.blur()` is used for all others.

[0123] Branch 3: The step where tool_name contains the substring "threshold" (case-insensitive) calls the _optimize_threshold() function to optimize the threshold binarization parameter. This function manages the parameter threshold (a binary threshold integer, ranging from 10 to 245).

[0124] Branch 4: Steps where `tool_name` contains the substring "contour" or "findcontours" (case-insensitive) call the `_optimize_contours()` function to optimize contour extraction parameters. This function manages five sets of parameters: `min_area` (minimum contour area integer), `morphology_type` (morphological operation type string enumeration: "close" / "open"), `kernel_size` (morphological kernel size integer), `iterations` (morphological operation iteration count integer), and `binarization_method` (binarization method string enumeration: "otsu" / "adaptive" / "fixed").

[0125] The remaining steps call the generic execution function _execute_generic_tool(), without performing iterative optimization, and directly return the default parameters.

[0126] Before performing initial parameter calculations, each parameter optimization function first queries the global parameter history memory instance, ParameterHistory. This instance stores historically successful parameter combinations and their corresponding scores, using an algorithm type string (such as "canny", "blur", "threshold", "contour") and an image feature vector (containing three floating-point values: brightness, contrast, and noise level) as the joint key. The current task's image feature vector is then matched against the historical records for similarity, with matching conditions that the absolute value of the brightness difference does not exceed 10% of the historical brightness value, the absolute value of the contrast difference does not exceed 15% of the historical contrast value, and the absolute value of the noise level difference does not exceed 20% of the historical noise level. If a historical record is matched, the parameters in that record are directly used as initial values, significantly reducing the number of iterations. If no matching record is found, the adaptive initial value calculation method built into each optimization function is used.

[0127] The initial value of the Canny edge detection parameter optimization algorithm is adaptively calculated as follows: If there are no historical matching records, the initial value of the low threshold (low_threshold) is calculated using the following formula: Where B is the image brightness value, This is an estimate of the noise level. The formula means that the signal quality of an image is estimated using the ratio of brightness to noise. The better the signal quality (the lower the noise ratio), the lower the initial threshold (the more sensitive the detection). The initial value of the high threshold, `high_threshold`, is set to 2.5 times `low_threshold`, i.e.:

[0128] ;

[0129] And guarantee the final (Maintain the dual threshold ratio required by the Canny algorithm). Based on this, fine-tune according to contrast: when contrast C < 30, multiply both thresholds by 0.8 (reduce the threshold to detect more edges on low-contrast images); when contrast C > 100, multiply both thresholds by 1.15 (increase the threshold to reduce noisy edges in high-contrast images).

[0130] The initial values ​​of the filtering and denoising parameter optimization algorithm are adaptively calculated: if there are no historical matching records, the initial kernel size is determined based on the noise level. When the kernel size is >20, kernel_size = 7 (large kernel for strong noise reduction); when the kernel size is <10, kernel_size = 7 (large kernel for strong noise reduction). When the kernel size is ≤20, kernel_size=5 (noise reduction in the middle kernel). When the kernel size is ≤10, kernel_size=3 (small kernel with weak noise reduction). The kernel size is always forced to be an odd number (if the calculated value is even, then add 1).

[0131] The initial values ​​for the threshold binarization parameter optimization algorithm are obtained using the Otsu automatic thresholding method: `cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)` is called to obtain the optimal Otsu threshold `T_otsu` (an integer). Then, the `adjust_parameters_with_otsu()` function is called to adaptively fine-tune the parameters based on image brightness and contrast: the adjustment factor `f` is determined according to the following rules: when brightness < 60, `f = 0.85`; when brightness > 200, `f = 1.15`; when contrast < 30, `f = 0.95`; and in other cases, `f = 1.0`. The final initial threshold is then calculated. (Limited to between 10 and 245).

[0132] In a preferred embodiment of the present invention, when the number of iterations of the parameter iterative optimization reaches a preset number and the comprehensive quality score does not meet the preset qualification condition, an aggressive reset strategy is triggered to jump the current parameter space to a preset intermediate value region in order to escape the local optimum.

[0133] In this embodiment, each parameter optimization function uses the same iteration termination logic, and the three termination conditions are connected by "OR" logic:

[0134] Termination condition 1: The quality score reaches the acceptable standard (is_acceptable=True). For contour detection, the combined conditions of (5≤contour_count≤20) and (avg_circularity≥0.5) and (score≥60) must also be met.

[0135] Termination condition two: The ConvergenceMonitor detects optimization stagnation. The stagnation judgment logic is: if the difference between the maximum and minimum quality scores for three consecutive iterations (patience=3) is less than min_delta (0.001), it means that parameter adjustments can no longer bring meaningful score improvement.

[0136] Termination condition three: The number of iterations reaches the maximum limit (MAX_ITERATIONS=20, which can be configured via environment variables).

[0137] When the iteration reaches the 10th iteration (iteration ≥ 10) and the quality evaluation result is still unacceptable (is_acceptable = False), an aggressive reset strategy is triggered. The principle of this strategy is: if the optimization process is more than halfway complete and the target is still not met, it indicates that the gradient descent direction in the current parameter space may be trapped in a local optimum, requiring a significant parameter reset to escape the local optimum. Specific reset rules are defined according to the detection type:

[0138] Radical reset for contour detection: min_area=500 (reset to the median value), morphology_type="close" (close operation), kernel_size=5, iterations=2. This reset scheme simultaneously adjusts the area filtering threshold and morphological processing intensity, jumping back from parameters that may be extremely large or small to the central region of the exploration space.

[0139] Aggressive edge detection reset: preprocessing="strong_blur" (enable strong preprocessing), blur_kernel=9 (9×9 large kernel Gaussian filter), low_threshold=80, high_threshold=200. This scheme significantly reduces noise through strong blur preprocessing before using a medium threshold for edge detection.

[0140] Aggressive reset for filtering and denoising: kernel_size=7 (medium to large kernel, balancing denoising and detail preservation).

[0141] Aggressive reset of threshold binarization: threshold=127 (the midpoint of the grayscale range, abandoning Otsu adaptive calculation and restarting the exploration using a neutral threshold).

[0142] The aggressive strategy is triggered only once (controlled by the global flag _aggressive_state["attempted"]). If it is still unqualified after resetting, the current result will be accepted and the iteration will be terminated in the next trigger.

[0143] After all detection steps have been optimized, the system collects the following three types of data and encapsulates them into a DetectionResult object for return:

[0144] The final parameters for each step are stored in the form of a dictionary, with the key being the name of the utility function (e.g., "Canny", "GaussianBlur") and the value being the dictionary of the optimal parameters finally determined for that step (e.g., {"low_threshold":50,"high_threshold":150,"preprocessing":"none","blur_kernel":0}).

[0145] Parameter iteration history: stored in list form, each element is a ParameterIteration object (containing 6 fields: iteration_number, parameter_used, result_metrics, evaluation, adjustment_reason, and next_action).

[0146] Quantization detection metrics: Stored as a dictionary, with the key being the tool function name and the values ​​being a dictionary of quantization results for that step. The quantization metrics differ depending on the detection type: edge detection includes `edge_density` and `edge_count`; denoising includes `noise_reduction_rate` and `residual_noise`; thresholding binarization includes `white_ratio`; and contour extraction includes `detected_count`, `avg_area`, `max_area`, and `min_area_detected`.

[0147] like Figure 4 As shown, since the code is generated through the template library strategy, all executable code originates from the pre-built, manually reviewed SAFE_TEMPLATES templates, making the possibility of it containing dangerous calls or malicious code zero. Therefore, only runtime exceptions need to be prevented; the security mechanism includes the following two mechanisms and one execution flow.

[0148] The system sets a timeout threshold `EXECUTION_TIMEOUT_SECONDS` (an integer in seconds, configurable via environment variables, default 30 seconds) for the complete iteration cycle of each detection step. When the total iteration optimization time of a single detection step exceeds this threshold, the parameter optimization loop of the current step is forcibly terminated through thread timeout control (a cross-platform solution using the `timeout` parameter of `threading.Timer` or `concurrent.futures`), returning a failure response containing the timeout duration and a description of the "timeout" reason, and using the highest-scoring set of parameters from the completed iterations as the final result of that step. This mechanism ensures that a single detection step will not block the entire detection process due to extreme computational demands caused by excessively large image sizes or exponential complexity triggered by algorithm parameters.

[0149] The parameter optimization loop for each detection step is wrapped in a try-except block to catch the following three types of exceptions:

[0150] OpenCV runtime exceptions (cv2.error): These are usually caused by invalid parameter values ​​(such as negative kernel size, threshold out of range) or incompatible image formats. Upon capture, the exception type and parameter context are recorded, the current iteration is skipped, and the next iteration continues using the valid parameters from the previous iteration. If cv2.error is triggered in three consecutive iterations, parameter optimization for this step is terminated, and the recorded optimal parameters are returned.

[0151] Numerical calculation errors (ValueError, ZeroDivisionError, OverflowError): These are usually caused by division by zero (e.g., edges.size is 0 in edge density calculation), numerical overflow, etc. Upon capture, the parameters that triggered the error are automatically corrected (e.g., the divisor for division by zero is replaced with the smallest positive value 1e-6), the correction operation is recorded, and iteration continues.

[0152] Memory Error: This error is usually caused by memory allocation failures triggered by excessively large images or kernel sizes. Immediately after capture, the parameter optimization of the current step is terminated, memory-sensitive parameters such as kernel size and image size are reset to safe default values, and a single detection is performed using these safe parameters as the final result.

[0153] In each iteration of parameter adaptive tuning, the system operates according to the following four sub-steps:

[0154] Sub-step 1 (Preparing the Execution Environment): Obtain the preprocessed PIL Image object from the global image state variable _current_image in the smart_metatools module. If _current_image is None (image not loaded), skip the current step and return an error response. Convert the PIL Image to OpenCV BGR format (np.array(pil_image) followed by cv2.cvtColor(..., COLOR_RGB2BGR)) and grayscale format as input data for the current iteration operation.

[0155] Sub-step two (passing the current parameter): The current parameter value calculated in this round of parameter optimization loop (e.g., {"low_threshold": 40, "high_threshold": 100, "preprocessing": "none", "blur_kernel": 0}) is directly passed as the parameter to the corresponding OpenCV function call (e.g., cv2.Canny(gray, low_threshold, high_threshold)). The parameter is passed in the native Python call stack, without code string injection or dynamic execution via exec(). This design utilizes the function call interface of template code, ensuring type-safe parameter passing with zero parsing overhead.

[0156] Sub-step 3 (Execution and Result Collection): Call the corresponding OpenCV function to perform the detection operation and obtain the return value. Process the detection results according to the function return type: if a binary image or a processed image (numpyndarray) is returned, calculate the corresponding detection metrics (edge ​​density, white area ratio, contour quantity and area statistics, etc.); if a list of contours is returned, pass it to evaluate_contour_quality() to calculate the shape quality metrics. All metrics are uniformly encapsulated into a result_metrics dictionary and passed to evaluate_detection_quality() for scoring.

[0157] Sub-step four (cleaning up intermediate states): Release references to large NumPy arrays created during this iteration (such as binary images, edge images, and intermediate filtering results) to avoid memory accumulation. Reset temporary variables specific to this iteration in the iteration record. The global state of _current_image remains unchanged during the iteration process (operations in each step do not modify the global image), ensuring that the execution environment is independent between different iterations of the same detection step and between different detection steps.

[0158] In a preferred embodiment of the present invention, before generating the structured detection strategy, the method further includes:

[0159] A query text is constructed based on the image features input by the user, and relevant domain knowledge is retrieved from a pre-constructed knowledge vector database using the query text.

[0160] The retrieved relevant domain knowledge is then injected into the strategy generation prompts, and the structured detection strategy is generated based on the injected strategy generation prompts.

[0161] In this embodiment, as Figure 5 As shown, the RAG (Retrieval Enhanced Generation) knowledge retrieval module improves the accuracy of large language model generation and detection strategies by converting external OpenCV algorithm knowledge bases into vector representations and dynamically retrieving and injecting them during inference. This module comprises two stages: offline vectorization and online retrieval.

[0162] The offline knowledge vectorization phase is executed once during system initialization and is implemented by the KnowledgeVectorizer class. It includes the following processing steps:

[0163] Load the knowledge base JSON file. The knowledge base JSON file contains three main categories of knowledge entries: the `algorithms` field (a dictionary of algorithm entries, with the algorithm name as the key; each algorithm entry contains a description, a category string, an `applicable_conditions` dictionary (e.g., {"min_brightness":30,"max_noise":20}), and a `typical_use_cases` list of typical use cases), the `decision_rules` field (a list of decision rules; each rule contains an `id` rule ID, a `condition` condition text, a `recommendation` recommended solution text, a `reason` reason text, and a `params` dictionary of recommended parameters), and the `pipeline_templates` field (a dictionary of process templates, with the template name as the key; each template contains a description and a list of steps).

[0164] Construct structured text descriptions, and build texts in different formats for each type of knowledge item.

[0165] The system uses an embedding model to convert structured text into floating-point vectors. Each text entry is independently converted into a vector representation using Zhipu AI's embedding-3 model (1024 dimensions). After conversion, the vectors are L2 normalized (each component is divided by the vector's L2 norm) to ensure that all vectors have a magnitude of 1, making subsequent cosine similarity calculations dependent only on the vector direction. If the embedding model is unavailable (e.g., API key not configured or network unreachable), the system uses simulated embedding: generating a 1024-dimensional standard normally distributed random vector and performing L2 normalization, outputting a warning log "Using simulated embedding vector".

[0166] Stored to a ChromaDB vector database. Create or recreate a ChromaDB collection named "opencv_knowledge" (delete it if it already exists), and add vectorized documents to the collection in batches. The collection is persisted to the directory specified in the configuration (default . / vector_db), using ChromaDB's PersistentClient to implement data persistence.

[0167] The online knowledge retrieval phase is executed each time a detection strategy is generated, and is completed collaboratively by the KnowledgeRetriever and KnowledgeInjector classes:

[0168] The query text is constructed based on image features. The `build_query_from_features()` method is called, which takes an `ImageFeatures` object as input and dynamically constructs the Chinese query text according to the following rules based on the values ​​of each dimension of the feature object: >At 20:00, add "High noise level, noise reduction required", 10< When C < 20, add "Medium noise level"; when C < 30, add "Low contrast, needs enhancement"; when U < 0.7, add "Uneven illumination"; when defect_scale.scale_category="small", add "Small defect scale"; when defect_scale.scale_category="large", add "Large defect scale"; when defect_features.defect_type is not "none", add "Defect type: <type>". If no condition is triggered, use the default query text "General image processing".

[0169] Vector similarity retrieval. The query text is transformed into a 1024-dimensional query vector using an embedding model (also L2 normalized), and the `query()` method of the ChromaDB collection is called to perform vector similarity retrieval. The retrieval parameters are: `query_embeddings=[query vector]`, `n_results=top_k` (default 5). ChromaDB returns a dictionary of results containing four fields: `ids`, `documents`, `metadatas`, and `distances`, where `distances` is the cosine distance (values ​​range from 0 to 2, where 0 indicates identical results and 2 indicates completely opposite results). The system converts the cosine distance to cosine similarity (similarity = 1 - distance) and filters out low-quality results with similarity below `similarity_threshold` (default 0.1).

[0170] Knowledge formatting. The `format_for_prompt()` method is called to convert the retrieved list of knowledge documents into structured Prompt text paragraphs. Documents are formatted by type: `algorithm` type is formatted as "###Recommendation Algorithm: <Name>-Description: <Description>-Applicable Conditions: <Conditions>-Typical Application: <Use Case>"; `rule` type is formatted as "###Decision Rule-<Content of Each Line of Rule Text>"; `template` type is formatted as "###Process Template: <Name>-Steps: <Step Sequence>". All document content is summarized under the heading "[KNOWLEDGE] Related Domain Knowledge (Retrieved from Knowledge Base):".

[0171] Knowledge Injection. The `inject_to_prompt()` method of the `KnowledgeInjector` class is called to inject the formatted knowledge paragraphs into the base prompt words generated by the strategy. The logic for selecting the injection position is as follows: The system searches for the "[TARGET]" marker string in the base prompt words. If found, the knowledge paragraph is inserted before the marker (content after the marker is unaffected). If the "[TARGET]" marker is not found but the prompt word contains the string "task:", the knowledge paragraph is inserted after the newline character at the end of the line containing "task:". If neither is found, the knowledge paragraph is inserted at the very beginning of the base prompt words. This injection strategy ensures that relevant knowledge is read and understood preferentially before the large language model processes core task instructions.

[0172] Degradation processing. When the ChromaDB client fails to initialize (e.g., the vector database directory does not exist or is corrupted), the embedding API is unavailable (the API key is not configured or the network is unreachable), or the retrieval result is empty, the system automatically degrades to use the static knowledge base. The static knowledge base works through the `build_knowledge_prompt()` function: it receives a list of standard algorithm names (e.g., `["GaussianBlur","Canny","findContours"]`), traverses each algorithm name, obtains the purpose description from the `ALGORITHM_PURPOSES` mapping table, obtains the applicable conditions and framework template from the algorithm registry `AlgorithmRegistry`, splices all the information into a plain text paragraph, and directly appends it to the end of the policy generation prompt as knowledge context.

[0173] As a preferred embodiment of the present invention, the comprehensive quality score is obtained by weighted calculation of the execution result based on a plurality of evaluation indicators. When the comprehensive quality score does not meet the preset qualification condition, the target indicator with the lowest score among the plurality of evaluation indicators is determined, a corresponding parameter adjustment suggestion is generated according to the type of the target indicator, and the parameter adjustment suggestion is fed back to the tool agent cluster to guide the next round of parameter iteration.

[0174] In this embodiment, the quality evaluation agent `QualityEvaluationAgent` adopts an independent weighted multi-index comprehensive scoring method to perform quantitative scoring on the detection result with a score ranging from 0 to 100. The core design principles of this agent are: the scoring logic and the parameter adjustment logic are completely decoupled, and the scoring standard is fixed, so as to ensure the objectivity and consistency of the evaluation result.

[0175] The quality evaluation agent uses five first-level evaluation indicators, and each indicator has a weight coefficient, a scoring function and a corresponding adjustment suggestion generation rule:

[0176] Contour number score (weight = 0.30): A piecewise constant scoring function is adopted. The scoring standard is: when the number of contours c = 0, the score = 10 (almost no detection result); when 1 ≤ c < 5, the score = 40 (too few detection results); when 5 ≤ c ≤ 20, the score = 90 (ideal range, the corresponding number of detection targets is most conducive to accurate quantitative analysis); when 20 < c ≤ 50, the score = 55 (too many, may contain noise contours); when 50 < c ≤ 200, the score = 35 (excessively many, obviously contains a large amount of noise); when c > 200, the score = 15 (severe over-detection). The design basis of this piecewise function is as follows: in industrial inspection scenarios, the number of detected valid contours is usually between 5 and 20 (covering the main target area and a small number of background structures). Too few contours indicate that the parameter is too strict, leading to missed detection, while too many contours indicate that the parameter is too loose, resulting in noise being also identified as targets.

[0177] Confidence score (weight = 0.25): scoring is performed based on the consistency of contour areas. First, the area list of all valid contours is obtained, the mean area and standard deviation of area are calculated, and then the coefficient of variation of area CVA is calculated as the ratio of the area standard deviation to the area mean. The scoring criteria are: when CVA < 0.5, = 90 (high area consistency, reliable detection result); when 0.5 ≤ CVA < 1.0, = 70 (relatively good area consistency); when 1.0 ≤ CVA < 2.0, = 50 (large area difference); when CVA ≥ 2.0, = 30 (extremely large area difference, which may contain multiple irrelevant targets). The coefficient of variation of area reflects the scale consistency of the detected contours. When detecting defects of the same type, the defect sizes usually fluctuate within a certain range, and a smaller coefficient of variation indicates higher reliability of the detection result.

[0178] Edge density score (weight = 0.20): the scoring criteria are: when edge density d < 0.01, = 20 (edges are too sparse, a large number of features may be missed); when 0.01 ≤ d ≤ 0.15, = 85 (ideal range, moderate edge density, which not only captures sufficient target edge features but also does not include excessive background noise); when 0.15 < d ≤ 0.30, = 60 (excessive edges, which may include some noise or texture edges); when d > 0.30, = 25 (excessive edges, a large amount of noise or texture is misidentified as valid edges). The design basis of this piecewise function is as follows: valid edges of typical industrial product images usually account for between 1% and 15% of the total pixels.

[0179] Shape feature score (weight = 0.15): the quality_score value (0-100) output by the contour shape quality evaluation function evaluate_contour_quality() is directly obtained from the detection result. This value has integrated four sub-dimensions: the average roundness (25%), average convexity (25%), average compactness (25%) and adjusted average aspect ratio (25%) of all contours. If no quality_score is included in the detection result (for non-contour detection scenarios), the default value 50 is used.

[0180] Uniform distribution score (weight = 0.10): A simplified evaluation is performed based on the number of contours. When c ≤ 5, that is = 50 (too few contours, the distribution evaluation is of limited significance); when 5 < c ≤ 20, that is = 80 (ideal range, assuming the contours are reasonably distributed in the image); when 20 < c ≤ 50, that is = 60 (acceptable); when c > 50, that is = 40 (too many contours, the distribution may be extremely non-uniform).

[0181] The comprehensive score is calculated by a weighted sum formula: . The calculation result is limited to the interval [0, 100] by the clip function. Each weight coefficient represents the importance of the corresponding indicator in the comprehensive evaluation. The number of contours has the highest weight (0.30), because the detected number directly reflects the rationality of parameter setting; the confidence degree takes the second place (0.25), because area consistency reflects the reliability of detection; edge density (0.20) and shape features (0.15) provide auxiliary reference; the uniform distribution has the lowest weight (0.10), because this indicator adopts a simplified evaluation method, and the accuracy is relatively limited.

[0182] Qualification judgment: When the comprehensive score ≥ QUALIFIED_THRESHOLD (60.0 points), the detection result is judged as qualified (is_qualified = True, need_retry = False), and the process enters the education report generation stage. < 60.0 points, the detection result is judged as unqualified (is_qualified = False, need_retry = True).

[0183] Adjustment suggestion generation: When the result is judged as unqualified, the system finds the indicator with the lowest original (unweighted) score among the five indicators, and selects the corresponding suggestion text from the predefined adjustment suggestion template library according to the type of the indicator. The suggestion template library contains 25 suggestions (5 indicators × 5 score intervals). For example, when the contour number score is < 20 points, the generated suggestion is "Too few contours, it is recommended to reduce the min_area threshold or use closing operation to connect broken edges"; when the edge density score is < 30 points, the generated suggestion is "Edge density is too low, it is recommended to reduce the Canny threshold or enhance image contrast"; when the confidence score is < 40 points, the generated suggestion is "The confidence of the detection result is low, it is recommended to check the contour area distribution or adjust the detection parameters".

[0184] Iteration Abort Handling: When the outer iteration loop reaches MAX_ITERATIONS (20 iterations) and the overall score is still below 60.0, the abort handling process is triggered. The system calls the generate_abort_report() function to generate an abort report containing the following four parts: an explanation of the abort reason ("Maximum number of iterations (20 iterations) has been reached, and the detection result still does not meet the passing standard") and the final score; an overview of the 20 complete iterations, displaying the iteration number, status flag (✓ success / ✗ failure), score value, and suggestion summary (the first 50 characters); the complete adjustment suggestion text for the last evaluation; and report header and footer markers.

[0185] The suspension report should be attached to the teaching report and shown to students.

[0186] In a preferred embodiment of the present invention, the generation of the structured teaching report includes:

[0187] The data generated by the teaching intelligent agent cluster and the tool intelligent agent cluster throughout the entire process are summarized, and the structured teaching report is generated according to a six-segment report template that includes an overview of the learning task, the design of the detection process, the technical link, the parameter adjustment process, the result analysis, and the practical suggestions.

[0188] When the number of detection tasks is greater than 1, a comprehensive detection function code integrating all detection tasks is attached to the structured teaching report.

[0189] In this embodiment, as Figure 6 As shown, the educational agent is activated after passing the quality assessment (overall score ≥ 60 points) and performs the following technical steps to generate a structured teaching report.

[0190] The educational agent collects the following data from other agents through a cross-cluster communication protocol:

[0191] Data source 1 (Teacher's questioning agent): product_description (product description string, such as "mineral water bottle cap"), detection_items (list of detection item strings, such as ["color detection", "crooked cap detection", "surface scratch detection"]), image_paths (image path dictionary), and algorithm_tree (algorithm tree structure dictionary, which is a list of algorithm nodes corresponding to each detection task and is used preferentially).

[0192] Data source 2 (tool intelligence agent): generated_tools tool code list (GeneratedTool object list), technical_pipeline technical link description list (string list, formatted as "step N:<function name>(<purpose>)").

[0193] Data source 3 (code parameter agent): parameter_iterations is a list of parameter iteration history records (a list of ParameterIteration objects, each containing 6 fields), final_parameters is a dictionary of final optimization parameters (the key is the tool function name, and the value is a dictionary of parameters), and quantitative_results is a dictionary of quantitative detection results.

[0194] Data source four (quality assessment agent): score (overall floating-point score), metrics (dictionary of scores for each indicator), is_qualified (boolean value indicating whether the test is qualified), and adjustment_suggestion (text string indicating adjustment suggestions).

[0195] The system generates a structured teaching report based on a predefined six-segment template, totaling approximately 1000 lines. It includes an overview of the learning tasks, testing process design, parameter adjustment procedures, testing result analysis, and practical suggestions.

[0196] When the testing process is terminated due to failure to meet the maximum number of iterations (20), the educational agent appends a termination report to the end of the generated local report. The termination report, generated by the `generate_abort_report()` function, includes the reason for termination ("Maximum number of iterations (20) reached, test results still do not meet the passing standard"), the final score (including a comparison with the passing threshold of 60.0 points), a summary list of scores and suggestions for all iteration records, and complete adjustment recommendations. The report uses prominent separators and warning labels to mark the termination information, allowing recipients to clearly distinguish between normally completed tests and tests terminated due to quality issues.

[0197] This invention also provides an industrial image detection teaching system based on multi-agent cooperation, the system comprising:

[0198] The teaching intelligent agent cluster is used to acquire the image to be detected and the description information input by the user, perform requirement parsing on the description information to determine at least one detection task, and generate a structured detection strategy based on the detection task according to the preset golden flow mapping table.

[0199] A cluster of intelligent tools is used to generate corresponding executable image processing code based on the detection steps in the structured detection strategy, and to perform parameter iterative optimization on the executable image processing code.

[0200] The teaching intelligent agent cluster also includes a quality assessment intelligent agent and an educational intelligent agent. The quality assessment intelligent agent is used to quantify and score the execution results after parameter optimization, and determine whether the preset qualification conditions are met based on the scoring results.

[0201] After the preset qualification conditions are met, the educational agent summarizes the process data of each agent and generates a structured teaching report.

[0202] In a preferred embodiment of the present invention, the teaching intelligent agent cluster further includes a teacher questioning intelligent agent and an algorithm architecture intelligent agent, and the tool intelligent agent cluster includes a tool intelligent agent, a code parameter intelligent agent and a sandbox running environment;

[0203] The teacher-asked agent is used to perform the requirement parsing and the generation of the structured detection strategy;

[0204] The algorithm architecture agent is used to generate a customized detection process based on the dynamic reasoning capability of a large language model when the detection task cannot be matched by the golden process mapping table.

[0205] The tool agent is used to generate the executable image processing code based on the secure template library;

[0206] The code parameter agent is used to perform the parameter iterative optimization;

[0207] The sandbox runtime environment is used to execute the executable image processing code in a restricted environment and produce execution results.

[0208] This invention automates the entire process through a collaborative architecture of teaching intelligent agent clusters (teacher questioning intelligent agent, algorithm architecture intelligent agent, quality assessment intelligent agent, and education intelligent agent) and tool intelligent agent clusters (tool intelligent agent, code parameter intelligent agent, and sandbox running environment). Students only need to upload images and text descriptions to obtain complete detection results and teaching reports, without needing to master the details of OpenCV programming, which greatly reduces the learning threshold of industrial vision inspection.

[0209] This invention designs a five-level progressive matching mechanism, covering standardized algorithm processes for more than 30 pre-verification detection tasks. If a matching fails at each level, it automatically degrades to the next level. Finally, the general default process of "preprocessing → edge detection → contour extraction" is adopted as a fallback, ensuring that a usable algorithm combination can be located from any perspective of describing the detection requirements, avoiding matching failures caused by differences in terminology.

[0210] This invention employs a pre-built template library strategy. All tool code originates from 15 commonly used OpenCV function templates that have undergone manual review. Functions not in the template library are intercepted during the code generation stage and return explicit failure information, fundamentally eliminating the uncertainty introduced by LLM-generated code. Simultaneously, the sandbox runtime environment executes code within a restricted function whitelist and performs runtime verification using synthesized images, further ensuring the usability and security of the code.

[0211] This invention introduces three innovative mechanisms in parameter optimization: First, a parameter history memory mechanism stores historically successful parameters using algorithm type and image feature vector as a joint key. New tasks can directly reuse historically optimal parameters as initial values ​​through similarity matching of brightness, contrast, and noise levels, significantly reducing the number of iterations. Second, a convergence stagnation detection mechanism automatically terminates the process when the score does not improve substantially after three consecutive iterations, avoiding ineffective iterations that waste computational resources. Third, an aggressive reset strategy actively resets the parameters to the central region of the exploration space when the optimization is more than halfway complete (the 10th iteration) and still fails to meet the target, in order to escape local optima and avoid continuous degradation caused by misjudgment of gradient direction.

[0212] This invention separates the quality assessment agent from the code parameter agent and uses a fixed five-dimensional weighted scoring system (contour quantity, confidence level, edge density, shape features, and distribution uniformity) to objectively and quantitatively evaluate the detection results. The scoring criteria do not change with the iteration process, ensuring that each evaluation result can be compared horizontally, and outputting specific adjustment suggestions for the lowest score indicator when it is unqualified.

[0213] This invention offline vectorizes and stores the OpenCV algorithm knowledge base (including algorithm entries, decision rules, and process templates) in ChromaDB. During online inference, it dynamically constructs relevant knowledge for query text retrieval based on image features (noise level, contrast, illumination uniformity, defect type and scale, etc.). The retrieval results are formatted by type and injected into the strategy to generate prompt words. This allows the LLM to refer to professional knowledge that is highly relevant to the current image features when generating detection strategies, thereby improving the targeting and accuracy of the strategies.

[0214] The teaching report in this invention fully records the entire process from requirements analysis to final testing, including the algorithm tree display for each testing task, detailed records and adjustment strategies for each round of parameter iteration, final quantitative testing indicators, and practical suggestions. When multiple testing tasks are involved, a comprehensive testing function code integrating all testing tasks is automatically generated and appended to the end of the report, enabling students to not only "know what" but also "know why".

[0215] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.< / format>

Claims

1. A teaching method for industrial image detection based on multi-agent collaboration, characterized in that, The method includes: Obtain the image to be detected and its description information input by the user, and perform requirement parsing on the description information to determine at least one detection task; A structured detection strategy is generated based on the detection task according to the preset golden process mapping table. Based on the detection steps in the structured detection strategy, corresponding executable image processing code is generated, and the executable image processing code is iteratively optimized for parameters. The execution results after parameter optimization are quantitatively scored, and the preset qualification conditions are determined based on the scoring results. After meeting the preset qualification conditions, the process data of each agent is summarized to generate a structured teaching report.

2. The teaching method for industrial image detection based on multi-agent collaboration according to claim 1, characterized in that, The step of parsing the description information to determine at least one detection task includes: The pre-configured multi-level matching mechanism determines the standard algorithm combination corresponding to the detection task. The multi-level matching mechanism attempts the following matching levels in order of priority: exact matching based on the detection task name, knowledge base matching based on product description keywords, mapping matching based on material type and defect type, mapping matching based on educational scenario, and default degradation strategy.

3. The teaching method for industrial image detection based on multi-agent collaboration according to claim 1, characterized in that, The step of generating corresponding executable image processing code based on the detection steps in the structured detection strategy includes: Obtain the name of the tool function in the detection step, and search for a code template that matches the name of the tool function in the preset security template library; If the search is successful, the executable image processing code is generated based on the code template; if the search fails, the current detection process is terminated and a failure message is returned.

4. The teaching method for industrial image detection based on multi-agent collaboration according to claim 1, characterized in that, The parameter iterative optimization of the executable image processing code includes: The corresponding parameter optimization strategy is invoked based on the algorithm type of the current detection step; During the parameter iterative optimization process, the historical memory instances of parameters are queried to obtain historically successful parameters that match the current image features as initial values. A convergence monitoring mechanism is used to determine whether the parameter adjustment has stalled. If it stalls, the iterative optimization of the current step is terminated.

5. The teaching method for industrial image detection based on multi-agent collaboration according to claim 1, characterized in that, When the number of iterations for parameter optimization reaches a preset number and the overall quality score fails to meet the preset qualification condition, an aggressive reset strategy is triggered to jump the current parameter space to a preset intermediate value region in order to escape the local optimum.

6. The teaching method for industrial image detection based on multi-agent collaboration according to claim 1, characterized in that, Before generating the structured detection strategy, the following is also included: A query text is constructed based on the image features input by the user, and relevant domain knowledge is retrieved from a pre-constructed knowledge vector database using the query text. The retrieved relevant domain knowledge is then injected into the strategy generation prompts, and the structured detection strategy is generated based on the injected strategy generation prompts.

7. The teaching method for industrial image detection based on multi-agent collaboration according to claim 1, characterized in that, The comprehensive quality score is obtained by weighting the execution result based on multiple evaluation indicators. When the comprehensive quality score does not meet the preset qualification conditions, the target indicator with the lowest score among the multiple evaluation indicators is determined, and a corresponding parameter adjustment suggestion is generated according to the type of the target indicator. The parameter adjustment suggestion is fed back to the tool intelligence agent cluster to guide the next round of parameter iteration.

8. The teaching method for industrial image detection based on multi-agent collaboration according to claim 1, characterized in that, The generation of structured teaching reports includes: The data generated by the teaching intelligent agent cluster and the tool intelligent agent cluster throughout the entire process are summarized, and the structured teaching report is generated according to a six-segment report template that includes an overview of the learning task, the design of the detection process, the technical link, the parameter adjustment process, the result analysis, and the practical suggestions. When the number of detection tasks is greater than 1, a comprehensive detection function code integrating all detection tasks is attached to the structured teaching report.

9. A teaching system for industrial image detection based on multi-agent collaboration, used to implement the teaching method for industrial image detection based on multi-agent collaboration as described in any one of claims 1-8, characterized in that, The system includes: The teaching intelligent agent cluster is used to acquire the image to be detected and the description information input by the user, perform requirement parsing on the description information to determine at least one detection task, and generate a structured detection strategy based on the detection task according to the preset golden flow mapping table. A cluster of intelligent tools is used to generate corresponding executable image processing code based on the detection steps in the structured detection strategy, and to perform parameter iterative optimization on the executable image processing code. The teaching intelligent agent cluster also includes a quality assessment intelligent agent and an educational intelligent agent. The quality assessment intelligent agent is used to quantify and score the execution results after parameter optimization, and determine whether the preset qualification conditions are met based on the scoring results. After the preset qualification conditions are met, the educational agent summarizes the process data of each agent and generates a structured teaching report.

10. The industrial image detection teaching system based on multi-agent collaboration according to claim 9, characterized in that, The teaching intelligent agent cluster also includes a teacher questioning intelligent agent and an algorithm architecture intelligent agent, and the tool intelligent agent cluster includes a tool intelligent agent, a code parameter intelligent agent and a sandbox runtime environment; The teacher-asked agent is used to perform the requirement parsing and the generation of the structured detection strategy; The algorithm architecture agent is used to generate a customized detection process based on the dynamic reasoning capability of a large language model when the detection task cannot be matched by the golden process mapping table. The tool agent is used to generate the executable image processing code based on the secure template library; The code parameter agent is used to perform the parameter iterative optimization; The sandbox runtime environment is used to execute the executable image processing code in a restricted environment and produce execution results.