A large model cooperative detection system and method based on three-stage cascade and prompt autonomous evolution

CN122821093APending Publication Date: 2026-09-25CHINA THREE GORGES RENEWABLES (GRP) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611004061.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

其一,传统单一模型检测方案难以兼顾全场景精度与实时性:轻量化小模型在复杂背景、罕见缺陷或低对比度条件下漏检率与误报率显著升高,而直接部署大模型则面临边缘设备算力不足、推理延迟不可接受的问题

Benefits of technology

[0014]本发明的有益效果:本发明通过构建基于生命阶段标签的三级级联检测架构,实现了检测路径与缺陷演变趋势的自适应匹配,在新生期利用完整三级链路确保高精度检出,在稳定期跳过昂贵的大模型调用大幅降低推理成本与延迟,在退化期直接启用大模型并注入历史未确认线索有效避免小模型误判,从而在检测精度、计算效率和资源利用率之间取得了优异平衡;同时,本发明基于历史劣质Prompt库的文本变异机制,能够从过往错误案例中自动提取对立样本并实施反义替换或否定插补,使模型在遇到相似困难场景时主动规避已知错误模式,实现了检测提示的持续自主进化,显著提升了系统在复杂光照、小目标、罕见缺陷等挑战性条件下的鲁棒性与泛化能力,并且通过将未确认线索作为残差提示逐级传递,有效增强了级联模型间的信息协同,最终达到了降低漏报率与误报率、节省边缘端算力开销、简化人工干预的技术效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821093A_ABST
    Figure CN122821093A_ABST
Patent Text Reader

Abstract

The application discloses a large and small model cooperative detection system and method based on three-level cascade and Prompt autonomous evolution, relates to the technical field of computer vision and intelligent detection, and realizes adaptive matching of a detection path and a defect evolution trend, ensures high-precision detection by using complete three-level links in a newborn period, skips expensive large model calling to greatly reduce reasoning cost and delay in a stable period, directly enables the large model and injects historical unconfirmed clues to effectively avoid small model misjudgment in a degradation period, so that an excellent balance between detection precision, calculation efficiency and resource utilization is achieved; meanwhile, the application realizes continuous autonomous evolution of detection prompts, significantly improves the robustness and generalization ability of the system under challenging conditions such as complex illumination, small targets and rare defects, and effectively enhances information cooperation between cascade models by transmitting unconfirmed clues as residual prompts level by level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and intelligent detection technology, and in particular to a collaborative detection system and method for large and small models based on three-level cascade and Prompt autonomous evolution. Background Technology

[0002] Computer vision object detection technology has evolved from traditional manual feature methods to deep convolutional neural networks, and now to multimodal large-scale models, achieving significant results in fields such as industrial defect identification and security monitoring. Traditional methods rely on manually designed feature operators, which are difficult to handle complex scenarios such as changes in lighting and differences in target scale. Breakthroughs in deep learning technology, especially object detection frameworks represented by the YOLO series and Faster R-CNN, have significantly improved detection accuracy and speed through end-to-end feature learning, becoming the mainstream technology route for industrial visual inspection. However, single models still lack generalization ability and robustness when facing changing on-site environments (such as sudden changes in lighting, weather interference, and target miniaturization) and new fault modes. In recent years, multimodal large-scale models (such as Qwen-VL and GPT-4V) have provided new ideas for understanding complex scenes with their massive pre-trained knowledge and vision-language joint reasoning capabilities, but their high computational cost and high inference latency make them difficult to deploy directly for real-time edge detection tasks. To address this, a collaborative detection architecture using small and large models has emerged. A lightweight small model performs rapid initial screening, while the large model conducts detailed verification of problematic samples, striking a balance between accuracy and efficiency. Simultaneously, prompting engineering techniques, through the design of task-specific guidance instructions, enable the large model to adapt to new tasks without updating parameters, providing a low-cost approach to dynamic model evolution. However, existing collaborative detection methods often employ fixed strategies in path selection, lacking the ability to adaptively adjust based on defect evolution trends, and failing to effectively utilize historical detection knowledge for continuous optimization of prompts.

[0003] Analysis of existing technologies reveals the following key technological bottlenecks in visual inspection systems for industrial scenarios. First, traditional single-model detection schemes struggle to balance accuracy and real-time performance across all scenarios: lightweight small models exhibit significantly higher false positive and false negative rates in complex backgrounds, rare defects, or low-contrast conditions, while directly deploying large models faces challenges such as insufficient computing power on edge devices and unacceptable inference latency. Second, existing methods for coordinating large and small models typically employ a static two-tier architecture of initial screening with a small model followed by verification with a large model. This architecture fails to dynamically adjust the cascading depth based on the evolving state of the inspected object (e.g., the trend of defect emergence and expansion), resulting in the continued use of expensive large models during defect stabilization periods, leading to resource waste; and the failure to activate large models in a timely manner during periods of rapid defect deterioration, missing opportunities for accurate diagnosis. Third, existing prompting engineering largely relies on manually designed static prompts, lacking an autonomous evolution mechanism based on historical error cases. When encountering inputs similar to historical failure patterns, the model will repeat the same errors, failing to continuously learn and improve from mistakes. Summary of the Invention

[0004] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this section, the abstract and title of the invention. Such simplifications or omissions shall not be used to limit the scope of the present invention.

[0005] In view of the aforementioned existing problems, the present invention is proposed.

[0006] To solve the above technical problems, the present invention provides the following technical solution: a collaborative detection system for large and small models based on three-level cascading and Prompt autonomous evolution, characterized in that it includes: a cascading path selection module, which acquires the identification identifier and historical detection records of the current detection object, determines the life stage label of the detection object based on the changing trend of continuous detection results, and determines the calling path of the three-level cascading based on the life stage label; a first-level small model initial detection module, which starts the first-level small model detection according to the calling path, and outputs the first-level detection conclusion and the first-level unconfirmed clue; and a second-level mutation detection module, which, if the calling path includes a second level, combines the first-level detection conclusion with the first-level unconfirmed clue. The first-level unconfirmed clues are concatenated into the second-level pre-prompts. After being input into the second-level model, the second-level detection conclusion and the second-level unconfirmed clues are output. Simultaneously, when the second-level detection conclusion matches the error type in the historical poor-quality Prompt library, an opposing sample is selected from the historical poor-quality Prompt library to perform text mutation on the current Prompt. The third-level final judgment storage module, if the call path includes the third level, inputs the unconfirmed clues output from the first two levels as residual prompts into the third-level large model to obtain the final detection conclusion. The final detection conclusion is then appended to the historical detection record, and the error cases in this detection are stored in the historical poor-quality Prompt library.

[0007] Furthermore, the present invention also includes a collaborative detection method for large and small models based on three-level cascading and Prompt self-evolution, characterized by: acquiring the identification identifier and historical detection records of the current detection object; determining the life stage label of the detection object based on the changing trend of continuous detection results; determining the calling path of the three-level cascading based on the life stage label; starting the first-level small model detection according to the calling path, outputting the first-level detection conclusion and the first-level unconfirmed clue; if the calling path includes a second level, then concatenating the first-level detection conclusion and the first-level unconfirmed clue into a second-level pre-prompt, inputting it into the second-level model, and outputting the second-level detection conclusion and the second-level unconfirmed clue; simultaneously, when the second-level detection conclusion matches the error type in the historical poor-quality Prompt library, selecting an opposing sample from the historical poor-quality Prompt library to perform text mutation on the current Prompt; if the calling path includes a third level, then inputting the unconfirmed clues output by the first two levels as residual prompts into the third-level large model to obtain the final detection conclusion; and appending the current detection conclusion to the historical detection record, while storing the error cases in the current detection in the historical poor-quality Prompt library.

[0008] As a preferred embodiment of the present invention, the determination of the life stage label of the detection object includes: reading the most recent n detection conclusions corresponding to the identification identifier, each detection conclusion including the defect type and defect area; if the number of detections is less than n, the life stage label is marked as the nascent stage; comparing the defect area in two adjacent detection conclusions, and calculating the direction of area change: if the defect area strictly increases sequentially in n consecutive detections, it is determined to be the degradation stage; if the defect area strictly decreases sequentially, it is determined to be the stable stage; if the direction of change is inconsistent, there are equal areas, or the number of detections is less than n, it is all marked as the nascent stage.

[0009] As a preferred embodiment of the present invention, the method of determining the three-level cascaded call path based on the life stage label includes: the newborn stage calls the complete three-level path of the first level → the second level → the third level; the stable stage calls the first level → the second level, with the second level outputting the final detection conclusion, skipping the third level; the degenerate stage directly calls the third level, skipping the first and second levels; the input residual prompt of the third level extracts the most recent unconfirmed clue from the history record, and if there is none, an empty prompt is used.

[0010] As a preferred embodiment of the present invention, the step of initiating the first-level small model detection includes: scaling the current detection image to the input size of the first-level small model, obtaining the category label, defect area estimate, and corresponding feature response heatmap for each candidate target after forward inference; defining the contrast mean as the arithmetic mean of the response values ​​of all pixels in the heatmap; extracting continuous pixel regions with response values ​​lower than the contrast mean, and recording the position coordinates of the continuous pixel regions as blurred regions; mapping the blurred regions back to the original detection image and cropping the corresponding image blocks; generating natural language descriptions based on the statistical characteristics of the image blocks as first-level unconfirmed clues; and outputting the category label and defect area estimate as the first-level detection conclusion, and the unconfirmed clue statement as the first-level unconfirmed clue, together to the buffer for the next level to call.

[0011] As a preferred embodiment of the present invention, the step of outputting the second-level detection conclusion and the second-level unconfirmed clue after inputting the second-level model includes: the second-level model outputting the following three fields: second-level detection conclusion, reason for unconfirmation, and error type keywords; the second-level detection result includes defect type, defect area, and confidence level; the reason for unconfirmation is natural language text, explaining the ambiguity or uncertainty of this level; the error type keywords are strings in a predefined set, generated by the model based on its own output features; the first-level unconfirmed clue and the reason for unconfirmation output by the second-level model are concatenated according to a format to form the second-level unconfirmed clue.

[0012] In a preferred embodiment of the present invention, the step of selecting an opposing sample to perform text mutation on the current Prompt includes: matching the error type keywords with the error type field of each record in the historical poor-quality Prompt library; if a match is successful, selecting the poor-quality Prompt with the highest semantic similarity to the current Prompt in terms of the error type keywords as the opposing sample; extracting the complete phrase containing the error type keywords in the opposing sample; searching for phrases containing the same keywords or synonyms in the current Prompt; if found, replacing the core adjectives or verbs in the corresponding phrases with corresponding words from a predefined antonym dictionary; if no suitable antonym is found, inserting a negative word before the corresponding phrase. The mutated string serves as the mutated Prompt, used to replace or supplement the original input of the second-level model.

[0013] As a preferred embodiment of the present invention, the final detection conclusion includes defect type, defect area and confidence score; the defect area, detection time and first-level unconfirmed clues of this detection are added to the historical record of the corresponding detection object; when the confidence of the final detection conclusion is lower than the preset confidence threshold, the input image of this detection, the final detection conclusion and the error type field are assembled into a record and added to the historical poor quality Prompt library.

[0014] The beneficial effects of this invention are as follows: By constructing a three-level cascaded detection architecture based on life stage labels, this invention achieves adaptive matching between the detection path and the defect evolution trend. During the nascent stage, the complete three-level link ensures high-precision detection; during the stable stage, it skips expensive large model calls, significantly reducing inference costs and latency; and during the degradation stage, it directly activates the large model and injects historical unconfirmed clues to effectively avoid misjudgments by small models. This achieves an excellent balance between detection accuracy, computational efficiency, and resource utilization. Simultaneously, based on the text mutation mechanism of a historically poor-quality Prompt library, this invention can automatically extract opposing samples from past error cases and implement antonymous replacement or negation imputation. This allows the model to proactively avoid known error patterns when encountering similar difficult scenarios, achieving continuous autonomous evolution of detection prompts. This significantly improves the system's robustness and generalization ability under challenging conditions such as complex lighting, small targets, and rare defects. Furthermore, by passing unconfirmed clues as residual prompts level by level, it effectively enhances information collaboration between cascaded models, ultimately achieving the technical effects of reducing false negative and false positive rates, saving edge computing power, and simplifying manual intervention. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart illustrating a method for collaborative detection of large and small models based on three-level cascade and Prompt autonomous evolution, as shown in this invention.

[0016] Figure 2 This is a structural diagram of a large-scale model collaborative detection system based on three-level cascade and Prompt autonomous evolution, as shown in this invention. Detailed Implementation

[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0018] Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort should fall within the scope of protection of this invention.

[0019] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0020] According to an embodiment of the present invention, in combination Figure 1 The diagram shown illustrates a collaborative detection system for large and small models based on a three-level cascade and Prompt autonomous evolution, comprising: The cascade path selection module obtains the identification mark and historical detection records of the current detection object, determines the life stage label of the detection object based on the changing trend of continuous detection results, and determines the three-level cascade calling path based on the life stage label. The first-level small model initial detection module starts the first-level small model detection according to the calling path, and outputs the first-level detection conclusion and the first-level unconfirmed clues; The secondary mutation detection module, if the call path includes the second level, concatenates the first-level detection conclusion and the first-level unconfirmed clue into a second-level pre-prompt, inputs it into the model in the second level, and outputs the second-level detection conclusion and the second-level unconfirmed clue; at the same time, when the second-level detection conclusion matches the error type in the historical poor Prompt library, it selects the opposite sample from the historical poor Prompt library to perform text mutation on the current Prompt; The three-level final judgment storage module, if the call path includes the third level, will input the unconfirmed clues output by the first two levels as residual prompts into the third-level large model to obtain the final detection conclusion; and will append the current detection conclusion to the historical detection record, while storing the erroneous cases in the current detection into the historical poor quality Prompt library.

[0021] like Figure 2 As shown, the present invention also includes a method for collaborative detection of large and small models based on three-level cascade and Prompt autonomous evolution, comprising: S1: Obtain the identification mark and historical detection records of the current detection object, determine the life stage label of the detection object based on the changing trend of continuous detection results, and determine the three-level cascaded calling path based on the life stage label.

[0022] S1.1: In response to the trigger command of the current detection task, first obtain the unique identification of the wind power component to be detected (such as tower, blade, bolt, etc.). The identification is pre-stored in the historical detection record table of the cloud-edge collaborative database and associated with the equipment ledger information of the component.

[0023] Based on the identification identifier, a query request is initiated to the database to read the n most recent (preferably 3 times in this embodiment) detection conclusions corresponding to that identifier. Each detection conclusion includes at least the following fields: detection timestamp, defect type (e.g., rust, damage, oil stains, etc.), and defect area (expressed as the area of ​​the connected components of the pixel-level segmented mask or the physical area after camera calibration, in mm). 2 If the number of records returned by the query is less than three (for example, the component is a newly installed device or the number of historical detections is less than 3), the current detection object is directly labeled as the newborn stage, and the label is written to the current task context cache. At the same time, the subsequent area trend calculation steps are skipped, and the process proceeds to S1.3 for path selection.

[0024] Furthermore, during the process of retrieving historical detection results, the system also needs to verify the data integrity of each record. If a record is missing a defect area field or the defect area is an invalid value (such as zero or a negative value), then that record is considered invalid, and the missing field is recursively supplemented from the three most recent valid records. If the total number of valid records is still less than three, then the life stage label is also determined to be in the nascent stage.

[0025] The defect area is calculated using instance segmentation masks based on the output of deep convolutional neural networks (such as YOLOv8-Seg or MaskR-CNN), and converted into physical size using pre-calibrated pixel equivalents to eliminate scale errors introduced by different UAV flight altitudes or fixed-point camera focal length variations.

[0026] The storage and retrieval of historical detection records are achieved through a two-level system: edge node caching in the cloud-edge collaborative architecture and persistent database in the cloud. The edge nodes cache the results of the five most recent detections to accelerate the reading process.

[0027] S1.2: Compare the defect areas in two adjacent inspection results and calculate the direction of area change: If in three consecutive inspections (denoted as t-2, t-1, and t, where t is the current inspection number), the defect area (denoted as...) Construct the adjacent comparison operator:

[0028]

[0029] The rules for determining the direction of area change are as follows: like and If it is a monotonically increasing trend, the corresponding life stage label is the degeneration period. like and If it is determined to be a monotonically decreasing trend, the corresponding life stage label is the stable period; like and If the sign is opposite (one positive and one negative), or either is 0, or there is any instance where the area value cannot be obtained, then it is determined to be a non-monotonic trend, and the corresponding life stage label is the nascent stage.

[0030] The above determination result serves as the life stage label for the current detection object and is stored in the task cache variable for use in the next sub-step.

[0031] To further improve the robustness of the discrimination, a threshold for ignoring area changes can be set. (For example =5mm 2 Or a relative rate of change of 1%, when or When there is no significant change, it is considered to be in a non-monotonic trend (i.e., the nascent stage) to avoid misjudgment due to measurement noise. Simultaneously, this invention supports grouping and determining defect types: if the defect types in the three most recent inspection results are inconsistent (e.g., the first is corrosion, the second is damage), it is directly determined to be in the nascent stage because the area values ​​of different defect types are not comparable. The consistency verification of the defect type is performed during calculation. , Previously executed. This trend determination logic is configurable, allowing different settings to be configured for different components of the wind turbine equipment (such as blades, towers, and gearboxes). Thresholds are set to accommodate the different degradation rates of different components.

[0032] S1.3: Based on the life stage labels output by S1.2, call the path selection mapping table pre-installed in the cloud-edge collaborative system to determine the three-level cascaded inference path of the current detection task.

[0033] The three-level cascaded reasoning path consists of the following three levels of models in sequence: Level 1: Lightweight visual small model, used for rapid target detection and preliminary defect localization; Level 2: Medium-sized model, used to receive the output of Level 1 and perform refined classification and area regression; Level 3: Multimodal large model, with visual-language joint reasoning ability, used to handle complex or ambiguous cases.

[0034] The mapping rules are as follows: During the nascent stage, a complete three-level path is selected, that is, the first level → second level → third level are executed sequentially, and the output of each level is passed to the next level in turn; During the stable stage, a truncated path is selected, that is, only the first level → second level are executed, and the second level outputs the final detection conclusion (including defect type, defect area and confidence level) and directly stores it in the result library, skipping the third level call; During the degradation stage, a skip path is selected, that is, the first and second levels are skipped, and the third level large model is called directly.

[0035] At this point, the input residual prompt for the third level needs to extract the most recent unconfirmed clue from the history (see S4.1 for details). If none is available, an empty string is used.

[0036] Furthermore, after path selection, the system needs to generate a path identifier and bind it with the identification identifier, lifecycle stage label, and timestamp of the currently detected object to form a task record, which is then written to the task cache queue of the edge node. Simultaneously, computing resources are pre-allocated according to the selected path: for nascent paths, edge GPU resources are reserved for first and second-level inference, and asynchronous requests are made to the cloud-based large model service; for stable paths, only edge node computing power is used; for degenerate paths, a synchronous call is directly initiated to the cloud-based large model service.

[0037] In addition, the degradation path requires an additional historical unconfirmed clue retrieval module. This module filters the unconfirmed clue fields (stored in JSON format) generated by the first-level small model in the most recent detection from the object's historical detection records. If they exist, they are appended to the end of the third-level input Prompt; otherwise, an empty string is appended. The resource allocation and clue retrieval mentioned above are completed synchronously in stage S1.3 to reduce the waiting time for subsequent cascading executions.

[0038] It should be noted that the dynamic path selection mechanism of this invention adaptively adjusts the model calling strategy according to the defect evolution trend: in the nascent stage, it uses the collaboration of large and small models across the entire link to ensure detection accuracy; in the stable stage, it skips the expensive large model call, saving about 60% of inference cost and latency; in the degradation stage, it directly uses the large model and injects historical unconfirmed clues to avoid misjudgment by the small model.

[0039] S2: Start the first-level small model detection according to the calling path, and output the first-level detection conclusion and the first-level unconfirmed clues.

[0040] S2.1: Provided that the determined call path includes the first-level small model (i.e., the nascent or stable phase path), the system reads the current detection image (usually with a resolution of 1920×1080 or higher) from the task cache queue and scales it to the fixed input size preset by the first-level lightweight small model.

[0041] The first-level small model adopts an improved YOLOv8n or YOLOv10-Nano architecture, with ≤10M parameters and inference latency ≤50ms / frame (tested on edge computing devices such as NVIDIA Jetson Orin).

[0042] The scaling operation employs bilinear interpolation to maintain the image aspect ratio, padding any insufficient areas with zeros (letterbox method). The scaled image tensor is then normalized and input into a small model for forward inference.

[0043] The reasoning process includes: extracting multi-scale feature maps through the backbone network, fusing features through the neck network, and finally outputting the detection results of candidate targets by the head.

[0044] The feature response heatmap is taken from the feature map of the last convolutional layer in the detection head and generated by the Grad-CAM or Grad-CAM++ method. The resolution is related to the size of the input image after downsampling (for example, the heatmap size is 20×20 when the original input is 640×640).

[0045] Furthermore, to adapt to the characteristics of multi-scale targets coexisting in wind power equipment defect detection (such as the simultaneous occurrence of large-scale tower corrosion and minor bolt loosening), the small model in this invention adopts a multi-scale detection mechanism: The target is predicted at three feature layers of different scales, each corresponding to a different receptive field. The output data structure of the small model is a list, with each element containing the following field: class id (Defect category index), confidence (confidence level, 0-1), bounding box (bounding box coordinates, including the center point of x and y, width, and height), area es (Defect area estimate, or pixel count if a segmentation mask is used), heatmap (response heatmap with the same width and height as the input image, obtained through upsampling interpolation). In addition, the model generates intermediate feature maps during inference, especially the output feature map of the last convolutional layer before the detection head (denoted as F, size...). (For example, 512×20×20). This feature map will be stored in memory.

[0046] To adapt to the characteristics of multiple scale targets coexisting in wind power equipment defect detection (such as the simultaneous occurrence of large-scale tower corrosion and minor bolt loosening), the small model in this step adopts a multi-scale detection mechanism: the target is predicted on three feature layers of different scales (such as 80×80, 40×40, and 20×20), and each feature layer corresponds to a different receptive field.

[0047] The small model outputs a list of detection results, with each element containing the fields mentioned above. When multiple candidate targets are detected, detection results with a confidence level greater than a preset threshold (e.g., 0.4) are retained, and redundant boxes are removed using non-maximum suppression (NMS, IoU threshold 0.5). If no targets are detected, an empty conclusion list is output and marked as True. After inference, the list of detection results and the feature map F of the last convolutional layer are temporarily stored in the memory cache of the edge nodes. The size of feature map F is... ,in The number of channels (e.g., 512). and The feature map height and width are (relative to the downsampling factor of the input image; for example, a 32x downsampling factor would correspond to a 20x20 feature map for a 640x640 input).

[0048] S2.2: For each candidate target in the output (a target with confidence > threshold and retained by NMS), the system generates the feature response heatmap H corresponding to the target based on the saved feature map F of the last convolutional layer using the Grad-CAM or Grad-CAM++ method.

[0049] Specifically, based on the confidence gradient of the target class c, the weights of each channel of the feature map F are calculated, weighted summed, and then activated by ReLU to obtain a result of the same size as the feature map. The heatmap is then upsampled to the original input image size (or detection box size) using bilinear interpolation.

[0050] First, calculate its average contrast. .definition The arithmetic mean of the response values ​​of all pixels in the heatmap is calculated using the following formula:

[0051] This mean represents the overall activation intensity level of the model for the target region.

[0052] Then, each pixel in the heatmap is traversed, and all pixels with response values ​​lower than the adjusted threshold are extracted.

[0053] To accommodate different target sizes and the non-uniformity of model response distribution, this invention sets an adjustable sensitivity coefficient. (default =1.0), adjust the threshold to .when When the threshold is less than 1, the threshold is lowered, and more low-response areas are extracted (increasing the range of the blurred area); when When the threshold is >1, the threshold is increased, and only regions with extremely low response are extracted (reducing blurred areas). Original extraction conditions ( =1) means below the mean. .

[0054] This coefficient can be dynamically adjusted based on feedback from subsequent cascaded detections (for example, if the second or third level frequently confirms that clues not confirmed by the first level are false alarms, then the coefficient will be increased). (To reduce sensitivity).

[0055] Furthermore, for cases where multiple candidate targets overlap or are adjacent, it is necessary to merge and deduplicate the blurred regions: if the blurred regions of two different targets overlap in the image space or the distance between them is less than a preset value (such as 5 pixels), they are merged into a larger blurred region. The coordinates of all blurred regions are finally uniformly transformed to the original input image coordinate system (through the inverse transformation of the scaling factor) and stored as a polygon point set or the minimum bounding rectangle.

[0056] S2.3: Map the obtained blurred region coordinates (located in the original input image coordinate system) back to the currently detected original resolution image (the unscaled source image).

[0057] The mapping formula is as follows: Let the original image width be... ,high Image width after scaling ,high Then any pixel coordinate in the blurred region ( , The coordinates mapped to the original image are:

[0058]

[0059] The mapped coordinates may be non-integers, and nearest neighbor or bilinear interpolation is used to round them.

[0060] Based on the mapped polygon point set or rectangle, the corresponding image patch is cropped from the original image. During cropping, a certain margin is extended outward (e.g., 10 pixels) to ensure that the context information of the blurred area edges is not lost. If the cropped image patch exceeds the image boundary, boundary padding is performed (fill value is 0 or edge pixels are copied).

[0061] For each image patch extracted from a blurred region, the system also needs to record metadata, including: the target category it belongs to (inherited from the candidate target category in S2.1), the position of the blurred region in the original image (center point coordinates and width and height), and the average response value of the region in the heatmap.

[0062] This metadata, along with the image patches themselves, is stored in the cache for subsequent natural language description generation. When multiple blurred regions exist for the same target, multiple image patches are extracted and sorted from largest to smallest region area, with larger blurred regions (which usually represent more significant blurred features) being processed first.

[0063] In addition, to prevent excessive image patches from causing excessive overhead in subsequent description generation, a maximum limit is set for the number of blurred regions. (For example, 5), if more than that, only the one with the largest area will be retained. Each region.

[0064] S2.4: For each captured image patch, call a lightweight visual-language description model (such as BLIP-2 or CLIP combined with a text template) to generate a natural language description.

[0065] The descriptive model was pre-tuned on industrial defect image blocks and can output statements such as blurred textures at the edges, insufficient local contrast, weak regional response, and high texture repetition leading to ambiguity.

[0066] The specific generation process is as follows: the image patch is scaled to the standard input size of the description model (e.g., 224×224), features are extracted by a visual encoder, and then short sentences (no more than 15 words in length) are generated by a text decoder. To reduce computational overhead, a rule-based template method can also be used: the statistical features of the image patch (grayscale histogram variance, Laplacian gradient energy, local binary mode entropy, etc.) are calculated, and the feature values ​​are mapped to a preset text description template library based on the interval they fall into.

[0067] For multiple ambiguous regions associated with the same target, their respective generated description statements are concatenated into a single string in the format "Region 1: {Description 1}; Region 2: {Description 2};...", which serves as the first-level unconfirmed clue for that target.

[0068] If no candidate targets are detected during the entire detection task (S2.1 output is empty), a special description is generated: No targets were detected across the entire field, which may indicate global illumination anomalies or model mismatch. All unconfirmed cue statements must be timestamped and include the model version number for subsequent recording by a poor-quality Prompt library. The generated descriptions are UTF-8 encoded and limited to 256 characters in length to ensure compatibility with the cue word templates in the subsequent second-level model. This description generation process is performed entirely at the edge, without relying on large cloud models, with a latency of ≤30ms per image patch.

[0069] Furthermore, the category label and defect area estimate obtained in S2.1 are used as the first-level detection conclusion, and the unconfirmed clue statement generated in S2.4 is used as the first-level unconfirmed clue. Both are encapsulated into a JSON-formatted data structure. This data structure contains the following fields: fixed value L1, ISO 8601 format, current detection object ID, task ID, etc.

[0070] If multiple candidate targets exist, each target independently stores its detection conclusion and corresponding unconfirmed clues. The encapsulated data object is written to the cache through the edge node's internal message bus (such as Redis or MQTT) and an expiration time (e.g., 30 minutes) is set for subsequent second-level (S3) or third-level (S4) calls.

[0071] While outputting to the buffer, the system also needs to compare the defect area in the first-level detection conclusion with the most recent area in the historical records (if any) and calculate the area change rate. This change rate is added to the output data as auxiliary information but does not participate in the decision-making of the current cascaded path. In addition, if the confidence level of the first-level small model output is lower than a preset threshold (e.g., 0.5), a system-level prompt needs to be added to the unconfirmed clues: the first-level confidence level is too low, and a review is recommended.

[0072] All output data is transmitted using TLS encryption to ensure data security. Finally, the time taken for this first-level detection (from image input to output encapsulation) is recorded in the performance monitoring log for subsequent model iteration and optimization.

[0073] After outputting the first-level detection conclusions and unconfirmed clues, post-processing steps are performed: the tensor cache of the current detected image in the edge GPU memory is released, the temporary heatmap matrix generated in S2.2 is destroyed, and the connection to the cache area is closed. Simultaneously, the first-level unconfirmed clues detected are persistently stored in the historical record table of the detected object for extraction of residual clues from S4.1 in subsequent degradation paths. Storage uses an append-only mode, retaining a maximum of five most recent records; records exceeding this limit are overwritten in a round-robin fashion.

[0074] If the current call path is in a stable period (only the first level + the second level are executed, and the third level is not executed), an asynchronous verification task needs to be triggered after the first level small model detection is completed: mark the case where the deviation between the defect area in the first level detection conclusion and the average area of ​​the three most recent historical data exceeds 30% as an abnormal fluctuation, and add this mark to the unconfirmed clues.

[0075] If it is a newborn path (requiring the execution of the third level), no additional operation is required. Finally, the system decides whether to unload the memory occupied by the first-level small model according to the preset scheduling strategy (unload to save resources if there are no new detection tasks in the short term; otherwise, keep it in a hot-load state).

[0076] S3: If the call path includes a second level, the first level detection conclusion and the first level unconfirmed clue are concatenated to form a second level pre-prompt. After inputting into the model in the second level, the second level detection conclusion and the second level unconfirmed clue are output. At the same time, when the second level detection conclusion matches the error type in the historical poor Prompt library, an opposing sample is selected from the historical poor Prompt library to perform text mutation on the current Prompt.

[0077] S3.1: When the determined call path contains a model in the second level (i.e., a path in the nascent or stable phase), the system reads the first-level detection conclusion and the first-level unconfirmed clues output by S2 from the cache.

[0078] Construct the input Prompt for the second-level model, with the following format: Concatenate the first-level detection conclusion (defect type and defect area) and the first-level unconfirmed clues according to the template to form a pre-prompt, and append it to the header of the second-level model's input Prompt. The specific template is: Conclusion: {Defect Type}, Area: {Defect Area}; Unconfirmed Clues: {First-Level Unconfirmed Clues}.

[0079] Here, {Defect Type} is taken from the category label output in S2 (e.g., tower corrosion), {Defect Area} is taken from the physical area or normalized pixel area, and {Level 1 Unconfirmed Clues} is taken from the descriptive statement generated in S2.4. This pre-prompt is immediately followed by the model's task instruction in the second level: Based on the above information, please perform fine-grained detection on the image and output the detection conclusion, the reason for unconfirmation, and error type keywords. The entire Prompt uses UTF-8 encoding and its length does not exceed 512 tokens.

[0080] The second-level model employs a lightweight multimodal architecture (such as MobileViT or EfficientNet + text encoder), with 100M to 500M parameters, supporting joint image and text input. The model receives two inputs: the original detection image (scaled to the model input size, e.g., 448×448); and the constructed text prompt. Internally, the model fuses visual and textual features through a cross-modal attention mechanism. If the unconfirmed cue in the first level is an empty string (e.g., no target was detected in the first level), the prompt only includes the conclusion: None; unconfirmed cue: None. Furthermore, to enhance the second level's utilization of historical information, the most recent detection conclusion (if any) for the current detected object can be appended to the end of the prompt as a historical reference: {historical defect type, area}. This historical information is extracted from the most recent (t-1) detection conclusion from the three most recent records read in S1.1.

[0081] S3.2: Forward inference is performed by the model in the second level, outputting three independent fields: second-level detection conclusion, reason for non-confirmation, and error type keywords. The second-level detection conclusion is structured data, including defect type (which may be the same as or modified from the type output in the first level) and defect area (in mm). 2 The `confidence` parameter is a natural language text generated by the model's text decoder, used to explain any ambiguity or uncertainty at this level. Examples include: target edge-to-background contrast below 0.3 leading to unreliable segmentation; strong reflection in the leaf area; and loss of local texture information. The `error type` keyword is a string from a predefined set, including but not limited to: missed detection of excessively small targets, edge segmentation errors, low-confidence misjudgments, interference from abnormal lighting, and confusion due to inter-class similarity. The model selects the best-matching error type from this set using an additional classification head (softmax output), or outputs "no significant error."

[0082] In the specific implementation, the output layer of the second-level model contains three branches: detection branch: outputs the defect type, area and confidence level through the regression head and classification head; reason generation branch: uses a lightweight Transformer decoder to generate natural language reasons (maximum length 50 words) based on the intermediate features of the detection branch; error type classification branch: inputs the feature vector of the detection branch into a fully connected layer, outputs the probability distribution of predefined error types, and takes the category corresponding to the highest probability as the error type keyword.

[0083] When the model determines that the detection result is reliable (confidence level ≥ 0.8) and there is no significant ambiguity, the reason for non-confirmation can be output as "no obvious uncertainties," and the error type keyword output as "no significant errors." All output fields are temporarily stored in the model inference engine's memory in JSON format.

[0084] S3.3: Concatenate the generated Level 1 unconfirmed clue with the generated Level 2 unconfirmed reason according to the specified format to form the Level 2 unconfirmed clue. The concatenation format is: Level 1 unconfirmed: {Level 1 clue}; Level 2 unconfirmed reason: {Level 2 reason}.

[0085] In this structure, {Level 1 Clues} are directly taken from the output string of S2.4, and {Level 2 Reasons} are taken from the Unconfirmed Reasons field of S3.2. If a certain level of clue or reason is empty, the corresponding part will be output as "None". The length of the concatenated string is limited to 512 characters and uses UTF-8 encoding.

[0086] The second-level unconfirmed clues are not only used for residual hints in the subsequent third-level (emergent path) process, but are also persistently stored in the historical records of the detected object for future research and analysis or expansion of a poorly designed prompt library. Simultaneously, the system calculates the semantic similarity (using cosine similarity or BERT embedding) between the first-level clues and the second-level reasons. If the similarity is below 0.3, it indicates a significant discrepancy between the judgments of the two levels of models. In this case, a system prompt is appended to the end of the unconfirmed clue: "The judgments of the two levels of models differ significantly; a thorough review is recommended." This similarity threshold is configurable.

[0087] S3.4: If it is a stable path, the second-level detection conclusion is directly used as the final detection conclusion and output to the result database. Before output, an integrity check is required: the defect type, defect area, and confidence level fields in the detection conclusion must all be non-empty, and the confidence level must be ≥0 (if the confidence level is lower than the preset threshold of 0.5, a warning log is appended, but the conclusion is still output). Subsequently, all subsequent cascading operations of the current detection task are terminated, the resources occupied by the second-level model are released, and the detection record (including the first and second-level conclusions and unconfirmed clues) is written to the historical database. The process ends, and the third level is no longer called.

[0088] If it is a newborn path, the second-level detection conclusion, the second-level unconfirmed clues, and (if any) the mutation prompt generated in S3.5 will be output to the buffer for use by the third level. The output data is in JSON format and includes the field detection conclusion, unconfirmed clues, mutation prompt, or an empty string, task ID, timestamp, etc. if none are present.

[0089] During the stabilization phase, the system performs an additional operation: comparing the defect area in the second-level detection results with the historical area trend calculated in S1.2. If the direction of area change does not match the historical trend (e.g., decreasing during the stabilization phase), it is marked as an anomaly, and this mark is recorded in the results database for subsequent analysis. During the nascent phase, the system asynchronously preloads the runtime environment of the third-level large model (e.g., Qwen-VL), including model weights and GPU memory allocation, to reduce subsequent call latency.

[0090] If the mutated Prompt is not empty, it is appended to the tail as a negative hint of the third-level input Prompt.

[0091] S3.5: Match the output error type keywords with each record in the historical poor Prompt library.

[0092] Historically poorly performing Prompt records are stored in the cloud or edge shared storage, and each record contains at least three fields: error type keywords, Prompt, etc. text (The complete input prompt that caused the error), timestamp. Matching rules use exact string matching (ignoring case) or WordNet-based synonym expansion matching. If a match is successful, all records with the same error type keyword are filtered from the database, and the prompt for each record is calculated. tex The semantic similarity between the current Prompt (i.e., the complete input Prompt constructed in S3.1) and the current Prompt. The similarity calculation method uses the cosine distance after Sentence-BERT embedding. The inferior Prompt with the highest semantic similarity is selected as the opposing sample.

[0093] Furthermore, after obtaining the opposing samples, a text mutation operation is performed, with the following steps: (1) In the Prompt of the opposing samples tex In the process, the complete phrase containing the error type keyword is located (for example, if the error type is "missed detection of a small target," the located phrase might be "failed to detect a tiny target" or "missed detection of a small target"). The location method uses regular expressions or dependency parsing to extract the smallest noun phrase or verb-object phrase containing the keyword.

[0094] (2) Find the corresponding phrase in the current Prompt: Search the current Prompt for a segment that is semantically similar to the phrase in step (1). If found, extract the core adjective or verb of the segment (e.g., small → large, clear → blurry). And replace the core word with its antonym according to the predefined antonym dictionary. If there is no suitable antonym, insert the negative word "non" before the phrase (e.g., edge segmentation error → non-edge segmentation error).

[0095] If the corresponding phrase is not found, append the following sentence to the end of the current Prompt: Note: Avoid errors such as '{error type keyword}'.

[0096] The mutated string serves as the mutated Prompt. This mutated Prompt can be used to replace the original input of the second-level model (i.e., to re-infer using the mutated Prompt), or it can be input as an additional negative prompt along with the original Prompt. This invention defaults to the latter (negative prompt method), that is, keeping the original Prompt unchanged and appending the mutated Prompt to the end of the input in the format of {mutated Prompt} (avoiding this format).

[0097] If no historical poor-quality Prompt record is matched, or if a match is successful but the highest semantic similarity value is below a preset threshold (e.g., 0.5), no text mutation is performed, and the mutated Prompt field outputs an empty string. All mutation operations are logged, including the original error type, the opposing sample ID, the replaced word pair, and the mutated Prompt fragment. This log is used for subsequent analysis of the mutation effect and can be fed back to the maintenance module of the historical poor-quality Prompt library.

[0098] S3.6: Encapsulate the second-level detection conclusion from S3.2, the second-level unconfirmed clues from S3.3, the variant Prompt generated from S3.5 (if any), and the metadata of the current task into a JSON object. If it is a nascent stage path, this object is written to the buffer via an edge message queue for use by the third stage in S4; if it is a stable stage path, this object is written directly to the result database, and an asynchronous task is triggered: the complete input Prompt of this second-level detection (constructed in S3.1) and the output results (including error type keywords) are appended to the historical poor-quality Prompt database, but only when the conditions are met—that is, the confidence level of the second-level detection conclusion is less than 0.7, or the second-level unconfirmed reason contains keywords such as vague or uncertain.

[0099] The asynchronous update module processes low-quality Prompt records awaiting database entry in batches at regular intervals (e.g., every hour), deduplicating them before writing them to the cloud database. Before entry, standardization is required: the specific area values ​​in the Prompt are normalized (e.g., replaced with {area} placeholders) to improve versatility. Simultaneously, the system maintains an online antonym dictionary update mechanism: if a text mutation uses negative word insertion instead of antonym replacement, and subsequent third-level detection confirms the mutation's effectiveness (i.e., correct detection after mutation), a manual review can be prompted to add the new antonym pair to the dictionary.

[0100] It can be seen that the asynchronous update mechanism avoids impacting the performance of the main detection process, while improving the reusability of inferior Prompt libraries through placeholder normalization. S4: If the call path includes a third level, the unconfirmed clues output by the first two levels are used as residual prompts and input into the third-level large model to obtain the final detection conclusion; and the conclusion of this detection is added to the historical detection record, while the erroneous cases in this detection are stored in the historical poor quality Prompt library.

[0101] S4.1: When the call path contains a third-level large model (i.e., a path in the nascent or degenerate phase), the system first performs branch processing based on the path identifier.

[0102] If it is a newborn path, then read the first-level unconfirmed clue generated by S2 and the second-level unconfirmed clue generated by S3 (S3.3) from the cache. Concatenate the two into a residual hint in the following format: First-level residual: {first-level unconfirmed clue}; Second-level residual: {second-level unconfirmed clue}.

[0103] The residual hint is attached to the end of the Level 3 large model input Prompt. For degradation paths, since Levels 1 and 2 are skipped, the system extracts the most recently stored Level 1 unconfirmed clue (written by S2.4) from the history of the detected object. If the clue exists, it is used as the unique residual, in the format: Historical Level 1 Residual: {Clues}; if no historical clue exists, an empty string is used. Degradation paths do not contain Level 2 unconfirmed clues.

[0104] While constructing residual prompts, the system also needs to truncate the residual content. The total length of the input prompt for the third-level large model (such as the Qwen-VL series, with ≥7B parameters) is limited (usually 2048 or 4096 tokens). If the concatenated residual prompt exceeds 512 tokens, truncation is performed based on the principle of prioritizing the preservation of key information: the complete content of the first-level residuals is retained, and the second-level residuals are truncated to within 256 tokens at the sentence level. The truncation operation is based on natural language separators such as periods and semicolons to avoid cutting off complete semantics. Furthermore, if unconfirmed historical first-level cues are used in the degradation path, a timestamp of when the cue was generated must be appended to the end of the cue to help the large model understand the timeliness of the information.

[0105] S4.2: Input the current detected image (original resolution, no scaling required) and the residual hint constructed in S4.1 into the third-level multimodal large model. The third-level large model uses the Qwen-VL series (such as Qwen2-VL-7B-Instruct) or an equivalent vision-language model, deployed on a cloud GPU server, supporting dynamic resolution input. The model's input is constructed as follows: system instructions, the original image (encoded as a visual token), and the residual hint (attached as a text token). The model performs forward inference and outputs the final detection conclusion in structured JSON format, containing at least the following fields: defect type, defect area, floating-point number, and unit mm. 2 Confidence level, 0-1 floating-point number, and reasoning basis.

[0106] S4.3: After obtaining the final detection conclusion at level three, the system appends key information about the detection to the historical record table of the corresponding detection object. The appended fields include: detection timestamp (ISO 8601 format), final defect area (the area value output at level three), unconfirmed clues at level one (taken from S2.4; empty if level one was not executed), unconfirmed clues at level two (taken from S3.3; empty if level two was not executed), and final confidence level.

[0107] The historical record table uses the object identification identifier as the primary key, stores records in reverse chronological order, retains the most recent 20 records, and automatically archives records beyond that to cold storage.

[0108] Meanwhile, if the current detection path is in the degradation period (i.e., skipping the first and second levels), no new unconfirmed first or second level clues will be generated. In this case, the unconfirmed first level clue field in the historical record will copy the corresponding value from the previous detection (if any) to avoid empty fields affecting future trend analysis. All historical record write operations are guaranteed to be consistent through distributed transactions, and a confirmation flag is returned after a successful write.

[0109] Finally, when the confidence level in the final detection conclusion of the third level is lower than the preset confidence threshold (default 0.7, configurable), the process of adding substandard cases to the database is triggered.

[0110] Furthermore, if text mutation occurs in S3.5 during this detection process (i.e., a mutated Prompt is generated), the original Prompt before mutation (i.e., the unmutated input Prompt constructed in S3.1) and its corresponding error type keywords (taken from the error type keywords output in S3.2) must also be stored as a separate record in the historical poor-quality Prompt database. The database entry operation uses asynchronous batch writing to avoid blocking the main detection process. Before database entry, images undergo desensitization processing (e.g., blurring faces, license plates, and other non-device areas) to comply with data security regulations.

[0111] After updating historical records and adding substandard data to the database, the system encapsulates the final detection conclusion of the third level into the final output result. This result is simultaneously written to two targets: a result database (for querying and display by business systems) and a message queue (for triggering subsequent alarms or work order processes). After successful writing, the cloud inference resources occupied by the third-level large model are released (session closed, video memory cache released).

[0112] The system also includes one or more processors and memory.

[0113] The memory is used to store operable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, including a flow of a large-scale model cooperative detection system based on a three-level cascade and Prompt autonomous evolution as described in the foregoing embodiments, particularly... Figure 1 The flowchart of the method is shown.

[0114] Other aspects disclosed in the embodiments of the present invention also propose a computer-readable medium for storing software including instructions executable by one or more computers, which, upon execution, cause the one or more computers to perform operations including a flow of a large-scale model collaborative detection system based on a three-level cascade and Prompt autonomous evolution as described in the foregoing embodiments, particularly... Figure 1 The flowchart of the method is shown.

[0115] It should be recognized that embodiments of the present invention may be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable storage medium.

[0116] The method can be implemented using standard programming techniques, including a non-transitory computer-readable storage medium configured with a computer program in the computer program, wherein the storage medium is configured such that the computer operates in a specific and predefined manner.

[0117] Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system; however, if required, the program can be implemented in assembly or machine language.

[0118] In any case, the language can be either compiled or interpreted.

[0119] Furthermore, for this purpose, the program can run on programmed application-specific integrated circuits.

[0120] The processes described herein (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that commonly executes on one or more processors. The computer program includes a plurality of instructions executable by one or more processors.

[0121] Furthermore, the method can be implemented in any suitable computing platform, including but not limited to personal computers, minicomputers, mainframes, workstations, networked or distributed computing environments, standalone or integrated computer platforms, or in communication with charged particle tools or other imaging devices.

[0122] Various aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether portable or integrated into a computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the processes described herein.

[0123] Furthermore, machine-readable code, or parts thereof, can be transmitted via wired or wireless networks.

[0124] When such media includes instructions or programs that combine with a microprocessor or other data processor to implement the steps described above, the invention described herein includes these and other different types of non-transitory computer-readable storage media.

[0125] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A large-scale model collaborative detection system based on three-level cascade and Prompt autonomous evolution, characterized in that: include: The cascade path selection module obtains the identification mark and historical detection records of the current detection object, determines the life stage label of the detection object based on the changing trend of continuous detection results, and determines the three-level cascade calling path based on the life stage label. The first-level small model initial detection module starts the first-level small model detection according to the calling path, and outputs the first-level detection conclusion and the first-level unconfirmed clues; The secondary mutation detection module, if the call path includes the second level, concatenates the first-level detection conclusion and the first-level unconfirmed clue into a second-level pre-prompt, inputs it into the model in the second level, and outputs the second-level detection conclusion and the second-level unconfirmed clue; at the same time, when the second-level detection conclusion matches the error type in the historical poor Prompt library, it selects the opposite sample from the historical poor Prompt library to perform text mutation on the current Prompt; The three-level final judgment storage module, if the call path includes the third level, will input the unconfirmed clues output by the first two levels as residual prompts into the third-level large model to obtain the final detection conclusion; and will append the current detection conclusion to the historical detection record, while storing the erroneous cases in the current detection into the historical poor quality Prompt library.

2. A method for collaborative detection of large and small models based on three-level cascade and Prompt autonomous evolution, based on the collaborative detection system for large and small models based on three-level cascade and Prompt autonomous evolution as described in claim 1, characterized in that: Also includes: Obtain the identification identifier and historical detection records of the current detection object, determine the life stage label of the detection object based on the changing trend of continuous detection results, and determine the three-level cascaded calling path based on the life stage label; The first-level small model detection is initiated according to the aforementioned call path, and the first-level detection conclusion and the first-level unconfirmed clues are output. If the call path includes a second level, the first level detection conclusion and the first level unconfirmed clue are concatenated to form a second level pre-prompt. After inputting into the model in the second level, the second level detection conclusion and the second level unconfirmed clue are output. At the same time, when the second level detection conclusion matches the error type in the historical poor Prompt library, an opposing sample is selected from the historical poor Prompt library to perform text mutation on the current Prompt. If the call path includes a third level, the unconfirmed clues output by the first two levels are used as residual prompts and input into the third-level large model to obtain the final detection conclusion. The results are then appended to the historical detection record based on the current detection conclusion, and the erroneous cases in this detection are stored in the historical poor-quality Prompt library.

3. The method for collaborative detection of large and small models based on three-level cascade and Prompt self-evolution as described in claim 2, characterized in that: The life stage tags for determining the detection object include: Read the n most recent detection results corresponding to the identification mark, each detection result includes the defect type and defect area; if the number of detections is less than n, mark the life stage label as the nascent stage; Compare the defect areas in two adjacent inspection results and calculate the direction of area change: If the defect area increases strictly in successive n consecutive tests, it is determined to be in the degradation period; If the defect area decreases strictly in sequence, it is considered to be in a stable period; If the direction of change is inconsistent, the areas are equal, or the number of tests is less than n, it is marked as the nascent stage.

4. The method for collaborative detection of large and small models based on three-level cascade and Prompt self-evolution as described in claim 3, characterized in that: The process of determining the three-level cascading call path based on the lifecycle stage tags includes: The neonatal period invokes a complete three-level path: Level 1 → Level 2 → Level 3. During the stable period, the first level is called → the second level, and the second level outputs the final detection result, skipping the third level; During the degradation phase, the third level is invoked directly, skipping the first and second levels; The third-level input residual prompt extracts the most recent unconfirmed clue from the history; if none is found, an empty prompt is used.

5. The method for collaborative detection of large and small models based on three-level cascade and Prompt self-evolution as described in claim 4, characterized in that: The initiation of the first-level small model detection includes: The current detected image is scaled to the input size of the first-level small model, and after forward inference, the category label, defect area estimate and corresponding feature response heatmap of each candidate target are obtained; The mean contrast is defined as the arithmetic mean of the response values ​​of all pixels in the heatmap. Extract continuous pixel regions with response values ​​lower than the average contrast value, and record the position coordinates of the continuous pixel regions as blurred regions; The blurred region is mapped back to the original detection image, and the corresponding image block is extracted. Based on the statistical characteristics of the image patches, a natural language description is generated as a first-level unconfirmed clue; The category label and the estimated defect area are used as the first-level detection conclusion, and the unconfirmed clue statements are used as the first-level unconfirmed clues. They are output together to the buffer for the next level to call.

6. The method for collaborative detection of large and small models based on three-level cascade and Prompt self-evolution as described in claim 5, characterized in that: The second-level detection conclusions and second-level unconfirmed clues output after inputting into the second-level model include: In the second level, the model outputs the following three fields: the second-level detection conclusion, the reason for non-confirmation, and the error type keywords; The second-level detection results include defect type, defect area, and confidence level; The reasons for non-confirmation are in natural language text, explaining any ambiguity or uncertainty at this level; The error type keywords are strings from a predefined set, which are generated by the model based on its own output features. The first-level unconfirmed clues are combined with the unconfirmed reasons output by the second-level model according to the specified format to form the second-level unconfirmed clues.

7. The method for collaborative detection of large and small models based on three-level cascade and Prompt self-evolution as described in claim 6, characterized in that: The step of selecting opposing samples to perform text mutation on the current Prompt includes: The error type keywords are matched against the error type field of each record in the historical poor-quality Prompt database. If a match is successful, the worst Prompt with the highest semantic similarity to the current Prompt in terms of the error type keyword is selected from the library as the opposing sample. Extract the complete phrase containing the error type keywords from the opposing samples; Find phrases containing the same keywords or synonyms in the current Prompt; If found, replace the core adjective or verb in the corresponding phrase with the corresponding word in the predefined antonym dictionary; If no suitable antonym is found, insert a negative word before the corresponding phrase; The mutated string, known as the mutated Prompt, is used to replace or supplement the original input of the second-level model.

8. The method for collaborative detection of large and small models based on three-level cascade and Prompt self-evolution as described in claim 7, characterized in that: The final detection conclusion includes the defect type, defect area, and confidence score. The defect area, detection time, and first-level unconfirmed clues of this detection will be added to the historical records of the corresponding detection object. When the confidence level of the final detection result is lower than the preset confidence threshold, the input image, the final detection result, and the error type field of this detection are assembled into a record and appended to the historical poor quality Prompt library.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the large-scale model collaborative detection system based on three-level cascade and Prompt autonomous evolution as described in any of claims 1.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the large-scale model collaborative detection system based on three-level cascade and Prompt autonomous evolution as described in any of claims 1.