Track defect image segmentation method based on visual knowledge base retrieval and verification
By employing a visual knowledge base-based retrieval and verification method, and utilizing dual logical verification of text instructions and visual evidence, pixel-level accurate segmentation of track defect images was achieved. This solves the problems of data scarcity and scene variability in track image segmentation, and improves the accuracy and robustness of segmentation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies for track image segmentation face challenges such as data scarcity, variable scenes, lighting variations, and significant background interference, resulting in insufficient accuracy and robustness in target segmentation. Furthermore, the general visual models lack domain adaptability, making it difficult to meet engineering accuracy requirements.
A visual knowledge base-based retrieval and verification method is adopted. By querying the image text command, relevant track defect image samples are retrieved from the visual knowledge base. Concept matching verification and location hypothesis testing are performed to generate spatial prompt information and drive the segmentation model to perform pixel-level accurate segmentation.
It improves the segmentation accuracy and robustness of track defect images, ensures that the selected visual evidence is highly consistent in semantics and space, and enhances the accuracy of segmentation and the quality of decision-making.
Smart Images

Figure CN121746715A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image segmentation, in particular to a track defect image segmentation method based on visual knowledge base retrieval and verification. BACKGROUND
[0002] The intelligent inspection of rail transportation systems relies on accurate and automated visual analysis of collected track images. Among them, instance segmentation of key components in track images, such as sleepers, fasteners, etc., or abnormal targets such as foreign matter, cracks, etc., is a basic task to achieve component state evaluation, disease identification and positioning.
[0003] However, track image segmentation based on deep learning faces several challenges in practical application. First, the track scene has high professionalism and structural characteristics. Target object classes are specific, and their appearance, size and spatial layout in the image are constrained by track physical structure, shooting angle and installation specifications. However, light changes, weather conditions and complex background interference result in large differences in the appearance characteristics of the same components. Second, high-quality labeled data is scarce. Obtaining pixel-level accurate segmentation labels requires deep domain knowledge and is costly, making it difficult to train fully supervised models for a large number of rare defects or new components, and limiting the generalization ability of the model. Third, the domain adaptability of general visual models is insufficient. When directly applied to track images, semantic confusion is likely to occur, making it difficult to meet the engineering precision requirements. Therefore, in order to improve the accuracy and robustness of target segmentation in track images with data scarcity and variable scenes, a new track image segmentation method is urgently needed. SUMMARY
[0004] The purpose of the present application is to provide a track defect image segmentation method based on visual knowledge base retrieval and verification, which can improve the segmentation accuracy of track defect images through concept matching verification and position hypothesis verification of visual evidence. When applied in computer hardware devices, it can improve the running speed of the hardware devices.
[0005] To achieve the above purpose, the present application provides the following solutions: In a first aspect, the present application provides a track defect image segmentation method based on visual knowledge base retrieval and verification, comprising: Based on the query image and the text instruction describing the defect in the query image, at least one track defect image sample related to the query image is retrieved from the visual knowledge base as visual evidence; wherein the visual knowledge base is constructed based on a plurality of track defect image samples; each track defect image sample includes a track defect image sample path, a text label of a defect in the track defect image sample, and a bounding box of the defect in the track defect image sample; The visual evidence is subjected to concept matching verification and location hypothesis testing. The concept matching verification is used to determine, based on the text instruction and the visual evidence, whether the text label of the defect in the visual evidence and the text instruction describing the defect in the query image are semantically consistent, and generates a first Boolean judgment result. The location hypothesis testing is used to determine, based on the query image, the text instruction and the bounding box of the defect in the visual evidence, whether the defect described by the text instruction is contained in the bounding box of the query image, and generates a second Boolean judgment result, when the concept matching verification passes. Based on the first Boolean judgment result and the second Boolean judgment result, spatial prompt information is generated; based on the spatial prompt information, a segmentation model is used to segment the query image for defects.
[0006] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a method for segmenting track defect images based on visual knowledge base retrieval and verification. It utilizes a query image and text instructions describing defects in the query image to retrieve the most relevant track defect image samples as visual evidence from a pre-built visual knowledge base containing a large number of track defect image samples. Furthermore, the most relevant visual evidence is subjected to dual logical verification of conceptual relevance and location rationality. Based on the verification results, a high-confidence spatial cue is generated, and this spatial cue is used to drive a segmentation model to achieve pixel-level accurate segmentation of defects in the query image. In addition, this application selects visual evidence based on feature cosine similarity and size similarity, which is equivalent to synergistically fusing semantic similarity and area similarity for comprehensive judgment. This ensures that the selected visual evidence is not only semantically consistent with the target scene in text description but also highly consistent in visual spatial proportion. This dual constraint mechanism effectively overcomes the one-sidedness of single-dimensional matching, making the final determined visual evidence more accurate and improving the quality of visual evidence-based decision-making. Attached Figure Description
[0007] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 A flowchart illustrating a method for segmenting track defect images based on visual knowledge base retrieval and verification, provided in an embodiment of this application; Figure 2A schematic diagram of the framework of a track defect image segmentation method based on visual knowledge base retrieval and verification provided in an embodiment of this application; Figure 3 A schematic diagram of the segmentation result of a track defect image segmentation method based on visual knowledge base retrieval and verification provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0009] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0010] To make the objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0011] In one exemplary embodiment, such as Figure 1 As shown, a method for segmenting track defect images based on visual knowledge base retrieval and verification is provided. This method includes steps 101 to 104. Wherein: Step 101: Based on the query image and the text instructions describing the defects in the query image, retrieve at least one track defect image sample related to the query image from the visual knowledge base as visual evidence; wherein, the visual knowledge base is constructed based on multiple track defect image samples; each track defect image sample includes a track defect image sample path, a text label of the defect in the track defect image sample, and a bounding box of the defect in the track defect image sample.
[0012] Step 101: Construct a visual knowledge base and perform feature retrieval to obtain visual evidence based on the query image and text instructions describing defects in the query image. The visual knowledge base is an external, non-parametric knowledge base constructed by integrating large-scale publicly available orbital image datasets. Each entry in the visual knowledge base... These are all structured visual examples, namely, track defect image samples, and their data structures... Includes: Storage path for orbital defect image samples Text labels of defects in orbital defect image samples And the bounding box calculated from the corresponding real segmentation mask to pinpoint the precise location of the segmentation target. .
[0013] The specific construction method of the visual knowledge base is as follows: First, define a list of metadata configurations for multiple datasets. For each track defect image sample, extract image features using a pre-trained visual language model and perform normalization processing to establish a feature index. The bounding box is determined by calculating the minimum bounding rectangle of the target region and can be expanded by a proportional coefficient to avoid region loss. In the retrieval stage, the query image and text command are encoded into a normalized query vector, and its similarity score with the feature vectors of samples in the visual knowledge base is calculated. Multiple most relevant track defect image samples are retrieved as a candidate visual evidence set, and the highest-ranked sample is selected as the best visual evidence.
[0014] Step 102: Perform concept matching verification and location hypothesis testing on the visual evidence; the concept matching verification is used to determine, based on the text instruction and the visual evidence, whether the text label of the defect in the visual evidence and the text instruction describing the defect in the query image are semantically consistent, and generate a first Boolean judgment result; the location hypothesis testing is used to determine, based on the query image, the text instruction and the bounding box of the defect in the visual evidence, whether the defect described by the text instruction is contained in the bounding box of the query image, and generate a second Boolean judgment result, when the concept matching verification passes.
[0015] Step 103: Generate spatial prompt information based on the first Boolean judgment result and the second Boolean judgment result.
[0016] Step 104: Based on the spatial hint information, perform defect segmentation on the query image using a segmentation model. Based on the spatial hint information generated in step 103, perform pixel-level defect segmentation on the query image using a segmentation model, and output the final defect mask.
[0017] By performing steps 101 to 104 above, this application can achieve pixel-level precise segmentation of defects in the query image. A schematic diagram of the overall framework is shown below. Figure 2 As shown, the segmentation result is as follows Figure 3 As shown. Step 101 above is replaced by steps 201 to 204: Step 201: Use a text encoder to extract features from the text instructions to obtain text feature vectors, and then normalize the text feature vectors.
[0018] Step 202: Use an image encoder to extract features from the query image to obtain an image feature vector, and then normalize the image feature vector.
[0019] Step 203: Perform feature fusion on the normalized text feature vector and the normalized image feature vector to obtain the query vector.
[0020] Step 204: Based on the query vector, retrieve at least one track defect image sample related to the query image from the visual knowledge base as visual evidence.
[0021] Specifically, text instructions Input a Transformer encoder to obtain text feature vectors; query images. The VisionTransformer encoder is input to obtain the image feature vector. The text feature vector and the image feature vector are fused using the Q-Former fusion method of the Bootstrapping Language-Image Pre-training (BLIP) model to form a unified, normalized query vector. This query vector simultaneously encodes the visual content of the query image and the semantic intent of the text instruction.
[0022] In another exemplary embodiment of this application, step 204 is replaced by the following steps 301 to 303: Step 301: Use the BLIP model to extract features from each track defect image sample to obtain the image semantic feature vector of each track defect image sample.
[0023] Specifically, to achieve rapid semantic retrieval, this application performs a one-time feature space indexing on the visual knowledge base. A pre-trained visual language model, BLIP, is selected, utilizing its feature encoder. The visual knowledge base... Each track defect image sample is processed by a visual preprocessor and then input into the BLIP model to extract high-dimensional features. The output feature vector is then normalized. The normalized image semantic feature vectors have geometric consistency, thus ensuring that the inner product of any two vectors is mathematically equivalent to their cosine similarity, significantly improving the computational efficiency of subsequent retrieval processes.
[0024] Step 302: Calculate the similarity between the query vector and the image semantic feature vector of each track defect image sample to obtain the feature cosine similarity of each track defect image sample.
[0025] Step 303: Based on the feature cosine similarity of each track defect image sample, retrieve at least one track defect image sample related to the query image from the visual knowledge base as visual evidence.
[0026] In another exemplary embodiment of this application, step 303 is replaced by the following steps 401 to 403: Step 401: Determine the size similarity of each track defect image sample based on the image area of the query image and the image area of each track defect image sample.
[0027] Step 402: Based on the feature cosine similarity and the size similarity of each track defect image sample, obtain the comprehensive evaluation value of each track defect image sample.
[0028] Step 403: Sort the comprehensive evaluation values of each track defect image sample from largest to smallest, and select the top_k best-matching track defect image samples from the visual knowledge base as visual evidence. The top-k best-matching track defect image samples are returned using the following formula: .
[0029] Where Sim() is the similarity function. To query images, For track defect image samples, the function This refers to a feature extraction function or an image encoder.
[0030] This retrieval process can be expressed by the following formula: .
[0031] in, From visual knowledge base Based on The retrieved set of visual evidence. The highest-ranked visual evidence is selected from this set. As the main basis for subsequent reasoning For text instructions, To query images.
[0032] The comprehensive evaluation value of the i-th track defect image sample is determined using the following formula: .
[0033] .
[0034] .
[0035] in, Let be the comprehensive evaluation value of the i-th track defect image sample. Let cosine similarity be the feature of the i-th track defect image sample. Let i be the size similarity of the i-th track defect image sample. For scale weight parameters, For query vector, Let i be the image semantic feature vector of the i-th track defect image sample. To query the area of an image, Let be the image area of the i-th track defect image sample.
[0036] In another exemplary embodiment of this application, step 102 involves executing a hierarchical reasoning strategy based on the comprehensive evaluation value of the retrieved visual evidence to verify its validity. Specifically, this includes determining the reasoning level based on the comprehensive evaluation value: if the comprehensive evaluation value is higher than a first threshold, a direct mapping strategy is executed, directly adopting the bounding box information of the visual evidence; if the comprehensive evaluation value is lower than a second threshold, the visual evidence is deemed invalid, and zero-sample reasoning is directly initiated; if the comprehensive evaluation value is between the first and second thresholds, the process enters the cautious reasoning engine, performing concept matching verification and location hypothesis testing.
[0037] The concept matching check uses a visual language model to determine whether the defect in the best visual evidence is consistent with the target described by the text instruction in terms of semantic concept or visual form, and whether it can provide effective inspiration for locating the defect in the query image, generating a first Boolean judgment result. The location hypothesis test, assuming the concept matching check passes, determines whether the target defect is contained within the bounding box of the best visual evidence if it is directly mapped to the same coordinate position in the query image, generating a second Boolean judgment result.
[0038] Specifically, in step 102, the concept matching verification includes the following steps 501 to 503: Step 501: Based on the text instructions and the visual evidence, obtain the first multimodal prompt information.
[0039] Step 502: Input the first multimodal prompt information into the visual language model, and drive the visual language model to analyze whether the defect type in the visual evidence is semantically consistent with the defect type described by the text instruction.
[0040] Step 503: If yes, the first Boolean judgment result generated is true; if no, the first Boolean judgment result generated is false.
[0041] In step 102, the location hypothesis test specifically includes the following steps 601 to 603: Step 601: When the concept matching verification passes, extract the position coordinates of the bounding box of the defect in the visual evidence and project them onto the query image to form an overlay image with position hypothesis boxes.
[0042] Step 602: Construct a second multimodal prompt message based on the overlaid image.
[0043] Step 603: Input the second multimodal prompt information into the visual language model, and drive the visual language model to determine whether the location hypothesis box contains a defect target that matches the description of the text instruction; if yes, the generated second Boolean judgment result is true; if no, the generated second Boolean judgment result is false.
[0044] Step 102 is the core innovation of this application, which lies in the process of using visual evidence. This process is executed by an inference engine that uses a visual language model as its inference engine. Step 103 designed a prompting strategy based on thought chains to guide the inference engine. Visual evidence Perform a two-step logic check.
[0045] Step 1: Concept matching verification. Inference engine First, assess visual evidence. It determines whether the semantic concept matches the current task. It judges visual evidence. Does the information contained in the query image help? The positioning is determined by the text instruction. Describe the defects in the query image and give the first Boolean judgment result. .in, If true, It is false.
[0046] Step 2: Location Hypothesis Testing. This occurs after successful concept matching verification. In this case, the inference engine continues with location hypothesis testing. It treats the bounding box in the visual evidence as a location hypothesis, tests the spatial plausibility of this hypothesis in the query image, and gives a second Boolean judgment result. .
[0047] In another exemplary embodiment of this application, step 103, generating spatial prompt information based on the first Boolean judgment result and the second Boolean judgment result, specifically includes steps 701 to 702: Step 701: If both the first Boolean judgment result and the second Boolean judgment result are true, then generate spatial prompt information based on the position coordinates of the bounding box in the visual evidence.
[0048] Step 702: If the first Boolean judgment result is true and the second Boolean judgment result is false, then the text label of the defect and the position coordinates of the bounding box of the defect in the visual evidence are used as visual cues to drive the visual language model to perform a new, targeted target search in the query image to generate spatial cues. Using the visual features in the visual evidence as "inspiration," the most similar object is searched again in the input image to generate a new bounding box.
[0049] Step 703: If the first Boolean judgment result is false, then discard the visual evidence, and generate spatial prompt information based on the query image and the text instruction describing the defects in the query image, according to the zero-shot capability of the visual language model.
[0050] Specifically, step 103 is based on the verification result in step 102, that is: based on and The logical combination results are used to select a strategy from a preset set of strategies to generate the final spatial prompt information. .
[0051] Strategy 1: Directly adopt the reference strategy. When When true, that is and If both conditions are true, the bounding box in the visual evidence will be used as spatial cue information. A value of true indicates that the retrieved visual evidence semantically matches the query image. The bounding box representing visual evidence is spatially plausible and credible on the query image.
[0052] Strategy Two: Guided Search Positioning Strategy. When When true, that is If true, False. This indicates that the retrieved visual evidence semantically matches the query image, but the bounding box provided by the visual evidence is spatially unreliable on the query image. The semantic information from the visual evidence will be used as strong guidance to drive the inference engine. In the image query Then, a new, targeted target search is performed to generate a new bounding box.
[0053] Strategy 3: Zero-shot inference strategy. When When true, it is A false result indicates that the visual evidence is irrelevant to the current task. This visual evidence will be ignored, and the external visual evidence will be discarded, relying directly on the inference engine. Its own zero-shot reasoning capability means that it can directly predict bounding boxes based solely on query images and text commands, utilizing the zero-shot capability of visual language models.
[0054] The goal of the entire reasoning and decision-making process is to generate a high-confidence spatial cue. This can be summarized by the following formula: This formula represents spatial hints. It is a reasoner based on and visual evidence generate.
[0055] In another exemplary embodiment of this application, step 104 is followed by steps 801 to 802: Step 801: If the spatial prompt information is valid, the spatial prompt information is input into the prompt encoder of the segmentation model to guide the segmentation model to perform segmentation within the area specified by the spatial prompt information, so as to generate a target segmentation mask.
[0056] Step 802: If the spatial hint information is invalid, the segmentation model will autonomously generate a target segmentation mask in the query image.
[0057] Specifically, spatial cue information generated by the hierarchical decision engine is input as a precise guiding signal to the cue encoder. In the SAM2 segmentation model, if a spatial cue box is received from an external detection model, the cue-guided segmentation strategy is activated; otherwise, a multi-mask generation and intersection-over-union (IoU) filtering mechanism is automatically executed to select the optimal segmentation result. The SAM2 segmentation model accurately understands the segmentation intent based on the spatial cue information, performs pixel-level segmentation within the specified region of the query image, and finally outputs the target segmentation mask through the mask decoder, completing the entire segmentation process.
[0058] In another exemplary embodiment of this application, comparative and ablation experiments are conducted to illustrate the technical effects of this application, including sensitivity (SE), specificity (SP), F1 score, accuracy (AC), and Dice coefficient (DC). The experimental results are shown in Tables 1 and 2. It can be seen that this application has superior performance and metrics compared to small and large models.
[0059] Table 1. Results of comparative experiments on the small model
[0060] Table 2 Comparative experimental results on the large model
[0061] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows.Figure 4 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores track defect image samples. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a track defect image segmentation method based on visual knowledge base retrieval and verification.
[0062] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0063] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A track defect image segmentation method based on visual knowledge base retrieval and verification, characterized in that, The method comprises: retrieving at least one track defect image sample related to the query image as visual evidence in a visual knowledge base based on the query image and the text instruction describing the defect in the query image; wherein the visual knowledge base is constructed based on a plurality of track defect image samples; each track defect image sample comprises a track defect image sample path, a text label of a defect in the track defect image sample, and a bounding box of the defect in the track defect image sample; performing concept matching verification and position hypothesis verification on the visual evidence; the concept matching verification is used to determine whether the text label of the defect in the visual evidence and the text instruction describing the defect in the query image are consistent in semantics according to the text instruction and the visual evidence, and generate a first Boolean judgment result by using a visual language model; the position hypothesis verification is used to determine whether the defect described by the text instruction is contained in the bounding box on the query image when the concept matching verification passes, and generate a second Boolean judgment result by using a visual language model according to the query image, the text instruction and the bounding box of the defect in the visual evidence; generating spatial prompt information according to the first Boolean judgment result and the second Boolean judgment result; and performing defect segmentation on the query image by using a segmentation model based on the spatial prompt information.
2. The method of claim 1, wherein the method is based on a visual knowledge bank search and verification of track defect image segmentation. retrieving at least one track defect image sample related to the query image as visual evidence in a visual knowledge base based on the query image and the text instruction describing the defect in the query image, specifically comprising: extracting features of the text instruction by using a text encoder to obtain a text feature vector, and performing normalization processing on the text feature vector; extracting features of the query image by using an image encoder to obtain an image feature vector, and performing normalization processing on the image feature vector; performing feature fusion on the normalized text feature vector and the normalized image feature vector to obtain a query vector; retrieving at least one track defect image sample related to the query image as visual evidence in a visual knowledge base based on the query vector.
3. The method of claim 2, wherein the method further comprises: retrieving at least one track defect image sample related to the query image as visual evidence in a visual knowledge base based on the query vector, specifically comprising: extracting features of each track defect image sample by using a BLIP model to obtain an image semantic feature vector of each track defect image sample; calculating the similarity between the query vector and the image semantic feature vector of each track defect image sample to obtain a feature cosine similarity of each track defect image sample; retrieving at least one track defect image sample related to the query image as visual evidence in a visual knowledge base based on the feature cosine similarity of each track defect image sample.
4. The method of claim 3, wherein the method further comprises: retrieving at least one track defect image sample related to the query image as visual evidence in a visual knowledge base based on the feature cosine similarity of each track defect image sample, specifically comprising: determine a size similarity of each track defect image sample based on an image area of the query image and an image area of each track defect image sample; obtain a comprehensive evaluation value of each track defect image sample based on the feature cosine similarity of each track defect image sample and the size similarity of each track defect image sample; sort the comprehensive evaluation values of each track defect image sample from large to small, and select the top_k most matched track defect image samples in the visual knowledge base as the visual evidence.
5. The method of claim 4, wherein the method further comprises: The comprehensive evaluation value of the i-th track defect image sample is determined by the following formula: ; ; ; wherein, is a comprehensive evaluation value of the i-th track defect image sample, is a characteristic cosine similarity of the i-th track defect image sample, is a size similarity of the i-th track defect image sample, is a scale weight parameter, is a query vector, is an image semantic feature vector of the i-th track defect image sample, is an image area of the query image, is an image area of the i-th track defect image sample.
6. The method of claim 4, wherein the method further comprises: perform concept matching verification and position hypothesis verification on the visual evidence, and generate spatial prompt information according to the first Boolean judgment result and the second Boolean judgment result, specifically including: if the comprehensive evaluation value corresponding to the visual evidence is greater than the first threshold value, then directly generate spatial prompt information based on the visual evidence; if the comprehensive evaluation value corresponding to the visual evidence is less than the second threshold value, then determine that the visual evidence is invalid, and generate spatial prompt information according to the zero-shot capability of the visual language model; otherwise, perform concept matching verification and position hypothesis verification on the visual evidence, and generate spatial prompt information according to the first Boolean judgment result and the second Boolean judgment result.
7. The method for segmenting track defect images based on visual knowledge base retrieval and verification according to claim 1, characterized in that, The concept matching verification specifically includes: obtain first multi-modal prompt information based on the text instruction and the visual evidence; input the first multi-modal prompt information into the visual language model to drive the visual language model to analyze whether the defect type in the visual evidence is consistent with the defect type described in the text instruction in semantics; if yes, the generated first Boolean judgment result is true; if no, the generated first Boolean judgment result is false.
8. The method of claim 7, wherein the method further comprises: The position hypothesis verification specifically includes: when the concept matching verification passes, extract the position coordinates of the bounding box of the defect in the visual evidence, and project them into the query image to form an overlay image with a position hypothesis box; construct second multi-modal prompt information based on the overlay image; input the second multi-modal prompt information into the visual language model to drive the visual language model to determine whether the position hypothesis box contains the defect target described in the text instruction; if yes, the generated second Boolean judgment result is true; if no, the generated second Boolean judgment result is false.
9. The method of claim 8, wherein the method further comprises: Generate spatial prompt information according to the first Boolean judgment result and the second Boolean judgment result, specifically including: if the first Boolean judgment result and the second Boolean judgment result are both true, generate spatial prompt information according to the position coordinates of the bounding box in the visual evidence; if the first Boolean judgment result is true and the second Boolean judgment result is false, use the text label of the defect and the position coordinates of the bounding box of the defect in the visual evidence as visual prompts to drive the visual language model to perform a new and targeted target search in the query image to generate spatial prompt information; If the first Boolean judgment result is false, the visual evidence is abandoned, and based on the query image and the text instruction describing the defect in the query image, spatial prompt information is generated according to the zero-shot capability of the visual language model.