Topology mask based data center size model collaborative visual retrieval method and system

CN122676313APending Publication Date: 2026-09-01WHALE CLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611131089.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0012]针对现有技术中存在的上述问题,本发明的目的在于提供基于拓扑掩码的数据中心大小模型协同视觉检索法及系统,该系统一方面在内部引入开集视觉大模型实现图像多目标解耦,另一方面面向数据中心大小模型协同质检系统提供训练数据的自动收集与过滤能力,以将无差别的全局特征抑制转化为受特定空间与逻辑约束的局部特征隔离,从而解决复杂视觉检索中排斥范围失控与无关目标混淆的技术难题

Benefits of technology

1.本发明提供了基于拓扑掩码的数据中心大小模型协同视觉检索法及系统,摒弃了传统的全局惩罚路线,通过统一坐标化实体定位与感兴趣区域对齐,在底层物理空间上强制剥离了不同目标的语义边界,实现了多目标特征的独立解耦,克服了多目标对象在全局特征表示下语义边界相互混叠、无法独立区分保留特征与排除特征的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676313A_ABST
    Figure CN122676313A_ABST
Patent Text Reader

Abstract

This invention discloses a data center large-scale model collaborative visual retrieval method and system based on topological masks. The system targets intelligent operation and maintenance scenarios in data centers, introducing an open-set large visual model to achieve multi-target decoupling of images and collecting and filtering training data for a large-scale model collaborative quality inspection system. The system decouples retrieval commands with negation semantics into positive, negative, relational feature vectors and topological order constraints, decoupling candidate images into target nodes and constructing a relational topology graph. Starting from a positive anchor point, a restricted breadth-first search is performed, identifying nodes traversed and rejected by negation semantics as locally affected nodes and projecting them to generate a topological validity mask. After Softmax, this mask is used to gate the rejection responses element-by-element, ensuring that rejection responses outside the masked region are zeroed before aggregating candidate features and outputting them in descending order. This invention transforms global suppression into topologically constrained local isolation, improving the accuracy and robustness of local negation retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image retrieval and multi-object relationship modeling technology, and in particular to a collaborative visual retrieval method and system based on topological masking for data center-sized models. Background Technology

[0002] With the evolution of computer vision technology, visual retrieval applications have been widely applied to various scenarios, including complex scene image understanding and medical image-assisted analysis. In practical applications, users' retrieval needs are no longer limited to the positive "what to find (positive semantics)," but increasingly include the complex logic of "what to exclude (negative semantics)." For example, in a data center intelligent operation and maintenance scenario, searching for monitoring frames that show "abnormal equipment indicator lights, but exclude red safety icons"; in a communication construction scenario, searching for "construction workers, but not wearing safety helmets"; and in medical auxiliary diagnosis, searching for "lung nodules, but excluding calcified areas locally adjacent to the nodule."

[0003] In the field of intelligent data center operation and maintenance, the collaborative architecture of small and large models (lightweight small model for rapid coarse screening + visual large model for fine recognition) has become the mainstream quality inspection solution. Its small model iterative training relies on a large number of scenario-based labeled samples, and there is an urgent need for an efficient visual retrieval system to automatically mine training data from massive monitoring image libraries.

[0004] However, visual retrieval with negative semantics faces an objective technical challenge: in complex images containing multiple similar objects, exclusion constraints should essentially only apply to local objects truly related to the retrieval target, and should not affect unrelated similar objects. For example, when retrieving "a person not wearing a hat," only the safety helmet associated with that person needs to be excluded, while idle safety helmets in the background and people wearing helmets correctly at the edge of the image should be ignored. This "boundary of action" is neither given by the search text nor determined solely by the object's category label, but depends on the spatial and logical relationships between the targets. How to accurately define the boundary of the negative semantics within an image and avoid indiscriminate global exclusion is a technical bottleneck that current visual retrieval urgently needs to overcome.

[0005] To address the aforementioned challenges, related technologies have evolved from global feature matching based on convolutional neural networks (CNNs) or visual transformers (ViTs) to multimodal models such as region features, scene graphs, and CLIPs, attempting to suppress negative category scores through exclusion penalties. However, a search revealed that the following existing technologies, closest to this invention, have failed to resolve the aforementioned boundary-of-effect problem:

[0006] Firstly, methods based on global features and negative semantic enhancement (such as NegCLIP and CLIP-Negation) achieve retrieval by comparing the global features of the text with those of the entire image. However, when the image contains both the "positive part" and the "part that should be excluded" of the target, the global features are deeply aliased, resulting in a reduction in scores that excludes coarse-grained features without any difference.

[0007] Secondly, fine-grained matching techniques based on region features (such as RegionCLIP and GLIP) extract region-level features and match them independently with the text; however, they view each local feature in isolation and lack modeling of the spatial and logical (topological) relationships between targets, making it impossible to handle local logic such as "excluding another object that is associated with a specific object".

[0008] Third, scene graph matching methods (such as SG-Match) model target relationships using "target node + relationship edge"; however, they usually rely on a pre-defined closed relationship category library and are mainly used for isomorphic matching of positive semantics. They lack a mechanism to combine "topological path" with "low-level feature rejection" and cannot accurately isolate irrelevant targets.

[0009] Fourth, based on the exclusion techniques of cross-attention masks and negative prompts (such as the negative constraints of Stable Diffusion), soft feature subtraction is used across the entire image. It cannot impose hard constraints based on the spatial and logical boundaries of objects, and in retrieval scenarios, it is easy to over-suppress independent similar objects in the background or at the edges, resulting in a high rate of missed detections.

[0010] In summary, while existing methods approach the problem from different angles, such as global features, region features, scene graphs, and attention subtraction, none of them address the same fundamental issue: in complex multi-target images, there is a lack of a mechanism to precisely limit exclusion constraints to a "local range with a specific spatial and logical relationship to the retrieval target." This deficiency manifests itself in two main ways at the underlying level: First, the semantics of multiple targets overlap and cannot be decoupled independently under global features: when an image contains multiple targets, the information of each object is compressed into a single global feature vector, making it impossible to distinguish and isolate the local features that should be retained and those that should be excluded at the feature level.

[0011] Secondly, the scope of the exclusion operation is out of control, leading to feature degradation and a high rate of missed detections: the exclusion operation is mostly performed at a coarse-grained level in the global space, and it is impossible to constrain its range based on the local boundaries of the object; once there are independent objects of the same kind that are unrelated to the target in the image, they will be downgraded along with the target and the entire image will be mistakenly suppressed, causing the target that truly matches the retrieval intent to be filtered out incorrectly, thus reducing the retrieval accuracy. Summary of the Invention

[0012] To address the aforementioned problems in existing technologies, the present invention aims to provide a data center-size model collaborative visual retrieval method and system based on topological masks. This system, on the one hand, introduces an open-set visual large model internally to achieve multi-target decoupling in images; on the other hand, it provides automatic collection and filtering capabilities for training data for data center-size model collaborative quality inspection systems, so as to transform the suppression of indiscriminate global features into the isolation of local features subject to specific spatial and logical constraints, thereby solving the technical problems of uncontrolled exclusion range and confusion of irrelevant targets in complex visual retrieval.

[0013] To achieve the above objectives, the present invention adopts the following technical solution, including: S1. Decouple the retrieval command with local negation semantics by structural means to obtain positive semantic vector, negative semantic vector, relation feature vector and topological order constraint; S2. Extract the bottom dense feature blocks of the candidate image to obtain the bottom key matrix and the bottom value matrix, and decouple the candidate image into multiple target nodes carrying spatial bounding boxes and local region features; S3. Construct a relational topology graph with the target node as the vertex, establish connected edges based on the spatial geometric relationship between nodes, and assign semantic interaction weights to the connected edges based on the cross-modal matching results between the relational feature vector and the visual interaction features aggregated from the local regional features of the nodes at both ends of the connected edges. S4. Based on the matching result of the local region features and the positive semantic vector, establish the positive anchor node and take it as the starting point. Under the joint constraints of the topological order constraint and the semantic interaction weight, perform a restricted breadth-first search. Determine the neighboring nodes that have been traversed and have passed through the semantic rejection of the negative semantic vector as the set of locally affected nodes, and generate the topological validity mask matrix accordingly. S5. Using the topological validity mask matrix as a hard gate signal, the rejection response between the bottom dense feature block and the rejection semantic vector is gated element-by-element, forcing the rejection response outside the region indicated by the mask matrix to zero, and accordingly... The weighted aggregation of the underlying value matrix yields candidate feature vectors that complete local feature suppression; S6. Calculate the comprehensive score of the candidate image based at least on the candidate feature vector and the positive semantic vector, and output the retrieval results in descending order of the comprehensive score.

[0014] Further, in step S4, generating a topology validity mask matrix based on the set of locally affected nodes includes: projecting the spatial coordinates of each node in the set of locally affected nodes onto the feature grid where the bottom layer dense feature block is located, to obtain the topology validity mask matrix; Step S5 obtains the candidate feature vector after local feature suppression, including: using the positive semantic vector as the query term, the underlying key matrix as the key, and the underlying value matrix as the value, calculating the positive semantic attention distribution by scaling dot product attention; calculating the negative correlation distribution by using the negative semantic vector and the underlying key matrix, normalizing the negative correlation distribution and then multiplying it element-wise with the topological validity mask matrix to obtain a restricted rejection attention score; subtracting the restricted rejection attention score weighted by the rejection strength coefficient from the positive semantic attention distribution, performing non-negative truncation and renormalization on the subtraction result, and using the obtained attention distribution to weighted aggregate the underlying value matrix to obtain the candidate feature vector.

[0015] Furthermore, step S1 involves structural decoupling of the retrieval command with local negation semantics, including: The search command is input into the natural language processing engine, and the negative trigger words and their syntactic scope are identified through dependency parsing and entity relation extraction. The search command is deconstructed into a structured triple containing positive semantic phrases, negative semantic phrases and relational constraint phrases. The positive semantic phrase, the negative semantic phrase, and the relational constraint phrase are respectively input into the text encoder, and the positive semantic vector, the negative semantic vector, and the relational feature vector are output accordingly. The topological order constraint is generated based on the syntactic dependency distance mapping between the negation trigger word and the positive semantic phrase, and when the retrieval instruction does not explicitly specify the distance boundary, the topological order constraint is assigned a preset constant.

[0016] Further, in step S2, the candidate image is decoupled into multiple target nodes, including: The open-set visual localization model is invoked, and the positive and negative semantic phrases decoupled in step S1 are reused as text prompt words to extract the semantic bounding boxes that match the text prompt words across modalities; at the same time, the category-independent instance segmentation model is invoked to obtain the pixel-level binary mask and its minimum bounding rectangle. The intersection-union ratio (IUR) is calculated for the semantic bounding box and the minimum bounding rectangle. When the IUR is greater than a preset threshold, the coordinates of the minimum bounding rectangle are retained and the semantic confidence of the semantic bounding box is inherited, and the two are merged into a single target entity. When the two have spatial intersection and the IUR is not greater than the preset threshold, the two are retained as independent target entities. Using the spatial bounding boxes of each target entity as indices, dense feature blocks falling within the spatial bounding boxes are extracted from the underlying key matrix using region of interest alignment and aggregated into the local region features. The spatial bounding boxes and the local region features are then encapsulated into corresponding target nodes.

[0017] Further, step S3 involves establishing connected edges and assigning semantic interaction weights to them, including: For any two target nodes, calculate the distance between the center points of their spatial bounding boxes and the intersection-union ratio (IU). If the distance between the center points is less than a preset distance threshold or the IU is greater than a preset IU threshold, establish a connecting edge between them; otherwise, determine that there is no connecting edge between them and prune and isolate them. For each connected edge, a minimum bounding rectangle enclosing the target nodes at both ends is generated as a joint region. Dense feature blocks falling into the joint region are extracted from the underlying key matrix and aggregated into the visual interaction feature. The similarity between the visual interaction feature and the relation feature vector is calculated and truncated non-negatively, and then used as the semantic interaction weight of the connected edge.

[0018] Furthermore, step S4 involves performing a restricted breadth-first search starting from the positive anchor node, including: All nodes in the relation topology graph whose local region features are more similar to the positive semantic vector than a preset positive recall threshold are identified as positive anchor nodes, in order to accommodate the situation where there are multiple similar positive targets in the same candidate image. The restricted breadth-first search is performed in parallel with each of the positive anchor nodes as an independent root node. The restricted breadth-first search is subject to the following two constraints: the depth of the traversal path does not exceed the maximum number of hops limited by the topological order constraint; and expansion is only performed along connected edges whose semantic interaction weight is greater than a preset edge activation threshold.

[0019] Further, step S4 involves determining the set of locally affected nodes and generating a topology validity mask matrix, including: For each neighboring node traversed, calculate the similarity between its local region features and the negation semantic vector, and push neighboring nodes with similarity greater than a preset semantic rejection threshold into the set of locally affected nodes. When the set of locally affected nodes is empty, the topological validity mask matrix with all zeros is output so that the subsequent rejection response outputs a zero tensor and completely preserves the feature response of the underlying dense feature block through residual bypass. When the set of locally affected nodes is not empty, the spatial bounding box coordinates of each node in the set of locally affected nodes are mapped to the index interval of the feature grid according to the downsampling step size of the feature grid. After setting the elements in the feature grid that are within the index interval, they are flattened to obtain the topology validity mask matrix.

[0020] Further, step S6 calculates the comprehensive score of the candidate image, including: Calculate the similarity between the candidate feature vector and the positive semantic vector, and use it as the implicit feature similarity. The maximum similarity between the local region features of each of the positive anchor nodes and the positive semantic vector is calculated as the maximum similarity of the anchor ontology. Extract the maximum value of the semantic interaction weight of the activated connected edges on the restricted breadth-first search path starting from the positive anchor node, and use it as the topological consistency score. Set the topological consistency score to zero when the path is empty. A local rejection penalty term is calculated based on the cross-modal semantic rejection similarity between each node in the locally affected node set and the negative semantic vector, and the local rejection penalty term is set to zero when the locally affected node set is empty. The comprehensive score is obtained by weighting and summing the implicit feature similarity, the maximum similarity of the anchor point ontology, and the topological relationship consistency score, and then subtracting the local exclusion penalty term after nonlinear scaling.

[0021] Furthermore, it also includes training data collection and filtering step S7: S7. The Top-K search results output in step S6 are imported into the data collection and filtering module as candidate training samples. After multi-level filtering, a training sample set is formed and fed back to the lightweight small model in the large and small model collaborative quality inspection system for iterative training. The multi-level filtering includes: similarity threshold screening, perceptual hash deduplication, hard example mining, category ratio balancing and confidence screening. Furthermore, step S7 performs multi-level filtering, including: For each candidate image output in step S6, the comprehensive score is used as the confidence screening criterion, and images with a comprehensive score greater than the preset inclusion threshold are retained as initial inclusion samples. Calculate the perceptual hash value of each initially included sample. When the perceptual hash Hamming distance between any two samples is less than the preset deduplication threshold, only the sample with the higher comprehensive score is retained, and the remaining duplicate samples are removed. Calculate the cross-modal semantic rejection similarity between each target node in the retained samples and the negative semantic vector, and mark the samples whose semantic rejection similarity is in the preset difficult example interval as difficult example samples and improve their inclusion priority. The difficult example interval is the similarity range that is greater than the preset lower limit of difficult examples and less than the preset hard blocking threshold. The number of samples in each scene category in the current training sample set is counted. When the deviation between the sample ratio of any category and the target ratio exceeds the preset balance tolerance, samples of that category are preferentially added from the initial collected samples until the category ratio balance condition is met. The filtered samples are packaged into a unified training data format. Each sample contains image data, retrieval command text, positive semantic labels and negative semantic labels, and is written into the training sample queue for iterative training by small models.

[0022] Furthermore, the search command supports both image-text combined search mode and plain text search mode; in the image-text combined search mode, steps S2 and S3 are also executed for the query image input by the user to obtain the set of query target nodes and the corresponding query topology map, and the comprehensive score is calculated by combining the query topology map and the relationship topology map of the candidate image.

[0023] A collaborative visual retrieval system for data center size models based on topology masks, used to implement the aforementioned collaborative visual retrieval method for data center size models based on topology masks, is characterized by comprising: The input parsing module is used to structurally decouple retrieval commands with local negation semantics, and obtain positive semantic vectors, negative semantic vectors, relation feature vectors, and topological order constraints. The feature extraction module is used to extract the bottom dense feature blocks of the candidate image to obtain the bottom key matrix and the bottom value matrix, and decouple the candidate image into multiple target nodes, each of which carries a spatial bounding box and local region features; The topology graph construction module is used to construct a relational topology graph with the multiple target nodes as vertices, establish connected edges based on the spatial geometric relationship between nodes, and assign semantic interaction weights to the connected edges based on the cross-modal matching results between the relational feature vector and the visual interaction features formed by the aggregation of the bottom-level dense feature blocks in the minimum bounded joint region of the target nodes at both ends of the connected edges. The mask generation module is used to establish positive anchor nodes based on the matching results of the local region features and the positive semantic vector. Starting from the positive anchor nodes, a restricted breadth-first search is performed under the joint constraints of the topological order constraint and the semantic interaction weight. The adjacent nodes that are traversed and pass through the cross-modal semantic rejection of the negative semantic vector are determined as the set of locally affected nodes. The spatial coordinates of the set of locally affected nodes are projected onto the feature grid where the bottom dense feature block is located to generate a topological validity mask matrix. The repulsive attention calculation module is used to use the topological validity mask matrix as a hard gate signal to perform element-wise gate on the probability-normalized repulsive response between the bottom dense feature block and the negative semantic vector, so that the repulsive response located outside the region indicated by the topological validity mask matrix is ​​forced to zero. The module also performs non-negative truncation and renormalization after subtracting the restricted repulsive attention score obtained by gate from the positive semantic attention distribution corresponding to the positive semantic vector. The module then uses the obtained attention distribution to weighted aggregate the bottom value matrix to obtain the candidate feature vector that completes local feature suppression. The scoring fusion module is used to calculate the comprehensive score of the candidate image based at least on the candidate feature vector and the positive semantic vector, and output the retrieval results in descending order of the comprehensive score.

[0024] Furthermore, it also includes: The data collection and filtering module is used to take the Top-K search results output by the scoring fusion module as candidate training samples, and perform multi-level filtering such as similarity threshold screening, perceptual hash deduplication, hard example mining, category ratio balancing and confidence screening to form a training sample set and feed it back to the lightweight small model in the large and small model collaborative quality inspection system for iterative training.

[0025] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention provides a collaborative visual retrieval method and system for data center size models based on topological masks. It abandons the traditional global penalty approach and forcibly strips the semantic boundaries of different targets in the underlying physical space by unifying coordinate entity localization and aligning with the region of interest. This achieves independent decoupling of multi-target features and overcomes the problem that the semantic boundaries of multi-target objects overlap under global feature representation and that it is impossible to independently distinguish between retained and excluded features.

[0026] 2. This invention provides a collaborative visual retrieval method and system for data center-sized models based on topological masks. By constructing a relational topology graph that integrates spatial geometric relationships and cross-modal interaction weights, and triggering a restricted breadth-first search with positive anchor points, it accurately locates locally affected nodes by combining topological order and semantic rejection conditions. This generates a topological validity mask matrix, strictly limiting the rejection weights to locally relevant feature regions and forcing rejection responses exceeding the topological constraints to zero. This fundamentally solves the technical problems of uncontrolled rejection range and confusion with irrelevant targets. Unlike the message passing mechanism of graph neural networks (GNN / GCN), restricted breadth-first search can provide the system with unambiguous and highly interpretable physical isolation boundaries through controllable topological order constraints, rigorously and accurately locating locally affected nodes without additional parameter tuning.

[0027] 3. This invention provides a collaborative visual retrieval method and system for data center size models based on topological masks. By placing the Hadamard product operation of the mask after Softmax probability normalization, unlike conventional pre-masking methods, it directly blocks the migration of the probability quality of the masked area to the unmasked area during the renormalization process. This avoids probability leakage and false suppression caused by abnormal amplification of local exclusion penalties, and achieves accurate protection of local irrelevant features.

[0028] 4. This invention provides a data center large-small model collaborative visual retrieval method and system based on topology masking. By adding training data collection and filtering steps, the Top-K results output by retrieval are filtered through multiple levels, including similarity thresholding, perceptual hash deduplication, hard example mining, category ratio balancing, and confidence filtering, to form a training sample set. This set is then fed back to the lightweight small model in the large-small model collaborative quality inspection system for iterative training. This achieves the closed-loop capability of automatically mining high-quality training data from massive monitoring image libraries, overcoming the problems of low efficiency and insufficient hard example coverage in traditional solutions where manual screening of training samples is inefficient. Attached Figure Description

[0029] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0030] Figure 1 A flowchart illustrating the method steps of a collaborative visual retrieval system for data center size models based on topology masks, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structured parsing and tensor encoding logic of the retrieval instruction provided in an embodiment of the present invention; Figure 3 A flowchart of multi-target feature decoupling and node aggregation provided in an embodiment of the present invention; Figure 4 This is a flowchart illustrating the two-stage construction process of a two-end relationship topology graph provided in an embodiment of the present invention. Figure 5 A flowchart of restricted graph traversal search and topology mask generation provided in an embodiment of the present invention; Figure 6 A schematic diagram of the topological constraint exclusionary cross-attention computation architecture provided in an embodiment of the present invention; Figure 7 This is a flowchart of the multidimensional retrieval score extraction and weighted fusion process provided in an embodiment of the present invention; Detailed Implementation

[0031] The technical solution of the present invention will be more clearly and completely explained below with reference to the accompanying drawings and through the description of preferred embodiments of the present invention.

[0032] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0033] like Figure 1 As shown in the embodiment of the present invention, the data center large-small model collaborative visual retrieval system based on topological mask is deployed on the data center intelligent operation and maintenance platform. On the one hand, it provides automatic collection and filtering of training data for the large-small model collaborative quality inspection system (small model fast coarse screening + large model fine recognition). On the other hand, it achieves multi-target decoupling of images by introducing an open-set visual large model within the system. The system as a whole includes steps S1 to S7 in sequence: after receiving an open domain retrieval request, it sequentially executes the following steps: retrieval system initialization and input parsing S1, image multi-target decoupling and feature extraction S2, construction of dual-end relation topology graph S3, generation of local affected node set and topology mask S4, calculation of exclusivity cross-attention of topological constraints S5, comprehensive score fusion and retrieval ranking output S6, and training data collection and filtering S7. Finally, it outputs Top-K target images and feeds the retrieval results back to the training data collection and filtering module. The specific implementation methods of each step are described in detail below.

[0034] S1. Retrieval System Initialization and Input Parsing. This step aims to provide structured feature representations and logical inputs for subsequent multimodal feature decoupling and topological constraint processing. For example... Figure 2 As shown, Figure 2 The logic and data flow of the structured parsing and tensor encoding of the retrieval instruction are illustrated. This step specifically includes the following sub-steps: S1.1 Initialization of the retrieval system environment and multimodal joint feature space. The system starts and instantiates an end-to-end multimodal encoder architecture, specifically including a CLIP-based text encoder and a visual encoder based on a visual Transformer (ViT). To ensure that cross-modal vectors can perform direct inner product operations, the system configures a unified multimodal joint semantic feature space, and uses a linear projection layer to forcibly align visual local feature blocks and text semantic vectors to the same dimensional space. .

[0035] S1.2 Receive open-domain query input containing partial negation semantics. The system front-end receives the user's search request. Based on the cross-modal architecture characteristics of this invention, the system supports dual modes: image-text joint search and fine-grained plain text search. Regardless of whether a query image is included, the core logic triggering relies on the natural language instruction Tquery with negation semantics.

[0036] S1.3. Structural decoupling and open-domain vectorization encoding of retrieval commands. The system inputs the text command Tquery into the natural language processing engine. Through dependency parsing and entity relation extraction, it identifies negation trigger words and their syntactic scope, deconstructing the unstructured text into structured triples. Here, Epos represents a positive semantic phrase, Eneg a negative semantic phrase, and Rneg a relational constraint phrase. Subsequently, the system discards the closed-set mapping and directly inputs these phrases independently into the initialized text encoder, outputting three types of tensor features for downstream network computation: (1) Positive semantic vector It is directly used in S4.1 to establish positive anchor points and in S5.1 to calculate positive attention distribution; (2) Negation semantic vector It is directly used for S4.2 cross-modal semantic rejection matching and S5.2 initial negation correlation calculation; (3) Local topological constraints It contains two sub-parameters: (3.1) Relational eigenvectors In order to The encoded text embedding vector is strictly aligned to d dimensions and used for cross-modal weight mapping with the joint visual region in S3.2; (3.2) The topological order constraint (Hop-count) is an integer parameter generated by the syntactic dependency distance mapping. If the text does not explicitly specify the distance boundary, the system assigns it a preset constant (1-hop), which is used as the depth blocking condition for breadth-first search in S4.1. The specific rule for generating the integer parameter by the syntactic dependency distance mapping is as follows: take the shortest dependency path length L between the negation trigger word and the core word of the positive semantic phrase on the dependency syntax tree, and let the topological order constraint Hop=min(L, Hmax), where Hmax is the preset maximum hop count upper limit (default is 3); when L cannot be calculated or the search instruction does not explicitly specify the distance boundary, take Hop=1 (i.e., 1-hop).

[0037] S2. Image Multi-Target Decoupling and Feature Extraction. To overcome the semantic aliasing of multiple targets caused by global feature representation, the system decouples the input image (containing all candidate images in the database) from the input image. And the query image entered by the user in the combined image and text mode. Parallel execution of low-level dense feature extraction and high-level explicit entity decoupling constructs a unified two-layer feature representation. For example... Figure 3 As shown, Figure 3 The process of multi-target feature decoupling and node aggregation is illustrated, and the specific steps include: S2.1, Low-level dense feature extraction based on a large visual model. The system extracts candidate images... The image is input into an instantiated visual encoder (CLIP-ViT), where it is divided into N non-overlapping low-level local feature blocks (PatchTokens). After self-attention computation through multiple layers of Transformer encoders, the system outputs a dense feature representation aligned in the feature space at the last layer. To support the repulsive cross-attention operation in the subsequent S5 step, the system extracts the linear projection result of this layer and explicitly outputs the low-level key matrix. With the underlying value matrix V ∈ ℝ (N×d) These two matrices constitute the fine-grained visual manifold space of the entire image.

[0038] S2.2 Unified Coordinated Entity Localization and Deduplication Based on Open-Set Visual Model. The system jointly calls Grounding DINO and SAM as an open-set visual model to explicitly locate visual entities in the image. To be compatible with downstream open-domain instructions and achieve closed-loop computation, this step reuses the positive semantic phrase Epos and the negative semantic phrase Eneg, which were decoupled from the structure in step S1.3, as text prompts for Grounding DINO. Specifically, Grounding DINO extracts image region features and calculates their cross-modal cosine similarity with the embedded vectors of the aforementioned text prompts. Boundaries with similarity greater than a preset text threshold (Ttext, default 0.25) are used as initially detected entities. At the same time, the SAM model performs class-independent instance segmentation across the entire image, obtains pixel-level binary masks, and calculates their minimum bounding rectangle.

[0039] To unify the coordinate results of the two and eliminate redundancy, the system performs deduplication based on the intersection-over-union ratio (IoU): for the semantic bounding box output by Grounding DINO and the minimum bounding rectangle of the mask output by SAM, the system calculates their IoU values ​​pairwise.

[0040] (1) Deduplication and fusion strategy: when IoU is greater than a preset determination threshold (e.g., IoU>0.7), the system determines that both detected results correspond to the same target. Since the geometric boundary of the bounding rectangle generated by SAM based on pixel-level masks fits the entity better, while Grounding DINO has clear semantic direction, the coordinates of SAM's minimum bounding rectangle are retained to ensure spatial accuracy, and the semantic confidence score of Grounding DINO is inherited at the same time, so that the two results are fused into one independent entity; (2) Conflict and scale difference adjudication: when there is a spatial intersection between the two detection results but the size difference is large (i.e., 0<IoU≤0.7), the system determines that the two may belong to an inclusion or juxtaposition relationship. In this case, no discarding operation is performed, and both are retained as independent entity nodes at the same time to prevent missed detection of multi-scale complex targets. (3) Processing of pure segmentation regions without semantic boxes: for the minimum bounding rectangle output by SAM that has an intersection-over-union of zero with all semantic bounding boxes, that is, has no spatial intersection, the system retains it as a candidate target node without explicit semantic labels, and its local region features still participate in the topological edge construction in subsequent step S3, the feature-based cross-modal semantic rejection check in step S4.2, and the feature aggregation in step S5; however, due to the lack of semantic confidence from Grounding DINO for such nodes, in the pseudo-label writing in step S7.2 and the confidence screening in step S7.3, the system assigns a preset minimum semantic confidence to them or excludes them from the backflow pseudo-labels, so as to ensure the reliability of the labels of backflow training samples. Finally, all entities in the image are uniformly represented as accurate spatial bounding box coordinates Bi=(x min , y min , x max , y max ).

[0041] S2.3 Entity feature aggregation and target node generation based on target region alignment. After obtaining the purely coordinated spatial bounding box Bi, the system performs a top-down space-to-feature mapping to extract independent semantic representations of each entity. Taking the bounding box Bi as the spatial index, the system calls the region of interest alignment algorithm (RoIAlign) to directly and accurately intercept the local feature blocks falling within the range of Bi from the underlying key matrix K output in step S2.1, then aggregates the intercepted feature blocks through average pooling to generate a local region feature vector with a constant dimension d . Finally, the system encapsulates the spatial coordinates and semantic vector of the entity to generate a unified set of candidate target nodes N={n1, n2, ..., n k , where each node in the set is defined as a binary tuple .

[0042] S3. Construction of a Two-End Relationship Topology Graph. Based on the multiple independent target nodes decoupled in step S2, the system further performs the construction of a two-end relationship topology graph. This aims to overcome the inherent defect of global features lacking local logical relationships, organizing discrete visual entities into a relationship graph with clear relative positions and semantic interaction logic. For example... Figure 4 As shown, Figure 4 The two-stage construction process of the dual-end relationship topology graph is illustrated, which specifically includes the following sub-steps: S3.1 Extraction of the dual-ended node set and graph initialization. The system obtains the query target node set N generated in step S2. q (In the image-text joint mode) and the candidate target node set N c Using these nodes as vertices of the graph, initialize an undirected attribute graph structure, defined as the query topology graph. and candidate topology map .

[0043] S3.2 Calculation of Multidimensional Relationships and Construction of Graph Edges. For the decoupled node set N, the system adopts a two-stage cascaded computation paradigm to construct the graph edge set E, using spatial location to determine the connectivity of the edges and semantic interaction to determine the weight attributes of the edges.

[0044] Phase 1: Establishing topological connectivity based on spatial geometry. For any two nodes n i With n j The system extracts its spatial bounding box and calculates the Euclidean distance D(i,j) between the center point and the intersection-union ratio (IoU(i,j)) of the bounding box. This is done if and only if the following condition is met: or Time (of which) and (As set by the preset spatial adjacency threshold), the system determines that two nodes have the spatial conditions to interact in physical space, within n. i With n j Establish an undirected connected edge e between them i,j Add it to the initial edge set In the middle; for node pairs that do not meet the above spatial thresholds, the system directly determines that there is no possibility of interaction and prunes and isolates them.

[0045] Phase Two: Establishing Semantic Weights Based on Vision-Language Cross-Modal Matching. After establishing spatial connectivity, the system only performs edge matching on the initial edge set. Deep semantic weight assignment is performed on the connected edges in the array: First, for existing connected edges e i,j The system generates surrounding node n i and n j The smallest bounding rectangle is defined as the joint region U.i,j Then, extract the key-value feature matrix K from the output of the visual encoder that falls into U. i,j The local feature blocks within the nodes are used to generate visual interaction feature vectors representing the overall interaction pattern between the two nodes through average pooling. Subsequently, the system obtains the relational feature vector generated by the text encoder in step S1.3. ,calculate With v rel The cosine similarity is calculated, and negative values ​​are truncated using the ReLU activation function to obtain the final semantic interaction weights: (1) The system will assign this weight Assign connected edges In the subsequent graph traversal step (S4.1), the system only allows the traversal algorithm to proceed along... Diffusion search is performed on weighted edges with a value greater than zero.

[0046] S4. Locally Affected Node Set and Topology Mask Generation. To upgrade "global exclusion" to "topologically constrained local exclusion," after constructing the dual-end topology graph, the system accurately defines the exclusion range based on the user's negation semantics and generates a mask matrix for subsequent attention calculations. For example... Figure 5 As shown, Figure 5 The process of restricted graph traversal search and topology mask generation is illustrated. This step specifically includes the following sub-steps: S4.1, Restricted Breadth-First Search (BFS) graph traversal based on multi-anchor triggering. The system receives the positive semantic vector output from step S1. And the candidate topology map constructed in step S3.2 The starting point for traversal is established within the graph. Specifically, the system extracts the local region feature vector Fi of all nodes in the graph and calculates its relationship with... The cosine similarity is used to determine the similarity score, which is greater than a preset positive recall threshold. Extract all nodes and form a set of positive anchor nodes. To accommodate scenes where multiple similar positive targets exist in an image. Subsequently, the system... Each node in the array is an independent root node. In a limited parallel execution of a breadth-first search, the diffusion of traversal is strictly subject to two joint constraints: (1) Order constraint: The depth of the traversal path cannot exceed the local topological constraint conditions defined in step S1.3. The maximum number of hops in the timeline (Hop-count); (2) Weight threshold constraint: The system only allows the semantic interaction weights calculated in step S3.2. Expand the connected edges (where (This is the edge activation threshold). The system will iterate through all adjacent nodes it touches and store them in the candidate affected queue.

[0047] S4.2 Determine the set of locally affected nodes based on a semantic rejection mechanism. For each node in the candidate affected queue, the system performs a final semantic identity verification: the system extracts the local region feature vector of that node. Calculate its relationship with the negation semantic vector. cosine similarity ,like Greater than the preset semantic rejection threshold If a node is determined to be both within the exclusion range in the spatial and logical network and belongs to the excluded object in the ontological semantics, the system will formally push the node that meets this condition into the local affected node set. If, after the traversal, no node in the queue satisfies the condition, then... empty set .

[0048] S4.3 Spatial Coordinate Downsampling Projection and One-Dimensional Topological Validity Mask Matrix Generation. To seamlessly sink high-dimensional node-level exclusion constraints to the underlying Transformer attention layer, the system performs spatial coordinate projection onto the feature mesh. The system initializes an all-zero one-dimensional topological validity mask matrix. (Where N is the total number of local feature blocks output by the ViT visual encoder, and N = H′ × W′, where H′ and W′ are the height and width of the feature map mesh), then different processing is performed based on the content of the node set: (1) Robust handling of the empty set: If The system directly outputs the current global zero matrix. And terminate this step, which will cause the subsequent repulsive attention operation module to output a zero tensor, and retain all feature responses completely through residual bypass; (2) Coordinate grid mapping: if For each node in the set, the system extracts its bounding box coordinates at the original image size. And using the overall downsampling stride S of the visual encoder, the original coordinates are mapped to the index boundaries of the feature map grid:

[0049]

[0050] The system then places the row index in the two-dimensional feature grid matrix at... And the column index is in Set all grid elements within the interval to 1, then flatten the 2D grid to 1×N dimensions, and output the final topology validity mask matrix. .

[0051] S5. Topology-Constrained Repulsive Cross-Attention Calculation. During the matching calculation phase, to accurately remove negated semantic features from the underlying feature manifold, the system configures and executes a topology-constrained repulsive cross-attention operator. For example... Figure 6 As shown, Figure 6 The underlying computational architecture and data flow of this operator are shown. The specific derivation and calculation process are as follows: S5.1 Calculation of Positive Semantic Cross-Attention Distribution. The system uses the positive semantic vector generated in step S1... As a query term, the low-level local feature block sequence matrix extracted in step S2 As a key matrix, with As the value matrix, the positive semantic similarity distribution matrix is ​​calculated using standard scaled dot product attention. :

[0052] S5.2, Global Relevance Calculation of Negation Semantics. To capture all visual features in the image related to the rejection intention, the system utilizes the negation semantic vector. Calculate the global correlation matrix with key term K. :

[0053] Here It contains all matches in the image. The original response value of the object.

[0054] S5.3, Rejection Penalty Activation Constrained by Topological Mask. To prevent global feature degradation caused by negation semantics, the system activates the one-dimensional topological validity mask matrix generated in step S4.3. As a hard-gated signal injection attention mechanism, a restricted repulsive attention score is calculated through the Hadamard product of the postmask. :

[0055] because Only within the feature regions constrained by local topology are the values ​​1, while in other regions they are 0. Therefore, the repulsion weights of irrelevant similar targets scattered outside the topologically restricted regions are forced to zero by the zero elements of the matrix, avoiding excessive suppression of irrelevant target features. Furthermore, this invention places the masking Hadamard product operation after Softmax probability normalization, unlike conventional pre-masking methods: the sum of the probability distributions output by Softmax is always 1. If a mask is applied before Softmax, the probability quality of the masked region will transfer to the unmasked region during the renormalization process, leading to an abnormal amplification of the local repulsion penalty, resulting in probability leakage and false suppression. This scheme performs binary gating truncation after Softmax, directly blocking the unexpected migration of repulsion weights and achieving precise protection of local irrelevant features.

[0056] S5.4 Candidate Aggregation Feature Extraction Based on Repulsion Subtraction. The system performs feature subtraction in the attention weight space. To prevent the global attention distribution from violating the probability axiom (i.e., the sum is not equal to 1) due to weight subtraction, the system introduces a non-negative truncation and renormalization mechanism based on the ReLU activation function:

[0057]

[0058] in, For the repulsion strength hyperparameter, To prevent the denominator from being zero, a minimum constant is achieved. Ultimately, the system utilizes the corrected attention distribution. The value matrix V is weighted and aggregated to extract the final candidate feature vectors. It is important to note that the feature subtraction mechanism of this invention is fundamentally different from the classifier-free guidance (CFG) mechanism in existing generative diffusion models. The two are not equivalent or simply interchangeable. The core differences are as follows: (1) The mask generation and constraint mechanisms are different. The negative repulsion of the diffusion model is a soft suppression with no difference across the entire graph domain and no clear spatial boundary, while the repulsion attention score of this scheme is different. Hard binary mask generated jointly by pre-constrained BFS topological constraints and cross-modal rejection Gated filtering provides precise spatial constraint boundaries; (2) The physical meaning of hyperparameters differs from their domain of action in the diffusion model CFG. The global guiding coefficients, which act across the entire feature domain, are prone to mistakenly suppressing irrelevant background targets, while the proposed scheme... It only applies to the repulsion penalty intensity within a local area of ​​topological constraints, and its effective range is precisely controllable. (3) The application scenarios and problems are different. The diffusion model is used to generate guidance and realize the regulation of the overall attributes of the image, while this solution is used for visual retrieval to solve the technical problem of modeling the topological relationship between targets and accurately isolating local features in complex scenarios.

[0059] S6. Comprehensive Score Fusion and Retrieval Ranking Output. After completing the underlying topology-constrained attention calculation, the system fuses multi-dimensional matching signals to generate the final retrieval ranking results. For example... Figure 7 As shown, Figure 7 The process of multidimensional retrieval score extraction and weighted fusion is shown, which specifically includes the following sub-steps: S6.1 Extraction and Calculation of Multi-Level Scoring Indicators. For each candidate image in the database, the system extracts independent matching scores in the following three dimensions: (1) Implicit feature similarity (Simfeat): The system obtains the candidate feature vectors output in step S5. Calculate its relationship with the positive semantic vector. The cosine similarity score implicitly incorporates the feature suppression effect of local repulsion penalty; (2) Maximum similarity of anchor points (Simobj), the system retrieves the set of forward nodes established in step S4.1 = To accommodate situations where multiple positive entities exist in an image, the system employs a max-pooling strategy to calculate... Features of all nodes and Maximum similarity:

[0060] (3) Topological consistency score ( To assess whether the local graph structure of the candidate image covers the logic of the retrieval instruction, the system reuses the graph edge information generated in steps S3.2 and S4.1. Extract the semantic interaction weights of all activated connected edges from the restricted graph traversal path. And take its maximum value as If the traversal path is empty, then =0. This calculation method enables the system to adaptively calculate topological consistency even in pure text retrieval mode without query images (only text retrieval instructions). It should be noted that the topological relationship consistency score is used to characterize whether the candidate image possesses the spatial relationship context described by the retrieval instruction, that is, whether the candidate image falls within the applicable scope of the retrieval instruction, and serves as a relevance measure for comprehensive scoring reference; while the suppression and exclusion of the rejected target itself are independently completed by the exclusionary cross-attention in step S5 and the local exclusion penalty term in step S6.2, with clear division of labor and no conflict between the two. In the image-text joint retrieval mode, the system further performs structural matching between the query topological graph and the relationship topological graph of the candidate image, and incorporates the graph structure similarity obtained by aligning nodes and connected edges into the topological relationship consistency score, thereby enhancing the matching and discrimination at the relationship level when a query image exists.

[0061] S6.2 Calculation of explicit local exclusion veto penalty. To compensate for potential numerical leakage in the underlying attention mechanism, the system introduces a strong penalty term at the logic layer. The system retrieves the set of locally affected nodes determined in step S4.2. :like ,but ;like The system extracts the cross-modal semantic rejection similarity of all nodes in the set. And calculate the maximum penalty value:

[0062] in For indicator functions, This is a hard blocking threshold. This design ensures that the severity of the penalty depends entirely on the precise match between the candidate object and the user's negative intent in the cross-modal feature space, rather than being limited to the detection confidence of a preset closed category.

[0063] S6.3, Comprehensive Scoring Formula Fusion and Final Ranking Output. The system uses a multi-dimensional weighted fusion formula to calculate the final comprehensive score (Score) for candidate images:

[0064] in, , , , The non-negative fusion weight coefficients preset for the system, This is a non-linear scaling term implemented using the Sigmoid activation function, used to smoothly map the penalty value to the [0, 1] interval, preventing extreme outliers from causing numerical overflow or rendering the search ranking index ineffective. Finally, the system sorts all candidate images in descending order of score and extracts the top-K images with the highest scores as the search results to be output to the user's front end.

[0065] S7. Training Data Collection and Filtering. To ensure that the visual retrieval system described in this invention forms a complete data loop within the data center-scale model collaborative quality inspection architecture, after outputting the Top-K retrieval results, the system further performs automatic collection and filtering of training data, transforming the retrieval results into a sample set for iterative training of lightweight small models. This step specifically includes the following sub-steps: S7.1 System Positioning and Data Interface in the Large and Small Model Collaborative Architecture. The visual retrieval system described in this invention is deployed on the intelligent operation and maintenance platform of the data center, and interfaces with the lightweight small model coarse screening module and the visual large model fine inspection module in the large and small model collaborative quality inspection system through standardized interfaces. Specifically, the small model coarse screening module performs frame-by-frame rapid classification on the real-time monitoring video stream, and writes frames suspected of being abnormal into the coarse screening result library; the retrieval system of this invention uses this coarse screening result library as a candidate image library, receives retrieval instructions with local negative semantics issued by operation and maintenance personnel (e.g., "find monitoring frames with abnormal device indicator lights but excluding red safety symbols"), and outputs Top-K retrieval results after processing in steps S1 to S6; the Top-K results are then imported into the data collection and filtering module, and after multi-level filtering, a training sample set is formed, which is then fed back into the small model iterative training process.

[0066] S7.2 Training Sample Collection Based on Top-K Retrieval Results. The system obtains the Top-K candidate images and their comprehensive scores from step S6 as the initial input for training sample collection. For each candidate image, the system performs the following collection operations: (1) Triggering condition determination: The system only activates the training data collection process when the false alarm rate or false negative rate of the small model coarse screening module exceeds the preset iteration trigger threshold, so as to avoid the waste of computing resources caused by redundant collection; (2) Sample format encapsulation: The system encapsulates each candidate image into a unified training data format, and each sample contains image data. Search command text Positive semantic tags and negative semantic tags The semantic tags directly reuse the triples obtained from the structural decoupling in step S1.3. No additional manual labeling is required; (3) Source of pseudo-labels: For each target entity detected by the open set visual large model in step S2.2 in the candidate image, the system assigns its semantic bounding box coordinates. The semantic confidence score is directly used as a pseudo-label and written into the sample, while the set of locally affected nodes determined in step S4.2 is also included. The nodes in the data are marked as negative pseudo-labels, enabling the integrated reuse of the retrieval and annotation processes.

[0067] S7.3 Specific Decision Logic of the Multi-Level Filtering Strategy. To ensure the quality and diversity of the returned training samples, the system sequentially performs the following five levels of filtering on the collected initial samples: (1) Similarity threshold screening: The system uses the comprehensive score calculated in step S6 as the confidence level screening criterion, and retains those with a score greater than the preset inclusion threshold. Images with a default score of 0.65 are used as the initial sample for inclusion, and low-quality search results with too low a score are removed; (2) Perceptual hash deduplication: The system calculates the perceptual hash value of each initially included sample. When the perceptual hash Hamming distance between any two samples is less than a preset deduplication threshold When the default score is 8, the two samples are considered to be nearly identical. Only the sample with the higher overall score is retained, and the remaining identical samples are removed to avoid redundancy in the training set. (3) Difficult example mining: The system calculates the target node and the negation semantic vector in the retained sample. Cross-modal semantic rejection similarity, placing semantic rejection similarity within a preset hard example range. Samples within the specified range are marked as difficult samples and given a higher priority for inclusion. The lower limit for difficult cases (default 0.45). This is the hard blocking threshold in step S6.2; the samples within this range are the boundary cases that the model is most likely to confuse at present, and prioritizing their inclusion can significantly improve the discrimination ability of the small model in difficult scenarios; (4) Category Balance: The system counts the number of samples in each scenario category (such as abnormal equipment indicator lights, personnel violations, messy cables, etc.) in the current training sample set. When the sample ratio of any category deviates from the target ratio, the system will determine the balance. When the default percentage is 15%, the system will prioritize supplementing samples of that category from the initial sample collection until the category balance condition is met, to prevent small model bias caused by the skewed distribution of the training set categories. (5) Confidence screening: The system performs a final verification on the semantic confidence of pseudo-labels in each sample, and removes all pseudo-labels whose confidence is lower than the preset minimum confidence threshold. (Default 0.3) samples are used to ensure the reliability of the labels on the re-trained samples.

[0068] S7.4, Closed Loop of Training Data Backflow for Small Model Iterative Training. After the above five-level filtering, the system writes the final retained sample set into the training sample queue and backflows it into the small model iterative training process in an incremental update manner. Specifically, the system merges the newly collected samples with the historical training set, performs mini-batch resampling according to the class ratio balancing strategy, and then inputs them into the small model for fine-tuning training. After training is completed, the model weights of the small model coarse screening module are updated and redeployed to the online inference environment, completing the data closed loop of "retrieval-collection-filtering-backflow training-redeployment". This closed loop enables the small model to continuously and automatically obtain high-quality training data covering difficult scenarios from a massive monitoring image library, without the need for manual image screening and labeling, significantly reducing the data acquisition cost of small model iterative training.

[0069] To further clarify the execution logic of this invention, the method steps S1 to S7 described above can be combined into a specific system architecture. Structurally, this system is divided from top to bottom into an input layer, a feature central layer (corresponding to S1 and S2), a logical constraint layer (corresponding to S3 and S4), a tensor operation and decision layer (corresponding to S5 and S6), and a data closed-loop layer (corresponding to S7). The modular division of this system and the signal flow between modules are as follows: (1) Input parsing module: corresponding to step S1, it serves as the entry point of the data bus and is responsible for receiving instructions and extracting positive semantics in tensor form. ), negation semantics ( ) and relational characteristics ( ); (2) Feature extraction module: Corresponding to step S2, it calls the visual Transformer and the open set visual large model to extract the underlying key-value feature matrix (K, V) and the target node set (N), and outputs them to the downstream. (3) Topology graph construction module: corresponding to step S3, serving as the logical constraint hub, calculating semantic edge weights based on spatial and cross-modal features. ), construct candidate topology graphs ( ); (4) Mask generation module: Corresponding to step S4, it triggers a restricted breadth-first search (BFS) based on the forward anchor point, and locks the affected nodes through logical gating. ), and map to generate a one-dimensional topological validity mask ( ); (5) Repulsive attention calculation module: corresponding to step S5, it acts as a tensor operation engine and introduces a mask after the Softmax probability calculation. Perform a hard Hadamard gating product and output candidate aggregate features. ); (6) Scoring fusion module: corresponding to step S6, it extracts matching scores from the four-dimensional signals and outputs the final Top-K search results.

[0070] (7) Data collection and filtering module: Corresponding to step S7, the Top-K search results output by the scoring fusion module are filtered through multiple levels of similarity threshold screening, perceptual hash deduplication, difficult example mining, category ratio balancing and confidence screening to form a training sample set, which is then fed back to the lightweight small model in the large and small model collaborative quality inspection system to perform iterative training, thereby realizing the data closed loop of retrieval-collection-filtering-backflow training.

[0071] To verify the effectiveness of the technical solution described in this invention and to quantitatively evaluate its technical advantages in overcoming the defects of existing technologies, this embodiment conducted a systematic comparative experiment on a dataset of construction scenarios in the communications field constructed in a real industrial setting. This experiment used 2239 images of construction scenarios in the communications field collected internally. These images generally exhibit characteristics such as multiple personnel working simultaneously, complex backgrounds, and haphazard placement of objects. To verify the generalization ability of this invention, the experiment set two typical retrieval tasks with local negative semantic constraints: Task 1 (retrieval of people not wearing safety helmets), retrieving images of "people not wearing safety helmets" (520 positive samples, 1719 negative samples), requiring the retrieval system to positively anchor "people" and locally exclude "safety helmets" associated with them; Task 2 (retrieval of people not wearing reflective vests), retrieving images of "people not wearing reflective vests" (1317 positive samples, 922 negative samples), requiring the retrieval system to positively anchor "people" and locally exclude "reflective vests" associated with them.

[0072] The baseline methods used for comparison in this experiment include: Baseline Method A (traditional global non-exclusion retrieval), which uses a standard global CLIP feature extractor, inputs mixed text, and does not include feature decoupling and independent exclusion modules; Baseline Method B (retrieval based on global exclusion), which uses negative text in the global feature space to perform indiscriminate feature subtraction and score reduction penalties on the features of the entire image; and the method of this invention (local topological exclusion retrieval), which performs unified coordinate entity localization (independent decoupling) and topological mask gating truncation (local exclusion).

[0073] Experiment 1 was used to verify the effectiveness of the present invention in overcoming the problem of "semantic boundaries of multiple target objects overlapping under global feature representation". The retrieval performance comparison of Task 1 is shown in Table 1.

[0074] Table 1

[0075] As shown in Table 1, the accuracy of baseline method A is only 41.2%, mainly because its underlying feature extraction method causes deep aliasing of attributes of different targets. When "person A wearing a helmet" and "person B not wearing a helmet" exist simultaneously in the image, baseline method A cannot independently distinguish between the two at the feature level, leading to system confusion. This invention introduces unified coordinate entity localization and RoIAlign region of interest feature extraction in step S2, forcibly stripping the semantic boundaries of different targets in the underlying physical space, achieving independent decoupling of multi-target features, and increasing the F1 score to 83.6%, fully demonstrating the technical superiority of this invention in complex multi-target scenarios.

[0076] Experiment 2 was used to verify the effectiveness of the present invention in overcoming the problem of "feature degradation caused by the global exclusion mechanism". Taking reflective clothing as an example, if a reflective clothing is draped on a chair in the background of the image, the existing global exclusion mechanism will suppress the features of the entire image, causing images containing the target to be searched (violationers who are not wearing reflective clothing) to be incorrectly filtered out by the system. The comparison and analysis of the false negative rate of the exclusion mechanism in Task 2 is shown in Table 2.

[0077] Table 2

[0078] As shown in Table 2, although baseline method B improves the accuracy from 45.3% to 59.1% by introducing a global exclusion mechanism, the cost is a surge in the false negative rate to 72.8%. This is because baseline method B can only understand "avoid reflective clothing features" and cannot define the local topological boundaries of the exclusion operation. When an image contains a worker violating regulations (who should be detected) but a reflective clothing is hanging on the scaffolding in the background, baseline method B will impose a penalty on the features of the entire image, resulting in severe feature degradation and causing the legitimate image to be directly discarded. This invention relies on the one-dimensional topological validity mask matrix generated in steps S4 and S5 to implement hard gating truncation, accurately defining the physical boundaries of the exclusion operation, and calculating exclusion attention only for local regions that are topologically associated with "people". Experiments show that this invention not only improves the accuracy to 86.4%, but also significantly reduces the false negative rate to 11.5%, effectively overcoming the global feature degradation problem and avoiding the excessive suppression of features and the false negative of positive samples caused by global exclusion triggered by unrelated background targets (such as reflective clothing placed at a distance).

[0079] To further verify the applicability of the present invention in the intelligent operation and maintenance scenario of data centers, a complete retrieval embodiment containing local negation semantics is given below, taking data center monitoring frame retrieval as the object.

[0080] The search command is set to "find monitoring frames where the device indicator light is abnormal but the red safety indicator is excluded". This command contains the positive semantic "device indicator light abnormal", the negative semantic "red safety indicator", and the implicit spatial relationship constraint "locally adjacent", which is a typical local negative search requirement.

[0081] Step S1 (Instruction Decoupling): The system identifies the negation trigger word "exclude" and its scope through dependency parsing, deconstructing the instruction into structured triples: the positive semantic phrase Epos = "device indicator light malfunction", the negative semantic phrase Eneg = "red safety indicator", and the relational constraint phrase Rneg = "locally adjacent". Each of the three phrases is then processed by a text encoder to output a positive semantic vector. Negation semantic vector Relationship eigenvectors Based on the syntactic dependency distance between the negation trigger word and the positive semantic phrase, the system assigns the topological order constraint to 1-hop (a preset constant).

[0082] Step S2 (Multi-target decoupling): The system performs low-level dense feature extraction and open-set entity localization on candidate monitoring frames. Grounding DINO and As text prompts, semantic bounding boxes for all "device indicator lights" and "safety symbols" in the image are detected; SAM performs instance segmentation on the entire image to obtain the minimum bounding rectangle. After IoU deduplication and fusion, each entity in the image is encapsulated as a target node. ,in For spatial bounding boxes, This represents the local region features after being truncated and averaged by RoIAlign.

[0083] Step S3 (Topology Graph Construction): The system uses the target node as the vertex and establishes connected edges based on the distance between the center points of the spatial bounding box and the intersection-union ratio. For each connected edge, the system generates the minimum bounding joint region U(i,j) surrounding the two end nodes. Feature blocks falling within U(i,j) are extracted from the underlying key matrix K and generated as visual interaction features through average pooling. The system then calculates the relationship feature vectors. The cosine similarity is truncated using ReLU and used as the semantic interaction weight of the connected edge.

[0084] Step S4 (Topology Mask Generation): The system first combines local region features with... Nodes with similarity greater than the positive recall threshold are established as positive anchors (i.e., each "device indicator light abnormal" node). A restricted BFS (depth not exceeding 1-hop, expanding only along connected edges with semantic interaction weights greater than the edge activation threshold) is performed starting from each anchor. For each traversed neighboring node, the system calculates its local region features and... Based on the similarity, nodes with a similarity greater than the semantic rejection threshold (i.e., nodes that are topologically associated with the anomaly indicator and semantically belong to the "red safety indicator") are pushed into the local affected node set. The system will then... The bounding box coordinates of each node are projected onto the feature mesh to generate a topological validity mask matrix. .

[0085] Step S5 (Calculation of Exclusive Cross-Attention): The system uses... Calculate the positive semantic attention distribution for the query terms, in order to After calculating the negative correlation distribution with the key matrix K and normalizing it using Softmax probability, it is then compared with... Element-wise multiplication forces the repulsion responses outside the topologically constrained regions to zero. Thus, the repulsion weights of red safety marker regions topologically associated only with the anomaly indicator are retained, while the repulsion weights of independently suspended red safety markers (without topological association with the anomaly indicator) are zeroed, preventing false suppression. The system subtracts the constrained repulsion attention score, weighted by the repulsion intensity coefficient, from the positive attention distribution, performs non-negative truncation and renormalization, and then weights and aggregates the value matrix V to obtain the candidate feature vector.

[0086] Step S6 (Score Output): The system integrates implicit feature similarity, anchor point ontology maximum similarity, topological relationship consistency score, and local exclusion penalty term to calculate the comprehensive score (Score) of each candidate monitoring frame and sort them in descending order, outputting the Top-K retrieval results. The above embodiments demonstrate that the present invention can accurately limit the exclusion scope of the negative semantic "red safety indicator" to a local area with a topological association with "device indicator light abnormality," rather than indiscriminately suppressing all red indicators in the image, effectively supporting the application scenario limitation of "data center" in the name.

[0087] Those skilled in the art should understand that some of the technical means in the above embodiments can be replaced by equivalent methods, specifically including: (1) Alternative solutions for basic large model and network architecture. First, replacement of multimodal encoder: In the above embodiment, step S1.1 adopts a dual-stream architecture based on CLIP and ViT. In practical applications, it can also be replaced by a single-stream architecture or other pre-trained basic models (such as BLIP, ALBEF, FLAVA, METER, etc.). Second, replacement of open-domain visual localization model: In step S2.2, Grounding DINO and SAM are used to extract the coordinates of entities without preset labels. In applications targeting specific vertical fields (such as industrial defect detection, specific medical image retrieval), if open-domain generalization ability is not required, it can be replaced by a closed-set object detection or segmentation network (such as YOLO series, Faster R-CNN, Mask R-CNN), which can also generate entity bounding boxes for downstream processing.

[0088] (2) Alternative solutions for instruction parsing and feature aggregation. First, natural language instruction deconstruction: In step S1.3, dependency parsing is used to extract triples. This step can be replaced by prompt word engineering extraction based on generative large models (such as ChatGPT, LLaMA). That is, by inputting specific instructions into the large model, it can directly output JSON format positive semantics, negative semantics and logical relation structured data. Second, local feature extraction operators: In step S2.3, RoIAlign and average pooling are used to extract regional features. This can be replaced by RoIPooling, precise mask pooling (directly using the pixel-level mask output by SAM to extract features, rather than the bounding rectangle), or attention-based feature aggregation to obtain more refined local node feature vectors.

[0089] (3) Alternative solutions for topological map construction and determination. First, spatial adjacency determination logic: In step S3.2, spatial connectivity is determined by Euclidean distance of the center point and bounding box. As an alternative, Gaussian distance kernel function can be introduced to calculate spatial decay weight, or in the case of RGB-D depth map data, the two-dimensional distance can be extended to three-dimensional spatial point cloud Euclidean distance to construct a 3D topological map. Second, semantic weight mapping algorithm: In stage two of step S3.2, cosine similarity combined with ReLU activation function is used to calculate weight. Alternatively, a learnable small multilayer perceptron (MLP) or bilinear mapping layer can be introduced to input joint visual features and textual relationship features and directly output nonlinear connectivity probability weights.

[0090] Furthermore, the system described in this invention is primarily applicable to intelligent operation and maintenance scenarios in data centers. As a training data collection and filtering source for collaborative quality inspection systems of large and small models, it automatically retrieves specific scene images from the massive monitoring image library accumulated by the small model for iterative training of the small model. It is also applicable to visual semantic retrieval scenarios containing local negation semantics, such as communication construction scenarios, fine-grained product retrieval, and medical image-assisted retrieval (e.g., searching for "lung nodules but excluding calcified lesions locally adjacent to the nodule").

[0091] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0092] The above-described specific embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Various modifications, substitutions, and improvements made by those skilled in the art to the technical solutions of the present invention based on the provided textual description and drawings, without departing from the design concept and spirit of the present invention, should all fall within the scope of protection of the present invention. The scope of protection of the present invention is determined by the claims.

Claims

1. A collaborative visual retrieval method for data center size models based on topology masks, characterized in that, include: S1. Decouple the retrieval command with local negation semantics by structural means to obtain positive semantic vector, negative semantic vector, relation feature vector and topological order constraint; S2. Extract the bottom dense feature blocks of the candidate image to obtain the bottom key matrix and the bottom value matrix, and decouple the candidate image into multiple target nodes carrying spatial bounding boxes and local region features; S3. Construct a relational topology graph with the target node as the vertex, establish connected edges based on the spatial geometric relationship between nodes, and assign semantic interaction weights to the connected edges based on the cross-modal matching results between the relational feature vector and the visual interaction features formed by the aggregation of the bottom-level dense feature blocks in the minimum bounded joint region of the nodes at both ends of the connected edges. S4. Based on the matching result of the local region features and the positive semantic vector, establish the positive anchor node and take it as the starting point. Under the joint constraints of the topological order constraint and the semantic interaction weight, perform a restricted breadth-first search. Determine the neighboring nodes that have been traversed and have passed through the semantic rejection of the negative semantic vector as the set of locally affected nodes, and generate the topological validity mask matrix accordingly. S5. Using the topological validity mask matrix as a hard gate signal, the rejection response between the bottom dense feature block and the rejection semantic vector is gated element by element, so that the rejection response outside the region indicated by the mask matrix is ​​forced to zero, and the bottom value matrix is ​​weighted and aggregated accordingly to obtain the candidate feature vector that completes local feature suppression. S6. Calculate the comprehensive score of the candidate image based at least on the candidate feature vector and the positive semantic vector, and output the retrieval results in descending order of the comprehensive score.

2. The collaborative visual retrieval method for data center size models based on topology masks according to claim 1, characterized in that: Step S4 generates a topology validity mask matrix based on the set of locally affected nodes, including: projecting the spatial coordinates of each node in the set of locally affected nodes onto the feature grid where the bottom dense feature block is located, to obtain the topology validity mask matrix; Step S5 obtains the candidate feature vector after local feature suppression, including: using the positive semantic vector as the query term, the underlying key matrix as the key, and the underlying value matrix as the value, calculating the positive semantic attention distribution by scaling dot product attention; calculating the negative correlation distribution by using the negative semantic vector and the underlying key matrix, normalizing the negative correlation distribution and then multiplying it element-wise with the topological validity mask matrix to obtain a restricted rejection attention score; subtracting the restricted rejection attention score weighted by the rejection strength coefficient from the positive semantic attention distribution, performing non-negative truncation and renormalization on the subtraction result, and using the obtained attention distribution to weighted aggregate the underlying value matrix to obtain the candidate feature vector.

3. The collaborative visual retrieval method for data center size models based on topology masks according to claim 1, characterized in that, Step S1 involves structural decoupling of search commands with local negation semantics, including: The search command is input into the natural language processing engine, and the negative trigger words and their syntactic scope are identified through dependency parsing and entity relation extraction. The search command is deconstructed into a structured triple containing positive semantic phrases, negative semantic phrases and relational constraint phrases. The positive semantic phrase, the negative semantic phrase, and the relational constraint phrase are respectively input into the text encoder, and the positive semantic vector, the negative semantic vector, and the relational feature vector are output accordingly. The topological order constraint is generated based on the syntactic dependency distance mapping between the negation trigger word and the positive semantic phrase, and when the retrieval instruction does not explicitly specify the distance boundary, the topological order constraint is assigned a preset constant.

4. The collaborative visual retrieval method for data center size models based on topology masks according to claim 3, characterized in that, Step S2 involves decoupling the candidate image into multiple target nodes, including: The open-set visual localization model is invoked, and the positive and negative semantic phrases decoupled in step S1 are reused as text prompt words to extract the semantic bounding boxes that match the text prompt words across modalities; at the same time, the category-independent instance segmentation model is invoked to obtain the pixel-level binary mask and its minimum bounding rectangle. The intersection-union ratio (IUR) is calculated for the semantic bounding box and the minimum bounding rectangle. When the IUR is greater than a preset threshold, the coordinates of the minimum bounding rectangle are retained and the semantic confidence of the semantic bounding box is inherited, and the two are merged into a single target entity. When the two have spatial intersection and the IUR is not greater than the preset threshold, the two are retained as independent target entities. Using the spatial bounding boxes of each target entity as indices, dense feature blocks falling within the spatial bounding boxes are extracted from the underlying key matrix using region of interest alignment and aggregated into the local region features. The spatial bounding boxes and the local region features are then encapsulated into corresponding target nodes.

5. The collaborative visual retrieval method for data center size models based on topology masks according to claim 1, characterized in that, Step S3 involves establishing connected edges and assigning semantic interaction weights to them, including: For any two target nodes, calculate the distance between the center points of their spatial bounding boxes and the intersection-union ratio (IU). If the distance between the center points is less than a preset distance threshold or the IU is greater than a preset IU threshold, establish a connecting edge between them; otherwise, determine that there is no connecting edge between them and prune and isolate them. For each connected edge, a minimum bounding rectangle enclosing the target nodes at both ends is generated as a joint region. Dense feature blocks falling into the joint region are extracted from the underlying key matrix and aggregated into the visual interaction feature. The similarity between the visual interaction feature and the relation feature vector is calculated and truncated non-negatively, and then used as the semantic interaction weight of the connected edge.

6. The collaborative visual retrieval method for data center size models based on topology masks according to claim 1, characterized in that, Step S4 involves performing a restricted breadth-first search starting from the positive anchor node, including: All nodes in the relation topology graph whose local region features are more similar to the positive semantic vector than a preset positive recall threshold are identified as positive anchor nodes, in order to accommodate the situation where there are multiple similar positive targets in the same candidate image. The restricted breadth-first search is performed in parallel with each of the positive anchor nodes as an independent root node. The restricted breadth-first search is subject to the following two constraints: the depth of the traversal path does not exceed the maximum number of hops limited by the topological order constraint; and expansion is only performed along connected edges whose semantic interaction weight is greater than a preset edge activation threshold.

7. The collaborative visual retrieval method for data center size models based on topology masks according to claim 2, characterized in that, Step S4 involves determining the set of locally affected nodes and generating a topology validity mask matrix, including: For each neighboring node traversed, calculate the similarity between its local region features and the negation semantic vector, and push neighboring nodes with similarity greater than a preset semantic rejection threshold into the set of locally affected nodes. When the set of locally affected nodes is empty, the topological validity mask matrix with all zeros is output so that the subsequent rejection response outputs a zero tensor and completely preserves the feature response of the underlying dense feature block through residual bypass. When the set of locally affected nodes is not empty, the spatial bounding box coordinates of each node in the set of locally affected nodes are mapped to the index interval of the feature grid according to the downsampling step size of the feature grid. After setting the elements in the feature grid that are within the index interval, they are flattened to obtain the topology validity mask matrix.

8. The collaborative visual retrieval method for data center size models based on topology masks according to claim 1, characterized in that, Step S6 calculates the comprehensive score of the candidate image, including: Calculate the similarity between the candidate feature vector and the positive semantic vector, and use it as the implicit feature similarity. The maximum similarity between the local region features of each positive anchor node and the positive semantic vector is calculated as the maximum similarity of the anchor ontology. Extract the maximum value of the semantic interaction weight of the activated connected edges on the restricted breadth-first search path starting from the positive anchor node, and use it as the topological consistency score. Set the topological consistency score to zero when the path is empty. A local rejection penalty term is calculated based on the cross-modal semantic rejection similarity between each node in the locally affected node set and the negative semantic vector, and the local rejection penalty term is set to zero when the locally affected node set is empty. The comprehensive score is obtained by weighting and summing the implicit feature similarity, the maximum similarity of the anchor point ontology, and the topological relationship consistency score, and then subtracting the local exclusion penalty term after nonlinear scaling.

9. The collaborative visual retrieval method for data center size models based on topology masks according to claim 1, characterized in that, It also includes training data collection and filtering step S7: S7. The search results output in step S6 are used as candidate training samples. Multi-level filtering, including similarity threshold screening, perceptual hash deduplication, hard example mining, category ratio balancing and confidence screening, is performed to form a training sample set. The sample set is then fed back to the lightweight small model in the large and small model collaborative quality inspection system for iterative training.

10. A collaborative visual retrieval system for data center size models based on topology masks, used to implement the collaborative visual retrieval method for data center size models based on topology masks as described in any one of claims 1-9, characterized in that, include: The input parsing module is used to structurally decouple retrieval commands with local negation semantics, and obtain positive semantic vectors, negative semantic vectors, relation feature vectors, and topological order constraints. The feature extraction module is used to extract the bottom dense feature blocks of the candidate image to obtain the bottom key matrix and the bottom value matrix, and decouple the candidate image into multiple target nodes, each of which carries a spatial bounding box and local region features; The topology graph construction module is used to construct a relational topology graph with the multiple target nodes as vertices, establish connected edges based on the spatial geometric relationship between nodes, and assign semantic interaction weights to the connected edges based on the cross-modal matching results between the relational feature vector and the visual interaction features formed by the aggregation of the bottom-level dense feature blocks in the minimum bounded joint region of the target nodes at both ends of the connected edges. The mask generation module is used to establish positive anchor nodes based on the matching results of the local region features and the positive semantic vector. Starting from the positive anchor nodes, a restricted breadth-first search is performed under the joint constraints of the topological order constraint and the semantic interaction weight. The adjacent nodes that are traversed and pass through the cross-modal semantic rejection of the negative semantic vector are determined as the set of locally affected nodes. The spatial coordinates of the set of locally affected nodes are projected onto the feature grid where the bottom dense feature block is located to generate a topological validity mask matrix. The repulsive attention calculation module is used to use the topological validity mask matrix as a hard gate signal to perform element-wise gate on the probability-normalized repulsive response between the bottom dense feature block and the negative semantic vector, so that the repulsive response located outside the region indicated by the topological validity mask matrix is ​​forced to zero. The module also performs non-negative truncation and renormalization after subtracting the restricted repulsive attention score obtained by gate from the positive semantic attention distribution corresponding to the positive semantic vector. The module then uses the obtained attention distribution to weighted aggregate the bottom value matrix to obtain the candidate feature vector that completes local feature suppression. The scoring fusion module is used to calculate the comprehensive score of the candidate image based at least on the candidate feature vector and the positive semantic vector, and output the retrieval results in descending order of the comprehensive score; The data collection and filtering module is used to take the retrieval results output by the scoring fusion module as candidate training samples, and perform multi-level filtering such as similarity threshold screening, perceptual hash deduplication, hard example mining, category ratio balancing and confidence screening to form a training sample set and feed it back to the lightweight small model in the large and small model collaborative quality inspection system for iterative training.