Interactive image-text retrieval method based on distinction degree optimization
By working together with dialogue rewriting, similarity optimization selection, and visual expansion modules, the stability and robustness issues of existing image retrieval technologies in complex scene-level retrieval are resolved, achieving more accurate and stable image retrieval results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN UNIV
- Filing Date
- 2026-01-22
- Publication Date
- 2026-05-01
AI Technical Summary
Existing image retrieval technologies suffer from insufficient stability and limited robustness in complex scene-level retrieval tasks. In particular, when users provide fuzzy sketches and brief descriptions, it is difficult to gradually refine the intent through interaction, resulting in inaccurate and unstable retrieval results.
An interactive image-text retrieval method based on discriminability optimization is adopted. The dialogue rewriting module generates concise declarative query statements, the similarity optimization selection module evaluates the global optimal discriminative power, and the Diffusion visual extension module generates high-quality synthetic images. Multiple rounds of interaction are conducted to optimize the retrieval process.
It achieves more robust and accurate scene-level image retrieval in complex interaction processes. The dialogue rewriting module extracts semantic information, the similarity optimization selection module maintains retrieval stability, and the Diffusion visual extension module enhances visual detail matching, thereby improving the stability and accuracy of retrieval.
Smart Images

Figure CN121958581A_ABST
Abstract
Description
Interactive image and text retrieval method based on discriminative optimization Technical Field
[0001] This invention relates to the field of image retrieval technology, and more specifically to an interactive image and text retrieval method based on discriminative optimization. Background Technology
[0002] With the rapid growth of multimedia data and the increasing demand for intelligent interaction, image retrieval technology has become a crucial bridge connecting user intent with massive amounts of visual content. Traditional image retrieval systems often follow a static paradigm of single query-result return. Users enter keywords or upload sample images to perform a search, and the interaction ends after the system returns similar results. However, in complex scene-level retrieval tasks, initial user queries are often vague, abstract, and incomplete. For example, a user may only provide a rough sketch of a scene with a brief text description. A single-round search is insufficient to fully capture the user's deeper intent, leading to inaccurate search results.
[0003] To address the aforementioned issues, existing technologies have proposed interactive text-image retrieval methods and scene-level image retrieval methods based on text and sketches.
[0004] Interactive image retrieval breaks through the traditional static paradigm of a single query and result return, reshaping the retrieval process into a multi-round iterative human-computer dialogue, allowing users to dynamically guide and refine their search intent based on preliminary results. However, existing interactive retrieval methods often rely on the assumption that the latest is the best, directly using the most recent round of information for retrieval, lacking dynamic evaluation and optimization of historical judgment states, which may lead to fluctuations or even degradation of search results during iterations. This ignores the complexity of real-world interaction: users have limited patience, feedback may contain noise, and not all users stop upon first seeing their target; their goal may be to find multiple suitable options, or they may be dissatisfied with the preliminary results but unsure if it's due to unclear descriptions, thus continuing the dialogue. In this case, the stability of the system's search results becomes crucial.
[0005] In scene-level image retrieval methods that combine text and sketches, text excels at describing the abstract semantic information of objects, while sketches are adept at describing the specific outlines and spatial layouts of objects. Combining the two achieves modality crossing, providing users with a more flexible and visually-oriented expression, particularly suitable for retrieval needs where users have ideas about scene composition but find it difficult to describe them precisely in words. While it has made progress in static retrieval tasks, it still relies on the assumption of single-round, one-time queries, failing to consider the inherent incompleteness and abstractness of user input. Especially when users can only provide rough sketches and vague descriptions, it lacks the ability to progressively refine intent through interaction, limiting its robustness in open scenarios.
[0006] In summary, existing image retrieval technologies suffer from insufficient stability and limitations in practicality and robustness. Summary of the Invention
[0007] To address the aforementioned shortcomings of existing technologies, this invention provides an interactive image and text retrieval method based on discriminability optimization, which solves the problems of insufficient stability and limited practicality and robustness of existing image retrieval technologies.
[0008] To achieve the above-mentioned objectives, the technical solution adopted by the present invention is as follows: An interactive image-text retrieval method based on discriminability optimization is provided, comprising the steps of: S1, obtaining the sketch to be retrieved. and its initial text description S2 serves as the initial query for multi-turn interactive dialogue in a visual language model; S3, indicating the number of interaction turns. The initial value is 1, and the global optimal discrimination score is initialized. and the global optimal similarity matrix S3, The visual language model performs the first step based on the input vector. Round interaction, obtain the current round including , and the historical sequence of dialogues in historical question-and-answer pairs And through the dialogue rewriting module Generate a declarative query statement S4, Calculation The text-image similarity matrix is compared with the image database, and the current round is calculated based on the text-image similarity matrix. Discrimination S5, Input The similarity optimization selection module makes its judgments. Is it greater than the global optimal discrimination score? If so, then let and order The text-image similarity matrix is equal to the current round's similarity matrix; otherwise, it remains unchanged. and S6, will and Input the Diffusion Vision Extensions module and generate synthetic images using its diffusion model. ,calculate S7. Calculate the graph similarity of each image in the search image database with the global optimal similarity. Then, weight and fuse the graph similarity of each image in the search image database to obtain the fused similarity of each image. Finally, sort all images in the search image database according to the fused similarity, and prioritize the top-ranked images. Zhang's image is used as the search result for the current round; S8, let And repeat steps S3 to S7 until... More than the preset number of rounds Alternatively, the user can actively terminate the conversation, outputting the final round of search results as the final answer to the multi-round interactive dialogue.
[0009] Furthermore, The expression is: in, and These represent the number of positive and negative images in the image database, respectively. and These represent the average and maximum similarity between the positive example image and the text image, respectively. and These represent the average and maximum similarity between the negative example image and the text image, respectively. The standard deviation of the image-text similarity for all images; and These are the standard deviations of the positive and negative images, respectively. This is the smoothing coefficient.
[0010] Furthermore, the expression for the fusion similarity obtained by weighted fusion is: in, To retrieve the image from the image library Zhang Image Fusion similarity; For synthesized images and Graph similarity between them; for and The global optimal image-text similarity between them; For fusion weighting coefficients.
[0011] Furthermore, The expression is: ;in, and The first The questions and answers in a round-robin interactive dialogue; and These are the questions and answers from the first round of interactive dialogue.
[0012] Furthermore, The expression is: ;in, It is a visual language model.
[0013] Furthermore, the globally optimal discrimination score The initial value is 0, and the global optimal similarity matrix is... The initial value is an empty matrix.
[0014] Compared with existing technologies, this invention has the following significant advantages: 1. In each round of interactive dialogue, this invention does not simply use the original dialogue directly for retrieval, but instead executes three key steps sequentially: First, the dialogue rewriting module condenses and reconstructs the accumulated dialogue history to generate a concise and informative cross-modal query description, abandoning the traditional assumption that the latest is the best, thus improving retrieval stability; Second, the similarity optimization selection module evaluates the global discriminative performance of the current round's generated query and compares and selects the best one with the historical global optimal discriminative score, ensuring that the global optimal similarity matrix driving the retrieval is always the feature representation with the strongest discriminative power; Finally, at the end of the interaction, the Diffusion visual expansion module combines the declarative query statement with the initial sketch to generate a high-quality synthetic image, and performs the final refined retrieval by fusing graph similarity with the global optimal image-text similarity. This series of designs enables this invention to synergistically utilize the semantic evolution of interactive dialogue, the relative discriminative power of similarity, and the visual prior of the diffusion model to achieve more robust and accurate scene-level image retrieval in complex interactive processes.
[0015] 2. Considering that although multi-turn interactive dialogues can gradually clarify user intent, the original dialogue sequences often contain redundant, ambiguous, or colloquial natural language expressions. If directly used as retrieval queries, this would introduce significant noise and computational burden to cross-modal alignment. Therefore, this invention designs a sketch-guided dialogue rewriting module. This module can identify and integrate key constraints and modifiers from historical question-and-answer pairs, while filtering out irrelevant conversation details, ultimately generating a declarative query statement that is highly consistent with the user's goal semantically and adapted to the image retrieval encoder in form.
[0016] 3. Existing interactive retrieval systems typically implicitly assume that the amount of information monotonically increases with each round of dialogue. Therefore, directly using the similarity calculated in the latest round is considered the optimal strategy. However, this invention discovers that the core of retrieval performance lies in the model's ability to effectively distinguish the target image from distractors, rather than the absolute increase in the similarity of the target image. Based on this key insight, this invention proposes a plug-and-play similarity optimization selection module. This module does not directly adopt the similarity of the current round after each interaction, but instead evaluates its overall discriminative efficiency and compares it with historical best results, thereby dynamically selecting the optimal similarity matrix for retrieval.
[0017] 4. Considering that although the dialogue rewriting and similarity optimization selection modules effectively extract semantic information and stabilize retrieval discriminative power, relying solely on the text modality for matching may still be limited by the incompleteness of language description and the gap between it and the visual modality. To further enhance the visual understanding and matching of users' complex and abstract intentions, this invention introduces a Diffusion visual extension module. The core idea of this module is to leverage the powerful text-to-image generation capabilities and rich visual prior knowledge of the pre-trained diffusion model to combine the optimized declarative query with a sketch to generate a high-quality synthetic image. By weighted fusion of graph similarity based on the synthetic image and text-image similarity, it cleverly combines the rich visual details provided by the generative model with the precise semantic information provided by the declarative query. This fusion strategy fully utilizes the advantages of different modalities and compensates for the shortcomings of a single modality, thereby obtaining a more comprehensive and accurate image ranking result.
[0018] 5. Fusion weighting coefficient As an adjustable parameter, it allows for dynamic adjustment of the proportion of visual and textual information in the final decision based on specific application scenarios or data characteristics, enhancing the flexibility and configurability of the method. Attached Figure Description
[0019] Figure 1 is a flowchart of the interactive image and text retrieval method based on discriminability optimization; Figure 2 shows the retrieval performance evaluation results of different models in multi-turn dialogue tests on the FS-COCO-Human-Like dataset; Figure 3 shows the retrieval performance evaluation results of different models in multi-turn dialogue tests on the FS-COCO-Rich dataset; Figure 4 shows the retrieval performance evaluation results of different models in multi-turn dialogue tests on the SketchyCOCO-Human-Like dataset; Figure 5 shows the retrieval performance evaluation results of different models in multi-turn dialogue tests on the SketchyCOCO-Rich dataset; Figure 6 is a visualization of retrieval result examples. Detailed Implementation
[0020] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0021] Example 1: Considering the instability, limited practicality, and limited robustness of existing image retrieval technologies, this example provides an interactive image-text retrieval method based on discriminability optimization. Referring to Figure 1, the method includes the following steps: S1, obtaining the sketch to be retrieved. and its initial text description , serving as the initial query for multi-turn interactive dialogue in a visual language model.
[0022] S2, Interaction Rounds The initial value is 1, and the global optimal discrimination score is initialized. and the global optimal similarity matrix Among them, the globally optimal discrimination score The initial value is 0, and the global optimal similarity matrix is... The initial value is an empty matrix.
[0023] S3, The visual language model performs the first step based on the input vector. Round interaction, obtain the current round including , and the historical sequence of dialogues in historical question-and-answer pairs And through the dialogue rewriting module Generate a declarative query statement .
[0024] Specifically, considering that although multi-turn dialogues can gradually clarify user intent, the original dialogue sequences often contain redundant, ambiguous, or colloquial natural language expressions. Directly using these as retrieval queries would introduce significant noise and computational burden to cross-modal alignment. Therefore, this embodiment designs a sketch-guided dialogue rewriting module, aiming to transform the dynamic dialogue context into a stable, concise, and information-dense retrieval query. The core idea of this module is to utilize a large visual language model... Powerful multimodal understanding and generation capabilities will enable the current round historical sequence As input, it automatically synthesizes a concise, accurate, and easy-to-retrieve declarative query. The model's task is not simply to concatenate text, but to deeply understand the entire interaction process, identify and integrate key constraints and modifiers, while filtering out irrelevant conversational details, ultimately generating a declarative query that is semantically highly consistent with the user's goal and formally adapted to the image retrieval encoder. , This module can compress lengthy multi-turn dialogues into a single, highly condensed description, thereby significantly improving the efficiency and robustness of subsequent text-image similarity calculations and providing clearer semantic-driven signals for the entire interaction process.
[0025] S4, Calculation The text-image similarity matrix is compared with the image database, and the current round is calculated based on the text-image similarity matrix. Discrimination .
[0026] S5, Input The similarity optimization selection module makes its judgments. Is it greater than the global optimal discrimination score? If so, then let and order The text-image similarity matrix is equal to the current round's similarity matrix; otherwise, it remains unchanged. and .
[0027] Specifically, existing interactive retrieval systems typically implicitly assume that the amount of information monotonically increases with each round of dialogue, thus making the use of the similarity calculated in the latest round the optimal strategy. However, the inventors discovered that the core of retrieval performance lies in the model's ability to effectively distinguish the target image from distractors, rather than the absolute increase in the similarity of the target image. Based on this key insight, this embodiment proposes a plug-and-play similarity optimization selection module. This module does not directly adopt the similarity of the current round after each interaction, but instead evaluates its overall discriminative efficiency and compares it with historical best states, thereby dynamically selecting the optimal similarity matrix for retrieval. Specifically, in the... In this embodiment, the round is based on the rewritten declarative query statement. The image database is searched, and a text-image similarity matrix is calculated using an image-text retrieval model (such as BLIP). The target image (which can be determined in labeled evaluations or obtained through user feedback in practical applications) is used as a positive example, and the remaining images are used as negative examples. The expression is: in, and These represent the number of positive and negative images in the image database, respectively. and These represent the average and maximum similarity between the positive example image and the text image, respectively. and These represent the average and maximum similarity between the negative example image and the text image, respectively. The standard deviation of the image-text similarity for all images; and These are the standard deviations of the positive and negative images, respectively. The smoothing coefficient is used. Step S5 ensures that the retrieval process is always driven by the most discriminative state among all interaction sequences, thereby fundamentally avoiding retrieval performance fluctuations or even degradation caused by ambiguity, noise, or invalid information introduced by a certain round of dialogue, and achieving a stable improvement in retrieval accuracy as the interaction progresses.
[0028] S6, will and Input the Diffusion Vision Extensions module and generate synthetic images using its diffusion model. ,calculate Image similarity with images in the retrieved image library.
[0029] S7. Weighted fusion of the image similarity and the global optimal similarity for each image in the retrieved image database to obtain the fused similarity for each image. Then, sort all images in the retrieved image database according to the fused similarity, and prioritize the top-ranked images. The images are used as the search results for the current round.
[0030] Specifically, considering that although the dialogue rewriting and similarity optimization selection modules effectively extract semantic information and stabilize retrieval discriminative power, relying solely on the text modality for matching may still be limited by the incompleteness of language description and the gap between it and the visual modality. To further enhance the visual understanding and matching of users' complex and abstract intentions, this embodiment introduces a Diffusion visual extension module. The core idea of this module is to leverage the powerful text-to-image generation capability and rich visual prior knowledge of the pre-trained diffusion model to combine the optimized declarative query with a sketch to generate a high-quality synthetic image. By weighted fusion of graph similarity based on the synthetic image and text-image similarity, it cleverly combines the rich visual details provided by the generative model with the precise semantic information provided by the declarative query. This fusion strategy fully utilizes the advantages of different modalities and compensates for the shortcomings of a single modality, thereby obtaining a more comprehensive and accurate image ranking result. The expression for the fused similarity obtained by weighted fusion is: in, To retrieve the image from the image library Zhang Image Fusion similarity; For synthesized images and Graph similarity between them; for and The global optimal image-text similarity between them; For fusion weighting coefficients.
[0031] S8, Order And repeat steps S3 to S7 until... More than the preset number of rounds Alternatively, the user can actively terminate the conversation, outputting the final round of search results as the final answer to the multi-round interactive dialogue.
[0032] In summary, the beneficial effects of this embodiment are as follows: In each round of interactive dialogue, this embodiment does not simply use the original dialogue directly for retrieval, but instead executes three key steps sequentially: First, the dialogue rewriting module condenses and reconstructs the accumulated dialogue history to generate a concise and informative cross-modal query description, abandoning the traditional assumption that the latest is the best, thus improving retrieval stability; second, the similarity optimization selection module evaluates the global discriminative performance of the current round's generated query and compares and selects the best one with the historical global optimal discriminative score, ensuring that the globally optimal similarity matrix driving the retrieval is always the feature representation with the strongest discriminative power; finally, at the end of the interaction, the Diffusion visual expansion module combines the declarative query statement with the initial sketch to generate a high-quality synthetic image, and performs the final refined retrieval by fusing graph similarity with the globally optimal image-text similarity. This series of designs enables this embodiment to synergistically utilize the semantic evolution of interactive dialogue, the relative discriminative power of similarity, and the visual prior of the diffusion model to achieve more robust and accurate scene-level image retrieval in complex interactive processes.
[0033] Example 2 This example is a further limitation based on Example 1. Its purpose is to provide an experiment of an interactive image and text retrieval method based on discriminability optimization (hereinafter referred to as IScene). Other parts not mentioned refer to Example 1 or the prior art.
[0034] Existing sketch retrieval datasets primarily focus on static, single-turn matching tasks, where a given query is directly retrieved as the corresponding image target. While these datasets have played a significant role in advancing sketch-image retrieval research, they cannot support the multi-turn dialogue process necessary for interactive retrieval. The core of interactive retrieval lies in multi-turn dialogue; the system needs to progressively understand and refine the user's intent based on historical dialogues, but existing datasets generally lack such dialogue annotation data. To fill this gap and provide an evaluation basis for the interactive framework of this research, this embodiment constructs the first multi-turn dialogue dataset specifically designed for text- and sketch-based interactive scene-level image retrieval tasks.
[0035] Specifically, this embodiment selects two widely used scene-level sketch retrieval benchmark datasets as the basis for creation: FS-COCO (refer to CHOWDHURY PN, SAIN A, BHUNIA AK, et al. FS-COCO: Towards Understanding of Freehand Sketches of Common Objects in Context[C] / / AVIDAN S, BROSTOW G, CISSÉ M, et al. Computer Vision – ECCV 2022.) and SketchyCOCO (GAO C, LIUQ, XU Q, et al. SketchyCOCO: Image Generation from Freehand Scene Sketches[C]. 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2020, pp. 5173-5182.). Both provide high-quality scene sketches, realistic images, and corresponding text descriptions, forming a sketch-image-text triplet, which provides a necessary starting point for interactive dialogue. However, their original test sets only support single-turn retrieval. Therefore, this embodiment expands its test set, selecting 3000 and 210 triples respectively as the starting point for interactive dialogues. Each triple contains a scene sketch, a real image, and a corresponding text description. To automatically generate high-quality multi-turn dialogues, this embodiment draws on the paradigm of ChatIR (LEVY M, BEN-ARI R, DARSHAN N, et al. Chatting Makes Perfect: Chat-based Image Retrieval[J]. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS '23), 2023,36: 61437-61449.), utilizing powerful visual language models (such as Qwen3-VL) to simulate the questioner and responder, thus simulating the human-computer interaction process. The questioner generates new follow-up questions based on the current sketch, the initial text description, and the accumulated dialogue history; the responder then provides a simulated user's answer based on the generated questions and the corresponding real image. Through this iterative process, this embodiment automatically generates a dialogue of up to 10 rounds for each triple.To fully capture the diversity of user interaction patterns in the real world, this embodiment specifically designs two distinct instruction strategies to guide the model, thereby generating multi-turn dialogues with different styles: one is a human style, simulating everyday, non-professional communication with casual and brief dialogues; the other is an expert style, simulating a rigorous intent clarification process, with richer and more detailed questions and answers in the dialogue. Furthermore, to ensure the final quality of the dataset, this embodiment introduces a manual polishing stage after the automatic generation process to improve the accuracy and standardization of the dialogues.
[0036] Ultimately, this embodiment constructs four new benchmark datasets: FS-COCO-Human-Like, FS-COCO-Rich, SketchyCOCO-Human-Like, and SketchyCOCO-Rich, as shown in Table 1.
[0037] During retrieval, the image search space was set to the complete image library of their respective original datasets (10,000 images for FS-COCO and 14,081 images for SketchyCOCO) to simulate real-world application scenarios.
[0038] This embodiment selects two representative baseline methods for comparison: 1) ZS (Zero-Shot Retrieval). This baseline only uses the refined text output by the dialogue rewriting module in the ICSene framework for querying; 2) DAR. This method combines the dialogue rewriting module and the Diffusion visual extension module to generate synthetic images in the initial round and fuse the similarity between the text and the generated images for retrieval.
[0039] This embodiment uses Qwen3-VL to construct the dataset, BLIP as the retrieval model, the dialogue rewriting module uses the BLIP-3 model, and the Diffusion visual extension module uses the T2I-Adapter to facilitate sketch-based control of generation. The fusion weight is set to 0.85. To balance retrieval overhead and performance, this embodiment only generates auxiliary synthetic images in the initial round (round 0). This embodiment uses Recall@K (R@K, K=1, 10), a common retrieval metric, as the core evaluation indicator. It measures the hit rate of containing the target image in the first K returned results of the current round. R@1 reflects the exact matching capability, and R@10 reflects the system's recall capability. This embodiment also reports the Hits@10 metric, commonly used in interactive retrieval, which is the proportion of the target image cumulatively appearing in the top 10 results across all rounds, used to comprehensively evaluate ranking quality.
[0040] To verify the effectiveness of the method in this embodiment, experiments were conducted on the four proposed datasets, and the experimental results are shown in Figures 2-5. The experimental results show that IScene significantly outperforms the baseline methods ZS and DAR on all datasets. Overall, IScene's retrieval accuracy shows a stable and continuous increase with the number of interaction rounds, especially in the Recall@1 metric, which is the primary measure of retrieval accuracy. On the larger datasets FS-COCO Human-Like and FS-COCO-Rich, IScene achieves approximately 14% to 17% improvement in Recall@1 in the final rounds compared to the two baseline methods. On the smaller SketchyCOCO dataset, IScene's performance advantage further expands, with a Recall@1 improvement of over 50% compared to the baseline methods on the SketchyCOCO-Rich dataset. This fully demonstrates that IScene can effectively address the retrieval challenges in small-sample scenarios.
[0041] Looking at the Recall@10 metric, which reflects the top 10 retrieval hits in the current round, IScene also leads across all datasets. Particularly on the SketchyCOCO dataset, its Recall@10 is significantly higher than the baseline method. This indicates that IScene not only maintains the target image at the top of the leaderboard but also ensures its stable appearance at the top of the retrieval list, resulting in higher overall ranking quality. It's worth noting that although the baseline method and IScene traded blows in the Hits@10 metric in some intermediate rounds on certain datasets, IScene consistently matched or nearly matched it in the final rounds. This demonstrates that IScene pursues high precision without sacrificing retrieval recall breadth.
[0042] Analyzing the performance curves across different dialogue rounds reveals a fundamental difference between IScene's performance growth pattern and that of baseline methods. Baseline methods ZS and DAR often experience slower or fluctuating growth in later rounds, reflecting that relying solely on the latest round information may introduce noise and lead to performance stagnation. IScene's curve, however, is smoother and exhibits a continuous upward trend. This non-decreasing performance improvement directly validates the effectiveness of the similarity optimization selection module proposed in this embodiment. This module, by dynamically evaluating and retaining the historically strongest discriminative search states, successfully avoids performance regression caused by low-quality dialogue in a single round, ensuring a stable increase in search results.
[0043] Further comparison of datasets with different dialogue styles reveals that IScene's improvement is more significant on the expert-style (Rich) dataset. For example, on FS-COCO-Rich, the final Recall@1 score is higher than on FS-COCO-Human-Like; the same is true on SketchyCOCO-Rich. This indicates that when the dialogue content is more formal and detailed, IScene's dialogue rewriting module can more effectively extract semantic information. Meanwhile, IScene also maintains its leading position on the Human-Like dataset, demonstrating its good adaptability to short, conversational interactions.
[0044] This embodiment selects a query case from FS-COCO-Human-Like to qualitatively and intuitively demonstrate IScene's retrieval performance. As shown in Figure 6, this embodiment displays the changes in the Top-1 image of the retrieval results from round 0 to round 10. Given an initial sketch and corresponding description, after several rounds of dialogue, IScene successfully retrieved the target image and maintained it in subsequent rounds, while the results of ZS and DAR showed significant fluctuations. This visualization strongly confirms IScene's core design philosophy: the traditional latest-is-best strategy exposes users to the risk of random fluctuations in retrieval results, and a single ambiguous feedback can lead to the loss of all previous efforts; while IScene's similarity optimization selection module maintains and continuously utilizes historical best discrimination states, ensuring the monotonic non-decreasing performance of the system. This means that once the target is successfully located in the top ranks, the system will not forget it due to its own instability. This not only theoretically guarantees retrieval efficiency, but also provides users with a predictable and steadily accumulating search process in actual experience, fundamentally improving the certainty of the interaction and user confidence.
[0045] In summary, the beneficial effects of this approach are as follows: 1. It constructs the first text-sketch scene retrieval dataset containing large-scale, multi-style, multi-turn dialogues, providing a crucial data foundation for interactive retrieval research.
[0046] 2. An interactive retrieval framework for IScene was proposed, which achieves a stable improvement in retrieval performance through the collaborative work of three modules: dialogue rewriting, similarity optimization selection, and visual expansion.
[0047] 3. Experiments on the newly created dataset show that the IScene framework outperforms existing methods in both retrieval accuracy and robustness, verifying the effectiveness and superiority of introducing interactive mechanisms into text-sketch scene retrieval.
Claims
1. An interactive image and text retrieval method based on discriminative optimization, characterized in that, The steps include: S1, obtaining the sketch to be retrieved. and its initial text description S2 serves as the initial query for multi-turn interactive dialogue in a visual language model; S3, indicating the number of interaction turns. The initial value is 1, and the global optimal discrimination score is initialized. and the global optimal similarity matrix S3, The visual language model performs the first step based on the input vector. Round interaction, obtain the current round including 、 and the historical sequence of dialogues in historical question-and-answer pairs And through the dialogue rewriting module Generate a declarative query statement S4, Calculation The text-image similarity matrix is compared with the image database, and the current round is calculated based on the text-image similarity matrix. Discrimination S5, Input Similarity optimization selection module, judgment Is it greater than the global optimal discrimination score? If so, then let and order The text-image similarity matrix is equal to the current round's similarity matrix; otherwise, it remains unchanged. and S6, will and Input the Diffusion Vision Extensions module and generate synthetic images using its diffusion model. ,calculate S7. Calculate the graph similarity of each image in the search image database with the global optimal similarity. Then, weight and fuse the graph similarity of each image in the search image database to obtain the fused similarity of each image. Finally, sort all images in the search image database according to the fused similarity, and prioritize the top-ranked images. Zhang's image is used as the search result for the current round; S8, let And repeat steps S3 to S7 until... More than the preset number of rounds Alternatively, the user can actively terminate the conversation, outputting the final round of search results as the final answer to the multi-round interactive dialogue.
2. The interactive image and text retrieval method based on discriminability optimization according to claim 1, characterized in that, The expression is: in, and These represent the number of positive and negative images in the image database, respectively. and These represent the average and maximum similarity between the positive example image and the text image, respectively. and These represent the average and maximum similarity between the negative example image and the text image, respectively. The standard deviation of the image-text similarity for all images; and These are the standard deviations of the positive and negative images, respectively. This is the smoothing coefficient.
3. The interactive image and text retrieval method based on discriminability optimization according to claim 1, characterized in that, The expression for the weighted fusion similarity is: in, To retrieve the image from the image library Zhang Image Fusion similarity; For synthesized images and Graph similarity between them; for and The global optimal image-text similarity between them; For fusion weighting coefficients.
4. The interactive image and text retrieval method based on discriminability optimization according to claim 2, characterized in that, The expression is: ;in, and The first The questions and answers in a round-robin interactive dialogue; and These are the questions and answers from the first round of interactive dialogue.
5. The interactive image and text retrieval method based on discriminability optimization according to claim 2, characterized in that, The expression is: ;in, It is a visual language model.
6. The interactive image and text retrieval method based on discriminability optimization according to claim 1, characterized in that, Global optimal discrimination score The initial value is 0, and the global optimal similarity matrix is... The initial value is an empty matrix.