Image retrieval method and system, electronic equipment and medium
By employing a multi-dimensional confidence-driven supplementary question strategy in the image retrieval system, the problem of low retrieval efficiency caused by ambiguous user input was solved, achieving fast and accurate image retrieval and improving user experience and interaction efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, when the natural search query entered by the user is complex and ambiguous, electronic devices cannot accurately understand the user's intent, resulting in low search efficiency. Furthermore, the fixed confirmation process of multi-turn dialogue systems reduces user experience and interaction efficiency.
A dynamic, multi-dimensional confidence-based supplementary question strategy is adopted to generate targeted clarifying supplementary questions. By updating the semantic elements in the dialogue state object, the system accurately guides users to input clear semantic elements, thereby improving interaction efficiency.
By employing a dynamic question-filling strategy, the system quickly converges to the user's final query target, improving the accuracy and efficiency of image retrieval, reducing unnecessary dialogue rounds, and enhancing the user experience.
Smart Images

Figure CN121808087A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of terminal technology, specifically relating to an image retrieval method, system, electronic device, and medium. Background Technology
[0002] Currently, terminals can retrieve images using a single-modal approach based on understanding image content (such as object detection and scene recognition).
[0003] Specifically, users can trigger electronic devices to perform image-based searches using natural language commands, such as "find photos of calico cats." The electronic devices can then perform searches based on "calico cats" and output all photos containing "calico cats."
[0004] However, if the user's natural search query is complex and ambiguous, the electronic device may not be able to accurately understand the user's intent. In such cases, the device can output a fixed re-question to allow the user to re-enter the natural search query. For example, if a user enters "Looking for a photo of Li Ming and me taken at the beach last summer, when the sun was shining brightly," the electronic device, unable to accurately identify the user's search intent, could output a re-question: "Unable to understand your intent, please re-enter." This fixed re-question format may prevent users from quickly and accurately entering precise search queries, resulting in low search efficiency. Summary of the Invention
[0005] The purpose of this application is to provide an image retrieval method, system, electronic device, and medium that can improve retrieval efficiency.
[0006] In a first aspect, embodiments of this application provide an image retrieval method, which includes: when a dialog state object corresponding to a user-inputted first image retrieval statement includes a first semantic element that does not meet a confidence condition, outputting a supplementary question statement for the first semantic element based on the first semantic element, wherein the first semantic element is slot information or retrieval intent information; updating the semantic elements in the dialog state object based on the user-inputted supplementary question response statement; and performing image retrieval on a first image set based on the slot information in the updated dialog state object to obtain image retrieval results.
[0007] Secondly, embodiments of this application provide an image retrieval system, including: a dialogue management module, an interactive presentation module, and a retrieval and reordering module; the dialogue management module is used to generate a supplementary question statement based on the first semantic element when the session state object corresponding to the user-inputted image retrieval statement includes a first semantic element that does not meet the confidence condition, wherein the first semantic element is slot information or retrieval intent information; the interactive presentation module is used to output the supplementary question statement; the dialogue management module is also used to update the semantic elements in the dialogue state object based on the supplementary question response statement input by the user; the retrieval and reordering module is used to perform image retrieval on a first image set based on the slot information in the updated dialogue state object to obtain image retrieval results.
[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0012] In this embodiment, since a supplementary question can be generated specifically based on the semantic elements in the dialogue state object corresponding to the user's input image retrieval statement that do not meet the confidence condition, the user can be accurately guided to input a supplementary question containing clear and accurate semantic elements. Therefore, the interaction efficiency of natural language-based image retrieval can be improved, and electronic devices can quickly and accurately understand the user's true retrieval intent, thereby improving image retrieval efficiency. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating image retrieval methods in related technologies;
[0014] Figure 2 This is a schematic flowchart of an image retrieval method provided in an embodiment of this application;
[0015] Figure 3 This is a flowchart illustrating the multi-turn dialogue method provided in the embodiments of this application;
[0016] Figure 4 This is a flowchart illustrating a retrieval and rearrangement method provided in an embodiment of this application;
[0017] Figure 5 This is a schematic diagram of the structure of an image retrieval system provided in an embodiment of this application;
[0018] Figure 6 This is a schematic diagram of the structure of an image retrieval system provided in an embodiment of this application;
[0019] Figure 7 This is one of the flowcharts of an electronic device provided in the embodiments of this application;
[0020] Figure 8 This is a second flowchart illustrating an electronic device provided in an embodiment of this application. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0022] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0023] The following explains the terms or terminology used in the embodiments of this application.
[0024] Retrieval-Augmented Generation (RAG) is a technical framework that enhances the generation capabilities of language models by retrieving external knowledge bases, reducing model illusions and improving accuracy.
[0025] Multimodal fusion is a processing method that integrates heterogeneous information such as images, text, and spatiotemporal metadata.
[0026] Quantized Approximate Nearest Neighbor (ANN) is a fast retrieval algorithm that reduces storage through vector compression (such as product quantization, PQ).
[0027] Federated Learning (FL) is a distributed model training framework where devices train locally and only upload parameters for updates.
[0028] Differential privacy (DP) is a privacy protection technique that adds noise to data to prevent the leakage of individual information.
[0029] Slot filling: The task of extracting structured parameters (such as time, people) from user queries.
[0030] Intent confidence threshold ( : A critical value used to determine the reliability of Natural Language Understanding (NLU) parsing results.
[0031] Short-Term Signed (URL): A temporary resource access link with an expiration date and an encrypted signature.
[0032] Hard filtering: Candidate results are directly excluded based on absolute conditions (such as time range).
[0033] Dual-Encoder Architecture: A model structure (such as CLIP) in which images and text are encoded separately into a shared vector space.
[0034] Multilayer Perceptron (MLP): A fully connected feedforward neural network used for the fusion of nonlinear features. MLPs, through their fully connected structure, provide the possibility of setting differentiated weights, and through backpropagation and gradient descent—a "learning engine"—automatically and data-drivenly find the differentiated weight values that best solve a specific task.
[0035] Hybrid Deployment Architecture: This refers to an architecture in which local devices and cloud servers work together.
[0036] Privacy & Security Layer: A security module used to control data encryption, transmission, and access.
[0037] Interaction Presentation Layer: A feedback module used for user interface rendering and interaction.
[0038] Hierarchical Navigable Small World (HNSW): A graph indexing data structure for efficient approximate nearest neighbor search. It constructs a hierarchical graph network, enabling extremely fast searching of most similar items in large-scale vector datasets, and is a core algorithm in many vector databases.
[0039] Recall@K: The proportion of relevant results among the top K results.
[0040] Top-K: Top-K is a sampling strategy used in text generation to select the next token (which can be understood as a word or character) from the model's vocabulary. Its working principle is very simple: when generating each new word, the model calculates the probability distribution of all possible words in the entire vocabulary.
[0041] Top-K Context: This section specifically emphasizes the Top-K approach used in the context of text generation (i.e., "context" generation). "Context" here refers to the sequence of text being generated. The model predicts the next word based on the preceding text (i.e., the current context).
[0042] Average Turns: The number of conversations required for a user to complete a query.
[0043] The image retrieval method, system, electronic device, and medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0044] The image retrieval method provided in this application can be applied to scenarios where image retrieval is performed using natural language.
[0045] In related technologies, accurate capture of user intent (Intent Detection) and extraction of slot information (Slot Filling) are core tasks in multi-turn dialogue systems. When a multi-turn dialogue system is unsure about the user's input search or query, confirmation or clarification is required. Specifically, this can be achieved using a fixed confirmation process or by re-asking based on simple Automatic Speech Recognition (ASR) or NLU score thresholds, such as "Are you sure that's correct?" or "I didn't understand you, could you please explain it more clearly?" This fixed strategy may lead to unnecessary dialogue turns, reducing user experience and interaction efficiency.
[0046] Furthermore, in related technologies, such as Figure 1 As shown, visual-language (VL) retrieval systems in related technologies, such as some pre-trained models, typically rely on separate model architectures, such as performing object detection first and then caption generation. This leads to complex implementation processes, high training costs, and low execution efficiency. When applied to large-scale personal photo album datasets, this separate, sequential processing approach has significant bottlenecks in ensuring real-time performance and accuracy. In contrast, this application uses vector-based retrieval, which eliminates the need to traverse the original image set, thus reducing retrieval latency and improving retrieval accuracy through joint retrieval of multimodal features.
[0047] Furthermore, existing RAG, multimodal fusion, and dialogue strategies have shortcomings: Retrieval-enhanced generation (RAG) technology has been introduced into the field of information retrieval to improve the accuracy and traceability of responses from large language models (LLMs) by leveraging external knowledge bases. However, applying a general RAG framework to multimodal, unstructured personal photo album data presents unique challenges:
[0048] First, there's the challenge of heterogeneity in multimodal fusion: Existing hybrid search methods typically employ simple weighted similarity fusion (Weighted Ranker) or inverse ranking (RRF Ranker) strategies to combine text and image retrieval results. While these methods achieve fusion within the vector space, they often neglect crucial low-dimensional structured metadata in photo album data, such as precise shooting timestamps (EXIF) and geolocation (GPS) information. For personal photo albums, shooting timestamps and geolocation are decisive factors in determining retrieval relevance, but non-linearly fusing these heterogeneous features with high-dimensional visual / text embedding similarity remains a challenge that current technologies haven't effectively addressed. Simply using time or geographic information as a filter sacrifices retrieval recall.
[0049] Second, dialogue management is inefficient: In multi-turn dialogue systems, accurately capturing user intent (IntentDetection) and extracting slot information (Slot Filling) are core tasks. When the system is uncertain about its understanding of user input, confirmation or clarification is required. Existing methods often employ fixed confirmation processes or re-question based on simple ASR / NLU score thresholds. This fixed strategy may lead to unnecessary dialogue turns, reducing user experience and interaction efficiency. There is a lack of a dynamic policy network based on multi-dimensional confidence (intent confidence, key slot position confidence) to accurately generate clarifying follow-up questions and quickly converge to the user's final query goal.
[0050] Third, the conflict between privacy protection and hybrid deployment: Personal photo album data is highly sensitive. While a hybrid cloud architecture (offloading some computing load to local devices) can provide low latency and a degree of privacy protection, when leveraging a robust cloud-based LLM for RAG generation, the retrieved context data still needs to be transferred to the cloud. This introduces serious data breach risks, including embedding leaks, query risks, and insecurity during transmission. Existing RAG systems require more advanced encryption and access control mechanisms, such as applying differential privacy or performing encrypted retrieval operations during the retrieval phase, to ensure the security and compliance of data during transmission and processing.
[0051] Therefore, this application provides a method that can efficiently integrate multimodal or heterogeneous metadata (also known as multimodal or heterogeneous slots), adopt intelligent dialogue strategies to improve interaction efficiency, and support a hybrid architecture with a strict privacy protection mechanism to achieve high-precision dialogue query of semantic albums.
[0052] The image retrieval method provided in this application uses a dynamic, multi-dimensional confidence-based (intent confidence, slot position confidence) supplementary question strategy to accurately generate clarifying supplementary questions, thereby quickly converging to the user's final query target. Specifically, when the dialog state object corresponding to the user's input image retrieval statement includes a first semantic element that does not meet the confidence condition, a targeted supplementary question can be output based on this first semantic element, where the first semantic element is slot information or retrieval intent information. Furthermore, the semantic elements in the dialog state object are updated based on the user's input supplementary question response. Finally, image retrieval is performed on the first image set based on the slot information in the updated dialog state object to obtain the image retrieval results. Thus, since supplementary questions can be generated specifically based on semantic elements that do not meet the confidence condition in the dialog state object corresponding to the user's input image retrieval statement, these supplementary questions can accurately guide the user to input supplementary questions containing clear and accurate semantic elements. This improves the interactive efficiency of natural language-based image retrieval, allowing electronic devices to quickly and accurately understand the user's true retrieval intent, thereby improving image retrieval efficiency.
[0053] The specific scenarios involved in the embodiments of this application will be described by way of example below.
[0054] Scenario 1: A user is using a semantic photo album app on their mobile device. They want to find a set of travel photos, but the query information is vague and fragmented. The image retrieval device can obtain semantic elements that do not meet the confidence criteria with a single query. Here, "semantic photo album app" refers to a photo album that supports natural language retrieval. In other words, the semantic photo album app is only used to explain that the images in the album can be retrieved using natural language search statements, rather than limiting the types, content, or formats of the images in the album.
[0055] First round of interaction: Fuzzy search and dynamic clarification. User inputs image search statement 1: "Help me find those photos that Li Fang and I took at the beach last summer, but I can't remember which beach it was."
[0056] 1. Input parsing and confidence assessment
[0057] Search statement 1 includes the following semantic elements:
[0058] Search intent information: Seaside image search;
[0059] Slot extraction results: Time: Last summer; Person: Li Fang; Location: Seaside.
[0060] The confidence levels of semantic elements are as follows:
[0061] Search intent information (seaside image retrieval): medium confidence (semantic ambiguity).
[0062] Person (Li Fang): High confidence level (from facial recognition tags).
[0063] Location (seaside): Medium confidence (semantic ambiguity).
[0064] Time (last summer): Low confidence level, large time span, and the definition of "summer" varies in different regions.
[0065] 2. Dynamic follow-up question strategy triggering
[0066] Because the confidence level of the time slot location is lower than the preset confidence threshold. Therefore, instead of immediately performing a search, it triggers a more precise follow-up question to quickly narrow down the query range.
[0067] The image retrieval device generates a clarification question based on the time slot "last summer":
[0068] "Okay, regarding 'last summer,' are you referring to our trip to Qingdao in August, or the photos taken in Hainan in June?"
[0069] Second round of interaction: Information convergence and high-precision retrieval
[0070] The user entered a clarification reply: "It's from Qingdao, and what I'm looking for is that really beautiful photo of the two of us facing the sea with the sunset in the background."
[0071] 1. Generation of slot update and retrieval commands
[0072] Slot update:
[0073] Time: Last August; Location: Qingdao; Person: Li Fang; Abstract semantic features: Facing the sea and with the sunset behind her (high-dimensional visual description).
[0074] The image retrieval device has all slots precisely filled (or the confidence level meets the confidence level condition), so it can generate retrieval instructions based on the extracted slot information, execute the retrieval in the personal album, and obtain the retrieval results.
[0075] Third round of interaction: Presentation of search results.
[0076] The image retrieval device outputs the retrieval results for users to view.
[0077] The image retrieval method provided in this application is executed by an image retrieval device, which can be an electronic device, or a functional module or entity within an electronic device. This application does not limit the specific implementation of this method. The following will use an electronic device as an example to illustrate the image retrieval method provided in this application.
[0078] This application provides an image retrieval method. Figure 2 A flowchart illustrating the image retrieval method provided in an embodiment of this application is shown, as follows: Figure 2 As shown, the image retrieval method may include the following steps 100 to 102.
[0079] Step 100: If the first image retrieval statement input by the user includes a first semantic element that does not meet the confidence condition, the image retrieval device outputs a supplementary question statement based on the first semantic element.
[0080] Step 101: The image retrieval device updates the semantic elements in the dialogue state object based on the user's input of the supplementary question and answer statement.
[0081] Step 102: The image retrieval device performs image retrieval on the first image set based on the slot information in the updated dialogue state object, and obtains the image retrieval results.
[0082] The first semantic element is slot information or retrieval intent information.
[0083] To simplify the description and avoid ambiguity, the "dialogue state object used when performing image retrieval" can be described as the "updated dialogue state object". The essential meaning of the "updated dialogue state object" is the "final dialogue state object", that is, "the dialogue state object in which all semantic elements satisfy the confidence condition".
[0084] In some embodiments of this application, "slot information" refers to a set of structured parameters extracted from the user-inputted image retrieval statement, used to specify and realize the retrieval intent. Slot information can be characterized as at least one slot-value pair, where "slot" is a predefined parameter category, and "value" is the specific content provided in the user statement corresponding to that category. Specifically, the image retrieval device extracts the "value" corresponding to the slot from the user-inputted image retrieval statement. For example, for time slot information: time slot - 2025.
[0085] In some embodiments of this application, slot information may include: abstract semantic slot information and specific semantic slot information.
[0086] Among them, abstract semantic slot information represents abstract concepts, while concrete semantic slot information represents concrete concepts.
[0087] In some embodiments of this application, the abstract semantic slot information may include, but is not limited to, at least one of the following: visual slot information, scene slot, emotion slot, lighting slot, texture slot, style slot, composition slot, etc.
[0088] In some embodiments of this application, specific semantic slot information may include: location slot information, time slot information, person slot information, animal slot information, etc.
[0089] In this embodiment of the application, the "value" in "slot information" can also be referred to as "key slot" or "slot". It means "slot-value" pair.
[0090] In some embodiments of this application, the search intent information indicates the user's search intent.
[0091] Search intent information and slot information are related as a whole to its parts, or as abstract to concrete. For example, search intent can be broken down into multiple slot values, or in other words, image search intent can guide the extraction of slots.
[0092] Understandably, to support subsequent dynamic dialogue strategies, a confidence score can be assigned to each identified intent and extracted slot. This score reflects the NLU model's confidence in the parsing results. The confidence score can be calculated based on various factors, including lexical likelihood, language model prediction features, or determined by analyzing the N-best candidate list.
[0093] For example, the confidence score of key slots (such as names or specific dates) is crucial in determining whether follow-up questions are needed.
[0094] In some embodiments of this application, the confidence condition can be: the confidence of the semantic element is greater than or equal to the confidence threshold.
[0095] In some embodiments of this application, the retrieval intent information corresponds to an intent confidence threshold, and the slot information corresponds to a slot position confidence threshold. The slot position confidence threshold and the intent confidence threshold may be the same or different; this application does not limit this.
[0096] For example, semantic elements that do not meet the confidence criteria may include: slot information with a confidence level less than the slot position confidence threshold, and intent information with a confidence level less than the intent confidence threshold.
[0097] It should be noted that the confidence level of slot information actually means the confidence level of the values in the slot information.
[0098] It can be understood that the "dialogue state object corresponding to the image retrieval statement" can be represented as a dynamic table.
[0099] This table can be represented as: Search Intent - Value; Slot - Value, as shown in Table 1.
[0100] Table 1
[0101] Search Intent Slot 1 Slot 2 …… Search intent information Slot value 1a Slot value 2a ……
[0102] As can be seen from Table 1, a dialogue state object contains at least one retrieval intent information, at least one slot associated with the retrieval intent information, and each slot corresponds to a slot value.
[0103] It should be noted that the retrieval intent and each slot in the dialogue state object can be considered a semantic element.
[0104] In some embodiments of this application, after a user inputs the first image search query in a search task, the image search device can establish or maintain a dialogue state object. It can then extract search intent information and slot information from the first image search query. Next, it can evaluate whether each semantic element in the dialogue state object meets the confidence criteria, and based on the evaluation results, perform an image search or trigger a targeted follow-up query.
[0105] In some embodiments of this application, the image retrieval device can... Figure 5 The dialogue management module 201 in the image retrieval system shown maintains the dialogue state object and dynamically adjusts the dialogue strategy based on the semantic element confidence of NLU.
[0106] In some embodiments of this application, such as Figure 5 As shown, the dialogue management module 201 may include an input parsing module and a supplementary question implementation module. The input parsing module is used to parse the semantic elements in the user-inputted statement, determine the confidence level of the semantic elements, and update the semantic elements in the dialogue state object. The supplementary question implementation module can output supplementary questions based on the confidence level of the semantic elements in the dialogue state object, and based on the semantic elements in the dialogue state object that do not meet the confidence level conditions.
[0107] In some embodiments of this application, after the user inputs the first image retrieval statement, the image retrieval device can create and maintain a dialogue state object to track the semantic elements in the user's input statement and the confidence scores of the semantic elements.
[0108] In some embodiments of this application, after receiving an image retrieval statement input by a user, the image retrieval device may first perform NLU parsing on the statement to extract slot information and retrieval intent information.
[0109] In some embodiments of this application, it can be achieved through... Figure 5 The input parsing module in the image retrieval system shown performs ASR recognition and / or NLU parsing on the image retrieval statement to convert the image retrieval statement (which can be natural language) into structured query components, i.e. semantic elements.
[0110] It should be noted that the dialog state object is created when the user enters the first image search statement for a search task and terminates after the search task is completed.
[0111] Specifically, during multiple rounds of interaction between the user and the image retrieval device, the image retrieval device can update the semantic elements in the dialogue state object based on the user's input supplementary questions (also known as clarification statements), and update the confidence level of the semantic elements until the confidence level of all semantic elements in the dialogue state object meets the confidence level condition. Then, it performs a retrieval based on the slot information in the dialogue state object to obtain the image retrieval. Therefore, the "updated dialogue state object" refers to the dialogue state object after all semantic elements meet the confidence level condition.
[0112] In other words, if a semantic element that meets the confidence condition cannot be extracted from the user's second-round input of the supplementary question and response statement, or if the supplementary question and response statement introduces a new semantic element that does not meet the confidence condition, the image retrieval device can continue to ask follow-up questions or supplementary questions about the semantic elements in the dialogue state object that do not meet the confidence condition, until all semantic elements in the dialogue state object meet the confidence condition.
[0113] In some embodiments of this application, each time a user inputs a statement, such as an initial search statement or a follow-up question response statement, the image retrieval device can... Figure 5 The input parsing module in the image retrieval system shown performs intent recognition, slot extraction, and confidence assessment on the search query.
[0114] Then, the image retrieval device can... Figure 5 The image retrieval system shown includes a question-filling module that determines whether the confidence scores of the obtained retrieval intent information and slot information meet confidence conditions, such as whether they exceed the corresponding confidence threshold. When the confidence score of any semantic element information is lower than the preset confidence threshold, the input parsing module can dynamically trigger the generation of precise clarification or question-filling statements based on semantic elements with confidence scores below the threshold, rather than simple fixed confirmations. This refines the query before retrieval execution and improves the efficiency of multi-round interactions.
[0115] In some embodiments of this application, after receiving a user's follow-up question and response statement, the image retrieval device can perform semantic element extraction on the statement to extract at least one semantic element and evaluate the confidence level of each semantic element. Then, the at least one semantic element can be used to update the semantic elements in the dialogue state object.
[0116] It is understood that the image retrieval device updates the semantic elements in the dialogue state object based on the question-and-answer statement, which may include at least one of the following: semantic element replacement, semantic element addition.
[0117] For example, in scenario 1 above, the image retrieval device can replace the original address slot "seaside" with "Qingdao" in the speech state object and update the original time slot "seaside" with "Qingdao" based on the semantic elements in the user's input question, and add an abstract semantic slot: "facing the sea and back to the sunset".
[0118] As can be seen, the dynamic clarification scheme based on intent and slot position confidence in this application can not only improve semantic elements that do not meet the confidence conditions, but also further supplement "new" semantic elements, thereby improving the accuracy of image retrieval, such as the abstract semantic slot added in the second round of interaction in the above scenario 1: "facing the sea and with the sunset behind."
[0119] It should be noted that the slots in the dialogue state object in this embodiment are dynamic, changing dynamically according to the image retrieval statement and the supplementary question statement, as long as they can represent or describe the user's true retrieval intent.
[0120] For example, the dialogue state object corresponding to a retrieval task includes time slots and location slots.
[0121] For example, the dialogue state corresponding to another retrieval task includes task slots and event slots.
[0122] In some embodiments of this application, the image retrieval device performs image retrieval based on the slot information in the dialogue state object only after all semantic elements in the aforementioned dialogue state object meet the confidence condition. It can be understood that the image retrieval device can sequentially update the semantic elements in the dialogue state object using all user-inputted statements to obtain a dialogue state object where all semantic elements meet the execution condition.
[0123] It can be understood that "all statements entered by the user" includes statements entered by the user in multiple rounds. For example, in scenario 1 above, the image retrieval statement entered by the user in the first round of interaction, and the supplementary question and answer statement entered in the second round of interaction.
[0124] In some embodiments of this application, the user input can be either a speech statement or a non-speech statement; this application is not limited to either. If it is a speech statement, it can first be processed by ASR recognition to convert it into a text statement, and then NLU parsing can be performed.
[0125] For example, if the image retrieval statement is speech, it can first be converted into text using automatic speech recognition technology. Natural language understanding is then performed on the text to perform intent recognition and slot extraction.
[0126] For example, if a user enters the image search query "find those photos we took in Dali, Yunnan last year", the intent is identified as "search for Dali travel photos last year". The slot information can include: time slot: "last year", location slot: "Dali, Yunnan".
[0127] The image retrieval method provided in this application embodiment can use a dialogue management strategy based on intent and slot location reliability to dynamically trigger precise follow-up questions, thereby improving the efficiency and query accuracy of multi-turn interactions.
[0128] In some embodiments of this application, the input parsing module may perform the following operations:
[0129] 1. Session state maintenance and slot tracking
[0130] The input parsing module can maintain a dialogue state object, also known as a session state recorder, to track the user's historical queries, current intent, and successfully filled slots and their corresponding confidence scores.
[0131] 1.1 Semantic element extraction, including slot information and retrieval intent information.
[0132] 1.2 Confidence Assessment: The input parsing module checks the confidence scores of all populated slots (e.g., time, person, location) in the dialogue state object. Simultaneously, assess the confidence level of the current intent recognition. .
[0133] In some embodiments of this application, the supplementary question implementation module executes the following dynamic supplementary question decision logic:
[0134] 2.1 Threshold Judgment: The slot position confidence level is... and confidence of intent Each is associated with its corresponding preset confidence threshold ( ) to make comparisons.
[0135] 2.2 Strategy Execution:
[0136] 2.11. If all slots in the dialog state object are... > ,and With sufficient precision, the query execution module can determine that the query intent is clear, and thus pass the instruction to... Figure 5 The image retrieval system 200 shown in the image retrieval system 200 performs the image retrieval operation, that is, performs the next step operation (Go-Ahead).
[0137] 2.22. If the dialog state object has at least one slot... < If a key slot is missing (determined by retrieval intent information), the session management module can trigger precise supplementary questions based on slots that do not meet confidence conditions or are missing. It should be noted that, unlike simple confirmations (such as "Are you sure?") in related technologies, the embodiments of this application dynamically generate precise clarification or supplementary questions based on the dialogue state, low-confidence slots, and low-confidence retrieval intent information.
[0138] For example, if the NLU has low confidence in resolving the "weekend" mentioned by the user, the session management module can generate: "Do you mean the photo taken last Saturday or last Sunday?".
[0139] In some embodiments of this application, the input parsing module may support NLU.
[0140] Thus, the dynamic supplementary question mechanism based on the confidence of semantic elements utilizes the inherent uncertainty of the NLU model to drive the dialogue process, effectively avoiding the problem of inefficient or failed retrieval due to initial semantic understanding errors, while reducing unnecessary confirmation rounds, significantly improving user experience and query success rate.
[0141] The flow of the confidence-driven multi-turn dialogue method based on semantic elements provided in the embodiments of this application will be described below with reference to the accompanying drawings.
[0142] For example, such as Figure 3 As shown, the multi-turn dialogue method includes the following steps:
[0143] Step A1: Receive user input.
[0144] Specifically, the received user input can be an initial search query or a follow-up question.
[0145] Step A2: Perform NLU parsing on the user input statement.
[0146] Step A3: Based on the NLU parsing results, update the semantic elements in the dialogue state object.
[0147] It should be noted that updating the semantic elements in the dialogue turntable object may include: updating the retrieval intent information and updating at least one slot information.
[0148] In some embodiments of this application, after updating the slot information, the slot status can be updated synchronously. The slot status can include any of the following: filled, unfilled. For example, if a time slot contains a time value, the status of that time slot is filled; otherwise, it is unfilled.
[0149] Step A4: Determine if the slot information is complete.
[0150] It is understandable that the completeness of slot information is determined by the search intent information. In other words, the retrieval of slot information can be guided by the search intent information.
[0151] It should be noted that the dialog state object does not need to have fixed slots. Instead, relevant slots can be generated specifically based on the search intent information. For example, in a person image retrieval task, slots for person, time, and location can be generated based on the search intent information.
[0152] In some embodiments of this application, "slot" may also be referred to as "critical slot".
[0153] In some embodiments of this application, if a critical slot is missing, step A5 is executed; if the critical slot is not missing, step A6 is executed.
[0154] Step A5: Generate targeted supplementary questions.
[0155] Different missing slots correspond to different filler questions.
[0156] For example, if the location slot is missing, the output could be: "Where would you like to find photos taken?".
[0157] For example, if a time slot is missing, the output could be: "What time period would you like to search for photos taken in?".
[0158] Step A6: Calculate the confidence level of the retrieval intent information. and the confidence level of slot information .
[0159] Step A7: Determine whether the confidence level of the retrieval intent information is greater than the first confidence threshold, and whether the threshold of the slot information is greater than the second confidence threshold, i.e., determine... ,and .
[0160] If the confidence level of the above semantic elements meets the condition of step A7, then proceed to step A8; otherwise, proceed to step A9.
[0161] Step A8: Perform image retrieval.
[0162] It is understandable that image retrieval can be performed based on the slot information filled in the dialogue state object.
[0163] Step A9: Identify low-confidence semantic elements.
[0164] Step A10: Generate precise supplementary question statements based on low-confidence semantic elements.
[0165] For example, assuming the time slot "last weekend" is less than the second confidence threshold, and the user took photos on both Saturday and Sunday of last week, the output could be: "Do you mean last Saturday or last Sunday?". It's understandable that the image retrieval device can first perform a rough search in the first image set using the "time" slot to arrive at the intermediate conclusion that "photos were taken on both Saturday and last Sunday."
[0166] Step A11: Receive the user's input of a follow-up question response statement.
[0167] For example, a user can enter a follow-up question reply: "It was last Sunday, and I took this photo in First Park."
[0168] After executing step A11, step A1 can be re-executed until all semantic elements in the dialogue state object meet the corresponding confidence threshold.
[0169] Thus, since the generation of precise supplementary questions can be driven by the confidence level of semantic elements, the accuracy of supplementary questions can be improved, the efficiency of interaction can be increased, and the number of interaction rounds can be reduced.
[0170] In some embodiments of this application, the image retrieval results may include at least one image in a first image set that is related to slot information in the dialogue state object.
[0171] In some embodiments of this application, the image retrieval results may further include at least one of the following:
[0172] A natural language description of the image content of at least one of the above images;
[0173] A summary of at least one of the aforementioned images, which can describe the association between the at least one image and the retrieval intent information;
[0174] The thumbnail of at least one of the above images; the index of the thumbnail image, which is used to open the image.
[0175] The thumbnail of at least one of the above images; the thumbnail image's shortcut operation identifier, such as a share identifier, a delete identifier, etc.
[0176] In some embodiments of this application, the aforementioned shortcut operation identifier includes at least one of the following:
[0177] Image thumbnail cards, "Show image X" operation icon, "Add these images to album" operation icon, share icon to the contact corresponding to the person in the photo, and delete image icon.
[0178] In some embodiments of this application, the image retrieval device can generate a natural language description or summary of the image based on a locally configured lightweight natural language generation device.
[0179] In some embodiments of this application, the image retrieval device can generate the final image retrieval result through a language server. For example, it can encrypt and send the vectors of images obtained by image retrieval from the first image set based on the slot information in the updated dialogue state object, along with all the sentences input by the user, to the language server, so that the language server can generate the final image retrieval result, such as generating a natural language response. Of course, the image retrieval device can also generate image retrieval results based on a lightweight natural language model on the terminal side.
[0180] In some embodiments of this application, step 100 may include steps 100A to 100C.
[0181] Step 100A: The image retrieval device retrieves multiple joint quantization vectors from the quantization vector index of the first image set based on the feature vector associated with the first semantic element.
[0182] The aforementioned quantization vector index is constructed from the joint quantization vector of each image in the first image set. The joint quantization vector of the images is generated based on the modal feature quantization vector of different modal features of the images.
[0183] Step 100B: The image retrieval device clusters multiple joint quantized vectors according to the element category corresponding to the first semantic element to obtain the vector clustering result.
[0184] Step 100C: The image retrieval device outputs supplementary questions based on the vector clustering results.
[0185] In some embodiments of this application, the feature vector associated with the first semantic element can be: a vector obtained by encoding the range features corresponding to the first semantic element.
[0186] For example, the first semantic element is "last summer". Since this semantic is relatively vague, it can be converted into a specific time range: June 1 to September 30. The vector obtained by encoding this time range can be used as the feature vector associated with "last summer".
[0187] In some embodiments of this application, the element category corresponding to the semantic element may include: person, location, time, object, scene, activity, emotion, style, brand, event, weather, lighting, composition, texture, etc. It is understood that if the semantic element is slot information, then the element category of the semantic element is the same as the category of the slot information.
[0188] For example, in scenario 1 above, the time slot "last summer" does not meet the confidence condition, and the user's album includes photos taken in Qingdao last June and in Hainan last August. Therefore, a follow-up question can be generated: "Okay, regarding 'last summer,' do you mean the trip we took to Qingdao in August? Or the photos taken in Hainan in June?" It can be seen that the follow-up question includes clustering results based on the time category of "last summer," clustering the photos taken last summer in the user's album according to time. This makes it easier for the user to remember and clarify the specific time range of the photos they are looking for. In essence, the follow-up question, based on the vector (or vector component) clustering results of the time metadata in the quantized vector index, presents two specific, time-separated travel events for the user to choose from.
[0189] In some embodiments of this application, the modal feature quantization vector of an image is a vector obtained by quantizing the modal feature vector of the image.
[0190] In some embodiments of this application, the modal feature vector can be an embedding vector of modal features.
[0191] In some embodiments of this application, the feature vector associated with semantic elements or slot information can be a vector of features associated with semantic elements or slot information, such as the embedding vector of features associated with semantic elements or slot information.
[0192] In some embodiments of this application, the above embodiments use the element category of the first semantic element for clustering. In actual implementation,
[0193] Thus, since the vectorized index of the first image set can be searched based on the feature vector associated with the first semantic element whose confidence level does not meet the condition, and the search results can be clustered, and supplementary questions can be output based on the clustering results, users can get a general understanding of the images related to the first semantic element based on the supplementary questions. This can guide users to say the correct semantic element, thereby facilitating the image retrieval device to quickly and accurately understand the user's search intent and extract the correct slot information. This can improve the efficiency of multi-pass interaction between the user and the image retrieval device and improve the query accuracy.
[0194] In some embodiments of this application, the first semantic element is slot information; step 100A may include step 100A1.
[0195] Step 100A1: The image retrieval device retrieves multiple joint quantization vectors from the quantization vector index of the first image set based on the feature vector associated with the first semantic element and the feature vector of the retrieval intent information in the first image statement.
[0196] In some embodiments of this application, the image retrieval device can obtain the feature vector associated with the first semantic element and the feature vector of the retrieval intent information respectively, and then perform joint processing on the two vectors to obtain a joint vector. Then, it can retrieve the image from the quantized vector index based on the joint vector to improve the retrieval recall rate.
[0197] In some embodiments of this application, retrieval intent information can assist the image retrieval device in quickly identifying the joint quantization vectors of images from the first image set that are related to the first semantic element and fall within the range of images the user wants to retrieve, thereby further improving the accuracy of the supplementary question. For example, information related to images that are related to the first semantic element but irrelevant to the user's intent can be filtered out, thereby improving the guidance accuracy of the supplementary question.
[0198] For example, referring to scenario 1 above, the confidence level of the time slot "last summer" is lower than the confidence threshold. Therefore, based on the feature vector of the time range corresponding to last summer, we can retrieve all images taken in the personal album last summer from the quantized vector index. Then, combined with the user's intent "beach trip," we can further filter the images from last summer, leaving images taken in Hainan in June last year and images taken in Qinghai in August last year. We can then classify these related images according to time and generate a follow-up question based on the clustering results: "Last summer," do you mean the trip we took to Qingdao in August? Or the photos taken in Hainan in June?
[0199] Of course, in actual implementation, a joint vector can be generated directly based on the feature vector of the retrieval intent information and the feature vector associated with the first semantic element, and then the retrieval can be further simplified by searching in the quantized vector index based on the joint vector.
[0200] Thus, since the user's search intent and purpose can be represented by the search intent information, the multiple joint quantization vectors retrieved in the quantization vector index using the feature vectors associated with the search intent information and the first semantic element are all related to the search intent and the first semantic element. Therefore, accurate supplementary questions can be generated based on these joint quantization vectors to improve the efficiency and accuracy of multi-round interactions.
[0201] In some embodiments of this application, step 102 may include steps 102A to 102D.
[0202] Step 102A: The image retrieval device retrieves a candidate joint quantization vector set from the quantization vector index of the first image set based on the feature vector of the slot information in the updated dialogue state object.
[0203] The quantization vector index is constructed from the joint quantization vector of each image in the first image set. The joint quantization vector of the images is generated based on the modal feature quantization vector of different modal features of the images.
[0204] Step 102B: The image retrieval device decouples each joint quantization vector in the above candidate joint quantization vector set to obtain the modal feature quantization vector corresponding to each joint quantization vector.
[0205] Step 102C: The image retrieval device performs weighted fusion processing on the modal feature quantization vectors corresponding to each joint quantization vector and the feature vectors associated with the corresponding slot information in the updated dialogue state object to obtain the score information of each joint quantization vector.
[0206] Step 102D: The image retrieval device outputs the image retrieval results based on the score information of each joint quantization vector and the candidate joint quantization vector set.
[0207] In some embodiments of this application, step 102A is used to perform coarse screening and contextual filtering on the first image set.
[0208] In some embodiments of this application, different modal features of an image may include, but are not limited to, at least two of the following: visual features, visual descriptive features, temporal metadata, location metadata, and tags.
[0209] In some embodiments of this application, time metadata and location metadata can be collectively referred to as spatiotemporal metadata, or spatiotemporal structured metadata.
[0210] In some embodiments of this application, image tags may include, but are not limited to: scene tags, event tags, person tags, location tags, time tags, etc.
[0211] For example, time tags can categorize the time metadata of an image into tags for day, month, quarter, or year.
[0212] For example, location tags can be labels obtained by classifying GPS information of an image into streets, towns, districts, counties, cities, and provinces.
[0213] In some embodiments of this application, the relationship between each modal feature quantization vector corresponding to each joint quantization vector and the slot information in the M slot information is any one of the following:
[0214] Each joint quantization vector corresponds to a visual feature quantization vector, which in turn corresponds to the visual slot information in the updated dialogue state object.
[0215] Each joint quantization vector corresponds to an abstract text feature quantization vector, which in turn corresponds to the abstract semantic slot information in the updated dialogue state object.
[0216] Each joint quantization vector corresponds to a time metadata quantization vector that is associated with the time slot information in the updated dialogue state object.
[0217] Each joint quantization vector corresponds to a location metadata quantization vector, which is in line with the location slot information in the updated dialogue state object.
[0218] Each joint quantization vector corresponds to an object-event feature quantization vector, which in turn corresponds to the object or event slot information in the updated dialogue state object.
[0219] For a description of the abstract semantic slot information, please refer to the relevant description in the above embodiments. To avoid repetition, it will not be repeated here.
[0220] In some embodiments of this application, the feature vector associated with slot information refers to the feature vector of the feature associated with slot information.
[0221] In some embodiments of this application, the feature associated with slot information can be any of the following:
[0222] The time range corresponding to the time slot information;
[0223] Location range corresponding to location slot information;
[0224] The coordinates of a reference point within the location range corresponding to the location slot information, such as the center location;
[0225] The reference time point within the time range corresponding to the time slot information, such as the center time point;
[0226] The time slot information indicates the inherent object-event features of the object-event clustered images in the first image set.
[0227] Specifically, the intrinsic object-event feature vectors corresponding to the intrinsic object-event features of each event-object cluster image in the first image set can be pre-set in the quantization vector index. Then, using event or object slot information, a retrieval (such as using an ANN) can be performed in the quantization vector index to obtain the intrinsic object-event feature vectors of the event-object cluster images indicated by the event or slot information.
[0228] In some embodiments of this application, the image retrieval results may include at least one of the following:
[0229] An image indicated by a joint quantization vector that satisfies the score information conditions;
[0230] The index of the thumbnail image of the image that satisfies the score information condition;
[0231] A description of an image indicated by a joint quantization vector that satisfies the score information conditions, such as a natural language description;
[0232] A summary of an image indicated by a joint quantization vector that satisfies the score information condition, such as a natural language summary;
[0233] The quick action identifier corresponding to the image indicated by the joint quantization vector that satisfies the score information conditions, such as the identifier of any possible quick action, such as share or delete.
[0234] Thus, since preliminary retrieval is performed based on the feature vector of slot information and the quantization vector index of the first image set, and the joint quantization vector obtained from the preliminary retrieval can be rearranged by weighted fusion of the quantization vectors of each modality feature corresponding to the joint quantization vector obtained from the preliminary retrieval with the feature vectors associated with the corresponding slot information, the retrieval accuracy and recall rate can be improved.
[0235] In some embodiments of this application, the quantization vector index is an ANN index. Step 102A may include steps 102A1 and 102A2.
[0236] Step 102A1: The image retrieval device performs ANN processing on the quantization vector index based on the feature vector of the abstract semantic slot information in the updated dialogue state object to obtain a coarsely screened joint quantization vector set.
[0237] Step 102A2: The image retrieval device performs context filtering on the coarse joint quantization vector set based on the feature vector of the specific semantic slot information in the updated dialogue state object to obtain the candidate joint quantization vector set.
[0238] In some embodiments of this application, step 102A1 can be used to coarsely screen the first image set. For example, an approximate nearest neighbor search is performed on the feature vector of the slot information in the updated dialogue state object and the joint quantization vector in the quantization vector index to quickly return a Top-K coarsely screened set containing hundreds of candidate photos, i.e., the coarsely screened joint quantization vector set. Then, step 102A2 is used to perform context filtering on the coarsely screened joint quantization vector set to narrow down the candidate set.
[0239] In some embodiments of this application, context filtering of the coarse-screened joint quantization vector set can also be referred to as "hard filtering".
[0240] For example, if the final dialogue state object includes time slots, location slots, or person slots (also known as a person whitelist), this information can be used to filter out joint quantization vectors in the coarse-screened joint quantization vector set that do not meet this information.
[0241] For example, if the final dialogue state object includes a time slot: "last August", then the candidate joint quantization vectors that were captured in August of last year can be filtered out.
[0242] For example, if the final dialogue state object includes a location slot: "Hainan", then joint quantization vectors that were clearly not filmed in Hainan can be filtered out.
[0243] For example, if the final dialogue state object includes the character slot "Li Fang", then joint quantization vectors that do not include images of Li Fang can be filtered out.
[0244] Thus, coarse screening of the quantized vector index using feature vectors of abstract semantic slot information can expand the search scope as much as possible, while contextual filtering of the coarse screening results based on feature vectors of specific semantic slot information can quickly and accurately narrow the search scope. The combined use of coarse screening and hard filtering can improve search accuracy.
[0245] In some embodiments of this application, step 102C may include steps 102C1 and 102C2.
[0246] Step 102C1: The image retrieval device calculates the similarity between the modal feature quantization vector corresponding to each joint quantization vector and the feature vector associated with the corresponding slot information in the updated dialogue state object, and obtains the similarity set corresponding to each joint quantization vector.
[0247] Step 102C2: The image retrieval device inputs the similarity set corresponding to each joint quantization vector into the nonlinear fusion model to perform nonlinear weighted fusion processing and outputs the score information of each joint quantization vector.
[0248] The parameters of the aforementioned nonlinear fusion model are used to characterize the weight distribution characteristics of the similarity set composed of similarities corresponding to different modal features of the image.
[0249] In some embodiments of this application, step 102C1 above may include at least one of the following steps:
[0250] Step 11: The image retrieval device calculates the cosine similarity between the visual feature quantization vector corresponding to each joint quantization vector and the feature vector associated with the visual slot information in the updated dialogue state object.
[0251] Step 12: The image retrieval device calculates the cosine similarity between the abstract text feature quantization vector corresponding to each joint quantization vector and the feature vector associated with the abstract semantic slot information in the updated dialogue state object.
[0252] Step 13: The image retrieval device calculates the temporal distance between the quantization vector of the associated temporal metadata corresponding to each joint quantization vector and the feature vector associated with the temporal slot information in the updated dialogue state object.
[0253] Step 14: The image retrieval device calculates the geographic distance between the location metadata quantization vector corresponding to each joint quantization vector and the feature vector associated with the location slot information in the updated dialogue state object.
[0254] Step 15: The image retrieval device calculates the object-event feature quantization vector corresponding to each joint quantization vector, and the object-event similarity between the feature vector associated with the object or event slot information in the updated dialogue state object.
[0255] In some embodiments of this application, the similarity between the quantized vector of the image's temporal metadata and the feature vector of the temporal slot information can be referred to as the image's normalized temporal distance (ΔT); correspondingly, the similarity between the quantized vector of the image's location metadata and the feature vector of the location slot information can be referred to as the image's normalized geographic distance (ΔG). The image's normalized temporal distance (ΔT) and normalized geographic distance (ΔG) are decisive factors in determining the relevance of the image to the retrieval intent information. By nonlinearly fusing these heterogeneous features with the vector similarity of high-dimensional visual / textual features, the retrieval recall rate can be improved.
[0256] For example, the normalized time distance (ΔT) of an image can be the value obtained after logarithmic or linear normalization of the difference between the feature vector of the EXIF timestamp of the image and the feature vector of the center time point corresponding to the time slot information.
[0257] For example, the normalized geographic distance (ΔG) of an image is the normalized value obtained by normalizing the distance between the GPS coordinates of the image capture and the coordinates of a reference point for location information (e.g., the great circle distance between the GPS coordinates of the image and the coordinates of the reference point).
[0258] In some embodiments of this application, object-event similarity can be referred to as object-event matching feature. Specifically, object-event similarity can be a Boolean value or a weighted score for object-event matching. For example, person-event similarity can be expressed as... .
[0259] In some embodiments of this application, the image retrieval device can perform nonlinear fusion of heterogeneous features such as visual similarity, text similarity, and similarity corresponding to spatiotemporal metadata to rearrange the joint quantization vectors in the candidate joint quantization vector set, thereby achieving retrieval accuracy.
[0260] In some embodiments of this application, the similarity set corresponding to the joint quantization vector can be referred to as the heterogeneous feature set corresponding to the joint quantization vector. The construction of this heterogeneous feature set ensures that both high-dimensional semantic information and low-dimensional structured information can be input into the nonlinear fusion model in a unified format, thereby reducing the difficulty and accuracy of nonlinear fusion of heterogeneous data.
[0261] To improve the accuracy of nonlinear fusion of multimodal heterogeneous data, modal gaps, and temporal / geographic metadata of images, a learning-to-rank (L2R) method can be used, employing a small neural network as the nonlinear fusion model.
[0262] In some embodiments of this application, the nonlinear fusion model can also be referred to as a nonlinear fusion network or a rearranger. This nonlinear fusion model takes the aforementioned heterogeneous feature set as input. It can be seen that the input feature set of the nonlinear fusion model includes high-dimensional vector similarity scores (visual and textual similarity) as well as low-dimensional, structured normalized temporal distance (ΔT) and normalized geographic distance (ΔG) and other heterogeneous metadata features. By learning a nonlinear function, the nonlinear fusion model can achieve more accurate ranking than traditional fixed-weighted fusion.
[0263] In some embodiments of this application, the nonlinear fusion model may include a small multilayer perceptron (MLP) or a feedforward network.
[0264] In some embodiments of this application, the MLP is trained using a heterogeneous feature set corresponding to the training image. This allows the MLP to learn the complex nonlinear relationships between these heterogeneous modal features. Training of the MLP can be terminated when it can output accurate score information corresponding to the heterogeneous feature set. Specifically, the heterogeneous feature set corresponding to the training image can be obtained based on the similarity between the joint quantization vector of the training image and the feature vector associated with the corresponding slot information in the preset training retrieval statement. In other words, based on the MLP in the nonlinear fusion network, weights can be adaptively set for the features in the heterogeneous feature set. Compared to the scheme of setting fixed weights for features in related technologies, the weight setting method of this application is more in line with the user's retrieval intent.
[0265] In some embodiments of this application, the training images can be images from a publicly available training image set, and the training retrieval statements can be retrieval statements collected from real users authorized by the user, or retrieval statements generated by the model (such as simulated retrieval scenarios).
[0266] It is understandable that by using a nonlinear fusion model, since the weights of different features can be adaptively determined based on the training data in a specific query scenario (e.g., involving time-sensitive queries), it is possible to assign higher scores to the joint quantization vectors most relevant to the user's search intent, thereby improving the accuracy of the reordering of joint quantization vectors in the candidate quantization vector set.
[0267] For example, nonlinear fusion models can assign higher weights to normalized time distance in time-sensitive retrieval.
[0268] For example, nonlinear fusion models can assign higher weights to normalized geographic distance in location-sensitive searches.
[0269] In some embodiments of this application, after the MLP determines the weights of the features in the heterogeneous feature set of the candidate joint quantization vector, the heterogeneous feature set can be nonlinearly fused based on the weights corresponding to the heterogeneous feature set to obtain the score information of the candidate joint quantization vector.
[0270] For example, taking a heterogeneous feature set of a candidate joint quantization vector, including visual similarity, label similarity, normalized temporal distance, and normalized geographic distance, as an example, the nonlinear fusion model can calculate the score information of the joint quantization vector according to the following formula:
[0271]
[0272] in, This represents the score information of the candidate joint quantization vector. This represents the visual similarity corresponding to the candidate joint quantization vectors. This represents the text similarity corresponding to the candidate joint quantization vector, such as label matching. This represents the time decay corresponding to the candidate joint quantization vector. This represents the geographical decay corresponding to the candidate joint quantization vector.
[0273] It should be noted that by using vector-based coarse screening, context filtering, and nonlinear fusion rearrangement of multimodal heterogeneous feature sets based on a nonlinear fusion model, the problem of integrating heterogeneous metadata is solved, thereby improving retrieval accuracy and recall.
[0274] Thus, by first calculating the similarity between the modal feature quantization vectors corresponding to each candidate joint quantization vector set and the feature vectors associated with the corresponding slot information, and then using a nonlinear fusion model to perform weighted fusion on the similarity corresponding to each joint quantization vector, it is possible to adaptively set weights for each similarity based on slot information, thereby accurately determining the score information of the joint quantization vectors in the candidate quantization vector set, which can improve the reordering accuracy and retrieval recall.
[0275] In some embodiments of this application, steps 102A to 102C described above can be performed by the retrieval and rearrangement module.
[0276] The complete process of steps 102A to 102C is briefly described below with reference to the accompanying drawings.
[0277] like Figure 4 As shown, the image retrieval device can input the slot information (hereinafter referred to as parsed slots) from the updated dialogue state object into the retrieval and rearrangement module, which then performs the following steps:
[0278] Step 21: Based on the parsed slots, the retrieval and rearrangement module performs coarse screening and context filtering of the local ANN index through the coarse screening module in the retrieval and rearrangement module to obtain the Top-K candidate set, which is the above-mentioned candidate joint quantized vector set.
[0279] Step 22: The retrieval and reordering module extracts heterogeneous feature sets from the joint quantization vectors in the Top-K candidate set through the feature extraction module. Each heterogeneous feature set includes visual similarity. Text similarity Time distance and geographical distance .
[0280] Step 23: The retrieval and reordering module inputs the heterogeneous feature set of each candidate joint quantization vector into the learning fusion network in the retrieval and reordering module. The weight of each feature in the heterogeneous feature set is determined by the multilayer perceptron (MLP) in the learning fusion network to obtain the weight set.
[0281] Step 24: Based on the heterogeneous feature set and corresponding weight set of each candidate joint quantization vector, perform weighted fusion through the nonlinear fusion module in the learning fusion network to obtain and output the score information of the candidate joint quantization vector set, and output the rearrangement result of the candidate joint quantization vector set based on the score information.
[0282] It can be understood that the rearrangement result is to arrange the candidate joint quantization vector set into a sequence according to the score information from lowest to highest. In this way, the first N candidate joint quantization vectors in this sequence can be used as candidate joint quantization vectors that satisfy the score information condition.
[0283] In some embodiments of this application, step 102D may include steps 102D1 to 102D4.
[0284] Step 102D1: The image retrieval device determines the retrieval context based on the score information of each joint quantization vector and the candidate joint quantization vector set.
[0285] The aforementioned retrieval context may include: a joint quantization vector that satisfies the score information condition, and a thumbnail image index of the image indicated by the joint quantization vector that satisfies the score information condition.
[0286] Step 102D2: The image retrieval device sends the first encrypted information to the large language server based on the short-term signature.
[0287] The first encrypted information may include the retrieval context and at least one of the following: all statements entered by the user and the updated dialogue state object.
[0288] Step 102D3: The image retrieval device receives the second encrypted information sent by the aforementioned large language server.
[0289] The second encrypted information may include at least one of the following: a natural language description or summary of the image indicated by the joint quantization vector that satisfies the score information condition, and a shortcut operation identifier of the image indicated by the joint quantization vector that satisfies the score information condition.
[0290] Step 102D4: The image retrieval device outputs the image indicated by the joint quantization vector that satisfies the score information conditions and the decrypted second encrypted information.
[0291] In some embodiments of this application, the joint quantization vectors in the first encrypted information that satisfy the score information conditions are arranged from high to low according to the score information, and can also be called serialized joint quantization vectors.
[0292] In some embodiments of this application, in terms of data security, the retrieved search context is securely transmitted to a language server, which may be referred to as a "large language server," through a short-term signature URL mechanism to generate natural responses.
[0293] In some embodiments of this application, since the retrieval context is desensitized metadata, the first encrypted information sent is encrypted, and the sending is based on a short-term signature, it is possible to protect user data security through multi-level privacy and security mechanisms.
[0294] It is understandable that the image retrieval device only transmits the context most relevant to the retrieval intent and which has been securely processed to the cloud LLM for natural response generation, thereby providing a coherent, efficient and secure album search experience.
[0295] For example, the top 5 photos with the highest scores in the candidate joint quantization vector set of an image retrieval device provide the context needed by the user. Therefore, to leverage the powerful generation capabilities of cloud-based LLM, the image retrieval device can generate short-term signed URLs to encrypt and temporarily transmit the descriptions, thumbnail links, and de-identified metadata (i.e., the joint quantization vectors of the images) of these 5 photos to the language server to prevent privacy breaches.
[0296] In some embodiments of this application, the LLM in the language server can receive user queries, session history, and serialized retrieval context as input to generate natural language responses, and its functions include at least one of the following:
[0297] (1) Content summary: Describe and summarize the retrieved photo set, for example, "I found three photos of you with Li Ming by Erhai Lake last summer".
[0298] (2) Interactive elements / shortcut icons: Generate interactive elements, such as cards containing photo thumbnails, or provide quick action links (such as “Show 3rd photo” or “Add these photos to album”), so that the response is not only text, but also an actionable dialog element.
[0299] For example, in conjunction with scenario 1, if the image retrieval device uploads: the joint quantization vector of 5 images arranged from highest to lowest score, the visual description information of the 5 images, and the user-inputted search query, then: Figure 6 As shown, the language server can generate a natural language response by combining semantic elements from the user's final dialogue state object and the retrieved context, using the generation and response module: "Found it. These 5 photos perfectly match your description. The first one shows you standing on the rocks with the sunset right on the horizon. I can also generate a poetic title for you, or share this set of photos directly with Li Fang." Then, the language server can encrypt the natural language response and the clickable thumbnail cards of the 5 photos to obtain a second encrypted message, which is then sent to the image retrieval device. Upon receiving the second encrypted message, the image retrieval device can decrypt it and display the image.
[0300] It is understandable that image retrieval devices use short-term signed URLs as a secure transmission mechanism. Only when a language server is needed, such as cloud-based RAG generation, the authorized language server, such as LLM, can temporarily, temporarily, and encryptedly query the context of the first few retrieved image results (such as descriptions, thumbnail links, slot information), thus avoiding the continuous exposure of long-term credentials or sensitive data in the cloud.
[0301] In some embodiments of this application, the image retrieval device can... Figure 5The privacy and security layer 204 in the image retrieval system 200 shown enables data encryption and decryption to prevent the leakage of privacy information.
[0302] It should be noted that, as Figure 5 As shown, the privacy and security layer 204 is deployed on the edge as a gatekeeper for the data flow, responsible for local data encryption, access control, maintaining audit logs, and generating short-term signed URLs to securely send the first few retrieved results.
[0303] In some embodiments of this application, the privacy and security layer can be a hardware structure or a software algorithm, and this application does not limit the embodiments.
[0304] In some embodiments of this application, Figure 5 The image retrieval system 200 shown can provide a clear informed consent and revocable authorization mechanism (regarding data upload, model training, and sharing), and maintains detailed audit logs to record all data access and transmission activities. Meanwhile, Figure 5 The image retrieval system shown supports a "forget me" function, allowing users to clear local indexes and residual data in the cloud.
[0305] In this embodiment of the application, the image retrieval device may adhere to the principle of local priority: the original image, high-dimensional embedding vector, and all sensitive EXIF metadata are stored encrypted in the electronic device by default.
[0306] For example, the aforementioned quantized vector index, the first image set, the metadata record of the first image set, and the text feature information of the images in the first image set are encrypted using a key or a secure storage area, and can only be accessed by authorized local processes.
[0307] In some embodiments of this application, the image retrieval device may initiate a secure transmission mechanism if and only if the user authorizes and requires cloud-based RAG services, specifically including:
[0308] (1) Application of Short-Term Signed URLs: To pass the security context (e.g., high-resolution thumbnails or detailed descriptions) of the retrieved Top-K results to the cloud-based LLM, the system uses Short-Term Signed URLs. (2) Mechanism Principle: The signed URL contains authentication information and a time-limited signature. Anyone can only temporarily and encrypted access (i.e., view) the resource within the specified validity period. This resolves the contradiction that the cloud-based LLM (i.e., language model) needs to access context information for generation, while avoiding the contradiction of persistent storage or long-term exposure of sensitive data in the cloud. Once the URL expires, even if the URL is leaked, user data can no longer be accessed.
[0309] For a description of the shortcut operation icon, please refer to the relevant description of the shortcut operation icon in the above embodiments.
[0310] The interaction process between the image retrieval device and the language server will be briefly described below with reference to the accompanying drawings.
[0311] For example, after obtaining the score information of the candidate joint quantization vector, the image retrieval device, such as Figure 6 As shown, the image retrieval method provided in this application embodiment may include the following steps:
[0312] Step B1: The local device (i.e., the image retrieval device) generates a retrieval context.
[0313] Step B2: The image retrieval device sends the encrypted retrieval context to the cloud LLM (i.e., language server) based on the short-term signature URL through the privacy and security layer.
[0314] In some embodiments of this application, the cloud-based LLM can receive encrypted search context sent by the image retrieval device.
[0315] Step B3: The cloud-based LLM decrypts the encrypted retrieval context and obtains an encrypted response (i.e., the second encrypted information) based on the decryption result.
[0316] In some embodiments of this application, the image retrieval device can receive encrypted response information that can be sent by the cloud-based LLM. After decrypting the encrypted response information, the image retrieval results can be displayed, such as a natural language response.
[0317] Step B4: The cloud-based LLM automatically destroys the encrypted search context after the URL expires.
[0318] Thus, since the retrieval context of the joint quantization vector that satisfies the score information conditions, the thumbnail image index of the image indicated by the joint quantization vector that satisfies the score information conditions, and at least one of the following: all statements entered by the user, the updated dialogue state object mentioned above, and user intent information are encrypted and sent to the language server, the leakage of user privacy information can be avoided, and a more natural and high-quality natural language response can be obtained by utilizing the language server. Therefore, the coherence, efficiency, and security of the retrieval can be improved, thereby enhancing the retrieval experience.
[0319] In the image retrieval method provided in this application embodiment, since a supplementary question can be generated specifically based on the semantic elements in the image retrieval statement input by the user that do not meet the confidence condition, the user can be accurately guided to input a supplementary question reply statement that includes clear and accurate semantic elements. Therefore, the interaction efficiency of image retrieval based on natural language can be improved, making it easier for electronic devices to quickly and accurately understand the user's true retrieval intent, thereby improving image retrieval efficiency.
[0320] In some embodiments of this application, the image retrieval method provided in this application may further include the following steps 103 to 106.
[0321] Step 103: The modal features of each image in the first image set of the image retrieval device are input into the multimodal index module for encoding to obtain the modal feature vector set corresponding to the first image set.
[0322] Step 104: The image retrieval device quantizes the modal feature vectors in the modal feature vector set to obtain the modal feature quantization vector set corresponding to the first image set.
[0323] Step 105: The image retrieval device performs joint processing on the modal feature quantization vectors corresponding to the same image in the above modal feature quantization vector set to obtain the joint quantization vector of each image in the above first image set.
[0324] Step 106: The image retrieval device constructs the quantization vector index, which uses the joint quantization vector of each image in the first image set as the index entry, based on the joint quantization vector of each image in the first image set.
[0325] In some embodiments of this application, the image retrieval device encodes the visual features of the image through an image encoder and encodes the remaining features of the image through a text encoder.
[0326] Specifically, (1) Image encoder: Using a lightweight vision model (such as Lightweight General-Purpose Vision Transformer (MobileViT)) or Efficient Network (EfficientNet) as the backbone network, in conjunction with a projection head, the original image is encoded into a fixed-dimensional visual embedding vector. To ensure efficient execution on mobile devices, the encoder uses an inference optimization framework for performance acceleration during deployment. For example, Open Neural Network Exchange (ONNX), Neural Networks API (NNAPI), or Metal graphics and computation API.
[0327] (2) Text / Metadata Encoder: Encodes the text information associated with the image (including automatically generated tags, user-generated annotations, and file names) and structured metadata (EXIF timestamps, GPS coordinates) into text embedding vectors. This metadata is high-value structured information unique to photo album retrieval.
[0328] In some embodiments of this application, the vector index is optimized considering the limitations of memory and computing resources in electronic devices.
[0329] 1) Vector quantization: Vector quantization techniques (such as product quantization, PQ) are used to compress high-dimensional embedded vectors, significantly reducing the size of the index file.
[0330] Specifically, the image retrieval device can employ a vision-language (VL) model based on a dual-tower or unidirectional coding architecture, such as drawing inspiration from the CLIP framework, to generate cross-modal consistent embedding vectors for images.
[0331] 2) Efficient ANN Structure: Utilizes a memory-efficient Approximate Nearest Neighbor (ANN) index structure, such as Hierarchical Navigable Small World (HNSW) based on quantized embeddings. Quantized HNSW can provide fast, low-latency coarse-grained retrieval capabilities on mobile devices, ensuring performance even with personal photo albums of tens of millions of images. For example, coarse-grained retrieval can be implemented through an AI similarity search engine (Faiss on-device) or a non-metric space library (nmslib).
[0332] 3) Construction of the joint retrieval vector: Image embeddings and text / metadata embeddings can be concatenated or mapped to a common joint retrieval vector space through a learned projection layer. This joint vector serves as an index entry, supporting subsequent cross-modal queries.
[0333] Thus, since a quantization vector index with the joint quantization vector of the images as the index entries can be constructed in advance for the first image set, when the image retrieval statement input by the user is received, the vector corresponding to the slot information extracted from the image retrieval statement can be directly used to perform the retrieval in the quantization vector, without having to retrieve the original data of the first image set again, thus achieving low-latency retrieval.
[0334] Furthermore, since the quantization vector index is constructed using the joint quantization vector of the image as the index entry, the space occupied by the quantization vector index can be greatly reduced, thereby saving storage space.
[0335] In some embodiments of this application, such as Figure 6 As shown, the image retrieval device can also iteratively improve the various modules deployed on the terminal side of the image retrieval system through federated learning and / or differential privacy technology, ensuring that the original data and privacy information of individual users are strictly protected when using multi-user data for model optimization.
[0336] For example, the parameters of the multimodal index module and the parameters of the dialogue management module can be updated.
[0337] Specifically, in order to continuously improve the visual encoder and dialogue policy network without sacrificing user privacy, the image retrieval device employs privacy-preserving computation techniques, such as... Figure 6 As shown, it specifically includes at least one of the following:
[0338] (1) FL-based module parameter updates: Allows modules to be trained on distributed user devices and aggregates only the locally trained model parameter updates (rather than the original data) to the cloud server. The FL framework ensures that the model is improved on diverse data while protecting the locality of the data.
[0339] (2) Data transmission method based on DP: When aggregating or analyzing statistical data, carefully calibrated noise is injected into the model parameter updates or statistical results to ensure that even the most powerful observer cannot infer the existence or absence of any individual user data from the aggregation results. This effectively resists reconstruction attacks and privacy leakage risks against the RAG system.
[0340] In some embodiments of this application, an image data storage method is provided, which uses a dynamic image data table. For example, a table is generated based on the relevant information of the extracted image. If there is scenery in the landscape photo, then the "landscape" label is filled in. However, if there are no people, the table items related to people are deleted, thereby reducing redundant information and false information. As shown in Table 4, the storage space occupied by the image and related data in this solution is greatly reduced.
[0341] Table 2 shows the performance improvement of semantic parsing capability in the image retrieval method provided in this application embodiment:
[0342] Table 2
[0343] Search scenarios Traditional solution This application Increase Specific object query 78% 92% +14% Abstract concept query 32% 85% +53% Multi-condition compound query 41% 89% +48% Cross-year event search 29% 76% +47%
[0344] As shown in Table 2, in the scenario of explicit object query, the retrieval accuracy of the traditional solution is 78%, while that of this application is 92%, an improvement of 14%; in the scenario of abstract concept query, the retrieval accuracy of the traditional solution is 32%, while that of this application is 85%, an improvement of 53%; in the scenario of multi-condition compound query, in the test set (500 complex queries), the semantic understanding accuracy improved from 41% of the traditional system to 89%, an improvement of +48%; in the scenario of cross-year event query, the retrieval accuracy of the traditional solution is 29%, while that of this application is 76%, an improvement of 47%, as shown in Table 2.
[0345] The image retrieval method provided in this application uses a supplementary question mechanism, namely an intelligent supplementary question mechanism driven by the confidence level of semantic elements. This mechanism reduces the average number of interaction rounds from 3.2 rounds in the traditional approach to 1.5 rounds, and improves the fuzzy query resolution rate by 68%. The test dataset consists of 1000 fuzzy queries.
[0346] The image retrieval method provided in this application embodiment has significantly improved retrieval time and album index construction time, and the improvement data is shown in Table 3.
[0347] Table 3
[0348] Operation type Traditional solutions are time-consuming This application took a long time. acceleration ratio Initial search (cold start) 1200-2500ms 80-150ms 15X Multi-round refined search 800-1500ms 50-100ms 12X Add photo index Full reconstruction required Incremental update (<100ms / image)
[0349] As can be seen from Table 3, this application uses vectors for retrieval throughout the process, which eliminates the need to search the original image set, thus significantly reducing the time spent on the first retrieval, the time spent on multiple rounds of retrieval, the time spent updating the album index, and the index size.
[0350] Specifically, for the initial retrieval scenario, the traditional solution takes 1200~2500ms, while the solution in this application takes only 80~150ms, achieving a speedup of 15 times; for multi-round refined retrieval, the traditional solution takes 800~1500ms, while the solution in this application takes only 50~100ms, achieving a speedup of 12 times; and for updating the album index, related technologies require full vector construction, while this application only needs to update a small portion of the vector index related to the added or removed images, with the index reconstruction time required for each changed image being <100ms / image.
[0351] The nonlinear fusion method in the image retrieval method provided in this application embodiment results in a significant improvement in retrieval accuracy. The test environment is as follows: a database of 100,000 photos, containing personal photos spanning 5 years (see Table 3).
[0352] The image retrieval method provided in this application embodiment improves the user's image retrieval experience through the edge-cloud integrated image retrieval system: the success rate of voice interaction is increased to 98% (compared to 72% in traditional systems), and the operation efficiency of visually impaired users is increased by 4.3 times (through voice broadcasting + simplified instructions).
[0353] The image retrieval method provided in this application improves the system performance of image retrieval, resulting in optimized response speed and significantly reduced resource consumption for locally deployed image retrieval algorithms on the terminal side. The optimization results are shown in Table 4.
[0354] Table 4
[0355] index Traditional solution This application Optimization effect Memory usage 2-3GB / 100,000 images 300-300MB / 100,000 images Reduced by 85% storage space 120% of the original data 35% of the original data Reduced by 65% Computing resources GPU server required Terminal side can run Cost reduction of 90%
[0356] It should be noted that, since this application uses an adaptive image data storage method, the storage space required for image-related data can be significantly reduced.
[0357] Specifically, related technologies store images and image data (such as image details) using fixed-format tables. For example, these tables include data items such as time, location, event, and scene. Taking a scene as an example, the table includes specific scene options such as people, animals, and landscapes. This results in many invalid, empty data items in the data table for a single image, which also require storage space, leading to a large storage requirement for the photo album. In contrast, this application uses an adaptive data table. For example, if image 1 is a landscape photo, the corresponding data table can be "time, location, landscape," without including invalid data items such as "event" or "people" that do not exist in the landscape photo. Therefore, the storage space required for a single image can be significantly reduced.
[0358] This application also provides an image retrieval system 200, such as... Figure 5 As shown, the image retrieval system 200 may include: a dialogue management module 201 and a retrieval and rearrangement module 203;
[0359] The dialogue management module 201 is used to generate a supplementary question based on the first semantic element when the session state object corresponding to the image retrieval statement input by the user includes a first semantic element that does not meet the confidence condition. The first semantic element is slot information or retrieval intent information.
[0360] Interactive presentation module 202 is used to output the supplementary question statement;
[0361] The dialogue management module 201 is also used to update the semantic elements in the dialogue state object according to the supplementary question and answer statement input by the user.
[0362] The retrieval and rearrangement module 203 is used to perform image retrieval on the first image set based on the slot information in the updated dialogue state object, and obtain image retrieval results.
[0363] In some embodiments of this application, the dialogue management module 201 is specifically used for:
[0364] Based on the feature vector associated with the first semantic element, multiple joint quantization vectors are retrieved from the quantization vector index of the first image set. The quantization vector index is constructed from the joint quantization vectors of each image in the first image set. The joint quantization vectors of the images are generated based on the modal feature quantization vectors of different modal features of the images.
[0365] Based on the element category corresponding to the first semantic element, the multiple joint quantized vectors are clustered to obtain the vector clustering result;
[0366] The supplementary question is generated based on the vector clustering results.
[0367] In some embodiments of this application, the dialogue management module 201 may include an input parsing module and a supplementary question implementation module. The input parsing module is used to parse the semantic elements in the user-inputted statement and determine the confidence level of the semantic elements, and update the semantic elements in the dialogue state object. The supplementary question implementation module can be used to output supplementary questions based on the confidence level of the semantic elements in the dialogue state object, and based on the semantic elements in the dialogue state object that do not meet the confidence level conditions.
[0368] In some embodiments of this application, the first semantic element is slot information; the dialogue management module 201 is specifically used to retrieve multiple joint quantization vectors from the quantization vector index of the first image set based on the feature vector associated with the first semantic element and the feature vector of the retrieval intent information in the first image statement.
[0369] In some embodiments of this application, the retrieval and rearrangement module 203 described above can be specifically used for:
[0370] Based on the feature vector of the updated dialogue state object, a candidate joint quantization vector set is retrieved from the quantization vector index of the first image set. The quantization vector index is constructed from the joint quantization vector of each image in the first image set. The joint quantization vector of the image is generated based on the modal feature quantization vector of different modal features of the image.
[0371] Decouple each joint quantization vector in the candidate joint quantization vector set to obtain the modal feature quantization vectors corresponding to each joint quantization vector;
[0372] The modal feature quantization vectors corresponding to each joint quantization vector and the feature vectors associated with the corresponding slot information in the updated dialogue state object are weighted and fused to obtain the score information of each joint quantization vector.
[0373] Based on the score information of each joint quantization vector and the candidate joint quantization vector set, an image retrieval result is output, which includes images indicated by joint quantization vectors that meet the score information conditions.
[0374] In some embodiments of this application, the quantization vector index is an ANN index;
[0375] The retrieval and rearrangement module 203 can be specifically used for:
[0376] Based on the feature vector of the abstract semantic slot information in the updated dialogue state object, ANN processing is performed on the quantization vector index to obtain a coarsely screened joint quantization vector set.
[0377] Based on the feature vector of the specific semantic slot information in the updated dialogue state object, the coarse joint quantization vector set is subjected to context filtering to obtain the candidate joint quantization vector set.
[0378] In some embodiments of this application, the retrieval and rearrangement module 203 described above can be specifically used for:
[0379] Calculate the similarity between each modal feature quantization vector corresponding to each joint quantization vector and the feature vector associated with the corresponding slot information in the updated dialogue state object, and obtain the similarity set corresponding to each joint quantization vector;
[0380] The similarity set corresponding to each joint quantization vector is input into the nonlinear fusion model to perform nonlinear weighted fusion processing, and the score information of each joint quantization vector is output.
[0381] The parameters of the nonlinear fusion model are used to characterize the weight distribution characteristics of the similarity set composed of the similarities corresponding to different modal features of the image.
[0382] In some embodiments of this application, the image retrieval system 200 described above may further include: a privacy and security layer 204 and a language server 205;
[0383] The retrieval and rearrangement module 203 is specifically used to determine the retrieval context based on the score information of each joint quantization vector and the candidate joint quantization vector set. The retrieval context includes: the joint quantization vector that satisfies the score information condition, and the thumbnail image index of the image indicated by the joint quantization vector that satisfies the score information condition.
[0384] The privacy and security layer 204 is used to send first encrypted information to the large language server 205 based on a short-term signature. The first encrypted information includes the retrieval context and at least one of the following: all statements entered by the user and the updated dialogue state object.
[0385] The large language server 205 is used to generate second encrypted information based on the first encrypted information and send the second encrypted information to the privacy and security layer 204. The second encrypted information includes at least one of the following: a natural language description or summary of the image indicated by the joint quantization vector that satisfies the score information condition, and a shortcut operation identifier of the image indicated by the joint quantization vector that satisfies the score information condition.
[0386] The privacy and security layer 204 is used to decrypt the second encrypted information and send the decrypted second encrypted information to the interactive presentation module 202.
[0387] The interactive presentation module 202 is also used to output the image indicated by the joint quantization vector that satisfies the score information conditions and the decrypted second encrypted information.
[0388] In some embodiments of this application, the image retrieval system 200 described above may further include a multimodal indexing module 206;
[0389] The multimodal indexing module 206 is used for:
[0390] Encode each modal feature of the images in the first image set to obtain the modal feature vector set corresponding to the first image set;
[0391] The modal feature vectors in the modal feature vector set are quantized to obtain the modal feature quantization vector set corresponding to the first image set;
[0392] The modal feature quantization vectors corresponding to the same image in the modal feature quantization vector set are jointly processed to obtain the joint quantization vector of each image in the first image set;
[0393] Based on the joint quantization vector of each image in the first image set, a quantization vector index is constructed, with the joint quantization vector of each image in the first image set as the index entry.
[0394] For other descriptions of the image retrieval system 200 and related effects, please refer to the corresponding descriptions in the above method embodiments.
[0395] In some embodiments of this application, the image retrieval system 200 may further include another multimodal indexing module 206 deployed in the language server 205. For ease of distinction, the multimodal indexing module 206 deployed on the terminal side is referred to as the terminal-side multimodal indexing module 206, and the multimodal indexing module 206 deployed on the server side is referred to as the cloud-side multimodal indexing module 206. The cloud-side multimodal indexing module 206 is used to generate a non-quantized, high-precision vector index for cloud photo albums to handle complex global queries, such as retrieving images from a user's cloud photo album.
[0396] In some embodiments of this application, the image retrieval system 200 may further include another privacy and security layer 204 deployed in a federated learning aggregation server. For ease of distinction, the privacy and security layer 204 deployed on the terminal side is referred to as the terminal-side privacy and security layer 204, and the privacy and security layer 204 deployed on the server side is referred to as the cloud-side privacy and security layer 204. The cloud-side multi-privacy and security layer 204 is used to aggregate model parameter updates transmitted from multiple user devices after differential privacy processing, in order to iteratively improve the encoder and dialogue policy model.
[0397] The following is combined with Figure 5 and Figure 6 This application provides a detailed description of the architecture of the image retrieval system 200 provided in the embodiments of this application, as well as the functions of each module in the system.
[0398] The image retrieval system 200 provided in some embodiments of this application adopts a distributed, hybrid deployment architecture. The necessity of this hybrid deployment strategy lies in balancing a low-latency user interaction experience, robust computing resource requirements, and strict privacy protection for sensitive personal data.
[0399] In some embodiments of this application, a secure and efficient operating environment is achieved by deploying the dialogue management module 201, multimodal indexing module 206, retrieval and rearrangement module 203, and privacy and security layer 204 in the image retrieval system 200 on the terminal side, and the generation and response module on the server side. This division of labor resolves the fundamental contradiction in personal album retrieval scenarios where users want high-quality, natural language responses from LLM (Local Mode Management) but also require the protection of local data. By leaving the retrieval (coarse screening, rearrangement) and decision-making (follow-up question strategy) locally, the system can quickly respond to users, transmitting only minimal, encrypted, and time-sensitive context data to the cloud for final response generation.
[0400] like Figure 5 As shown, the image retrieval system 200 is initiated by the input parsing module, and the retrieval is performed collaboratively by the dialogue management module 201, the coordinated retrieval and rearrangement module 203, and the generation / response module. The multimodal indexing module 206 serves as a knowledge base supporting the retrieval locally. All local data streams are strictly controlled by the privacy and security layer 204, and the final results are returned to the user through the interactive presentation module 202.
[0401] In some embodiments of this application, all modules deployed on the terminal side in the image retrieval system 200 can be collectively referred to as local / edge computing modules (Edge / Local Device Components). These modules are mainly deployed on user devices, ensuring the locality of user data and low-latency interaction, and are the basis for the system to implement a "local-first" privacy strategy.
[0402] In some embodiments of this application, such as Figure 5 As shown, the modules deployed on the terminal side in the image retrieval system 200 can be called cloud computing modules (Cloud Server Components).
[0403] The functions of each module deployed on the terminal side and the server side are described below.
[0404] I. Local / Edge Computing Module
[0405] 1. Input parsing module: Responsible for ASR / text input parsing, intent recognition, slot extraction, and calculating the confidence score of NLU results.
[0406] It is responsible for converting the user's natural language input into structured query components, namely semantic elements.
[0407] For example, the input parsing module performs speech and text processing: if the input is speech, the system first converts it into text using Automatic Speech Recognition (ASR) technology. The module then performs Natural Language Understanding (NLU) on the text, including intent recognition and slot extraction. For instance, for the query "Find those photos we took in Dali, Yunnan last year," the intent is recognized as a search term, and the slots include time: last year and location: Dali, Yunnan. (People slots are missing.)
[0408] For example, the input parsing module calculates a confidence score. To support subsequent dynamic dialogue strategies, this module assigns a confidence score to each identified intent and extracted key slot. This score reflects the NLU model's confidence in the parsing results. The confidence score can be calculated based on various factors, including lexical likelihood, language model prediction features, or determined by analyzing an N-best candidate list. For instance, for a key slot (such as a name or a specific date), its confidence score is a crucial input for determining whether a follow-up question is needed.
[0409] 2. Multimodal Indexing Module: Responsible for core data storage, data encryption, deployment of lightweight visual / text encoders, and memory-efficient quantized approximate nearest neighbor (ANN) vector indexing services (e.g., quantized HNSW). The multimodal indexing module is the foundation of local retrieval, responsible for efficiently converting multimodal data from the personal photo album, i.e., the first image set, into searchable vector representations.
[0410] In some embodiments of this application, the multimodal indexing module may include: a vectorization submodule, a vector index optimization and implementation module, and a joint retrieval vector construction module; wherein,
[0411] 2.1 Vectorization Submodule: This submodule employs a vision-language (VL) model based on a dual-tower or unidirectional encoding architecture, such as drawing inspiration from the CLIP framework, to generate cross-modal consistent embedding vectors. The vectorization submodule includes:
[0412] (1) Image Encoder: A lightweight vision model (such as Mobile ViT or Efficient Net) is used as the backbone network, in conjunction with a projection head, to encode the raw image into a fixed-dimensional visual embedding vector. To ensure efficient execution on mobile devices, the encoder is accelerated for performance using an inference optimization framework (such as ONNX, NNAPI or Metal) during deployment.
[0413] (2) Text / Metadata Encoder: Encodes the text information associated with the image (including automatically generated tags, user-generated annotations, and file names) and structured metadata (EXIF timestamps, GPS coordinates) into text embedding vectors. This metadata is high-value structured information unique to photo album retrieval.
[0414] (3) Vector Index Optimization and Implementation Module: Considering the limitations of memory and computing resources on user devices, this embodiment optimizes the vector index. The vector index optimization and implementation module can perform the following steps:
[0415] Vector quantization: Vector quantization techniques (such as product quantization, PQ) are used to compress high-dimensional embedded vectors, significantly reducing the size of the index file.
[0416] Create efficient ANN indexes: Use memory-efficient Approximate Nearest Neighbor (ANN) index structures, such as HNSW (Hierarchical Navigable Small World) based on quantized embeddings. Quantized HNSW can provide fast, low-latency coarse-screen retrieval capabilities on mobile devices (such as through Faisson-device or nmslib), ensuring performance even with personal photo albums of tens of millions of images.
[0417] (4) Joint Retrieval Vector Construction Module: Image embeddings and text / metadata embeddings can be concatenated by the joint retrieval vector construction module or mapped to a common joint retrieval vector space through the learned projection layer. This joint vector serves as an index entry, supporting subsequent cross-modal queries.
[0418] 3. Retrieval and Reordering Module: This module is responsible for performing ANN coarse-grained retrieval and context filtering using vector indexes. It is the core of achieving high-precision image retrieval, especially in solving the challenge of integrating heterogeneous metadata through its unique learning-based fusion reordering network.
[0419] To address the issues of heterogeneity in multimodal data, modal gaps, and the neglect of time / geographic metadata, the retrieval and reordering module in this application adopts a learning-to-rank (L2R) method, using a small neural network as the reorderer.
[0420] In some embodiments of this application, the retrieval and rearrangement module may include a coarse screening and context filtering module, a feature extraction module, a multilayer perceptron (MLP), and a nonlinear fusion module, wherein the MLP and the nonlinear fusion module may be referred to as a fusion network module.
[0421] (1) Coarse screening and context filtering module, used for:
[0422] (1.1) Coarse screening (ANN Retrieval): Perform an approximate nearest neighbor search (ANN) on the text embedding of the user query (from module 110) and the joint retrieval vector in the local quantized vector index (module 120) to quickly return a Top-K coarse screening set containing hundreds of candidate photos.
[0423] (1.2) Context filtering: Apply hard filtering based on session state to the coarse screening results. For example, if the dialogue management module has determined a specific time range or a whitelist of people, these constraints are applied to narrow down the candidate set.
[0424] (2) Heterogeneous Feature Extraction Module: This module extracts heterogeneous features for each candidate photo in the coarse screening set. For each photo in the coarse screening set, the module extracts a heterogeneous feature set containing multimodal and structured information, which is then input into the rearrangement network. These heterogeneous features include not only the similarity scores of high-dimensional vectors but also low-dimensional metadata features; specifically,
[0425] (2.1) Visual similarity score ( ): Cosine similarity between visual embedding and query text embedding.
[0426] (2.2) Text similarity score ( ): Cosine similarity between image label / description embedding and query text embedding.
[0427] (2.3) Normalized time distance (ΔT): The difference between the photo EXIF timestamp (vector) and the time specified by the user query (vector), after logarithmic or linear normalization.
[0428] (2.4) Normalized geographic distance (ΔG): The distance (e.g., great circle distance) between the GPS coordinates of the photo and the location specified by the user, after normalization.
[0429] (2.5) Person / Event Matching Features ): Represents a Boolean value or weighted score for matching key people or events.
[0430] It is understandable that the construction of the above heterogeneous feature set ensures that both high-dimensional semantic information and low-dimensional structured information can be input into the subsequent learning model in a unified format.
[0431] (3) Fusion Network (MLP Reranker): A small multilayer perceptron (MLP) or feedforward network is used as input to the heterogeneous feature set mentioned above. The MLP, as a fusion network, can learn the complex nonlinear relationships between these heterogeneous modal features and output an accurate final reranking score (S_Final). This learning-based fusion mechanism is superior to simple fixed-weight or ranking-based fusion strategies because it can adaptively determine, based on the training data, that time or geographic features should have higher weights in specific query scenarios (e.g., time-sensitive queries). Finally, the system accurately ranks the candidate set according to S_Final and outputs the most relevant Top-K retrieval result set.
[0432] 4. Input parsing module: Responsible for session state management and multi-turn follow-up question strategy decision-making (based on confidence threshold). Dynamically trigger clarification questions, i.e., follow-up questions.
[0433] For example, the input parsing module is responsible for maintaining the session state and dynamically adjusting the dialogue strategy based on the confidence level of NLU.
[0434] In some embodiments of this application, the input parsing module may include: a session state and slot tracking module (i.e., the input parsing module described above) and a dynamic question filling strategy implementation module (i.e., the question filling implementation module described above). Among them: (1) Session state and slot tracking module: used to maintain a session state recorder to track the user's historical queries, current intent, and successfully filled slots and their corresponding confidence scores.
[0435] (2) Dynamic Question Supplement Implementation Module. After receiving and parsing new user input, the question supplement implementation module can execute the following dynamic decision logic:
[0436] a. Threshold determination: The slot location confidence level... and intention confidence Compare with their respective preset confidence thresholds.
[0437] b, Follow-up question strategy execution:
[0438] If all and If the value is high enough, the system determines that the query intent is clear and passes the instruction to the retrieval module for execution (Go-Ahead).
[0439] If any critical slot exists < If a key slot is missing, the system will trigger a replacement strategy.
[0440] c. Generating Clarification Questions, or Follow-up Questions: Unlike traditional simple confirmations (such as "Are you sure?"), this system dynamically generates precise clarification or follow-up questions based on the dialogue state and low-confidence slots. For example, if the NLU has low confidence in the user's mention of "weekend," the system will generate: "Do you mean the photo taken last Saturday or last Sunday?" This allows for precise follow-up questions that guide the user to input a response containing the correct semantic elements, reducing interaction rounds and improving retrieval efficiency.
[0441] Thus, by using a confidence-based follow-up question mechanism and leveraging the inherent uncertainty of the NLU model to drive the dialogue process, inefficient or failed retrievals caused by initial misunderstandings are effectively avoided, while unnecessary confirmation rounds are reduced, significantly improving user experience and query success rate.
[0442] 5. Interactive Presentation Module: Responsible for controlling the output of the electronic device's user interface, chat window, voice broadcast, and interactive thumbnails.
[0443] 6. Privacy and Security Layer: Deployed on the edge, it acts as a gatekeeper for the data flow, responsible for local data encryption, access control, maintaining audit logs, and generating short-term signed URLs to securely authorize cloud access to retrieved Top-K results.
[0444] In some embodiments of this application, the privacy and security layer is used for at least one of the following:
[0445] (6.1) Local Priority and Data Encryption: Adhere to the default local priority principle: the original image, high-dimensional embedding vector, and all sensitive EXIF metadata are stored encrypted on the user device by default. Local vector indexes and metadata records are encrypted using device keys or secure storage areas, and can only be accessed by authorized local processes.
[0446] (6.2) Secure Transmission Mechanism: The system activates the secure transmission mechanism only when the user authorizes and requires the cloud-based RAG service. Specifically:
[0447] (6.21) Application of short-term signed URLs: In order to pass the security context of the retrieved Top-K results (e.g., high-resolution thumbnails or detailed descriptions) to the cloud LLM, the system uses short-term signed URLs.
[0448] (6.22) Mechanism Principle: The signed URL contains authentication information and a time-limited signature. Anyone can only temporarily and encryptedly access the resource within the specified validity period. This resolves the contradiction between cloud-based LLMs requiring access to context information for generation and persistent storage or long-term exposure of sensitive data in the cloud. Once the URL expires, even if the URL is leaked, user data can no longer be accessed.
[0449] (6.3) Privacy Protection in Model Iteration: In order to continuously improve the visual encoder and dialogue policy network without sacrificing user privacy, the system adopts advanced privacy computing technology:
[0450] (6.31) Federated Learning (FL): Allows models to be trained on distributed user devices and only updates to the locally trained model parameters (rather than the original data) are aggregated to the cloud server. The FL framework ensures that models are improved on diverse data while protecting data locality.
[0451] (6.32) Differential Privacy (DP): When aggregating or analyzing statistical data, carefully calibrated noise is injected into model parameter updates or statistical results to ensure that even the most powerful observer cannot infer the existence or absence of any individual user data from the aggregation results. This effectively resists reconstruction attacks and privacy breaches against RAG systems.
[0452] 4. User Control and Compliance
[0453] The system provides a clear informed consent and revocable authorization mechanism (regarding data upload, model training, and sharing), and maintains detailed audit logs to record all data access and transmission activities. Additionally, the system supports a "forget me" function, allowing users to clear local indexes and residual data in the cloud.
[0454] II. Cloud Computing Module (Cloud Server Components)
[0455] 1. Generation / Response Module (150): This module primarily contains a powerful but computationally expensive Large Language Model (LLM) RAG generator. This module receives a signed retrieval context transmitted through a secure channel and generates the final natural language response.
[0456] The generation / response module primarily leverages the powerful capabilities of cloud-based LLM, combining retrieved context information to generate high-quality, natural conversational responses. Specifically:
[0457] (1) Retrieval enhancement and context transfer.
[0458] (1.1) Serialization Context: The rearranged Top-K search result set (usually a small number of results, such as 5 to 10 photos) is serialized into a search context that LLM can process. This context includes de-identified descriptions of the photos, metadata fragments, and temporary access links to thumbnails.
[0459] (1.2) Secure transmission protocol: In order to strictly comply with the privacy protocol, this context is transmitted to the cloud LLM interface through the secure transmission mechanism provided by the privacy and security layer (170), that is, using a short-signed URL.
[0460] (2) Natural Language Response Generation: The LLM receives user queries, session history, and serialized retrieval context as input. Based on this information, the LLM generates natural language responses, including the following functions:
[0461] (2.2) Content summary: Describe and summarize the retrieved photo set, for example, "I found three photos of you with Li Ming by Erhai Lake last summer".
[0462] (2.2) Generation of interactive elements: Generate interactive elements, such as cards containing photo thumbnails, or provide quick action links (such as “Show 3rd photo” or “Add these photos to album”), so that the response is not only text, but also an actionable dialog element.
[0463] 2. Cloud-based multimodal index module: used to deploy non-quantized high-precision vector indexes and handle complex global queries.
[0464] 3. Cloud-based Privacy and Security Layer: Deployed in the federated learning aggregation server, this layer aggregates differentially privacy-processed model parameter updates from multiple user devices to iteratively improve the encoder and dialogue policy models. The privacy and security layer is crucial for ensuring compliance when processing personal photo album data.
[0465] In the image retrieval system provided in this application embodiment, since supplementary questions can be generated specifically based on semantic elements in the image retrieval statement input by the user that do not meet the confidence condition, the supplementary questions can accurately guide the user to input supplementary questions that include clear and accurate semantic elements. Therefore, the interaction efficiency of image retrieval based on natural language can be improved, making it easier for electronic devices to quickly and accurately understand the user's true retrieval intent, thereby improving image retrieval efficiency.
[0466] The image retrieval device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0467] The image retrieval device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0468] The image retrieval device provided in this application embodiment can achieve... Figure 2 The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.
[0469] Optionally, such as Figure 7 As shown, this application embodiment also provides an electronic device 600, including a processor 601 and a memory 602. The memory 602 stores a program or instructions that can run on the processor 601. When the program or instructions are executed by the processor 601, they implement the various steps of the above-described image retrieval method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0470] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0471] Figure 8 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0472] The electronic device 1500 includes, but is not limited to, components such as: radio frequency unit 1501, network module 1502, audio output unit 1503, input unit 1504, sensor 1505, display unit 1506, user input unit 1507, interface unit 1508, memory 1509, and processor 1510.
[0473] Those skilled in the art will understand that the electronic device 1500 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 8 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0474] The processor 1510 is configured to generate a supplementary question based on the first semantic element when the session state object corresponding to the image retrieval statement input by the user includes a first semantic element that does not meet the confidence condition. The first semantic element is slot information or retrieval intent information.
[0475] The audio output unit 1503 or the display unit 1506 is used to output the supplementary question statement;
[0476] The processor 1510 is used to update the semantic elements in the dialogue state object based on the user's input of a follow-up question response statement;
[0477] The processor 1510 is configured to perform image retrieval on the first image set based on the slot information in the updated dialogue state object, and obtain image retrieval results.
[0478] For the remaining description of the electronic device, please refer to the relevant description in the above system embodiments.
[0479] In the electronic device provided in this application embodiment, since a supplementary question can be generated specifically based on the semantic elements in the image retrieval statement input by the user that do not meet the confidence condition, the user can be accurately guided to input a supplementary question containing clear and accurate semantic elements. Therefore, the interaction efficiency of image retrieval based on natural language can be improved, making it easier for the electronic device to quickly and accurately understand the user's true retrieval intent, thereby improving image retrieval efficiency.
[0480] It should be understood that, in this embodiment, the input unit 1504 may include a graphics processing unit (GPU) 15041 and a microphone 15042. The GPU 15041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1506 may include a display panel 15061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1507 includes at least one of a touch panel 15071 and other input devices 15072. The touch panel 15071 is also called a touch screen. The touch panel 15071 may include a touch detection device and a touch controller. Other input devices 15072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0481] The memory 1509 can be used to store software programs and various data. The memory 1509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0482] Processor 1510 may include one or more processing units; optionally, processor 1510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1510.
[0483] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image retrieval method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0484] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0485] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image retrieval method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0486] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0487] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the image retrieval method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0488] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0489] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0490] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An image retrieval method, characterized in that, The method includes: If the dialog state object corresponding to the image retrieval statement entered by the user includes a first semantic element that does not meet the confidence condition, the supplementary question statement is output according to the first semantic element, where the first semantic element is slot information or retrieval intent information. Update the semantic elements in the dialogue state object based on the user's input of the follow-up question and answer statement; Based on the slot information in the updated dialogue state object, perform image retrieval on the first image set to obtain the image retrieval results.
2. The method according to claim 1, characterized in that, The supplementary question statement output based on the first semantic element includes: Based on the feature vector associated with the first semantic element, multiple joint quantization vectors are retrieved from the quantization vector index of the first image set. The quantization vector index is constructed from the joint quantization vectors of each image in the first image set. The joint quantization vectors of the images are generated based on the modal feature quantization vectors of different modal features of the images. Based on the element category corresponding to the first semantic element, the multiple joint quantized vectors are clustered to obtain the vector clustering result; Based on the vector clustering results, output the supplementary question statement.
3. The method according to claim 2, characterized in that, The first semantic element is slot information; The step of retrieving multiple joint quantization vectors from the quantization vector index of the first image set based on the feature vector associated with the first semantic element includes: Based on the feature vector associated with the first semantic element and the feature vector of the retrieval intent information in the first image statement, multiple joint quantization vectors are retrieved from the quantization vector index of the first image set.
4. The method according to claim 1, characterized in that, The step of performing image retrieval on the first image set based on the slot information in the updated dialogue state object to obtain image retrieval results includes: Based on the feature vector of the slot information in the updated dialogue state object, a candidate joint quantization vector set is retrieved from the quantization vector index of the first image set. The quantization vector index is constructed from the joint quantization vector of each image in the first image set. The joint quantization vector of the image is generated based on the modal feature quantization vector of different modal features of the image. Decouple each joint quantization vector in the candidate joint quantization vector set to obtain the modal feature quantization vectors corresponding to each joint quantization vector; The modal feature quantization vectors corresponding to each joint quantization vector and the feature vectors associated with the corresponding slot information in the updated dialogue state object are weighted and fused to obtain the score information of each joint quantization vector. Based on the score information of each joint quantization vector and the candidate joint quantization vector set, the image retrieval result is output.
5. The method according to claim 4, characterized in that, The quantized vector index is a quantized approximate nearest neighbor search ANN index; The step of performing image retrieval on the first image set based on the slot information in the updated dialogue state object to obtain image retrieval results includes: Based on the feature vector of the abstract semantic slot information in the updated dialogue state object, ANN processing is performed on the quantization vector index to obtain a coarsely screened joint quantization vector set. Based on the feature vector of the specific semantic slot information in the updated dialogue state object, the coarse joint quantization vector set is subjected to context filtering to obtain the candidate joint quantization vector set.
6. The method according to claim 4, characterized in that, The weighted fusion process is performed on the modal feature quantization vectors corresponding to each joint quantization vector and the feature vectors associated with the corresponding slot information in the updated dialogue state object to obtain the score information of each joint quantization vector, including: Calculate the similarity between each modal feature quantization vector corresponding to each joint quantization vector and the feature vector associated with the corresponding slot information in the updated dialogue state object, and obtain the similarity set corresponding to each joint quantization vector; The similarity set corresponding to each joint quantization vector is input into the nonlinear fusion model to perform nonlinear weighted fusion processing, and the score information of each joint quantization vector is output. The parameters of the nonlinear fusion model are used to characterize the weight distribution characteristics of the similarity set composed of the similarities corresponding to different modal features of the image.
7. The method according to claim 4, characterized in that, The step of outputting image retrieval results based on the score information of each joint quantization vector and the candidate joint quantization vector set includes: Based on the score information of each joint quantization vector and the candidate joint quantization vector set, a retrieval context is determined. The retrieval context includes: the joint quantization vector that satisfies the score information condition, and the thumbnail image index of the image indicated by the joint quantization vector that satisfies the score information condition. Based on the short signature, a first encrypted message is sent to the language server, the first encrypted message including the retrieval context and at least one of the following: all statements entered by the user, the updated dialogue state object; The system receives second encrypted information sent by the large language server, the second encrypted information including at least one of the following: a natural language description or summary of the image indicated by the joint quantization vector that satisfies the score information condition, and a shortcut operation identifier of the image indicated by the joint quantization vector that satisfies the score information condition; Output the image indicated by the joint quantization vector that satisfies the score information condition and the decrypted second encrypted information.
8. The method according to claim 2 or 4, characterized in that, The method further includes: The modal features of each image in the first image set are input into the multimodal indexing module for encoding to obtain the modal feature vector set corresponding to the first image set; The modal feature vectors in the modal feature vector set are quantized to obtain the modal feature quantization vector set corresponding to the first image set; The modal feature quantization vectors corresponding to the same image in the modal feature quantization vector set are jointly processed to obtain the joint quantization vector of each image in the first image set; Based on the joint quantization vector of each image in the first image set, a quantization vector index is constructed, with the joint quantization vector of each image in the first image set as the index entry.
9. An image retrieval system, characterized in that, include: The module includes a dialogue management module, an interactive presentation module, and a search and rearrangement module. The dialogue management module is used to generate a supplementary question based on the first semantic element when the session state object corresponding to the image retrieval statement input by the user includes a first semantic element that does not meet the confidence condition. The first semantic element is slot information or retrieval intent information. The interactive presentation module is used to output the supplementary question statement; The dialogue management module is used to update the semantic elements in the dialogue state object based on the user's input of supplementary question and answer statements; The retrieval and rearrangement module is used to perform image retrieval on the first image set based on the slot information in the updated dialogue state object, and obtain the image retrieval results.
10. The system according to claim 9, characterized in that, The dialogue management module is specifically used for: Based on the feature vector associated with the first semantic element, multiple joint quantization vectors are retrieved from the quantization vector index of the first image set. The quantization vector index is constructed from the joint quantization vectors of each image in the first image set. The joint quantization vectors of the images are generated based on the modal feature quantization vectors of different modal features of the images. Based on the element category corresponding to the first semantic element, the multiple joint quantized vectors are clustered to obtain the vector clustering result; The supplementary question is generated based on the vector clustering results.
11. The system according to claim 10, characterized in that, The first semantic element is slot information; The dialogue management module is specifically used to retrieve multiple joint quantization vectors from the quantization vector index of the first image set based on the feature vector associated with the first semantic element and the feature vector of the retrieval intent information in the first image statement.
12. The system according to claim 9, characterized in that, The retrieval and rearrangement module is specifically used for: Based on the feature vector in the updated dialogue state object, a candidate joint quantization vector set is retrieved from the quantization vector index of the first image set. The quantization vector index is constructed from the joint quantization vector of each image in the first image set. The joint quantization vector of the image is generated based on the modal feature quantization vector of different modal features of the image. Decouple each joint quantization vector in the candidate joint quantization vector set to obtain the modal feature quantization vectors corresponding to each joint quantization vector; The modal feature quantization vectors corresponding to each joint quantization vector and the feature vectors associated with the corresponding slot information in the updated dialogue state object are weighted and fused to obtain the score information of each joint quantization vector. Based on the score information of each joint quantization vector and the candidate joint quantization vector set, an image retrieval result is output, which includes images indicated by joint quantization vectors that meet the score information conditions.
13. The system according to claim 12, characterized in that, The system also includes: a privacy and security layer and a language server; The retrieval and rearrangement module is specifically used to determine the retrieval context based on the score information of each joint quantization vector and the candidate joint quantization vector set. The retrieval context includes: the joint quantization vector that satisfies the score information condition, and the thumbnail image index of the image indicated by the joint quantization vector that satisfies the score information condition. The privacy and security layer is used to send first encrypted information to the large language server based on a short-term signature. The first encrypted information includes the retrieval context and at least one of the following: all statements entered by the user, the M slot information, and the updated dialogue state object. The language server is configured to generate second encrypted information based on the first encrypted information and send the second encrypted information to the privacy and security layer. The second encrypted information includes at least one of the following: a natural language description or summary of the image indicated by the joint quantization vector that satisfies the score information condition, and a shortcut operation identifier of the image indicated by the joint quantization vector that satisfies the score information condition. The privacy and security layer is used to decrypt the second encrypted information and send the decrypted second encrypted information to the interactive presentation module. The interactive presentation module is also used to output the image indicated by the joint quantization vector that satisfies the score information conditions and the decrypted second encrypted information.
14. The system according to claim 10 or 12, characterized in that, The system also includes a multimodal indexing module; The multimodal indexing module is used for: Encode each modal feature of the images in the first image set to obtain the modal feature vector set corresponding to the first image set; The modal feature vectors in the modal feature vector set are quantized to obtain the modal feature quantization vector set corresponding to the first image set; The modal feature quantization vectors corresponding to the same image in the modal feature quantization vector set are jointly processed to obtain the joint quantization vector of each image in the first image set; Based on the joint quantization vector of each image in the first image set, a quantization vector index is constructed, with the joint quantization vector of each image in the first image set as the index entry.
15. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the image retrieval method as described in any one of claims 1-9.