Similar picture retrieval method

By constructing a hierarchical semantic-visual representation framework and a semantic-visual bidirectional mapping framework, the problem of insufficient term ambiguity and mapping adaptability in professional fields is solved, and high precision of professional image retrieval is achieved, and dynamic adjustment of different professional backgrounds and visual data is adapted.

CN120256660AActive Publication Date: 2025-07-04JIANGSU YUNLAN INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510733788.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

When handling professional field query, the existing similar image search technology lacks a comprehensive contextual understanding and field background knowledge of professional terms, resulting in inaccurate ambiguity analysis of terminology, and traditional semantic-visual mapping lacks adaptability, and cannot optimize mapping relationships based on actual image data, affecting the retrieval accuracy.

Method used

Build a hierarchical semantic-visual representation framework, adopt a professional term understanding system of context perception, generate context feature vectors and hierarchical weight vectors, establish a semantic-visual bidirectional mapping and optimization framework, and generate comprehensive similarity scores through multimodal similarity calculation and iterative optimization to achieve accurate retrieval.

Benefits of technology

It improves the accuracy of retrieval of similar pictures in professional fields, can correctly understand professional terms based on professional background, and adjusts semantic understanding through visual feedback, solving the problem of not being able to dynamically adjust semantic understanding in traditional methods, and improving retrieval accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256660A_ABST
    Figure CN120256660A_ABST
Patent Text Reader

Abstract

The invention provides a similar picture retrieval method, which comprises the following steps of: constructing a hierarchical semantic-visual representation framework, and decomposing query and images into three levels, namely a core concept, attribute modification and relation description; a context-aware professional term understanding system is achieved, the professional field background and purpose of query are analyzed, context feature vectors are generated, and term ambiguity is eliminated; establishing a semantic-vision bidirectional mapping and optimization framework, realizing forward mapping from semantics to vision and reverse mapping from vision to semantics, and generating an optimization mapping matrix and a candidate set through iterative optimization; and performing multi-modal similarity calculation, converting the context feature vector into a hierarchical weight vector, fusing the hierarchical weight vector with the multi-level similarity into a comprehensive similarity score, and sorting and outputting a retrieval result set according to the comprehensive similarity score. According to the method, the accuracy of similar picture retrieval in the professional field is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to information similarity comparison and retrieval, especially a method for retrieving similar images. Background Art

[0002] The technology of retrieving similar images is a core part of modern information processing systems and is widely used in fields such as paper plagiarism checking, medical diagnosis, security monitoring, e-commerce, etc. With the increasing requirements for the accuracy of image retrieval in professional fields, traditional retrieval methods based on visual features can no longer meet the needs. In some fields, users usually use queries containing a large number of professional terms to find similar images, which requires the system to not only understand visual content but also accurately grasp professional semantics, understand the query background and purpose, and achieve accurate retrieval of professional images.

[0003] Current similar image retrieval technologies are mainly divided into two categories: content-based image retrieval (CBIR) and semantic-based image retrieval. Content-based retrieval methods mainly use low-level visual features such as color, texture, and shape for matching. For example, the QBIC system proposed by Flickner et al. uses a global feature vector to represent image content. With the development of deep learning, feature extraction methods such as CNN and SIFT have greatly improved the expression ability of visual features. At the semantic level, researchers have proposed visual-semantic embedding models to narrow the semantic gap, such as the deep visual-semantic alignment model proposed by Karpathy et al., the hierarchical attention network developed by Yang et al., and recently emerging cross-modal retrieval methods such as large visual-language pre-training models like CLIP.

[0004] However, there are still several key problems to be solved in the existing technology. First, when dealing with queries in professional fields, the existing methods have serious deficiencies in understanding professional terms, lack domain background knowledge and context awareness capabilities, resulting in inaccurate parsing of term ambiguities. For example, in medical image retrieval, the same term may have completely different visual correspondence relationships under different diagnostic purposes, and existing systems cannot dynamically adjust their understanding according to the query purpose and domain background. Second, traditional semantic-visual mapping is mostly a one-way process, and visual features cannot feedback to adjust semantic understanding, lacking adaptability, resulting in the retrieval system being unable to optimize the mapping relationship according to the actual image data distribution. Summary of the Invention

[0005] Object of the Invention: To provide a method for retrieving similar images in order to solve at least one technical problem existing in the prior art.

[0006] Technical Solution: The method for retrieving similar images includes:

[0007] Receiving user query data and constructing a hierarchical semantic-visual representation framework, including hierarchical semantics and visual representation;

[0008] Based on the hierarchical semantic representation, a context-aware professional term understanding system is adopted to generate context feature vectors, hierarchical weight vectors, and context-enhanced query representations;

[0009] Based on the context-enhanced query representation and hierarchical visual representation, a semantic-visual bidirectional mapping and optimization framework is adopted to generate an optimized mapping matrix and an optimized candidate set, and combined with the context feature vectors, multi-modal similarity calculation is performed to obtain multi-level similarities;

[0010] The multi-level similarities and the hierarchical weight vectors are fused into a comprehensive similarity score, and based on this, the optimized candidate set is sorted to output the retrieval result set.

[0011] Beneficial effects: The present invention provides comprehensive context for term disambiguation, enabling the system to correctly understand professional terms according to the professional background; not only mapping from semantics to vision, but also adjusting semantic understanding from visual feedback, solving the problem that semantic understanding cannot be adjusted according to actual visual data in traditional one-way mapping; improving the accuracy of similar image retrieval in professional fields. Brief Description of the Drawings

[0012] Figure 1 It is a step flowchart of a method for retrieving similar images provided by an embodiment of the present application.

[0013] Figure 2 It is a step flowchart of generating context feature vectors, hierarchical weight vectors, and context-enhanced query representations provided by an embodiment of the present application.

[0014] Figure 3 It is a step flowchart of obtaining a list of professional terms provided by an embodiment of the present application.

[0015] Figure 4 It is a step flowchart of constructing a hierarchical semantic-visual representation framework provided by an embodiment of the present application. Detailed Embodiments

[0016] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0017] It should be noted that for clearly showing the step flow of this application, serial numbers are marked for each step in the specification. These serial numbers are only for the convenience of explanation and do not limit the execution order of the steps. In actual operation, according to the technical requirements of specific implementation scenarios, the steps can be executed in an order different from that shown in the specification, and in some cases, parallel processing between steps can also be achieved.

[0018] During the research process, it is found that the current systems mostly adopt holistic feature representation, lacking fine-grained modeling of different semantic levels such as objects, attributes, and relationships, and unable to perform accurate matching and similarity calculation for the hierarchical semantic structures in professional fields, which further limits the retrieval accuracy and applicability.

[0019] As Figure 1 shown, a method for retrieving similar pictures is proposed, including the following steps:

[0020] S1. Receive user query data and construct a hierarchical semantic-visual representation framework, including hierarchical semantics and visual representation;

[0021] Specifically, the user query data can be text, pictures, or a combination of both. Hierarchical semantics is the analysis of text information, which divides the content into different levels, such as main topics, sub-topics, detailed information, etc., enabling the system to more clearly understand the organizational structure of the information. Visual representation is the processing of pictures or visual content, which may include identifying different parts in the image and analyzing the relationships between them, making it easier to understand the content of the picture.

[0022] S2. Based on the hierarchical semantic representation, adopt a context-aware professional term understanding system to generate context feature vectors, hierarchical weight vectors, and context-enhanced query representations;

[0023] Specifically, a context-aware professional term understanding system means that the system not only recognizes individual words but also combines the context to understand professional terms. For example, "deep learning" may just be a concept in ordinary conversations, but in technical papers, it may involve specific algorithms or models. The system ensures accurate understanding of terms by analyzing the context.

[0024] S3. Based on the context-enhanced query representation and hierarchical visual representation, adopt a semantic-visual bidirectional mapping and optimization framework to generate an optimized mapping matrix and an optimized candidate set, and combine the context feature vectors to perform multi-modal similarity calculation to obtain multi-level similarities;

[0025] Specifically, semantic-visual bidirectional mapping means making text and images correlate with each other, enabling language to understand visual information and visual information to be mapped to language expressions. By continuously adjusting the matching method, the mapping effect is improved, making the conversion between text and images more accurate.

[0026] S4. Integrate the multi-level similarity and the hierarchical weight vector into a comprehensive similarity score, sort the optimized candidate set based on the comprehensive similarity score, and output the retrieval result set.

[0027] Specifically, by integrating the multi-level similarity and the hierarchical weight, calculate a comprehensive score to represent the matching degree between a certain candidate item and the query. The higher the comprehensive score, the stronger the relevance between the candidate item and the query. Sort all candidate items according to the comprehensive similarity score, and the candidate item with the highest relevance will be ranked at the front to ensure that the user can see the most matching results first.

[0028] As Figure 2 shown, according to one aspect of the present application, generating a context feature vector, a hierarchical weight vector, and a context-enhanced query representation includes:

[0029] Identify professional terms from the hierarchical semantic representation to obtain a list of professional terms; combine the list of professional terms with the hierarchical semantic representation, analyze the professional field background and purpose of the query, generate domain features, purpose features, and knowledge background vectors, and fuse them into a context feature vector;

[0030] Input the context feature vector into a weight generation network to dynamically generate a hierarchical weight vector according to the professional field characteristics and purpose of the query;

[0031] Based on the list of professional terms and the context feature vector, perform term ambiguity resolution to obtain a disambiguated semantic representation; fuse the disambiguated semantic representation with the context feature vector to generate a context-enhanced query representation.

[0032] As Figure 3 shown, according to one aspect of the present application, obtaining a list of professional terms includes:

[0033] Match the vocabulary in the hierarchical semantic representation with a pre-trained domain-specific professional term vocabulary to generate a preliminary candidate set of professional terms;

[0034] Perform domain adaptability verification on each term in the preliminary candidate set of professional terms, calculate the domain membership probability of the term, and generate a verified candidate set of professional terms;

[0035] For each term in the verified candidate set of professional terms, evaluate its relevance in the current query context based on semantic link relationships and sentence structures, calculate the context importance score, and filter and sort to obtain a list of professional terms.

[0036] According to one aspect of the present application, obtaining a context feature vector includes:

[0037] Based on the hierarchical semantic representation and the list of professional terms, perform the following steps:

[0038] Identify the set of professional fields to which the query belongs, calculate the probability distribution of attribution for each field, and generate a field feature vector;

[0039] Analyze the query purpose and consider the application scenario to generate a purpose feature vector;

[0040] Retrieve relevant professional background knowledge from a pre-built professional knowledge base, extract key concepts and relationships, and generate a knowledge background vector;

[0041] Integrate the field features, purpose features, and knowledge background vector into a context feature vector.

[0042] As Figure 4 shown, according to one aspect of the present application, construct a hierarchical semantic-visual representation framework, including:

[0043] Encode the user query data to obtain a basic query vector representation; split the basic query vector representation into a core concept representation, an attribute modification representation, and a relationship description representation, and integrate them into a hierarchical semantic representation;

[0044] Extract object-level visual features, attribute-level visual features, and relationship-level visual features from the images in the image library, and integrate them into a hierarchical visual representation;

[0045] Based on the hierarchical semantic representation and the hierarchical visual representation, establish a strict correspondence relationship to form a hierarchical semantic-visual representation framework.

[0046] According to one aspect of the present application, integrate it into a hierarchical semantic representation, including:

[0047] For the basic query vector representation, identify and verify the subject and key entities, extract the core concepts and assign weight values to generate a core concept representation;

[0048] For the basic query vector representation, identify the adjectives, state words, and characteristic descriptions that modify the core concepts, construct attribute triples, and generate an attribute modification representation;

[0049] For the basic query vector representation, identify the relational verbs, state changes, and temporal information between the core concepts, construct relational quadruples, and generate a relationship description representation;

[0050] Integrate the core concept, attribute modification, and relationship description representations into a hierarchical semantic representation.

[0051] According to one aspect of the present application, integrate it into a hierarchical visual representation, including:

[0052] Identify the objects to be inspected from the image, extract the feature vectors of each object to be inspected, and generate object-level visual features;

[0053] For the object to be inspected, identify the domain attributes, calculate the attribute confidence scores, construct attribute triples, and generate attribute-level visual features;

[0054] Based on the object to be inspected and its attributes, identify the spatial and functional relationships between the objects to be inspected, construct a relationship graph between the objects to be inspected, and generate relationship-level visual features;

[0055] Integrate the object-level visual features, attribute-level visual features, and relationship-level visual features into a hierarchical visual representation that strictly corresponds to the semantic hierarchy.

[0056] According to one aspect of the present application, generating an optimized mapping matrix and an optimized candidate set includes:

[0057] Convert the context-enhanced query representation into a visual query representation, generate a multi-level visual attention map, and adjust the multi-level visual attention map based on the context feature vector to guide visual feature matching and calculate the forward mapping similarity matrix;

[0058] Based on the forward mapping similarity matrix, select a candidate image set and analyze the distribution characteristics of its visual features, construct a visual-semantic bridging network, project the visual features back to the semantic space, identify semantic differences and dynamically adjust semantic understanding, and generate a reverse mapping similarity matrix;

[0059] Based on the forward and reverse mapping similarity matrices, perform two-way mapping optimization to generate an optimized mapping matrix, and obtain an optimized candidate set by screening according to the candidate image set.

[0060] Specifically, calculate the object-level similarity between the query object representation and the object-level visual features, the attribute-level similarity between the query attribute representation and the attribute-level visual features, and the relationship-level similarity between the query relationship representation and the relationship-level visual features;

[0061] Combine the object-level similarity, attribute-level similarity, and relationship-level similarity to generate an optimized mapping matrix, and obtain an optimized candidate set by screening according to the preliminary matching results.

[0062] According to one aspect of the present application, generating an optimized mapping matrix and an optimized candidate set includes:

[0063] Read the query object representation, query attribute representation, and query relationship representation corresponding to the visual features;

[0064] And correspondingly match with the object-level visual features, attribute-level visual features, and relationship-level visual features in the image library,

[0065] Calculate the object-level similarity, attribute-level similarity, and relationship-level similarity respectively; accordingly, generate an optimized mapping matrix, and obtain an optimized candidate set by screening according to the preliminary matching results.

[0066] According to one aspect of the present application, calculating the forward mapping similarity matrix includes:

[0067] Input the context-enhanced query representation into the semantic-visual conversion network to generate visual query representations at the object level, attribute level, and relationship level;

[0068] Based on the visual query representation, calculate the similarity distribution of the hierarchical visual features in the image library and generate multi-level visual attention maps;

[0069] Combine the context feature vector, adaptively adjust the multi-level visual attention maps and guide visual feature matching, and calculate the forward mapping similarity matrix.

[0070] According to one aspect of the present application, generating the reverse mapping similarity matrix includes:

[0071] Perform visual feature distribution analysis on the images in the candidate image set to identify the dominant visual patterns;

[0072] Based on the dominant visual patterns, construct a visual-semantic bridging network, project the visual features into the semantic space, and generate reverse semantic descriptions;

[0073] Compare the reverse semantic descriptions with the original query semantics to identify potential semantic differences; combine the semantic differences with the context feature vector, and dynamically adjust the parameters of the semantic understanding model to generate the reverse mapping similarity matrix.

[0074] According to one aspect of the present application, performing multi-modal similarity calculation and outputting the final retrieval result set includes:

[0075] Obtain the object-level similarity, attribute-level similarity, and relationship-level similarity from the optimized candidate set to form multi-level similarities;

[0076] Multiply the multi-level similarities by the hierarchical weight vector and sum them to generate the comprehensive similarity score for each candidate image;

[0077] Sort the images in the optimized candidate set in descending order according to the comprehensive similarity score, select the top N images with the highest scores as the final retrieval result set, and generate similarity explanation information for each result image.

[0078] According to one aspect of the present application, the steps of generating the comprehensive similarity score include:

[0079] Based on the context feature vector, analyze the focus of attention and importance distribution of the current query in the professional field;

[0080] Apply a dynamic weight generator to generate object-level weights, attribute-level weights, and relationship-level weights according to the focus of attention, importance distribution, and query purpose, and form a hierarchical weight vector;

[0081] Identify and smooth outliers and noise in the multi-level similarity to ensure the stability of the similarity distribution;

[0082] Adopt a weighted fusion algorithm to multiply and sum the smoothed multi-level similarity with the hierarchical weight vector, and at the same time consider the interactive effects between levels to generate a comprehensive similarity score.

[0083] The present invention provides a method for retrieving similar pictures, including: constructing a hierarchical semantic-visual representation framework, decomposing queries and images into three levels: core concepts, attribute modifiers, and relationship descriptions; implementing a context-aware professional term understanding system to analyze the professional field background and purpose of the query, generate context feature vectors, and resolve term ambiguities; establishing a semantic-visual bidirectional mapping and optimization framework to achieve forward mapping from semantics to vision and reverse mapping from vision to semantics, and generating an optimized mapping matrix and a candidate set through iterative optimization; performing multi-modal similarity calculation, converting the context feature vector into a hierarchical weight vector, fusing it with the multi-level similarity into a comprehensive similarity score, and sorting and outputting a retrieval result set accordingly. The present invention improves the accuracy of retrieving similar pictures in a professional field.

[0084] According to one aspect of the present application, a method for retrieving similar pictures is proposed, and the implementation process of its context-aware semantic-visual bidirectional mapping framework is specifically as follows:

[0085] S1: Construct a semantic-visual bimodal hierarchical representation framework. Receive the user's professional query text, split it into three levels: core concepts, attribute modifiers, and relationship descriptions through a hierarchical semantic decomposition algorithm, and at the same time apply a corresponding visual feature extraction network to the image library to establish a strict correspondence between the semantic level and the visual features, generating a hierarchical semantic representation and a hierarchical visual representation. Convert the traditional holistic feature representation into multi-level visual features that precisely correspond to the semantic structure, and solve the fundamental problem of unclear semantic-visual correspondence in the professional field.

[0086] S11: Receive the user's professional query text, and perform preliminary encoding using the Transformer encoder model to obtain a basic query vector representation.

[0087] S12: Apply a three-level structured decomposition algorithm to the basic query vector representation to split the query into three semantic levels:

[0088] S121: Core concept extraction. Read the basic query vector representation, apply the attention mechanism to identify the subject and key medical entities; use the medical ontology knowledge base to verify the identified entities and filter out non-core concepts; apply the semantic importance scoring algorithm to assign weight values to each identified core concept; output the core concept representation sorted by weight, including the entity vector and its associated weight.

[0089] S122: Attribute modification extraction. Read the basic query vector representation, identify adjectives, state words, and characteristic descriptions that modify the core concepts; apply dependency syntactic analysis to identify the association relationships between these modifiers and the core concepts; construct (attribute type, attribute value, associated core concept) triples for each modifier; output the integrated attribute modification representation, including the modifier vector and its association relationship with the core concept.

[0090] S123: Relationship description extraction. Read the basic query vector representation, identify relationship verbs, state changes, and temporal information between core concepts; apply semantic role labeling to identify the action subject, object, and their relationship types; construct (subject concept, relationship type, object concept, temporal information) quadruples; apply a specialized relationship recognition model for specific medical relationships (such as treatment effect, pathological reaction, etc.); output the integrated relationship description representation, including the relationship vector and the concepts it connects.

[0091] S13: Integrate the core concept representation, attribute modification representation, and relationship description representation into a hierarchical semantic representation, maintaining the association structure between levels.

[0092] S131: Construct a hierarchical association graph. Use the entities in the core concept representation as nodes to establish an initial concept graph; based on the attribute modification representation, add attribute edges and attribute values to each node; based on the relationship description representation, add relationship edges and relationship attributes between concept nodes; generate a complete semantic hierarchical association graph and save the vector representations of the nodes and edges.

[0093] S132: Optimize the compatibility of the hierarchical representation. Standardize the vector representations of the three levels to ensure dimension consistency; apply the hierarchical fusion algorithm to make the representations at different levels comparable in the feature space; construct an attention mechanism between levels so that high-level semantics can focus on low-level relevant features; output the final hierarchical semantic representation, including the vectors of the three levels and their association structure.

[0094] S14: For each medical image in the image library, apply the hierarchical visual feature extraction network to achieve a visual feature decomposition that strictly corresponds to the semantic hierarchy.

[0095] S141: Object-level feature extraction. Read the medical image, apply the Faster R-CNN model to identify medical objects (cells, tissues, organs, etc.) in the image; for each detected object, use a domain-specific feature extractor to extract feature vectors; apply a medical object recognition enhancement module to fuse texture, shape, and context information; apply multi-scale feature fusion to process medical targets of different sizes; output the position information and feature vectors of each object to form object-level visual features.

[0096] S142: Attribute-level feature extraction. Read each detected medical object region, apply an attribute classification network to identify medical attributes (such as cell morphology, tissue density, etc.); calculate a confidence score for each attribute and construct a (attribute type, attribute value, confidence) triple; apply a multi-attribute joint reasoning module to handle the correlation and mutual exclusion between attributes; combine global image features to enhance the accuracy of local attribute recognition; output the attribute representation vector of each object to form attribute-level visual features.

[0097] S143: Relationship-level feature extraction. Read the detected medical objects and their attributes, apply a relationship reasoning network to identify the spatial and functional relationships between objects; construct a relationship graph between objects to represent their spatial position relationships, containment relationships, and functional relationships; apply a temporal analysis module (for videos or multi-frame images) to identify state change relationships; apply medical prior knowledge to guide relationship reasoning and improve the accuracy of professional relationship recognition; output the representation vector of the relationships between objects to form relationship-level visual features.

[0098] S15: Organize the object-level visual features, attribute-level visual features, and relationship-level visual features of each image into a hierarchical visual representation and establish a visual feature index library.

[0099] S151: Construct an image hierarchy. Integrate the three-layer features of each image to construct a visual hierarchy corresponding to the semantic hierarchy; apply feature standardization to ensure the scale consistency of features at different levels; construct attention connections between levels so that high-level features can guide the focus of low-level features; output the image hierarchy of each image.

[0100] S152: Establish an efficient index structure. Apply the locality-sensitive hashing algorithm to construct a fast retrieval index for the hierarchical representations of all images; establish an index for each semantic level separately to support hierarchical retrieval; apply the inverted index technology to support fast lookup based on concepts and attributes; output a complete visual feature index library for subsequent retrieval.

[0101] S2: Implement a context-aware professional term understanding system. Process professional terms in hierarchical semantic representations. By analyzing the domain, purpose, and implicit knowledge background of the query, construct a context feature vector, and based on this, perform term ambiguity resolution to generate a context-enhanced query representation. The core lies in going beyond literal understanding, interpreting the query in a professional context, solving the key problem of the changing meanings of professional terms in different scenarios, and enhancing the system's ability to understand complex professional queries.

[0102] S21: Receive the hierarchical semantic representation, identify the professional terms therein, and obtain a list of professional terms.

[0103] S211: Professional term detection. Read all the words and phrases in the hierarchical semantic representation; apply a medical term recognition model and match it with a professional term knowledge base; use the context word co-occurrence feature to enhance the accuracy of term boundary recognition; output all the detected professional terms and their position information in the query.

[0104] S212: Polysemous term marking. Read the detected professional terms, query the professional term ambiguity library; mark the terms with multiple possible meanings and their possible meaning lists; calculate the ambiguity degree of each term and assign different processing priorities for subsequent resolution; output a list of professional terms with ambiguity marks, including the set of possible meanings for each term.

[0105] S22: Analyze the professional domain background and purpose of the query.

[0106] S221: Professional domain identification. Read the hierarchical semantic representation and the list of professional terms; extract the term co-occurrence feature and compare it with a pre-trained domain feature template; apply a domain classification model to calculate the probability distribution of the query belonging to each professional domain; apply hierarchical domain classification, from coarse-grained (such as "medicine") to fine-grained (such as "tumor pathology"); output the domain feature vector of the query, including the main domain and possible sub-domain information.

[0107] S222: Query purpose analysis. Read the hierarchical semantic representation, identify the intention indicator words in the query (such as "diagnosis", "analysis", "comparison"); analyze the syntactic structure and question type of the query to infer the query purpose; apply a query intention classification model to identify whether it is a diagnostic requirement, research analysis, teaching purpose, etc.; analyze the detail level and professional depth of the query to evaluate the user's professional level; output the purpose feature vector of the query, including the main purpose and additional purpose information.

[0108] S223: Knowledge background extraction. Read the hierarchical semantic representation and the list of professional terms; apply knowledge graph reasoning to identify the implicit prerequisite knowledge requirements of the query; analyze the professional degree of term usage and evaluate the depth of domain knowledge required; construct a knowledge dependency graph of the query to represent the knowledge structure required to understand the query; output the knowledge background vector of the query, representing the knowledge premise and background assumptions of the query.

[0109] S23: Integrate the domain feature vector, purpose feature vector, and knowledge background vector into a context feature vector.

[0110] S231: Feature vector standardization. Read the three feature vectors, apply dimension unification and range standardization processing; apply a feature selection algorithm to retain the most discriminative feature dimensions; output the three standardized feature vectors.

[0111] S232: Context feature fusion. Apply an attention-weighted fusion mechanism to dynamically adjust the weights of the three vectors according to the query characteristics; construct an interaction layer between features to capture the mutual influence of the domain, purpose, and knowledge background; apply context consistency checking to ensure the rationality of the fusion result; output the final context feature vector, containing complete context information.

[0112] S24: Perform term ambiguity resolution based on the list of professional terms and the context feature vector.

[0113] S241: Construct a term conditional probability model. Read the list of professional terms and the context feature vector; for each polysemous term, retrieve the usage frequency statistics in different contexts; construct a Bayesian conditional probability model to calculate P(meaning is the context feature); apply domain-specific prior knowledge to adjust the probability distribution; output the meaning probability distribution of each term in the current context.

[0114] S242: Context feature matching. Read the list of possible meanings of each polysemous term and the context feature vector; construct a feature representation for each possible meaning and calculate the similarity with the query context feature; apply hierarchical similarity calculation, considering domain similarity, purpose relevance, and knowledge consistency respectively; output the context matching scores of various meanings of each term.

[0115] S243: Term network collaborative disambiguation. Read all terms in the query and their meaning probability distributions and context matching scores; construct a semantic association network between terms to represent the dependency and mutual support relationships between terms; apply a graph reasoning algorithm to make semantically related terms tend to select semantically consistent meanings; apply disambiguation iterative optimization until the term meaning network reaches a stable state; output the optimal meaning selection and its confidence of each term.

[0116] S25: Integrate the disambiguated hierarchical semantic representation with the context feature vector to generate a context-enhanced query representation.

[0117] S251: Update the semantic representation. Read the original hierarchical semantic representation and the optimal meaning selection of terms; update the vector representations of terms in each level, replacing them with the professional representations of the selected meanings; adjust the representations of other concepts related to these terms to ensure semantic consistency; output the updated precise semantic representation.

[0118] S252: Context-enhanced fusion. Read the precise semantic representation and the context feature vector; apply context-conditioned feature transformation to adjust each dimension of the semantic representation according to the context; apply different context-enhanced strategies for different semantic levels to construct a hierarchical context representation; embed context information into the semantic structure to enhance the environmental adaptability of the query representation; output the final context-enhanced query representation for subsequent mapping.

[0119] S3: Establish a semantic-visual bidirectional mapping and optimization framework. Based on the context-enhanced query representation and the hierarchical visual representation, construct a two-way interaction mechanism between semantic and visual features. First, implement the forward mapping of semantic-guided visual feature activation, and then reverse-adjust semantic understanding based on the visual feature distribution. Through iterative optimization, generate an optimized mapping matrix and an optimized candidate set. The core lies in breaking the limitations of traditional one-way mapping, realizing the dynamic mutual feedback between semantic understanding and visual expression, enabling the system to adjust query understanding according to actual image data, and improving retrieval accuracy and adaptability.

[0120] S31: Receive the context-enhanced query representation and the hierarchical visual representation, and initialize the bidirectional mapping matrix.

[0121] S311: Hierarchical mapping initialization. Read the structural information of the context-enhanced query representation and the hierarchical visual representation; construct initial mapping matrices for the three semantic levels respectively to establish the initial correspondence between semantic-visual features; apply domain transfer learning to initialize the mapping weights using pre-trained medical semantic-visual correspondence knowledge; output the initial set of hierarchical mapping matrices, including three matrices: core concept mapping, attribute mapping, and relationship mapping.

[0122] S312: Establish the mapping association between levels. Read the set of hierarchical mapping matrices and construct a joint mapping mechanism between levels; construct a hierarchical mutual feedback channel to enable high-level semantic mapping to guide the attention allocation of low-level mapping; apply the overall consistency constraint to ensure the coordination of mapping results at different levels; output the complete bidirectional mapping matrix, including the intra-level mapping and the inter-level association structure.

[0123] S32: Implement the forward mapping from semantics to vision.

[0124] S321: Generate multi-level visual attention maps. Read the context-enhanced query representation and the bidirectional mapping matrix; for the core concept layer, apply mapping transformation to generate target region attention maps to highlight relevant medical objects; for the attribute modification layer, generate feature attention maps to highlight relevant visual attribute regions; for the relationship description layer, generate relationship attention maps to highlight the interaction regions between objects; output three-layer visual attention maps, representing the mapping of different semantic levels in the visual space.

[0125] S322: Adjust attention weights based on context. Read the visual attention maps and the context feature vectors; according to the query domain features, adjust the attention weights of different types of medical objects; according to the query purpose features, adjust the attention allocation of attribute features (e.g., more attention to abnormal features for diagnostic purposes); according to the knowledge background features, adjust the professional sensitivity of relationship attention; output the context-weighted attention maps, reflecting the attention allocation in the current professional context.

[0126] S323: Perform visual feature activation and retrieval. Read the context-weighted attention maps and the visual feature index library; apply attention-guided feature activation to assign activation intensities to each feature in the visual library; calculate the weighted feature similarities at three levels to generate preliminary hierarchical similarity scores; apply hierarchical weighted fusion to synthesize the similarities at three levels into an overall similarity; rank based on the overall similarity and select the top N most similar images; output the candidate image set and its initial similarity scores.

[0127] S33: Implement the reverse mapping from vision to semantics.

[0128] S331: Visual feature distribution analysis. Read the hierarchical visual features of the candidate image set; apply clustering analysis to identify the visual patterns and feature distribution rules in the candidate set; identify the high-frequency visual features, which may represent the core visual manifestations of the query; identify the differential visual features, which may represent the ambiguous manifestations of the query; construct a visual feature distribution map, representing the visual feature space distribution of the retrieval results.

[0129] S332: Feature space mapping analysis. Read the visual feature distribution map and the current bidirectional mapping matrix; analyze the mapping effect from semantic features to visual features, and identify the regions where the mapping is inaccurate or incomplete; detect the problem of semantic over-mapping (one semantic concept maps to too many irrelevant visual features); detect the problem of semantic under-mapping (the semantic concept fails to map to relevant visual features); output a mapping quality assessment report, including the mapping regions that need to be adjusted and the adjustment directions.

[0130] S333: Semantic Understanding Dynamic Adjustment. Read the visual feature distribution map and the mapping quality assessment report; Based on the visual feature distribution, dynamically adjust the importance weights of each concept in the hierarchical semantic representation; Increase the weight for semantic concepts with insufficient mapping and decrease the weight for over-mapped concepts; Introduce new semantic concepts to represent features found in visual features but not explicitly in the original query; Suppress semantic concepts not reflected in visual features and reduce their influence in the mapping; Output the adjusted adaptive semantic representation to better match the actual visual feature distribution.

[0131] S334: Update the Bidirectional Mapping Relationship. Read the adaptive semantic representation and the original bidirectional mapping matrix; Apply the gradient descent algorithm to optimize the mapping matrix parameters according to the semantics; Adjust the mapping weights at different semantic levels to reflect their actual importance in visual expression; Optimize the association structure between levels to enhance the overall consistency of the semantic-visual mapping; Output the updated bidirectional mapping matrix to reflect the impact of visual feedback on semantic understanding.

[0132] S34: Execute the Iterative Optimization Loop.

[0133] S341: Recalculate the Semantic-Visual Similarity. Read the updated bidirectional mapping matrix and the adaptive semantic representation; Repeat step S323 to calculate the new similarity score using the updated mapping; Compare with the initial similarity score to evaluate the optimization effect; Output the optimized similarity score and the updated candidate image set.

[0134] S342: Evaluate the Optimization Effect. Calculate the change in the similarity distribution before and after optimization; Analyze the update degree of the candidate set, calculate the set similarity and ranking changes; Evaluate the convergence stability of the mapping matrix, calculate the parameter change amplitude; Output the optimization effect evaluation report, including improvement metrics and convergence status.

[0135] S343: Determine Whether to Continue Iteration. Read the optimization effect evaluation report and check if the preset convergence condition is met; Check if the number of iterations reaches the upper limit; Check if the similarity improvement is below the threshold; If the continuation condition is satisfied, return to S32 for the next iteration; If the convergence condition is reached, output the final optimized mapping matrix and the optimized candidate set.

[0136] S4: Perform multi-modal similarity calculation and result ranking. Using the optimized mapping matrix, the optimized candidate set, and the context feature vector, calculate the similarities between the three semantic levels and the corresponding visual features to obtain the core concept similarity, the attribute modification similarity, and the relationship description similarity. Dynamically generate a hierarchical weight vector based on the context feature vector, weighted-fuse the three-layer similarities into a comprehensive similarity score, and output the final retrieval result set after ranking. The context understanding directly affects the result ranking strategy, making the retrieval results more in line with the user's expectations in a specific professional scenario, while saving the retrieval information for the system's continuous learning.

[0137] S41: Receive the optimized mapping matrix, the optimized candidate set, and the context feature vector. Prepare the final similarity calculation resources. Read the optimized mapping matrix, the optimized candidate set, and the context feature vector; verify the data integrity to ensure that all necessary information is available; optimize the memory allocation and computing resources for large-scale similarity calculation; prepare the similarity normalization parameters to ensure the comparability of scores at different levels; output the calculation-ready state, including the optimized calculation parameters.

[0138] S42: For each candidate image, calculate the feature similarities at three levels.

[0139] S421: Core concept similarity calculation. Read the object-level visual features of each candidate image and the core concept representation of the query; apply the concept mapping part in the optimized mapping matrix to align the semantic space and the visual feature space; calculate the cosine similarity after alignment as the basic similarity; apply the structural similarity enhancement to consider the matching degree of the concept hierarchy; apply the professional concept weighting to highlight the importance of the matching of medical key concepts; output the core concept similarity of each image, representing the object-level matching degree.

[0140] S422: Attribute modification similarity calculation. Read the attribute-level visual features of each candidate image and the attribute modification representation of the query; apply the attribute mapping part in the optimized mapping matrix to align the attribute semantics and the visual features; construct an attribute matching matrix and calculate the similarity between each pair of attributes; apply the Hungarian algorithm to find the optimal attribute matching combination to maximize the overall matching degree; consider the difference in attribute importance and give higher weights to key medical attributes; output the attribute modification similarity of each image, representing the attribute-level matching degree.

[0141] S423: Relationship description similarity calculation. Read the relationship-level visual features of each candidate image and the relationship description representation of the query; apply the relationship mapping part in the optimized mapping matrix to align the relationship semantics and visual relationship features; construct a relationship graph matching problem to compare the structural similarity between the query relationship graph and the image relationship graph; apply a graph matching algorithm, considering node similarity and edge similarity; apply expert rules to medical-specific relationships (such as pathological reaction chains, treatment effect relationships) to enhance the matching accuracy; output the relationship description similarity of each image, indicating the relationship-level matching degree.

[0142] S43: Dynamically determine the weight coefficients of the three-level similarity based on the context feature vector to generate a hierarchical weight vector.

[0143] S431: Contextual condition weight generation. Read the context feature vector, analyze the domain, purpose, and knowledge background features of the query; apply the weight generation strategy library to determine the basic weight according to the best practices in different professional fields; for example, assign a higher weight to the attribute layer for pathological diagnosis queries and a higher weight to the relationship layer for anatomical structure queries; adjust the weight according to the query purpose, with research purposes paying more attention to relationships and teaching purposes paying more balanced attention to each layer; adjust the professional sensitivity according to the knowledge background, and professional user queries paying more attention to detailed attribute matching; output the preliminary contextual condition weight.

[0144] S432: Weight self-adaptive optimization. Read the contextual condition weight and the three-level similarity score distribution of the candidate images; analyze the discrimination ability of the three-level similarity, and increase the weight of the level with strong discrimination ability; detect the deviation caused by extreme weight distribution, and apply a balance factor to ensure that each layer has an appropriate influence; apply task adaptability adjustment and adjust according to the historical optimal weight distribution of similar tasks; output the final hierarchical weight vector, including the optimal weight distribution of the three levels.

[0145] S44: Perform weighted fusion of the core concept similarity, attribute modification similarity, and relationship description similarity with the hierarchical weight vector to obtain a comprehensive similarity score.

[0146] S441: Weighted fusion calculation. Read the three-level similarity scores of each candidate image and the hierarchical weight vector; apply linear weighted fusion to calculate the preliminary weighted score; apply non-linear adjustment to handle the complementary and redundant relationships between layers; for example, when the core concept matching degree is extremely high, appropriately reduce the weight requirement for attribute matching; apply threshold constraints to ensure the minimum matching requirements for key layers; output the comprehensive similarity score of each candidate image.

[0147] S442: Confidence assessment. Analyze the consistency of the three - layer similarity, calculate the confidence level of the similarity judgment; apply entropy measurement to evaluate the certainty of the similarity distribution; attach a confidence index to each comprehensive similarity score to indicate the reliability of the similarity judgment; output the final similarity score with confidence.

[0148] S45: Sort the optimized candidate set according to the comprehensive similarity score to obtain the final retrieval result set.

[0149] S451: Similarity sorting and filtering. Read the final similarity scores of all candidate images; apply a sorting algorithm to sort them from high to low similarity; apply a similarity threshold filter to remove results with too low similarity; apply diversity optimization to ensure that the result set contains relevant images with different visual representations; output the sorted candidate result set.

[0150] S452: Result set enhancement processing. Read the candidate result set, apply a result grouping algorithm to cluster according to visual feature similarity; select the most representative image for each group as the main result; add result metadata, including the matching key concepts, prominent visual features, and confidence information; generate result interpretation data to explain the matching reasons and key features of each result; output the structured final retrieval result set.

[0151] S46: Return the final retrieval result set to the user, and at the same time save the context and feedback information of this retrieval for the system's continuous learning.

[0152] S461: Result presentation preparation. Read the final retrieval result set, optimize the result display format; adjust the professional depth of the result interpretation according to the professional level in the context feature vector; generate a result summary to highlight the common features and key differences; prepare visual enhancement, such as highlighting the matching area, annotating key features, etc.; output the enhanced retrieval result for the user, ready to be presented to the user.

[0153] S462: Retrieval session recording and learning. Record the complete retrieval session data, including queries, context, mapping matrix, and result set; store the intermediate optimization data during the retrieval for subsequent analysis; prepare a user feedback collection mechanism to obtain result relevance evaluation; update the system knowledge base, including term mapping, visual features, and common query patterns; output the retrieval session record for the system's continuous optimization and personalized adaptation.

[0154] Case 1: Taking the duplicate check and retrieval of academic paper images as the application scenario. With the large number of academic papers published, the phenomenon of repeated use of images in papers is becoming increasingly serious, including unintentional management negligence and intentional academic misconduct. Traditional image retrieval methods are difficult to accurately identify similar images in the academic professional field, especially in the case of involving professional terms and complex charts, with insufficient accuracy. This embodiment demonstrates how to use the method of the present invention to achieve high-precision similar retrieval of academic images and effectively identify the problem of image duplication in papers. This embodiment is implemented in the following environment: the processor is an Intel Xeon E5-2680 v4 CPU, the memory is 128GB DDR4, the GPU is an NVIDIA Tesla V100, the storage device is a 1TB NVMe SSD, the operating system is Ubuntu 20.04 LTS, and the main development frameworks are Python 3.8 and PyTorch 1.9.0.

[0155] Step 1: Build a hierarchical semantic-visual representation framework.

[0156] 1.1. Receive the user query. The user submits the query: "Find pictures similar to this fluorescence staining image of apoptosis, mainly focusing on the change of mitochondrial membrane potential". The query contains the image Q and the text description T.

[0157] 1.2. Hierarchically decompose the query. Process the query text T through BiLSTM and the attention mechanism to obtain the basic query vector representation: BQ = BiLSTM(T) = [0.72, 0.35, 0.91, 0.42, 0.58, 0.77, 0.29, 0.84, 0.61, 0.45]; where: BiLSTM() is the bidirectional long short-term memory network function; T is the query text "Find pictures similar to this fluorescence staining image of apoptosis, mainly focusing on the change of mitochondrial membrane potential"; BQ is the ten-dimensional basic query vector representation.

[0158] 1.2.1. Core concept extraction. The core concept representation CC = ExtractConcepts(BQ) = {(c1, w1),(c2, w2),..., (c n , w n )}; where: c1 = "apoptosis", w1 = 0.92; c2 = "fluorescence staining", w2 = 0.87; c3 = "mitochondria", w3 = 0.94; c4 = "membrane potential", w4 = 0.91; ExtractConcepts() is the core concept extraction function; BQ is the basic query vector; w is the concept weight, indicating the importance of the concept in the query, and the value range is [0,1].

[0159] Through verification by the ontology knowledge base, it is confirmed that these concepts are all valid professional terms in the field of cell biology.

[0160] 1.2.2. Attribute Modification Extraction. The attribute modification is expressed as AM = ExtractAttributes(BQ) = {(a1, v1, c1), (a2, v2, c2),...}; where: a1 = "staining type", v1 = "fluorescence", c1 = "apoptosis"; a2 = "change type", v2 = "potential change", c2 = "mitochondrial membrane"; a3 = "image type", v3 = "microscopic", c3 = "apoptosis"; ExtractAttributes() is the attribute extraction function; BQ is the basic query vector; a represents the attribute type; v represents the attribute value; c represents the associated core concept.

[0161] 1.2.3. Relationship Description Extraction. The relationship description is expressed as RD = ExtractRelations(BQ) = {(c1, r, c2, t),...}; where: c1 = "mitochondria", r = "shows", c2 = "membrane potential change", t = "during the process"; ExtractRelations() is the relationship extraction function; BQ is the basic query vector; c1 is the subject concept; r is the relationship type; c2 is the object concept; t is the timing information.

[0162] 1.3. Integrate Hierarchical Semantic Representation. Integrate the core concept representation CC, the attribute modification representation AM, and the relationship description representation RD into a hierarchical semantic representation HSR.

[0163] 1.4. Extract Image Visual Features. Apply hierarchical visual feature extraction to the query image Q and all images {I1, I2,..., I k} in the image library:

[0164] 1.4.1. Object-Level Feature Extraction. The object-level feature OLF = FastRCNN(I) = {(o1, v1), (o2, v2),...}; where: o1 = "cell nucleus", v1 = [0.72, 0.35, 0.91, 0.42]; o2 = "mitochondria", v2 = [0.58, 0.77, 0.29, 0.84]; o3 = "cytoplasm", v3 = [0.61, 0.45, 0.33, 0.79]; FastRCNN() is the improved fast region convolutional neural network function; I is the input image; o is the detected medical object; v is the corresponding feature vector. 7 cell objects are detected in image Q, and each object is represented by a 128-dimensional feature vector.

[0165] 1.4.2, Attribute-level feature extraction. The attribute-level feature ALF = ExtractVisualAttr(OLF) = {(a1, v1, o1, s1), ...}; where: a1 = "morphology", v1 = "round", o1 = "cell nucleus", s1 = 0.93; a2 = "intensity", v2 = "highlighted", o2 = "mitochondria", s2 = 0.87; a3 = "distribution", v3 = "peripherally aggregated", o3 = "mitochondria", s3 = 0.82; ExtractVisualAttr() is the visual attribute extraction function; OLF is the object-level feature; a is the attribute type; v is the attribute value; o is the associated object; s is the attribute confidence score, with a value range of [0,1].

[0166] 1.4.3, Relationship-level feature extraction. The relationship-level feature RLF = ExtractVisualRel(OLF, ALF) ={(o1, r, o2, s), ...}; where: o1 = "mitochondria", r = "close to", o2 = "cell nucleus", s = 0.89; o3 = "cytoplasm", r = "contains", o4 = "mitochondria", s = 0.94; ExtractVisualRel() is the visual relationship extraction function; OLF is the object-level feature; ALF is the attribute-level feature; o1, o2 are the relationship objects; r is the relationship type; s is the relationship confidence score, with a value range of [0,1].

[0167] 1.5, Construct a hierarchical visual representation. Integrate the object-level feature OLF, the attribute-level feature ALF, and the relationship-level feature RLF into a hierarchical visual representation HVR.

[0168] Step 2, Implement a context-aware professional term understanding system.

[0169] 2.1, Identify professional terms. Extract professional terms from the hierarchical semantic representation HSR: The professional term list PTL =ExtractTerms(HSR)={(t1, d1, c1), ...}; where t1 = "apoptosis", d1 = 0.12, c1 = 0.95; t2 = "fluorescent staining", d2 = 0.08, c2 = 0.93; t3 = "mitochondrial membrane potential", d3 = 0.19, c3 = 0.96; ExtractTerms() is the professional term extraction function; HSR is the hierarchical semantic representation; t is the extracted professional term; d is the term ambiguity, the larger the value, the higher the term ambiguity; c is the term domain relevance, the higher the value, the higher the relevance to the current domain.

[0170] 2.2, Analyze the professional field background and purpose of the query.

[0171] 2.2.1. Domain feature recognition. The domain feature vector DF = DomainAnalysis(HSR, PTL) = [0.94, 0.87, 0.23, 0.12, 0.08]; where DomainAnalysis() is the domain analysis function; HSR is the hierarchical semantic representation; PTL is the professional term list; DF[0]=0.94 represents the probability that the query belongs to the "Cell Biology" domain; DF[1]=0.87 represents the probability that the query belongs to the "Apoptosis Research" sub-domain; DF[2]=0.23 represents the probability that the query belongs to the "Drug Research" domain; DF[3]=0.12 represents the probability that the query belongs to the "Pathology" domain; DF[4]=0.08 represents the probability that the query belongs to the "Molecular Biology" domain.

[0172] 2.2.2. Query purpose analysis. The purpose feature vector PF = PurposeAnalysis(HSR, PTL) = [0.85, 0.62, 0.12, 0.09, 0.05]; where PurposeAnalysis() is the purpose analysis function; HSR is the hierarchical semantic representation; PTL is the professional term list; PF[0]=0.85 represents the probability that the query purpose is "Research and Analysis"; PF[1]=0.62 represents the probability that the query purpose is "Comparison and Verification"; PF[2]=0.12 represents the probability that the query purpose is "Teaching and Demonstration"; PF[3]=0.09 represents the probability that the query purpose is "Repeated Detection"; PF[4]=0.05 represents the probability that the query purpose is "Literature Research".

[0173] 2.2.3. Knowledge background extraction. The knowledge background vector KF = KnowledgeExtract(HSR, PTL) = [0.92, 0.88, 0.75, 0.67, 0.43]; where: KnowledgeExtract() is the knowledge background extraction function; HSR is the hierarchical semantic representation; PTL is the professional term list; KF[0]=0.92 represents the knowledge confidence of "Mitochondria play a key role in apoptosis"; KF[1]=0.88 represents the knowledge confidence of "Membrane potential change is an early marker of apoptosis"; KF[2]=0.75 represents the knowledge confidence of "Fluorescent staining is used to observe the state of living cells"; KF[3]=0.67 represents the knowledge confidence of "JC-1 dye is commonly used for mitochondrial membrane potential detection"; KF[4]=0.43 represents the knowledge confidence of "Morphological changes occur in cells during apoptosis".

[0174] 2.3. Integrate context feature vectors. The context feature vector CF = ContextFusion(DF, PF, KF) = [0.94, 0.85, 0.92, 0.88, 0.75, 0.62, 0.67]; where: ContextFusion() is the context feature fusion function; DF is the domain feature vector; PF is the purpose feature vector; KF is the knowledge background vector; CF is the integrated context feature vector, which retains the features with higher weights in each vector.

[0175] 2.4. Term ambiguity resolution. "Mitochondrial membrane potential" may refer to the measured value of membrane potential, the process of membrane potential change, or the method of membrane potential detection in different contexts. The disambiguation function DisambiguateTerms(PTL, CF) = {(t1, m1),...}; where: t1 = "mitochondrial membrane potential", m1 = "change process"; t2 = "apoptosis", m2 = "programmed death process"; t3 = "fluorescent staining", m3 = "method for detecting living cells"; DisambiguateTerms() is the term ambiguity resolution function; PTL is the professional term list; CF is the context feature vector; t is the professional term; m is the precise meaning after disambiguation.

[0176] 2.5. Generate context-enhanced query representation. The context-enhanced query representation CEQ = EnhanceQuery(HSR, CF) = [0.87, 0.92, 0.78, 0.93, 0.68, 0.82, 0.91, 0.74, 0.89, 0.85]; where: EnhanceQuery() is the query enhancement function; HSR is the hierarchical semantic representation; CF is the context feature vector; CEQ is the ten-dimensional context-enhanced query representation.

[0177] Step 3. Establish a semantic-visual bidirectional mapping and optimization framework.

[0178] 3.1. Initialize the mapping matrix. The initial mapping matrix IMM = InitializeMapping(CEQ, HVR) = [M1, M2, M3]; where: InitializeMapping() is the mapping initialization function; CEQ is the context-enhanced query representation; HVR is the hierarchical visual representation; M1 is the 128×128-dimensional core concept mapping matrix; M2 is the 64×64-dimensional attribute mapping matrix; M3 is the 32×32-dimensional relationship mapping matrix.

[0179] 3.2. Semantic-to-visual forward mapping. Visual Attention Map VAM = ForwardMapping(CEQ, IMM, HVR) = [A1, A2, A3]; where: ForwardMapping() is the forward mapping function; CEQ is the context-enhanced query representation; IMM is the initial mapping matrix; HVR is the hierarchical visual representation; A1 is the object attention map with a size of 7×7, mainly highlighting the mitochondrial region (value 0.92) and the nucleus region (value 0.85); A2 is the attribute attention map with a size of 7×7, mainly highlighting the region of fluorescence intensity change (value 0.88); A3 is the relationship attention map with a size of 7×7, mainly highlighting the interaction region between mitochondria and the nucleus (value 0.79). Use the visual attention map to perform a preliminary retrieval on the image library to obtain the candidate image set CS = {I1, I2, ..., I 50}, which contains 50 most similar images.

[0180] 3.3. Visual-to-semantic backward mapping. Backward mapping similarity RMS = BackwardMapping(CS, CEQ, IMM) = [S1, S2, ..., S 50 ; where: BackwardMapping() is the backward mapping function; CS is the candidate image set; CEQ is the context-enhanced query representation; IMM is the initial mapping matrix; S1 = 0.89 represents the backward mapping similarity of the first candidate image; S2 = 0.87 represents the backward mapping similarity of the second candidate image;...; S 50 = 0.61 represents the backward mapping similarity of the fiftieth candidate image.

[0181] 3.4. Calculate the three-level similarity and optimize the mapping matrix. Optimized Mapping Matrix OMM = OptimizeMapping(CEQ, CS, IMM, RMS) = [M'1, M'2, M'3]; where: OptimizeMapping() is the mapping optimization function; CEQ is the context-enhanced query representation; CS is the candidate image set; IMM is the initial mapping matrix; RMS is the backward mapping similarity; M'1 is the optimized core concept mapping matrix; M'2 is the optimized attribute mapping matrix; M'3 is the optimized relationship mapping matrix. Taking mitochondrial object recognition as an example, the corresponding value in the initial weight IMM is 0.72, and the corresponding value in the optimized OMM is adjusted to 0.86, enhancing the sensitivity to mitochondrial features. Through optimization, the optimized candidate set OCS = {I1, I2,..., I 30} is obtained, which contains 30 most similar images after optimization.

[0182] Step 4: Perform multi-modal similarity calculation and result ranking.

[0183] 4.1. Calculate multi-level similarity. The multi-level similarity MLS = MultiLevelSimilarity(OCS, CEQ, OMM) = {(I1, OLS1, ALS1, RLS1),...}; where: MultiLevelSimilarity() is the multi-level similarity calculation function; OCS is the optimized candidate set; CEQ is the context-enhanced query representation; OMM is the optimized mapping matrix; I1 is the first candidate image; OLS1 = 0.91 is the object-level similarity of I1; ALS1 = 0.87 is the attribute-level similarity of I1; RLS1 = 0.83 is the relationship-level similarity of I1.

[0184] 4.2. Generate hierarchical weight vectors based on context feature vectors. The hierarchical weight vector LWV = GenerateWeights(CF) = [w1, w2, w3]; where: GenerateWeights() is the weight generation function; CF is the context feature vector; w1 = 0.45 is the object-level weight; w2 = 0.35 is the attribute-level weight; w3 = 0.20 is the relationship-level weight. Since this query focuses on mitochondrial membrane potential changes, the analysis shows that object recognition (mitochondria) and attribute features (membrane potential changes) are more important, so the weights are adjusted accordingly.

[0185] 4.3. Fuse to generate a comprehensive similarity score. The comprehensive similarity score CSS = FuseSimilarities(MLS, LWV) = [s1, s2,..., s 30 ; where: FuseSimilarities() is the similarity fusion function; MLS is the multi-level similarity; LWV is the hierarchical weight vector; s1 = 0.45×0.91 + 0.35×0.87 + 0.20×0.83 = 0.88 is the comprehensive similarity score of the first candidate image; s2 = 0.45×0.88 + 0.35×0.85 + 0.20×0.79 = 0.86 is the comprehensive similarity score of the second candidate image;...; s 30 = 0.45×0.62 + 0.35×0.58 + 0.20×0.55 = 0.59 is the comprehensive similarity score of the thirtieth candidate image.

[0186] 4.4. Sort and output the retrieval results. The final retrieval result FRS = RankResults(OCS, CSS) = [(I1, s1), (I2, s2),...,(I 20 , s 20)]; where RankResults() is the result sorting function; OCS is the optimized candidate set; CSS is the comprehensive similarity score; sort in descending order according to the comprehensive similarity score, and select the top 20 images as the final retrieval results.

[0187] This embodiment conducts experiments on an academic paper image duplicate checking dataset (including 10,000 biomedical images) and compares with the existing state-of-the-art methods: the precision (P@10) of the traditional CBIR method (only based on visual features) is 0.68, and the recall (R@20) is 0.72; the precision of the vision-semantic embedding method (such as CLIP) is 0.76, and the recall is 0.79; the precision of this embodiment is 0.92, and the recall is 0.89. In terms of professional term ambiguity resolution, the accuracy of the traditional method is 67%, while this embodiment reaches 93%, improving the ability to understand professional terms. In practical applications, this embodiment has successfully detected the reuse of images in multiple papers, including some slightly modified images that are difficult to directly identify through visual features, demonstrating significant advantages in solving similar image retrieval problems in professional fields. This embodiment combines domain features, purpose features, and knowledge background to generate context feature vectors, solve the problem of professional term ambiguity; achieve the forward mapping from semantics to vision and the reverse mapping from vision to semantics, and generate an optimized mapping matrix through iterative optimization, enabling the system to optimize query understanding according to the actual image data distribution; dynamically generate hierarchical weight vectors based on context feature vectors to achieve the adaptive fusion of object-level, attribute-level, and relationship-level similarities, and improve the retrieval accuracy.

[0188] Case 2: Taking medical image retrieval as the background, enabling users to retrieve similar medical images that match the query semantics through natural language queries containing professional terms. Based on the context-aware semantic-visual bidirectional mapping framework, precise retrieval is achieved through hierarchical representation, professional term understanding, bidirectional mapping optimization, and multimodal similarity calculation. Specifically:

[0189] Step 1: Construct a hierarchical semantic-visual representation framework. Input the user query text "There is a 2-cm ground-glass nodule shadow in the upper right lobe of the lung, with blurred edges and traction signs around" and a medical image library (including 10,000 chest CT images).

[0190] 1.1. Receive the user query text and perform preliminary encoding on the query using a Transformer encoder. Use a pre-trained medical BERT encoder with a word vector dimension d = 768; the input word embedding sequence: E = [e1, e2, ..., e_n], where n = 15 (the number of words in the query); process through the multi-head self-attention layer: Att(Q,K,V) = softmax(QK T / sqrt(d))·V; where Q is the query matrix, K is the key matrix, and finally obtain the basic query vector representation Q_base ∈ R 768 , with values [0.21, 0.35, -0.14,..., 0.52].

[0191] 1.2. Apply a three-level structured decomposition algorithm to the basic query vector representation:

[0192] 1.2.1. Core concept extraction. Use a medical entity recognition model to identify the subject and key entities: "lung", "right upper lobe", "ground-glass nodule shadow"; verify the entities through a medical ontology knowledge base (such as UMLS) to obtain confidence scores: C("lung") = 0.98; C("right upper lobe") = 0.95; C("ground-glass nodule shadow") = 0.96; apply a semantic importance scoring algorithm to calculate the weight values: W("lung") = 0.75; W("right upper lobe") = 0.85; W("ground-glass nodule shadow") = 0.95; output the core concept representation matrix C_core∈R 3×768 , where each row corresponds to the embedding vector of a core concept.

[0193] 1.2.2. Attribute modification extraction. Identify the modifiers: "2cm", "fuzzy margin". Construct attribute triples: A1 = ("size", "2cm", "ground-glass nodule shadow"), confidence 0.92; A2 = ("edge characteristic", "fuzzy", "ground-glass nodule shadow"), confidence 0.88; output the attribute modification representation matrix A_att ∈ R 2×(2×768+64) , which contains the attribute type, attribute value vector, and the index and association strength of the associated core concept, and 64 dimensions are used to represent the association information.

[0194] 1.2.3. Relationship description extraction. Identify the relationship "there are traction signs around"; construct a relationship quadruple: R1 = ("ground-glass nodule shadow", "around", "traction sign", "current"), confidence 0.86; output the relationship description representation matrix R_rel∈R 1×(3×768+64) , which contains the subject concept, relationship type, object concept, and time series information vector, and 64 dimensions represent the time series information.

[0195] 1.3. Integrate the core concept representation, attribute modification representation, and relationship description representation into a hierarchical semantic representation: Construct a hierarchical association graph \(G_{sem}=(V, E)\), where \(V\) is the node set (core concepts) and \(E\) is the edge set (attributes and relationships); Use a graph attention network (GAT) to process the association graph and calculate the node feature update: \(h'_v=\sigma(\sum_{j\in N(v)}\alpha_{vj}\cdot W\cdot h_j)\); The attention coefficient \(\alpha_{vj} = softmax(LeakyReLU(a T [W\cdot h_v\parallel W\cdot h_j]))\); Finally, obtain the hierarchical semantic representation \(H_{sem}\in R (3+2+1)×1024 \), which integrates the semantic representations of three levels. Where \(\sigma\) is the activation function, \(N(v)\) is the set of neighbor nodes, \(W\) is the weight matrix, \(h_j\) is the neighbor node feature vector, \(a\) is the attention mechanism parameter, and \(h_v\) is the current node feature vector.

[0196] 1.4. For each medical image in the image library, apply a hierarchical visual feature extraction network:

[0197] 1.4.1. Object-level feature extraction. Use the Faster R-CNN model to detect objects in the medical image; Take the first image as an example, 3 objects are detected: O1 ("lung"), O2 ("right upper lobe"), O3 ("nodular opacity"); Extract a feature vector (2048 dimensions) for each object and record the location information (4 dimensions: x, y, w, h); Output the object-level visual feature matrix \(O_{vis}\in R 3 ×2052 。

[0198] 1.4.2. Attribute-level feature extraction. For each detected medical object region, use an attribute classification network to identify medical attributes. Take O3 ("nodular opacity") as an example, identify the attributes: Size: 1.8 cm (confidence 0.91); Morphology: irregular (confidence 0.87); Margin: blurred (confidence 0.85); Density: ground-glass (confidence 0.93); Output the attribute-level visual feature matrix \(A_{vis}\in R 8×(512+4) \), where 512 is the attribute feature dimension and 4 is the association information.

[0199] 1.4.3. Relationship-level feature extraction. Identify the spatial and functional relationships between objects; Construct a relationship graph between objects: \(R_{vis1}=(\text{"nodular opacity"}, \text{"is located in"}, \text{"right upper lobe"})\), confidence 0.94; \(R_{vis2}=(\text{"nodular opacity"}, \text{"has around"}, \text{"traction change"})\), confidence 0.82; Output the relationship-level visual feature matrix \(R_{vis}\in R 2×1024 。

[0200] 1.5. Organize the object-level, attribute-level, and relationship-level visual features of each image into a hierarchical visual representation: construct an image hierarchy to form a complete hierarchical visual representation \(H_{vis}\in\mathbb{R}\) (3+8+2)×1024 ; establish a visual index library based on Locality-Sensitive Hashing (LSH) to accelerate the retrieval process.

[0201] Step 2. Implement a context-aware professional term understanding system.

[0202] 2.1. Identify the list of professional terms.

[0203] 2.1.1. Professional term detection. Extract words from the hierarchical semantic representation and match them with the medical term library; identify the professional terms: "lung", "right upper lobe", "ground-glass nodule opacity", "indistinct margin", "retraction sign"; calculate the term specificity scores: \(S(\text{"lung"}) = 0.70\) (general term); \(S(\text{"right upper lobe"}) = 0.85\) (region-specific term); \(S(\text{"ground-glass nodule opacity"}) = 0.95\) (highly specific term); \(S(\text{"indistinct margin"}) = 0.80\) (descriptive term); \(S(\text{"retraction sign"}) = 0.90\) (sign-specific term).

[0204] 2.1.2. Polysemous term marking. Query the professional term ambiguity library to mark polysemous terms: "ground-glass nodule opacity": possible meaning 1 = "pulmonary interstitial lesion" (in the context of pneumonia), possible meaning 2 = "sign of early lung cancer" (in the context of tumor screening); "retraction sign": possible meaning 1 = "fibrotic contraction" (in the context of interstitial lung disease), possible meaning 2 = "tumor invasion" (in the context of malignant tumor). Calculate the ambiguity degree: \(Amb(\text{"ground-glass nodule opacity"}) = 0.65\); \(Amb(\text{"retraction sign"}) = 0.72\); output the list of professional terms \(T\_list\) with ambiguity markings.

[0205] 2.2. Analyze the professional field background and purpose of the query.

[0206] 2.2.1. Professional field identification. Calculate the probability distribution \(D\_field\) of the query belonging to different medical professional fields using a multi-label classification model: \(P(\text{"thoracic radiology"}) = 0.92\); \(P(\text{"pulmonology"}) = 0.85\); \(P(\text{"oncology"}) = 0.73\); \(P(\text{"general internal medicine"}) = 0.45\). Construct a domain feature vector \(F\_domain\in\mathbb{R}\) 256 , and adopt weighted fusion: \(F\_domain=\sum_{i}P(i)\cdot V\_i\), where \(V\_i\) is the standardized feature vector of each field.

[0207] 2.2.2. Query purpose analysis. Identify the query intention and calculate the probability distribution for different purposes: P("Diagnosis") = 0.87; P("Screening") = 0.65; P("Follow-up") = 0.32; P("Teaching") = 0.15. Determine the main query purpose: "Diagnosis". Construct the purpose feature vector F_purpose ∈ R 128 , and encode the query intention information.

[0208] 2.2.3. Knowledge background extraction. Retrieve relevant professional background knowledge from the medical knowledge graph: "Ground-glass nodule corresponds to stage IA in the TNM staging of lung cancer"; "A nodule with blurred margins may indicate invasive growth"; "Retraction signs are common in tumors or fibrosis". Extract key concepts and relationships to generate the knowledge background vector F_knowledge ∈ R 384 .

[0209] 2.3. Integrate the domain, purpose, and knowledge background features to generate the context feature vector. Apply feature standardization to unify the dimensions and ranges; use the attention-weighted fusion mechanism to calculate the fusion weights: w_domain = 0.45; w_purpose = 0.35; w_knowledge = 0.20; Generate the context feature vector F_context = w_domain·F_domain + w_purpose·F_purpose + w_knowledge·F_knowledge; Finally, F_context ∈ R 512 , which encodes the complete query context information. Here, w_domain is the domain feature weight, w_purpose is the query purpose weight, and w_knowledge is the knowledge background weight.

[0210] 2.4. Perform term ambiguity resolution based on the professional term list and the context feature vector.

[0211] 2.4.1. Construct the term conditional probability model. Calculate the conditional probability P(meaning|context) for each polysemous term: For "ground-glass nodule": P("Pulmonary interstitial lesion"|F_context) = 0.25; P("Early lung cancer sign"|F_context) = 0.75; For "retraction sign": P("Fibrotic contraction"|F_context) = 0.30; P("Tumor invasion"|F_context) = 0.70; The calculation formula: P(meaning|context) = softmax(f_θ(meaning, F_context)), where f_θ is a trained deep neural network.

[0212] 2.4.2. Context Feature Matching. Calculate the matching degree M_score between the meaning of each term and the query context: M_score("ground-glass nodule shadow", "early lung cancer sign") = cos(V("early lung cancer sign"), F_context) = 0.82; M_score("retraction sign", "tumor invasion") = cos(V("tumor invasion"), F_context) = 0.78; where V() is the vector representation of the term meaning, and cos() is the cosine similarity function.

[0213] 2.4.3. Term Network Collaborative Disambiguation. Construct a relationship network between terms to represent the semantic dependencies between terms; apply a graph reasoning algorithm to update the meaning probability: P'("ground-glass nodule shadow" = "early lung cancer sign"|F_context, Network) = 0.85; P'("retraction sign" = "tumor invasion"|F_context, Network) = 0.82; Update formula: P'(m_i|F_context, Network) = P(m_i|F_context) + α·∑_j w_ij·P(m_j|F_context), where α = 0.3 is the balance factor and w_ij is the association weight between terms.

[0214] 2.5. Integrate the disambiguated semantic representation and the context feature vector. Update the term vectors in the hierarchical semantic representation and replace them with the professional representations of the selected meanings; apply context-conditioned feature transformation to enhance the core concepts, attribute modifications, and relationship descriptions; obtain the context-enhanced query representation Q_enhanced ∈ R 1536 。

[0215] Step Three: Establish a Semantic-Visual Bidirectional Mapping and Optimization Framework.

[0216] 3.1. Initialize the bidirectional mapping matrix: Use the pre-trained semantic-visual mapping weights in the medical field as the initial values; construct initial mapping matrices for core concepts, attribute modifications, and relationship descriptions respectively: W_concept ∈ R 1024×1024 ; W_attribute ∈ R 1024×1024 ; W_relation ∈ R 1024×1024 ; Establish an inter-level association matrix W_inter ∈ R 3×3 , with the initial value of [[0.5, 0.3, 0.2], [0.3, 0.5, 0.2], [0.2, 0.3, 0.5]].

[0217] 3.2. Implement the forward mapping from semantics to vision.

[0218] 3.2.1. Generate multi-level visual attention maps. Decompose the context-enhanced query representation Q_enhanced into query components corresponding to hierarchical visual representations: Q_concept ∈ R 3×1024 ; Q_attribute ∈ R 2×1024 ; Q_relation ∈ R 1×1024 . Generate three-level attention maps: A_concept = softmax(Q_concept·W_concept·O_vis T ) ∈ R 3×3 ; A_attribute = softmax(Q_attribute·W_attribute·A_vis T ) ∈ R 2×8 ; A_relation = softmax(Q_relation·W_relation·R_vis T ) ∈ R 1×2 .

[0219] 3.2.2. Adjust attention weights based on context. Use the context feature vector F_context to adjust attention weights: A_concept' = A_concept * sigmoid(F_context_c·O_vis); A_attribute' = A_attribute * sigmoid(F_context_a·A_vis); A_relation' = A_relation * sigmoid(F_context_r·R_vis); where F_context_c, F_context_a, and F_context_r are the projections of F_context at different levels.

[0220] 3.2.3. Perform visual feature activation and retrieval. Calculate the similarity matrix of the forward mapping: S_concept = A_concept'·O_vis·Q_concept T ∈ R 3×3 ; S_attribute = A_attribute'·A_vis·Q_attribute T ∈ R 2×2 ; S_relation = A_relation'·R_vis·Q_relation T ∈ R 1×1Calculate the initial forward mapping similarity: S_forward = (w_c·tr(S_concept) + w_a·tr(S_attribute) + w_r·tr(S_relation)) / (w_c + w_a + w_r), where w_c = 0.4, w_a = 0.3, w_r = 0.3 are the initial hierarchical weights. Select the top 50 similar images as the candidate set C_initial based on the forward mapping similarity S_forward.

[0221] 3.3. Implement the reverse mapping from vision to semantics.

[0222] 3.3.1. Visual feature distribution analysis. Perform clustering analysis on the image features in the candidate set C_initial to identify visual patterns; use the k-means algorithm (k = 3) to cluster the object-level features to obtain the centroids: Centroid1 = [0.72, 0.31,..., 0.55]; Centroid2 = [0.45, 0.62,..., 0.28]; Centroid3 = [0.81, 0.27,..., 0.63]; Identify the dominant visual pattern: "nodular opacity in the right upper lobe", with an appearance frequency of 80%.

[0223] 3.3.2. Feature space mapping analysis. Evaluate the quality of the semantic-to-visual mapping: The semantic concept "ground-glass nodule" maps to multiple visual manifestations (over-mapping), quantified as 0.75; the semantic concept "fuzzy margin" maps insufficiently (under-mapping), quantified as 0.65; Generate a mapping quality assessment report to determine the mapping areas that need to be adjusted.

[0224] 3.3.3. Dynamic adjustment of semantic understanding. Based on the visual feature distribution, adjust the semantic concept weights: W'("ground-glass nodule") = W("ground-glass nodule") * 0.85 = 0.95 * 0.85 = 0.81 (reduce the weight of the over-mapped concept); W'("fuzzy margin") = W("fuzzy margin") * 1.25 = 0.80 * 1.25 = 1.00 (increase the weight of the under-mapped concept); Introduce a new semantic concept: "ground-glass density" (discovered from visual features but not explicitly in the original query); Generate an adaptive semantic representation S_adaptive to better match the actual visual feature distribution.

[0225] 3.3.4. Update the bidirectional mapping relationship. Apply the gradient descent algorithm to optimize the mapping matrix parameters according to the mapping error: W_concept' = W_concept - η·▽L(W_concept), where η = 0.01 is the learning rate; W_attribute' = W_attribute - η·▽L(W_attribute); W_relation' = W_relation - η·▽L(W_relation); update the hierarchical association structure W_inter'; generate the reverse mapping similarity matrix S_backward. Here, W_concept' is the updated concept mapping matrix, ▽L is the gradient of the loss function, W_relation' is the updated relation mapping matrix, and W_attribute' is the updated attribute mapping matrix.

[0226] 3.4. Execute the iterative optimization loop.

[0227] 3.4.1. Recalculate the semantic-visual similarity. Use the updated mapping matrix and the adaptive semantic representation to recalculate the similarity; obtain the optimized similarity S_optimized = 0.6·S_forward + 0.4·S_backward; update the candidate set C_optimized (the top 30 most similar images).

[0228] 3.4.2. Evaluate the optimization effect. Calculate the change in similarity distribution: the average similarity increases from 0.73 to 0.81; analyze the change in ranking: 75% of the images have their rankings improved, with an average improvement of 3.2 positions; evaluate the magnitude of parameter change: the average parameter change is 2.8%, which is lower than the threshold of 5%.

[0229] 3.4.3. Judge the iteration termination condition. The similarity improvement < 0.01 (lower than the threshold of 0.02); the number of iterations = 3 (not reaching the upper limit of 5); the termination condition is satisfied, and output the final optimized mapping matrix and the optimized candidate set.

[0230] Step Four. Execute the multi-modal similarity calculation and result ranking.

[0231] 4.1. Prepare the resources for the final similarity calculation. Verify that all necessary information is complete; optimize the calculation resource allocation; set the normalization parameters to ensure the comparability of scores at different levels.

[0232] 4.2. Calculate the feature similarities at three levels.

[0233] 4.2.1. Core concept similarity calculation. Calculate the core concept similarity for each image i: S_concept(i) = ∑_j∑_k A_concept'(j,k)·cos(Q_concept(j),O_vis(i,k)); for example, for the first image: S_concept(1) = 0.92*cos([0.42,0.68,...,0.55],[0.45,0.66,...,0.58])+... = 0.87.

[0234] 4.2.2. Attribute modification similarity calculation. Calculate the attribute-level similarity: S_attribute(i)=∑_j∑_k A_attribute'(j,k)·cos(Q_attribute(j),A_vis(i,k)); for example, for the first image: S_attribute(1)=0.85*cos([0.32,0.71,...,0.48],[0.35,0.68,...,0.51])+... = 0.79.

[0235] 4.2.3. Relationship description similarity calculation. Calculate the relationship-level similarity: S_relation(i)=∑_j∑_k A_relation'(j,k)·cos(Q_relation(j),R_vis(i,k)); for example, for the first image: S_relation(1)=0.80 * cos([0.62,0.41,...,0.38],[0.58,0.44,...,0.35]) = 0.76.

[0236] 4.3. Dynamically generate hierarchical weight vectors based on context feature vectors.

[0237] 4.3.1. Context condition weight generation. Analyze the context feature vector F_context to determine the basic weights: According to the characteristics of the "Chest Radiology" field, the initial weight distribution is: w_c = 0.30, w_a = 0.45, w_r = 0.25; Adjust according to the "diagnosis" purpose: w_c' = 0.25, w_a' = 0.50, w_r' = 0.25 (increase the weight of the attribute layer); Calculation formula: w_x' = w_x + γ·f_purpose(x), where γ = 0.15 is the adjustment factor and f_purpose is the purpose adjustment function.

[0238] 4.3.2. Weight Adaptive Optimization. Analyze the discrimination ability of the three-layer similarity: Var(S_concept) = 0.034 (low discrimination ability); Var(S_attribute) = 0.087 (high discrimination ability); Var(S_relation) = 0.052 (medium discrimination ability); Apply discrimination ability enhancement: w_x'' = w_x' * (1 + β·Var(S_x)), where β = 0.5 is the amplification factor. The final hierarchical weight vector W_level = [0.22, 0.56, 0.22].

[0239] 4.4. Integrate multi-level similarities and calculate the comprehensive similarity score.

[0240] 4.4.1. Weighted Fusion Calculation. Calculate the linear weighted score: S_linear(i) = W_level(1)·S_concept(i)+W_level(2)·S_attribute(i)+W_level(3)·S_relation(i); For example, for the first image: S_linear(1) = 0.22·0.87 + 0.56·0.79 + 0.22·0.76 = 0.80. Apply the non-linear adjustment function to handle the complementary relationship between levels: S_nonlinear(i) = S_linear(i) * (1 + Δ·min(S_concept(i), S_attribute(i), S_relation(i))); where Δ = 0.1 is the non-linear factor. For example, for the first image: S_nonlinear(1) = 0.80 * (1 + 0.1·0.76) = 0.86.

[0241] 4.4.2. Confidence Evaluation. Calculate the consistency of the three-layer similarities: C(i) = 1 - σ(S_concept(i), S_attribute(i), S_relation(i)); where σ is the standard deviation function. For example, for the first image: C(1) = 1 - σ(0.87, 0.79, 0.76) = 1 - 0.057 = 0.943. The final similarity score: S_final(i) = S_nonlinear(i) * C(i). For example, for the first image: S_final(1) = 0.86 * 0.943 = 0.81.

[0242] 4.5. Sort and output the final retrieval result set.

[0243] 4.5.1 Similarity Sorting and Filtering. Sort all candidate images in descending order according to the final similarity score S_final, apply the similarity threshold for filtering, and remove the results where S_final < 0.65; apply diversity optimization to ensure that the result set contains relevant images with different visual representations.

[0244] 4.5.2 Result Set Enhancement Processing. Cluster the results into 5 groups according to visual feature similarity; select the most representative image for each group as the main result; add result metadata, including the key concepts of the match, prominent visual features, and confidence information; generate result explanation data, such as "Reason for match: ground-glass nodule in the right upper lobe, with blurred edges and traction signs".

[0245] 4.6 Return the final retrieval result set to the user, and at the same time save the retrieval session data for the continuous learning of the system.

[0246] In this embodiment, through the three-dimensional vector representation of the domain background, query purpose, and knowledge background, a comprehensive context is provided for term disambiguation, enabling the system to correctly understand professional terms based on the professional background. It not only maps from semantics to vision but also adjusts semantic understanding from visual feedback, solving the problem in traditional one-way mapping that semantic understanding cannot be adjusted according to actual visual data. The query and image are decomposed into three levels: core concepts, attribute modifiers, and relationship descriptions to achieve fine-grained semantic matching, and at the same time, the hierarchical weights are dynamically adjusted through the context.

[0247] In another embodiment of the present application, generating a list of professional terms specifically involves: matching the vocabulary in the hierarchical semantic representation with a pre-trained domain-specific professional term vocabulary. The input hierarchical semantic representation HSR = [0.87, 0.92, 0.78, 0.93, 0.68, 0.82, 0.91, 0.74, 0.89, 0.85], and extracting the vocabulary sequence WS = ["apoptosis", "fluorescent staining", "mitochondria", "membrane potential", "change", "image", "similar", "detection"] from it. The vocabulary match score MS = VocabMatch(WS, PDV) = {(w1, m1, p1), (w2, m2, p2),...}; where: w1 = "apoptosis", m1 = 0.95, p1 = 0.92; w2 = "fluorescent staining", m2 = 0.89, p2 = 0.87; w3 = "mitochondria", m3 = 0.97, p3 = 0.94; w4 = "membrane potential", m4 = 0.91, p4 = 0.88; w5 = "change", m5 = 0.43, p5 = 0.25; w6 = "image", m6 = 0.52, p6 = 0.31; w7 = "similar", m7 = 0.38, p7 = 0.22; w8 = "detection", m8 = 0.67, p8 = 0.45; VocabMatch() is the vocabulary match function; WS is the vocabulary sequence; PDV is the pre-trained professional term vocabulary in the field of cell biology; w is the word; m is the matching confidence, indicating the degree of match between the vocabulary and the professional term; p is the professionalism score, indicating the intensity of the professional term attribute of the vocabulary. Set the matching threshold to 0.8, and filter out the preliminary professional term candidate set IPTC = {("apoptosis", 0.95), ("fluorescent staining", 0.89), ("mitochondria", 0.97), ("membrane potential", 0.91)}.

[0248] Based on the preliminary professional term candidate set, calculate the attribution probability of each term in the current query domain. The domain attribution probability DAP = DomainVerify(IPTC, DF)={(t1,dp1),(t2, dp2),...}; where: t1 = "apoptosis", dp1 = 0.97; t2 = "fluorescent staining", dp2 = 0.93; t3 = "mitochondria", dp3 = 0.96; t4 = "membrane potential", dp4 = 0.89; DomainVerify() is the domain verification function; IPTC is the preliminary professional term candidate set; DF is the domain feature vector [0.94,0.87, 0.23, 0.12, 0.08]; t is the professional term; dp is the domain attribution probability, and the calculation formula is dp = Σ i (DF i×term_domain_weight i ), where term_domain_weight i is the weight distribution of the term in the i-th domain. The verified set of professional term candidates VPTC = {("Apoptosis", 0.97), ("Fluorescent staining", 0.93), ("Mitochondria", 0.96), ("Membrane potential", 0.89)}, and the domain membership probabilities of all terms exceed the 0.85 threshold.

[0249] Based on semantic link relationships and sentence structures, evaluate the importance of terms in the current query context. The context importance score CIS = ContextRelevance(VPTC, HSR, SG) = {(t1, cis1), (t2, cis2),...}; where: t1 = "Apoptosis", cis1 = 0.92; t2 = "Fluorescent staining", cis2 = 0.87; t3 = "Mitochondria", cis3 = 0.94; t4 = "Membrane potential", cis4 = 0.91; ContextRelevance() is the context relevance evaluation function; VPTC is the verified set of professional term candidates; HSR is the hierarchical semantic representation; SG is the sentence structure graph; t is the professional term; cis is the context importance score, and the calculation formula is cis = α × semantic_link_score + β × structure_position_score + γ × co_occurrence_score, where α = 0.4, β = 0.3, γ = 0.3 are the weight coefficients of semantic links, structural positions, and co-occurrence relationships respectively. The final list of professional terms PTL = {("Mitochondria", 0.94), ("Apoptosis", 0.92), ("Membrane potential", 0.91), ("Fluorescent staining", 0.87)}, sorted in descending order of the context importance score.

[0250] In another embodiment of the present application, generating the reverse mapping similarity matrix is specifically as follows: Based on the candidate image set CS = {I1, I2,..., I 50}, perform reverse mapping analysis to generate the reverse mapping similarity matrix. Analyze the visual feature distribution of each image in the candidate image set to identify the dominant visual pattern. The visual feature distribution characteristic VFD = AnalyzeVisualDist(CS = {(I i , mv i , kv i),...}; where: For I1, mv1 = ["mitochondrial aggregation", "fluorescence highly bright", "cell outline clear"], kv1 = [0.89, 0.84, 0.76]; for I2, mv2 = ["mitochondrial dispersion", "fluorescence medium", "cell outline blurred"], kv2 = [0.72, 0.68, 0.54]; for I3, mv3 = ["mitochondrial swelling", "fluorescence weak", "cell morphology abnormal"], kv3 = [0.81, 0.45, 0.69]; AnalyzeVisualDist() is the visual distribution analysis function; CS is the candidate image set; I i is the i-th candidate image; mv i is the list of dominant visual patterns; kv i is the key visual feature intensity vector, and the numerical range [0, 1] represents the significance level of the feature. Through cluster analysis, three main visual patterns are identified: normal state pattern (accounting for 45%), early apoptosis pattern (accounting for 35%), and late apoptosis pattern (accounting for 20%).

[0251] Based on the identified dominant visual patterns, a visual-semantic bridging network is constructed to project visual features into the semantic space. The bridging network outputs VSB = BridgeNetwork(VFD, SEM) = {(I i , rs i ),...}; where: For I1, rs1 = ["cells are in a normal state", "mitochondrial function is active", "membrane potential is stable"]; for I2, rs2 = ["cells start to apoptosis", "mitochondrial function weakens", "membrane potential slightly drops"]; for I3, rs3 = ["cells are deeply apoptotic", "mitochondria are severely damaged", "membrane potential significantly decreases"]; BridgeNetwork() is the visual-semantic bridging function; VFD is the visual feature distribution characteristic; SEM is the semantic embedding matrix; I i is the i-th candidate image; rs i is the reverse semantic description, which is the semantic expression deduced reversely through visual features. The bridging network adopts a three-layer transformation structure: visual feature layer → intermediate mapping layer → semantic representation layer, and the dimension of the transformation matrix for each layer is 128×64, 64×32, 32×10 respectively.

[0252] Compare the reverse semantic description with the original query semantics to identify potential differences. The semantic difference matrix SDM = IdentifyDifference(VSB, CEQ) = {(I i , sd i),...}; where: The sd1 of I1 = [0.15, 0.08, 0.12, 0.06, 0.11, 0.09, 0.07, 0.04, 0.13, 0.10]; The sd2 of I2 = [0.23, 0.19, 0.25, 0.17, 0.22, 0.18, 0.14, 0.21, 0.26, 0.20]; The sd3 of I3 = [0.41, 0.38, 0.44, 0.35, 0.42, 0.37, 0.33, 0.39, 0.45, 0.40]; IdentifyDifference() is the semantic difference recognition function; VSB is the output of the bridging network; CEQ is the context-enhanced query representation [0.87, 0.92, 0.78, 0.93, 0.68, 0.82, 0.91, 0.74, 0.89, 0.85]; I i is the i-th candidate image; sd i is the semantic difference vector, and the calculation formula is sd i = |CEQ - reverse_semantic_vector i |, representing the absolute difference of each dimension. Where reverse_semantic_vector i is the reverse semantic vector of the i-th candidate image. Calculate the average semantic difference: The average difference of I1 is 0.095, the average difference of I2 is 0.205, and the average difference of I3 is 0.394. The smaller the difference, the closer the image is to the query semantics.

[0253] According to the identified semantic differences, dynamically adjust the parameters of the semantic understanding model to generate the final reverse mapping similarity matrix. The reverse mapping similarity matrix RMS = DynamicAdjust(SDM, CF, MP) = [s1, s2,..., s 50 ; where: s1 = 1 - (0.4 × avg_semantic_diff1 + 0.3 × context_penalty1 + 0.3 × visual_penalty1) = 1 - (0.4 × 0.095 + 0.3 × 0.12 + 0.3 × 0.08) = 1 - 0.098 = 0.902; s2 = 1 - (0.4 × 0.205 + 0.3 × 0.18 + 0.3 × 0.15) = 1 - 0.181 = 0.819; s3 = 1 - (0.4 × 0.394 + 0.3 × 0.35 + 0.3 × 0.32) = 1 - 0.359 = 0.641; DynamicAdjust() is the dynamic adjustment function; SDM is the semantic difference matrix; CF is the context feature vector; MP is the model parameter; avg_semantic_diffi is the average semantic difference of the i-th image; context_penalty i is the context consistency penalty term; visual_penalty i is the visual feature consistency penalty term. Through dynamic adjustment, the model parameters are adjusted from the initial state θ0 = [0.72, 0.68, 0.75, 0.71, 0.69] to the optimized state θ1 = [0.78, 0.74, 0.81, 0.77, 0.75], improving the accuracy of semantic understanding. Finally, the reverse mapping similarity matrix RMS = [0.902, 0.819, 0.641, 0.785, 0.856,..., 0.623], which contains the reverse mapping similarity scores of all 50 candidate images.

[0254] In another embodiment of the present application, the steps of bidirectional mapping optimization are specifically as follows: Based on the forward mapping similarity matrix FMS = [0.89, 0.87, 0.61, 0.78, 0.85,..., 0.62] and the reverse mapping similarity matrix RMS = [0.902, 0.819, 0.641, 0.785, 0.856,..., 0.623], bidirectional mapping optimization is performed. The optimized mapping matrix OMM = OptimizeMapping(FMS, RMS, IMM, λ) = [M'1, M'2, M'3]; where M'1 is the optimized core concept mapping matrix, and the adjustment amplitude Δ1 = λ × (FMS - RMS) = 0.3 × ([0.89, 0.87, 0.61,...] - [0.902, 0.819, 0.641,...]) = 0.3 × [-0.012, 0.051, -0.031,...] = [-0.0036, 0.0153, -0.0093,...]; M'2 is the optimized attribute mapping matrix, and the adjustment amplitude Δ2 = 0.2 × (FMS - RMS); M'3 is the optimized relationship mapping matrix, and the adjustment amplitude Δ3 = 0.1 × (FMS - RMS); OptimizeMapping() is the mapping optimization function; FMS is the forward mapping similarity matrix; RMS is the reverse mapping similarity matrix; IMM is the initial mapping matrix; λ is the learning rate constant; M' i is the optimized mapping matrix of the i-th layer. Through bidirectional consistency constraint Loss = Σ i (α i |FMS i - RMS i | 2 + β i ||M' i - IMM i || 2), where α1 = 0.5, α2 = 0.3, α3 = 0.2 are the consistency weights of each layer, β1 = 0.1, β2 = 0.1, β3 = 0.1 are the regularization weights, and the final loss function converges to 0.034. After optimization by bidirectional mapping, the accuracy of the system in understanding professional terms is improved from the initial 67% to 93%, and the consistency index in reverse mapping is improved from 0.72 to 0.91. Compared with the traditional unidirectional mapping method, in this embodiment, when dealing with the professional field image retrieval task, the precision is improved from 0.76 to 0.92, and the recall rate is improved from 0.79 to 0.89, verifying the effectiveness of the bidirectional mapping optimization mechanism.

[0255] The present invention constructs a context feature vector by combining domain features, purpose features, and knowledge background, and performs term ambiguity resolution based on the context features, thereby improving the ability to accurately understand professional terms. In the scenario of academic paper image retrieval, the system can identify the specific meaning differences of "mitochondrial membrane potential" under different research backgrounds (such as measurement values, change processes, or detection methods), and increase the accuracy rate of term ambiguity resolution from 67% of traditional methods to 93%. This enables the system to accurately understand the professional query intentions of researchers, correctly identify images that are visually similar but have different semantics or research purposes, and effectively solves the misjudgment problem caused by inaccurate understanding of professional terms during image duplicate checking in academic papers. When researchers query for cell morphology images under specific experimental conditions, the system can accurately distinguish visually similar images with different experimental purposes, reducing the false positive rate in academic review. By establishing a two-way interaction mechanism, the system can dynamically adjust semantic understanding to adapt to the actual visual feature distribution. In academic image retrieval, traditional one-way mapping methods are easily affected by visual noise and variations, while the present invention can obtain feedback from the visual feature distribution through reverse mapping to optimize the mapping matrix parameters. For example, when detecting apoptosis images, the system improves the weight of mitochondrial object recognition from 0.72 to 0.86 through iterative optimization, enhancing the sensitivity to key biomarkers. This enables the system to still capture the core semantic features when dealing with images deliberately modified (such as adjusting brightness, contrast, cropping, or rotating) to avoid duplicate checking, and the retrieval accuracy rate is increased by 21%. For fine-tuned academic images, the detection rate of traditional methods is only 58%, while the present invention reaches 87%, effectively preventing academic misconduct. By decomposing queries and images into three levels: core concepts, attribute modifications, and relationship descriptions, and establishing a strictly corresponding hierarchical representation framework, more refined semantic and visual feature matching can be performed. In the application of academic paper image retrieval, this hierarchical representation enables the system to distinguish images with similar objects but different relationships, such as distinguishing the subtle differences between "mitochondria close to the nucleus" and "mitochondria far from the nucleus", which is crucial for judging experimental results in cell biology research. Traditional holistic feature representation methods are difficult to capture such fine-grained differences, while the present invention extracts and matches object-level, attribute-level, and relationship-level features respectively, and improves the discrimination accuracy of complex images from 72% to 89%. Especially when identifying redrawn or re-typeset academic charts, the consistency of core data relationships can be identified through relationship-level features, effectively preventing behavior of evading duplicate checking by redrawing charts. By converting the context feature vector into a hierarchical weight vector and dynamically fusing it with multi-level similarity to generate a comprehensive similarity score, the retrieval strategy can be adaptively adjusted according to different professional fields and query purposes.In the academic image duplicate checking scenario, when researchers focus on "mitochondrial membrane potential changes", the system automatically increases the weights of the object level (weight 0.45) and the attribute level (weight 0.35), while reducing the influence of the relationship level (weight 0.20), because the membrane potential changes are mainly manifested as changes in the fluorescence intensity (attribute) of mitochondria (object). The precision rate in academic paper image retrieval reaches 0.92, which is 24 percentage points higher than the traditional CBIR method and 16 percentage points higher than the visual-semantic embedding method. When facing different types of academic images (such as microscopic images, charts, flowcharts), the system can automatically adjust the weights of the corresponding levels, optimize the retrieval strategy, so that the system maintains high precision in diverse academic image retrieval tasks, and the retrieval speed is increased by 38% compared with the traditional method. By matching professional terms with a pre-trained domain term list, verifying the domain adaptability of terms, evaluating context relevance, identifying the set of professional fields to which the query belongs, analyzing the query purpose, and retrieving relevant professional background knowledge, a comprehensive context feature vector is constructed. In the application of academic paper image retrieval, it can accurately distinguish different interpretations of the same visual phenomenon in different research fields. For example, the same cell morphological change represents programmed cell death in apoptosis research, while it represents the cell differentiation process in differentiation research. By identifying that the query comes from the "cell biology" field (probability 0.94) and the "cell apoptosis research" sub-field (probability 0.87), and determining the query purpose as "research analysis" (probability 0.85), the specific context of the term can be correctly understood. The term understanding accuracy rate is increased from 71% of the traditional method to 95%, reducing the semantic confusion problem in academic image retrieval, especially in interdisciplinary research fields, making the retrieval results more in line with the professional expectations of researchers. By analyzing the focus of attention and importance distribution of the current query in the professional field, the system applies a dynamic weight generator to generate a hierarchical weight vector adapted to different professional scenarios according to the focus of attention, importance distribution, and query purpose. In the application of academic paper image duplicate checking, it can accurately identify different focuses of attention on image features in different research types. For example, for cell morphology research, the system will increase the weight of object structure features; for molecular localization research, it will increase the weight of spatial relationship features; for kinetics research, it will emphasize the weight of temporal change features. Through targeted weight adjustment, when comprehensively evaluating a large academic image library (10,000 biomedical images), the system achieves an average precision rate of 92%, which is 15 percentage points higher than the fixed weight method. Especially when processing interdisciplinary research images, the system can adapt to the evaluation criteria of different disciplines, reduce misjudgments caused by disciplinary differences, and improve the applicability and fairness of the academic image duplicate checking system in a multi-disciplinary environment. By constructing a visual-semantic bridging network, projecting visual features into the semantic space, identifying potential semantic differences, and dynamically adjusting the semantic understanding model according to actual image data.In the scenario of academic paper image retrieval, when the system detects that mitochondria exhibit a specific fluorescence pattern in a large number of candidate images, it will automatically enhance the semantic understanding of this visual pattern, even if this feature is not explicitly stated in the original query. When processing an image set containing a specific professional visual pattern, it can automatically discover and strengthen the recognition of key visual features, improving the retrieval accuracy from the initial 76% to 89%. Especially when detecting common image series in academic literature (such as time-series experimental images or dose-gradient response images), the system can identify the common visual patterns and changing trends in the series of images, effectively recognize the same experimental data used in different papers, even if these images have undergone different processing and presentation methods, and the successful detection rate has increased from 62% to 91%, enhancing the ability of academic image traceability.

[0256] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A method for retrieving similar images, characterized in that, Including: Receiving user query data and constructing a hierarchical semantic-visual representation framework, including hierarchical semantics and visual representation; Based on the hierarchical semantic representation, adopting a context-aware professional term understanding system to generate context feature vectors, hierarchical weight vectors, and context-enhanced query representations; Based on the context-enhanced query representation and hierarchical visual representation, adopting a semantic-visual bidirectional mapping and optimization framework to generate an optimized mapping matrix and an optimized candidate set, and combining with context feature vectors to perform multimodal similarity calculation to obtain multi-level similarities; Fusing the multi-level similarities with the hierarchical weight vectors into a comprehensive similarity score, sorting the optimized candidate set accordingly, and outputting the retrieval result set.

2. The method according to claim 1, characterized in that Generating context feature vectors, hierarchical weight vectors, and context-enhanced query representations, including: Identifying professional terms from the hierarchical semantic representation to obtain a list of professional terms; combining it with the hierarchical semantic representation to analyze the professional field background and purpose of the query, generating domain feature, purpose feature, and knowledge background vectors, and fusing them into context feature vectors; Inputting the context feature vectors into a weight generation network to dynamically generate hierarchical weight vectors according to the professional field characteristics and purposes of the query; Based on the list of professional terms and context feature vectors, performing term ambiguity resolution, obtaining the disambiguated semantic representation and fusing it with the context feature vectors to generate a context-enhanced query representation.

3. The method according to claim 2, wherein Obtaining a list of professional terms, including: Matching the vocabulary in the hierarchical semantic representation with a pre-trained domain professional term vocabulary to generate a preliminary candidate set of professional terms; performing domain adaptability verification on each term in it, calculating the domain membership probability of the term, and generating a verified candidate set of professional terms; For each term in the verified candidate set of professional terms, evaluating its relevance in the current query context based on semantic link relationships and sentence structures, calculating the context importance score, and screening and sorting to obtain a list of professional terms.

4. The method according to claim 2, wherein Obtaining context feature vectors, including: Based on the hierarchical semantic representation and the list of professional terms, performing the following steps: Identifying the set of professional fields to which the query belongs, calculating the distribution of membership probabilities for each field, and generating domain feature vectors; Analyzing the query purpose and considering the application scenario to generate purpose feature vectors; Retrieving relevant professional background knowledge from a pre-constructed professional knowledge base, extracting key concepts and relationships, and generating knowledge background vectors; Integrating the domain feature, purpose feature, and knowledge background vectors into context feature vectors.

5. The method according to claim 1, wherein Constructing a hierarchical semantic-visual representation framework, including: Encoding the user query data to obtain a basic query vector representation; splitting it into core concept, attribute modification, and relationship description representations, and integrating them into a hierarchical semantic representation; Extracting object-level, attribute-level, and relationship-level visual features from the images in the image library and integrating them into a hierarchical visual representation; Based on the hierarchical semantics and visual representation, establishing a strict correspondence relationship to form a hierarchical semantic-visual representation framework.

6. The method according to claim 5, wherein Integrating into a hierarchical semantic representation, including: For the basic query vector representation, identifying and verifying the subject and key entities, extracting core concepts and assigning weight values to generate core concept representations; For the basic query vector representation, identify adjectives, state words, and characteristic descriptions that modify the core concepts, construct attribute triples, and generate attribute modification representations; For the basic query vector representation, identify relational verbs, state changes, and temporal information among the core concepts, construct relational quadruples, and generate relational description representations; Integrate the core concept, attribute modification, and relational description representations into a hierarchical semantic representation.

7. The method according to claim 5, characterized in that, Integrate into a hierarchical visual representation, including: Identify the objects to be inspected from the image, extract the feature vectors of each object to be inspected, and generate object-level visual features; For the objects to be inspected, identify domain attributes, calculate attribute confidence scores, construct attribute triples, and generate attribute-level visual features; Based on the objects to be inspected and their attributes, identify the spatial and functional relationships among the objects to be inspected, construct a relationship graph among the objects to be inspected, and generate relationship-level visual features; Integrate the object-level, attribute-level, and relationship-level visual features into a hierarchical visual representation that strictly corresponds to the semantic hierarchy.

8. The method according to claim 1, characterized in that, Generate an optimized mapping matrix and an optimized candidate set, including: Convert the context-enhanced query representation into a visual query representation, generate multi-level visual attention maps, and adjust them based on the context feature vectors to guide visual feature matching and calculate the forward mapping similarity matrix; Based on the forward mapping similarity matrix, select a candidate image set and analyze the distribution characteristics of its visual features, construct a visual-semantic bridging network, project the visual features back to the semantic space, identify semantic differences, and dynamically adjust semantic understanding to generate a reverse mapping similarity matrix; Based on the forward and reverse mapping similarity matrices, perform two-way mapping optimization to generate an optimized mapping matrix, and obtain an optimized candidate set by screening the candidate image set.

9. The method according to claim 8, wherein Calculate the forward mapping similarity matrix, including: Input the context-enhanced query representation into the semantic-visual conversion network to generate object-level, attribute-level, and relationship-level visual query representations; Based on the visual query representation, calculate the similarity distribution of the hierarchical visual features in the image library and generate multi-level visual attention maps; Combine the context feature vectors, adaptively adjust the multi-level visual attention maps and guide visual feature matching to calculate the forward mapping similarity matrix.

10. The method according to claim 8, wherein Generate the reverse mapping similarity matrix, including: Analyze the visual feature distribution of the images in the candidate image set to identify the dominant visual patterns; accordingly, construct a visual-semantic bridging network, project the visual features into the semantic space, and generate reverse semantic descriptions; Compare the reverse semantic descriptions with the original query semantics to identify potential semantic differences; combine them with the context feature vectors and dynamically adjust the parameters of the semantic understanding model to generate the reverse mapping similarity matrix.

Citation Information

Patent Citations

  • Method for screening useful images from retrieved images

    CN103778227A

  • Semantic image retrieval method based on attention mechanism

    CN111782853A

  • Common service system cross-domain security interaction method and system based on general large model

    CN118761852A

  • Knotarization intelligent question and answer customer service method and system based on knowledge graph

    CN119938816A

  • Multimodal semantic analysis and image retrieval

    US20240354336A1

Cited By

  • Medical image data management method and system based on large language model

    CN121148734A

  • Medical image data management method and system based on large language model

    CN121148734B

  • Matching method of fields in table, electronic equipment and computer readable storage medium

    CN121279278A

  • Intelligent retrieval accurate matching method and system based on multi-level semantic decomposition

    CN121434472A

  • Image semantic coding and retrieval system based on visual truly

    CN121636742A