Similar image retrieval method

By constructing a hierarchical semantic-visual representation framework and semantic-visual bidirectional mapping, the problems of insufficient term understanding and insufficient mapping adaptability in professional fields are solved, and high accuracy and adaptability of professional image retrieval is achieved.

CN120256660BActive Publication Date: 2025-08-22JIANGSU YUNLAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510733788.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-08-22
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

When handling professional field query, the existing similar image search technology lacks the ability to understand professional terms and contextually perceive professional terms, resulting in inaccurate ambiguous analysis of terminology, and traditional semantic-visual mapping lacks adaptability and cannot optimize mapping relationships based on the distribution of actual image data.

Method used

Build a hierarchical semantic-visual representation framework, adopt a professional term understanding system of context perception, generate context feature vectors and hierarchical weight vectors, and realize multimodal similarity calculation and result sorting through semantic-visual bidirectional mapping and optimization framework, and output the search result set.

Benefits of technology

It improves the accuracy of retrieval of similar pictures in professional fields, ensures that the system can correctly understand professional terms based on professional background, and adjusts semantic understanding from visual feedback, improving the adaptability and accuracy of retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256660B_ABST
    Figure CN120256660B_ABST
Patent Text Reader

Abstract

The present invention provides a similar image retrieval method, comprising: constructing a hierarchical semantic-visual representation framework to decompose queries and images into three levels: core concepts, attribute modifications, and relationship descriptions; implementing a context-aware professional terminology understanding system to analyze the professional domain background and purpose of the query, generate context feature vectors, and resolve term ambiguity; establishing a semantic-visual bidirectional mapping and optimization framework to achieve forward mapping from semantics to visuals and reverse mapping from visuals to semantics, and generating an optimized mapping matrix and candidate set through iterative optimization; performing multimodal similarity calculations to convert context feature vectors into hierarchical weight vectors, which are then merged with multi-level similarities to form a comprehensive similarity score, which is then sorted and output as a set of search results. The present invention improves the accuracy of similar image retrieval in professional domains.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to information similarity comparison and retrieval, in particular to a similar image retrieval method. Background Art

[0002] Similar image retrieval technology is a core component of modern information processing systems and is widely used in areas such as paper duplication checking, medical diagnosis, security monitoring, and e-commerce. As professional fields demand higher precision in image retrieval, traditional retrieval methods based on visual features are no longer sufficient. In some fields, users often use queries containing a large number of specialized terms to find similar images. This requires the system to not only understand the visual content but also accurately grasp the specialized semantics and understand the query context and purpose to accurately retrieve specialized images.

[0003] Current similar image retrieval technologies are mainly divided into two categories: content-based image retrieval (CBIR) and semantic-based image retrieval. Content-based retrieval methods mainly use low-level visual features such as color, texture, and shape for matching. For example, the QBIC system proposed by Flickner et al. uses global feature vectors to represent image content. With the development of deep learning, feature extraction methods such as CNN and SIFT have greatly improved the expressive power of visual features. At the semantic level, researchers have proposed visual-semantic embedding models to narrow the semantic gap, such as the deep visual-semantic alignment model proposed by Karpathy et al., the hierarchical attention network developed by Yang et al., and the recently emerging cross-modal retrieval methods, such as large-scale visual-language pre-training models such as CLIP.

[0004] However, existing technologies still have several key issues that need to be addressed. First, when processing queries in professional fields, existing methods have a serious lack of understanding of professional terminology and lack of domain background knowledge and context perception capabilities, resulting in inaccurate resolution of term ambiguity. For example, in medical image retrieval, the same term may have completely different visual correspondences for different diagnostic purposes, and existing systems are unable to dynamically adjust their understanding based on the query purpose and domain background. Second, traditional semantic-visual mapping is mostly a one-way process, and visual features cannot provide feedback to adjust semantic understanding, lacking adaptability, resulting in the retrieval system being unable to optimize the mapping relationship based on the actual image data distribution. Summary of the Invention

[0005] Purpose of the invention: To provide a similar image retrieval method in order to solve at least one technical problem existing in the prior art.

[0006] Technical solution: Similar image retrieval method, including:

[0007] Receive user query data and build a hierarchical semantic-visual representation framework, including hierarchical semantic and visual representations;

[0008] Based on hierarchical semantic representation, a context-aware terminology understanding system is used to generate context feature vectors, hierarchical weight vectors, and context-enhanced query representations.

[0009] Based on context-enhanced query representation and hierarchical visual representation, a semantic-visual bidirectional mapping and optimization framework is adopted to generate an optimized mapping matrix and an optimized candidate set. Combined with the context feature vector, multimodal similarity calculation is performed to obtain multi-level similarity.

[0010] The multi-level similarity and the level weight vector are fused into a comprehensive similarity score, and the optimized candidate set is sorted based on it, and the retrieval result set is output.

[0011] Beneficial effects: The present invention provides comprehensive context for term disambiguation, enabling the system to correctly understand professional terms based on professional backgrounds; it not only maps from semantics to visuals, but also adjusts semantic understanding based on visual feedback, solving the problem of traditional one-way mapping that cannot adjust semantic understanding based on actual visual data; and it improves the accuracy of similar image retrieval in professional fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 A flowchart of the steps of a similar image retrieval method provided in an embodiment of the present application.

[0013] Figure 2 A flowchart of the steps for generating a context feature vector, a hierarchical weight vector, and a context-enhanced query representation provided in an embodiment of the present application.

[0014] Figure 3 A flowchart of the steps for obtaining a list of professional terms provided in an embodiment of the present application.

[0015] Figure 4 A flowchart of the steps for constructing a hierarchical semantic-visual representation framework provided in an embodiment of the present application. DETAILED DESCRIPTION

[0016] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0017] It should be noted that to clearly illustrate the steps of this application, serial numbers are assigned to each step in the specification. These serial numbers are for illustrative purposes only and do not limit the order in which the steps must be executed. In actual operation, depending on the technical requirements of the specific implementation scenario, the steps may be executed in a different order than shown in the specification, and in some cases, parallel processing between steps may be implemented.

[0018] During the research process, it was found that the current systems mostly use holistic feature representation, lack of detailed modeling of different semantic levels such as objects, attributes and relationships, and cannot accurately match and calculate similarity for the hierarchical semantic structures in professional fields, which further limits the retrieval accuracy and applicability.

[0019] like Figure 1 As shown in FIG, a similar image retrieval method is proposed, which includes the following steps:

[0020] S1. Receive user query data and build a hierarchical semantic-visual representation framework, including hierarchical semantic and visual representations;

[0021] Specifically, user query data can be text, images, or a combination of both. Hierarchical semantics analyzes textual information, dividing it into different levels, such as main topics, subtopics, and detailed information, allowing the system to more clearly understand the organizational structure of information. Visual representation processes images or visual content, which may include identifying different parts of an image and analyzing the relationships between them, making it easier to understand the content of the image.

[0022] S2. Based on hierarchical semantic representation, a context-aware terminology understanding system is used to generate context feature vectors, hierarchical weight vectors, and context-enhanced query representations.

[0023] Specifically, a context-aware terminology understanding system means that the system not only recognizes individual words but also understands the terminology in context. For example, "deep learning" may be just a concept in ordinary conversation, but in a technical paper, it may refer to a specific algorithm or model. By analyzing the context, the system ensures accurate understanding of the terminology.

[0024] S3, based on context-enhanced query representation and hierarchical visual representation, adopts a semantic-visual bidirectional mapping and optimization framework to generate an optimized mapping matrix and an optimized candidate set, and combines them with the context feature vector to perform multimodal similarity calculation to obtain multi-level similarity;

[0025] Specifically, bidirectional semantic-visual mapping connects text and images, allowing language to understand visual information and vision to be mapped to language. By continuously adjusting the matching method, the mapping effect is improved, making the conversion between text and image more accurate.

[0026] S4. Fusion of multi-level similarities and level weight vectors into a comprehensive similarity score, sorting of the optimized candidate set based on the comprehensive similarity score, and output of the search result set.

[0027] Specifically, by integrating multi-level similarity and level weights, a comprehensive score is calculated to indicate the degree of match between a candidate and the query. A higher comprehensive score indicates a stronger relevance of the candidate to the query. All candidates are sorted based on the comprehensive similarity score, with the most relevant candidates placed first, ensuring that users see the most relevant results first.

[0028] like Figure 2 As shown, according to one aspect of the present application, generating a context feature vector, a hierarchical weight vector, and a context-enhanced query representation includes:

[0029] Identify professional terms from the hierarchical semantic representation and obtain a professional term list; combine the professional term list with the hierarchical semantic representation to analyze the professional domain background and purpose of the query, generate domain feature, purpose feature and knowledge background vectors, and fuse them into a context feature vector;

[0030] The context feature vector is input into the weight generation network to dynamically generate a hierarchical weight vector based on the professional domain characteristics and purpose of the query;

[0031] Based on the professional term list and context feature vector, term ambiguity is resolved to obtain the disambiguated semantic representation; the disambiguated semantic representation is fused with the context feature vector to generate a context-enhanced query representation.

[0032] like Figure 3 As shown, according to one aspect of the present application, a list of professional terms is obtained, including:

[0033] Match the vocabulary in the hierarchical semantic representation with the pre-trained domain terminology vocabulary to generate a preliminary set of domain terminology candidates;

[0034] Verify the domain adaptability of each term in the preliminary professional term candidate set, calculate the domain belonging probability of the term, and generate a verified professional term candidate set;

[0035] For each term in the verified professional term candidate set, its relevance in the current query context is evaluated based on the semantic link relationship and sentence structure, the context importance score is calculated, and the term list is filtered and sorted to obtain the professional term list.

[0036] According to one aspect of the present application, obtaining a context feature vector includes:

[0037] Based on the hierarchical semantic representation and the term list, the following steps are performed:

[0038] Identify the set of professional fields to which the query belongs, calculate the probability distribution of belonging to each field, and generate the field feature vector;

[0039] Analyze the query purpose and consider the application scenario to generate the purpose feature vector;

[0040] Retrieve relevant professional background knowledge from the pre-built professional knowledge base, extract key concepts and relationships, and generate knowledge background vectors;

[0041] The domain features, purpose features and knowledge background vectors are integrated into the context feature vector.

[0042] like Figure 4 As shown, according to one aspect of the present application, a hierarchical semantic-visual representation framework is constructed, including:

[0043] Encode user query data to obtain a basic query vector representation; split the basic query vector representation into core concept representation, attribute modification representation, and relationship description representation, and integrate them into a hierarchical semantic representation;

[0044] Extracting object-level visual features, attribute-level visual features, and relationship-level visual features from images in an image library and integrating them into a hierarchical visual representation;

[0045] Based on hierarchical semantic representation and hierarchical visual representation, a strict correspondence is established to form a hierarchical semantic-visual representation framework.

[0046] According to one aspect of the present application, the hierarchical semantic representation is integrated, including:

[0047] Based on the basic query vector representation, identify and verify the subject and key entities, extract core concepts and assign weight values ​​to generate core concept representation;

[0048] Based on the basic query vector representation, identify the adjectives, status words and feature descriptions that modify the core concept, construct attribute triples, and generate attribute modification representations;

[0049] For the basic query vector representation, identify the relational verbs, state changes, and time sequence information between core concepts, construct relation quadruples, and generate relational description representations;

[0050] Integrate core concepts, attribute modifications and relationship descriptions into a hierarchical semantic representation.

[0051] According to one aspect of the present application, the hierarchical visual representation is integrated, including:

[0052] Identify the objects to be inspected from the image, extract the feature vector of each object to be inspected, and generate object-level visual features;

[0053] For the object to be inspected, identify domain attributes, calculate attribute confidence scores, construct attribute triples, and generate attribute-level visual features;

[0054] Based on the objects to be inspected and their attributes, the spatial and functional relationships between the objects to be inspected are identified, a relationship graph between the objects to be inspected is constructed, and relation-level visual features are generated;

[0055] Object-level visual features, attribute-level visual features, and relationship-level visual features are integrated into a hierarchical visual representation that strictly corresponds to the semantic hierarchy.

[0056] According to one aspect of the present application, generating an optimized mapping matrix and an optimized candidate set includes:

[0057] Convert the context-enhanced query representation into a visual query representation, generate a multi-level visual attention map, and adjust the multi-level visual attention map based on the context feature vector to guide visual feature matching and calculate the forward mapping similarity matrix;

[0058] Based on the forward mapping similarity matrix, a candidate image set is selected and its visual feature distribution characteristics are analyzed. A visual-semantic bridging network is constructed to project the visual features back into the semantic space, identify semantic differences, and dynamically adjust the semantic understanding to generate a reverse mapping similarity matrix.

[0059] Based on the forward and reverse mapping similarity matrices, bidirectional mapping optimization is performed to generate an optimized mapping matrix, and the optimized candidate set is obtained by screening the candidate image set.

[0060] Specifically, the object-level similarity between the query object representation and the object-level visual features, the attribute-level similarity between the query attribute representation and the attribute-level visual features, and the relationship-level similarity between the query relationship representation and the relationship-level visual features are calculated;

[0061] The object-level similarity, attribute-level similarity and relationship-level similarity are combined to generate an optimized mapping matrix, and the optimized candidate set is obtained based on the preliminary matching results.

[0062] According to one aspect of the present application, generating an optimized mapping matrix and an optimized candidate set includes:

[0063] Read the query object representation, query attribute representation and query relationship representation corresponding to the visual features;

[0064] And match the object-level visual features, attribute-level visual features, and relationship-level visual features in the image library.

[0065] The object-level similarity, attribute-level similarity and relationship-level similarity are calculated respectively; based on this, an optimized mapping matrix is ​​generated, and the optimized candidate set is obtained by screening according to the preliminary matching results.

[0066] According to one aspect of the present application, calculating a forward mapping similarity matrix includes:

[0067] The context-enhanced query representation is fed into a semantic-visual transformation network to generate visual query representations at the object level, attribute level, and relation level.

[0068] Based on the visual query representation, the similarity distribution of the hierarchical visual features in the image library is calculated and a multi-level visual attention map is generated;

[0069] Combined with the context feature vector, the multi-level visual attention map is adaptively adjusted to guide visual feature matching and calculate the forward mapping similarity matrix.

[0070] According to one aspect of the present application, generating a reverse mapping similarity matrix includes:

[0071] Perform visual feature distribution analysis on images in the candidate image set to identify dominant visual patterns;

[0072] Based on the dominant visual pattern, a visual-semantic bridging network is constructed to project visual features into the semantic space and generate reverse semantic descriptions.

[0073] The reverse semantic description is compared with the original query semantics to identify potential semantic differences; the semantic differences are combined with the context feature vector, the parameters of the semantic understanding model are dynamically adjusted, and a reverse mapping similarity matrix is ​​generated.

[0074] According to one aspect of the present application, performing multimodal similarity calculation and outputting a final search result set includes:

[0075] Obtain object-level similarity, attribute-level similarity, and relationship-level similarity from the optimized candidate set to form a multi-level similarity;

[0076] Multiply the multi-level similarities with the level weight vector and sum them to generate a comprehensive similarity score for each candidate image;

[0077] The images in the optimized candidate set are sorted in descending order according to the comprehensive similarity scores, the N images with the highest scores are selected as the final retrieval result set, and the similarity explanation information of each result image is generated.

[0078] According to one aspect of the present application, the step of generating a comprehensive similarity score includes:

[0079] Based on the context feature vector, analyze the focus and importance distribution of the current query in the professional field;

[0080] Applying a dynamic weight generator, object-level weights, attribute-level weights, and relationship-level weights are generated according to the focus, importance distribution, and query purpose to form a hierarchical weight vector.

[0081] Identify and smooth outliers and noise in multi-level similarities to ensure the stability of similarity distribution;

[0082] A weighted fusion algorithm is used to multiply the smoothed multi-level similarities with the level weight vector and sum them up, while considering the interaction between levels to generate a comprehensive similarity score.

[0083] The present invention provides a similar image retrieval method, comprising: constructing a hierarchical semantic-visual representation framework to decompose queries and images into three levels: core concepts, attribute modifications, and relationship descriptions; implementing a context-aware professional terminology understanding system to analyze the professional domain background and purpose of the query, generate context feature vectors, and resolve term ambiguity; establishing a semantic-visual bidirectional mapping and optimization framework to achieve forward mapping from semantics to vision and reverse mapping from vision to semantics, and generating an optimized mapping matrix and candidate set through iterative optimization; performing multimodal similarity calculation to convert the context feature vectors into hierarchical weight vectors, which are then merged with the multi-level similarity to form a comprehensive similarity score, and the retrieval result set is sorted and output accordingly. The present invention improves the accuracy of similar image retrieval in professional domains.

[0084] According to one aspect of the present application, a similar image retrieval method is proposed, and the implementation process of its context-aware semantic-visual bidirectional mapping framework is as follows:

[0085] S1: Construct a hierarchical semantic-visual bimodal representation framework. This framework receives user query text and uses a hierarchical semantic decomposition algorithm to decompose it into three levels: core concepts, attribute modifications, and relationship descriptions. Simultaneously, a corresponding visual feature extraction network is applied to the image library to establish a strict correspondence between semantic levels and visual features, generating both hierarchical semantic representations and hierarchical visual representations. This framework transforms the traditional holistic feature representation into multi-level visual features that precisely correspond to the semantic structure, resolving the fundamental issue of unclear semantic-visual correspondence in specialized domains.

[0086] S11: Receive the user's professional query text and use the Transformer encoder model for preliminary encoding to obtain the basic query vector representation.

[0087] S12: Apply a three-level structured decomposition algorithm to the basic query vector representation, splitting the query into three semantic levels:

[0088] S121: Core Concept Extraction. This process reads the base query vector representation and applies an attention mechanism to identify the subject and key medical entities. The identified entities are verified using a medical ontology knowledge base, filtering out non-core concepts. A semantic importance scoring algorithm is applied to assign weights to each identified core concept. The system then outputs a weighted core concept representation, consisting of entity vectors and their associated weights.

[0089] S122: Attribute Modifier Extraction. This process reads the base query vector representation and identifies adjectives, status terms, and feature descriptions that modify the core concept. It then applies dependency parsing to identify the relationships between these modifiers and the core concept. For each modifier, it constructs a triplet of (attribute type, attribute value, associated core concept). Finally, it outputs an integrated attribute modification representation, including the modifier vector and its relationship to the core concept.

[0090] S123: Relation Description Extraction. This process reads the basic query vector representation and identifies the relational verbs, state changes, and temporal information between core concepts. It applies semantic role annotation to identify the subject, object, and their relationship type. It constructs a quadruple of (subject concept, relationship type, object concept, temporal information). It applies specialized relationship recognition models to specific medical relationships (such as treatment effect and pathological response). It then outputs an integrated relational description representation, including the relationship vector and its connected concepts.

[0091] S13: Integrate the core concept representation, attribute modification representation, and relationship description representation into a hierarchical semantic representation, maintaining the association structure between the levels.

[0092] S131: Construct a hierarchical association graph. Using entities in the core concept representation as nodes, establish an initial concept graph. Based on the attribute modification representation, add attribute edges and attribute values ​​to each node. Based on the relationship description representation, add relationship edges and relationship attributes between concept nodes. Generate a complete semantic hierarchical association graph and save the vector representations of the nodes and edges.

[0093] S132: Optimize hierarchical representation compatibility. Standardize the vector representations of the three layers to ensure dimensionality consistency; apply a hierarchical fusion algorithm to make representations at different layers comparable in the feature space; construct an inter-layer attention mechanism to enable high-level semantics to focus on relevant features at lower layers; and output a final hierarchical semantic representation consisting of the three layers of vectors and their associated structure.

[0094] S14: For each medical image in the image library, a hierarchical visual feature extraction network is applied to achieve visual feature decomposition that strictly corresponds to the semantic level.

[0095] S141: Object-Level Feature Extraction. This module reads medical images and applies the Faster R-CNN model to identify medical objects (cells, tissues, organs, etc.) within the image. For each detected object, a domain-specific feature extractor is used to extract a feature vector. The Medical Object Recognition Enhancement Module is applied to fuse texture, shape, and context information. Multi-scale feature fusion is applied to handle medical objects of varying sizes. The module outputs the location information and feature vector of each object, forming an object-level visual feature.

[0096] S142: Attribute-level feature extraction. Each detected medical object region is read and an attribute classification network is applied to identify medical attributes (such as cell morphology and tissue density). A confidence score is calculated for each attribute, constructing a triplet of (attribute type, attribute value, confidence score). A multi-attribute joint reasoning module is applied to handle correlations and mutual exclusivity between attributes. Global image features are combined to enhance the accuracy of local attribute recognition. The attribute representation vector for each object is output to form an attribute-level visual feature.

[0097] S143: Relationship-level feature extraction. The detected medical objects and their attributes are read, and a relational reasoning network is applied to identify the spatial and functional relationships between objects. A relationship graph is constructed between objects, representing their spatial positional relationships, inclusion relationships, and functional relationships. A temporal analysis module (for video or multi-frame images) is applied to identify state change relationships. Medical prior knowledge is used to guide relationship reasoning and improve the accuracy of professional relationship recognition. A representation vector of the relationship between objects is output to form a relation-level visual feature.

[0098] S15: Organize the object-level visual features, attribute-level visual features, and relationship-level visual features of each image into a hierarchical visual representation and establish a visual feature index library.

[0099] S151: Constructing an image hierarchy. This involves integrating the three layers of features for each image to construct a visual hierarchy corresponding to the semantic level. Feature normalization is applied to ensure scale consistency across different layers of features. Attention connections are constructed between layers, enabling high-level features to guide the attention of lower-level features. The image hierarchy is then output for each image.

[0100] S152: Establish an efficient index structure. Apply a locality-sensitive hashing algorithm to the hierarchical representation of all images to construct a fast retrieval index. Establish separate indexes for each semantic level to support hierarchical retrieval. Apply inverted indexing technology to support fast concept- and attribute-based searches. Output a complete visual feature index library for subsequent retrieval.

[0101] S2: Implements a context-aware terminology understanding system. This system processes terminology within a hierarchical semantic representation, analyzes the query domain, purpose, and implicit knowledge context, constructs a contextual feature vector, and uses this to resolve terminology ambiguity and generate a context-enhanced query representation. This system goes beyond literal understanding and interprets queries within a professional context. This addresses the critical issue of terminology changing meaning across different scenarios, improving the system's ability to understand complex professional queries.

[0102] S21: Receive the hierarchical semantic representation, identify the professional terms therein, and obtain a professional term list.

[0103] S211: Terminology Detection. Read all words and phrases in the hierarchical semantic representation; apply a medical terminology recognition model to match them against the terminology knowledge base; use contextual word co-occurrence features to enhance the accuracy of term boundary identification; and output all detected terminology and its location in the query.

[0104] S212: Marking Polysemous Terms. This process reads detected professional terms and queries a professional term ambiguity database. It marks terms with multiple possible meanings and lists their possible meanings. It calculates the degree of ambiguity for each term and assigns different processing priorities for subsequent resolution. It then outputs a list of professional terms with ambiguity markers, including the set of possible meanings for each term.

[0105] S22: Analyze the professional field background and purpose of the query.

[0106] S221: Domain Identification. Read the hierarchical semantic representation and the domain term list; extract term co-occurrence features and compare them with pre-trained domain feature templates; apply the domain classification model to calculate the probability distribution of the query belonging to each domain; apply hierarchical domain classification, from coarse-grained (such as "medicine") to fine-grained (such as "oncology"); output the query domain feature vector, which contains the main domain and possible subdomain information.

[0107] S222: Query Purpose Analysis. This step reads the hierarchical semantic representation and identifies query intent indicators (e.g., "diagnosis," "analysis," and "comparison"). The query's grammatical structure and question type are analyzed to infer the query's purpose. A query intent classification model is applied to identify queries with diagnostic needs, research and analysis, or educational purposes. The query's level of detail and expertise is analyzed to assess the user's expertise. The query's purpose feature vector is output, containing information about the primary and additional purposes.

[0108] S223: Knowledge Background Extraction. This process reads the hierarchical semantic representation and the list of specialized terms; applies knowledge graph reasoning to identify the underlying knowledge requirements of the query; analyzes the degree of specialized terminology used to assess the depth of required domain knowledge; constructs a knowledge dependency graph for the query, representing the knowledge structure required to understand the query; and outputs a knowledge background vector for the query, representing the knowledge premises and background assumptions of the query.

[0109] S23: Fusion of domain feature vector, purpose feature vector and knowledge background vector into context feature vector.

[0110] S231: Feature Vector Normalization. Read the three feature vectors, apply dimension unification and range normalization, apply a feature selection algorithm to retain the most discriminative feature dimensions, and output the three normalized feature vectors.

[0111] S232: Contextual Feature Fusion. This approach applies an attention-weighted fusion mechanism to dynamically adjust the weights of the three vectors based on query characteristics. It also constructs an interaction layer between features to capture the mutual influence of domain, purpose, and knowledge background. It applies context consistency checks to ensure the rationality of the fusion results. Finally, it outputs a final contextual feature vector containing complete context information.

[0112] S24: Disambiguate terminology based on the professional term list and context feature vector.

[0113] S241: Build a term conditional probability model. Read a list of professional terms and context feature vectors; for each polysemous term, retrieve its usage frequency statistics in different contexts; build a Bayesian conditional probability model and calculate P (meaning is context feature); apply domain-specific prior knowledge to adjust the probability distribution; and output the probability distribution of each term's meaning in the current context.

[0114] S242: Contextual Feature Matching. Read the list of possible meanings and the contextual feature vector for each polysemous term; construct a feature representation for each possible meaning and calculate similarity with the query context features; apply a hierarchical similarity calculation, considering domain similarity, purpose relevance, and knowledge consistency; and output a contextual matching score for each meaning of the term.

[0115] S243: Collaborative Term Network Disambiguation. This process reads all terms in the query, along with their meaning probability distributions and contextual matching scores. It constructs a semantic association network between terms, representing their dependencies and mutual support. It applies a graph reasoning algorithm to ensure that semantically related terms tend to have semantically consistent meanings. It iteratively optimizes the disambiguation process until the term meaning network reaches a stable state. It then outputs the optimal meaning for each term, along with its confidence score.

[0116] S25: The disambiguated hierarchical semantic representation is fused with the context feature vector to generate a context-enhanced query representation.

[0117] S251: Update Semantic Representation. Read the original hierarchical semantic representation and the optimal meaning selection for the terms; update the vector representations of the terms in each hierarchy, replacing them with specialized representations of the selected meanings; adjust the representations of other concepts related to these terms to ensure semantic consistency; and output the updated, precise semantic representation.

[0118] S252: Context-enhanced fusion. This process reads the precise semantic representation and context feature vectors; applies context-conditioned feature transformation to adjust the dimensions of the semantic representation based on the context; applies different context-enhanced strategies to different semantic levels to construct a hierarchical context representation; embeds context information into the semantic structure to enhance the query representation's adaptability to different environments; and outputs the final context-enhanced query representation for subsequent mapping.

[0119] S3: Establish a semantic-visual bidirectional mapping and optimization framework. Based on context-enhanced query representation and hierarchical visual representation, a bidirectional interaction mechanism between semantics and visual features is constructed. First, a forward mapping is implemented, where semantics guides visual feature activation. Then, semantic understanding is reversely adjusted based on the distribution of visual features. Through iterative optimization, an optimized mapping matrix and an optimized candidate set are generated. The key is to break the limitations of traditional one-way mapping and achieve dynamic feedback between semantic understanding and visual representation. This allows the system to adjust query understanding based on actual image data, improving retrieval accuracy and adaptability.

[0120] S31: Receive the context-enhanced query representation and hierarchical visual representation, and initialize the bidirectional mapping matrix.

[0121] S311: Hierarchical Mapping Initialization. This step reads the structural information of the context-enhanced query representation and the hierarchical visual representation; constructs initial mapping matrices for each of the three semantic levels, establishing initial correspondences between semantic and visual features; applies domain transfer learning, initializing mapping weights using pre-trained medical semantic-visual correspondence knowledge; and outputs an initial set of hierarchical mapping matrices, including three matrices: core concept mapping, attribute mapping, and relationship mapping.

[0122] S312: Establish inter-layer mapping associations. Read the hierarchical mapping matrix group and build a joint mapping mechanism between layers. Build a hierarchical feedback channel so that high-level semantic mappings can guide the attention allocation of low-level mappings. Apply overall consistency constraints to ensure that the mapping results of different layers are coordinated. Output a complete bidirectional mapping matrix, including the intra-layer mapping and inter-layer association structure.

[0123] S32: Implementing semantic-to-visual forward mapping.

[0124] S321: Generate a multi-level visual attention map. This process reads the context-enhanced query representation and bidirectional mapping matrix. For the core concept layer, a mapping transformation is applied to generate a target region attention map, highlighting relevant medical objects. For the attribute modification layer, a feature attention map is generated, highlighting relevant visual attribute regions. For the relationship description layer, a relationship attention map is generated, highlighting interaction regions between objects. The system then outputs a three-layer visual attention map, representing the mapping of different semantic levels in visual space.

[0125] S322: Adjust attention weights based on context. Read the visual attention map and context feature vector; adjust the attention weights for different types of medical objects based on query domain characteristics; adjust the attention allocation of attribute features based on query purpose characteristics (e.g., focusing more on abnormal features for diagnostic purposes); adjust the professional sensitivity of relational attention based on knowledge background characteristics; and output a context-weighted attention map reflecting the attention allocation in the current professional context.

[0126] S323: Perform visual feature activation and retrieval. Read the context-weighted attention map and visual feature index library; apply attention-guided feature activation to assign activation strength to each feature in the visual library; calculate weighted feature similarities across three levels to generate a preliminary hierarchical similarity score; apply hierarchical weighted fusion to combine the three similarities into an overall similarity score; sort based on the overall similarity and select the top N most similar images; output a set of candidate images and their initial similarity scores.

[0127] S33: Implementing reverse mapping from vision to semantics.

[0128] S331: Visual Feature Distribution Analysis. This involves reading the hierarchical visual features of a candidate image set; applying cluster analysis to identify visual patterns and feature distribution patterns within the candidate set; identifying high-frequency visual features that may represent the core visual representation of the query; identifying differential visual features that may represent ambiguous representations of the query; and constructing a visual feature distribution map to represent the spatial distribution of visual features in the search results.

[0129] S332: Feature Space Mapping Analysis. This function reads the visual feature distribution map and the current bidirectional mapping matrix; analyzes the mapping of semantic features to visual features, identifying areas of inaccurate or incomplete mapping; detects semantic over-mapping (where a semantic concept is mapped to too many irrelevant visual features); and detects semantic under-mapping (where a semantic concept is not mapped to relevant visual features). The function then outputs a mapping quality assessment report, including any mapping areas that require adjustment and the direction of such adjustments.

[0130] S333: Dynamic Adjustment of Semantic Understanding. This system reads the visual feature distribution map and mapping quality assessment report; based on the visual feature distribution, it dynamically adjusts the importance weights of each concept in the hierarchical semantic representation. It increases the weights of under-mapped semantic concepts and decreases the weights of over-mapped concepts. It introduces new semantic concepts to represent features found in the visual features but not explicitly defined in the original query. It suppresses semantic concepts not reflected in the visual features to reduce their influence in the mapping. It then outputs an adjusted adaptive semantic representation that better matches the actual visual feature distribution.

[0131] S334: Update the bidirectional mapping relationship. Read the adaptive semantic representation and the original bidirectional mapping matrix; apply the gradient descent algorithm to adjust and optimize the mapping matrix parameters based on the semantics; adjust the mapping weights of different semantic levels to reflect their actual importance in visual representation; optimize the association structure between levels to enhance the overall consistency of the semantic-visual mapping; output the updated bidirectional mapping matrix, which reflects the impact of visual feedback on semantic understanding.

[0132] S34: Execute an iterative optimization loop.

[0133] S341: Recalculate semantic-visual similarity. Read the updated bidirectional mapping matrix and adaptive semantic representation; repeat step S323, using the updated mapping to calculate a new similarity score; compare it with the initial similarity score to evaluate the optimization effect; output the optimized similarity score and the updated candidate image set.

[0134] S342: Evaluate optimization results. Calculate changes in similarity distribution before and after optimization; analyze the update level of the candidate set, calculate set similarity and ranking changes; evaluate the convergence stability of the mapping matrix and calculate the magnitude of parameter changes; and output an optimization results evaluation report, including improvement indicators and convergence status.

[0135] S343: Determine whether to continue iteration. Read the optimization effect evaluation report and check whether the preset convergence conditions have been met; check whether the number of iterations has reached the upper limit; and check whether the similarity improvement is below the threshold. If the continuation conditions are met, return to S32 and enter the next round of iteration. If the convergence conditions are met, output the final optimized mapping matrix and the optimized candidate set.

[0136] S4: Perform multimodal similarity calculation and result ranking. Using the optimized mapping matrix, the optimized candidate set, and the contextual feature vector, the similarities between the three semantic levels and the corresponding visual features are calculated, resulting in core concept similarity, attribute modification similarity, and relationship description similarity. Based on the contextual feature vector, a hierarchical weight vector is dynamically generated. The three similarities are weighted and fused into a comprehensive similarity score, which is then sorted and output as the final search result set. This contextual understanding directly influences the result ranking strategy, ensuring that search results better meet user expectations in specific professional scenarios. The search information is also stored for continuous learning by the system.

[0137] S41: Receive the optimized mapping matrix, the optimized candidate set, and the context feature vector. Prepare the final similarity calculation resources. Read the optimized mapping matrix, the optimized candidate set, and the context feature vector; verify data integrity to ensure all necessary information is available; optimize memory allocation and computing resources for large-scale similarity calculations; prepare similarity normalization parameters to ensure comparability of scores at different levels; and output a calculation-ready status, including the optimized calculation parameters.

[0138] S42: For each candidate image, calculate the feature similarity of three levels.

[0139] S421: Core Concept Similarity Calculation. Read the object-level visual features of each candidate image and the core concept representation of the query; apply the concept mapping portion of the optimized mapping matrix to align the semantic space and the visual feature space; calculate the cosine similarity after alignment as the base similarity; apply structural similarity enhancement to consider the matching degree of the concept hierarchy; apply professional concept weighting to highlight the matching importance of key medical concepts; output the core concept similarity of each image, indicating the object-level matching degree.

[0140] S422: Attribute Modification Similarity Computation. This algorithm reads the attribute-level visual features of each candidate image and the attribute modification representation of the query. It applies the attribute mapping portion of the optimization mapping matrix to align the attribute semantics with the visual features. It constructs an attribute matching matrix and calculates the similarity between each pair of attributes. It applies the Hungarian algorithm to find the optimal attribute matching combination to maximize the overall matching score. It considers differences in attribute importance and assigns higher weights to key medical attributes. It outputs the attribute modification similarity for each image, indicating the degree of attribute-level matching.

[0141] S423: Relational Description Similarity Calculation. Read the relational-level visual features of each candidate image and the relational description representation of the query; apply the relational mapping portion of the optimized mapping matrix to align the relational semantics and visual relational features; construct a relational graph matching problem and compare the structural similarity between the query relational graph and the image relational graph; apply a graph matching algorithm, considering both node and edge similarity; apply expert rules to enhance matching accuracy for specific medical relationships (e.g., pathological reaction chains, treatment effect relationships); output the relational description similarity for each image, indicating the degree of relational-level matching.

[0142] S43: Based on the context feature vector, dynamically determine the weight coefficients of the three levels of similarity and generate a level weight vector.

[0143] S431: Contextual Condition Weight Generation. This process reads the contextual feature vector and analyzes the query domain, purpose, and knowledge background. A weight generation strategy library is applied to determine basic weights based on best practices in different professional fields. For example, pathology diagnosis queries assign higher weights to the attribute layer, while anatomical structure queries assign higher weights to the relationship layer. Weights are adjusted based on the query purpose; research queries focus more on relationships, while educational queries pay more attention to all layers. Professional sensitivity is adjusted based on the knowledge background; professional user queries place greater emphasis on detailed attribute matching. Preliminary contextual condition weights are then output.

[0144] S432: Weight Adaptive Optimization. This process reads the contextual weights and the three-layer similarity score distribution of the candidate images; analyzes the discriminative power of the three-layer similarity scores and increases the weights of layers with strong discriminative power; detects deviations caused by extreme weight distributions and applies balancing factors to ensure that each layer has an appropriate impact; applies task-adaptive adjustment based on the historically optimal weight distributions for similar tasks; and outputs a final layer weight vector containing the optimal weight distributions for the three layers.

[0145] S44: The core concept similarity, attribute modification similarity and relationship description similarity are weightedly fused with the hierarchical weight vector to obtain a comprehensive similarity score.

[0146] S441: Weighted Fusion Calculation. Read the three-layer similarity scores and layer weight vectors for each candidate image; apply linear weighted fusion to calculate a preliminary weighted score; apply nonlinear adjustments to address inter-layer complementarity and redundancy; for example, when the core concept matches extremely well, appropriately reduce the weight requirements for attribute matching; apply threshold constraints to ensure minimum matching requirements for key layers; and output a comprehensive similarity score for each candidate image.

[0147] S442: Confidence Assessment. Analyze the consistency of the three similarity levels and calculate the confidence level of the similarity judgment. Apply entropy measurement to assess the certainty of the similarity distribution. Add a confidence indicator to each composite similarity score to indicate the reliability of the similarity judgment. Output the final similarity score with the confidence level.

[0148] S45: Sort the optimized candidate set according to the comprehensive similarity score to obtain the final retrieval result set.

[0149] S451: Similarity Sorting and Filtering. Read the final similarity scores of all candidate images; apply a sorting algorithm to sort them from high to low similarity; apply a similarity threshold to filter and remove results with low similarity; apply diversity optimization to ensure that the result set contains relevant images with different visual representations; and output the sorted candidate result set.

[0150] S452: Result set enhancement processing. Read the candidate result set and apply the result grouping algorithm to cluster them by visual feature similarity. Select the most representative image for each group as the main result. Add result metadata, including the matched key concepts, prominent visual features, and confidence information. Generate result explanation data to explain the matching reasons and key features of each result. Output the structured final search result set.

[0151] S46: Return the final search result set to the user, and save the context and feedback information of this search for continuous learning of the system.

[0152] S461: Prepare for result presentation. Read the final search result set and optimize the result display format. Adjust the depth of expertise in the result interpretation based on the expertise level in the context feature vector. Generate a summary of the results, highlighting common features and key differences. Prepare for visualization enhancements, such as highlighting matching regions and labeling key features. Output user-oriented enhanced search results and prepare for presentation to the user.

[0153] S462: Retrieval Session Recording and Learning. Record complete retrieval session data, including query, context, mapping matrix, and result set; store intermediate optimization data during the retrieval process for subsequent analysis; prepare a user feedback collection mechanism to obtain result relevance evaluation; update the system knowledge base, including term mappings, visual features, and common query patterns; and output retrieval session records for continuous system optimization and personalized adaptation.

[0154] Case 1: The application scenario is the duplicate retrieval of academic paper images. With the publication of a large number of academic papers, the phenomenon of repeated use of images in papers has become increasingly serious, including unintentional management negligence and intentional academic misconduct. Traditional image retrieval methods have difficulty in accurately identifying similar images in academic professional fields, especially when involving professional terms and complex charts, the accuracy is insufficient. This example demonstrates how to use the method of the present invention to achieve high-precision academic image similarity retrieval and effectively identify image duplication problems in papers. This example is implemented in the following environment: the processor is Intel Xeon E5-2680 v4CPU, the memory is 128GB DDR4, the GPU is NVIDIA Tesla V100, the storage device is 1TB NVMe SSD, the operating system is Ubuntu 20.04 LTS, and the main development frameworks are Python 3.8 and PyTorch 1.9.0.

[0155] Step 1: Construct a hierarchical semantic-visual representation framework.

[0156] 1.1. Receive a user query. The user submits the query: "Find images similar to this fluorescent staining image of apoptosis, focusing on changes in mitochondrial membrane potential." The query consists of an image Q and a text description T.

[0157] 1.2. Hierarchical decomposition of the query. The query text T is processed using a BiLSTM and attention mechanism to obtain the basic query vector representation: BQ = BiLSTM(T) = [0.72, 0.35, 0.91, 0.42, 0.58, 0.77, 0.29, 0.84, 0.61, 0.45]; where BiLSTM() is a bidirectional long short-term memory network function; T is the query text "Find images similar to this fluorescent staining image of apoptosis, focusing on changes in mitochondrial membrane potential"; and BQ is the 10-dimensional basic query vector representation.

[0158] 1.2.1、Core concept extraction. Core concept representation CC = ExtractConcepts(BQ) = {(c1, w1),(c2, w2), ..., (c n , w n )}; where: c1="apoptosis", w1=0.92; c2="fluorescence staining", w2=0.87; c3="mitochondria", w3=0.94; c4="membrane potential", w4=0.91; ExtractConcepts() is the core concept extraction function; BQ is the basic query vector; w is the concept weight, which indicates the importance of the concept in the query, and its value range is [0,1].

[0159] Verification through the ontology knowledge base confirmed that these concepts are valid professional terms in the field of cell biology.

[0160] 1.2.2. Attribute modification extraction. Attribute modification is represented by AM = ExtractAttributes(BQ) = {(a1, v1, c1), (a2, v2, c2), ...}; where: a1 = "staining type", v1 = "fluorescence", c1 = "apoptosis"; a2 = "change type", v2 = "potential change", c2 = "mitochondrial membrane"; a3 = "image type", v3 = "microscopic", c3 = "apoptosis"; ExtractAttributes() is the attribute extraction function; BQ is the base query vector; a represents the attribute type; v represents the attribute value; and c represents the associated core concept.

[0161] 1.2.3. Relation Description Extraction. Relation description representation RD = ExtractRelations(BQ) = {(c1, r, c2, t), ...}; where: c1 = "mitochondria", r = "display", c2 = "membrane potential change", t = "in process"; ExtractRelations() is the relationship extraction function; BQ is the basic query vector; c1 is the subject concept; r is the relationship type; c2 is the object concept; and t is the time series information.

[0162] 1.3. Integrate the hierarchical semantic representation: Integrate the core concept representation CC, attribute modification representation AM, and relationship description representation RD into the hierarchical semantic representation HSR.

[0163] 1.4. Extract image visual features. For the query image Q and all images in the image library {I1, I2, ..., I k Apply hierarchical visual feature extraction:

[0164] 1.4.1. Object-level feature extraction. Object-level feature OLF = FastRCNN(I) = {(o1, v1), (o2, v2), ...}; where: o1 = "nucleus", v1 = [0.72, 0.35, 0.91, 0.42]; o2 = "mitochondria", v2 = [0.58, 0.77, 0.29, 0.84]; o3 = "cytoplasm", v3 = [0.61, 0.45, 0.33, 0.79]; FastRCNN() is the improved fast regional convolutional neural network function; I is the input image; o is the detected medical object; v is the corresponding feature vector. Seven cellular objects are detected in image Q, each represented by a 128-dimensional feature vector.

[0165] 1.4.2. Attribute-level feature extraction. Attribute-level feature ALF = ExtractVisualAttr(OLF) = {(a1,v1,o1,s1),...}; where: a1 = "morphology", v1 = "circular", o1 = "nucleus", s1 = 0.93; a2 = "intensity", v2 = "highlight", o2 = "mitochondria", s2 = 0.87; a3 = "distribution", v3 = "peripheral aggregation", o3 = "mitochondria", s3 = 0.82; ExtractVisualAttr() is a visual attribute extraction function; OLF is an object-level feature; a is the attribute type; v is the attribute value; o is the associated object; s is the attribute confidence score, ranging from [0,1].

[0166] 1.4.3. Relationship-level feature extraction. Relationship-level feature RLF = ExtractVisualRel(OLF, ALF) = {(o1, r, o2, s), ...}; where: o1 = "mitochondria", r = "near", o2 = "nucleus", s = 0.89; o3 = "cytoplasm", r = "contains", o4 = "mitochondria", s = 0.94; ExtractVisualRel() is the visual relationship extraction function; OLF is the object-level feature; ALF is the attribute-level feature; o1 and o2 are the relationship objects; r is the relationship type; s is the relationship confidence score, ranging from [0, 1].

[0167] 1.5. Construct hierarchical visual representation. Integrate object-level features (OLF), attribute-level features (ALF), and relation-level features (RLF) into a hierarchical visual representation (HVR).

[0168] Step 2: Implement a context-aware professional terminology understanding system.

[0169] 2.1. Identify professional terms. Extract professional terms from the hierarchical semantic representation (HSR): Professional term list (PTL) = ExtractTerms(HSR) = {(t1, d1, c1), ...}; where t1 = "apoptosis", d1 = 0.12, c1 = 0.95; t2 = "fluorescence staining", d2 = 0.08, c2 = 0.93; t3 = "mitochondrial membrane potential", d3 = 0.19, c3 = 0.96; ExtractTerms() is the professional term extraction function; HSR is the hierarchical semantic representation; t is the extracted professional term; d is the term ambiguity; a larger value indicates higher term ambiguity; c is the term domain relevance; a higher value indicates higher relevance to the current domain.

[0170] 2.2. Analyze the professional background and purpose of the query.

[0171] 2.2.1. Domain feature identification. Domain feature vector DF = DomainAnalysis(HSR, PTL) = [0.94, 0.87, 0.23, 0.12, 0.08]; where DomainAnalysis() is the domain analysis function; HSR is the hierarchical semantic representation; PTL is the professional term list; DF[0] = 0.94 indicates the probability that the query belongs to the "cell biology" domain; DF[1] = 0.87 indicates the probability that the query belongs to the "apoptosis research" sub-domain; DF[2] = 0.23 indicates the probability that the query belongs to the "drug research" domain; DF[3] = 0.12 indicates the probability that the query belongs to the "pathology" domain; DF[4] = 0.08 indicates the probability that the query belongs to the "molecular biology" domain.

[0172] 2.2.2. Query Purpose Analysis. Purpose feature vector PF = PurposeAnalysis(HSR, PTL) = [0.85, 0.62, 0.12, 0.09, 0.05]; where PurposeAnalysis() is the purpose analysis function; HSR is the hierarchical semantic representation; PTL is the professional term list; PF[0] = 0.85 indicates the probability that the query purpose is "research and analysis"; PF[1] = 0.62 indicates the probability that the query purpose is "comparison and verification"; PF[2] = 0.12 indicates the probability that the query purpose is "teaching demonstration"; PF[3] = 0.09 indicates the probability that the query purpose is "repeated detection"; PF[4] = 0.05 indicates the probability that the query purpose is "literature research".

[0173] 2.2.3. Knowledge background extraction. Knowledge background vector KF = KnowledgeExtract(HSR, PTL) = [0.92, 0.88, 0.75, 0.67, 0.43]; where: KnowledgeExtract() is the knowledge background extraction function; HSR is the hierarchical semantic representation; PTL is the professional term list; KF[0]=0.92 represents the knowledge confidence of "mitochondria play a key role in cell apoptosis"; KF[1]=0.88 represents the knowledge confidence of "membrane potential changes are an early sign of cell apoptosis"; KF[2]=0.75 represents the knowledge confidence of "fluorescent staining is used to observe the state of living cells"; KF[3]=0.67 represents the knowledge confidence of "JC-1 dye is often used to detect mitochondrial membrane potential"; KF[4]=0.43 represents the knowledge confidence of "cells undergo morphological changes during apoptosis".

[0174] 2.3. Fusion context feature vector. Context feature vector CF = ContextFusion(DF, PF, KF) = [0.94, 0.85, 0.92, 0.88, 0.75, 0.62, 0.67]; where: ContextFusion() is the context feature fusion function; DF is the domain feature vector; PF is the purpose feature vector; KF is the knowledge background vector; CF is the fused context feature vector, which retains the features with higher weights in each vector.

[0175] 2.4. Term Disambiguation. "Mitochondrial membrane potential" may refer to a membrane potential measurement, a membrane potential change process, or a membrane potential detection method in different contexts. The disambiguation function DisambiguateTerms(PTL, CF) = {(t1, m1),...}; where: t1 = "mitochondrial membrane potential", m1 = "change process"; t2 = "apoptosis", m2 = "programmed cell death process"; t3 = "fluorescence staining", m3 = "live cell detection method"; DisambiguateTerms() is the term disambiguation function; PTL is a list of specialized terms; CF is a context feature vector; t is a specialized term; and m is the precise meaning after disambiguation.

[0176] 2.5. Generate a context-enhanced query representation. Context-enhanced query representation CEQ = EnhanceQuery(HSR, CF) = [0.87, 0.92, 0.78, 0.93, 0.68, 0.82, 0.91, 0.74, 0.89, 0.85], where EnhanceQuery() is the query enhancement function; HSR is the hierarchical semantic representation; CF is the context feature vector; and CEQ is the ten-dimensional context-enhanced query representation.

[0177] Step 3: Establish a semantic-visual bidirectional mapping and optimization framework.

[0178] 3.1. Initialize the mapping matrix. The initial mapping matrix IMM = InitializeMapping(CEQ, HVR) = [M1, M2, M3]; where InitializeMapping() is the mapping initialization function; CEQ is the context-enhanced query representation; HVR is the hierarchical visual representation; M1 is the 128×128 core concept mapping matrix; M2 is the 64×64 attribute mapping matrix; and M3 is the 32×32 relationship mapping matrix.

[0179] 3.2. Semantic to visual forward mapping. Visual attention map VAM = ForwardMapping(CEQ, IMM, HVR) = [A1, A2, A3]; where: ForwardMapping() is the forward mapping function; CEQ is the context-enhanced query representation; IMM is the initial mapping matrix; HVR is the hierarchical visual representation; A1 is the object attention map, with a size of 7×7, mainly highlighting the mitochondrial region (value 0.92) and the cell nucleus region (value 0.85); A2 is the attribute attention map, with a size of 7×7, mainly highlighting the fluorescence intensity change region (value 0.88); A3 is the relationship attention map, with a size of 7×7, mainly highlighting the interaction region between mitochondria and cell nucleus (value 0.79). Use the visual attention map to perform a preliminary search on the image library and obtain the candidate image set CS = {I1, I2, ..., I 50}, containing the 50 most similar images.

[0180] 3.3, Visual to semantic reverse mapping. Backward mapping similarity RMS = BackwardMapping(CS, CEQ,IMM) = [S1, S2, ..., S 50 ]; where: BackwardMapping() is the reverse mapping function; CS is the candidate image set; CEQ is the context-enhanced query representation; IMM is the initial mapping matrix; S1=0.89 represents the reverse mapping similarity of the first candidate image; S2=0.87 represents the reverse mapping similarity of the second candidate image; ...; S 50 =0.61 represents the reverse mapping similarity of the fiftieth candidate image.

[0181] 3.4. Calculate the three-level similarity and optimize the mapping matrix. Optimize mapping matrix OMM = OptimizeMapping(CEQ, CS, IMM, RMS) = [M'1, M'2, M'3]; where: OptimizeMapping() is the mapping optimization function; CEQ is the context-enhanced query representation; CS is the candidate image set; IMM is the initial mapping matrix; RMS is the reverse mapping similarity; M'1 is the optimized core concept mapping matrix; M'2 is the optimized attribute mapping matrix; M'3 is the optimized relationship mapping matrix. Taking mitochondrial object recognition as an example, the corresponding value in the initial weight IMM is 0.72, and the corresponding value in the optimized OMM is adjusted to 0.86, which enhances the sensitivity to mitochondrial features. Through optimization, the optimized candidate set OCS = {I1, I2,..., I 30}, containing the 30 most similar images after optimization.

[0182] Step 4: Perform multimodal similarity calculation and result sorting.

[0183] 4.1. Calculate multi-level similarity. Multi-level similarity MLS = MultiLevelSimilarity(OCS, CEQ, OMM) = {(I1, OLS1, ALS1, RLS1), ...}; where: MultiLevelSimilarity() is the multi-level similarity calculation function; OCS is the optimized candidate set; CEQ is the context-enhanced query representation; OMM is the optimized mapping matrix; I1 is the first candidate image; OLS1 = 0.91 is I1's object-level similarity; ALS1 = 0.87 is I1's attribute-level similarity; RLS1 = 0.83 is I1's relation-level similarity.

[0184] 4.2. Generate a hierarchical weight vector based on the context feature vector. The hierarchical weight vector LWV = GenerateWeights(CF) = [w1, w2, w3]; where GenerateWeights() is the weight generation function; CF is the context feature vector; w1 = 0.45 is the object-level weight; w2 = 0.35 is the attribute-level weight; and w3 = 0.20 is the relationship-level weight. Since this query focuses on changes in mitochondrial membrane potential, analysis shows that object identification (mitochondria) and attribute features (membrane potential changes) are more important, so the weights are increased accordingly.

[0185] 4.3. Fusion generates a comprehensive similarity score. Comprehensive similarity score CSS = FuseSimilarities(MLS,LWV) = [s1, s2, ..., s 30 ]; where: FuseSimilarities() is the similarity fusion function; MLS is the multi-level similarity; LWV is the level weight vector; s1=0.45×0.91+0.35×0.87+0.20×0.83=0.88 is the comprehensive similarity score of the first candidate image; s2=0.45×0.88+0.35×0.85+0.20×0.79=0.86 is the comprehensive similarity score of the second candidate image; ...; s 30 =0.45×0.62+0.35×0.58+0.20×0.55=0.59 is the comprehensive similarity score of the 30th candidate image.

[0186] 4.4、Rank and output the search results. Final search result FRS = RankResults(OCS, CSS) = [(I1,s1), (I2, s2), ..., (I 20 , s 20)]; RankResults() is the result ranking function; OCS is the optimized candidate set; CSS is the comprehensive similarity score; sort in descending order according to the comprehensive similarity score, and select the first 20 images as the final retrieval results.

[0187] This embodiment was experimented on an academic paper image duplication check dataset (containing 10,000 biomedical images) and compared with the most advanced existing methods: the traditional CBIR method (based only on visual features) had a precision (P@10) = 0.68 and a recall (R@20) = 0.72; the visual-semantic embedding method (CLIP, etc.) had a precision of 0.76 and a recall of 0.79; the precision of this embodiment was 0.92 and the recall was 0.89. In terms of professional terminology ambiguity resolution, the accuracy of the traditional method was 67%, while this embodiment reached 93%, improving the ability to understand professional terminology. In actual applications, this embodiment successfully detected the reuse of images in multiple papers, including some slightly modified images that were difficult to identify directly through visual features, demonstrating its significant advantages in solving the problem of similar image retrieval in professional fields. This embodiment generates a contextual feature vector by combining domain features, purpose features and knowledge background to resolve the problem of professional terminology ambiguity; realizes forward mapping from semantics to vision and reverse mapping from vision to semantics, and generates an optimized mapping matrix through iterative optimization, so that the system can optimize query understanding according to the actual image data distribution; dynamically generates a hierarchical weight vector based on the contextual feature vector to realize the adaptive fusion of object-level, attribute-level and relationship-level similarities, thereby improving retrieval accuracy.

[0188] Case 2: Using medical image retrieval as the background, users can retrieve similar medical images that match the query semantics through natural language queries containing professional terms. A context-aware semantic-visual bidirectional mapping framework achieves accurate retrieval through hierarchical representation, professional terminology understanding, bidirectional mapping optimization, and multimodal similarity calculation. Specifically:

[0189] Step 1: Build a hierarchical semantic-visual representation framework. Input the user query text "2 cm ground-glass nodule in the right upper lobe of the lung, with blurred edges and surrounding traction signs" and a medical image library (containing 10,000 chest CT images).

[0190] 1.1. Receive user query text and use Transformer encoder to preliminarily encode the query. Use pre-trained medical BERT encoder with word embedding dimension d=768; input word embedding sequence: E = [e1, e2, ..., e_n], where n=15 (the number of words in the query); process through multi-head self-attention layer: Att(Q,K,V) = softmax(QK T / sqrt(d))·V; where Q is the query matrix, K is the key matrix, and i finally obtains the basic query vector representation Q_base ∈ R 768 , the values ​​are [0.21, 0.35, -0.14, ..., 0.52].

[0191] 1.2. Apply a three-level structured decomposition algorithm to the basic query vector representation:

[0192] 1.2.1. Core concept extraction. Use the medical entity recognition model to identify the subject and key entities: "lung", "right upper lobe", and "ground-glass nodule shadow". Verify the entities using a medical ontology knowledge base (such as UMLS) to obtain confidence scores: C("lung") = 0.98; C("right upper lobe") = 0.95; C("ground-glass nodule shadow") = 0.96. Apply the semantic importance scoring algorithm to calculate weights: W("lung") = 0.75; W("right upper lobe") = 0.85; W("ground-glass nodule shadow") = 0.95. Output the core concept representation matrix C_core∈R 3×768 , each row corresponds to the embedding vector of a core concept.

[0193] 1.2.2. Attribute modification extraction. Identify modifiers: "2cm" and "edge blur". Construct attribute triples: A1 = ("size", "2cm", "ground glass nodule shadow"), with a confidence level of 0.92; A2 = ("edge characteristics", "blur", "ground glass nodule shadow"), with a confidence level of 0.88; output the attribute modification representation matrix A_att ∈ R 2×(2×768+64) , including attribute type, attribute value vector, index of associated core concepts and association strength, and 64 dimensions are used to represent association information.

[0194] 1.2.3. Relationship description extraction. Identify the relationship "surrounded by traction signs"; construct the relationship quadruple: R1 = ("ground glass nodule shadow", "surrounded by", "traction signs", "current"), with a confidence level of 0.86; output the relationship description representation matrix R_rel∈R 1×(3×768+64) , including subject concept, relationship type, object concept, and time series information vector, with 64 dimensions representing time series information.

[0195] 1.3. Integrate the core concept representation, attribute modification representation, and relationship description representation into a hierarchical semantic representation: Construct a hierarchical association graph G_sem = (V, E), where V is the node set (core concept) and E is the edge set (attributes and relationships); use the graph attention network (GAT) to process the association graph and calculate the node feature update: h′v = σ(∑{j∈N(v)} α_{vj}·W·h_j); attention coefficient α_{vj} = softmax(LeakyReLU(a T [W·h_v ∥ W·h_j])); finally obtain the hierarchical semantic representation H_sem ∈ R (3+2+1)×1024 , which integrates three levels of semantic representation. Where σ is the activation function, N(v) is the set of neighbor nodes, W is the weight matrix, h_j is the feature vector of the neighbor node, a is the attention mechanism parameter, and h_v is the feature vector of the current node.

[0196] 1.4. For each medical image in the image library, apply a hierarchical visual feature extraction network:

[0197] 1.4.1. Object-level feature extraction. Use the Faster R-CNN model to detect objects in medical images. Taking the first image as an example, three objects are detected: O1 ("lung"), O2 ("right upper lobe"), and O3 ("nodule"). For each object, a feature vector (2048 dimensions) is extracted and its position information (4 dimensions: x, y, w, h) is recorded. The object-level visual feature matrix O_vis ∈ R is output. 3 ×2052 .

[0198] 1.4.2. Attribute-level feature extraction. For each detected medical object region, an attribute classification network is used to identify medical attributes. Taking O3 ("nodule shadow") as an example, the following attributes are identified: size: 1.8 cm (confidence 0.91); shape: irregular (confidence 0.87); edge: fuzzy (confidence 0.85); density: ground glass (confidence 0.93). The output attribute-level visual feature matrix A_vis ∈ R 8×(512+4) , where 512 are attribute feature dimensions and 4 are associated information.

[0199] 1.4.3. Relationship-level feature extraction. Identify the spatial and functional relationships between objects; construct the relationship graph between objects: R_vis1 = ("nodule shadow", "located", "right upper lobe"), confidence level 0.94; R_vis2 = ("nodule shadow", "surrounded by", "traction change"), confidence level 0.82; output the relationship-level visual feature matrix R_vis ∈ R 2×1024 .

[0200] 1.5. Organize the object-level, attribute-level, and relation-level visual features of each image into a hierarchical visual representation: Construct an image hierarchy to form a complete hierarchical visual representation H_vis ∈ R (3+8+2)×1024 ; Establish a visual index library based on locality sensitive hashing (LSH) to accelerate the retrieval process.

[0201] Step 2: Implement a context-aware professional terminology understanding system.

[0202] 2.1. Identify the list of professional terms.

[0203] 2.1.1. Term Detection. Vocabulary was extracted from the hierarchical semantic representation and matched against the medical terminology database. The following terminology was identified: "lung," "right upper lobe," "ground-glass opacity," "fuzzy margins," and "traction sign." Term specificity scores were calculated: S("lung") = 0.70 (general term); S("right upper lobe") = 0.85 (region-specific term); S("ground-glass opacity") = 0.95 (highly specific term); S("fuzzy margins") = 0.80 (descriptive term); and S("traction sign") = 0.90 (sign-specific term).

[0204] 2.1.2. Marking Polysemous Terms. Query the professional term ambiguity database and mark polysemous terms: "ground-glass nodule": possible meaning 1: "interstitial lung disease" (in the context of pneumonia), possible meaning 2: "sign of early lung cancer" (in the context of cancer screening); "traction sign": possible meaning 1: "fibrosis contraction" (in the context of interstitial lung disease), possible meaning 2: "tumor invasion" (in the context of malignant tumors). Calculate the degree of ambiguity: Amb("ground-glass nodule") = 0.65; Amb("traction sign") = 0.72. Output the ambiguous term list T_list.

[0205] 2.2. Analyze the professional background and purpose of the query.

[0206] 2.2.1. Identification of professional fields. Calculate the probability distribution D_field of the query belonging to different medical professional fields using a multi-label classification model: P("Chest Radiology") = 0.92; P("Pulmonology") = 0.85; P("Oncology") = 0.73; P("General Internal Medicine") = 0.45. Construct the domain feature vector F_domain ∈ R 256 , weighted fusion is adopted: F_domain = ∑_i P(i)·V_i, where V_i is the standardized feature vector of each domain.

[0207] 2.2.2. Query Purpose Analysis. Identify query intent and calculate the probability distribution of different purposes: P("diagnosis") = 0.87; P("screening") = 0.65; P("follow-up") = 0.32; P("teaching") = 0.15. Determine the primary query purpose: "diagnosis." Construct the purpose feature vector F_purpose ∈ R 128 , encoding query intent information.

[0208] 2.2.3. Knowledge background extraction. Retrieve relevant professional background knowledge from the medical knowledge graph: "Ground-glass nodules correspond to stage IA in the TNM staging of lung cancer"; "Nodules with blurred edges suggest possible invasive growth"; "Stretching signs are common in tumors or fibrosis." Extract key concepts and relationships to generate a knowledge background vector F_knowledge ∈ R 384 .

[0209] 2.3. Fusion of domain, purpose, and knowledge background features to generate a context feature vector. Apply feature normalization to unify dimensions and scope; use the attention-weighted fusion mechanism to calculate the fusion weights: w_domain = 0.45; w_purpose = 0.35; w_knowledge = 0.20; generate the context feature vector F_context = w_domain·F_domain + w_purpose·F_purpose + w_knowledge·F_knowledge; finally, F_context ∈ R 512 , which encodes the complete query context information. Among them, w_domain is the domain feature weight, w_purpose is the query purpose weight, and w_knowledge is the knowledge background weight.

[0210] 2.4. Based on the professional term list and context feature vector, term ambiguity resolution is performed.

[0211] 2.4.1. Construct a conditional probability model for terms. Calculate the conditional probability P(meaning|context) for each polysemous term: for "ground-glass nodules": P("pulmonary interstitial lesions"|F_context) = 0.25; P("early lung cancer signs"|F_context) = 0.75; for "traction signs": P("fibrosis contraction"|F_context) = 0.30; P("tumor invasion"|F_context) = 0.70. Calculation formula: P(meaning|context) = softmax(f_θ(meaning, F_context)), where f_θ is a trained deep neural network.

[0212] 2.4.2. Contextual Feature Matching. Calculate the matching degree M_score between each term's meaning and the query context: M_score("ground-glass nodules", "early signs of lung cancer") = cos(V("early signs of lung cancer"), F_context) = 0.82; M_score("traction signs", "tumor invasion") = cos(V("tumor invasion"), F_context) = 0.78; where V() is the vector representation of the term's meaning, and cos() is the cosine similarity function.

[0213] 2.4.3. Collaborative Disambiguation of Term Networks. A relationship network was constructed between terms to represent their semantic dependencies. A graph reasoning algorithm was applied to update the meaning probabilities: P'("ground-glass nodule" = "early lung cancer sign" | F_context, Network) = 0.85; P'("traction sign" = "tumor invasion" | F_context, Network) = 0.82. The update formula was: P'(m_i|F_context, Network) = P(m_i|F_context) + α·∑_j w_ij·P(m_j|F_context), where α = 0.3 is the balancing factor and w_ij is the association weight between terms.

[0214] 2.5. Fusion of the resolved semantic representation and contextual feature vector. Update the term vectors in the hierarchical semantic representation and replace them with the professional representation of the selected meaning; apply contextual conditional feature transformation to enhance the core concepts, attribute modifications, and relationship descriptions; obtain the context-enhanced query representation Q_enhanced ∈ R 1536 .

[0215] Step 3: Establish a semantic-visual bidirectional mapping and optimization framework.

[0216] 3.1. Initialize the bidirectional mapping matrix: Use the pre-trained medical domain semantic-visual mapping weights as the initial value; construct the initial mapping matrix for the core concept, attribute modification and relationship description respectively: W_concept ∈ R 1024×1024 ; W_attribute ∈ R 1024×1024 ;W_relation ∈ R 1024×1024 ; Establish the inter-level correlation matrix W_inter ∈ R 3×3 , with initial values ​​of [[0.5, 0.3, 0.2], [0.3, 0.5, 0.2], [0.2, 0.3, 0.5]].

[0217] 3.2. Realize the forward mapping from semantics to vision.

[0218] 3.2.1. Generate multi-level visual attention map. Decompose the context-enhanced query representation Q_enhanced into query components corresponding to the hierarchical visual representation: Q_concept ∈ R 3×1024 ; Q_attribute ∈ R 2×1024 ; Q_relation∈R 1×1024 Generate a three-layer attention map: A_concept = softmax(Q_concept·W_concept·O_vis T )∈R 3×3 ;A_attribute = softmax(Q_attribute·W_attribute·A_vis T ) ∈ R 2×8 ;A_relation = softmax(Q_relation·W_relation·R_vis T ) ∈ R 1×2 .

[0219] 3.2.2. Adjust attention weights based on context. Use the context feature vector F_context to adjust attention weights: A_concept' = A_concept * sigmoid(F_context_c·O_vis); A_attribute' = A_attribute * sigmoid(F_context_a·A_vis); A_relation' = A_relation * sigmoid(F_context_r·R_vis); where F_context_c, F_context_a, and F_context_r are projections of F_context at different levels.

[0220] 3.2.3. Perform visual feature activation and retrieval. Calculate the similarity matrix of the forward mapping: S_concept = A_concept'·O_vis·Q_concept T ∈R 3×3 ;S_attribute = A_attribute'·A_vis·Q_attribute T ∈R 2×2 ;S_relation = A_relation'·R_vis·Q_relation T ∈ R 1×1Calculate the initial forward mapping similarity: S_forward = (w_c tr(S_concept) + w_a tr(S_attribute) + w_r tr(S_relation)) / (w_c + w_a + w_r), where w_c = 0.4, w_a = 0.3, and w_r = 0.3 are the initial layer weights. Based on the forward mapping similarity S_forward, select the top 50 similar images as the candidate set C_initial.

[0221] 3.3. Realize the reverse mapping from vision to semantics.

[0222] 3.3.1 Visual Feature Distribution Analysis. Perform cluster analysis on the image features in the candidate set C_initial to identify visual patterns. Use the k-means algorithm (k=3) to cluster object-level features, obtaining the following centroids: Centroid1 = [0.72, 0.31, ..., 0.55]; Centroid2 = [0.45, 0.62, ..., 0.28]; Centroid3 = [0.81, 0.27, ..., 0.63]; and identify the dominant visual pattern: "right upper lobe nodule," with a frequency of 80%.

[0223] 3.3.2 Feature Space Mapping Analysis. Evaluate the quality of semantic-to-visual mapping: The semantic concept "ground-glass nodules" is mapped to multiple visual representations (over-mapping), quantified as 0.75; the semantic concept "edge blurring" is insufficiently mapped (under-mapping), quantified as 0.65. Generate a mapping quality assessment report to identify mapping areas that require adjustment.

[0224] 3.3.3 Dynamic Adjustment of Semantic Understanding. Based on the distribution of visual features, adjust the weights of semantic concepts: W'("ground glass nodule") = W("ground glass nodule") * 0.85 = 0.95 * 0.85 = 0.81 (reducing the weight of over-mapped concepts); W'("edge blur") = W("edge blur") * 1.25 = 0.80 * 1.25 = 1.00 (increasing the weight of under-mapped concepts). Introduce a new semantic concept: "ground glass density" (discovered from visual features but not explicitly stated in the original query). Generate an adaptive semantic representation, S_adaptive, to better match the actual visual feature distribution.

[0225] 3.3.4. Update the bidirectional mapping relationship. Apply the gradient descent algorithm to optimize the mapping matrix parameters based on the mapping error: W_concept' = W_concept - η·▽L(W_concept), where η = 0.01 is the learning rate; W_attribute' = W_attribute - η·▽L(W_attribute); W_relation' = W_relation - η·▽L(W_relation); update the inter-hierarchical association structure W_inter'; and generate the reverse mapping similarity matrix S_backward. Where W_concept' is the updated concept mapping matrix, ▽L is the gradient of the loss function, W_relation' is the updated relationship mapping matrix, and W_attribute' is the updated attribute mapping matrix.

[0226] 3.4. Perform iterative optimization loop.

[0227] 3.4.1. Recalculate semantic-visual similarity. Recalculate similarity using the updated mapping matrix and the adapted semantic representation; obtain the optimized similarity S_optimized = 0.6·S_forward + 0.4·S_backward; and update the candidate set C_optimized (the top 30 most similar images).

[0228] 3.4.2. Evaluate the optimization results. Calculate the change in similarity distribution: the average similarity increased from 0.73 to 0.81; analyze the ranking change: 75% of images improved in ranking, with an average increase of 3.2 places; evaluate the parameter change magnitude: the average parameter change was 2.8%, which is below the threshold of 5%.

[0229] 3.4.3. Determine the termination conditions for iteration. Similarity improvement < 0.01 (below the threshold of 0.02); number of iterations = 3 (below the upper limit of 5); if the termination conditions are met, output the final optimized mapping matrix and the optimized candidate set.

[0230] Step 4: Perform multimodal similarity calculation and result sorting.

[0231] 4.1. Prepare the final similarity calculation resources. Verify that all necessary information is complete; optimize the allocation of computing resources; and set normalization parameters to ensure comparability of scores at different levels.

[0232] 4.2. Calculate the feature similarity of the three levels.

[0233] 4.2.1. Calculation of core concept similarity. For each image i, calculate the core concept similarity: S_concept(i) =∑_j∑_k A_concept'(j,k)·cos(Q_concept(j),O_vis(i,k)); for example, for the first image: S_concept(1) = 0.92*cos([0.42,0.68,...,0.55],[0.45,0.66,...,0.58])+... = 0.87.

[0234] 4.2.2 Attribute modification similarity calculation. Calculate attribute-level similarity: S_attribute(i)=∑_j∑_k A_attribute'(j,k)·cos(Q_attribute(j),A_vis(i,k)); for example, for the first image: S_attribute(1)=0.85*cos([0.32,0.71,...,0.48],[0.35,0.68,...,0.51])+...=0.79.

[0235] 4.2.3. Calculation of relational similarity. Calculate relation-level similarity: S_relation(i)=∑_j∑_k A_relation'(j,k)·cos(Q_relation(j),R_vis(i,k)); for example, for the first image: S_relation(1)=0.80 * cos([0.62,0.41,...,0.38],[0.58,0.44,...,0.35]) = 0.76.

[0236] 4.3. Dynamically generate hierarchical weight vectors based on context feature vectors.

[0237] 4.3.1. Contextual Condition Weight Generation. Analyze the contextual feature vector F_context to determine the base weights: Based on the characteristics of the "chest radiology" domain, the initial weights are: w_c = 0.30, w_a = 0.45, w_r = 0.25. Based on the "diagnosis" purpose, the weights are adjusted to: w_c' = 0.25, w_a' = 0.50, w_r' = 0.25 (increasing the attribute layer weights). The calculation formula is: w_x' = w_x + γ·f_purpose(x), where γ = 0.15 is the adjustment factor and f_purpose is the purpose adjustment function.

[0238] 4.3.2. Weight Adaptive Optimization. Analyzing the discriminative power of the three-level similarity: Var(S_concept) = 0.034 (low discriminative power); Var(S_attribute) = 0.087 (high discriminative power); Var(S_relation) = 0.052 (medium discriminative power). We then applied discriminative power enhancement: w_x'' = w_x' * (1 + β·Var(S_x)), where β = 0.5 is the amplification factor. The final layer weight vector W_level = [0.22, 0.56, 0.22].

[0239] 4.4. Integrate multi-level similarities and calculate the comprehensive similarity score.

[0240] 4.4.1. Weighted fusion calculation. Calculate the linear weighted score: S_linear(i) = W_level(1)·S_concept(i)+W_level(2)·S_attribute(i)+W_level(3)·S_relation(i); for example, for the first image: S_linear(1) = 0.22·0.87 + 0.56·0.79 + 0.22·0.76 = 0.80. Apply a nonlinear adjustment function to process the complementary relationship between levels: S_nonlinear(i) = S_linear(i) * (1 + Δ·min(S_concept(i), S_attribute(i), S_relation(i))); where Δ=0.1 is the nonlinear factor. For example, for the first image: S_nonlinear(1) = 0.80 * (1 + 0.1·0.76) = 0.86.

[0241] 4.4.2 Confidence Assessment. Calculate the consistency of the three-layer similarity: C(i) = 1 - σ(S_concept(i), S_attribute(i), S_relation(i)); where σ is the standard deviation function. For example, for the first image: C(1) = 1 - σ(0.87, 0.79, 0.76) = 1 - 0.057 = 0.943. The final similarity score: S_final(i) = S_nonlinear(i) * C(i). For example, for the first image: S_final(1) = 0.86 * 0.943 = 0.81.

[0242] 4.5. Sort and output the final search result set.

[0243] 4.5.1. Similarity Sorting and Filtering: Sort all candidate images in descending order by their final similarity score S_final. Apply a similarity threshold filter and remove results with S_final < 0.65. Apply diversity optimization to ensure that the result set contains relevant images with different visual representations.

[0244] 4.5.2. Result Set Enhancement. Cluster the results into five groups based on visual feature similarity. Select the most representative image for each group as the primary result. Add result metadata, including key concepts of the match, prominent visual features, and confidence information. Generate result explanation data, such as "Reason for match: Ground-glass nodule in the right upper lobe, with blurred margins and traction signs."

[0245] 4.6. Return the final search result set to the user and save the search session data for continuous learning of the system.

[0246] This embodiment provides comprehensive context for term disambiguation through a three-dimensional vector representation of domain context, query purpose, and knowledge background, enabling the system to correctly understand professional terminology based on the professional context. This not only maps semantics to visuals but also adjusts semantic understanding based on visual feedback, resolving the issue of traditional one-way mapping, where semantic understanding cannot be adjusted based on actual visual data. Queries and images are decomposed into three levels: core concepts, attribute modifications, and relationship descriptions, achieving refined semantic matching while dynamically adjusting the weights of these levels based on context.

[0247] In another embodiment of the present application, generating a list of professional terms is specifically as follows: matching the vocabulary in the hierarchical semantic representation with the pre-trained domain professional terminology vocabulary. Input the hierarchical semantic representation HSR = [0.87, 0.92, 0.78, 0.93, 0.68, 0.82, 0.91, 0.74, 0.89, 0.85], and extract the vocabulary sequence WS = ["apoptosis", "fluorescence staining", "mitochondria", "membrane potential", "change", "image", "similar", "detection"]. Vocabulary matching score MS = VocabMatch(WS, PDV) = {(w1, m1, p1), (w2, m2, p2), ...}; where: w1="apoptosis", m1=0.95, p1=0.92; w2="fluorescence staining", m2=0.89, p2=0.87; w3="mitochondria", m3=0.97, p3=0.94; w4="membrane potential", m4=0.91, p4=0.88; w5="change", m5=0.43, p5=0.25; w6="image", m6=0.52, p6= 0.31; w7 = "similarity", m7 = 0.38, p7 = 0.22; w8 = "detection", m8 = 0.67, p8 = 0.45; VocabMatch() is the vocabulary matching function; WS is the vocabulary sequence; PDV is the pre-trained vocabulary of cell biology terminology; w is a word; m is the match confidence, indicating the degree of match between the vocabulary and the term; p is the professional score, indicating the strength of the vocabulary's terminology attribute. Setting the matching threshold to 0.8, the preliminary candidate term set IPTC = {("apoptosis", 0.95), ("fluorescence staining", 0.89), ("mitochondria", 0.97), ("membrane potential", 0.91)} was selected.

[0248] Based on the preliminary candidate set of professional terms, the probability of each term belonging to the current query domain is calculated. Domain attribution probability DAP = DomainVerify(IPTC, DF)={(t1,dp1),(t2, dp2),...}; where: t1="apoptosis", dp1=0.97; t2="fluorescence staining", dp2=0.93; t3="mitochondria", dp3=0.96; t4="membrane potential", dp4=0.89; DomainVerify() is the domain verification function; IPTC is the preliminary candidate set of professional terms; DF is the domain feature vector [0.94, 0.87, 0.23, 0.12, 0.08]; t is the professional term; dp is the domain attribution probability, calculated as dp = Σ i (DF i×term_domain_weight i ), where term_domain_weight i is the weight distribution of terms in the i-th field. The verified professional term candidate set VPTC = {("apoptosis", 0.97), ("fluorescence staining", 0.93), ("mitochondria", 0.96), ("membrane potential", 0.89)}, and the field belonging probability of all terms exceeds the 0.85 threshold.

[0249] Based on semantic link relationships and sentence structure, the importance of terms in the current query context is evaluated. Contextual importance CIS = ContextRelevance(VPTC, HSR, SG) = {(t1, cis1), (t2, cis2), ...}; where: t1 = "cell apoptosis", cis1 = 0.92; t2 = "fluorescence staining", cis2 = 0.87; t3 = "mitochondria", cis3 = 0.94; t4 = "membrane potential", cis4 = 0.91; ContextRelevance() is the contextual relevance evaluation function; VPTC is the verified candidate set of professional terms; HSR is the hierarchical semantic representation; SG is the sentence structure graph; t is the professional term; and cis is the contextual importance score. The calculation formula is cis = α × semantic_link_score + β × structure_position_score + γ × co_occurrence_score, where α = 0.4, β = 0.3, and γ = 0.3 are the weight coefficients for semantic link, structural position, and co-occurrence, respectively. The final professional term list PTL = {("mitochondria", 0.94), ("apoptosis", 0.92), ("membrane potential", 0.91), ("fluorescence staining", 0.87)}, which is sorted in descending order of contextual importance score.

[0250] In another embodiment of the present application, the reverse mapping similarity matrix is ​​generated as follows: based on the candidate image set CS={I1,I2,...,I 50}, perform reverse mapping analysis to generate a reverse mapping similarity matrix. Perform visual feature distribution analysis on each image in the candidate image set to identify the dominant visual pattern. Visual feature distribution characteristic VFD = AnalyzeVisualDist(CS = {(I i ,mv i ,kv i),...}; where: mv1 of I1 = ["mitochondrial aggregation", "high fluorescence", "clear cell outline"], kv1 = [0.89, 0.84, 0.76]; mv2 of I2 = ["mitochondrial dispersion", "medium fluorescence", "fuzzy cell outline"], kv2 = [0.72, 0.68, 0.54]; mv3 of I3 = ["mitochondrial swelling", "weak fluorescence", "abnormal cell morphology"], kv3 = [0.81, 0.45, 0.69]; AnalyzeVisualDist() is the visual distribution analysis function; CS is the candidate image set; I i is the i-th candidate image; mv i is a list of dominant visual modes; kv i is the key visual feature intensity vector, with a value range of [0,1] indicating the significance of the feature. Through cluster analysis, three main visual patterns were identified: normal state pattern (accounting for 45%), early apoptosis pattern (accounting for 35%), and late apoptosis pattern (accounting for 20%).

[0251] Based on the identified dominant visual pattern, a visual-semantic bridge network is constructed to project the visual features into the semantic space. The bridge network output VSB = BridgeNetwork(VFD, SEM) = {(I i , rs i ), ...}; where: rs1 of I1 = ["cells are in normal state", "mitochondrial function is active", "membrane potential is stable"]; rs2 of I2 = ["cells begin to apoptosis", "mitochondrial function is weakened", "membrane potential decreases slightly"]; rs3 of I3 = ["cells are in deep apoptosis", "mitochondria are severely damaged", "membrane potential decreases significantly"]; BridgeNetwork() is the visual-semantic bridge function; VFD is the visual feature distribution characteristic; SEM is the semantic embedding matrix; I i is the i-th candidate image; rs i Inverse semantic description is a semantic expression derived by reverse engineering visual features. The bridging network uses a three-layer transformation structure: visual feature layer → intermediate mapping layer → semantic representation layer. The dimensions of the transformation matrix in each layer are 128×64, 64×32, and 32×10, respectively.

[0252] Compare the reverse semantic description with the original query semantics to identify potential differences. Semantic difference matrix SDM = IdentifyDifference(VSB, CEQ) = {(I i , sd i), ...};where: sd1 of I1=[0.15, 0.08, 0.12, 0.06, 0.11, 0.09, 0.07, 0.04, 0.13, 0.10];sd2 of I2=[0.23, 0.19, 0.25, 0.17, 0.22, 0.18, 0.14, 0.21, 0.26, 0.20];sd3 of I3=[0.41, 0.38, 0.44, 0.35, 0.42, 0.37, 0.33, 0.39, 0.45, 0.40];IdentifyDifference() is the semantic difference recognition function;VSB is the output of the bridging network;CEQ is the context enhanced query representation[0.87, 0.92, 0.78, 0.93, 0.68, 0.82,0.91, 0.74, 0.89, 0.85]; I i is the i-th candidate image; sd i is the semantic difference vector, and the calculation formula is sd i = |CEQ - reverse_semantic_vector i |, represents the absolute difference of each dimension. i is the reverse semantic vector of the i-th candidate image. Calculate the average semantic difference: the average difference of I1 is 0.095, the average difference of I2 is 0.205, and the average difference of I3 is 0.394. The smaller the difference, the closer the image is to the query semantics.

[0253] According to the semantic differences identified, the semantic understanding model parameters are dynamically adjusted to generate the final reverse mapping similarity matrix. Reverse mapping similarity matrix RMS=DynamicAdjust(SDM,CF,MP)=[s1,s2,...,s 50 ]; where: s1=1-(0.4×avg_semantic_diff1+0.3×context_penalty1 + 0.3×visual_penalty1) = 1-(0.4×0.095+0.3×0.12+0.3×0.08)=1-0.098=0.902; s2=1-(0.4×0.205+0.3×0.18+0.3×0.15)=1-0.181=0.819; s3=1-(0.4×0.394+0.3×0.35+0.3×0.32)=1- 0.359= 0.641; DynamicAdjust() is the dynamic adjustment function; SDM is the semantic difference matrix; CF is the context feature vector; MP is the model parameter; avg_semantic_diffi is the average semantic difference of the i-th image; context_penalty i is the context consistency penalty; visual_penalty i is a visual feature consistency penalty. Through dynamic adjustment, the model parameters were adjusted from the initial state θ0 = [0.72, 0.68, 0.75, 0.71, 0.69] to the optimized state θ1 = [0.78, 0.74, 0.81, 0.77, 0.75], improving the accuracy of semantic understanding. The final reverse mapping similarity matrix RMS = [0.902, 0.819, 0.641, 0.785, 0.856, ..., 0.623], which contains the reverse mapping similarity scores of all 50 candidate images.

[0254] In another embodiment of the present application, the steps of bidirectional mapping optimization are specifically as follows: bidirectional mapping optimization is performed based on the forward mapping similarity matrix FMS=[0.89, 0.87, 0.61, 0.78, 0.85, ..., 0.62] and the reverse mapping similarity matrix RMS=[0.902, 0.819, 0.641, 0.785, 0.856, ..., 0.623]. Optimize mapping matrix OMM = OptimizeMapping (FMS, RMS, IMM, λ) = [M'1, M'2, M'3]; where M'1 is the optimized core concept mapping matrix, which is adjusted from the initial value IMM1 by Δ1 = λ × (FMS - RMS) = 0.3 × ([0.89, 0.87, 0.61, ...] - [0.902, 0.819, 0.641, ...]) = 0.3 × [-0.012, 0.051, -0.031, ...] = [-0.0036, 0.0153, -0.0093, ...]; M'2 is the optimized attribute mapping matrix, which is adjusted by Δ2 = 0.2 × (FMS - RMS); M'3 is the optimized relationship mapping matrix, which is adjusted by Δ3 = 0.1 × (FMS - RMS); OptimizeMapping() is the mapping optimization function; FMS is the forward mapping similarity matrix; RMS is the reverse mapping similarity matrix; IMM is the initial mapping matrix; λ is the learning rate constant; M' i is the optimized mapping matrix of the i-th layer. Through the bidirectional consistency constraint Loss=Σ i (α i |FMS i -RMS i | 2 + β i ||M' i -IMM i || 2), where α1=0.5, α2=0.3, α3=0.2 are the consistency weights of each layer, β1=0.1, β2=0.1, β3=0.1 are regularization weights, and the final loss function converges to 0.034. After bidirectional mapping optimization, the accuracy of the system in professional terminology understanding increased from the initial 67% to 93%, and the reverse mapping consistency index increased from 0.72 to 0.91. Compared with the traditional one-way mapping method, this embodiment improves the precision from 0.76 to 0.92 and the recall from 0.79 to 0.89 when processing professional field image retrieval tasks, verifying the effectiveness of the bidirectional mapping optimization mechanism.

[0255] This invention improves the ability to accurately understand professional terminology by combining domain features, purpose features, and knowledge background to construct contextual feature vectors and performing term disambiguation based on contextual features. In the context of academic paper image retrieval, the system can identify the specific meaning differences of "mitochondrial membrane potential" in different research contexts (such as measured value, change process, or detection method), improving the accuracy of term disambiguation from 67% with traditional methods to 93%. This enables the system to accurately understand researchers' professional query intent and correctly identify images that are visually similar but have different semantics or research purposes, effectively addressing the problem of misidentification caused by inaccurate understanding of professional terminology during image duplication checks in academic papers. When researchers query for cell morphology images under specific experimental conditions, the system can accurately distinguish visually similar images with different experimental purposes, reducing false positives in academic review. By establishing a bidirectional interactive mechanism, the system can dynamically adjust its semantic understanding to adapt to the actual visual feature distribution. In academic image retrieval, traditional one-way mapping methods are susceptible to visual noise and variation. However, this invention uses reverse mapping to obtain feedback from the visual feature distribution and optimize the mapping matrix parameters. For example, when detecting apoptosis images, the system increased the weight of mitochondrial object recognition from 0.72 to 0.86 through iterative optimization, enhancing sensitivity to key biomarkers. This enabled the system to capture core semantic features even when processing images that had been intentionally modified (such as adjusting brightness, contrast, cropping, or rotation) to circumvent duplication detection, improving retrieval accuracy by 21%. For fine-tuned academic images, traditional methods achieved a detection rate of only 58%, while the proposed method achieved an 87% detection rate, effectively preventing academic misconduct. By decomposing the query and image into three levels: core concepts, attribute modifications, and relationship descriptions, a strictly corresponding hierarchical representation framework was established, enabling more refined semantic and visual feature matching. In academic paper image retrieval applications, this hierarchical representation enables the system to distinguish images with similar objects but different relationships, such as distinguishing the subtle difference between "mitochondria close to the nucleus" and "mitochondria far from the nucleus," which is crucial for interpreting experimental results in cell biology research. Traditional holistic feature representation methods have difficulty capturing such fine-grained differences, while the present invention improves the discrimination accuracy of complex images from 72% to 89% by extracting and matching object-level, attribute-level, and relationship-level features respectively. In particular, when identifying academic charts that have been redrawn or reformatted, the consistency of core data relationships can be identified through relationship-level features, effectively preventing the behavior of evading duplicate checks by redrawing charts. By converting context feature vectors into hierarchical weight vectors and dynamically fusing them with multi-level similarities to generate a comprehensive similarity score, the retrieval strategy can be adaptively adjusted according to different professional fields and query purposes.In an academic image duplication check scenario, when researchers focus on "mitochondrial membrane potential changes," the system automatically increases the weights of the object level (weight 0.45) and attribute level (weight 0.35) while reducing the influence of the relationship level (weight 0.20), because membrane potential changes primarily manifest as changes in the fluorescence intensity (attribute) of the mitochondria (objects). The system achieves a precision of 0.92 in academic paper image retrieval, 24 percentage points higher than traditional CBIR methods and 16 percentage points higher than visual-semantic embedding methods. When faced with different types of academic images (such as microscopic images, diagrams, and flowcharts), the system automatically adjusts the weights of the corresponding levels and optimizes the retrieval strategy, maintaining high accuracy across a wide range of academic image retrieval tasks and increasing retrieval speed by 38% compared to traditional methods. A comprehensive contextual feature vector is constructed by matching professional terms with a pre-trained domain terminology vocabulary, verifying the domain adaptability of the terms, assessing contextual relevance, identifying the query's domain set, analyzing the query intent, and retrieving relevant professional background knowledge. In academic paper image retrieval applications, it can accurately distinguish between different interpretations of the same visual phenomenon across different research fields. For example, the same cell morphological changes in apoptosis research may represent programmed cell death, while in differentiation research, they represent the process of cell differentiation. By identifying the query as originating from the "cell biology" domain (with a probability of 0.94) and the "apoptosis research" subdomain (with a probability of 0.87), and determining the query intent as "research analysis" (with a probability of 0.85), the system correctly understands the specific context of the term. This improves the accuracy of term understanding from 71% with traditional methods to 95%, reducing semantic confusion in academic image retrieval, particularly in interdisciplinary research, and ensuring that retrieval results better meet researchers' professional expectations. By analyzing the focus and importance distribution of the query within a specific domain, the system applies a dynamic weight generator to generate hierarchical weight vectors tailored to different professional scenarios based on the focus, importance distribution, and query intent. In academic paper image duplication detection applications, the system accurately identifies the different focuses of image features across different research types. For example, for cell morphology research, the system prioritizes object structural features; for molecular localization research, it prioritizes spatial relationship features; and for kinetic research, it emphasizes temporal variation features. Through targeted weight adjustments, the system achieved an average accuracy of 92% when comprehensively evaluating a large academic image library (10,000 biomedical images), 15 percentage points higher than the fixed-weight method. Especially when processing interdisciplinary research images, the system is able to adapt to the evaluation standards of different disciplines, reducing misjudgments due to disciplinary differences and improving the applicability and fairness of the academic image duplication detection system in a multidisciplinary environment. By constructing a visual-semantic bridging network, visual features are projected into the semantic space, potential semantic differences are identified, and the semantic understanding model is dynamically adjusted based on the actual image data.In the academic paper image retrieval scenario, when the system finds that mitochondria in a large number of candidate images exhibit a specific fluorescence pattern, it automatically enhances its semantic understanding of this visual pattern, even if this feature is not explicitly stated in the original query. When processing image sets containing specific professional visual patterns, it can automatically discover and strengthen the understanding of key visual features, improving the retrieval accuracy from the initial 76% to 89%. In particular, when detecting image series commonly found in academic literature (such as time series experimental images or dose gradient response images), the system can identify common visual patterns and changing trends in the series of images, and effectively identify the same experimental data used in different papers, even if these images have been processed and presented in different ways. The successful detection rate has increased from 62% to 91%, improving the ability to trace the origin of academic images.

[0256] The preferred embodiments of the present invention are described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the scope of protection of the present invention.

Claims

1. A similar image retrieval method, characterized in that: include: Receive user query data and build a hierarchical semantic-visual representation framework, including hierarchical semantic and visual representations; Based on hierarchical semantic representation, a context-aware terminology understanding system is used to generate context feature vectors, hierarchical weight vectors, and context-enhanced query representations. Based on context-enhanced query representation and hierarchical visual representation, a semantic-visual bidirectional mapping and optimization framework is adopted to generate an optimized mapping matrix and an optimized candidate set. Combined with the context feature vector, multimodal similarity calculation is performed to obtain multi-level similarity. The multi-level similarity and the level weight vector are combined into a comprehensive similarity score, and the optimized candidate set is sorted based on it, and the search result set is output; Construct a hierarchical semantic-visual representation framework, including: Encode user query data to obtain a basic query vector representation; split it into core concepts, attribute modifications, and relationship description representations, and integrate them into a hierarchical semantic representation; Extract object-level, attribute-level, and relation-level visual features from images in an image library and integrate them into a hierarchical visual representation; Based on hierarchical semantics and visual representation, a strict correspondence is established to form a hierarchical semantic-visual representation framework; Generate the optimized mapping matrix and optimized candidate set, including: Convert the context-enhanced query representation into a visual query representation, generate a multi-level visual attention map, and adjust it based on the context feature vector to guide visual feature matching and calculate the forward mapping similarity matrix; Based on the forward mapping similarity matrix, a candidate image set is selected and its visual feature distribution characteristics are analyzed. A visual-semantic bridging network is constructed to project the visual features back into the semantic space, identify semantic differences, and dynamically adjust the semantic understanding to generate a reverse mapping similarity matrix. Based on the forward and reverse mapping similarity matrices, bidirectional mapping optimization is performed to generate an optimized mapping matrix, and the optimized candidate set is obtained by screening the candidate image set; Generate a reverse mapping similarity matrix, including: Perform visual feature distribution analysis on the images in the candidate image set to identify the dominant visual pattern; based on this, construct a visual-semantic bridging network to project the visual features into the semantic space and generate reverse semantic descriptions; The reverse semantic description is compared with the original query semantics to identify potential semantic differences; it is combined with the context feature vector to dynamically adjust the semantic understanding model parameters and generate a reverse mapping similarity matrix.

2. The method according to claim 1, characterized in that Generate contextual feature vectors, hierarchical weight vectors, and context-enhanced query representations, including: Identify professional terms from the hierarchical semantic representation and obtain a professional term list; combine it with the hierarchical semantic representation to analyze the professional domain background and purpose of the query, generate domain feature, purpose feature and knowledge background vectors, and fuse them into a context feature vector; The context feature vector is input into the weight generation network to dynamically generate a hierarchical weight vector based on the professional domain characteristics and purpose of the query; Based on the professional term list and context feature vector, term ambiguity is resolved, and the disambiguated semantic representation is obtained and fused with the context feature vector to generate a context-enhanced query representation.

3. The method according to claim 2, characterized in that Get a list of specialized terms, including: Match the vocabulary in the hierarchical semantic representation with the pre-trained domain terminology vocabulary to generate a preliminary set of domain term candidates; verify the domain adaptability of each term, calculate the domain belonging probability of the term, and generate a verified set of domain term candidates; For each term in the verified professional term candidate set, its relevance in the current query context is evaluated based on the semantic link relationship and sentence structure, the context importance score is calculated, and the term list is filtered and sorted to obtain the professional term list.

4. The method according to claim 2, characterized in that Get the context feature vector, including: Based on the hierarchical semantic representation and the term list, the following steps are performed: Identify the set of professional fields to which the query belongs, calculate the probability distribution of belonging to each field, and generate the field feature vector; Analyze the query purpose and consider the application scenario to generate the purpose feature vector; Retrieve relevant professional background knowledge from the pre-built professional knowledge base, extract key concepts and relationships, and generate knowledge background vectors; The domain features, purpose features and knowledge background vectors are integrated into the context feature vector.

5. The method according to claim 1, wherein Integrate into a hierarchical semantic representation, including: Based on the basic query vector representation, identify and verify the subject and key entities, extract core concepts and assign weight values ​​to generate core concept representation; Based on the basic query vector representation, identify the adjectives, status words and feature descriptions that modify the core concept, construct attribute triples, and generate attribute modification representations; For the basic query vector representation, identify the relational verbs, state changes, and time sequence information between core concepts, construct relation quadruples, and generate relational description representations; Integrate core concepts, attribute modifications and relationship descriptions into a hierarchical semantic representation.

6. The method according to claim 1, characterized in that Integrate into a hierarchical visual representation, including: Identify the objects to be inspected from the image, extract the feature vector of each object to be inspected, and generate object-level visual features; For the object to be inspected, identify domain attributes, calculate attribute confidence scores, construct attribute triples, and generate attribute-level visual features; Based on the objects to be inspected and their attributes, the spatial and functional relationships between the objects to be inspected are identified, a relationship graph between the objects to be inspected is constructed, and relation-level visual features are generated; Object-level, attribute-level, and relationship-level visual features are integrated into a hierarchical visual representation that strictly corresponds to the semantic hierarchy.

7. The method according to claim 1, characterized in that Calculate the forward mapping similarity matrix, including: The context-enhanced query representation is fed into a semantic-visual transformation network to generate visual query representations at the object level, attribute level, and relation level. Based on the visual query representation, the similarity distribution of the hierarchical visual features in the image library is calculated and a multi-level visual attention map is generated; Combined with the context feature vector, the multi-level visual attention map is adaptively adjusted to guide visual feature matching and calculate the forward mapping similarity matrix.

Citation Information

Patent Citations

  • Multimodal semantic analysis and image retrieval

    US20240354336A1

  • Large model-based search method and system, device and medium

    WO2025097484A1