A visual content identification and categorization system and method thereof

WO2026202959A1PCT designated stage Publication Date: 2026-10-01EICON VISION PVT LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IN2026/050546
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-26
Publication Date
2026-10-01

Smart Images

  • Figure IN2026050546_01102026_PF_FP_ABST
    Figure IN2026050546_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a system 100 for visual content identification and retrieval The visual content identification and retrieval system 100 includes an input validator and preprocessor module 112 for validating and preprocessing the received input; a multi-segment generation module 121 configured to partition the at least one input into a plurality of discrete segments; a feature extraction module 122 configured to extract a plurality of raw feature vectors from each discrete segment; a normalization module 124 configured to rescale the plurality of raw feature vectors into a standardized range to generate a plurality of normalized feature vectors; an aggregation module 126 configured to perform mean pooling over the plurality of feature vectors to generate a unique identifier; a search module 130 configured to identify and retrieve the input by comparing the unique identifier with stored unique identifiers. Also disclosed is a method for visual content identification and retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A VISUAL CONTENT IDENTIFICATION AND CATEGORIZATION SYSTEM AND METHOD THEREOF FIELD OF THE DISCLOSURE

[0002] The present disclosure generally relates to the field of image processing, artificial intelligence and machine learning, more particularly, to a system and method for an image-based artwork identification, retrieval, and classification system.

[0003] BACKGROUND

[0004] This section is intended to provide information relating to the technical field and thus, any approach or functionality described below should not be assumed to be qualified as prior art merely by its inclusion in this section.

[0005] The identification and categorization of visual content items have long been important tasks in the art world, given the cultural, historical, and monetary value of such pieces. Cataloguing and correctly attributing artworks requires tracing their history, known as provenance, which includes documenting previous ownership, exhibitions, and sales transactions. Art experts and curators play a central role in this process. Their deep knowledge of an artist's style, techniques, and historical context allows them to identify and classify works. However, when gaps or inconsistencies in provenance arise, such as missing ownership records or unclear exhibition histories, it can be difficult to accurately identify and categorize the piece.

[0006] In recent years, the rise of digital technologies has introduced new tools to supplement expert analysis. Human experts use their experience and knowledge, combined with data from scientific techniques such as ultra-violet light fluorescence, X-ray analysis, infrared reflectography and pigment analysis to study artworks in detail. These tools are useful for revealing hidden details, identifying materials, and understanding an artwork's construction. For example, X-ray analysis can show hidden layers or changes made by the artist, while infrared reflectography uncovers preliminary drawings beneath the surface. Nevertheless, the above-mentioned tools reveal only certain aspects of an artwork such as hidden layers or preliminary sketches and fail to provide a full picture ofthe artwork's identity and classification. They also depend on experts for interpretation, and if the expert is not sufficiently familiar with the artist's techniques or the specific historical context, their conclusions might be unreliable.

[0007] In addition to traditional methods, reverse image search engines like TinEye, Google Reverse Image Search, and Yandex have emerged as digital tools that can quickly search for similar images across large online databases. These tools identify images of artworks or photographs that are visually similar to the query image but do not verify the identity of the specific artwork. Further, when an artwork is imaged under different lighting conditions or orientations, it may appear different and may not yield the correct match from the database. Furthermore, these tools lack the ability to account for stylistic and contextual features of artworks, such as artistic technique, period, or medium, which are necessary for accurate identification and categorization. As a result, without capturing these distinguishing features, the tools may struggle to accurately identify and classify artworks.

[0008] Therefore, there is a need for an improved technology that is more efficient, cost effective, reliable and capable of identifying and categorizing artwork in real time.

[0009] SUMMARY

[0010] This section is provided to introduce certain objects and aspects of the present disclosure in a simplified form that are further described below in the detailed description. This summary is not intended to identify the key features or the scope of the claimed subject matter.

[0011] It is an object of the present disclosure to develop a system for identifying, categorizing and authenticating visual content items accurately.

[0012] It is another object of the present disclosure to develop a user-friendly system for identifying and categorizing visual content items with improved accessibility and compatibility across various devices and platforms.

[0013] It is another object of the present disclosure to develop a system for identifying and categorizing visual content items that learns and improves over time, becoming moreaccurate at identifying, retrieving, and classifying artworks. It learns from image data by extracting and using multiple pieces of visible information such as colour, texture, artist signatures, and materials used in the painting. Physical dimensions may also be inferred where 3D reconstruction of the visual content items is available.

[0014] It is another object of the present disclosure to develop a system with improved preprocessing that eliminates the effects of non-uniform lighting, camera poses and reflections.

[0015] In one aspect of the present disclosure, a visual content identification and retrieval system is provided. The system includes an input validator and preprocessor module for validating and preprocessing the received input, a multi-segment generation module configured to partition the at least one input into a plurality of discrete segments, a feature extraction module configured to extract a plurality of raw feature vectors from each discrete segment, a normalization module configured to rescale the plurality of raw feature vectors into a standardized range to generate a plurality of normalized feature vectors, an aggregation module configured to perform mean pooling over the plurality of feature vectors to generate a unique identifier, a reconstruction module configured to regenerate an approximate visual representation from the compact feature embedding and store in the storage device , a search module configured to identify and retrieve the input by comparing the unique identifier with stored unique identifiers.

[0016] In a preferred embodiment, the multi-segment generation module may perform segmentation using a multi-scale crop strategy, wherein a first crop is extracted as a centre-aligned crop, and one or more additional crops are extracted at progressively reduced spatial scales from varying positions within the image.

[0017] In a preferred embodiment, visual content authentication system as claimed in claim, wherein the feature extraction module includes a pre-feature extraction module configured to normalize each discrete segment, a content extraction module configured to extract structural and semantic information associated with spatial arrangement of each partitioned segment from the pre-processed input, a style extraction module configured to extract appearance-based features including texture patterns and creation characteristics of the image data.In an embodiment, the aggregation module comprises a compression module configured to reduce the size of the unique identifier.

[0018] In an embodiment, the search module is configured to calculate a cosine similarity score between the compressed feature signature and the stored reference signatures and identify a match only when the cosine similarity score exceeds a predetermined similarity threshold.

[0019] In an embodiment, the reconstruction module is configured to regenerate an approximate visual representation from the compact feature embedding and store in the storage device.

[0020] In another aspect, method for visual content identification and retrieval, comprising the steps of validating the received input followed by preprocessing by an input validator and preprocessor module, partitioning the at least one input into a plurality of discrete segments by a multi-segment generation module, extracting a plurality of raw feature vectors from each discrete segment by a feature extraction module, rescaling the plurality of raw feature vectors into a standardized range to generate a plurality of normalized feature vectors by a normalization module, performing mean pooling over the plurality of feature vectors to generate a unique identifier by an aggregation module, regenerating an approximate visual representation from the compact feature embedding by a reconstruction module, identifying and retrieving the input by comparing the unique identifier with stored unique identifiers by a search module.

[0021] BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings, which are incorporated herein, and constitute a part of this disclosure, illustrate exemplary embodiments of the disclosed methods and systems in which like reference numerals refer to the same parts throughout the different drawings. Components in the drawings are not necessarily to scale; emphasis instead being placed upon clearly illustrating the principles of the present disclosure.

[0023] Although exemplary connections between sub-components have been shown in the accompanying drawings, it will be appreciated by those skilled in the art, that other connections may also be possible, without departing from the scope of the disclosure. Allsub-components within a component may be connected to each other, unless otherwise indicated.

[0024] Figure 1 illustrates a block diagram of an image-based artwork identification and categorization system according to the present disclosure.

[0025] Figure 2 illustrates a block diagram of the processor according to the present disclosure.

[0026] Figure 3 illustrates a schematic illustration of the end-to-end processing pipeline of the image-based artwork identification and categorization system according to the present disclosure.

[0027] Figure 4 is a flowchart illustrating a method of identifying and categorizing an artwork according to the present disclosure.

[0028] Figure 5 is a flowchart illustrating the different steps involved in identifying and categorizing an artwork according to the present disclosure.

[0029] The foregoing shall be more apparent from the following more detailed description of the disclosure.

[0030] DESCRIPTION OF THE DISCLOSURE

[0031] Exemplary embodiments now will be described with reference to the accompanying drawings. The disclosure may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey its scope to those skilled in the art. The terminology used in the detailed description of the particular exemplary embodiments illustrated in the accompanying drawings is not intended to be limiting. In the drawings, like numbers refer to like elements.

[0032] The specification may refer to “an”, “one” or “some” embodiment s) in several locations. This does not necessarily imply that each such reference is to the same embodiment s), or that the feature only applies to a single embodiment. Single features of different embodiments may also be combined to provide other embodiments.As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless expressly stated otherwise. It will be further understood that the terms “include”, “comprises”, “including” and / or “comprising” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It will be understood that when an element is referred to as being “connected” or “coupled” to another element, it can be directly connected or coupled to the other element or intervening elements may be present. Furthermore, “connected” or “coupled” as used herein may include wirelessly connected or coupled. As used herein, the term “and / or” includes any and all combinations and arrangements of one or more of the associated listed items.

[0033] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure pertains. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0034] In the following description, for the purposes of explanation, numerous specific details have been set forth in order to provide a description of the disclosure. It will be apparent, however, that the disclosure may be practiced without these specific details and features.

[0035] The present disclosure provides a visual content identification and retrieval system. As used herein, visual content refers to any data or information capable of being represented or displayed in a visual format including but not limited to digital images, and artworks such as paintings, sculptures, illustrations, graphic designs, three-dimensional models and architectural renderings. The present disclosure analyses visual content through their visual (image-based) features, such as patterns, colours, shapes, or textures, and then expressing those features mathematically for example, through formulas, models, or algorithms.The system converts an image or series of images of visual content items into a numerical or mathematical representation (hereinafter referred to as a "fingerprint"), which serves as a unique identifier for that artwork. In an example, the initial fingerprint is generated from the image features of an artwork, and may include data such as the artist's signature, colour palette of the artwork, and materials used in the painting, all of which can be inferred from images. Physical dimensions and surface attributes of the visual content may be added to supplement the fingerprint at a later stage, following optional 3D reconstruction of the artwork's surface as described in further detail below.

[0036] As illustrated in Fig. 1, the visual content identification and retrieval system 100 receives an image or series of images of a visual content item from one or more user devices 102a-n. According to an embodiment, the image may be an index image that may represent specific features, characteristics or classifications of data. For example, images of the original artwork may be captured using professional cameras in high-definition or may be acquired from secondary sources like digital or printed catalogues or books.

[0037] The system 100 includes an input validator and preprocessor module 112 for validating and preprocessing the received input.

[0038] According to an embodiment, the input validator and preprocessor module 112 performs image quality validation before preprocessing. The process of quality validation includes a darkness check to reject underexposed images, blur detection by way of edge density analysis, focus assessment by way of local variance analysis, and resolution validation against a minimum threshold. Images that fail to satisfy any of these quality validation checks are rejected and not forwarded for further processing. The quality gating step ensures that only images satisfying the required quality criteria are processed for fingerprint generation, thereby improving the robustness and reliability of downstream identification and categorization operations.

[0039] The input validator and preprocessor module 112 preprocesses the received images from the user device. According to an embodiment, the input validator and preprocessor module 112 preprocesses the images using image processing techniques that involve applying various techniques to enhance the quality by removing artifacts and optimizing the image for further processing. These techniques may include noise reduction, dust andscratch removal, colour correction, contrast enhancement, image imperfections and histogram equalization. According to an embodiment, the input validator and preprocessor module 112 may apply an image processing technique such as Contrast Limited Adaptive Histogram Equalization (CLAHE) to improve local contrast in images where lighting conditions are uneven, thereby preserving detail in both bright and dark regions of the artwork image. However, the embodiment does not limit the scope of the present disclosure.

[0040] According to an embodiment, the input validator and preprocessor module 112 preprocesses the received images using computer vision techniques. The techniques may include background and picture-frame removal, warping and equalization. For example, as illustrated in Fig. 3, high-definition images or index images may be pre-processed using computer vision techniques to adjust distorted or skewed documents or enhance contrast to improve character visibility. The description of the particular embodiment is illustrative and should not be construed as limiting the scope of the present disclosure.

[0041] Alternatively, the system may be configured to receive a query image or a query image fingerprint from one or more user devices 102a.n. If query images are received from the user device 102a.n, the input validator and preprocessor module 112 preprocesses the received query images.

[0042] The system 100 includes a multi-segment generation module 121 configured to partition the at least one input into a plurality of discrete segments.

[0043] In a preferred embodiment, the multi-segment generation module 121 performs segmentation using a multi-scale crop strategy. A primary crop is extracted from the centre of the input image at its full, original resolution. One or more additional crops are extracted at progressively reduced spatial scales from varying positions within the image using random spatial sampling. These crops capture localised details or features at different regions of the input.

[0044] In an embodiment, the multi-segment generation module 121 extracts additional crops at progressively reduced spatial scales from varying positions within the image using random spatial sampling.In an embodiment, the multi-segment generation module 121 may perform segmentation using spatial partitioning techniques, grid-based segmentation, region-of-interest detection, or other segmentation methodologies known in the art.

[0045] Referring to Fig.2, the system 100 includes a feature extraction module 122 configured to extract a plurality of raw feature vectors from each discrete segment. According to an embodiment, for improved handling of partial views, varying zoom levels, and cropping variations, the system optionally generates vectors from discrete segments, which are multiple crops of different sizes taken from the query image. The feature extraction module 122 identifies and extracts relevant features from raw data and converts them into a manageable or informative representation. Various pieces of metadata such as the artist's signature, colour palette of the artwork, materials used in the painting, and image size are extracted from the images. The patterns, shapes, colours, textures, edges, corners, or other visual elements of the artwork may also be identified and represented in a simplified form.

[0046] According to an embodiment, the feature extraction module 122 employs a Vision Transformer-based deep learning model with dual projection heads. The model includes a style branch and a content branch that separately extract style features and content features into high-dimensional vectors. The style branch captures features such as colour palette, texture, and artistic technique, while the content branch captures features such as subject matter, composition, and shapes. Both sets of vectors are L2-normalized to ensure consistent comparison across different artworks.

[0047] In a preferred embodiment, each discrete segment may be normalized using channel-wise mean and standard deviation values derived from a large reference image dataset, ensuring that pixel intensity distributions are consistently scaled across all inputs irrespective of the original capture conditions.

[0048] In an embodiment, the feature extraction module 122 employs a backbone neural network for transforming discrete segment into high-dimensional feature representations.

[0049] According to an embodiment, the feature extraction module 122 employs a multi-crop strategy to generate vectors from a plurality of crops, each crop being extracted from thequery image at two or more configurable spatial scales. A first crop is extracted as a centre-aligned crop at full input resolution, capturing a global view of the artwork. One or more additional crops are extracted at reduced spatial scales from varying positions within the image, capturing localised detail at different regions.

[0050] In some embodiments, the feature extraction module 122 employs a backbone neural network including a large-scale Vision Transformer architecture. Each crop is processed independently through the Vision Transformer-based model to yield a corresponding pair of style and content vectors.

[0051] The backbone neural network processes each discrete segment and generates the corresponding feature embeddings that capture spatial, semantic and contextual visual information. The backbone neural network may be pre-trained on image-text pairs, operating on input images resized to a fixed input resolution. The backbone neural network produces a high-dimensional feature representation of the input image, which is subsequently projected through two independent learned linear projection heads to yield the style vector and the content vector respectively. Each projection head maps the backbone feature representation into a fixed-dimensional fingerprint vector.

[0052] In an embodiment, the style and content features are projected into compact fixeddimensional vectors, for example 768-dimensional and 768-dimensional vectors respectively, by way of learned linear projections. This dimensionality reduction preserves the discriminative information needed for accurate matching while reducing storage and computation requirements.

[0053] The system 100 includes a normalization module 124 which is configured to rescale the plurality of raw feature vectors into a standardized range to generate a plurality of normalized feature vectors. In an embodiment, the normalized feature vectors may be style vector and the content vector. In a preferred embodiment, both the style vector and the content vector may be independently subjected to L2-normalization, whereby each vector is divided by its own Euclidean norm, such that all stored and query fingerprints lie on the surface of a unit hypersphere. This normalization ensures that subsequent similarity comparisons are based solely on the angular relationship between vectors and are invariant to differences in vector magnitude. As used herein, L2-normalization refersto a mathematical operation in which a feature vector is divided by its Euclidean (L2) norm so as to produce a unit-length vector for providing comparison between feature vectors irrespective of their original magnitude. The system 100 includes an aggregation module 126 configured to perform mean pooling over the plurality of feature vectors to generate a unique identifier. The aggregation module 126 computes an aggregated representation by averaging the normalized feature vectors. The normalized feature vectors from discrete segments of a visual content and averages them by mean pooling to create a single fingerprint. The fingerprint generated by aggregating the individual crop vectors by way of mean pooling is more tolerant of framing differences between the query image and the stored index images. This averaged representation reduces the effect of differences in image cropping or positioning between the query image and stored images, improving matching robustness. The per-crop vectors for each branch are aggregated by computing the element-wise arithmetic mean across all crop vectors.

[0054] According to an embodiment the generated fingerprint is normalized by L2-normalization of the resulting mean vector. This aggregation produces a single composite fingerprint for each branch that is robust to partial views, varying zoom levels, cropping variations, and differences in framing between the query image and the stored index images. In an embodiment, the aggregation module 126 includes a compression module 127 configured to reduce the size of the unique identifier. The compression module 127 reduces the dimensionality while retaining the most relevant information. According to an embodiment, the compression techniques might be used in combination with decomposition. As used herein, decomposition of image refers to separation of image into different components or layers based on some intrinsic properties like reflectance, shading, or even segmentation into foreground and background etc. The image may be split into reflectance (albedo), which captures the true colour of the artwork and shading (illumination) which accounts for light and shadow effects.

[0055] The visual content is processed through the input validator and preprocessor module 112, a multi-segment generation module 121, a feature extraction module 122 and a normalization module 124, to generate a unique fingerprint of an artwork based on the image features. The fingerprint relevant to the artwork may be stored in a repository.The system 100 includes a reconstruction module 128 for regenerating an approximate visual representation from the compact feature embedding. The reconstruction module 128 may utilize one or more computational imaging, photometric analysis, or computer vision techniques to derive a three-dimensional representation of the artwork's surface structure from multiple images captured under different viewpoints, illumination conditions, or focus settings.

[0056] According to an optional embodiment directed to further strengthening the fingerprinting process, the image reconstruction module 128 may be configured to reconstruct the 3D geometry of an artwork using multiple index images captured from different angles and / or illumination conditions. This optional reconstruction process may use techniques such as photogrammetry, photometric stereo, shape from shading and shape from defocus, which involve capturing overlapping photographs from various perspectives and illumination conditions and applying computer vision algorithms to generate a 3D model of the artwork's surface structure. According to a further optional embodiment, the image reconstruction module 128 may incorporate photometry, which focuses on capturing fine surface details and textures by analysing lighting variations from a single viewpoint. The 3D surface information obtained through these optional techniques may be used to supplement the fingerprint with physical attributes such as surface texture and dimensions, thereby strengthening the identification and categorization capability of the system.

[0057] The reconstruction module 128 (as shown in Fig. 2), when employed as described above, may further supplement and strengthen the stored fingerprint with additional physical attributes derived from 3D reconstruction.

[0058] According to an embodiment, reconstruction module 128 may store the artwork fingerprints in a vector database that is structured for approximate nearest neighbor search using Hierarchical Navigable Small World (HNSW) graph indexing. Each fingerprint entry in the database is stored alongside a metadata payload that includes the artwork title, artist name, period, medium, source museum or gallery, and geographic coordinates of the holding institution. This indexing structure allows for fast retrieval even when the database contains a large number of fingerprints.In some embodiments, the vector database may maintain multiple independent vector spaces for each stored fingerprint entry, including a first vector space corresponding to a style representation and a second vector space corresponding to a content representation, each defined within a common fixed-dimensional embedding space. The database may further maintain associated metadata attributes and payload indexes. In an example, the payload indexes may be defined on selected metadata fields, such as geographic attributes including the city or country associated with a holding institution, thereby enabling prefiltering of candidate entries prior to performing vector similarity computations.

[0059] In certain implementations, stored vector representations may be maintained using fullprecision numerical formats in order to preserve matching accuracy. The vector database may further employ an approximate nearest-neighbour indexing structure configured to organize stored vectors based on proximity. In an example, the stored vectors may be organized using a proximity-based graph structure that facilitates faster similarity searches. In an example, the approximate nearest-neighbour graph structure is constructed using Hierarchical Navigable Small World indexing, which organises vectors into a layered proximity graph enabling sub-linear search time complexity. The search depth parameter of the graph traversal is adaptively configured as described herein.

[0060] The system 100 includes a search module 130 that identifies and retrieves the input by comparing the generated unique identifier with stored unique identifiers maintained in a vector database.

[0061] In an example, (as illustrated in Fig.4) any subsequently acquired image or a query image of an artwork can then be compared to its fingerprint in the database for identification, retrieval, and / or categorization. The input validator and preprocessor module lllreceives an artwork image, the multi-segment generation module 121 partitions the image input into discrete segments, the feature extraction module 122 extracts a plurality of raw feature vectors from each discrete segment, the normalization module 124 extracts feature vectors and the aggregation module 126 performs mean pooling over the plurality of feature vectors to generate a unique identifier. The fingerprints of artworks, derived from image-based analysis, can be used to identify, retrieve, categorize, and perform other tasks related to an artwork. The search module 130 identifies the artwork upon matching the query image fingerprint with an artwork fingerprint already stored in the database.The search module 130 recognizes a specific artwork from the database or collection of stored fingerprints. According to an embodiment, the identity of an artwork can be found by a straightforward search using the fingerprint of the artwork. For example, if an image of the Mona Lisa is scanned, the system will identify it as Leonardo da Vinci's Mona Lisa by matching it with a stored fingerprint in the database.

[0062] In another example, the system 100 scans an artwork of 'Warli' painting to generate a fingerprint and searches the database for other Warli’ paintings. Alternatively, the system may receive an input fingerprint of a 'Warli' painting and search for similar paintings. According to an embodiment, the search module 130 computes cosine similarity between the query fingerprint vector and the stored index fingerprint vectors. Results that exceed a configurable similarity threshold are returned as matches, ranked by their similarity score. This allows the system to return both exact matches and visually similar artworks in order of relevance.

[0063] According to an embodiment, all fingerprint vectors — both at the time of storage and at the time of query — are L2-normalised prior to any comparison operation. As a consequence, the cosine similarity between a query fingerprint vector q and a stored index fingerprint vector d reduces to their dot product, expressed as: sim (q, d) = q • d, where both q and d are unit vectors. This value ranges from 0 to 1, where a value of 1 indicates identical orientation, corresponding to a perfect match, and a value approaching 0 indicates orthogonality, corresponding to dissimilar artworks.

[0064] According to an embodiment, the search module 130 may employ a dual -branch approach such that two independent cosine similarity scores are obtained for each candidate, a style similarity score and a content similarity score. The style similarity score is derived from the comparison of style vectors, and a content similarity score derived from the comparison of content vectors. These two scores are combined into a single weighted composite score according to the formula: combined score = w style x sim style + w_content x sim_content, where w_style and w_content are configurable weighting parameters whose values sum to 1.0. According to a preferred embodiment, w style is assigned a greater value than w content, reflecting the relatively higher discriminative contribution of stylistic features such as colour palette, texture, and artistic technique in distinguishing between artworks. The weighting parameters may be adjusted based onthe nature of the query or the composition of the database without departing from the scope of the present disclosure.

[0065] According to an embodiment, the style similarity score is additionally retained as a standalone gating signal independent of the composite score, since it is calibrated on a consistent scale and serves as a reliable confidence indicator for threshold filtering. The candidate pool for each search is expanded by a configurable multiplier prior to score merging and re-ranking, ensuring sufficient recall before the final ranked list is produced. The search module 130 returns the highest-ranked result as the primary identification match and up to a configurable number of additional results as alternative matches, all ranked in descending order of their composite similarity score.

[0066] According to an embodiment, the similarity threshold applied by the search module 130 to accept or reject a candidate match is not a single fixed value but is adaptively selected based on contextual parameters associated with the search request, reflecting the principle that geographic proximity and source reliability independently contribute to identification confidence and therefore permit a lower vector similarity requirement to still yield a correct match.

[0067] According to an embodiment, a scope-based threshold is applied in accordance with the geographic search tier in which a candidate result is found. For searches constrained to a local or nearby radius, a lower threshold is applied, reflecting the additional confidence conferred by geographic proximity. For searches conducted at a regional or national scope, an intermediate threshold is applied. For unconstrained global searches in which no geographic pre-filtering is applied, a higher threshold is applied to maintain identification precision. These threshold values reflect the empirically determined relationship between spatial confidence and the minimum vector similarity required to achieve reliable identification and may be adjusted through configuration without departing from the scope of the present disclosure.

[0068] According to an embodiment, when the query originates from a previously stored image rather than a live camera capture, a reduced threshold may be applied to account for systematic differences in image processing between pre-stored and live-captured inputs. According to an embodiment, platform-specific thresholds may be applied to account forsystematic differences in image quality between device types, with higher-quality capture platforms admitting a higher baseline threshold and lower-quality capture platforms admitting a correspondingly lower threshold.

[0069] According to an embodiment, the similarity threshold applied to secondary or alternative results is set lower than the threshold applied to the primary top match, permitting retrieval of visually similar artworks that fall below the primary confidence requirement, thereby enabling comparative analysis and discovery of related works. The primary threshold gate and the alternative threshold gate are applied independently to their respective result slots.

[0070] According to an embodiment, the system employs an adaptive search depth parameter that governs the number of candidate nodes explored during approximate nearest-neighbour graph traversal in the vector database. A higher search depth yields greater recall at the cost of increased search latency. A default search depth is assigned for unconstrained global searches, with reduced values applied for geographically constrained searches that operate over smaller candidate pools. In the event that an initial search returns fewer than a minimum number of candidate results, the search depth is automatically increased by a configurable increment and the search is repeated, ensuring minimum result coverage under sparse database conditions. This adaptive mechanism allows the system to balance identification accuracy against response time across varying database densities and search contexts.

[0071] The search module 130 categorizes or classifies an artwork into one or more predefined categories based on features extracted by feature extraction module 122, such as its artistic style (e.g., Impressionism, Baroque), subject matter (e.g., portrait, landscape), or medium (e.g., oil painting, sculpture). The search module 130 retrieves artworks that are similar to an artwork based on specific characteristics such as style, composition, or colour palette from the database or collection of stored fingerprints.

[0072] According to an embodiment, the system incorporates geographic context from the user device to identify their current position. The system is configured to execute a sequential search algorithm such that when a user submits a query, the system performs search in expanding location tiers, starting from local, then nearby, regional, national, and finallyglobal scope. In an example, the system employs a geospatial search optimization strategy for searching an artwork for reducing the search space dynamically and prioritizing local results to improve user experience.

[0073] According to an embodiment, location-based pre-filtering is applied at the database level to reduce the search space, and proximity-based score boosting is applied to prioritize artworks held in nearby museums or galleries. This location-aware approach reduces response time and improves the relevance of results for users who are physically near the artwork they are querying.

[0074] In an example, a user at a physical location such as a gallery in Bihar, India, uses a mobile device to scan a physical artwork, such as a ‘Madhubani’ painting. The system 100 processes the image to generate a unique fingerprint. The system first searches the local tier (e.g., the specific gallery's collection) for matching or related ‘Madhubani’ fingerprints. While 'Madhubani' paintings may exist in global databases (e.g., a museum in London), the system 100 applies a proximity-based score boost to results held in nearby institutions.

[0075] According to an embodiment, artworks are categorized based on metadata associated with matched fingerprints, including artistic style, period, medium, and artist. The system also enables retrieval of visually similar artworks based on fingerprint proximity in the vector space, allowing users to discover related works that share stylistic or compositional characteristics with the queried artwork.

[0076] As illustrated in Fig. 5, the method 200 for identifying and categorizing artwork includes the following steps:

[0077] At a step 202, an input validator and preprocessor module receives artwork fingerprint or images from multiple user devices 102a-n and validates the received input followed by preprocessing.

[0078] At a step 204, a multi-segment generation module 121 partitions the input into a plurality of discrete segments. At a step 206, a feature extraction module 122 extracts a plurality of raw feature vectors from each discrete segment. At a step 208, a normalization module 124rescales the plurality of raw feature vectors into a standardized range to generate aplurality of normalized feature vectors. At a step 210, an aggregation module 126 performs mean pooling over the plurality of feature vectors to generate a unique identifier by an aggregation module. At a step 212, the reconstruction module 128 regenerating an approximate visual representation from the compact feature embedding. At a final step 214, a search module 130 identifies and retrieves the input by comparing the unique identifier with stored unique identifiers.

[0079] The system is configured to be updated with additional training data over time. According to an embodiment, the system employs a human-in-the-loop regression training process in which domain experts review and annotate identification and categorization results produced by the system. The expert feedback is used to generate corrective training signals, which are fed back into the model to refine its feature extraction and matching capabilities. This iterative training cycle allows the system to continuously improve its accuracy in identifying, retrieving, and categorizing artworks as more annotated data becomes available. By training on large datasets of artwork, the system can objectively identify unique characteristics, such as textures, colours, colour palette of the artwork, materials used in the painting, image size or structural details, and use this knowledge to improve identification, retrieval, and categorization.

[0080] The present disclosure offers numerous advantages related to the image-based identification and categorization system. A few of the advantages achieved using the features of the present disclosure are provided below:

[0081] • The present disclosure provides a system having a dual-branch fingerprint architecture with separate style and content projection heads that enable independent analysis of stylistic attributes that enables more granular and accurate similarity matching and categorization compared to conventional singlefingerprint image retrieval approaches.

[0082] • The present disclosure provides a system incorporating a location-aware tiered search mechanism configured to reduce search space and improve response time. The system further enhances retrieval relevance by boosting similarity scores associated with geographically proximate collections, thereby improving both efficiency and contextual accuracy of artwork identification.• The present disclosure provides non-contact methods of identifying and searching a visual content item that preserves the integrity of the artwork while conducting detailed analysis.

[0083] • The present disclosure provides that analyses multiple cropped regions of a query image and aggregates the resulting features, thereby improving the recognition accuracy, noise resistance and robustness to varied, real-world conditions.

[0084] • The present disclosure provides a system with improved accuracy in tasks such as identification, retrieval, and categorization by learning from large datasets of artwork.

[0085] • The present disclosure provides a flexible and device-agnostic system that ensures compatibility across multiple platforms and devices.

[0086] • The present disclosure provides enhanced reliability by combining advanced machine learning algorithms with data-driven decision-making.

[0087] • The present disclosure provides a system having a quality-gated preprocessing pipeline that performs automated validation of captured images through darkness detection, blur detection, focus assessment, and resolution verification. This preprocessing stage ensures that only images meeting predefined quality thresholds are processed for fingerprint generation, thereby reducing false matches and improving overall system reliability and accuracy.

[0088] The present disclosure provides faster processing of large datasets, enabling quicker identification and analysis of artworks compared to traditional methods. While the present disclosure has been described with reference to certain preferred embodiments and examples thereof, other embodiments, equivalents, and modifications are possible and are also encompassed by the scope of the present disclosure.

Claims

1. WE CLAIM1. A visual content identification and retrieval system (100), comprising:an input validator and preprocessor module (112) for validating and preprocessing the received input;a multi-segment generation module (121) configured to partition the at least one input into a plurality of discrete segments;a feature extraction module (122) configured to extract a plurality of raw feature vectors from each discrete segment;a normalization module (124) configured to rescale the plurality of raw feature vectors into a standardized range to generate a plurality of normalized feature vectors;an aggregation module (126) configured to perform mean pooling over the plurality of feature vectors to generate a unique identifier;a search module (130) configured to identify and retrieve the input by comparing the unique identifier with stored unique identifiers.

2. The visual content authentication system (100) as claimed in claim 1, wherein the feature extraction module (122) comprises:a pre-feature extraction module (132) configured to normalize each discrete segment.

3. The visual content authentication system (100) as claimed in claim 1, wherein the feature extraction module (122) comprises:a content extraction module (134) configured to extract structural and semantic information associated with spatial arrangement of each partitioned segment from the pre-processed input;a style extraction module (136) configured to extract appearance-based features including texture patterns and creation characteristics of the image data.

4. The visual content identification and retrieval system (100) as claimed in claim 1, wherein the search module (130) is configured to calculate a cosine similarity score between the compressed feature signature and the stored reference signatures; andidentify a match only when the cosine similarity score exceeds a predetermined similarity threshold.

5. The visual content identification and retrieval system (100) as claimed in claim 1, wherein the reconstruction module (128) is configured to regenerate an approximate visual representation from the compact feature embedding and store in the storage device (104).

6. A method for visual content identification and retrieval, comprising the steps of validating the received input followed by preprocessing by an input validator and preprocessor module (112);partitioning the at least one input into a plurality of discrete segments by a multi-segment generation module (121);extracting a plurality of raw feature vectors from each discrete segment by a feature extraction module (122);rescaling the plurality of raw feature vectors into a standardized range to generate a plurality of normalized feature vectors by a normalization module (124);performing mean pooling over the plurality of feature vectors to generate a unique identifier by an aggregation module (126);regenerating an approximate visual representation from the compact feature embedding by a reconstruction module (128);identifying and retrieving the input by comparing the unique identifier with stored unique identifiers by a search module (130).

7. The method for visual content identification and retrieval as claimed in claim 6, wherein the feature extraction module (122) comprises:pre-feature extraction module (132) configured to normalize each discrete segment.

8. The method for visual content identification and retrieval as claimed in claim 6, wherein the feature extraction module (122) comprises:a content extraction module (134) configured to extract structural and semantic information associated with spatial arrangement of each partitioned segment from the pre-processed input;a style extraction module (136) configured to extract appearancebased features including texture patterns and creation characteristics of the image data.

9. The method for visual content identification and retrieval as claimed in claim 6, wherein the search module (130) is configured to calculate a cosine similarity score between the compressed feature signature and the stored reference signatures; andidentify a match only when the cosine similarity score exceeds a predetermined similarity threshold.

10. The method for visual content identification and retrieval as claimed in claim 6, wherein the reconstruction module (128) is configured to regenerate an approximate visual representation from the compact feature embedding and store in the storage device (104).