A system and method for storing and validating digital representations of physical objects

Hyperspectral imaging and machine learning with distributed ledger technology enable reliable authentication and preservation of artistic works and cultural objects, addressing the challenges of distinguishing originals and detecting changes.

WO2026020203A1PCT designated stage Publication Date: 2026-01-29UNIVERSITY OF MELBOURNE
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/AU2025/050791
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-24
Filing Date
2025-07-24
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing methods for authenticating and preserving artistic works and cultural objects are unreliable and lack statistical surety, as they often fail to distinguish originals from copies, and changes in these items are difficult to detect over time.

Method used

A system and method using hyperspectral imaging and machine learning to create multi-dimensional arrays of artwork data, which are compared to reference datasets, and stored on a distributed ledger with non-fungible tokens for authentication and preservation.

Benefits of technology

Provides reliable and repeatable authentication of artistic works and cultural objects, ensuring their integrity and value by detecting alterations and maintaining a secure, decentralized record of their condition and ownership.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AU2025050791_29012026_PF_FP_ABST
    Figure AU2025050791_29012026_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for identifying, validating or authenticating an artwork using one or more digital representations is provided which includes acquiring image data of an artwork in one or more regions of the electromagnetic (EM) spectrum at one or more sample regions of the artwork; forming a multi-dimensional array of the image data; creating a dataset from the multi-dimensional array, wherein the artwork may be identified from the dataset. The created dataset is compared to one or more previously generated datasets relating to previously acquired image data obtained from a previous encounter of the artwork to evaluate if the artwork is the same as the previous encounter of the artwork.
Need to check novelty before this filing date? Find Prior Art

Description

A SYSTEM AND METHOD FOR STORING AND VALIDATING DIGITAL REPRESENTATIONS OF PHYSICAL OBJECTSTechnical Field

[0001] The invention generally relates to a system for storing and identifying digital representations of physical objects for the purpose of authentication and / or validation of those objects, and more specifically to artistic works, high value artifacts and objects of cultural importance.Background of Invention

[0002] Artwork and artistic work are broad terms that encompasses various forms of human expression and creativity. Artworks can take countless forms, including paintings, sculptures, drawings and photographs. Each medium offers unique opportunities for artists to explore and express their visions, utilizing different techniques, materials, and styles.

[0003] Among the earliest documented expressions of art are visual forms that depict images and objects. Visual artistic creations encompass a range of mediums, including paintings, sculptures, printmaking, photography and other visual media. Paintings stand out as particularly prevalent in the realm of visual or fine arts. Typically, a painting is created by applying paint, pigment, or other colour mediums onto a surface, such as fabric or a wooden panel. Additionally, in some cases, protective coatings are applied to the top layers of a painting to enhance colour saturation and safeguard the underlying paint from dirt, abrasion, and moisture.

[0004] Artworks such as paintings are often valuable for their artistic merit, historical significance, rarity and cultural and symbolic importance. Because of this, artworks are often reproduced or 'forged' to such an extent that they become difficult to authenticate against an original. Additionally, changes in artworks, for example, damage or deterioration can be difficult to detect over time by simply relying on the human eye. Such damage or deterioration if left untreated can alter artworks and subsequently devalue them, both from a monetary perspective and a cultural perspective.

[0005] Unlike artworks, which are primarily valued for their aesthetic or expressive qualities, cultural objects are valued for their role in preserving or representing aspects of a culture's heritage, identity, beliefs, or practices. They too are often reproduced or 'forged' to the extent that it is difficult to determine an original from a reproduction by the naked eye.

[0006] As a consequence, challenges arise in protecting and preserving these artistic works or cultural objects, which can affect their value.

[0007] Some recent efforts have been made to try to use computational imaging techniques including artificial intelligence to analyse and compare visible spectrum images to identify art. For example, high-resolution cameras can capture minute details of the artwork, including brushstrokes, texture, and patterns that are not visible to the naked eye. These details can be compared to known authentic pieces. However, these efforts have not resulted in reliable repeatable results. Furthermore, these efforts do not provide statistical surety that an artwork is not a copy.

[0008] It would be desirable to provide a system and method for identifying, validating or authenticating artistic works or cultural objects.

[0009] It would also be desirable to provide a system and method for creating and storing digital representations of artistic works or cultural objects such that they can be studied over time or used for the purpose of acting as reference datasets for future authentication of respective artworks.

[0010] A reference herein to a patent document or other matter which is given as prior art is not to be taken as an admission or a suggestion that the document or matter was known or that the information it contains was part of the common general knowledge as at the priority date of any of the claims.Summary of Invention

[0011] According to an aspect of the present invention, there is provided a method for identifying, validating or authenticating an artwork using one or more digital representations, comprising: acquiring image data of an artwork in one or more regions of the electromagnetic (EM) spectrum at one or more sample regions of the artwork; forming a multi-dimensional array of the image data; creating a dataset from the multi-dimensional array, wherein the artwork may be identified from the dataset; and comparing the created dataset to one or more previously generated datasets relating to previously acquired image data obtained from a previous encounter of the artwork to evaluate if the artwork is the same as the previous encounter of the artwork. Depending on the application, in an embodiment, the system may or may not require registration of the created datasets before comparing the created dataset to one or more previously generated datasets.

[0012] In one or more embodiments, the one or more regions of the EM spectrum includes an infrared (IR) spectrum, a visible light spectrum, an ultraviolet (UV) spectrum, or an X-ray spectrum, or a combination thereof.

[0013] In one or more embodiments, the multi-dimensional array of the image data comprises a hyperspectral data cube.

[0014] In one or more embodiments, the multi-dimensional array of the image data comprises data from two spatial dimensions and one spectral dimension.

[0015] In one or more embodiments, the spectral dimension comprises a range of wavelengths in the image data.

[0016] In one or more embodiments, the method further comprises cropping the multidimensional array of the image data to one or more of the size of the artwork, or a region within the artwork. It will be appreciated that depending on the artwork or the size of the artwork, a region within the artwork may instead be of interest rather than the entire artwork.

[0017] In one or more embodiments, the cropping includes wavelength cropping to remove wavelengths from the multi-dimensional array of the image data at wavelengths above or below a predetermined threshold.

[0018] In one or more embodiments, the method further comprises registering the multidimensional array of the image data and / or created dataset with at least one of a location and an orientation of the artwork.

[0019] In one or more embodiments, the method further comprises reducing a dimension of the multi-dimensional array.

[0020] In one or more embodiments, reducing the dimension of the multi-dimensional array of the image data is performed via Principal Component Analysis (PCA) or a manifold learning method.

[0021] In one or more embodiments, the step of creating a dataset from the multi-dimensional array comprises creating a new dataset (which may take the form of an image-level similarity map) from the starting dataset of the cropped multi-dimensional array of the image data, the method comprising the steps of: first (a) identifying a primary pixel in the starting dataset at a spatial position; then (b) identifying a secondary pixel in the starting dataset at a different spatial position; then (c) determining the similarity of the spectral dimension of the primary pixel with the spectral dimension of the secondary pixel, using a spectral similarity function; then (d) repeating steps (b) and (c) for each secondary pixel in the starting dataset, to create the new intermediate dataset (which may take the form of a pixel-level similarity map) containing spectral similarity values for every secondary pixel identified in step (b) when compared to the primary pixel identified in step (a). As will be appreciated, these are iterative steps to identify a primary and in turn a secondary pixel in the created dataset at a spatial position and to determine the similarity of the spectral dimension of the primary pixel with the spectral dimension of the secondary pixel.

[0022] The method may further comprise, determining one value from the intermediate dataset (which may take the form of a pixel-level similarity map), which will be attached to the primary pixel: equal to the proportion of secondary pixels with a spectral similarity measure assessed against a predetermined threshold. A full analytic dataset is then created with the same spatial dimensions of the starting dataset, by repeating steps (a) to (e) for each primary pixel in the spatial dimension; each spatial position of the new dataset (which may take the form of an image-level similarity map) is populated with the value calculated in step (e) when that spatial position was the primary pixel.

[0023] In an embodiment, the intermediate dataset (pixel-level similarity map) is a derived subset of the hyperspectral data. In an embodiment, the new dataset (image-level similarity map) is a derived subset of the intermediate dataset (pixel-level similarity map).

[0024] Advantageously, the similarity map does not require dimension reduction and / or feature extraction processes. It will be appreciated that the spectral similarity function may take any suitable form. For example, Spectral Angle Mapping (SAM) or another suitable spectral similarity function with improved accuracy or improved spectral comparisons.

[0025] In a further advantage, the pixel-level similarity map and image-level similarity map, and in turn, comparison decisions may be derived from a subset of the hyperspectral data - obviating the need to operate on hyperspectral raw data. The subset may be derived in any suitable manner such as a proportion calculation and then a thresholding (min, max) calculation. Advantageously, smaller derived datasets can be more easily stored, compared and transferred.

[0026] In an embodiment, the wavelengths and / or spatial dimensions may be manipulated and / or cropped. This may be carried out in a manual or automated fashion. For example, the edges of the cropping space may be manually determined - spatially and by wavelength and then the pixels and wavelengths from the hypercube are removed in making a suitable dataset.

[0027] In one or more embodiments, the method further comprises using one or more feature extraction techniques to localise features in the multi-dimensional array of the image data.

[0028] In one or more embodiments, the method further comprises extracting spectral features by applying a pre-trained machine learning model to the multi-dimensional array of the image data; and determining characteristics of the artwork based on a comparison of the spectral features from the multi-dimensional array of the image data.

[0029] In one or more embodiments, the machine learning model is trained using a training set of hyperspectral images and is configured to classify a plurality of spectral features in a plurality of hyperspectral images.

[0030] In one or more embodiments, the machine learning model includes using a supervised classifier method.

[0031] In one or more embodiments, the supervised classifier method is a multi-layer perceptron method.

[0032] In one or more embodiments, the artwork comprises a painting, a sculpture, a drawing, a photograph, a cultural object, an object or a printed material.

[0033] According to an aspect of the present invention, there is provided a method for creating and storing a digital representation of an artwork, comprising: acquiring image data of an artwork in one or more regions of the electromagnetic (EM) spectrum at one or more sample regions of the artwork; forming a multi-dimensional array of the image data; creating one or more datasets from the multi-dimensional array, wherein the artwork may be identified from the created dataset; and storing on a distributed ledger one or more of the created dataset or a reference to the created dataset. The distributed ledger may include a database, such as a blockchain. Advantageously, the use of a "distributed ledger" means that any records (e.g., records relating to an artwork) are authenticated by a federated consensus protocol. A distributed ledger may be a blockchain. Multiple computer systems within the distributed ledger, referred to herein as "nodes" or "compute nodes," each comprise a copy of the entire ledger of records. In some embodiments, the node may also be a "light node" in that it is similar to a node in terms of its functions. However, instead of storing a copy of the entire ledger of records or blockchain in its memory, it only stores parts of the ledger or blockchain that are relevant to the transaction being performed. Nodes may write a data "block" to the distributed ledger, the block comprising data regarding an electronic event, said blocks further comprising data and / or metadata. In some embodiments, only miner nodes may write electronic events to the distributed ledger. In other embodiments, all nodes have the ability to write to the distributed ledger. In some embodiments, the block may further comprise a time stamp and a pointer to the previous block in the chain (e.g., a "hash"). In some embodiments, the block may further comprise metadata indicating the node that was the originator of the electronic event. In this way, the entire record of electronic events is not dependent on a single database which may serve as a single point of failure; the distributed ledger will persist so long as the nodes on the distributed ledger persist.

[0034] In one or more embodiments, the method further comprises generating a hash based on at least a part of the dataset; and storing on a distributed ledger the generated hash.

[0035] In one or more embodiments, storing on the distributed ledger comprises minting a non- fungible token (NFT).

[0036] In one or more embodiments, the method further comprises maintaining, by a set of node devices, the distributed ledger that stores the NFT.

[0037] In one or more embodiments, the distributed ledger is a blockchain.

[0038] In one or more embodiments, the set of node devices host the blockchain.

[0039] In one or more embodiments, the method further comprises receiving information relating to the artwork. The information may include metadata relating to the artwork, including equipment used for the capture, distance of imaging tool from artwork, exposure settings, light source, distance of light source from artwork and angle and the like.

[0040] In one or more embodiments, the hash is generated based on at least a part of the dataset and the information relating to the artwork.

[0041] In one or more embodiments, the method further comprises calculating an original validation value associated with the information relating to the artwork via at least one rule from a rules engine.

[0042] In one or more embodiments, at least one rule from the rules engine is a smart contract. A smart contract in the context of a Non-Fungible Token (NFT) is a self-executing piece of code on a blockchain that defines and enforces the ownership, transfer, and unique attributes of the NFT. Advantageously, it ensures that each NFT is distinct and verifiable, governing the creation, distribution, and trading of these digital assets. The smart contract automates the verification of ownership and the execution of transactions, enabling artists, creators, and buyers to interact in a decentralized and trustless environment. This enhances security, transparency, and efficiency in managing digital representations of the related artistic works, high value artifacts and objects of cultural importance.

[0043] In one or more embodiments, the hash is generated based on at least a part of the dataset and one or more of an owner name of the artwork, an author or artist name of the artwork, or a state of condition of the artistic work.

[0044] According to an aspect of the present invention, there is provided a system for identifying, validating or authenticating an artwork, comprising: at least one imaging platform to acquire image data of an artwork in one or more regions of the electromagnetic (EM) spectrum at one or more sample regions of the artwork; at least one processor; and a memory coupled to the processor, the memory containing instructions that, when executed by the processor, configure the system to: form a multidimensional array of the image data; create one or more datasets from the multi-dimensional array, wherein the artwork may be identified from the created dataset; and compare the created dataset to one or more previously generated datasets relating to previously acquired image data obtained from a previous encounter of the artwork to evaluate if the artwork is the same as the previous encounter of the artwork.

[0045] According to an aspect of the present invention, there is provided a system for creating and storing a digital representation of an artwork, comprising: at least one imaging platform to acquire image data of an artwork in one or more regions of the electromagnetic (EM) spectrum at one or more sample regions of the artwork; at least one processor; and a memory coupled to the processor, the memory containing instructions that, when executed by the processor, configure the system to: form a multi-dimensional array of the image data; create one or more datasets from the multi-dimensional array, wherein the artwork may be identified from the created dataset; and storing on a distributed ledger one or more of the dataset or a reference to the created dataset.

[0046] Advantageously, in this particular form of the invention the created dataset may be stored "on chain" or "off chain". Storing the created dataset "off chain" in the context of the present invention means keeping the created dataset or file outside the blockchain network while maintaining a reference or link to that dataset on the blockchain. This approach is often used for storing large files or datasets that are impractical to store directly on the blockchain due to size, cost, or efficiency considerations.

[0047] The actual created dataset (e.g., large files, media content) may be stored on an external storage system, such as a centralized server, decentralized storage network (e.g., IPFS - Interplanetary File System), or cloud storage.

[0048] A cryptographic hash or a URL pointing to the off-chain dataset is recorded on the blockchain. This hash or link ensures that the integrity and authenticity of the off-chain dataset can be verified. Users can access the off-chain dataset through the link stored on the blockchain. The hash ensures that any alteration of the dataset would result in a different hash, thus maintaining the dataset's integrity and security.

[0049] Storing the created dataset "on chain" in the context of the present invention means keeping the created dataset or file directly on the blockchain network itself. This involves recording all information, such as transaction details, digital assets, or smart contract data, within the blockchain's immutable and distributed ledger.

[0050] Here, the dataset is included in the transactions that are processed and validated by the blockchain network. This could be any type of information, such as ownership records, metadata for digital assets, or the contents of smart contracts. Once the transaction containing the data is validated, it is bundled into a block. This block, along with its data, is then added to the blockchain, making the information a permanent part of the chain.

[0051] Advantageously, data stored on-chain benefits from the blockchain's properties of immutability and decentralization. Once the data is added, it cannot be altered or deleted, ensuring a secure and tamper-proof record. Further, since the data is part of the blockchain, it can be accessed and verified by any participant in the network. This transparency ensures that anyone can independently verify the integrity and authenticity of the data.Brief Description of Drawings

[0052] The invention will now be described in further detail by reference to the accompanying drawings. It is to be understood that the particularity of the drawings does not supersede the generality of the preceding description of the invention.

[0053] FIG. 1 shows a block diagram of an embodiment of a system for identifying, validating or authenticating an artwork;

[0054] FIG. 2a shows a flow chart of an embodiment of a method of comparing datasets pertaining to an artwork;

[0055] FIG. 2b shows an example of the selection of a primary pixel in of an embodiment of the present invention;

[0056] FIG. 2c shows the 'pixel-level similarity map' for the primary pixel in the embodiment of FIG. 2b;

[0057] FIG. 2d shows the map of pixels over a predetermined threshold with respect to the primary pixel of FIG. 2b;

[0058] FIG. 2e shows the 'image-level similarity map' at a threshold for a cropped and normalised image according to an embodiment;

[0059] FIG. 2f shows an example of a stacked histogram of similarity scores according to an embodiment;

[0060] FIG. 3a shows a flow chart of an embodiment of a method of comparing datasets pertaining to an artwork according to an embodiment;

[0061] FIG. 3b shows an example of a Red Green Blue (RGB) preview of a hyperspectral image capture; the fraction of explained variance from each principal component; the spectra of the first three principal components; and a false colour images found by projecting the full hyperspectral cube into each of the first three principal components;

[0062] FIG. 3c shows an example of comparison scores between hyperspectral images according to an embodiment;

[0063] FIG. 4a shows a flow chart of an embodiment of a method of comparing datasets pertaining to an artwork;

[0064] FIG. 4b shows initial dimension reduction results for two example hypercubes from the same painting according to an embodiment; and

[0065] FIG. 5 shows a flow chart of one embodiment of a method of creating and storing a digital representation of an artwork according to an embodiment.Detailed Description

[0066] The invention is suitable for identifying, validating or authenticating artistic works, and it will be convenient to describe the invention in relation to that exemplary, but non-limiting, application. However, it will be appreciated that the same approach is applicable to physical objects more generally, including, but not limited to, high value artifacts, printed material, drawings and objects of cultural importance.

[0067] FIG. 1 shows a block diagram of one embodiment of a system 100 for identifying and validating or authenticating an artwork 102. The system 100 includes an imaging platform 104 to acquire image data of the artwork 102 in one or more regions of the electromagnetic (EM) spectrum at one or more sample regions of the artwork 102 and / or obtain information pertaining to the artwork 102. The system includes a processor 106, in communication with the imaging platform 104, to processthe acquired image data of the artwork 102 for analysis (and the creation of datasets from that image data for subsequent analysis). The system 100 includes data storage devices 108 in communication with the processor 106 to store data. The system 100 includes an output device or terminal 110 in communication with the processor 106 and the data storage device 108 to deliver and / or present the acquired data or information to a user operating the output device or to another device in communication with the device 110.

[0068] In one or more embodiments, the imaging platform 104 includes one or more hyperspectral sensors that collect and processes information pertaining to the artwork 102 from across the electromagnetic spectrum (for example, an infrared (IR) spectrum, a visible light spectrum, an ultraviolet (UV) spectrum, or an X-ray spectrum, or a combination thereof). The hyperspectral sensors collect information as a set of 'images'. Each image represents a wavelength range of the electromagnetic spectrum, also known as a spectral band. These 'images' are combined to form a three- dimensional (x, y, X) hyperspectral data cube or 'hypercube' for processing and analysis, where x and y represent two spatial dimensions of the scene, and lambda, X, represents the spectral dimension (comprising a range of wavelengths).

[0069] Specifically, the artwork 102 may be captured by many individual images, each taken within a different narrow wavelength band. The result is a 'hypercube' of the combined images. For example, surface pixels may be identified by a two-dimensional coordinate W (width) x H (height), where W is pixels across the artwork 102 surface and H is pixels vertically down the artwork 102 surface, where the origin point is top left. Each surface pixel has a list (or vector) of values attached to it, this list is a spectrum lambda, X, and the spectrum has 204 values, one each a narrow wavelength filter centred on different wavelengths. It will be appreciated that the hypercube can be any size- depending on the image size and / or the range of wavelengths captured.

[0070] A variety of different imaging devices with multispectral / hyperspectral imaging capabilities are commercially available. Examples of commercially available imaging devices with these capabilities are the hyperspectral cameras commercially available under the SpecimIQ. series trade designation from Specim Spectral Imaging Oy Lt., Finland, for example.

[0071] In a number of embodiments, the resultant 'hypercube' is stored in the data storage device 108 where it may be accessed by the processor 106 and compared with other hypercubes (or datasets derived from hypercubes) to determine the validity or authenticity of the artwork 102. The methodologies applied to make this comparison will be described in greater detail with reference to Figures 2a, 3a, 4a and 5.

[0072] Those skilled in the art will appreciate that there are several reasons as to why someone would like to authenticate an artwork, for example: Authentication can significantly impact the financial value of an artwork. An authenticated piece by a renowned artist can be worth a significant amount of money, whereas a forgery or misattributed work has substantially less value. Collectors, sellers, and buyers all rely on authentication to ensure that they are dealing with genuine works; Authenticating an artwork also establishes its provenance, or history of ownership. This can provide important context and historical significance, often making the piece more valuable or interesting. Museums and collectors value provenance as it enhances the narrative and cultural importance of the artwork.

[0073] One situation where it may be desirable to validate an artwork, rather than authenticate, may occur when a museum plans to exhibit an artwork on loan from a private collector, for example. Validation can protect the collector's interests and provide assurance to potential insurers that the returned artwork is in the same condition as before it went on loan. Furthermore, studying the condition of an artwork over time can aid in its preservation, as it allows conservators to monitor and address any signs of deterioration, such as fading, cracking, or damage from environmental factors like light, humidity, and temperature. By regularly assessing the artwork's condition, conservators can implement preventive measures and necessary restorations to maintain its integrity and aesthetic value. Authentication may be considered a subset of validation.

[0074] The purpose of each methodology explained in the following sections is to show examples of analysing the hypercubes, that can separately identify if two hyperspectral datasets (hypercubes) are of the same artwork e.g., the same painting. The general pattern being is (i) how the hypercubes are analysed, (ii) how comparisons are made, then (iii) how a 'threshold', or other means of identifying same-different artworks is created. It will be appreciated that the cropping and registration described with reference to the embodiment of Figure 2a may equally be applicable to the embodiments described with reference to Figure 3a, 4a and 5. It will further be appreciated that depending on the application or methodology used, the system may or may not require registration of the created datasets before comparing the created dataset to one or more previously generated datasets. In one embodiment, registration is needed since the location of an artwork within the frame of a space will be different when one data capture is taken at a different time to another that it is being compared to (i.e. so that similarity maps line up with multiple reference positions). Alternatively, other applications or methodologies simply don't need the registration and those that do may have the registration as part of the initial steps of the method.

[0075] Referring to Figure 2a, there is shown a flow chart of an embodiment of a method 200 of comparing datasets pertaining to an artwork (for example, a set including painting 102, and one ormore paintings bearing resemblance to painting 102). The method 200 may be carried out on processor 106 with reference to Figure 1, for example. The method 200 may be defined in computer-executable instructions (software), thereby providing flexibility for potentially changing or altering the method 200. The method 200 outlines a 'similarity map' concept which centres around an attempt to preserve the 'pattern' of the pixels with similar spectra, without needing to exactly identify their locations / dimensions in future images. As will be appreciated by those skilled in the art, comparing hyperspectral cubes directly, pixel-by-pixel, is challenging. Small changes in painting orientation, lighting and camera systematics can lead to large and noisy difference measurements. Each methodology has attempted to ameliorate this key limitation.

[0076] The method 200 starts at block 202 where two hypercubes are captured each pertaining to an artwork (for example a painting 102 and another painting bearing resemblance to painting 102). The two hypercubes are then cropped twice. Note, these captures may happen at different times. Capturing the same artwork at different times may be useful for insurance or conservation purposes as outlined above.

[0077] First: 'surface' cropping. All data associated with the 'background' is removed, retaining as much of the artwork as possible. It is preferable that this is done without creating 'edges' that are not within the actual artwork surface. The crop points may be manually identified or automatically identified i.e., intelligently identifying and removing the background or unnecessary or unwanted parts of the hypercube. It will be appreciated that surface cropping may not be needed if a smaller section of a large piece of art is imaged. It also will be appreciated that the 'background' refers to the surface that the artwork is laying on or against (and may take the form of a reference panel that is placed adjacent to the artwork). The reference panel may be white in colour so as to assist data acquisition.

[0078] Crops may be restricted to being aligned with the axes of the image for simplicity. However, other crops are envisaged, including those not restricted to being aligned with the axes of the image, for example. It will be appreciated, where the artwork is large in size, the hypercube may be of a section of the artwork - and in that case, spatial and / or surface cropping may not be required.

[0079] Second: 'wavelength' cropping. The data for wavelength >950nm is excluded to reduce influence by thermal noise. As will be appreciated by those skilled in the art, thermal noise is a type of noise that occurs in imaging sensors due to the random generation of charge carriers caused by thermal vibrations in the sensor material. This noise is present even when the sensor is not exposed to any light. The random nature of it makes it challenging to completely eliminate and is hence mitigated in thisinstance by eliminating data outside a wavelength threshold, here >950nm, but other thresholds are envisaged.

[0080] At block 204, a set of pixel-level similarities maps is created for each cropped hypercube, the dimensions of which, for example may be 425 (W) x 318 (H) x 186 (X) comprising 135150 (= 425 x 318) 'surface pixels'. It will be appreciated that these dimensions are exemplary, but other dimensions can be used. The dimensions are largely governed by the imaging platform 104 chosen i.e., cropped down from a 512x512px image. On that basis, for block 204, the hypercube may be conceptualised as a three-dimensional table of data or array, where each pixel on the surface of the (cropped) artwork image has an integer value list, for example a 186-value list (representing the spectrum) behind it. An interim pixel-level 'similarity map' is created for each 'surface' pixel in the cropped hypercube— for the example above, 135150 datasets are created. This dataset has the same dimensions as the cropped artwork surface, and each cell is populated with a measure of how similar the spectrum of the pixel at that location is to a 'primary pixel' spectrum. It will be appreciated that each pixel can serve as a 'primary pixel' and also in turn a 'secondary pixel' at different points of time. It will be appreciated creating the dataset includes iteratively identifying a primary and in turn a secondary pixel in the created dataset at a spatial position and determining the similarity of the spectral dimension of the primary pixel with the spectral dimension of the secondary pixel. The similarity score chosen is [1 - Spectral Angle Mapping (SAM)]; the higher this score, the more similar the two spectra. As an example, a pixel at [W=309, 1-1=65] may be chosen, as will be discussed in greater detail with reference to Figure 2b.

[0081] In the above example, the dataset is a table or array of 135150 numbers and is easier to understand if it is visualised by colouring the pixels where the darker the colour the higher the measure (i.e., the more similar the pixel spectrum to the primary pixel). Figure 2c is a visualisation of the pixellevel similarity measures for primary pixel [W=309, H=65].

[0082] At block 206, the pixel-level similarity maps are summarised into an image-level similarity map— one per hypercube (i.e. an image-level similarity map may be constructed for both a reference hypercube and a sample hypercube that will be compared to the reference). Handling so many datasets per hypercube is infeasible, and therefore the first step toward the image-level map is for each pixellevel dataset to be distilled to a single number. This number is the proportion of pixels in the pixel-level similarity map with a similarity score assessed against a particular threshold (for example, in this case, above a particular threshold). For example, x number of the pixels have a similarity measure over the particular threshold, which represents a certain percentage of the whole pixel surface. As an example, these are mapped in Figure 2d, where a coloured cell identifies each of the pixels above the threshold(only one colour is used, because this is a binary result, either above the threshold or not) - the pixels being shown in black which are above the threshold. The threshold used in this example is (l-SAM) > 0.98. In one or more embodiments, a user can select one pixel in the hyperspectral image, and a threshold, and the user will then be presented with all pixels with a (l-SAM) above that threshold highlighted in the image. In operation the method may start with a "primary pixel" and compare every other pixel in a cropped surface using a similarity measure (e.g. SAM as noted above). Ordinarily, lower SAM values indicate a higher degree of similarity, so the method of the present invention may invert the calculation (essentially 1 minus SAM), so that higher values mean a higher degree of similarity — since this is a more intuitive expression of the SAM value. In the example above, the method checks which pixels have a similarity score above 0.98. Each pixel is assigned a "yes" if it passes that threshold and these are counted and divided by the total number of pixels in the crop surface to get a proportion — which is a single number that indicates how much of the surface is highly similar to the primary pixel. This number is then assigned to the location of that primary pixel, creating a simplified map of similarity.

[0083] As noted above, when this step is undertaken for each pixel in the surface of a cropped hypercube, the result is one value— the proportion of pixels with a spectrum similarity measure above the threshold— for each pixel.

[0084] In one or more embodiments, for each cropped hypercube, a single dataset (with the same dimensions as the cropped artwork surface) is created. Each cell is populated with the pixel-level proportional value, calculated using the above method, of the pixel at that location. This is the imagelevel similarity map. This data can also be visualised, which will be described with reference to Figure 2e.

[0085] The method continues to block 208, where same-different boundaries are developed by: creating a set of paintings that were as similar to each other as possible; taking hyperspectral images of each painting many times under slightly varying conditions; creating a similarity map for each hypercube; for each pair of hypercubes in an artwork set, registering and comparing the similarity maps using image-similarity measures; combining all available comparisons scores, along with the comparison 'ground truth' (i.e., it was known if the similarity maps being compared were of the same or different paintings), to establish a same-different threshold.

[0086] As will be appreciated here, registration is essentially resizing, rotating, skewing, and / or translating the spatial dimensions of one image-level similarity map to match another and may take any suitable form.

[0087] The intent is to enable comparison of two similarity maps, to use calculated imagesimilarity scores, to identify whether the two hyperspectral images (which created the similarity maps) are of the same artwork. This comparison is done for all artworks in the same artwork set e.g., two or more paintings bearing resemblance to painting 102; the image-level similarity map for each comparison pair is registered, and then one or more image-similarity measures are calculated. The comparison may be made by known methods for image quality assessment using a variety of methods and those skilled in the art will recognise suitable metrics for implementing the comparison, including: Mean Squared Error (MSE), Root Mean Squared Error (RMSE), Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), Universal Quality Image Index (UQI), Multi-scale Structural Similarity Index (MS-SSIM), Erreur Relative Globale Adimensionnelle de Synthese (ERGAS), Spatial Correlation Coefficient (SCC), Relative Average Spectral Error (RASE), Spectral Angle Mapper (SAM), Spectral Distortion Index (DJambda), Spatial Distortion Index (D_S), Quality with No Reference (QNR), Visual Information Fidelity (VIF), Block Sensitive - Peak Signal-to-Noise Ratio (PSNR-B).

[0088] In some embodiments, it is possible that when two image-level similarity maps are compared, for example Image 1 and Image 2, the order in which they are fed into the registration function will impact the outcome. Therefore, a comparison between two similarity maps may require a combination of two sets of comparisons, Image 1 to Image 2 and Image 2 to Image 1. Further, the registration process has a small degree of inherent variability, due to a random element. To mitigate against this, each registration can be run a number of times and the median for each measure taken, for example 10 or 20 times. Thus, to compare two similarity maps, the registration and calculation is done a number of times for Image 1 to Image 2, and then done that same number of times for Image 2 to Image 1, and the median of the scores for each set is then averaged and this is the comparison score.

[0089] In an embodiment, each set of artwork comparison scores may then be assessed to determine if a pattern can be observed, and if there is clear delineation between the comparison scores for 'same' painting pairs and 'different' painting comparison pairs. In this stage of the creation of the same-different threshold, paintings are known to be the same or different.

[0090] In one or more embodiments, when analysing a set of image-level similarity map comparisons, a stacked histogram is created for each score type, to visualise the data and identify the existence / scale of the 'same-difference boundary'. Using the graphed SSIM measure as an example in Figure 2f, a successful boundary is possible if there is a gap between the highest score for different pair comparisons and the lowest score for same painting comparisons. In this example, there is a gap between the maximum different score (0.3932) and the minimum same score (0.64815), and therefore a same-different boundary can be reasonably confidently derived from the SSIM score for this examplepainting set. As will be appreciated, in the above example, a set of artworks were created for the purpose of the above example - where there was a notional 'original' artwork and a notional set of 'copies'. Knowing the set of artworks allowed for controlled comparison image-level similarity maps of the same artwork, or if they were derived from different artworks for testing. In this manner, SSIM scores were assessed for 'same' and 'different' comparisons, and a clear margin was found. As will be appreciated, in practice, it is unknown if two encounters of artworks that resemble each other are 'same' or 'different', and thus the establishment of assessment boundaries will differ.

[0091] Referring to Figure 2b, there is shown a cropped 425 (w) x 318 (h) x 186 (X) artwork 102; therefore, there are 135150 (=425x318) surface pixels selected primary pixel 210 for this example is outlined w=309, h=65. An example of the selection of the primary pixel occurs at block 204. In this methodology all pixels are treated as the primary pixel as some point in the process (i.e. there are multiple primary pixels and each pixel in turn will become a primary pixel).

[0092] Referring to Figure 2c, there is shown an example of a pixel-level similarity map 212 for pixel (w=309, h=65) of a cropped and normalised image, for example artwork 102. The pixel-level similarity map 212 is a convenient visual representation of the dataset created at block 204. The darker the colour the more similar each pixel spectrum is to the primary pixel.

[0093] Referring to Figure 2d, there is shown an example of a map of pixels over a threshold (here (l-SAM) > 0.98) or similarity score to primary pixel (w=309, h=65) in cropped and normalised image 214, for example artwork 102. This map is one of the possible outputs that may occur at step 206. Those skilled in the art will recognise that such a map may be generated in software, including proprietary software (for example, Specim IQ Studio software designed for spectral imaging and analysis, specifically tailored for use with the Specim IQ hyperspectral camera as referenced above) here a user can select one pixel in the hyperspectral image. It will also be appreciated that the threshold is preferably set to a suitable level - between too low (capturing too many pixels) and too high (not capturing enough) for example, a degree of similarity in the range of between 0.95 and 0.98 has been found to provide a suitable middle ground for this measure, though a different threshold range may be different for other artworks and may be assessed by visualisation of threshold options.

[0094] Referring to Figure 2e, there is shown an 'image-level similarity map' at particular threshold (here (l-SAM) > 0.98) for cropped and normalised image 216, for example artwork 102. In one or more embodiments, for each cropped hypercube, a single dataset (with the same dimensions as the cropped artwork surface) is created. Each cell is populated with the pixel-level proportional value(i.e. the proportion of pixels that are above the threshold), calculated using the method described above, of the pixel at that location.

[0095] Referring to Figure 2f, there is shown an example of a stacked histogram for SSIM similarity scores, of a selected set of artworks, for example artworks related to (or bearing resemblance to) artwork 102. The stacked histogram shows 'Exact' = self-same image comparison 218, should give the outside boundary of the score (blue; subset of 'same'), where the remaining stacked histograms represent the following scenarios SSsS = Same artwork, same standard image collection protocol, same imaging session (orange; subset of 'same') 220; SSsD = Same artwork, same standard image collection protocol, different imaging session (green; subset of 'same') 222; scenarios SSnS = Same artwork, same non-standard image collection protocol, same imaging session (red; subset of 'same') 224; SSnD = Same artwork, same non-standard image collection protocol, different imaging session (purple; subset of 'same') 226; SDS = Same artwork, different image comparison protocol, same set (brown; subset of 'same') 228; SDD = Same artwork, different image comparison protocol, different set (pink; subset of 'same') 230; DIFF = different artwork 232 (grey). As is clear from the gap between the 'different' score 232 and the minimum 'same' score (lowest of all 218-230), the same-different boundaries can be reasonably confidently derived from the SSIM score, for example. As will be appreciated, a number of stacked histograms may be provided, as part of a library for example, each representing the image similarity score distributions between image-level similarity maps of a known artwork - such that a new histogram (which may relate to a new encounter of the same or similar artwork) can be compared to those in the library.

[0096] Referring to Figure 3a, there is shown a flow chart of an embodiment of a method 300 of comparing datasets pertaining to an artwork 102 (for example a set of paintings). The method 300 may be carried out on processor 106 with reference to Figure 1, for example. The method 300 relies on Principal Component Analysis, or PCA, which is a dimensionality reduction method that is often used to reduce the dimensionality of large data sets, by transforming a large set of variables into a smaller one that still contains most of the information in the large set. The technique involves identifying linear transformations of the data that capture the maximum fraction of the original variance. If each pixel in the hypercube is imagined as being somewhere in a multidimensional space e.g., a 204-dimensional space: by finding a small number of linear combinations of vectors (approximately 3) that capture the majority of variance and expressing the position of each pixel as a sum of these vectors, the dimensionality of the comparison space is vastly reduced. This has several advantages: (i) the main features of the data are captured, while the less important, noisy features are discarded; (ii) the dimensionality of the data is reduced, making it easier to visualise and compare; (iii) the computational complexity of the comparison is reduced. Accordingly, the method 300 can be applied efficiently andmay operate quickly (or in substantially real-time) on relatively low-end hardware that is, for example, constrained by processor speed and memory.

[0097] The PCA is applied in a process that is constructed as a comparison of two hypercubes. In one or more embodiments, the hypercubes being compared are given specifical assignations: the 'original' is 'known', and the 'capture' is the artwork being compared to the original to test if they are the same artwork, for example artwork 102.

[0098] The method 300 starts at block 302 where two hypercubes are captured each pertaining to an artwork. The two hypercubes are then cropped twice as discussed with reference to Figure 2a.

[0099] At block 304, the two cropped hypercubes are registered. The exact cropped-and- registered hypercube for any original hyperspectral dataset will vary depending on the hypercube it is being compared to.

[0100] As described above, here registration is essentially resizing, rotating, skewing, and / or translating the spatial dimensions of one hypercube to match another. Because registration functions only accept two-dimensional input, it is necessary to reduce the three-dimensional hypercubes to two dimensions. This is done by applying a three-component PCA to the cropped 'original' hypercube (output from block 304) (analysis proved that three was sufficient to cover the majority of the variance for hypercubes in tests), then creating three two-dimensional images (for example, set 310 in Figure 3b) for the 'original' hypercube. The same projection is then used to create three two-dimensional images for the 'capture' hypercube. Each of the three transformed images are then separately registered, and the one that results in the best agreement is applied to all spectral bands of the cropped 'capture' hypercube. The result is a registered 'capture' cube that is the same size as the 'original' cropped hypercube and can be directly compared. However, after registration there are often empty pixels at the borders of the transformed result; therefore, further cropping of both hypercubes is undertaken to remove this element (for it is not a 'real' feature of the artwork).

[0101] At block 306, the principal components of each cropped and registered hypercube are determined. Each component is expressed as a linear combination of weights along the wavelength dimension (such as plot C 310 in Figure 3b) and alternately visualised (such as plot D 312 in Figure 3b).

[0102] In one or more embodiments, a PCA is carried out on the cropped 'original' hypercube. This will differ to the initial PCA carried out during the registration step by (i) the extent the additional cropping may impact the available information, and (ii) the number of components is not limited to three, but allowed to extend to the count that recovers at least 99.5% of the total variance in theoriginal hypercube. The projection is then applied to both cropped and registered hypercubes, creating between three and 16 images per hypercube (as will be appreciated by those skilled in the art, this will largely depend on the propertied of the original artwork), which are compared using image-similarity measures.

[0103] At block 308, the same-different boundaries are developed: for each pair of hypercubes, each PCA feature map is compared using image-similarity measure MS-SSIM, and the average used as the overall measure; combine all available comparisons scores, along with known 'ground truth' (as described above with reference to Figure 2a), to find a same-different threshold.

[0104] In a number of embodiments, each set of projection images are compared, and the average of these scores assigned as the final comparison score of the two hypercubes. The score used is Multiscale Structural Similarity Index (MS-SSIM; also successfully applied in with reference to the 'similarity map' method outlined in Figure 2a).

[0105] The threshold for same / different artwork assessment is approached in the same way as in the similarity map embodiment - hypercubes from the same artwork set are compared to each other, and a distribution of scores for each known 'ground truth' is investigated, looking for a gap between 'different' and 'same'.

[0106] Referring to Figure 3b, there is shown an RGB preview of a hyperspectral image capture 320; the fraction of explained variance from each principal component 312; the spectra of the first three principal components (numbered here 0, 1, 2) 310; and false colour images found by projecting the full hyperspectral cube into each of the first three principal components 322 RED (R), 324 GREEN (G), 326 BLUE (B).

[0107] Referring to Figure 3c, the threshold for same / different artwork assessment is approached in the same way as similarity map - hypercubes from the same painting set are compared to each other, and a distribution of scores for each known 'ground truth' is investigated, looking for a gap between 'different' and 'same'. In Figure 3c, there is shown the distribution of comparison scores for every permutation of comparison between all 24 hyperspectral images of an artwork series (i.e., several artworks each bearing resemblance to each other). The histogram 314 shows the distribution of scores for comparisons where the original and capture are different artworks. The histogram 316 shows the distribution of scores when different hyperspectral images of the same artwork are being compared. The black dashed line 318 indicates the resulting threshold comparison score that can be used to delineate an authentic hyperspectral image of this composition from a reproduction, copy or the like.

[0108] Referring to Figure 4a, there is shown a flow chart of an embodiment of a method 400 of comparing datasets pertaining to an artwork 102 (for example a set of paintings). The method 400 may be carried out on processor 106 with reference to Figure 1, for example. The method 400 outlines an approach which centres around the use of a pre-trained machine learning model for feature extraction and using those features to draw a comparison between artworks.

[0109] The method 400 starts at block 402 where two hypercubes are captured each pertaining to an artwork (either the same artwork or a different artwork). The two hypercubes are then cropped twice as discussed with reference to Figure 2a. However, in this embodiment the images are only cropped to exclude the background, no cropping >950nm is undertaken in this approach.

[0110] At block 404, the dimensions of each hypercube are reduced. Like the PCA methodology described with reference to Figure 3a (and leaning on its verification that the first three components are often sufficient to cover a significant portion of the overall variance), each hypercube is first transformed into a three-dimensional dataset with the same surface dimensions as the original hypercube but now with three values (transformed via PCA) instead of the original amount of spectral bands, which in some embodiments is in the order of 204 (note that the cropping >950nm is not undertaken in this approach); for example in Figure 4b. These reduced-size dataset are next investigated using pre-trained models that are built for images with RGB (red, green, blue) 410.

[0111] At block 406, each cropped-reduced data cube is pushed through a pre-trained machine learning model, to extract a set of features. The purpose of such a model is to identify exactly what makes each image unique, to extract its features. Those skilled in the art will recognise ResNet50 as one such model, being a specific type of Convolutional Neural Network (CNN). ResNet50 has been trained on a dataset of images collected from public domain and held in a dataset called 'ImageNet', which contains over 14 million images, which have been allocated to over 1000 classes. The model has 50 'convolution layers', and 175 intermediate feature layers. The usual way of using ResNet50 is to push through an image and for the output to be one of the classes - i.e., the model assesses what 'kind' of image it has been presented with. However, here this is not of interest, but instead the aim is to extract the feature layers and use these to create a unique means of characterising the cropped-reduced data. When a 3-band image (i.e., an RGB image) is fed through the ResNet50 model, the output selected is an 'embedding', a one-dimensional vector, per layer. So, for each cropped reduced hypercube, the output is a vector per layer - which can be organised as a matrix, with each layer being a row.

[0112] At block 408, a Multi-Layer Perceptron (MLP) is constructed, which is trained on all sameartwork comparisons and a (same-sized) subset of different-artwork comparisons, to predict if the setsof features for the two hypercubes being compared are of the same artwork. As those skilled in the art will appreciate, a MLP is a kind of abbreviated or lightweight fully-connected neural network (NN) - it has fewer layers, and therefore can be trained on a smaller set of images / data - and is a deep learning model / algorithm that learns the characteristics of the data (i.e., assumptions or parameters do not need to be input at the start by the model maker).

[0113] When comparing the 'feature matrix' from two hypercubes, the vectors for each layer are compared separately using Euclidean distance (i.e., the smaller the measure, the more similar the vector). A comparison therefore has the output of a large list of measures, 175 for example. These numbers are not simple for a human to discriminate or assess, but numeric comparisons can be applied, and one means of achieving this is by building another model.

[0114] The dataset for this step is built from this list of comparison measures for a subset of all possible comparisons from image data sets containing any number of artworks; each two-hypercube comparison is a row in this dataset. As per 'ground truth' discussion above, there are a number of two- hypercube comparisons where the artwork is known to be the same. For example, 80 comparisons may be randomly selected from the SSD set (same artwork, same protocol, different set). From the comparisons of different artworks, may be randomly selected (ensuring these were comparisons within the same painting set), this is done so that the dataset has equal counts of 'same' and 'different' comparisons (required by the binomial character of, and the need to avoid bias in, the next model).

[0115] The data set may then be split into training and testing / validation subsets - this is usual model creation protocol [80% / 20% train and test, and the 80% train split further into 80% / 20% train and validate] - for building the MLP. The predictions from the MLP can then be visualised using a tool, for example, 'Deep SHAP' (SHapley Additive exPlanations), which is specifically designed to interpret machine learning models' output. However, the use of other explainability techniques that can be used for models with a neural network-based architecture are also envisaged. Those skilled in the art will appreciate that model explainability refers to the process whereby outputs produced by machine learning models are explained in terms of how and which features influence the model's actual output.

[0116] In summary, the work flow may follow the following steps: crop two hypercubes 402; reduce each with PCA to three-bands 404; push each cropped-reduced dataset through a model e.g., ResNet50406, the result being a two matrices; compare the two matrices, line by line, the result being a list of similarity measures; this list can then be pushed through the bespoke MLP, which will output (i) the probability the two hypercubes are of the same artwork and (ii) the probability the twohypercubes are of different artworks 408; the most likely prediction is then visualised (enabled by SHAP) detailing the contribution of each feature (layer) to the prediction.

[0117] Figure 4b shows an example of PCA results 410, for two hypercubes from the same artwork.

[0118] In a number of embodiments, running the hypercubes through ResNet50 showed that some layers in the model were particularly powerful in characterising the artworks, for example features from layer 120 may be more significant than features from layer 140. The nature of the ResNet50 model is such that it is not intuitive to identify or understand what each of these models 'actually' describes about each hyperspectral image of a painting, but each is finding something in the data that is unique and can be extracted from the data to characterise the images.

[0119] In a number of embodiments, each two-hypercube comparison can be fed into SHAP (or another tool) to provide more information about the drivers of the prediction. The influence of each ResNet50 layer visualised as it contributes to the overall score of confidence of the prediction. For example, the prediction is that a model is confident up to a level that the hypercubes are of the same painting.

[0120] This can be visualised to identify which are the most influential layers in this determination; with a representation of increasing the confidence in 'same' and another representation of eroding the confidence (i.e., driving towards 'different'). Those skilled in the art will recognise that this may be visualised on a sliding scale or similar linear representation.

[0121] Referring to Figure 5, there is shown a flow chart of a method 500 for creating and storing a digital representation of an artwork 102 according to an embodiment of the present invention. The method 500 begins at block 502 where an imaging platform 104 is used to acquire image data of the artwork 102 in one or more regions of the electromagnetic (EM) spectrum at one or more sample regions of the artwork 102 and / or obtain information pertaining to the artwork 102.

[0122] At block 504, a multi-dimensional array of the image data is formed. As discussed with reference to Figure 1, the imaging platform 104 typically includes one or more hyperspectral sensors that collect and processes information pertaining to the artwork 102 from across the electromagnetic spectrum (for example, an infrared (IR) spectrum, a visible light spectrum, an ultraviolet (UV) spectrum, or an X-ray spectrum, or a combination thereof). The hyperspectral sensors collect information as a set of 'images'. Each image represents a wavelength range of the electromagnetic spectrum, also known as a spectral band. These 'images' are combined to form a multi-dimensional (x, y, X) hyperspectral datacube or 'hypercube' for processing and analysis, where x and y represent two spatial dimensions of the scene, and X represents the spectral dimension (comprising a range of wavelengths).

[0123] At block 506, one or more datasets from the multi-dimensional array are created. The artwork 103 may be identified from the created dataset. The dataset may be stored in a proprietary database or the like.

[0124] At block 508 in some implementations, some or all of the data in the dataset (or the raw data from the hypercube) may be also stored in encrypted form on a distributed ledger, in which the limited and encrypted version of the dataset can be disseminated across a network of computers (e.g., the Internet). For example, peer-to-peer data sharing requires the use of web services to communicate, and thus the proprietary data of the dataset is never fully transmitted to the client node, but instructions and / or results associated with the digital identification of an artwork may be permitted to transfer in such manner. As will be appreciated by those skilled in the art, storing some or all of the dataset directly on the distributed ledger (e.g., a blockchain) is typical of on-chain storage. This involves embedding the dataset into transactions that are then permanently recorded on the blockchain's distributed ledger. On-chain storage ensures decentralization, immutability, and transparency, as data once stored on the blockchain cannot be altered or deleted, and it is accessible to all participants in the network.

[0125] In one or more embodiments, a reference to the dataset may be stored on the distributed ledger. As will be appreciated by those skilled in the art, this is typical of off-chain storage in that it involves storing data outside of the blockchain, typically on separate databases or storage systems. While off-chain data is not recorded on the blockchain's ledger, it can be referenced or linked to from within the blockchain. Off-chain storage offers scalability, reduced transaction costs, and privacy advantages, making it suitable for storing large files as may be the case with artwork, or data that does not need to be publicly accessible on the blockchain.

[0126] In one or more embodiments, a hash based on at least a part of the dataset may be stored on the distributed ledger. Those skilled in the art will appreciate that a hash is a numerical representation of an image (for example artwork 102) that is generated by applying a specific algorithm to the image data pertaining to an artwork, for example. This hash is typically a fixed-length string of numbers or characters that uniquely represents the content of the image. In the context of the present invention, the use of image hashing may help to quickly identify, duplicate or near-duplicate the image in a large dataset without needing to compare the actual pixel values of each image directly. In one ormore embodiments, the hash may be used as a reference to the entire dataset in the context of off- chain storage as described above.

[0127] In one or more embodiments, storing on the distributed ledger includes minting a non- fungible token (NFT) related to the artwork 102. Minting a (NFT) is the process of creating a unique digital asset on a blockchain network. Unlike fungible tokens, NFTs and can be exchanged on a one-to- one basis, NFTs are unique and indivisible, each representing a distinct artwork, item or the like. Advantageously, in this particular form of the invention, the NFT may be transferred to another party together with the physical artwork, upon change of ownership or the like. Furthermore, each NFT maintains its uniqueness and ownership history, making it distinguishable from other tokens thereby ensuring its authenticity and scarcity.

[0128] In a number of embodiments, the NFT may also be attributed to a smart contract or rules engine on a blockchain platform. This smart contract or rules engine defines the characteristics of the NFT, including its metadata (such as title, description, and creator information) and its ownership rules.

[0129] The smart contract may mint a unique token representing the digital asset. This token is the NFT itself and is stored on the blockchain. It has a unique identifier that distinguishes it from other tokens on the same blockchain. Once the NFT is minted, it can be transferred to another party through a blockchain transaction. Each transfer of the NFT is recorded on the blockchain, providing a transparent and immutable record of ownership.

[0130] As will be appreciated by those skilled in the art, minting an NFT involves creating a unique digital asset on a blockchain. Any suitable block chain that supports NFTs, like Ethereum may be used. Here, a digital representation of an artwork may be minted and uploaded to a platform, to register this digital representation as a digital asset. During minting, the digital representation of the artwork may be converted into a unique token on the blockchain, which includes metadata like the name, description, and further properties of the artwork including condition or the date on which the digital representation was capture, which may be useful for conservation purposes. This process involves smart contracts that ensure the token's uniqueness and ownership. Once minted, the NFT may be stored in the creator's digital wallet and can be sold, traded, or held, with the blockchain providing a transparent and immutable record of ownership and provenance.

[0131] Advantageously, the authenticity and ownership of the NFT can be verified by checking its transaction history on the blockchain. This ensures that the NFT is genuine and has not been tampered with.

[0132] Advantageously, users may access the distributed ledger to verify an artwork's provenance or the like. For example, users may access one or more computing devices that perform and / or obtain research directly about or relevant to the artistic work, such as the artistic work's provenance, history (including ownership, handling, modifications, restorations, etc.), legal documentation, or other.

[0133] In one or more embodiments, a "zero-knowledge proof" (ZKP) a cryptographic method by which one party (the prover) can prove to another party (the verifier) that they know a value or have performed a computation correctly, without revealing any information about the value itself or the computation may be employed. Advantageously, zero-knowledge proofs enhance privacy and security in the blockchain application of the present invention. ZKPs allow transactions to be verified without revealing the underlying dataset. This is especially useful in blockchain for maintaining confidentiality of sensitive information while ensuring its validity.

[0134] In a number of embodiments, the prover can convince the verifier that they know a secret (e.g., a password or a cryptographic key) without disclosing the secret itself. The verifier can be confident that the prover's claim is true without gaining any additional knowledge about the data or the process.

[0135] In the context of the present invention, ZKPs can be used to keep transaction details relating to an artwork private, while still proving that transactions are valid and conform to the rules of the blockchain. Users can prove their identity or credentials without revealing the actual information, enhancing privacy and security in identity management.

[0136] The term 'distributed ledger,' as used herein, refers to a decentralised electronic ledger of data records which are authenticated by a federated consensus protocol. A distributed ledger may be a blockchain. Multiple computer systems within the distributed ledger, referred to herein as 'nodes', 'light nodes' or 'compute nodes,' each comprise a copy of the entire ledger of records (or in the case of a 'light node' instead of storing the entirety of the blockchain in its memory, it only stores parts of the blockchain that are relevant to the transaction being performed). Nodes may write a data 'block' to the distributed ledger, the block comprising data regarding an electronic event, said blocks further comprising data and / or metadata. In some embodiments, only miner nodes may write electronic events to the distributed ledger. In other embodiments, all nodes have the ability to write to the distributed ledger. In some embodiments, the block may further comprise a time stamp and a pointer to the previous block in the chain (e.g., a 'hash'). In some embodiments, the block may further comprise metadata indicating the node that was the originator of the electronic event. In this way, the entirerecord of electronic events is not dependent on a single database which may serve as a single point of failure; the distributed ledger will persist so long as the nodes on the distributed ledger persist.

[0137] Furthermore, as used herein the term 'user device' may refer to any device that employs a processor and memory and can perform computing functions, such as a personal computer or a mobile device, wherein a mobile device is any mobile communication device, such as a cellular telecommunications device (i.e., a cell phone or mobile phone), personal digital assistant (PDA), a mobile Internet accessing device, or other mobile device.

[0138] It will be appreciated that some embodiments may be comprised of one or more generic or specialised controllers or processors (or 'processing devices') such as microcontrollers, microprocessors, digital signal processors, customised processors and field programmable gate arrays (FPGAs) and unique stored program instructions (including both software and firmware) that control the one or more processors to implement, in conjunction with certain non-processor circuits, some, most, or all of the functions of the method and / or apparatus described herein. Alternatively, some or all functions could be implemented by a state machine that has no stored program instructions, or in one or more application specific integrated circuits (ASICs), in which each function or some combinations of certain of the functions are implemented as custom logic. Of course, a combination of the two approaches could be used.

[0139] As discussed above, the various embodiments can be implemented in a wide variety of operating environments, which in some cases can include one or more user computers, computing devices, or processing devices which can be used to operate any of a number of applications. User or client devices can include any of a number of general-purpose personal computers, such as desktop or laptop computers running a standard operating system, as well as cellular, wireless, and handheld devices running mobile software and capable of supporting a number of networking and messaging protocols. Such a system also can include a number of workstations running any of a variety of commercially-available operating systems and other known applications for purposes such as development and database management. These devices also can include other electronic devices, such as dummy terminals, thin-clients, gaming systems, and other devices capable of communicating via a network.

[0140] Various aspects also can be implemented as part of at least one service or Web service, such as can be part of a service-oriented architecture. Services such as Web services can communicate using any appropriate type of messaging, such as by using messages in extensible markup language (XML) format and exchanged using an appropriate protocol such as SOAP (derived from the "SimpleObject Access Protocol"). Processes provided or executed by such services can be written in any appropriate language, such as the Web Services Description Language (WSDL). Using a language such as WSDL allows for functionality such as the automated generation of client-side code in various SOAP frameworks.

[0141] Some embodiments may utilise at least one network that would be familiar to those skilled in the art for supporting communications using any of a variety of commercially-available protocols, such as TCP / IP, OSI, FTP, UPnP, NFS, and CIFS. The network can be, for example, a local area network, a wide-area network, a virtual private network, the Internet, an intranet, an extranet, a public switched telephone network, an infrared network, a wireless network, and any suitable combination thereof.

[0142] The environment can include a variety of data stores and other memory and storage media as discussed above. These can reside in a variety of locations, such as on a storage medium local to (and / or resident in) one or more of the computers or remote from any or all of the computers across the network. In a particular set of embodiments, the information can reside in a storage-area network ("SAN") familiar to those skilled in the art. Similarly, any necessary files for performing the functions attributed to the computers, servers, or other network devices can be stored locally and / or remotely, as appropriate. Where a system includes computerised devices, each such device can include hardware elements that can be electrically coupled via a bus, the elements including, for example, at least one central processing unit (CPU), at least one input device (e.g., a mouse, keyboard, controller, touch screen, or keypad), and at least one output device (e.g., a display device, printer, or speaker). Such a system can also include one or more storage devices, such as disk drives, optical storage devices, and solid-state storage devices such as random access memory ("RAM") or read-only memory ("ROM"), as well as removable media devices, memory cards, flash cards, etc.

[0143] Such devices also can include a computer-readable storage media reader, a communications device (e.g., a modem, a network card (wireless or wired), an infrared communication device, etc.), and working memory as described above. The computer-readable storage media reader can be connected with, or configured to receive, a computer-readable storage medium, representing remote, local, fixed, and / or removable storage devices as well as storage media for temporarily and / or more permanently containing, storing, transmitting, and retrieving computer-readable information. The system and various devices also typically will include a number of software applications, modules, services, or other elements located within at least one working memory device, including an operating system and application programs, such as a client application or Web browser. It should be appreciated that alternate embodiments can have numerous variations from that described above. For example, customised hardware might also be used and / or particular elements might be implemented inhardware, software (including portable software, such as applets), or both. Further, connection to other computing devices such as network input / output devices can be employed.

[0144] Storage media and computer readable media for containing code, or portions of code, can include any appropriate media known or used in the art, including storage media and communication media, such as but not limited to volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage and / or transmission of information such as computer readable instructions, data structures, program modules, or other data, including RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a system device.

[0145] Where the terms "comprise", "comprises", "comprised" or "comprising" are used in this specification (including the claims) they are to be interpreted as specifying the presence of the stated features, integers, steps or components, but not precluding the presence of one or more other features, integers, steps or components, or group thereof.

[0146] While the invention has been described in conjunction with a limited number of embodiments, it will be appreciated by those skilled in the art that many alternative modifications and variations in light of the foregoing description are possible. Accordingly, the present invention is intended to embrace all such alternative, modifications and variations as may fall within the spirit and scope of the invention as disclosed.

Claims

The claims defining the invention are as follows:

1. A method for identifying, validating or authenticating an artwork using one or more digital representations, comprising: acquiring image data of an artwork in one or more regions of the electromagnetic (EM) spectrum at one or more sample regions of the artwork; forming a multi-dimensional array of the image data; creating a dataset from the multi-dimensional array, wherein the artwork may be identified from the dataset; and comparing the created dataset to one or more previously generated datasets relating to previously acquired image data obtained from a previous encounter of the artwork to evaluate if the artwork is the same as the previous encounter of the artwork.

2. The method of claim 1, wherein the one or more regions of the EM spectrum includes an infrared (IR) spectrum, a visible light spectrum, an ultraviolet (UV) spectrum, or an X-ray spectrum, or a combination thereof.

3. The method of claim 1 or 2, wherein the multi-dimensional array of the image data comprises a hyperspectral data cube.

4. The method of claim 3, wherein the multi-dimensional array of the image data comprises data from two spatial dimensions and one spectral dimension.

5. The method of claim 4, wherein the spectral dimension comprises a range of wavelengths in the image data.

6. The method of any one of claims 1 to 5, further comprising cropping the multi-dimensional array of the image data to one or more of the size of the artwork, or a region within the artwork.

7. The method of claim 6, wherein cropping includes wavelength cropping to remove wavelengths from the multi-dimensional array of the image data at wavelengths above or below a predetermined threshold.

8. The method of any one of claims 1 to 7 , further comprising registering the multi-dimensional array of the image data and / or created dataset with at least one of a location and an orientation of the artwork.

9. The method of any one of claims 1 to 8, further comprising reducing a dimension of the multidimensional array.

10. The method of any one of claims 1 to 9, wherein reducing the dimension of the multi-dimensional array of the image data is performed via Principal Component Analysis (PCA) or a manifold learning method.

11. The method of any one of claims 1 to 10, wherein the step of creating a dataset from the multidimensional array comprises creating a new dataset from the created dataset from the multidimensional array of the image data, the method comprising the steps of: a. identifying a primary pixel in the created dataset at a spatial position; b. identifying a secondary pixel in the created dataset at a different spatial position; c. determining the similarity of the spectral dimension of the primary pixel with the spectral dimension of the secondary pixel, using a spectral similarity function; and d. repeating steps b and c for each secondary pixel in the created dataset, to create an intermediate dataset containing spectral similarity values for every secondary pixel identified in step b when compared to the primary pixel identified in step a.

12. The method of claim 11, further comprising the steps of: e. from the intermediate dataset, determining one value which will be attached to the primary pixel equal to the proportion of secondary pixels with a spectral similarity measure assessed against a predetermined threshold; and f. repeating steps a to e for each primary pixel in the spatial dimension; each spatial position of the new dataset is populated with the value calculated in step e when that spatial position was the primary pixel.

13. The method of claim 11, wherein the intermediate dataset is a derived subset of the hyperspectral data.

14. The method of any one of claims 11 to 13, wherein the new dataset is a derived subset of the intermediate dataset.

15. The method of any one of claims 1 to 14, wherein the wavelengths are manipulated and / or cropped.

16. The method of any one of claims 1 to 15, further comprising using one or more feature extraction techniques to localise features in the multi-dimensional array of the image data.

17. The method of any one of claims 1 to 16, further comprising extracting spectral features by applying a pre-trained machine learning model to the multi-dimensional array of the image data; and determining characteristics of the artwork based on a comparison of the spectral features from the multi-dimensional array of the image data.

18. The method of claim 17, wherein the machine learning model is trained using a training set of hyperspectral images and is configured to classify a plurality of spectral features in a plurality of hyperspectral images.

19. The method of claim 17, wherein the machine learning model includes using a supervised classifier method.

20. The method of claim 19, wherein the supervised classifier method is a multi-layer perceptron method.

21. The method of any one of claims 1 to 20, wherein the artwork comprises a painting, a sculpture, a drawing, a photograph, a cultural object, an object or a printed material.

22. A method for creating and storing a digital representation of an artwork, comprising: acquiring image data of an artwork in one or more regions of the electromagnetic (EM) spectrum at one or more sample regions of the artwork; forming a multi-dimensional array of the image data; creating one or more datasets from the multi-dimensional array, wherein the artwork may be identified from the created dataset; and storing on a distributed ledger one or more of the created dataset or a reference to the created dataset.

23. The method of claim 22, further comprising: generating a hash based on at least a part of the dataset; and storing on a distributed ledger the generated hash.

24. The method of claim 22 or 23, wherein storing on the distributed ledger comprises minting a non- fungible token (NFT).

25. The method of claim 24, further comprising maintaining, by a set of node devices, the distributed ledger that stores the NFT.

26. The method of any one of claims 22 to 25, wherein the distributed ledger is a blockchain.

27. The method of claim 26, wherein the set of node devices host the blockchain.

28. The method of any one of claims 22 to 27, further comprising receiving information relating to the artwork.

29. The method of claim 28, wherein the hash is generated based on at least a part of the dataset and the information relating to the artwork.

30. The method of claim 29, further comprising calculating an original validation value associated with the information relating to the artwork via at least one rule from a rules engine.

31. The method of claim 30, wherein the at least one rule from the rules engine is a smart contract.

32. The method of any one of claims 23 to 31, wherein the hash is generated based on at least a part of the dataset and one or more of an owner name of the artwork, an author or artist name of the artwork, or a state of condition of the artistic work.

33. A system for identifying, validating or authenticating an artwork, comprising: at least one imaging platform to acquire image data of an artwork in one or more regions of the electromagnetic (EM) spectrum at one or more sample regions of the artwork; at least one processor; and a memory coupled to the processor, the memory containing instructions that, when executed by the processor, configure the system to:form a multi-dimensional array of the image data; create one or more datasets from the multi-dimensional array, wherein the artwork may be identified from the created dataset; and compare the created dataset to one or more previously generated datasets relating to previously acquired image data obtained from a previous encounter of the artwork to evaluate if the artwork is the same as the previous encounter of the artwork.

34. A system for creating and storing a digital representation of an artwork, comprising: at least one imaging platform to acquire image data of an artwork in one or more regions of the electromagnetic (EM) spectrum at one or more sample regions of the artwork; at least one processor; and a memory coupled to the processor, the memory containing instructions that, when executed by the processor, configure the system to: form a multi-dimensional array of the image data; create one or more datasets from the multi-dimensional array, wherein the artwork may be identified from the created dataset; and storing on a distributed ledger one or more of the dataset or a reference to the created dataset.

Citation Information

Patent Citations

  • System and method to optically authenticate physical objects

    US20200143032A1

  • Systems and methods for identifying and authenticating artistic works

    US20200311452A1

  • Optical acquisition system and probing method for object matching

    US20220351492A1

  • System and Method for Certified Digitization of Physical Objects

    US20230171116A1