Sound reconstruction method

WO2026167373A1PCT designated stage Publication Date: 2026-08-13ROSE MICHAEL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-06
Publication Date
2026-08-13

Smart Images

  • Figure GB2026050171_13082026_PF_FP_ABST
    Figure GB2026050171_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Sound Reconstruction Method A method of storing information relating to at least one grooved media. The method comprises receiving at least one image relating to a physical audio track on a grooved medium. The method comprises generating a hash in dependence on the at least one image to produce at least one identifying value. The method comprises storing at least one electronic audio section in a data structure in dependence on the at least one identifying value.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Sound Reconstruction Method

[0002] Field of invention

[0003] The present disclosure relates to recording, replaying, and improving grooved media.

[0004] Background

[0005] Grooved media, notably vinyl records or Edison cylinders are used to physically store an audio signal for later playback. Whilst these forms of audio storage were popular in the past, they were largely superseded by CDs and digital storage at the end of the 20thcentury. Recently however, grooved media has experienced a notable increase in popularity due to their high-fidelity audio and aesthetics.

[0006] Users of grooved media however may still wish to digitise their grooved media so that they can have easy access to the audio content, for instance when travelling as the necessary hardware to play a grooved medium can be cumbersome. Presently analogue-to-digital convertors can be used which take the analogue electrical signal output by a grooved media playing device and convert it to a digital signal which may be stored for instance on a memory stick or in a cloud environment. When the grooved medium is a vinyl record this process is commonly known as ‘ripping vinyl’

[0007] These traditional methods however may result in poor quality audio conversions as some detail in the analogue signal may be lost in the conversion due to improperly configured audio equipment and / or dirty or worn styli and / or grooved media. Because traditional methods may rely on physical styli to make contact with the grooved medium there is also the risk that either will get worn or damaged in the process.

[0008] In trying to overcome these problems and improve audio quality, non-contact optical recording and playback techniques have been devised that record visual data from the grooves or physical audio tracks on a grooved medium to reconstruct a digital audio signal. These methods however conventionally require extremely high magnification of the grooves or physical audio tracks on a grooved medium, are unreliable at detecting the edges of a groove or physical audio track on the grooved medium (particularly when physical defects are present) and so may be unable to recover the audio data with high fidelity and require the processing of vast quantities of data.

[0009] The present invention aims to solve these problems, amongst others.

[0010] Summary of the invention

[0011] Aspects of the disclosure are as set out in the independent claims and optional features are set out in the dependent claims. Aspects of the disclosure may be provided in conjunction with each other and features of one aspect may be applied to other aspects.

[0012] The solution offered by this invention is a non-contact optical method of playing grooved media. The invention is primarily directed towards the playback of phonograph discs, but the methods described herein may also be applied to phonograph cylinders. A device is also provided which implements the methods described herein. Users can play their grooved media with improved sound quality while preserving them as the optical sensor used to play the media requires no physical contact with it.According to a first aspect of the disclosure, there is provided a method of storing information relating to at least one grooved media, the method comprising receiving at least one image relating to a physical audio track on a grooved medium, generating a hash in dependence on the at least one image to produce at least one identifying value; and storing at least one electronic audio section in a data structure in dependence on the at least one identifying value.

[0013] The at least one image may relate to a physical audio track in the sense that at least a section of the physical audio track is depicted in the at least one image, and / or that the at least one image is ‘of’ the physical audio track. The hash may be generated in dependence on the at least one image in the sense that the hash is generated based on the at least one image. The image may be received in any computer-readable format, such as Joint Photographic Experts Group (.jpg or .jpeg), bitmap (.bmp), Portable Network Graphic (.png), Tagged Image File Format (.tiff) or raw image sensor data. The image may be processed before hashing by an image processing algorithm, such as by use of an image deblurring algorithm or a feature extraction algorithm or another algorithm. The hash may be calculated / generated in dependence on the image. For example, generating a hash in dependence on the at least one image may involve (directly) hashing the at least one image. Alternatively, generating a hash in dependence on the at least one image may involve processing (including hashing) the image to generate an intermediate value, and then hashing the intermediate value to produce the at least one identifying value.

[0014] The method may further comprise receiving an electronic audio track corresponding to a physical audio track on a grooved medium, the electronic audio track comprising a plurality of electronic audio sections. The electronic audio track may be obtained by playing the actual grooved medium in part or whole and recording the audio reproduced from the grooved medium. The at least one image may relate to at least one section of the (physical and / or electronic) audio track. The physical audio track may comprise a plurality of physical audio sections.

[0015] Advantageously this may allow for an improved method of storing an audio track recorded from a grooved medium. By storing audio sections in a data structure being indexed by hashes based on the images of the respective track sections, a more useable, smaller size, and potentially more feature-rich (by way of the auxiliary data that will be described) data structure may be provided (as compared e.g. to a simple audio file relating to a digitised audio track). Additionally, storing audio data in this manner may be advantageous as it facilitates fast storage and lookup, and faster extraction or retrieval processes (in contrast to storing the audio data independently of the hash or identifying value). The latter approach would likely require the hash or identifying value to be stored with each corresponding electronic audio section, which would not only require additional storage space, but also more time would be required to search for the hash or identifying value before the electronic audio section associated with it could be extracted. The latter approach would likely require an exhaustive search of the data structure or some other computationally intensive strategy.

[0016] Furthermore, the method may allow audio data to be stored in a data structure in an essentiallyrandomised order (because the data structure is effectively indexed by the hashes) which may prevent access to the audio data without the correct grooved medium which may be used to access the entries in the data structure. This may therefore prevent accessing stored audio data without playing the correct grooved medium.

[0017] The method may also reduce the amount of data required to be stored whilst digitising grooved media. Additionally, the method may overcome the limitations of traditional grooved media playback technology and ensure the characteristic sounds of a grooved medium is reproduced in the stored audio, meaning the grooved media sounds new or un-played even where the grooved medium is worn. Further advantages may include improved speed of storage and retrieval of digitised audio data.

[0018] The term ‘grooved medium’ or ‘grooved media’ is used herein to refer to a product or component that stores in physical form an analogue audio signal. Non-limiting examples include vinyl records which comprise a disc with a groove on a surface or an Edison cylinder which comprise a cylinder with a groove on an external surface. A stylus rests in the groove and as the cylinder or disc is rotated, the stylus traverses the groove, and its motion is converted into an electrical signal. The groove therefore may be a physical representation of the electric signal being produced. Examples of grooved disc media include 12-inch Long-Playing, 7-inch Single, and Extended-Play records.

[0019] Grooved media typically include a plurality of audio tracks. These are individually identifiable continuous pieces of music or audio. The term ‘physical audio track’ is therefore used herein to refer to a single continuous piece or movement of audio reproduced in physical form as a groove on a grooved medium. As an example, a vinyl record (itself a grooved medium) may comprise three distinct physical audio tracks on each surface. The physical audio tracks are typically joined together by a groove containing no audio information, such that the end of one physical audio track automatically leads to the beginning of the next.

[0020] By way of specific example, ‘The The’s 1983 album ‘Soul Mining’ was released as a Long-Playing record containing seven physical audio tracks: four audio tracks on a first surface and three audio tracks on a second surface. Each audio track is identifiable by a name denoted by the artists on the record label or album cover (in the example given, the first physical audio track on the first surface is ‘I've Been Waitin' For Tomorrow (All Of My Life)1).

[0021] When played using an apparatus, the apparatus converts the physical audio track on a grooved medium into an analogue electronic audio track. This is an electronic signal that may be amplified then output to a loudspeaker or other electronic device. This signal may also be converted into a digital electronic audio track using an analogue-to-digital converter. This turns an analogue signal into a digital signal, thereby allowing computer devices to interact with the signal. Electronic audio track is used herein to refer to electronic signals corresponding to a physical audio track.

[0022] The methods described herein may be performed using a single or plurality of computer devices.

[0023] A physical or electronic audio track may be subdivided into smaller sections. For instance, a three-minute-long audio track may be split into three one-minute sections, such that, whencombined, the one-minute sections join to form a complete audio track. Specifically, as will be described in further detail, the method involves dividing an image of a physical audio track into smaller sections using image processing techniques.

[0024] The term ‘physical audio section’ is therefore used herein to refer to a subdivision of a physical audio track. The term ‘electronic audio section’ is used herein to refer to a subdivision of an electronic audio track.

[0025] Hashing is used herein to refer to the process of applying a hash function to data. A hash function takes data of either fixed or arbitrary size and maps it to a fixed sized value, for instance a numerical value. The data that is input into the hash function is known as the preimage. The output of the hash function is known as the hash or digest. Since hash functions always produce the same results when acting on the same data, hash functions may equivalently be viewed as a map, mapping a preimage (or “key”) onto a hash (or “value”). Different hashing algorithms may be used to implement hash functions, for example Message Digest Algorithm 5 (MD5), Secure Hash Algorithm (SHA), Perceptual Hash (pHash), Difference Hash (dHash) or Average Hash (aHash). To express the above in a different way, a hash function is a function on binary data (i.e. , image pixels) for which the length of the output is fixed. The output of a hash function is a hash code that may be used to index a data structure containing data or pointers to data. Herein, the terms ‘hashing algorithm(s)’ and ‘hash function(s)’ are used interchangeably as they have broadly the same purpose namely hashing data.

[0026] As used herein, the terms a hashed image and an identifying value are used interchangeably. A hashed image is a fixed-length value produced by applying a hash function to an image. As such, the hashed image is an identifying value as it may be used to identify the image that was used to create the identifying value (as the hash function maps the same data to the same fixed value output).

[0027] The invention does not have to permanently store any images but hashes may optionally be permanently stored to implement some features of the invention.

[0028] The electronic audio track corresponds to a physical audio track by being the electronic audio signal generated by playing the physical audio track. This may be achieved for instance by using a turntable or similar apparatus which uses a small stylus configured to rest in the physical audio track groove and convert the vibrations generated by the stylus traversing the groove into electronic signals.

[0029] As described herein, an electronic audio section may be an analogue or digital signal, and the methods provided herein may further comprise digitising an analogue electronic audio section. For instance, the electronic audio sections received from the apparatus playing a grooved media are analogue electronic audio sections, and the electronic audio sections stored in a data structure are digital electronic audio sections. Advantageously this may improve communication with a computer device used to perform the method.

[0030] An image of the physical audio track may be an image in which the physical audio track can be seen. The image may be sufficiently detailed to capture the minute groove details that generate notable audio differences, or the image may be a high-level view of the audio trackshowing a greater portion of the track. Of course, any magnification between these two extremes may be used.

[0031] The image may be in any format able to be processed by a computer device. Non-limiting examples include Joint Photographic Experts Group (.jpg or .jpeg), bitmap (.bmp), Portable Network Graphic (.png), Tagged Image File Format (.tiff), or raw image sensor data.

[0032] A perceptual hashing algorithm may be used to perform the hashing operation on the at least one image. This may require lower resolution image data and reduce the volume of data required to digitise the grooved medium.

[0033] A similarity preserving or locality sensitive hashing algorithm may also be used to perform the hashing operation on the at least one image. Advantageously, similarity preserving or locality sensitive hashing algorithms tend to maximize collisions for visually similar images. This means that if an audio data structure is created using physically pristine grooved media, but the user presents non-pristine grooved media for playback, the hashing process may effectively cause some types of groove defects or microscopic debris in the grooves to be ignored so that the resulting audio quality will not be adversely affected.

[0034] The at least one image may be split into multiple subimages. Each subimage may be hashed using the same or different hash functions. The hashes from all the hashed subimages may be combined using methods described herein or known in the art. The result (the combined hashes) may be used in place of a hash of the at least one image.

[0035] The image may be taken continuously of the physical audio track and separated into sections, or a number of images may be taken along the length of the physical audio track (where the ‘length’ of a physical audio track refers to the length along the entire path of the track, and e.g. length in a single dimension). The images are hashed which may reduce data usage and increase security of the method. These images (or their representations or the hashes of these images) may be treated as visual words.

[0036] If taken continuously, the computer device may perform a sectioning step in which the image of a physical audio track is sectioned into images of physical audio track sections. The collection of such images or their representations or the collection of their hashes corresponding to a physical audio track or part of a physical audio track may be treated as a document, comprising a collection of visual words, and a hash derived from the collection may be treated as a document identifier. An electronic audio section duration may first be defined by the computer device, for instance by separating the corresponding electronic audio track into equal or approximately equal sub-second sections. The computer device may then separate the image of the physical audio track into the same number of equal or approximately equal sections (along the length of the groove) as the electronic audio track. In other words, along the length of the groove the image sections have the same proportions as their corresponding sections in the electronic audio track.

[0037] Intermediate data structures may be used in the storage and processing of data relating to the methods described herein. Electronic audio sections and / or pointers to the locations in other data structures of electronic audio sections may be stored in the intermediate data structures.The at least one image may show a view of a physical audio section, and so the at least one image may relate to at least one section of the audio track. Advantageously, by taking an image of a smaller section of the physical audio track this may allow for a higher resolution of image as the image can contain a smaller field of view, thereby improving the hashing of the image. The image may be captured in a manner such that the groove in the image appears substantially straightened or “unwrapped”. For example, a suitably mounted line scan camera fitted with a microscope objective lens may be used to acquire the image. If the camera lens is sharply focused on a groove and remains so while tracking the groove in the manner of a conventional lineartracking tonearm whilst the grooved medium is rotating, then the groove in the resulting image may appear substantially straightened or “unwrapped”.

[0038] A groove segment or physical audio section shown in the image may be identified using an image segmentation technique. Segmentation divides an image into its constituent regions or objects. Segmentation allows extracting objects in images. In this case the objects of interest are the groove segments or physical audio sections shown in the image.

[0039] Image segmentation techniques are described in ‘A Survey on Threshold Based Segmentation Techniques in Image processing’ in the November 2014 Special Issue of the International Journal of Innovative Research and Development. Pages two and four of this journal issue describe global and adaptive thresholding techniques which may be implemented by this invention to identify groove segments or physical audio sections.

[0040] These techniques may produce binary images as a first step which clearly show the boundaries of the groove segment or physical audio section and makes their recognition and subsequent sectioning prior to hashing easier. Neither the binary image nor the final segmented image is hashed. Instead, they are used to identify the correct regions of the original image to isolate, section and subsequently hash. They are used to identify at least approximately the pixels and / or pixel coordinates in the original image that form the edges or boundaries of the groove segment or physical audio section shown in the image.

[0041] Some variations of the invention may use different image segmentation techniques. For example, they may use Otsu’s method of thresholding. Other variations of the invention may employ a trained neural network or vision model to perform the image segmentation step. Any other method of identifying groove segments or physical audio sections in images may be used.

[0042] Image segmentation may be combined with a template matching step where images of groove segments or physical audio sections are matched against those from successive and / or previous time steps in order to track and / or measure the lateral movement or drift or shift of a groove segment or physical audio section across multiple images, for example as the grooved medium is rotating on the platter of the playback device in the case of phonograph records. In the case of phonograph cylinders the medium may be mounted on a mandrel.

[0043] The invention may use the data collected from the template matching step to control an electric motor responsible for positioning the image sensor over the surface of the grooved medium. The number of electronic audio sections and the number of hashed images may be equal.Advantageously this may make the storage process of the electronic audio sections in the data structure easier.

[0044] This may be achieved for instance by the computer device controlling an imaging device and triggering the image capture at the start of each electronic audio section. This may also be achieved by post-processing in which the computer device separates a recorded electronic audio track and an image of a physical audio track into the same number of smaller sections.

[0045] The method may further comprise associating an identifying value with an electronic audio section corresponding to an audio output of a physical audio section contained in the image that was used to produce the identifying value. The audio output of a physical audio section may refer to the audio information / data that is encoded into the physical audio section, and which may be output when the grooved medium to which the image relates is played using suitable apparatus. The physical audio section is ‘contained in the image’ in the sense that it is depicted in the image. The physical image section may be a subsection of the aforementioned physical audio track. That is, the physical audio track comprises a plurality of physical audio sections. Advantageously this may facilitate easier and faster storage of the electronic audio sections in the data structure as the correct electronic audio section to be stored can be identified by its corresponding image.

[0046] This linking or association process may be carried out using a computer device. The computer device may receive an electronic audio section and an image of the physical audio section that produced the (same) electronic audio section at the same time. The computer device can therefore associate the electronic audio section with the image. This process may be iterated for each electronic audio section and image.

[0047] The electronic audio section may be stored at a location in the data structure in dependence on the identifying value.

[0048] The electronic audio section may be stored at an index of the data structure that is equal to the identifying value. The electronic audio section may be stored in the data structure at an index of the data structure that is numerically equal to the numerical value of the hashed image. Alternatively, the electronic audio section may be stored in the data structure at an index that is numerically equal to the numerical value of the hashed image minus an offset (the offset for example being the value of the lowest generated hash). This may speed up the storage process of the electronic audio sections in the data structure as the data structure can be made smaller. Advantageously the offset may further reduce the time needed to populate the data structure.

[0049] The data structure may further be copied or inserted into a database for persistent storage and later extracted from or copied out of the database when required.

[0050] The at least one image of the grooved medium may contain depth information of the groove. Advantageously, this may provide more information to the hashing algorithm and may reduce the number of errors during the hashing process.

[0051] The grooved medium may comprise or may be modified by making at least two physical marks that are used to identify the start and / or end of a physical or electronic audio track.Advantageously this may make detecting the start and / or end of a physical or electronic audio track easier.

[0052] The method may further comprise identifying the at least two physical marks using image processing techniques thereby identifying the start and / or end of a physical audio track or section. The method may further comprise identifying the start or end of an electronic audio track or section by detecting in the electronic audio section a signal corresponding to the physical mark.

[0053] The physical marks may be made transverse to an unmodulated or silent part of the groove before the start of an audio track and after the end of the same audio track. The physical marks may be made in a manner that prevents groove skipping by a playback apparatus as the groove is traversed. Herein the physical marks may be known as track boundary marks as they serve the purpose of marking the boundaries of a physical audio track.

[0054] The start or end of an electronic audio track or section may be identified by detecting a signal corresponding to the physical mark. As the physical mark is traversed by a phonograph stylus of a playback device, it may produce a transient noise pulse in a resulting recorded electronic signal recorded from the audio generated by the playback device. This transient noise pulse may be detected using an algorithm (such as an algorithm making use of transient detection techniques or audio onset detection techniques, or a combination thereof).

[0055] The transient noise pulse may also be detected by cross correlating the resulting recorded electronic signal with a sample recorded transient noise pulse. The resulting third signal has peaks at sample positions corresponding to the transient noise pulse. This enables the start time or position of the transient noise pulse in the recorded electronic signal to be determined. By detecting transient noise pulses caused by physical marks, the computer device may detect the start or end of a physical or electronic audio track.

[0056] By detecting the presence of physical marks in an image using image processing techniques, the computer device may detect the start or end of a physical audio track.

[0057] If the at least one electronic audio section is an analogue signal, the method may further comprise digitising the electronic audio section.

[0058] A plurality of hashed images or identifying values may be created from a single image of a physical audio track or section. A single image may be hashed multiple times, optionally using different hashing algorithms or a single image may be separated into sections with each section hashed separately.

[0059] The single image of a physical audio track may be separated into a plurality of images each containing a physical audio section. This may be achieved by identifying physical boundary marks in the image.

[0060] The provided method may further comprise a step in which auxiliary data related to the grooved medium is stored. Advantageously this data may improve the storage and subsequent retrieval and playback processes as more data is known about the grooved medium (for instancechoosing a particular hashing algorithm or data structure based on the known technical characteristics of the recording such as groove modulation type and disc equalization requirements for correct reproduction). Optionally, a user may configure at least some of the auxiliary data.

[0061] Auxiliary data may also be known herein as instruction or information codes, former of which relate to data configured to instruct an apparatus to perform a certain task. For instance, the instruction code may instruct the apparatus to use a particular hashing algorithm or play back the grooved medium at a certain speed. Information codes are auxiliary data used to provide additional information about the grooved media or physical audio track which may be interesting to the user or used by the method to improve the quality of the output. For instance, the auxiliary data may inform that the grooved medium should be rotated at 45RPM (revolutions per minute) (with 16, 33 and 78RPM being alternatives) to yield the best results.

[0062] Each physical audio section may have a number of different information and / or instruction codes associated with it, meaning that as an audio track is played there is a stream of additional data associated with the track which may be accessed (where this data stream relates to sections of the track, and not merely the track as a whole). This may provide a more featurerich way of storing and accessing audio data.

[0063] The auxiliary data may be stored in separate data structures to the data structures used to store the electronic audio sections. These data structures may be copied or inserted into a persistent storage database which may allow them to be accessed later.

[0064] The auxiliary data may be stored in and accessed from the data structures using the methods described herein to store and access electronic audio sections. During playback, the auxiliary data is preferably accessed before the electronic audio section is retrieved as this may allow for additional time for the computer to process the data.

[0065] The auxiliary data may be associated with the physical audio sections. The auxiliary data may have information and instruction codes associated with them.

[0066] The auxiliary data may comprise any of the following: a groove type; a hash function type; a genre; a disc speed; a track start code; a track end code; a time value; a silent groove; an image dimension; a groove modulation type; a colour; a groove radius; a groove number; a hash or identifying value; a randomly generated number; and a pseudo-randomly generated number.

[0067] The auxiliary data may comprise instructions to do any of the following (during storage, replay, etc.): change a hash function to be applied; select a different digitised audio data structure; get current time; start a timer; stop a timer; get elapsed time; mute audio; generate an audio tone; modify an audio signal and set a digital output port to high or low. It will be appreciated that the use of ‘instruction codes’ in the auxiliary data may further improve the functionality involved in storing, accessing, and processing audio data and controlling grooved media. Many example use cases can be contemplated - e.g. a user may use instruction codes to mute or skip a section of track containing profanity during replaying, or to adjust audio settings (e.g. to boost bass) in certain sections of a track during replaying.The auxiliary data may also provide genre related information of the audio track. The genre of an audio track may be determined by analysis of the auxiliary data relating to the physical audio sections (in particular, auxiliary data related to the genre of the relevant section) and assigning a genre to the track if the particular genre is found to be in the majority throughout the sections of the physical audio track (i.e. if the majority of the auxiliary data related to the genre of the various sections show a particular genre). Alternatively, if there is no clear majority then a genre may be assigned to an audio track on the basis of the longest uninterrupted sequence of auxiliary data containing a particular genre. Determining the genre of an audio track may also / alternatively involve counting instances of auxiliary data relevant to genre in relation to a physical audio section; preferably counting sequences of such instances; and more preferably, if multiple instances of such sequences are present, determining the genre of an audio track by determining the longest such sequence (and / or the sequence having the higher count value).

[0068] The method may further comprise creating electronic waveforms from the auxiliary data. In a variation, a user may label a portion of the physical audio track with an instruction code. The method may further comprise modifying an electronic audio section based on the auxiliary data.

[0069] The method may further comprise associating each at least one physical or electronic audio section and / or each at least one identifying value with auxiliary data. That is, the physical or electronic audio section and / or identifying value is associated with (multiple data fields of different types of) auxiliary data in the data structure. In this sense, a data structure is created which provides information (by way of the auxiliary data) in relation to each physical or electronic audio section.

[0070] The data structure may be conceptualised as having a first ‘column’ consisting of the identifying value and further columns relating to auxiliary data. The ‘rows’ of the data structure relate to the different physical audio sections.

[0071] The method may further comprise associating each at least one electronic audio section and / or each at least one identifying value with auxiliary data.

[0072] The locations or indexes of the electronic audio sections in the data structure may be stored in a second data structure. Advantageously this may speed up the retrieval process of the electronic audio sections from the first data structure as their locations are clearly accessible in an organised structure.

[0073] The method may further comprise applying an additional hashing process to at least one of: the at least one identifying value, and the electronic audio sections.

[0074] According to another aspect of the disclosure, there is provided a method of retrieving an electronic audio section stored in a data structure, the method comprising receiving at least one identifying value in dependence on at least one image of a grooved medium (specifically of a physical audio section), indexing the data structure in dependence on the at least one hashed image, and extracting the electronic audio section from the data structure at the same index.Alternatively, the data structure may be indexed in dependence on the at least one hashed image minus an offset (the offset for example being the value of the lowest generated hash mentioned previously).

[0075] Advantageously this may provide a fast and reliable method of extracting an electronic audio section.

[0076] Each image used to generate the at least one identifying value may show a physical audio track section corresponding to an electronic audio signal stored in the data structure.

[0077] A plurality of electronic audio sections may be extracted and combined to form an electronic audio track.

[0078] The hashed image may index the data structure according to a stored electronic audio section corresponding to the electronic audio section of the physical audio section contained in the image that was used to produce the identifying value. Advantageously this may simplify and reduce the time required to access and retrieve the electronic audio section from the data structure.

[0079] The extracted electronic audio section may be converted to an analogue signal using a digital-to-analogue converter and played through a speaker using an amplifier. Advantageously this may allow for the method to be implemented in a live playback device and may reduce the burden on the user in transforming the retrieved electronic audio track or electronic audio section into playable audio.

[0080] The method may further comprise extracting auxiliary data associated with the physical audio section from either the data structure or an additional data structure. The method may further comprise presenting extracted auxiliary data to a user. The method may further comprise performing, in response to receiving the auxiliary data from the data structure, any one or combination of the following actions: changing a hash function to be applied to further images of the grooved medium; selecting a different data structure for retrieving electronic audio sections; get current time; starting a timer; stopping a timer; getting elapsed time; muting audio; generating an audio tone; modifying the audio signal; and setting a digital output port to high or low. The method may further comprise determining, in response to the auxiliary data from the data structure, a genre of the audio track. The method may further comprise counting auxiliary data associated with a plurality of physical audio sections or a plurality of identifying values thereby to determine further data in respect of the physical audio track; preferably wherein the further data comprises a groove type; a genre; a disc speed; a groove modulation type; and a disc colour. It will be appreciated that such further data relates to the track as a whole, and not to the individual sections (which make up the track). The further data may be stored in a separate data structure. The method may further comprise modifying an electronic audio section based on the auxiliary data.

[0081] Generating a hash in dependence on the at least one image may comprise: applying a local binary pattern feature extraction algorithm to the image, producing a representation of the image, and producing a hash that is based at least in part on the representation of the image. Hashes produced in this manner may also be treated as visual words. Advantageously thismay reduce errors during the hashing process and may reduce the frequency of hash collisions between multiple images.

[0082] A local binary pattern feature extraction algorithm may refer to an algorithm configured to turn an image into a local representation of texture. This algorithm accentuates changes in texture or gradient of the image, thereby allowing changes to be more easily identified. When applied to an image of a physical audio track or section on a grooved medium, this accentuates features of the groove making them easier to detect.

[0083] An example of how a hash may be produced that is based at least in part on an image representation is given in the academic paper titled “Aggregated Bidirectional Local Binary Pattern for Robust Perceptual Image Hashing" by Qasim Abbas, Jannat Shirazi and Yi-Ping Phoebe Chen (2022).

[0084] The at least one image may be a plurality of images.

[0085] The method may further comprise cropping a particular electronic audio section to which a particular image relates based on a comparison of the particular image with an adjacent image. The method may further comprise determining an image overlap area between two adjacent images. Specifically, the method may further comprise determining an area of overlap between two adjacent images (corresponding to a previous timestep and a current timestep) from the plurality of images, determining an overlap between two adjacent electronic audio sections (corresponding to the previous timestep and the current timestep) corresponding to the two adjacent images in dependence on an image overlap ratio, and removing the overlap between the two adjacent electronic audio sections from the electronic audio section at the current timestep.

[0086] The image overlap area may be determined using a template matching process. The method may further comprise calculating an image overlap ratio in dependence on the image overlap area and the non-overlap area of the image. The method may further comprise removing a portion of the at least one electronic audio section in dependence on the image overlap ratio. When successive images of a grooved medium are taken, it is inevitable that each image will not show only exclusive image data - there will likely be a part of the image that is shared with another image taken immediately before or after the present image. T o determine the amount of each image that shows a physical audio section not shown in other images, an image overlap ratio may be calculated.

[0087] The overlap ratio of an image may be calculated by determining an area of overlap between two adjacent images of a physical groove on a grooved medium and dividing the overlap area by the overlap area plus the non-overlap area in the image for which the ratio is to be calculated. This will return a value between zero and one corresponding to the proportion of the image that overlaps with the other image. It will be appreciated that two such overlap ratios could be calculated for an image immediately after and immediately before other images; one corresponding to an image overlap at one end of the image, and one corresponding to an image overlap at the end of the image.Alternatively, the overlap ratio of an image could be calculated similarly as above but using the width of the image that is overlapped or not overlapped with other images rather than the area, thereby simplifying the calculation.

[0088] The image overlap ratio may then be used to determine the equivalent overlap that will exist in the electronic audio sections corresponding to the physical audio sections shown in each image. This may allow for example the overlapped audio data to be removed during a processing operation, thereby improving the audio quality of a complete electronic audio track formed from these electronic audio sections.

[0089] Template matching may be used to perform the comparison to determine the overlapping area between the images, where the corresponding amount of overlap may be removed from the particular electronic audio section. A template matching process involves comparing two images of physical audio sections (corresponding to the electronic audio sections) captured at successive time steps (making the images adjacent to each other) in which one image is used as a template and the other a reference to identify the overlap area in an image of a physical audio section. The image may also be cropped. The term ‘adjacent image’ in this context refers to an image that is taken immediately before or after the particular image.

[0090] Generating a hash in dependence on the plurality of images may comprise utilising a localitysensitive hashing method.

[0091] The locality-sensitive hashing method may comprise: hashing each image with a perceptual image hash function to produce a plurality of hash digests, each hash digest comprising a plurality of positions, each position in each hash digest comprising a binary digit; for each position, summing the binary digit at the position across all the hash digests to produce a position sum for each position; computing the average value of the position sums; and setting a result hash, wherein for each position in the result hash, when the position sum is greater than the average value of the position sums, then the position is set to 1 , and when the position sum is not greater than the average value of the position sums, then the position is set to 0. Alternatively, other methods of combining hashes that are known in the art may be used. The number of hashed images and the number of electronic audio sections may be equal. Advantageously this may make the association process of the electronic audio sections with the hashed images easier.

[0092] The electronic audio section may be stored in a data structure at an index dependent on the hashed image. Advantageously this may allow for a securely stored electronic audio signal and may reduce the chance of the electronic audio section being accessed without permission.

[0093] The method may further comprise producing a plurality of identifying values from a single image of the grooved medium.

[0094] The method may further comprise separating the hash into one or more pixel triplets (for instance RGB (Red-Blue-Green) or BGR (Blue-Green-Red) triplets). The method may further comprise appending bits to the hash to form the one or more pixel triplets. The one or more pixel triplets may be communicated and / or transmitted over a wired or wireless network and / ordisplayed on an electronic display. The method may also treat the hash as image pixels. The start or end of a physical audio track may be detected using a groove state change detection algorithm.

[0095] Herein a groove may have one of three states that may change along the path of the groove. These states are: unmodulated, modulated, and modified.

[0096] A groove state change detection algorithm may be an algorithm performed by circuitry or a computer device that detects a change of the groove state which may be used to detect a physical mark for instance track boundary mark or the start or end of a physical audio track. The groove state change detection algorithm may search a groove for a change of the groove state by comparing windows of groove image data or their hashes with successive ones until a difference is detected, using image processing techniques. The search may begin on an unmodulated or silent groove segment and progress towards a change of the groove state. For example, assume that the groove state change detection algorithm starts on an unmodulated or silent part of the groove (before audio starts) and while moving toward the start of audio it compares four successive image windows along the groove: A, B, C, and D (where A is furthest from the start of audio in comparison to the other windows). If B is considered identical to A, and C is considered identical to B, but D is not considered identical to C, then a change of the groove state has been detected in window D.

[0097] Groove state change detection may involve searching for a point on the groove where one groove state ends and another one begins. For example, if the current state is unmodulated and the next state is modified then a change of groove state has occurred and implies that a track boundary mark has been detected. For another example, if the current state is unmodulated and the next state is modulated then a change of groove state has occurred and implies that the start of audio has been detected.

[0098] Advantageously, this may allow for accurately locating a change of the groove state.

[0099] In a first stage of the groove state change detection algorithm, windows of image data or their hashes may be taken from the processed images and compared with successive windows of image data or their hashes to determine the difference between them. For image data this may be by mean-squared error analysis, wherein a change of the groove state may be detected when the mean squared error exceeds a threshold, by a structural similarity index measure, a Manhattan distance, a sum of absolute differences method, a sum of squared differences method or a Euclidean distance or by any other known method. For hashes this may be by Hamming distance, wherein a change of the groove state may be detected when the Hamming distance between the hashes exceeds a threshold or by any other suitable method.

[0100] In a second stage, at least one dimension of the windows and / or a space between them may be adjusted after the first stage.

[0101] Stages one and two may be repeated starting at the groove position at which a change was last detected. In other words, the window showing the change of the groove state may befurther sectioned and stages one and two repeated over the new windows.

[0102] Advantageously, this may allow for accurately locating a change of the groove state.

[0103] In a third stage an edge detection method (for instance using a Canny operator) may be used to detect the edges of a track boundary mark or applied coating with subpixel accuracy. Advantageously, this may allow for accurately locating a track boundary mark or applied coating.

[0104] A window of image data may include the image of a physical audio section itself or a section of it.

[0105] The images under analysis may be overlapping or non-overlapping, and accordingly the corresponding electronic audio sections may be overlapping or non-overlapping. An overlapping electronic audio section contains a part of another electronic audio section.

[0106] If the images under analysis are overlapping the duplicated or extra audio data may be removed from the corresponding electronic audio sections before being passed to the input of the digital to analogue converter.

[0107] This may be done by identifying the overlap and non-overlap areas present in an image of a physical audio section (for instance using a template matching process as previously described), using the identified overlap and non-overlap areas to compute an image overlap ratio; using the computed overlap ratio or quantity of overlapped data present in an image to compute the number of overlapping audio data samples present in an electronic audio section, and removing the computed number of overlapping audio data samples from the appropriate position within the electronic audio section.

[0108] The term audio data samples is used herein to refer to discrete numbers that make up an electronic audio section. Analogue audio signals must be sampled at discrete intervals to produce digital audio data; therefore an electronic audio section is simply a sequence of audio data samples. For example, a common audio sampling frequency of 44.1kHz would mean that a 1 -second-long electronic audio section contains 44,100 audio data samples, each sample of which is a discrete number.

[0109] According to another aspect of the disclosure there is also provided a system for digitising grooved media, the system comprising means for receiving an electronic audio track corresponding to a physical audio track on the grooved medium, the electronic audio track comprising a plurality of electronic audio sections, means for receiving at least one image of the physical audio track or audio section, means for separating and hashing the at least one image of the physical audio track or audio section, and means for storing the at least one electronic audio section in a data structure, in dependence on the at least one hashed image.

[0110] An image of a physical audio track may be separated into images of physical audio track sections.

[0111] The means for receiving the electronic audio track may be an audio recording system (forinstance a record player comprising a turntable and pickup) which may further comprise an analogue to digital converter, an amplifier, a transmitting device (for instance wires or wireless communication devices) and a computer device.

[0112] The audio recording system may support multi-channel channel audio recording. In a first step, the individual audio tracks are recorded, then in a second step each channel is separated and saved as uncompressed Waveform Audio File format (WAV). For example, a stereo recording is split into left and right channels and a four channel or quadraphonic recording is split into left-back, left-front, right-back, and right-front channels. This procedure may be repeated. Each saved audio channel may be corrected for wow and flutter and resaved as a raw uncompressed audio file. In other words, wow and flutter are removed from the recording and it is stored as a raw uncompressed audio file.

[0113] The means for receiving at least one image of the physical audio track may be an imaging apparatus (for instance a camera, optionally mounted to the playback apparatus), a transmitting device and a computer device. The terms ‘imaging apparatus’ and ‘imaging device’ are used interchangeably herein.

[0114] The imaging apparatus or playback device may include or incorporate a plurality of sources of optical energy configured to illuminate part of a surface of the grooved medium such as the area being imaged by the imaging apparatus. For example, a plurality of light emitting diodes may be used.

[0115] The images may be two dimensional or contain depth information. The groove in the at least one image may appear substantially straight or ‘unwrapped’. The at least one image may be stored and accessed as files on storage media (such as an external USB device).

[0116] The start and end boundaries of a physical audio track may be found by searching images of the physical audio track sections with a groove state change detection algorithm. The detected boundary may be a row in the image. If a detected boundary is a row in the image, then the other detected boundary may also be a row in the image. If a detected boundary is a column in the image, then the other detected boundary may also be a column in the image. In the case where the start and end boundaries are set by physical marks (as opposed to the normal start or end of groove modulation) they may not be precisely perpendicular to the groove if they are made by hand. In that case the pixels forming the boundary in the image may not form a straight path that aligns with a row or column in the image. The row position and / or column position of pixels on detected boundaries may be real numbers.

[0117] An image of a physical audio track may be split into multiple, equal sized (along the length of the groove), consecutively numbered images and stored as individual files with the image number in the file name. The image file with the lowest number in its file name may contain the image data corresponding to the start of the physical audio track and the image file with the highest number in its file name may contain the image data corresponding to the end of the physical track. The collection of such image files may represent a complete physical audio track.

[0118] The means for hashing may be a computer device carrying out a hashing algorithm usingstored instructions or may be other electronic circuitry configured to perform the same process. The means for storing may be a computer device with an external or internal memory.

[0119] According to another aspect of the disclosure there is provided a system for retrieving an electronic audio section stored in a data structure, the system comprising a means for receiving at least one hashed image of a grooved medium comprising a physical audio track section corresponding to the electronic audio signal stored in the data structure, a means for indexing the data structure in dependence on the hash of the at least one image, and a means for extracting the electronic audio section from the data structure. The hash of the at least one image may be a hashed image or a hash of a plurality of images.

[0120] The system may also comprise a means for receiving and hashing an image of a grooved medium comprising a physical audio track section corresponding to the electronic audio signal stored in the data structure.

[0121] According to another aspect of the disclosure there is provided a method of detecting, in a grooved medium, the beginning or ending of the modulation of a groove (such as a track boundary mark or a change in the modulation of a groove) using image processing techniques. Using image processing techniques may comprise using an edge detection method. The method may further comprise determining a similarity between pixels in a first window along a groove and pixels in a second window along a groove; and when the similarity is above or below a threshold, detecting a change of the groove state.

[0122] In a variation the method may further comprise determining a difference between pixels in a first window along a groove and pixels in a second window along a groove; and when the difference is equal to or above a threshold, detecting a change of the groove state.

[0123] The method may further comprise determining a similarity or difference comprising using one of: a mean squared error function; a root mean square error function; structural similarity index measure; a Hamming distance associated with hashes of the first window and the second window; a Manhattan distance associated with hashes of the first window and the second window; and a Euclidean distance associated with hashes of the first window and the second window. The method may further comprise translating the first and second window along the groove.

[0124] The method may further comprise determining a difference comprises using one of: a sum of squared differences method; a sum of absolute differences method; a mean squared error function; and a root mean square error function. The method may further comprise translating the first and second window along the groove.

[0125] According to another aspect of the disclosure there is provided a method for detecting movement of a grooved medium depicted in a plurality of images, the method comprising: receiving a plurality of images of the grooved medium; and determining whether a first property associated with a first image of the plurality of images differs from a second property associated with a second image of the plurality of images thereby to detect movement of the grooved medium.The first and second properties may be first and second hashes generated in dependence on respective first and second images of the plurality of images.

[0126] The detected movement may relate to the spin of the grooved medium; preferably the presence or absence of spinning and / or a reversal in direction of the spinning.

[0127] According to another aspect of the disclosure, there is provided a method of storing information relating to grooved media, the method comprising receiving a plurality of images relating to a physical audio track on a grooved medium, generating a hash in dependence on each image to produce a plurality of identifying values, and storing the plurality of identifying values in a data structure.

[0128] According to another aspect of the disclosure there is provided a computer-readable medium storing a data structure; the data structure comprising a first data field related to sections of a physical audio track on a grooved medium; and a plurality of further data fields related to properties of: the sections of the physical audio track and / or electronic audio sections corresponding to the sections of the physical audio track.

[0129] The first data field may relate to identifying values associated with the sections of the physical audio track; preferably wherein the identifying values are produced by generating respective hashes of respective images of the physical audio track.

[0130] The data structure may further comprise at least one data field related to properties of the physical audio track and / or the corresponding electronic audio track; preferably wherein said at least one data field is populated by counting entries in at least one of the further data fields. The data structure may further comprise a plurality of yet further data fields related to instructions to be performed.

[0131] The means for performing the receiving, the indexing, and the extracting may be a computer device.

[0132] The methods described herein may also store a data structure, hash collision statistics, image dimensions, details about the hash function used to populate the data structure, array hash digests, details about the physical audio tracks used to populate the array such as genre, track title and its duration, the number of audio sections from each track used, details about the audio equipment, settings and engineers used for the audio capture, copyright information and other data or metadata in a data structure such as a file. The data structure may have a field which stores a value which is set by a recording engineer to a value indicating locked or open. The method may further comprise a step in which a value in this field is read and based on its value allows or denies the user to perform certain actions in relation to the data inside the data structure. The data structure may also be inserted into a database for persistent storage or later exportation.

[0133] It will be appreciated that features described in reference to the storing of data relating to a grooved media may also be applied to the method of retrieving data relating to grooved media, and vice versa.Brief Description of the Drawings

[0134] Figure 1 shows a computer device on which aspects of the disclosed system are implemented.

[0135] Figure 2 shows a grooved medium comprising a physical audio track and a physical audio section.

[0136] Figure 3 shows a grooved medium on a playing device with a detachable stylus and a detachable imaging apparatus.

[0137] Figure 4 shows a method for digitising and storing electronic audio sections in a data structure. Figure 5 shows a method for retrieving stored electronic audio sections from a data structure.

[0138] Figure 6 shows a method for storing electronic audio sections and their indexes in two data structures.

[0139] Figure 7 shows a flow of information for a method for digitising and storing electronic audio sections in a data structure.

[0140] Figure 8 shows a flow of information for a different method for digitising and storing electronic audio sections in a data structure.

[0141] Figure 9 shows a pair of overlapping images with a highlighted overlap area.

[0142] Detailed Description of the Embodiments

[0143] Embodiments of the claims relate to a method 4 for digitising and storing physical audio tracks contained on a grooved medium. In particular, embodiments of the claims relate to a method 4 that uses hashed images of a physical audio section on a grooved medium to store digitised electronic audio sections in a data structure.

[0144] An image showing a section of the physical audio track on the grooved medium can be associated with the corresponding electronic audio section produced by recording or playing (for instance using a stylus) the section of the physical audio track in the image. This image is hashed, and the resulting hash is used as an identifier for the electronic audio section produced by recording or playing the section of physical audio track in the image. The electronic audio section may then be stored in a data structure in dependence on the hashed image.

[0145] Other embodiments of the claims relate to a method 5 for retrieving stored electronic audio sections relating to a grooved medium from a data structure. In particular, embodiments of the claims relate to a method 5 that receives hashed images of a grooved medium, indexes a data structure in dependence of the hashed images, extracts an electronic audio section from the data structure, and concatenates the electronic audio sections to form an electronic audio track.

[0146] Electronic audio sections may later be retrieved from the data structure firstly by retrieving hashed images of the grooved medium. Each hashed image initially was an image of a section of physical audio track on the grooved medium that has been hashed. The data structure maythen be indexed in dependence on this data to access the corresponding electronic audio section. A plurality of electronic audio sections may be retrieved in this manner and then concatenated to produce an electronic audio track corresponding to the physical audio track present on the grooved medium. This electronic audio track (now being in digital form) may then be converted to an analogue form and played using an audio amplifier and loudspeakers. Referring to Figure 1 , there is shown a computer device 1000 on which aspects of the systems and methods disclosed herein may be implemented. The computer device may be a part of the device, e.g. a computer device that is integrated into the device or the computer device may be a standalone computer device, such as a laptop or a personal computer. The methods disclosed herein may be implemented using a plurality of (different) computer devices. A computer device may also be referred to as a ‘computing device’.

[0147] The computer device 1000 comprises a processor 1002, a communications interface 1004, a memory 1006, storage 1008, and a user interface 1010 coupled to one another by a bus 1012. The processor 1002 executes instructions, including instructions stored in the memory 1006 and / or the storage 1008. The processor typically comprises a central processing unit (CPU) and / or a graphical processing unit (GPU).

[0148] The communications interface 1004 enables the computer device to communicate with other computer devices. The communications interface may, for example, be an Ethernet network adaptor coupling the bus 1012 to an Ethernet socket, which Ethernet socket is coupled to a network, such as the Internet. It will be appreciated that any other communication medium may be used by the communications interface, such as area networks, infrared communication, and Bluetooth®.

[0149] The memory 1006 stores instructions and other information for use by the CPU 1002. The memory typically comprises Random Access Memory (RAM) and / or Read Only Memory (ROM).

[0150] The storage 1008 provides mass storage for the computer device 1000. In different implementations, the storage is an integral storage device in the form of a hard disk device, a flash memory or some other similar solid state memory device, or an array of such devices. The storage 1008 may comprise removable storage. In various implementations, the storage may comprise an optical disk, for example a Compact Disc Read Only Memory or Digital Versatile Disk, a portable flash drive or some other similar portable solid state memory device, or an array of such devices. In other embodiments, the removable storage is remote from the rest of the computer device 1000 and comprises a network storage device or a cloud-based storage device.

[0151] The user interface 1010 enables a user to interact with the computer device. The user interface is arranged to receive an input from a user and / or to provide an output to a user. For example, the user interface may comprise a touchscreen, a keyboard, and / or a mouse.

[0152] The computer device 1000 is arranged to implement aspects of the methods disclosed herein and may be configured for this purpose. For example, the communications interface 1004 maycomprise a plurality of interfaces to provide redundancy and to enable external devices to communicate with the computer device.

[0153] A computer program product is provided that includes instructions for carrying out aspects of the method(s) described herein. The computer program product is stored, at different stages, in any one of the memory 1006, the storage 1008 and / or a removable storage. Where a removeable storage is used, the removable storage is typically removable from the computer device 1000, such that the computer program product may be held separately from the computer device from time to time. Different computer program products, or different aspects of a single overall computer program product, may be present on multiple different computer devices.

[0154] Referring to Figure 2, there is shown a grooved medium 20 comprising a physical audio track 22 and a physical audio section 24.

[0155] The grooved medium 20 comprises one or more physical audio tracks 22; each physical audio track 22 comprises a groove comprising left and right walls, and a floor. The walls and floor are shaped such that when a stylus or similar implement traverses the groove, the stylus vibrates. The frequencies of vibration are dictated by the geometry of the groove such that music (comprising numerous frequencies) may be physically stored on a grooved medium by shaping the groove appropriately.

[0156] A grooved medium 20 will typically rotate in use, thus causing a stationary stylus (which has been placed into the groove) to traverse the groove. As it does so, the groove forces the stylus to oscillate at frequencies which are converted into an analogue electrical signal.

[0157] A grooved medium 20 may contain a plurality of physical audio tracks 22 on each surface each joined by a groove containing little or no audio information. This ensures that once a physical audio track 22 has finished being traversed by the stylus, the stylus is allowed to progress to the next physical audio track 22 and begins playing it automatically. Each surface of a grooved medium 20 therefore typically contains one continuous groove.

[0158] In the present disclosure however, it is beneficial to separate the groove into discrete physical audio tracks for processing and eventual playback. Furthermore, it is beneficial to separate each electronic or physical audio track 22 into smaller electronic or physical audio sections 24 to improve the storage of data and allow for a sufficient image resolution for the methods disclosed herein.

[0159] A start boundary mark 26a and end boundary mark 26b may therefore be used to separate a physical audio track 22 contained in a continuous groove from other physical audio tracks contained in the same continuous groove. The boundary marks 26 may be either an indentation, recess, or coating applied to the grooved medium 20 such that when the stylus traverses the mark it produces a transient noise pulse that may be identified above the electrical audio signal produced from the stylus traversing the groove. The marks are applied across an unmodulated or silent groove as close as possible to the start or end of the physical audio track or as close as possible to the start of the next physical audio track. This signal may be identified by its larger amplitude for instance.The boundary marks 26 may be formed by marking the grooved medium 20 with the blade of a sharp instrument such as a scalpel or by applying a coating. The boundary marks 26 are made perpendicular to the groove. The audio track 22 may be split into any suitable number of audio sections 24, however preferably each section is less than 0.6 seconds in length and further preferably less than one millisecond (ms). The duration of each audio section 24 or number of audio sections 24 may be left to the user to select using the computer device 1000. In a variation, only the beginning and end of the physical audio track 22 are identified using boundary marks 26 with the rest of the physical audio track 22 remaining unmarked. This can increase the quality of the audio output from the grooved medium 20 by preventing unnecessary noises during media playback.

[0160] Referring to Figure 3, there is shown a grooved medium 20 on a playing device 30 with a detachable stylus 32 and a detachable imaging apparatus 34.

[0161] The playing device 30 may be controlled by the computer device 1000 and is configured to rotate the grooved medium 20, thereby moving the detachable stylus 32 and the detachable imaging apparatus 34 along the groove on the grooved medium 20. The movement of the detachable stylus 32 relative to the groove on the grooved medium 20 is therefore approximately linear or parallel. The detachable stylus 32 reproduces the electrical signal contained in the groove as the physical audio track 22. The detachable imaging apparatus 34 may be controlled by the computer device 1000 and is arranged with a viewing window 36 that may be focused on a groove when the detachable stylus 32 is detached. The detachable imaging apparatus 34 is configured to capture images of the groove as the grooved medium 20 is rotated on the playing device 30. The detachable imaging apparatus 34 therefore produces images with each image containing a separate physical audio section 24 of the physical audio track 22. The images are in a digital format to enable further processing by the computer device 1000.

[0162] The detachable imaging apparatus 34 comprises a camera with a sufficiently high resolution and capture frame rate that it can produce clear images of the physical audio track 22 when the grooved medium 20 is rotated at a rotation speed determined and set by the computer device 1000.

[0163] The detachable imaging apparatus 34 may be configured to capture images at a set rate (for instance every 50ms or 20 Hertz (Hz) or more frequently every 1ms or IkiloHz) or may be configured to separate a physical audio track 22 into a set number of images. For example, the imaging apparatus 34 may be configured to produce N images per track thereby requiring an image every X ms for a three-minute audio track 22.

[0164] In a variation, the detachable imaging apparatus 34 is configured to capture an image of the physical audio track 22 which may later be separated into a plurality of images, each containing a physical audio section 24.

[0165] Alternatively, the detachable imaging apparatus 34 can continually record the surface of the grooved medium 20 and later the computer device 1000 segments the recording into individual images of physical audio sections 24. Using image processing techniques, the physical audio sections 24 are therefore identifiable and may be converted into electronic audio sections.If there is overlap of physical audio sections 24 shown in the images, there may be a need to remove a corresponding overlap of audio data in their corresponding electronic audio sections prior to processing of the audio sections.

[0166] Using two successive overlapping image sections (image one and image two) overlapping image data in image two may be identified using template matching which computes the similarity between a sliding window of image data from image one with image two using the sum of squared differences or the sum of absolute differences method.

[0167] An image overlap ratio may be computed as described previously. This ratio is multiplied by the number of digital audio samples in the electronic audio section corresponding to image 2 and the resulting number of digital audio samples are removed from the start of the electronic audio section before optional crossfading or presentation to the input of the digital to analogue converter.

[0168] In a variation, this ratio is multiplied by the duration of the electronic audio section being processed and the audio signal corresponding to the resulting duration is removed from the start of the electronic audio section before optional crossfading or presentation to the input of the digital to analogue converter.

[0169] In another example variation the section of overlapping or duplicated audio data in an electronic audio section is identified by cross-correlating it with the previously retrieved electronic audio section. Once identified, the redundant audio data is removed from the electronic audio section being processed before being passed to the input of the digital to analogue converter.

[0170] The computer device 1000 is then able to associate each image produced by the detachable imaging apparatus 34 with the electronic audio section corresponding to the physical audio section 24 contained within each image.

[0171] The start boundary mark 26a can be used by the computer device 1000 to identify the start of each audio track 22. This information can be used later by the computer device to associate the images and their corresponding electronic audio sections. As each boundary mark 26 will produce a transient noise pulse when the stylus traverses it, the computer device 1000 can identify the start and end of each physical audio track 22 by identifying these pulses. The boundary marks 26 may for example produce a brief electrical signal with a larger amplitude than the electronic audio signal making the boundary easy to identify.

[0172] The image may be separated into a plurality of smaller images by arbitrarily splitting the image into equal sized portions and examining the images for boundary marks 26. A mean square error method may be used to locate the boundary marks by comparing the difference between pixels in two corresponding images.

[0173] The electronic audio track or electronic audio sections produced by the detachable stylus 32 are passed through an analogue-to-digital converter to allow the data to be read by the computer device 1000 or further processed.Referring to Figure 4, there is shown a method 4 for digitising and storing electronic audio sections in a data structure. The method 4 is performed by the computer device 1000.

[0174] In a first step 40, a physical audio track 22 is recorded (thus becoming an electronic audio track) and images are taken of the physical audio track 22. This method is as described previously with reference to Figures 2 and 3.

[0175] In a second step 41 , the electronic audio track is separated into electronic audio sections. The method is as described with reference to Figures 2 and 3. As the physical audio track 22 was recorded as a single file, the electronic audio track may be separated into electronic audio sections before storing in a data structure.

[0176] At the end of the second step 41 , there are therefore digitised images, each showing a physical audio section 24 and corresponding digitised electronic audio sections.

[0177] In a third step 42, the images of the physical audio section 24 are hashed.

[0178] A hashing algorithm is performed by the computer device 1000 on each image of the physical audio sections 24. This operation may be performed for instance using the Difference Hash function available in the ImageHash Python coding library. Ideally a 28-bit hash algorithm is used to give a suitable number of possible values to reduce the possibility of collisions, although other number of bits may be used to achieve the same result.

[0179] The computer device 1000 applies the hashing algorithm in turn to each of the images. Each image may therefore be turned into a numerical value. The computer device 1000 also records the highest and lowest hash values generated after all the images have been hashed and computes the absolute difference between them as the range and stores this value.

[0180] In a fourth step 43, the electronic audio sections are associated with their hashed images.

[0181] Prior to hashing, each image contained a view of a physical audio section 24 of the physical audio track 22. In method step 41, a plurality of electronic audio sections were produced with each one corresponding to an image. In the fourth step 43, the computer device 1000 associates each electronic audio section with its corresponding hashed image. As the hashed images are simply numerical values, this means that, for each physical audio track, each corresponding electronic audio section is assigned a value.

[0182] In a fifth and final step 44, the electronic audio sections are stored in a data structure.

[0183] The computer device 1000 first computes the number of possible hash values for the chosen hash size and determines if it has or has access to sufficient free storage resources (memory and / or disk storage capacity) to create a data structure with the same number of storage locations as possible hash values. If it does not, then the computer device 1000 may suggest that the user increases available free storage resources or chooses a smaller hash size for image rehashing. If the computer device 1000 does have or has access to sufficient free storage resources to create a data structure with the same number of storage locations as possible hash values, then the computer device 1000 creates a data structure containing at least as many storage locations as the computed number of possible hash values for thechosen hash size. In this example given previously with a 28-bit hash algorithm, this would require a data structure with 228storage locations. In a variation, the computer device 1000 creates a data structure containing at least as many storage locations as the range of generated hash values computed in step 42. Advantageously, this may significantly minimise the size of the data structure needed to store the electronic audio sections.

[0184] The computer device 1000 may create or build any data structure disclosed herein using its internal and / or external memory and storage resources.

[0185] The electronic audio sections are then stored in the data structure in dependence on their associated values (assigned in step 43). For example, the electronic audio sections may be stored in the data structure at an index equal to their assigned value, or at their assigned value minus an offset (the offset for example being the value of the lowest generated hash - which may reduce the time needed to populate the data structure), however other mapping methods may be used as required. Advantageously, this may significantly reduce the time needed to populate the data structure. This produces a populated data structure containing electronic audio sections that, when suitably concatenated, form the electronic audio track corresponding to the original physical audio track 22.

[0186] Referring to Figure 5, there is shown a method 5 for retrieving stored electronic audio sections 24 from a data structure.

[0187] In a first step 50, the computer device 1000 receives hashed images of a grooved medium 20. As previously described in reference to Figure 4, hashed images of a grooved medium 20 are generated by applying a hashing algorithm to images containing a view of physical audio sections 24 of a physical audio track 22. The computer device 1000 receives these hashed images for instance from the memory 1006, an external device or from the detachable imaging apparatus 34 which produces images of the grooved medium 20 which are then hashed by the computer device 1000.

[0188] In a second step 51 , the computer device 1000 indexes a data structure in dependence on the hashed images.

[0189] The indexing values are derived from the hashed image data then used by the computer device 1000 to index the data structure. For example, the computer device 1000 may index the data structure to a storage index equal to the numerical value of the hashed image, or at their assigned value minus an offset (the offset for example being the value of the lowest generated hash - where this may reduce the storage space needed for the data structure), although other mapping may be used. Advantageously, this may significantly reduce the storage space needed for the data structure.

[0190] In a third step 52, the computer device 1000 extracts electronic audio sections stored in the data structure. The data structure is previously populated with electronic audio sections corresponding to at least one physical audio track 22. The hashed image corresponding to the indexing value that was used by the computer device 1000 to index the data structure is associated with the electronic audio signal stored in the data structure at the same index. The hashed image therefore acts as a key to extract the electronic audio section corresponding tothe physical audio section shown in the original image prior to the image being hashed.

[0191] Alternatively, the data structure may directly store the audio data at an index corresponding to the hash value corresponding to the hash of the image of the relevant section of the physical audio track. For example, the hash value may be a pointer pointing directly to the audio data. In a fourth step 53, the electronic audio sections are extracted by the computer device 1000 and concatenated to form a complete electronic audio track.

[0192] The complete electronic audio track may be converted to analogue audio and played through an audio amplifier connected to a plurality of loudspeakers or headphones.

[0193] The hashed images originally received in step 50 corresponded to a physical audio track 22 with each hashed image containing a physical audio section 24 of the track. The physical audio track 22 is therefore contained in its entirety in the hashed images. As the same hashing algorithm is used by the computer device 1000 to produce the indexing values, the computer device 1000 extracts the electronic audio section corresponding to the hashed image.

[0194] The hashed images are processed sequentially with the image containing the first physical audio section 24 of the physical audio track 22 processed first, then the next hashed image until the final hashed image containing the last physical audio section 24 of the physical audio track 22 is processed. This process may also be reversed (in which the last image is processed first). The data structure is therefore indexed each time and the electronic audio section stored at that index extracted. By performing this process sequentially and concatenating the electronic audio sections, a continuous electronic audio track is formed.

[0195] The continuous electronic audio track may be converted to analogue audio and played through an audio amplifier connected to a plurality of loudspeakers or headphones.

[0196] It will be appreciated that by storing electronic audio sections in a data structure indexed by hashes based on corresponding physical audio sections, data security may be improved in that the hashed images must be processed in the correct sequence in order for a correct audio track to be formed - where the hashed images will, in usual circumstances, be received in step 50 in the correct sequence (corresponding to the usual way in which the grooved medium is played, which leads to the images being produced in sequential order as previously described). Receiving hashed images corresponding to the correct grooved medium (including the physical audio track corresponding to the relevant stored audio data) may allow the relevant electronic audio track to be quickly extracted from the data structure (by way of the hashing functionality) and assembled, whereas without the correct hashed images (in the correct order) it may be difficult for a third party to access and retrieve the electronic audio sections (due to both the hashing and the fact that the audio track is disassembled into (scrambled and out of order) sections).

[0197] In an alternative, rather than being processed sequentially, the hashed images are processed in a different order according to instructions from the computer device 1000.

[0198] In a variation, the extracted electronic audio sections are stored in a separate data structure in the correct order corresponding to the physical audio track 22. The first electronic audio sectionis stored at index 0, the next electronic audio section at index 1 and so on until the final electronic audio section is stored. The separate data structure may then be accessed by the computer device 1000 in index order to produce the electronic audio track.

[0199] In a variation, each electronic audio section is extracted from the data structure and output without any intermediate storage.

[0200] In a fifth and final step 54, the electronic audio track is output from the computer device 1000. If the electronic audio sections have been stored in a separate data structure in index order, the computer device 1000 sequentially accesses the separate storage and outputs the electronic audio section stored in each index of the data structure. Crossfading may be applied as part of this process to smooth the transition between each output electronic audio section when forming the electronic audio track.

[0201] In a variation, the electronic audio track is stored first as a single file prior to being output. In this instance crossfading may be applied whilst concatenating the electronic audio sections to form the electronic audio track.

[0202] During the output process the electronic audio track is transmitted via a wired or wireless system. For example, the electronic audio track may be passed to a digital-to-analogue converter which converts the electronic audio track into an analogue signal which can be passed to an audio amplifier. This resultant amplified signal may then be played using a loudspeaker. Separately, the digital audio track may be transmitted by wires or wirelessly for instance via Bluetooth™ or Wi-Fi to an external device for further processing.

[0203] The method 5 therefore takes images of a grooved medium 20, hashes the images to determine indexing values, and traverses a data structure to retrieve data stored at the relevant index value. The method then extracts electronic audio sections from the data structure, concatenates the electronic audio sections to form an electronic audio track, and outputs the electronic audio track.

[0204] Figure 6 shows another method 6 for digitising and storing electronic audio sections in a data structure.

[0205] In a first step 60, a physical audio track 22 is recorded and images are taken of the physical audio track 22. This method 60 is as previously described method step 40.

[0206] In a second step 61 , the electronic audio track is separated into electronic audio sections. The method 61 is as previously described method step 41.

[0207] In a third step 62, the images of the physical audio track 22 or physical audio sections 24 are hashed. The method 62 is as previously described method step 42.

[0208] In a fourth step 63, the electronic audio sections are duplicated, and a hash algorithm applied to the one set of duplicated electronic audio sections.

[0209] The electronic audio sections are duplicated by the computer device 1000 or by a separatedevice configured to duplicate electrical signals. The electronic audio sections are initially digitised using an analogue-to-digital converter and then passed through the duplicating device. This produces two identical output electronic audio sections for each input electronic audio section.

[0210] This process is carried out sequentially such that as the duplicating device receives each input electronic audio section it duplicates the signal and passes one version to the hashing algorithm before duplicating the next input electronic audio section. In an alternative, this process is performed concurrently with a plurality of electronic audio sections being duplicated at one time.

[0211] One set of the duplicated electronic audio sections is then passed to a hashing algorithm. Preferably the hashing algorithm was designed for audio data and is applied by the computer device 1000. The hashing algorithm takes each input electronic audio section and hashes it, resulting in a hash value at the output of the hashing algorithm. This hash value is then associated with the corresponding electronic audio section output from the duplicating device that was not fed into the hashing algorithm. The computer device 1000 also records the highest and lowest hash values generated after all the electronic audio sections have been hashed and computes the absolute difference between them as the range and stores this value. In a fifth step 64, the set of electronic audio sections that were not hashed are stored in a first data structure in dependence on the hashed audio sections. The collection of electronic audio sections may be treated as a dataset, and the collection of images and / or image sections may also be treated as a dataset.

[0212] The computer device 1000 first creates a first data structure containing as many storage locations as there were possible hash values. In the examples given previously with a 28-bit hash algorithm, this would require a data structure with 228storage locations. In an alternative, the computer device 1000 creates a first data structure containing at least as many storage locations as the range of generated hash values computed in step 63 (so as to reduce the size of the data structure).

[0213] The minimum required hash length or size (in bits) required to size a data structure for a target dataset may be determined after deduplicating the dataset. For any such deduplicated dataset the base two logarithm of the smallest power of two that is greater than or equal to the number of sections in the dataset may be chosen as the minimum required hash length or size (in bits).

[0214] The electronic audio sections that were not hashed are then stored in the data structure in dependence on their associated hashed audio sections. For example, the electronic audio sections that were not hashed may be stored in the first data structure at an index equal to the hashed value of the same duplicated electronic audio section. If the first data structure was created using the range of generated hash values, the electronic audio sections that were not hashed are stored in the first data structure at an index equal to the hashed value of the same duplicated electronic audio section minus an offset (so as to reduce the time needed to populate the data structure). As an example, this offset could be the value of the lowest generated hash that was recorded in step 63. Advantageously, this may significantly reduce the size required for the data structure and / or reduce the time needed to populate the data structure. This produces a populated first data structure containing electronic audio sectionsstored at indexes related to the hashed values of the same electronic audio sections.

[0215] In a sixth and final step 65, the hashed audio sections are stored in a second data structure in dependence on the hashed images.

[0216] The computer device 1000 creates a second data structure containing at least as many storage locations as there are hashed images. In an alternative, the computer device 1000 creates a second data structure containing at least as many storage locations as the range of generated hash values computed in step 62. Advantageously this may significantly minimise the size of the data structure needed to store the electronic audio sections.

[0217] Each hashed image produced at step 62 is associated with each hashed electronic audio section. Each image that was originally used to create each hashed image contained a view of a physical audio section 24. The physical audio sections 24 have in previous steps of the method 6 been digitised, duplicated, and hashed therefore each hashed image has a corresponding hashed audio section.

[0218] Each hashed electronic audio section is therefore stored in the second data structure using the corresponding hashed image. For example, each hashed electronic audio section may be stored at an index in the second data structure equal to the value of the corresponding hashed image. If the second data structure was created using the range of generated hash values, each hashed electronic audio section is stored at an index in the second data structure equal to the value of the corresponding hashed image minus an offset (so as to reduce the time needed to populate the data structure). As an example, this offset could be the value of the lowest generated image hash that was recorded in step 62. Advantageously, this may significantly reduce the time needed to populate the data structure.

[0219] Flow of information

[0220] Figure 7 shows the flow of information as described previously with reference to method 4.

[0221] Images taken of the physical audio section 24 are transmitted 71 to a hashing algorithm 73. Meanwhile, the physical audio sections 24 are played back and turned into analogue electronic audio sections, and then transmitted 72 to an Analogue-to-Digital Converter 74. The Analogue-to-Digital Converter 74 converts the analogue electronic audio sections into digital electronic audio sections.

[0222] The digital electronic audio sections are transmitted 76 to a data structure 77 containing N (a positive integer) storage locations (where N is the last storage location). The hashed images are also transmitted 75 to the data structure 77 and the digital audio sections are stored in the data structure 77 in dependence on the hashed images. As previously described with reference to figure 4, the hashed images are simply values which may be used as index values for storing the digital audio sections.

[0223] The index value for each data structure 77 location is denoted by numeral 78. In the example shown in Figure 7, the image being hashed by the hashing algorithm 73 has produced a hash value of two, so the corresponding digital audio section is stored at index two in the data structure 77.Transmitted is used herein to refer to the passing of data, information, or electrical signals between two entities. This process may occur for example within a computer device, via a wired connection, or via a wireless connection.

[0224] Figure 8 shows the flow of information as described previously with reference to method 6. The flow of information is as described with reference to Figure 7 with the additional information flows:

[0225] Once produced by the Analogue-to-Digital Converter 74, the digital audio sections are duplicated with one signal fed into an additional hashing algorithm 82 and the other stored in a second data structure 86.

[0226] The hashed digital audio sections are transmitted 84 to the second data structure 86, and the digital audio section from the Analogue-to-Digital Converter 74 are stored in the second data structure 86 in dependence on the hashed digital audio sections. The figure shows the data structure 86 having exemplary index values 88 for each data storage position.

[0227] The hashed digital audio sections are also transmitted 84 to the first data structure 77 and stored in dependence on the hashed images transmitted 75 from the hashing algorithm 73.

[0228] Figure 9 shows duplicated views of the same pair of successive overlapping images of physical audio sections 24; image 1 and image 2, where image 1 was captured before image 2. The methods disclosed herein identify the respective non-overlap areas 90, 92 and overlap area 91 associated with the two images. The methods disclosed herein also identify image dimensions that may be used to compute an overlap ratio for image 2. These include overlap length 93, non-overlap length 94 and their sum; image length 95.

[0229] An overlap area 91 may be identified in an image of a physical audio section 24 (for instance by using a template matching process as previously described), and the overlap area 91 and the non-overlap area 92 may be used to compute an image overlap ratio which herein may be defined as the overlap area 91 divided by: the overlap area 91 plus the non-overlap area 92, or alternatively the image overlap ratio may be calculated using the overlap length 93 and the image length 95, this calculation defined as the overlap length 93 divided by the image length 95. The image overlap ratio may be used to compute the duration or number of overlapping audio data samples in a corresponding electronic audio section (by multiplying the duration or number of audio data samples in an electronic audio section by the image overlap ratio) and removing the computed duration or number of overlapping audio data samples (segment of overlapping audio data) from the appropriate position in the electronic audio section.

[0230] Auxiliary Data

[0231] Whilst auxiliary data has been previously described, this section aims to provide more specific information on how auxiliary data may be generated and used.

[0232] As used herein, application may refer to an entity performed on a computer device carrying out the methods described herein. For instance, an application may provide a user-interface for performing the method steps described herein. As such, if the application performs stepsthese are to be considered an extension of the methods described herein.

[0233] Auxiliary data may be derived by the computer device 1000 during the processing of the grooved medium 20, read directly from the grooved medium 20, or presented to the computer device by the user by means of manual data entry, a list or computer script or another device. The auxiliary data provides additional information or instructions to the computer device 1000. The auxiliary data is in the form of Boolean bits, single or multibyte binary codes, decimal codes, binary coded decimal codes, characters, or hexadecimal codes. Auxiliary data in this form is described herein as information or instruction codes.

[0234] The computer device 1000 can create a data structure to store the auxiliary data (an auxiliary data structure) and in an embodiment the auxiliary data structure is at least the same size as the data structure used for storing digital electronic audio sections and their indexes are derived from the same hash function. The auxiliary data is read, derived, and / or stored for each physical audio section, thereby generating a host of additional data for an entire audio track. These data structures are described herein as control data structures.

[0235] The auxiliary data can be used to alter the method carried out by the computer device 1000 for example by changing the order in which method steps are carried out, causing the application to change the type of hash function being used, start or stop using an additional hash function, switch to using a different digitised audio data structure, start or stop using an additional digitised audio data structure, start or stop a timer, get elapsed time, get current time, perform disc rotation direction change detection, provide information to the user, turn on or off status indicators, present statistical data on genre information codes, perform groove modulation classification and / or track genre classification, perform waveform generation and provide synchronizing or triggering signals for internal use or external equipment usage. Whilst a general overview of the types of auxiliary data have been given previously, examples of Boolean or numerical auxiliary data as could be used by the application are as follows:

[0236] Information Code Value - Meaning

[0237] 0 - Mono groove left plus right in-phase modulation

[0238] 1 - Stereo groove left plus right anti-phase modulation

[0239] 2 - Left groove wall modulation

[0240] 3 - Right groove wall modulation

[0241] 4 - Quadraphonic disc

[0242] 5 - Micro groove disc

[0243] 6 - Groove radius

[0244] 7 - Groove number

[0245] 8 - Cylinder

[0246] 9 - Hash function type

[0247] 10 - Genre Jazz

[0248] 11 - Genre Blues

[0249] 12 - Genre Classical

[0250] 13 - Disc speed 33 rpm

[0251] 14 - Disc speed 45 rpm

[0252] 15 - Disc speed 78 rpm

[0253] 16 - T rack start17 - T rack end

[0254] 18 - Mono silent groove

[0255] 19 - Stereo silent groove

[0256] 20 - Image row dimension

[0257] 21 - Image column dimension

[0258] 22 - Randomly generated number

[0259] 23 - Pseudo randomly generated number

[0260] 24 - Colour

[0261] 25 - Time

[0262] Instruction Code Value - Instruction

[0263] 26 - Change hash function

[0264] 27 - Add hash function

[0265] 28 - Select audio data structure type 1

[0266] 29 - Select audio data structure type 2

[0267] 30 - Change audio data structure

[0268] 31 - Add audio data structure type 3

[0269] 32 - Get current time

[0270] 33 - Start a timer

[0271] 34 - Stop a timer

[0272] 35 - Get elapsed time

[0273] 36 - Mute audio

[0274] 37 - Generate audible beep

[0275] 38 - Waveform generation on - Set a digital output port high

[0276] 39 - Waveform generation on - Set a digital output port low

[0277] Each of these data are associated with a unique value, such that the application is able to use the transmission or reception of the values to determine auxiliary data.

[0278] As the auxiliary data is produced for each physical audio section, it is possible to display or further derive information about an entire audio track; for instance, how the groove radius changes throughout the track or whether a change in rotation speed has occurred.

[0279] When performing groove modulation classification, the application may keep and update a counter for each of the different groove modulation information codes retrieved from a control data structure during a period or for the entire audio track duration and compare them. These counters are called basic counters.

[0280] In a variant the groove modulation information code whose basic counter has the largest value may be taken as the groove modulation classification for that physical audio track. Alternatively, the groove modulation information code whose basic counter passes a predetermined threshold value and has the largest value may be taken as the groove modulation classification for that physical audio track.

[0281] The application may also keep and update a counter of the uninterrupted sequence length for each of the different groove modulation information codes retrieved from a control data structure in a period or for the entire audio track duration and compare them. These countersare called sequence length counters.

[0282] If the basic counter value of two or more different groove modulation information codes Is the same, then the groove modulation information code whose sequence length counter has the largest value may be taken as the groove modulation classification for that physical audio track. This analysis may be performed for determining the overall value of other information codes at the level of the audio track (rather than for each audio section), for instance track genre classification.

[0283] The results of this analysis may be displayed on an electronic display in the form of bargraphs. As auxiliary data may also contain instruction codes, characteristics relating to the grooved media may be changed throughout the playback of a track. For example, the audio balance may be altered throughout the track, or an audible beep generated at a pre-determined time. The auxiliary data can be output to a user on a visual display or stored to provide additional information about the grooved medium 20.

[0284] During the playback process, the computer device 1000 reads the stored auxiliary data from the auxiliary data structure using the hashes of the images of the physical audio sections 24. The computer device 1000 can compare the auxiliary data values at each timestep to determine facts relating to the grooved medium or to provide further information to the user; for instance by temporarily storing auxiliary data values from previous timesteps in a storage area and then comparing the auxiliary data at the current reading timestep with stored readings from previous timesteps, the computer device 1000 can determine whether the grooved medium 20 is rotating or if the direction of rotation has changed (for instance if no change in any auxiliary data occurs). To acknowledge this condition the computer device 1000 changes the value of a rotation detection or direction change Boolean variable. To inform a user, the computer device 1000 may further cause a status indicator such as a light emitting diode to turn on or off. The computer device 1000 may also take an additional pre-determined course of action.

[0285] Auxiliary data or instruction codes may be used to create custom electronic waveforms. In one variation a user identifies part of an audio track that may require modification. An image hash code or the combined hash codes of a sequence of images that correspond to the start of this part of the audio track has one or multiple waveform generation instruction codes associated with it. When these hash codes are processed the associated waveform generation instruction codes are also processed and cause GPIO (general purpose input output) ports to be turned on or off. This feature may be used to enable or disable electronic circuits which modify the audio signals according to some target characteristic and / or provide synchronising or triggering signals for internal and / or external use.

[0286] In another variation the application initialises a Boolean variable to zero or logical false which indicates that no motion or rotation change has been detected and stores at least one previously generated hashed image in a data structure.

[0287] In another variation the data structure is an array, and the previously generated hashed imageis stored at index zero.

[0288] The latest hashed image is compared with the previous one stored in the array at index zero.

[0289] The comparison may be made using a distance metric such as Hamming Distance. If the result of the comparison is below a threshold the two hash codes are considered equal. If the result of the comparison is above a threshold the two hash codes are considered not equal.

[0290] If they are not considered equal the application assumes that media motion or rotation has occurred. The application acknowledges this condition by setting the Boolean variable to ONE or TRUE. It also overwrites the previously generated hash code at array index zero with the latest groove image hash code. It may also cause a status indicator such as a light emitting diode to turn on. The application may also take an additional pre-determined course of action.

[0291] If the latest hashed image is considered equal to the previous one stored in the array at index zero and other auxiliary data do not indicate a silent groove the application assumes that no media motion or rotation has occurred. The application acknowledges this condition by setting the Boolean variable to ZERO or FALSE. It may also cause a status indicator such as a light emitting diode to turn off. The application may also take an additional pre-determined course of action.

[0292]

[0293] When imaging or playing back phonograph cylinders the playback device may be based on a modified cylinder phonograph.

[0294] The playback device uses a motor driven mandrel over which the cylinder is slid and such that the cylinder and mandrel rotate in the same direction, rotate simultaneously and share the same longitudinal axis of rotation. The motor driving the mandrel may be controlled by the computer device 1000. Illumination is provided to illuminate the relevant cylinder grooves inscribed in the exterior surface of the cylinder that will be imaged or played. The imaging device may be attached to a motor driven and electronically controlled arm or carriage in a manner that prevents unintentional physical contact with the cylinder or damage to it. The said arm or carriage is controlled by the computer device 1000 and is configured to move the arm or carriage with the imaging device along the length of the cylinder while keeping the imaging device at a constant distance from the cylinder exterior. This helps to keep the image produced by the imaging device in focus.

[0295] At least one optical displacement sensor may be incorporated in the body housing the imaging device or elsewhere on the arm or carriage to monitor any change in distance between the imaging device and cylinder. The at least one sensor may be interfaced with the computer device 1000 which uses signals or data from the at least one sensor to adjust a position of the arm or carriage to maintain a constant distance between the imaging device and cylinder.

[0296] As with the playback device described for the phonograph discs, the imaging device may be detachable from the arm or carriage so that a conventional phonograph cartridge capable of playing cylinders may be fitted in its place.Advantageously this arrangement allows for the capture of cylinder groove images using the imaging device or capture of audio from cylinder grooves using a phonograph cartridge or the non-contact imaging device.

[0297] Media rotation direction reversal detection or media scan direction reversal detection using groove image hash codes

[0298] A method of rotation detection is provided, which works by using the application to compare hashed image seguences comprised of single hashed images from past time steps with current hashed images to determine if the direction of media rotation has changed.

[0299] The term control array is used herein to refer to a data structure containing data that may cause the application to take an action. Examples include changing the audio characteristics of a playback device (for instance the amount of treble or the left-right balance) or changing the type of hash used in analysing the grooved media.

[0300] In another variation the application uses a first hash function f1 and a second hash function f2, a first array and a second array. The second hash function (f2) generates a hash (h2) of an electronic audio section and the first hash function (f1) generates a hash (hi) from the image of its corresponding physical audio section. The hash (h2) of the electronic audio section is stored in the first array at the index given by the hash (hi) of the image of its corresponding physical audio section. The electronic audio section is stored in the second array at the index egual to the hash (h2).

[0301] Two hash functions are reguired only for populating the arrays. At playback time only one hash function is needed. This is the first hash function (f1) and it is used to generate hashes from input images of physical audio sections to access the first array and retrieve the index of the corresponding electronic audio sections stored in the second array for retrieval and playback. Compact disc (CD) technology uses another track on the disc that stores information such as the number of tracks on the disc, their timings and location on the disc. This is known as the Table of Contents and enables a CD player to move guickly to a selected track.

[0302] Some of this functionality may be provided thereby enabling users to enter and store details within the application about individual phonograph album discs such as user preferences and / or disc or track playback characteristics such as RIAA or NAB egualization reguirements, groove modulation type, the number of tracks, track numbers, track titles, track durations, track playback order and other metadata associated with the disc.

[0303] These details and / or settings may be stored in the application settings database and / or a file. The file may be exported from the application and distributed and shared among application users so that other application users may import them and update their internal application settings databases.

[0304] Non limiting examples of suitable file formats include text and comma separated values files. Stored settings are used to instruct the application to take specific actions at specific time points for a specified duration. For example, the user settings may instruct the application tochange to a different data structure at a specific time for a specified duration and for certain tracks.

[0305] These details and stored settings may be recalled by the application and applied under user control or automatically when the target disc or grooved media object is detected by the application or selected by the user for playback.

[0306] In a variation the user may capture a digital picture of the disc itself, album cover, sleeve, insert or disc labels and present the image to the application which hashes the image data, stores the hashes in a database. The user may configure the application to associate the hashes with the stored details and settings for that album disc or grooved media object.

[0307] If the user subsequently presents a digital image of the album cover, sleeve, insert or disc labels to the application the application again hashes the image and searches its internal database for a matching hash using the Hamming Distance metric for the comparison. If the application finds a matching hash it assumes that a target disc or grooved media object has been detected and retrieves the stored application settings associated with the hash and applies them.

[0308] For example, stored album track information may be displayed on a visual display unit. The user may also enter and store audio parameters that the application may use to modify electronic audio sections before or after they are crossfaded and joined.

[0309] Non limiting examples of such audio parameters may include volume, bass, treble, gain, frequency, and phase.

[0310] Such parameters may be used by the application to modify entire audio tracks or parts of audio tracks at specific times for specific durations. This feature allows audio tracks to be dynamically modified or remastered.

[0311] Multiple sequences of hashes or combined hashes of images of physical audio sections may also be entered and stored and used as reference points to identify multiple time points or multiple points along a groove or audio track. The generation or computation of these hashes may be performed or automated by the application.

[0312] These reference points may be used to identify the side of a phonograph disc for example Side 1 or Side 2 or Side A or Side B so that the application loads the correct stored data.

[0313] Time periods and / or offset times relative to another position such as the start or end of an audio track or reference points in different grooves may also be computed by the application, stored, and associated with the reference points.

[0314] These reference points may have media playback speeds associated with them. This means that a particular reference point is expected to be detected by the application at particular times depending on the speed of media playback.

[0315] When a reference point is detected by the application during media playback the time of detection may be compared by the application with time data stored for that reference point. Ifthe time difference is equal to a predetermined value or within a predetermined range the media is assumed to be playing at a speed associated with that signature. The terms ‘signature(s)’ and ‘reference point(s)’ are used interchangeably herein.

[0316] The application may create and maintain multiple data structures such as linked lists which store data mapping reference points to other reference points .

[0317] Other facts or data may also be associated with reference points , for example groove number, groove radius and groove modulation type. This information may be used to check or confirm the accuracy of information codes stored in the application databases or maybe used internally by the application and / or presented to the user.

[0318] In a variation, electronic audio sections are concatenated or joined together at a zero crossing.

[0319] A zero crossing is a point where an electronic signal crosses the zero-amplitude level axis. In other words, a zero crossing is where an electronic signal (such as an audio signal) changes from positive to negative or vice versa.

[0320] When processing non-overlapping electronic audio sections, the computer device 1000 may store information related to zero crossing points within it in a control array or auxiliary data structure. This information may include the positions of the zero crossing points within the electronic audio section and algebraic signs of the audio signal as it enters or exits the zero crossing points.

[0321] The computer device 1000 may later use this information to choose which audio data samples are removed from the electronic audio section prior to concatenation or joining with other electronic audio sections.

[0322] The computer device 1000 may also remove audio data samples from either end or both ends of an electronic audio section so that the first and last audio data samples in the resulting cropped electronic audio section are zero crossing points.

[0323] The computer device 1000 may perform this step on overlapping electronic audio sections after overlapping data has been removed.

[0324] Advantageously, this method of processing and joining electronic audio sections may eliminate or significantly minimize popping or clicking sounds that may occur due to signal discontinuities that may be formed when audio signals are concatenated or joined together without modification and provides an effective alternative to crossfading.

[0325] For correct sound reproduction of a grooved medium using this invention, relevant audio data stored in the data structure is retrieved and then played back in a particular order. The order is naturally determined by the image or optical pattern of the groove or physical audio track being played back. However, the computer device 1000 may use a data structure to map a numerical position of a groove segment within a groove or physical audio track to a location in the data structure storing audio data. This data structure may be a one-dimensional array or lookup table. Herein these data structures are called ‘virtual grooves’. In this arrangement anarray index is treated as the numerical position of a groove segment within a groove or physical audio track and each element is a location in the data structure storing audio data.

[0326] By using virtual grooves, the computer device 1000 may optionally change the order of play back of retrieved audio data. This would enable one grooved medium to impersonate another or to virtually replace specific tracks on the grooved medium.

[0327] Virtual grooves may be constructed for specific grooved media such as phonograph or vinyl records using computer software executed on the computer device 1000 or another computer and stored as files on electronic storage devices such as memory sticks or hard disk drives. They may also be shared, streamed or distributed over computer or radio networks.

[0328] The user may obtain a virtual groove as a file stored on a memory stick and proceed to plug it into a Universal Serial Bus (USB) port on the computer device 1000 or download the virtual groove onto the computer device 1000 using a network interface on the device that was previously described. The computer device 1000 may then copy the virtual groove into an internal memory at which point it is ready to use by the central processor unit in the computer device 1000.

[0329] The computer device 1000 may maintain a counter of segments of a physical audio track detected since the start of playback and increment the counter as each new segment is detected. For each new segment detected the computer device 1000 may use the current value of the counter as an index into the virtual groove to retrieve the location in the data structure of the substitute audio data.

[0330] Alternatively, the computer device 1000 may access and ‘play’ the virtual groove independently of a physical grooved medium allowing the computer device 1000 to function as a media player. For example, the computer device 1000 may be configured as a streaming network client that receives virtual grooves or hash streams from a network server

[0331] The method described herein may enable a physical audio track or groove on a grooved medium to be represented by a list or sequence of hashes or as a data structure in a memory device or data storage device storing such a list or sequence of hashes. The list or sequence of hashes may be stored in a data structure or a file and / or transmitted or streamed over a computer network or radio network from a computer device 1000 configured as a server to a plurality of computer devices 1000 configured as clients. The data structure may be an array storing the sequence of hashes.

[0332] If this array is used in conjunction with a data structure populated with electronic audio sections, an index position of the array of hashes may be viewed as representing the stylus of a phono cartridge and the array of hashes may be viewed as representing a physical audio track.

[0333] Incrementing or decrementing the index position may be viewed as forward or reverse playing the physical audio track.

[0334] The hash stored at the indexed location of the array may be used to index the data structure populated with electronic audio sections. This data structure may be viewed as representingthe generator inside a phono cartridge whose electrical output has been digitised.

[0335] The term ‘indexing’ is used herein to refer to the process of locating a particular address or location in a data structure. For example, a data structure may be able to store eight pieces of data with the location of each piece of data assigned to a respective position, each position having a number from 0-7, wherein ‘0’ corresponds to the first location in the data structure and 7’ refers to the final location.

[0336] In a variation the hash functions may be based on or implemented in neural networks, and / or a neural network may be used to supplement the methods described herein. In various examples described herein, the neural network is an artificial neural network. However, other types of neural network may be used in some variations.

[0337] In another variation a neural network may be configured and trained to predict a target hash (centre hash or non-centre hash) given a fixed-size window of context hashes surrounding it. The neural network may be a language model.

[0338] The training data for the neural network may consist of hashes of images of physical audio sections.

[0339] The language model may be a masked language model that has been adapted to treat sequences of hashes of images of physical audio sections in similar a manner to if it were text. For example, masked language models are normally trained on text to process text. These models are trained on datasets of text where certain words are purposely concealed during training. The goal of the model is to guess the concealed word based on the surrounding context. In this invention an adapted masked language model may be implemented on or in the computer device 1000 and trained on a dataset of hashes of images of groove segments or hashes of images of physical audio sections where certain hashes are purposely concealed during training. The goal of the model is to guess the concealed hash based on the surrounding context.

[0340] Preferably, the hashes in the dataset are correctly ordered or sequenced as would be required for correct audio playback if they were being used to access electronic audio sections stored in a data structure.

[0341] The neural network may be used to ensure that anomalies or imperfections in the physical audio sections shown in the images (such as groove damage, dust, or microscopic debris) do not result in electronic audio sections or auxiliary data being stored in or retrieved from the incorrect data structure locations by the computer device 1000. The computer device 1000 may substitute hashes predicted by the neural network for hashes generated by the hash function.

[0342] Alternatively, the computer device 1000 may first compare predicted hashes with generated hashes using a hash comparison method, for example Hamming Distance. If the result is above a threshold the computer device 1000 may substitute hashes predicted by the neural network for hashes generated by the hash function and store or retrieve electronic audiosections or auxiliary data in or from the correct data structure locations.

[0343] In another variation the computer device 1000 may treat hashes and / or identifying values as image data or convert them to a form that may be treated as image data. For example, a 24-bit hash may be treated as a RGB (Red-Green-Blue) or BGR (Blue-Green-Red) triplet by separating it into three 8-bit numbers each representing the intensity of a colour channel. For another example, a 22-bit hash may have two extra bits appended to it to form a 24-bit identifying value that may be treated as a RGB or BGR triplet as previously described.

[0344] The computer device 1000 may process this image data and / or transfer the image data to an electronic display or transmit the image data over a network. Another computer device 1000 may receive the transmitted image data and process the image data and / or transfer the image data to an electronic display. The receiving computer device 1000 may also treat the received image data as image hashes and / or identifying values and use them to access electronic audio sections stored in a data structure as previously described.

[0345] Advantageously this method allows hashes to be visualised, and groove segments to be represented on an electronic display by colours in a consistent way. The method creates an alternative and visual representation of a groove that may indicate 1) that groove playback is occurring, 2) that the media is rotating and 3) the direction of media rotation due to displayed changing colours on the electronic display. When the media is stationary the colours shown on the electronic display may be static. When the media is rotating the colours shown on the electronic display may be changing due to the changing received data.

[0346] It will be understood that the present invention has been described above purely by way of example, and modifications of detail can be made within the scope of the invention.

[0347] Each feature disclosed in the description, and (where appropriate) the claims and drawings may be provided independently or in any appropriate combination.

[0348] Reference numerals appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims.

Claims

Claims1. A method of storing information relating to at least one grooved media, the method comprising;receiving at least one image relating to a physical audio track on a grooved medium;generating a hash in dependence on the at least one image to produce at least one identifying value; andstoring at least one electronic audio section in a data structure in dependence on the at least one identifying value.

2. A method as claimed in claim 1, further comprising receiving an electronic audio track corresponding to a physical audio track on a grooved medium, the electronic audio track comprising a plurality of electronic audio sections, preferably wherein receiving the electronic audio track comprises capturing the electronic audio track.

3. A method as claimed in claims 1 or 2, wherein the at least one image relates to at least one section of the audio track.

4. A method as claimed in any preceding claim, further comprising associating an identifying value with an electronic audio section corresponding to an audio output of a physical audio section contained in the image that was used to produce the identifying value.

5. A method as claimed in any preceding claim, wherein the electronic audio section is stored at a location in the data structure in dependence on the identifying value.

6. A method as claimed in claim 5, wherein the electronic audio section is stored at an index of the data structure that is equal to the identifying value.

7. A method as claimed in any preceding claim, wherein the grooved medium comprises at least two physical marks used to identify the start and / or end of a physical audio track.

8. A method as claimed in claim 7, further comprising identifying the at least two physical marks using image processing techniques thereby identifying the start and / or end of a physical audio track.

9. A method as claimed in claims 7 or 8, further comprising identifying the start or end of an electronic audio track or section by detecting in the electronic audio section a signal corresponding to the physical mark.

10. A method as claimed in any preceding claim, wherein at least one electronic audio section is an analogue signal, wherein the method further comprises digitising the electronic audio section.

11. A method as claimed in any preceding claim, further comprising storing auxiliary data related to the grooved medium.

12. A method as claimed in claim 11, wherein the auxiliary data comprises any of the following: a groove type; a hash function type; a genre; a disc speed; a track start code; a track end code; a time value; a silent groove code; an image dimension; a groove modulation type; a colour; a groove radius; a groove number; a hash or identifying value; a randomly generated number; and a pseudo-randomly generated number.

13. A method as claimed in claims 11 or 12, wherein the auxiliary data comprises instructions to do any of the following: change a hash function to be applied; select a data structure; get current time; start a timer; stop a timer; get elapsed time; mute audio; generate an audio tone; modify an audio signal and set a digital output port to high or low.

14. A method as claimed in any of claims 11 to 13, further comprising associating each at least one electronic audio section and / or each at least one identifying value with auxiliary data.

15. A method as claimed in any preceding claim, further comprising storing, in a second data structure, locations of the at least one electronic audio section in the data structure.

16. A method as claimed in any preceding claim, further comprising applying an additional hashing process to at least one of: the at least one identifying value, and the electronic audio sections.

17. A method as claimed in any preceding claim, wherein the electronic audio track corresponds to the physical audio track.

18. A method of retrieving an electronic audio section stored in a data structure, the method, comprising:receiving at least one identifying value produced by generating a hash in dependence on at least one image of a grooved medium;indexing the data structure in dependence on the at least one identifying value; andextracting the electronic audio section from the data structure at the same index.

19. A method as claimed in claim 18, wherein each image used to generate the at least one identifying value shows a physical audio track section corresponding to an electronic audio signal stored in the data structure.

20. A method as claimed in claim 18 or 19, further comprising extracting a plurality ofelectronic audio sections and combining the plurality of electronic audio sections to form an electronic audio track.

21. A method as claimed in any of claims 18 to 20, further comprising indexing the data structure according to a stored electronic audio section corresponding to the audio output of the physical audio section contained in the image that was used to produce the identifying value.

22. A method as claimed in any of claims 18 to 21 , further comprising: extracting auxiliary data associated with the physical audio section from an auxiliary data structure.

23. A method as claimed in claim 22, further comprising performing, in response to receiving the auxiliary data from the auxiliary data structure, one or a combination of the following actions: changing a hash function to be applied to further images of the grooved medium; selecting a data structure; get current time; starting a timer; stopping a timer; getting elapsed time; muting audio; generating an audio tone; modifying the audio signal; and setting a digital output port to high or low.

24. A method as claimed in claims 22 or 23, further comprising determining, in response to the auxiliary data from the data structure, a genre of the audio track.

25. A method as claimed in claims 22 to 24, further comprising counting auxiliary data associated with a plurality of physical audio sections or a plurality of identifying values thereby to determine further data in respect of the physical audio track; preferably wherein the further data comprises a groove type; a genre; a groove modulation type; and a disc colour.

26. A method as claimed in any of claims 22 to 25, further comprising modifying an electronic audio section based on the auxiliary data.

27. A method as claimed in any preceding claim, wherein generating a hash in dependence on the at least one image comprises: applying a local binary pattern feature extraction algorithm to the image, producing a representation of the image, and producing a hash that is based at least in part on the representation of the image.

28. A method as claimed in any preceding claim, wherein the at least one image is a plurality of images.

29. A method as claimed in claim 28, further comprising determining an image overlap area between two adjacent images.

30. A method as claimed in claim 29, wherein the image overlap area is determined using a template matching process.

31. A method as claimed in claim 29 or 30, further comprising calculating an image overlap ratio in dependence on the image overlap area and the non-overlap area of the image.

32. A method as claimed in claim 31, further comprising removing a portion of the at least one electronic audio section in dependence on the image overlap ratio.

33. A method as claimed in any of claims 28 to 32, wherein generating a hash in dependence on the plurality of images comprises using a locality-sensitive hashing method; preferably wherein the locality-sensitive hashing method comprises: hashing each image with a perceptual image hash function to produce a plurality of hash digests, each hash digest comprising a plurality of positions, each position in each hash digest comprising a binary digit; for each position, summing the binary digit at the position across all the hash digests to produce a position sum for each position; computing the average value of the position sums; and setting a result hash, wherein for each position in the result hash, when the position sum is greater than the average value of the position sums, then the position is set to 1, and when the position sum is not greater than the average value of the position sums, then the position is set to 0.

34. A method as claimed in any preceding claim, further comprising producing a plurality of identifying values from a single image of the grooved medium.

35. A method as claimed in any preceding claim, further comprising: separating the hash into one or more pixel triplets.

36. A method as claimed in claim 35, further comprising appending bits to the hash to form the one or more triplets.

37. A method as claimed in any of claims 35 or 36, wherein the one or more triplets are communicated and / or displayed on an electronic display.

38. A method as claimed in any of claims 18 to 37, comprising playing back the extracted electronic audio section.

39. A method of detecting, in a grooved medium, a change of the state of a groove using image processing techniques.

40. A method according to claim 39, wherein using image processing techniques comprises using an edge detection method.

41. A method according to claim 39 or 40, comprising:determining a difference between pixels in a first window along a groove and pixels in a second window along a groove; andwhen the difference is equal to or above a threshold, detecting a change of the groove state.

42. A method according to claim 41 , wherein determining a difference comprises using one of: a mean squared error function; a root mean square error function; a structural similarity index measure; a sum of absolute differences function; a sum of squareddifferences function; a Hamming distance associated with hashes of the first window and the second window; a Manhattan distance associated with hashes of the first window and the second window; and a Euclidean distance associated with hashes of the first window and the second window.

43. A method according to claim 41 or 42, further comprising translating the first and second window along the groove.

44. A method for detecting movement of a grooved medium depicted in a plurality of images, the method comprising:receiving a plurality of images of the grooved medium; anddetermining whether a first property associated with a first image of the plurality of images differs from a second property associated with a second image of the plurality of images thereby to detect movement of the grooved medium.

45. A method according to claim 44, wherein the first and second properties are first and second hashes generated in dependence on respective first and second images of the plurality of images.

46. A method according to claim 44 or 45, wherein the detected movement relates to the spin of the grooved medium; preferably the presence or absence of spinning and / or a reversal in direction of the spinning.

47. A method of storing information relating to grooved media, the method comprising;receiving a plurality of images relating to a physical audio track on a grooved medium;generating a hash in dependence on each image to produce a plurality of identifying values; andstoring the plurality of identifying values in a data structure.

48. A method as claimed in any preceding claim, wherein the grooved medium is a vinyl record.

49. A method as claimed in any of claims 1 to 47, wherein the grooved medium is an Edison cylinder.

50. A computer device arranged to implement a method as claimed in any preceding claim.

51. A computer program product including instructions for carrying out a method as claimed in any preceding claim.