System and method for recording and verifying voice clones, photo and video clones, and / or copies of text

The system addresses security and verification challenges in voice, photo, and text data management by using cryptographic hashing and watermarking techniques, ensuring secure storage and reliable verification.

WO2026072395A1PCT designated stage Publication Date: 2026-04-02BLINDER INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing systems for managing and licensing voice, photo, and text data lack robust security measures, transparent licensing processes, and reliable verification methods, leading to concerns about intellectual property rights and potential misuse.

Method used

A system involving multiple devices and distributed networks with cryptographic hashing and watermarking techniques to securely store and verify the authenticity of voice, photo, and text data, using methods like noise reduction, normalization, segmentation, and embedding unique characteristics to create secure and reliable verification processes.

Benefits of technology

Provides secure storage and reliable verification of voice, photo, and text data, ensuring authenticity and authorized use while protecting intellectual property rights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025046679_02042026_PF_FP_ABST
    Figure US2025046679_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Recording and verifying a voice clone includes submitting a voice recording in a binary code record to a device, conducting a pre-cryptographic processing on the binary code record to produce a voice clone record, cryptographically hashing the binary code record and the voice clone record based on a cryptographic algorithm, submitting the binary code record and the voice clone record to a distributed network, generating storage and access control record(s) for the voice clone record, cryptographically hashing the storage and access control records based on a cryptographic algorithm, generating watermark record(s) and associating them with the voice clone record, cryptographically hashing the watermark record(s) based on a cryptographic algorithm, receiving an external voice clone record, evaluating the external voice clone record for presence of the watermark record(s), and recording and storing a verified transaction indicative of the presence of the watermark record(s) in the external voice clone record in database(s).
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR RECORDING AND VERIFYING VOICECLONES, PHOTO AND VIDEO, CLONES AND / OR COPIES OF TEXTCROSS REFERENCES AND PRIORITIES

[0001] This Application claims priority from United States Provisional Application No.63 / 698,452 filed on 24 September 2024, the teachings of which are incorporated byreference herein in their entirety.BACKGROUND

[0002] Voice cloning technology has rapidly advanced in recent years, allowing for thecreation of highly realistic artificial voices based on samples of real human speech. Thistechnology has numerous applications in entertainment, accessibility, andcommunication industries. However, the ease of creating and distributing voice cloneshas raised significant concerns regarding intellectual property rights, unauthorized use,and potential for misuse for fraudulent activities. Similar concerns exist for photographs,videos, and texts.

[0003] Existing systems for managing and licensing voice clone data, and copies ofphotos, videos, and text often lack robust security measures, transparent licensingprocesses, and reliable verification methods. Additionally, current solutions struggle toprovide a balance between protecting the rights of voice, photo, video, and / or text ownersand facilitating legitimate use of voice cloning technology.

[0004] The need exists, therefore, for an improved system and method that can securelystore voice data, photos, videos, and / or text, manage licensing agreements, and providea reliable method for verifying the authenticity and authorized use of voice clones, and / orcopies of photos, videos, and text.SUMMARY

[0005] Described herein is a system for recording and verifying a voice clone. The systemincludes a first device, a first distributed network, a second device, a second distributednetwork, a third device, a third distributed network, and a database. The first device beingfor receiving a voice recording in a form of a binary code record or for converting thevoice recording into the binary code record, conducting a pre-cryptographic processingon the binary code record to produce a voice clone record, and cryptographically hashingthe binary code record and the voice clone record based on a cryptographic algorithm.The binary code record and the voice clone record are submitted to the first distributednetwork. The second device generates one or more storage and access control records forthe voice clone record, generates one or more watermark records associated with the voiceclone record, and transmits the one or more watermark records to the third device. Theone or more watermark records and the one or more storage and access control recordsare submitted to the second distributed network. The third device receives the one ormore watermark records from the second device, receives an external voice clone record,and evaluates the external voice clone record for the presence of the one or morewatermark records. The watermark records and the external voice clone record aresubmitted to the third distributed network. The database records and stores a verifiedtransaction indicative of the presence of the one or more watermark records in theexternal voice clone record.

[0006] In some embodiments, the first device may conduct a pre-binary code recordprocessing on the voice recording. The pre-binary code record processing may includeconducting a noise reduction on the voice recording to reduce background noise andproduce a noise reduced voice recording, conducting a normalization on the noisereduced voice recording to produce a normalized voice recording having a consistentvolume of speech, conducting a segmentation on the normalized voice recording to splitthe normalized voice recording into a segmented voice recording, conducting a Short-Time Fourier Transform (STFT) on the segmented voice recording to create aspectrogram representative of constituent frequencies of the segmented voice recordingover time, creating a mel-frequency cepstral coefficient from the spectrogram, estimatinga pitch of a voice in the spectrogram at one or more individual time segments, extractingformants from the spectrogram with said formants representing one or more resonantfrequencies of a vocal tract, extracting a duration of one or more of phonemes, syllables,and pauses from the spectrogram, extracting energy patterns representing an amplitudeof a speech signal from the spectrogram, processing sequential data from the spectrogramto generate a time-series representation of individual voice features, and creating anembedding of a compact representation of one or more unique characteristics of a voice.

[0007] In certain embodiments of the pre-binary code record processing, the noisereduction may be conducted by a filter. In other embodiments of the pre-binary coderecord processing, the noise reduction may be conducted by a spectral subtraction noisereduction algorithm.

[0008] In some embodiments of the pre-binary code record processing, the pitch of thevoice may be estimated by an autocorrelation. In other embodiments of the pre-binarycode record processing the pitch of the voice may be estimated by a harmonic productspectrum (HPS).

[0009] In certain embodiments of the pre-binary code record processing, the time-seriesrepresentation of individual voice features may be generated by a recurrent neuralnetwork (RNN). In other embodiments of the pre-binary code record processing, the time-series representation of individual voice features may be generated by a transformer.

[0010] In some embodiments of the pre-binary code record processing, converting thevoice recording into the binary record may include converting a plurality of floating-pointnumbers into a binary format by quantizing continuous values of the embedding to afixed number of bits to create a quantized embedding, organizing data from the quantizedembedding into a binary format to create a serialized binary file, compressing theserialized binary file, and associating headers comprising metadata related to the one ormore individual voice segments. The serialized binary file may have one or moreindividual voice segments associated with one or more corresponding voice features. Incertain such embodiments, the serialized binary file may be compressed by Huffmancoding. In other such embodiments, the serialized binary file may be compressed by a run-length encoding (RLE). The metadata may include a metadata record corresponding toone or more of a sampling rate, a bit depth, and a length of an embedding vector.

[0011] In certain embodiments, the pre-binary code record processing may includealigning a text input transcription with one or more corresponding audio features tocreate a mapping between one or more text segments and one or more voice embeddings,and converting the text input transcription into voice embeddings extracted and storedin the binary code record. In some such embodiments, aligning the text inputtranscription may be conducted by a Montreal Forced Aligner. In certain suchembodiments, converting the text input transcription may be conducted using asequence-to-sequence model.

[0012] In some embodiments, the pre-cryptographic processing may include aligning atext input transcription with one or more corresponding audio features to create amapping between one or more text segments and one or more voice embeddings, andconverting one or more predicted embeddings into one or more actual speech waveforms.In certain such embodiments, aligning the text input transcription may be conducted bya Montreal Forced Aligner. In some such embodiments, converting one or morepredicted embeddings may be conducted using a vocoder.

[0013] In certain embodiments, the one or more storage and access control record maybe stored in a smart contract.

[0014] In some embodiments, generating the one or more watermark records mayinclude embedding one or more audio signals into the voice clone record, andmodulating one or both of the amplitude and / or the phase of at least one harmonicassociated with the one or more audio signals. In such embodiments, evaluating theexternal voice clone record for the presence of the one or more watermark records mayinclude filtering the external voice clone record for the presence of the one or more audiosignals, and correlating the one or more audio signals to corresponding audio signals inthe external voice clone record.

[0015] In certain embodiments, generating the one or more watermark records mayinclude embedding a delayed version of the voice recording into the voice clone record.The delayed version of the voice recording may have a sound pressure of less than 130 dB.In such embodiments, evaluating the external voice clone record for the presence of theone or more watermark records may include filtering the voice clone record for thepresence of the delayed version of the voice recording.

[0016] In some embodiments, generating the one or more watermark records mayinclude identifying one or more least significant bits (LSBs) within the voice recording,and altering at least one parameter of at least one of the LSBs to create one or morealtered LSBs. The at least one parameter may be selected from the group consisting ofamplitude and frequency. In such embodiments, evaluating the external voice clonerecord for the presence of the one or more watermark records may include examining atleast one parameter of the LSBs within the external voice clone record for the presenceor absence of the one or more altered LSBs.

[0017] In certain embodiments, generating the one or more watermark records mayinclude conducting a Discrete Fourier Transform (DFT) or a Discrete Cosine Transform(DCT) on the voice recording, modifying one or more frequency coefficients within thevoice recording to generate one or more modified frequency coefficients, and embeddingthe one or more modified frequency coefficients into the voice clone record. In suchembodiments, evaluating the external voice clone record for the presence of the one ormore watermark records may include filtering the external voice clone record for thepresence of the one or more modified frequency coefficients, and correlating the one ormore modified frequency coefficients to corresponding modified frequency coefficientsin the external voice clone record.

[0018] Further described herein is a method to record and verify a voice clone. Themethod includes submitting a voice recording in a form of a binary code record to a firstdevice or converting the voice recording to the binary code record using the first device.The method further includes conducting a pre-cryptographic processing on the binarycode record to produce a voice clone record. The method further includes providing afirst cryptographic algorithm to hash the binary code record and the voice clone record.The method also includes cryptographically hashing the binary code record and the voiceclone record based on the first cryptographic algorithm. The method also includessubmitting the binary code record and the voice clone record to a first distributednetwork. The method further includes generating one or more storage and access controlrecords for the voice clone record. The method further includes providing a secondcryptographic algorithm to hash the storage and access control records. The method alsoincludes cryptographically hashing the storage and access control records based on thesecond cryptographic algorithm. The method also includes generating one or morewatermark records and associating the one or more watermark records with the voiceclone record. The method further includes providing a third cryptographic algorithm tohash the one or more watermark records. The method further includes cryptographicallyhashing the one or more watermark records based on the third cryptographic algorithm.The method also includes receiving an external voice clone record. The method alsoincludes evaluating the external voice clone record for the presence of the one or morewatermark records. The method further includes recording and storing a verifiedtransaction indicative of the presence of the one or more watermark records in theexternal voice clone record in one or more databases.

[0019] In some embodiments, the method may further include conducting a noisereduction on the voice recording to reduce background noise and produce a noisereduced voice recording. In certain such embodiments, the method may also includeconducting a normalization on the noise reduced voice recording to produce anormalized voice recording having a consistent volume of speech. In some suchembodiments, the method may further include conducting a segmentation on thenormalized voice recording to split the normalized voice recording into a segmented voicerecording. In certain such embodiments, the method may also include conducting aShort-Time Fourier Transform (STFT) on the segmented voice recording to create aspectrogram representative of constituent frequencies of the segmented voice recordingover time. In some such embodiments, the method may further include creating a mel-frequency cepstral coefficient from the spectrogram. In certain such embodiments, themethod may further include estimating a pitch of a voice in the spectrogram at one ormore individual time segments. In some such embodiments, the method may also includeextracting formants from the spectrogram with said formants representing one or moreresonant frequencies of a vocal tract. In certain such embodiments, the method may alsoinclude extracting a duration of one or more of phonemes, syllables, and pauses from thespectrogram. In some such embodiments, the method may also include extracting energypatterns representing an amplitude of a speech signal from the spectrogram. In certainsuch embodiments, the method may further include processing sequential data from thespectrogram to generate a time-series representation of individual voice features. Incertain such embodiments, the method may also include creating an embedding of acompact representation of one or more unique characteristics of a voice. Each of theseoptional features of the method may be conducted by the first device prior to conductingthe pre-cryptographic processing on the binary code record.

[0020] In certain embodiments of the method, the noise reduction may be conducted bya filter. In other embodiments of the method, the noise reduction may be conducted bya spectral subtraction noise reduction algorithm.

[0021] In some embodiments of the method, the pitch of the voice may be estimated byan autocorrelation. In other embodiments of the method, the pitch of the voice may beestimated by a harmonic product spectrum (HPS).

[0022] In certain embodiments of the method, the time-series representation ofindividual voice features may be generated by a recurrent neural network (RNN). In otherembodiments of the method, the time-series representation of individual voice featuresmay be generated by a transformer.

[0023] In some embodiments of the method, converting the voice recording to the binarycode record may include converting a plurality of floating-point numbers into a binaryformat by quantizing continuous values of the embedding to a fixed number of bits tocreate a quantized embedding, organizing data from the quantized embedding into abinary format to create a serialized binary file, compressing the serialized binary file, andassociating headers comprising metadata related to the one or more individual voicesegments. The serialized binary file may have one or more individual voice segmentsassociated with one or more corresponding voice features. Each of these optional featuresof the method may be conducted by the first device.

[0024] In certain embodiments of the method, compressing the serialized binary file maybe conducted by Huffman coding. In other embodiments of the method, compressingthe serialized binary file may be conducted by a run-length encoding (RLE).

[0025] In some embodiments of the method, the metadata record may correspond toone or more of a sampling rate, a bit depth, and a length of an embedding vector.

[0026] In certain embodiments of the method, the pre-cryptographic processing mayinclude aligning a text input transcription with one or more corresponding audio featuresto create a mapping between one or more text segments and one or more voiceembeddings, and converting the text input transcription into voice embeddings extractedand stored in the binary code record. In some such embodiments, aligning the text inputtranscription may be conducted by a Montreal Forced Aligner. In certain suchembodiments, converting the text input transcription may be conducted using asequence-to-sequence model.

[0027] In some embodiments of the method, the pre-cryptographic processing mayinclude aligning a text input transcription with one or more corresponding audio featuresto create a mapping between one or more text segments and one or more voiceembeddings, and converting one or more predicted embeddings into one or more actualspeech waveforms. In certain such embodiments, aligning the text input transcriptionmay be conducted by a Montreal Forced Aligner. In some such embodiments, theconverting one or more predicted embeddings may be conducted using a vocoder.

[0028] In certain embodiments of the method, the one or more storage and accesscontrol records may be stored in a smart contract.

[0029] In some embodiments of the method, generating one or more watermark recordsmay include embedding one or more audio signals into the voice clone record, andmodulating one or both of the amplitude and / or the phase of at least one harmonicassociated with the one or more audio signals. In such embodiments, evaluating theexternal voice clone record for the presence of the one or more watermark records mayinclude filtering the external voice clone record for the presence of the one or more audiosignals, and correlating the one or more audio signals to corresponding audio signals inthe external voice clone record.

[0030] In certain embodiments of the method, generating the one or more watermarkrecords may include embedding a delayed version of the voice recording into the voiceclone record. The delayed version of the voice recording may have a sound pressure ofless than 130 dB. In such embodiments, evaluating the external voice clone record forthe presence of the one or more watermark records may include filtering the voice clonerecord for the presence of the delayed version of the voice recording.

[0031] In some embodiments of the method, generating the one or more watermarkrecords may include identifying one or more least significant bits (LSBs) within the voicerecording, and altering at least one parameter of the at least one of the LSBs to create oneor more altered LSBs. The at least one parameter may be selected from the groupconsisting of amplitude and frequency. In such embodiments, evaluating the externalvoice clone record for the presence of the one or more watermark records may includeexamining at least one parameter of the LSBs within the external voice clone record forthe presence or absence of the one or more altered LSBs.

[0032] In certain embodiments of the method, generating the one or more watermarkrecords may include conducting a Discrete Fourier Transform (DFT) or a Discrete CosineTransform (DCT) on the voice recording, modifying one or more frequency coefficientswithin the voice recording to generate one or more modified frequency coefficients, andembedding the one or more modified frequency coefficients into the voice clone record.In such embodiments, evaluating the external voice clone record for the presence of theone or more watermark records may include filtering the external voice clone record forthe presence of the one or more modified frequency coefficients, and correlating the oneor more modified frequency coefficients to corresponding modified frequencycoefficients in the external voice clone record.

[0033] Also described herein is a system for recording and verifying authenticity of aphoto or video. The system includes a first device, a first distributed network, a seconddevice, a second distributed network, a third device, a third distributed network, and adatabase. The first device receive a photo or vide file in a form of a binary code record orconverts the photo or video file into the binary code record, conducts a pre-cryptographicprocessing on the binary code record to produce a clone record, and cryptographicallyhashes the binary code record and the clone record based on a cryptographic algorithm.The binary code record and the clone record are submitted to the first distributednetwork. The second device generates one or more storage and access control records forthe clone record, generates one or more watermark records associated with the clonerecord, and transmits the one or more watermark records to the third device. The one ormore watermark records and the one or more storage and access control records aresubmitted to the second distributed network. The third device receives the one or morewatermark records from the second device, receives an external voice clone record, andevaluates the external clone record for the presence of the one or more watermark records.The watermark records and the external clone record are submitted to the thirddistributed network. The database records and stores a verified transaction indicative ofthe presence of the one or more watermark records in the external clone records.

[0034] In some embodiments, generating one or more watermark records may includeidentifying one or more least significant bits (LSBs) of a plurality of pixel values of theclone record, and embedding a watermark image in at least one of the LSBs to create oneor more altered LSBs. In such embodiments, evaluating the external clone record for thepresence of the one or more watermark records may include examining at least one of theLSBs within the external clone record for the presence or absence of the watermark image.

[0035] In certain embodiments, generating one or more watermark records may includemodifying a parameter selected from the group consisting of texture, color, andcombinations thereof within a plurality of pixels in the clone record. In suchembodiments, evaluating the external clone record for the presence of the one or morewatermark records may include examining the plurality of pixels within the external clonerecord for the presence or absence of the parameter.

[0036] In some embodiments, generating one or more watermark records may includetransforming the clone record into a frequency domain by Discrete Cosine Transform(DCT), and embedding a watermark into a mid-frequency coefficient of the frequencydomain. In such embodiments, evaluating the external clone record for the presence ofthe one or more watermarks may include transforming the external clone record into anexternal clone record frequency domain by Discrete Cosine Transform (DCT), andanalyzing an external clone record mid-frequency coefficient for the presence or absenceof the watermark.

[0037] In certain embodiments, generating the one or more watermark records mayinclude decomposing an image in the clone record into a plurality of frequencycomponents by Discrete Wavelet Transform (DWT), and embedding a watermark into alow-frequency band and / or a mid-frequency band. In such embodiments, evaluating theexternal clone record for the presence of the one or more watermark records may includedecomposing the external clone record into a plurality of external clone record frequencycomponents by Discrete Wavelet Transform (DWT), and analyzing the external clonerecord frequency components for the presence or absence of the watermark in an externalclone record low-frequency band and / or an external clone record mid-frequency band.

[0038] In some embodiments, generating the one or more watermark records mayinclude embedding a watermark in a plurality of pixels in an image in the clone record byNeural Network-Based Embedding. In such embodiments, evaluating the external clonerecord for the presence of the one or more watermark records may include analyzing theexternal clone record for the presence or absence of the watermark.

[0039] In certain embodiments, generating the one or more watermark records mayinclude embedding a watermark in a plurality of pixels in an image in the clone record byAutoencoder-Based Watermarking. In such embodiments, evaluating the external clonerecord for the presence of the one or more watermark records may include analyzing theexternal clone record for the presence or absence of the watermark.

[0040] In some embodiments, the one or more storage and access control records maybe stored in a smart contract.

[0041] Also described herein is a method to record and verify an image clone or videoclone. The method includes submitting a photo or video file in a form of a binary coderecord to a first device or converting the photo or video file to the binary code recordusing the first device. The method further includes conducting a pre-cryptographicprocessing on the binary code record to produce a clone record. The method also includesproviding a first cryptographic algorithm to hash the binary code record and the clonerecord. The method also includes cryptographically hashing the binary code record andthe clone record based on the first cryptographic algorithm. The method further includessubmitting the binary code record and the clone record to a first distributed network. Themethod further includes generating one or more storage and access control records forthe clone record. The method also includes providing a second cryptographic algorithmto hash the storage and access control records. The method also includescryptographically hashing the storage and access control records based on the secondcryptographic algorithm. The method further includes generating one or more watermarkrecords and associating the one or more watermark records with the clone record. Themethod further includes providing a third cryptographic algorithm to hash the one ormore watermark records. The method further includes cryptographically hashing the oneor more watermark records based on the third cryptographic algorithm. The method alsoincludes receiving an external clone record. The method also includes evaluating theexternal clone record for the presence of the one or more watermark records. The methodalso includes recording and storing a verified transaction indicative of the presence of theone or more watermark records in the external clone record in one or more databases.

[0042] In certain embodiments of the method, generating one or more watermarkrecords may include identifying one or more least significant bits (LSBs) of a plurality ofpixel values of the clone record, and embedding a watermark image in at least one of theLSBs to create one or more altered LSBs. In such embodiments, evaluating the externalclone record for the presence of the one or more watermark records may includeexamining at least one of the LSBs within the external clone record for the presence orabsence of the watermark image.

[0043] In certain embodiments of the method, generating one or more watermarkrecords may include modifying a parameter selected from the group consisting of texture,color, and combinations thereof within a plurality of pixels in the clone record. In suchembodiments, evaluating the external clone record for the presence of the one or morewatermark records may include examining the plurality of pixels within the external clonerecord for the presence or absence of the parameter.

[0044] In some embodiments of the method, generating one or more watermark recordsmay include transforming the clone record into a frequency domain by Discrete CosineTransform (DCT), and embedding a watermark into a mid-frequency coefficient of thefrequency domain. In such embodiments, evaluating the external clone record for thepresence of the one or more watermarks may include transforming the external clonerecord into an external clone record frequency domain by Discrete Cosine Transform(DCT), and analyzing an external clone record mid-frequency coefficient for the presenceor absence of the watermark.

[0045] In certain embodiments of the method, generating the one or more watermarkrecords may include decomposing an image in the clone record into a plurality offrequency components by Discrete Wavelet Transform (DWT), and embedding awatermark into a low-frequency band and / or a mid-frequency band. In suchembodiments, evaluating the external clone record for the presence of the one or morewatermark records may include decomposing the external clone record into a plurality ofexternal clone record frequency components by Discrete Wavelet Transform (DWT), andanalyzing the external clone record frequency components for the presence or absence ofthe watermark in an external clone record low-frequency band and / or an external clonerecord mid-frequency band.

[0046] In some embodiments of the method, generating the one or more watermarkrecords may include embedding a watermark in a plurality of pixels in an image in theclone record by Neural Network-Based Embedding. In such embodiments, evaluating theexternal clone record for the presence of the one or more watermark records may includeanalyzing the external clone record for the presence or absence of the watermark.

[0047] In certain embodiments of the method, generating the one or more watermarkrecords may include embedding a watermark in a plurality of pixels in an image in theclone record by Autoencoder-Based Watermarking. In such embodiments, evaluating theexternal clone record for the presence of the one or more watermark records may includeanalyzing the external clone record for the presence or absence of the watermark.

[0048] In some embodiments of the method, the one or more storage and access controlrecords may be stored in a smart contract.

[0049] Further described herein is a system for recording and verifying authenticity of atext. The system includes a first device, a first distributed network, a second distributednetwork, a second device, a second distributed network, a third device, a third distributednetwork, and a database. The first device receives a text file in a form of a binary coderecord or converts the text file to the binary code record, conducts a pre-cryptographicprocessing on the binary code record to produce a copied text record, andcryptographically hashes the binary code record and the copied text record based on acryptographic algorithm. The binary code record and the copied text record are submittedto the first distributed network. The second device generates one or more storage andaccess control records for the copied text record, generates one or more watermarkrecords associated with the copied text record, and transmits the one or more watermarkrecords to the third device. The one or more watermark records and the one or morestorage and access control records are submitted to the second distributed network. Thethird device receives the one or more watermark records from the second device, receivesan external copied text record, and evaluates the external copied text record for thepresence of the one or more watermark records. The watermark records and the externalcopied text record are submitted to the third distributed network. The database recordsand stores a verified transaction indicative of the presence of the one or more watermarkrecords in the external copied text records.

[0050] In some embodiments, generating one or more watermark records may includeidentifying one or more repeated words in the copied text record, and replacing at leastone of the repeated words with a synonym for the repeated word. In such embodiments,evaluating the external copied text record for the presence of the one or more watermarkrecords may include examining the external copied text record for the presence or absenceof the synonym for the repeated word.

[0051] In certain embodiments, generating one or more watermark records may includeintroducing a patterned alteration in sentence structure to at least one sentence in thecopied text record. In such embodiments, evaluating the external copied text record forthe presence of the one or more watermark records may include examining the externalcopied text record for the presence or absence of the patterned alteration in sentencestructure.

[0052] In some embodiments, generating one or more watermark records may includeintroducing a patterned systematic variation in punctuation and / or a patternedsystematic variation in capitalization to the copied text record. In such embodiments,evaluating the external copied text record for the presence of the one or more watermarkrecords may include examining the external copied text record for the presence or absenceof the patterned systemic variation in punctuation and / or the patterned systemicvariation in capitalization.

[0053] In certain embodiments, generating one or more watermark records may includeinserting a contextually appropriate patterned word or phrase into the copied text record.In such embodiments, evaluating the external copied text record for the presence of theone or more watermark records may include examining the external copied text recordfor the presence or absence of the contextually appropriate patterned word or phrase.

[0054] In some embodiments, generating the one or more watermarks records mayinclude identifying one or more sentences in the copied text record, and replacing at leastone of the sentences with a patterned paraphrased sentence. In such embodiments,evaluating the external copied text record for the presence of the one or more watermarkrecords may include examining the external copied text record for the presence or absenceof the patterned paraphrased sentence.

[0055] In certain embodiments, generating one or more watermark records may includeintroducing a patterned ambiguity and / or a patterned dual meaning to at least oneportion of the copied text record. In such embodiments, evaluating the external copiedtext record for the presence of the one or more watermark records may include examiningthe external copied text record for the presence or absence of the patterned ambiguityand / or the patterned dual meaning.

[0056] In some embodiments, the one or more storage and access control records maybe stored in a smart contract.

[0057] Further disclosed herein is a method to record and verify a copied text. Themethod includes submitting a text file in a form of a binary code record to a first deviceor converting the text file to the binary code record using the first device. The methodfurther includes conducting a pre-cryptographic processing on the binary code record toproduce a copied text record. The method also includes providing a first cryptographicalgorithm to hash the binary code record and the copied text record. The method furtherincludes cryptographically hashing the binary code record and the copied text recordbased on the first cryptographic algorithm. The method also includes submitting thebinary code record and the copied text record to a first distributed network. The methodfurther includes generating one or more storage and access control records for the copiedtext record. The method also includes providing a second cryptographic algorithm to hashthe storage and access control records. The method further includes cryptographicallyhashing the storage and access control records based on the second cryptographicalgorithm. The method also includes generating one or more watermark records andassociating the one or more watermark records with the copied text record. The methodfurther includes providing a third cryptographic algorithm to hash the one or morewatermark records. The method also includes cryptographically hashing the one or morewatermark records based on the third cryptographic algorithm. The method furtherincludes receiving an external copied text record. The method also includes evaluatingthe external copied text record for the presence of the one or more watermark records.The method further includes recording and storing a verified transaction indicative ofthe presence of the one or more watermark records in the external copied text record inone or more databases.

[0058] In some embodiments of the method, generating one or more watermark recordsmay include identifying one or more repeated words in the copied text record, andreplacing at least one of the repeated words with a synonym for the repeated word. Insuch embodiments, evaluating the external copied text record for the presence of the oneor more watermark records may include examining the external copied text record for thepresence or absence of the synonym for the repeated word.

[0059] In certain embodiments of the method, generating one or more watermarkrecords may include introducing a patterned alteration in sentence structure to at leastone sentence in the copied text record. In such embodiments, evaluating the externalcopied text record for the presence of the one or more watermark records may includeexamining the external copied text record for the presence or absence of the patternedalteration in sentence structure.

[0060] In some embodiments of the method, generating one or more watermark recordsmay include introducing a patterned systematic variation in punctuation and / or apatterned systematic variation in capitalization to the copied text record. In suchembodiments, evaluating the external copied text record for the presence of the one ormore watermark records may include examining the external copied text record for thepresence or absence of the patterned systemic variation in punctuation and / or thepatterned systemic variation in capitalization.

[0061] In certain embodiments of the method, generating one or more watermarkrecords may include inserting a contextually appropriate patterned word or phrase intothe copied text record. In such embodiments, evaluating the external copied text recordfor the presence of the one or more watermark records may include examining theexternal copied text record for the presence or absence of the contextually appropriatepatterned word or phrase.

[0062] In some embodiments of the method, generating the one or more watermarksrecords may include identifying one or more sentences in the copied text record, andreplacing at least one of the sentences with a patterned paraphrased sentence. In suchembodiments, evaluating the external copied text record for the presence of the one ormore watermark records may include examining the external copied text record for thepresence or absence of the patterned paraphrased sentence.

[0063] In certain embodiments of the method, generating one or more watermarkrecords may include introducing a patterned ambiguity and / or a patterned dual meaningto at least one portion of the copied text record. In such embodiments, evaluating theexternal copied text record for the presence of the one or more watermark records mayinclude examining the external copied text record for the presence or absence of thepatterned ambiguity and / or the patterned dual meaning.

[0064] In some embodiments of the method, the one or more storage and access controlrecords may be stored in a smart contract.BRIEF DESCRIPTION OF FIGURES

[0065] FIG. 1 is an illustration of a flow chart of an exemplary embodiment for recordingand verifying works.DETAILED DESCRIPTION

[0066] Described herein are various systems and methods for recording and verifyingworks. The works may be in the form of a voice clone, a photo, a video, or a textdocument.

[0067] When the work is a voice clone, the system may include a first device, a firstdistributed network, a second device, a second distributed network, a third device, a thirddistributed network, and a database. The first device, the second device, and the thirddevice may each independently be in the form of a network device selected from the groupconsisting of a laptop computer, a desktop computer, a server, a mobile device (cellularphone, tablet, or the like), an Internet-of-Things (IoT) device, an endpoint, a virtualmachine, a cloud based server, a cloud based storage device, and combinations thereof.The first distributed network, the second distributed network, and the third distributednetwork may each independently be in the form of a blockchain network. Blockchainreferring to a growing list of records, or blocks, that are cryptographically linked. Theblocks commonly include a transaction data, a time stamp, and a cryptographic hash ofthe prior block in the chain that confirms the integrity of the prior block. Blockchainscan be used to implement a shared digital ledger for recording user transactions betweenmultiple devices – often stored in a cloud-based environment. When a user transaction isexecuted and validated, it is appended to the end of the blockchain, making theblockchain an immutable history of all valid transactions. The database may be in theform of a blockchain database of executed and validated transactions (or blocks).

[0068] The first device receives a voice recording in the form of a binary code record. Insome embodiments, the first device may generate the binary code from the voicerecording. The first device may generate the binary code record by first converting aplurality of floating-point numbers into a binary format. This may be accomplished byquantizing continuous values of the embedding to a fixed number of bits to create aquantized embedding.

[0069] Next, the first device may organize data from the quantized embedding into abinary format to create a serialized binary file. The serialized binary file may have one ormore individual voice segments associated with one or more corresponding voice features.The first device may then compress the serialized binary file, and associate headerscomprising metadata related to the one or more individual voice segments. Compressingthe serialized binary file may be conducted by Huffman coding and / or a run-lengthencoding (RLE). A Huffman coding referring to an optimal prefix code used for losslessdata compression in which the output from an algorithm conducting the Huffman codingis capable of being viewed as a variable-length code table for encoding a source symbol. Arun-length encoding (RLE) referring to a form of lossless data compression in whichconsecutive occurrences of the same data value are stored as a single occurrence of thatdata value and a count of its consecutive occurrences. The metadata may include ametadata record corresponding to one or more of a sampling rate which is the reductionof a continuous-time signal to a discrete-time signal, a bit depth which is the number ofbits of information in each sample, and a length of an embedding vector which is anumerical representation of data points in the compressed serialized binary file.

[0070] In some embodiments, the first device may first conduct a pre-binary code recordprocessing on the voice recording before generating the binary code. The pre-binary coderecord processing may include conducting a noise reduction on the voice recording toreduce background noise. The noise reduction may be conducted by a filter and / or by aspectral subtraction noise reduction algorithm. A filter referring to a frequency dependentcircuit which can amplify (boost), pass, or attenuate (cut) certain frequency ranges in anaudio file. A spectral subtraction noise reduction algorithm referring to an algorithmwhich estimates the noise spectrum during speech pauses and subtracts from the noisyspeech spectrum to estimate the clean speech. The resulting output being a noise reducedvoice recording.

[0071] The pre-binary code record processing may further include conducting anormalization on the noise reduced voice recording. Normalization referring to theprocess of applying a constant amount of gain to an audio recording to bring theamplitude to a target level. The resulting output being a normalized voice recordinghaving a consistent volume of speech.

[0072] The pre-binary code record processing may also include conducting asegmentation on the normalized voice recording. Segmentation referring to the processof separating different types of sounds – such as speech, environmental sounds, silence,music, and combinations thereof – from within the normalized voice recording. Theresulting output being a segmented voice recording.

[0073] The pre-binary code record processing may also include conducting a Short-TimeFourier Transform (STFT) on the segmented voice recording. The Short-Time FourierTransform being a Fourier-related transform for determining the sinusoidal frequencyand phase content of local sections of an audio signal as it changes over time. Theresulting output being a spectrogram representative of constituent frequencies of thesegmented voice recording over time.

[0074] The pre-binary code record processing may further include creating a mel-frequency cepstral coefficient from the spectrogram. A mel-frequency cepstral coefficientreferring to a representation of the short-term power spectrum of a sound based on linearcosine transform of a log power spectrum on a nonlinear mel scale of frequency.

[0075] The pre-binary code record processing may also include estimating a pitch of avoice in the spectrogram at one or more individual time segments. The pitch of the voicemay be estimated by an autocorrelation and / or a harmonic product spectrum (HPS). Anautocorrelation referring to the correlation of an audio signal with a delayed copy of itselfas a function of delay. A harmonic product spectrum (HPS) referring to a measure of themaximum coincidence for harmonics for each spectral frame in an audio signal.

[0076] The pre-binary code record processing may further include extracting formantsfrom the spectrogram. The formants will represent one or more resonant frequencies ofa vocal tract. Formant referring to the broad spectral maximum that results from anacoustic resonance of the human vocal tract.

[0077] The pre-binary code record processing may also include extracting a duration ofone or more of phonemes, syllables, and pauses from the spectrogram. This can beachieved by utilizing forced alignment tools such as the Montreal Forced Aligner, whichalign phonetic transcripts to audio data. By mapping each phoneme in the transcript toits corresponding time segment in the audio, the system can determine the start and endtimes of each phoneme. For syllable detection, algorithms that detect peaks in the energycontour of the speech signal may be used, as syllables often correspond to peaks inamplitude. Pause detection involves identifying regions of low energy or silence in thespectrogram by setting an energy threshold below which the signal is considered a pause.

[0078] The pre-binary code record processing may further include extracting energypatterns from the spectrogram. The energy patterns will represent an amplitude of aspeech signal. Extracting an energy pattern from the spectrogram may include computingthe short-time energy for each frame in the spectrogram, which may include calculatingthe sum of the squared magnitudes of the spectral coefficients. This gives a representationof the signal’s energy over time. The energy values may then be normalized and smoothedusing window functions such as Hamming or Hanning windows to reduce spectralleakage and make the energy contour suitable for further analysis, such as detecting stresspatterns, intonation, or identifying speech segments.

[0079] The pre-binary code record processing may further include processing sequentialdata from the spectrogram to generate a time-series representation of individual voicefeatures. The time-series representation of individual voice features may be generated bya recurrent neural network (RNN) and / or by a transformer. A recurrent neural network(RNN) referring to a class of artificial neural networks for sequential data processingacross multiple time steps for modelling and processing text, speech, and time series. Atransformer referring to a deep learning model which receives a text or audio input andgenerates various embeddings which can be read by a convolutional neural network(CNN).

[0080] The pre-binary code record processing may further include creating an embeddingof a compact representation of one or more unique characteristics of a voice. Embeddinga compact representation of one or more unique characteristics of a voice may entail usinga speaker embedding model, such as a speaker encoder trained on a large dataset ofdiverse voices. The process may involve feature extraction, speaker encoder network,training with triplet loss, and generating the embedding. Feature extraction extractsacoustic features from the voice recording, such as Mel-frequency cepstral coefficients(MFCCs) or filter bank features. Speaker encoder network inputs the acoustic featuresinto a deep neural network designed to produce fixed-dimensional embeddings with thenetwork being trained to map variable-length speech segments to embeddings thatrepresent the speaker’s identity. Training with triplet loss involves embeddings of thesame speaker being pulled closer together while embeddings of different speakers arepushed apart. After said training, generating the embedding may involve passing the voicerecording through the network to obtain the speaker embedding which is a compactrepresentation of the unique voice characteristics.

[0081] The first device further conducts a pre-cryptographic processing on the binarycode record to produce a voice clone record, and converts the text input transcriptioninto voice embeddings extracted and stored in the binary code record. The pre-cryptographic processing may include aligning a text input transcription with one or morecorresponding audio features to create a mapping between one or more text segments andone or more voice embeddings. Aligning the text transcription may be conducted by aMontreal Forced Aligner. A Montreal Forced Aligner referring to a program for timealigning orthographic and phonological forms from a pronunciation dictionary toorthographically transcribed audio files. Converting the text input transcription intovoice embeddings may be conducted using a sequence-to-sequence model. A sequence-to-sequence model referring to a network operating on a sequence using its own output asinput for subsequent steps.

[0082] Alternatively, the pre-cryptographic processing may include aligning a text inputtranscription with one or more corresponding audio features to create a mapping betweenone or more text segments and one or more voice embeddings, and converting one ormore predicted embeddings into one or more actual speech waveforms. Aligning the texttranscription may be conducted by a Montreal Forced Aligner. Converting one or morepredicted embeddings may be conducted using a vocoder. A vocoder referring to speechcoding which analyzes and synthesizes the human voice signal for audio data compression,multiplexing, voice encryption, and / or voice transformation.

[0083] The first device further cryptographically hashes the binary code record and thevoice clone record based on a cryptographic algorithm. The cryptographic algorithm maybe in the form of an asymmetric-key algorithm or a hash function. Once the first devicecompletes the cryptographic hashing on the binary code record and the voice clonerecord, the first device submits the binary code record and the voice clone record to thefirst distributed network.

[0084] The second device generates one or more storage and access control records forthe voice clone record. The one or more storage and access control records may be storedin a smart contract. A smart contract is a computer program or a transaction protocolwhich automatically executes, controls, and / or documents events and actions accordingto the terms of a contract or agreement by sending a transaction from a wallet for theblockchain which includes compiled code and a special receiver address, the transactionbeing included in a block added to the blockchain such that the smart contract’s codewill execute to establish the initial state of the smart contract.

[0085] The second device further generates one or more watermark records associatedwith the voice clone record. Generating the one or more watermark records may includeembedding one or more audio signals into the voice clone record, and modulating oneor both of the amplitude and / or the phase of at least one harmonic associated with theone or more audio signals. Embedding the audio signals may be achieved using a varietyof techniques including spread spectrum technique, least significant bit (LSB)modification, phase coding, and / or quantization index modulation (QIM). Spreadspectrum technique may involve spreading the watermark audio signal over a widefrequency band, making it less susceptible to interference and more difficult to removewithout degrading the original signal. Least significant bit (LSB) modification may involvealtering the least significant bits of the audio samples to encode the watermark signal,ensuring minimal impact on audio quality. Phase coding may involve modifying the phaseof the original audio signal according to the watermark signal, which is less perceptible tohuman listeners. Quantization index modulation (QIM) may involve quantizing the hostaudio signal and adjusting quantization indices to encode the watermark. Modulating theamplitude and / or phase of at least one harmonic associated with the one or more audiosignals may include frequency domain conversion, harmonic selection, amplitudeadjustment, phase shift application, and / or inverse FFT. Frequency domain conversionmay include transforming the audio signal into the frequency domain using a Fast FourierTransform (FFT). Harmonic selection may include identifying the fundamental frequencyand its harmonics in the frequency domain. Amplitude adjustment may includemultiplying the amplitude of the selected harmonics by a modulation factor to embed thewatermark. Phase shift application may include adding a constant phase offset to thephase component of the selected harmonics. Inverse FFT may include applying theinverse FFT to convert the modified frequency-domain signal back to the time domain.

[0086] Alternatively, the second device may generate the one or more watermark recordsby embedding a delayed version – sometimes referred to as an “echo” – of the voicerecording into the voice clone record. Preferably, the delayed version of the voicerecording will have a sound pressure outside of that which is audibly detectable by thehuman ear – such as a sound pressure of less than 130 dB. Embedding the delayed versionof the voice recording into the voice clone record may include delay line creation,amplitude adjustment, superimposition, and / or frequency masking. Delay line creationmay include creating a delayed version of the original signal by applying a delay (of as littleas two to three milliseconds) to the original signal). Amplitude adjustment may includereducing the amplitude of the delayed signal to a level that is imperceptible to humanlisteners, but which can be detected algorithmically. Superimposition may include addingthe delayed, attenuated signal back into the original signal to create a composite signal.Frequency masking may include embedding the delayed signal into the frequency bandsthat are less perceptible to humans to further reduce audibility.

[0087] As another alternative, the second device may generate the one or morewatermark records by identifying one or more least significant bits (LSBs) within the voicerecording. The least significant bits being the bit position in a binary integer representingthe binary 1s place of the integer. Once identified, at least one parameter of at least oneof the LSBs may be altered to create one or more altered LSBs. Said parameter may beselected from the group consisting of amplitude and frequency. Altering said parametermay include modifying the amplitude values represented by the LSBs to encode thewatermark data, ensuring that changes are minimal to preserve audio quality.

[0088] A further alternative for generating the one or more watermark records mayinclude conducting a Discrete Fourier Transform (DFT) or a Discrete Cosine Transform(DCT) on the voice recording. A Discrete Fourier Transform referring to a process forconverting a finite sequence of equally-spaced samples of a function into a same-lengthsequence of equally-spaced samples of the discrete-time Fourier transform (DTFT). ADiscrete Cosine Transform referring to a finite sequence of data points in terms of a sumof cosine functions oscillating at different frequencies using only real numbers. Onceconducted, one or more frequency coefficients within the voice recording may bemodified to generate one or more modified frequency coefficients. This may beconducted by slightly adjusting the magnitude of selected frequency components in a waythat represents the watermark data. The one or more modified frequency coefficients maythen be embedded into the voice clone record.

[0089] Once the second device generates the one or more storage and access controlrecords and the one or more watermark records, the second device transmits the one ormore watermark records to the third device, and submits the one or more watermarkrecords and the one or more storage and access control records to the second distributednetwork. The second distributed network is then capable of receiving a voice clonerequest from a third party licensee and receiving a completed license contract from thethird party licensee for storage in the storage and access control records.

[0090] The third device receives an external voice clone record. Upon receiving theexternal voice clone record, the third device evaluates the external voice clone record forthe presence of the one or more watermark records. The third device then submits thewatermark records and the external voice clone record to the third distributed network.The evaluation of the external voice clone record for the presence of the one or morewatermark records may involve filtering, correlation analysis, and / or frequency analysis.Filtering may include applying a matched filter using the known watermark signal as thefilter to maximize the signal-to-noise ratio for detecting the watermark. Correlationanalysis may include computing the cross-correlation between the external voice clonerecord and the known watermark signal to detect similarity peaks. Frequency analysis mayinclude performing spectral analysis to identify frequency components associated with thewatermark signal.

[0091] When generating the one or more watermark records includes embedding one ormore audio signals into the voice clone record, and modulating one or both of theamplitude and / or the phase of at least one harmonic associated with the one or moreaudio signals, evaluating the external voice clone record for the presence of the one ormore watermark records may include filtering the external voice clone record for thepresence of the one or more audio signals. This may include frequency domaintransformation, amplitude / phase modulation detection, and / or watermark harmonicsidentification. Frequency domain transformation may include converting the externalvoice clone record back into the frequency domain using a Fast Fourier Transform (FFT)thereby allowing access to the harmonic components of the signal. Amplitude / phasemodulation detection may include comparing the amplitude and phase of the harmonicsin the external voice clone record to those of the original voice signal to detectmodifications in amplitude and / or phase indicative of the presence of the watermark.Watermark harmonics identification may include analyzing which specific harmonicshave been modulated in the external voice clone record based on differences in amplitudeor phase and determining whether the detected modulations match the pattern usedduring watermark embedding. Once the external voice clone record is filtered, the one ormore audio signals in the voice clone record may be correlated to corresponding audiosignals in the external voice clone record.

[0092] When generating the one or more watermark records includes embedding adelayed version of the voice recording into the voice clone record, evaluating the externalvoice clone record for the presence of the one or more watermark records may includefiltering the voice clone record for the presence of the delayed version of the voice record.This may be conducted by echo detection and / or signal processing. Echo detection mayinclude analyzing the audio signal for characteristic patterns of echoes by examining theautocorrelation function of the signal. Signal processing may including the use ofalgorithms that detect low-amplitude delayed replicas of the original signal within thecomposite signal.

[0093] When generating the one or more watermark records includes identifying one ormore least significant bits (LSBs) within the voice recording and alternating at least oneparameter of at least one of the LSBs to create an altered LSB, evaluating the externalvoice clone record for the presence of the one or more watermark records may includeexamining at least one parameter of the LSBs within the external voice clone record forthe presence of the one or more altered LSBs. This may be conducted by LSB extraction,pattern matching, and / or statistical analysis. LSB extraction may include extracting theleast significant bits from the audio samples of the external voice clone record. Patternmatching may include comparing the extracted LSBs with the expected watermark patternor sequence. Statistical analysis may include performing statistical tests to determine ifthe LSBs match the watermark, considering potential noise or modifications.

[0094] When generating the one or more watermark records includes conducting a DFTor a DCT on the voice recording, modifying one or more frequency coefficients withinthe voice recording, and embedding one or more modified frequency coefficients into thevoice clone record, then evaluating the external voice clone record for the presence of theone or more watermark records may include filtering the external voice clone record forthe presence of the one or more modified frequency coefficients. This may be conductedby frequency domain conversion, coefficient comparison, and / or thresholding.Frequency domain conversion may include applying DFT or DCT to the external voiceclone record to obtain frequency domain coefficients. Coefficient comparison mayinclude retrieving the frequency coefficients that were used for embedding and comparingthem with expected values. Thresholding may include using predefined thresholds todetermine if the coefficients have been modified according to the watermark encodingscheme. Next, the one or more modified frequency coefficients may be correlated tocorresponding modified frequency coefficients in the external voice clone record.

[0095] Once the third device identifies the presence – or absence – of the one or morewatermark records in the external voice clone record, the database records and stores averified transaction. This verified transaction being indicative of the presence – or absence– of the one or more watermark records in the external voice clone record. The presenceof the one or more watermark records in the external voice clone record indicates thatthe external voice clone record is an authentic clone of the voice recording. Conversely,the absence of the one or more watermark records in the external voice clone recordindicates that the external voice clone record is not an authentic clone of the voicerecording.

[0096] When the work is an image clone or a video clone, the system may include a firstdevice, a first distributed network, a second device, a second distributed network, a thirddevice, a third distributed network, and a database. The first device, the second device,and the third device may each independently be in the form of a network device selectedfrom the group consisting of a laptop computer, a desktop computer, a server, a mobiledevice (cellular phone, tablet, or the like), an Internet-of-Things (IoT) device, an endpoint,a virtual machine, a cloud based server, a cloud based storage device, and combinationsthereof. The first distributed network, the second distributed network, and the thirddistributed network may each independently be in the form of a blockchain network.Blockchain referring to a growing list of records, or blocks, that are cryptographicallylinked. The blocks commonly include a transaction data, a time stamp, and acryptographic hash of the prior block in the chain that confirms the integrity of the priorblock. Blockchains can be used to implement a shared digital ledger for recording usertransactions between multiple devices – often stored in a cloud-based environment.When a user transaction is executed and validated, it is appended to the end of theblockchain, making the blockchain an immutable history of all valid transactions. Thedatabase may be in the form of a blockchain database of executed and validatedtransactions (or blocks).

[0097] The first device receives a photo or video file in the form of a binary code record.In some embodiments, the first device may generate the binary code from the photo orvideo file. The first device may generate the binary code record by first converting aplurality of floating-point numbers into a binary format. This may be accomplished byquantizing continuous values of the embedding to a fixed number of bits to create aquantized embedding.

[0098] The first device further conducts a pre-cryptographic processing on the binarycode record to produce a clone record. The pre-cryptographic processing may include oneor more of image / video compression, feature extraction, metadata association,normalization, and / or noise reduction and enhancement. Image / video compressionreferring to compressing the media using standardized codecs (such as JPEG for imagesor H.264 for videos) to reduce their size while maintaining quality. Feature extractionmay include extracting key features using techniques such as Scale-Invariant FeatureTransform (SIFT) or Speeded-Up Robust Features (SURF) to represent the visual content.Metadata Association may include attaching metadata such as timestamps, GPScoordinates, device information, and creator identity. Normalization may includestandardizing image sizes, color spaces, and resolution to ensure consistency acrossdifferent devices and platforms. Noise reduction and enhancement may include applyingfilters to reduce noise and enhance important visual elements.

[0099] The first device further cryptographically hashes the binary code record and theclone record based on a cryptographic algorithm. The cryptographic algorithm may be inthe form of an asymmetric-key algorithm or a hash function. Once the first devicecompletes the cryptographic hashing on the binary code record and the clone record, thefirst device submits the binary code record and the clone record to the first distributednetwork.

[0100] The second device generates one or more storage and access control records forthe clone record. The one or more storage and access control records may be stored in asmart contract. A smart contract is a computer program or a transaction protocol whichautomatically executes, controls, and / or documents events and actions according to theterms of a contract or agreement by sending a transaction from a wallet for the blockchainwhich includes compiled code and a special receiver address, the transaction beingincluded in a block added to the blockchain such that the smart contract’s code willexecute to establish the initial state of the smart contract.

[0101] The second device further generates one or more watermark records associatedwith the clone record. Generating the one or more watermark records may includeidentifying one or more least significant bits (LSBs) of a plurality of pixel values of theclone record. The least significant bits being the bit position in a binary integerrepresenting the binary 1s place of the integer. Once the LSBs have been identified, awatermark image may be embedded into at least one of the LSBs to create one or morealtered LSBs.

[0102] When generating the one or more watermark records includes identifying one ormore least significant bits (LSBs) within the clone record and embedding a watermarkimage into at least one of the LSBs to create an altered LSB, evaluating the external clonerecord for the presence of the one or more watermark records may include examining atleast one of the LSBs within the external clone record for the presence of the one or morealtered LSBs. This may be conducted by extracting and examining the LSBs within theexternal clone record for the presence of the watermark image.

[0103] Alternatively, generating the one or more watermark records may includemodifying one or more parameters within a plurality of pixels in the clone records. Saidparameters may be selected from the group consisting of texture, color, and combinationsthereof.

[0104] When generating the one or more watermark records includes modifying one ormore parameters within a plurality of pixels in the clone record, evaluating the externalclone record for the presence of the one or more watermark records may includeexamining the plurality of pixels within the external clone record for the presence orabsence of the altered parameter – such as alterations in texture, color, and combinationsthereof. This may be conducted by analyzing the pixels in the external clone record formodifications in texture or color parameters which match the watermarking scheme.

[0105] Another alternative for generating the one or more watermark records mayinclude transforming the clone record into a frequency domain by Discrete CosineTransform (DCT). Discrete Cosine Transform referring to a finite sequence of data pointsin terms of a sum of cosine functions oscillating at different frequencies using only realnumbers. Once the DCT has been conducted, a watermark may be embedded into a mid-frequency coefficient of the frequency domain.

[0106] When generating the one or more watermark records includes transforming theclone record into a frequency domain by DCT, then evaluating the external clone recordfor the presence of the one or more watermark records may include first transforming theexternal clone record into an external clone record frequency domain by Discrete CosineTransform (DCT). Following which the external clone record mid-frequency coefficientmay be analyzed for the presence or absence of the watermark. This may be conducted byapplying DCT to the external clone record and analyzing the mid-frequency coefficientsfor the embedded watermark.

[0107] A further alternative for generating the one or more watermark records mayinclude decomposing an image in the clone record into a plurality of frequencycomponents by Discrete Wavelet Transform (DWT). Discrete Wavelet Transform (DWT)referring to a wavelet transform in which the wavelets are discretely sampled to captureboth frequency and location in time information. Following the DWT, a watermark maybe embedded into a low-frequency band and / or a mid-frequency band.

[0108] When generating the one or more watermark records includes decomposing animage in the clone record into a plurality of frequency components by DWT, andembedding a watermark into a low-frequency band and / or a mid-frequency band, thenevaluating the external clone record for the presence of the one or more watermarkrecords may include first decomposing the external clone record into a plurality ofexternal clone record frequency components by Discrete Wavelet Transform (DWT).Following which the external clone record frequency components may be analyzed forthe presence or absence of the watermark in an external clone record low-frequency bandand / or an external clone record mid-frequency band. This analysis may include usingDWT to decompose the external clone record and examine the frequency componentsfor the watermark.

[0109] In still another alternative for generating the one or more watermark records, awatermark may be embedded into a plurality of pixels in an image in the clone record byNeural Network-Based Embedding. Neural Network-Based Embedding referring to alearned low-dimensional representation of discrete data as continuous vectors.

[0110] When generating the one or more watermark records includes embedding awatermark into a plurality of pixels in an image in the clone record by Neural Network-Based Embedding, then evaluating the external clone record for the presence of the oneor more watermark records may include analyzing the external clone record for thepresence or absence of the watermark. This analysis may be conducted utilizing thetrained neural network to detect the presence of the watermark in the external clonerecord.

[0111] In yet a further alternative for generating the one or more watermark records, awatermark may be embedded into a plurality of pixels in an image in the clone record byAutoencoder-Based Watermarking. Autoencoder-Based Watermarking referring to anartificial neural network which learns an encoding function that transforms input datafrom the image and a decoding function that recreates the input data from the encodedrepresentation.

[0112] When generating the one or more watermark records includes embedding awatermark into a plurality of pixels in an image in the clone record by Autoencoder-BasedWatermarking, then evaluating the external clone record for the presence of the one ormore watermark records may include analyzing the external clone record for the presenceor absence of the watermark. This analysis may be conducted by loading into the clonerecord a pre-trained Autoencoder used to embed the watermark to allow the Autoencoderto extract latent features from the external clone record. The latent features may then becompared with the expected watermark pattern. When analyzing a video, this process mayneed to be repeated for a plurality – and in some occasions all – of the frames in thevideo.

[0113] Once the second device generates the one or more storage and access controlrecords and the one or more watermark records, the second device transmits the one ormore watermark records to the third device and submits the one or more watermarkrecords and the one or more storage and access control records to the second distributednetwork. The second distributed network is then capable of receiving a external clonerequest from a third party licensee and receiving a completed license contract from thethird party licensee for storage in the storage and access control records.

[0114] The third device receives an external clone record. Upon receiving the externalclone record, the third device evaluates the external clone record for the presence of theone or more watermark records. The third device then submits the watermark recordsand the external voice clone record to the third distributed network.

[0115] Once the third device identifies the presence – or absence – of the one or morewatermark records in the external clone record, the database records and stores a verifiedtransaction. This verified transaction being indicative of the presence – or absence – ofthe one or more watermark records in the external clone record. The presence of the oneor more watermark records in the external clone record indicates that the external clonerecord is an authentic clone of the photo or video. Conversely, the absence of the one ormore watermark records in the external clone record indicates that the external clonerecord is not an authentic clone of the photo or video.

[0116] When the work is a text, the system may include a first device, a first distributednetwork, a second device, a second distributed network, a third device, a third distributednetwork, and a database. The first device, the second device, and the third device mayeach independently be in the form of a network device selected from the group consistingof a laptop computer, a desktop computer, a server, a mobile device (cellular phone,tablet, or the like), an Internet-of-Things (IoT) device, an endpoint, a virtual machine, acloud based server, a cloud based storage device, and combinations thereof. The firstdistributed network, the second distributed network, and the third distributed networkmay each independently be in the form of a blockchain network. Blockchain referring toa growing list of records, or blocks, that are cryptographically linked. The blockscommonly include a transaction data, a time stamp, and a cryptographic hash of the priorblock in the chain that confirms the integrity of the prior block. Blockchains can be usedto implement a shared digital ledger for recording user transactions between multipledevices – often stored in a cloud-based environment. When a user transaction is executedand validated, it is appended to the end of the blockchain, making the blockchain animmutable history of all valid transactions. The database may be in the form of ablockchain database of executed and validated transactions (or blocks).

[0117] The first device receives a text file in the form of a binary code record. In someembodiments, the first device may generate the binary code from the text file. The firstdevice may generate the binary code record by first converting a plurality of floating-pointnumbers into a binary format. This may be accomplished by quantizing continuous valuesof the embedding to a fixed number of bits to create a quantized embedding.

[0118] The first device further conducts a pre-cryptographic processing on the binarycode record to produce a copied text record. The pre-cryptographic processing mayinclude one or more of text normalization, tokenization, linguistic analysis, featureextraction, and / or metadata attachment. Text normalization may include converting textto a standard format, including case normalization, removal of extra whitespace, andstandardizing punctuation. Tokenization may include splitting the text into tokens (suchas words and / or sentences) for further processing. Linguistic analysis may includeperforming part-of-speech tagging, syntactic parsing, and semantic analysis to understandthe structure and meaning of the text. Feature extraction may include extracting keyfeatures like word frequencies, n-grams, or embeddings using models such as Word2Vecor BERT. Metadata attachment may involve including information such as authorship,creation date, and version history.

[0119] The first device further cryptographically hashes the binary code record and thecopied text record based on a cryptographic algorithm. The cryptographic algorithm maybe in the form of an asymmetric-key algorithm or a hash function. Once the first devicecompletes the cryptographic hashing on the binary code record and the copied textrecord, the first device submits the binary code record and the copied text record to thefirst distributed network.

[0120] The second device generates one or more storage and access control records forthe copied text record. The one or more storage and access control records may be storedin a smart contract. A smart contract is a computer program or a transaction protocolwhich automatically executes, controls, and / or documents events and actions accordingto the terms of a contract or agreement by sending a transaction from a wallet for theblockchain which includes compiled code and a special receiver address, the transactionbeing included in a block added to the blockchain such that the smart contract’s codewill execute to establish the initial state of the smart contract.

[0121] The second device further generates one or more watermark records associatedwith the copied text record. Generating the one or more watermark records may includeidentifying one or more repeated words in the copied text record. Once the repeatedword(s) are identified, at least one of the repeated words may be replaced with a synonymfor the repeated word.

[0122] When generating the one or more watermark records includes identifying arepeated word in the copied text record and replacing it with a synonym, then evaluatingthe external copied text record for the presence of the one or more watermark recordsmay include examining the external copied text record for the presence or absence of thesynonym for the repeated word. This analysis may include analyzing the text for synonymreplacements of repeated words, comparing with the original text patterns.

[0123] Alternatively, the second device may generate the one or more watermark recordsby introducing a patterned alteration in sentence structure to at least one sentence in thecopied text record. For example, a sentence in the copied text record may have itsstructure altered from subject, verb, object to subject, object, vert.

[0124] When generating the one or more watermark records includes introducing apatterned alteration in sentence structure, then evaluating the external copied text recordfor the presence of the one or more watermark records may include examining theexternal copied text record for the presence or absence of the patterned alteration insentence structure. This analysis may include using syntactic parsing to detect alterationsin sentence structure that match the watermarking pattern.

[0125] As a further alternative, the second device may generate the one or morewatermark records by introducing a patterned systematic variation in punctuation and / ora patterned systematic variation in capitalization to the copied text record. For example,commas may be replaced with semi-colons in all or a part of the text.

[0126] When generating the one or more watermark records includes introducing apatterned systematic variation in punctuation and / or capitalization, then evaluating theexternal copied text record for the presence of the one or more watermark records mayinclude examining the external copied text record for the presence or absence of thepatterned systemic variation in punctuation and / or the patterned systemic variation incapitalization. This analysis may include examining the text for systematic variations inpunctuation and capitalization.

[0127] As another alternative, the second device may generate the one or morewatermark records by inserting a contextually appropriate patterned word or phrase intothe copied text record. By “contextually appropriate” it is meant that the word or phrasedoes not alter the overall meaning of the section of text into which it is inserted.

[0128] When generating the one or more watermark records includes inserting acontextually appropriate patterned word or phrase into the copied text record, thenevaluating the external copied text record for the presence of the one or more watermarkrecords may include examining the external copied text record for the presence or absenceof the contextually appropriate patterned word or phrase. This analysis may includelooking for inserted words or phrases that align with the watermark pattern.

[0129] Still another alternative for generating the one or more watermark records mayinclude identifying one or more sentences in the copied text record. Once said sentencesare identified, at least one of the sentences may be replaced with a patterned paraphrasedsentence. In doing so, the paraphrased sentence should not alter the overall meaning ofthe sentence.

[0130] When generating the one or more watermark records includes replacing at leastone sentence with a patterned paraphrased sentence, then evaluating the external copiedtext record for the presence of the one or more watermark records may include examiningthe external copied text record for the presence or absence of the patterned paraphrasedsentence. This analysis may include utilizing semantic similarity measures or paraphrasedetection models to identify paraphrased sentences.

[0131] Yet another alternative for generating the one or more watermark records mayinclude introducing a patterned ambiguity and / or a patterned dual meaning to at leastone portion of the copied text record. The patterned ambiguity or patterned dualmeaning resulting in the portion of the text being open to multiple interpretations.

[0132] When generating the one or more watermark records includes introducing apatterned ambiguity and / or a patterned dual meaning into the copied text record, thenevaluating the external copied text record for the presence of the one or more watermarkrecords may include examining the external copied text record for the presence or absenceof the patterned ambiguity and / or the patterned dual meaning. This analysis may includeidentifying sections with ambiguities or dual meanings introduced as part of thewatermark.

[0133] Once the second device generates the one or more storage and access controlrecords and the one or more watermark records, the second device transmits the one ormore watermark records to the third device, and submits the one or more watermarkrecords and the one or more storage and access control records to the second distributednetwork. The second distributed network is then capable of receiving a text copy requestfrom a third party licensee and receiving a completed license contract from the third partylicensee for storage in the storage and access control records.

[0134] The third device receives an external copied text record. Upon receiving theexternal copied text record, the third device evaluates the external copied text record forthe presence of the one or more watermark records. The third device then submits thewatermark records and the external copied text record to the third distributed network.

[0135] Once the third device identifies the presence – or absence – of the one or morewatermark records in the external copied text record, the database records and stores averified transaction. This verified transaction being indicative of the presence – or absence– of the one or more watermark records in the external copied text record. The presenceof the one or more watermark records in the external copied text record indicates thatthe external copied text record is an authentic copy of the text. Conversely, the absenceof the one or more watermark records in the external copied text record indicates thatthe external copied text record is not an authentic copy of the text.

[0136] The tri-blockchain clone / copy licensing and verification systems and methodsdisclosed herein represent an improvement over prior systems and methods for managingand licensing voice clone data, copies of photos, copies of videos, and / or copies of text.The systems and methods disclosed herein provide robust security measures via themultiple blockchain storage steps. Additionally, transparent licensing is provided via thestorage and access control records which are subject to rapid and reliable verificationusing the watermark system.

[0137] While the systems and methods have been described as having one or moreexemplary embodiments, the present systems and methods may be further modifiedwithin the spirit and scope of this disclosure. This application is therefore intended tocover any variations, uses, or adaptations of the systems and methods using their generalprinciples.

Claims

CLAIMSWhat is claimed is:

1. A system for recording and verifying a voice clone, said system comprising:a first device for receiving a voice recording in a form of a binary coderecord or converting the voice recording into the binary code record,conducting a pre-cryptographic processing on the binary code record toproduce a voice clone record, and cryptographically hashing the binarycode record and the voice clone record based on a cryptographicalgorithm; afirst distributed network to which the binary code record and thevoice clone record are submitted;a second device for generating one or more storage and access controlrecords for the voice clone record, generating one or more watermarkrecords associated with the voice clone record, and transmitting the oneor more watermark records to a third device;a second distributed network to which the one or more watermarkrecords and the one or more storage and access control records aresubmitted; the third device for receiving the one or more watermark records fromthe second device, receiving an external voice clone record, and evaluatingthe external voice clone record for the presence of the one or morewatermark records;a third distributed network to which the watermark records and theexternal voice clone record are submitted; anda database for recording and storing a verified transaction indicativeof the presence of the one or more watermark records in the external voiceclone record.

2. The system of claim 1, wherein the first device conducts a pre-binary code recordprocessing on the voice recording, said pre-binary code record processingcomprising: conducting a noise reduction on the voice recording to reducebackground noise and produce a noise reduced voice recording;conducting a normalization on the noise reduced voice recording toproduce a normalized voice recording having a consistent volume ofspeech; conducting a segmentation on the normalized voice recording to splitthe normalized voice recording into a segmented voice recording;conducting a Short-Time Fourier Transform (STFT) on the segmentedvoice recording to create a spectrogram representative of constituentfrequencies of the segmented voice recording over time;creating a mel-frequency cepstral coefficient from the spectrogram;estimating a pitch of a voice in the spectrogram at one or moreindividual time segments;extracting formants from the spectrogram, said formants representingone or more resonant frequencies of a vocal tract;extracting a duration of one or more of phonemes, syllables, andpauses from the spectrogram;extracting energy patterns representing an amplitude of a speech signalfrom the spectrogram;processing sequential data from the spectrogram to generate a time-series representation of individual voice features; andcreating an embedding of a compact representation of one or moreunique characteristics of a voice.

3. The system of claim 2, wherein the noise reduction is conducted by a filter.

4. The system of claim 2, wherein the noise reduction is conducted by a spectralsubtraction noise reduction algorithm.

5. The system of any of claims 2 to 4, wherein the pitch of the voice is estimated byan autocorrelation.

6. The system of any of claims 2 to 4, wherein the pitch of the voice is estimated bya harmonic product spectrum (HPS).

7. The system of any of claims 2 to 6, wherein the time-series representation ofindividual voice features is generated by a recurrent neural network (RNN).

8. The system of any of claims 2 to 6, wherein the time-series representation ofindividual voice features is generated by a transformer.

9. The system of any of claims 1 to 8, wherein converting the voice recording intothe binary code record comprises:converting a plurality of floating-point numbers into a binary formatby quantizing continuous values of the embedding to a fixed number ofbits to create a quantized embedding;organizing data from the quantized embedding into a binary formatto create a serialized binary file having one or more individual voicesegments associated with one or more corresponding voice features;compressing the serialized binary file; andassociating headers comprising metadata related to the one or moreindividual voice segments.

10. The system of claim 9, wherein compressing the serialized binary file is conductedby Huffman coding.

11. The system of claim 9, wherein compressing the serialized binary file is conductedby a run-length encoding (RLE).

12. The system of any of claims 9 to 11, wherein the metadata includes a metadatarecord corresponding to one or more of a sampling rate, a bit depth, and a lengthof an embedding vector.

13. The system of any of claims 1 to 12, wherein the pre-crytographic processingcomprises: aligning a text input transcription with one or more correspondingaudio features to create a mapping between one or more text segmentsand one or more voice embeddings; andconverting the text input transcription into voice embeddingsextracted and stored in the binary code record.

14. The system of claim 13, wherein aligning the text input transcription is conductedby a Montreal Forced Aligner.

15. The system of any of claims 13 to 14, wherein converting the text inputtranscription is conducted using a sequence-to-sequence model.

16. The system of any of claims 1 to 12, wherein the pre-cryptographic processingcomprises:aligning a text input transcription with one or more correspondingaudio features to create a mapping between one or more text segmentsand one or more voice embeddings; andconverting one or more predicted embeddings into one or more actualspeech waveforms.

17. The system of claim 16, wherein aligning the text input transcription is conductedby a Montreal Forced Aligner.

18. The system of any of claims 16 to 17, wherein converting one or more predictedembeddings is conducted using a vocoder.

19. The system of any of claims 1 to 18, wherein the one or more storage and accesscontrol records are stored in a smart contract.

20. The system of any of claims 1 to 19, wherein the generating one or morewatermark records comprises:embedding one or more audio signals into the voice clone record; andmodulating one or both of the amplitude and / or the phase of at leastone harmonic associated with the one or more audio signals.

21. The system of claim 20, wherein the evaluating the external voice clone record forthe presence of the one or more watermark records comprises filtering theexternal voice clone record for the presence of the one or more audio signals, andcorrelating the one or more audio signals to corresponding audio signals in theexternal voice clone record.

22. The system of any of claims 1 to 19, wherein the generating one or morewatermark records comprises embedding a delayed version of the voice recordinginto the voice clone record wherein the delayed version of the voice recording hasa sound pressure of less than 130 dB.

23. The system of claim 22, wherein the evaluating the external voice clone record forthe presence of the one or more watermark records comprises filtering the voiceclone record for the presence of the delayed version of the voice recording.

24. The system of any of claims 1 to 19, wherein the generating one or morewatermark records comprises:identifying one or more least significant bits (LSBs) within the voicerecording; andaltering at least one parameter of at least one of the LSBs to create oneor more altered LSBs, said at least one parameter selected from the groupconsisting of amplitude and frequency.

25. The system of claim 24, wherein the evaluating the external voice clone record forthe presence of the one or more watermark records comprises examining at leastone parameter of the LSBs within the external voice clone record for the presenceor absence of the one or more altered LSBs.

26. The system of any of claims 1 to 19, wherein the generating one or morewatermark records comprises:conducting a Discrete Fourier Transform (DFT) or a Discrete CosineTransform (DCT) on the voice recording;modifying one or more frequency coefficients within the voicerecording to generate one or more modified frequency coefficients; andembedding the one or more modified frequency coefficients into thevoice clone record.

27. The system of claim 26, wherein the evaluating the external voice clone record forthe presence of the one or more watermark records comprises filtering theexternal voice clone record for the presence of the one or more modifiedfrequency coefficients, and correlating the one or more modified frequencycoefficients to corresponding modified frequency coefficients in the external voiceclone record.

28. A method to record and verify a voice clone, said method comprising:submitting a voice recording in a form of a binary code record to afirst device or converting the voice recording to the binary code recordusing the first device;conducting a pre-cryptographic processing on the binary code recordto produce a voice clone record;providing a first cryptographic algorithm to hash the binary coderecord and the voice clone record;cryptographically hashing the binary code record and the voice clonerecord based on the first cryptographic algorithm;submitting the binary code record and the voice clone record to a firstdistributed network;generating one or more storage and access control records for the voiceclone record;providing a second cryptographic algorithm to hash the storage andaccess control records;cryptographically hashing the storage and access control records basedon the second cryptographic algorithm;generating one or more watermark records and associating the one ormore watermark records with the voice clone record;providing a third cryptographic algorithm to hash the one or morewatermark records;cryptographically hashing the one or more watermark records basedon the third cryptographic algorithm;receiving an external voice clone record;evaluating the external voice clone record for the presence of the oneor more watermark records; andrecording and storing a verified transaction indicative of the presenceof the one or more watermark records in the external voice clone recordin one or more databases.

29. The method of claim 28, further comprising:conducting a noise reduction on the voice recording to reducebackground noise and produce a noise reduced voice recording;conducting a normalization on the noise reduced voice recording toproduce a normalized voice recording having a consistent volume ofspeech; conducting a segmentation on the normalized voice recording to splitthe normalized voice recording into a segmented voice recording;conducting a Short-Time Fourier Transform (STFT) on the segmentedvoice recording to create a spectrogram representative of constituentfrequencies of the segmented voice recording over time;creating a mel-frequency cepstral coefficient from the spectrogram;estimating a pitch of a voice in the spectrogram at one or moreindividual time segments;extracting formants from the spectrogram, said formants representingone or more resonant frequencies of a vocal tract;extracting a duration of one or more of phonemes, syllables, andpauses from the spectrogram;extracting energy patterns representing an amplitude of a speech signalfrom the spectrogram;processing sequential data from the spectrogram to generate a time-series representation of individual voice features; andcreating an embedding of a compact representation of one or moreunique characteristics of a voice; andeach of which being conducted by the first device prior to conducting the pre-cryptographic processing on the binary code record.

30. The method of claim 29, wherein the noise reduction is conducted by a filter.

31. The method of claim 29, wherein the noise reduction is conducted by a spectralsubtraction noise reduction algorithm.

32. The method of any of claims 29 to 31, wherein the pitch of the voice is estimatedby an autocorrelation.

33. The method of any of claims 29 to 31, wherein the pitch of the voice is estimatedby a harmonic product spectrum (HPS).

34. The method of any of claims 29 to 33, wherein the time-series representation ofindividual voice features is generated by a recurrent neural network (RNN).

35. The method of any of claims 29 to 33, wherein the time-series representation ofindividual voice features is generated by a transformer.

36. The method of any of claims 28 to 35, wherein converting the voice recording tothe binary code record comprises:converting a plurality of floating-point numbers into a binary formatby quantizing continuous values of the embedding to a fixed number ofbits to create a quantized embedding;organizing data from the quantized embedding into a binary formatto create a serialized binary file having one or more individual voicesegments associated with one or more corresponding voice features;compressing the serialized binary file; andassociating headers comprising metadata related to the one or moreindividual voice segments; andeach of which being conducted by the first device.

37. The method of claim 36, wherein compressing the serialized binary file isconducted by Huffman coding.

38. The method of claim 36, wherein compressing the serialized binary file isconducted by a run-length encoding (RLE).

39. The method of any of claims 36 to 38, wherein the metadata includes a metadatarecord corresponding to one or more of a sampling rate, a bit depth, and a lengthof an embedding vector.

40. The method of any of claims 28 to 39, wherein the pre-cryptographic processingcomprises: aligning a text input transcription with one or more correspondingaudio features to create a mapping between one or more text segmentsand one or more voice embeddings; andconverting the text input transcription into voice embeddingsextracted and stored in the binary code record.

41. The method of claim 40, wherein aligning the text input transcription isconducted by a Montreal Forced Aligner.

42. The method of any of claims 40 to 41, wherein converting the text inputtranscription is conducted using a sequence-to-sequence model.

43. The method of any of claims 28 to 39, wherein the pre-cryptographic processingcomprises: aligning a text input transcription with one or more correspondingaudio features to create a mapping between one or more text segmentsand one or more voice embeddings; andconverting one or more predicted embeddings into one or more actualspeech waveforms.

44. The method of claim 43, wherein aligning the text input transcription isconducted by a Montreal Forced Aligner.

45. The method of any of claims 43 to 44, wherein converting one or more predictedembeddings is conducted using a vocoder.

46. The method of any of claims 28 to 45, wherein the one or more storage and accesscontrol records are stored in a smart contract.

47. The method of any of claims 28 to 46, wherein the generating one or morewatermark records comprises:embedding one or more audio signals into the voice clone record; andmodulating one or both of the amplitude and / or the phase of at leastone harmonic associated with the one or more audio signals.

48. The method of claim 47, wherein the evaluating the external voice clone recordfor the presence of the one or more watermark records comprises filtering theexternal voice clone record for the presence of the one or more audio signals, andcorrelating the one or more audio signals to corresponding audio signals in theexternal voice clone record.

49. The method of any of claims 28 to 46, wherein the generating one or morewatermark records comprises embedding a delayed version of the voice recordinginto the voice clone record wherein the delayed version of the voice recording hasa sound pressure of less than 130 dB.

50. The method of claim 49, wherein the evaluating the external voice clone recordfor the presence of the one or more watermark records comprises filtering thevoice clone record for the presence of the delayed version of the voice recording.

51. The method of any of claims 28 to 46, wherein the generating one or morewatermark records comprises:identifying one or more least significant bits (LSBs) within the voicerecording; andaltering at least one parameter of the at least one of the LSBs to createone or more altered LSBs, said at least one parameter selected from thegroup consisting of amplitude and frequency.

52. The method of claim 51, wherein the evaluating the external voice clone recordfor the presence of the one or more watermark records comprises examining atleast one parameter of the LSBs within the external voice clone record for thepresence or absence of the one or more altered LSBs.

53. The method of any of claims 28 to 46, wherein the generating one or morewatermark records comprises:conducting a Discrete Fourier Transform (DFT) or a Discrete CosineTransform (DCT) on the voice recording;modifying one or more frequency coefficients within the voicerecording to generate one or more modified frequency coefficients; andembedding the one or more modified frequency coefficients into thevoice clone record.

54. The method of claim 53, wherein the evaluating the external voice clone recordfor the presence of the one or more watermark records comprises filtering theexternal voice clone record for the presence of the one or more modifiedfrequency coefficients, and correlating the one or more modified frequencycoefficients to corresponding modified frequency coefficients in the external voiceclone record.

55. A system for recording and verifying authenticity of a photo or video, said systemcomprising: afirst device for receiving a photo or video file in a form of a binarycode record or converting the photo or video file into the binary coderecord, conducting a pre-cryptographic processing on the binary coderecord to produce a clone record, and cryptographically hashing the binarycode record and the clone record based on a cryptographic algorithm;a first distributed network to which the binary code record and theclone record are submitted;a second device for generating one or more storage and access controlrecords for the clone record, generating one or more watermark recordsassociated with the clone record, and transmitting the one or morewatermark records to a third device;a second distributed network to which the one or more watermarkrecords and the one or more storage and access control records aresubmitted; the third device for receiving the one or more watermark records fromthe second device, receiving an external clone record, and evaluating theexternal clone record for the presence of the one or more watermarkrecords; athird distributed network to which the watermark records and theexternal clone record are submitted; anda database for recording and storing a verified transaction indicativeof the presence of the one or more watermark records in the external clonerecord.

56. The system of claim 55, wherein the generating one or more watermark recordscomprises: identifying one or more least significant bits (LSBs) of a plurality ofpixel values of the clone record; andembedding a watermark image in at least one of the LSBs to createone or more altered LSBs.

57. The system of claim 56, wherein the evaluating the external clone record for thepresence of the one or more watermark records comprises examining at least oneof the LSBs within the external clone record for the presence or absence of thewatermark image.

58. The system of claim 55, wherein the generating one or more watermark recordscomprises modifying a parameter selected from the group consisting of texture,color, and combinations thereof within a plurality of pixels in the clone record.

59. The system of claim 58, wherein the evaluating the external clone record for thepresence of the one or more watermark records comprises examining the pluralityof pixels within the external clone record for the presence or absence of theparameter.

60. The system of claim 55, wherein the generating one or more watermark recordscomprises: transforming the clone record into a frequency domain by DiscreteCosine Transform (DCT); andembedding a watermark into a mid-frequency coefficient of thefrequency domain.

61. The system of claim 60, wherein the evaluating the external clone record for thepresence of the one or more watermark records comprises:transforming the external clone record into an external clone recordfrequency domain by Discrete Cosine Transform (DCT); andanalyzing an external clone record mid-frequency coefficient for thepresence or absence of the watermark.

62. The system of claim 55, wherein the generating one or more watermark recordscomprises: decomposing an image in the clone record into a plurality of frequencycomponents by Discrete Wavelet Transform (DWT); andembedding a watermark into a low-frequency band and / or a mid-frequency band.

63. The system of claim 62, wherein evaluating the external clone record for thepresence of the one or more watermark records comprises:decomposing the external clone record into a plurality of externalclone record frequency components by Discrete Wavelet Transform(DWT); andanalyzing the external clone record frequency components for thepresence or absence of the watermark in an external clone record low-frequency band and / or an external clone record mid-frequency band.

64. The system of claim 55, wherein the generating one or more watermark recordscomprises embedding a watermark in a plurality of pixels in an image in the clonerecord by Neural Network-Based Embedding.

65. The system of claim 64, wherein evaluating the external clone record for thepresence of the one or more watermark records comprises analyzing the externalclone record for the presence or absence of the watermark.

66. The system of claim 55, wherein the generating one or more watermark recordscomprises embedding a watermark in a plurality of pixels in an image in the clonerecord by Autoencoder-Based Watermarking.

67. The system of claim 66, wherein evaluating the external clone record for thepresence of the one or more watermark records comprises analyzing the externalclone record for the presence or absence of the watermark.

68. The system of any of claims 55 to 67, wherein the one or more storage and accesscontrol records are stored in a smart contract.

69. A method to record and verify an image clone or video clone, said methodcomprising: submitting a photo or video file in a form of a binary code record to afirst device or converting the photo or video file to the binary code recordusing the first device;conducting a pre-cryptographic processing on the binary code recordto produce a clone record;providing a first cryptographic algorithm to hash the binary coderecord and the clone record;cryptographically hashing the binary code record and the clone recordbased on the first cryptographic algorithm;submitting the binary code record and the clone record to a firstdistributed network;generating one or more storage and access control records for theclone record;providing a second cryptographic algorithm to hash the storage andaccess control records;cryptographically hashing the storage and access control records basedon the second cryptographic algorithm;generating one or more watermark records and associating the one ormore watermark records with the clone record;providing a third cryptographic algorithm to hash the one or morewatermark records;cryptographically hashing the one or more watermark records basedon the third cryptographic algorithm;receiving an external clone record;evaluating the external clone record for the presence of the one ormore watermark records; andrecording and storing a verified transaction indicative of the presenceof the one or more watermark records in the external clone record in oneor more databases.

70. The method of claim 69, wherein the generating one or more watermark recordscomprises: identifying one or more least significant bits (LSBs) of a plurality ofpixel values of the clone record; andembedding a watermark image in at least one of the LSBs to createone or more altered LSBs.

71. The method of claim 70, wherein the evaluating the external clone record for thepresence of the one or more watermark records comprises examining at least oneof the LSBs within the external clone record for the presence or absence of thewatermark image.

72. The method of claim 69, wherein the generating one or more watermark recordscomprises modifying a parameter selected from the group consisting of texture,color, and combinations thereof within a plurality of pixels in the clone record.

73. The method of claim 72, wherein the evaluating the external clone record for thepresence of the one or more watermark records comprises examining the pluralityof pixels within the external clone record for the presence or absence of theparameter.

74. The method of claim 69, wherein the generating one or more watermark recordscomprises: transforming the clone record into a frequency domain by DiscreteCosine Transform (DCT); andembedding a watermark into a mid-frequency coefficient of thefrequency domain.

75. The method of claim 74, wherein the evaluating the external clone record for thepresence of the one or more watermark records comprises:transforming the external clone record into an external clone recordfrequency domain by Discrete Cosine Transform (DCT); andanalyzing an external clone record mid-frequency coefficient for thepresence or absence of the watermark.

76. The method of claim 69, wherein the generating one or more watermark recordscomprises: decomposing an image in the clone record into a plurality of frequencycomponents by Discrete Wavelet Transform (DWT); andembedding a watermark into a low-frequency band and / or a mid-frequency band.

77. The method of claim 76, wherein evaluating the external clone record for thepresence of the one or more watermark records comprises:decomposing the external clone record into a plurality of externalclone record frequency components by Discrete Wavelet Transform(DWT); andanalyzing the external clone record frequency components for thepresence or absence of the watermark in an external clone record low-frequency band and / or an external clone record mid-frequency band.

78. The method of claim 69, wherein the generating one or more watermark recordscomprises embedding a watermark in a plurality of pixels in an image in the clonerecord by Neural Network-Based Embedding.

79. The method of claim 78, wherein evaluating the external clone record for thepresence of the one or more watermark records comprises analyzing the externalclone record for the presence or absence of the watermark.

80. The method of claim 69, wherein the generating one or more watermark recordscomprises embedding a watermark in a plurality of pixels in an image in the clonerecord by Autoencoder-Based Watermarking.

81. The method of claim 80, wherein evaluating the external clone record for thepresence of the one or more watermark records comprises analyzing the externalclone record for the presence or absence of the watermark.

82. The method of any of claims 69 to 81, wherein the one or more storage and accesscontrol records are stored in a smart contract.

83. A system for recording and verifying authenticity of a text, said system comprising:a first device for receiving a text file in a form of a binary code recordor converting the text file to the binary code record, conducting a pre-cryptographic processing on the binary code record to produce a copiedtext record, and cryptographically hashing the binary code record and thecopied text record based on a cryptographic algorithm;a first distributed network to which the binary code record and thecopied text record are submitted;a second device for generating one or more storage and access controlrecords for the copied text record, generating one or more watermarkrecords associated with the copied text record, and transmitting the oneor more watermark records to a third device;a second distributed network to which the one or more watermarkrecords and the one or more storage and access control records aresubmitted;a third device for receiving the one or more watermark records fromthe second device, receiving an external copied text record, and evaluatingthe external copied text record for the presence of the one or morewatermark records;a third distributed network to which the watermark records and theexternal copied text record are submitted; anda database for recording and storing a verified transaction indicativeof the presence of the one or more watermark records in the externalcopied text records.

84. The system of claim 83, wherein the generating one or more watermark recordscomprises: identifying one or more repeated words in the copied text record; andreplacing at least one of the repeated words with a synonym for therepeated word.

85. The system of claim 84, wherein the evaluating the external copied text record forthe presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the synonym for therepeated word.

86. The system of claim 83, wherein the generating one or more watermark recordscomprises introducing a patterned alteration in sentence structure to at least onesentence in the copied text record.

87. The system of claim 86, wherein the evaluating the external copied text record forthe presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the patterned alterationin sentence structure.

88. The system of claim 83, wherein the generating one or more watermark recordscomprises introducing a patterned systematic variation in punctuation and / or apatterned systematic variation in capitalization to the copied text record.

89. The system of claim 88, wherein the evaluating the external copied text record forthe presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the patterned systemicvariation in punctuation and / or the patterned systemic variation in capitalization.

90. The system of claim 83, wherein the generating one or more watermark recordscomprises inserting a contextually appropriate patterned word or phrase into thecopied text record.

91. The system of claim 90, wherein the evaluating the external copied text record forthe presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the contextuallyappropriate patterned word or phrase.

92. The system of claim 83, wherein the generating one or more watermark recordscomprises: identifying one or more sentences in the copied text record; andreplacing at least one of the sentences with a patterned paraphrasedsentence.

93. The system of claim 92, wherein the evaluating the external copied text record forthe presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the patternedparaphrased sentence.

94. The system of claim 83, wherein the generating one or more watermark recordscomprises introducing a patterned ambiguity and / or a patterned dual meaning toat least one portion of the copied text record.

95. The system of claim 94, wherein the evaluating the external copied text record forthe presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the patterned ambiguityand / or the patterned dual meaning.

96. The system of any of claims 83 to 95, wherein the one or more storage and accesscontrol records are stored in a smart contract.

97. A method to record and verify a copied text, said method comprising:submitting a text file in a form of a binary code record to a first deviceor converting the text file to the binary code record using the first device;conducting a pre-cryptographic processing on the binary code recordto produce a copied text record;providing a first cryptographic algorithm to has the binary code recordand the copied text record;cryptographically hashing the binary code record and the copied textrecord based on the first cryptographic algorithm;submitting the binary code record and the copied text record to a firstdistributed network;generating one or more storage and access control records for thecopied text record;providing a second cryptographic algorithm to has the storage andaccess control records;cryptographically hashing the storage and access control records basedon the second cryptographic algorithm;generating one or more watermark records and associating the one ormore watermark records with the copied text record;providing a third cryptographic algorithm to has the one or morewatermark records;cryptographically hashing the one or more watermark records basedon the third cryptographic algorithm;receiving an external copied text record;evaluating the external copied text record for the presence of the oneor more watermark records; andrecording and storing a verified transaction indicative of the presenceof the one or more watermark records in the external copied text recordin one or more databases.

98. The method of claim 97, wherein the generating one or more watermark recordscomprises: identifying one or more repeated words in the copied text record; andreplacing at least one of the repeated words with a synonym for therepeated word.

99. The method of claim 98, wherein the evaluating the external copied text recordfor the presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the synonym for therepeated word.

100. The method of claim 97, wherein the generating one or more watermark recordscomprises introducing a patterned alteration in sentence structure to at least onesentence in the copied text record.

101. The method of claim 100, wherein the evaluating the external copied text recordfor the presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the patterned alterationin sentence structure.

102. The method of claim 97, wherein the generating one or more watermark recordscomprises introducing a patterned systematic variation in punctuation and / or apatterned systematic variation in capitalization to the copied text record.

103. The method of claim 102, wherein the evaluating the external copied text recordfor the presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the patterned systemicvariation in punctuation and / or the patterned systemic variation in capitalization.

104. The method of claim 97, wherein the generating one or more watermark recordscomprises inserting a contextually appropriate patterned word or phrase into thecopied text record.

105. The method of claim 104, wherein the evaluating the external copied text recordfor the presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the contextuallyappropriate patterned word or phrase.

106. The method of claim 97, wherein the generating one or more watermark recordscomprises: identifying one or more sentences in the copied text record; andreplacing at least one of the sentences with a patterned paraphrasedsentence.

107. The method of claim 106, wherein the evaluating the external copied text recordfor the presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the patternedparaphrased sentence.

108. The method of claim 97, wherein the generating one or more watermark recordscomprises introducing a patterned ambiguity and / or a patterned dual meaning toat least one portion of the copied text record.

109. The method of claim 108, wherein the evaluating the external copied text recordfor the presence of the one or more watermark records comprises examining theexternal copied text record for the presence or absence of the patterned ambiguityand / or the patterned dual meaning.

110. The method of any of claims 97 to 109, wherein the one or more storage andaccess control records are stored in a smart contract.

Citation Information

Patent Citations

  • Image work copyright protection method and device, electronic equipment and storage medium

    CN114359010A

  • Video copyright control method and device and medium

    CN117412145A

  • Apparatus and method for transmitting secure and / or copyrighted digital video broadcasting data over internet protocol network

    US20090132825A1

  • Method and apparatus for controlling digital evidence

    US20200294163A1

  • System, method, and computer program for secure authentication of live video

    US20210344498A1