Method for intelligently modifying digital content
The method uses machine-trained models to dynamically modify digital content with non-deterministic rules, addressing the limitations of existing watermarking systems by ensuring each content distribution is uniquely identifiable and resistant to attacks, facilitating efficient proof of infringement.
Patent Information
- Application Number
- PCT/IB2025/057974
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-16
- Filing Date
- 2025-08-05
- Publication Date
- 2026-03-19
AI Technical Summary
Existing digital content watermarking methods are easily circumvented, lack adaptability to various content formats, and are impractical for enforcing intellectual property rights due to their deterministic nature and reliance on binary structure, making them ineffective for textual searches and prone to collusive attacks.
A method utilizing machine-trained models, such as neural networks or transformer models, to dynamically and semantically modify digital content with non-deterministic rules, ensuring the content is uniquely identifiable and resistant to watermark removal techniques, with tracking and validation steps to verify infringement.
The method provides robust, adaptable, and imperceptible modifications that make digital content uniquely identifiable, resistant to attacks, and facilitates efficient proof of infringement by ensuring each content distribution has distinct modifications, enabling rapid identification of illicit copies.
Smart Images

Figure IB2025057974_19032026_PF_FP_ABST
Abstract
Description
[0001] DESCRIPTION
[0002] METHOD FOR INTELLIGENTLY MODIFYING DIGITAL CONTENT
[0003] Technical field
[0004] The present invention relates to a method for the intelligent modification of a digital content in order to make the digital content itself unique, a related system, and a related computer program.
[0005] State of the art
[0006] In the field concerning the production of digital content, it is known that such content is protected by various legal frameworks which, depending on the generated digital content, may guarantee protections of different nature. In particular, in addition to the well-known system of protection of technical creations through patents, other instruments are known such as copyright and trade secret protection, which in Italy is protected by Art. 98 of the Industrial Property Code (CPI), but which, in other jurisdictions, also present substantially comparable regulations.
[0007] A common issue when attempting to enforce the aforementioned rights is that of demonstrating that the content of an alleged infringement actually derives from the proprietary content of the creator. This applies both to copyright, where it is necessary to prove being the first creator of the work, and especially to trade secrets, where, in the case of massive archives of secrets, it is very complex to identify correspondences between files that can demonstrate, quickly and effectively, the actual derivation of the infringing material from the proprietary material.
[0008] Likewise, in the field of trademarks and distinctive signs, it is important to verify whether any trademarks used online or also in secondary markets are authorized or illicitly used.
[0009] Solutions are known, commonly referred to as watermarking procedures, which allow the insertion of additional content within digital contents, indicating that such contents are proprietary and can eventually be recognized later.
[0010] However, such solutions suffer from disadvantages that, also in light of the evolution of artificial intelligence, are becoming increasingly relevant.
[0011] In particular, the watermarking that is carried out is additional with respect to the digital content itself and, as such, it is very easy to identify it in the document by watermark search software. Furthermore, very often, in the copying operations carried out by alleged copiers, it is plausible that those who are aware of potential watermarking simply copy the digital content into a new document, devoid of watermarking.
[0012] Moreover, watermarking has a static and repeatable nature and, as such, once the criterion with which it is carried out is identified (in the case of hidden watermarking), it is quite easy to find it also in other documents.
[0013] Another important disadvantage lies in the identification of correspondence when it is necessary to prove infringement. Current watermarking systems are not suitable for textual searches such as those used in most litigations of this nature.
[0014] This makes current systems easily circumvented and therefore of little protective value. Furthermore, current solutions are impractical for the subsequent and possible enforcement of rights.
[0015] A method for watermarking is illustrated, for example, in document US2008028474A1 , which discloses a system for watermarking digital contents, in particular software, aimed at protecting copyright and countering unauthorized distribution. The proposed techniques are based on a combination of static and dynamic watermarking, code obfuscation, anti-debugging techniques, self-checking, and customization of each distributed copy. A model called “priming and stamping” is introduced, in which the software is initially prepared (priming) with container structures for the watermark, and subsequently marked (stamping) with specific data for each user. Detection methodologies for watermarks even on modified programs are also proposed, based on the statistical analysis of known fragments, as well as collusion-resistant coding algorithms. The document includes models for the secure distribution of software and mechanisms for verifying the integrity of markings even under attack conditions. Finally, the robustness of the system is analyzed through probabilistic formalisms, with explicit reference to collusive attacks and multi-level coding strategies.
[0016] The present document shows the following disadvantages: dependence on the binary structure of the content because the system is mainly designed for software (executables), resulting in poor adaptability to other formats (audio, video, images, natural texts); poor semantics of the watermark, since the modifications are at bit-level or instruction-level, without relation to the meaning of the content; limited cross-media flexibility, because the analyzed system is hardly adaptable to contents other than software.
[0017] Summary of the invention
[0018] The purpose of the present invention is to provide a method for the intelligent modification of a digital content, a system, and a computer program that solve the problems of the prior art mentioned above.
[0019] Said purpose is fully achieved by the method, the system, and the computer program objects of the present invention, which are characterized by what is contained in the claims below.
[0020] According to one aspect of the present description, the present invention provides a method for the intelligent modification of a digital content in order to make the digital content itself unique.
[0021] The method comprises a step of receiving input data, representative of the proprietary digital content.
[0022] The method comprises a step of evaluating one or more modification rules, identifying a rule with which to modify the input data. Said evaluation step is preferably carried out on the basis of the input data themselves. Said evaluation step may lead to the selection of one or more modification rules to be adopted.
[0023] In one embodiment, said one or more modification rules are selected from a machine-trained model, which takes into account the context of the multimedia content and the semantics. In one embodiment, said one or more modification rules are updated on the basis of the context and semantics of the multimedia content.
[0024] In this way, the modification rule is not deterministic and cannot be derived with reverse engineering techniques to circumvent the watermarking, making it much more robust. The non-deterministic nature of artificial intelligence models makes the watermarking difficult to detect.
[0025] Preferably, the models for selecting and updating modification rules are implemented through neural networks, transformer models, or sequential models trained on annotated datasets for semantic understanding of the content.
[0026] This approach allows injecting different modifications for each content even with the same structure, since the selection of rules depends on the contextual and semantic understanding performed by the model.
[0027] This guarantees high resistance to watermark removal or spoofing techniques, even in scenarios where the attacker has multiple modified versions of the content, making differential analysis typical of collusive attacks impracticable.
[0028] The method comprises a step of identifying modification data, represented by a group of input data. In other words, the modification data are parts of the input data, which are selected as the data to be modified. Their choice is preferably carried out on the basis of the input data themselves, according to one or more selection criteria.
[0029] The method comprises a step of applying said one or more modification rules to said modification data, so as to generate modified data.
[0030] The method comprises a step of replacing, within the input data, the modification data with the modified data, so as to generate unique input data.
[0031] The method comprises a step of making available the unique input data.
[0032] The method proposed above allows the digital content itself to be intelligently modified, according to appropriate criteria, so as to make it unique because characterized by modifications of the content that are present without a reason linked to the information to be conveyed but only to identify the owner of the digital content. In this case, therefore, the modification assumes the function of a signature of the owner himself.
[0033] In particular, according to one embodiment, the method of the present invention provides that the modification data are established through the use of a machine-trained model. Furthermore, the method provides that the modified data are determined through a machine-trained model.
[0034] By way of example, but not limited thereto, the machine-trained model may be a LLM (Large Language Model) that identifies portions of the multimedia content suitable to be modified and then modifies them appropriately, considering the context and semantics.
[0035] Such an advanced watermarking system provides that the modifications (understood as intelligent errors) are not inserted according to a deterministic scheme, but are dynamically generated by an artificial intelligence, which adapts them to the content on the basis of its semantic and contextual structure.
[0036] This artificial intelligence evaluates the meaning of the digital content and selects the portions most suitable for modification, adopting minimal and imperceptible alterations for a human being, but sufficiently structured to make the content unique in a robust way and not detectable by conventional analysis techniques.
[0037] The non-deterministic and semantic-contextual nature of the error makes it possible to create different watermarks for each content or even for each distribution of the content, making techniques of detection based on binary comparison or bit-level patterns, as employed in prior art systems, inapplicable.
[0038] The system also takes into account the perceptual levels of the human user, adapting the injection of errors so as to be completely transparent to the enjoyment of the content, while maintaining high resistance to collusive, compressive, or transformation attacks.
[0039] Therefore, the insertion of defects in the digital content, intelligent so as not to alter its informational content, makes them unique and unequivocally linked to the owner. This ensures the possibility, in the future, of identifying such uniqueness in third-party documents, also identifying, in that case, other documents in which the content has been copied (for example to avoid watermarking).
[0040] Preferably, the method comprises a tracking step.
[0041] The tracking step comprises a step of associating an identification code with the digital content.
[0042] The tracking step comprises a step of saving, in an archive, the modified data in association with the unique identification code.
[0043] The tracking step is very important because it allows the tracking of the modifications carried out, which facilitates their identification in future cases in which it may be necessary to seek evidence or proof of infringement.
[0044] Preferably, in the saving step, the modified data are associated with a timestamp, in order to certify the date on which said input data were modified by integrating the modified data. Alternatively, the modified data are saved in a decentralized archive, in which the saving date has a certification value, given the intrinsic characteristics of decentralized archives (for example, blockchain).
[0045] In this way, it is also possible to certify a date on which the possible modification occurred, so as to overcome objections of modification subsequent to the discovery of the allegedly infringing digital content.
[0046] In one embodiment, the method comprises a plurality of intelligent modification steps, temporally spaced apart and each including respective different modification data and / or respective different modification rules.
[0047] Each intelligent modification step is associated with a respective timestamp or is also saved in a decentralized database, to certify the respective date of execution.
[0048] The multiplicity of intelligent modifications with their respective dates of execution also makes it possible to identify in which period the misappropriation of files was possibly carried out. This information, in certain situations, may be useful to reconstruct events and support the arguments of the injured party.
[0049] The method may also comprise a reporting step, in which print data are generated, representing a list including at least pairs of values. Each pair includes the identification code of the proprietary digital content and the related modified data.
[0050] In one embodiment, the method comprises a validation step. The validation step comprises a step of receiving verification data, representative of an external digital content (i.e. identified by a third party). The validation step comprises a step of analyzing the verification data and searching, within the verification data, for one or more modified data saved in the archive. The validation step comprises a step of identifying a number (i.e. even just a presence) of correspondences between the verification data and the modified data saved in the archive. Each of said correspondences is associated with a corresponding identification code.
[0051] For each digital content, the evaluation step comprises a step of diagnosing an illicit derivation. The step of illicit derivation is identified if:
[0052] (a) at least one correspondence is identified between the verification data and the modified data saved in the archive;
[0053] (b) a number of correspondences greater than a predetermined number of correspondences is identified.
[0054] With the validation step, it is possible to quickly validate the allegedly misappropriated content, since the software verifies the intelligent errors that have been saved and, if it identifies one or more error correspondences, determines the presence of a derived content, also indicating the identification code of the document that was used for the derivation.
[0055] In one embodiment, the validation step is a targeted validation step. The targeted validation step comprises a step of receiving verification data, representative of an external digital content. The targeted validation step comprises a step of receiving a comparison code, representative of a unique code of a digital content presumed to have been used. In one embodiment, the receiving step is a step of receiving one or more comparison codes, each representative of a unique code of a respective digital content presumed to have been used.
[0056] The targeted validation step comprises a step of analyzing the verification data and searching, within the verification data, for one or more modified data saved in the archive in association with the entered comparison code (or the entered comparison codes).
[0057] The targeted validation step comprises a step of identifying the number of correspondences between the verification data and the modified data saved in the archive in association with the entered comparison code (or the entered comparison codes).
[0058] The targeted validation step comprises a step of diagnosing an illicit derivation in the event that:
[0059] (a) at least one correspondence is identified between the verification data and the modified data saved in the archive in association with the entered comparison code (or the entered comparison codes);
[0060] (b) a number of correspondences is identified between the verification data and the modified data saved in the archive in association with the entered comparison code (or the entered comparison codes) greater than a predetermined number of correspondences.
[0061] The targeted validation step makes it possible to search only in a subgroup of saved modified data, when there are specific suspicions about the identified documents, also limiting the effort and computational power required for the search.
[0062] In such targeted validation steps, a keyword search (preferably multiple) may also be envisaged, in which the external digital content is indexed and in which queries including the modified data are carried out, which define unique strings presumably found in the external digital content (i.e. under validation). In one embodiment, the method comprises a counterfeiting control step. The counterfeiting control step comprises a step of receiving verification data, representative of an external digital content. The method comprises a step of receiving a comparison code (or more comparison codes), representative of a unique code (or more unique codes) of a digital content (or more digital contents) presumed to have been used.
[0063] The method comprises a step of analyzing the verification data and searching, within the verification data, for one or more modified data saved in the archive in association with the entered comparison code (or the entered comparison codes).
[0064] The method comprises a step of identifying the number of correspondences between the verification data and the modified data saved in the archive in association with the entered comparison code (or the entered comparison codes).
[0065] The method comprises a step of identifying a counterfeiting of the external digital content with respect to the proprietary digital content associated with the comparison code, for a number of correspondences between the verification data and the modified data saved in the archive in association with the entered comparison code (or the entered comparison codes) lower than a predetermined number of correspondences or in case of absence of correspondence between the verification data and the modified data saved in the archive in association with the entered comparison code (or the entered comparison codes).
[0066] The method can therefore serve as an important counterfeiting tool, using an opposite logic. In fact, if the presence of the modification data identifies a correspondence with a proprietary document, the total absence of modification data on a photo of a trademark or a distinctive sign in general indicates that the one using it has illicitly produced this distinctive sign. For a practical example, the modification data of a logo could be represented by specific pixels of the logo image that have a certain color different from the surrounding one. If such specific color in such specific pixel were not detected in analogous logos, this would constitute proof that such logo is not the one authorized and lawfully disclosed by the owner. This verification can therefore be carried out quickly with the method of the present invention.
[0067] In one embodiment, the method comprises a prior art control step. The prior art control step comprises, following the generation of the modified data, a step of searching for correspondences between the generated modified data and the modified data already saved in the archive. In other words, it is verified that identical modified data already saved in the archive are not present, so as to maintain each modified data as a unique string.
[0068] In case of positive outcome, represented by the presence of duplicate modified data among the generated modified data corresponding to the modified data already saved in the archive, the method comprises the following steps:
[0069] (a) removal of said duplicate modified data;
[0070] (b) new generation of substitute modified data;
[0071] (c) execution of the prior art control step on the substitute modified data.
[0072] Therefore, the method provides for iterative generation until the substitute modified data satisfy the prior art verification.
[0073] Therefore, in the case of a negative outcome of the prior art verification step, represented by the absence of duplicate modified data, the method comprises a step of continuing the method with the replacement of the modification data of the input data with the modified data (first generation or substitute) that have passed the prior art control step.
[0074] This step makes it possible to have modification data that are unique and therefore even more focused and allow, with greater certainty, to identify the misappropriated document and not other documents that merely constitute noise in the search and prolong search times.
[0075] In one embodiment, the modified data saved in the archive are associated with a proprietary code, uniquely associated with an owner of the digital content. In this case, in the prior art control step, said step of searching for correspondences is carried out on the modified data already saved in the archive in association with a specific proprietary code.
[0076] This makes it possible to have a relative uniqueness, i.e. a uniqueness with respect to a specific owner, since absolute uniqueness may not be necessary and therefore the process of generation and control of modification data becomes more agile and less burdensome.
[0077] In one embodiment, the input data are representative of one or more of the following multimedia contents:
[0078] • texts;
[0079] • images;
[0080] • audio tracks;
[0081] • data archives;
[0082] • source code of computer programs.
[0083] In one embodiment, the modification data include one or more of the following data:
[0084] • one or more text characters of a text;
[0085] • one or more pixels of an image;
[0086] • one or more audio windows of an audio track;
[0087] • one or more attributes of tables of an archive;
[0088] • one or more instructions of source code.
[0089] In one embodiment, said one or more modification rules include one or more of the following rules:
[0090] • replacement of a text character with a text character graphically similar;
[0091] • addition of text characters representative of a content devoid of logical function or real relevance with respect to the text in which it is included;
[0092] • replacement of a punctuation character with a punctuation character graphically similar but improper in the position of the replaced punctuation character; • replacement of a text character representing a blank space with multiple text characters representing a blank space;
[0093] • repetition of a text character where unnecessary;
[0094] • modification of the color of a single pixel in the image;
[0095] • imperceptible variation of the color of an image area, through variation of the opacity linked to the color;
[0096] • insertion of instructions in the source code not linked to the function of the code itself;
[0097] • variation of frequency or amplitude of the sound in a specific window of the audio track.
[0098] In one embodiment, the method comprises a step of selecting modification data. Said step of selecting modification data comprises a step of applying selection rules, configured to receive the input data as input and to provide, as output, a subset of the input data that defines the modification data.
[0099] In one embodiment, the selection rules comprise one or more of the following rules (which are exemplary and not exhaustive examples of usable selection rules):
[0100] • modification data corresponding to strings of prepositions;
[0101] • modification data corresponding to strings of punctuation;
[0102] • modification data corresponding to a string of text repeated more than a minimum number of occurrences;
[0103] • modification data representing pixels close to the edge of an image. According to one aspect of the present description, the present invention provides a system for the intelligent modification of a digital content in order to make the digital content itself unique, comprising a processor configured to execute the steps of the method according to any of the steps described in the method of the present invention.
[0104] According to one embodiment of the present system, the processor is a processor on a remote server configured to:
[0105] • receive the input data from a first terminal; • process the input data to generate the unique input data;
[0106] • send the unique input data to the first terminal;
[0107] • delete the unique input data from the temporary memory of the remote server.
[0108] In this way, the confidentiality of the processed data is ensured, since they are not saved in places other than the owner’s terminal, except for the time necessary to execute the method.
[0109] In one embodiment, the present invention provides a computer program including instructions to execute the steps of the method including any of the steps described in the present invention.
[0110] Brief description of drawings
[0111] These and other features will be further highlighted by the following description of a preferred embodiment, illustrated purely by way of nonlimiting example in the attached drawing tables, in which:
[0112] • Figure 1 illustrates an embodiment of a method for the intelligent modification of a digital content in order to make the digital content itself unique according to the present invention;
[0113] • Figure 2 illustrates a further embodiment of a method for the intelligent modification of a digital content in order to make the digital content itself unique according to the present invention;
[0114] • Figure 3 illustrates a further embodiment of a method for the intelligent modification of a digital content in order to make the digital content itself unique according to the present invention;
[0115] • Figure 4 illustrates a further embodiment of a method for the intelligent modification of a digital content in order to make the digital content itself unique according to the present invention;
[0116] • Figure 5 illustrates a further embodiment of a method for the intelligent modification of a digital content in order to make the digital content itself unique according to the present invention.
[0117] Detailed description of the invention
[0118] With reference to the attached figures, a block diagram of the method for the intelligent modification of a digital content is illustrated. In the following, two indicative examples of the method will be described in detail without thereby limiting the nature of the digital contents that can be processed with the present method, which, as recalled, range from text contents, image contents, audio and video contents, contents relating to software code, and other contents defined by a set of digital data, which, with reference to the attached figures, are indicated with reference D1. Another example that is certainly pertinent to such method concerns technical construction drawings created for the manufacture of mechanical parts, electrical diagrams, or others. These can be considered equivalent to text documents because, in addition to the graphic part, which cannot be modified for obvious reasons, they include text specifications that can instead be appropriately modified without losing their informational content.
[0119] The method will therefore be described both in the case of digital content corresponding to text files and in the case of image content.
[0120] Therefore, the method comprises a step of receiving F1 input data D1 that identify, i.e. represent, the text that one wishes to make unique or the image that one wishes to make unique.
[0121] The method provides for a step of evaluating F2 one or more modification rules RM. The evaluation step F2 of the modification rules RM is carried out on the basis of the input data themselves, i.e. from the nature of the text and / or the image.
[0122] More in detail, the method comprises a set of modification rules which may be one or more of the following:
[0123] • replacement of a text character with a text character graphically similar;
[0124] • addition of text characters representative of a content devoid of logical function or real relevance with respect to the text in which it is included;
[0125] • elimination of a normally double letter in the correct word; • replacement of a punctuation character with a punctuation character graphically similar but improper in the position of the replaced punctuation character;
[0126] • replacement of a text character representing a blank space with multiple text characters representing a blank space;
[0127] • repetition of a text character where unnecessary;
[0128] • modification of the color of a single pixel in the image;
[0129] • imperceptible variation of the color of an image area, through variation of the opacity linked to the color;
[0130] • variation of frequency or amplitude of the sound in a specific window of the audio track.
[0131] Therefore, starting from this list, the method can select only some of said modification rules to define the modification rules RM that it will use for the method.
[0132] Such identification, indeed, depends on the input data themselves. Firstly, it depends on the type of digital content data (text, image, code), but other selection parameters may also be used, for example, but not limited to:
[0133] • length of the text;
[0134] • subject matter of the text;
[0135] • presence and / or number of nouns, prepositions, adverbs, and punctuation;
[0136] • index of modifiability of the document (i.e. a parameter entered by the user that indicates an expected level of modifiability of the document, for example high, medium, or low);
[0137] • chromatic variety of the image;
[0138] • percentage of the image background with respect to the image foreground.
[0139] These are just some examples of parameters that can be used to understand which rules to use. It is also part of the present invention a solution in which the identification is not necessarily a selection, but simply the rules available to the processor are fixed. The method then includes a step of identifying F3 modification data D2. The modification data D2 are a subset of the input data D1. In this case, therefore, the identification step F3 is a selection step.
[0140] The selection step F3 is regulated by selection rules RS which, on the basis of the input data D1 , make it possible to identify corresponding modification data D2.
[0141] By way of purely exemplary but not exhaustive example, the selection rules comprise one or more of the following rules:
[0142] • modification data corresponding to strings of prepositions;
[0143] • modification data corresponding to strings of punctuation;
[0144] • modification data corresponding to a string of text repeated more than a minimum number of occurrences;
[0145] • modification data corresponding to words having a length greater than a minimum length;
[0146] • modification data representing pixels close to the edge of an image;
[0147] • modification data representing pixels surrounded by pixels having the same color.
[0148] The identification step F3 of the modification data D2 is preferably also carried out on the basis of the selected modification rules RM. Preferably, the selected modification rules RM are greater than 1 , preferably greater than 2 (to allow a difficult determination of the unique modification criterion).
[0149] Following the application of the selection rules RS to the input data D1 , the modification data D2 are obtained.
[0150] Therefore, at this point, the processor includes the modification data D2 and the modification rules RM to be applied to them.
[0151] Let us see it applied to our examples. In the case of text, the processor may have identified an associative array in which each element comprises, for example, a position value in the text of the modification data D2 and a value representing the string of the modification data D2.
[0152] Modification data D2 Text case: [“position”:5, “value”:”ferro”; “position”:”3”, “value”:”:”]
[0153] Image case: [“positional 0,100), “value”:”rgb(10,10,10)”;
[0154] “position”:”40,300”, “value”:”rgba(100,232,344, 0.5)”] Modification rules:
[0155] Text case: [elimination of a normally double letter in the correct word, replacement of a punctuation character with a punctuation character graphically similar but improper in the position of the replaced punctuation character]
[0156] Image case: [modification of the color of a single pixel in the image; imperceptible variation of the color of an image area, through variation of the opacity linked to the color]
[0157] Once the modification data D2 and the corresponding modification rules RM have been identified, the method provides for a step of applying F4 said one or more modification rules to said modification data D2, so as to generate modified data D2’.
[0158] Thus, after the application step F4, the following data are obtained with respect to the practical examples illustrated here:
[0159] Modified data D2’
[0160] Text case: [“position”:5, “value”:”fero”; “position”:”3”, “value”:”;”]
[0161] Image case: [“positional 0,100), “value”:”rgb(11 ,11 ,11 )”;
[0162] “position”:”40,300”, “value”:”rgba(100,232,344, 0.7)”]
[0163] Once the modified data D2’ are obtained, the method provides that these are replaced, within the input data D1 , in place of the corresponding modification data D2. This is possible because the array saves the position of the corresponding modification data D2. The method therefore comprises a replacement step F5 of the modification data D2 with the modified data D2’, so as to generate unique input data D1 ’.
[0164] These modified input data DT, which define a unique digital content, are made available F6, so that the user can circulate the digital content thus made unique and “signed” with the appropriate and intelligent errors.
[0165] For a significant efficiency improvement of the method, a tracking step F7 of the above is also provided.
[0166] More in detail, when the processor receives the input data D1 , the latter associates F71 with the input data D1 a unique code Cl, which is uniquely associated with the digital content.
[0167] Furthermore, the processor saves F72 the modified data D2’ in association with the unique identifier Cl.
[0168] The archive may be a local or remote archive, in any case connectable to the processor to exchange data with normal CRUD operations (Create, Read, Update, Delete).
[0169] This step is very important for future checks to be carried out on digital contents potentially misappropriated illicitly.
[0170] According to a further advantageous embodiment, when the modified data D2’ are saved in the archive, the processor associates a timestamp with such saving that certifies that, from that moment onwards, the document contains such modified data D2’. This allows not only to identify a possible misappropriation but also to identify at least a period in which this occurred (i.e. after the timestamp).
[0171] In this regard, it is particularly relevant to note that the method also provides for a multiple intelligent modification step, in which the injection of the intelligent error is updated over time. In particular, there are a plurality of intelligent modification steps, temporally spaced apart and each including respective different modified data D2’ and / or respective different modification rules RM.
[0172] This makes it possible to increase the resilience of the method, preventing the modification criteria from being identified and the injected error from being easily identified.
[0173] Preferably, each intelligent modification step is associated with a respective timestamp.
[0174] In this way, a tighter temporal tracking is obtained, because certain modified data D2’ are associated with each period. Therefore, by finding the misappropriated document with its modified data D2’, it is possible to determine exactly the period of misappropriation.
[0175] The method provides for a reporting step F8, in which print data DS are generated, representative of a list including at least pairs of values, each pair including the identification code Cl of the proprietary digital content and the related modified data D2’.
[0176] In one embodiment, the method provides that in the saving step F72, the processor also associates with the modified data D2’ an identification code of the owner of the digital content. This makes it possible to facilitate the subsequent validation and infringement identification steps as will be better explained below.
[0177] To increase the reliability of the method and its results, the method comprises a prior art control step F9.
[0178] The prior art control step provides, following the generation of the modified data D2’, a search F91 for correspondences between the generated modified data D2’ and the modified data already saved D2S in the archive. This verification makes it possible to avoid the repetition of errors that could, for example, create false positives in the checks, i.e. cases where the error corresponds but is not the same original digital content.
[0179] At the outcome of this verification, in case of positive outcome, i.e. in the event that duplicate modified data are present among the generated modified data D2’ corresponding to the modified data already saved D2S in the archive, the processor removes F911 the duplicate modified data and proceeds to a new generation F4 of modified data D2’, selecting new modification data D2 or varying the modification rules RM. Iteratively, the prior art control step F91 is repeated to verify that the newly regenerated modified data are at this point unique. Obviously, this process is iterative until modified data D2’ that are unique are identified.
[0180] Instead, in case of negative outcome, i.e. in the event that no duplicate modified data are found, the method continues F92 with the replacement step F5 of the modification data D2 of the input data D1 with the modified data D2’ that have passed the prior art control step D91 . As anticipated earlier, when the modified data D2’ are also associated with the code of an owner, it is possible to filter the prior art control step F91 to only the modified data D2’ associated with the owner, so as to make the uniqueness relative to the specific owner.
[0181] Once the description of the method that actually carries out the modification of the data and tracks it has been completed, we proceed to the description of the further part of the method, particularly important, which corresponds to the validation step of texts that may possibly have been illicitly misappropriated.
[0182] In this regard, the method may provide for the validation of the modified contents in two, paradoxically opposite, modes. In fact, for those digital contents that are kept secret or placed on the market as proprietary productions, the presence of the error may indeed indicate that such contents are the same as those produced by the owner. However, according to another control criterion, for those digital contents that are placed on the market as distinctive signs (for example graphic trademarks), the modifications can be used as a certificate of authenticity, i.e. in their absence, such distinctive sign is not official and therefore may correspond to a potential counterfeiter. This could facilitate the identification and demonstration of a site that illicitly exploits the trademarks of a certain company.
[0183] On this front, the method comprises a validation step V1 , which could also be defined as generic validation.
[0184] In the validation step V1 , the processor receives V11 verification data D3, representative of an external digital content. The verification data D3 are nothing more than counterparts of the input data D1 and which are presumably derived from the input data D1. The validation step V1 comprises a step of analyzing the verification data D3. In particular, the validation step V1 comprises a search step V12, in which the processor searches, in the verification data D3, for one or more of the modified data D2’ that are saved within the archive. Following said search step V12, there is therefore an identification step V13 of a number of correspondences COR between the verification data D3 and the modified data D2’ saved in the archive. The number of correspondences COR is clearly an indicative parameter of a degree of derivation GD which is then estimated by the processor.
[0185] In particular, therefore, the processor carries out a diagnosis step V14, in which, on the basis of the number of correspondences COR, it determines the degree of derivation GD of the external digital content with respect to the proprietary digital contents. More in detail, each correspondence COR will clearly be associated with a proprietary digital content, since the related modified data D2’ are associated with a unique code Cl of the digital content. Therefore, at the end of such verification, the processor identifies the potential document or documents of derivation (which may also be multiple and combined with each other).
[0186] In the diagnosis, clearly the degree of derivation GD is determined as probable for a number of correspondences COR greater than a predetermined number, which may even be zero. In other words, in some cases a document may be considered as probable derivation even if there is only one occurrence present.
[0187] Such validation step may also be a targeted validation step VM1. The targeted validation step VM1 comprises the same steps as the validation step V1 , namely the reception of verification data VM11 , representative of an external digital content, and the analysis of the verification data and the search VM13, within the verification data, for one or more modified data D2’ saved in the archive.
[0188] What changes is the search perimeter on the modified data D2’. In fact, in this case, the method provides for a step of receiving a comparison code VM12. The comparison code may be representative of:
[0189] • a unique code of a digital content presumed to have been used, and / or a code of the owner of the digital content presumed to have been misappropriated.
[0190] In this way, one can search only among the modified data D2’ of a specific user or even of a specific digital content. These steps make it possible to speed up identification but require knowledge by the user about the presumed digital content that has been misappropriated.
[0191] Therefore, the search step VM13 is carried out on the modified data D2’ associated with the entered comparison code.
[0192] The method then continues also with the same identification steps VMM of the number of correspondences COR between the verification data D3 and the modified data D2’ saved in the archive in association with the entered comparison code and the diagnosis VM15 of an illicit degree of derivation GD for a number of correspondences COR greater than a predetermined number of correspondences.
[0193] On the other front, that of the use of an unauthorized distinctive sign, the method reverses the evaluation criterion of the infringement. Therefore, the method comprises a counterfeiting control step (C1).
[0194] In principle, the initial steps are exactly the same and we briefly repeat them below for conciseness:
[0195] • reception C11 of the verification data D3, representative of an external digital content;
[0196] • reception C12 of a comparison code, representative of a unique code of a digital content presumed to have been used;
[0197] • analysis of the verification data and search C13, within the verification data D3, for one or more modified data D2’ saved in the archive in association with the entered comparison code;
[0198] • identification C14 of the number of correspondences COR between the verification data D3 and the modified data D2’ saved in the archive in association with the entered comparison code.
[0199] What changes is the final diagnosis step, for which it is the absence of the modified data D2’ that verifies a presumed counterfeiting or illicit use of the distinctive sign. More in detail, the method therefore comprises an identification step C15 of a counterfeiting of the external digital content with respect to the proprietary digital content associated with the comparison code, for a number of correspondences COR between the verification data D3 and the modified data D2’ saved in the archive in association with the entered comparison code lower than a predetermined number of correspondences. For a practical example, such method could be used to intelligently “soil” an image of a company logo, in ways that are not detectable to the human eye and with a criterion that is not easily determinable by processing the image on a processor, since no exact modification criterion is followed. In this case, the company logo (proprietary digital content) undergoes the modification step, in which some pixels of the image (modification data D2) are modified by varying, for example, one or more color parameters (not perceptible to the human eye) so as to generate new pixels modified in color (modified data D2’).
[0200] The new modified logo DT is then placed on the market to authorized resellers only. When one wishes to evaluate the lawfulness of a subject advertising said trademark, it will be sufficient to validate it in order to verify that the color of the modified pixel corresponds to the modified color D2’.
Claims
CLAIMS1. Method for the intelligent modification of a digital content in order to make the digital content itself unique, the method comprising:• receiving (F1 ) input data (D1 ), representative of the proprietary digital content;• on the basis of the input data (D1 ), evaluating (F2) one or more modification rules (RM), identifying a rule with which to modify the input data (D1 );• identifying (F3) modification data (D2), represented by a group of input data (D1 );• applying (F4) said one or more modification rules (RM) to said modification data (D2), so as to generate modified data (D2’);• within the input data (D1 ), replacing (F5) the modification data (D2) with the modified data (D2’), so as to generate unique input data (D1 ’);• making available (F6) the unique input data (D1 ’).
2. Method according to claim 1 , comprising a tracking step (F7), comprising the following steps:• associating (F71 ) an identification code (Cl) with the digital content;• saving (F72), in an archive, the modified data (D2’) in association with the unique identification code (Cl).
3. Method according to claim 2, wherein, in the saving step (F72), the modified data (D2’) are associated with a timestamp, in order to certify the date on which said input data (D1 ) were modified by integrating the modified data (D2’).
4. Method according to claim 3, comprising a plurality of intelligent modification steps, temporally spaced apart and each including respective different modification data (D2’) and / or respective different modification rules (RM), wherein each intelligent modification step is associated with a respective timestamp.
5. Method according to claim 2, 3 or 4, comprising a validation step (V1 )and / or a targeted validation step (VM1 ), wherein the validation step (V1 ) comprises the following steps:• receiving (V11) verification data (D3), representative of an external digital content;• analyzing the verification data (D3) and searching (V12), within the verification data (D3), for one or more modified data (D2’) saved in the archive;• identifying (V13) the number of correspondences (COR) between the verification data (D3) and the modified data (D2’) saved in the archive, each of said correspondences being associated with a corresponding identification code (Cl);• for each digital content, diagnosing (V14) a degree of derivation (GD), in the event that a number of correspondences (COR) greater than a predetermined number of correspondences is identified. wherein the targeted validation step (VM1 ) comprises the following steps:• receiving (VM11 ) verification data (D3), representative of an external digital content;• receiving (VM12) a comparison code (CC), representative of a unique code (Cl) of a digital content presumed to have been used, or of an identification code of a user who is owner of a plurality of digital contents;• analyzing the verification data and searching (VM13), within the verification data (D3), for one or more modified data (D2’) saved in the archive in association with the entered comparison code (CC);• identifying (VMM) the number of correspondences (COR) between the verification data (D3) and the modified data (D2’) saved in the archive in association with the entered comparison code (CC);• diagnosing (VM15) a degree of derivation (GD) for a number of correspondences (COR) greater than a predetermined number of correspondences.
6. Method according to any of the preceding claims, comprising a counterfeiting control step (C1 ), comprising the following steps:• receiving (C11) verification data (D3), representative of an external digital content;• receiving (C12) a comparison code (CC), representative of a unique code (Cl) of a digital content presumed to have been used;• analyzing the verification data and searching (C13), within the verification data (D3), for one or more modified data (D2’) saved in the archive in association with the entered comparison code (CC);• identifying (C14) the number of correspondences (COR) between the verification data (D3) and the modified data (D2’) saved in the archive in association with the entered comparison code (CC);• identifying (C15) a counterfeiting of the external digital content with respect to the proprietary digital content associated with the comparison code (CC), for a number of correspondences (COR) between the verification data (D3) and the modified data (D2’) saved in the archive in association with the entered comparison code (CC) lower than a predetermined number of correspondences.
7. Method according to any of the preceding claims, comprising a prior art control step (F9), including the following steps:• following the generation of the modified data (D2’), searching (F91 ) for correspondences between the generated modified data (D2’) and the modified data (D2S) already saved in the archive;• in case of positive outcome, represented by the presence of duplicate modified data among the generated modified data (D2’) corresponding to the modified data (D2S) already saved in the archive:(a) removing (F911 ) said duplicate modified data;(b) generating again (F4) substitute modified data (D2’);(c) executing the prior art control step (F9) on the substitute modified data;• in case of negative outcome, represented by the absence of duplicate modified data, continuing (F92) the method with the replacement (F5) of the modification data (D2) of the input data (D1 ) with the modified data (D2’) that have passed the prior art control step (F9).
8. Method according to claim 7, wherein the modified data (D2S) saved in the archive are associated with a proprietary code, uniquely associated with an owner of the digital content, and wherein, in the prior art control step (F9), said search step (F91 ) for correspondences is carried out on the modified data (D2S) already saved in the archive in association with a specific proprietary code.
9. Method according to any of the preceding claims, wherein the input data (D1 ) are representative of one or more of the following multimedia contents:• texts;• images;• audio tracks;• data archives;• source code of computer programs, and wherein the modification data (D2) include one or more of the following data:• one or more text characters of a text;• one or more pixels of an image;• one or more audio windows of an audio track;• one or more attributes of tables of an archive, wherein said one or more modification rules (RM) include one or more of the following rules:• replacement of a text character with a text character graphically similar;• addition of text characters representative of a content devoid of logical function or real relevance with respect to the text in which it is included;• replacement of a punctuation character with a punctuation character graphically similar but improper in the position of the replaced punctuation character;• replacement of a text character representing a blank space withmultiple text characters representing a blank space;• repetition of a text character where unnecessary;• modification of the color of a single pixel in the image;• imperceptible variation of the color of an image area, through variation of the opacity linked to the color;• variation of frequency or amplitude of the sound in a specific window of the audio track.
10. Computer program comprising instructions to execute the steps of the method according to any of claims 1 to 8.
Citation Information
Patent Citations
Systems and Methods for Watermarking Software and Other Media
US20080028474A1