Prevention of data leakage of an artificial intelligence (AI) training set
By using a modification AI algorithm and obfuscator to transform sensitive data and adjusting model weights, the AI algorithm is prevented from leaking sensitive information, addressing data leakage and copyright concerns effectively.
Patent Information
- Application Number
- PCT/US2024/025630
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-10-30
AI Technical Summary
Artificial Intelligence (AI) algorithms face challenges in preventing the leakage of sensitive training data, which can include copyrighted information or proprietary source code, and identifying the source of the training set, leading to potential copyright infringement and data exposure.
The implementation of a modification AI algorithm and an obfuscator to transform sensitive data into modified data, combined with a weight pattern AI algorithm to adjust model weights, ensures that the AI algorithm does not leak sensitive information by training on modified data, thereby maintaining data privacy and compliance with copyright laws.
This approach effectively prevents the leakage of sensitive data by altering the structure and obfuscating the content, making it difficult for the AI algorithm to generate similar outputs, thus protecting intellectual property and maintaining data anonymity.
Smart Images

Figure US2024025630_30102025_PF_FP_ABST
Abstract
Description
PREVENTION OF DATA LEAKAGE OF AN ARTIFICIAL INTELLIGENCE (Al) TRAINING SETFIELD
[0001] The disclosure relates generally to Artificial Intelligence (Al) algorithms and particularly to modification of a training set of an Al algorithm to prevent leakage of training set data.BACKGROUND
[0002] One of the problems with Al algorithms is that the training data may be leaked. For example, by using specific types of prompts, actual snippets of training data may be released by the Al algorithm. This can be an issue if the training set contains sensitive data and / or if the training set contains copyrighted information or other key intellectual property. For example, based on specific prompts, the Al algorithm may generate a key portion of source code of a proprietary application that is part of the training set.
[0003] Another problem is that a source of the training set may not want to be revealed or needs to be anonymous. For example, an author, a company, or an owner of data in the training set may not want to be revealed.
[0004] In addition, the leaking of copyrightable data may lead to copyright infringement litigation. For example, the New York Times® has filed a lawsuit with Open Al and Microsoft® alleging copyright infringement based on the training set and the fact that the Al algorithm can be triggered to produce copyrightable works as an output. Part of this litigation is based on where specific prompts actually produce most or all of a copyrighted article.
[0005] Because of the complexity of Al algorithms (e.g., a Large Language Model (LLM)), it is extremely difficult, if not impossible, to identify if and how information in a training set will be leaked based on inputs to the Al algorithm. This is because LLMs typically have millions of nodes, thousands of layers, and have billions even trillions of weights.SUMMARY
[0006] These and other needs are addressed by the various embodiments and configurations of the present disclosure. The present disclosure can provide a number of advantages depending on the particular configuration. These and other advantages will be apparent from the disclosure contained herein.
[0007] Sensitive data is received. For example, the sensitive data may be sensitive source code (e.g., proprietary source code). The sensitive data is information that is not wanted to be leaked by an Al algorithm. For example, this information may need to be protected regardless of the prompts provided to the Al algorithm. The sensitive data is passed to at least one of: a modification Al algorithm or an obfuscator to generate modified data. For example, if the sensitive data is sensitive source code, the modification Al algorithm may change a structure of the sensitive source code and / or the obfuscator may obfuscate the sensitive source code to generate the modified source code. The Al algorithm is then trained using the generated modified data. This makes it more likely that that sensitive data is no longer leaked by the Al algorithm.
[0008] The phrases "at least one", "one or more", “or,” and "and / or" are open-ended expressions that are both conjunctive and disjunctive in operation. For example, each of the expressions "at least one of A, B and C", "at least one of A, B, or C", "one or more of A, B, and C", "one or more of A, B, or C", "A, B, and / or C", and "A, B, or C" means A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B and C together.
[0009] The term "a" or "an" entity refers to one or more of that entity. As such, the terms "a" (or "an"), "one or more" and "at least one" can be used interchangeably herein. It is also to be noted that the terms “comprising,” “including,” and “having” can be used interchangeably.
[0010] The term “automatic” and variations thereof, as used herein, refers to any process or operation, which is typically continuous or semi-continuous, done without material human input when the process or operation is performed. However, a process or operation can be automatic, even though performance of the process or operation uses material or immaterial human input, if the input is received before performance of the process or operation. Human input is deemed to be material if such input influences howthe process or operation will be performed. Human input that consents to the performance of the process or operation is not deemed to be “material.”
[0011] Aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a “circuit,” “module” or “system.” Any combination of one or more computer readable medium(s) may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium.
[0012] A computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0013] A computer readable signal medium may include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium may be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0014] The terms “determine,” “calculate” and “compute,” and variations thereof, as used herein, are used interchangeably, and include any type of methodology, process, mathematical operation, or technique.
[0015] T he term “means” as used herein shall be given its broadest possible interpretation in accordance with 35 U.S.C., Section 112(f) and / or Section 1 12, Paragraph 6. Accordingly, a claim incorporating the term “means” shall cover all structures, materials, or acts set forth herein, and all of the equivalents thereof Further, the structures, materials or acts and the equivalents thereof shall include all those described in the summary, brief description of the drawings, detailed description, abstract, and claims themselves
[0016] The term “blockchain” as described herein and in the claims refers to a growing list of records, called blocks, which are linked using cryptography. The blockchain is commonly a decentralized, distributed and public digital ledger that is used to record transactions across many computers so that the record cannot be altered retroactively without the alteration of all subsequent blocks and the consensus of the network. Each block contains a cryptographic hash of the previous block, a timestamp, and transaction data (generally represented as a merkle tree root hash). For use as a distributed ledger, a blockchain is typically managed by a peer-to-peer network collectively adhering to a protocol for inter-node communication and validating new blocks. Once recorded, the data in any given block cannot be altered retroactively without alteration of all subsequent blocks, which requires consensus of the network majority. In verifying or validating a block in the blockchain, a hashcash algorithm generally requires the following parameters: a service string, a nonce, and a counter. The service string can be encoded in the block header data structure, and include a version field, the hash of the previous block, the root hash of the merkle tree of all transactions (or information or data) in the block, the current time, and the difficulty level. The nonce can be stored in an extraNonce field, which is stored as the left most leaf node in the merkle tree. The counter parameter is often small at 32-bits so each time it wraps the extraNonce field must be incremented (or otherwise changed) to avoid repeating work. When validating or verifying a block, the hashcash algorithm repeatedly hashes the block header while incrementing the counter & extraNonce fields. Incrementing the extraNonce field entailsrecomputing the merkle tree, as the transaction or other information is the left most leaf node. The body of the block contains the transactions or other information. These are hashed only indirectly through the Merkle root.
[0017] The preceding is a simplified summary to provide an understanding of some aspects of the disclosure. This summary is neither an extensive nor exhaustive overview of the disclosure and its various embodiments. It is intended neither to identify key or critical elements of the disclosure nor to delineate the scope of the disclosure but to present selected concepts of the disclosure in a simplified form as an introduction to the more detailed description presented below. As will be appreciated, other embodiments of the disclosure are possible utilizing, alone or in combination, one or more of the features set forth above or described in detail below. Also, while the disclosure is presented in terms of exemplary embodiments, it should be appreciated that individual aspects of the disclosure can be separately claimed.BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Fig. l is a block diagram of a first illustrative system for prevention of Al algorithm data leakage.
[0019] Fig. 2 is a block diagram of a second illustrative system for prevention of Al algorithm data leakage.
[0020] Fig. 3 is a block diagram of a third illustrative system for prevention of Al algorithm data leakage using a weight pattern Al algorithm.
[0021] Fig. 4 is a block diagram of a fourth illustrative system for prevention of source code leakage in a code generation Al algorithm.
[0022] Fig. 5 is a block diagram of a fifth illustrative system for prevention for prevention of source code leakage in a code generation Al algorithm.
[0023] Fig. 6 is a block diagram of a sixth illustrative system for prevention of source code leakage in a code generation Al algorithm using a weight pattern Al algorithm.
[0024] Fig. 7 is a flow diagram of a process for prevention of Al algorithm data leakage.
[0025] Fig. 8 is a flow diagram of a process for handling rules based on whether the initial training set comprises sensitive source code, images, or other data that has specific rules.
[0026] Fig. 9 is a flow diagram of a process for generating vectors of snippets of generated data.
[0027] Fig. 10 is a block diagram of a blockchain that is used to store information associated with training an Al algorithm that is used to prevent leakage of training data.
[0028] Fig. 11 is a block diagram of a seventh illustrative system for training a weight pattern Al algorithm to identify patterns in weights.
[0029] Fig. 12 is a block diagram of an exemplary neural network of an Al algorithm.
[0030] Fig. 13 is a block diagram of an eighth illustrative system for training a weight pattern Al algorithm by changing Al inputs.
[0031] Fig. 14 is a block diagram of a ninth illustrative system for training a weight pattern Al algorithm by using a plurality of training sets.
[0032] Fig. 15 is a block diagram of a tenth illustrative system for training a weight pattern Al algorithm to produce weights for a plurality of Al algorithms that have different neural network architectures.
[0033] Fig. 16 is a block diagram of an eleventh illustrative system for creating a new set of weights for an Al algorithm.
[0034] Fig. 17 is a flow diagram of a process for training a weight pattern Al algorithm by using base Al inputs.
[0035] Fig. 18 is a flow diagram of a process for training a weight pattern Al algorithm by using a plurality of training sets.
[0036] In the appended figures, similar components and / or features may have the same reference label. Further, various components of the same type may be distinguished by following the reference label by a letter that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.DETAILED DESCRIPTION
[0037] Fig. 1 is a block diagram of a first illustrative system 100 for prevention of Al algorithm 121 data leakage. The first illustrative system 100 comprises communication devices 101A-101N, a network 110, and a server 120. In addition, users 102A-102N are shown for convenience.
[0038] The communication devices 101A-101N can be or may include any user device that can communicate on the network 110, such as a Personal Computer (PC), a cellular telephone, a Personal Digital Assistant (PDA), a tablet device, a notebook device, a laptop computer, a smartphone, and / or the like. As shown in Fig. 1, any number of communication devices 101A-101N may be connected to the network 110, including only a single communication device 101. The communication devices 101A-101N are used by the users 102A-102N to access services provided by the server 120.
[0039] The network 110 can be or may include any collection of communication equipment that can send and receive electronic communications, such as the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), a packet switched network, a circuit switched network, a cellular network, a combination of these, and the like. The network 110 can use a variety of electronic protocols, such as Ethernet, Internet Protocol (IP), Hyper Text Transfer Protocol (HTTP), Web Real-Time Protocol (Web RTC), and / or the like. Thus, the network 110 is an electronic communication network configured to carry messages via packets and / or circuit switched communications.
[0040] The server 120 may be any type of server 120 that can be used host / manage the Al algorithm 121, such as an application server, a cloud service, and / or the like. The server 120 may comprise a number of servers / computers / computer cores that are used to host the Al algorithm 121. The server 120 comprises the Al algorithm 121, a final training set 122, modified data 123, a data modifier 124, a weight pattern Al algorithm 127, an initial training set 128, generated data 131, released generated data 132, rules 133, and a vector Al algorithm 134.
[0041] The Al algorithm 121 may be any type of Al algorithm that uses a training set such as a supervised Al algorithm, a linear regression Al algorithm, a neural network Al algorithm, a reinforcement learning Al algorithm, and / or the like. The Al algorithm 121is used to generate data based in input parameters / prompts. The AT algorithm 121 is trained using the final training set 122.
[0042] The final training set 122 is a training set that is produced based on modified data 123 and optionally non-modified data (i.e., the non-sensitive data 129). The final training set 122 is created by the data modifier 124.
[0043] The modified data 123 is sensitive data 130 that is modified by the data modifier 124. The modified data 123 may be modified by the modification Al algorithm 125 and / or the obfuscator 126. The modified data 123 becomes part of the final training set 122.
[0044] The data modifier 124 is used to change the sensitive data 130 of the initial training set 128. The data modifier 124 further comprises the modification Al algorithm 125 and the obfuscator 126.
[0045] The modification Al algorithm 125 is used to modify the sensitive data 130 in various ways. For example, the modification Al algorithm 125 may create a new version of a document or a new version of source code.
[0046] The obfuscator 126 is used to modify the sensitive data 130 in various ways. For example, the obfuscator 126 may modify a document by removing headers, footers, links, and / or the like. The obfuscator 126 may obfuscate source code. For example, the obfuscator 126 may obfuscate source code by changing variable names / function names (e g., the function name namePrintToPrinter(document);) to simple names (e g., X(z);).
[0047] The weight pattern Al algorithm 127 is an Al algorithm that is trained using weights of the Al algorithm 121. The weight pattern Al algorithm 127 changes the weights of the Al algorithm 121 based on input parameters / prompts.
[0048] The initial training set 128 is data that is initially selected to use for training the Al algorithm 121. The initial training set 128 comprises non-sensitive data 129 and sensitive data 130. The non-sensitive data 129 is data that is okay to be leaked by the Al algorithm 121. For example, the non-sensitive data 129 may be data that is not subject to copyrights, public domain data, open-source code, and / or the like. The sensitive data 130 is data that is not wanted to be leaked by the Al algorithm 121, such as copyrighted data, proprietary data, data where a source of the data is not wanted to be disclosed, data that is subject to a particular license, and / or the like.
[0049] The generated data 131 is data that is initially generated by the Al algorithm 121. The generated data 131 may be scanned / processed to determine if it should be released (the released generated data 132). If the generated data 131 contains leaks of the sensitive data 130, the generated data 131 may be reprocessed by the Al algorithm 121 to produce the released generated data 132.
[0050] The rules 133 are rules that are used to determine how the modification Al algorithm 125 / obfuscator 126 change the sensitive data 130. The rules 133 may be different based on the type of data in the initial training set 128. For example, if the initial training set 128 is data scraped from the Internet, the rules 133 may be different from rules 133 where the initial training set 128 is source code.
[0051] The vector Al algorithm 134 is an Al algorithm that is used to generate vectors (e g., floating point vectors) based on snippets of the generated data 131. Snippets of the generated data 131 may be compared to snippets of the sensitive data 130, data modified by the modification Al algorithm 125, data modified by the obfuscator 126, and / or the like.
[0052] Fig. 2 is a block diagram of a second illustrative system 200 for prevention of Al algorithm 121 data leakage. Fig. 2 shows a system 200 that obfuscates / transforms the sensitive data 130 in the initial training set 128 in ways that can reduce / eliminate the leakage of sensitive data 130 from the initial training set 128. Fig. 2 is described where the initial training set 128 is based on generic data. For example, the initial training set 128 may be based on information scraped from the Internet. In other embodiments, the initial training set 128 may be a specific type of training data, such an initial training set 128 that comprises only images, only specific types of documents, only music, only source code, etc. For example, the initial training set 128 may be a group of images and the Al algorithm 121 would be an image creation Al algorithm 121. If the Al algorithm 121 is trained on images, the rules 133 would be different for the specific types of initial training sets 128. For example, the rules 133 for the obfuscator 126 (e.g., for all or individual images) may be to remove text from image(s), to remove specific color(s) from the image(s), to add specific color(s) to the image, to change a background of the image(s), to grey scale the image(s), and / or the like.
[0053] The initial training set 128 is divided into two types of data 1) one that contains non-sensitive data 129, and 2) one that contains sensitive data 130 (e.g., copyrightable data). The data in the initial training set 128 may have associated metadata that indicates if it is sensitive data 130 or non-sensitive 129. Alternatively, the non-sensitive 129 may be stored in different directories than the sensitive / copyrightable data 130. In one embodiment, the sensitive data 130 / non-sensitive data 129 may be user selectable. The non-sensitive data 129 becomes part of the final training set 122. In one embodiment, there may not be any non-sensitive data 129. Thus, there would not be any non-sensitive data 129 in the final training set 122.
[0054] The sensitive data 130 (e.g., copyrightable data / proprietary data) is modified in different ways depending on user selectable rules 133 that control the obfuscator 126 and / or a modification Al algorithm 125. The modification Al algorithm 125 / obfuscator 127 generate new versions of the sensitive data 130. The obfuscator 126 can change the sensitive data 130 (e.g., documents) in different ways, such as removing footnote(s), removing reference(s), removing embedded image(s), removing name(s), removing header(s), removing link(s), removing page number(s), and / or the like.
[0055] The modification Al algorithm 125 can also modify the sensitive data 130 in various ways, such as modifying a structure of one or more documents, rewriting the one or more documents, replacing name(s) in the one or more documents, replacing word(s) in the one or more documents with synonym(s), reformatting an image, modifying an image based in Al input data, creating new images based on input data / prompts, reformatting a webpage, and / or the like.
[0056] The rules 133 may apply to all of the sensitive data 130 or could be applied at an individual document level, an individual image level, an individual file level, an individual section level, a group level, at a folder level, an individual section level, an individual area level (e.g., a portion of an image), and / or the like. The rules 133 may be predefined and / or user selectable.
[0057] To illustrate consider the following example. The user 102 selects to have the obfuscator 126 to remove footnotes, references, embedded images, and links in a set of documents. The user 102 also selects to have the modification Al algorithm 125 to modify the structure of two selected documents in the sensitive data 130, to rewrite alldocuments in the sensitive data 130, and replace words with synonyms (could be all or specific words or specific words in specific documents) in all the documents in the sensitive data 130. The obfuscator 126 then changes the documents according to the rules 133. The output of the obfuscator 126 is then input into the modification Al algorithm 125. The modification Al algorithm 125 then modifies the obfuscated data based on the rules 133 for the modification Al algorithm 125 to produce the modified data 123. Thus, the sensitive data 130 has been dramatically changed and is different from the modified data 123. This prevents leakage of the sensitive data 130 from initial training set 128 in the generated data 131.
[0058] Another feature of the modifying of the sensitive data 130 is that a source (e.g., who wrote / created the sensitive data) of the sensitive data 130 may be removed. For example, by removing names, links, headers, footers, the source of the sensitive data 130 can be removed.
[0059] The final training set 122 that is comprised of the non-sensitive data 129 and the modified data 123 becomes the final training set 122 that is used to train the Al algorithm 121. Because the final training set 122 does not actually have the sensitive data 130, input parameters / prompts will no longer be able to cause the Al algorithm 121 to leak the sensitive data 130.
[0060] In addition, a feedback loop can be added to identify any potential leaks that are still in the generated data 131. In Fig. 2 the vector Al algorithm 134 provides feedback by identifying leaks of sensitive data 130 in the generated data 131. The vector Al algorithm 134 generates vectors of snippets (e.g., floating point or integer vectors) of the generated data 131 and compares them to vectors of snippets from one or more of 1) the modified data 123, 2) the modified data 123 from the obfuscator 126 (e.g., if the obfuscator 126 and the modification Al algorithm 125 are in series), and 3) the sensitive data 130. The vectors of snippets can determine if there is too much of sensitive data 130, too much obfuscated data, and / or too much modified data 123 is in the generated data 131.
[0061] If there is too much of the sensitive data 130, the obfuscated data, and / or modified data 123 (there may be separate thresholds for each) the matched snippets may be provided as input to the modification Al algorithm 125. For example, a snippet of thesensitive data 130 may be provided to the modification AT algorithm 125 as an input / prompt to generate modified data 123 that does not contain data similar to the snippet of the sensitive data 130 and / or generated data 131. The input snippet may be user selectable the next time the training process completes.
[0062] Although Fig. 2 shows the obfuscator 126 and the modification Al algorithm 125 in parallel, the obfuscator 126 / modification Al algorithm 125 may work in series. For example, the obfuscator 126 may perform the obfuscation function and the resulting obfuscated data is then feed as an input into the modification Al algorithm 125. In another embodiment, the modification Al algorithm 125 may be executed first and the obfuscator 126 may be executed second.
[0063] Fig. 3 is a block diagram of a third illustrative system 300 for prevention of Al algorithm 121 data leakage using a weight pattern Al algorithm 127. The difference between Fig. 2 and Fig. 3 is that instead of the output of the matching of threshold(s) going to the modification Al algorithm 125, the output of the matching of the threshold(s) (e.g., matched snippets) goes to the weight changing algorithm 127 as input / prompts. In this example, the weight pattern Al algorithm 127 changes the weights of the Al algorithm 121 instead of requiring the Al algorithm 121 to be retrained. The changing of the weights causes the Al algorithm 121 to no longer produce output that is similar to the sensitive data 130.
[0064] Fig. 4 is a block diagram of a fourth illustrative system 400 for prevention of source code leakage in a code generation Al algorithm 421. The fourth illustrative system 400 comprises the communication devices 101A-101N, the network 110, and the server 120. In addition, the users 102A-102N are shown for convenience.
[0065] In Fig. 4, the server 120 comprises the code generation Al algorithm 421, a final training set 422, modified source code 423, a source code modifier 424, a weight pattern Al algorithm 427, an initial training set 428, generated source code 431, released generated source code 432, rules 433, and a vector Al algorithm 434.
[0066] The code generation Al algorithm 421 is an Al algorithm that is trained using various kinds of source code. The types of source code may include source code that is written in various programming languages, such as Java, C, C++, C#, Pearl, JavaScript,Hyper Text Markup Language (HTTP), assembly language, machine language, and / or the like.
[0067] The final training set 422 is a final training set that is based on source code. The final training set 422 comprises the modified source code 423 and the non-sensitive source code 429 (if there is any).
[0068] The modified source code 423 is sensitive source code 430 that has been modified by the modification Al algorithm 425 and / or the obfuscator 426.
[0069] The source code modifier 424 modifies the sensitive source code 430 in the initial training set 428. The source code modifier 424 comprises a modification Al algorithm 425 and the obfuscator 426. The source code modifier 424 modifies the sensitive source code 430 in the initial training set 428.
[0070] The modification Al algorithm 425 is used to modify the sensitive source code 430 based on the rules 433. The obfuscator 426 is used to modify the sensitive source code 430 based on the rules 433.
[0071] The weight pattern Al algorithm 427 is used to modify weights of the code generation Al algorithm 421. The weight pattern Al algorithm 427 is trained using sets of weights that were generated when training the code generation Al algorithm 421.
[0072] The initial training set 428 is initial source code that is initially selected to be used to train the code generation Al algorithm 421. The initial training set 428 comprises non-sensitive source code 429 and sensitive source code 430. The non-sensitive source code 429 is source code that can be leaked by the code generation Al algorithm 421. For example, the non-sensitive source code 429 may be open-source code, public domain source code, and / or the like. The sensitive source code 430 may be any source code that the user 102 does not want to be leaked by the code generation Al algorithm 421, such as, copyrightable source code, open-source code, proprietary source code, licensed source code, and / or the like.
[0073] The generated source code 431 is source code that is generated by the code generation Al algorithm 421. The generated source code 431 is generated based on specific input parameters / prompts to the code generation Al algorithm 421.
[0074] The released generated source code 432 is source code that has been approved to be released. For example, the released generated source code 432 may be softwarethat is released in a commercial software application. The released generated source code 432 may the same or similar to the generated source code 431.
[0075] The rules 433 are rules that are specific to source code. The rules apply to the modification Al algorithm 425 and / or the obfuscator 426. The rules 433 may include rules for obfuscating the sensitive source code 430, removing comments from the sensitive source code 430, changing a format of the sensitive source code 430, creating a new version of the sensitive source code 430, changing a sensitive source code naming attribute, changing a naming attribute of comments in the sensitive source code, optimizing the sensitive source code 430, changing a structure of the sensitive source code 430, and / or the like.
[0076] The vector Al algorithm 434 creates vectors of the generated source code 431 and compares the vectors of the generated source code to vectors of the sensitive source code 430, vectors of an output of the obfuscator 426, and / or vectors from the modified source code 423.
[0077] Fig. 5 is a block diagram of a fifth illustrative system 500 for prevention for prevention of source code leakage in the code generation Al algorithm 421. In Fig. 5, the initial training set 428 is divided into two sets of training source code: 1) one that contains the non-sensitive source code 429, and 2) one that contains the sensitive source code 430. The source code in the initial training set 428 may have associated metadata that indicates if it is sensitive source code 430 or non-sensitive source code 429.Alternatively, the non-sensitive source code 429 may be stored in different directories than the sensitive source code 430. The non-sensitive source code 429 (e.g., open-source code / public domain source code) becomes part of the final training set 422 (assuming that there is non-sensitive source code 429).
[0078] The sensitive source code 430 is modified in different ways depending on the rules 433 that control the obfuscator 426 and / or the modification Al algorithm 425 that generates modified source code 423 of the sensitive source code 430 (e g., with a different structure). The obfuscator 426 can perform various functions, such as 1) to obfuscate the sensitive source code 430, 2) to remove comments from the sensitive source code 430, 3) to change format / layout of sensitive source code 430 (e.g., like a C beautifier), and / or the like. The modification Al algorithm 425 can perform variousfunctions, such as: 1) to create a new version of the sensitive source code 430, 2) to change sensitive source code attributes (e.g., change / rename the variable names, function names, and / or the like) using things like synonyms, 3) to change the comments (e.g., to modify the comments using analogous language), 4) to optimize the sensitive source code 430, and / or the like. The rules 433 may apply to all of the sensitive source code 430 or could be applied at an individual software component level, fde level, a package level, a group level, a function level, a directory level, and / or the like.
[0079] To illustrate consider the following example. The user 102 selects to remove the comments from the sensitive source code 430, to create a modified version of the sensitive source code 430, and to use synonyms for the function / variable names in the sensitive source code 430. The obfuscator 426 first removes the comments from the sensitive source code 430. Based on the obfuscated source code, the modification Al algorithm 425 then generates a modified version of the sensitive source code 430 that performs the same function. For example, new prompts are given to the modification Al algorithm 425 that tell the modification Al algorithm 425 to generate new version of source code that performs the same function but uses a different code structure than the obfuscated source code. The modification Al algorithm 425 also changes the variable names / function call names with synonym s / synonymous language. As a result, the modified sensitive source code 423 is now dramatically changed but still has the same functionality as the sensitive source code 423.
[0080] For the modified sensitive source code 430, it could be argued that it is a new work (not a derivative work) and therefore not a violation of the copyright law. In many ways, it is similar to how someone would create a new source code based on existing copyrightable source code. The modified source code 423 will have a different structure, different variable names, different function names, and different comments. Likewise, in this example, the sensitive source code 430 in the modified source code 423 would likely be considered a new work instead of a derivative work because all the artistic content has been changed and is different from the original work. Even if there were minor similarities the minor similarities would likely be considered fair use under the copyright laws.
[0081] Another feature of the modifying of the sensitive source code 430 is that the source (who generated the sensitive source code 430) of the sensitive source code 430 may be removed. For example, by changing the variable / function call naming, modifying comments, and / or removing comments, a particular user 102 / company who produced the sensitive source code 430 may no longer be able to be identified.
[0082] The final training set 422 is comprised of the non-sensitive source code 429 and the modified source code 423. The final training set 422 is used to train the code generation Al algorithm 421. Because the final training 422 set does not actually have the sensitive source code 430, input / prompts will likely no longer be able to generate source code 431 that has the sensitive source code 430.
[0083] Another option would be to provide a feedback loop to identify any potential leaks that are still in the generated source code 431. One way to deal with this would be to use the vector Al algorithm 434. The vector Al algorithm 434 takes vectors of snippets (e.g., floating point or integer vectors) of the generated source code 431 and compares them to vectors of snippets from one or more of: 1) the modified source code 423, 2) the output of the obfuscator 426, 3) the sensitive source code 430 and / or 4) the output of the modification Al algorithm 425. This information can determine if there is too much of sensitive source code 430, too much obfuscated source code, and / or too much modified source code 423 in the generated source code 431.
[0084] If there is too much of the sensitive source code 430, the obfuscated source code, and / or the modified source code 423 (there may be separate thresholds for each) the matched source code snippets may be provided as an input parameter / prompt to the modification Al algorithm 425. For example, a snippet of the sensitive source code 430 or the generated source code 431 may be provided to the modification Al algorithm 425 as an input to generate source code that does not contain source code similar to the snippet of sensitive source code 430 / generated source code 431. This may be user selectable the next time the training process completes.
[0085] Fig. 6 is a block diagram of a sixth illustrative system 600 for prevention of source code leakage in a code generation Al algorithm 421 using a weight pattern Al algorithm 427. The difference between Fig. 5 and Fig. 6 is that instead of the output of the matching of threshold(s) / identified snippets going to the modification Al algorithm425, the output of the matching of the threshold(s) goes to the weight pattern AT algorithm 427. In this example, the weight pattern Al algorithm 427 changes the weights of the code generation Al algorithm 421 instead of requiring the code generation Al algorithm 421 to be retrained. The changing of the weights causes the code generation Al algorithm 421 to no longer produce an output that is similar to the sensitive source code 430. This can be accomplished by providing input / prompts to the weight pattern Al algorithm 427. For example, one or more vectors of snippets of the generated source code 431 / sensitive source code 430 can be used in a prompt provided to the weight pattern Al algorithm 427. The prompt can indicate to the weight pattern Al algorithm 427 to modify the weights of the code generation Al algorithm 421 to produce modified source code 423 that does not have source code similar to the matched snippet(s).
[0086] Fig. 7 is a flow diagram of a process for prevention of Al algorithm data leakage. Illustratively, the communication devices 101A-101N, the server 120, the Al algorithm 121, the data modifier 124, the modification Al algorithm 125, the obfuscator 126, the weight pattern Al algorithm 127, the vector Al algorithm 134, the code generation Al algorithm 421, the source code modifier 424, the modification Al algorithm 425, the obfuscator 426, the weight pattern Al algorithm 427, and the vector Al algorithm 434 are stored-program-controlled entities, such as a computer or microprocessor, which performs the method of Figs. 7-10 and the processes described herein by executing program instructions stored in a computer readable storage medium, such as a memory (i.e., a computer memory, a hard disk, and / or the like). Although the methods described in Figs. 7-10 are shown in a specific order, one of skill in the art would recognize that the steps in Figs. 7-10 may be implemented in different orders and / or be implemented in a multi-threaded environment. Moreover, various steps may be omitted or added based on implementation.
[0087] While Fig. 7 is described based on Fig. 1, the process of Fig. 7 also applies to Fig. 4 where the sensitive 130 data is or may comprise sensitive source code 430. The process starts in step 700. The data modifier 124 waits, in step 702, to receive sensitive data 130 that will be used to train the Al algorithm 121. If no sensitive data 130 is received in step 702, the process of step 702 repeats.
[0088] Otherwise, if sensitive data 130 is received for training the Al algorithm 121 in step 702, the data modifier 124 gets the rules 133 for the sensitive data 130 in step 704. The rules 133 may be different if the sensitive data 130 comprises sensitive source code 430, images, documents, text, database records, user interfaces, webpages, etc. The data modifier 124 passes the sensitive data 130 to the modification Al algorithm 125 and / or the obfuscator 126 in step 706. For example, the sensitive data 130 may be first passed to the obfuscator 126. The obfuscator 126 then passes the obfuscated data to the modification Al algorithm 125. In combination or individually, the obfuscator 126 and the modification Al algorithm 125 generate the modified data 123 based on the rules 133 in step 708.
[0089] The Al algorithm 121 is trained using the modified data 123 in step 710. For example, the modified data 123 along with the non-sensitive data 129 are combined to produce the final training set 122. The final training set 122 is then used to train the Al algorithm 121 in step 710.
[0090] The data modifier 124 determines, in step 712, if the process is complete. If the process is not complete in step 712, the process goes back to step 702. Otherwise, if the process is complete in step 712, the process ends in step 714.
[0091] Fig. 8 is a flow diagram of a process for handling rules 133 / 433 based on whether the initial training set 128 comprises sensitive source code 423, images, or other data that has specific rules. The process of Fig. 8 goes between steps 702 and 710 of Fig. 7.
[0092] After receiving sensitive data 130 in step 702, the data modifier 124 / 424 determines, in step 800, if there is sensitive source code 430, images, or other data in the sensitive data 130 that has unique rules 133. If there is not any sensitive source code 430, images, or other data in the sensitive data 130, the process goes to step 704. Otherwise, if there is sensitive source code 430, images, or other data in the sensitive data 130 that has unique rules 133, the data modifier 124 determines, in step 802, a type of sensitive data 130.
[0093] If the sensitive data 130 is source code in step 802, the source code modifier 424 gets, in step 804, the rules 433 for the sensitive source code 430 (e.g., the rules 433 shown in Fig. 5). The source code modifier 424 passes the sensitive source code 430 tothe obfuscator 426 and / or the modification Al algorithm 425 based on the rules 433 for the sensitive source code 430 in step 806. The modified source code 423 is generated in step 808 based on the rules 433. The process then goes to step 710.
[0094] If the sensitive data 130 is image(s), the data modifier 124 gets, in step 810, the rules 133 for the images. For example, the rules 133 for images may include removing a text from the image(s), removing a color(s) from the image(s), changing a background of the image(s), grey scaling the image(s), reformatting the image(s), modifying the image(s) based on Al input data, creating a new image(s) based on the Al input data, and / or the like. The data modifier 124 passes the images to the obfuscator 426 and / or the modification Al algorithm 425 based on the rules 133 for the images in step 812. The modified data 123 is generated in step 814 based on the rules 133 for the images. The process then goes to step 710.
[0095] If the sensitive data 130 is other data, the data modifier 124 gets, in step 816, the rules 133 for the other data. The rules 133 for the other data (e.g., a website data, video data, music data, audio data), may be different. For example, if the sensitive data 130 is video data, the rules may be similar to image data where the rules are applied to each frame. If the sensitive data 130 is music data or audio data, the rules 133 may be to change the voice of a speaker in the audio / music, remove a voice in the audio / music, modify a voice in the audio / music, add a new voice in the audio / music, change a musical rhythm, create a new Al generated musical recording based on Al input data, change an intensity of an instrument, remove an instrument, add an instrument, change a pitch, and / or the like.
[0096] If the sensitive data is website data / graphical user interface data, the rules 133 may be based on changing / moving graphical objects, changing the text of graphical objects, changing colors of graphical objects, adding a tab, removing a tab, reorganizing a tab, changing a button, removing a button, adding a button, adding an animation, removing an animation, changing an animation, changing a menu, and adding a menu item, removing a menu item, and / or the like. If the sensitive data is a presentation data (e.g., slides of a presentation), the rules 133 may change background colors, change background formats, changing of text (e g., using synonyms), change fonts, and / or the like.
[0097] The data modifier 124 passes the other data to the obfuscator 426 and / or the modification Al algorithm 425 based on the rules 133 for the other data in step 818. The modified data 123 is generated in step 820 based on the rules 133 for the other data. The process then goes to step 710.
[0098] As can be seen in Fig. 8 the sensitive data 130 may contain sensitive source code 430 along with sensitive data 130 that is not source code (e.g., images / website data / documents). In this case, there would be an individual flow for the sensitive source code 430 / other data (steps 804-808, 810-814, and / or 816-820) and one flow for the sensitive source data 130 (steps 704-708), each flow using different rules 133 / 433.
[0099] Fig. 9 is a flow diagram of a process for generating vectors of snippets of generated data 131. While the process of Fig. 9 is described for generic sensitive data 130, the process of Fig. 9 will also work for sensitive source code 430 and / or other data.
[0100] The process starts in step 900. The vector Al algorithm 134 / 434 determines, in step 902, if it is supposed to generate vectors of snippets of the generated data 131 / 431. If the vector Al algorithm 134 / 434 is not to generate snippets of the generated data 131 / 431 in step 902, the process of step 902 repeats.
[0101] Otherwise, if the vector Al algorithm 134 / 434 is to generate vectors of snippets (e.g., floating point vectors of snippets that are fifty lines of code, that are based on function calls, that are based on paragraphs, that based on portions of an image, and / or the like) of the generated data 131 / 431, the vector Al algorithm 134 / 434 generates the vectors of the snippets of the generated data 131 / 431 in step 904. The vector Al algorithm 134 / 434 generates vectors of snippets of the sensitive data 130 / 430, the obfuscated data (data from the obfuscator 126 / 426) and / or the modified data 123 / 423 in step 906. The vectors are then compared in step 908. For example, the vectors of the generated data 131 / 431 may be compared to the vectors of the sensitive data 130 / 430 in step 908.
[0102] If there are similarities (e.g., close floating-point vectors) or matches (e.g., match floating-point vectors / integer vectors) in step 910, the snippets of the generated data 131 / 431 / sensitive data 130 / 430 are flagged in step 912. The flagged snippets can be used to provide feedback to the obfuscator 126 / 426, the modification Al algorithm 125 / 424 and / or to the weight pattern Al algorithm 127 / 427. For example, the feedbackmay be to provide to the obfuscator 126 / 426 to change one of the rules 133 / 433. Another example would be to use the snippet of the generated data 131 / 431 as an input parameter to the modification Al algorithm 125 / 425 to indicate not to generate data that is similar to the snippet of the generated data 131 / 431. The process then goes to step 914.
[0103] Otherwise, if there are not any similarities / matches in step 910, the data modifier 124 / 424 determines, in step 914, if the process is complete. If the process is not complete in step 914, the process goes back to step 902. Otherwise, the process ends in step 916.
[0104] Fig. 10 is a block diagram of a blockchain 1000 that is used to store information associated with training an Al algorithm 121 / 421 that is used to prevent leakage of training data. The blockchain 1000 comprises a genesis block 1001, an Al algorithm block 1002, a training set block 1003, a modification Al algorithm block 1004, an obfuscator block 1005, a snippet block 1006, and a change block 1007. The blockchain 1000 comprises links 1010A-1010F. The links 1010A-1010F link the blocks 1001-1007 together as is traditionally done with standard blockchains. While the order of the blocks 1002-1007 are shown in a specific order, the order of the blocks 1002-1007 may vary depending on implementation.
[0105] The genesis block 1001 is the first block in the blockchain 1000. The genesis block 1001 is created at a point in time to start the blockchain 1000.
[0106] The Al algorithm block 1002 is used to track information associated with the Al algorithm 121 / 421. The Al algorithm block 1002 comprises version information (version 1.3 of the Al algorithm 121 / 421), a release date of the Al algorithm 121 / 421 (4-3-2024), a size of the Al algorithm 121 / 421, a hash of the Al algorithm 121 / 421, a hash of the weight(s) of the Al algorithm 121 / 421. The weight hash may comprise individual hashes of each weight of the Al algorithm 121 / 421, a hash based on all the weights of the Al algorithm 121 / 421, and / or the like. In addition, source code of the Al algorithm 121 / 421 may be stored in the Al algorithm block (or a pointer to the source code of the Al algorithm 121 / 421).
[0107] The training set block 1003 comprises information about the training set(s) 122 / 422 / 128 / 428. In Fig. 10, the training set block 1003 comprises information about the initial training set 128 / 428, modified data 123 / 423 information, and / or the final trainingset 122 / 422. The initial training set 128 / 428 information may comprise each component (e.g., source code / document or a pointer to the source code / document) in the initial training set 128 / 428, hashes of each component in the initial training set 128 / 428, and / or the like. The modified data 123 / 423 information may comprise each component (e.g., source code / document or a pointer to the source code / document) of the modified data 123 / 423, hashes of each component of the modified data 123 / 423, changes in the modified data 123 / 423, and / or the like. The final training set 122 / 422 information may comprise each component (e.g., source code / document or a pointer to the source code / document) in the final training set 122 / 422, hashes of each component in the final training set 122 / 422, and / or the like.
[0108] The modification Al algorithm block 1004 comprises information associated with the modification Al algorithm 125 / 425. In Fig. 10, the modification Al algorithm block 1004 comprises version information (version 4.1), a release date of the modification Al algorithm 125 / 425 (1-2-2024), a size of the modification Al algorithm 125 / 425, a hash of the modification Al algorithm 125 / 425, a weight hash of weights of the modification Al algorithm 125 / 425 (e.g., a hash for each weight, a hash of all the weights, and / or the like). In addition, the modification Al algorithm block 1004 may contain source code of the modification Al algorithm 125 / 425 or a pointer to the source code of the modification Al algorithm 125 / 425.
[0109] The obfuscator block 1005 is used to store information about the obfuscator 126 / 426. The obfuscator block 1005 comprises a version of the obfuscator 126 / 426 (version 2.0.1), a release date of the obfuscator 126 / 426 (9-11-2023), a size of the obfuscator 126 / 426, a hash of the obfuscator 126 / 426. In addition, the obfuscator block 1005 may include source code of the obfuscator 126 / 426 or a pointer to the source code of the obfuscator 126 / 426.
[0110] The snippet block 1006 is used to store information associated with the snippets generated by the vector Al algorithm 134 / 434. The snippets may be snippets of the final training set 122 / 422, snippets of the modified data 123 / 423, snippets of the output of the obfuscator 126 / 426, snippets of the sensitive data 130 / 430, snippets of the generated data 131 / 431, and / or the like. In addition, instead of the source code of the snippets, pointers can be used in the snippet block 1006.
[0111] The change block(s) 1007 are used to track any changes to the environment of the Al algorithm 121 / 421, such as if a new version of the Al algorithm 121 / 421, a new version of the modification Al algorithm 125 / 425, a new version of the vector Al algorithm 134 / 434, a new version of the final training set 122 / 422, a new version of the modified data 123 / 423, a new version of the obfuscator 126 / 426, new sensitive data 130 / 430, new released generated data 132 / 432, and / or the like.
[0112] An alternative to the change block 1007 is where new block 1002-1006 is created when there is a change. For example, if a new version of the obfuscator 126 / 426 is used, a new obfuscator block 1005 may be added to the blockchain 1000.
[0113] Although not shown, other blocks may be added to the blockchain 1000, such as a vector Al algorithm block that has information associated with the vector Al algorithm 134 / 434 such as the vector Al algorithm 134 / 434 (or a pointer to the vector Al algorithm 134 / 434), a hash of the vector Al algorithm 134 / 434, a version of the vector Al algorithm 134 / 434, a release date of the vector Al algorithm 134 / 434, and / or the like. Likewise, the blockchain 1000 may include a weight pattern Al algorithm block that contains similar information (e.g., like described above for the vector Al algorithm block).
[0114] In addition, the processes described herein may be part of a Software as a Service (SaaS) cloud service. Different users 102 / organizations would be able to have their own instance of the Al algorithm 121 / 421 that is trained on different final training sets 122 / 422 and uses the processes / portions of the processes described above.
[0115] Various information associated with the above processes may be stored in the blockchain 1000 to keep a record of how the modified source code 423 was used to produce the generated source code 431. For example, the final training set 422, the sensitive source code 430, the modified source code 423, and / or the generated source code 431 may be stored in different blocks in the blockchain 1000. This can be used to prove how the sensitive source code 430 was modified in order to prove that it is a new work instead of a derivative work.
[0116] In addition, the blockchain 1000 may include other information in different blocks 1002-1007 in the blockchain 1000, such as the snippets of source code provided to the modification Al algorithm 425 if there is a match of a snippet by the vector Al algorithm 434, the thresholds used to determine a match, a hash of the code generation Alalgorithm 421, a hash of the modification Al algorithm 125 / 425, a hash of the vector Al algorithm 134 / 434, a hash of the weight pattern Al algorithm 127 / 427, copies of the source code / binaries of the different Al algorithms 121 / 421 / 125 / 425 / 127 / 427 / 134 / 434, the user selected rules 133 / 433 that were used to generate the modified source code 423, sensitive data 130 / 430, and / or the like.
[0117] Fig. 11 is a block diagram of a seventh illustrative system 1100 for training a weight pattern Al algorithm 127 / 427 to identify patterns in weights. The seventh illustrative system 1100 comprises the communication devices 101A-101N, the network 110, and a server 1120.
[0118] The server 1120 may be any hardware device that can host the Al algorithm 123, such as an application server, a cloud service, a communications server, and / or the like. The server 1120 may comprise multiple servers / processing cores that host the Al algorithm 121. The server 1120 further comprises a weight pattern Al algorithm 127 / 427, an Al algorithm manager 1122, the Al algorithm 121, training set(s) 1124, sets of weights 1125, base Al inputs 1126, Al output data 1127, neural network architecture information 1128, and an input scope checker 1129.
[0119] The weight pattern Al algorithm 127 / 427 is an Al algorithm that is trained on different versions of the sets of weights 1125 of the Al algorithm 121. The weight pattern Al algorithm 127 / 427 identifies patterns in the sets of weights 1125 in relation to the training set(s) 1124. In addition, the weight pattern Al algorithm 127 / 427 can identify how changing weights of the Al algorithm 121 affect how Al output data 1127 by using the known base Al inputs 1126.
[0120] The Al algorithm manager 1122 is used to manage the weight pattern Al algorithm 121, the Al algorithm 121, and processes associated with these, such as providing the base Al inputs 1126 to the Al algorithm 121, changing the weights of the Al algorithm 121, storing off information in memory, and / or the like.
[0121] The Al algorithm 121 may be any type of Al algorithm that uses sets of weights 1125, such as a linear regression Al algorithm, a gradient descent Al algorithm, a random forest Al algorithm, a Generative Adversarial Neural Network (GANN) Al algorithm, a Large Language Model (LLM) Al algorithm, neural network Al algorithm, and / or the like. The Al algorithm 121 is initially trained using the training set(s) 1124.
[0122] The Al algorithm 121 may comprise multiple Al algorithms 121. For example, there may be multiple types / sizes of Al algorithms 121 that are trained using the training sets 1124 to produce multiple sets of weights 1125 for use by the weight pattern Al algorithm 127 / 427. The Al algorithms 121 may be part of a library of Al algorithms.
[0123] The training sets 1124 may comprises different types of information / data (e.g., text, images, audio, video, etc.). The training sets 1124 may comprise specific training sets 1124 and / or general training sets 1124. For example, a specific training set 1124 may only include source code for software development. A general training set 1124 may comprise information scraped from the Internet or some other generic source of information.
[0124] The sets of weights 1125 are generated when the Al algorithm 121 is initially trained using a training set 1124. The sets of weights 1125 are associated with nodes in a neural network. The sets of weights 1125 are numbers that are used as an input into a node in the Al algorithm 121. The set of weights 1125 may be a set of weights 1125 that are generated by the weight pattern Al algorithm 127 / 427.
[0125] The base Al inputs 1126 are a series of defined prompts (e.g., text images, numeric tabular data, audio, video, etc.) that are provided as input to the Al algorithm 121. For example, the base Al inputs 1126 may comprise ten thousand different questions / groups of queries / prompts that are provided to the Al algorithm 121. The base Al inputs 1126 are used to create a baseline of Al output data 1127. The baseline Al output data 1127 is compared to Al output data 1127 that is generated when one or more of the weights in the set of weights 1125 are changed. For example, the one or more weights in the set of weights 1125 may be changed in real-time while the Al algorithm 121 is running to determine how the changes affect the Al output data 1127.
[0126] The Al output data 1127 is data / information that is generated by the Al algorithm 121 based on the Al input data. The Al output data 1127 may comprises a set of Al output data 1127 that is stored off based on different Al input data.
[0127] The neural network architecture information 1128 is information that is used to determine the structure of the neural network of the Al algorithm 121. The neural network architecture information 1128 defines the nodes of the Al algorithm 121 and how the nodes of the Al algorithm 121 are connected. The neural network architectureinformation 1128 is used by the weight pattern Al algorithm 127 / 427 for training Al algorithms 121 that have different types / sizes of neural networks. For example, a large LLM can have billions of nodes / weights in the set of weights 1125 and a small GANN Al algorithm 121 may only have millions of nodes / weights in the set of weights 1125.
[0128] The input scope checker 1129 is used to check a scope of the Al algorithm inputs to make sure that the Al algorithm inputs are within the same scope of the training set 1124 that the Al algorithm 121 was trained on. The input scope checker 1129 is used to help the Al algorithm 121 to not have as many hallucinations. The input scope checker 1129 may flag what Al inputs are out of scope and allow a user to proceed or not with the out-of-scope Al inputs. The input scope checker 1129 may provide suggestions on alternative Al inputs.
[0129] The malicious weight pattern database 1130 comprises different compromised sets of weights 1125 for different types of compromised Al algorithms 121. In other words, the malicious weight pattern database 1130 is a dictionary of known compromised sets of weights 1125. The malicious weight pattern database 1130 may be used by multiple parties (e.g., different companies) to lookup different compromised weight patterns associated with different types of compromised Al algorithms 121. For example, there may be multiple sets of weights 1125 for each of ten different types of compromised Al algorithms 121. The malicious weight pattern database 1130 may contain the compromised sets of weights 1125, information about how the compromised weight sets of weights 1125 affect the Al algorithm 121, and / or the like.
[0130] Fig. 12 is a block diagram of an exemplary neural network 200 of an Al algorithm 121. While Fig. 12 is an exemplary fully connected feedforward LLM, other types of architectures may be used, such as convolutional neural network (CNN), recurrent neural network (RNN), transformer architecture, etc. The neural network 1200 comprises nodes 1201A1, 1201B1, 1201BN, and 1201N1. The nodes 1201 are linked together by links 1202B1, 1202BN, 1202N1, and 1202NN. The links 1202 are basically function calls between the nodes 1201 (i.e., software of the nodes 1201). Where a link 1202 comes into a node 1201, there is an associated weight 1205. In Fig. 12, the node 1201B1 has the weight 1205B1, the node 1201BN has the weight 1205BN, and the node 1201N1 has the weights 1205N1 / 1205NN.
[0131] The neural network 1200 of the Al algorithm 121 has three node layers 1203: 1) an input layer 1203 A, a hidden layer 1203B, and an output layer 1203N. The links 1202 between the nodes 1201 form the link layers 1204B / 1204N.
[0132] While Fig. 12 is an exemplary embodiment of a simple feedforward neural network, any practical neural network 1200 is typically much larger. For example, a neural network 1200 of a LLM may comprise millions of nodes 1201, thousands of link layers 1204, and billions or even trillions of weights 1205. Because of the extremely large number of nodes 1201 / weights 1205 in a typical neural network 1200, the complexity of identifying changes and how training affects weights 1205 is an extremely complex process. In addition, determining how specific weights 1205 and Al inputs affect the Al output data 1127 of the Al algorithm 121 is far too complex a process to do manually for a neural network 1200 that has even a few hundred nodes 1201, much less for a neural network 1200 that has millions of nodes 1201.
[0133] Fig. 13 is a block diagram of an eighth illustrative system 1300 for training a weight pattern Al algorithm 127 / 427 by changing Al inputs. One of the key issues with Al algorithms 121 is that it is difficult to identify how the different weights 1205 in an Al algorithm 121 affect the Al output data 1127 of the Al algorithm 121. One way to deal with this is to use the weight pattern Al algorithm 127 / 427 to learn how changes in weights 1205 affect the Al output data 1127 from the Al algorithm 121. Fig. 13 illustrates an exemplary embodiment of where the weight pattern Al algorithm 127 / 427 is trained to determine effects that happen to the Al output data 1127 of the Al algorithm 121 when different weight(s) 1205 of the Al algorithm 121 are changed (e.g., in realtime). Fig. 13 assumes that the Al algorithm 121 has already been trained.
[0134] Once the Al algorithm 121 is initially trained, the process provides the base Al inputs 1126 and then identifies the corresponding initial Al output data 1127 associated with the base Al inputs 1126 using the weights 1205 from the initially trained Al algorithm 121. The base Al inputs 1126 may be a series and / or a group of specific types of base Al inputs 1126 depending on the initial training set 1124.
[0135] The Al algorithm manager 1122 then changes one or more weight(s) 1205 (e.g., in real-time). The weight pattern Al algorithm 127 / 427 then provides the same base Al inputs 1126 as input to the Al algorithm 121 and then gets the corresponding Al outputdata 1127. The Al output data 1127 (where the weights 1205 are changed) is then compared to the original / initial Al output data 1127 (generated using the original weights 1205) to determine the effect of the change in the weight(s) 1205. Another option would be to also compare the current Al output data 1127 to the previous Al output data 1127. The process of changing the weights 205 in real-time is repeated multiple times using the base Al inputs 1126 to learn the effect of changing the weight(s) 1205 in relation to the specific base Al inputs 1126.
[0136] The changing of the weight(s) 1205 may include changing individual weight(s) 1205 (e.g., weight 1205B1), changing groups of weights 1205 (e.g., weights 1205B1, 1205BN, and 1125NN), changing layers of weights 1205 (e.g., the weights 1205B1 / 1205BN of the nodes 1201B1 / 201BN of the hidden layer 1203B), changing partial layers of weights 1205 (e.g., weight 1205B1), changing flow(s) of weights 1205 (e.g., weight 1205B1 and 1205N1 that are linked together by link 1202N1), changing partial flows of weights 205 (e.g., weight 1205NN), a combination of these, and / or the like.
[0137] For example, using the next base Al input 1126, the various weight(s) 1205 are changed and the Al output data 1127 for each changed weight 1205 is compared to the original Al output data 1127 for that particular base input parameter 1126 to determine the effect of the change to the weight(s) 1205. Thus, the weight pattern Al algorithm 127 / 427 can learn over time how changes to the weights 1205 affect Al output data 1127 of the Al algorithm 121. The base Al inputs 1126 could be grouped in specific areas of questions / data or a range of questions / data that progressively changes the base Al inputs 1126.
[0138] The comparisons of the various weight changes / base Al inputs 1126 can be used to learn over time to learn how to change the Al algorithm 121 without having to retrain the Al algorithm 121. For example, a simple Al algorithm 121 that has one hundred weights 1205 and has ten base Al inputs 1126 is assumed. After learning the initial Al output data 127 for the ten base Al inputs 1126 (using the original weights 1205) the Al algorithm manager 1122 can individually change each of the one hundred weights 1205 (could be multiple changes to each individual weight 1205 (e.g., using steps of change of .1 (assuming a maximum weight value of 1)) with the same base Al input1126) with the first of the ten base Al inputs 1126. The Al algorithm manager 1 122 can change different weight settings as well (e.g., a group or flow of weights 1205). The weight pattern Al algorithm 127 / 427 can then learn how the changes to the weights 1205 affect the Al output data 1127. This process is repeated for the remaining nine of the ten base Al inputs 1126 / remaining one hundred weights 1205. Thus, the weight pattern Al algorithm 127 / 427 learns how the changes to the weights 1205 affect the Al output data 1127. As this process occurs, the weight pattern Al algorithm 127 / 427 is trained on what happens to the Al output data 1127 when the weights 1205 are changed. The trained weight pattern Al algorithm 127 / 427 can then be used to create a completely new set of weights 1125 for the Al algorithm 121 (e.g., as described in Fig. 16).
[0139] This is different from fine tuning an Al algorithm 121 where a new training set 1124 is used to further train an already trained Al algorithm 121. In Fig. 13, the base Al inputs 1126 are not being used to fine-tuning an existing Al algorithm 121 but are instead being used to learn how changes to the weights 1205 affect the Al output data 1127 so that a newly trained Al algorithm 121 can be created without retraining the Al algorithm 121. For example, the process of Fig. 13 will likely create a completely new set of weights 1125 where fine-tuning typically only changes a small portion of the weights 1205. If all the weights 1205 are changed with fine tuning, it typically takes an extremely large amount of processing power / time versus the process described in Fig. 13.
[0140] Fig. 14 is a block diagram of a ninth illustrative system 1400 for training a weight pattern Al algorithm 127 / 427 by using a plurality of training sets 1124. One of the primary advantages of the weight pattern Al algorithm 127 / 427 is the ability to identify patterns in data. The embodiment described in Fig. 14 leverages the ability of an Al algorithm (i.e., the weight pattern Al algorithm 127 / 427) to identify patterns that occur based on the weights 1205 that are created each time an Al algorithm 121 is trained.
[0141] The weight pattern Al algorithm 127 / 427 uses weight patterns of the Al algorithm 121 learned over time as a way to produce a new set of weights 1125 of the Al algorithm 121 rather than retraining the Al algorithm 121.
[0142] In Fig. 14, the weight pattern Al algorithm 127 / 427 identifies patterns in the weights 1205 as the Al algorithm 121 is trained using different training sets 1124. Thetraining sets 1124 may vary and / or may be similar. For example, a first training set 1124 produces a first set of weights 1125 and a second training set 1124 produces a different set of weights 1125. If the two training sets 1124 are both used, the process will be able to identify the weight patterns in comparison to the data in the two training sets 1124. The weight patterns could be groups of weight patterns that produce a specific type of trained Al algorithm 121. The weight patterns are learned in conjunction with the data in the training sets 1124.
[0143] To give a simple illustrative example, assume that the there are three training sets 1124: 1) a training set 1124 to identify cats, 2) a training set 1124 to identify dogs, and 3) a training set 1124 to identify sheep. The weight pattern Al algorithm 127 / 427 would learn the weight patterns for the three training sets 1124. To produce a new training set 1124 for identifying cows, the Al inputs may include a training set 1124 for cows and / or information about cows (e.g., pictures of different cows). Based on the patterns learned from the three training sets 1124 and the Al inputs (the training set 1124 for cows), the trained weight pattern Al algorithm 127 / 427 will build a new set of weights 1125 that are used to identify cows using the input data and the correlations in the three training sets 1124.
[0144] The learning process could be implemented on a smaller Al algorithm 121 (e.g., an Al algorithm 121 that has a smaller neural network 1200) and then applied to a larger Al algorithm 121. For example, the weight pattern Al algorithm 127 / 427 may learn specific patterns based on specific training sets 1124 used on a smaller Al algorithm 121. Based on the patterns in the smaller Al algorithm 121, the weight pattern Al algorithm 127 / 427 could then apply the weight patterns learned from the smaller Al algorithm 121 to a larger Al algorithm 121.
[0145] Fig. 15 is a block diagram of a tenth illustrative system 1500 for training a weigh pattern Al algorithm 127 / 427 to produce weights 1205 for a plurality of Al algorithms 121 that have different neural network architectures. In Fig. 15, the weight pattern Al algorithm 127 / 427 is built first using supervised learning. The training sets 1124 consists of training sets 1124 from already trained Al algorithms 121 using standard gradient descent-based backpropagation Al algorithms 121 trained on various tasks. The training sets 1124 for these Al algorithms 121 (with different neural networkarchitectures) and the generated set of weights 1125 for each of the input neural network architectures may be the same or different. The weight pattern Al algorithm 127 / 427 is trained on training sets 1124 using the processes described herein and evaluated on a test training set 1124. The commonly used neural network architectures are added to a neural network architecture library to be used when weight pattern Al algorithm 127 / 427 is deployed.
[0146] Once a trained weight pattern Al algorithm 127 / 427 is available for a specific type of neural network 1200, the trained weight pattern Al algorithm 127 / 427 can be used for generating new sets of weights 1125 for new tasks for the specific type of neural network 1200. The training sets 1124 for a task and a corresponding neural network architecture (e.g., chosen from the neural network architecture information 1128) is input to the weight pattern Al algorithm 127 / 427 to generate the new sets of weights 1125.
[0147] Fig. 16 is a block diagram of an eleventh illustrative system 1600 for creating a new set of weights 1125 for an Al algorithm 121. The ultimate goal is to produce a trained Al algorithm 121 that does not require the traditional training processes. Fig. 16 shows how the weight pattern Al algorithm 127 / 427 may be used once the weight pattern Al algorithm 127 / 427 has been trained to identify weight patterns (e.g., as described above).
[0148] A user can provide to the Al algorithm 121, Al algorithm inputs 1601 that are used to define the scope of a new Al algorithm 121. For example, the Al algorithm inputs 1 601 may be the images of cows as described above. The Al algorithm inputs 1601 may be checked by the input scope checker 1129 that uses the training sets 1124 to make sure that the Al algorithm inputs 1601 are within the scope of the training sets 1124 (e.g., images) used to train the weight pattern Al algorithm 127 / 427. This can include filtering out any Al algorithm inputs 1601 that are out of scope, providing alternative Al algorithm inputs 1601 (e.g., learned Al algorithm inputs 1601), not allowing an Al algorithm input parameter 1601, and / or the like.
[0149] The output of the input scope checker 1129 (e.g., the training set 1124 for cows) is input to the trained weight pattern Al algorithm 127 / 427. The trained weight pattern Al algorithm 127 / 427 then generates a new set of weights 1125 that are used to producethe newly trained Al algorithm 121 . Thus, a newly trained Al algorithm 121 is created that does not use traditional training methods.
[0150] Fig. 17 is a flow diagram of a process for training a weight pattern Al algorithm 127 / 427 by using base Al inputs 126. Illustratively, the communication devices 101A- 101N, the server 1120, the weight pattern Al algorithm 127 / 427, the Al algorithm manager 1122, the Al algorithm 121, the input scope checker 1129, and the neural network 1200 are stored-program-controlled entities, such as a computer or microprocessor, which performs the method of Figs. 17-18 and the processes described herein by executing program instructions stored in a computer readable storage medium, such as a memory (i.e., a computer memory, a hard disk, and / or the like). Although the methods described in Figs. 17-18 are shown in a specific order, one of skill in the art would recognize that the steps in Figs. 17-18 may be implemented in different orders and / or be implemented in a multi-threaded environment. Moreover, various steps may be omitted or added based on implementation.
[0151] The process starts in step 1700. The Al algorithm 121 is trained using an initial training set 1124 to produce an initial set of weights 1125 in step 1702. The initial set of weights 1125 are retrieved and saved off in step 1704. For example, the initial set of weights 1125 may be saved off in a database. The Al algorithm manager 1122 provides, in step 1706, a series of base Al inputs 1126 to the Al algorithm 121 to produce an initial set of Al output data 1127. Although not shown in Fig. 17, the initial set of Al output data 1127 may also be stored off in step 1706.
[0152] The Al algorithm manager 1122 changes one or more individual weights 1205 of the Al algorithm 121 in step 1708. For example, the Al algorithm manager 1122 may change a single weight 1205 of the initial set of weights 1125 of the Al algorithm 121 in step 1708. The Al algorithm manager 1122 provides the series of base Al inputs 1126 to the Al algorithm 121 with the changed weight(s) 1205 to produce corresponding Al output data 1127. Although not shown in step 1710, the corresponding Al output data 127 may be stored off in step 1710.
[0153] The Al algorithm manager 1122 determines, in step 1712, if there are more weight changes. If there are more weight changes in step 1712, the process goes back to step 1708. For example, if the neural network 1200 of the Al algorithm 121 has onebillion weights 1205 and the Al algorithm manager 122 changes each weight 1205 ten times, there would be a total of ten billion changes to the set of weights 1125. In other words, in this example, steps 1708 and 1710 would be repeated ten billion times.
[0154] If there are not any more weight changes in step 1712, the weight pattern Al algorithm 1121 is trained by determining how the changes in the weights 1205 affect the corresponding Al output data 1127 in relation to the base Al inputs 1126 in step 1714.
[0155] The Al algorithm manager 1122 determines, in step 1716, if the process is complete. If the process is not complete in step 1716, the process goes back to step 1702. Otherwise, the process ends in step 1718.
[0156] Fig. 18 is a flow diagram of a process for training a weight pattern Al algorithm 127 / 427 by using a plurality of training sets 1124. The process starts in step 1800. The Al algorithm 121 is trained using a training set 1124 to produce a corresponding set of weights 1125 in step 1802. The weight pattern Al algorithm 127 / 427 retrieves and saves off the corresponding set of weights 1125 in step 1804. The Al algorithm manager 1122 determines, in step 1806, if more training sets 1124 are going to be used to train the Al algorithm 121. If there are more training sets 1124 in step 1806, the process goes back to step 1802. For example, there may be tens of thousands of training sets 1124 that are used to train the Al algorithm 121 to produce tens of thousands of corresponding sets of weights 125.
[0157] Otherwise, if there are no more training sets 124 in step 1806, the weight pattern Al algorithm 127 / 427 is trained based on the relationship between information in each of the training sets 1124 and the corresponding set of weights 1125 in step 1808. For example, the weight pattern Al algorithm 127 / 427 learns how different information in the training sets 1124 cause different weights 1205 to have different values.
[0158] The weight pattern Al algorithm 127 / 427 determines, in step 1810, if there are any Al algorithm inputs 1601 with new scope in step 1810. If there are not any received Al algorithm inputs 1601 with new scope in step 1810, the process of step 1810 repeats. Otherwise, if there are received Al algorithm inputs 1601 with new scope in step 1810, the weight pattern Al algorithm 127 / 427 generates, in step 1812, based on the received Al algorithm inputs 1601 with the new scope, a new set of weights 1125 for the new version of the Al algorithm 121.
[0159] The weight pattern Al algorithm 127 / 427, determines, in step 1814, if the process is complete. If the process is not complete in step 1814, the process goes back to step 1802 (or optionally could go back to step 1810). Otherwise, if the process is complete in step 1814, the process ends in step 1816.
[0160] Examples of the processors as described herein may include, but are not limited to, at least one of Qualcomm® Snapdragon® 800 and 801, Qualcomm® Snapdragon® 610 and 615 with 4G LTE Integration and 64-bit computing, Apple® A7 processor with 64-bit architecture, Apple® M7 motion coprocessors, Samsung® Exynos® series, the Intel® Core™ family of processors, the Intel® Xeon® family of processors, the Intel® Atom™ family of processors, the Intel Itanium® family of processors, Intel® Core® i5- 4670K and i7-4770K 22nm Haswell, Intel® Core® i5-3570K 22nm Ivy Bridge, the AMD® FX™ family of processors, AMD® FX-4300, FX-6300, and FX-8350 32nm Vishera, AMD® Kaveri processors, Texas Instruments® Jacinto C6000™ automotive infotainment processors, Texas Instruments® OMAP™ automotive-grade mobile processors, ARM® Cortex™-M processors, ARM® Cortex-A and ARM926EJ-S™ processors, other industry-equivalent processors, and may perform computational functions using any known or future-developed standard, instruction set, libraries, and / or architecture.
[0161] Any of the steps, functions, and operations discussed herein can be performed continuously and automatically.
[0162] However, to avoid unnecessarily obscuring the present disclosure, the preceding description omits a number of known structures and devices. This omission is not to be construed as a limitation of the scope of the claimed disclosure. Specific details are set forth to provide an understanding of the present disclosure. It should however be appreciated that the present disclosure may be practiced in a variety of ways beyond the specific detail set forth herein.
[0163] Furthermore, while the exemplary embodiments illustrated herein show the various components of the system collocated, certain components of the system can be located remotely, at distant portions of a distributed network, such as a LAN and / or the Internet, or within a dedicated system. Thus, it should be appreciated, that the components of the system can be combined in to one or more devices or collocated on aparticular node of a distributed network, such as an analog and / or digital telecommunications network, a packet-switch network, or a circuit-switched network. It will be appreciated from the preceding description, and for reasons of computational efficiency, that the components of the system can be arranged at any location within a distributed network of components without affecting the operation of the system. For example, the various components can be located in a switch such as a PBX and media server, gateway, in one or more communications devices, at one or more users’ premises, or some combination thereof. Similarly, one or more functional portions of the system could be distributed between a telecommunications device(s) and an associated computing device.
[0164] Furthermore, it should be appreciated that the various links connecting the elements can be wired or wireless links, or any combination thereof, or any other known or later developed element(s) that is capable of supplying and / or communicating data to and from the connected elements. These wired or wireless links can also be secure links and may be capable of communicating encrypted information. Transmission media used as links, for example, can be any suitable carrier for electrical signals, including coaxial cables, copper wire and fiber optics, and may take the form of acoustic or light waves, such as those generated during radio-wave and infra-red data communications.
[0165] Also, while the flowcharts have been discussed and illustrated in relation to a particular sequence of events, it should be appreciated that changes, additions, and omissions to this sequence can occur without materially affecting the operation of the disclosure.
[0166] A number of variations and modifications of the disclosure can be used. It would be possible to provide for some features of the disclosure without providing others.
[0167] In yet another embodiment, the systems and methods of this disclosure can be implemented in conjunction with a special purpose computer, a programmed microprocessor or microcontroller and peripheral integrated circuit element(s), an ASIC or other integrated circuit, a digital signal processor, a hard-wired electronic or logic circuit such as discrete element circuit, a programmable logic device or gate array such as PLD, PLA, FPGA, PAL, special purpose computer, any comparable means, or the like. In general, any device(s) or means capable of implementing the methodology illustratedherein can be used to implement the various aspects of this disclosure. Exemplary hardware that can be used for the present disclosure includes computers, handheld devices, telephones (e.g., cellular, Internet enabled, digital, analog, hybrids, and others), and other hardware known in the art. Some of these devices include processors (e.g., a single or multiple microprocessors), memory, nonvolatile storage, input devices, and output devices. Furthermore, alternative software implementations including, but not limited to, distributed processing or component / object distributed processing, parallel processing, or virtual machine processing can also be constructed to implement the methods described herein.
[0168] In yet another embodiment, the disclosed methods may be readily implemented in conjunction with software using object or object-oriented software development environments that provide portable source code that can be used on a variety of computer or workstation platforms. Alternatively, the disclosed system may be implemented partially or fully in hardware using standard logic circuits or VLSI design. Whether software or hardware is used to implement the systems in accordance with this disclosure is dependent on the speed and / or efficiency requirements of the system, the particular function, and the particular software or hardware systems or microprocessor or microcomputer systems being utilized.
[0169] In yet another embodiment, the disclosed methods may be partially implemented in software that can be stored on a storage medium, executed on programmed general-purpose computer with the cooperation of a controller and memory, a special purpose computer, a microprocessor, or the like. In these instances, the systems and methods of this disclosure can be implemented as program embedded on personal computer such as an applet, JAVA® or CGI script, as a resource residing on a server or computer workstation, as a routine embedded in a dedicated measurement system, system component, or the like. The system can also be implemented by physically incorporating the system and / or method into a software and / or hardware system.
[0170] Although the present disclosure describes components and functions implemented in the embodiments with reference to particular standards and protocols, the disclosure is not limited to such standards and protocols. Other similar standards and protocols not mentioned herein are in existence and are considered to be included in thepresent disclosure. Moreover, the standards and protocols mentioned herein, and other similar standards and protocols not mentioned herein are periodically superseded by faster or more effective equivalents having essentially the same functions. Such replacement standards and protocols having the same functions are considered equivalents included in the present disclosure.
[0171] The present disclosure, in various embodiments, configurations, and aspects, includes components, methods, processes, systems and / or apparatus substantially as depicted and described herein, including various embodiments, sub combinations, and subsets thereof. Those of skill in the art will understand how to make and use the systems and methods disclosed herein after understanding the present disclosure. The present disclosure, in various embodiments, configurations, and aspects, includes providing devices and processes in the absence of items not depicted and / or described herein or in various embodiments, configurations, or aspects hereof, including in the absence of such items as may have been used in previous devices or processes, e.g., for improving performance, achieving ease and\or reducing cost of implementation.
[0172] The foregoing discussion of the disclosure has been presented for purposes of illustration and description. The foregoing is not intended to limit the disclosure to the form or forms disclosed herein. In the foregoing Detailed Description for example, various features of the disclosure are grouped together in one or more embodiments, configurations, or aspects for the purpose of streamlining the disclosure. The features of the embodiments, configurations, or aspects of the disclosure may be combined in alternate embodiments, configurations, or aspects other than those discussed above. This method of disclosure is not to be interpreted as reflecting an intention that the claimed disclosure requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment, configuration, or aspect. Thus, the following claims are hereby incorporated into this Detailed Description, with each claim standing on its own as a separate preferred embodiment of the disclosure.
[0173] Moreover, though the description of the disclosure has included description of one or more embodiments, configurations, or aspects and certain variations and modifications, other variations, combinations, and modifications are within the scope ofthe disclosure, e.g., as may be within the skill and knowledge of those in the art, after understanding the present disclosure. It is intended to obtain rights which include alternative embodiments, configurations, or aspects to the extent permitted, including alternate, interchangeable and / or equivalent structures, functions, ranges or steps to those claimed, whether or not such alternate, interchangeable and / or equivalent structures, functions, ranges or steps are disclosed herein, and without intending to publicly dedicate any patentable subject matter.
Claims
CLAIMSWhat is claimed is:
1. A system comprising: a microprocessor; and a computer readable medium, coupled with the microprocessor and comprising microprocessor readable and executable instructions that, when executed by the microprocessor, cause the microprocessor to: receive sensitive data, wherein the sensitive data is information that is not wanted to be leaked by an Al algorithm; pass the sensitive data to at least one of: a modification Al algorithm and an obfuscator, to generate modified data; and train the Al algorithm using the generated modified data.
2. The system of claim 1, wherein sensitive data is sensitive source code and wherein the obfuscator modifies the sensitive source code based on at least one of: obfuscating the sensitive source code, removing comments from the sensitive source code, and changing a format of the sensitive source code.
3. The system of claim 1, wherein sensitive data is sensitive source code and wherein the modification Al algorithm modifies the sensitive source code based on at least one of: changing a sensitive source code naming attribute, changing a naming attribute of comments in the sensitive source code, and optimizing the sensitive source code.
4. The system of claim 1, wherein the sensitive data is sensitive source code, wherein the modified data is modified source code, wherein the modification Al algorithm modifies the sensitive source code based on one or more prompts that are provided to the modification algorithm, and wherein the one or more prompts that are provided to the modification Al algorithm tell the modification Al algorithm to generate the modified source code that has a same function as the sensitive source code but has a different code structure than the sensitive source code.
5. The system of claim 1, wherein the microprocessor readable and executable instructions further cause the microprocessor to: generate vectors of snippets of data generated by the Al algorithm; and compare the generated vectors of the snippets of the data generated by the Al algorithm to vectors of snippets based on at least one of: the sensitive data, obfuscated data, and the generated modified data.
6. The system of claim 5, wherein, based on a match of at least one of the vectors of snippets of the data generated by the Al algorithm, a weight pattern Al algorithm is used to change weights in the Al algorithm.
7. The system of claim 6, wherein a prompt provided to the weight pattern Al algorithm is based on the matched at least one of the vectors of snippets of data generated by the Al algorithm.
8. The system of claim 1, wherein the sensitive data comprises one or more documents and wherein the obfuscator modifies the one or more documents by doing at least one of: removing a footnote, removing a reference, removing an embedded image, removing a name, removing a header, removing a page number, and removing a link.
9. The system of claim 1, wherein the sensitive data comprises one or more documents and wherein the Al modification Al algorithm modifies the one or more documents data by doing at least one of: modifying a structure of the one or more documents, rewriting the one or more documents, replacing a name in the one or more documents, and replacing a word in the one or more documents with a synonym.
10. The system of claim 1, wherein the sensitive data comprises sensitive source code and sensitive non-source code data and wherein the modification Al algorithm and / or the obfuscator apply different rules based the sensitive source code and the sensitive non-source code data.
11. The system of claim 1, wherein the sensitive data comprises an image and wherein the obfuscator modifies the image by doing at least one of: removing atext from the image, removing a color from the image, adding a color to the image, changing a background of the image, and grey scaling the image.
12. The system of claim 1, wherein the sensitive data comprises an image and wherein the modification Al algorithm modifies the image by doing at least one of: reformatting the image, modifying the image based on Al input data, and creating a new image based on the Al input data.
13. The system of claim 1, wherein the sensitive data comprises music data and / or audio data and wherein the modification Al algorithm and / or the obfuscator modifies the music data and / or the audio data by at least one of: changing a voice of a speaker in the audio data and / or music data, removing a voice in the audio data and / or music data, adding a new voice in the audio data and / or music data, modifying a voice in the audio data and / or music data, changing a musical rhythm, creating a new Al generated musical recording based on Al input data, changing an intensity of an instrument, removing an instrument, adding an instrument, and changing a pitch.
14. The system of claim 1, wherein the sensitive data comprises website data and / or graphical user interface data and wherein the modification Al algorithm and / or the obfuscator modifies the website data and / or the graphical user interface data by at least one of: changing a graphical object, moving a graphical object, changing a color of a graphical object, changing a text of a graphical object, adding a tab, removing a tab, reorganizing a tab, changing a button, removing a button, adding a button, adding an animation, removing an animation, changing an animation, changing a menu, adding a menu item, and removing a menu item.
15. The system of claim 1, wherein information associated with the Al algorithm is stored in a blockchain and wherein the information associated with the Al algorithm that is stored in the blockchain comprises at least one of: the sensitive data, the generated modified data, a version of the Al algorithm, a release date of the Al algorithm, a size of the Al algorithm, a hash of the Al algorithm, a hash of weights of the Al algorithm, an initial training set, a final training set, snippets ofdata generated by the Al algorithm, a hash of the modification Al algorithm, a hash of the obfuscator, rules for the modification Al algorithm, and rules for the obfuscator.
16. The system of claim 1, wherein information associated with the Al algorithm is stored in a blockchain and wherein the blockchain comprises at least one of: a training set block, a modification Al algorithm block, an obfuscator block, a snippet block, a change block, a vector Al algorithm block, and a weight pattern Al algorithm block.
17. A method comprising: receiving, by a microprocessor, sensitive data, wherein the sensitive data is information that is not wanted to be leaked by an Al algorithm; passing, by the microprocessor, the sensitive data to at least one of: a modification Al algorithm and an obfuscator, to generate modified data; and training, by the microprocessor, the Al algorithm using the generated modified data.
18. The method of claim 17, wherein sensitive data is sensitive source code and wherein the obfuscator modifies the sensitive source code based on at least one of: obfuscating the sensitive source code, removing comments from the sensitive source code, and changing a format of the sensitive source code.
19. The method of claim 17, wherein sensitive data is sensitive source code and wherein the modification Al algorithm modifies the sensitive source code based on at least one of: changing a sensitive source code naming attribute, changing a naming attribute of comments in the sensitive source code, and optimizing the sensitive source code.
20. The method of claim 17, wherein the sensitive data is sensitive source code, wherein the modified data is modified source code, wherein the modification Al algorithm modifies the sensitive source code based on one or more prompts that are provided to the modification algorithm, and wherein the one or more prompts that are provided to the modification Al algorithm tell the modification Alalgorithm to generate the modified source code that has a same function as the sensitive source code but has a different code structure than the sensitive source code.
21. The method of claim 17, further comprising: generating vectors of snippets of data generated by the Al algorithm; and comparing the generated vectors of the snippets of the data generated by the Al algorithm to vectors of snippets based on at least one of: the sensitive data, obfuscated data, and the generated modified data.
22. The method of claim 21 , wherein, based on a match of at least one of the vectors of snippets of the data generated by the Al algorithm, a weight pattern Al algorithm is used to change weights in the Al algorithm.
23. The method of claim 22, wherein a prompt provided to the weight pattern Al algorithm is based on the matched at least one of the vectors of snippets of data generated by the Al algorithm.
24. The method of claim 17, wherein the sensitive data comprises one or more documents and wherein the obfuscator modifies the one or more documents by doing at least one of: removing a footnote, removing a reference, removing an embedded image, removing a name, removing a header, removing a page number, and removing a link.
25. The method of claim 17, wherein the sensitive data comprises one or more documents and wherein the Al modification Al algorithm modifies the one or more documents by doing at least one of: modifying a structure of the one or more documents, rewriting the one or more documents, replacing a name in the one or more documents, and replacing a word in the one or more documents with a synonym.
26. The method of claim 17, wherein the sensitive data comprises sensitive source code and sensitive non-source code data and wherein the modification Al algorithm and / or the obfuscator apply different rules based the sensitive source code and the sensitive non-source code data.
27. The method of claim 17, wherein the sensitive data comprises an image and wherein the obfuscator modifies the image by doing at least one of: removing a text from the image, removing a color from the image, adding a color to the image, changing a background of the image, and grey scaling the image.
28. The method of claim 17, wherein the sensitive data comprises an image and wherein the modification Al algorithm modifies the image by doing at least one of: reformatting the image, modifying the image based on Al input data, and creating a new image based on the Al input data.
29. The method of claim 17, wherein the sensitive data comprises music data and / or audio data and wherein the modification Al algorithm and / or the obfuscator modifies the music data and / or the audio data by at least one of: changing a voice of a speaker in the audio data and / or music data, removing a voice in the audio data and / or music data, adding a new voice in the audio data and / or music data, modifying a voice in the audio data and / or music data, changing a musical rhythm, creating a new Al generated musical recording based on Al input data, changing an intensity of an instrument, removing an instrument, adding an instrument, and changing a pitch.
30. The method of claim 17, wherein the sensitive data comprises website data and / or graphical user interface data and wherein the modification Al algorithm and / or the obfuscator modifies the website data and / or the graphical user interface data by at least one of: changing a graphical object, moving a graphical object, changing a color of a graphical object, changing a text of a graphical object, adding a tab, removing a tab, reorganizing a tab, changing a button, removing a button, adding a button, adding an animation, removing an animation, changing an animation, changing a menu, adding a menu item, and removing a menu item.
31. A non-transient computer readable medium having stored thereon instructions that cause a processor to execute a method, the method comprising instructions to: receive sensitive data, wherein the sensitive data is information that is not wanted to be leaked by an Al algorithm;pass the sensitive data to at least one of: a modification Al algorithm and an obfuscator, to generate modified data; and train the Al algorithm using the generated modified data.
Citation Information
Patent Citations
Code vulnerability remediation
US20210124830A1
Systems, methods, and storage media for creating secured transformed code from input code using a neural network to obscure a transformation function
US20210303662A1
System architecture for providing privacy by design
US20220067204A1
Deidentifying code for cross-organization remediation knowledge
US20230153459A1
Systems and methods for sanitizing sensitive data and preventing data leakage using on-demand artificial intelligence models
US20240111891A1