Method of protecting content from unauthorized access
By employing modular arithmetic to rearrange database or file fields within a database or file, the method effectively addresses the vulnerabilities of existing data security methods, providing a robust and computationally efficient solution to protect data from unauthorized access.
Patent Information
- Application Number
- JP2025026627
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-02-24
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-02-24
Smart Images

Figure 2025084820000003 
Figure 2025084820000004 
Figure 2025084820000005
Abstract
Description
Cross - reference to related applications
[0001]
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 153,352, filed on February 24, 2021, entitled "Protection of Databases, Data Transmissions and Files without Encryption" under 35 U.S.C. § 119(e), and incorporates by reference all of its disclosure herein.
Technical Field
[0002]
[0002] This disclosure generally relates to computer data security, and more particularly, to protecting databases and other types of computer files by using modular arithmetic to rearrange fields within a column or row without involving content encryption or individual data changes in the database. Background
[0003]
[0003] Protecting computer data from access by malicious actors such as hackers is of utmost importance. An increasing amount of the world's data is being stored online (e.g., in addition to databases of companies, governments, universities, and other organizations, accounts of companies and individuals on various cloud - based storage services). Despite data encryption and the use of other conventional security protocols of varying effectiveness (such as firewalls and two - factor authentication), hacking of data stored online often succeeds. Private information as well as information of governments, companies, and individuals, which is often highly confidential, is often illegally accessed, stolen, sold, and / or used for improper purposes such as espionage, ransomware, extortion, financial fraud, and other criminal acts.
[0004]
[0004] Data stored online is typically protected by encryption. Encryption is the process of encoding information. In encryption, the original data, called plaintext, is converted into an unreadable form called ciphertext. Encryption is usually based on using a pseudo-random encryption key to encrypt the plaintext into ciphertext. The ciphertext can be decrypted back into plaintext by using a decryption key given to an authorized person. There are many different techniques for encryption. Modern encryption typically uses public key (asymmetric) or symmetric key methods.
[0005]
[0005] Hackers use various techniques to decrypt ciphertext without having the decryption key. These techniques usually require a significant amount of computing power, but powerful computing resources are readily available and not prohibitively expensive, even at the level required to break modern commercial-grade encryption. Hackers and computer security experts are constantly engaged in a "cat and mouse" game. Hackers try to stay one step ahead of security experts, and security experts try to "improve technology" to catch up with or, preferably, get ahead of hackers.
[0006]
[0006] As a characteristic of encryption, when the decryption of ciphertext is successful, the result becomes immediately obvious. While the ciphertext is fragmented, the plaintext can be clearly read. When a hacker attempts to decrypt ciphertext, it becomes immediately clear when they succeed.
[0007]
[0007] It is considered desirable to address these problems. Overview
[0008]
[0008] Confidential content such as personal information or financial information in a database or other types of files is protected from unauthorized access by hiding the relationship between the confidential content and other related data. For example, in the case of a database, the relationship between the fields containing the protected information and their respective related database records is hidden. To that end, a permutation algorithm is applied to the cells of one or more specific fields of the database using modular arithmetic. This permutation changes the order of the cells of the (one or more) specific fields without changing the content of any individual cell. Since the permuted fields still contain all of the original cells in the permuted order, the relationship between the cells of the (one or more) permuted fields and the other information of the related records is hidden. Usually, a database stores fields in columns and records in rows, and in this case, the permutation of the cells of a specific field of the database is in the form of applying a permutation algorithm to the cells of the column containing that field. In one embodiment where the database stores fields in rows and records in columns, the permutation of the fields is in the form of permuting the corresponding rows.
[0009]
[0009] The permutation algorithm may be in the form of a bijective function using modular arithmetic, for example, by using modular addition and modular subtraction in either order. In different embodiments, another permutation algorithm using a single parameter or different numbers of multiple parameters may be applied. As the number of parameters used increases, the security level increases, but generally at the cost of using more computing resources. In different embodiments, different trade-offs between security and computing resource utilization may be applied as needed in different scenarios.
[0010] In the step of applying the sorting algorithm, one or more parameters used may be applied in the modular operation so that a specific cell may be located in the sorting field. It is understood that the specific cell to be located is associated with a specific record as a result of the sorting, but is not arranged within the specific record. To locate the cell, one or more parameters, the identification information of the specific record with which the specific cell in the sorting field is associated, and the identification information of the specific sorting field are applied in the modular operation, whereby the specific cell is located in the sorting field.
[0011]
[0011] By the inverse modular operation, a specific record with which a specific cell in the sorting field is associated but in which the specific cell is not arranged may be obtained. More specifically, through one or more parameters used in the step of applying the sorting algorithm, the identification information of the specific record in which the specific cell in the sorting field is arranged and the identification information of the specific sorting field are applied in the modular operation, whereby the specific record with which the specific cell in the sorting field is associated is located.
[0012] In some embodiments, by hiding the relationships between units of content of certain types of files other than databases, the content in the files is protected from unauthorized access. To that end, one or more segment sizes of certain types of files may be determined. And the file is split as a linear composition of segments of the (one or more) determined sizes. The linear composition may be processed as a series of rows and columns, and the segments may fill the cells of the columns of the rows. Segments that are cells of different columns may have different sizes. The cells of a particular column or particular row (or multiple columns or rows) may be arranged by applying a sorting algorithm using the modular operations described above to the cells of the particular column or particular row, and this sorting changes the order of the cells without changing the content of any individual cell. Since the (one or more) sorted columns or (one or more) sorted rows still contain all of the original cells in the sorted order, the relationship between the cells of the (one or more) sorted columns or (one or more) sorted rows and the other cells of the file is hidden.
[0013]
[0013] Even with the sorting function described in this specification, the protected database (or other types of files) does not change the content of any cell in the unprotected ("clear") database. For this reason, if a cell records the number of products purchased by a specific customer, someone's birthday, or a doctor's prognosis description, these values do not change even when protected. By themselves, birthdays are not data that can cause damage even if they fall into the hands of malicious people, unless the birthday is associated with its context (whose birthday it is). Since protection is achieved by changing the position of the content of this cell, it can no longer be associated with a name or other background information. Since moving the birthday to the "diagnosis" column does not gain anything, the position of the content of any given cell is changed within the column. Once protection is applied, only the birthday is included in the "birthday" column, only the operator's name is included in the "operator" column, and only the credit limit is included in the "credit limit" column. By changing the position within the column in this way, [a] the number of rows is maintained, [b] a specific position (usually different from the original position) is assigned to the cell, and [c] each destination position is unique (for example, it is not possible to place the content of {row 4, column 17} and the content of {row 27, column 17} both in {row 62, column 17}). These features are satisfied when the movement of cell data within a column is a sorting (changing the order of an ordered set). In fact, any method of moving the content of cells within a column that satisfies [a], [b], and [c] is a sorting.
[0014]
[0014] Systematic sorting of the content of cells within a column requires modular arithmetic, which means that the arithmetic is configured so that the content of the cells does not move to rows outside the set of rows in the unprotected database. In the case of sorting when the number of rows is 10,000, the content of any cell needs to move to at least the row numbered 0 (the top row) and at most the row numbered 9,999 (the bottom row). If the range of downward or upward movement of the cell content is random, the randomization is controlled so that the above [c] is satisfied.
[0015]
[0015] With this method, it is possible to obtain from the protected database directly the data necessary for a response to a query (e.g., "Who among Dr. X's patients has also visited Dr. Y?" or "What was the total frozen food sales last week from customers with zip code 90084?"). This is based on the fact that in the protected database, the zip code at {row 27, column 17} is not related to the frozen food sales at {row 27, column 43}, but rather to the frozen food sales at, for example, {row 196, column 43} (where 196 is obtained from 27 by modular arithmetic used for database protection).
[0016]
[0016] However, assume that the reason file protection is desired is that it is not a database for which the organization requires frequent and secure queries, but rather a certain type of file that is digitally stored for possible future use and contains highly confidential and proprietary information. Conceivable are any type of files that require digital storage while also needing to be securely protected, such as top-secret contracts, date files of images that can prove the owner owned them at a certain date, call records, etc. The rearrangement function described herein is available for protecting any type of file, making the protected file more secure than encryption.
[0017]
[0017] Any type of file, such as an image file, audio file, text file, etc., can be protected as described herein. To a computer device, a file may appear as a very long single line of hexadecimal characters in some cases. After splitting the line into segments (perhaps each 4 bytes (or another selected length)), these segments can be converted into a rectangle (perhaps 1,000 rows spanning each 2,400 columns), where each "cell" at {row R, column C} consists of, for example, 4 bytes. Here, when using modular arithmetic to rearrange the columns, a protected file is generated that is much more difficult to decrypt (or even difficult to recognize as, for example, a call record) than in the case of protection by encryption.
[0018]
[0018] The features and advantages described in this summary and the following detailed description are not all-inclusive. In particular, many additional features and advantages may become apparent to those skilled in the relevant art in light of the drawings, specification, and claims of this specification. Further, the expressions used in this specification are mainly selected for readability and teaching purposes and may not be selected for defining or limiting the subject matter of the present invention. It should be noted that for determining the subject matter of such a present invention, it is necessary to rely on the claims.
Brief Description of the Drawings
[0019]
Figure 1
[0019] FIG. 1 is a block diagram of an exemplary network architecture in which a rearrangement-based data protection system according to some embodiments may be implemented.
[0020]
Figure 2
[0020] FIG. 2 is a block diagram of the operation of a rearrangement-based data protection system in the context of database protection according to some embodiments.
[0021]
Figure 3
[0021] FIG. 3 is a flowchart showing the steps of the operation of a rearrangement-based data protection system in the context of file protection of a non-database format according to some embodiments.
[0022]
Figure 4
[0022] FIG. 4 is a block diagram of a computer system suitable for implementing a rearrangement-based data protection system according to some embodiments. Detailed Description
[0023]
[0023] The above drawings illustrate various embodiments but are for illustrative purposes only. Those skilled in the art will readily recognize that other embodiments of the structures and methods shown in this specification can be adopted without departing from the principles described herein.
[0024]
[0024] FIG. 1 is a high-level block diagram showing an exemplary network architecture 100 in which a rearrangement-based data protection system 101 can be implemented. As will be described in more detail below, the rearrangement-based data protection system 101 protects a database 113 (and computer files in other formats) without any individual data changes and can be used to protect data transmission by adding unrelated data without any required data changes. The individual data in the database 113 usually has no meaning by itself and is only useful in the context of its association with other associated data. For this reason, even if a hacker obtains individual cells of data from the database 113, no damage can be caused. For example, obtaining an individual credit card number without any background information is useless to a hacker. What is desirable to protect is the association between data points, such as the name of the person associated with the specific credit card number, the associated three-digit code (CVV), the person's postal code, the billing address, etc. As another example, the measurement results of a given medical test performed on a specific patient (e.g., a single cell in the database 113) or a group of patients (e.g., an entire column) are useless without understanding how the raw measurement data is associated with other data (e.g., the timing of the test, the practitioner, how the results are compared with the most recent measurement results of the same patient, etc.).
[0025]
[0025] Database 113 typically organizes data by storing data types (e.g., names, credit card numbers, addresses, etc.) in columns (which may also be referred to as fields of the database), and data associated with a given entity (e.g., the name, credit card number, address, etc. of a specific person) in rows (which may also be referred to as records). That is, a record contains all the fields of data for a given entry in the database. Thus, the fact that a given credit card number is associated with a given name, CVV, and postal code is only grasped by accessing the entire record of the given person. Hackers cannot also grasp the association between the credit card and the person by accessing individual cells (or the entire column) of the credit card column. Conventionally, records are stored as rows and fields as columns, but this is arbitrary, and it is also possible to store records as columns and fields as rows. Furthermore, it is possible to associate multiple tables with each other through common fields. For example, a table containing patient name, address, and credit card information and a table containing medical history data can be associated by a common "patient ID" field in both tables.
[0026]
[0026] The sorting-based data protection system 101 protects the relationship between fields within a record by rearranging the data in one or more fields (e.g., columns) of the database 113 or a file, hiding the relationship between different fields of the data of a given record, thereby preventing damage by hackers. As a result, for example, the measurement results are separated from the patient name, and the supplier name is separated from the supplier's bank account information.
[0027]
[0027] Sorting is the rearrangement of the order of elements in a set. Considering the individual cells of a database column as a set, the possible permutations of that set are all the cases where the column content can be ordered. For illustrative purposes, using a three-row column containing the positive integers 1, 2, and 3 as an (obvious) example, the set is {1, 2, 3}, and the possible permutations are (1, 2, 3), (1, 3, 2), (2, 1, 3), (2, 3, 1), (3, 1, 2), (3, 2, 1). The number of permutations of a set consisting of n elements is n factorial (n!). Therefore, the number of possible permutations of a set increases factorially as the size of the set grows. An exemplary set consisting of three elements has 3! (6) permutations. However, simply adding one element increases the number of permutations to 4! (24), adding two elements increases it to 5! (120), and so on.
[0028]
[0028] Even if the data in the cells of the database is sorted, the data does not change; only the order of the data changes. Therefore, instead of changing the data in the cells of the database 113, the sorting-based data protection system 101 sorts the data within a column (or, in a more typical configuration where fields are column-based and records are stored in rows, within a row in the case of a field-based database). The database 113 protected (secured) by the sorting-based data protection system 101 still appears to be like the database 113. In other words, the unencrypted column data still appears in columns, and the unencrypted row data still appears in rows. Further, even if an attempt is made to restore the association by restoring the data to its original unprotected order for a permutation of columns, perhaps randomly selected, a result that still appears to be a valid database 113 is generated. Therefore, even if an attempt to unprotect the database 113 sorted by the sorting-based data protection system 101 is successful, both the success and failure results appear to be a valid database 113, making it indistinguishable from failure. In contrast, if an attempt to decrypt the ciphertext is successful, this is obvious. This is because the resulting plaintext can be read by humans while the ciphertext cannot.
[0029]
[0029] In this way, the rearrangement-based data protection system 101 protects data by "hiding it in plain sight", leaving only the content of data points (e.g., cells in database 113 or configuration data table) and removing the associations between them by rearrangement. This is not only significantly safer than encrypting data, but also, in principle, more computationally efficient. This is because executing a single rearrangement operation requires less computation than executing commercial-grade encryption. However, since the success of decrypting the rearranged database 113 cannot be distinguished from failure, the data protected by the rearrangement-based data protection system 101 is not practically affected even if the computing power available to hackers increases significantly. As described above, to attempt rearrangements possible for a column consisting of n rows, it is necessary to execute n! permutations. For a database of just 25 rows (orders of magnitude smaller than a normal deployment in the real world), any one of more than 10 25 rearrangements could be the only correct one (for comparison, one trillion is 10 12 ). This already makes the hacker's task more difficult than for a 10-million-row database protected by the longest encryption key used by Amazon AWS. This level of security implies relative permanence compared to encryption-based methods. Given the increasing computing power available to hackers beyond skills and patience, many companies are currently spending huge amounts of money repeating "technology improvements" to stay one step ahead of hackers. As described herein, such repeated expensive efforts are unnecessary if the data is protected by the rearrangement-based data protection system 101.
[0030]
[0030] Referring to FIG. 1, the illustrated network architecture 100 includes a plurality of clients 103A, 103B, and 103N (which may also be collectively referred to as "clients 103"), as well as a plurality of servers 105A and 105N (which may also be collectively referred to as "servers 105"). In FIG. 1, a rearrangement-based data protection system 101 exists on server 105A, a database management system 111 and a corresponding database 113 exist on server 105N, and client agents 109 are shown as operating on each of clients 103A to 103N. It should be understood that this is merely an example. In various embodiments, it is also possible to instantiate various functions of the rearrangement-based data protection system 101 on servers 105 and clients 103, or to distribute them among a plurality of servers 105 and / or clients 103. Also, although the database management system 111 is shown as existing on a single server 105B, it should be understood that the database management system 111 and / or the database 113 can be distributed among a plurality of computer devices and / or storage devices. As will be discussed in more detail below in conjunction with FIG. 3, in some embodiments, the database management system 111 is not used, and instead (or in addition), the rearrangement-based data protection system 101 operates in cooperation with data in file formats other than the database management system 111.
[0031]
[0031] Client 103 can also be in the form of a computer device operated by a user of the database management system 111 (or a user who accesses other forms / types of data). The client agent 109 can be in the form of an application that includes endpoint-level functions for the use and / or interaction with the rearrangement-based data protection system 101 and / or the database system 101. In some embodiments, the client agent 109 is not used, and the functions of the rearrangement-based data protection system 101 and / or the database system 101 are accessed by other means (e.g., via a browser (not shown)).
[0032]
[0032] Client 103 and server 105 can be realized using a computer system 610 as shown in FIG. 4 and described hereinafter. Client 103 and server 105 are communicatively coupled to network 107 via a network interface 248 as shown in FIG. 4 and described hereinafter, for example. Client 103 can access applications and / or data on server 105 using other client software such as, for example, a web browser or client agent 109. Client 103 can be in the form of other types of computers / computer devices including mobile computer devices having a laptop, desktop, and / or portable computer system capable of connecting to network 107 to execute applications (e.g., smartphones, tablet computers, wearable computer devices, etc.). Server 105 can be in the form of, for example, a rack-mounted computer device disposed in a data center.
[0033] [
[0033] ]FIG. 1 shows three clients 103 and two servers 105 as an example, but in reality, it is also possible to increase (or decrease) the number of clients 103 and / or servers 105 to be deployed. In one embodiment, the network 107 is in the form of the Internet. In other embodiments, other networks 107 or network-based environments can be used.
[0034] [
[0034] ]FIG. 2 shows the operation of a rearrangement-based data protection system 101 that operates on the server 105 and communicates with a plurality of client agents 109 according to some embodiments. As described above, the functions of the rearrangement-based data protection system 101 can also exist on the server 105 or other specific computers 610, or can be distributed among a plurality of computer systems 610. Examples of cloud-based computer environments in which the functions of the rearrangement-based data protection system 101 are provided as cloud-based services on the network 107. Although the rearrangement-based data protection system 101 is shown as a single entity in FIG. 2, it should be understood that the rearrangement-based data protection system 101 represents a set of functions that can be instantiated as a single module or multiple modules as needed. In some embodiments, different modules of the rearrangement-based data protection system 101 can exist on different computer devices 610 as needed. Each client agent 109 can be instantiated as an application configured to operate under an operating system such as Windows®, OS X, Linux®, etc. or an app for a given mobile operating system (e.g., Android®, iOS, Windows 10, etc.), and different client agents 109 are clearly implemented for different types of operating environments used by different end users.
[0035]
[0035] It is to be understood that the components and modules of the rearrangement-based data protection system 101 can be instantiated (e.g., as object code or executable image) within the system memory 617 (e.g., RAM, ROM, flash memory) of any computer system 610 such that when the processor 614 of the computer system 610 processes the modules, the computer system 610 performs the associated functions. As used herein, the terms "computer system", "computer", "client", "client computer", "server", "server computer", and "computing device" mean one or more computers configured and / or programmed to perform the functions described above. Also, the program code for implementing the functions of the rearrangement-based data protection system 101 can be stored on a computer-readable storage medium. In this context, any form of tangible computer-readable storage medium such as magnetic, optical, flash, and / or semiconductor storage media, or any other type of media, etc., can be used. As used herein, the term "computer-readable storage medium" does not mean an electrical signal separate from the underlying physical medium.
[0036]
[0036] As shown in FIG. 2, in one embodiment, this “hide as it appears” is achieved by rearranging the cells within a given column of the database 113. As will be described in more detail below in conjunction with FIG. 3, in other embodiments, similar functionality is utilized in the context of files of other formats. In the instantiation of the database 113 of FIG. 2, the data in the cells of the rearranged column is still of the same type as before the rearrangement and represents the same type of object (e.g., first visit date, social security number, etc.), but the order is different and no longer matches that of the other columns (i.e., non-rearranged columns (referred to herein as “clear”)). Thus, when a given column is rearranged, the cells of that column are no longer associated with the other columns in that given row, which represents different data points for a particular entry in the database 113, such as a given patient, client, customer, supplier, cardholder, member, employee, etc. For example, if the records regarding a hospital patient are stored in the rows of the database 113 and the column of the database 113 where the patient's social security number is stored is rearranged, the social security number of a given patient will no longer be placed in the given row of the database 113 associated with that given patient, making it impossible for a malicious person who has obtained access to the database 113 to identify it.
[0037]
[0037] It is to be understood that the rearrangement-based data protection system 101 may further hide the relationships between units within the database 113 by rearranging a plurality of columns. As described above, these rearranged columns, except that they still do not contain any proprietary information or highly confidential classified information by themselves because the relationship with a specific patient cannot be identified, contain the same type of content (e.g., a patient's social security number or specific medical test results) as the non-protected database 113. CLEAR and the same type of content (e.g., a patient's social security number or specific medical test results).
[0038]
[0038] To rearrange the cells of a column, the rearrangement-based data protection system 101 applies a rearrangement algorithm that uses modular arithmetic. As will be described in more detail below, in different embodiments, the rearrangement-based data protection system 101 may apply different rearrangement algorithms based on modular arithmetic with different numbers and / or orders of parameters to the column(s) to be protected. More specifically, the rearrangement-based data protection system 101 may apply a bijective function that uses modular arithmetic, including, for example, modular subtraction and modular addition in either order. A bijective function is a mathematical function between the elements of two sets, where each element of the first set is paired with exactly one element of the second set, and vice versa, with no unpaired elements. In the case of the rearrangement-based data protection system 101, a bijective function from a set to itself is applied, which is a rearrangement. In this context, what a bijective function from a set to itself means is a rearrangement from the set of elements of the clear column to the same set of elements of the rearranged column. Modular arithmetic is a system of integer arithmetic where, when a given value called the modulus is reached, the numerical value "wraps around" leaving a remainder. A typical example of this is a 12-hour clock, where 12 is the modulus. When the time exceeds 12, it wraps around and the remainder becomes the new time. For example, in a 12-hour clock, 9 + 5 = 2.
[0039]
[0039] As described above, the specific number of parameters used in the mathematical bijective function applied to the rearrangement of the cells of a column is a design choice that can vary. As the number of parameters used increases, the level of security increases, but so do the required processing resources. In different scenarios, different people may select different balances between these factors as needed.
[0040]
[0040] In the rearrangement of the cells of one or more columns of the database 113 (or other form of data table), the clear database 113 to be rearranged CLEARand clearing columns to create a protected database 113 having one or more columns in which the cells have been rearranged. PROTECTED It is understood that the Clear Database 113 CLEAR The content of the database is 113 CLEAR whereas the protected database 113 PROTECTED The contents of the sorted columns of the database 113 are hidden because the relationships between the cells of the sorted columns and the associated database records are hidden, as described herein. PROTECTED Thus, the conversion of databases into secure content without encryption and with the reordering-based data protection system 101 is a major improvement in the field of computer security.
[0041]
[0041] As an example, we will discuss an instantiation where the owner of a confidential file cares so much about security that performance cost is not an issue. It is understood that this is a theoretical scenario and that in practical examples there will likely be a trade-off between security and processing cost. In any case, in this hypothetical example, a parameter-heavy procedure may be chosen for the desired column permutation of a database 113 consisting of R rows and C columns, as follows: The clear database 113 is considered to be a rectangular array U (unprotected array) with R rows and C columns. In this example, we assume that column 0 is not to be permuted (e.g., the unpermuted data of column 0 does not pose a security risk in itself). The other columns 1 to C-1 are to be permuted. Thus, a protected array P is generated by the following steps: Generate three two-dimensional arrays S, FC, and FR, each with R+1 rows and C-1 columns, and initially fill FC and FR with all 0s, and work your way down the rows filling each column of S with integers 0, 1, 2, . . . , R (e.g., row 3 is made up of all 3s). Here, we protect column c=1. We remove the leftmost column from S to obtain a one-dimensional array T cLet it be {0, 1, 2, ···, R}. First, in step cr where r = 0, at column c = 1, determine the location to move the content of cell r = 0. At this time, a random number X rc is taken from the set Z = {0, 1, ···, R - r}. T c Remove X rc from it, reducing the number of elements in the array by one (for example, when X rc = 4, T c becomes {0, 1, 2, 3, 5, 6, ···, R}). Here, set FC 0c equal to the element X rc th in the set Z, set FR's X rc = 0, and put the content of U rc into the element X rc of P. As we go down the column like this, each cell of U enters a randomly selected cell (from the still-empty cells in the same column) of P. Furthermore, for each column on the right side of P, it is determined in the same way by a new randomization.
[0042]
[0042] Although randomization is used in parameter selection in the exemplary instantiation described in this specification, it should be understood that the randomness used is only accidental selection. For example, in a hypothetical example where the parameters are not random but intentionally selected, the selected parameters may not use all or almost any numbers ending with 5 or 0 (for example, due to a certain preference to avoid them). In this case, if the rarity of parameters ending with 5 or 0 is somehow inferred by hackers, the factorial problem becomes a little smaller. In the case of 1 million lines, if 0 and 5 are not used at all, the number of permutations will decrease from 56.5 million digits to 2 million digits, and if it is disproportionately less than 1 - 4 and 6 - 9 of those used, it will decrease to 5.2 million digits. In other words, even without random parameter selection, the level of protection remains very high.
[0043]
[0043] By using two subroutines provided by the rearrangement-based data protection system 101, it is possible to obtain specific desired content by querying the protected database 113 PROTECTED The query to may obtain specific desired content. In this specification, these two subroutines are referred to as FindCell201 and FindRecord203. In the above rearrangement example, the array FC provides R*(C - 1) parameters of FindCell201, and the array FR of the same dimension provides R*(C - 1) parameters of FindRecord203. The operations of FindCell201 and FindRecord203 will be discussed in more detail below.
[0044]
[0044] Hereinafter, the application of different exemplary algorithms by the rearrangement-based data protection system 101 will be described. However, since this is close to other theoretical limits and the performance cost is relatively high, the performance cost is minimized by applying a one-parameter algorithm, sacrificing a decrease in the robustness of the provided security. In a specific example of an embodiment where the rearrangement-based data protection system 101 applies a one-parameter algorithm, similarly, among N rows, a random integer A between, for example, 11 and N - 12 is taken out. Here, using modular arithmetic, the data point at row r and column c is placed at row r' = mod(r + c*A, N) while keeping column c. If the data point at row r and column c satisfies the query and the corresponding data point in row 0 is required, the said data point is in the row r = mod(r + N - c*A, N) (the inverse function of the preceding function). In this specification, * indicates multiplication.
[0045]
[0045] In the application of this exemplary algorithm, the protection data table is satisfied by a set of projection fields, which are mathematical terms. By definition, two sets A and B that are a partition of a given set S form a projection field if any element of A crosses all elements of B and any element of B crosses all elements of A. Thus, the rows and columns are projection fields. Here, A consists of the elements defined respectively by the cells of the leftmost column and the cells of each other column obtained by applying the formula mod(r + c*A, N), and B can consist of the set of columns of the data table. In this context, projection fields can be used in other instantiations of the permutation algorithm with other numbers of parameters as needed.
[0046]
[0046] Here, an example of a permutation-based data protection system 101 applying a two-parameter algorithm is provided with parameters (x, y). Here, to protect any column c of the unprotected database 113, the following steps are performed. If c is odd, the data points in row r are placed in row r = mod(r + x, N) and if c is even, the data points in row r are placed in row
Number
[0047]
[0047] In other examples, other numbers of parameters can be used as needed. In the example of four parameters for column C, a random integer D is taken from [0.25C, 0.6C]. After taking D, a random integer D is taken from [D + 1, 0.9C]. For the column to the left of D, a random integer A is used as in the one-parameter example, and for the columns to the right of D and D a separately randomized B using the same formula is used for the column to the left of. D And for the column on its right side, use G separately randomized with the same formula.
[0048]
[0048] In the example of the three parameters in column C, a random integer D is taken from [0.3C, 0.8C]. Similar to the example of the two parameters, for the column on the left side of D, when c is odd, the data point of row r is placed in row r =mod(r + x, N) and when c is even, the data point of row r is placed in row
Number
[0049]
[0049] As described above, the specific number of parameters used in the bijective function used for the rearrangement of specific columns of the database 113 is a variable design choice. As the number of parameters increases, the security improves, but more processing power will be used. Furthermore, it should be understood that the specific modular operations to be applied and the order of their application may vary for each embodiment. Also, which columns and how many columns to rearrange in the database 113 are variable and are based on the fields of the database 113 that contain confidential information to be separated from the related records.
[0050]
[0050] Therefore, as described above, the rearrangement is controlled by the application of an algorithm that implements one of a group of mathematical functions. This can be considered appropriately expressed as a random rearrangement in two senses. That is, (a) A malicious person who has obtained the protected database 113 but cannot obtain the algorithm or the function parameters observes the rearrangement of all or most columns of the data table that cannot be distinguished as random. (b) In practice, it is (not necessary but) possible to randomly select the parameters (of the specific function used for cell detachment) themselves, resulting in a random permutation. When the database 113 PERMUTED is rearranged, as described below, since the desired data can be queried and accessed from it, it is understood that there is no need to retain the non-rearranged database 113 CLEAR . However, if necessary, for example, in a place where hackers cannot access it offline, it is possible to maintain a copy of the non-rearranged database 113 CLEAR .
[0051]
[0051] Hereinafter, the acquisition of desired content from the rearranged database 113 according to several embodiments will be described. By using the functions described in this specification, it is possible to process a query to the rearranged database and obtain the requested data without unlocking the protection of the database 113 PERMUTED or unlocking the protection of any specific column or each specific column that may be related to satisfying the query. When the database 113 PERMUTED is protected, the data itself is not changed and is only rearranged, so the data still exists in the protected database 113 PERMUTED and is available for query processing. The data is not in the same location as the unprotected database 113 CLEAR . Therefore, in contrast to the protected state, the rearranged database 113 PERMUTEDThe performance cost associated with locating and retrieving specific data in it is minimal. However, the performance overhead is minimal compared to decryption. More specifically, to read the data transmitted as a response to a query, the correct rows of the queried columns are found by using a preliminary step of modular arithmetic. The sum of such modular arithmetic steps for all aspects of processing the query results in a performance cost that is significantly lower than the decryption of the encrypted database 113 ENCRYPTED whether it is done for the entire database or separately for each field (i.e., when "structure-preserving encryption" is used), the encrypted database 113 ENCRYPTED has a much lower performance cost than decryption. Note that queries on the sorted database 113 PERMUTED do not involve reverting the entire sorted database or any of its sorted columns.
[0052] After data responsive to a query is found, the transmission to the query computer (e.g., client 103 interacting with database management system 111) may be protected by essentially the same logic as described above. For example, for a query of specific data for a specific provider, assume that one record containing eight fields is obtained. And, to protect data transmission, a ratio of unrelated data to actual data may be employed. The ratio used is a design feature that can vary and may be selected, for example, by the database license, the database administrator, the issuer of the sorting-based data protection system 101, etc. In an exemplary scenario where this ratio is 4:1, four records of unrelated data (or completely fictional data) are collected, and each of these four records contains the same number and type of fields as the actual record (e.g., in this example, eight fields per record). Here, the same sorting algorithm using time-dependent parameters may be applied to the resulting grid of five records, each consisting of eight fields. Thus, a single record containing the response to the query is sorted with the four unrelated records. After sorting, this five-record grid is sent to the query computer. Since the actual result of the query is sorted with the other four records, the transmission is protected. And the query computer can use the same parameters to reverse the modular operation to unprotect the five records and select one requested record, which may then be displayed on the requesting device as needed or may be used as needed.
[0053]
[0053] In some embodiments, typically, for the protection or unprotection of any kind of file that is in a rectangular structure (e.g., the data table of database 113) or is adapted to a rectangular structure (e.g., by treating each byte of the file as a single piece of data in database 113 as will be described later in conjunction with FIG. 3), it can be realized by repeatedly using two basic subroutines. The discovery, modification, addition, and removal of data may all rely on the same two subroutines. The subroutines, as described above, provide results determined by various parameters used in the corresponding application sorting algorithm, which are typically specific to the situation. In this specification, these two subroutines are referred to as FindCell201 and FindRecord203, but this naming is not important, and it is understood that in other embodiments, the subroutines may have different names (and / or more subroutines or different subroutines may be used).
[0054]
[0054] The subroutine FindCell201 locates a specific cell associated with a specific record in a sorted field. Since the field is sorted, the specific cell is not located within the associated specific record. In one embodiment where the field is stored in columns and the records are stored in rows, FindCell201 locates the specific cell in the sorted column associated with a given row, but it is not located in that row due to the sorting. More specifically, FindCell201 performs a modular operation by applying one or more parameters used in the application of the sorting algorithm, the identifier of the specific record in which the specific cell of the sorted field is located, and the identifier of the specific sorted field, to locate the specific cell in the sorted field.
[0055]
[0055] A rectangular structure (e.g., a database or other file type) will (usually) have a "main field" (although not necessarily). In some embodiments, this will be the leftmost column (in the instantiation where fields are stored in columns). The subroutine FindCell201 takes as input an array of (one or more) parameters used for sorting the database 113 (which will have different configurations in different embodiments), the row number of the main column (or other form of identifier of the record to which a particular cell of the sorting field is associated), and the desired column expressed, for example, as a column number or column header (or other form of identifier of the desired field). The subroutine FindCell201 applies the (one or more) parameters and performs one modular operation to output the corresponding row of the desired column. For example, assume that the shipment quantities to various customers last month are in the "Qm-1" column, and this data is requested for customer 1234. Here, 1234 appears in row 567 of the main column. Then, the subroutine FindCell201 takes the (one or more) relevant parameters (in the example of one parameter, let it be 8910), adds 567 + 8910 = 9477 and outputs 9477, which is the desired row of the "Qm-1" column where the shipment for this customer last month is found. For example, if the number of rows in the rectangular structure is only 7000 rows, the desired row will be 9477 - 7000 = 2477, which is automatically obtained by adding modulo 7000.
[0056]
[0056] Subroutine FindRecord203 locates the specific record associated with a specific cell in the sorted field. Since the field is sorted, the specific cell is not located within the associated specific record. In one embodiment where the field is stored in a column and the records are stored in rows, FindRecord203 locates the row associated with the specific cell in the sorted column. The specific cell is associated with the specific row but is not located there due to the sorting. More specifically, FindRecord203 performs modular arithmetic to apply one or more parameters used in the sorting algorithm, the identification information of the specific record where the specific cell of the sorted field is located but not associated, and the identification information of the specific sorted field, to locate the specific record associated with the sorted field.
[0057]
[0057] Assume, as the requested data, the customer numbers and the distances from the shipping factory locations of all customers who shipped 100 to 200 units last month. Further assume that the content of the cell in row 9477 of column "Qm-1" falls within this range. In this type of scenario, subroutine FindRecord203 is used, which is essentially the inverse function of subroutine FindCell201. FindRecord203 takes as input the same array of (one or more) parameters, column "Qm-1", and row 9477 where the shipped quantity that matches the range was found. In an example of one parameter, FindRecord203 takes the same parameter and outputs 567 by subtracting 9477 - 8910, which is the row of the main column where customer number 1234 of this customer was found. In the case of this example, since the shipped distance is not considered confidential information unless the customer is identified, the data is in row 9477 of column "Distance", indicated by the parameter of that column matching the parameter of column "Qm-1".
[0058]
[0058] In almost all databases 113, the most frequently used operation is information query. By using the sorting-based data protection system 101 as described above, it is possible to respond to queries without releasing the protection of the database 113. On the other hand, in a conventional encryption-based system, it is necessary to decrypt the encrypted database 113 for querying. Since there is no need to decrypt the encrypted database 113 into an unprotected state where it can be queried, the performance cost is significantly reduced. Although it is less frequent than querying the database 113 in general, the update of the content of existing records (i.e., to update one or more specific fields of a given record to indicate information such as whether a given person's subscription has been updated, whether a given customer has made a payment of a specific amount, etc.) is still supported regularly.
[0059]
[0059] Although relatively rare, another supported database operation is the addition of new records, which is executed, for example, when a new customer opens an account at a bank. Deleting existing records is also another operation that is relatively rarely supported, and is executed, for example, when a patient decides to request medical treatment from a different doctor outside a given facility, provides information to the new doctor, and then requests deletion. For example, when a new product is launched on the market, when a new franchisee is added, etc., it is also possible to edit the existing schema of the database 113 by adding one or more new fields. In some embodiments, these additions and deletions to the existing database 113 in the context of the sorting-based data protection system 101 are performed when the database 113 is sorted for the first time.
[0060] More specifically, before sorting, a plurality of additional fields can be added to the database 113. One is for system use as described later, and the others are for enabling new fields to be added later without having to resort the database 113 and then re-sort it. Any data can be initially stored in the introduced fields for future use. In one embodiment, by storing data that replicates or mimics existing fields, it can be made more difficult for unauthorized persons to interpret. The specific number of fields added for future use is a design choice that can vary. Also, it is possible to add many unused records to the database before the first sorting. This is to enable new records to be added later without having to resort and re-sort the database 113 each time a record is added. Similar to the unused fields, the number of unused records to be added and the content stored therein are design choices that can vary, but for security enhancement, existing records may be replicated in the unused records. The new fields for the aforementioned system use may be used to encode whether a record contains actual data or is just a placeholder for adding a new record. Placeholder records are sorted together with the actual data records of the database 113, but all of them contain placeholder fields.
[0061] When a new record is added, the actual data of the record is replaced with the dummy placeholder data of the placeholder record, and the coded cell indicating the dummy data is switched to the coding indicating the actual data. In the case of removing an actual record, the reverse operation is performed (i.e., the actual data is replaced with the dummy data and the coded field is updated accordingly). Similar to other records, the placeholder record where replacement with a new record is found by an additional operation, or the actual record where replacement with dummy data is found by a deletion operation, is sorted, and since it is not stored straight in the entire row of the database 113, it is located using modular arithmetic as described above.
[0062]
[0062] As shown in FIG. 3, in some embodiments, the rearrangement-based data protection system 101 may be adapted to protect files in formats other than the database 113. In such embodiments, any file of any type, including proprietary or confidential information, can be protected by applying the corresponding functions described herein. This function can be used to protect document files, text files, or any other file containing alphanumeric content (e.g., words, abbreviations, whitespace, numbers, etc.), image files, audio files, animations, movies, or video files in other forms, any one of a variety of file types with wide or narrow usage (e.g.,.docx,.txt,.html,.xls,.pdf,.gif,.jpg,.mp3,.alac,.wav,.flac,.mp4,.mov,.avi, etc.). Also, by applying the functions described herein, any file that can be represented as a header followed by a series of binary and / or hexadecimal numbers (essentially, any file type) can be protected. Note that, unlike the above exemplary embodiments that protect the database 113, for protected files in other formats, when decryption is successful, the hacker can be visually identifiable. This is because all attempts to reverse the data are fragmented, while successful attempts using the correct parameters result in content that appears as a specific type of file. Nevertheless, despite such specific advantages not being present, the resulting protection is still very strong compared to encryption, and the computational overhead is still low.
[0063]
[0063] To rearrange any type of file, the rearrangement-based data protection system 101 may be configured to determine the “cell size” of a particular file of a given type (301). As the cell size, for example, one character in a document, one character (phonetic character), a number, a punctuation mark, a space, a paragraph end indicator, one hexadecimal character, one byte, a group consisting of consecutive n bytes, one grid of n×n pixels, a voice or animation equivalent to 0.2 seconds, or other size specifications are possible. The rearrangement-based data protection system 101 may split the file as a linear configuration of cells of the size determined in step 301 respectively (303). Thereafter, the rearrangement-based data protection system 101 arranges the linear configuration in a rectangular shape, first filling the top row and then the next row, repeating this for all data in the file (305). Thereafter, the rearrangement-based data protection system 101 may, if necessary, apply a specific mathematical bijective function to rearrange the cells within a given column for all columns, almost all columns, or a given subset of columns (307) to remove the association between adjacent cells within a row and protect this structure. Although a file that is not a database has been described as being organized into column-based fields and row-based records, it is understood that in other embodiments, row-based fields and column-based records may be used. In any case, the subroutines FindCell201 and FindRecord203 are usable in the context of file types that are not databases, as described above in the context of database embodiments.
[0064]
[0064] Here, as a specific example of improving security by protecting the most frequent types of data transmission using the rearrangement-based data protection system 101, a digital transaction in which a credit card number is transmitted for authentication purposes will be described. More specifically, when a purchaser conducts a credit card transaction, the credit card number, other identification information regarding the cardholder (e.g., the purchaser's name, three-digit card verification value (CVV), expiration date), and information regarding the transaction are transmitted to the card-issuing company or a third-party authentication service, and these approve or reject the transaction. Conventionally, the credit card number and CCV in such a transmission have only been hashed, which is not very secure. This is because by accessing multiple hashed credit card numbers, the hash algorithm being used can be identified and the actual information can be accessed (after all, a credit card has only 16 digits in decimal and a CVV has only 3 digits).
[0065]
[0065] The sorting-based data protection system 101 system described in this specification can achieve much higher security by treating the information sent for authentication as records, adding records irrelevant to the data, and sorting the actual records of interest with the irrelevant records. As a simple example, assume that the format of the data sent for credit card authentication includes a name, a dollar-cent amount, a 16-digit card number, a CVV code, and perhaps the expiration date (month / year) of the card (the actual format may vary). Using a ratio of 4:1 as an example, for each actual record containing information about the transaction to be authenticated, four irrelevant records of the same format are added, and each field of the total five records can be sorted. Since the data sent for credit card authentication is not very much even with the addition of irrelevant records, the additional performance cost is very small. However, it is much more difficult for a hacker or other malicious person attempting to intercept the transaction to bypass the protection of the data and access the purchaser's credit card information. And, as described above, the authentication service can rearrange the received data back.
[0066]
[0066] The method described in this specification for protecting both the database 113 and other types of files achieves a very robust level of protection against infringement. For a hacker, the difficulty of bypassing the protection of the sorted database or file as described in this specification far exceeds that of encryption even when the encryption key is quite long. The hacking task is proportional to the length of the encryption key. Larger files provide more clues for narrowing down the possible encryption keys. The bypass task of the functions described in this specification is the factorial of the number of rows in the rectangular representation of any file type. When the number of rows is as few as 25, for a single column, any one of more than 10 25 permutations can be the only correct permutation (for comparison, one trillion is 10 12) This already makes the hacker's task more difficult than a 10-million-line database 113 protected by the longest encryption key used by Amazon AWS. At 50 lines (which is still a very small file), there are over 10 64 permutations, and at 100 lines, there are over 10 157 permutations. Applying this protection method to a medium-sized database 113, a large high-resolution image, or a video clip of a few minutes easily reaches 1 million lines, for which there are over 8.26*10 5,565,708 permutations (which is approximately {the number of hydrogen atoms estimated to exist in the universe} 50,000 and seems to be the largest integer calculated for purposes other than abstract number theory). Since this level of security is unattainable even with encryption, applying the permutation-based function described herein to secure the content stored in the database 113 and / or files in other formats against unauthorized access by malicious actors is not only a major improvement in the field of computer security but also an improvement in the operation of server farms, data centers, and generally secure storage technologies.
[0067]
[0067] FIG. 4 is a block diagram of an exemplary computer system 610 suitable for implementing the sorting-based data protection system 101. Both the client 103 and the server 105 can be implemented in the form of such a computer system 610. As shown in the figure, a bus 612 is included as a component of the computer system 610. The bus 612 communicatively couples at least one processor 614, a system memory 617 (e.g., random access memory (RAM), read only memory (ROM), flash memory), an input / output (I / O) controller 618, an audio output interface 622 communicatively coupled to an audio output device such as a speaker 620, a display adapter 626 communicatively coupled to a video output device such as a display screen 624, a universal serial bus (USB) receptacle 628, a serial port 630, one or more interfaces such as a parallel port (not shown), a keyboard controller 633 communicatively coupled to a keyboard 632, a storage interface 634 communicatively coupled to one or more hard disks 644 (or (one or more) other forms of storage media), a host bus adapter (HBA) interface card 635A configured to be connected to a fiber channel (FC) network 690, an HBA interface card 635B configured to be connected to a SCSI bus 639, an optical disk drive 640 configured to receive an optical disk 642, a mouse 646 (or other pointing device) coupled to the bus 612 via, for example, the USB receptacle 628, a modem 647 coupled to the bus 612 via, for example, the serial port 630, and one or more other components of the computer system 610 such as one or more wired and / or wireless network interfaces 648 directly coupled to, for example, the bus 612.
[0068]
[0068] Other components (not shown) (e.g., a document scanner, a digital camera, a printer, etc.) may be similarly connected. Conversely, not all of the components shown in FIG. 4 need to be present (e.g., smartphones and tablets typically do not have an optical disk drive 640, an external keyboard 632, or an external pointing device 646 either, but various external components can be coupled to a mobile computing device via, for example, a USB receptacle 628). Also, the various components can be interconnected in a manner different from that shown in FIG. 4.
[0069]
[0069] The bus 612 enables data communication between the processor 614 and the system memory 617 which may include RAM in addition to ROM and / or flash memory as described above. The RAM is typically the main memory into which the operating system 650 and application programs are loaded. The ROM and / or flash memory may include, among other codes, a BIOS (Basic Input-Output System) that controls certain basic hardware operations. The application programs can be stored on a local computer-readable medium (e.g., a hard disk 644, an optical disk 642) and loaded into the system memory 617 for execution by the processor 614. Also, the application programs can be loaded into the system memory 617 from a remote location (i.e., a remotely located computer system 610) via, for example, a network interface 648 or a modem 647. In FIG. 4, the rearrangement-based data protection system 101 is shown as being present in the system memory 617, but in some embodiments, a part of the system 101 may be located elsewhere, e.g., on the hard disk 644 or other storage mechanisms.
[0070]
[0070] The storage interface 634 is coupled to one or more hard disks 644 (and / or other standard storage media). The (one or more) hard disks 644 may be part of the computer system 610 or may be physically separated and adapted to be accessed through other interface systems.
[0071]
[0071] The network interface 648 and / or the modem 647 can be communicatively coupled directly or indirectly to a network 107 such as the Internet. Such a coupling can be wired or wireless.
[0072]
[0072] As will be understood by those skilled in the art, the subject matter described herein may be embodied in other specific forms without departing from its spirit or overall characteristics. Similarly, the particular naming and partitioning of parts, modules, agents, managers, components, functions, procedures, operations, layers, features, attributes, methods, data structures, and other aspects are not essential or important, and the entities used in the implementation of the subject matter described herein may have different names, partitioning, and / or formats. The above description has been presented with reference to specific embodiments. However, the above exemplary discussion is neither exhaustive nor intended to limit the disclosure to the detailed forms disclosed. Many modifications and variations are possible in light of the above teachings. The above embodiments have been selected and described in order to best enable others skilled in the art to make the best use of various embodiments, regardless of the presence or absence of various improvements suitable for a particular contemplated use, by best explaining the relevant principles and their respective practical applications.
[0073]
[0073] In some cases, various embodiments may be presented herein with respect to algorithms and symbolic representations of operations on data bytes in a computer memory. An algorithm is generally considered here to be a set of self-consistent operations leading to a desired result. Operations require physical manipulation of physical quantities. Typically, these quantities are, although not necessarily, in the form of electrical or magnetic signals capable of being stored, transformed, combined, compared, or otherwise manipulated. It has been found that these signals are sometimes conveniently represented, mainly for reasons of common usage, in terms of bits, bytes, values, elements, symbols, characters, terms, numbers, etc.
[0074]
[0074] However, it should be borne in mind that all of these terms and similar terms are merely convenient labels associated with appropriate physical quantities and applied to these quantities. Unless otherwise specifically stated as will be apparent from the following discussion, of course, throughout the present disclosure, discussions using terms including "processing", "computing", "calculating", "configuring", "determining", "displaying", etc. represent the operations and processes of a computer system or similar electronic device that manipulates data represented as physical (electronic) quantities in the registers and memory of the computer system and converts it into other data similarly represented as physical quantities in the memory or registers of the computer system or other such information storage, transmission, or display device.
[0075]
[0075] Finally, the structures, algorithms, and / or interfaces presented in this specification are essentially not related to any specific computer or other device. Along with the programs according to the teachings of this specification, various general-purpose systems may be used, or it may be found convenient to configure more special devices for executing method blocks. The structures for these various systems will be apparent from the above description. Also, the description in this specification does not refer to any specific programming language. Naturally, various programming languages can be used to implement the teachings as described in this specification.
[0076]
[0076] From the above, this disclosure is merely illustrative and is not intended to be limiting in any way. [Items of the Invention] [Item 1] A method for protecting confidential content of a database from unauthorized access by hiding the relationship between the confidential content and other fields of a database record related to the confidential content, comprising: rearranging cells of a specific field of the database by applying a rearrangement algorithm using modular arithmetic to cells of the specific field of the database, the rearrangement changing the order of the cells of the specific field without changing the content of any individual cell; hiding the relationship between the cells of the rearranged field and a database record related to the cells of the rearranged field by the rearranged field still containing all of the original cells in the rearranged order. [Item 2] the database stores fields in columns and records in rows, the method according to item 1, further comprising applying the rearrangement algorithm using modular arithmetic to cells of a specific column of the database, the specific column containing cells of the specific field. [Item 3] The database stores fields in rows and records in columns, The step of sorting cells of the specific field of the database is to apply the sorting algorithm using modular arithmetic to the cells of a specific row of the database, the specific row including the cells of the specific field, further including the step of applying, the method according to item 1. [Item 4] The step of sorting the cells of each of the plurality of specific fields of the database by applying the sorting algorithm using modular arithmetic, the sorting changing the order of the cells of each specific field without changing the content of any individual cell, further including the step of, By each sorted field still including all of the original cells in the sorted order, hiding the relationship between the cells of each sorted field and the database records associated therewith, the method according to item 1. [Item 5] The step of applying a sorting algorithm using modular arithmetic is Further including sorting the cells of the specific field of the database by applying a bijective function using modular arithmetic, the method according to item 1. [Item 6] The step of applying a sorting algorithm using modular arithmetic is Further including applying a sorting algorithm using modular addition and modular subtraction in either order, the method according to item 1. [Item 7] The step of applying a sorting algorithm using modular arithmetic is Further including applying a sorting algorithm using modular arithmetic with a plurality of parameters, the method according to item 1. [Item 8] The step of applying a sorting algorithm using modular arithmetic is The method according to item 1, further comprising applying a sorting algorithm using modular arithmetic with a single parameter. [Item 9] The step of applying a sorting algorithm using modular arithmetic comprises The method according to item 1, further comprising applying a sorting algorithm using modular arithmetic that satisfies the cells of the field sorted through a projection field. [Item 10] The step of applying a sorting algorithm using modular arithmetic comprises The method according to item 1, further comprising applying a sorting algorithm using modular arithmetic with at least one pseudo-random selection parameter. [Item 11] The step of applying a sorting algorithm using modular arithmetic comprises The method according to item 1, further comprising applying a sorting algorithm using modular arithmetic with at least one intentional selection parameter. [Item 12] The method according to item 1, further comprising locating, in the sorted field, a specific cell that is associated with a specific record but not arranged within the specific record, by applying one or more parameters used in the step of applying the sorting algorithm in a modular arithmetic. [Item 13] The step of locating a specific cell in the sorted field by applying one or more parameters used in the step of applying the sorting algorithm in a modular arithmetic comprises The method according to item 12, further comprising locating the specific cell in the sorted field by applying, in the modular arithmetic, the one or more parameters, the identification information of the specific record with which the specific cell of the sorted field is associated, and the identification information of the specific sorted field. [Item 14] A step of identifying a specific record associated with a specific cell in a sorted field by applying one or more parameters used in the step of applying the sorting algorithm in a modular operation, further including a step where the specific cell is not arranged in the specific record with which it is associated, the method according to item 1. [Item 15] A step of identifying a specific record associated with a specific cell in a sorted field by applying one or more parameters used in the step of applying the sorting algorithm in a modular operation further includes applying, in the modular operation, the one or more parameters, identification information of the specific record that is where the specific cell of the sorted field is arranged but not associated, and identification information of the specific sorted field, to identify the specific record associated with the specific cell in the sorted field, the method according to item 14. [Item 16] Before the step of sorting, a step of adding a plurality of placeholder records to the database; Before the step of sorting, a step of adding a status field to each record in the database, where the status field of a given record includes content indicating whether the given record is a placeholder record; After the step of sorting, a step of identifying the sorted placeholder records without reversing or re-sorting the database, replacing the content of the fields of the sorted placeholder records with the content of the fields of new records, and updating the status field of the new records to indicate that they are not placeholder records, thereby adding the new records to the database; further includes, the method according to item 1. [Item 17] Prior to the step of sorting, adding a status field to each record of the database, wherein the status field of a given record contains content indicating whether the given record is a placeholder record; After the step of sorting, identifying the existing sorted records of the database without reversing or re-sorting the database, replacing the content of the fields of the existing sorted records with placeholder content, and updating the status field of the existing records to indicate that they are placeholder records, thereby deleting the existing sorted records; The method according to item 1, further comprising. [Item 18] Prior to the step of sorting, adding a plurality of placeholder fields to each record of the database; After the step of sorting, adding the new fields to the database by adding the content of the new fields to the placeholder fields of each record of the database without reversing or re-sorting the database; The method according to item 1, further comprising. [Item 19] A method for protecting the content of a file from unauthorized access by hiding the relationship between units of the content of a specific type of file, comprising: Determining one or more segment sizes of the specific type of file; Dividing the file into a linear configuration of segments of the one or more determined sizes; Processing the linear configuration as a series of rows and columns, wherein the segments fill the cells of the columns of the rows; The step of rearranging the cells of the specific column or the specific row by applying a rearrangement algorithm using modular arithmetic to the cells of the specific column or the specific row, wherein the rearrangement changes the order of the cells without changing the content of any individual cell; comprising a method of hiding the relationship between the cells of the rearranged column or the rearranged row and the other cells of the file by the rearranged column or the rearranged row still containing all of the original cells in the rearranged order. [Item 20] The step of further comprising rearranging the cells of the plurality of specific columns or specific rows of the file by applying the rearrangement algorithm using modular arithmetic to the cells of each of the plurality of specific columns or specific rows, wherein the rearrangement changes the order of the cells of each specific column or specific row without changing the content of any individual cell; The method according to item 19, wherein each rearranged column or rearranged row still contains all of the original cells in the rearranged order, thereby hiding the relationship between the cells of each rearranged column or rearranged row and the other content of the file. [Item 21] The step of applying the rearrangement algorithm using modular arithmetic further comprises rearranging the cells of the specific column or specific row of the file by applying a bijective function using modular arithmetic, the method according to item 19. [Item 22] The step of applying the rearrangement algorithm using modular arithmetic further comprises applying a rearrangement algorithm using modular addition and modular subtraction in either order, the method according to item 19. [Item 23] The step of applying the rearrangement algorithm using modular arithmetic further comprises The method according to item 19, further comprising applying a sorting algorithm using modular arithmetic with a plurality of parameters. [Item 24] The step of applying a sorting algorithm using modular arithmetic further comprises The method according to item 19, further comprising applying a sorting algorithm using modular arithmetic with a single parameter. [Item 25] The step of applying a sorting algorithm using modular arithmetic further comprises The method according to item 19, further comprising applying a sorting algorithm using modular arithmetic that satisfies cells of sorted columns or sorted rows via a projection field. [Item 26] The one or more segment sizes The method according to item 19, further comprising different sizes of segments that are cells of different columns. [Item 27] A method for protecting content transmitted from a first computer device to a second computer device from unauthorized access by hiding the relationship between units of the content, comprising: The step of obtaining, by the first computer device, a data record to be transmitted to the second computer device, the data record including a plurality of fields; The step of obtaining a specific number of additional data records of the same type as the transmitted data record, the additional records including excessive data for the transmitted record; The step of processing the transmitted record and the additional records as a data grid including a plurality of records, each record of the data grid including a plurality of fields; Sorting the cells of the at least one field by applying a sorting algorithm using modular arithmetic to the cells of at least one field of the data grid, the sorting changing the order of the cells without changing the content of any individual cell; Sending the sorted data grid to the second computing device; comprising; a method, wherein the at least one sorted field still contains all of the original cells in the sorted order, hiding the relationship between the cells of the sorted field and the other cells of the data grid. [Item 28] The step of applying a sorting algorithm using modular arithmetic further comprises sorting the cells of the specific field of the data grid by applying a bijective function using modular arithmetic, according to the method of item 27. [Item 29] The step of applying a sorting algorithm using modular arithmetic further comprises applying a sorting algorithm using modular addition and modular subtraction in either order, according to the method of item 27. [Item 30] The data record sent to the second computing device further comprises a credit card number and related data sent for authenticating a credit card transaction, according to the method of item 27.
Claims
1. A method implemented by a database server comprising a processor, program code, and an unsorted database, the program code, when processed by the processor, causing the processor to perform the method; The method further comprising: converting the unsorted database into a secured database by sorting cells of specific fields of the database; Sorting cells of a particular field of the database Sorting the cells of a particular field of the database by applying a sorting algorithm using modular arithmetic to the cells of the particular field; the reordering involves changing the order of the cells for the particular field across multiple records without changing the contents of any individual cells; the sorted field still contains all of the original cells in sorted order, thereby hiding the relationship between the cells of the sorted field and the associated database records; a transforming step, wherein the transformed and protected database protects the sensitive content from unauthorized access by hiding relationships between the sensitive content and other fields of the associated database record; retrieving, from said secured database, a data record designated for transmission to a second computing device, said data record including a plurality of fields; obtaining a certain number of additional data records of the same type as the data record designated for transmission, the additional data records containing data not responsive to the query that returned the record designated for transmission; processing the data record and the additional data record designated for transmission as a data grid including a plurality of records, each record of the data grid including a plurality of fields; converting the data grid into a secured data grid by reordering cells of at least one field of the data grid; Sorting cells of at least one field of the data grid includes: applying a sorting algorithm using modular arithmetic to the cells of the at least one field; a transforming step, wherein the reordering changes the order of cells of the data grid across multiple records without changing the content of any individual cells; transmitting the sorted data grid to the second computing device; A method comprising:
2. The database stores fields in columns and records in rows; 2. The method of claim 1 , further comprising: applying the sorting algorithm using modular arithmetic to the cells of a particular column of the database, the particular column including the cells of the particular field.
3. The database stores fields in rows and records in columns; 2. The method of claim 1 , further comprising: applying the sorting algorithm using modular arithmetic to the cells of a particular row of the database, the particular row containing the cells of the particular field.
4. reordering the cells of a plurality of particular fields of the database by applying the reordering algorithm using modular arithmetic to the cells of each of the plurality of particular fields, the reordering changing the order of the cells of each particular field without altering the content of any individual cell; 2. The method of claim 1, wherein each sorted field still contains all of the original cells in sorted order, thereby hiding relationships between the cells of each sorted field and the associated database records.
5. applying a sorting algorithm using modular arithmetic, 2. The method of claim 1, further comprising reordering the cells of the particular field of the database by application of a bijective function using modular arithmetic.
6. applying a sorting algorithm using modular arithmetic, The method of claim 1 , further comprising applying a reordering algorithm that uses modular addition and modular subtraction, in any order.
7. applying a sorting algorithm using modular arithmetic, The method of claim 1 , further comprising applying a sorting algorithm that uses modular arithmetic involving multiple parameters.
8. applying a sorting algorithm using modular arithmetic, The method of claim 1 , further comprising applying a sorting algorithm that uses modular arithmetic with a single parameter.
9. applying a sorting algorithm using modular arithmetic, The method of claim 1 , further comprising applying a reordering algorithm that uses modular arithmetic to fill cells of the reordered field via the projection field.
10. applying a sorting algorithm using modular arithmetic, The method of claim 1 , further comprising applying a reordering algorithm that uses modular arithmetic with at least one pseudo-random selection parameter.
11. applying a sorting algorithm using modular arithmetic, The method of claim 1 , further comprising applying a sorting algorithm that uses modular arithmetic with at least one deliberately selected parameter.
12. 2. The method of claim 1, further comprising the step of locating in the sorted field particular cells that are associated with, but not located within, a particular record by applying in a modular operation one or more parameters used in the step of applying the sorting algorithm.
13. Locating a particular cell in the sorted field by applying in a modular operation one or more parameters used in the step of applying the sorting algorithm, 13. The method of claim 12, further comprising locating the particular cell in the sorted field by applying in the modular operation the one or more parameters, an identification of the particular record with which the particular cell of the sorted field is associated, and an identification of the particular sorted field.
14. 2. The method of claim 1, further comprising the step of locating a particular record with which a particular cell in the sorted field is associated by applying one or more parameters used in the step of applying the sorting algorithm in a modular operation, the particular cell not being located in the particular record with which it is associated.
15. locating a particular record with which a particular cell in the sorted field is associated by applying one or more parameters used in the step of applying the sorting algorithm in a modular operation, 15. The method of claim 14, further comprising locating the particular record with which the particular cell in the sorted field is associated by applying in the modular operation the one or more parameters, an identification of the particular record in which the particular cell of the sorted field is located but is not associated, and an identification of the particular sorted field.
16. adding a plurality of placeholder records to the database prior to the sorting step; prior to said sorting step, adding a status field to each record of said database, said status field of a given record having content indicating whether said given record is a placeholder record; adding the new record to the database after the reordering step, without back-sorting or re-sorting the database, by locating the reordered placeholder record, replacing the contents of the fields of the reordered placeholder record with the contents of the fields of a new record, and updating the status field of the new record to indicate that it is not a placeholder record; The method of claim 1 further comprising:
17. prior to said sorting step, adding a status field to each record of said database, said status field of a given record having content indicating whether said given record is a placeholder record; after said reordering step, without unsorting or reordering said database, locating an existing reordered record in said database, replacing the contents of the fields of said existing reordered record with placeholder content, and erasing said existing reordered record by updating said status field of said existing record to indicate that it is a placeholder record; The method of claim 1 further comprising:
18. adding a plurality of placeholder fields to each record in the database prior to the sorting step; adding the new field to the database by adding the contents of the new field to a placeholder field of each record of the database without back-sorting or re-sorting the database after the sorting step; The method of claim 1 further comprising:
Citation Information
Patent Citations
Reading image processor
JP1996147329A
How to make configuration changes in relational database tables
JP2002507795A
Data protection program and data protection method
JP2006185096A
Data management device, data management method, data processing method, data storage method, and program
JP2007034423A
File protection system, method and program
JP2009251748A