Multi-source data fusion calculation method and multi-source data fusion calculation system
By constructing the address mapping between the data encoding sequence and the database during the system initialization stage, and querying the function function value based on the mapping during the system operation stage, the problem of large calculation and communication overhead in multi-source data fusion calculation is solved, efficiency is improved and data privacy is guaranteed.
Patent Information
- Application Number
- CN202510233660.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The calculation overhead and communication overhead are large during the multi-source data fusion calculation process, resulting in low efficiency.
In the system initialization stage, an address mapping between the data encoding sequence and the database is constructed, and the target address is calculated based on the data encoding and address mapping of the secret data of each data party, and the function value is stored in the target address. During the system operation phase, the database is queryed based on address mapping and data encoding to obtain functional function values.
By focusing on performing multi-source data fusion computing tasks in the system initialization stage, the computing and communication overhead in the system operation stage is reduced, the efficiency of multi-source data fusion computing is improved, and data privacy is guaranteed.
Smart Images

Figure CN119719159B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computers. Specifically, the present application relates to a multi-source data fusion calculation method, a multi-source data fusion calculation system, a computer-readable storage medium, and an electronic device. Background Art
[0002] Multi-source data fusion calculation refers to integrating and analyzing multiple data from different sources in order to extract valuable information from them. It involves a very wide range of fields, including intelligent transportation, healthcare, environmental monitoring, and financial analysis, etc. Although multi-source data fusion calculation has broad application prospects, it also faces some challenges, such as data heterogeneity, computational complexity, privacy protection, and real-time requirements, etc.
[0003] In order to improve the implementation performance of the privacy calculation technology for multi-source data fusion calculation, some solutions have proposed a series of improvement and optimization solutions from the perspectives of algorithm design or software and hardware implementation. Although these improvement and optimization measures have achieved certain results, they still have limitations in practical applications. The main reason is that the computational overhead and communication overhead in the process of multi-source data fusion calculation are relatively large, resulting in low efficiency of multi-source data fusion calculation. Summary of the Invention
[0004] Embodiments of the present application provide a multi-source data fusion calculation method, a multi-source data fusion calculation system, a computer-readable storage medium, and an electronic device, so as to at least solve the problem that the computational overhead and communication overhead in the process of multi-source data fusion calculation in related technologies are relatively large, resulting in low efficiency of multi-source data fusion calculation.
[0005] The present application provides a multi-source data fusion calculation method. In the system initialization stage, an address mapping between a data coding sequence and a database is constructed, and a target address is calculated based on the data coding of the secret data of each data party and the address mapping. The function function value obtained by performing a privacy calculation task based on the secret data of each data party is stored in the target address, where the function function value corresponds to the data coding sequence, and the data coding sequence is an ordered vector composed of multiple data codings of all data parties; in the system operation stage, based on the address mapping and the data coding of each data party, the address of the function function value corresponding to the current secret data of each data party in the database is calculated, and the function function value is obtained by querying the database according to the address, where the data coding of the current secret data is an element in the data coding sequence corresponding to the current address.
[0006] The present application also provides another embodiment according to the present application, which provides a multi-source data fusion computing system. The multi-source data fusion computing system includes a data provider, a requester, and a privacy computing system. The privacy computing system includes multiple computing parties. Wherein, in the system initialization stage, the requester constructs an address mapping between the data coding sequence and the database based on the data of each data provider; each data provider is used to randomly code the held secret data to obtain corresponding data coding, and send the data coding of the secret data to the requester; the requester calculates the function function value, and calculates the address of the function function value corresponding to the data coding sequence in the database based on the received data coding of each data provider and the address mapping, and then stores the function function value in the corresponding address of the database; in the system operation stage, the data provider is used to send the data coding of the current secret data to the requester, so that the requester can query the database according to the address mapping to obtain the function function value corresponding to the current secret data of each data provider.
[0007] The present application also provides yet another embodiment according to the present application, and also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. Wherein, the computer program is set to execute the steps in any of the above method embodiments when running.
[0008] The present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory. The processor is set to run the computer program to execute the steps in any of the above method embodiments.
[0009] Through this application, in the system initialization stage, an address mapping between the data coding sequence and the database is constructed, and the target address is calculated based on the data coding of the secret data of each data party and the address mapping. The function function value obtained by performing the privacy calculation task based on the secret data of each data party is stored in the target address, where the function function value corresponds to the data coding sequence, and the data coding sequence is an ordered vector composed of multiple data codings of all data parties; in the system operation stage, based on the address mapping and the data coding of each data party, the address of the function function value corresponding to the current secret data of each data party in the database is calculated, and the function function value is obtained by querying the database according to the address, where the data coding of the current secret data is an element in the data coding sequence corresponding to the current address. In this way, the construction of the address mapping, the privacy calculation task, and the storage of the function function value are completed in the system initialization stage, and only the corresponding function function value needs to be obtained from the corresponding address according to the address mapping in the system operation stage. That is to say, this solution concentrates the multi-source data fusion calculation task in the system initialization stage, and only performs simple calculations on a small amount of data in the system operation stage. The calculation overhead and communication overhead in the multi-source data fusion calculation process are reduced, and the efficiency of the multi-source data fusion calculation is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a schematic flowchart of a multi-source data fusion calculation method according to an embodiment of the present application;
[0011] Figure 2 is a schematic architecture diagram of a first multi-source data fusion calculation system according to an embodiment of the present application;
[0012] Figure 3 is a schematic architecture diagram of a second multi-source data fusion calculation system according to an embodiment of the present application;
[0013] Figure 4 is a schematic diagram of a tree-structured database of a multi-source data fusion calculation system according to an embodiment of the present application;
[0014] Figure 5 is a schematic diagram of the construction of a database based on the inverse of the address mapping of a multi-source data fusion calculation system according to an embodiment of the present application;
[0015] Figure 6 is a schematic diagram of the construction of a database based on multiple loops of a multi-source data fusion calculation system according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the protection scope of the present application.
[0017] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0018] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] A multi-source data fusion calculation method is provided in this embodiment. Figure 1 It is a schematic flowchart of a multi-source data fusion calculation method according to an embodiment of the present application, as Figure 1 shown. The above method includes the following steps:
[0020] Step S101: In the system initialization stage, construct an address mapping between the data coding sequence and the database, and calculate the target address based on the data coding of the secret data of each data party and the above address mapping. Store the functional function value obtained by performing the privacy calculation task based on the secret data of each data party in the above target address, where the functional function value corresponds to the data coding sequence, and the data coding sequence is an ordered vector composed of multiple data codings of all data parties;
[0021] Among them, one of the above data coding sequences is obtained by arranging the data codings of all data parties in a certain order, and one data coding sequence includes one data coding of each data party;
[0022] Specifically, a privacy computing task is executed based on a privacy computing protocol to calculate a functional function value. The privacy computing protocol is a key technology to ensure that data will not be leaked during the fusion computing process, such as Homomorphic Encryption (HE), Multi-Party Computation (MPC), etc. The functional function can be any function that performs fusion computing on multi-source data, such as statistical average, model prediction results, etc., but its specific definition depends on application requirements.
[0023] First, an address mapping is constructed to associate the data encoding sequence with a specific address in the database. Based on the constructed address mapping, the target address corresponding to each data encoding sequence is calculated. The calculation of the target address ensures that the functional function value can be correctly stored and retrieved. The functional function value is the result obtained after executing the privacy computing task based on the secret data of each data party. In the system initialization stage, based on the data encoding of the secret data sent by the data parties, the privacy computing task is executed. After the calculation is completed, these functional function values are stored in the database, specifically in the target address corresponding to the data encoding sequence. Since the functional function value corresponds one-to-one with the data encoding sequence, this storage mechanism ensures that the functional function value can be quickly located and retrieved according to the queried data encoding sequence.
[0024] In the system initialization stage, by constructing an address mapping between the data encoding sequence and the database, and calculating the target address based on the data encoding of the secret data of each data party and the address mapping, and storing the functional function value in the target address, the present application provides an efficient and secure data processing and storage method. This mechanism ensures the orderliness of the data encoding sequence, the correct storage of the functional function value, and the rapidity and privacy of database queries, providing a solid foundation for data queries and real-time data fusion in the subsequent system operation stage.
[0025] Step S102: In the system operation stage, based on the above address mapping and the data encoding of each of the above data parties, calculate the address of the functional function value corresponding to the current secret data of each of the above data parties in the above database, and query the database according to the above address to obtain the above functional function value, where the data encoding of the above current secret data is an element in the data encoding sequence corresponding to the current address.
[0026] Among them, the data encoding of the current secret data of the data party is encrypted and generated by the data party in the initialization stage, and corresponds one-to-one with the actual secret data.
[0027] Based on the data encoding of the current secret data of the data provider, using the address mapping constructed in the system initialization phase, which enables the calculation of the address in the database from the data encoding without knowing the actual secret data corresponding to the data encoding. The design of the address mapping ensures that each data encoding sequence corresponds to a unique address in the database, thus enabling the accurate positioning of the functional function value.
[0028] Directly query the database based on the address calculated from the address mapping. What is stored in the database is the functional function value calculated in the system initialization phase, and each value corresponds to a specific data encoding sequence and is stored at a predefined address in the database. Without performing complex privacy calculation tasks, simply querying the database can obtain the required functional function value.
[0029] Finally, the functional function value corresponding to the data encoding of the current secret data of the data provider can be obtained from the database through a simple query operation. These functional function values are obtained through the fusion calculation of multi-source data without revealing the original data, satisfying data privacy protection while providing the functions of data analysis and fusion calculation. Moreover, since each element in the data encoding sequence corresponds to a specific secret data, the data encoding of the current secret data is the element in the data encoding sequence corresponding to the current address, which ensures that the queried functional function value is directly related to the current secret data of the data provider.
[0030] In summary, by constructing an address mapping between the data encoding sequence and the database in the system initialization phase, the core in the system operation phase is to utilize the address mapping and database query, which can quickly and directly obtain the functional function value corresponding to the data encoding of the current secret data of the data provider without performing real-time privacy calculations, greatly improving the system's response speed, reducing communication overhead, and protecting the data privacy of the data provider throughout the process. This method provides an efficient and privacy-secure solution for real-time data fusion scenarios.
[0031] In summary, in the system initialization stage of the multi-source data fusion calculation method proposed in this application, an address mapping between the data coding sequence and the database is constructed, and based on the data coding and address mapping of each data party, the address of the function value corresponding to the data coding sequence in the database is calculated, and the calculated function value is stored at the corresponding address; in the system operation stage, based on the above address mapping and the data coding of each of the above data parties, the address of the function value corresponding to the current secret data of each of the above data parties in the above database is calculated, and the function value is obtained by querying the database according to the above address, where the data coding of the current secret data is an element in the data coding sequence corresponding to the current address. In this way, the construction of the address mapping, the privacy calculation task, and the storage of the function value are completed in the system initialization stage, and in the system operation stage, only the corresponding function value needs to be obtained from the corresponding address according to the address mapping, that is, this solution concentrates the multi-source data fusion calculation task in the system initialization stage, and only performs simple calculations on a small amount of data in the system operation stage. The calculation overhead and communication overhead in the multi-source data fusion calculation process are reduced, the efficiency of the multi-source data fusion calculation is improved, and the privacy of each party's data is guaranteed at the same time.
[0032] In one embodiment of this application, before storing the function value obtained by performing the privacy calculation task based on the secret data of each of the above data parties in the above target address, the method further includes: encrypting the secret data held by each of the above data parties to obtain ciphertext; performing the above privacy calculation task based on the ciphertext obtained by the encryption process to obtain an intermediate result; calculating the above function value according to the above intermediate result.
[0033] In a specific implementation, the way to calculate the function value depends on the specific privacy calculation protocol adopted. For example: if it is a protocol based on additive secret, the function value is obtained by summing all the intermediate results; if it is a protocol based on Shamir secret sharing, the function value can be obtained by calculating the Lagrange interpolation polynomial or solving a system of linear equations.
[0034] Specifically, before the data fusion calculation starts, the secret data held by the data party is encrypted. Encryption is the process of converting the original data into ciphertext using cryptographic techniques, with the aim that even if the data is intercepted by a third party during transmission or storage, the third party cannot interpret the original content of the data. Encryption can use secure encryption algorithms such as the Advanced Encryption Standard (AES) or techniques based on homomorphic encryption (HE) to ensure the confidentiality of the data when it is processed in the privacy calculation system.
[0035] Based on the ciphertext obtained after encryption processing, perform specific privacy computing tasks. Privacy computing tasks generally include, but are not limited to, homomorphic encryption computing, secure multi-party computing (MPC), computing in a trusted execution environment (TEE), etc. These computing technologies can perform operations on data without revealing the data itself and obtain intermediate results. There are multiple intermediate results, and each computing party corresponds to one intermediate result. The intermediate result is a key step in the privacy computing process, containing partial computing results or encrypted-form numerical values for subsequent computing.
[0036] Process and calculate the intermediate results obtained from the privacy computing tasks to obtain the functional function values. In this step, various mathematical methods or computing protocols can be used to combine or process the intermediate results to obtain the functional function values associated with the fused data. The functional function values reflect certain characteristics or results of the data of the data parties after the fusion calculation, but do not contain any specific information that can be traced back to the original data, thus providing the function of data fusion calculation while ensuring data privacy.
[0037] In summary, the purpose of the above steps is to provide a secure data processing environment for the subsequent database construction and system operation phases. The application of data encryption and privacy computing technologies ensures that the data privacy of the data parties is maximally protected during the data fusion calculation process, while ensuring the accuracy and effectiveness of the calculation results. This method provides a basis for the efficient fusion calculation of multi-source data while protecting data privacy.
[0038] In the embodiment of determining the data encoding of the present application, the above method further includes: The first method: In the above system initialization phase, traverse the addresses of the above database, and based on the above address mapping, determine the above data encoding sequence corresponding to the traversed address; The second method: In the above system initialization phase, based on the data encoding sets of the above data parties, traverse the data encoding sequences in a nested loop manner. In the above system initialization phase, traverse the above data encodings of each of the above data parties respectively based on the nested loop traversal method to obtain the corresponding data encoding sequences. That is to say, the above two methods can be used to implement the traversal of the data encoding sequences. The first method is based on traversing the addresses of the database, and the second method is to directly traverse the data encoding sets of the data parties.
[0039] In a more specific embodiment, in the above system initialization phase, traversing the data encoding sequences in a nested loop manner based on the data encoding sets of the above data parties includes:
[0040] Data party From 1 to Traverse the data encoding ; For Data party From 1 to Traverse the data encoding ; For each two-dimensional sequence traversed , the data party From 1 to Traverse the data encoding ; Until, for each m - 1 dimensional sequence traversed From 1 to Traverse the data encoding , to obtain the data encoding sequence ; Among them, the data party From 1 to Traverse the data encoding , indicating that the data party Will Be assigned in sequence as , where Indicates the data volume of the data party , Indicates the number of the above data parties;
[0041] In the above system initialization stage, based on the data encoding sets of each of the above data parties, traverse the data encoding sequence in a nested loop manner. After that, the above method further includes: In the above system initialization stage, calculate the database address corresponding to the above data encoding sequence according to k , where Indicates the data encoding of each of the above data parties.
[0042] The above, the data encoding set refers to the set composed of all possible values of the encoding mapping of each data party. For example, the data encoding set of the data party Is , here . And all data encoding sets involved in this application represent this meaning.
[0043] The above, in the second method, for each traversed data encoding sequence, the requester needs to calculate the database address according to the data encoding sequence, which is different from the idea of directly traversing the database address in the first method.
[0044] It can be understood that a key operation in the first method is to traverse the addresses in the database. The database is designed to be able to store the functional function values corresponding to each data coding sequence. To build a complete database, it is necessary to traverse every address in the database, and this process ensures that all possible data coding sequences are taken into account. When traversing each database address, an address mapping is used to determine the data coding sequence corresponding to the current address. The address mapping is an algorithm or function that performs a one-to-one matching of the data coding sequence with a specific address in the database. Through the address mapping, the above-mentioned data coding sequence corresponding to the currently traversed address can be obtained.
[0045] For each data coding sequence, it is necessary to determine its corresponding actual secret data. These secret data are provided by the data provider, and once the secret data are determined, they can be used for subsequent privacy calculations. Among them, the data coding of the secret data is obtained by means of random coding. Random coding is a mechanism that converts the original data into a coded form. This coding does not directly disclose any information about the data, but can be used for subsequent calculations.
[0046] Regarding the second method, in the system initialization stage, the data provider traverses the data coding sequences based on multiple loops. This usually involves multiple data providers cooperating with each other to traverse their respective data codings to build a complete data coding sequence. Each data provider has its own data coding, and these codings form a data coding sequence through specific combination rules. Based on the multiple-loop traversal, it can be ensured that the functional function values stored in the database cover all combinations of data coding sequences, providing comprehensive data support for the subsequent system operation stage.
[0047] The above steps are executed in the system initialization stage with the aim of building a complete database that stores the functional function values corresponding to all possible data coding sequences. Through random coding and address mapping, a rapid traversal of the data coding sequences can be achieved; by traversing the data coding sequences in a multiple-loop-based manner, the comprehensiveness and accuracy of the database can be ensured. This process provides an important foundation for the subsequent system operation stage, enabling the rapid acquisition of the functional function values corresponding to the data coding sequences without having to re-execute the privacy calculation tasks.
[0048] In a specific embodiment of determining the data coding, traversing the addresses of the above-mentioned database and determining the above-mentioned data coding sequence corresponding to the traversed address based on the above-mentioned address mapping includes: calculating the address parameter , where , , represents the address of the above-mentioned database, represents the number of the above-mentioned data providers, Indicates the second data party of the data volume, indicating the th data party of the data volume; set , , and determine , and judge whether it can be divided evenly by , and determine the th data encoding in the data encoding sequence according to the result of whether it can be divided evenly, indicating the th data encoding in the above data encoding sequence. Among them, is specifically expressed as: .
[0049] Specifically, the address parameter is used to determine the data encoding sequence corresponding to the current database address. In the calculation of the address parameter, the database address , the number of data parties , and the data scale of each data party (for example, the data volume of the second data party and the data volume of the th data party ) are used. The calculation of the address parameter is based on the address mapping mechanism and is used to associate the database address with the data encoding sequence.
[0050] After calculating the address parameter, set two parameters and . is set to initialize a variable for subsequent calculations, while is set to record the position of the current data encoding in the sequence during the traversal. The initialization of these two parameters is the starting point of the traversal process, and through these two parameters, the data encoding sequence corresponding to the current address can be gradually determined.
[0051] Next, it is necessary to determine a parameter , and make a judgment: judge whether it can be divided evenly by . The purpose of this judgment is to determine how the data encodings of each data party should be selected and calculated in the current calculation process.
[0052] Through the above steps, it is ensured that when traversing the database addresses, the data encoding sequence corresponding to each address can be correctly and securely determined. This process not only takes into account the privacy protection of the data, but also utilizes the data scale of the data provider and the parametric calculations during the traversal process to ensure that the construction of the database is both efficient and secure. After completing these calculations during the system initialization phase, the database will be able to store the function values corresponding to all possible data encoding sequences, providing the necessary data support for the subsequent system operation phase while protecting the data privacy of all parties.
[0053] In summary, the above steps are executed during the system initialization phase to traverse the database addresses and determine the data encoding sequence based on the address mapping. By calculating the address parameters, setting the initialization parameters, performing the divisibility judgment, and determining the data encoding, a secure and comprehensive database can be constructed, which stores the function values corresponding to all possible data encoding sequences. These steps ensure the protection of data privacy and also provide the necessary data support for subsequent multi-source data fusion calculations.
[0054] In a specific alternative embodiment for determining the data encoding, the th data encoding in the data encoding sequence is determined according to the result of whether it can be divided evenly, including: the first judgment and determination step: if , calculate ; the second judgment and determination step: if , calculate ; the calculation step: ; where means divisible, means not divisible, means modulo operation, for example means divided by remainder; the data provider repeatedly executes the above first judgment and determination step, the above second judgment and determination step, and the above calculation step to obtain the data encoding , where .
[0055] In this embodiment, it is further elaborated in detail how to determine the specific data encoding in the data encoding sequence according to the result of the divisibility judgment. The following is a detailed explanation of this process:
[0056] In the first judgment and determination step, the data provider will first execute a judgment to check whether can be divided evenly by . If this condition holds (i.e., the divisibility judgment is true), then a calculation will be executed to determine the A data encoding. The result of this calculation will directly affect the generation of the data encoding, ensuring that it matches the data scale of the data party and the structure of the entire database.
[0057] In the second determination step, if the first determination condition is not satisfied (i.e., not divisible ), then another calculation will be performed to determine the th data encoding in the data encoding sequence. Similarly, the result of this calculation will be used to generate the data encoding, ensuring that the construction of the data encoding sequence follows the rules set in the system initialization phase.
[0058] In the calculation step, whether it is the first determination step or the second determination step, a specific calculation process will be involved. The calculation formula is .
[0059] The data party repeatedly executes the above first determination step, second determination step, and calculation step to obtain the th data encoding in the data encoding sequence . This repeated execution process ensures that for each database address, the corresponding data encoding can be accurately calculated based on the address mapping and the data scale of each data party. The generation of the data encoding depends on the execution of these steps and the corresponding calculation results.
[0060] The above steps are executed in the system initialization phase to determine the data encoding sequence corresponding to each address in the database address mapping. By performing the divisibility judgment and related calculations, the data encoding related to the current address can be generated, ensuring that the construction of the database not only follows the design principles but also protects the data privacy of the participating parties. This process involves mathematical operations such as divisibility judgment and modulo operation, making the construction of the data encoding sequence accurate and secure while meeting the system requirements.
[0061] In a specific embodiment of determining the data encoding, traverse the addresses of the above database, and based on the above address mapping, determine the above data encoding sequence corresponding to the traversed address, including: set , , , , for , represents the address of the above database, represents the number of the above data parties, represents the th data party 's data volume, represents the th data party The amount of data, the data provider Calculate Whether it can be divided evenly , and determine the data encoding in the data encoding sequence according to the result of whether it can be divided evenly , where the floor operation is Cannot be divided evenly in the case of; the data provider Calculate .
[0062] It can be understood that this embodiment further details the steps of traversing the database address and determining the data encoding sequence corresponding to the traversed address in the system initialization stage. First, initialize , assign the database address k to the variable , as the starting point of the traversal and calculation process. Next, define a sequence, where , that is, starting from the th data provider 's data volume , to the th data provider 's data volume product of all data volumes. The definition of the sequence is used in subsequent calculations as the basis for judging whether it can be divided evenly.
[0063] For 's range, it means that the traversal will start from the first data provider and end at the th data provider. For each data provider , the following calculations will be performed:
[0064] Calculate : In each loop, calculate the value of the variable , that is . In this calculation, is the current variable value. Here is obtained through the floor operation divided by 's integer part of the quotient. This operation ensures that even when cannot be divided evenly by , an integer result can still be obtained for subsequent calculations. The floor operation plays a key role in integerization and normalization in the traversal and calculation process. The above calculations ensure that when traversing the database address, according to the address mapping rule and the state of the data encoding sequence, 's calculation result is used to update the current address mapping state during the traversal, providing a basis for subsequent data encoding calculations.
[0065] Data provider Judge Whether it can be divided evenly This judgment is very crucial because it is used to determine the data encoding in the data encoding sequence If Can be divided evenly Then the data provider Can directly determine ; If Cannot be divided evenly Then calculate through floor operation And based on The result of further determine This process ensures that the construction of the data encoding sequence follows the preset address mapping rules, and also considers the role of the data volume of the data provider in the calculation
[0066] At the end of the traversal and calculation process, the data provider Will calculate its data encoding That is This is because when the traversal reaches the last data provider Will directly correspond to the last data encoding, without the need for additional divisibility judgment or adjustment. This step ensures that the last data encoding in the data encoding sequence can also be accurately calculated and corresponds to the database address, completing the construction of the entire data encoding sequence
[0067] The entire process through loops and conditional judgments ensures that the complete data encoding sequence corresponding to each address in the database can be correctly calculated. This sequence contains the random encodings of all data providers, associated with each address in the database through the address mapping mechanism, providing an accurate data basis for subsequent functional function calculations and database construction, while protecting the data privacy of the data providers
[0068] In summary, this embodiment describes the key calculation steps of traversing the database address and determining the data encoding sequence in the system initialization stage of the multi-source data fusion calculation method. By initializing variables, defining sequences, performing traversal calculations, making divisibility judgments, and determining data encodings, it ensures that the construction of the database follows both the address mapping rules and considers the data volume differences between data providers, providing accurate and secure data support for the subsequent system operation stage
[0069] In a specific embodiment of determining the data encoding, determine the data encoding in the data encoding sequence according to the result of whether it can be divided evenly The data provider Calculate Including: Execute the steps: Let For , the data party executes the first calculation sub-step, the second calculation sub-step, and the third calculation sub-step to calculate and obtain the above data encoding ; the above first calculation sub-step: if , calculate , and no longer execute the above second calculation sub-step, the above third calculation sub-step, and the calculation step, where , represents divisible by , represents the -th data encoding in the above data encoding sequence, represents the -th data encoding in the above data encoding sequence; the above second calculation sub-step: if , calculate , where represents not divisible by , represents rounding down; the above third calculation sub-step: calculate ; calculation step: the data party calculates .
[0070] Specifically, in this embodiment, the calculation process of the data encoding in the database initialization stage is further refined, and the steps for determining each data encoding in the data encoding sequence according to the database address are specifically explained. First, let , where is the address of the database, and assign the address of this database to variable as the starting point for traversal and calculation. Traverse the data encoding sequence for the range of , and the data party executes a series of calculations to determine each data encoding in the data encoding sequence. This traversal process ensures that the complete data encoding sequence corresponding to the database address can be calculated.
[0071] First calculation sub-step: If during the traversal process, is divisible by , then the first calculation sub-step executed by the data party is to calculate to obtain the -th data encoding in the data encoding sequence. In addition, for each data party from to , its data encoding will be directly assigned its data volume . The first calculation sub-step is applied when is divisible by , simplifying subsequent calculations and no longer executing the second calculation sub-step, the third calculation sub-step, and the calculation step.
[0072] Second calculation sub-step: If is not divisible by , the data party executes the second calculation sub-step which is to calculate . This step is applied when is not divisible by , ensuring the accuracy of data encoding through floor operation and considering the impact of the data volume of the data party on data encoding.
[0073] Third calculation sub-step: After executing the second calculation sub-step, the data party will also execute the third calculation sub-step to calculate , updating the traversal variable to prepare a new state for the next loop. This update step ensures that the traversal process can correctly transition from one data encoding to the next until the entire data encoding sequence is traversed.
[0074] Calculation step: When the traversal reaches the last data party , its data encoding is directly assigned to , that is . This is because is the last data party, and its data volume is directly related to the result of the traversal process , and there is no other data party to consider, simplifying the calculation process.
[0075] Through initialization, traversal, conditional judgment, and calculation update, the entire process ensures that the data encoding of each data party can be correctly calculated according to its data volume and address mapping rules and stored in the data encoding sequence. The calculation step of the data party ensures that the data encoding of the last data party can be directly determined, completing the construction of the entire data encoding sequence and providing accurate data support for subsequent database construction and functional function calculation.
[0076] Generally speaking, this embodiment details how to calculate each data encoding in the data encoding sequence corresponding to the database address according to the address mapping and the data volume of the data party through the first calculation sub-step, the second calculation sub-step, the third calculation sub-step, and the calculation step. This calculation process not only considers the protection of data privacy but also ensures the accuracy and efficiency of database construction.
[0077] In yet another specific embodiment of determining the data encoding, traverse the addresses of the above database, and determine the above data encoding sequence corresponding to the traversed address based on the above address mapping, including: traversing the addresses of the above database and calculating the inverse of the above address with respect to the above address mapping to obtain the above data encoding sequence.
[0078] Specifically, the inverse operation of the address mapping is used, which means that for each address in the database, through a reverse calculation process, the address mapping can be applied in reverse to determine the data encoding sequence corresponding to the address. The specific steps are as follows:
[0079] Traverse the database addresses: Start from the first address of the database and gradually traverse to each address. This traversal process is the basis of the system initialization phase, ensuring that each address can be accessed and processed.
[0080] Inverse mapping calculation: For each traversed database address, calculate the inverse of the address with respect to the address mapping. This inverse mapping calculation is achieved by applying the address mapping rules in reverse, which can convert the address into the corresponding data encoding sequence.
[0081] Importance of the inverse mapping calculation The inverse mapping calculation ensures the correct correspondence between the data encoding sequence and the database address, which is the core link in the database construction process. Through the inverse mapping, the data encoding sequence corresponding to each address in the database can be accurately determined, ensuring the correct calculation and storage of the function values, as well as the accuracy of subsequent data queries.
[0082] The inverse mapping calculation also takes into account the data scale differences of the data parties, so that even if the data volumes between the data parties are different, the data encoding sequences can be correctly associated through the address mapping mechanism. This step is the key to realizing data privacy protection and efficient database construction, because it allows the calculation of function values and the storage of results without exposing the original data.
[0083] In yet another specific embodiment of determining the data encoding, perform random encoding on the above secret data to obtain the corresponding data encoding, including: setting a data encoding set, and randomly selecting a data encoding from the above data encoding set to encode any one of the above secret data of the data party, where the above data encoding corresponds to the above secret data one by one, and the encoding mapping is used to represent the mapping between the above secret data and the above data encoding.
[0084] Specifically, in the data encoding stage, the data provider will first set up a data encoding set, which contains a series of predefined data encodings that will be used to encode the secret data. The size and element selection of the data encoding set should ensure the randomness and security of the encoding, preventing the reverse inference of the original secret data through the encoding. For each piece of secret data held by the data provider, the data provider randomly selects a data encoding from the data encoding set. This random selection process is confidential to ensure that external parties cannot infer the value of the specific secret data from the data encoding. The purpose of random encoding is to enhance the intensity of data privacy protection, so that even during the data fusion calculation process, the real data of the data provider will not be leaked. A one-to-one encoding mapping relationship is established between the randomly selected data encoding and the secret data.
[0085] Each data encoding corresponds to a specific secret data. This one-to-one encoding mapping ensures the integrity and traceability of the data, and also provides a necessary basis for secure multi-party computing. To represent the mapping relationship, the data provider will use an encoding mapping to represent the correspondence between the secret data and the data encoding. This mapping ensures that the generation of the data encoding is performed according to the attributes of the secret data, and the confidentiality of the mapping guarantees the privacy of the data.
[0086] Through the above steps, without disclosing the secret data itself, it can be securely transformed into a data encoding, which can participate in the subsequent fusion calculation process. At the same time, due to the uniqueness and confidentiality of the encoding mapping, even if the data encoding is known to other participating parties, the original secret data cannot be reverse-inferred, thus realizing the effective protection of data during the fusion calculation process.
[0087] In another embodiment of the present application, the above method further includes: secretly storing the above encoding mapping between the held secret data and the above data encoding, where the above encoding mapping is a bijective mapping; the above method further includes: in the case where the above data encoding is known and the above secret data is unknown, calculating the inverse of the above encoding mapping with respect to the above data encoding to obtain the above secret data corresponding to the above data encoding.
[0088] Specifically, secretly store the encoding mapping between the held secret data and the data encoding. This encoding mapping is a bijective mapping, that is, each secret data corresponds to a specific data encoding, and this correspondence is unique, that is, no two different secret data will be mapped to the same data encoding. Similarly, each data encoding also corresponds to only one specific secret data, ensuring the uniqueness and reversibility of the data encoding, thereby enhancing the security of the system.
[0089] During the data fusion calculation process, data encoding is used to replace the secret data for secure multi-party calculation or storage in the database. Since the mapping relationship between the data encoding and the secret data is confidential, this ensures that even if the data encoding is known to other parties, the original secret data cannot be inferred reversely, thus protecting data privacy.
[0090] Inverse operation of the encoding mapping In the case where the data encoding is known and the secret data is unknown, the inverse operation of the encoding mapping can be utilized to calculate the secret data corresponding to the known data encoding. This is because the encoding mapping is a bijective function, ensuring the existence and uniqueness of its inverse operation. The inverse operation allows the data party to convert the data encoding back to the original secret data when needed without the participation of other parties.
[0091] In summary, this embodiment emphasizes the confidentiality and reversibility of the encoding mapping, that is, the encoding mapping needs to be secretly stored, and in the case where the data encoding is known, the corresponding secret data can be calculated through the inverse operation. This feature is the core of realizing secure data fusion calculation, ensuring the protection of data privacy, and at the same time providing a means for the data party to decode the data, so that the original data can be restored when needed. Through this encoding mapping mechanism, the effective fusion and calculation of data can be achieved without revealing the original secret data, improving the security and efficiency of data fusion calculation.
[0092] In an embodiment of constructing the address mapping of this application, the above address mapping is a mapping based on a tree-structured database, where the tree-structured database is a level-i linear partition table. The first level divides the above database into m first-level intervals, and the i-th level divides each i-th level interval into n i + 1-level intervals, , and each i-th level interval contains k addresses, The i-th level interval corresponds to the address of the above database, m represents the number of the above data parties, n_i represents the data volume of the i-th data party , and n_j represents the data volume of the j-th data party .
[0093] Specifically, this embodiment emphasizes the implementation method of the address mapping, adopting a method of constructing a tree-structured database, which further optimizes the structure of the database and the storage of function values, and also takes into account the data volume differences of data parties, improving the overall efficiency of the system and the level of data privacy protection.
[0094] This embodiment adopts an address mapping based on a tree structure, constructing the entire database into a multi-level linear partition table. The core of this structure lies in using the data scale (data volume) of the data party to perform hierarchical partitioning of the database, so that the address mapping can more effectively locate and store function values.
[0095] The first-level partition: The database is divided into first-level intervals, which is the data volume of the first data party . That is to say, the top layer of the database is determined by the data scale, and each first-level interval contains a certain amount of function values or mapping results of data coding sequences.
[0096] The (i + 1)-th level partition: For the -th level, each -th level interval is further divided into -th level intervals, where is the data volume of the -th data party . This level of partition is carried out on the basis of the previous level of partition, reflecting the gradual refinement of the data volume of the data party, ensuring that the address mapping can accurately correspond to each address in the database.
[0097] In the tree-structured database, the number of addresses contained in each -th level interval is , until the -th level interval corresponds to the address of the database. This means that as the partition level increases, the number of addresses is determined by the data volume of the subsequent data parties, forming an address space based on the data scale of the data parties.
[0098] The -th level interval directly corresponds to the address of the database, and is the number of data parties, which means that each leaf node of the database (i.e., the
[0099] -th level interval) has a specific address, and this address corresponds to the data coding sequence through the rules of address mapping, directly pointing to the storage location of the function value in the database.By adopting a tree - structured address mapping, this embodiment can efficiently divide the database into multiple levels of linear partitions according to the data volume of the data providers. This address mapping mechanism not only optimizes the storage of the database structure and functional function values, making the query and storage of functional function values faster and more accurate, but also considers the differences in data volumes among data providers through hierarchical partitioning, ensuring the protection of data privacy in data fusion calculations. In the system initialization stage, through the construction of address mapping, a foundation can be laid for subsequent database queries and the calculation of functional function values, enabling efficient data fusion calculations and the acquisition of functional function values without revealing the original data.
[0100] In a specific embodiment of the implementation method of an address mapping, the above - mentioned address mapping adopts meta - function: , for representation, where represents the data coding sequence, and the above - mentioned data coding sequence includes data codings, represents the data coding of data provider . The data provider holds secret data items, which are respectively , , represents the coding mapping between the above - mentioned secret data and the above - mentioned data coding.
[0101] Specifically, this embodiment further elaborates in detail the mathematical implementation method of the address mapping. By using a meta - function to specifically describe the mapping relationship between the data coding sequence and the database address, the construction of the database address and the storage of functional function values are realized.
[0102] The address mapping is represented by a meta - function , where represents the data coding sequence, and is the data coding corresponding to data provider . This meta - function converts the data coding sequence into a specific address in the database through specific calculation rules.
[0103] Data provider holds secret data items, which are respectively For each piece of secret data, the data provider uses an encoding mapping to convert it into a data encoding. The encoding mapping is a bijective mapping between the data provider and the secret data, ensuring a one-to-one correspondence between the secret data and the data encoding. Thus, without exposing the original data, the data encoding can participate in subsequent fusion calculations and address mapping processes.
[0104] In summary, it has been explained how to construct a database address based on a data encoding sequence and how the data encoding is generated from secret data through an encoding mapping. This tree-structure-based address mapping combines the randomness of the data encoding and the bijective property of the encoding mapping, not only ensuring the protection of data privacy but also optimizing the structure of the database and the storage of functional function values, improving the efficiency and security of multi-source data fusion calculations.
[0105] In an embodiment of determining the size of the database in this application, the above method further includes: in the above system initialization stage, according to the formula to determine the size of the above database, where N represents the size of the above database, represents the data scale of the i-th above data provider, , represents the number of the above data providers.
[0106] Specifically, the determination of the database size is carried out in the system initialization stage, and the size N of the database is determined according to the characteristics of data fusion calculations and the data scales of each data provider. The size of the database determines how many functional function values can be stored, and the number of functional function values is directly related to the types and quantities of data encoding sequences. Therefore, all possible combinations of data encoding sequences of all data providers need to be taken into account.
[0107] According to the formula to determine the size of the database, that is to say, the size of the database is determined by the product of the data scales of all data providers. Specifically, the size of the database is equal to the data scale of the first data provider multiplied by the data scale of the second data provider , the data scale of the third data provider ... until the data scale of the m-th data provider . For example, if there are 3 data providers , , , holding 10, 20, and 30 pieces of data respectively, then the size of the database is: 10×20×30 = 6000, which means the database needs to be able to store 6000 different functional function values to cover all combinations of data encoding sequences.
[0108] The rationality and efficiency of this calculation method are ensured by setting the database size as the product of the data scales of all data parties, which can ensure that the database can accommodate all possible corresponding relationships between data coding sequences and functional function values. Therefore, in subsequent database queries, the functional function value corresponding to the data coding sequence can be accurately found. At the same time, since the data is encoded and an address mapping is constructed during the system initialization phase, during the system operation phase, the plaintext data coding sequence is directly used for querying without performing privacy calculations again, significantly reducing the computational and communication overhead.
[0109] In another embodiment of the present application, the above-mentioned requester is further configured to construct an address mapping between the above-mentioned data coding sequence and the above-mentioned database based on the data scales of each of the above-mentioned data parties during the above-mentioned system initialization phase.
[0110] Specifically, the requester constructs a mapping relationship between the data coding sequence and the database address based on the data scales of each data party. This mapping process ensures that each data coding sequence can uniquely correspond to a specific address in the database, so that when it is necessary to query the functional function value corresponding to a certain data coding sequence, the corresponding position in the database can be quickly located. The design and construction of this mapping are crucial for achieving efficient data query and privacy protection.
[0111] During the process of constructing the address mapping, the size of the database is determined according to the data scales of each data party. The size of the database is set as the product of the data scales of all data parties. Setting the size of the database in this way ensures that the database can accommodate all possible data coding sequences and corresponding functional function values, meeting the requirements of data fusion calculation.
[0112] Based on the above settings, the following steps are performed during the system initialization phase to construct an address mapping between the data coding sequence and the database:
[0113] Determine the size of the database according to the data scales of each data party.
[0114] Construct an injective address mapping to ensure that each data coding sequence corresponds to a unique database address, and different data coding sequences correspond to different database addresses.
[0115] Adopt an algorithm or formula that is easy to calculate forward to implement the address mapping, so that the corresponding database address can be directly calculated according to the data coding sequence without complex post-processing or searching. The address mapping can be constructed based on a tree structure, dividing the database into different levels of intervals, and each level of interval corresponds to the data scale of a data party, thus forming a hierarchical and partitioned database structure.
[0116] Ensure that the inverse operation of the address mapping can be performed at the data side, that is, the data side can reverse-calculate the corresponding data coding sequence according to the database address, which is crucial for subsequent database query and decoding processes.
[0117] The significance of constructing the address mapping: Constructing the mapping between the data coding sequence and the database address not only provides a guide for the storage and retrieval of functional function values, but also ensures the security of data privacy by avoiding direct storage and processing of secret data. At the same time, this mapping mechanism reduces the computational and communication overhead during the system operation phase, improving the system's response speed and efficiency.
[0118] In a specific embodiment, the above address mapping is an injective mapping, and there is a one-to-one correspondence between the above data coding sequence and the address in the above database.
[0119] Specifically, in this embodiment, the property of the injective mapping is reflected in the address mapping, ensuring the one-to-one correspondence between the data coding sequence and the database address. In data fusion calculation, the data coding sequence is a sequence formed by arranging the coded data of each data side in order, and the database address is the specific location where the functional function value is stored in the database. The address mapping is injective, which means that different data coding sequences will generate different database addresses, that is, there is no situation where two different data coding sequences are mapped to the same database address.
[0120] Each data coding sequence can uniquely determine a database address through the address mapping function, and the corresponding data coding sequence can be calculated in reverse through the address mapping function.
[0121] The construction method of the address mapping can ensure that all possible data coding sequences have a unique address in the database corresponding to them, so that no matter how the data coding sequence changes, the accurate storage and rapid retrieval of the functional function value can be ensured.
[0122] In multi-source data fusion calculation, the injective mapping address mapping is the key to ensuring data privacy and calculation efficiency. Through the injective mapping, this application can ensure:
[0123] Data privacy protection: The one-to-one correspondence between the data coding sequence and the database address makes it impossible to directly infer the corresponding data coding sequence even if the database address is leaked, thus protecting the secret data privacy of the data side.
[0124] Computational and communication efficiency: Since each data coding sequence corresponds to a unique database address, when querying the database, the address can be directly calculated according to the data coding sequence, avoiding complex search and data matching processes, and significantly reducing computational and communication overhead.
[0125] By ensuring that the address mapping is injective, the present application can achieve efficient storage and retrieval of functional function values without compromising data privacy, thereby improving the performance of the entire data fusion computing system. This feature is the basis for building a secure, efficient, and practical multi-source data fusion computing system.
[0126] An embodiment of the present application also provides a multi-source data fusion computing system. Refer to Figure 2 , the above-mentioned multi-source data fusion computing system includes a data party, a demand party, and a privacy computing system. The above-mentioned privacy computing system includes multiple computing parties. Among them, in the system initialization stage, the above-mentioned demand party constructs an address mapping between the data coding sequence and the database based on the data of each above-mentioned data party; each above-mentioned data party is used to randomly encode the held secret data to obtain the corresponding data coding, and send the data coding of the above-mentioned secret data to the above-mentioned demand party; the above-mentioned demand party calculates the functional function value, and calculates the address of the functional function value corresponding to the data coding sequence in the above-mentioned database based on the received data coding of each above-mentioned data party and the above-mentioned address mapping, and then stores the above-mentioned functional function value in the corresponding address of the above-mentioned database; in the system operation stage, the above-mentioned data party is used to send the data coding of the current secret data to the above-mentioned demand party, so that the above-mentioned demand party can query the above-mentioned database according to the above-mentioned address mapping to obtain the functional function value corresponding to the current secret data of each above-mentioned data party, where the data coding of the above-mentioned current secret data is an element in the data coding sequence corresponding to the current address.
[0127] Among them, as long as the demand party knows the data scale of each data party , and the coding mapping of each data party is a one-to-one correspondence from the secret data to , the demand party can construct the address mapping. The construction of the address mapping by the demand party and the construction of the coding mapping by the data party can actually be carried out simultaneously;
[0128] Specifically, this embodiment describes a multi-source data fusion computing system including a data party, a demand party, and a privacy computing system, where the privacy computing system is composed of multiple computing parties.
[0129] In the system initialization stage, the multi-source data fusion computing system performs the following key operations:
[0130] Address mapping construction: The demand party constructs an address mapping between a data coding sequence and a database based on the data scale of all data parties, that is, the number of secret data held by each data party. This address mapping is injective, ensuring a one-to-one correspondence between the data coding sequence and the database address, thus providing an accurate address location for subsequent storage of functional function values.
[0131] Data Encoding and Transmission: The data provider randomly encodes the secret data it holds to generate data encodings. This encoding process not only provides an alternative representation of the data but also enhances the protection of data privacy through randomness. Then, the data provider sends the data encodings to the requester for subsequent address mapping and functional function value calculation.
[0132] Functional Function Value Calculation and Storage: After the requester receives the data encoding sequences from all data providers, it calculates the functional function values based on these data encodings and the constructed address mapping. The key to the functional function values lies in their execution based on privacy computing technology, ensuring privacy protection during the data processing. The calculated functional function values are then stored at specific addresses in the database, which are obtained through the address mapping function and correspond one-to-one with the original data encoding sequences.
[0133] During the system operation phase, which is the phase when the multi-source data fusion computing system actually processes data fusion requests, the following operations are involved:
[0134] Data Encoding Transmission: The data provider sends the data encodings of the current secret data to the requester. This data encoding is the encoded representation of the current data and depends on the specific content and changes of the current data.
[0135] Functional Function Value Query: After receiving the data encodings, the requester uses the address mapping function to calculate the address of the functional function value corresponding to the data encoding sequence in the database. Based on this address, the requester can directly retrieve the functional function value corresponding to the data encoding sequence from the database without having to perform complex privacy computing processes again.
[0136] Real-time Data Processing and Response: The focus of this phase is to respond to real-time or updated data requests. Through the fast query mechanism of the database, the system can quickly provide the functional function values corresponding to the data encoding sequences, thus achieving the real-time and efficiency of data fusion computing.
[0137] By constructing an address mapping between the data encoding sequence and the database during the system initialization phase, the core during the system operation phase is to utilize the address mapping and database query to quickly and directly obtain the functional function values corresponding to the data encodings of the current secret data of the data provider without having to perform real-time privacy computing. This greatly improves the system's response speed, reduces communication overhead, and protects the data privacy of the data provider throughout the process. This method provides an efficient and privacy-secure solution for real-time data fusion scenarios. In summary, the multi-source data fusion computing system of this embodiment can meet the real-time and data privacy requirements in multi-source data fusion computing and provide a practical and efficient solution.
[0138] In a specific embodiment, during the above system initialization phase, each of the above data parties is further configured to encrypt the above secret data to obtain ciphertext and send the ciphertext to each of the above computing parties; the multiple above computing parties are configured to receive the ciphertext input by each of the above data parties, execute a privacy computing task to obtain an intermediate result, and send the intermediate result to the above requester, so that the above requester can calculate the above functional function value based on the above intermediate result.
[0139] Specifically, during the system initialization phase, the data party not only needs to perform random encoding on the secret data to generate data codes, but also needs to further encrypt these data codes to generate ciphertext. This encryption process is usually implemented using secret sharing, homomorphic encryption, or other privacy protection technologies to ensure that even if the data is intercepted by a third party during transmission and calculation, the original content of the data cannot be directly interpreted. The data party sends the encrypted ciphertext to each computing party in the privacy computing system as the input of the privacy computing task.
[0140] After receiving the ciphertext input by each data party, the computing party executes a privacy computing task to calculate the intermediate result of the functional function. Privacy computing technologies, such as secure multi-party computing (MPC), homomorphic encryption (HE), or trusted execution environment (TEE), can perform calculations without decrypting the data to protect data privacy. The communication and calculation between computing parties follow specific privacy computing protocols to ensure the security of the calculation process and the accuracy of the results.
[0141] After the computing party completes the privacy computing task, it will send the calculated intermediate result to the requester. Based on these intermediate results and the address mapping and database structure constructed during the system initialization phase, the requester performs subsequent calculations and processing to finally calculate the functional function value. The functional function value will then be stored at a specific address in the database, which is calculated based on the address mapping function and corresponds one-to-one with the data coding sequence.
[0142] In summary, through data encryption and privacy computing technologies, it is possible to accurately calculate the functional function value and store it in the database without revealing data privacy, providing a basis for fast data query during the system operation phase. This mechanism ensures the security, efficiency, and practicality of the system. By combining data coding, encryption, and privacy computing, it is possible to achieve privacy protection and efficient fusion calculation of sensitive multi-source data, meeting the dual requirements of data privacy protection and data fusion calculation.
[0143] In another specific embodiment, the above privacy computing system further includes other relevant parties, and the above other relevant parties include a trusted third party that assists in the execution of the privacy computing task.
[0144] Specifically, refer to Figure 3, in this embodiment, the architecture of the multi-source data fusion computing system is further extended, and a trusted third party is introduced as other relevant parties to assist in the execution of privacy computing tasks. This mechanism aims to enhance the security, reliability, and performance of the system, especially when dealing with complex privacy computing scenarios.
[0145] A trusted third party is introduced to assist in the execution of secure multi-party computing, homomorphic encryption, or other privacy computing tasks. The trusted third party can provide additional trust mechanisms, assist in the initialization of protocols, coordinate the computing process, handle disputes, or provide additional computing resources. Since privacy computing tasks may involve complex cryptographic operations and protocols, the participation of the trusted third party can simplify the execution process of these tasks, improve computing efficiency, and ensure data security and privacy protection.
[0146] The ways in which the trusted third party participates can be diverse, depending on the specific design and requirements of the privacy computing system. For example, the trusted third party can assist in generating and distributing keys to ensure the secure distribution of keys; in some MPC protocols, the trusted third party can distribute preprocessed shared information to reduce communication and computing overhead in the online phase; during the execution of the protocol, the trusted third party can act as a neutral coordinator to ensure that all participating parties correctly execute the computing tasks according to the protocol; in case of disputes or failures, the trusted third party can intervene to help resolve disputes or restore the normal operation of the system.
[0147] As part of the privacy computing system, the introduction of a trusted third party can bring the following advantages:
[0148] Simplify the protocol: The trusted third party can undertake part of the computing tasks or protocol initialization work, thereby reducing the computing overhead of the participating parties and simplifying the complexity of the protocol.
[0149] Enhance security: The trusted third party can provide an additional layer of trust to ensure fairness and security in the data processing process and prevent attacks by malicious participating parties.
[0150] Dispute resolution: During the multi-source data fusion computing process, if data disputes or inconsistencies in protocol execution occur, the trusted third party can intervene and provide a dispute resolution mechanism.
[0151] Resource coordination: The trusted third party can coordinate the resource allocation of the computing parties to ensure the efficient execution of computing tasks, especially in resource-constrained environments.
[0152] In summary, a trusted third party is introduced into the multi-source data fusion computing system as other relevant parties to assist in the execution of privacy computing tasks. This design enhances the security, reliability, and performance of the system, especially when dealing with complex privacy protection computing scenarios. Through the participation of the trusted third party, the privacy computing protocol can be simplified, the computing efficiency can be improved, and at the same time, the fairness and security in the data processing process can be ensured, providing stronger privacy computing protection for multi-source data fusion computing.
[0153] In an embodiment of the present application, the requester constructs an address mapping between the data coding sequence and the database based on the data of each of the above data parties during the system initialization phase. The above address mapping is a mapping constructed based on a tree-structured database. For specific explanations, please refer to the embodiment of address mapping construction and the embodiment of the implementation method of address mapping in the method embodiment.
[0154] In an embodiment of the present application, during the system initialization phase, the above requester constructs an address mapping between the data coding sequence and the database based on the data of each of the above data parties, including: during the above system initialization phase, the above requester constructs the address mapping between the data coding sequence and the above database based on the data scale of each of the above data parties.
[0155] Specifically, the requester constructs a mapping relationship between the data coding sequence and the database address based on the data scale of each data party. This mapping process ensures that each data coding sequence can uniquely correspond to a specific address in the database, so that when it is necessary to query the function value corresponding to a certain data coding sequence, the corresponding position in the database can be quickly located. The design and construction of this mapping are crucial for realizing efficient data query and privacy protection.
[0156] During the process of constructing the address mapping, the size of the database is determined according to the data scale of each data party. The size of the database is set to the product of the data scales of all data parties. Setting the size of the database in this way ensures that the database can accommodate all possible data coding sequences and corresponding function values, meeting the requirements of data fusion computing.
[0157] Based on the above settings, the following steps are performed during the system initialization phase to construct the address mapping between the data coding sequence and the database:
[0158] Determine the size of the database according to the data scale of each data party.
[0159] Construct an injective address mapping to ensure that each data coding sequence corresponds to a unique database address, and different data coding sequences correspond to different database addresses.
[0160] Implement address mapping using an algorithm or formula that is easy to calculate forward, so that the corresponding database address can be directly calculated based on the data coding sequence without complex post - processing or searching. The address mapping can be constructed based on a tree structure, dividing the database into intervals at different levels, and each level of interval corresponds to the data scale of a data party, thus forming a hierarchical partitioned database structure.
[0161] Ensure that the inverse operation of the address mapping can be performed at the data party, that is, the data party can calculate the corresponding data coding sequence inversely based on the database address, which is crucial for subsequent database query and decoding processes.
[0162] The significance of constructing the address mapping. Constructing the mapping between the data coding sequence and the database address not only provides a guide for the storage and retrieval of functional function values, but also ensures data privacy security by avoiding direct storage and processing of secret data. At the same time, this mapping mechanism reduces the computational and communication overhead during the system operation phase, improving the system's response speed and efficiency.
[0163] In an embodiment of the present application, the above - mentioned requester is also used to determine the size of the above - mentioned database according to the formula during the above - mentioned system initialization phase. For specific explanations, please refer to the embodiment of determining the database size in the method embodiment.
[0164] In an embodiment of the present application, the above - mentioned requester is also used to determine the above - mentioned data party, the data - party - related parameters, each of the above - mentioned computing parties participating in the calculation, and the privacy computing protocol used during the above - mentioned system initialization phase based on the scenario requirements.
[0165] Specifically, during the system initialization phase, the requester first needs to determine the data party based on the scenario requirements, that is, determine which data sources will participate in the data fusion calculation. Scenario requirements come from various application fields, such as financial analysis, healthcare, intelligent transportation, etc. Each scenario may require fusing different types of source data. After determining the data party, the requester also needs to determine the parameters related to the data party, including the number of data parties, the scale of data held by each data party, the data type, etc. These parameters are crucial for subsequent construction of the address mapping, determination of the database size, and execution of the computing tasks.
[0166] During the system initialization phase, the requester also needs to determine each computing party participating in the calculation, that is, decide which computing nodes will execute the privacy computing tasks. The selection of computing parties is based on factors such as computing resources, geographical location, computing power, and reputation. Since privacy computing tasks may involve a large amount of computing, selecting high - performance and trustworthy computing parties is crucial for the overall efficiency and data security of the system.
[0167] In the system initialization phase, the requester is responsible for determining the privacy computing protocol used in the system, which is the core of privacy protection. Privacy computing protocols include technologies such as Homomorphic Encryption (HE), Secure Multi-Party Computation (MPC), Trusted Execution Environment (TEE), or Federated Learning (FL). The requester needs to select the most suitable privacy computing protocol based on the scenario requirements, the parameters of the data provider, and the capabilities of the computing party to ensure data privacy while meeting the computational performance requirements.
[0168] In summary, in the system initialization phase, the responsibilities of the requester include determining the data provider and the data provider-related parameters, selecting the computing parties participating in the computation, and determining the privacy computing protocol used. This series of decisions constitutes the architectural foundation of the multi-source data fusion computing system and directly affects the system's privacy protection capabilities, computational efficiency, and data processing capabilities. By reasonably configuring these parameters, the requester can build an efficient computing system that not only meets the application requirements but also protects data privacy. This mechanism ensures the security and usability of the system during the data fusion computation process and provides the necessary preparation for subsequent data processing and the calculation of functional function values.
[0169] In an embodiment of the present application, the above address mapping is injective, and the above data coding sequence has a one-to-one correspondence with the address in the above database.
[0170] Specifically, in this embodiment, the property of the injective mapping is reflected in the address mapping, ensuring the one-to-one correspondence between the data coding sequence and the database address. In data fusion computation, the data coding sequence is a sequence formed by arranging the encoded data of each data provider in order, and the database address is the specific location where the functional function value is stored in the database. The address mapping is injective, which means that different data coding sequences will generate different database addresses, that is, there is no situation where two different data coding sequences are mapped to the same database address.
[0171] Each data coding sequence can uniquely determine a database address through the address mapping function, and the corresponding data coding sequence can be calculated inversely through the address mapping function.
[0172] The construction method of the address mapping can ensure that all possible data coding sequences have a unique address in the database corresponding to them, so that no matter how the data coding sequence changes, the accurate storage and rapid retrieval of the functional function value can be ensured.
[0173] In multi-source data fusion computing, the address mapping of injective mapping is the key to ensuring data privacy and computing efficiency. Through injective mapping, this application can ensure that:
[0174] Data privacy protection: The one-to-one correspondence between the data coding sequence and the database address ensures that even if the database address is leaked, it is impossible to directly infer the corresponding data coding sequence, thus protecting the secret data privacy of the data party.
[0175] Computing and communication efficiency: Since each data coding sequence corresponds to a unique database address, when querying the database, the address can be directly calculated based on the data coding sequence, avoiding complex search and data matching processes, and significantly reducing computing and communication overhead.
[0176] By ensuring that the address mapping is injective, this application can achieve efficient storage and retrieval of functional function values without affecting data privacy, thereby improving the performance of the entire data fusion computing system. This feature is the basis for building a secure, efficient, and practical multi-source data fusion computing system.
[0177] In an embodiment of this application, during the above system initialization phase, the addresses of the above database are traversed, and the data coding sequence corresponding to the traversed address is determined based on the above address mapping; the above data party is also used to traverse the data coding sequence in a multiple-loop manner based on the data coding sets of the above data parties during the above system initialization phase. For the specific explanation and the specific implementation manner of traversing the data coding sequence in a multiple-loop manner, refer to the embodiment of determining the data coding in the method embodiment, which will not be elaborated here.
[0178] In an embodiment of this application, the data party is also used to secretly store the above coding mapping between the held secret data and the above data coding, where the above coding mapping is a bijective mapping; in the case where the above data coding is known and the above secret data is unknown, calculate the inverse of the above coding mapping with respect to the above data coding to obtain the above secret data corresponding to the above data coding.
[0179] Specifically, the data party secretly stores the coding mapping between the held secret data and the data coding, and this coding mapping is a bijective mapping. That is to say, each secret data corresponds to a specific data coding, and this correspondence is unique, that is, no two different secret data will be mapped to the same data coding. Similarly, each data coding also only corresponds to a specific secret data, ensuring the uniqueness and reversibility of the data coding, thereby increasing the security of the system.
[0180] During the data fusion calculation process, data encoding is used to replace the secret data for secure multi-party calculation or storage in the database. Since the mapping relationship between the data encoding and the secret data is confidential, it ensures that even if the data encoding is known to other parties, the original secret data cannot be inferred reversely, thus protecting data privacy.
[0181] Inverse operation of the encoding mapping In the case where the data encoding is known and the secret data is unknown, the data party can utilize the inverse operation of the encoding mapping to calculate the secret data corresponding to the known data encoding. This is because the encoding mapping is a bijective function, ensuring the existence and uniqueness of its inverse operation. The inverse operation allows the data party to convert the data encoding back to the original secret data when needed without the participation of other parties.
[0182] In summary, this embodiment emphasizes the confidentiality and reversibility of the encoding mapping, that is, the data party needs to secretly store the encoding mapping and can calculate the corresponding secret data through the inverse operation when the data encoding is known. This feature is the core of realizing secure data fusion calculation, ensuring the protection of data privacy and also providing the data party with a means of data decoding to restore the original data when needed. Through this encoding mapping mechanism, the effective fusion and calculation of data can be achieved without revealing the original secret data, enhancing the security and efficiency of data fusion calculation.
[0183] To enable those skilled in the art to more clearly understand the technical solution of this application, the implementation process of the multi-source data fusion calculation of this application will be described in detail below in combination with specific embodiments. Specifically, the multi-source data fusion calculation system involves five parts: system setting, data encoding, address mapping construction, database construction, and database query.
[0184] 1. System setting
[0185] This stage is completed by the requester, mainly determining the data parties and related parameters based on the scenario requirements, including the number m of data parties, the data scale of each data party, and the scale of the database, etc.; determining each participant in the privacy calculation system, including the calculation participant and the trusted third party (optional); the security model, such as honest majority / dishonest majority, semi-honest security / malicious security, etc., and the privacy calculation protocol used; determining the functional function based on the data of each data party , where is the data held by each data party.
[0186] 2. Data encoding
[0187] Each data party performs random encoding on the held secret data. Taking data party as an example, without loss of generality, let data party hold a piece of data Let , then Select to a random bijective mapping on (referred to as the encoding mapping), and the encoding of the data is . Among them, the symbol means traverses from from to , and the symbol represents the set . The corresponding relationship (encoding mapping) between the data of each party and the encoding is secretly kept by each data party.
[0188] 3. Address mapping construction
[0189] The requester determines the size of the database based on the data scale of each data party and constructs an address mapping from the data encoding sequence to the database. The above address mapping must meet the following conditions: First, the address mapping is injective, that is, for any data encoding sequence, there is a corresponding address in the database, and different data encoding sequences correspond to different addresses; Second, the forward calculation of the address mapping is easy. Optionally, this application provides an address mapping constructed based on a tree structure, which is as follows.
[0190] Regard the database as a linear partition table of level, where the first level divides the database into primary intervals from left to right, and each primary interval contains addresses; the second level divides each primary interval into secondary intervals from left to right, and each secondary interval contains addresses; the third level divides each secondary interval into tertiary intervals from left to right, and each tertiary interval contains addresses; until the th level divides each level interval into addresses. Equivalently, as shown in Figure 4 , the database can be regarded as the sequential arrangement of the leaf nodes of a multi-way tree with primary nodes at the root node, each primary node has secondary nodes, each secondary node has tertiary nodes, until each level node has leaf nodes.
[0191] Based on the above idea, the requester constructs a reversible mapping (referred to as address mapping) from the data coding sequence to the set where represents the product of the data scales of each data party, that is, the size of the database. Specifically, without loss of generality, let the data coding sequences of the data parties be , then the function value corresponding to this coding sequence is located in the th first-level interval of the database, the th second-level interval of the current first-level interval,..., until the th-level interval's th address. Thus, the above address mapping can be a -ary function
[0192] where is the data coding of data party .
[0193] 4. The first method for constructing the database
[0194] Database construction means that for each address of the partitioned table (each leaf node of the multi-way tree), each data party reversely calculates the corresponding data coding sequence based on the address mapping, and reversely calculates the corresponding secret data based on the coding mapping, and then calculates the function value of the functional function with respect to the secret data of each party based on privacy computing technology, and stores the above functional function value into the current address, so as to obtain the database of the functional function value with respect to the data coding sequence while ensuring the data privacy of each party.
[0195] As Figure 5 shown, for each address of the database, each data party calculates the data coding sequence based on the address mapping, and inputs the corresponding data into the privacy computing system to calculate the functional function value, and then stores the functional function value into the current address of the database. Specifically, each data party executes steps 4.1 - 4.4 to calculate the data coding sequence corresponding to the current address, that is, to calculate the inverse of the address mapping.
[0196] Step 4.1: Each data party calculates .
[0197] Step 4.2: For , data party loops through steps 4.2.1 - 4.2.3 to calculate the data coding .
[0198] Step 4.2.1: If , calculate .
[0199] Step 4.2.2: If , calculate .
[0200] Step 4.2.3: Calculate .
[0201] Among them, the symbols and respectively represent the meanings of integer division and non - integer division. The symbol represents the modulo operation. For example, represents divided by the remainder.
[0202] Optionally, each data party can also execute Step 4.3 to calculate the data encoding sequence.
[0203] Step 4.3: Let . For , the data party executes Steps 4.3.1 - 4.3.3 to calculate the data encoding .
[0204] Step 4.3.1: If , calculate , and no longer execute Steps 4.3.2 - 4.3.3 and Step 4.4. Among them, (the same below), the expression represents integer division of .
[0205] Step 4.3.2: If , calculate . Among them, the expression represents non - integer division of , the expression represents the floor function, that is, the largest integer less than or equal to .
[0206] Step 4.3.3: Calculate .
[0207] Step 4.4: The data party calculates .
[0208] Step 4.5: For , the data party calculates the inverse of the encoding mapping with respect to the data encoding , that is, calculates .
[0209] Step 4.6: For , the data party For data Process it through cryptographic techniques such as secret sharing, and input the processing result into the privacy computing system.
[0210] Step 4.7: Each computing party of the privacy computing system calculates the functional function based on the input of each data party, and sends the calculation result to the requester.
[0211] Step 4.8: The requester calculates the functional function value based on the output of each computing party 。
[0212] Step 4.9: The requester stores the functional function value into the th address of the database.
[0213] 5. The second method of constructing the database
[0214] As Figure 6 shown, for the data coding sequence , each data party sends the data coding to the requester in plaintext form or after being encrypted by traditional methods. The requester calculates the corresponding address of the data coding sequence based on the address mapping. At the same time, each data party inputs the corresponding data into the privacy computing system to calculate the functional function. The requester calculates the functional function value based on the output result of the privacy computing system and stores the above functional function value into the current address.
[0215] Specifically, for , data party executes steps 5.1 - 5.3.
[0216] Step 5.1: Data party sends the data coding to the requester in plaintext form or after being encrypted by traditional methods.
[0217] Step 5.2: Data party calculates the inverse of the coding mapping with respect to the data coding , that is, calculates .
[0218] Step 5.3: Data party processes the data through cryptographic techniques such as secret sharing, and sends the processing result to each computing party of the privacy computing system.
[0219] Step 5.4: Each computing party of the privacy computing system calculates the functional function based on the input of each data party.
[0220] The requester executes steps 5.5 - 5.7.
[0221] Step 5.5: The requester calculates the address of the address mapping with respect to the data coding sequence based on the data coding of all parties. 。
[0222] Step 5.6: The requester calculates the functional function value based on the output result of the privacy computing system. 。
[0223] Step 5.7: The requester stores in the th address of the database.
[0224] 6. Database Query
[0225] Step 6.1: For , the data provider sends the coding of the current data to the requester in plaintext or after being encrypted by traditional methods.
[0226] Step 6.2: The requester calculates the address of the address mapping with respect to the data coding sequence , then the element at the th address of the database is the corresponding functional function value 。
[0227] It should be noted that in the embodiments of this application, the " " in is unordered, indicating that the value range of is the " " in is ordered, indicating that traverses from in sequence, here it refers to the data provider executing from to of the loop body. In similar expressions in the text, for example, 、 the " " and " " in 、 have the same meaning as the " " and " " in
[0228] The above-mentioned multi-source data fusion computing system can be applied to the transportation field. In the transportation field, especially the road traffic management under severe weather conditions, the multi-source data fusion computing system can significantly improve the overall traffic management and travel service efficiency through the application of privacy computing technology, while ensuring data privacy security. In related technologies, although the combination of meteorological warning and speed control systems has improved the road traffic efficiency in severe weather and guaranteed travel needs, the traffic, meteorological and other data involving quasi-all-weather travel are subject to privacy security restrictions. At present, there is no unified data service, and the overall efficiency and accuracy need to be improved. But the overall efficiency needs to be improved. Privacy computing has become a research hotspot at home and abroad. The core technologies include secure multi-party computing, homomorphic encryption, trusted execution environment, etc. Its applications are mainly concentrated in the fields of finance, Internet services, etc., but its application in the transportation field needs further exploration. By applying the multi-source data fusion computing system of the present application, traffic and meteorological data from multiple data sources can be integrated in real time, such as vehicle location, speed, direction and other information, which can be used to evaluate traffic flow and road conditions; real-time weather data, such as temperature, humidity, wind speed, rainfall, etc., can be used to predict weather changes and road slipperiness; video data from road monitoring cameras can be used to monitor road conditions and traffic incidents in real time. The system performs fusion calculations on these data under privacy protection, extracts valuable information for road safety and efficiency, and provides support for traffic management decisions. For example, the system can calculate real-time visibility prediction maps, road slipperiness distribution maps, etc., to help traffic management departments formulate more accurate speed limits and road traffic strategies, improve traffic efficiency, and ensure travel safety.
[0229] In summary, the application of multi-source data fusion computing system in the transportation field not only improves the efficiency and safety of road traffic management in severe weather, but also promotes data sharing and service unification under data privacy protection, providing more comprehensive, accurate and safer data support for traffic management and travel.
[0230] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to some solutions, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the above methods of each embodiment of the present application.
[0231] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any of the above method embodiments when running.
[0232] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0233] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.
[0234] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device. Wherein, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0235] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary embodiments, and details are not described herein again.
[0236] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.
[0237] The above has introduced in detail a multi-source data fusion calculation method, a multi-source data fusion calculation system, a computer-readable storage medium, and an electronic device provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A multi-source data fusion calculation method, characterized in that: In the system initialization phase, an address mapping between a data coding sequence and a database is constructed, and a target address is calculated based on the data coding of the secret data of each data party and the address mapping, and a function value obtained by performing a privacy computing task based on the secret data of each data party is stored in the target address, wherein the function value corresponds to a data coding sequence, and the data coding sequence is an ordered vector composed of multiple data codes of all data parties; During the system operation stage, based on the address mapping and the data encoding of each data party, the address of the function value corresponding to the current secret data of each data party in the database is calculated, and the database is queried according to the address to obtain the function value, wherein the data encoding of the current secret data is an element in the data encoding sequence corresponding to the current address.
2. The multi-source data fusion calculation method according to claim 1 is characterized in that: Before storing the function value obtained by performing the privacy computing task based on the secret data of each data party in the target address, the method further includes: Encrypting the secret data held by each of the data parties to obtain a ciphertext; Execute the privacy computing task based on the ciphertext obtained by encryption to obtain an intermediate result; The functional function value is calculated according to the intermediate result.
3. The multi-source data fusion calculation method according to claim 1 is characterized in that: The method further comprises: In the system initialization phase, traversing the addresses of the database, and determining the data encoding sequence corresponding to the traversed addresses based on the address mapping; or, During the system initialization phase, based on the data coding sets of each of the data cubes, the data coding sequence is traversed in a multiple-cycle manner.
4. The multi-source data fusion calculation method according to claim 3 is characterized in that: Traversing the address of the database and determining the data encoding sequence corresponding to the traversed address based on the address mapping, comprising: Calculate the address parameter f m , where f m =k+N2…N m +…+N m , k = 1..N, k represents the address of the database, m represents the number of data cubes, N2 represents the amount of data in the second data cube D2, N m Represents the mth data cube D m The amount of data; Set i = 1...m, j = m..i, and determine And judge N j Is it divisible by f? j , determine the jth data code in the data code sequence according to whether it can be divided integerly, Γ j Represents the jth data code in the data code sequence.
5. The multi-source data fusion calculation method according to claim 4 is characterized in that: Determining the jth data code in the data code sequence according to the result of whether it is divisible, including: The first step of determination: If N j |f j , calculate Γ j =N j ; Second judgment step: If Calculate Γ j =f j mod N j ; Calculation steps: Among them, | means divisible by integers, Indicates that it is not divisible, and mod indicates the modular operation, such as amod b means the remainder when a is divided by b; Data Cube D i The first determination step, the second determination step and the calculation step are executed cyclically to obtain the data code Γ i , where i=1...m.
6. The multi-source data fusion calculation method according to claim 3 is characterized in that: Traversing the address of the database and determining the data encoding sequence corresponding to the traversed address based on the address mapping, comprising: Set f1 = k, M j =N j+1 …N m , f j+1 =f j -γ j ·M j , For i=1...m-1, j=1..i, k represents the address of the database, m represents the number of the data cube, N j+1 Indicates the j+1th data cube D j+1 The amount of data, N m Represents the mth data cube D m The amount of data, data D i Calculate M j Is it divisible by f? j , determine the data code Γ in the data coding sequence according to whether it can be divided integerly i , where the floor operation is M j Cannot divide f j carried out under the circumstances; Data Cube D m Calculate Γ m =f m .
7. The multi-source data fusion calculation method according to claim 6 is characterized in that: Determine the data code Γ in the data code sequence according to whether the result is divisible i , data D m Calculate Γ m =f m ,include: Execution steps: Let f1 = k, for i = 1...m-1, j = 1..i, data D i Execute the first calculation sub-step, the second calculation sub-step and the third calculation sub-step to calculate the data code Γ i ; The first calculation sub-step: if M j |f j ,calculate Γ ι =N l , l=j+1, ..., m, and the second calculation sub-step, the third calculation sub-step and the calculation step are no longer executed, wherein M j =N j+1 …N m , M j |f j Indicates M j Divisibility f j , Γ j represents the jth data code in the data code sequence, Γ l represents the lth data code in the data code sequence; The second calculation sub-step: if calculate Γ j =γ j +1, among which, Indicates M j Does not divide f j , Indicates rounding down; The third calculation sub-step: calculate f j+1 =f j -γ j ·M j ; Calculation steps: Data cube D m Calculate Γ m =f m .
8. The multi-source data fusion calculation method according to claim 3 is characterized in that: In the system initialization phase, based on the data encoding set of each data entity, the data encoding sequence is traversed in a multiple-cycle manner, including: Data cube D1 traverses data code Γ1 from 1 to N1; for each one-dimensional sequence Γ1 traversed, data cube D2 traverses data code Γ2 from 1 to N2; for each two-dimensional sequence (Γ1, Γ2) traversed, data cube D3 traverses data code Γ3 from 1 to N3; until, for each m-1 dimensional sequence (Γ1, Γ2, ..., Γ m-1 ), data D m From 1 to N m Traversal Data Encoding Γ m , get the data encoding sequence (Γ1, Γ2, ..., Γ m ), where data D i From 1 to N i Traversal Data Encoding Γ i , indicating data D i Γ i Assign values 1...N in sequence i , where i = 1...m, N i Indicates data D i The amount of data, m represents the number of data parties; In the system initialization phase, based on the data encoding set of each data cube, the data encoding sequence is traversed in a multiple-cycle manner, and then the method further includes: In the system initialization phase, according to The database address corresponding to the data encoding sequence is calculated, where: Denotes the address mapping, Γ1, Γ2, ..., Γ m Represents the data encoding of each of the data entities.
9. The multi-source data fusion calculation method according to claim 3, characterized in that: Traversing the address of the database and determining the data encoding sequence corresponding to the traversed address based on the address mapping, including: traversing the address of the database, calculating the inverse of the address with respect to the address mapping to obtain the data encoding sequence; The method also includes: setting a data coding set, and randomly selecting a data coding from the data coding set to encode any secret data of the data party, wherein the data coding corresponds to the secret data one-to-one, and a coding mapping is used to characterize the mapping between the secret data and the data coding.
10. The multi-source data fusion calculation method according to claim 9, characterized in that: The method further comprises: secretly preserving the encoding mapping between the held secret data and the data encoding, wherein the encoding mapping is a bijection; The method further includes: when the data encoding is known and the secret data is unknown, calculating the inverse of the encoding mapping with respect to the data encoding to obtain the secret data corresponding to the data encoding.
11. The multi-source data fusion calculation method according to claim 1, characterized in that: The address mapping is a mapping constructed based on a tree-structured database, wherein the tree-structured database is an m-level linear partition table, the first level divides the database into N1 first-level intervals, and the i+1th level divides each i-level interval into N i+1 There are i+1 level intervals, 1≤i≤m-1, and each i-level interval contains N i+1 ...N m The m-level interval corresponds to the address of the database, m represents the number of data cubes, N i+1 Indicates the i+1th data cube D i+1 The amount of data, N m Represents the mth data cube D m The amount of data.
12. The multi-source data fusion calculation method according to claim 11, characterized in that: The address mapping uses an m-ary function: Represented by, where (Γ1, Γ2, ..., Γ m ) represents a data coding sequence, wherein the data coding sequence includes m data codes, Indicates data D i Data encoding, data square D i Hold N i The secret data are Γ ij =σ i (X ij )∈[N i ],σ i A coding mapping between the secret data and the data coding is represented.
13. The multi-source data fusion calculation method according to claim 1, characterized in that: The method further comprises: in the system initialization stage, according to the formula Determine the size of the database, where N represents the size of the database, N i represents the data size of the i-th data cube, i=1...m, and m represents the number of the data cubes.
14. The multi-source data fusion calculation method according to claim 1, characterized in that: During the system initialization phase, the address mapping between the data encoding sequence and the database is constructed, including: During the system initialization phase, an address mapping between the data encoding sequence and the database is constructed based on the data size of each data cube.
15. The multi-source data fusion calculation method according to any one of claims 1 to 14, characterized in that: The address mapping is injective, and the data encoding sequence and the address in the database are in a one-to-one correspondence.
16. A multi-source data fusion computing system, characterized in that: The multi-source data fusion computing system includes a data party, a demand party and a privacy computing system. The privacy computing system includes multiple computing parties, wherein: In the system initialization stage, the demand party constructs an address mapping between a data coding sequence and a database based on the data of each of the data parties; each of the data parties is used to randomly encode the secret data it holds to obtain a corresponding data code, and send the data code of the secret data to the demand party; the demand party calculates a functional function value, and calculates the address of the functional function value corresponding to the data coding sequence in the database based on the received data codes of each of the data parties and the address mapping, and then stores the functional function value in the corresponding address of the database; During the system operation phase, the data party is used to send the data encoding of the current secret data to the demand party, so that the demand party can query the database according to the address mapping to obtain the functional function value corresponding to the current secret data of each data party.
17. The multi-source data fusion computing system according to claim 16, characterized in that: In the system initialization phase, each of the data parties is further configured to encrypt the secret data to obtain a ciphertext, and send the ciphertext to each of the computing parties; The multiple computing parties are used to receive the ciphertext input by each data party, and perform privacy computing tasks to obtain intermediate results and send them to the demand party, so that the demand party can calculate the functional function value based on the intermediate results.
18. The multi-source data fusion computing system according to claim 17, characterized in that: The privacy computing system also includes other related parties, including a trusted third party that assists in the execution of privacy computing tasks.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the multi-source data fusion calculation method according to any one of claims 1 to 15.
20. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the multi-source data fusion calculation method according to any one of claims 1 to 15 are implemented.
Citation Information
Patent Citations
Multi-party fusion computing system, multi-party fusion computing method and readable storage medium
CN114944935A
Multi-source heterogeneous data fusion storage system
CN116992103A
Cited By
Data center PMO supervision knowledge graph construction and multi-source data fusion method and system
CN121365717A