Data alignment method and device, electronic equipment and readable storage medium
Patent Information
- Application Number
- CN202210428731.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-22
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2042-04-22
AI Technical Summary
由于隐私计算技术比较前沿,企业内部相对保持谨慎,在前期阶段,如果各家企业采用全量数据进行对齐,则需要进行亿级别的数据计算,效率低下,而根据信息安全最小化原则,以某个区域的数据进行对齐,当数据本身具有区域属性的情况下,则会存在区域信息的泄漏风险
[0041]本发明实施例提供的数据对齐方法、装置、电子设备和可读存储介质,在获取参与数据对齐的对象数据集后,根据第二区域标签或第一区域标签,对对象数据集进行划分,得到至少一个目标数据集,在得到目标数据集后,若各目标数据集的目标区域标签包括第二区域标签,对各目标数据集中第一区域标签与各目标数据集对应的目标区域标签未匹配的对象数据进行过滤,若各目标数据集的目标区域标签包括第一区域标签,则对各目标数据集中对应对象的第二区域标签与各目标数据集对应的目标区域标签未匹配的对象数据进行过滤,如此,使得过滤后的目标数据集中对象数据所属的区域与对象数据对应的对象的区域是相同的,在根据过滤后的目标数据集进行数据对齐时,避免了区域信息的泄露,又可保证数据对齐效率。
Smart Images

Figure CN116975366B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to a data alignment method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] With the introduction of data security laws and other related policies, companies need to align their data in a secure and compliant manner when collaborating on data projects.
[0003] Currently, the industry is using privacy-preserving computing technology to align data among companies. Because privacy-preserving computing is a relatively cutting-edge technology, companies are being cautious. In the early stages, if companies were to align their entire datasets, it would require computations on the scale of hundreds of millions of data points, which is inefficient. Furthermore, based on the principle of minimizing information security, aligning data based on a specific region carries the risk of leakage of regional information, especially if the data itself has regional attributes. Summary of the Invention
[0004] Based on the above research, the present invention provides a data alignment method, apparatus, electronic device and readable storage medium, which avoids the leakage of regional information and ensures data alignment efficiency.
[0005] Embodiments of the present invention can be implemented in the following ways:
[0006] In a first aspect, embodiments of the present invention provide a data alignment method, the method comprising:
[0007] Obtain the object dataset for data alignment; the object dataset includes object data, a first region label to which the object data belongs, and a second region label of the object corresponding to the object data.
[0008] The object dataset is divided according to the second region label or the first region label to obtain at least one target dataset; each target dataset has a target region label; the target region label includes at least one second region label or includes at least one first region label.
[0009] If the target region label of each target dataset includes a second region label, filter out object data in each target dataset whose first region label does not match the target region label corresponding to each target dataset; if the target region label of each target dataset includes a first region label, filter out object data in each target dataset whose second region label of the corresponding object does not match the target region label corresponding to each target dataset.
[0010] Based on the filtered target datasets, data alignment is performed.
[0011] In an optional implementation, the step of dividing the object dataset according to the second region label or the first region label to obtain at least one target dataset includes:
[0012] According to the second region label or the first region label, the object data in the object dataset is aggregated to obtain at least one initial dataset; each initial dataset corresponds to a second region label or a first region label.
[0013] Based on the set number of datasets, the initial datasets are combined to obtain at least one target dataset. Each target dataset has a target region label, which includes at least one second region label or at least one first region label.
[0014] In an optional implementation, the step of participating in data alignment based on the filtered target datasets includes:
[0015] Based on the preset alignment data volume, target object data is selected from each of the filtered target datasets, and data alignment is performed based on the target object data.
[0016] If the amount of data to be aligned does not meet the preset alignment data amount, then according to the preset increment, incremental object data is selected from the remaining object data of each of the filtered target datasets, and the incremental object data is added to the target object data;
[0017] The data alignment is performed based on the increased target object data. If the amount of data that has been aligned still does not meet the preset alignment quantity, incremental object data is selected again from the remaining object data of each of the filtered target datasets according to the preset increment. This process is repeated until the amount of data that has been aligned meets the preset alignment quantity or the object data in each of the filtered target datasets has been selected.
[0018] In an optional implementation, the step of participating in data alignment based on the filtered target datasets includes:
[0019] Determine the hash value of each object data in each of the filtered target datasets;
[0020] Based on the target character in the hash value, the target string of each object data in each of the filtered target datasets is determined;
[0021] Data alignment is performed based on the target strings of each object data in each of the filtered target datasets.
[0022] In an optional implementation, determining the target string for each object data in each filtered target dataset based on the target character in the hash value includes:
[0023] Based on the set character length, determine the hash prefix of the hash value of each object data in each of the filtered target datasets;
[0024] Based on the hash prefix, the target string of each object data in each of the filtered target datasets is obtained.
[0025] In an optional implementation, after data alignment based on the filtered target datasets, the method further includes:
[0026] Retrieve data information during data alignment;
[0027] The data information is encrypted using the set key to obtain ciphertext.
[0028] The set key is encrypted using the public key to obtain the ciphertext key;
[0029] The encrypted data and the encrypted key are sent to the blockchain.
[0030] In an optional implementation, encrypting the data information according to a set key to obtain ciphertext includes:
[0031] Determine the hash values of all object data in the data information;
[0032] Based on the hash values of all object data, determine the root hash value that yields the hash values of all object data;
[0033] The root hash value is encrypted using the set key to obtain ciphertext data.
[0034] In a second aspect, embodiments of the present invention provide a data alignment device, the data alignment device comprising:
[0035] The data acquisition module is used to acquire the object dataset for data alignment; the object dataset includes object data, a first region label to which the object data belongs, and a second region label of the object corresponding to the object data;
[0036] A first processing module is configured to divide the object dataset according to the second region label or the first region label to obtain at least one target dataset; each target dataset has a target region label, the target region label including at least one second region label or including at least one first region label;
[0037] The second processing module is used to filter object data in each target dataset whose first region label does not match the target region label of each target dataset if the target region label of each target dataset includes a second region label, and to filter object data in each target dataset whose second region label of the corresponding object does not match the target region label of each target dataset if the target region label of each target dataset includes a first region label;
[0038] The data alignment module is used to participate in data alignment based on the filtered target datasets.
[0039] Thirdly, embodiments of the present invention provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the data alignment method described in any of the foregoing embodiments.
[0040] Fourthly, embodiments of the present invention provide a readable storage medium, the readable storage medium including a computer program, wherein the computer program, when running, controls the calibration device where the readable storage medium is located to perform the data alignment method described in any of the foregoing embodiments.
[0041] The data alignment method, apparatus, electronic device, and readable storage medium provided in this invention, after obtaining the object dataset for data alignment, divides the object dataset according to a second region label or a first region label to obtain at least one target dataset. After obtaining the target dataset, if the target region label of each target dataset includes a second region label, object data in each target dataset whose first region label does not match the target region label corresponding to each target dataset is filtered. If the target region label of each target dataset includes a first region label, object data in each target dataset whose second region label does not match the target region label corresponding to each target dataset is filtered. In this way, the region to which the object data belongs in the filtered target dataset is the same as the region of the object corresponding to the object data. When performing data alignment based on the filtered target dataset, leakage of region information is avoided, and data alignment efficiency is guaranteed. Attached Figure Description
[0042] The technical solution and other beneficial effects of the present invention will become apparent from the following detailed description of specific embodiments of the invention, in conjunction with the accompanying drawings.
[0043] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of the present invention.
[0044] Figure 2 This is a schematic diagram of an electronic device provided in an embodiment of the present invention.
[0045] Figure 3 This is a flowchart illustrating a data alignment method provided in an embodiment of the present invention.
[0046] Figure 4 This is a schematic diagram of a hash calculation process provided in an embodiment of the present invention.
[0047] Figure 5 This is a block diagram of a data alignment device provided in an embodiment of the present invention. Detailed Implementation
[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0049] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0050] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.
[0051] The following disclosure provides many different embodiments or examples for implementing various structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. In addition, examples of various specific processes and materials are provided in this invention, but those skilled in the art will recognize the application of other processes and / or the use of other materials.
[0052] As described in the background section, the industry has adopted privacy-preserving computation technology to achieve data alignment among enterprises. With the development of privacy-preserving computation technology, more and more enterprises are exploring its practical applications, enabling joint marketing and risk control through vertical federated learning. Data ID alignment, as a crucial step in vertical federated learning, has become a hot research topic due to its high efficiency and security.
[0053] The industry generally uses unique identifiers such as mobile phone numbers and ID cards as data IDs for alignment. However, mobile phone numbers and ID cards are considered sensitive user data, and to protect this sensitive data, companies typically encrypt and store it in their databases. Since the encrypted data varies between companies, data alignment is a key issue that needs to be addressed in business collaborations. Taking mobile phone numbers as an example, companies first need to decrypt the encrypted mobile phone numbers and then perform data alignment using the Private Set Intersection (PSI) algorithm.
[0054] Because privacy computing technology is relatively cutting-edge, companies are relatively cautious. In the early stages, if companies use full data for alignment, it would require hundreds of millions of data calculations, which is inefficient. According to the principle of minimizing information security, aligning data based on a certain region would pose a risk of leakage of regional information if the data itself has regional attributes.
[0055] The data alignment method, apparatus, electronic device, and readable storage medium provided in this invention, after obtaining the object dataset for data alignment, divides the object dataset according to a second region label or a first region label to obtain at least one target dataset. After obtaining the target dataset, if the target region label of each target dataset includes a second region label, object data in each target dataset whose first region label does not match the target region label corresponding to each target dataset is filtered. If the target region label of each target dataset includes a first region label, object data in each target dataset whose second region label does not match the target region label corresponding to each target dataset is filtered. In this way, the region to which the object data belongs in the filtered target dataset is the same as the region of the object corresponding to the object data. When performing data alignment based on the filtered target dataset, leakage of region information is avoided, and data alignment efficiency is guaranteed.
[0056] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of the data alignment method provided in this embodiment. For example... Figure 1 As shown in the schematic diagram of the application scenario provided in this embodiment, a data alignment system is included, which includes a matching platform 200, a data requester 300, and a data provider 100.
[0057] Among them, data requester 300 refers to the party that needs to introduce external data and share it through privacy-preserving computation, serving internal business needs; data provider 100 refers to the party that provides data to data requester; and matching platform 200 refers to the platform that performs privacy-preserving computation.
[0058] In this embodiment, the matching platform 200 performs privacy calculations using the Privacy Set Intersection (PSI) technique. The purpose of PSI is to complete the intersection calculation of datasets while protecting the data privacy of both communicating parties.
[0059] In this embodiment, the matching platform 200 can be deployed on either the data requester 300 or the data provider 100, or on a third party. It can be set according to actual needs, and this embodiment does not make any specific limitations.
[0060] In this embodiment, the matching platform 200 can provide a communication interface to the data requester 300 and the data provider 100. The matching platform 200 can use the communication interface to call the data of the data requester and the data provider for privacy calculation. The data requester 300 and the data provider 100 can also use the communication interface to publish privacy calculation tasks on the matching platform 200. After the data requester 300 and the data provider 100 publish privacy calculation tasks on the matching platform 200, the matching platform 200 will perform privacy calculations.
[0061] To ensure data privacy and security, in this embodiment, both the data requester 300 and the data provider 100 are equipped with electronic devices. The data requester 300 and the data provider 100 process the data through the deployed electronic devices, such as performing encryption and decryption calculations on the data participating in data alignment.
[0062] In this embodiment, please refer to the following: Figure 2 The electronic devices of the data requester 300 and the data provider 100 may include a memory 10, a processor 20 and a communication unit 30. The memory 10 stores machine-readable instructions that can be executed by the processor 20. When the electronic device is running, the processor 20 and the memory 10 communicate via a bus. The processor 20 executes the machine-readable instructions and performs a data alignment method.
[0063] The memory 10, processor 20, and communication unit 30 are electrically connected directly or indirectly to each other to achieve signal transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory 10 stores software function modules. The processor 20 is used to execute the software function modules stored in the memory 10.
[0064] The memory 10 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0065] In some embodiments, processor 20 is used to perform one or more functions described in this embodiment. In some embodiments, processor 20 may include one or more processing cores (e.g., a single-core processor (S) or a multi-core processor (S)).
[0066] By way of example only, processor 20 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a graphics processing unit (GPU), a physical processing unit (PPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic device (PLD), a controller, a microcontroller unit, a reduced instruction set computing (RISC) computer, or a microprocessor, or any combination thereof.
[0067] For ease of explanation, only one processor is described in the electronic device. However, it should be noted that the electronic device in this embodiment may also include multiple processors, and therefore the steps performed by one processor as described in this embodiment may also be performed jointly or individually by multiple processors. For example, if the processor of the electronic device performs steps A and B, it should be understood that steps A and B may also be performed jointly by two different processors or individually by one processor. For example, one processor performs step A, and a second processor performs step B, or the first processor and the second processor jointly perform steps A and B.
[0068] In this embodiment, the memory 10 is used to store the program, and the processor 20 is used to execute the program after receiving the execution instruction. The process definition method disclosed in any implementation of this embodiment can be applied to the processor 20, or implemented by the processor 20.
[0069] The communication unit 30 is used to establish a communication connection between the electronic device and other devices via a network, and to send and receive data via the network.
[0070] In some implementations, the network can be any type of wired or wireless network, or a combination thereof. By way of example only, the network may include wired networks, wireless networks, fiber optic networks, telecommunications networks, intranets, the Internet, local area networks (LANs), wide area networks (WANs), wireless local area networks (WLANs), metropolitan area networks (MANs), public switched telephone networks (PSTNs), Bluetooth networks, ZigBee networks, or near field communication (NFC) networks, or any combination thereof.
[0071] In this embodiment, the electronic device can be an ultra-mobile personal computer (UMPC), a physical server, or a service cluster composed of multiple physical servers. This embodiment does not impose any restrictions on the specific type of electronic device.
[0072] When the matching platform 200 is deployed on the data provider 100, it can be deployed on any data processing device within the data provider 100 that communicates with the electronic devices of the data provider 100. This device also communicates with the electronic devices in the data requester 300. Thus, either the data provider 100 or the data requester 300 can publish a privacy computation task on the matching platform 200 through this device. The matching platform then sends the task to the electronic devices in both the data provider 100 and the data requester 300. Upon receiving the task, the electronic devices in both parties select object data to participate in data alignment, i.e., they send the selected object data to the matching platform. The matching platform then performs privacy computation to determine the objects shared by the data provider 100 and the data requester 300.
[0073] Accordingly, when the matching platform 200 is deployed on the data requester 300, it can also be deployed on any data processing device within the data requester 300 that communicates with the electronic devices of the data requester 300, and this device also communicates with the electronic devices in the data provider 100. It is understood that when the matching platform 200 is deployed on a third-party device, the third-party device needs to communicate with the electronic devices of both the data requester 300 and the data provider 100. In this way, the data requester 300 and the data provider 100 can publish privacy-preserving computation tasks on the matching platform 200, and the matching platform 200 can also access the data from the data requester and the data provider for privacy-preserving computation.
[0074] Understandably, Figure 2 The structure shown is for illustrative purposes only. Electronic devices may also have more advanced features. Figure 2 Showing more or fewer components, or having with Figure 2 The different configurations shown. Figure 2 The components shown can be implemented using hardware, software, or a combination thereof.
[0075] based on Figure 1 and Figure 2 Please refer to the implementation architecture shown below. Figure 3 This embodiment provides a data alignment method. This data alignment method is applied to the data provider, who executes the method. The execution steps of the data alignment method provided in this embodiment are described in detail below. Figure 3 As shown, the data alignment method provided in this embodiment includes steps S101 to S104.
[0076] Step S101: Obtain the dataset of objects participating in data alignment.
[0077] The object dataset includes object data, the first region label to which the object data belongs, and the second region label of the object corresponding to the object data.
[0078] In this embodiment, the object can be the target user for whom data alignment is needed. For example, assuming there are user 1, user 2, and user 3, if the data requester needs data from user 1 and user 2, they will select user 1 and user 2 for data alignment. Therefore, user 1 and user 2 are the objects selected by the data requester. Correspondingly, if the data provider needs data alignment from user 1, user 2, and user 3, then user 1, user 2, and user 3 are the objects selected by the data provider.
[0079] In this model, the data requester is the party that needs to introduce external data for privacy-preserving computation and data sharing, while the data provider is the party that provides data to the data requester. The data requester possesses relatively little feature data about its users, while the data provider possesses a greater amount of feature data about its users (such as behavioral data, informational data, etc.). The data requester requires the data provider to provide more feature data about its users. Therefore, the data requester can identify the users who need to provide feature data based on its actual business needs. These users represent the target users for data alignment, meaning the data requester can identify the objects based on its actual business needs and obtain the object dataset based on these identified objects. For the data provider, possessing a large amount of feature data about its users, it needs to match these users with the objects provided by the data requester to determine which of its own users can provide feature data to the data requester. Therefore, for the data provider, all its own users represent the target users for data alignment, meaning the data provider can treat all its own users as objects to obtain the object dataset.
[0080] To avoid providing feature data of other objects (i.e., objects other than those provided by the data requester) to the data requester and to prevent data privacy leaks, the data provider will perform data alignment calculations with the data requester using the object data of the objects before providing feature data to the data requester. This will identify the objects that are common to both parties. The objects that are common to both parties are the objects for which feature data needs to be provided. Only then can the data provider send the feature data of the common objects to the data requester.
[0081] Among them, object data can be the object's identification data, i.e., ID data, such as mobile phone number, ID card number, etc.
[0082] For data providers, possessing vast amounts of data, aligning object data with data requesters using the entire dataset would involve extensive computation, resulting in extremely low efficiency. Therefore, adhering to the principle of minimizing information security, data providers typically offer object data for only a specific region at a time for alignment. This region-based alignment is determined by selecting objects whose actual addresses correspond to the regions where those objects reside. For instance, if data alignment requires objects from region A, only object data from objects whose actual addresses belong to region A will be selected for alignment.
[0083] However, according to the principle of information security minimization, when aligning object data for objects in a specific region, the object data itself has a regional attribute. Since the region to which the object data belongs may differ from the region to which the object's actual address belongs, the following situations may occur: the object data does not belong to that region, but the actual address of the corresponding object does; or the object data belongs to that region, and the actual address of the corresponding object also belongs to that region. During data alignment, this could lead to the leakage of the address region information of objects whose object data does not belong to that region, but whose actual addresses do.
[0084] For example, when the object data is a mobile phone number, the region to which the mobile phone number belongs can be obtained based on the first 7 digits of the mobile phone number. Suppose that the object data (mobile phone number) of the object in region A is selected for data alignment. If there are two situations for the selected object data, the first is that the region obtained by analyzing the first 7 digits of the object data does not belong to region A, and the second is that the region obtained by analyzing the first 7 digits of the object data belongs to region A.
[0085] In the first scenario, if the analyzed region of most object data belongs to region A, then the data alignment will be performed on object data belonging to objects in region A. For object data that does not belong to region A, it will be inferred that the actual address of the corresponding object is in region A, thus leaking the address region information (region A) of that object data. In the second scenario, since the region analyzed from the object data is the same as the selected region A, both the data provider and the data requester can determine the region of the object data. Therefore, there is no leakage of the object's address region information for this part of the object data.
[0086] To avoid the leakage of address area information, in this embodiment, the object data participating in data alignment can be filtered according to the first area label to which the object data belongs and the second area label of the object corresponding to the object data. This avoids the leakage of address area information and ensures that the data is more controllable when solving the problem of decryption for all users.
[0087] In this implementation, each region corresponds to a region label, which can be a unique label such as the region's name or code. The first region label of the object data represents the label of the region to which the object data itself belongs. This label can be obtained from the object data itself. For example, for a mobile phone number, the region to which the mobile phone number belongs can be determined from the first 7 digits, and then the label of the region can be obtained. For an ID card number, the region to which the ID card number belongs can be determined from the first two digits, and then the label of the region can be obtained. The second region label of the object corresponding to the object data represents the label of the region where the actual address of the object is located. This label can be obtained from the address of the object, determining the region where the object is located, and then the label of the region.
[0088] In this embodiment, the second region label of an object may be different from the first region label to which the object data belongs, that is, the actual address of the object is not located in the region to which the object data belongs. Alternatively, the second region label of an object may be the same as the first region label to which the object data belongs, that is, the actual address of the object is located in the region to which the object data belongs.
[0089] Step S102: Divide the object dataset according to the second region label or the first region label to obtain at least one target dataset.
[0090] Since the object dataset contains a large number of objects, directly filtering the object data based on the first and second region labels would be inefficient. To improve data alignment efficiency, this embodiment first divides the object data in the object dataset into multiple regions, and then filters the object data in each region.
[0091] Therefore, in this embodiment, after obtaining the object dataset, it can be divided according to the second region label or the first region label to obtain the target dataset. The obtained target dataset has a target region label, which includes the second region label or includes the first region label. It can be understood that when the object dataset is divided according to the second region label, the target region label of the obtained target dataset includes the second region label; when the object dataset is divided according to the first region label, the target region label of the obtained target dataset includes the first region label.
[0092] When partitioning an object dataset based on its second region label, one can first identify object data with the same second region label, and then group these object data together to obtain the target dataset. Understandably, the target dataset's target region labels include the second region labels of the objects corresponding to the object data within that target dataset.
[0093] For example, the second region label of the object corresponding to object data 'a' is A, the second region label of the object corresponding to object data 'b' is A, the second region label of the object corresponding to object data 'c' is B, the second region label of the object corresponding to object data 'd' is C, and the second region label of the object corresponding to object data 'e' is B. The second region labels of the objects corresponding to object data 'a' and object data 'b' are the same, and the second region labels of the objects corresponding to object data 'c' and object data 'e' are the same. Therefore, object data 'a' and object data 'b' can be grouped into one target dataset, object data 'c' and object data 'e' into another target dataset, and object data 'd' into yet another target dataset. Specifically, the target dataset including object data 'a' and object data 'b' has a target region label including the second region label 'A', the target dataset including object data 'c' and object data 'e' has a target region label including the second region label 'B', and the target dataset including object data 'd' has a target region label including the second region label 'C'.
[0094] In an optional implementation, after grouping object data with the same second region label together, object data with different second region labels can also be selected and combined to obtain a target dataset. Thus, the target dataset can also have multiple second region labels as its target region labels. For example, following the above example, after grouping object data a and object data b together to obtain object data corresponding to second region label A, and grouping object data c and object data e together to obtain object data corresponding to second region label B, the object data corresponding to second region label A and the object data corresponding to second region label B can be combined into a single target dataset. This target dataset then has target region labels including second region label A and second region label B.
[0095] Based on this, in this embodiment, the target region label of the target dataset may include one or more second region labels; that is, the target region label of the target dataset includes at least one second region label. When the target region label of the target dataset includes one second region label, it indicates that there is only one region of object data in the target dataset. When the target region label of the target dataset includes multiple second region labels, it indicates that there is object data containing multiple regions in the target dataset. It should be noted that since the target dataset is obtained by partitioning the object dataset according to the second region labels, the number of second region labels included in the target region label of the target dataset is less than the number of second region labels in the object dataset. That is, the number of second region labels included in the target region label is less than the sum of the number of types of second region labels corresponding to all object data in the object dataset.
[0096] Understandably, when all objects in the object dataset have the same second region label, that is, when there is only one second region label in the object dataset, the target region label of the target dataset obtained by partitioning will also include only one second region label. That is, the target region label of this target dataset includes only one second region label, but there are multiple data objects with the same second region label in this target dataset.
[0097] Understandably, when partitioning an object dataset based on the first region label of the object data, one can first find object data with the same first region label, and then group the object data with the same first region label together to obtain the target dataset. The target region label of the target dataset includes the first region label of the object data in the target dataset. The specific process can be referred to as the process of partitioning an object dataset based on the second region label, which will not be elaborated on here.
[0098] Step S103: If the target region label of each target dataset includes a second region label, filter the object data in each target dataset whose first region label does not match the target region label of the corresponding target dataset. If the target region label of each target dataset includes a first region label, filter the object data in each target dataset whose second region label does not match the target region label of the corresponding object.
[0099] In this scenario, when the target region label of the target dataset includes a second region label, it indicates that the target dataset was partitioned according to the second region label of the object corresponding to the object data. If the object data in the target dataset is then used for data alignment, and the region to which a certain object data belongs differs from its second region label, the region information of the object corresponding to that object data will be leaked. Therefore, after partitioning the target dataset, to avoid the leakage of the second region labels of the objects corresponding to the object data in the target dataset, the object data in the target dataset can be filtered.
[0100] Therefore, when the target region label of the target dataset includes a second region label, it is necessary to filter out object data in each target dataset whose first region label does not match the corresponding target region label. That is, for each target dataset, the first region label of each object data in that target dataset can be compared with the target region label of that target dataset to identify object data whose first region label does not match the target region label of that target dataset.
[0101] In this embodiment, when the target region label of the target dataset includes only one second region label, the first region label of the object data in the target dataset does not match the target region label of the target dataset. This means that the first region label of the object data in the target dataset is different from the second region label included in the target region label of the target dataset. It can be understood that when the first region label of the object data in the target dataset is the same as the second region label included in the target region label of the target dataset, it indicates that the first region label of the object data in the target dataset matches the target region label of the target dataset.
[0102] When the target region label of the target dataset includes multiple second region labels, a mismatch between the first region label of the object data in the target dataset and the target region label of the target dataset means that the first region label of the object data in the target dataset is different from any of the second region labels included in the target region label of the target dataset. Conversely, a match between the first region label of the object data in the target dataset and the target region label of the target dataset is indicated by the first region label of the object data being the same as any of the second region labels included in the target region label of the target dataset.
[0103] Since the second region label included in the target region label of the target dataset represents the region label of the object corresponding to all object data in the target dataset, when the first region label of a certain object data in the target dataset does not match the target region label of the target dataset, it means that the region to which the object data belongs is different from the region of the object corresponding to the object data. If the object data is used for data alignment, it may cause leakage of the region of the object corresponding to the object data. Therefore, the object data needs to be removed to avoid leakage of the region information of the object corresponding to the object data.
[0104] Based on this, in this embodiment, if the target region label of each target dataset includes a second region label, when the first region label of the object data in the target dataset does not match the target region label of the target dataset, the object data is filtered, that is, the object data is removed, and the filtered target dataset is obtained.
[0105] In this embodiment, when the target region label of the target dataset includes the first region label, it indicates that the target dataset is divided according to the first region label to which the object data belongs. When dividing the object data according to the first region label to which the object data belongs, the region of the object corresponding to the object data is not involved. Therefore, the leakage of the region information of the object corresponding to the object data can be avoided. Therefore, in order to improve efficiency and save computing power, if the target region label of each target dataset includes the first region label, data alignment can be directly performed based on each target dataset.
[0106] To further enhance security during the data alignment process and prevent the leakage of object region information, if the target region label of each target dataset includes a first region label, object data whose second region label of the corresponding object in each target dataset does not match the target region label of the target dataset can also be filtered. That is, for each target dataset, the second region label of the object corresponding to each object data in the target dataset can be compared with the target region label of the target dataset to determine the object data whose second region label of the corresponding object does not match the target region label of the target dataset.
[0107] When the second region label of an object in the target dataset does not match the target region of the target dataset, meaning the second region label of the object is different from the first region label of the object, the object data can be filtered out. This ensures the consistency between the first region label and the second region label of the object, further preventing the leakage of region information of the object. The specific process can be referenced from the process of filtering object data in the target dataset whose first region label does not match the target region label of the target dataset; it will not be elaborated upon here.
[0108] Step S104: Perform data alignment based on the filtered target datasets.
[0109] In this process, after filtering the target dataset obtained by the above method, the filtered target dataset can be used for data alignment.
[0110] During data alignment, the data requester can send the object data to be aligned to the matching platform. After the data provider sends the filtered target dataset to the matching platform, the matching platform performs privacy calculations on the object data sent by the data requester and the object data in the filtered target dataset sent by the data provider to obtain the common object data of the data provider and the data requester. Then, the common object data is sent to the data provider and the data requester to complete the data alignment.
[0111] After receiving the common object data, the data provider can send the characteristic data of the corresponding object to the data requester to facilitate the data requester's business development.
[0112] In this embodiment, the object data in the target dataset is filtered by the target region label of the target dataset, the first region label of each object data in the target dataset, and the second region label of the object corresponding to each object data in the target dataset. This ensures that the region to which the object data in the filtered target dataset belongs is the same as the region of the object corresponding to the object data. Thus, when data alignment is performed based on the filtered target dataset, the region information of objects whose region is different from their own region will not be leaked.
[0113] It should be noted that, in this embodiment, after determining the object data to be aligned, the data requester can also filter the object data based on the first region label and the second region label of the object data. The specific process can be referenced from the data provider's object data filtering process, which will not be elaborated upon here. After obtaining the filtered object data, the data requester can send it to the matching platform to participate in data alignment. Thus, even for objects whose region differs from the data requester's own region, the data requester will not cause leakage of their region information.
[0114] Given that in existing technologies, data alignment is performed by providing only object data from a single region at a time, based on the principle of minimizing information security, this method is inefficient and prone to leaking regional information. To improve data alignment efficiency and prevent information leakage caused by single-region alignment, this embodiment combines object data from multiple regions for data alignment. Therefore, in this embodiment, the object dataset is divided according to a second region label or a first region label to obtain at least one target dataset, including:
[0115] The object data in the object dataset is aggregated according to the second region label or the first region label to obtain at least one initial dataset; each initial dataset corresponds to a second region label or a first region label.
[0116] Based on the set number of datasets, the initial datasets are combined to obtain at least one target dataset. Each target dataset has a target region label, which includes at least one second region label or at least one first region label.
[0117] When aggregating object data in an object dataset according to the second region label, the process can begin by identifying object data with the same second region label, and then aggregating these object data together to obtain an initial dataset. Each initial dataset contains object data with the same second region label; therefore, each initial dataset corresponds to one second region label.
[0118] Accordingly, when aggregating object data in an object dataset according to the first region label, one can first find object data with the same first region label, and then aggregate the object data with the same first region label together to obtain an initial dataset. Each initial dataset includes object data with the same first region label; therefore, each initial dataset corresponds to one first region label.
[0119] After obtaining the initial dataset, each initial dataset represents object data in the same region. In order to improve the efficiency of data alignment and prevent regional information leakage caused by a single region, the obtained initial datasets can be combined according to the set number of datasets to obtain the target dataset.
[0120] The number of datasets can be set according to actual needs, and this embodiment does not impose a specific limitation. For example, when the number of datasets is set to 3, 3 datasets can be selected from each initial dataset and combined, that is, 3 initial datasets result in one target dataset. However, it should be noted that if the object data in the object dataset is aggregated according to the second region label, the number of datasets must be less than the number of second region labels in the object dataset, that is, the number of datasets must be less than the sum of the number of types of second region labels corresponding to all object data in the object dataset. If the object data in the object dataset is aggregated according to the first region label, the number of datasets must be less than the number of first region labels in the object dataset, that is, the number of datasets must be less than the sum of the number of types of first region labels belonging to all object data in the object dataset.
[0121] Understandably, each target dataset has a different initial dataset, and the target region label of each target dataset includes the region label (first region label or second region label) corresponding to the initial dataset. If the region label corresponding to the initial dataset is the first region label, then the target region label of the target dataset includes the first region label corresponding to each initial dataset in the target dataset; if the region label corresponding to the initial dataset is the second region label, then the target region label of the target dataset includes the second region label corresponding to each initial dataset in the target dataset.
[0122] In an optional implementation, this embodiment can also set the priority of region labels, and combine the initial datasets according to the priority of the region labels and the set number of datasets. Specifically, the priority of the second region label can be set according to indicators such as the number of users in each region and the range of each region. Then, according to the priority of the region labels and the set number of datasets, the initial datasets are selected sequentially from the highest priority for combination. For example, assuming the region label corresponding to the initial dataset is the second region label, initial datasets a, b, c, and d, sorted by the priority of the second region labels from high to low, are initial dataset a, initial dataset c, initial dataset d, and initial dataset b, respectively. Assuming a dataset 2 is set, by selecting the initial datasets sequentially from the highest priority according to the priority of the second region labels and the set number of datasets, initial datasets a and c can be combined into a target dataset, and initial datasets d and b can be combined into a target dataset.
[0123] After combining the initial datasets to obtain the target dataset, the object data in the target dataset can be filtered, and then data alignment can be performed based on the filtered target datasets.
[0124] The data alignment method provided in this embodiment combines an initial dataset with a set number of datasets to obtain a target dataset. The target dataset includes object data in the same number of regions as the set number of datasets. In this way, when data alignment is performed based on the filtered target dataset, the problem of regional information leakage caused by a single region is avoided, and the alignment efficiency can also be improved.
[0125] Since the data provider possesses a large amount of data, in order to meet the principle of minimizing security and improve the security of the data alignment process, this embodiment can adopt a multi-round incremental approach for data alignment. Based on this, in this embodiment, the steps involved in data alignment, according to the filtered target datasets, may further include:
[0126] Based on the preset alignment data volume, target object data is selected from each filtered target dataset, and data alignment is performed based on the target object data.
[0127] If the amount of data to be aligned does not meet the preset alignment amount, then incremental object data is selected from the remaining object data of each target dataset after filtering, according to the preset increment, and the incremental object data is added to the target object data.
[0128] The data alignment is performed based on the increased target object data. If the amount of data that has been aligned still does not meet the preset alignment quantity, incremental object data is selected again from the remaining object data of each filtered target dataset according to the preset increment. This process is repeated until the amount of data that has been aligned meets the preset alignment quantity or all object data in each filtered target dataset has been selected.
[0129] The preset alignment data volume refers to the amount of data that the data requester and the data provider agree on to achieve alignment, i.e., the amount of shared object data that needs to be identified. For example, if the data requester provides 2000 object data items, and the preset alignment data volume agreed upon by the data requester and the data provider is 1800, then the number of object data items that the data provider and the data requester need to align is 1800. In other words, the data provider and the data requester need to identify 1800 shared objects from the 2000 object data items through data alignment.
[0130] When a data provider selects target object data from filtered target datasets based on a preset alignment data volume, it selects target object data with the same volume as the preset alignment data volume. It's important to note that the selected target object data only represents the object data provided by the data provider for data alignment. This target object data may include object data identical to or different from the object data provided by the data requester. Therefore, the amount of aligned object data achieved through target object data alignment may be less than the preset alignment data volume. For example, if the preset alignment data volume is 1000, and 1000 target object data are selected from the target datasets, only a portion of these 1000 target object data may be aligned with the object data provided by the data requester, failing to meet the agreed-upon requirement of a sufficient number of aligned data.
[0131] Based on this, in this embodiment, after filtering out the target object data in each target dataset, the target object data can be sent to the matching platform. The matching platform will then perform the first round of privacy calculations based on the target object data sent by the data provider and the object data sent by the data requester to obtain the first round of shared object data.
[0132] If the amount of shared object data in the first round does not meet the preset alignment data amount, that is, the amount of data that has reached alignment does not meet the preset alignment quantity, the data provider will select incremental object data from each target dataset that has been filtered out, that is, select incremental object data from the remaining object data in each target dataset, and then add the incremental object data to the target object data to become the target object data. The added target object data will then participate in data alignment, that is, the added target object data will be sent to the matching platform. The matching platform will then perform a second round of privacy calculation based on the added target object data and the object data sent by the data requester to obtain the second round of shared object data.
[0133] If the amount of shared object data in the second round does not meet the preset alignment data amount, that is, if the amount of data that has reached alignment does not meet the preset alignment quantity, the data provider needs to select incremental object data again from each target dataset that has been filtered for target object data and incremental object data, according to the preset increment. That is, incremental object data is selected again from the remaining object data in each target dataset. The newly selected incremental object data is added to the target object data to become the target object data. The target object data is then used for data alignment again. That is, the target object data is sent to the matching platform again, and the matching platform performs a third round of privacy calculation with the object data sent by the data requester. This process is repeated until the amount of data that has reached alignment meets the preset alignment quantity or all object data in each target dataset has been selected after filtering.
[0134] For example, if the preset alignment data volume is 1000 and the preset increment is 200, the data provider first selects 1000 target object data from each target dataset and sends them to the matching platform for the first round of privacy calculation. If the data volume of the shared object data in the first round is 300, which is less than 1000, then 200 incremental object data are selected from the remaining object data in each target dataset and added to the target object data, bringing the total number of target object data to 1200. Then, the 1200 target object data are sent to the matching platform for the second round of privacy calculation. If the data volume of the shared object data in the second round is 350, which is less than 1000, then 200 incremental object data are selected from the remaining object data in each target dataset and added to the target object data, bringing the total number of target object data to 1400. Then, the 1400 target object data are sent to the matching platform for the third round of privacy calculation, and so on, until the data volume of the shared object data reaches 1000, or all object data in each target dataset has been selected.
[0135] In an optional implementation, to improve the efficiency of data alignment, the data provider can sequentially select object data from each target dataset, and after selecting object data from each target dataset, select object data from the next target dataset. Understandably, when a target dataset is obtained by combining multiple initial datasets, when the data provider sequentially selects object data from each filtered target dataset, it is difficult to guess the corresponding region label for the target dataset because the selected object data belongs to different first region labels. This avoids the problem of region information leakage caused by a single region.
[0136] In an optional implementation, to improve the efficiency of data alignment, the data provider can assign priorities to each target dataset and then select object data based on the priority of each target dataset. The priority of each target dataset can be set based on the target region label or the amount of object data it possesses; there are no specific limitations, and it can be set according to actual requirements.
[0137] The data alignment method provided in this embodiment performs data alignment in a multi-round incremental manner, which improves privacy and security in data alignment while meeting the principle of security minimization.
[0138] Since there is a large amount of object data in each target dataset and it is complete identification data of the objects, in order to reduce the leakage of object data and improve the data alignment efficiency, this embodiment can perform hash calculation on the object data in the target dataset, then select some characters in the hash value obtained by hash calculation for pre-alignment, and then perform data alignment with the complete object data based on the pre-alignment result.
[0139] Based on this, in this embodiment, the steps involved in data alignment, according to the filtered target datasets, may include:
[0140] (1) Determine the hash value of each object data in each target dataset after filtering.
[0141] (2) Determine the target string of each object data in each filtered target dataset based on the target character in the hash value.
[0142] (3) Based on the target strings of each object data in each filtered target dataset, participate in data alignment.
[0143] Among them, a preset hash algorithm can be used to perform hash calculations on the object datasets in each filtered target dataset to obtain the hash value of each object data in each filtered target dataset.
[0144] In this embodiment, the hash algorithm can be, but is not limited to, SHA256, MD5, MD4, MD3, etc. For example, when the object data is 12345678901, using the SM3 algorithm, its hash value can be obtained as follows:
[0145] 8da28dd7ed4357331a0d05202acb5b6bc7be4b7530e24e1cadb1086d4deb7ce5.
[0146] After determining the hash value of each object data in each filtered target dataset, the target character can be determined from the hash value of each object data in each filtered target dataset, and then the target string of the hash value of the object data can be obtained based on the target character.
[0147] In this embodiment, when determining the target character from the hash value of the object data, it can start from the beginning of the hash value and then determine the target character according to the set character length, that is, select the prefix of the hash value according to the set character length; it can also start from the end of the hash value and then determine the target character according to the set character length, that is, select the suffix of the hash value according to the set character length; or it can start from the middle of the hash value and then determine the target character according to the set character length, that is, select the middle part of the hash value according to the set character length.
[0148] In this embodiment, to facilitate data alignment, a hash prefix is determined based on a set character length, and this hash prefix is used as the target string corresponding to the object data. Therefore, in this embodiment, the target character can be determined from the hash value of the object data through the following steps, and the target string of the object data can be obtained based on the target character in the hash value:
[0149] Based on the set character length, determine the hash prefix of the hash value of each object data in each filtered target dataset.
[0150] Based on the hash prefix, the target string of each object data in each filtered target dataset is obtained.
[0151] In this embodiment, the character length can be set according to actual needs; however, this embodiment does not impose such a requirement. For example, when the set character length is 4, and the object data is the mobile phone number 12345678901, the hash value of this object data is:
[0152] If the hash value of the object data is 8da28dd7ed4357331a0d05202acb5b6bc7be4b7530e24e1cadb1086d4deb7ce5, then the hash prefix of the object data is 4 bits: 8da2.
[0153] In this embodiment, after determining the hash prefix of the hash value of each object data in each filtered target dataset, the data provider can directly set the hash prefix of the hash value of each object data in each filtered target dataset as the target string of each object data in each filtered target dataset, or encrypt the hash prefix of the hash value of each object data in each filtered target dataset to obtain the target string of each object data in each filtered target dataset.
[0154] Once the data provider obtains the target strings of each object data in each filtered target dataset, it can participate in pre-alignment based on the target strings of each object data in each filtered target dataset.
[0155] It should be noted that if the data provider uses the target string of the object data for data alignment, the data requester also needs to use the target string of the object data for pre-alignment. Specifically, the data requester performs a hash calculation on the object data that needs alignment, obtains a hash value, and then selects a hash prefix from the hash value of the object data with the same character length as the data provider. Based on the hash prefix, the target string of the object data is then obtained.
[0156] It should be noted that the data provider and the data requester use the same key when encrypting the hash prefix. The specific key can be set according to actual needs, and this embodiment does not impose any specific restrictions.
[0157] Once the data requester receives the target string of the object data, they can directly send the target string to the data provider. The data provider can then filter out the pre-aligned data based on the target string sent by the data requester.
[0158] Optionally, when the data provider filters out pre-aligned data based on the target string sent by the data requester, it can match the target string sent by the data requester with the target strings of each object data in each target dataset to obtain a matching target string that is the same as the target string sent by the data requester. After obtaining the matching target string, the object data corresponding to the matching target string can be used as the pre-aligned data.
[0159] Since the target string is determined from the hash value of the object data, after the data provider obtains the same target string as the target string sent by the data requester, it can match the target string with the hash value of each object data in each target dataset. That is, for each target string, it checks whether there is a hash value that includes the target string. If so, the object data corresponding to the hash value that includes the target string is used as the pre-aligned data.
[0160] In this embodiment, when matching the target string with the hash value of the object data, a fuzzy matching method is used. As a result, one target string may match multiple hash values, and thus one target string may match multiple object data.
[0161] In this embodiment, after receiving a matching target string that is identical to the target string sent by the data requester, the data provider can send the matching target string to the data requester to indicate that the matching target string is a target string shared by both parties. Upon receiving the matching target string, the data requester can then match it with the hash values of its own object data. Specifically, for each matching target string, it checks if a hash value containing that matching target string exists. If so, it uses the object data corresponding to the hash value containing that matching target string as pre-aligned data.
[0162] In one alternative implementation, to improve the security and fairness of data alignment, the data requester and the data provider may send the target string to the matching platform for alignment after obtaining the target string of the object data.
[0163] After the matching platform receives the target strings sent by the data requester and the data provider, it matches the target strings sent by the data requester and the data provider, finds the same target strings in the target strings sent by the data requester and the data provider, obtains the intersection of the same target strings of the data requester and the data provider, and then returns the intersection to the data requester and the data provider.
[0164] Once the data requester and data provider receive the intersection, they can match the target string in the intersection with the hash value of their own object data to obtain pre-aligned data.
[0165] Once both the data requester and the data provider have determined the pre-aligned data, the shared object data between the two parties can be determined based on their pre-aligned data.
[0166] To improve the security and fairness of data alignment, after receiving the pre-aligned data, both the data requester and the data provider can send the pre-aligned data to the matching platform, which will then determine the shared object data between the two parties.
[0167] After receiving the pre-aligned data sent by the data requester and the data provider, the matching platform performs privacy calculations on the object data in the pre-aligned data, that is, it finds the same object data in the pre-aligned data sent by the data requester and the data provider, and then sends the identified same object data to the data requester and the data provider.
[0168] Once the data requester and data provider receive the same object data, they confirm that they have obtained object data shared by the other party, thus completing data alignment. Understandably, once the data provider confirms that it has obtained the shared object data, it can send the characteristic data of the object corresponding to the shared object data to the data requester.
[0169] Understandably, data providers can also send object data from pre-aligned data to the matching platform in a multi-round incremental manner. The specific process can be referred to the above description, and will not be elaborated on here.
[0170] This embodiment performs pre-alignment using the target string of the object data. Based on the pre-aligned data, the common object data is determined, avoiding the need to align all of its own object data. This reduces the amount of data involved in data alignment, thereby reducing the risk of data leakage and improving data alignment efficiency.
[0171] To ensure the authenticity and traceability of the data alignment process, this embodiment includes a monitoring node in the data alignment system. This monitoring node can co-build a blockchain with the data provider and the data requester. The data requester and data provider need to upload data information from the data alignment process to the blockchain for monitoring purposes. Based on this, after aligning the filtered target datasets, the data alignment method provided in this embodiment further includes the following steps:
[0172] (a) Obtain data information in data alignment.
[0173] (b) The data information is encrypted according to the set key to obtain the ciphertext.
[0174] (c) Encrypt the set key using the public key to obtain the ciphertext key.
[0175] (d) Send the encrypted data and the encrypted key to the blockchain.
[0176] In this context, the data information in data alignment refers to all the object data involved in the data alignment process (including object data in the object dataset, the first region label to which the object data belongs, the second region label of the object corresponding to the object data, the hash value of the object data, the hash prefix, and the object data in each round of the multi-round incremental alignment process, etc.) as well as log information and code information involved in the data alignment process.
[0177] By analyzing log information, code information, and object data information during the data alignment process, the calculation process and object data involved in the alignment can be obtained, thereby enabling traceability of the data alignment process and monitoring of the data alignment.
[0178] In this embodiment, when encrypting data information, either a symmetric key or an asymmetric key can be used for encryption; no specific restriction is imposed.
[0179] After encrypting the data to obtain the ciphertext, the public key of the monitoring node can be used to encrypt the set key to generate the ciphertext key. After obtaining the ciphertext and the ciphertext key, these two pieces of data are uploaded to the blockchain composed of the data requester, the data provider, and the monitoring node.
[0180] It should be noted that data requesters also need to encrypt the data information during their own data alignment process to obtain ciphertext. Then, they need to use the public key of the monitoring node to encrypt the set key to generate ciphertext key. After obtaining the ciphertext and ciphertext key, they need to upload these two pieces of data to the blockchain.
[0181] When a data dispute arises between the data requester and the data provider, the regulatory node can use its private key to decrypt the set key, and then use the decrypted set key to decrypt the encrypted data to obtain all the data information in the data alignment. This allows for regulatory auditing, obtaining audit results, and resolving the data dispute between the data requester and the data provider.
[0182] To save on data transmission volume, in this embodiment, the steps of encrypting data information according to a set key to obtain ciphertext include:
[0183] Determine the hash values of all object data in the data information.
[0184] Determine the root hash value that yields the hash values of all object data based on the hash values of all object data.
[0185] The root hash value is encrypted using the set key to obtain the ciphertext data.
[0186] The data provider can determine the hash value of all object data using a predefined hash algorithm, and then use a Merkle tree method to calculate the parent node hash value for each pair of hash values of all object data until the root hash value is calculated. For example... Figure 4 As shown, suppose there are object data 1, object data 2, object data 3, and object data 4. The hash value of object data 1 is hash value 1, the hash value of object data 2 is hash value 2, the hash value of object data 3 is hash value 3, and the hash value of object data 4 is hash value 4. After determining the hash values of object data 1, object data 2, object data 3, and object data 4, the parent node hash value can be calculated for each pair of hash values. This can be done by calculating the parent node hash value for the hash values of object data 1 and object data 2, resulting in parent node hash value 'a', and then calculating the parent node hash value for the hash values of object data 3 and object data 4, resulting in parent node hash value 'b'. Then, based on parent node hash values 'a' and 'b', the root hash value can be calculated. Understandably, the data requester can also use the same encryption method as the data provider.
[0187] After obtaining the root hash value, the root hash value and other information in the data can be encrypted using the set key to obtain the data ciphertext. Then, the set key is encrypted using the public key of the supervisory node to obtain the ciphertext key. Finally, the ciphertext key and the data ciphertext are uploaded to the consortium blockchain.
[0188] This embodiment introduces blockchain technology and adds supervisory nodes to ensure the authenticity of data during the data alignment process.
[0189] The data alignment method provided in this embodiment filters the object data in the target dataset by using the second region label of the target dataset and the first region label of each object data in the target dataset. This ensures that the region to which the object data in the filtered target dataset belongs is the same as the region of the object corresponding to the object data. Thus, when performing data alignment based on the filtered target dataset, there will be no leakage of region information for objects whose region is different from their own region.
[0190] Based on the same inventive concept, this embodiment also provides a data alignment device 40, applied to the electronic equipment of the data provider, such as... Figure 5 As shown, the data alignment device 40 includes a data acquisition module 41, a first processing module 42, a second processing module 43, and a data alignment module 44.
[0191] The data acquisition module 41 is used to acquire the object dataset that participates in data alignment; the object dataset includes object data, the first region label to which the object data belongs, and the second region label of the object corresponding to the object data.
[0192] The first processing module 42 is used to divide the object dataset according to the second region label or the first region label to obtain at least one target dataset; each target dataset has a target region label; the target region label includes a second region label or includes at least one first region label.
[0193] The second processing module 43 is used to filter object data in each target dataset whose first region label does not match the target region label of each target dataset if the target region label of each target dataset includes a second region label, and to filter object data in each target dataset whose second region label of the corresponding object does not match the target region label of each target dataset if the target region label of each target dataset includes a first region label.
[0194] Data alignment module 44 is used to participate in data alignment based on the filtered target datasets.
[0195] In an optional implementation, the first processing module 42 is used to:
[0196] The object data in the object dataset is aggregated according to the second region label or the first region label to obtain at least one initial dataset; each initial dataset corresponds to a second region label or a first region label.
[0197] Based on the set number of datasets, the initial datasets are combined to obtain at least one target dataset. Each target dataset has a target region label, which includes at least one second region label or at least one first region label.
[0198] In an optional implementation, the data alignment module 44 is used for:
[0199] Based on the preset alignment data volume, target object data is selected from each filtered target dataset, and data alignment is performed based on the target object data.
[0200] If the amount of data to be aligned does not meet the preset alignment amount, then incremental object data is selected from the remaining object data of each target dataset after filtering, according to the preset increment, and the incremental object data is added to the target object data.
[0201] The data alignment is performed based on the increased target object data. If the amount of data that has been aligned still does not meet the preset alignment quantity, incremental object data is selected again from the remaining object data of each filtered target dataset according to the preset increment. This process is repeated until the amount of data that has been aligned meets the preset alignment quantity or the object data in each filtered target dataset has been selected.
[0202] In an optional implementation, the data alignment module 44 is used for:
[0203] Determine the hash value of each object in each of the filtered target datasets.
[0204] Based on the target character in the hash value, determine the target string for each object in each filtered target dataset.
[0205] Data alignment is performed based on the target strings of each object data in each filtered target dataset.
[0206] In an optional implementation, the first processing module 42 is used to:
[0207] Based on the set character length, determine the hash prefix of the hash value of each object data in each filtered target dataset.
[0208] Based on the hash prefix, the target string of each object data in each filtered target dataset is obtained.
[0209] In an optional implementation, after data alignment based on the filtered target datasets, the data alignment module 44 is used to:
[0210] Retrieve data information during data alignment.
[0211] The data is encrypted using the set key to obtain the ciphertext.
[0212] The ciphertext key is obtained by encrypting the set key with the public key.
[0213] Send the encrypted data and the encrypted key to the blockchain.
[0214] In an optional implementation, the data alignment module 44 is used for:
[0215] Determine the hash values of all object data in the data information.
[0216] Determine the root hash value that yields the hash values of all object data based on the hash values of all object data.
[0217] The root hash value is encrypted using the set key to obtain the ciphertext data.
[0218] The data alignment device provided in this embodiment, after obtaining the object datasets to be aligned, divides the object datasets according to the second region label or the first region label to obtain at least one target dataset. After obtaining the target datasets, if the target region label of each target dataset includes the second region label, the object data in each target dataset whose first region label does not match the target region label of the corresponding target dataset is filtered. If the target region label of each target dataset includes the first region label, the object data in each target dataset whose second region label of the corresponding object does not match the target region label of the corresponding target dataset is filtered. In this way, the region to which the object data in the filtered target dataset belongs is the same as the region of the object corresponding to the object data. When performing data alignment based on the filtered target dataset, the leakage of region information is avoided, and the data alignment efficiency is guaranteed.
[0219] Based on the above, this embodiment provides a readable storage medium, which includes a computer program. When the computer program is executed, it controls the electronic device where the readable storage medium is located to perform the data alignment method described in any of the foregoing embodiments.
[0220] The readable storage medium can be, but is not limited to, USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, and other media capable of storing program code.
[0221] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the readable storage medium described above can be referred to the corresponding process in the aforementioned data alignment system, and will not be elaborated further here.
[0222] In summary, the data alignment method, apparatus, electronic device, and readable storage medium provided in this embodiment of the invention, after obtaining the object datasets participating in data alignment, divide the object datasets according to a second region label or a first region label to obtain at least one target dataset. After obtaining the target datasets, if the target region label of each target dataset includes a second region label, object data in each target dataset whose first region label does not match the target region label corresponding to that target dataset is filtered. If the target region label of each target dataset includes a first region label, object data in each target dataset whose second region label of the corresponding object does not match the target region label corresponding to that target dataset is filtered. In this way, the region to which the object data belongs in the filtered target dataset is the same as the region of the object corresponding to the object data. Thus, when performing data alignment based on the filtered target dataset, leakage of region information is avoided, and data alignment efficiency is guaranteed.
[0223] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0224] The data alignment method, apparatus, electronic device, and readable storage medium provided by the embodiments of the present invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the technical solutions and core ideas of the present invention. Those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data alignment method, characterized in that, The method includes: Obtain the object dataset for data alignment; the object dataset includes object data, a first region label to which the object data belongs, and a second region label of the object corresponding to the object data; the first region label represents the label of the region to which the object data itself belongs, and the second region label represents the label of the region where the actual address of the object is located; According to the second region label or the first region label, the object data in the object dataset is aggregated to obtain at least one initial dataset, and each initial dataset corresponds to a second region label or a first region label. Based on the priority of the region labels and the set number of datasets, the initial datasets are combined to obtain at least one target dataset; each target dataset has a target region label; the target region label includes at least one second region label or includes at least one first region label; If the target region label of each target dataset includes a second region label, filter out object data in each target dataset whose first region label does not match the target region label corresponding to each target dataset; if the target region label of each target dataset includes a first region label, filter out object data in each target dataset whose second region label of the corresponding object does not match the target region label corresponding to each target dataset. Based on the filtered target datasets, data alignment is performed.
2. The data alignment method according to claim 1, characterized in that, The step of aligning the data based on the filtered target datasets includes: Based on the preset alignment data volume, target object data is selected from each of the filtered target datasets, and data alignment is performed based on the target object data. If the amount of data to be aligned does not meet the preset alignment data amount, then according to the preset increment, incremental object data is selected from the remaining object data of each of the filtered target datasets, and the incremental object data is added to the target object data; The data alignment is performed based on the increased target object data. If the amount of data that has been aligned still does not meet the preset alignment quantity, incremental object data is selected again from the remaining object data of each of the filtered target datasets according to the preset increment. This process is repeated until the amount of data that has been aligned meets the preset alignment quantity or the object data in each of the filtered target datasets has been selected.
3. The data alignment method according to claim 1, characterized in that, The step of aligning the data based on the filtered target datasets includes: Determine the hash value of each object data in each of the filtered target datasets; Based on the target character in the hash value, the target string of each object data in each of the filtered target datasets is determined; Data alignment is performed based on the target strings of each object data in each of the filtered target datasets.
4. The data alignment method according to claim 3, characterized in that, The step of determining the target string for each object data in each filtered target dataset based on the target character in the hash value includes: Based on the set character length, determine the hash prefix of the hash value of each object data in each of the filtered target datasets; Based on the hash prefix, the target string of each object data in each of the filtered target datasets is obtained.
5. The data alignment method according to any one of claims 1-4, characterized in that, After data alignment based on the filtered target datasets, the method further includes: Retrieve data information during data alignment; The data information is encrypted using the set key to obtain ciphertext. The set key is encrypted using the public key to obtain the ciphertext key; The encrypted data and the encrypted key are sent to the blockchain.
6. The data alignment method according to claim 5, characterized in that, The step of encrypting the data information according to the set key to obtain ciphertext includes: Determine the hash values of all object data in the data information; Based on the hash values of all object data, determine the root hash value that yields the hash values of all object data; The root hash value is encrypted using the set key to obtain ciphertext data.
7. A data alignment device, characterized in that, The data alignment device includes: The data acquisition module is used to acquire the object dataset participating in data alignment; the object dataset includes object data, a first region label to which the object data belongs, and a second region label of the object corresponding to the object data; the first region label represents the label of the region to which the object data itself belongs, and the second region label represents the label of the region where the actual address of the object is located; A first processing module is configured to aggregate object data in the object dataset according to the second region label or the first region label to obtain at least one initial dataset, each initial dataset corresponding to a second region label or a first region label; and to combine the initial datasets according to the priority of the region labels and a set number of datasets to obtain at least one target dataset; each target dataset has a target region label, the target region label including at least one second region label or including at least one first region label. The second processing module is used to filter object data in each target dataset whose first region label does not match the target region label of each target dataset if the target region label of each target dataset includes a second region label, and to filter object data in each target dataset whose second region label of the corresponding object does not match the target region label of each target dataset if the target region label of each target dataset includes a first region label; The data alignment module is used to participate in data alignment based on the filtered target datasets.
8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the data alignment method of any one of claims 1 to 6.
9. A readable storage medium, characterized in that, The readable storage medium includes a computer program that, when executed, controls the calibration device containing the readable storage medium to perform the data alignment method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Block chain-based data alignment method and device, equipment and medium
CN113032817A
Data binning processing method, device and equipment based on confusion box and storage medium
CN113704800A