A method and device for implementing ID mapping based on the Spark framework
Through the Spark framework-based method, massive user data are processed and unified ID representation is realized using unified identifiers, which solves the problem of inefficient ID Mapping in the prior art and improves the ability to deal with complex ID networks.
Patent Information
- Application Number
- CN201910199055.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-03-15
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-03-15
AI Technical Summary
The prior art cannot quickly and accurately implement ID Mapping when processing massive user data, resulting in difficult to handle the complexity of the ID network.
Using a Spark framework-based method, the two-dimensional ID relationship table is obtained, ID splitting and aggregation is performed, and a unified identifier is used to realize the unified representation of ID. The specific steps include: obtaining the initial number-ID relationship table, performing ID relationship aggregation, obtaining the initial numbered aggregate subset, further implementing the numbered relationship aggregation through iterative operations, and finally using a unified identifier to number the result.
It improves the efficiency, accuracy and reliability of the ID Mapping algorithm, can effectively handle complex ID networks, and realizes fast and accurate unified representation of user IDs.
Smart Images

Figure CN111694876B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data, and in particular to a method and device for implementing ID mapping based on the Spark framework, a computer storage medium, and a computing device. Background Art
[0002] IDMapping is a basic and very crucial technology in the field of big data. Briefly speaking, ID-Mapping is to identify several data from different sources as the same user or entity through some technical means. For example, for a user Zhang San, he uses AA mobile assistant on the first mobile phone, Baidu Map on the second mobile phone, watches iQIYI videos on the tablet computer, and uses AA browser on the personal computer. The first mobile phone, the second mobile phone and the tablet computer often share the same wifi, and the second mobile phone is often connected to the personal computer through a data cable. Then, how to determine that these 4 objects are the same user based on the behaviors of the objects on these 4 devices and the connections between them is the main problem to be solved by ID Mapping.
[0003] IDMapping has a wide range of application scenarios and commercial values. The behavior information and attribute data of a user are scattered on many different sources of data. Analyzing the data from a single source only shows a certain part of the characteristics of this user. Through IDMapping, the fragmented partial characteristics of the user can be all connected in series to provide a complete user portrait. For example, based on Zhang San's behavior of watching iQIYI videos on the tablet computer in the above example, iQIYI app and related movies can be recommended to him on the mobile phone side. Another example is programmatic trading. An important link of it is to match the user of the current advertising request with the user's historical interest data in the first-party DMP (Data Management Platform). Without IDMapping, programmatic trading will be relatively blind and unable to achieve real-time bidding and precise placement.
[0004] Currently, the main technical bottleneck of ID Mapping is that it is impossible to quickly and accurately obtain the mapping result when processing a large amount of user data. Because the number of user IDs is huge, and there are certain connections between different IDs, these IDs and their connections form an ID network with complex relationships. How to extract a sub-network from the complex ID network, and then effectively separate or process the sub-network to obtain a reliable sub-network is a main problem to be solved in the ID Mapping project. Therefore, there is an urgent need for a method that can efficiently and reliably implement ID Mapping. Summary of the Invention
[0005] In view of the above problems, the present invention is proposed to provide a method and apparatus for implementing ID mapping based on the Spark framework, a computer storage medium, and a computing device that overcome the above problems or at least partially solve the above problems.
[0006] According to one aspect of the embodiments of the present invention, a method for implementing ID mapping based on the Spark framework is provided, including:
[0007] Step S1: Obtain a two-dimensional ID relationship table including multiple ID pairs, number each ID pair, and obtain an initial number - ID pair relationship table;
[0008] Step S2: Using the ID as the key, split and aggregate the initial number - ID pair relationship table to obtain multiple initial number first-aggregation subsets, where each initial number first-aggregation subset is composed of initial numbers;
[0009] Step S3: Using the initial number as the key, split and aggregate the multiple initial number first-aggregation subsets to obtain an initial number aggregation subset result, where there is no intersection between any two initial number aggregation subsets in the initial number aggregation subset result; number the initial number aggregation subset result using a unified identifier to obtain a unified identifier - initial number aggregation subset relationship table;
[0010] Step S4: According to the unified identifier - initial number aggregation subset relationship table and the initial number - ID pair relationship table, obtain the corresponding relationship between the unified identifier and the ID, and realize the unified representation of the ID.
[0011] Optionally, using the initial number as the key, splitting and aggregating the multiple initial number first-aggregation subsets to obtain an initial number aggregation subset result includes:
[0012] Using the initial number as the key, using the multiple initial number first-aggregation subsets as the objects for splitting and aggregating in the first iterative operation, perform iterative operations of splitting and aggregating to obtain the initial number aggregation subset result; where in each iterative operation, output the initial number aggregation subset that has no intersection with other initial number aggregation subsets after aggregation, and use the remaining initial number aggregation subsets as the objects for splitting and aggregating in the next iterative operation; until the remaining initial number aggregation subsets cannot be aggregated anymore in an iterative operation, output the remaining initial number aggregation subsets, terminate the iterative operation, and integrate the initial number aggregation subsets output in each iterative operation to obtain the initial number aggregation subset result.
[0013] Optionally, step S2 specifically includes:
[0014] Split the initial serial number - ID pair relationship table into an initial serial number - ID relationship table;
[0015] Aggregate the initial serial number - ID relationship table with ID as the key to obtain multiple initial serial number first - level aggregation subsets, where each initial serial number first - level aggregation subset is composed of initial serial numbers;
[0016] Renumber the initial serial number first - level aggregation subsets to obtain a secondary serial number - initial serial number first - level aggregation subset relationship table.
[0017] Optionally, after aggregating to obtain multiple initial serial number first - level aggregation subsets, step S2 further includes:
[0018] Determine whether each subset in the multiple initial serial number first - level aggregation subsets is an isolated subset that has no intersection with other initial serial number first - level aggregation subsets;
[0019] If so, output the initial serial number first - level aggregation subset as the target initial serial number aggregation subset, and renumber the remaining initial serial number first - level aggregation subsets.
[0020] Optionally, determining whether each subset in the multiple initial serial number first - level aggregation subsets is an isolated subset that has no intersection with other initial serial number first - level aggregation subsets includes:
[0021] Count the occurrence times and the number of elements contained in each initial serial number first - level aggregation subset;
[0022] Judge the initial serial number first - level aggregation subset with an occurrence time of 2 and an element number of 1 as an isolated subset.
[0023] Optionally, step S3 specifically includes:
[0024] Step S31: Split the secondary serial number - initial serial number first - level aggregation subset relationship table into a secondary serial number - initial serial number - initial serial number first - level aggregation subset relationship table;
[0025] Step S32: Aggregate the secondary serial number - initial serial number - initial serial number first - level aggregation subset relationship table with the initial serial number as the key to obtain one or more initial serial number second - level aggregation subsets;
[0026] Step S33: Filter and output the initial serial number second - level aggregation subsets that have no intersection with other initial serial number second - level aggregation subsets as the target initial serial number aggregation subsets;
[0027] Step S34: Remove duplicates from the remaining initial serial number second - level aggregation subsets and renumber them to obtain a tertiary serial number - initial serial number second - level aggregation subset relationship table;
[0028] Step S35: And so on, repeat steps S31 to S34 for n iterative operations until the number of the remaining initial-numbered (n + 2)-th aggregated subsets is 0 or 1, and output the remaining initial-numbered (n + 2)-th aggregated subsets as the target initial-numbered aggregated subsets, where n is a natural number;
[0029] Step S36: Integrate the target initial-numbered aggregated subsets output in the foregoing steps to obtain the result of the initial-numbered aggregated subsets, and number the result of the initial-numbered aggregated subsets using a unified identifier to obtain a unified-identifier - initial-numbered aggregated subset relationship table.
[0030] Optionally, step S32 specifically includes:
[0031] Aggregate the secondary-number - initial-number - initial-number first-aggregated subset relationship table with the initial number as the key to obtain one or more initial-number second-aggregated subsets and corresponding secondary-number aggregated subsets, where each secondary-number aggregated subset is composed of secondary numbers, and each initial-number second-aggregated subset is formed by merging the initial-number first-aggregated subsets corresponding to each secondary number in its corresponding secondary-number aggregated subset;
[0032] Step S33 specifically includes:
[0033] Determine whether each subset in the one or more secondary-number aggregated subsets is an isolated subset that has no intersection with other secondary-number aggregated subsets. If so, output the initial-number second-aggregated subset corresponding to this secondary-number aggregated subset after deduplication as the target initial-number aggregated subset.
[0034] Optionally, determining whether each subset in the one or more secondary-number aggregated subsets is an isolated subset that has no intersection with other secondary-number aggregated subsets includes:
[0035] Count the number of elements included in each secondary-number aggregated subset;
[0036] Determine the secondary-number aggregated subset with the number of elements being 1 as an isolated subset.
[0037] Optionally, after counting the number of elements included in each secondary-number aggregated subset, step S33 further includes:
[0038] Determine whether the number of elements included in each secondary-number aggregated subset is greater than a given threshold;
[0039] If so, output the initial-number second-aggregated subset corresponding to this secondary-number aggregated subset after deduplication as the target initial-number aggregated subset.
[0040] Optionally, step S4 specifically includes:
[0041] Split the unified identifier - initial number aggregation subset relationship table into an initial number - unified identifier relationship table;
[0042] Obtain a unified identifier - ID relationship table based on the initial number - unified identifier relationship table and the initial number - ID pair relationship table;
[0043] Aggregate the unified identifier - ID relationship table with the unified identifier as the key to obtain a unified representation table of IDs.
[0044] Optionally, obtaining a unified identifier - ID relationship table based on the initial number - unified identifier relationship table and the initial number - ID pair relationship table includes:
[0045] Execute the leftOuterJoin command on the initial number - unified identifier relationship table and the initial number - ID pair relationship table, and then separate the unified identifier and ID through the map command to obtain the unified identifier - ID relationship table.
[0046] Optionally, obtaining a two - dimensional ID relationship table including multiple ID pairs includes:
[0047] Integrate multiple two - dimensional ID relationship source data tables into the two - dimensional ID relationship table including multiple ID pairs.
[0048] Optionally, the numbering operation is performed through the zipWithUniqueId command.
[0049] Optionally, the aggregation operation is performed through the reduceByKey command.
[0050] According to another aspect of the embodiments of the present invention, there is also provided a device for implementing ID mapping based on the Spark framework, including:
[0051] A data pre - processing module, adapted to obtain a two - dimensional ID relationship table including multiple ID pairs, number each ID pair to obtain an initial number - ID pair relationship table;
[0052] An ID relationship aggregation module, adapted to split and aggregate the initial number - ID pair relationship table with the ID as the key to obtain multiple initial number first - level aggregation subsets, where each initial number first - level aggregation subset is composed of initial numbers;
[0053] The number relationship aggregation module is adapted to use the initial number as the key to split and aggregate the multiple initial number one-time aggregation subsets, obtaining the initial number aggregation subset result, where there is no intersection between any two initial number aggregation subsets in the initial number aggregation subset result; number the initial number aggregation subset result using a unified identifier to obtain a unified identifier - initial number aggregation subset relationship table; and
[0054] The ID unified representation module is adapted to obtain the corresponding relationship between the unified identifier and the ID according to the unified identifier - initial number aggregation subset relationship table and the initial number - ID pair relationship table, realizing the unified representation of the ID.
[0055] Optionally, the number relationship aggregation module is further adapted to:
[0056] Use the initial number as the key, and use the multiple initial number one-time aggregation subsets as the objects for splitting and aggregating in the first iterative operation, performing iterative operations of splitting and aggregating to obtain the initial number aggregation subset result; where in each iterative operation, output the initial number aggregation subsets that have no intersection with other initial number aggregation subsets after aggregation, and use the remaining initial number aggregation subsets as the objects for splitting and aggregating in the next iterative operation; until the remaining initial number aggregation subsets cannot be aggregated any more in an iterative operation, output the remaining initial number aggregation subsets, terminate the iterative operation, and integrate the initial number aggregation subsets output in each iterative operation to obtain the initial number aggregation subset result.
[0057] Optionally, the ID relationship aggregation module includes:
[0058] The first splitting unit is adapted to split the initial number - ID pair relationship table into an initial number - ID relationship table;
[0059] The first aggregation unit is adapted to use the ID as the key to aggregate the initial number - ID relationship table, obtaining multiple initial number one-time aggregation subsets, where each initial number one-time aggregation subset is composed of initial numbers; and
[0060] The first numbering unit is adapted to re-number the initial number one-time aggregation subsets to obtain a secondary number - initial number one-time aggregation subset relationship table.
[0061] Optionally, the ID relationship aggregation module further includes:
[0062] The first filtering and output unit is adapted to, after the first aggregation unit aggregates to obtain multiple initial number one-time aggregation subsets, determine whether each subset in the multiple initial number one-time aggregation subsets is an isolated subset that has no intersection with other initial number one-time aggregation subsets;
[0063] If so, output the initial numbered first-level aggregated subset as the target initial numbered aggregated subset, and trigger the first numbering unit to renumber the remaining initial numbered first-level aggregated subsets.
[0064] Optionally, the first filtering and output unit is further adapted to:
[0065] Count the occurrence times and the number of elements included in each initial numbered first-level aggregated subset;
[0066] Determine the initial numbered first-level aggregated subset with an occurrence time of 2 and a number of elements of 1 as an isolated subset.
[0067] Optionally, the numbering relationship aggregation module includes:
[0068] A second splitting unit, adapted to split the secondary numbering - initial numbered first-level aggregated subset relationship table into a secondary numbering - initial number - initial numbered first-level aggregated subset relationship table;
[0069] A second aggregation unit, adapted to aggregate the secondary numbering - initial number - initial numbered first-level aggregated subset relationship table with the initial number as the key, to obtain one or more initial numbered second-level aggregated subsets;
[0070] A second filtering and output unit, adapted to filter and output the initial numbered second-level aggregated subset that has no intersection with other initial numbered second-level aggregated subsets as the target initial numbered aggregated subset;
[0071] A second numbering unit, adapted to renumber the remaining initial numbered second-level aggregated subsets to obtain a tertiary numbering - initial numbered second-level aggregated subset relationship table;
[0072] An analog iterative unit, adapted to trigger the second splitting unit, the second aggregation unit, the second filtering and output unit, and the second numbering unit to perform n times of iterative operations by analogy until the number of the remaining initial numbered (n + 2)-level aggregated subsets is 0 or 1, and output the remaining initial numbered (n + 2)-level aggregated subsets as the target initial numbered aggregated subset, where n is a natural number; and
[0073] An output result integration unit, adapted to integrate the foregoing output target initial numbered aggregated subsets to obtain the initial numbered aggregated subset result, and number the initial numbered aggregated subset result with a unified identifier to obtain a unified identifier - initial numbered aggregated subset relationship table.
[0074] Optionally, the second aggregation unit is further adapted to:
[0075] Using the initial number as the key, aggregate the secondary number - initial number - initial number first - level aggregation subset relationship table to obtain one or more initial number secondary aggregation subsets and corresponding secondary number aggregation subsets, where each secondary number aggregation subset is composed of secondary numbers, and each initial number secondary aggregation subset is formed by merging the initial number first - level aggregation subsets corresponding to each secondary number in its corresponding secondary number aggregation subset;
[0076] The second filtering and output unit is further adapted to:
[0077] Determine whether each subset in the one or more secondary number aggregation subsets is an isolated subset that has no intersection with other secondary number aggregation subsets. If so, perform deduplication on the initial number secondary aggregation subset corresponding to this secondary number aggregation subset and output it as the target initial number aggregation subset.
[0078] Optionally, the second filtering and output unit is further adapted to:
[0079] Count the number of elements contained in each secondary number aggregation subset;
[0080] Judge the secondary number aggregation subset with the number of elements being 1 as an isolated subset.
[0081] Optionally, the second filtering and output unit is further adapted to:
[0082] After counting the number of elements contained in each secondary number aggregation subset, determine whether the number of elements contained in each secondary number aggregation subset is greater than a given threshold;
[0083] If so, perform deduplication on the initial number secondary aggregation subset corresponding to this secondary number aggregation subset and output it as the target initial number aggregation subset.
[0084] Optionally, the ID unified representation module includes:
[0085] A third splitting unit, adapted to split the unified identifier - initial number aggregation subset relationship table into an initial number - unified identifier relationship table;
[0086] A relationship connection unit, adapted to obtain a unified identifier - ID relationship table according to the initial number - unified identifier relationship table and the initial number - ID pair relationship table; and
[0087] A third aggregation unit, adapted to aggregate the unified identifier - ID relationship table with the unified identifier as the key to obtain an ID unified representation table.
[0088] Optionally, the relationship connection unit is further adapted to:
[0089] Execute the leftOuterJoin command on the initial number-unified identifier relationship table and the initial number-ID pair relationship table, and then separate the unified identifier and the ID through the map command to obtain the unified identifier-ID relationship table.
[0090] Optionally, the data preprocessing module is further adapted to:
[0091] Integrate multiple two-dimensional ID relationship source data tables into the two-dimensional ID relationship table including multiple ID pairs.
[0092] Optionally, the numbering operation is performed through the zipWithUniqueId command.
[0093] Optionally, the aggregation operation is performed through the reduceByKey command.
[0094] According to another aspect of the embodiments of the present invention, there is also provided a computer storage medium storing computer program code, which, when run on a computing device, causes the computing device to execute the method for implementing ID mapping based on the Spark framework described in any one of the above.
[0095] According to yet another aspect of the embodiments of the present invention, there is also provided a computing device, including:
[0096] A processor; and
[0097] A memory storing computer program code;
[0098] When the computer program code is run by the processor, it causes the computing device to execute the method for implementing ID mapping based on the Spark framework described in any one of the above.
[0099] The method and device for implementing ID mapping based on the Spark framework proposed in the embodiments of the present invention, after numbering each ID pair in the obtained two-dimensional ID relationship table to obtain the initial numbering-ID pair relationship table, first splits and aggregates the ID relationships in the initial numbering-ID pair relationship table with the ID as the key to obtain multiple initial numbering first-aggregation subsets, and then splits and aggregates the numbering relationships of the multiple initial numbering first-aggregation subsets with the initial number as the key to obtain the initial numbering aggregation subset result. Furthermore, the initial numbering aggregation subset result is numbered with a unified identifier to obtain the unified identifier-initial numbering aggregation subset relationship table. Finally, the corresponding relationship between the unified identifier and the ID is obtained according to the unified identifier-initial numbering aggregation subset relationship table and the initial numbering-ID pair relationship table, thereby realizing the unified representation of the user ID. The present invention implements the ID mapping algorithm based on the Spark distributed computing framework, and uses the idea of set theory in mathematics to implement operations such as storage, filtering, splitting, and aggregation of a large amount of user data sets, thereby improving the efficiency, accuracy, and reliability of the ID mapping algorithm.
[0100] Further, the process of aggregating the numbering relationships with the initial number as the key is implemented through iterative operations. By outputting, in each iterative operation, the initial numbering aggregation subset that has no intersection with other initial numbering aggregation subsets after aggregation (that is, no further aggregation is required), the number of remaining subsets (that is, the subsets to be split and aggregated in the next iterative operation) in each iterative operation can be made smaller and smaller, thereby further reducing the memory overhead and improving the operation efficiency.
[0101] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below.
[0102] According to the following detailed description of the specific embodiments of the present invention in conjunction with the drawings, those skilled in the art will become more clear about the above and other purposes, advantages, and features of the present invention. Description of the Drawings
[0103] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0104] Figure 1 Shows a flowchart of a method for implementing ID mapping based on the Spark framework according to an embodiment of the present invention;
[0105] Figure 2 The flowchart of the ID relationship aggregation step of the method for implementing ID mapping based on the Spark framework according to an embodiment of the present invention is shown;
[0106] Figure 3 The flowchart of the ID relationship aggregation step of the method for implementing ID mapping based on the Spark framework according to another embodiment of the present invention is shown;
[0107] Figure 4 The flowchart of the number relationship aggregation step of the method for implementing ID mapping based on the Spark framework according to an embodiment of the present invention is shown;
[0108] Figure 5 The flowchart of the number relationship aggregation step of the method for implementing ID mapping based on the Spark framework according to another embodiment of the present invention is shown;
[0109] Figure 6 The flowchart of the ID unified representation step of the method for implementing ID mapping based on the Spark framework according to an embodiment of the present invention is shown;
[0110] Figure 7 The structural schematic diagram of the device for implementing ID mapping based on the Spark framework according to an embodiment of the present invention is shown; and
[0111] Figure 8 The structural schematic diagram of the device for implementing ID mapping based on the Spark framework according to another embodiment of the present invention is shown. Detailed Embodiment
[0112] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0113] Currently, when facing a large amount of user data, a large number and types of IDs, and complex relationships between IDs, ID Mapping cannot effectively extract the ID relationship network, making it difficult to effectively implement in engineering.
[0114] The Spark framework is a fast and general-purpose cluster computing platform designed for large-scale data processing. It enables in-memory distributed datasets and has great advantages in processing massive data. Therefore, based on the Spark framework, using a distributed computing system to implement operations such as storage, filtering, splitting, and merging of massive data sets is expected to efficiently implement the ID Mapping process.
[0115] To solve the above technical problems, an embodiment of the present invention proposes a method for implementing ID mapping based on the Spark framework. Figure 1 The flowchart of the method for implementing ID mapping based on the Spark framework according to an embodiment of the present invention is shown. Refer to Figure 1 This method may at least include the following steps S1 to S4.
[0116] Step S1, data preprocessing: Obtain a two-dimensional ID relationship table including multiple ID pairs, number each ID pair, and obtain an initial number - ID pair relationship table.
[0117] Step S2, ID relationship aggregation: Using the ID as the key, split and aggregate the initial number - ID pair relationship table to obtain multiple initial number first-aggregation subsets, where each initial number first-aggregation subset is composed of initial numbers.
[0118] Step S3, number relationship aggregation: Using the initial number as the key, split and aggregate multiple initial number first-aggregation subsets to obtain an initial number aggregation subset result, where there is no intersection between any two initial number aggregation subsets in the initial number aggregation subset result; use a unified identifier to number the initial number aggregation subset result to obtain a unified identifier - initial number aggregation subset relationship table.
[0119] Step S4, unified ID representation: According to the unified identifier - initial number aggregation subset relationship table and the initial number - ID pair relationship table, obtain the corresponding relationship between the unified identifier and the ID, and realize the unified representation of the ID.
[0120] The method for implementing ID mapping based on the Spark framework proposed by the embodiment of the present invention uses the idea of set theory in mathematics to implement operations such as storage, filtering, splitting, and aggregation of massive user data sets, thereby improving the efficiency, accuracy, and reliability of the ID mapping algorithm.
[0121] In the above step S1, the two-dimensional ID relationship table including multiple ID pairs can be obtained by integrating multiple two-dimensional ID relationship source data tables obtained from various channels. Among them, each two-dimensional ID relationship source data table may include multiple user IDs that appear in pairs.
[0122] After obtaining the two-dimensional ID relationship table, a corresponding RDD (Resilient Distributed Datasets) is created for this two-dimensional ID relationship table as the operation object of the Spark framework.
[0123] Preferably, the zipWithUniqueId command of the Spark framework can be used to number each ID pair in the two-dimensional ID relationship table, so as to generate a unique initial number for each ID pair.
[0124] In the above step S2, with the ID as the key, the initial number - ID pair relationship table is split and the ID relationships are aggregated.
[0125] In an alternative embodiment, as Figure 2 shown, step S2 may include the following steps:
[0126] Step S21: Split the initial number - ID pair relationship table into an initial number - ID relationship table.
[0127] Specifically, in the manner of splitting the corresponding relationship between each ID pair and the initial number into the correspondence between each ID in each ID pair and the initial number of this ID pair respectively, the initial number - ID pair relationship table is split into an initial number - ID relationship table.
[0128] Step S22: Aggregate the initial number - ID relationship table with the ID as the key to obtain multiple initial number first-aggregation subsets, where each initial number first-aggregation subset is composed of initial numbers.
[0129] Preferably, with the ID as the key, the reduceByKey command of the Spark framework is used to aggregate the initial number - ID relationship table.
[0130] Step S23: Renumber the multiple initial number first-aggregation subsets to obtain a secondary number - initial number first-aggregation subset relationship table.
[0131] Preferably, the zipWithUniqueId command of the Spark framework can be used to renumber the multiple initial number first-aggregation subsets.
[0132] In a preferred embodiment, as Figure 3 shown, after performing step S22 to aggregate the initial number - ID relationship table with the ID as the key to obtain multiple initial number first-aggregation subsets, step S2 may further include:
[0133] Step S24: Determine whether each subset in the multiple initial number first-aggregation subsets is an isolated subset that has no intersection with other initial number first-aggregation subsets;
[0134] If so, output the initial number first aggregation subset as the target initial number aggregation subset.
[0135] At this time, step S23 is correspondingly adjusted to:
[0136] Renumber the remaining initial number first aggregation subsets to obtain a secondary number - initial number first aggregation subset relationship table.
[0137] Furthermore, the operation of determining whether each subset in multiple initial number first aggregation subsets is an isolated subset that has no intersection with other initial number first aggregation subsets in step S24 above can be implemented in the following way:
[0138] Count the occurrence times and the number of elements contained in each subset in the multiple initial number first aggregation subsets;
[0139] If the occurrence times of a certain initial number first aggregation subset is 2 and the number of elements contained is 1, then determine that this initial number first aggregation subset is an isolated subset.
[0140] Since the isolated subset has no intersection with other subsets and there is no need to perform aggregation anymore, it can be directly output. By filtering and outputting the isolated initial number first aggregation subsets, the computational amount of subsequent splitting and aggregation can be reduced, and the memory overhead can be reduced.
[0141] In step S3 above, the initial number can be used as the key, and the reduceByKey command of the Spark framework can be used to aggregate multiple initial number first aggregation subsets.
[0142] Preferably, iterative operations are used to implement the splitting and aggregation of the number relationship for multiple initial number first aggregation subsets with the initial number as the key. At this time, step S3 can be implemented in the following way:
[0143] Using the initial number as the key and taking multiple initial number first aggregation subsets as the objects for splitting and aggregation in the first iterative operation, perform iterative operations of splitting and aggregation to obtain the initial number aggregation subset result. Among them, in each iterative operation, output the initial number aggregation subset that has no intersection with other initial number aggregation subsets after aggregation, and use the remaining initial number aggregation subsets as the objects for splitting and aggregation in the next iterative operation. Until the remaining initial number aggregation subsets cannot be aggregated anymore in an iterative operation, output the remaining initial number aggregation subsets, terminate the iterative operation, and integrate the initial number aggregation subsets output in each iterative operation to obtain the initial number aggregation subset result. In this way, there is no intersection between any two initial number aggregation subsets in the obtained initial number aggregation subset result.
[0144] By outputting, in each iteration operation, the initially numbered aggregated subset that has no intersection with other initially numbered aggregated subsets after aggregation (i.e., no further aggregation is required), the number of remaining subsets in each iteration operation (i.e., the subsets to be split and aggregated in the next iteration operation) can be made smaller and smaller, thereby further reducing memory overhead and improving operation efficiency.
[0145] In an optional embodiment, as Figure 4 shown, step S3 may specifically include the following steps:
[0146] Step S31: Split the secondary number - initially numbered first - aggregated subset relationship table obtained in step S23 into a secondary number - initially numbered - initially numbered first - aggregated subset relationship table.
[0147] Specifically, in the manner of splitting the correspondence between each initially numbered first - aggregated subset and the secondary number into the correspondence between each initially numbered in each initially numbered first - aggregated subset and the secondary number of this initially numbered first - aggregated subset respectively, split the secondary number - initially numbered first - aggregated subset relationship table into a secondary number - initially numbered - initially numbered first - aggregated subset relationship table.
[0148] Step S32: Aggregate the secondary number - initially numbered - initially numbered first - aggregated subset relationship table with the initially numbered as the key to obtain one or more initially numbered second - aggregated subsets.
[0149] Preferably, with the initially numbered as the key, use the reduceByKey command of the Spark framework to aggregate the secondary number - initially numbered - initially numbered first - aggregated subset relationship table.
[0150] Step S33: Filter and output the initially numbered second - aggregated subsets that have no intersection with other initially numbered second - aggregated subsets as the target initially numbered aggregated subsets.
[0151] Step S34: Remove duplicates from the remaining initially numbered second - aggregated subsets and re - number them to obtain a tertiary number - initially numbered second - aggregated subset relationship table.
[0152] Preferably, use the zipWithUniqueId command of the Spark framework to re - number the remaining initially numbered second - aggregated subsets after removing duplicates.
[0153] Step S35: And so on, repeat steps S31 to S34 for n iteration operations until the number of remaining initially numbered n + 2 - aggregated subsets is 0 or 1, and output the remaining initially numbered n + 2 - aggregated subsets as the target initially numbered aggregated subsets, where n is a natural number.
[0154] Step S36: Integrate the target initial number aggregation subsets output from the foregoing steps to obtain the initial number aggregation subset result, and number the initial number aggregation subset result using a unified identifier to obtain a unified identifier - initial number aggregation subset relationship table.
[0155] In this step, the unified identifier may be denoted as dmid for use as the unified number for the obtained initial number aggregation subset result.
[0156] In a preferred embodiment, as Figure 5 shown, step S32 can be further implemented as:
[0157] Aggregate the secondary number - initial number - initial number first aggregation subset relationship table with the initial number as the key to obtain one or more initial number secondary aggregation subsets and corresponding secondary number aggregation subsets, where each secondary number aggregation subset is composed of secondary numbers, and each initial number secondary aggregation subset is formed by merging the initial number first aggregation subsets corresponding to each secondary number in its corresponding secondary number aggregation subset.
[0158] Correspondingly, step S33 can be further implemented as:
[0159] Determine whether each subset in one or more secondary number aggregation subsets is an isolated subset that has no intersection with other secondary number aggregation subsets. If so, perform deduplication on the initial number secondary aggregation subset corresponding to the secondary number aggregation subset and output it as the target initial number aggregation subset.
[0160] Furthermore, the operation of determining whether each subset in one or more secondary number aggregation subsets is an isolated subset that has no intersection with other secondary number aggregation subsets in step S33 can be implemented in the following manner:
[0161] Count the number of elements contained in each secondary number aggregation subset;
[0162] If the number of elements in a certain secondary number aggregation subset is 1, then determine that the secondary number aggregation subset is an isolated subset.
[0163] When a secondary number aggregation subset contains only 1 element, the initial number secondary aggregation subset corresponding to the secondary number aggregation subset cannot be aggregated (or merged) with other initial number secondary aggregation subsets anymore. Therefore, the initial number secondary aggregation subset corresponding to the secondary number aggregation subset can be directly output after deduplication.
[0164] Even further, after counting the number of elements contained in each secondary number aggregation subset in step S33, the following steps can also be performed:
[0165] Determine whether the number of elements included in each secondary number aggregation subset is greater than a given threshold;
[0166] If so, after removing duplicates from the initial number secondary aggregation subset corresponding to the secondary number aggregation subset, output it as the target initial number aggregation subset.
[0167] In this step, the given threshold can be set to 20, 30, 50, etc. according to actual needs. By introducing the "threshold" of the number of elements in the subset, the number of IDs owned by each user can be restricted, and at the same time, data skew in the operation can be prevented, resulting in an excessive number of IDs in a certain computing node and memory overflow.
[0168] In the above step S4, according to the unified identifier-initial number aggregation subset relationship table obtained in step S3 and the initial number-ID pair relationship table obtained in step S1, the mapping result of the final user ID is obtained.
[0169] In an alternative embodiment, as Figure 6 shown, step S4 may specifically include the following steps:
[0170] Step S41: Split the unified identifier-initial number aggregation subset relationship table obtained in step S3 into an initial number-unified identifier relationship table.
[0171] Specifically, according to the correspondence between each initial number aggregation subset and the unified identifier, split it into the correspondence between each initial number in each initial number aggregation subset and the unified identifier of the initial number aggregation subset, and split the unified identifier-initial number aggregation subset relationship table into an initial number-unified identifier relationship table.
[0172] Step S42: According to the initial number-unified identifier relationship table and the initial number-ID pair relationship table, obtain a unified identifier-ID relationship table.
[0173] Preferably, by executing the leftOuterJoin command on the initial number-unified identifier relationship table and the initial number-ID pair relationship table, and then separating the unified identifier and ID through the map command, a unified identifier-ID relationship table is obtained.
[0174] Step S43: Aggregate the unified identifier-ID relationship table with the unified identifier as the key to obtain a unified representation table of IDs.
[0175] Preferably, with the unified identifier as the key, the reduceByKey command of the Spark framework can be used to aggregate the unified identifier-ID relationship table.
[0176] The above has introduced Figure 1For various implementation methods of each link of the illustrated embodiment, the implementation process of the method for implementing ID mapping based on the Spark framework of the present invention will be introduced in detail through specific embodiments below.
[0177] In this specific embodiment, it is assumed that all IDs in the source data include: imei (International Mobile Equipment Identity), aid (Android ID), sn (Serial Number), mac (Media Access Control), and tel (Telephone). These IDs are all obtained from various channels, and they appear in pairs, constituting the entire two-dimensional ID relationship source data table. For 5 IDs, there are at most source data tables). For simplicity, in this embodiment, natural numbers are used to represent the values of these IDs. Then, it is assumed that all two-dimensional ID relationship source data tables are as follows:
[0178]
[0179] The method for implementing ID mapping based on the Spark framework according to the specific embodiment of the present invention includes the following steps:
[0180] (1) Data preprocessing: Integrate multiple two-dimensional ID relationship source data tables into a two-dimensional ID relationship table including multiple ID pairs, create a corresponding RDD for this two-dimensional ID relationship table, and use the zipWithUniqueId command to number each ID pair in this two-dimensional ID relationship table to obtain an initial numbered-ID pair relationship table. The initial numbered-ID pair relationship table includes three columns, one of which is the initial number, and the other two are IDs, representing the relationship between the initial number and the ID.
[0181] Specifically, in this embodiment, the source data tables of Table 1 to Table 6 above are preprocessed to obtain the initial numbered-ID pair relationship table in the form of (num1, id1, id2) as shown in Table 7 below, where the num1 column is the initial number, and the id1 column and the id2 column are the two IDs in each ID relationship pair.
[0182] Table 7 Initial numbered-ID pair relationship table
[0183] num1 id1 id2 1 imei_1 aid_1 2 imei_1 sn_1 3 imei_2 sn_3 4 imei_3 tel_1 5 aid_2 sn_3 6 aid_1 mac_1 7 sn_2 tel_1
[0184] (2) Aggregation of ID relationships: Split the serialized initial numbered-ID pair relationship table as shown in Table 7, and use the reduceByKey command to aggregate with the ID as the key.
[0185] Specifically, the aggregation step of ID relationships includes the following sub-steps:
[0186] (2.1) Split the initial number - ID pair relationship table: Split the initial number - ID pair relationship table obtained in step (1) into an initial number - ID relationship table. In fact, it disassembles the two IDs in each row of Table 7 and splits one row of data into two rows of data. After splitting the initial number - ID pair relationship table shown in Table 7, the initial number - ID relationship table shown in Table 8 below is obtained.
[0187] Table 8 Initial number - ID relationship table
[0188]
[0189]
[0190] (2.2) Aggregate with ID as the key: Aggregate the initial number - ID relationship table with ID as the key using the reduceByKey command to obtain multiple initial number first - level aggregation subsets, where each initial number first - level aggregation subset is composed of initial numbers.
[0191] In this embodiment, after aggregating the initial number - ID relationship table shown in Table 8 with ID as the key using the reduceByKey command, the ID - initial number first - level aggregation subset relationship table shown in Table 9 below is obtained. Among them, the num1sets column in Table 9 shows multiple initial number first - level aggregation subsets obtained after aggregation.
[0192] Table 9 ID - initial number first - level aggregation subset relationship table
[0193] id num1 sets imei_1 1,2 aid_1 1,6 sn_1 2 imei_2 3 sn_3 3,5 imei_3 4 tel_1 4,7 aid_2 5 mac_1 6 sn_2 7
[0194] (2.3) Filter and output isolated subsets: Use the Spark command to parallelly judge, filter, and output the isolated subsets in the multiple initial number first - level aggregation subsets obtained in step (2.2).
[0195] The method for determining whether each subset in multiple initial number first - level aggregation subsets is an isolated subset is as follows: Distributively count the occurrence times and the number of elements included in each subset in multiple initial number first - level aggregation subsets. If the number of elements in a certain initial number first - level aggregation subset is 1 and it appears 2 times, then determine that this initial number first - level aggregation subset is an isolated subset. The isolated initial number first - level aggregation subset has no intersection with other initial number first - level aggregation subsets and can be directly output as the target initial number aggregation subset. The reason for being able to use this method to determine the isolated initial number first - level aggregation subset is as follows: Taking Table 8 as an example, in the num1 column of Table 8, each initial number appears twice. If a certain initial number has no intersection with other initial numbers, then after aggregating using reduceByKey with the ID as the key, this initial number alone forms an initial number first - level aggregation subset after aggregation, and this initial number first - level aggregation subset will appear twice.
[0196] In this embodiment, the occurrence times of each initial number first - level aggregation subset in the num1 sets column of Table 9 are counted. It is found that each initial number first - level aggregation subset appears only once. Therefore, there are no isolated subsets, and all initial number first - level aggregation subsets can continue to be aggregated (or merged).
[0197] (2.4) Renumber the remaining initial number first - level aggregation subsets: Use the zipWithUniqueId command to renumber the remaining initial number first - level aggregation subsets to obtain a secondary number - initial number first - level aggregation subset relationship table.
[0198] In this embodiment, since there are no isolated subsets in the multiple initial number first - level aggregation subsets shown in Table 9, the remaining initial number first - level aggregation subsets are all the initial number first - level aggregation subsets in Table 9. After renumbering the initial number first - level aggregation subsets in the num1 sets column of Table 9, the following secondary number - initial number first - level aggregation subset relationship table as shown in Table 10 is obtained. Among them, the num2 column in Table 10 represents the secondary number.
[0199] Table 10 Secondary number - initial number first - level aggregation subset relationship table
[0200] num2 num1 sets 1 1,2 2 1,6 3 2 4 3 5 3,5 6 4 7 4,7 8 5 9 6 10 7
[0201] (3) Aggregation of numbering relationships: The aggregation result of Table 10 above is split and aggregated, and this step is achieved through iteration. In each iteration step, first split the result aggregated in the previous step, then use the reduceByKey command to aggregate with the initial number as the key, output the aggregated isolated subsets, and then split and aggregate the remaining subsets. This cycle continues until there are no remaining subsets to output and the iteration terminates. Integrate the output results of each step, and uniformly number the integrated output results with a unified identifier to obtain the aggregation result (that is, the unified identifier - initial number aggregation subset relationship table, which actually represents the relationship between the unified identifier and the initial number).
[0202] Specifically, the aggregation steps of the numbering relationship include the following sub-steps:
[0203] (3.1) Split the secondary numbering - initial number first - aggregation subset relationship table: Split each row of data in the secondary numbering - initial number first - aggregation subset relationship table shown in Table 10 obtained in step (2.4) into multiple rows of data. Specifically, split the data in the num1 sets column in Table 10 and place it in the first column of the new table (as the key) to obtain the secondary numbering - initial number - initial number first - aggregation subset relationship table shown in Table 11 below. Among them, the num1 column in Table 11 represents the initial numbers split from the num1 sets column data in Table 10, and the num2 and num1 sets columns respectively represent the corresponding secondary numbers and the initial number first - aggregation subsets.
[0204] Table 11 Secondary numbering - initial number - initial number first - aggregation subset relationship table
[0205] num1 num2 num1 sets 1 1 1,2 2 1 1,2 1 2 1,6 6 2 1,6 2 3 2 3 4 3 3 5 3,5 5 5 3,5 4 6 4 4 7 4,7 7 7 4,7 5 8 5 6 9 6 7 10 7
[0206] (3.2) Aggregate with the initial number as the key: Use the reduceByKey command to aggregate the secondary numbering - initial number - initial number first - aggregation subset relationship table with the initial number as the key to obtain multiple initial number second - aggregation subsets and the corresponding secondary number aggregation subsets. Among them, each secondary number aggregation subset is composed of secondary numbers, and each initial number second - aggregation subset is formed by merging the initial number first - aggregation subsets corresponding to each secondary number in the corresponding secondary number aggregation subset.
[0207] In this embodiment, after aggregating Table 11, the initial number - secondary number aggregation subset - initial number second - aggregation subset relationship table shown in Table 12 below is obtained. Among them, the data in the num2 sets and num1 sets columns are respectively multiple secondary number aggregation subsets and the corresponding initial number second - aggregation subsets obtained after aggregation.
[0208] Table 12 Initial Number - Secondary Number Aggregation Subset - Initial Number Secondary Aggregation Subset Relationship Table
[0209] num1 num2 sets num1 sets 1 1,2 1,2,6 2 1,3 1,2 3 4,5 3,5 4 6,7 4,7 5 5,8 3,5 6 2,9 1,6 7 7,10 4,7
[0210] (3.3) Filter and output the sets that do not need to be further aggregated: Use Spark commands to parallelly count the number of elements in each of the multiple secondary number aggregation subsets (i.e., subsets of the num2 sets column in Table 12) obtained after aggregation in step (3.2). If the number of elements in a secondary number aggregation subset is 1, then determine that this secondary number aggregation subset is an isolated subset. The isolated secondary number aggregation subset has no intersection with other secondary number aggregation subsets, and the initial number secondary aggregation subset corresponding to this isolated secondary number aggregation subset is also an isolated subset (i.e., has no intersection with other initial number secondary aggregation subsets). Therefore, the initial number secondary aggregation subset corresponding to the secondary number aggregation subset determined to be an isolated subset can be de-duplicated and output as the target initial number aggregation subset. Further, after counting the number of elements in each secondary number aggregation subset, it is also determined whether the number of elements in each secondary number aggregation subset is greater than a given threshold (such as 50). If so, the initial number secondary aggregation subset corresponding to this secondary number aggregation subset is de-duplicated and output as the target initial number aggregation subset.
[0211] In this embodiment, there are no isolated subsets in the num2 sets column of Table 12.
[0212] (3.4) Re-number the remaining sets: Use Spark commands to de-duplicate the remaining initial number secondary aggregation subsets, and then use the zipWithUniqueId command to re-number them to obtain a three - number - initial number secondary aggregation subset relationship table.
[0213] In this embodiment, since there are no isolated subsets in the num2 sets column of Table 12, the remaining initial number secondary aggregation subsets are all the data in the num2 sets column. After de-duplicating and re-numbering the data in the num2 sets column of Table 12, the three - number - initial number secondary aggregation subset relationship table shown in Table 13 below is obtained. Among them, in Table 13, the num2 column is the three - number data (for the convenience of iteration, still denoted as num2), and the num1 sets column is the initial number secondary aggregation subset data.
[0214] Table 13 Three - number - Initial Number Secondary Aggregation Subset Relationship Table
[0215] num2 num1 sets 1 1,2,6 2 1,2 3 3,5 4 4,7 5 1,6
[0216] It should be noted that after steps (3.1)-(3.4), the number of initial number aggregation subsets has been reduced from 10 in Table 10 to 5 in Table 13.
[0217] (3.5) Analogy iteration: And so on, repeat steps (3.1) to (3.4) for n iterations until the number of remaining initial number n+2 times aggregation subsets is 0 or 1.
[0218] In this embodiment, the specific iterative operation is as follows:
[0219] Repeat step (3.1) for the three-number - initial number secondary aggregation subset relationship table shown in Table 13 to obtain the three-number - initial number - initial number secondary aggregation subset relationship table shown in Table 14 below. Among them, in Table 14, the num1, num2, and num1 sets columns are the initial number, three-number, and initial number secondary aggregation subset data respectively.
[0220] Table 14 Three-number - initial number - initial number secondary aggregation subset relationship table
[0221]
[0222]
[0223] Repeat step (3.2) for Table 14 to obtain the initial number - three-number aggregation subset - initial number three aggregation subset relationship table shown in Table 15 below. Among them, the data in the num2 sets and num1 sets columns are multiple three-number aggregation subsets obtained after aggregation and the corresponding initial number three aggregation subsets respectively.
[0224] Table 15 Initial number - three-number aggregation subset - initial number three aggregation subset relationship table
[0225] num1 num2 sets num1 sets 1 1,2,5 1,2,6 2 1,2 1,2,6 3 3 3,5 4 4 4,7 5 3 3,5 6 1,5 1,2,6 7 4 4,7
[0226] Execute step (3.3) on Table 15. Note that the three-number aggregation subsets {3}, {4} in the num2 sets column of Table 15 have only 1 element, so the initial number three aggregation subsets in the corresponding num1 sets column cannot be aggregated with other initial number three aggregation subsets anymore. Screen them out and output them to obtain the target initial number aggregation subset shown in Table 16 below.
[0227] Table 16 Target initial number aggregation subset
[0228] num2 sets num1 sets 3 3,5 4 4,7
[0229] Perform step (3.4) on the remaining initial number triple aggregation subsets in Table 15 to obtain the quadruple number-initial number triple aggregation subset relationship table as shown in Table 17 below. Among them, in Table 17, the num2 column is the quadruple number data (for the convenience of iteration, still denoted as num2), and the num1 sets column is the initial number triple aggregation subset data.
[0230] Table 17 Quadruple number-initial number triple aggregation subset relationship table
[0231] num2 num1 sets 1 1,2,6
[0232] In this way, only 1 initial number triple aggregation subset {1, 2, 6} remains in the num1 sets column, and no further aggregation is required. The loop ends, and this remaining initial number triple aggregation subset is output as the target initial number aggregation subset.
[0233] (3.6) Integrate the output results: Integrate the target initial number aggregation subsets output in steps (2.3), (3.3), and (3.5), and use a unified identifier (denoted by dmid) to number them uniformly to obtain the unified identifier-initial number aggregation subset relationship table as shown in Table 18 below. Among them, the dmid column is the unified identifier data, and the num1 sets column is the initial number aggregation subset data.
[0234] Table 18 Unified identifier-initial number aggregation subset relationship table
[0235] dmid num1 sets 1 1,2,6 2 3,5 3 4,7
[0236] (4) Unified representation of ID: According to the unified identifier-initial number aggregation subset relationship table obtained in step (3) and the initial number-ID pair relationship table obtained in step (1), obtain the corresponding relationship between the unified identifier and the ID, and realize the unified representation of the user ID.
[0237] Specifically, the steps for the unified representation of ID include the following sub-steps: ·
[0238] (4.1) Split the unified identifier-initial number aggregation subset relationship table: Disassemble the column of the initial number aggregation subset in the unified identifier-initial number aggregation subset relationship table shown in Table 18, so as to split one row of data in Table 18 into multiple rows of data, and split the (dmid, initial number) relationship pair into (initial number, dmid) relationship pairs, thereby obtaining the initial number-unified identifier relationship table as shown in Table 19 below. Among them, the num1 column in Table 19 is the initial number data, and the dmid column is the unified identifier data.
[0239] Table 19 Initial number-unified identifier relationship table
[0240] num1 dmid 1 1 2 1 6 1 3 2 5 2 4 3 7 3
[0241] (4.2) Obtain the unified identifier - ID relationship table based on the initial number - unified identifier relationship table obtained in step (4.1) and the initial number - ID pair relationship table obtained in step (1).
[0242] In this embodiment, by executing the leftOuterJoin command on the initial number - unified identifier relationship table shown in Table 19 and the initial number - ID pair relationship table shown in Table 7, and then distributively separating the two fields of dmid and ID through the map command, the unified identifier - ID relationship table shown in Table 20 below is obtained.
[0243] Table 20 Unified identifier - ID relationship table
[0244] dmid id 1 imei_1 1 aid_1 1 sn_1 1 mac_1 2 imei_2 2 sn_3 2 aid_2 3 imei_3 3 tel_1 3 sn_2
[0245] (4.3) Generate the unified representation table of IDs: For the unified identifier - ID relationship table shown in Table 20, with dmid as the key, perform aggregation through the reduceByKey command to obtain the unified representation table of IDs shown in Table 21 below. Among them, in Table 21, the idsets column is the aggregated subset data of IDs.
[0246] Table 21 Unified representation table of IDs
[0247] dmid id sets 1 imei_1,aid_1,sn_1,mac_1 2 imei_2,sn_3,aid_2 3 imei_3,tel_1,sn_2
[0248] Equivalently, Table 21 can also be organized into the form shown in Table 22 below:
[0249] Table 22 Unified representation table of IDs
[0250] dmid imei aid sn mac tel 1 1 1 1 1 2 2 2 3 3 3 2 1
[0251] In this way, the final ID mapping result is obtained, realizing the unified representation of user IDs.
[0252] Based on the same inventive concept, an embodiment of the present invention also provides a device for implementing ID mapping based on the Spark framework, which is used to support the method for implementing ID mapping based on the Spark framework provided by any one of the above embodiments or a combination thereof. Figure 7 The structural schematic diagram of a device 700 for implementing ID mapping based on the Spark framework according to an embodiment of the present invention is shown. Refer to Figure 7 , this device may at least include: a data preprocessing module 710, an ID relationship aggregation module 720, a number relationship aggregation module 730, and an ID unified representation module 740.
[0253] Now, the functions of the components or devices of the apparatus for implementing ID mapping based on the Spark framework in the embodiments of the present invention and the connection relationships between the various parts are introduced:
[0254] The data preprocessing module 710 is adapted to obtain a two-dimensional ID relationship table including a plurality of ID pairs, number each ID pair, and obtain an initial number - ID pair relationship table.
[0255] The ID relationship aggregation module 720 is connected to the data preprocessing module 710 and is adapted to split and aggregate the initial number - ID pair relationship table with the ID as the key, and obtain a plurality of initial number first-aggregated subsets, where each initial number first-aggregated subset is composed of initial numbers.
[0256] The number relationship aggregation module 730 is connected to the ID relationship aggregation module 720 and is adapted to split and aggregate the plurality of initial number first-aggregated subsets with the initial number as the key, and obtain an initial number aggregated subset result, where there is no intersection between any two initial number aggregated subsets in the initial number aggregated subset result; use a unified identifier to number the initial number aggregated subset result, and obtain a unified identifier - initial number aggregated subset relationship table.
[0257] The ID unified representation module 740 is respectively connected to the number relationship aggregation module 730 and the data preprocessing module 710, and is adapted to obtain the corresponding relationship between the unified identifier and the ID according to the unified identifier - initial number aggregated subset relationship table and the initial number - ID pair relationship table, and realize the unified representation of the ID.
[0258] In an optional embodiment, the number relationship aggregation module 730 is further adapted to:
[0259] Take the initial number as the key, take the plurality of initial number first-aggregated subsets as the objects for splitting and aggregating in the first iterative operation, perform iterative operations of splitting and aggregating, and obtain an initial number aggregated subset result; where in each iterative operation, output the initial number aggregated subsets that have no intersection with other initial number aggregated subsets after aggregation, and use the remaining initial number aggregated subsets as the objects for splitting and aggregating in the next iterative operation; until the remaining initial number aggregated subsets cannot be aggregated any more in an iterative operation, output the remaining initial number aggregated subsets, terminate the iterative operation, and integrate the initial number aggregated subsets output in each iterative operation to obtain an initial number aggregated subset result.
[0260] In an optional embodiment, as Figure 8 shown, the ID relationship aggregation module 720 may include:
[0261] The first splitting unit 721 is adapted to split the initial serial number - ID pair relationship table into an initial serial number - ID relationship table;
[0262] The first aggregation unit 722 is connected to the first splitting unit 721 and is adapted to aggregate the initial serial number - ID relationship table with ID as the key to obtain a plurality of initial serial number first - level aggregation subsets, where each initial serial number first - level aggregation subset is composed of initial serial numbers; and
[0263] The first numbering unit 723 is connected to the first aggregation unit 722 and is adapted to re - number the plurality of initial serial number first - level aggregation subsets to obtain a secondary serial number - initial serial number first - level aggregation subset relationship table.
[0264] Preferably, still referring to Figure 8 As shown, the ID relationship aggregation module 720 may further include a first filtering and output unit 724. The first filtering and output unit 724 can be respectively connected to the first aggregation unit 722 and the first numbering unit 723, and is adapted to, after the first aggregation unit 722 aggregates to obtain a plurality of initial serial number first - level aggregation subsets, determine whether each subset in the plurality of initial serial number first - level aggregation subsets is an isolated subset that has no intersection with other initial serial number first - level aggregation subsets; if so, output the initial serial number first - level aggregation subset as the target initial serial number aggregation subset, and trigger the first numbering unit 723 to re - number the remaining initial serial number first - level aggregation subsets.
[0265] Furthermore, the first filtering and output unit 724 is further adapted to:
[0266] Count the occurrence times and the number of elements included in each subset in the plurality of initial serial number first - level aggregation subsets;
[0267] If the occurrence times of a certain initial serial number first - level aggregation subset is 2 and the number of elements is 1, then determine that the initial serial number first - level aggregation subset is an isolated subset.
[0268] In an alternative embodiment, still referring to Figure 8 As shown, the numbering relationship aggregation module 730 may include:
[0269] The second splitting unit 731 is adapted to split the secondary serial number - initial serial number first - level aggregation subset relationship table into a secondary serial number - initial serial number - initial serial number first - level aggregation subset relationship table;
[0270] The second aggregation unit 732 is connected to the second splitting unit 731 and is adapted to aggregate the secondary serial number - initial serial number - initial serial number first - level aggregation subset relationship table with the initial serial number as the key to obtain one or more initial serial number second - level aggregation subsets;
[0271] The second filtering and output unit 733, connected to the second aggregation unit 732, is adapted to filter and output the initial-number second-aggregation subsets that have no intersection with other initial-number second-aggregation subsets as the target initial-number aggregation subsets;
[0272] The second numbering unit 734, connected to the second filtering and output unit 733, is adapted to re-number the remaining initial-number second-aggregation subsets to obtain a three-numbering - initial-number second-aggregation subset relationship table;
[0273] The analog iterative unit 735, which can be respectively connected to the second numbering unit 734 and the second splitting unit 731, is adapted to trigger the second splitting unit 731, the second aggregation unit 732, the second filtering and output unit 733, and the second numbering unit 734 to perform n iterative operations by analogy until the number of the remaining initial-number (n + 2)-aggregation subsets is 0 or 1, and output the remaining initial-number (n + 2)-aggregation subsets as the target initial-number aggregation subsets, where n is a natural number; and
[0274] The output result integration unit 736, which can be respectively connected to the second filtering and output unit 733 and the analog iterative unit 735, is adapted to integrate the aforementioned output target initial-number aggregation subsets to obtain an initial-number aggregation subset result, and number the initial-number aggregation subset result with a unified identifier to obtain a unified identifier - initial-number aggregation subset relationship table.
[0275] Preferably, the second aggregation unit 732 is further adapted to:
[0276] Using the initial number as the key, aggregate the two-numbering - initial-number - initial-number first-aggregation subset relationship table to obtain one or more initial-number second-aggregation subsets and corresponding two-numbering aggregation subsets, where each two-numbering aggregation subset is composed of two numbers, and each initial-number second-aggregation subset is formed by merging the initial-number first-aggregation subsets corresponding to each two-number in its corresponding two-numbering aggregation subset.
[0277] Correspondingly, the second filtering and output unit 733 is further adapted to:
[0278] Judge whether each subset in one or more two-numbering aggregation subsets is an isolated subset that has no intersection with other two-numbering aggregation subsets. If so, de-duplicate the initial-number second-aggregation subset corresponding to the two-numbering aggregation subset and output it as the target initial-number aggregation subset.
[0279] Furthermore, the second filtering and output unit 733 is further adapted to:
[0280] Count the number of elements contained in each two-numbering aggregation subset;
[0281] Judge a quadratic numbering aggregation subset with the number of elements being 1 as an isolated subset.
[0282] Furthermore, the second filtering output unit 733 is also adapted to:
[0283] After counting the number of elements included in each quadratic numbering aggregation subset, judge whether the number of elements included in each quadratic numbering aggregation subset is greater than a given threshold;
[0284] If so, perform deduplication on the initial numbering quadratic aggregation subset corresponding to the quadratic numbering aggregation subset and output it as the target initial numbering aggregation subset.
[0285] In an alternative embodiment, still referring to Figure 8 as shown, the ID unified representation module 740 may include:
[0286] A third splitting unit 741, adapted to split the unified identifier - initial numbering aggregation subset relationship table into an initial numbering - unified identifier relationship table;
[0287] A relationship connection unit 742, connected to the third splitting unit 741, adapted to obtain a unified identifier - ID relationship table according to the initial numbering - unified identifier relationship table and the initial numbering - ID pair relationship table; and
[0288] A third aggregation unit 743, connected to the relationship connection unit 742, adapted to aggregate the unified identifier - ID relationship table with the unified identifier as the key to obtain a unified representation table of IDs.
[0289] Preferably, the relationship connection unit 742 is also adapted to:
[0290] Execute the leftOuterJoin command on the initial numbering - unified identifier relationship table and the initial numbering - ID pair relationship table, and then separate the unified identifier and the ID through the map command to obtain the unified identifier - ID relationship table.
[0291] In an alternative embodiment, the data pre - processing module 710 is also adapted to:
[0292] Integrate multiple two - dimensional ID relationship source data tables into a two - dimensional ID relationship table including multiple ID pairs.
[0293] In an alternative embodiment, the numbering operations of the above - mentioned modules are performed through the zipWithUniqueId command.
[0294] In an alternative embodiment, the aggregation operations of the above - mentioned modules are performed through the reduceByKey command.
[0295] Based on the same inventive concept, an embodiment of the present invention further provides a computer storage medium. The computer storage medium stores computer program code, which, when running on a computing device, causes the computing device to execute the method for implementing ID mapping based on the Spark framework according to any one of the above embodiments or a combination thereof.
[0296] Based on the same inventive concept, an embodiment of the present invention further provides a computing device. The computing device may include:
[0297] a processor; and
[0298] a memory storing computer program code;
[0299] When the computer program code is run by the processor, it causes the computing device to execute the method for implementing ID mapping based on the Spark framework according to any one of the above embodiments or a combination thereof.
[0300] According to any one of the above optional embodiments or a combination of multiple optional embodiments, the embodiments of the present invention can achieve the following beneficial effects:
[0301] The method and apparatus for implementing ID mapping based on the Spark framework proposed by the embodiments of the present invention, after numbering each ID pair in the obtained two-dimensional ID relationship table to obtain an initial number - ID pair relationship table, first split and aggregate ID relationships in the initial number - ID pair relationship table with the ID as the key to obtain multiple initial number first-aggregated subsets, then split and aggregate number relationships for the multiple initial number first-aggregated subsets with the initial number as the key to obtain the initial number aggregated subset result, and then number the initial number aggregated subset result with a unified identifier to obtain a unified identifier - initial number aggregated subset relationship table. Finally, according to the unified identifier - initial number aggregated subset relationship table and the initial number - ID pair relationship table, the corresponding relationship between the unified identifier and the ID is obtained, thus realizing the unified representation of user IDs. The present invention implements the ID mapping algorithm based on the Spark distributed computing framework, and uses the idea of set theory in mathematics to implement operations such as storage, filtering, splitting, and aggregation of a massive user data set, thereby improving the efficiency, accuracy, and reliability of the ID mapping algorithm.
[0302] Furthermore, the process of aggregating number relationships with the initial number as the key is realized through iterative operations. In each iterative operation, an initial number aggregation subset that has no intersection with other initial number aggregation subsets after aggregation (that is, no further aggregation is required) is output. This can make the number of remaining subsets in each iterative operation (that is, the subsets to be split and aggregated in the next iterative operation) smaller and smaller, thereby further reducing memory overhead and improving operation efficiency.
[0303] Those skilled in the art can clearly understand the specific working processes of the above-described systems, devices, and units. They can refer to the corresponding processes in the foregoing method embodiments. For the sake of brevity, they will not be described in detail here.
[0304] In addition, in each embodiment of the present invention, the functional units can be physically independent of each other, or two or more functional units can be integrated together, or all functional units can be integrated in a processing unit. The above-mentioned integrated functional units can be implemented in the form of hardware, or in the form of software or firmware.
[0305] Those of ordinary skill in the art can understand that if the above-mentioned integrated functional unit is implemented in the form of software and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention essentially or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computing device (such as a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention when the instructions are run. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0306] Alternatively, all or part of the steps of implementing the foregoing method embodiments can be completed by hardware related to program instructions (such as a computing device such as a personal computer, a server, or a network device). The program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the computing device, the computing device executes all or part of the steps of the method described in each embodiment of the present invention.
[0307] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that within the spirit and principle of the present invention, it is still possible to modify the technical solutions described in the foregoing embodiments, or to equivalently replace some or all of the technical features therein; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the protection scope of the present invention.
Claims
1. A method for implementing ID mapping based on the Spark framework, comprising: Step S1: Obtain a two-dimensional ID relationship table including multiple ID pairs, number each ID pair, and obtain an initial number - ID pair relationship table; Step S2: Using the ID as the key, split and aggregate the initial number - ID pair relationship table to obtain multiple initial number first-aggregation subsets, where each initial number first-aggregation subset is composed of initial numbers; Step S3: Using the initial number as the key, split and aggregate the multiple initial number first-aggregation subsets to obtain an initial number aggregation subset result, where there is no intersection between any two initial number aggregation subsets in the initial number aggregation subset result; number the initial number aggregation subset result using a unified identifier to obtain a unified identifier - initial number aggregation subset relationship table; Step S4: According to the unified identifier - initial number aggregation subset relationship table and the initial number - ID pair relationship table, obtain the correspondence between the unified identifier and the ID, and realize the unified representation of the ID.
2. The method according to claim 1, wherein, Using the initial number as the key, splitting and aggregating the multiple initial number first-aggregation subsets to obtain an initial number aggregation subset result, including: Using the initial number as the key, taking the multiple initial number first-aggregation subsets as the objects of splitting and aggregating in the first iteration operation, and performing iterative operations of splitting and aggregating to obtain the initial number aggregation subset result; where in each iterative operation, output the initial number aggregation subset that has no intersection with other initial number aggregation subsets after aggregation, and take the remaining initial number aggregation subsets as the objects of splitting and aggregating in the next iterative operation; until the remaining initial number aggregation subsets cannot be aggregated anymore in an iterative operation, output the remaining initial number aggregation subsets, terminate the iterative operation, and integrate the initial number aggregation subsets output in each iterative operation to obtain the initial number aggregation subset result.
3. The method according to claim 1, wherein, Step S2 specifically includes: Split the initial number - ID pair relationship table into an initial number - ID relationship table; Using the ID as the key, aggregate the initial number - ID relationship table to obtain multiple initial number first-aggregation subsets, where each initial number first-aggregation subset is composed of initial numbers; Renumber the initial number first-aggregation subsets to obtain a secondary number - initial number first-aggregation subset relationship table.
4. The method according to claim 3, wherein After aggregating to obtain multiple initial number first-aggregation subsets, Step S2 further includes: Judge whether each subset in the multiple initial number first-aggregation subsets is an isolated subset that has no intersection with other initial number first-aggregation subsets; If so, output the initial number first-aggregation subset as the target initial number aggregation subset, and renumber the remaining initial number first-aggregation subsets.
5. The method according to claim 4, wherein Judging whether each subset in the multiple initial number first-aggregation subsets is an isolated subset that has no intersection with other initial number first-aggregation subsets includes: Count the occurrence times and the number of elements included in each initial number first-aggregation subset; An initial numbered once aggregated subset with an occurrence count of 2 and an element count of 1 is judged as an isolated subset.
6. The method according to any one of claims 3-5, wherein, Step S3 specifically includes: Step S31: Split the secondary numbered - initial numbered once aggregated subset relationship table into a secondary numbered - initial numbered - initial numbered once aggregated subset relationship table; Step S32: Aggregate the secondary numbered - initial numbered - initial numbered once aggregated subset relationship table with the initial number as the key to obtain one or more initial numbered secondary aggregated subsets; Step S33: Filter and output the initial numbered secondary aggregated subsets that have no intersection with other initial numbered secondary aggregated subsets as the target initial numbered aggregated subsets; Step S34: Remove duplicates from the remaining initial numbered secondary aggregated subsets and re - number them to obtain a tertiary numbered - initial numbered secondary aggregated subset relationship table; Step S35: By analogy, repeat steps S31 to S34 for n iterative operations until the number of remaining initial numbered n + 2 - times aggregated subsets is 0 or 1, and output the remaining initial numbered n + 2 - times aggregated subsets as the target initial numbered aggregated subsets, where n is a natural number; Step S36: Integrate the target initial numbered aggregated subsets to obtain the initial numbered aggregated subset result, and number the initial numbered aggregated subset result using a unified identifier to obtain a unified identifier - initial numbered aggregated subset relationship table.
7. The method according to claim 6, wherein Step S32 specifically includes: Aggregate the secondary numbered - initial numbered - initial numbered once aggregated subset relationship table with the initial number as the key to obtain one or more initial numbered secondary aggregated subsets and their corresponding secondary numbered aggregated subsets, where each secondary numbered aggregated subset is composed of secondary numbers, and each initial numbered secondary aggregated subset is formed by merging the initial numbered once aggregated subsets corresponding to each secondary number in its corresponding secondary numbered aggregated subset; Step S33 specifically includes: Judge whether each subset in the one or more secondary numbered aggregated subsets is an isolated subset that has no intersection with other secondary numbered aggregated subsets. If so, de - duplicate the initial numbered secondary aggregated subset corresponding to this secondary numbered aggregated subset and output it as the target initial numbered aggregated subset.
8. The method according to claim 7, wherein, Judging whether each subset in the one or more secondary numbered aggregated subsets is an isolated subset that has no intersection with other secondary numbered aggregated subsets includes: Count the number of elements contained in each secondary numbered aggregated subset; Judge the secondary numbered aggregated subset with an element count of 1 as an isolated subset.
9. The method according to claim 8, wherein After counting the number of elements contained in each secondary numbered aggregated subset, step S33 further includes: Judge whether the number of elements contained in each secondary numbered aggregated subset is greater than a given threshold; If so, de - duplicate the initial numbered secondary aggregated subset corresponding to this secondary numbered aggregated subset and output it as the target initial numbered aggregated subset.
10. The method according to claim 1, wherein, Step S4 specifically includes: Split the unified identifier - initial numbered aggregated subset relationship table into an initial numbered - unified identifier relationship table; Obtain a unified identifier - ID relationship table based on the initial numbered - unified identifier relationship table and the initial numbered - ID pair relationship table; Aggregate the unified identifier - ID relationship table with the unified identifier as the key to obtain the unified representation table of IDs.
11. The method according to claim 10, wherein, According to the initial number - unified identifier relationship table and the initial number - ID pair relationship table, obtain the unified identifier - ID relationship table, including: Execute the leftOuterJoin command on the initial number - unified identifier relationship table and the initial number - ID pair relationship table, and then use the map command to separate the unified identifier and the ID to obtain the unified identifier - ID relationship table.
12. The method according to claim 1, wherein, Obtain a two-dimensional ID relationship table including multiple ID pairs, including: Integrate multiple two-dimensional ID relationship source data tables into the two-dimensional ID relationship table including multiple ID pairs.
13. The method according to claim 1, wherein The operation of numbering is performed through the zipWithUniqueId command.
14. The method according to claim 1, wherein, The operation of aggregation is performed through the reduceByKey command.
15. An apparatus for implementing ID mapping based on the Spark framework, including: A data preprocessing module, adapted to obtain a two-dimensional ID relationship table including multiple ID pairs, number each ID pair to obtain an initial number - ID pair relationship table; An ID relationship aggregation module, adapted to use the ID as the key, split and aggregate the initial number - ID pair relationship table to obtain multiple initial number first-aggregation subsets, where each initial number first-aggregation subset is composed of initial numbers; A number relationship aggregation module, adapted to use the initial number as the key, split and aggregate the multiple initial number first-aggregation subsets to obtain an initial number aggregation subset result, where there is no intersection between any two initial number aggregation subsets in the initial number aggregation subset result; number the initial number aggregation subset result with a unified identifier to obtain a unified identifier - initial number aggregation subset relationship table; and An ID unified representation module, adapted to obtain the corresponding relationship between the unified identifier and the ID according to the unified identifier - initial number aggregation subset relationship table and the initial number - ID pair relationship table, and implement the unified representation of the ID.
16. The device according to claim 15, wherein, The number relationship aggregation module is further adapted to: Use the initial number as the key, and use the multiple initial number first-aggregation subsets as the objects for splitting and aggregating in the first iteration operation, perform iterative operations of splitting and aggregating to obtain the initial number aggregation subset result; where in each iteration operation, output the initial number aggregation subset that has no intersection with other initial number aggregation subsets after aggregation, and use the remaining initial number aggregation subsets as the objects for splitting and aggregating in the next iteration operation; until the remaining initial number aggregation subsets cannot be aggregated any more in an iteration operation, output the remaining initial number aggregation subsets, terminate the iteration operation, and integrate the initial number aggregation subsets output in each iteration operation to obtain the initial number aggregation subset result.
17. The apparatus according to claim 15, wherein The ID relationship aggregation module includes: A first splitting unit, adapted to split the initial number - ID pair relationship table into an initial number - ID relationship table; A first aggregation unit, adapted to aggregate the initial number - ID relationship table with ID as the key, to obtain a plurality of initial number first - level aggregation subsets, where each initial number first - level aggregation subset is composed of initial numbers; and A first numbering unit, adapted to re - number the initial number first - level aggregation subsets to obtain a secondary number - initial number first - level aggregation subset relationship table.
18. The device according to claim 17, wherein The ID relationship aggregation module further includes:[[]]END A first filtering and output unit, adapted to, after the first aggregation unit aggregates to obtain a plurality of initial number first - level aggregation subsets, determine whether each subset in the plurality of initial number first - level aggregation subsets is an isolated subset that has no intersection with other initial number first - level aggregation subsets; If so, output the initial number first - level aggregation subset as the target initial number aggregation subset, and trigger the first numbering unit to re - number the remaining initial number first - level aggregation subsets.
19. The device according to claim 18, wherein, The first filtering and output unit is further adapted to:[[]]END Statistically count the occurrence times and the number of elements included in each initial number first - level aggregation subset; Determine an initial number first - level aggregation subset with an occurrence time of 2 and a number of elements of 1 as an isolated subset.
20. The apparatus according to any one of claims 17-19, wherein, The numbering relationship aggregation module includes:[[]]END A second splitting unit, adapted to split the secondary number - initial number first - level aggregation subset relationship table into a secondary number - initial number - initial number first - level aggregation subset relationship table; A second aggregation unit, adapted to aggregate the secondary number - initial number - initial number first - level aggregation subset relationship table with the initial number as the key, to obtain one or more initial number second - level aggregation subsets; A second filtering and output unit, adapted to filter and output an initial number second - level aggregation subset that has no intersection with other initial number second - level aggregation subsets as the target initial number aggregation subset; A second numbering unit, adapted to re - number the remaining initial number second - level aggregation subsets to obtain a tertiary number - initial number second - level aggregation subset relationship table; An analog iterative unit, adapted to trigger the second splitting unit, the second aggregation unit, the second filtering and output unit, and the second numbering unit to perform n - times of iterative operations in this way, until the number of remaining initial number (n + 2) - level aggregation subsets is 0 or 1, and output the remaining initial number (n + 2) - level aggregation subsets as the target initial number aggregation subset, where n is a natural number; and An output result integration unit, adapted to integrate the target initial number aggregation subsets to obtain the initial number aggregation subset result, and number the initial number aggregation subset result with a unified identifier to obtain a unified identifier - initial number aggregation subset relationship table.
21. The apparatus according to claim 20, wherein, The second aggregation unit is further adapted to:[[]]END Aggregate the secondary number - initial number - initial number first - level aggregation subset relationship table with the initial number as the key, to obtain one or more initial number second - level aggregation subsets and corresponding secondary number aggregation subsets, where each secondary number aggregation subset is composed of secondary numbers, and each initial number second - level aggregation subset is formed by merging the initial number first - level aggregation subsets corresponding to each secondary number in its corresponding secondary number aggregation subset; The second filtering and output unit is further adapted to:[[]]END Determine whether each subset in the one or more secondary number aggregation subsets is an isolated subset that has no intersection with other secondary number aggregation subsets. If so, perform deduplication on the initial number secondary aggregation subset corresponding to this secondary number aggregation subset and output it as the target initial number aggregation subset.
22. The apparatus according to claim 21, wherein, The second filtering and output unit is further adapted to: Count the number of elements included in each secondary number aggregation subset; Judge the secondary number aggregation subset with the number of elements being 1 as an isolated subset.
23. The device according to claim 22, wherein, The second filtering and output unit is further adapted to: After counting the number of elements included in each secondary number aggregation subset, determine whether the number of elements included in each secondary number aggregation subset is greater than a given threshold; If so, perform deduplication on the initial number secondary aggregation subset corresponding to this secondary number aggregation subset and output it as the target initial number aggregation subset.
24. The apparatus according to claim 15, wherein The ID unified representation module includes: A third splitting unit, adapted to split the unified identifier - initial number aggregation subset relationship table into an initial number - unified identifier relationship table; A relationship connection unit, adapted to obtain a unified identifier - ID relationship table according to the initial number - unified identifier relationship table and the initial number - ID pair relationship table; and A third aggregation unit, adapted to aggregate the unified identifier - ID relationship table with the unified identifier as the key to obtain a unified representation table of IDs.
25. The apparatus according to claim 24, wherein The relationship connection unit is further adapted to: Execute the leftOuterJoin command on the initial number - unified identifier relationship table and the initial number - ID pair relationship table, and then separate the unified identifier and the ID through the map command to obtain the unified identifier - ID relationship table.
26. The apparatus according to claim 15, wherein, The data preprocessing module is further adapted to: Integrate multiple two - dimensional ID relationship source data tables into the two - dimensional ID relationship table including multiple ID pairs.
27. The apparatus according to claim 15, wherein, The operation of numbering is performed through the zipWithUniqueId command.
28. The apparatus according to claim 15, wherein The operation of aggregation is performed through the reduceByKey command.
29. A computer storage medium, storing computer program code, which when running on a computing device, causes the computing device to execute the method for implementing ID mapping based on the Spark framework according to any one of claims 1 - 14.
30. A computing device, comprising: A processor; And A memory storing computer program code; When the computer program code is run by the processor, it causes the computing device to execute the method for implementing ID mapping based on the Spark framework according to any one of claims 1 - 14.
Citation Information
Patent Citations
User ID (Identification) recognition method and device
CN105099729A
User identifier processing method and apparatus
CN105224606A