Data processing method and device, equipment and storage medium
By acquiring data from real-time data sources and generating and updating real-time tables, the problem of low connectivity and accuracy caused by offline data delays in data integration systems is solved, enabling more efficient and accurate ID querying and recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BAIDU (CHINA) CO LTD
- Filing Date
- 2023-04-20
- Publication Date
- 2026-04-24
AI Technical Summary
The data integration system suffers from low integration rate and ID recognition accuracy due to offline data processing delays, making it unable to meet the needs of application scenarios with high timeliness requirements.
By acquiring data to be processed from real-time data sources, generating and updating real-time tables, and supplementing data to improve the data connectivity of systems that were missing due to the poor timeliness of offline tables, the efficiency and accuracy of real-time ID queries are improved.
It improved the overall integration rate and real-time ID query efficiency of the data integration system, enhanced the accuracy of ID recognition, and strengthened the integrity and completeness of the data integration system.
Smart Images

Figure CN116701796B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to the field of artificial intelligence technologies such as data mining, data analysis, and data querying. Background Technology
[0002] With the continuous development of big data and artificial intelligence, data processing methods have become increasingly diversified. Among related technologies, data integration systems primarily provide offline data to business users. However, these systems often involve numerous upstream data product lines, resulting in massive and disorganized data volumes that require multiple processing rounds, leading to delays of several days or even a week for offline data. Furthermore, the use of these product lines generates a large number of new small text files (cookies), which are highly variable identity documents (IDs). Due to these delays, these IDs are not retrieved. Consequently, the integration rate and ID recognition accuracy of data integration systems are low, failing to meet the demands of applications with high timeliness requirements. Summary of the Invention
[0003] This disclosure provides a data processing method, apparatus, device, and storage medium.
[0004] According to a first aspect of this disclosure, a data processing method is provided, comprising:
[0005] Obtain the data to be processed, which includes a first-class ID and the first attribute information of the first-class ID. The first-class ID is a randomly generated ID.
[0006] Based on the data to be processed, obtain target ID pairs, which include first-type IDs and second-type IDs;
[0007] The real-time table is updated based on the target ID. The real-time table includes the first connection record of the first type of ID. The first connection record includes at least the first correspondence between the first type of ID and the first attribute information of the first type of ID.
[0008] The real-time table is mounted to the offline table of each cluster in the data integration system. The offline table includes the second integration record of the second type ID. The second integration record includes at least the second correspondence between the second type ID and the second attribute information of the second type ID.
[0009] According to a second aspect of this disclosure, a data processing apparatus is provided, comprising:
[0010] The first acquisition module is used to acquire data to be processed, which includes a first type ID and the first attribute information of the first type ID. The first type ID is a randomly generated ID.
[0011] The second acquisition module is used to acquire target ID pairs based on the data to be processed, the target ID pairs including a first type of ID and a second type of ID;
[0012] The update module is used to update the real-time table based on the target ID. The real-time table includes the first connection record of the first type of ID. The first connection record includes at least the first correspondence between the first type of ID and the first attribute information of the first type of ID.
[0013] The mounting module is used to mount real-time tables to offline tables in each cluster of the data integration system. The offline table includes a second integration record with a second type of ID. The second integration record includes at least a second correspondence between the second type of ID and the second attribute information of the second type of ID.
[0014] According to a third aspect of this disclosure, an electronic device is provided, comprising:
[0015] At least one processor;
[0016] Memory that is communicatively connected to at least one processor;
[0017] The memory stores instructions that can be executed by at least one processor to enable the at least one processor to perform the methods of any embodiment of this disclosure.
[0018] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause a computer to perform a method according to any embodiment of this disclosure.
[0019] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program stored on a storage medium, which, when executed by a processor, implements a method according to any embodiment of this disclosure.
[0020] According to the scheme disclosed herein, by obtaining target ID pairs from the data to be processed, updating the real-time table based on the target ID pairs, and mounting the real-time table to the offline tables of each cluster in the data integration system, the data integration system can supplement the missing part of the integration rate due to the poor timeliness of the offline tables by generating and updating the real-time table, thereby improving the efficiency of real-time ID query and the accuracy of ID recognition.
[0021] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of this application will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description
[0022] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in this application and should not be construed as limiting the scope of this application.
[0023] Figure 1 This is a schematic flowchart of a data processing method according to an embodiment of the present disclosure;
[0024] Figure 2 This is a schematic diagram of the processing flow of the real-time data processing module in the data integration system according to an embodiment of the present disclosure;
[0025] Figure 3 This is a schematic diagram of the process for obtaining target ID pairs according to an embodiment of this disclosure;
[0026] Figure 4 This is a flowchart illustrating the process of determining the confidence level of a second type of candidate ID pair according to an embodiment of this disclosure;
[0027] Figure 5 This is a schematic diagram of the process of updating the real-time table based on the target ID according to an embodiment of the present disclosure;
[0028] Figure 6 This is a schematic diagram of the processing flow of the data injection module in the data integration system according to an embodiment of the present disclosure;
[0029] Figure 7 This is a schematic diagram of the processing flow of the online query module in the data integration system according to an embodiment of the present disclosure;
[0030] Figure 8 This is a schematic diagram of the structure of a data processing apparatus according to an embodiment of the present disclosure;
[0031] Figure 9 This is a schematic diagram of a data processing scenario according to an embodiment of the present disclosure;
[0032] Figure 10 This is a schematic diagram of the structure of an electronic device used to implement the data processing method of the embodiments of this disclosure. Detailed Implementation
[0033] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0034] The terms "first," "second," and "third," etc., used in the embodiments, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0035] In related technologies, data integration systems primarily provide offline data to business users. However, these systems often involve numerous upstream data product lines, resulting in massive and disorganized data volumes that require multiple processing rounds, leading to delays of several days or even a week. Furthermore, each time a user opens a webpage, a cookie is generated, resulting in a large number of new cookies daily. Cookies are highly variable IDs. Due to these delays, some of these IDs are not retrieved, leading to a low integration rate for the data integration system and failing to meet the needs of applications with high timeliness requirements.
[0036] This disclosure proposes a data processing method to at least partially address one or more of the aforementioned problems and other potential issues. By acquiring data to be processed from a real-time data source, generating and updating a real-time table, it supplements the data integration system that suffers from the poor timeliness of offline tables, thereby improving the overall integration rate of the data integration system, and thus enhancing the efficiency of real-time ID queries and the accuracy of real-time ID recognition.
[0037] This disclosure provides a data processing method applied to a data integration system. Figure 1 This is a schematic flowchart of a data processing method according to an embodiment of the present disclosure. This data processing method can be applied to a data processing device. The data processing device is located on an electronic device. The electronic device includes, but is not limited to, fixed devices and / or mobile devices. For example, fixed devices include, but are not limited to, servers, which can be cloud servers or ordinary servers. For example, mobile devices include, but are not limited to, mobile phones, tablets, and vehicle-mounted terminals. In some possible implementations, the data processing method can also be implemented by a processor calling computer-readable instructions stored in memory. Figure 1 As shown, the data processing method includes:
[0038] S101: Obtain data to be processed from a real-time data source. The data to be processed includes a first type ID and the first attribute information of the first type ID. The first type ID is a randomly generated ID.
[0039] S102: Obtain target ID pairs based on the data to be processed. The target ID pairs include first-type IDs and second-type IDs.
[0040] S103: Update the real-time table based on the target ID pair. The real-time table includes the first connection record of the first type ID. The first connection record includes at least the first correspondence between the first type ID and the first attribute information of the first type ID.
[0041] S104: Mount the real-time table to the offline tables of each cluster in the data integration system. The offline table includes the second integration record of the second type ID. The second integration record includes at least the second correspondence between the second type ID and the second attribute information of the second type ID.
[0042] In this embodiment of the disclosure, the data to be processed may include multiple access behavior records, such as webpage access behavior records and application (APP) access behavior records. Each access behavior record includes a first type of ID, and may also include one or more of the following: timestamp, Internet Protocol (IP) address, device attributes, and operating system (OS) attributes. The above is merely an illustrative example and is not intended to limit all possible contents included in the data to be processed; it is simply not an exhaustive list.
[0043] In this embodiment of the disclosure, the first type of ID can be a cookie generated by opening a webpage in a computer's browser, a cookie generated by opening a search page in a shopping app on an electronic device, or a cookie generated by opening a research topic page on an electronic device. The above is merely illustrative and is not intended to limit all possible sources of the first type of ID; it is simply not an exhaustive list.
[0044] Here, a cookie is a small text file, which can be understood as data stored on the user's local terminal to distinguish users during session tracking. It is information that is temporarily or permanently stored by the user's client computer.
[0045] In this embodiment of the disclosure, the content stored in the real-time table may include a first type of ID and a first correspondence between the first type of ID and its first attribute information. The first type of ID may include a cookie, and the first attribute information may include information such as a timestamp, IP address, device model, and operating system model. For example, if an access behavior record includes information such as a cookie, timestamp, IP address, device model, and operating system model, then the correspondence between the cookie and its IP address and device model can be stored in the real-time table.
[0046] In this embodiment, the data integration system may include a data extraction module, a real-time data processing module, and a data injection module. The data extraction module performs deduplication on recurring data to be processed obtained from a real-time data source, thereby avoiding repeated processing of identical data and saving computational resources. The real-time data processing module generates a real-time table based on the data to be processed. This real-time table records a first correspondence between a first type of ID and its first attribute information. The real-time data processing module is highly scalable, capable of introducing more new data to be processed when resources are sufficient.
[0047] In practical applications, ID pairs are extracted from real-time data sources, aggregated to the main cluster for real-time computation, and then updated and distributed to other clusters. For example, a large number of ID pairs extracted from real-time data sources are aggregated to cluster A of the data integration system for real-time computation. After cluster A updates its local real-time table through real-time computation, the update results are synchronously distributed to other clusters of the data integration system, such as clusters B, C, and D.
[0048] In this embodiment, the real-time calculation process may include: Step 1, directly filtering out access behavior records related to IPs located in the IP blacklist; wherein the IP blacklist stores IPs with abnormal behavior. Step 2, querying the cookie and username (User ID) in the ID pair (cookie, User ID) of an access behavior record within a preset time period in a cached real-time table; wherein the preset time period can be two hours, half a day, or a day, and the length of the preset time period can be set or adjusted according to requirements. Step 3, if the User ID is found in the real-time table, the cookie is added to the ID list connected to the User ID, and the attribute information is updated; if the cookie is found in the real-time table, the User ID is added to the ID list connected to the cookie, and the attribute information is updated; wherein the attribute information includes product line, update time, statistical information, etc.
[0049] In this embodiment, the data injection module may include a parsing submodule and a processing submodule. The parsing submodule converts the first attribute information of the first type of ID in the data to be processed into a target data exchange format. It generates a key for the first type of ID through hashing, shifting, and other operations, and stores the first attribute information of the first type of ID in the value corresponding to the key. The target data exchange format may be a Protocol Buffers (protobuf) format. The first attribute information may include timestamps, IP addresses, device attributes, operating system attributes, product lines, active time, frequency of occurrence, etc. The first attribute information in the value is in the target data exchange format, such as protobuf format. In practical applications, the parsing submodule may first store the first correspondence between the first type of ID and its first attribute information on a temporary path. The processing submodule converts the first correspondence in the target data exchange format into a first correspondence in a relational database format and stores it on a formal path. The relational database used in the data integration system is a simpleDB data management system. The formal path is the address used for online query integration services. For example, after obtaining the first correspondence between the first type of ID and the first attribute information of the first type of ID through the real-time data processing module, the data injection module performs the injection operation to generate a real-time table in the format of a final relational database (such as simpleDB) for business users to query.
[0050] Figure 2 The diagram illustrates the processing flow of the real-time data processing module in the data integration system, such as... Figure 2 As shown, data to be processed is extracted from a real-time data source, and task processing based on rules or strategies is performed on the data to obtain updated connection records. The updated connection records are then aggregated. Data updates and database injection are performed through a data update interface, and finally, the data is written to different clusters for caching, generating a real-time table in SimpleDB format. The real-time data source can be a basic dynamic web service system (BigPipe). The real-time data source can include data to be processed from multiple product lines, which in practical applications can be understood as application software (APP). Figure 2 In this example, clusters A, B, C, and D are clusters located in different regions. If an ID pair cannot be directly formed from the access behavior records obtained from the real-time data source, it is established according to a policy; otherwise, if an ID pair is directly formed from the access behavior records obtained from the real-time data source, it is established according to rules. Choosing different methods to obtain the target ID pair based on different situations helps improve the accuracy of ID pair identification.
[0051] In this embodiment, the data update interface executes the update commands from the real-time computing logic within the cluster, updating the aggregated data in the cluster's cache. Simultaneously, it allows for a degree of filtering, such as data pruning based on the importance of the integrated product lines and the integration time.
[0052] In this embodiment, policy integration uses machine learning to determine whether two IDs belong to the same target object. Policy integration primarily involves introducing cookies and device IDs from multiple product lines and their attribute information, and using IP address and geographic location information to store this information in IP-based buckets. This significantly saves computational resources compared to pairwise combinations of all ID data.
[0053] For example, a webpage access behavior record is retrieved from a real-time data source; a target ID pair is obtained based on this record; the real-time table is updated based on the target ID pair, and the real-time table includes a first connection record for the first type of ID, which at least includes a first correspondence between the first type of ID and its first attribute information; the real-time table is then mounted to offline tables in each cluster of the data integration system, and the offline tables include a second connection record for the second type of ID, which at least includes a second correspondence between the second type of ID and its second attribute information. The data injection module performs a database injection operation to generate the final real-time table for querying by business users.
[0054] The technical solution of this disclosure involves obtaining data to be processed from a real-time data source. The data to be processed includes a first type of ID and first attribute information of the first type of ID, where the first type of ID is a randomly generated ID. Based on the data to be processed, target ID pairs are obtained. These target ID pairs include a first type of ID and a second type of ID. A real-time table is updated based on the target ID pairs. The real-time table includes a first connection record for the first type of ID, where the first connection record includes at least a first correspondence between the first type of ID and its first attribute information. The real-time table is then mounted to offline tables in each cluster of the data integration system. These offline tables include a second connection record for the second type of ID, where the second connection record includes at least a second correspondence between the second type of ID and its second attribute information. Thus, by obtaining target ID pairs from the data to be processed, updating the real-time table based on these target ID pairs, and mounting the real-time table to offline tables in each cluster of the data integration system, the missing connection rate due to the poor timeliness of offline tables in the data integration system can be supplemented by generating and updating the real-time table. This improves the efficiency of real-time ID queries and the accuracy of ID recognition. Furthermore, it enhances the completeness and robustness of the data integration system.
[0055] In some embodiments, S102 may include:
[0056] S102a: In response to the first type of candidate ID pair that can form a first type of ID based on the data to be processed, the first type of candidate ID pair is used as the target ID pair of the first type of ID.
[0057] Here, in response to the first type of candidate ID pairs that can form a first type of ID based on the data to be processed, using the first type of candidate ID pairs as the target ID pairs of the first type of ID constitutes rule association. In this case, the confidence level of the target ID pair can reach 100%.
[0058] In this embodiment of the disclosure, a user behavior record is obtained from a real-time data source. If information such as cookie, User ID, timestamp, and product line name is obtained by analyzing the user behavior record, a first-type candidate ID pair is obtained. This first-type candidate ID pair can be denoted as (cookie, User ID). Here, User ID can represent the username under the product line name. Since the cookie and User ID belong to the same user behavior record, the confidence level of this first-type candidate ID pair is very high, and in some cases, the confidence level can reach 100%.
[0059] Thus, the target ID pairs obtained from the first type of ID based on the same access behavior record can provide a highly confident method for obtaining target ID pairs for the data integration system, which helps to improve the efficiency of ID pair acquisition.
[0060] Figure 3 The flowchart illustrating the process of obtaining target ID pairs is shown below. Figure 3 As shown, S102 may include:
[0061] S102b: In response to the inability to form a first-class candidate ID pair based on the data to be processed, obtain at least one second-class candidate ID pair of the first-class ID from the IP bucket corresponding to the IP value of the first-class ID;
[0062] S102c: Analyze the confidence level of at least one second-class candidate ID pair;
[0063] S102d: Select the second type of candidate ID pairs with a confidence level not less than a preset threshold as the target ID pairs of the first type of IDs.
[0064] In this embodiment, all IDs within the corresponding IP bucket are retrieved based on the IP address of the data to be processed. For IDs that meet the matching threshold, a random forest model is used to filter out a second type of candidate ID pairs based on the discriminative power of a preset combination, thus filtering out abnormal data. The filtered second type of candidate ID pairs are data that are consistent across dimensions of the preset combination. Conversely, abnormal data refers to data that is inconsistent across dimensions of the preset combination. In this embodiment, if the number of IDs associated with an ID exceeds the matching threshold, then that ID is considered an abnormal ID. An abnormal ID refers to an ID triggered by abnormal user behavior. The preset threshold can be set or adjusted as needed. For example, the matching threshold can be 200. Another example is that the matching threshold can be 500.
[0065] Here, an ID pair refers to a pairing relationship between a first-class ID and a known ID within an IP bucket. IP buckets are divided by geographical region. For example, IPs from Beijing are placed in the first IP bucket, IPs from Tianjin in the second IP bucket, and IPs from Guangzhou in the third IP bucket. This preset combination can be based on features such as operating system, operating system version (OS-Version), brand, and browser.
[0066] In this embodiment of the disclosure, IP addresses are divided into a blacklist and a whitelist based on user behavior. The blacklist includes IP addresses that have engaged in abnormal access behavior, such as accessing a certain page 100 times consecutively. The whitelist may include IP addresses that have engaged in normal access behavior.
[0067] For example, the cookie, i.e., the first type of ID, is determined from the first record obtained from the real-time data source, and this first record does not contain other valid information. This first type of ID is then matched against all IDs in the IP bucket. Specifically, if there are M IDs in the IP bucket, the first type of ID is paired with each of the M IDs. If the first type of ID is consistent with x IDs in a preset combination dimension, then the first type of ID forms x ID pairs, where 0 < x ≤ M. If the first type of ID is inconsistent with any of the M IDs in the preset combination dimension, it means that the first type of ID cannot form an ID pair relationship from within the IP bucket. Further, if the first type of ID is inconsistent with any of the M IDs in the preset combination dimension, a new connection record is generated based on the first type of ID.
[0068] In this embodiment of the disclosure, a first prediction model and a second prediction model from the model library are used to predict the probability that real-time issued ID pairs belong to the same target object, and a confidence score is returned. The ID pair is a target ID pair obtained based on a second type of candidate ID pairs. The model library contains at least two prediction models used to predict the confidence score of real-time issued ID pairs belonging to the same target object. Thus, by fusing the results of multiple models to obtain the confidence score, the generalization ability of the data integration system is improved.
[0069] In this embodiment of the disclosure, the first prediction model predicts whether the ID pair belongs to the same target object, and the prediction confidence result is C1; the second prediction model predicts whether the ID pair belongs to the same target object, and the prediction confidence result is C2; when (C1+C2) / 2 is greater than a preset threshold, it is determined that the ID pair belongs to the same target object. The target object can be understood as a natural person or a robot with a unique identifier.
[0070] In this embodiment, the training process of the first and second prediction models specifically includes: First, labeling positive and negative samples based on login behavior and calculating sample features. Second, learning to predict sample features using a first initial model, and obtaining the first prediction model based on the predicted and true values of the sample features; learning to predict sample features using a second initial model, and obtaining the second prediction model based on the predicted and true values of the sample features. The sample features may include multi-granularity features such as IP overlap, IP frequency similarity during working / non-working hours, IP frequency similarity on working / non-working days, search term (query) similarity, device model similarity, and browser similarity. The first initial model may be a Deep Neural Network (DNN) model, and the second initial model may be a Naive Bayesian Model (NBM). Finally, storing the trained first and second prediction models in a model library. Both the first and second prediction models can be used to predict the confidence level of ID pairs belonging to the same target object.
[0071] In this embodiment of the disclosure, IP overlap refers to the degree of overlap between the IPs corresponding to the first type of ID and the second type of ID in the target ID pair. For example, if three ID pairs are formed from the IP bucket, denoted as [cookie, ID3], [cookie, ID4], and [cookie, ID5], and if the IP of cookie is {1, 2, 3, 4, 5}, and the IP of ID3 is {3, 4, 5, 6, 7}, then the IP overlap between cookie and ID3 is 0.6.
[0072] In this embodiment of the disclosure, similarity at different granularities refers to the similarity between the predicted value and the true value at each granularity. For example, device model similarity refers to the similarity between the predicted device model and the true device model. Browser similarity refers to the similarity between the predicted browser value and the true browser value. To ensure more accurate sample features, different similarity calculation methods can be used to calculate the similarity at different granularities. Preferably, the similarity at different granularities is calculated using the Tanimoto coefficient similarity calculation method and cosine similarity. Specifically, the search term similarity, device model similarity, and browser similarity are calculated using the Tanimoto coefficient similarity calculation method. The similarity of IP address frequency during working / non-working hours and the similarity of IP address frequency on working / non-working days are calculated using the cosine similarity calculation method.
[0073] Thus, in response to the inability to form a first-class candidate ID pair based on the data to be processed, at least one second-class candidate ID pair of the first-class ID is obtained from the IP bucket corresponding to the Internet Protocol (IP) value of the first-class ID; the confidence level of at least one second-class candidate ID pair is analyzed; and the second-class candidate ID pair with a confidence level not less than a preset threshold is used as the target ID pair of the first-class ID. This provides a policy-based method for obtaining target ID pairs, which helps improve the efficiency of ID pair identification. Simultaneously, by setting a preset threshold, anti-fraud processing can be performed, further improving the accuracy of target ID pairs.
[0074] Figure 4 A flowchart illustrating the process of determining the confidence level of second-class candidate ID pairs is shown, as follows: Figure 4 As shown, it specifically includes:
[0075] S401: Use the first prediction model to predict the confidence level of each second-class candidate ID pair, and obtain the first confidence level of each second-class candidate ID pair;
[0076] S402: Use the second prediction model to predict the confidence level of each second-class candidate ID pair, and obtain the second confidence level of each second-class candidate ID pair;
[0077] S403: Based on the first confidence level and the second confidence level of each second-class candidate ID pair, obtain the confidence level of each second-class candidate ID pair.
[0078] In this embodiment of the disclosure, a first prediction model is used to predict whether the ID pair belongs to the same target object, and a confidence result of C1 is obtained; a second prediction model is used to predict whether the ID pair belongs to the same target object, and a confidence result of C2 is obtained; if (C1+C2) / 2 is greater than a preset threshold, then it is determined that the ID pair belongs to the same target object. The target object can be understood as a natural person or a robot with a unique identifier.
[0079] Thus, by first fusing the results of the first and second prediction models and then outputting the confidence score, the generalization ability of the data integration system can be enhanced, thereby improving the accuracy of ID recognition.
[0080] In some embodiments, the data processing method may further include:
[0081] S105: In response to the inability to obtain the target ID pair of the first type ID from the IP bucket corresponding to the IP value, generate the first breakthrough record of the first type ID in the real-time table.
[0082] In this embodiment of the disclosure, the data to be processed may include multiple access behavior records. By analyzing a certain access behavior record, information such as cookies, timestamps, and product line names are obtained. Based on the ID of the cookie and the M IDs in the IP bucket corresponding to the IP value of the access behavior record, x second-type candidate ID pairs are obtained, where 0 < x ≤ M. If the confidence of all x second-type candidate ID pairs is less than a preset threshold, it is determined that the second-type target ID pair of the first-type ID cannot be obtained in the IP bucket corresponding to the IP value, that is, the target ID pair of the cookie ID cannot be obtained.
[0083] In this way, in response to the inability to obtain the target ID pair of the first type of ID from the IP bucket corresponding to the IP value, the first breakthrough record of the first type of ID is generated in the real-time table, which can provide data support for subsequent real-time calculation and processing, help provide the business with more timely query results, and also help the business learn and train the model to build a complete user profile.
[0084] Figure 5 The diagram illustrates the process of updating the real-time table based on the target ID, as shown below. Figure 5 As shown, S103 may include:
[0085] S103a: Perform query processing in the real-time table based on the target ID, including the first type of ID and the second type of ID respectively;
[0086] S103b: In response to finding the first type ID in the real-time table, add the second type ID to the first access list of the first type ID and update the attribute information of the first access list;
[0087] S103c: In response to finding a second type ID in the real-time table, add the first type ID to the second access list of the second type ID in the real-time table, and update the attribute information of the second access list;
[0088] S103d: In response to the failure to find the first type ID and the second type ID in the real-time table, generate a third access list based on the first type ID and the second type ID, and update the attribute information of the third access list.
[0089] In some embodiments, the first type of ID is an unstable ID. For example, a cookie generated when browsing a webpage. The second type of ID is a stable ID. For example, a User ID, a mobile phone number, etc.
[0090] In some embodiments, the first connection list includes at least one first connection record. After adding the second type ID to the first connection list of the first type ID, the first connection list also includes at least a third correspondence between the second type ID and the first attribute information of the first type ID.
[0091] In this embodiment of the disclosure, the second connection list includes at least one second connection record. After adding the first type ID to the second connection list of the second type ID in the real-time table, the second connection list also includes at least a fourth correspondence between the second attribute information of the first type ID and the second type ID.
[0092] In this embodiment of the disclosure, the third access list includes at least the association relationship between the first type of ID and the second type of ID included in the target ID pair. Specifically, it includes the association relationship between the first type of ID and the first attribute information of the first type of ID, and between the second type of ID and the second attribute information of the second type of ID.
[0093] In this embodiment of the disclosure, if the User ID can be found in the real-time table, the cookie is added to the ID list linked to the User ID, and the attribute information is updated; if the cookie can be found in the real-time table, the User ID is added to the ID list linked to the cookie, and the attribute information is updated. The attribute information may include product line, update time, statistical information, etc.
[0094] In this way, by querying the first and second types of IDs included in the target ID pair in the real-time table, a highly efficient query method can be provided for the business, improving the data integration rate and thus enhancing the accuracy of ID recognition.
[0095] In some embodiments, obtaining a target ID pair based on the data to be processed may include: sending the ID data to be processed to a first cluster in the data integration system that matches the identification information based on the identification information of the data to be processed; and obtaining the target ID pair through the first cluster. Updating the real-time table based on the target ID pair includes: updating the real-time table of the first cluster based on the target ID pair through the first cluster.
[0096] In some embodiments, the identification information may include an IP address and a geographic location. The above is merely illustrative and is not intended to limit all possible contents of the identification information; it is simply not an exhaustive list.
[0097] In the embodiments of this disclosure, all computations are performed on the main cluster to obtain the results. The local cache is read, modified, and the results are synchronized to other clusters to save computing resources.
[0098] In embodiments of this disclosure, in response to the IP address of the data to be processed being Beijing, the ID data to be processed is sent to the Beijing cluster in the data integration system. The target ID pair is obtained through the Beijing cluster; wherein, updating the real-time table based on the target ID pair includes: updating the real-time table of the Beijing cluster based on the target ID pair through the Beijing cluster.
[0099] Figure 6 The diagram illustrates the processing flow of the data injection module in the data integration system, such as... Figure 6 As shown, the process involves obtaining the aggregated real-time ID data; reading the real-time data through the parsing submodule of the data injection module, converting the data into a storage unit (Peta Byte, PB) in the online service format according to the specified format through hashing / shifting; reading the pre-parsed PB result through the processing submodule of the data injection module, and producing the final data for simpleDB injection; finally, the real-time ID SimpleDB is obtained.
[0100] Thus, based on the identification information of the data to be processed, the ID data to be processed is sent to the first cluster in the data integration system that matches the identification information, and the real-time table is updated based on the target ID. Selecting the corresponding cluster for matching based on the identification information can save computing power and time costs, thereby helping to improve the data integration rate and further improve the accuracy of ID recognition.
[0101] In some embodiments, S104 may include:
[0102] S104a: Mount the real-time table generated by the first cluster to the offline table of the first cluster; and
[0103] S104b: Send the real-time table to the second cluster of the data integration system so that the second cluster can mount the real-time table to the offline table of the second cluster. The second cluster is a cluster in the data integration system other than the first cluster.
[0104] In the embodiments of this disclosure, both the first cluster and the second cluster include an offline ID data table and a real-time ID data table. Compared to the high timeliness of the real-time table, the offline table, although less timely, has more comprehensive data. Therefore, by mounting the real-time table to the offline table, the missing parts of the offline table due to timeliness can be supplemented. Through the combination of the real-time table and the offline table, a target object with a more complete ID can ultimately be obtained.
[0105] By mounting the real-time table to the offline table, the real-time table and the offline table can complement each other, making the data integration system more complete and comprehensive, and significantly improving the accuracy of ID recognition.
[0106] In some embodiments, the data processing method may further include:
[0107] S106: Obtain the query ID and query instruction parameters input through the query interface;
[0108] S107: In response to detecting that the query indication parameter is a real-time query, retrieve the query result of the ID to be queried from the real-time table;
[0109] S108: In response to detecting that the query indication parameter is a non-real-time query or that the query result for the ID to be queried is not obtained from the real-time table, retrieve the query result for the ID to be queried from the offline table.
[0110] Here, the query interface is equivalent to the online integration interface of the data integration system.
[0111] In some embodiments, the data integration system further includes an online query module for providing an online query interface to offer query services. Users can use this online query interface to access the ID integration query service. Specifically, the user inputs a sample ID to be queried and sets a query indication parameter. Specifically, if the query indication parameter is 1, it indicates a query from the real-time table; if the query indication parameter is 0, it indicates a query from the offline table. The real-time indication parameter is the ID query parameter.
[0112] In some embodiments, obtaining the query result of the ID to be queried from the real-time table includes: converting the ID type and the ID value to be queried into a key; querying the real-time table based on the key; if the key exists, using the ID included in the key as the associated ID of the ID to be queried; further, using the relevant attribute information of the ID included in the value of the key as the associated attribute information of the ID to be queried. Here, the key can be a string of numbers generated by hash shifting the ID to be queried.
[0113] In some embodiments, a preset fixed algorithm can be used to convert the ID type and the ID value to be queried into a string of numbers, which serves as the key. For example, the preset fixed algorithm may include hashing, bit shifting, or other operations. Different types are defined for different IDs. For instance, ID types include 1 for cookies, 2 for devices, and 3 for User IDs. In the real-time table, the ID-related attribute information included in the key can exist in a target data format, such as protobuf.
[0114] In some embodiments, key-value storage refers to a type of non-relational (Not Only SQL, NoSQL) database that accesses data through key-value pairs. It has a novel storage structure that differs from relational database management systems (RDMS) and is well-suited for managing unstructured data on the network.
[0115] In some embodiments, Figure 7 The diagram illustrates the processing flow of the online query module in the data integration system, such as... Figure 7 As shown, the user inputs the ID to be queried through the online connection interface (also known as the query interface) and instructs to query in real time. The ID type and value of the ID to be queried are converted into a key, which is then used to query simpleDB. If the key exists in simpleDB, the connection record with the same input ID value is searched in the value of the key, and the query result of the ID to be queried is returned through the online connection interface. If the key does not exist in simpleDB, the query result of not finding the ID to be queried is returned through the online connection interface.
[0116] Iterate through the real-time table until a record with the same ID value and type as the input ID to be queried is found, and return the output result. If no record with the same ID value and type as the input ID to be queried is found after iterating through the real-time table, it is determined that no query result for the ID to be queried was obtained from the real-time table.
[0117] In this way, the system obtains the ID to be queried and the query instruction parameters input through the query interface; and selects between real-time query and offline query based on the instruction parameters, providing business users with two query methods. Business users can choose the appropriate query method according to their needs, which can improve the diversity of ID recognition options.
[0118] In some embodiments, S107 may include:
[0119] S107a: Obtain the penetration depth N;
[0120] S107b: Based on the integration depth N, perform N integration queries in the real-time table;
[0121] S107c: Use the result of the Nth query as the query result for the ID to be queried.
[0122] In some embodiments, since the parsing logic of the real-time part is not as complex as that of offline mining, the confidence of the query results is not as high as that of the query results based on offline mining. When querying the connected ID in real-time, the connected results will be queried N times based on the input connected depth N. That is, the connected results will be used to query the real-time table. The confidence of the connected results is the product of the confidence of the connected path. Here, N is a configurable parameter.
[0123] In some embodiments, the penetration depth is used to characterize the depth of the query. The penetration depth can be preset by the system or input by the business party through an online query interface. Specifically, if N=1, the query result for the ID to be queried is retrieved from the real-time table based on the ID to be queried. If N>1, the ID to be queried is retrieved from the real-time table based on the ID to be queried, obtaining the ID result of the first query; based on the ID result of the first query, the query result for the ID to be queried is retrieved from the real-time table, obtaining the ID result of the second query; and so on, based on the ID result of the (N-1)th query, the query result for the ID to be queried is retrieved from the real-time table, obtaining the ID result of the Nth query; the ID result of the Nth query is determined as the query result for the ID to be queried.
[0124] For example, the target key is used to query simpleDB. SimpleDB stores real-time tables. If the corresponding key is found, it means that the record corresponding to the target key contains the user-input ID and its related attribute information.
[0125] In the embodiments of this disclosure, querying the offline ID pair table for the ID to be queried includes: if the target key is not found in the real-time ID pair table, it means that there is no record for that ID in the real-time table; querying the offline ID pair table based on the target key; if the target key exists in the offline table, using the ID information associated with the target key as the ID information of the ID to be queried; specifically, using the relevant attribute information of the ID information associated with the target key as the relevant attribute information of the ID to be queried. If the target key does not exist in the offline table, returning a query result indicating that there is no record for the ID to be queried in the data processing system.
[0126] Thus, the integration depth N is obtained; based on the integration depth N, N integration queries are performed in the real-time table; the result of the Nth integration query is used as the query result for the ID to be queried; by introducing real-time ID query, the problem of poor timeliness of offline query is supplemented, and the timeliness and accuracy of ID query are improved.
[0127] It should be understood that Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 7The schematic diagrams shown are merely illustrative and not limiting, and are scalable; those skilled in the art can use them as a basis. Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 7 Even with various obvious changes and / or substitutions to the examples, the resulting technical solutions still fall within the scope of this disclosure.
[0128] This disclosure provides a data processing apparatus, such as... Figure 8 As shown, the data processing device may include: a first acquisition module 801, used to acquire data to be processed from a real-time data source, the data to be processed including a first type ID and first attribute information of the first type ID, the first type ID being a randomly generated ID; a second acquisition module 802, used to acquire target ID pairs based on the data to be processed, the target ID pairs including a first type ID and a second type ID; an update module 803, used to update a real-time table based on the target ID pairs, the real-time table including a first connection record of the first type ID, the first connection record including at least a first correspondence between the first type ID and the first attribute information of the first type ID; and a mounting module 804, used to mount the real-time table to offline tables of each cluster in the data connection system, the offline table including a second connection record of the second type ID, the second connection record including at least a second correspondence between the second type ID and the second attribute information of the second type ID.
[0129] In some embodiments, the second acquisition module 802 may include: a first acquisition submodule, configured to, in response to a first type of candidate ID pair that can form a first type of ID based on the data to be processed, use the first type of candidate ID pair as the target ID pair of the first type of ID.
[0130] In some embodiments, the second acquisition module 802 may further include: a second acquisition submodule, configured to, in response to the inability to form a first type of candidate ID pair based on the data to be processed, acquire at least one second type of candidate ID pair of the first type of ID from the IP bucket corresponding to the IP value of the first type of ID; an analysis submodule, configured to analyze the confidence level of at least one second type of candidate ID pair; and a third acquisition submodule, configured to use the second type of candidate ID pair with a confidence level not less than a preset threshold as the target ID pair of the first type of ID.
[0131] In some embodiments, the analysis submodule is specifically configured to: predict the confidence level of each second-class candidate ID pair using a first prediction model to obtain a first confidence level of each second-class candidate ID pair; predict the confidence level of each second-class candidate ID pair using a second prediction model to obtain a second confidence level of each second-class candidate ID pair; and obtain the confidence level of each second-class candidate ID pair based on the first and second confidence levels of each second-class candidate ID pair.
[0132] In some embodiments, the data processing apparatus may further include: a generation module 805 ( Figure 8 (not shown in the table), used to generate the first breakthrough record of the first type of ID in the real-time table in response to the inability to obtain the target ID pair of the first type of ID from the IP bucket corresponding to the IP value.
[0133] In some embodiments, the update module 803 may include: a query submodule, configured to perform query processing in a real-time table for the included first type ID and second type ID based on the target ID respectively; a first update submodule, configured to, in response to finding a first type ID in the real-time table, add the second type ID to a first access list of the first type ID and update the attribute information of the first access list; a second update submodule, configured to, in response to finding a second type ID in the real-time table, add the first type ID to a second access list of the second type ID in the real-time table and update the attribute information of the second access list; and a third update submodule, configured to, in response to not finding a first type ID and second type ID in the real-time table, generate a third access list based on the first type ID and second type ID and update the attribute information of the third access list.
[0134] In some embodiments, the second acquisition module 802 includes: a sending submodule, configured to send the ID data to be processed to a first cluster in the data integration system that matches the identification information based on the identification information of the data to be processed; and a fourth acquisition submodule, configured to acquire a target ID pair through the first cluster; wherein, the update module 803 includes: a fourth update submodule, configured to update the real-time table of the first cluster based on the target ID pair through the first cluster.
[0135] In some embodiments, the mounting module 804 includes: a first mounting submodule for mounting a real-time table generated by a first cluster to an offline table of the first cluster; and a second mounting submodule for sending the real-time table to a second cluster of the data integration system, so that the second cluster can mount the real-time table to an offline table of the second cluster, wherein the second cluster is a cluster in the data integration system other than the first cluster.
[0136] In some embodiments, the data processing apparatus may further include: a third acquisition module 806. Figure 8(Not shown in the image), used to obtain the query ID and query indication parameters input through the query interface; the fourth acquisition module 807 ( Figure 8 (not shown in the image), used to retrieve the query result of the ID to be queried from the real-time table in response to detecting that the query indication parameter is a real-time query; the fifth acquisition module 808 ( Figure 8 (not shown in the table), used to retrieve the query result of the ID to be queried from the offline table in response to detecting that the query indication parameter is a non-real-time query or that the query result of the ID to be queried is not obtained from the real-time table.
[0137] In some embodiments, the fourth acquisition module 807 ( Figure 8 (Not shown in the table), may include: a fifth acquisition submodule, used to acquire the connection depth N; a connection submodule, used to perform N connection queries in the real-time table based on the connection depth N; and a sixth acquisition submodule, used to use the result of the Nth connection query as the query result of the ID to be queried.
[0138] Those skilled in the art should understand that the functions of each processing module in the data processing apparatus of this disclosure embodiment can be understood with reference to the relevant description of the data processing method described above. Each processing module in the data processing apparatus of this disclosure embodiment can be implemented by an analog circuit that implements the functions of this disclosure embodiment, or by running software that executes the functions of this disclosure embodiment on an electronic device.
[0139] The data processing apparatus of this disclosure can supplement the missing connection rate due to timeliness in offline parts through real-time ID query, thereby improving the accuracy of user ID recognition.
[0140] This disclosure provides a schematic diagram of a data processing scenario, such as... Figure 9 As shown.
[0141] As previously described, the data processing method provided in this disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.
[0142] Specifically, the electronic device may perform the following operations:
[0143] Data to be processed is obtained from a real-time data source. The data to be processed includes a first type ID and the first attribute information of the first type ID. The first type ID is a randomly generated ID.
[0144] Based on the data to be processed, obtain target ID pairs, which include first-type IDs and second-type IDs;
[0145] The real-time table is updated based on the target ID. The real-time table includes the first connection record of the first type of ID. The first connection record includes at least the first correspondence between the first type of ID and the first attribute information of the first type of ID.
[0146] The real-time table is mounted to the offline table of each cluster in the data integration system. The offline table includes the second integration record of the second type ID. The second integration record includes at least the second correspondence between the second type ID and the second attribute information of the second type ID.
[0147] The real-time data source can be obtained from a data source. The data source can be various forms of data storage devices, such as laptops, desktop computers, workstations, personal digital assistants (PDAs), servers, blade servers, mainframes, and other suitable computers. The data source can also represent various forms of mobile devices, such as PDAs, cellular phones, smartphones, wearable devices, and other similar computing devices. Furthermore, the data source and the user terminal can be the same device.
[0148] It should be understood that Figure 9 The scene diagrams shown are merely illustrative and not restrictive; those skilled in the art can interpret them based on... Figure 9 Even with various obvious changes and / or substitutions to the examples, the resulting technical solutions still fall within the scope of this disclosure.
[0149] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0150] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0151] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0152] like Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0153] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0154] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the data processing methods described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform a data processing method by any other suitable means (e.g., by means of firmware).
[0155] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0156] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0157] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0158] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0159] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), and the Internet.
[0160] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0161] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0162] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A data processing method applied to a data integration system, comprising: Acquire data to be processed, the data to be processed including a first type of ID and a first attribute information of the first type of ID, wherein the first type of ID is a randomly generated ID; Based on the data to be processed, target ID pairs are obtained, and the target ID pairs include the first type of ID and the second type of ID; The real-time table is updated based on the target ID, and the real-time table includes a first connection record of the first type ID, and the first connection record includes at least a first correspondence between the first type ID and the first attribute information of the first type ID; The real-time table is attached to the offline tables of each cluster in the data integration system. The offline table includes a second integration record of the second type ID. The second integration record includes at least a second correspondence between the second type ID and the second attribute information of the second type ID. The step of obtaining the target ID pair based on the data to be processed includes: In response to the inability to form a first type of candidate ID pair based on the data to be processed, at least one second type of candidate ID pair of the first type of ID is obtained from the IP bucket corresponding to the IP value of the first type of ID, based on the Internet Protocol IP value of the first type of ID, according to a preset matching threshold, a preset combination dimension and a random forest model. Analyze the confidence level of the at least one second-class candidate ID pair; The second type of candidate ID pair with a confidence level not less than a preset threshold is used as the target ID pair of the first type of ID; The method further includes: Retrieve the ID to be queried and the query instruction parameters input through the query interface; In response to detecting that the query indication parameter is a real-time query, based on the connection depth N, N connection queries are performed in the real-time table; the result of the Nth connection query is used as the query result of the ID to be queried; In response to detecting that the query indication parameter is a non-real-time query or that no query result for the ID to be queried is obtained from the real-time table, the query result for the ID to be queried is obtained from the offline table.
2. The method according to claim 1, wherein, The step of obtaining the target ID pair based on the data to be processed includes: In response to the fact that a first type of candidate ID pair can be formed based on the data to be processed, the first type of candidate ID pair is used as the target ID pair of the first type of ID.
3. The method according to claim 1, wherein, The analysis of the confidence level of the at least one second-class candidate ID pair includes: The confidence level of each second-class candidate ID pair is predicted using the first prediction model, thus obtaining the first confidence level of each second-class candidate ID pair; The confidence level of each second-class candidate ID pair is predicted using the second prediction model, thus obtaining the second confidence level of each second-class candidate ID pair; The confidence level of each second-class candidate ID pair is obtained based on the first confidence level and the second confidence level of each second-class candidate ID pair.
4. The method according to claim 1, further comprising: In response to the inability to obtain the target ID pair of the first type of ID from the IP bucket corresponding to the IP value, the first connection record of the first type of ID is generated in the real-time table.
5. The method according to claim 1, wherein, The update of the real-time table based on the target ID includes: Based on the target ID, query processing is performed on the real-time table for the first type of ID and the second type of ID, respectively. In response to finding the first type of ID in the real-time table, the second type of ID is added to the first access list of the first type of ID, and the attribute information of the first access list is updated; In response to finding the second type of ID in the real-time table, the first type of ID is added to the second access list of the second type of ID in the real-time table, and the attribute information of the second access list is updated; In response to the absence of the first type ID and the second type ID in the real-time table, a third access list is generated based on the first type ID and the second type ID, and the attribute information of the third access list is updated.
6. The method according to claim 1, wherein, The step of obtaining the target ID pair based on the data to be processed includes: Based on the identification information of the data to be processed, the ID data to be processed is sent to the first cluster in the data integration system that matches the identification information; The target ID pair is obtained through the first cluster; The step of updating the real-time table based on the target ID includes: The real-time table of the first cluster is updated based on the target ID through the first cluster.
7. The method according to claim 6, wherein, The step of mounting the real-time table to the offline tables of each cluster in the data integration system includes: The real-time table generated by the first cluster is mounted to the offline table of the first cluster; and The real-time table is sent to the second cluster of the data integration system so that the second cluster can mount the real-time table to the offline table of the second cluster. The second cluster is a cluster in the data integration system other than the first cluster.
8. A data processing apparatus, applied to a data integration system, comprising: The first acquisition module is used to acquire data to be processed, the data to be processed includes a first type ID and first attribute information of the first type ID, wherein the first type ID is a randomly generated ID; The second acquisition module is used to acquire target ID pairs based on the data to be processed, wherein the target ID pairs include the first type of ID and the second type of ID; The update module is used to update the real-time table based on the target ID. The real-time table includes a first connection record of the first type ID, and the first connection record includes at least a first correspondence between the first type ID and the first attribute information of the first type ID. The mounting module is used to mount the real-time table to the offline tables of each cluster in the data integration system. The offline table includes a second integration record of the second type ID. The second integration record includes at least a second correspondence between the second type ID and the second attribute information of the second type ID. The second acquisition module includes: The second acquisition submodule is used to respond to the fact that a first type of candidate ID pair cannot be formed based on the data to be processed, and to acquire at least one second type of candidate ID pair of the first type of ID from the IP bucket corresponding to the IP value of the first type of ID, based on the Internet Protocol IP value of the first type of ID, according to a preset matching threshold, a preset combination dimension and a random forest model. An analysis submodule is used to analyze the confidence level of the at least one second-type candidate ID pair; The third acquisition submodule is used to take the second type of candidate ID pair with a confidence level not less than a preset threshold as the target ID pair of the first type of ID; The device further includes: The generation module is used to obtain the query ID and query indication parameters input through the query interface; in response to detecting that the query indication parameters are real-time queries, it performs N interconnection queries in the real-time table based on the interconnection depth N; and uses the result of the Nth interconnection query as the query result of the query ID; in response to detecting that the query indication parameters are non-real-time queries or that the query result of the query ID is not obtained from the real-time table, it obtains the query result of the query ID from the offline table.
9. The apparatus according to claim 8, wherein, The second acquisition module includes: The first acquisition submodule is configured to, in response to the first type of candidate ID pair that can be formed based on the data to be processed, use the first type of candidate ID pair as the target ID pair of the first type of ID.
10. The apparatus according to claim 8, wherein, The analysis submodule is used for: The confidence level of each second-class candidate ID pair is predicted using the first prediction model, thus obtaining the first confidence level of each second-class candidate ID pair; The confidence level of each second-class candidate ID pair is predicted using the second prediction model, thus obtaining the second confidence level of each second-class candidate ID pair; The confidence level of each second-class candidate ID pair is obtained based on the first confidence level and the second confidence level of each second-class candidate ID pair.
11. The apparatus according to claim 8, wherein, The generation module is further configured to generate the first connection record of the first type of ID in the real-time table in response to the inability to obtain the target ID pair of the first type of ID from the IP bucket corresponding to the IP value.
12. The apparatus according to claim 8, wherein, The update module includes: The query submodule is used to perform query processing on the real-time table based on the target ID, including the first type of ID and the second type of ID; The first update submodule is used to respond to the query of the first type ID in the real-time table, add the second type ID to the first access list of the first type ID, and update the attribute information of the first access list; The second update submodule is used to respond to the query of the second type ID in the real-time table, add the first type ID to the second access list of the second type ID in the real-time table, and update the attribute information of the second access list; The third update submodule is used to generate a third access list based on the first type ID and the second type ID in response to the fact that the first type ID and the second type ID are not found in the real-time table, and to update the attribute information of the third access list.
13. The apparatus according to claim 8, wherein, The second acquisition module includes: The sending submodule is used to send the ID data to be processed to the first cluster in the data integration system that matches the identification information, based on the identification information of the data to be processed. The fourth acquisition submodule is used to acquire the target ID pair through the first cluster; The update module includes: The fourth update submodule is used to update the real-time table of the first cluster based on the target ID through the first cluster.
14. The apparatus according to claim 13, wherein, The mounting module includes: A first mounting submodule is used to mount the real-time table generated by the first cluster to the offline table of the first cluster; and The second mounting submodule is used to send the real-time table to the second cluster of the data integration system, so that the second cluster can mount the real-time table to the offline table of the second cluster. The second cluster is a cluster in the data integration system other than the first cluster.
15. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.
17. A computer program product comprising a computer program stored on a storage medium, the computer program implementing the method according to any one of claims 1-7 when executed by a processor.
Citation Information
Patent Citations
Data processing method and device, data query method and device, equipment and computer readable medium
CN112115154A
ID access method and device, and equipment
CN112835872A