Method, device and equipment for identity punch-through and identity punch-through result query, and medium

By acquiring user behavior data in real time and recording and identifying related relationships, and combining this data with offline data tables to generate query results, the problem of missing query results under the offline integration method is solved, achieving higher accuracy and efficiency.

CN116467306BActive Publication Date: 2026-02-17BAIDU (CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310333089.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2026-02-17
Estimated Expiration
2043-03-30

AI Technical Summary

Technical Problem

In existing technologies, the offline user identification method results in missing query results during the update interval, which reduces the accuracy of the query results.

Method used

By acquiring user behavior data in real time, recording and identifying relationships, and combining it with offline data tables to generate query results, the system uses real-time data tables to supplement missing content in offline data tables, and employs extended ID verification and lightweight strategy prediction models to improve accuracy.

Benefits of technology

It improves the accuracy and efficiency of querying results by integrating identifiers, reduces resource consumption, and enhances the timeliness and comprehensiveness of query results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116467306B_ABST
    Figure CN116467306B_ABST
Patent Text Reader

Abstract

The disclosure provides an identity breakthrough and identity breakthrough result query method, device, equipment and medium, and relates to the fields of artificial intelligence such as big data processing, deep learning and cloud computing. The identity breakthrough method can include: for each user behavior data acquired in real time, the following processing is performed respectively: extracting the identity of the user from the user behavior data, if the number of extracted identities is greater than one, each identity is combined with each other identity to form an identity pair; determining a target identity pair from the formed identity pairs and saving it to a real-time data table, the two identities in the target identity pair correspond to the same user, the real-time data table is used to generate a query result corresponding to a query request in combination with an offline data table, the query request is a query request for an identity breakthrough result, and the identity breakthrough result generated offline is saved in the offline data table. By applying the scheme of the disclosure, the accuracy of the query result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of artificial intelligence, in particular to an identification connection method and an identification connection result query method, device, equipment and medium in the fields of big data processing, deep learning and cloud computing. BACKGROUND

[0002] An identification (ID) connection service provides related capabilities for unified user personalized data, and plays an important role in business advertisement attribution and user behavior aggregation scenarios. SUMMARY

[0003] The present disclosure provides an identification connection method and an identification connection result query method, device, equipment and medium.

[0004] An identification connection method comprises:

[0005] Real-time user behavior data is acquired, and for each user behavior data acquired, the following processing is performed respectively:

[0006] From the user behavior data, the identification of a user is extracted, and in response to the number of extracted identifications being greater than one, each identification and each other identification other than itself forms an identification pair;

[0007] A target identification pair is determined from the formed identification pairs, and the target identification pair is saved to a real-time data table, the two identifications in the target identification pair correspond to the same user, the real-time data table is used to generate a query result corresponding to a query request in combination with an offline data table, the query request is a query request for an identification connection result, the offline data table saves an offline generated identification connection result, and the offline data table is periodically updated.

[0008] An identification connection result query method comprises:

[0009] A query request for an identification connection result is acquired;

[0010] In response to determining that the query request is a first type of query request, a corresponding query result is generated and returned according to the identification connection result in a real-time data table and an offline data table, the offline data table saves an offline generated identification connection result, the offline data table is periodically updated, the real-time data table saves a target identification pair determined from the formed identification pairs, the two identifications in the same target identification pair correspond to the same user, the identification pairs are identification pairs formed two by two from the identifications extracted from each user behavior data acquired in real time and the identifications of a user, and the number of extracted identifications is greater than one.

[0011] An identification connection device comprises a data acquisition module and an identification connection module.

[0012] The data acquisition module is configured to acquire user behavior data in real time.

[0013] The identifier connection module is configured to, for each user behavior data acquired, perform the following processing: extract an identifier of a user from the user behavior data, in response to the number of extracted identifiers being greater than one, form identifier pairs each comprising one of the extracted identifiers and each of the other extracted identifiers, determine a target identifier pair from the formed identifier pairs, the two identifiers in the target identifier pair corresponding to the same user, and save the target identifier pair into a real-time data table, the real-time data table being used to generate a query result corresponding to a query request in combination with an offline data table, the query request being a query request for an identifier connection result, the offline data table storing an offline generated identifier connection result, and the offline data table being periodically updated.

[0014] An identifier connection result query apparatus, comprising: a request acquisition module and a result generation module.

[0015] The request acquisition module is configured to acquire a query request for an identifier connection result.

[0016] The result generation module is configured to, in response to determining that the query request is a first type of query request, generate a corresponding query result according to an identifier connection result in a real-time data table and an offline data table and return the query result, the offline data table storing an offline generated identifier connection result, the offline data table being periodically updated, the real-time data table storing a target identifier pair determined from formed identifier pairs, the two identifiers in the target identifier pair corresponding to the same user, the identifier pairs being formed by using two extracted identifiers of a user from each of real-time acquired user behavior data, and the number of extracted identifiers being greater than one.

[0017] An electronic device, comprising:

[0018] at least one processor; and

[0019] a memory communicatively connected to the at least one processor; wherein

[0020] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0021] A non-transitory computer readable storage medium storing computer instructions for causing a computer to perform the method described above.

[0022] A computer program product comprising computer programs / instructions which, when executed by a processor, implement the method as described above.

[0023] It should be understood that the contents described in this part are not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0024] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:

[0025] Figure 1 Flow chart of the ID connection method embodiment described in the present disclosure;

[0026] Figure 2 Flow chart of the ID connection result query method embodiment described in the present disclosure;

[0027] Figure 3 Schematic diagram of the overall implementation process of the ID connection and ID connection result query method described in the present disclosure;

[0028] Figure 4 Schematic diagram of the component structure of the ID connection device embodiment 400 described in the present disclosure;

[0029] Figure 5 Schematic diagram of the component structure of the ID connection result query device embodiment 500 described in the present disclosure;

[0030] Figure 6 A schematic block diagram of an electronic device 600 that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0031] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.

[0032] In addition, it should be understood that the term "and / or" herein merely describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it.

[0033] Figure 1A flowchart of the ID connection method embodiment of the present disclosure is shown in FIG. 1. As shown in FIG. 1, the following specific implementation is included. Figure 1

[0034] In step 101, user behavior data is acquired in real time, and each user behavior data is processed according to the manner shown in steps 102-103.

[0035] In step 102, the ID of a user is extracted from the user behavior data, and in response to the number of extracted IDs being greater than one, each ID forms an ID pair with each of the other IDs.

[0036] In step 103, a target ID pair is determined from the formed ID pairs, and the target ID pair is saved to a real-time data table. The two IDs in the target ID pair correspond to the same user, and the real-time data table is used to generate a query result corresponding to a query request in combination with an offline data table. The query request is a query request for an ID connection result, the offline data table stores an offline-generated ID connection result, and the offline data table is periodically updated.

[0037] In the conventional manner, an offline connection method is usually used, such as updating the offline data table once by offline calculation based on all user IDs captured at present at a predetermined time interval. Since there are many upstream product lines, the data is large and disordered, and the task flow is long, the value of the predetermined time interval is usually not too short, such as 3 days or 5 days. The offline data table can also be referred to as a core user data table, etc., in which the ID connection results of different users are recorded, i.e., different types of IDs of the same user are connected, such as can be represented in the form of a topological graph, etc. Different types of IDs of the same user can include a login account, a device ID, and a web cookie, etc.

[0038] However, this method at least has the following problems: during the interval between two updates, when a user queries the ID connection result, some results will be missing, such as the results between the last update and the current time, thereby reducing the accuracy of the query result.

[0039] However, the scheme of the above method embodiment can record the newly added ID association relationship in the real-time data table according to the user behavior data acquired in real time, thereby supplementing the missing content in the offline data table, i.e., supplementing the IDs not captured by the offline delay, and accordingly, the query result corresponding to the query request can be generated in combination with the real-time data table and the offline data table, thereby improving the accuracy of the query result, etc.

[0040] ​There is no limitation on how to obtain the user behavior data. For example, the user behavior data can be obtained from a real-time data source. For example, when a user logs in a shopping website and browses a product, a user behavior data can be generated. Similarly, other behaviors of the user can generate corresponding user behavior data.

[0041] For each user behavior data obtained, the user behavior data can be processed according to the steps 102 and 103.

[0042] First, the ID of the user can be extracted from the user behavior data. The number of extracted IDs can be greater than one or equal to one. Accordingly, different processing methods can be used subsequently. In addition, preferably, before extracting the ID of the user from the user behavior data, the user behavior data can be identified as abnormal. In response to identifying the user behavior data as abnormal data, the user behavior data can be filtered out. For example, the user behavior data corresponding to an Internet Protocol (IP) not in an IP whitelist can be identified as abnormal data, or the user behavior data corresponding to an abnormal timestamp can be identified as abnormal data.

[0043] Through the above processing, abnormal data can be filtered out in advance, thereby reducing the workload of subsequent processing and reducing resource consumption.

[0044] If the number of extracted IDs is greater than one, each ID can be paired with other IDs except itself to form an ID pair. For example, assuming that three different types of IDs are extracted, namely ID1, ID2 and ID3, three ID pairs can be formed, namely an ID pair of ID1 and ID2, an ID pair of ID1 and ID3, and an ID pair of ID2 and ID3.

[0045] Then, a target ID pair can be determined from the formed ID pairs, and the target ID pair can be saved to a real-time data table. In theory, each formed ID pair can be directly used as a target ID pair. Preferably, for each formed ID pair, the following processing can be performed: according to the direct association relationship between the IDs saved in the real-time data table, the extended IDs of the two IDs in the ID pair are determined, the ID pair is verified according to the extended IDs, in response to determining that the verification is passed, it is determined that the two IDs in the ID pair have a direct association relationship, and the two IDs having the direct association relationship are saved to the real-time data table, and the two IDs having the direct association relationship correspond to the same user.

[0046] In actual applications, due to various reasons, the two IDs appearing in the same user behavior data can correspond to different users. Accordingly, the formed ID pairs can be verified, and only the ID pairs that pass the verification can be saved to the real-time data table, thereby improving the accuracy of the content in the real-time data table.

[0047] In addition, if it is determined that any composed ID pair already exists in the real-time data table, subsequent processing can be performed without saving repeatedly.

[0048] Preferably, for any ID pair of the composition, the determination of the extended IDs of the two IDs in the ID pair can include: determining the IDs having a direct association relationship with the first ID and the second ID respectively as one degree extended IDs according to the direct association relationship between the IDs saved in the real-time data table, the first ID and the second ID being the two IDs in the ID pair, and then determining the IDs having a direct association relationship with each one degree extended ID as two degree extended IDs, and further taking the one degree extended IDs and the two degree extended IDs as the required extended IDs.

[0049] Through the above processing, the required extended IDs can be efficiently and accurately obtained, thereby laying a good foundation for subsequent processing.

[0050] According to the obtained extended IDs, the ID pair can be verified. Preferably, the ID confidence of each extended ID can be obtained, and each extended ID can be sorted in descending order of ID confidence. In response to determining that the extended IDs in the first M positions after sorting correspond to the same CoreID, it can be determined that the verification of the ID pair is passed, M being a positive integer and M being less than or equal to the number of extended IDs, and the CoreID being the highest priority ID among all the IDs of the corresponding user (the user corresponding to the CoreID), and correspondingly, in response to determining that the verification is passed, it can be determined that the two IDs in the ID pair have a direct association relationship, and can be saved to the real-time data table. Specifically, the ID pair can be saved to the real-time data table, and the corresponding CoreID can be recorded, i.e. the corresponding relationship with the CoreID is established, and the specific saving and recording manner is not limited.

[0051] For example, assuming that one degree expansion is performed using the first ID in an ID pair, 2 one-degree expanded IDs are obtained, one degree expansion is performed using the second ID, 2 one-degree expanded IDs are also obtained, and further, two degrees of expansion are performed using the 4 one-degree expanded IDs, assuming that 2 two-degree expanded IDs are obtained respectively, thus, 4+8=12 expanded IDs are obtained in total, further, ID confidence of the 12 expanded IDs can be obtained respectively, and the 12 expanded IDs can be sorted in descending order of ID confidence, in response to determining that the expanded IDs in the top 3 positions correspond to the same CoreID after sorting, it can be determined that the ID pair passes the verification, accordingly, it can be determined that the two IDs in the ID pair have a direct association relationship, and can be saved to the real-time data table, in addition, the corresponding CoreID can also be recorded.

[0052] Different priorities can be set for different types of IDs in advance, such as the highest priority for login accounts, followed by device IPs, and so on. For example, in the offline data table, different types of IDs of the same user can be recorded in the form of a topology diagram in descending order of priority, and the ID with the highest priority is the CoreID. The required CoreID can be obtained by querying the offline data table, or the CoreID corresponding to the same user in the real-time data table can be directly used.

[0053] In addition, for any ID pair, in addition to saving the two IDs as having a direct association relationship to the real-time data table, the attribute information corresponding to the two IDs respectively can also be saved to the real-time data table, and the attribute information specifically includes which content can be determined according to actual needs, such as including the corresponding product line, timestamp, and browser model, and accordingly, when generating a query result for a query request, the attribute information can be returned together with the corresponding ID.

[0054] Through the above processing, the newly added ID association relationship can be recorded in the real-time data table, and can be mounted to the stable natural person (user) ID connection result mined offline, so that the real-time data table can be used to supplement the missing content in the offline data table.

[0055] In actual application, the content in the real-time data table can also be filtered periodically, such as old information can be eliminated, or part of the repeated content in the real-time data table can be eliminated according to the content in the offline data table, and the specific manner is not limited, so as to save storage resources.

[0056] Preferably, the ID confidence of each extended ID can be obtained in the following manner: for each extended ID, the following is performed: according to attribute information corresponding to the extended ID, an initial confidence of the extended ID is determined, and a product of the initial confidence and the ID confidence of the previous level ID of the extended ID is obtained, and a ratio of the product to 100 is taken as the ID confidence of the extended ID, wherein if the extended ID is a first degree extended ID, the previous level ID of the extended ID is the corresponding first ID or second ID, and if the extended ID is a second degree extended ID, the previous level ID of the extended ID is the corresponding first degree extended ID, and the ID confidence of the first ID and the second ID is 100.

[0057] The attribute information specifically includes which content can be determined according to actual needs, and how to determine the initial confidence of the extended ID according to the attribute information can also be determined according to actual needs, for example, the initial confidence can be calculated according to a pre-set calculation rule.

[0058] Suppose for a certain ID pair, which includes two IDs IDa and IDb, one degree extension is performed using IDa, and 2 first degree extended IDs are obtained, which are IDc and IDd, one degree extension is also performed using IDb, and 2 first degree extended IDs are also obtained, which are IDe and IDf, further, two degree extension is performed using the 4 first degree extended IDs respectively, and 2 second degree extended IDs are obtained, wherein it is assumed that the second degree extended IDs obtained using IDc are IDh and IDi.

[0059] Then, taking IDc as an example, the initial confidence thereof can be obtained first, then the product of the initial confidence of IDc and the ID confidence of IDa can be obtained, and then the ratio of the product to 100 can be taken as the ID confidence of IDc, and taking IDh as an example, the initial confidence thereof can be obtained first, then the product of the initial confidence of IDh and the ID confidence of IDc can be obtained, and then the ratio of the product to 100 can be taken as the ID confidence of IDh.

[0060] The ID confidence of the first ID and the second ID is set to 100, the ID confidence of the extended ID is not more than 100, and the initial confidence of the next level extended ID needs to be multiplied by the ID confidence of the previous level extended ID, and then divided by 100 to be taken as the ID confidence of the next level extended ID.

[0061] That is, dst_id_confidence = src_id_confidence * initial_confidence / 100; (1)

[0062] Wherein, dst_id_confidence represents the ID confidence of the next level extended ID, and src_id_confidence represents the ID confidence of the previous level extended ID.

[0063] The above describes the processing manner when the number of IDs extracted from the user behavior data is greater than one.

[0064] Preferably, if the number of extracted IDs is equal to one, the extracted ID and corresponding attribute information, which is also extracted from the user behavior data, can be added to the prediction set. Then, for the extracted ID, it can be paired with other IDs in the prediction set that meet predetermined requirements to form ID pairs, and for each ID pair, the following processing can be performed: according to the attribute information corresponding to the two IDs in the ID pair, a lightweight strategy prediction model pre-trained is used to determine whether the two IDs correspond to the same user. If the determination result is yes, the ID pair can be saved to the real-time data table, and the corresponding CoreID can be recorded.

[0065] In addition, preferably, for the extracted ID, the way of pairing it with other IDs that meet predetermined requirements to form ID pairs can include: pairing it with each of the other IDs in the prediction set, respectively, or pairing it with each of the other IDs in the IP bucket where it is located, respectively. The IDs in the prediction set belong to at least two different IP buckets.

[0066] The IP bucket can refer to grouping the IPs in Beijing into a group, grouping the IPs in Shanghai into a group, etc., that is, grouping according to cities. Generally speaking, for two IDs, if the IPs corresponding to them belong to Beijing and Shanghai respectively, the likelihood that the two IDs correspond to the same user is small. Therefore, for the extracted ID, it can only be paired with each of the other IDs in the IP bucket where it is located to save the workload of subsequent processing, etc., or to avoid omissions, it can also be paired with each of the other IDs in the prediction set to improve the comprehensiveness and accuracy of the processing results, etc. The specific way to be adopted can be determined according to actual needs, which is very flexible and convenient.

[0067] For each ID pair formed, according to the attribute information corresponding to the two IDs in the ID pair, a lightweight strategy prediction model can be used to determine whether the two IDs correspond to the same user. For example, the two IDs and the attribute information corresponding to them respectively can be taken as the input of the lightweight strategy prediction model to obtain the confidence of the output, and then the obtained confidence can be compared with a predetermined threshold. If it is greater than the threshold, it can be determined that the two IDs correspond to the same user, otherwise, it can be determined that the two IDs correspond to different users. The specific value of the threshold can be determined according to actual needs.

[0068] If it is determined that the two IDs correspond to the same user, they can be saved to the real-time data table as the ID pair that has a direct association relationship, and the corresponding CoreID can be recorded.

[0069] Through the above processing, the case that the two possible IDs belong to the same user can be solved in real time, helping to recall more correct user IDs, improving the ID connection rate, etc.

[0070] In addition, it can be seen that, in the scheme of the disclosure, different processing methods can be used for different numbers of IDs extracted from user behavior data, so that the processing is more targeted, and the accuracy of the processing result is further improved, etc.

[0071] In actual application, the lightweight strategy prediction model can be trained based on the constructed training sample, and the accuracy of the lightweight strategy prediction model can be improved through the diversity of the training sample.

[0072] For example, for each ID in different IP buckets, two-by-two combination can be performed, then according to the discrimination of system version, browser, brand and other features, a logistic regression (LR) model can be used to screen out ID pairs that are roughly likely to correspond to the same user, then through artificial correction and annotation, positive and negative training samples can be constructed, and a lightweight strategy prediction model can be trained based on the set feature selection.

[0073] The real-time data table obtained according to the scheme of the disclosure can be used to generate a query result corresponding to a query request, the query request is a query request for the ID connection result, the query result is a query result generated in combination with an offline data table, the offline data table stores the offline generated ID connection result, and the offline data table is periodically updated.

[0074] In addition, preferably, the scheme of the disclosure can further include: in response to the cluster being a master cluster in the set N clusters, synchronizing the real-time data table to each non-master cluster, N being a positive integer greater than one, and the master cluster and the non-master cluster being respectively used for generating a query result for a query request sent by a user in a subordinate region range.

[0075] For example, for the overall region range (such as the whole country), a plurality of different clusters can be set, each cluster can be subordinate to a different region range, such as North China, Northeast China, etc., and one of the clusters can be determined as the master cluster, and correspondingly, the other clusters are non-master clusters. The execution subject of the scheme of the disclosure can be located in the master cluster, and the real-time data table can be synchronized to each non-master cluster. The master cluster and each non-master cluster can be respectively used for generating a query result for a query request sent by a user in a subordinate region range, that is, only users in the subordinate region range can be provided with query services, without involving cross-cluster queries, so that the transmission delay can be reduced, and the response speed can be improved, etc.

[0076] Correspondingly, Figure 2 The flowchart of the ID connection result query method embodiment of the disclosure is shown in FIG. 1.Figure 2 The specific implementation is shown below.

[0077] In step 201, a query request for ID connection results is acquired.

[0078] In step 202, in response to determining that the query request is a first type of query request, a corresponding query result is generated and returned according to ID connection results in a real-time data table and an offline data table, the offline data table stores offline generated ID connection results, the offline data table is periodically updated, and the real-time data table stores target ID pairs determined from composed ID pairs, two IDs in the same target ID pair correspond to the same user, the ID pairs are ID pairs composed of each user behavior data acquired in real time and IDs of the user extracted from the same user behavior data, and the number of extracted IDs is greater than one.

[0079] According to the above method embodiment, the newly added ID association relationship can be recorded in the real-time data table according to the user behavior data acquired in real time, so that the real-time data table can be used to supplement the missing content in the offline data table, that is, as a supplement to the ID not captured by the offline delay, and accordingly, the query result corresponding to the query request can be generated in combination with the real-time data table and the offline data table, thereby improving the accuracy of the query result and the like.

[0080] Preferably, in response to determining that the query request is a second type of query request, a corresponding query result can be generated and returned according to the ID connection results in the offline data table.

[0081] That is, the user who issues the query request can select which type of query to perform according to his own needs, that is, different query modes can be provided for the user to select, which is very flexible and convenient.

[0082] In addition, preferably, the manner of determining that the query request is the first type of query request can include reading a predetermined parameter carried in the query request, and in response to determining that the value of the predetermined parameter is a first value, determining that the query request is the first type of query request, and the manner of determining that the query request is the second type of query request can include reading a predetermined parameter carried in the query request, and in response to determining that the value of the predetermined parameter is a second value, determining that the query request is the second type of query request.

[0083] The predetermined parameter can be a RealTime ID query parameter. If a user wants to obtain a more timely and more abundant query result, the value of RealTime can be set to True, such as 1, that is, the real-time ID is selected to be used to query, and correspondingly, the query result can be generated and returned by combining the ID in the real-time data table and the ID in the offline data table. Otherwise, the value of RealTime can be set to 0, that is, the real-time ID is not selected to be used to query, and correspondingly, the query result can be generated and returned according to the ID in the offline data table.

[0084] In addition, the information carried in the query request can be determined according to actual needs. For example, in addition to the predetermined parameter, the query object and the query type can be included, such as the query object can be one / multiple cookies, and the query type can be used to specify which type of ID is returned, and correspondingly, the ID of the query type of the same user as the cookie(s) can be returned (which can also include corresponding attribute information, etc.).

[0085] As can be seen, by means of the predetermined parameter, the user can conveniently select the query mode, and no matter which query mode is selected, the required query result can be quickly obtained, and the accuracy of the query result is ensured.

[0086] In combination with the above, Figure 3 The whole implementation process of the ID connection and the ID connection result query method of the present disclosure is shown in the schematic diagram.

[0087] As Figure 3 shown, the user behavior data can be obtained in real time, and the ID of the user can be extracted from each user behavior data, and the extracted ID can be saved to the middleware such as the data center (datehub) after being processed such as de-duplication. Then, real-time calculation can be performed, that is, the ID pairs with direct association relationship generated by the real-time calculation can be directly filled into the real-time data table according to the different number of IDs extracted from the user behavior data, and the data can be filled into the real-time data table after being converted into the required format by means of hash / shifting according to the predetermined format, and the predetermined format is not limited, such as simpleDB.

[0088] In addition, as Figure 3 shown, the above operation can be completed in the master cluster, and the real-time data table can also be updated to each non-master cluster, and correspondingly, each cluster (such as Figure 3The real-time data table exists in each of the cluster A, the cluster B and the cluster C shown in the figure, and can provide a query service through a related interface, and then can generate a corresponding query result and return for different types of query requests.

[0089] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the disclosure is not limited by the action sequence described, because according to the disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily necessary for the disclosure. In addition, the parts not described in detail in a certain embodiment can refer to the related description in other embodiments.

[0090] The above is the introduction of the method embodiment, and the scheme described in the disclosure will be further described through the device embodiment.

[0091] Figure 4 The constituent structure schematic diagram of the ID punch-through device embodiment 400 described in the disclosure is shown in the figure. Figure 4 As shown, it includes a data acquisition module 401 and an ID punch-through module 402.

[0092] The data acquisition module 401 is used to acquire user behavior data in real time.

[0093] The ID punch-through module 402 is used to perform the following processing for each user behavior data acquired: extract the ID of the user from the user behavior data, in response to the number of extracted IDs being greater than one, respectively form ID pairs with each ID and other IDs except itself, determine a target ID pair from the formed ID pairs, the two IDs in the target ID pair correspond to the same user, save the target ID pair to a real-time data table, the real-time data table is used to generate a query result corresponding to a query request in combination with an offline data table, the query request is a query request for the ID punch-through result, the offline data table saves the ID punch-through result generated offline, and the offline data table is updated periodically.

[0094] The scheme described in the above device embodiment can record the newly added ID association relationship in the real-time data table according to the user behavior data acquired in real time, so as to supplement the missing content in the offline data table by using the real-time data table, that is, as a supplement to the ID not captured by the offline delay, and correspondingly, the query result corresponding to the query request can be generated in combination with the real-time data table and the offline data table, and the accuracy of the query result is improved.

[0095] For each piece of acquired user behavior data, first, the identification and connection module 402 can extract the ID of the user therefrom, and the number of extracted IDs can be greater than one or equal to one, and accordingly, different processing methods can be used subsequently. In addition, preferably, the identification and connection module 402 can perform abnormality identification on the user behavior data before extracting the ID of the user therefrom, and in response to identifying the user behavior data as abnormal data, the user behavior data can be filtered out.

[0096] If the number of extracted IDs is greater than one, the identification and connection module 402 can form ID pairs respectively from each ID and other IDs other than itself, and the identification and connection module 402 can determine a target ID pair from the formed ID pairs and save the target ID pair to the real-time data table. In theory, each of the formed ID pairs can be directly used as a target ID pair, but preferably, the following processing can be performed respectively for each of the formed ID pairs: determining the extended IDs of the two IDs in the ID pair according to the direct association relationship between the IDs saved in the real-time data table, verifying the ID pair according to the extended IDs, in response to determining that the verification is passed, determining that the two IDs in the ID pair have a direct association relationship and saving them to the real-time data table, and the two IDs having a direct association relationship correspond to the same user.

[0097] Preferably, the identification and connection module 402 can determine the extended IDs of the two IDs in any ID pair in the following manner: determining the ID having a direct association relationship with the first ID and the ID having a direct association relationship with the second ID as one-degree extended IDs according to the direct association relationship between the IDs saved in the real-time data table, the first ID and the second ID being the two IDs in the ID pair, and then determining the IDs having a direct association relationship with each one-degree extended ID as two-degree extended IDs according to the direct association relationship between the IDs saved in the real-time data table, and further using the one-degree extended IDs and the two-degree extended IDs as the required extended IDs.

[0098] According to the obtained each extended ID, the identity breaking through module 402 can check the ID pair. Preferably, the identity breaking through module 402 can obtain the ID confidence of each extended ID respectively, and can sort each extended ID according to the order from high to low of the ID confidence, and in response to determining that the extended ID in the first M positions after sorting corresponds to the same CoreID, it can be determined that the check of the ID pair is passed, M is a positive integer, and M is less than or equal to the number of extended IDs, CoreID is the highest priority ID among all IDs corresponding to the user obtained, accordingly, in response to determining that the check is passed, it can be determined that there is a direct association relationship between the two IDs in the ID pair, and can be saved to the real-time data table. Specifically, the ID pair can be saved to the real-time data table, and the corresponding CoreID can be recorded, that is, the corresponding relationship between the CoreID is established.

[0099] Preferably, the identity breaking through module 402 can obtain the ID confidence of each extended ID in the following way: for each extended ID, the following processing is performed respectively: according to the attribute information corresponding to the extended ID, the initial confidence of the extended ID is determined, and the product of the initial confidence and the ID confidence of the upper level ID of the extended ID is obtained. The ratio of the product to 100 is taken as the ID confidence of the extended ID, wherein if the extended ID is a first extended ID, the upper level ID thereof is the corresponding first ID or second ID, and if the extended ID is a second extended ID, the upper level ID thereof is the corresponding first extended ID. The ID confidence of the first ID and the second ID is 100.

[0100] The above describes the processing method when the number of IDs extracted from the user behavior data is greater than one. Preferably, if the number of extracted IDs is equal to one, the identity breaking through module 402 can add the extracted ID and the corresponding attribute information to the prediction set, and the attribute information is also extracted from the user behavior data. Then, for the extracted ID, it can be combined with other IDs in the prediction set that meet the predetermined requirements to form an ID pair, and for each ID pair formed, the following processing can be performed respectively: according to the attribute information corresponding to the two IDs, using the lightweight strategy prediction model trained in advance, it is determined whether the two IDs correspond to the same user. In response to the determination result being yes, the ID pair can be saved to the real-time data table, and the corresponding CoreID can be recorded.

[0101] In addition, preferably, for the extracted ID, the identity breaking through module 402 can form an ID pair with other IDs in the prediction set that meet the predetermined requirements in the following way: forming an ID pair with each ID in the prediction set respectively, or forming an ID pair with each ID in the IP bucket where it is located respectively. The IDs in the prediction set belong to at least two different IP buckets.

[0102] In addition, preferably, the identity punch-through module 402 is further configured to synchronize the real-time data table to each non-primary cluster in response to the cluster where the identity punch-through module 402 is located being a primary cluster in the set of N clusters, N being a positive integer greater than 1, the primary cluster and the non-primary cluster being used to generate query results for query requests sent by users within a subordinate region range, respectively.

[0103] Figure 5 A constituent structure schematic diagram of an embodiment 500 of the ID punch-through result query apparatus disclosed in the present disclosure is shown in FIG. 5. As shown in FIG. 5, the embodiment 500 comprises a request obtaining module 501 and a result generating module 502. Figure 5

[0104] The request obtaining module 501 is configured to obtain a query request for an identity punch-through result.

[0105] The result generating module 502 is configured to, in response to determining that the query request is a first type of query request, generate and return a corresponding query result according to an ID punch-through result in a real-time data table and an offline data table, the offline data table storing offline-generated ID punch-through results, the offline data table being periodically updated, the real-time data table storing a target ID pair determined from composed ID pairs, two IDs in the same target ID pair corresponding to the same user, the ID pair being an ID pair composed of each two of user behavior data obtained in real time and an ID of a user extracted from the same user behavior data, the number of extracted IDs being greater than one.

[0106] By using the above-mentioned apparatus embodiment and scheme, the newly added ID association relationship can be recorded in the real-time data table according to the user behavior data obtained in real time, so that the real-time data table can be used to supplement the missing content in the offline data table, i.e., as a supplement to the ID not captured by the offline delay, and accordingly, the query result corresponding to the query request can be generated in combination with the real-time data table and the offline data table, thereby improving the accuracy of the query result and the like.

[0107] Preferably, the result generating module 502 is configured to, in response to determining that the query request is a second type of query request, generate and return a corresponding query result according to the ID punch-through result in the offline data table.

[0108] In addition, preferably, the manner in which the result generating module 502 determines that the query request is the first type of query request can comprise reading a predetermined parameter carried in the query request, and determining that the query request is the first type of query request in response to determining that the value of the predetermined parameter is a first value, and the manner in which the result generating module 502 determines that the query request is the second type of query request can comprise reading the predetermined parameter carried in the query request, and determining that the query request is the second type of query request in response to determining that the value of the predetermined parameter is a second value.

[0109] ​The predetermined parameter can be RealTime. If a user wants to obtain a query result with higher timeliness and richer content, the value of RealTime can be set to True, for example, 1, that is, the real-time ID is selected to be used to query, and correspondingly, the query result can be generated and returned by combining the ID in the real-time data table and the ID in the offline data table. Otherwise, the value of RealTime can be set to 0, that is, the real-time ID is not selected to be used to query, and correspondingly, the query result can be generated and returned according to the ID in the offline data table.

[0110] Figure 4 and Figure 5 The specific working process of the device embodiment shown in the figure can be referred to the related description in the foregoing method embodiment, and will not be described here.

[0111] The scheme described in the disclosure can be applied to the field of artificial intelligence, and particularly relates to the fields of big data processing, deep learning, and cloud computing. Artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) of people. Artificial intelligence has both hardware technology and software technology. Artificial intelligence hardware technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, and several other directions.

[0112] In addition, the user behavior data and the like in the embodiments of the disclosure are not for a specific user, and cannot reflect the personal information of a specific user. In the technical scheme of the disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with the provisions of relevant laws and regulations, and do not violate public order and good customs.

[0113] According to the embodiments of the disclosure, the disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0114] Figure 6 A schematic block diagram of an electronic device 600 that can be used to implement embodiments of the disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, servers, blades, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit implementations of the present disclosure described and / or claimed in this document.

[0115] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0116] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0117] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as those described in this disclosure. For example, in some embodiments, the methods described in this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the methods described in this disclosure can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the methods described in this disclosure by any other suitable means (e.g., by means of firmware).

[0118] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0119] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.

[0120] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0121] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0122] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0123] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0124] It should be understood that various forms of flow shown above can be used, with steps reordered, added, or removed. For example, steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without limitation, so long as the desired results of the technology disclosed in the present disclosure are achieved.

[0125] The specific embodiments described above are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that various modifications, combinations, sub-combinations, and alternatives can be made to the specific embodiments without departing from the spirit and principles of the present disclosure. Any further modifications, changes, improvements, and the like that come within the spirit and principles of the present disclosure should be considered within the scope of the present disclosure.

Claims

1. An identity breakthrough method, comprising: obtaining user behavior data in real time, and for each obtained user behavior data, performing the following processing: extracting identities of a user from the user behavior data, and in response to the number of extracted identities being greater than one, forming each identity with each other identity except itself into an identity pair; determining a target identity pair from the formed identity pairs, comprising: for each formed identity pair, determining extended identities of two identities in the identity pair according to direct association relationships between identities saved in a real-time data table, verifying the identity pair according to identity confidence of the extended identities, and in response to determining that the verification passes, determining that the identity pair is the target identity pair; and saving the target identity pair into the real-time data table, the two identities in the target identity pair corresponding to a same user, the real-time data table being used to generate a query result corresponding to a query request in combination with an offline data table, the query request being a query request for an identity breakthrough result, the offline data table saving offline generated identity breakthrough results, and the offline data table being periodically updated; wherein for each extended identity, an initial confidence of the extended identity is determined according to attribute information corresponding to the extended identity, and a product of the initial confidence and identity confidence of an upper-level identity of the extended identity is obtained, and a ratio of the product to 100 is taken as the identity confidence of the extended identity.

2. The method of claim 1, wherein, The determining of the extended identities of the two identities in the identity pair comprises: determining, according to the direct association relationships between identities saved in the real-time data table, identities having a direct association relationship with a first identity and identities having a direct association relationship with a second identity as one-degree extended identities, the first identity and the second identity being the two identities in the identity pair; determining, according to the direct association relationships between identities saved in the real-time data table, identities having a direct association relationship with each one-degree extended identity as two-degree extended identities; taking the one-degree extended identities and the two-degree extended identities as the extended identities.

3. The method of claim 2, wherein: the verifying of the identity pair according to the identity confidence of the extended identities comprises: obtaining identity confidence of each extended identity, and sorting each extended identity in order of identity confidence from high to low, and in response to determining that the extended identities in the top M positions after sorting correspond to a same core identity, determining that the verification of the identity pair passes, M being a positive integer and M being less than or equal to the number of extended identities, the core identity being a highest priority identity among all identities corresponding to the user obtained; the saving of the target identity pair into the real-time data table comprises: saving the target identity pair into the real-time data table and recording the corresponding core identity.

4. The method of claim 3, wherein: If the extended identifier is the one-time extended identifier, the upper-level identifier is the corresponding first identifier or the second identifier, and if the extended identifier is the two-time extended identifier, the upper-level identifier is the corresponding one-time extended identifier, and the identification confidence of the first identifier and the second identifier is 100.

5. The method of claim 1, further comprising: in response to the number of extracted identifiers being equal to one, adding the extracted identifier and corresponding attribute information to a prediction set, the attribute information also being extracted from the user behavior data; and, for the extracted identifier, respectively forming an identifier pair with other identifiers in the prediction set that meet predetermined requirements, and for each formed identifier pair, respectively performing the following processing: determining whether the two identifiers correspond to the same user according to the attribute information corresponding to the two identifiers using a pre-trained lightweight strategy prediction model, and in response to the determination result being yes, saving the identifier pair to the real-time data table.

6. The method of claim 5, wherein, The forming of the identifier pair with other identifiers in the prediction set that meet predetermined requirements comprises: respectively forming an identifier pair with each of the other identifiers in the prediction set; or, respectively forming an identifier pair with each of the other identifiers in an Internet protocol bucket in which the identifier belongs, the identifiers in the prediction set belonging to at least two different Internet protocol buckets.

7. The method of claim 1, further comprising: for any acquired user behavior data, before extracting the identifier of a user from the user behavior data, performing anomaly identification on the user behavior data, and in response to identifying the user behavior data as abnormal data, filtering out the user behavior data.

8. The method of any one of claims 1-7, further comprising: in response to the cluster being a master cluster in the set of N clusters, synchronizing the real-time data table to each non-master cluster, N being a positive integer greater than one, the master cluster and the non-master cluster being respectively used to generate query results for query requests sent by users within a subordinate region.

9. An identifier breakthrough result query method, comprising: acquiring a query request for an identifier breakthrough result; In response to determining that the query request is a first type of query request, a corresponding query result is generated and returned according to an online data table and an offline data table, the offline data table storing offline generated identification connection results, the offline data table being periodically updated, the online data table storing a target identification pair determined from the identification pairs, two identifications in the same target identification pair corresponding to the same user, the identification pairs being identification pairs formed by each two of identifications extracted from real-time acquired user behavior data, the number of the extracted identifications being greater than one, the target identification pair being an identification pair determined from the identification pairs according to the following steps: expanding identifications in the identification pair according to a direct correlation between the identifications stored in the online data table, and verifying the identification pair according to an identification confidence of the expanded identifications, and determining the identification pair as passing the verification, the identification confidence of any expanded identification being determined according to the following steps: determining an initial confidence of the expanded identification according to attribute information corresponding to the expanded identification, and obtaining a product of the initial confidence and an identification confidence of a previous level identification of the expanded identification, and then obtaining a ratio of the product to 100.

10. The method of claim 9, further comprising: In response to determining that the query request is a second type of query request, a corresponding query result is generated and returned according to the identification connection results in the offline data table.

11. The method of claim 10, wherein: the determining that the query request is the first type of query request comprises reading a predetermined parameter carried in the query request, and in response to determining that a value of the predetermined parameter is a first value, determining that the query request is the first type of query request; the determining that the query request is the second type of query request comprises reading the predetermined parameter carried in the query request, and in response to determining that a value of the predetermined parameter is a second value, determining that the query request is the second type of query request.

12. An identification punch-through device comprising: a data acquisition module and an identification connection module; the data acquisition module is configured to acquire user behavior data in real time; The identity breakthrough module is configured to, for each user behavior data obtained, perform the following processing: extracting an identity of a user from the user behavior data, in response to the number of extracted identities being greater than one, forming each identity with each other identity except itself into an identity pair, determining a target identity pair from the formed identity pairs, including: for each identity pair, determining an extended identity of each identity in the identity pair according to a direct association relationship between identities saved in a real-time data table, verifying the identity pair according to an identity confidence of the extended identity, and in response to determining that the verification is passed, determining that the identity pair is the target identity pair; and saving the target identity pair into the real-time data table, the two identities in the target identity pair corresponding to a same user, the real-time data table being used to generate a query result corresponding to a query request in combination with an offline data table, the query request being a query request for an identity breakthrough result, the offline data table saving an offline generated identity breakthrough result, and the offline data table being periodically updated; wherein for each extended identity, an initial confidence of the extended identity is determined according to attribute information corresponding to the extended identity, and a product of the initial confidence and an identity confidence of an upper-level identity of the extended identity is obtained, and a ratio of the product to 100 is taken as the identity confidence of the extended identity.

13. The apparatus of claim 12, wherein, The identity breakthrough module determines, according to the direct association relationship between identities saved in the real-time data table, an identity having a direct association relationship with the first identity and an identity having a direct association relationship with the second identity as first-degree extended identities, the first identity and the second identity being the two identities in the identity pair, determines, according to the direct association relationship between identities saved in the real-time data table, an identity having a direct association relationship with each first-degree extended identity as a second-degree extended identity, and takes the first-degree extended identities and the second-degree extended identity as the extended identities.

14. The apparatus of claim 13, wherein, The identity breakthrough module obtains an identity confidence of each extended identity, and sorts each extended identity in an order from high to low according to the identity confidence, in response to determining that the extended identities in the first M positions after the sorting correspond to a same core identity, determines that the verification of the identity pair is passed, M being a positive integer and M being less than or equal to the number of extended identities, and the core identity being a highest-priority identity among all identities corresponding to the user obtained; The identity breakthrough module, in response to determining that the verification is passed, saves the identity pair into the real-time data table and records the corresponding core identity.

15. The apparatus of claim 14, wherein, If the extended identity is the first-degree extended identity, the upper-level identity is the corresponding first identity or the second identity, and if the extended identity is the second-degree extended identity, the upper-level identity is the corresponding first-degree extended identity, and the identity confidence of the first identity and the second identity is 100.

16. The apparatus of claim 12, wherein, the identity bridging module is further configured to, in response to the number of extracted identities being equal to one, add the extracted identity and corresponding attribute information to a prediction set, the attribute information also being extracted from the user behavior data, and for each of the extracted identities, form an identity pair with each of the other identities in the prediction set that meets a predetermined requirement, and for each of the formed identity pairs, determine whether the two identities in the identity pair correspond to the same user according to attribute information corresponding to the two identities using a pre-trained lightweight strategy prediction model, and in response to a determination result being yes, save the identity pair to the real-time data table.

17. The apparatus of claim 16, wherein, the identity bridging module forms an identity pair for each of the extracted identities with each of the other identities in the prediction set, or forms an identity pair for each of the extracted identities with each of the other identities in an internet protocol bucket in which the extracted identities belong to, the identities in the prediction set belonging to at least two different internet protocol buckets.

18. The apparatus of claim 12, wherein, the identity bridging module is further configured to, for any of the obtained user behavior data, perform anomaly identification on the user behavior data before extracting an identity of a user from the user behavior data, and in response to identifying the user behavior data as abnormal data, filter out the user behavior data.

19. The apparatus of any of claims 12-18, wherein, the identity bridging module is further configured to, in response to the cluster being a master cluster of the set of N clusters, synchronize the real-time data table to each of the non-master clusters, N being a positive integer greater than one, the master cluster and the non-master clusters each being configured to generate a query result for a query request sent by a user in a range of the master cluster and the non-master clusters.

20. An apparatus for identifying a break-out result query, comprising: the request obtaining module and the result generating module; the request obtaining module is configured to obtain a query request for an identity bridging result. The result generation module is configured to, in response to determining that the query request is a first type of query request, generate and return a corresponding query result according to an online data table and an offline data table. The offline data table stores offline generated identity connection results. The offline data table is periodically updated. The online data table stores a target identity pair determined from the composed identity pairs. Two identities in the same target identity pair correspond to the same user. The identity pairs are composed of two identities of each user behavior data obtained in real time and a user identity extracted from the same user behavior data. The number of extracted identities is greater than one. The target identity pair is determined from the composed identity pairs as follows: according to a direct association relationship between the identities stored in the online data table, an extended identity of each identity in the identity pair is determined, and after the identity pair is verified according to the identity confidence of the extended identity, the identity pair that passes the verification is determined. The identity confidence of any extended identity is determined as follows: according to attribute information corresponding to the extended identity, an initial confidence of the extended identity is determined, and after the initial confidence and the identity confidence of the previous level identity of the extended identity are multiplied, a ratio of the product to 100 is obtained.

21. The apparatus of claim 20, wherein, The result generation module is further configured to, in response to determining that the query request is a second type of query request, generate and return a corresponding query result according to the identity connection results in the offline data table.

22. The apparatus of claim 21, wherein, The result generation module reads a predetermined parameter carried in the query request, determines that the query request is the first type of query request in response to determining that the predetermined parameter has a first value, and determines that the query request is the second type of query request in response to determining that the predetermined parameter has a second value.

23. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.

24. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-11.

25. A computer program product comprising computer programs / instructions, which, when executed by a processor, implement the method of any one of claims 1-11.

Citation Information

Patent Citations

  • User ID (Identification) recognition method and device

    CN105099729A