A data processing system for obtaining a wobble user identification
Patent Information
- Application Number
- CN202410592518.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2044-05-14
AI Technical Summary
[0004]用户行为数据较为复杂,人工对用户行为数据进行分析确定摇摆用户具有不可控性,并且摇摆用户的用户行为数据可能会与使用APP的新用户、偶尔使用APP的用户等用户的用户行为数据的相似度较高,存在分析结果错误的情况,因此,通过上述方法获取到的摇摆用户的精准度较低
[0017]本发明提供了一种获取摇摆用户标识的数据处理系统,所述数据处理系统能够根据第一预设用户标识列表和第二预设用户标识列表获取第三预设用户标识列表,根据第三预设用户标识列表获取目标APP对应的用户标识匹配度和目标APP对应的数据相关系数,当用户标识匹配度不小于预设标识匹配度且数据相关系数不小于预设相关系数时,获取目标APP对应的第一预设用户标识列表中的第一预设用户标识,根据第一预设用户标识获取目标APP对应的摇摆用户标识,无需人工对用户行为数据进行分析获取摇摆用户标识,有利于提高获取摇摆用户标识的精准度。
Smart Images

Figure CN118349709B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of APP technology, and in particular to a data processing system for obtaining the identifiers of swing users. Background Technology
[0002] For an app, a "swing user" refers to a user who constantly switches between the app and other apps that are different from the app or have partially overlapping customer groups. Identifying these swing users and analyzing their behavioral data allows for a better understanding of user behavior, facilitating refined app operations and ultimately improving user retention. Currently, most methods for identifying swing users involve acquiring user behavior data from app users, and then analyzing this data to determine if a user is indeed a swing user for that app.
[0003] However, the above method also has the following technical problems:
[0004] User behavior data is complex, and manually analyzing it to identify swing users is uncontrollable. Furthermore, the user behavior data of swing users may be highly similar to that of new users of the app or occasional users, which may lead to errors in the analysis results. Therefore, the accuracy of swing users obtained through the above methods is low. Summary of the Invention
[0005] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows:
[0006] This invention provides a data processing system for obtaining swing user identifiers. The data processing system includes a processor and a memory storing a computer program. When the computer program is executed by the processor, the following steps are implemented:
[0007] S1. Obtain the user identifier matching degree BS corresponding to the target APP and the data correlation coefficient XS corresponding to the target APP. Step S1 also includes the following sub-steps S11-S15:
[0008] S11. Obtain the first preset user identifier list B corresponding to the target APP. B includes several first preset user identifiers. The first preset user identifiers are user identifiers of users who have used the target APP collected by the target SDK.
[0009] S12. Obtain the second preset user identifier list A corresponding to the target APP. A includes several second preset user identifiers, which are user identifiers of users who have used the target APP stored in the database.
[0010] S13. Use the intersection of B and A as the third preset user identifier list C = {C1, C2, ..., C...} i , ..., C m}, C i Let i be the i-th third preset user identifier, where i ranges from 1 to m, and m is the number of third preset user identifiers.
[0011] S14. Obtain BS based on A and C, where BS satisfies the following conditions:
[0012] BS = m / A 0 , where A 0 The number of second preset user identifiers in A.
[0013] S15, When BS < BS 0 When XS = 0, and when BS ≥ BS0, according to C i Obtain XS, where BS 0 This is the preset identifier matching degree.
[0014] S2, when BS≥BS 0 And XS≥XS 0 When, obtain B = {B1, B2, ..., B} e , ..., B f}, where XS 0 To preset the correlation coefficient, B e Let e be the first preset user identifier corresponding to the target APP, where e ranges from 1 to f, and f is the number of first preset user identifiers corresponding to the target APP.
[0015] S3, according to B e Obtain the swing user identifier corresponding to the target APP.
[0016] The present invention has at least the following beneficial effects:
[0017] This invention provides a data processing system for obtaining swing user identifiers. The system can obtain a third preset user identifier list based on a first preset user identifier list and a second preset user identifier list. It then obtains the user identifier matching degree and the data correlation coefficient corresponding to the target app based on the third preset user identifier list. When the user identifier matching degree is not less than a preset identifier matching degree and the data correlation coefficient is not less than a preset correlation coefficient, it obtains the first preset user identifier from the first preset user identifier list corresponding to the target app. Finally, it obtains the swing user identifier corresponding to the target app based on the first preset user identifier. This eliminates the need for manual analysis of user behavior data to obtain swing user identifiers, thus improving the accuracy of obtaining swing user identifiers. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating how a data processing system, as provided in an embodiment of the present invention, executes a computer program to obtain a swing user identifier. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar tasks and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0022] Embodiments of the present invention provide a data processing system for obtaining swing user identifiers. The data processing system includes a processor and a memory storing a computer program. When the computer program is executed by the processor, it performs the following steps: Figure 1 As shown:
[0023] Specifically, the "swing user" identifier is a unique identifier for a swing user. A swing user is a user who constantly switches between two or more apps whose customer groups partially overlap. For example, if the customer groups of the first XX app and the second XX app are partially the same and the user constantly switches between the first XX app and the second XX app, using the first XX app one day and the second XX app the other day, then the user is a swing user corresponding to the first XX app and the second XX app.
[0024] S1. Obtain the user identifier matching degree BS corresponding to the target APP and the data correlation coefficient XS corresponding to the target APP.
[0025] Specifically, S1 includes the following sub-steps S11-S15:
[0026] S11. Obtain the first preset user identifier list B corresponding to the target APP. B includes several first preset user identifiers. The first preset user identifiers are user identifiers of users who have used the target APP collected by the target SDK.
[0027] S12. Obtain the second preset user identifier list A corresponding to the target APP. A includes several second preset user identifiers. The second preset user identifiers are user identifiers of users who have used the target APP stored in the database; they can be understood as user identifiers provided by the target APP's server.
[0028] S13. Use the intersection of B and A as the third preset user identifier list C = {C1, C2, ..., C...} i , ..., C m}, C i Let i be the i-th third preset user identifier, where i ranges from 1 to m, and m is the number of third preset user identifiers. As those skilled in the art know, any method in the prior art for obtaining the intersection of two lists is within the protection scope of this invention, and will not be elaborated here.
[0029] S14. Obtain BS based on A and C, where BS satisfies the following conditions:
[0030] BS = m / A 0 , where A 0 The number of second preset user identifiers in A.
[0031] Specifically, the higher the user identifier matching degree, the more the user identifiers collected by the SDK match the user identifiers stored in the database, which can be understood as being more similar.
[0032] S15, When BS < BS 0 When XS = 0, and when BS ≥ BS0, according to C i Obtain XS, where BS 0 This is the preset identifier matching degree.
[0033] Specifically, BS 0 The value range is [0, 1].
[0034] Through the above steps, a first and second preset user identifier list are obtained. A third preset user identifier list is then obtained based on these lists. Furthermore, the user identifier matching degree is calculated. A higher matching degree indicates a better match between the user identifiers collected by the SDK and those stored in the database. Therefore, if the matching degree is less than the preset matching degree, it means the user identifiers collected by the SDK do not match the database. Consequently, the accuracy of obtaining the target app's corresponding swing user identifiers based on the SDK-collected user identifiers is low, and there is no need to continue obtaining the data correlation coefficient. If the matching degree is not less than the preset matching degree, it indicates a good match between the SDK-collected user identifiers and the database. In this case, the data correlation coefficient is obtained based on the third preset user identifier to further determine the accuracy of obtaining the target app's corresponding swing user identifiers based on the SDK-collected user identifiers, which helps improve the accuracy of obtaining the target app's corresponding swing user identifiers.
[0035] Specifically, step S15 includes the following sub-steps S151-S156:
[0036] S151, Obtain C i The corresponding first usage time period list SY i ={SY i1 SY i2 , ..., SY ia , ..., SY ic(i)}, where SY ia C i The corresponding first usage time period of the 'a'th time period, where 'a' takes values from 1 to c(i), and c(i) is C i The corresponding number of the first usage time period, SY ia The starting time point is SY 1 ia SY ia The end time point is SY 2 ia The start and end times of the first usage period are: the times when the user corresponding to the third preset user identifier stored in the database opens the target APP and the times when the user closes the target APP within the preset time period; it can be understood as: the times when the user corresponding to the third preset user identifier provided by the server of the target APP opens the target APP and the times when the user closes the target APP within the preset time period. The preset time period is a time period pre-set by those skilled in the art, and will not be elaborated here.
[0037] Specifically, SY 2 ia >SY 1ia SY 2 ia <SY 1 i(a+1) SY 1 i(a+1) For SY i(a+1) The starting time point, SY i(a+1) C i The corresponding (a+1)th first usage time period.
[0038] S152, when SY 2 ia -SY 1 ia When ≥△t1, SY ia The length of C i The corresponding first active duration to obtain C i The corresponding first active duration list D i ={D i1 D i2 , ..., D ij , ..., D in(i)}, D ij C i The corresponding j-th first active duration, where j takes values from 1 to n(i), and n(i) is C. i The corresponding number of first active durations, where △t1 is the first preset active duration. As those skilled in the art know, the specific value of the first preset active duration is determined by those skilled in the art according to actual needs, and will not be elaborated here.
[0039] S153, according to D i , get C i The first adjustment parameter C of the corresponding data correlation coefficient 0 i C 0 i The following conditions must be met:
[0040] C 0 i =(C 1 i +C 2 i ) / 2, where C 1 i C i The corresponding adjustment for C 0 i The first parameter, C 1 i The following conditions must be met:
[0041] C 1 i=(n(i)-min(n(1),n(2),…,n(i),…,n(n))) / (max(n(1),n(2),…,n(i),…,
[0042] n(n))-min(n(1),n(2),……,n(i),……,n(n))), where main() is the function to get the minimum value and max() is the function to get the maximum value;
[0043] C 2 i C i The corresponding adjustment for C 0 i The second parameter, C 2 i The following conditions must be met:
[0044] C 2 i =(∑ n(i) j=1 D ij -min(∑ n(1) j=1 D 1j , ∑ n(2) j=1 D 2j , ..., ∑ n(i) j=1 D ij , ..., ∑ n(n) j= 1D nj )) / (max(∑ n(1) j=1 D 1j ,
[0045] ∑ n(2) j=1 D 2j , ..., ∑ n(i) j=1 D ij , ..., ∑ n(n) j=1 D nj )-min(∑ n(1) j=1 D 1j , ∑ n(2) j= 1D 2j , ..., ∑ n(i) j=1 D ij , ..., ∑ n(n) j=1 D nj )).
[0046] S154, Obtain C i The corresponding second usage time period list DE i ={DE i1 DE i2 , ..., DE i(ag) , ..., DE i(ah(i))}, where DE i(ag) C i The corresponding ag-th second usage time period, where ag takes values from 1 to ah(i), and ah(i) is C. i The corresponding number of second usage time periods, DE i(ag) The starting time point is DE 1 i(ag) DE i(ag) The end time point is DE 2 i(ag) The start and end times of the second usage period are: the times when the user corresponding to the third preset user identifier collected by the target SDK opens the target APP and the times when the target APP closes within the preset time period.
[0047] Specifically, DE 2 i(ag) >DE 1 i(ag) DE 2 i(ag) <DE 1 i(ag+1) DE 1 i(ag+1) For DE i(ag+1) The starting time point, DE i(ag+1) C i The corresponding ag+1th second usage time period.
[0048] S155, according to DE 2 i(ag) and DE 1 i(ag) Get C i The second adjustment parameter E of the corresponding data correlation coefficient 0 i , where, according to DE 2 i(ag) and DE 1 i(ag) Get E 0 i The methods and steps in S152-S153 are based on SY 1 ia and SY 2 ia Get C 0 iThe methods are basically the same, and those skilled in the art can refer to steps S152-S153 according to DE. 2 i(ag) and DE 1 i(ag) Get E 0 i This will not be elaborated further here; it can be understood as: SY in step S152 2 ia Replace with DE 2 i(ag) SY 1 ia Replace with DE 1 i(ag) And execute steps S152-S153 to obtain E 0 i .
[0049] S156, According to C 0 i and E 0 i Obtain XS, where XS meets the following conditions:
[0050] XS=∑ m i=1 ((C 0 i -∑ m i=1 C 0 i / m)×(E 0 i -∑ m i=1 E 0 i / m)) / ((∑ m i=1 (C 0 i -∑ m i=1 C 0 i / m) 2 ) 1 / 2 ×(∑ m i=1 (E 0 i -∑
[0051] m i=1 E 0 i / m) 2 ) 1 / 2 ).
[0052] Specifically, the higher the data correlation coefficient, the stronger the correlation between the data collected by the SDK and the data stored in the database.
[0053] Through the above steps, a first usage time period list corresponding to the third preset user identifier is obtained. Based on the first usage time period list, a first active duration list corresponding to the third preset user identifier is obtained. Based on the first active duration list, a first adjustment parameter for the data correlation coefficient is obtained. A second usage time period list corresponding to the third preset user identifier is obtained. Based on the second usage time period list, a second adjustment parameter for the data correlation coefficient is obtained. Based on the first correlation coefficient adjustment parameter and the second adjustment parameter for the data correlation coefficient, the data correlation coefficient is obtained. The higher the data correlation coefficient, the stronger the correlation between the data collected by the SDK and the data stored in the database. The stronger the data correlation, the higher the accuracy of obtaining the swing user identifier corresponding to the target APP based on the user identifier collected by the SDK. Therefore, when the data correlation coefficient is not less than the preset correlation coefficient, the obtained swing user identifier corresponding to the target APP is conducive to improving the accuracy of obtaining the swing user identifier corresponding to the target APP.
[0054] S2, when BS≥BS 0 And XS≥XS 0 When, obtain B = {B1, B2, ..., B} e , ..., B f}, where XS 0 To preset the correlation coefficient, B e Let e be the first preset user identifier corresponding to the target APP, where e ranges from 1 to f, and f is the number of first preset user identifiers corresponding to the target APP.
[0055] Specifically, XS 0 The value range is [0, 1].
[0056] S3, according to B e Obtain the swing user identifier corresponding to the target APP.
[0057] Specifically, step S3 includes the following steps S31-S37:
[0058] S31, Obtain B e The corresponding first associated APP identifier list F e ={F e1 F e2 , ..., F eg , ..., F eh(e)}, where F eg For B e The corresponding g-th first associated APP identifier, where g takes values from 1 to h(e), and h(e) is B. eThe number of corresponding first associated APP identifiers. The first associated APP identifier is a unique identifier of the first associated APP. The first associated APP is an APP that integrates the target SDK on the mobile terminal device used by the user corresponding to the first preset user identifier. The mobile terminal device can be understood as a mobile phone or a tablet computer. For example, if the first XX APP is installed on the mobile phone used by the user corresponding to the first preset user identifier, and the first XX APP integrates the target SDK, then the first XX APP is the first associated APP corresponding to the first preset user identifier.
[0059] S32. Obtain the list of historical time slices T = {T1, T2, ..., T...} corresponding to the historical time period. x , ..., T p}, where T x Let x be the x-th historical time slice corresponding to the historical time period, where x ranges from 1 to p, and p is the number of historical time slices corresponding to the historical time period. The length of the historical time slice is 24 hours. The length of the historical time period is determined by those skilled in the art based on actual needs, and will not be elaborated here.
[0060] S33, Obtain F eg In T x The corresponding third usage time period list G xeg ={G 1 xeg G 2 xeg , ..., G y xeg , ..., G q(xeg) xeg}, G y xeg For F eg In T x The corresponding y-th third usage time period, where y ranges from 1 to q(xeg), and q(xeg) is F. eg In T x The number of times G corresponds to the third usage time period. y xeg The starting time point is G1 y xeg G y xeg The end time point is G2 y xeg The start and end times of the third usage period are: the time when the user corresponding to the first preset user identifier, which is the first associated APP identifier collected by the target SDK, opens the APP corresponding to the first associated APP identifier and the time when the user closes the APP corresponding to the first associated APP identifier in the historical time slice.
[0061] Specifically, G2y xeg >G1 y xeg G2 y xeg <G1 y+1 xeg G1 y+1 xeg For F eg In T x The corresponding y+1th third usage time period.
[0062] S34, when G2 y xeg -G1 y xeg When ≥△t1, G y xeg The length of F eg In T x The corresponding third active duration is used to obtain F. eg In T x The corresponding third active duration list FT_xeg = {FT_xeg1, FT_xeg2, ..., FT_xeg} k , ..., FT_xeg t(xeg)}, FT_xeg k For F eg In T x The corresponding k-th third active duration, where k ranges from 1 to t(xeg), and t(xeg) is F. eg In T x The number of the third active durations corresponding to this.
[0063] S35, when ∑ h(e) g=1 t(xeg)≥△PC and ∑ h(e) g=1 ∑ t(xeg) k=1 FT_xeg k When ≥△t2, T x As B e The corresponding active time slice to obtain B e Corresponding active time slice list HY e Where △PC is the preset usage frequency value, △t2 is the second preset active duration, and HY e It includes several active time slices. As those skilled in the art know, the specific values of the preset usage frequency value and the second preset active duration are set by those skilled in the art according to actual needs, and will not be elaborated here.
[0064] S36, when HY eWhen the number of active time slices in the HY is not less than the preset number of time slices, HY will... e Corresponding B e As the active user identifiers corresponding to the target APP, obtain the list of active user identifiers H = {H1, H2, ..., H...} for the target APP. u H v}, where H u Let u be the active user identifier corresponding to the target APP, where u ranges from 1 to v, and v is the number of active user identifiers corresponding to the target APP. As those skilled in the art know, the specific value of the preset number of time slices is set by those skilled in the art according to actual needs, and will not be elaborated here.
[0065] S37. Based on H, obtain the swing user identifier corresponding to the target APP.
[0066] Through the above steps, a list of first associated APP identifiers corresponding to the first preset user identifier is obtained. A third time period list is obtained based on the first associated APP identifier list and the historical time slice list. A third active duration list is obtained based on the third time period list. An active time slice list corresponding to the first preset user identifier is obtained based on the third active duration list. When the number of active time slices in the active time slice list is not less than the preset number of time slices, it indicates that the first preset user identifier corresponding to the active time slice list is an active user identifier, not a fake user identifier. Therefore, through the above steps, fake user identifiers can be filtered out, which helps improve the accuracy of obtaining the active user identifier list. Obtaining the swing user identifier corresponding to the target APP based on the active user identifier list helps improve the accuracy of obtaining the swing user identifier corresponding to the target APP. Furthermore, since it does not obtain the swing user identifier corresponding to the target APP based on all the first preset user identifiers, the computational load is reduced, which also helps improve the efficiency of obtaining the swing user identifier corresponding to the target APP.
[0067] Specifically, step S37 includes the following sub-steps S371-S377:
[0068] S371, Obtain H from the xth historical time slice. u List of the first historical usage time period corresponding to the target APP J ux ={J ux1 J ux2 , ..., J uxγ , ..., J uxθ(ux)}, where J uxγ For the xth historical time slice H u J is the γth first historical usage time period corresponding to the target APP. uxγ The starting time point is J1 uxγ J uxγThe end time point is J2 uxγ The start and end times of the first historical usage period are: the times when the user corresponding to the active user identifier collected by the target SDK opens the target APP and the times when the target APP closes in the historical time slice.
[0069] Specifically, J2 uxγ >J1 uxγ J2 uxγ <J1 ux(γ+1) J1 ux(γ+1) For J ux(γ+1) The starting time point, J ux(γ+1) For the xth historical time slice H u The (γ+1)th first historical usage time period corresponding to the target APP.
[0070] S372, according to J ux Get H from the xth historical time slice u K, the first historical usage duration list corresponding to the target app ux ={K ux1 K ux2 , ..., K ux(ai) , ..., K ux(am(x))}, where K ux(ai) For the xth historical time slice, H u The ai-th first historical usage duration corresponding to the target APP, where ai ranges from 1 to am(x), and am(x) is the H in the x-th historical time slice. u The number of first historical usage times corresponding to the target APP.
[0071] Specifically, step S372 includes the following sub-steps:
[0072] S3721. When ai = 1, proceed to step S3722; when ai ≠ 1, set J... 1 ux((ai)-1) As J ux Then proceed to step S3722, J 1 ux((ai)-1) For K ux((ai)-1) The corresponding updated first historical usage time period list, K ux((ai)-1) For the xth historical time slice, H u The ai-1th first historical usage duration corresponding to the target APP.
[0073] S3722, when J2 uxγ -J1 ux1 ≤△t3 and J1 ux(γ+1) -J1 ux1 When Δt3 > 0, determine K. ux(ai) =J0 ux1 +J 0 ux2 +……J 0 uxη +……+J 0 uxγ and J ux1 J ux2 , ..., J uxη , ..., J uxγ From J ux Delete to get K ux(ai) The corresponding updated first historical usage time period list J 1 ux(ai) Among them, J1 ux1 For J ux1 The starting time point, △t3 is the third preset active duration, J 0 uxη For J uxη Duration, J uxη For the xth historical time slice H u The ηth first historical usage time period corresponding to the target APP, where the value of η ranges from 1 to γ.
[0074] S373, according to K ux Get H in the historical time period u The total historical usage frequency value L corresponding to the target APP u and H in historical time period u Total historical usage time M corresponding to the target APP u , where L u The following conditions must be met:
[0075] L u =∑ p x=1 am(x); M u The following conditions must be met:
[0076] M u =∑ p x=1 ∑ am(x) ai=1 K ux(ai) .
[0077] S374. Obtain the preset APP identifier combination list R = {R1, R2, ..., R...} z , ..., R w}, where R z For the z-th preset APP identifier combination, z takes values from 1 to w, w is the number of preset APP identifier combinations, and R z By R z1 and R z2Composition, R z1 R is the first preset APP identifier in the z-th preset APP identifier combination. z2 For R z1 The corresponding second preset APP identifier, wherein the customer groups of the first preset APP corresponding to the first preset APP identifier and the second preset APP corresponding to the second preset APP identifier partially overlap, which can be understood as: the business areas of the first preset APP and the second preset APP are the same or similar. For example: the first preset APP provides food delivery services, and the second preset APP also provides food delivery services.
[0078] Specifically, the first preset APP identifier is the unique identifier of the first preset APP, and the second preset APP identifier is the unique identifier of the second preset APP.
[0079] Specifically, the preset APP identifier combination list is a list of APP identifier combinations pre-set in the database.
[0080] S375, Traverse R, when R... z1 When = MB, R z1 The corresponding R z2 As the second associated APP identifier corresponding to MB, when R z2 When = MB, R z2 The corresponding R z1 As the identifier of the second associated APP corresponding to MB, we can obtain the list of the second associated APPs corresponding to MB, U = {U1, U2, ..., U...} α , ..., U β}, where MB is the target app identifier, which is the unique identifier of the target app, and U α Let α be the αth second associated APP identifier corresponding to MB, where α ranges from 1 to β, and β is the number of second associated APP identifiers corresponding to MB.
[0081] Specifically, the second associated APP identifier is the unique identity identifier of the second associated APP.
[0082] S376, according to H u and U α Get H in the historical time period u and U α The corresponding historical total usage frequency value P uα and H in historical time period u and U α The corresponding total historical usage time Q uα , where, according to H u and U α Get P uα and Q uαThe methods and steps in S371-S373 are based on H u Get L from the target APP u and M u The methods are basically the same, and those skilled in the art can refer to steps S371-S373 according to H u and U α Get P uα and Q uα This will not be elaborated further; it can be understood as: replacing the target APP in step S371 with U. α The corresponding second associated APP is executed, and steps S371-S373 are performed to obtain P. uα and Q uα .
[0083] S377, when L u M u P uα Q uα Simultaneously satisfying L u ≥count、P uα ≥count、M u ≥△t4、Q uα ≥△t4、BL1≤L u / P uα ≤BL2、BL1≤M u / Q uα When ≤BL2, determine H u The target app is identified by the swing user identifier, where count is the preset total usage frequency value, △t4 is the fourth preset active duration, BL1 is the first preset ratio, and BL2 is the second preset ratio. As those skilled in the art know, the specific values of the preset total usage frequency value and the fourth preset active duration are set by those skilled in the art according to actual needs, and will not be elaborated here.
[0084] Specifically, determine H u Also for U α The corresponding swing user identifier for the second associated app can be understood as: H u The corresponding users are in the target APP and U α The corresponding second-related apps are constantly switched and used.
[0085] Specifically, the value range of BL1 is [0, 1].
[0086] Specifically, the value range of BL2 is [1, 2].
[0087] Through the above steps, historical usage time periods are obtained; a historical usage duration list is obtained based on the historical usage time periods; the historical total usage frequency and historical total usage duration corresponding to the active user identifier and the target APP are obtained based on the historical usage duration; a preset APP identifier combination list is obtained; a second associated APP identifier is obtained based on the preset APP identifier combination list; and the historical total usage frequency and historical total usage duration corresponding to the active user identifier and the second associated APP identifier are obtained. When the historical total usage frequency corresponding to the active user identifier and the target APP is not less than a preset total usage frequency value, and the historical total usage duration corresponding to the active user identifier and the target APP is not less than a fourth preset active duration, it indicates that the user corresponding to the active user identifier frequently uses the target APP. When the historical total usage frequency corresponding to the active user identifier and the second associated APP identifier is not less than a preset total usage frequency value, and the historical total usage duration corresponding to the active user identifier and the second associated APP identifier is not less than a fourth preset active duration, it indicates that the user corresponding to the active user identifier frequently uses the second associated APP identifier. The second associated APP, when the ratio of the total historical usage frequency value corresponding to the active user identifier and the target APP to the total historical usage frequency value corresponding to the active user identifier and the second associated APP identifier is not less than a first preset ratio and not greater than a second preset ratio, indicates that the active user uses the target APP and the second associated APP corresponding to the second associated APP identifier approximately the same number of times. When the ratio of the total historical usage duration corresponding to the active user identifier and the target APP to the total historical usage duration corresponding to the active user identifier and the second associated APP identifier is not less than a first preset ratio and not greater than a second preset ratio, it indicates that the active user uses the target APP and the second associated APP corresponding to the second associated APP identifier for approximately the same duration. Therefore, it can be determined that the active user identifier is both a swing user identifier corresponding to the target APP and a swing user corresponding to the second associated APP corresponding to the second associated APP. Therefore, obtaining the swing user identifier corresponding to the target APP through the above steps helps to improve the accuracy of obtaining the swing user identifier corresponding to the target APP.
[0088] In one specific embodiment, the following steps are included after step S1:
[0089] S4, when BS < BS 0 Or XS < XS 0 At that time, obtain the target series of APP identifier list V = {V1, V2, ..., V...} corresponding to the target APP. (ae) , ..., V (af)}, where V (ae)Let a be the ath target series APP identifier corresponding to the target APP, where a takes values from 1 to af, and af is the number of target series APP identifiers corresponding to the target APP. The target series APP identifier is a unique identifier for the target series APP. The target series APP is an APP that belongs to the same operating entity as the target APP and integrates the target SDK. As those skilled in the art know, any method in the prior art for obtaining other APPs that belong to the same operating entity as the APP is within the protection scope of this invention, and will not be elaborated here.
[0090] S5. Obtain the swing user identifier corresponding to V based on V.
[0091] Specifically, step S5 includes the following sub-steps S51-S57:
[0092] S51. Obtain the list of key user identifiers corresponding to V. The list of key user identifiers includes several key user identifiers, which are user identifiers collected by the target SDK for users using any one of the target series of APPs.
[0093] S52. Obtain the list of active user identifiers corresponding to V based on the key user identifier: HYV = {HYV1, HYV2, ..., HYV} (ar) , ..., HYV (as)}, HYV (ar) Let ar be the active user identifier corresponding to V, where ar ranges from 1 to as, and as is the number of active user identifiers corresponding to V. The method and steps S31-S36 for obtaining HYV based on the key user identifier are based on B. e The method for obtaining H is basically the same. Those skilled in the art can refer to steps S31-S36 to obtain HYV based on the key user identifier, which will not be repeated here; it can be understood as: taking B in step S31 e Replace with the key user identifier and perform steps S31-S36 to obtain the HYV.
[0094] S53, according to HYV (ar) and V (ae) Get HYV in historical time period (ar) and V (ae) The corresponding historical total usage frequency value AB (ar) (ae) HYV in historical time periods (ar) and V (ae) Corresponding total historical usage time AC (ar) (ae) Among them, according to HYV (ar) and V (ae) Get AB (ar) (ae) and AC(ar) (ae) The methods and steps in S371-S373 are based on H u Get L from the target APP u and M u The methods are basically the same, and those skilled in the art can refer to steps S371-S373 according to HYV (ar) and V (ae) Get AB (ar) (ae) and AC (ar) (ae) This will not be elaborated further here; it can be understood as: H in step S371 u Replace with HYV (ar) Replace the target app with V (ae) The corresponding target series of APPs are obtained by executing steps S371-S373. (ar) (ae) and AC (ar) (ae) .
[0095] S54, according to AB (ar) (ae) and AC (ar) (ae) Get HYV (ar) The historical cumulative usage frequency value AE corresponding to V ar Historical cumulative usage time AF ar , among which, AE ar The following conditions must be met:
[0096] AE ar =∑ af ae=1 AB (ar) (ae) ;AFar meets the following conditions:
[0097] AF ar =∑ af ae=1 AB (ar) (ae) .
[0098] S55. Obtain the set of associated APP identifiers corresponding to the target APP, AG = {AG1, AG2, ..., AG...} (ak) , ..., AG (at)}, where AG (ak)This refers to the list of the ak-th associated series APP identifiers corresponding to the target APP, where ak ranges from 1 to at, and at is the number of associated series APP identifiers corresponding to the target APP. The list of associated series APP identifiers includes several associated series APP identifiers, which serve as unique identifiers for associated series APPs. The associated series APPs corresponding to the same associated series APP identifier in the same list all integrate the target SDK, belong to the same operating entity, and have some overlap with the customer groups of the operating entity to which the target APP belongs. This can be understood as: all associated series APPs corresponding to all associated series APP identifiers in the list belong to the same operating entity, and the operating entity has the same or similar business scope as the operating entity to which the target APP belongs.
[0099] S56, according to HYV (ar) and AG (ak) Get HYV (ar) and AG (ak) The corresponding historical cumulative usage frequency value AH (ar) (ak) AR and historical cumulative usage time (ar) (ak) Among them, according to HYV (ar) and AG (ak) Get AH (ar) (ak) and AR (ar) (ak) The methods and steps in S53-S54 are based on HYV (ar) and V (ae) Get After Effects ar and AF ar The methods are basically the same, and those skilled in the art can refer to steps S53-S54 according to HYV (ar) and AG (ak) Get AH (ar) (ak) and AR (ar) (ak) This will not be elaborated further here; it can be understood as: V in step S53 (ae) Replace with AG (ak) And execute steps S53-S54 to obtain AH (ar) (ak) and AR (ar) (ak) .
[0100] S57, when AE ar AF ar AH (ar) (ak) AR (ar) (ak)Simultaneously satisfy AE ar ≥AJ、AH (ar) (ak) ≥AJ、AF ar ≥AT、AR (ar) (ak) ≥AT、BL1≤AE ar / AH (ar) (ak) ≤BL2、BL1≤AF ar / AR (ar) (ak) When ≤BL2, determine HYV (ar) V represents the swing user identifier, where AJ is the preset cumulative usage frequency value and AT is the preset cumulative usage duration. As those skilled in the art know, the specific values of the preset cumulative usage frequency value and the preset cumulative usage duration are set by those skilled in the art according to actual needs, and will not be elaborated here.
[0101] Specifically, determine HYV (ar) Also for AG (ak) The corresponding swing user identifier can be understood as: HYV (ar) The corresponding user in V is identified by the target series APP and AG. (ak) The associated series of apps in the middle are constantly switched and used.
[0102] Through the above steps, when the user identifier matching degree is less than the preset identifier matching degree or the data correlation coefficient is less than the preset correlation coefficient, it indicates that the user identifiers collected by the SDK and the user identifiers stored in the database do not match well or the correlation between the data collected by the SDK and the data stored in the database is weak. At this time, the accuracy of obtaining the swing user identifiers corresponding to the target APP is low. In this case, obtain the target series APP identifier list corresponding to the target APP, and obtain the swing user identifiers corresponding to the target series APP identifier list based on the target series APP identifier list. By expanding the scope of obtaining swing user identifiers, the accuracy of obtaining the swing user identifiers corresponding to the target series APP identifier list can be improved.
[0103] In one specific embodiment, when the computer program is executed by the processor, the following steps are also performed:
[0104] S10. Obtain the target series APP identifier list V = {V1, V2, ..., V...} corresponding to the target APP. (ae) , ..., V (af)}
[0105] S20, Obtain V (ae) The corresponding user identifier matching degree AK(ae) and V (ae) The corresponding data correlation coefficient AL (ae) Among them, obtaining AK (ae) and AL (ae) The method is basically the same as the method for obtaining BS and XS in step S1. Those skilled in the art can refer to step S1 to obtain AK. (ae) and AL (ae) This will not be elaborated further; it can be understood as: replacing the target APP in step S1 with V. (ae) The corresponding target series of apps will be executed and step S1 will be performed to obtain the AK. (ae) and AL (ae) .
[0106] S30, according to AK (ae) and AL (ae) Obtain the user identifier matching degree AM and data correlation coefficient AN corresponding to V, where AM meets the following conditions:
[0107] AM = ∑ af ae=1 AK (ae) / af;AN meets the following conditions:
[0108] AN=∑ af ae=1 AL (ae) / af.
[0109] S40, when AM ≥ BS 0 And AN≥XS 0 When V is obtained, the swing user identifier corresponding to V is obtained. As those skilled in the art know, the step of obtaining the swing user identifier corresponding to V is referred to steps S51-S57, and will not be repeated here.
[0110] By taking the above steps, a list of target series APP identifiers corresponding to the target APP is obtained, and the user identifier matching degree and data correlation coefficient corresponding to the target series APP identifiers are obtained. The user identifier matching degree and data correlation coefficient are compared to further obtain the swing user identifiers corresponding to the target series APP identifier list. This can avoid obtaining the swing user identifiers corresponding to the target series APP identifier list that are not accurate enough when the user identifier matching degree or data correlation coefficient is small, which is conducive to improving the accuracy of obtaining the swing user identifiers corresponding to the target series APP identifier list.
[0111] This invention provides a data processing system for obtaining swing user identifiers. The system can obtain a third preset user identifier list based on a first preset user identifier list and a second preset user identifier list. It then obtains the user identifier matching degree and the data correlation coefficient corresponding to the target app based on the third preset user identifier list. When the user identifier matching degree is not less than a preset identifier matching degree and the data correlation coefficient is not less than a preset correlation coefficient, it obtains the first preset user identifier from the first preset user identifier list corresponding to the target app. Finally, it obtains the swing user identifier corresponding to the target app based on the first preset user identifier. This eliminates the need for manual analysis of user behavior data to obtain swing user identifiers, thus improving the accuracy of obtaining swing user identifiers.
[0112] While specific embodiments of the invention have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention.
Claims
1. A data processing system for acquiring swing user identifiers, characterized in that, The "swing user" identifier is a unique identifier for a swing user, who is a user who continuously switches between two or more apps where there is partial overlap in their customer base; the data processing system includes a processor and a memory storing a computer program, which, when executed by the processor, performs the following steps: S1. Obtain the user identifier matching degree BS corresponding to the target APP and the data correlation coefficient XS corresponding to the target APP. Step S1 also includes the following sub-steps S11-S15: S11. Obtain the first preset user identifier list B corresponding to the target APP. B includes several first preset user identifiers. The first preset user identifiers are user identifiers of users who have used the target APP collected by the target SDK. S12. Obtain the second preset user identifier list A corresponding to the target APP. A includes several second preset user identifiers, which are user identifiers of users who have used the target APP stored in the database. S13. Use the intersection of B and A as the third preset user identifier list C = {C1, C2, ..., C}. i , ..., C m }, C i Let i be the i-th third preset user identifier, where i ranges from 1 to m, and m is the number of third preset user identifiers; S14. Obtain BS based on A and C, where BS satisfies the following conditions: BS=m / A 0 , where A 0 The number of second preset user identifiers in A; S15, When BS < BS 0 When XS=0, and when BS≥BS0, according to C i Obtain XS, where BS 0 Preset identifier matching degree; S2, when BS≥BS 0 And XS≥XS 0 When, obtain B = {B1, B2, ..., B} e , ..., B f }, where XS 0 To preset the correlation coefficient, B e Let e be the first preset user identifier corresponding to the target APP, where e ranges from 1 to f, and f is the number of first preset user identifiers corresponding to the target APP; S3, according to B e Obtain the swing user identifiers corresponding to the target app; including: S31, Obtain B e The corresponding first associated APP identifier list F e ={F e1 F e2 , ..., F eg , ..., F eh(e) }, where F eg For B e The corresponding g-th first associated APP identifier, where g takes values from 1 to h(e), and h(e) is B. e The number of corresponding first associated APP identifiers, where the first associated APP identifier is a unique identity identifier of the first associated APP, and the first associated APP is an APP that integrates the target SDK in the mobile terminal device used by the user corresponding to the first preset user identifier; S32. Obtain the list of historical time slices corresponding to the historical time period T = {T1, T2, ..., T...} x , ..., T p }, where T x Let x be the x-th historical time slice corresponding to the historical time period, where x ranges from 1 to p, and p is the number of historical time slices corresponding to the historical time period. S33, Obtain F eg In T x The corresponding third usage time period list G xeg ={G 1 xeg G 2 xeg , ..., G y xeg , ..., G q (xeg) xeg }, G y xeg For F eg In T x The corresponding y-th third usage time period, where y ranges from 1 to q(xeg), and q(xeg) is F. eg In T x The number of times G corresponds to the third usage time period. y xeg The starting time point is G1 y xeg G y xeg The end time point is G2 y xeg The start and end times of the third usage period are: the time when the user corresponding to the first preset user identifier, which is the first associated APP identifier collected by the target SDK, opens the APP corresponding to the first associated APP identifier and closes the APP corresponding to the first associated APP identifier in the historical time slice. S34, when G2 y xeg -G1 y xeg When ≥△t1, G y xeg The length of F eg In T x The corresponding third active duration is used to obtain F. eg In T x The corresponding third active duration list FT_xeg={FT_xeg1, FT_xeg2, ..., FT_xeg k , ..., FT_xeg t(xeg) }, FT_xeg k For F eg In T x The corresponding k-th third active duration, where k ranges from 1 to t(xeg), and t(xeg) is F. eg In T x The number of the third active durations corresponding to this; S35, when ∑ h(e) g=1 t(xeg)≥△PC and ∑ h(e) g=1 ∑ t(xeg) k=1 FT_xeg k When ≥△t2, T x As B e The corresponding active time slice to obtain B e Corresponding active time slice list HY e Where △PC is the preset usage frequency value, △t2 is the second preset active duration, and HY e It includes several active time slices; S36, when HY e When the number of active time slices in the HY is not less than the preset number of time slices, HY will... e Corresponding B e As the active user identifiers corresponding to the target APP, obtain the list of active user identifiers H={H1, H2, ..., H...} for the target APP. u H v }, where H u Let u be the identifier of the uth active user corresponding to the target APP, where u ranges from 1 to v, and v is the number of active user identifiers corresponding to the target APP. S37. Based on H, obtain the swing user identifier corresponding to the target APP.
2. The data processing system for obtaining swing user identifiers according to claim 1, characterized in that, Step S15 also includes the following steps: S151, Obtain C i The corresponding first usage time period list SY i ={SY i1 SY i2 , ..., SY ia , ..., SY ic(i) }, where SY ia C i The corresponding first usage time period of the 'a'th time period, where 'a' takes values from 1 to c(i), and c(i) is C i The corresponding number of the first usage time period, SY ia The starting time point is SY 1 ia SY ia The end time point is SY 2 ia The start and end times of the first usage period are: the time when the user corresponding to the third preset user identifier stored in the database opens the target APP and the time when the user closes the target APP within the preset time period; S152, when SY 2 ia -SY 1 ia When ≥△t1, SY ia The length of C i The corresponding first active duration to obtain C i The corresponding first active duration list D i ={D i1 D i2 , ..., D ij , ..., D in(i) }, D ij C i The corresponding first active duration is j, where j ranges from 1 to n(i), and n(i) is C. i The corresponding number of first active durations, where △t1 is the first preset active duration; S153, according to D i , get C i The first adjustment parameter C of the corresponding data correlation coefficient 0 i ; S154, Obtain C i The corresponding second usage time period list DE i ={DE i1 DE i2 , ..., DE i(ag) , ..., DE i(ah(i)) }, where DE i(ag) C i The corresponding ag-th second usage time period, where ag takes values from 1 to ah(i), and ah(i) is C. i The corresponding number of second usage time periods, DE i(ag) The starting time point is DE 1 i(ag) DE i(ag) The end time point is DE 2 i(ag) The start and end times of the second usage period are: the time when the user corresponding to the third preset user identifier collected by the target SDK opens the target APP and the time when the user closes the target APP within the preset time period; S155, according to DE 2 i(ag) and DE 1 i(ag) Get C i The second adjustment parameter E of the corresponding data correlation coefficient 0 i ; S156, According to C 0 i and E 0 i Obtain XS, where XS meets the following conditions: XS=∑ m i=1 ((C 0 i -∑ m i=1 C 0 i / m)×(E 0 i -∑ m i=1 E 0 i / m)) / ((∑ m i=1 (C 0 i -∑ m i=1 C 0 i / m) 2 ) 1 / 2 ×(∑ m i=1 (E 0 i -∑ m i=1 E 0 i / m) 2 ) 1 / 2 )。 3. The data processing system for obtaining swing user identifiers according to claim 2, characterized in that, In step S153, C 0 i The following conditions must be met: C 0 i =(C 1 i +C 2 i ) / 2, where C 1 i C i The corresponding adjustment for C 0 i The first parameter, C 1 i The following conditions must be met: C 1 i =(n(i)-min(n(1),n(2),……,n(i),……,n(n))) / (max(n(1),n(2),……,n(i),……,n(n))-min(n(1),n(2),……,n(i),……,n(n))), where main() is the function to get the minimum value and max() is the function to get the maximum value; C 2 i C i The corresponding adjustment for C 0 i The second parameter, C 2 i The following conditions must be met: C 2 i =(∑ n(i) j=1 D ij -min(∑ n(1) j=1 D 1j ,∑ n(2) j=1 D 2j ,……,∑ n(i) j=1 D ij ,……,∑ n(n) j=1 D nj )) / (max(∑ n(1) j=1 D 1j ,∑ n(2) j=1 D 2j ,……,∑ n(i) j=1 D ij ,……,∑ n(n) j=1 D nj )-min(∑ n(1) j=1 D 1j ,∑ n(2) j= 1D 2j ,……,∑ n(i) j=1 D ij ,……,∑ n(n) j=1 D nj ))。 4. The data processing system for obtaining swing user identifiers according to claim 1, characterized in that, Step S37 also includes the following steps: S371, Obtain H from the xth historical time slice. u List of the first historical usage time period corresponding to the target APP J ux ={J ux1 J ux2 , ..., J uxγ , ..., J uxθ(ux) }, where J uxγ For the xth historical time slice H u J is the γth first historical usage time period corresponding to the target APP. uxγ The starting time point is J1 uxγ J uxγ The end time point is J2 uxγ The start and end times of the first historical usage period are: the time when the user corresponding to the active user identifier collected by the target SDK opens the target APP and the time when the target APP closes in the historical time slice; S372, according to J ux Get H from the xth historical time slice u K, the first historical usage duration list corresponding to the target app ux ={K ux1 K ux2 , ..., K ux(ai) , ..., K ux(am(x)) }, where K ux(ai) For the xth historical time slice, H u The ai-th first historical usage duration corresponding to the target APP, where ai ranges from 1 to am(x), and am(x) is the H in the x-th historical time slice. u The number of first historical usage times corresponding to the target app; S373, according to K ux Get H in the historical time period u The total historical usage frequency value L corresponding to the target APP u and H in historical time period u Total historical usage time M corresponding to the target APP u , where L u The following conditions must be met: L u =∑ p x=1 am(x); M u The following conditions must be met: M u =∑ p x=1 ∑ am(x) ai=1 K ux(ai) ; S374. Obtain the preset APP identifier combination list R={R1, R2, ..., R...} z , ..., R w }, where R z For the z-th preset APP identifier combination, z takes values from 1 to w, w is the number of preset APP identifier combinations, and R z By R z1 and R z2 Composition, R z1 R is the first preset APP identifier in the z-th preset APP identifier combination. z2 For R z1 The corresponding second preset APP identifier; S375, Traverse R, when R... z1 When =MB, R z1 The corresponding R z2 As the second associated APP identifier corresponding to MB, when R z2 When =MB, R z2 The corresponding R z1 As the identifier of the second associated APP corresponding to MB, we can obtain the list of the second associated APPs corresponding to MB, U={U1, U2, ..., U...} α , ..., U β }, where MB is the target app identifier, which is the unique identifier of the target app, and U α Let α be the αth second associated APP identifier corresponding to MB, where α ranges from 1 to β, and β is the number of second associated APP identifiers corresponding to MB; S376, according to H u and U α Get H in the historical time period u and U α The corresponding historical total usage frequency value P uα and H in historical time period u and U α The corresponding total historical usage time Q uα ; S377, when L u M u P uα Q uα Simultaneously satisfying L u ≥count、P uα ≥count、M u ≥△t4、Q uα ≥△t4、BL1≤L u / P uα ≤BL2、BL1≤M u / Q uα When ≤BL2, determine H u This refers to the identifier of the swing user corresponding to the target app.
5. The data processing system for obtaining swing user identifiers according to claim 4, characterized in that, Step S372 includes the following sub-steps: S3721. When ai=1, proceed to step S3722; when ai≠1, set J... 1 ux((ai)-1) As J ux And proceed to step S3722, J 1 ux((ai)-1) For K ux((ai)-1) The corresponding updated first historical usage time period list, K ux((ai)-1) For the xth historical time slice, H u The (ai-1)th first historical usage duration corresponding to the target app; S3722, when J2 uxγ -J1 ux1 ≤△t3 and J1 ux(γ+1) -J1 ux1 When Δt3 > 0, determine K. ux(ai) =J 0 ux1 +J 0 ux2 +……J 0 uxη +……+J 0 uxγ and J ux1 J ux2 , ..., J uxη , ..., J uxγ From J ux Delete to get K ux(ai) The corresponding updated first historical usage time period list J 1 ux(ai) Among them, J1 ux1 For J ux1 The starting time point, △t3 is the third preset active duration, J 0 uxη For J uxη Duration, J uxη For the xth historical time slice H u The ηth first historical usage time period corresponding to the target APP, where the value of η ranges from 1 to γ.
6. The data processing system for obtaining swing user identifiers according to claim 1, characterized in that, BS 0 The value range is [0, 1].
7. The data processing system for obtaining swing user identifiers according to claim 1, characterized in that, XS 0 The value range is [0, 1].
8. The data processing system for obtaining swing user identifiers according to claim 4, characterized in that, The value range of BL1 is [0, 1].
9. The data processing system for obtaining swing user identifiers according to claim 4, characterized in that, The value range of BL2 is [1, 2].
Citation Information
Patent Citations
Method for acquiring target interval duration, electronic equipment and storage medium
CN118012726A
Method, computer device and readable medium for user's intent mining
US20200034431A1