Entity ID fusion method, electronic equipment and storage medium
By receiving and classifying user ID data information sets, the preset entity ID category mapping table is used to classify and integrate entity IDs, solving the problem of uncommon user IDs and realizing the integration and query of user data resources.
Patent Information
- Application Number
- CN202510342703.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
It is difficult for the existing technology to effectively integrate and query discrete IDs obtained by different channels and applications, resulting in disagreement of user information and affecting the integration and query of data resources.
By receiving the user ID data information set uploaded by the third-party platform, the initial entity ID is extracted, and according to the preset entity ID category mapping table, the initial entity ID is classified into the first, second and third entity IDs, and match and fusion are performed separately, and the user ID fusion is finally realized.
It realizes the accurate integration of user IDs and forms a complete ID integration system to facilitate the integration and query of user data resources.
Smart Images

Figure CN120179704A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method for fusing entity IDs, an electronic device, and a storage medium. Background Art
[0002] At present, with the increasing number and types of user data collection channels, the number of ID identifiers related to people is also increasing. For example, in social life, offline data collection of users, including express information, household registration, vehicle registration, etc., can obtain multiple IDs related to users. Through the registration information of each app, several IDs related to users can also be obtained. Moreover, through the usage data of the app, several usage behaviors of users can be obtained, including shopping records of shopping software, sending records of emails, device trajectories obtained by base stations, etc., all of which can reflect the dynamic behaviors and trajectory information of users, making the perception data of users in society more and more, which is of great significance for characterizing user behaviors. However, the IDs obtained from different channels and applications are all discrete, resulting in non-common user information, which is not conducive to user data integration and query. Summary of the Invention
[0003] In view of the above technical problems, the present invention provides a method for fusing entity IDs, an electronic device, and a storage medium, which can accurately unify several categories of IDs corresponding to users, facilitating user data resource integration and query.
[0004] According to a first aspect of the present invention, there is provided a method for fusing entity IDs, including the following steps:
[0005] S100, receiving a set of user ID data information uploaded by several third-party platforms, and extracting an initial entity ID set A = {A1, A2,..., A i ,..., A m} from the set of user ID data information, where A i is the initial entity ID list of the i-th initial user, i = 1, 2,..., m, and m is the number of initial entity ID lists.
[0006] S200, according to a preset entity ID category mapping table, mapping j initial entity IDs in A i to any one of a preset first type of entity ID, a second type of entity ID, and a third type of entity ID, to obtain the entity ID classification result corresponding to A i .
[0007] S300. From A, filter out the initial entity ID list with the first type of entity ID, the initial entity ID list without the first type of entity ID but with the second type of entity ID, and the initial entity ID list with only the third type of entity ID, and respectively determine them as the first entity ID list, the second entity ID list, and the third entity ID list in sequence.
[0008] S400. Match the first type of initial entity IDs in several first entity ID lists, add all the initial entity IDs in the first entity ID list corresponding to the matched first type of initial entity IDs to the unique target ID group of the corresponding target user, and add the first entity ID lists corresponding to the unmatched first type of initial entity IDs to their respective corresponding unique target ID groups of the target user.
[0009] S500. Match the second type of entity IDs in each second entity ID list with the second type of entity IDs in several unique target ID groups, add all the initial entity IDs in the second entity ID list corresponding to the matched second type of entity IDs to the corresponding unique target ID group, and use the second entity ID list corresponding to the unmatched second type of entity IDs as the pending entity ID list.
[0010] S600. Filter out the target entity ID list with the third type of entity ID from the pending entity ID list, calculate the similarity between the third type of entity IDs in the third entity ID list and the target entity ID list and the third type of entity IDs in several unique target ID groups, and add the third entity ID list and the target entity ID list to the corresponding unique target ID groups respectively according to the similarity calculation results to achieve the final ID fusion.
[0011] According to the second aspect of the present invention, there is provided a non-transitory computer-readable storage medium, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the above entity ID fusion method.
[0012] According to the third aspect of the present invention, there is provided an electronic device, including a processor and the above non-transitory computer-readable storage medium.
[0013] The present invention has at least the following beneficial effects:
[0014] The present invention provides a method for entity ID fusion. First, a number of initial entity ID lists are extracted from the received user ID data information set. According to a preset entity ID category mapping table, each initial entity ID in the initial entity ID list is mapped to any one of a preset first type of entity ID, a second type of entity ID, and a third type of entity ID. According to the classification results, a first entity ID list, a second entity ID list, and a third entity ID list are determined, and the three entity ID lists are processed in different ways. The selected entity ID lists are added to the unique target ID group of the corresponding target user to achieve the final user ID fusion. Through the above method, several categories of IDs corresponding to the user can be accurately unified to form a complete ID fusion system, which is convenient for user data resource integration and query. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0016] Figure 1 It is a flowchart of the entity ID fusion method provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0018] The embodiment of the present invention provides a method for entity ID fusion, as Figure 1 shown, the method includes the following steps:
[0019] S100, receive a user ID data information set uploaded by a number of third-party platforms, and extract an initial entity ID set A = {A1, A2,..., A i ,..., A m}, where A iis the list of initial entity IDs for the i-th initial user, where i = 1, 2, ……, m, and m is the number of lists of initial entity IDs; for example, the initial entity ID can be an ID card number, a mobile phone number, a license plate number, an address, a QQ number, etc. Among them, the user ID data information set can be several IDs obtained according to the user registration information of the app, such as a name, an app account string, etc., or can be several IDs collected offline, such as the license plate number collected during vehicle registration, the address collected during household registration, etc. It should be noted that different lists of initial entity IDs refer to lists of IDs collected at different times, that is, the initial users corresponding to different lists of initial entity IDs may be the same user.
[0020] S200, according to the preset entity ID category mapping table, map the j initial entity IDs in A i to any one of the preset first-class entity IDs, second-class entity IDs, and third-class entity IDs respectively, to obtain the entity ID classification result corresponding to A i
[0021] Furthermore, the entity ID category mapping table refers to a table including several preset entity IDs and the mapping relationships between each preset entity ID and entity IDs of different preset categories.
[0022] Specifically, the first-class entity ID is the entity ID representing the unique identity of the user; for example, the ID card number, and there is a one-to-one correspondence between the user and the ID card number.
[0023] Specifically, the second-class entity ID is the entity ID that is not the unique identity of the user but can represent a unique user; for example, the license plate number, the mobile phone number, etc. A user can have multiple license plate numbers and multiple mobile phone numbers, but a license plate number and a mobile phone number can only correspond to one user.
[0024] Specifically, the third-class entity ID is other entity IDs that are not the first-class entity ID and not the second-class entity ID; for example, the address, the name, etc. A user can have multiple addresses, and one address may also be inhabited by multiple users.
[0025] As described above, by classifying different initial entity IDs, different categories of entity IDs can be screened out. Since different categories of entity IDs have different degrees of reliability when determining users, ID stratification is achieved through ID classification, which is beneficial to the processing of user ID fusion and the improvement of accuracy.
[0026] S300. From A, filter out the initial entity ID list with the first type of entity ID, the initial entity ID list without the first type of entity ID but with the second type of entity ID, and the initial entity ID list with only the third type of entity ID, and respectively determine them as the first entity ID list, the second entity ID list, and the third entity ID list in sequence.
[0027] S400. Match the first type of initial entity IDs in several first entity ID lists. Add all the initial entity IDs in the first entity ID list corresponding to the matched first type of initial entity IDs to the unique target ID group of the corresponding target user, and add the first entity ID lists corresponding to the unmatched first type of initial entity IDs to their respective corresponding unique target ID groups of the target users. It can be understood that for the first entity ID lists corresponding to the unmatched first type of initial entity IDs, fuse all the initial entity IDs within the first entity ID list into the same ID group to achieve the preliminary fusion of entity IDs.
[0028] As described above, by filtering out different categories of entity ID lists, it is possible to process different categories of entity ID lists specifically. And for the first type of entity ID list, the first type of entity ID lists including the same first type of initial entity ID can be regarded as several ID lists of the same user, and then fuse the initial entity IDs corresponding to the same user, which is conducive to achieving the accurate fusion of IDs.
[0029] S500. Match the second type of entity IDs in each second entity ID list with the second type of entity IDs in several unique target ID groups. Add all the initial entity IDs in the second entity ID list corresponding to the matched second type of entity IDs to the corresponding unique target ID group, and regard the second entity ID list corresponding to the unmatched second type of entity IDs as the pending entity ID list.
[0030] As described above, since the second type of entity ID can represent the corresponding target user, when the second type of entity ID is matched, it indicates that the second entity ID list and the user corresponding to the matched unique target ID group are the same user. Therefore, combine the entity IDs corresponding to the same user to achieve the further fusion of entity IDs.
[0031] S600. Filter out the target entity ID list with the third type of entity ID from the pending entity ID list. Calculate the similarity between the third type of entity IDs in the third entity ID list and the target entity ID list and the third type of entity IDs in several unique target ID groups, and add the third entity ID list and the target entity ID list to the corresponding unique target ID groups respectively according to the similarity calculation results to achieve the final ID fusion.
[0032] In a specific embodiment, step S600 further includes the following steps:
[0033] S601. For any third entity ID list, several third - type entity IDs in the third entity ID list are merged to generate a first entity ID vector. It should be noted that the processing method of the target entity ID list is the same as that of the third entity ID list. In this embodiment, the third entity ID list is taken as an example for illustration.
[0034] S602. Third - type entity IDs are screened out from any one of the unique target ID groups and merged to generate a second entity ID vector.
[0035] S603. The first entity ID vector and the second entity ID vector are subjected to data alignment processing. Based on the preset ID weights corresponding to each third - type entity ID in the first entity ID vector and the second entity ID vector, the similarity between the aligned first entity ID vector and the second entity ID vector is calculated to obtain the similarity between the first entity ID vector and the second entity ID vector corresponding to each unique target ID group. It can be understood that data alignment processing means keeping the dimensions of the first entity ID vector and the second entity ID vector consistent. For example, for a vector with fewer dimensions, it is achieved by supplementing 0. In addition, it should be noted that when the third - type entity ID is an address, the address needs to be subjected to a consistency check first. When the check shows the same target address, the address dimension in the vector is represented as the same value, which is conducive to accurate calculation of the similarity.
[0036] Specifically, the preset ID weight refers to the weight preset for each third - type entity ID. For example, when calculating the similarity between two vectors, the similarity of each dimension in the vector is first calculated, and then a weighted operation is performed to obtain the final similarity.
[0037] Furthermore, the similarity can be obtained by calculating the Euclidean distance between the first entity ID vector and the second entity ID vector. Those skilled in the art know the calculation method of the Euclidean distance, which will not be elaborated here.
[0038] S604. When the maximum similarity is greater than the preset similarity threshold, the unique target ID group corresponding to the second entity ID vector with the maximum similarity is determined as the final target ID group corresponding to the first entity ID vector, and the third entity ID list corresponding to the first entity ID vector is added to the final target ID group to achieve the final ID fusion.
[0039] As described above, since there are many third - type entity IDs and none of them can accurately determine the corresponding user, it is necessary to calculate the similarity to screen out the closest user. When the similarity degree is greater than the similarity threshold, they can be considered as the same user, and the corresponding several third - type entity IDs are added to the corresponding unique target ID group to achieve reliable ID fusion.
[0040] Further, the method further includes the following steps:
[0041] S10, screen out the initial entity ID list that has not been added to the corresponding unique target ID group from A and use it as the list of entities to be processed.
[0042] S20, add several initial entity IDs in each list of entities to be processed to the preset ID group corresponding to the list of entities to be processed itself, and add a supplement - required flag to each preset ID group.
[0043] As described above, the initial entity ID list that has not been added to the corresponding unique target ID group is the list of entity IDs for which the corresponding target user has not been found. Some of the entity IDs in it have only achieved preliminary fusion, but the corresponding target user has not been determined. Therefore, a supplement - required flag needs to be added to facilitate subsequent ID supplementation.
[0044] In another specific embodiment, the method further includes the following steps:
[0045] P100, based on the first - type entity IDs corresponding to each initial abnormal object obtained in advance, and according to the unique target ID group corresponding to each first - type entity ID, screen out several target app usage information corresponding to each initial abnormal object from several unique target ID groups.
[0046] In another embodiment, the initial abnormal object is determined through the following steps before step P100:
[0047] P001, based on the historical video recordings of the target geographical area in the obtained historical time period, obtain several target images after frame - by - frame processing, input the several target images into the image feature extraction model, and obtain the set of behavior feature vectors corresponding to each initial acquisition object in the historical video recordings; the set of behavior feature vectors includes at least one behavior feature vector; it can be understood that: the behavior features include but are not limited to any one of appearance features and several action features, for example, behaviors such as looking up, staying duration, looking down, walking, etc.
[0048] Specifically, the target images are obtained through the following steps:
[0049] P0011, divide the historical video recordings into several video segments according to a preset duration. For example, the preset duration is 1s.
[0050] P0012, Use an image frame extraction tool to perform frame division on each video clip, and use the frame image at the middle moment corresponding to each video clip as the target image.
[0051] P002, Align the set of behavior feature vectors corresponding to each initial acquisition object, and perform k-means clustering on the processed sets of behavior feature vectors to obtain a preset number of clusters of sets of behavior feature vectors; it can be understood that aligning the set of behavior feature vectors means keeping the dimensions of the set of behavior feature vectors consistent and setting the missing behavior feature dimensions to 0.
[0052] Furthermore, obtain the preset number of clusters through the following steps:
[0053] P0021, Obtain the number of initial acquisition objects corresponding to each behavior feature vector.
[0054] P0022, For any behavior feature vector, when the number of initial acquisition objects corresponding to the behavior feature vector is greater than the preset object number threshold, use the behavior feature vector as the target feature vector; those skilled in the art set the preset object number threshold according to actual needs, for example, 70%, which will not be elaborated here.
[0055] P0023, Calculate the preset number of clusters D based on the number of target feature vectors and the number of all behavior feature vectors corresponding to several initial acquisition objects; the preset number of clusters D meets the following conditions:
[0056] D = D1 - D2 + 1, where D1 is the number of all behavior feature vectors corresponding to several initial acquisition objects, and D2 is the number of target feature vectors.
[0057] As described above, when the number of people with the same behavior feature vector is large, it can be considered an important feature vector. Each important feature vector corresponds to a large number of people, and it is inclined to consider that several important feature vectors are the feature vectors of the same type of object. Therefore, it is hoped to cluster these important features into one cluster, thereby reducing the number planning of clusters and making the clustering result more reasonable.
[0058] P003, Obtain the common features corresponding to each cluster of sets of behavior feature vectors, and determine the key unique features corresponding to each cluster of sets of behavior feature vectors according to the common features corresponding to each cluster of sets of behavior feature vectors; the common features refer to the features whose corresponding set of behavior feature vectors accounts for a proportion greater than the preset proportion threshold in all sets of behavior feature vectors in the cluster of sets of behavior feature vectors; those skilled in the art set the preset proportion threshold according to actual needs, for example, 60%.
[0059] Specifically, step P003 further includes the following steps:
[0060] P0031. For any cluster of behavioral feature vectors, use any common feature corresponding to the cluster of behavioral feature vectors as the first behavioral feature, and use all common features corresponding to several other clusters of behavioral feature vectors as the second behavioral features; the several other clusters of behavioral feature vectors refer to all clusters of behavioral feature vectors except the any cluster of behavioral feature vectors.
[0061] P0032. Traverse all the second behavioral features. When there is no second behavioral feature that matches the first behavioral feature, determine the first behavioral feature as the key exclusive feature corresponding to the cluster of feature vectors; it can be understood that the key exclusive feature corresponding to the cluster of behavioral feature vectors refers to the one that only exists in several common features corresponding to the cluster of behavioral feature vectors, but does not exist in several common features corresponding to several other clusters of behavioral feature vectors.
[0062] As described above, by obtaining the common features corresponding to the cluster of behavioral feature vectors, several behavioral features that can characterize the cluster of behavioral feature vectors can be obtained. Then, through the comparison of the common features corresponding to each cluster of behavioral feature vectors, the exclusive behavioral features that can individually characterize the cluster of behavioral feature vectors can be found, so as to provide more reasonable samples for the subsequent training of the preset classification model.
[0063] P004. Obtain the key object from several initial acquisition objects in the historical video; the key object refers to the initial acquisition object corresponding to any one of the key exclusive features.
[0064] P005. After performing face recognition on each key object, obtain the unique identity identifier corresponding to each key object, and screen out the preset abnormal objects from several key objects as the initial abnormal objects according to the preset information comparison table; the information comparison table includes several preset abnormal objects and the unique identity identifier corresponding to each preset abnormal object.
[0065] As described above, the obtained key objects are objects containing exclusive features. Since the abnormal objects to be searched will have relatively similar behavioral feature manifestations and are different from the behavioral features of other normal objects, but it is not known specifically what the exclusive behavioral features are, so the features are classified into different categories and the required feature categories are found through comparison with face recognition, and these are used as the seed population, which improves the reliability of sample selection and makes the model trained subsequently have a better prediction effect.
[0066] P200. For any initial abnormal object, obtain the behavioral data features in the corresponding several target app usage information of the initial abnormal object, and screen out several trajectory information reported by each target app within a preset time period from the several behavioral data features; the trajectory information includes trajectory points and the reporting time of each trajectory point. For example, the usage data features of a shopping software include browsing records, shopping records, etc., and the map software has search locations, etc., and each app corresponds to its own reported trajectory data.
[0067] Furthermore, the preset time period is the historical time period corresponding to the historical video, and considering the error in the reported location information of the app, the two ends of the historical time period should be extended to obtain the preset time period to obtain the reported location information within a longer time period.
[0068] P300. Integrate the several trajectory information corresponding to each target app to obtain the final trajectory route of the initial abnormal object.
[0069] Specifically, in step P300, the final trajectory route of the initial abnormal object is obtained through the following steps:
[0070] P301. According to the trajectory points uploaded by each target app within the preset time period and the reporting time of each trajectory point, draw the initial trajectory route corresponding to each target app respectively; it can be understood that: a straight line is connected between two trajectory points to obtain the route between the two trajectory points.
[0071] P302. For any initial trajectory route, obtain the position information corresponding to each moment within the preset time period of the initial trajectory route; it can be understood that: the position information is represented by the longitude and latitude mapped by the initial trajectory route.
[0072] P303. According to the position information corresponding to each moment within the preset time period of the initial trajectory route, calculate the first target position point corresponding to several initial trajectory routes at each moment; it can be understood that at any moment, the center point of the corresponding several position information is taken as the first target position point, or the average longitude and average latitude are calculated to obtain the first target position point.
[0073] P304. Calculate the distance difference between the position information corresponding to each moment within the preset time period of the initial trajectory route and the first target position point at each moment, and screen out the position information with the corresponding distance difference greater than the preset distance threshold to obtain several remaining position information corresponding to each moment.
[0074] P305. For any moment, calculate the second target position point corresponding to any moment according to the several remaining position information corresponding to any moment, and form the final trajectory route through several second target position points.
[0075] Furthermore, the method for acquiring the second target location point is the same as the method for acquiring the first target location point, which will not be described in detail herein.
[0076] As mentioned above, each app will periodically report the location information of the device, and different apps will have inconsistent reporting intervals and upload time points due to inconsistent setting rules or inconsistent user login time periods. In addition, the reported location information of each app has certain errors. Through the above process, the trajectory points with large errors are screened out, and the remaining trajectory points are integrated, which can reduce the error of the final trajectory route, thereby improving the accuracy of the final trajectory route.
[0077] P400, when the error between the first time interval in which the initial abnormal object itself is located in the target geographical area and the second time interval in which the initial abnormal object itself appears in the historical video recording corresponding to the target geographical area is less than the time difference threshold, the initial abnormal object is determined as a key abnormal object; technical personnel in this field set the time difference threshold according to actual needs, wherein, since each time point of the app's trajectory point reporting has a certain interval, when the object stays for a short time, the first time interval may also be a time point. In a specific implementation, when the distance between any trajectory point corresponding to the initial abnormal object and the location point of the target geographical area is less than the preset target distance, it can be considered that the initial abnormal object is in the target geographical area at the reporting time point corresponding to any trajectory point.
[0078] Further, it is determined that the error between the first time interval in which the initial abnormal object itself is located in the target geographical area and the second time interval in which the initial abnormal object itself appears in the historical video recording corresponding to the target geographical area is less than the time difference threshold through the following steps:
[0079] P401, when the first time interval and the second time interval overlap, it is determined that the error between the first time interval in which the initial abnormal object itself is located in the target geographic area and the second time interval in which the initial abnormal object itself appears in the historical video recording is less than a time difference threshold.
[0080] P402, when the first time interval and the second time interval do not overlap, obtain the time interval between the first time interval and the second time interval, and when the time interval is less than a preset interval threshold, determine that the error between the first time interval when the initial abnormal object itself is located in the target geographic area and the second time interval when the initial abnormal object itself appears in the historical recorded video is less than the time difference threshold.
[0081] As described above, considering that face recognition may have errors due to the influence of video clarity or other reasons resulting in inaccurate initial abnormal objects obtained, it is necessary to further confirm the obtained abnormal objects. Since there are certain errors in the determined final trajectory route, by searching for the final trajectory route of the initial abnormal object to see if there is overlap or proximity with the time interval when it appears in the target geographical area, the final abnormal objects are screened out. By combining the two, the screening accuracy of abnormal objects is improved, and at the same time, the reliability of the subsequent samples used is improved.
[0082] Further, after step P400, the following steps are also included:
[0083] P500, combine a number of determined key abnormal objects into a seed population.
[0084] P600, use the behavior feature vectors corresponding to the initial collection objects in the seed population as positive samples, and use the behavior feature vectors corresponding to the initial collection objects other than the seed population as negative samples to train a preset classification model to obtain a trained target classification model; it can be understood that: the preset classification model is a binary classification model for distinguishing abnormal objects and non-abnormal objects.
[0085] Specifically, the preset classification model is trained through the following steps:
[0086] P601, for any behavior feature vector in the positive samples, determine the ratio of the number of initial collection objects corresponding to the behavior feature vector in the positive samples to the total number of all initial collection objects corresponding to the positive samples as the first initial weight of the behavior feature vector.
[0087] P602, for any behavior feature vector in the negative samples, determine the ratio of the number of initial collection objects corresponding to the behavior feature vector in the negative samples to the total number of all initial collection objects corresponding to the negative samples as the second initial weight of the behavior feature vector.
[0088] P603, obtain the reference weight coefficient θ corresponding to each behavior feature vector; the reference weight coefficient θ corresponding to the behavior feature vector refers to the ratio of the number of initial collection objects corresponding to the behavior feature vector to the total number of all initial collection objects.
[0089] P604, for any behavior feature vector, when the behavior feature vector exists in both the positive samples and the negative samples, take the smaller value of θ and 1 - θ as the target weight coefficient corresponding to the behavior feature vector; otherwise, take the larger value of θ and 1 - θ as the target weight coefficient corresponding to the behavior feature vector.
[0090] P605. Adjust each first initial weight and each second initial weight according to the corresponding target weight coefficients, and perform normalization processing on the adjusted first initial weights and second initial weights respectively to obtain the final weights corresponding to each behavior feature vector.
[0091] P606. Input the positive samples, negative samples, and the final weights corresponding to each behavior feature vector into a preset classification model for training to obtain a trained target classification model.
[0092] As described above, since each initial acquisition object corresponds to several behavior feature vectors, considering that there may be overlapping behavior feature vectors between the positive samples and the negative samples, which will affect the training results, the weights of the behavior feature vectors shared by both the positive and negative samples are reduced, and conversely, the weights are increased accordingly. After adjusting the weights of each feature behavior vector, training the preset classification model can improve the prediction accuracy of the target classification model.
[0093] Furthermore, the method also obtains the target classification model through the following steps:
[0094] P607. Obtain a new historical video recording and input the new historical video recording into the target classification model to obtain several abnormal objects identified by the target classification model.
[0095] P608. When the accuracy rate corresponding to several abnormal objects is greater than the accuracy rate threshold, take the target classification model as the final target classification model; otherwise, obtain new positive samples and new negative samples according to the new historical video recording, and perform iterative training on the target classification model until the accuracy rate of the identified abnormal objects is greater than the accuracy rate threshold or reaches the preset number of iterative rounds to obtain a new target classification model. Those skilled in the art set the accuracy rate threshold according to actual needs, which will not be elaborated here.
[0096] P700. When receiving a video recording to be processed, input the video recording to be processed into the trained target classification model to identify several target abnormal objects.
[0097] As described above, by obtaining a new historical video recording to verify the accuracy rate of the target classification model, when the accuracy rate does not meet the requirements, update the samples and perform further iterative training, so that the prediction accuracy of the finally trained target classification model is higher, and thus when obtaining a new video recording to be processed, more accurate target abnormal objects can be identified.
[0098] An embodiment of the present invention further provides a non-transitory computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to a method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the entity ID fusion method provided in the foregoing embodiment.
[0099] An embodiment of the present invention further provides an electronic device, including a processor and the foregoing non-transitory computer-readable storage medium.
[0100] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. An entity ID fusion method, characterized in that: The method comprises the following steps: S100, receiving user ID data information sets uploaded by several third-party platforms, and extracting an initial entity ID set A={A1, A2, . . . , A i , ..., A m }, where A i is the initial entity ID list of the i-th initial user, i = 1, 2, ..., m, m is the number of initial entity ID lists; S200, according to the preset entity ID category mapping table, A i The j initial entity IDs in are mapped to any of the preset first-category entity IDs, second-category entity IDs, and third-category entity IDs, and A is obtained. i The corresponding entity ID classification result; S300, screening out from A an initial entity ID list containing first-category entity IDs, an initial entity ID list containing second-category entity IDs but not first-category entity IDs, and an initial entity ID list containing only third-category entity IDs, and determining them as a first entity ID list, a second entity ID list, and a third entity ID list, respectively; S400, matching the first-category initial entity IDs in the plurality of first-category initial entity ID lists, adding all the initial entity IDs in the first-category initial entity ID lists corresponding to the matched first-category initial entity IDs to the unique target ID groups of the corresponding target users, and adding the first-category initial entity ID lists corresponding to the unmatched first-category initial entity IDs to the unique target ID groups of the corresponding target users respectively; S500, matching the second-category entity ID in each second-category entity ID list with the second-category entity IDs in a plurality of unique target ID groups, adding all the initial entity IDs in the second-category entity ID list corresponding to the matched second-category entity IDs to the corresponding unique target ID group, and treating the second-category entity ID list corresponding to the unmatched second-category entity IDs as a pending entity ID list; S600, filter out the target entity ID list with the third type entity ID from the pending entity ID list, calculate the similarity between the third type entity ID in the third entity ID list and the target entity ID list and the third type entity ID in several unique target ID groups, and add the third entity ID list and the target entity ID list to the corresponding unique target ID groups respectively according to the similarity calculation results to achieve the final ID fusion.
2. The entity ID fusion method according to claim 1, characterized in that: The first type of entity ID is an entity ID that represents a user's unique identity; The second type of entity ID refers to an entity ID that is not a user's unique identity but can represent a unique user; The third type of entity ID is other entity IDs that are not the first type of entity ID and not the second type of entity ID.
3. The entity ID fusion method according to claim 1, characterized in that: Step S600 also includes the following steps: S601, for any third entity ID list, merging several third-category entity IDs in the third entity ID list to generate a first entity ID vector; S602, filtering out a third type of entity ID from any of the unique target ID groups and merging them to generate a second entity ID vector; S603, performing data alignment processing on the first entity ID vector and the second entity ID vector, and calculating the similarity between the first entity ID vector and the second entity ID vector after the alignment processing based on the preset ID weight corresponding to each third-category entity ID in the first entity ID vector and the second entity ID vector, so as to obtain the similarity between the first entity ID vector and the second entity ID vector corresponding to each unique target ID group; S604, when the maximum similarity is greater than the preset similarity threshold, the unique target ID group corresponding to the second entity ID vector corresponding to the maximum similarity is determined as the final target ID group corresponding to the first entity ID vector, and the third entity ID list corresponding to the first entity ID vector is added to the final target ID group to achieve the final ID fusion.
4. The entity ID fusion method according to claim 1, characterized in that: The method further comprises the steps of: S10, filtering out the initial entity ID list that has not been added to the corresponding unique target ID group from A and using them as the entity ID list to be processed; S20, adding a number of initial entity IDs in each to-be-processed entity ID list to the preset ID group corresponding to the to-be-processed entity ID list itself, and adding a to-be-supplemented identifier to each preset ID group.
5. The entity ID fusion method according to claim 1, characterized in that: The method further comprises the steps of: P100, based on the first-category entity ID corresponding to each initial abnormal object obtained in advance, and according to the unique target ID group corresponding to each first-category entity ID, filtering out several target app usage information corresponding to each initial abnormal object from several unique target ID groups; P200, for any initial abnormal object, obtain the behavior data features in the usage information of several target apps corresponding to the initial abnormal object, and filter out several track information reported by each target app within a preset time period from the several behavior data features; the track information includes track points and the reporting time of each track point; P300, integrates several trajectory information corresponding to each target app to obtain the final trajectory route of the initial abnormal object; P400, when the final trajectory route of the initial abnormal object represents that the error between the first time interval in which the initial abnormal object itself is located in the target geographical area and the second time interval in which the initial abnormal object itself appears in the historical video recording corresponding to the target geographical area is less than the time difference threshold, the initial abnormal object is determined as a key abnormal object.
6. The entity ID fusion method according to claim 5, characterized in that: In step P300, the final trajectory of the initial abnormal object is obtained through the following steps: P301, drawing the initial trajectory route corresponding to each target app according to the trajectory points uploaded by each target app within a preset time period and the reporting time of each trajectory point; P302, for any initial trajectory route, obtain the corresponding position information of the initial trajectory route at each moment within a preset time period; P303, calculating the first target position points corresponding to a number of initial trajectory routes at each moment according to the position information corresponding to the initial trajectory route at each moment within a preset time period; P304, calculating the distance difference between the position information corresponding to each moment of the initial trajectory route and the first target position point at each moment within the preset time period, and filtering out the position information whose corresponding distance difference is greater than the preset distance threshold, to obtain a number of remaining position information corresponding to each moment; P305, for any moment, the corresponding second target position point at any moment is calculated according to the corresponding plurality of remaining position information at any moment, and the final trajectory route is formed by the plurality of second target position points.
7. The entity ID fusion method according to claim 5, characterized in that: Determine that the error between the first time interval in which the initial abnormal object itself is located in the target geographical area and the second time interval in which the initial abnormal object itself appears in the historical video corresponding to the target geographical area is less than the time difference threshold by the following steps: P401, when the first time interval and the second time interval overlap, determining that the error between the first time interval in which the initial abnormal object itself is located in the target geographic area and the second time interval in which the initial abnormal object itself appears in the historical video recording is less than a time difference threshold; P402, when the first time interval and the second time interval do not overlap, obtain the time interval between the first time interval and the second time interval, and when the time interval is less than a preset interval threshold, determine that the error between the first time interval when the initial abnormal object itself is located in the target geographic area and the second time interval when the initial abnormal object itself appears in the historical recorded video is less than the time difference threshold.
8. A non-transitory computer-readable storage medium, wherein at least one instruction or at least one program is stored in the storage medium, characterized in that: The at least one instruction or the at least one program is loaded and executed by the processor to implement the entity ID fusion method as described in any one of claims 1-7.
9. An electronic device, characterized in that: The invention comprises a processor and the non-transitory computer-readable storage medium as claimed in claim 8.