Anonymous file merging method and device and electronic equipment

Through big data analysis and trajectory data matching, combined with portrait image feature similarity screening and merging, the problems of inefficient and high labor costs of anonymous archive merging in the existing technology are solved, efficient and accurate archive merging is achieved, and practical application effects are improved.

CN120067230APending Publication Date: 2025-05-30XIAMEN MEIYABAIKE INFORMATION SECURITY RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510097136.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-30
Filing Date
2025-01-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology is inefficient and labor-intensive in the process of merging anonymous files, making it difficult to effectively solve the problem of multiple files in the portrait archive.

Method used

Through big data analysis, the trajectory data matching analysis of the full archives in the portrait archive library is carried out, more and more accurate candidate archive pairs are recalled, and screened and merged based on the similarity of the portrait image feature.

Benefits of technology

It improves the efficiency and accuracy of anonymous file merging, effectively solves the problem of multiple files in the portrait archive, improves the practical effect of subsequent applications, and provides more powerful support for security work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067230A_ABST
    Figure CN120067230A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of big data analysis, and discloses an anonymous archive merging method and device and electronic equipment, and the method comprises the steps: obtaining first track data corresponding to a real-name archive and second track data corresponding to an anonymous archive in a portrait archive library; determining a first grid identifier of a geosphere grid where the longitude and latitude are located in the first trajectory data and a second grid identifier of a geosphere grid where the longitude and latitude are located in the second trajectory data; matching the real-name archives with the anonymous archives, and matching different anonymous archives to obtain a first candidate archive pair and a second candidate archive pair; screening the first candidate archive pair and the second candidate archive pair based on portrait image feature similarity to obtain a screened first candidate archive pair and a screened second candidate archive pair; and carrying out file merging on the screened first candidate file pair and the screened second candidate file pair. The anonymous file merging efficiency and accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of big data analysis, and particularly relates to a method, device and electronic device for merging anonymous files. Background Art

[0002] The portrait file clustering technology usually converts multiple collected portrait images into feature vectors respectively through a deep neural network model to obtain the feature vectors corresponding to each portrait image; then clusters multiple feature vectors through a clustering algorithm, and clusters the portrait images belonging to the same object into the same type of portrait images; and matches the clustered portrait images with the portrait images included in the portrait file library; if the clustered portrait images match the portrait images included in a certain file, the clustered portrait images are merged into the file, otherwise a new file is generated.

[0003] In practical applications, the portrait files in the portrait file library will rely on the local identity library for identity confidence. The files that can determine the portrait identity are called real-name files, and the files that cannot determine the portrait identity are determined as anonymous files. However, due to the diversity and wide distribution of image acquisition devices, the quality of images of the same object collected by different devices is unstable, making it difficult to accurately identify images of the same object in different situations. This results in a target object in the archive library may be split into multiple files in practical applications, and this phenomenon is called one person with multiple files.

[0004] To reduce the occurrence of the situation of one person with multiple files, it is necessary to analyze anonymous files and merge anonymous files including the same target object. The technical solution adopted by the prior art is to manually review multiple anonymous files in the file library, and judge whether multiple anonymous files include the same object by observing images and comparing information, etc. If it is judged that multiple anonymous files include the same object, these multiple anonymous files are merged. However, this method of merging anonymous files in the prior art has the problems of low efficiency and high labor cost. Summary of the Invention

[0005] The purpose of the present invention is to carry out trajectory data matching analysis on all files in the portrait file library based on big data analysis, improve the comprehensiveness of the analysis, recall more and more accurate candidate file pairs, and improve the efficiency and accuracy of anonymous file merging.

[0006] In a first aspect, an embodiment of the present invention provides a method for merging anonymous files, the method comprising:

[0007] Obtain the first trajectory data corresponding to the real-name archives and the second trajectory data corresponding to the anonymous archives in the portrait archive. Among them, the trajectory data of an archive includes the acquisition time of the portrait image in the archive, the longitude and latitude of the area corresponding to the portrait image, the age data and gender data of the portrait in the portrait image, and the number of days when the trajectory data appears;

[0008] Respectively determine the abnormal trajectory data in the first trajectory data and the second trajectory data, and delete the abnormal trajectory data in the first trajectory data and the second trajectory data;

[0009] Respectively perform geocoding on the longitude and latitude in the first trajectory data and the longitude and latitude in the second trajectory data to obtain the first grid identifier of the earth grid where the longitude and latitude in the first trajectory data are located, and the second grid identifier of the earth grid where the longitude and latitude in the second trajectory data are located;

[0010] Based on the first grid identifier and the second grid identifier, the acquisition time of the portrait images in the first trajectory data and the second trajectory data, the age data and gender data in the first trajectory data and the second trajectory data, and the number of days corresponding to the first trajectory data and the second trajectory data respectively, match the real-name archives and the anonymous archives, and match different anonymous archives to obtain the first candidate archive pairs composed of real-name archives and anonymous archives, and the second candidate archive pairs composed of different anonymous archives;

[0011] Based on the similarity of portrait image features, screen the first candidate archive pairs and the second candidate archive pairs to obtain the screened first candidate archive pairs and the screened second candidate archive pairs; for any candidate archive pair in the screened first candidate archive pairs and the screened second candidate archive pairs, the similarity of the portrait image features of the two archives included in the candidate archive pair is greater than the preset similarity;

[0012] Merge the screened first candidate archive pairs and the screened second candidate archive pairs.

[0013] Optionally, the trajectory data further includes the storage time of the portrait image and the device identifier of the portrait image acquisition device. The step of respectively determining the abnormal trajectory data in the first trajectory data and the second trajectory data, and deleting the abnormal trajectory data in the first trajectory data and the second trajectory data includes:

[0014] For any one of the first trajectory data and the second trajectory data, if the time difference between the acquisition time of the portrait image in the trajectory data and the storage time of the portrait image is greater than the first preset time difference, determine that the acquisition time of the portrait image is abnormal, and delete the acquisition time of the portrait image from the trajectory data;

[0015] Compare the longitude and latitude in the trajectory data with the longitude and latitude range of the city where the portrait image acquisition device is located. If the longitude and latitude in the trajectory data are not within the longitude and latitude range, determine that the longitude and latitude in the trajectory data are abnormal, and delete the abnormal longitude and latitude from the trajectory data.

[0016] Optionally, the separately determining the abnormal trajectory data in the first trajectory data and the second trajectory data, and deleting the abnormal trajectory data in the first trajectory data and the second trajectory data includes:

[0017] For any one of the first trajectory data and the second trajectory data, compare the multiple acquisition times corresponding to the same device identifier in the trajectory data;

[0018] If the time difference between the multiple acquisition times is less than the second preset time difference, compare the storage times corresponding to the multiple acquisition times respectively, and retain the target acquisition time corresponding to the latest storage time, and delete the acquisition times other than the target acquisition time among the multiple acquisition times.

[0019] Optionally, the method further includes:

[0020] For the second trajectory data corresponding to each anonymous file in the portrait archive, count the number of trajectories formed by the second trajectory data, the number of days when the second trajectory data appears, and the number of device identifiers of the portrait image acquisition devices included in the second trajectory data;

[0021] Judge whether the number of trajectories is within the first preset number range, whether the number of days is within the preset number of days range, and whether the number of device identifiers is within the second preset number range;

[0022] If the number of trajectories, the number of days, and the number of target device identifiers do not meet the preset conditions, delete the second trajectory data, where the preset conditions are that the number of trajectories is within the first preset number range, the number of days is within the preset number of days range, and the number of device identifiers is within the second preset number range.

[0023] Optionally, based on the first grid identifier and the second grid identifier, the acquisition times of portrait images in the first trajectory data and the second trajectory data, the age data and gender data in the first trajectory data and the second trajectory data, and the number of days of appearance corresponding to the first trajectory data and the second trajectory data respectively, obtaining a first candidate file pair composed of a real-name file and an anonymous file, including:

[0024] For any real-name file and any anonymous file, match the first grid identifier in the first trajectory data corresponding to the real-name file with the second grid identifier in the second trajectory data corresponding to the anonymous file, match the acquisition time in the first trajectory data with the acquisition time in the second trajectory data, and match the gender data in the first trajectory data with the gender data in the second trajectory data;

[0025] If the matching conditions are met, determine the real-name file and the anonymous file as a pre-candidate file pair; the matching conditions are that the first grid identifier matches the second grid identifier, the acquisition times match, and the gender data matches;

[0026] For the real-name file and the anonymous file included in the pre-candidate file pair, calculate the age difference between the first age data and the second age data, where the first age data is the age data in the first trajectory data corresponding to the real-name file, and the second age data is the age data in the second trajectory data corresponding to the anonymous file;

[0027] If the age difference is less than a preset age difference, count the number of days the first trajectory data corresponding to the real-name file appears and the number of first grid identifiers in the first trajectory data, and the number of days the second trajectory data corresponding to the anonymous file appears and the number of second grid identifiers in the second trajectory data;

[0028] If the number of days the first trajectory data appears and the number of days the second trajectory data appears are both within a preset number of days range, and the number of first grid identifiers and the number of second grid identifiers are both within a third preset quantity range, determine the pre-candidate file pair as the first candidate file pair.

[0029] Optionally, the method further includes:

[0030] After obtaining the first candidate file pair, determine whether there is a situation where one target anonymous file matches multiple real-name files among the multiple first candidate file pairs;

[0031] If there is a situation where one target anonymous file matches multiple real-name files, sort the multiple real-name files in descending order of the number of first grid identifiers in the first trajectory data corresponding to the multiple real-name files, and / or in descending order of the number of days when the first trajectory data appears, and determine the real-name file with the sorting serial number of 1 as the target real-name file;

[0032] Delete, from the multiple first candidate file pairs formed by one target anonymous file and multiple real-name files, the first candidate file pairs other than the target candidate file pair, where the target candidate file pair is the candidate file pair formed by the target anonymous file and the target real-name file.

[0033] Optionally, the method further includes:

[0034] After obtaining the first candidate file pairs and the second candidate file pairs, determine whether the first anonymous file and the second anonymous file included in any second candidate file pair appear in different first candidate file pairs;

[0035] If the first anonymous file and the second anonymous file appear in different first candidate file pairs, delete the second candidate file pair formed by the first anonymous file and the second anonymous file.

[0036] Optionally, performing file merging on the filtered first candidate file pairs and the filtered second candidate file pairs includes:

[0037] For any filtered first candidate file pair, update the file identifier of the anonymous file in the first candidate file pair to the file identifier of the real-name file in the first candidate file pair, update the file information of the anonymous file to the file information of the real-name file, and delete the file information of the anonymous file.

[0038] For any filtered second candidate file pair, use the anonymous file with a larger number of trajectory quantities in the trajectory formed by the trajectory data in the second candidate file pair as the main file, use the other anonymous file in the second candidate file pair as the secondary file, update the file identifier of the secondary file to the file identifier of the main file, update the file information of the secondary file to the file information of the main file, and delete the file information of the secondary file.

[0039] In a second aspect, an embodiment of the present invention provides an anonymous file merging device, and the device includes:

[0040] A trajectory data acquisition module, configured to acquire first trajectory data corresponding to real-name archives and second trajectory data corresponding to anonymous archives in a portrait archive library, wherein the trajectory data of an archive includes the acquisition time of the portrait image in the archive, the longitude and latitude of the area corresponding to the portrait image, the age data and gender data of the portrait in the portrait image, and the number of days when the trajectory data appears;

[0041] An abnormal trajectory data deletion module, configured to respectively determine the abnormal trajectory data in the first trajectory data and the second trajectory data, and delete the abnormal trajectory data in the first trajectory data and the second trajectory data;

[0042] A grid identification determination module, configured to respectively perform geocoding on the longitude and latitude in the first trajectory data and the longitude and latitude in the second trajectory data to obtain a first grid identification of the earth grid where the longitude and latitude in the first trajectory data are located, and a second grid identification of the earth grid where the longitude and latitude in the second trajectory data are located;

[0043] A candidate archive pair determination module, configured to match real-name archives and anonymous archives, and match different anonymous archives based on the first grid identification and the second grid identification, the acquisition time of the portrait images in the first trajectory data and the second trajectory data, the age data and gender data in the first trajectory data and the second trajectory data, and the number of days when the first trajectory data and the second trajectory data respectively appear, so as to obtain a first candidate archive pair composed of a real-name archive and an anonymous archive, and a second candidate archive pair composed of different anonymous archives;

[0044] A candidate archive pair screening module, configured to screen the first candidate archive pair and the second candidate archive pair based on the similarity of portrait image features to obtain a screened first candidate archive pair and a screened second candidate archive pair; for any candidate archive pair in the screened first candidate archive pair and the screened second candidate archive pair, the similarity of the portrait image features of the two archives included in the candidate archive pair is greater than a preset similarity;

[0045] An archive merging module, configured to merge the screened first candidate archive pair and the screened second candidate archive pair.

[0046] In a third aspect, an embodiment of the present invention provides an electronic device, including:

[0047] At least one processor;

[0048] A memory for storing executable instructions of the at least one processor;

[0049] Wherein, the at least one processor is configured to execute the instructions to implement the method described in the first aspect.

[0050] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in the first aspect.

[0051] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, which implements the method described in the first aspect when executed by a processor.

[0052] It can be seen that the present invention is based on big data analysis, conducts trajectory data matching analysis on all files in the portrait file library, improves the comprehensiveness of the analysis, can recall more and more accurate candidate file pairs. Compared with the prior art, the efficiency of anonymous file merging is higher, and the accuracy of anonymous file merging is relatively high. It can effectively solve the problem of multiple files for one person in the portrait file library, improve the actual combat effect of subsequent applications based on portrait file trajectories, and provide more powerful support for security work. Description of the Drawings

[0053] Figure 1 Schematic diagram of multiple files corresponding to one person;

[0054] Figure 2 Flowchart of the overall technical solution provided by the embodiment of the present invention;

[0055] Figure 3 Relationship graph composed of all candidate file pairs provided by the embodiment of the present invention;

[0056] Figure 4 Flowchart of a method for merging anonymous files provided by the embodiment of the present invention;

[0057] Figure 5 Schematic structural diagram of a device for merging anonymous files provided by the embodiment of the present invention;

[0058] Figure 6 Schematic structural diagram of an electronic device provided by the embodiment of the present invention. Detailed Embodiments

[0059] The present invention will be described in detail below through embodiments.

[0060] Portrait clustering technology usually converts multiple collected portrait images into feature vectors respectively through a deep neural network model to obtain the feature vectors corresponding to each portrait image; then uses a clustering algorithm to cluster multiple feature vectors, clustering the portrait images belonging to the same object into the same type of portrait images; and matches the clustered portrait images with the portrait images included in the portrait archive in the portrait archive; if the clustered portrait images match the portrait images included in a certain archive, the clustered portrait images are merged into that archive, otherwise a new archive is generated.

[0061] In practical applications, the portrait archives in the portrait archive will use the local identity database to conduct identity confidence. The archives that can determine the portrait identity are called real-name archives, and the archives that cannot determine the portrait identity are determined as anonymous archives. And, if the portrait identities of the portraits included in two real-name archives are the same, then these two real-name archives are merged into the same real-name archive, that is to say, real-name archives are unique, each real-name archive corresponds to a portrait identity, and different real-name archives correspond to different portrait identities.

[0062] However, due to the diversity and wide distribution of image acquisition devices, the quality of images of the same object collected by different devices is unstable. For example, affected by factors such as light, angle, and resolution, it is difficult to accurately identify images including the same object as the same object under different circumstances. This leads to the situation that a target object in practical applications may be split into multiple archives in the archive, and this phenomenon is called one person with multiple files. Specifically, there may be one real-name file for a person in the local identity database, but there may also be multiple anonymous files. And for people not in the local identity database, there may be multiple anonymous files. For example, as Figure 1 shown, a person corresponds to multiple files such as a real-name file, anonymous file A, anonymous file B, and anonymous file C.

[0063] The situation of one person with multiple files will lead to incomplete and fragmented file trajectories in the archive, bringing many problems to subsequent various applications based on file trajectory analysis and seriously affecting the actual combat effect. For example, in the tracking and control of a target object, because the target object corresponds to multiple files, it may cause errors in the trajectory analysis of the target object and make it impossible to accurately judge the action route and activity range of the target object. In security warning and prevention, multiple files may cause confusion in the system's recognition of personnel and reduce the accuracy and timeliness of warnings.

[0064] To reduce the occurrence of multiple files for one person, it is necessary to analyze anonymous files and merge anonymous files that include the same target object. Specifically, the anonymous files that include the target object can be merged into the real-name file that includes this anonymous file, or multiple anonymous files that include the same target object can be merged into one anonymous file. Currently, the technical solution adopted by the prior art is to arrange professional personnel to manually review multiple files in the file library, and determine whether multiple files include the same object by observing images and comparing information. If it is determined that multiple files include the same object, these multiple files will be merged. However, this method of file merging in the prior art has the problems of low efficiency and high labor costs, and is not applicable to file libraries that include a large number of files.

[0065] Another file merging solution in the prior art is: starting from the similarity of image features for algorithm design. This method calculates the similarity of image features of portrait images included in different files through a similarity algorithm. This method has a high computing power requirement for calculating the similarity of image features and cannot calculate all files, resulting in a low coverage rate.

[0066] Specifically, this method needs to calculate the similarity of portrait image features between different files. However, in the scenario of a city-level file library with tens of millions of files, if all files need to be analyzed, that is, if all files in the file library are analyzed, it will inevitably involve a huge amount of calculations, which is unrealistic. In practical applications, only strategies can be adopted to select files with high priorities for analysis. These files with high priorities are used as retrieval files, and candidate files with higher similarity to the retrieval files are retrieved according to the similarity of picture features. Further, combined with other information, such as spatio-temporal information, basic file information, etc., it is more precisely determined whether the candidate files can be merged with the retrieval files. It can be seen that the computing power required by the file merging solution is relatively high and is not applicable to file libraries that include a large number of files.

[0067] To solve the above technical problems, the embodiment of the present invention is based on a big data analysis method to perform matching analysis on the trajectory data of all files, and recall candidate file pairs to be merged. Then, the candidate file pairs to be merged are secondarily verified by using the similarity of image features, thereby improving the accuracy of file merging, having a high efficiency of file merging, and having a small computing power required during the file merging process.

[0068] For the sake of clear description of the solution, the overall technical solution of the embodiment of the present invention will be elaborated in detail below. As Figure 2 described, the overall technical solution of the embodiment of the present invention may include the following steps:

[0069] S210, obtain the trajectory data of all files.

[0070] Among them, the full - volume archives are all the archives in the portrait archive library, which include real - name archives and anonymous archives.

[0071] Using big data technologies such as Spark, read the trajectory data of the full - volume archives in the recent N days. Among them, the Spark big data technology is a fast, general - purpose, and scalable big data analysis and computing engine based on memory. The value of N can be set according to business experience. Generally, N takes 30 days or 60 days to meet the analysis requirements. Of course, the size of N can also be adjusted according to actual needs. The embodiments of the present invention do not specifically limit the size of N.

[0072] For each archive in the full - volume archives, the trajectory data of the archive may include the archive ID, the acquisition time of the portrait image in the archive, the longitude and latitude of the area where the portrait image belongs, the device ID of the portrait image acquisition device, the device address of the portrait image acquisition device, the age of the portrait in the portrait image, the gender of the portrait in the portrait image, and the storage time of each portrait image in the library.

[0073] Among them, if an archive is a real - name archive, the archive ID of the real - name archive is its corresponding portrait identity ID. If an archive is an anonymous archive, the archive ID is the ID of the anonymous archive. The age of the portrait in the above - mentioned portrait image is the age obtained by analyzing the portrait image through a face model; the gender of the portrait in the portrait image is the gender obtained by analyzing the portrait image through a face model.

[0074] S220, pre - process the trajectory data. Among them, the following pre - processing can be performed on the trajectory data:

[0075] 1. For the trajectory data of any archive, the trajectory data with abnormal acquisition time and abnormal longitude and latitude can be removed. Among them, the method for judging abnormal acquisition time can be: judging according to the time difference between the acquisition time of the portrait image and the storage time of the portrait image; if the time difference between the acquisition time and the storage time of the portrait image is greater than the preset time difference, judge that the acquisition time is abnormal and delete the abnormal acquisition data from the trajectory data.

[0076] The method for judging abnormal longitude and latitude can be: comparing the longitude and latitude in the trajectory data with the longitude and latitude range of the area where the portrait image acquisition device is located. If the longitude and latitude in the trajectory data are not within the longitude and latitude range of the city where the portrait image belongs, judge that the longitude and latitude range is abnormal and delete the abnormal longitude and latitude from the trajectory data.

[0077] 2. For the trajectory data of any file, the trajectory data with an age lower than the preset threshold recognized by the face model can be excluded. In practical applications, the preset threshold can be set to a relatively small value to screen out low ages through the preset threshold. Since the current face model has a poor effect on feature extraction of young children, the trajectory data of young children is not analyzed.

[0078] 3. For the trajectory data of any file, the acquisition time of the portrait image is accurate to the minute level. For multiple portrait images of a file, the repeated acquisition time and storage time can be deleted according to the acquisition time of each portrait image and the device ID of the portrait image acquisition device, and only the latest storage time and the acquisition time corresponding to the storage time are retained.

[0079] Compare the multiple acquisition times corresponding to the same device identifier in the trajectory data. If the time difference between the multiple acquisition times is small, it means that the multiple portrait images may be portrait images repeatedly acquired by the same image acquisition device. At this time, the repeated acquisition times can be excluded. Specifically, the storage times corresponding to the multiple acquisition times can be compared, and the target acquisition time corresponding to the latest storage time is retained, and the acquisition times other than the target acquisition time among the multiple acquisition times are deleted.

[0080] 4. For any file in the portrait archive, calculate the gender data of the portrait image of the file. Here, the gender data refers to the gender data recognized by the face model by inputting the portrait image into the face model. In practical applications, the gender mode of multiple portrait images in the file can be used as the gender data of the file. The gender mode refers to the gender with a relatively large proportion among multiple portrait images. For example, when multiple portrait images of the file are input into the face recognition model and 90% of the genders output from the face recognition model are female, then the gender data of the file is female.

[0081] 5. For any file in the portrait archive, calculate the age data of the portrait image of the file. Here, the age data refers to the age data recognized by the face model by inputting the portrait image into the face model. In practical applications, the age mode of multiple portrait images in the file can be used as the age data of the file, where the age mode refers to the age with a relatively large proportion among multiple portrait images. If there is no obvious distinguishable age mode, the median of the multiple ages corresponding to the multiple portrait images is taken.

[0082] 6. For the second trajectory data corresponding to each anonymous file in the portrait archive, count the number of trajectories formed by the second trajectory data, the number of days when the second trajectory data appears, and the number of device identifiers of the portrait image acquisition devices included in the second trajectory data; determine whether the number of trajectories is within the first preset number range, whether the number of days is within the preset number of days range, and whether the number of device identifiers is within the second preset number range. If all three meet the threshold requirements, perform subsequent analysis on the second trajectory data of the anonymous file; otherwise, stop further analysis. This is because files outside the threshold range are often abnormal files, and the significance of analysis is small. Instead, they are interference items.

[0083] Among them, the first preset number range, the preset number of days range, and the second preset number range can all be set according to actual situations, and the embodiments of the present invention do not make specific limitations in this regard. For example, for the preset number of days range, according to business experience, it can be within 30 days or within 60 days.

[0084] 7. Use the GeoHash algorithm to perform geocoding on the longitude and latitude in the trajectory data, and convert the longitude and latitude in the trajectory data into a one-dimensional string. In practical applications, the first 8 digits of the one-dimensional string can be extracted as the grid ID of the longitude and latitude, that is, the grid identifier of the longitude and latitude; or, the first 7 digits can also be extracted as the grid ID of the longitude and latitude; or, both levels can be extracted, which is all acceptable.

[0085] Among them, the basic principle of the GeoHash algorithm is: through the method of binary search, the earth's surface is divided into multiple rectangular regions, and each region is assigned a unique code. This division is based on the longitude and latitude range, where the longitude is [-180 degrees, 180 degrees], and the latitude range is [-90 degrees, 90 degrees]. By recursively dividing the current interval into smaller sub-intervals until the required accuracy is reached.

[0086] The encoding process of the GeoHash algorithm can include the following steps, namely steps a to c:

[0087] Step a, initialize the interval. Specifically, set the initial longitude and latitude interval to the entire range of the earth.

[0088] Step b, recursive division. In each iteration, divide the current interval into 4 sub-intervals according to the midpoint, determine which sub-interval the target point belongs to according to the longitude and latitude of the target point, and update the code accordingly.

[0089] Step c, cross encoding. Alternately arrange the binary codes of longitude and latitude to form the final GeoHash code. Even bits represent longitude, and odd bits represent latitude. It can be understood that those skilled in the art should be able to understand the encoding process of the GeoHash algorithm, and it will not be elaborated in too much detail here.

[0090] S230, Trajectory data matching analysis. The specific implementation of trajectory data matching analysis is as follows:

[0091] The following will introduce in detail the trajectory data matching analysis of any real-name file and any anonymous file in the portrait file library.

[0092] The first step is to first perform trajectory matching according to the 8-digit grid identifier, the same acquisition time (in minutes), and the same file gender data. If the matching conditions are met, the real-name file and the anonymous file are determined as pre-candidate file pairs. Among them, the matching conditions are that the first grid identifier matches the second grid identifier, the acquisition time matches, and the gender data matches;

[0093] The second step is to calculate the age difference between the first age data corresponding to the anonymous file and the second age data corresponding to the real-name file for the pre-candidate file pair composed of <anonymous file, real-name file>, and eliminate the pre-candidate file pair composed of <anonymous file, real-name file> whose age difference is not within the preset threshold range. The remaining pre-candidate file pairs composed of <anonymous file, real-name file> with age differences within the preset threshold range are left.

[0094] The third step is to, for any pre-candidate file pair selected in the second step, count the number of days the first trajectory data corresponding to the real-name file appears in the pre-candidate file, the number of first grid identifiers in the first trajectory data, the number of days the second trajectory data corresponding to the anonymous file appears, and the number of second grid identifiers in the second trajectory data.

[0095] The fourth step is to screen out the pre-candidate file pairs whose number of grid identifiers in the trajectory data and the number of days the trajectory data appears are both within the preset threshold range as the <anonymous file, real-name file> candidate file pairs.

[0096] If the same anonymous file matches multiple real-name files, then, according to the number of days the trajectory data appears and the number of grid identifiers, sort the multiple real-name files in descending order, and use the file pair composed of the real-name file with the sorting serial number 1 and the anonymous file as the candidate file pair.

[0097] In practical applications, in the first step, the grid identifier can be changed to a 7-digit grid, and the time can be changed to the 10-minute level, and then the trajectory data matching analysis is performed. At this time, the preset threshold range in the fourth step can be modified to a strict value according to actual needs. Repeat the first step to the fourth step, and merge and deduplicate the obtained results with the results already obtained in the above first step to the fourth step. The deduplication gives priority to retaining the results with more refined grid identifiers and time levels, and more candidate file pairs can be recalled without losing accuracy.

[0098] The specific implementation method of obtaining the first candidate file pair composed of real-name files and anonymous files is introduced above. The specific implementation method of obtaining the second candidate file pair composed of different anonymous files is similar and will not be elaborated here.

[0099] After obtaining the first candidate file pair and the second candidate file pair, that is, all candidate file pairs are obtained. All candidate file pairs will constitute a relationship graph, as Figure 3 shown. From Figure 3 it can be seen that there will be two anonymous files connected (anonymous file 2 and anonymous file 3), but anonymous file 2 and anonymous file 3 point to different real-name files (real-name file 1 and real-name file 2 respectively). At this time, the association between anonymous file 2 and anonymous file 3 will be removed, that is, it is not allowed that the associated anonymous files point to multiple real-name files. Then, from the relationship graph, two types of matching associations will be extracted:

[0100] The first type is: the first candidate file pair of <multiple anonymous files, real-name file>, such as <anonymous file 4, anonymous file 2, anonymous file 1, real-name file 1>.

[0101] The second type is: the association of anonymous files that do not point to real-name files, such as the second candidate file pair composed of <anonymous file 5, anonymous file 6>.

[0102] S240, perform secondary verification on the candidate file pair determined in S230 through the image feature similarity of portrait images.

[0103] For the candidate file pair determined in S230, the present invention will perform secondary verification based on the image feature similarity. For the first candidate file pair pointing to the real-name file, take the representative portrait image of the anonymous file or the central feature vector of the anonymous file, and compare the similarity with the identity image feature of the real-name file. If the similarity is greater than the preset similarity, then the anonymous file in this first candidate file pair will be merged into the real-name file. Among them, the preset similarity can be determined according to the actual situation. The central feature vector is the average vector of the feature vectors of multiple portrait images in the anonymous file.

[0104] For the second candidate file pair without pointing to the real-name file, take the representative pictures of the two anonymous files in this second candidate file pair or the central feature vectors of the anonymous files to compare the similarity. If the similarity is greater than the preset similarity, then the file merging of this second candidate file pair will be performed next.

[0105] Moreover, in practical applications, other dimension information can be fused on the basis of the feature similarity, and a machine learning classification model can be used for discrimination.

[0106] S250, file merging and updating.

[0107] Specifically, after obtaining the first candidate file pair and the second candidate file pair to be merged through S240, for any first candidate file pair, update the file identifier to which the trajectory data of the anonymous file belongs to the file identifier of the corresponding real-name file, then update the file information of the anonymous file to the file information of the real-name file, and finally delete the file information of the anonymous file.

[0108] For any second candidate file pair, take the anonymous file with more trajectory counts as the main file, take the other anonymous file as the secondary file, update the file identifier to which the trajectory data of the secondary file belongs to the file identifier of the main file, then update the file information of the secondary file to the file information of the main file, and finally delete the file information of the secondary file.

[0109] It can be seen that the present invention is based on big data analysis, conducts trajectory data matching analysis on all files in the portrait file library, improves the comprehensiveness of the analysis, can recall more and more accurate candidate file pairs. Compared with the prior art, the efficiency of merging anonymous files is higher, and the accuracy of merging anonymous files is relatively high, which can effectively solve the problem of multiple files for one person in the portrait file library, improve the actual combat effect of subsequent applications based on portrait file trajectories, and provide more powerful support for security work.

[0110] Next, a method for merging anonymous files provided by an embodiment of the present invention will be elaborated in detail.

[0111] As Figure 4 shown, a method for merging anonymous files provided by an embodiment of the present invention may include the following steps:

[0112] S410, obtain the first trajectory data corresponding to the real-name files and the second trajectory data corresponding to the anonymous files in the portrait file library.

[0113] Among them, the trajectory data of a file includes the acquisition time of the portrait image in the file, the longitude and latitude of the area corresponding to the portrait image, the age data and gender data of the portrait in the portrait image, and the number of days when the trajectory data appears.

[0114] Specifically, the embodiment of the present invention conducts trajectory data analysis on all files in the portrait file library. The all files are all files in the portrait file library, which include real-name files and anonymous files. Big data technologies such as Spark can be used to read the trajectory data of all files in the recent N days. The value of N can be set according to business experience. Generally, N is taken as 30 days or 60 days.

[0115] For each file in the full set of files, the trace data of the file is used to characterize the attribute information of the file in multiple dimensions. The trace data may include file identification, the acquisition time of the portrait image in the file, the longitude and latitude of the region to which the portrait image belongs, the device identification of the portrait image acquisition device, the device address of the portrait image acquisition device, the age of the portrait in the portrait image, the gender of the portrait in the portrait image, and the storage time of each portrait image, etc. Of course, in practical applications, the trace data of the file may also include other data, and the embodiments of the present invention do not make specific limitations on this.

[0116] S420. Respectively determine the abnormal trace data in the first trace data and the second trace data, and delete the abnormal trace data in the first trace data and the second trace data.

[0117] Specifically, the acquisition time, longitude and latitude, and age data in the trace data may be abnormal, and multiple acquisition times may be repeated acquisition times with a small time difference. Multiple repeated acquisition times also belong to abnormal trace data. In order to improve the accuracy of the candidate file pairs determined subsequently, the abnormal trace data in the first trace data and the second trace data can be deleted.

[0118] For the sake of clear description of the solution, the specific implementation manner of S420 will be elaborated in detail in the following embodiments.

[0119] S430. Respectively perform geocoding on the longitude and latitude in the first trace data and the longitude and latitude in the second trace data to obtain the first grid identifier of the longitude and latitude in the first trace data in the earth grid where they are located, and the second grid identifier of the longitude and latitude in the second trace data in the earth grid where they are located.

[0120] Specifically, the GeoHash algorithm can be used to perform geocoding on the longitude and latitude in the trace data to convert the longitude and latitude in the trace data into a one-dimensional string. In practical applications, the first 8 digits of the one-dimensional string can be extracted as the grid ID of the longitude and latitude, that is, the grid identifier of the longitude and latitude; or, the first 7 digits can also be extracted as the grid ID of the longitude and latitude; or, both levels can be extracted, which is all acceptable.

[0121] The principle and encoding steps of the GeoHash algorithm have been introduced in the above embodiments, and will not be elaborated here.

[0122] S440. Based on the first grid identifier and the second grid identifier, the acquisition times of the portrait images in the first trajectory data and the second trajectory data, the age data and gender data in the first trajectory data and the second trajectory data, and the number of days of appearance corresponding to the first trajectory data and the second trajectory data respectively, match the real-name file and the anonymous file, and match different anonymous files to obtain a first candidate file pair composed of the real-name file and the anonymous file, and a second candidate file pair composed of different anonymous files.

[0123] Specifically, files in the portrait file library can be matched through the grid identifier, the acquisition time of the portrait image, the age data and gender data, and the number of days of appearance in the trajectory data. Since each real-name file corresponds to one identity information, there is no need to merge real-name files. Anonymous files may need to be merged into real-name files, or two different anonymous files may need to be merged. Therefore, the real-name file and the anonymous file can be matched, and different anonymous files can be matched to obtain a first candidate file pair composed of the real-name file and the anonymous file, and a second candidate file pair composed of different anonymous files.

[0124] In one implementation, S440. Based on the first grid identifier and the second grid identifier, the acquisition times of the portrait images in the first trajectory data and the second trajectory data, the age data and gender data in the first trajectory data and the second trajectory data, and the number of days of appearance corresponding to the first trajectory data and the second trajectory data respectively, obtaining a first candidate file pair composed of the real-name file and the anonymous file may include the following steps, namely steps a1 to a5:

[0125] Step a1. For any real-name file and any anonymous file, match the first grid identifier in the first trajectory data corresponding to the real-name file and the second grid identifier in the second trajectory data corresponding to the anonymous file, match the acquisition time in the first trajectory data and the acquisition time in the second trajectory data, and match the gender data in the first trajectory data and the gender data in the second trajectory data.

[0126] Step a2. If the matching conditions are met, determine the real-name file and the anonymous file as a pre-candidate file pair. The matching conditions are that the first grid identifier matches the second grid identifier, the acquisition times match, and the gender data match.

[0127] Step a3. For the real-name file and the anonymous file included in the pre-candidate file pair, calculate the age difference between the first age data and the second age data. The first age data is the age data in the first trajectory data corresponding to the real-name file, and the second age data is the age data in the second trajectory data corresponding to the anonymous file.

[0128] Step a4, if the age difference is less than the preset age difference, count the number of days when the first trajectory data corresponding to the real-name file appears and the number of first grid identifiers in the first trajectory data, as well as the number of days when the second trajectory data corresponding to the anonymous file appears and the number of second grid identifiers in the second trajectory data.

[0129] Step a5, if the number of days when the first trajectory data appears and the number of days when the second trajectory data appears are both within the preset number of days range, and the number of first grid identifiers and the number of second grid identifiers are both within the third preset quantity range, determine the pre-candidate file pair as the first candidate file pair.

[0130] Among them, steps a1 to a5 have been elaborated in detail in the overall technical solution embodiment and will not be repeated here. Also, the determination method of the second candidate file pair is similar to that of the first candidate file pair and will not be repeated here either.

[0131] S450, screen the first candidate file pair and the second candidate file pair based on the similarity of portrait image features to obtain the screened first candidate file pair and the screened second candidate file pair.

[0132] Among them, for any candidate file pair in the screened first candidate file pair and the screened second candidate file pair, the similarity of the portrait image features of the two files included in the candidate file pair is greater than the preset similarity.

[0133] Specifically, after determining the first candidate file pair and the candidate file pair through S410 to S450, in order to improve the accuracy of the finally merged file pair, the present invention will perform a secondary verification on the first candidate file pair and the second candidate file pair based on the image feature similarity.

[0134] For the first candidate file pair pointing to the real-name file, take the representative portrait image of the anonymous file or the central feature vector of the anonymous file, and perform a similarity comparison with the identity image features of the real-name file. If the similarity is greater than the preset similarity, next, merge the anonymous file in the first candidate file pair into the real-name file. Among them, the preset similarity can be determined according to the actual situation. The central feature vector is the average vector of the portrait image feature vectors in the anonymous file.

[0135] For the second candidate file pair without pointing to the real-name file, take the representative pictures of the two anonymous files in the second candidate file pair or the central feature vector of the anonymous file for similarity comparison. If the similarity is greater than the preset similarity, next, merge the second candidate file pair.

[0136] S460, perform file merging on the screened first candidate file pair and the screened second candidate file pair.

[0137] In one embodiment, S460, merging the screened first candidate file pairs and the screened second candidate file pairs may include the following steps, namely step b1 and step b2:

[0138] Step b1, for any screened first candidate file pair, update the file identifier of the anonymous file in the first candidate file pair to the file identifier of the real-name file in the first candidate file pair, update the file information of the anonymous file to the file information of the real-name file, and delete the file information of the anonymous file.

[0139] Step b2, for any screened second candidate file pair, use the anonymous file with a larger number of track counts in the track formed by the track data in the second candidate file pair as the main file, use the other anonymous file in the second candidate file pair as the secondary file, update the file identifier of the secondary file to the file identifier of the main file, update the file information of the secondary file to the file information of the main file, and delete the file information of the secondary file.

[0140] Specifically, for any screened first candidate file pair, update the file identifier to which the track data of the anonymous file belongs to the corresponding file identifier of the real-name file, then update the file information of the anonymous file to the file information of the real-name file, and finally delete the file information of the anonymous file.

[0141] For any screened second candidate file pair, take the anonymous file with a larger number of track counts as the main file, use the other anonymous file as the secondary file, and update the file identifier to which the track data of the secondary file belongs to the file identifier of the main file, then update the file information of the secondary file to the file information of the main file, and finally delete the file information of the secondary file.

[0142] It can be seen that the present invention is based on big data analysis, conducts track data matching analysis on all files in the portrait file library, improves the comprehensiveness of the analysis, can recall more and more accurate candidate file pairs. Compared with the prior art, the efficiency of anonymous file merging is higher, and the accuracy of anonymous file merging is relatively high, which can effectively solve the problem of multiple files for one person in the portrait file library, improve the actual combat effect of subsequent applications based on portrait file tracks, and provide more powerful support for security work.

[0143] Based on the above embodiments, in one embodiment, S420, the track data further includes the storage time of the portrait image and the device identifier of the portrait image acquisition device. Respective determination of the abnormal track data in the first track data and the second track data, and deletion of the abnormal track data in the first track data and the second track data may include the following steps, namely step c1 and step c2:

[0144] Step c1, for any one of the first trajectory data and the second trajectory data, if the time difference between the acquisition time of the portrait image in the trajectory data and the storage time of the portrait image is greater than the first preset time difference, determine that the acquisition time of the portrait image is abnormal, and delete the acquisition time of the portrait image from the trajectory data;

[0145] Step c2, compare the longitude and latitude in the trajectory data with the longitude and latitude range of the city where the portrait image acquisition device is located. If the longitude and latitude in the trajectory data are not within the longitude and latitude range, determine that the longitude and latitude in the trajectory data are abnormal, and delete the abnormal longitude and latitude from the trajectory data.

[0146] For the trajectory data of any file, the trajectory data with abnormal acquisition time and abnormal longitude and latitude can be excluded. Among them, the method for judging abnormal acquisition time can be: judge according to the time difference between the acquisition time of the portrait image and the storage time of the portrait image; if the time difference between the acquisition time and the storage time of the portrait image is greater than the preset time difference, judge that the acquisition time is abnormal, and delete the abnormal acquisition data from the trajectory data.

[0147] The method for judging abnormal longitude and latitude can be: compare the longitude and latitude in the trajectory data with the longitude and latitude range of the area where the portrait image acquisition device is located. If the longitude and latitude in the trajectory data are not within the longitude and latitude range of the city to which the portrait image belongs, then judge that the longitude and latitude range is abnormal, and delete the abnormal longitude and latitude from the trajectory data.

[0148] It can be seen that by deleting the abnormal data in the trajectory data, it helps to select accurate candidate file pairs in the subsequent steps, thereby improving the efficiency and accuracy of anonymous file merging.

[0149] Based on the above embodiments, in one implementation manner, S420, respectively determine the abnormal trajectory data in the first trajectory data and the second trajectory data, and delete the abnormal trajectory data in the first trajectory data and the second trajectory data, which may include the following steps, namely step d1 and step d2:

[0150] Step d1, for any one of the first trajectory data and the second trajectory data, compare the multiple acquisition times corresponding to the same device identifier in the trajectory data;

[0151] Step d2, if the time difference between the multiple acquisition times is less than the second preset time difference, compare the storage times corresponding to the multiple acquisition times respectively, and retain the target acquisition time corresponding to the latest storage time, and delete the acquisition times other than the target acquisition time among the multiple acquisition times.

[0152] Specifically, for the trajectory data of any file, the acquisition time of the portrait image is accurate to the minute level. For multiple portrait images of a file, the duplicate acquisition time and storage time can be deleted according to the acquisition time of each portrait image and the device ID of the portrait image acquisition device, and only the latest storage time and the acquisition time corresponding to the storage time are retained.

[0153] Among them, the multiple acquisition times corresponding to the same device identifier in the trajectory data are compared. If the time difference between the multiple acquisition times is small, it means that multiple portrait images may be portrait images repeatedly acquired by the same image acquisition device. At this time, the duplicate acquisition times can be eliminated. Specifically, the storage times corresponding to the multiple acquisition times can be compared, and the target acquisition time corresponding to the latest storage time is retained, and the acquisition times other than the target acquisition time among the multiple acquisition times are deleted.

[0154] It can be seen that by deleting the duplicate acquisition times in the trajectory data, it helps to select accurate candidate file pairs in the subsequent steps, thereby improving the efficiency and accuracy of anonymous file merging.

[0155] Based on the above embodiments, in one implementation manner, the method may further include the following steps, namely steps e1 to e3:

[0156] Step e1, for the second trajectory data corresponding to each anonymous file in the portrait file library, count the number of trajectories formed by the second trajectory data, the number of days when the second trajectory data appears, and the number of device identifiers of the portrait image acquisition devices included in the second trajectory data.

[0157] Step e2, determine whether the number of trajectories is within the range of the first preset number, whether the number of days is within the preset number of days, and whether the number of device identifiers is within the range of the second preset number.

[0158] Step e3, if the number of trajectories, the number of days, and the number of the target device identifiers do not meet the preset conditions, delete the second trajectory data, where the preset conditions are that the number of trajectories is within the range of the first preset number, the number of days is within the preset number of days, and the number of device identifiers is within the range of the second preset number.

[0159] Specifically, for the second trajectory data corresponding to each anonymous file in the portrait archive, count the number of trajectories formed by the second trajectory data, the number of days when the second trajectory data appears, and the number of device identifiers of the portrait image acquisition devices included in the second trajectory data; determine whether the number of trajectories is within the first preset number range, whether the number of days is within the preset number of days range, and whether the number of device identifiers is within the second preset number range. If all three meet the threshold requirements, perform subsequent analysis on the second trajectory data of the anonymous file. Otherwise, stop further analysis because files outside the threshold range are often abnormal anonymous files, and the significance of analysis is small, and they are instead interference items. Therefore, deleting the second trajectory data of abnormal anonymous files helps to select accurate candidate file pairs in subsequent steps, thereby improving the efficiency and accuracy of anonymous file merging.

[0160] Based on the above embodiments, the method may further include the following steps, namely steps f1 to f3:

[0161] Step f1, after obtaining the first candidate file pairs, determine whether there is a situation where one target anonymous file matches multiple real-name files among the multiple first candidate file pairs;

[0162] Step f2, if there is a situation where one target anonymous file matches multiple real-name files, sort the multiple real-name files in descending order according to the number of first grid identifiers in the first trajectory data corresponding to the multiple real-name files, and / or, in descending order according to the number of days when the first trajectory data appears, and determine the real-name file with the sorting serial number 1 as the target real-name file;

[0163] Step f3, delete the first candidate file pairs other than the target candidate file pair among the multiple first candidate file pairs formed by one target anonymous file and multiple real-name files, where the target candidate file pair is the candidate file pair formed by the target anonymous file and the target real-name file.

[0164] Specifically, if the same anonymous file matches multiple real-name files, then it is impossible to determine which real-name file to merge the anonymous file into. Therefore, the multiple real-name files can be sorted in descending order according to the number of days when the trajectory data appears and the number of grid identifiers, and the file pair formed by the real-name file with the sorting serial number 1 and the anonymous file is used as the final target candidate file pair, and the first candidate file pairs other than the target candidate file pair among the multiple first candidate file pairs are deleted.

[0165] It can be seen that through this embodiment, accurate candidate file pairs can be selected, which helps to improve the efficiency and accuracy of anonymous file merging.

[0166] Based on the above embodiments, in one implementation, the method may further include step g1 and step g2:

[0167] Step g1, after obtaining the first candidate file pair and the second candidate file pair, determine whether the first anonymous file and the second anonymous file included in any second candidate file pair appear in different first candidate file pairs;

[0168] Step g2, if the first anonymous file and the second anonymous file appear in different first candidate file pairs, delete the second candidate file pair composed of the first anonymous file and the second anonymous file.

[0169] Specifically, as Figure 3 shown, there will be two anonymous files connected (anonymous file 2 and anonymous file 3), that is, anonymous file 2 and anonymous file 3 form a second candidate file pair, but anonymous file 2 and anonymous file 3 point to different real-name files (real-name file 1 and real-name file 2 respectively), that is, anonymous file 2 and anonymous file 3 appear in different first candidate file pairs. At this time, delete the second candidate file pair composed of the first anonymous file and the second anonymous file, that is, delete the association between anonymous file 2 and anonymous file 3, that is, do not allow the associated anonymous files to point to multiple real-name files.

[0170] It can be seen that through this implementation, accurate candidate file pairs can also be selected, which helps to improve the efficiency and accuracy of anonymous file merging.

[0171] In a second aspect, an embodiment of the present invention provides an anonymous file merging device, as Figure 5 shown, the device includes:

[0172] A trajectory data acquisition module 510, configured to acquire first trajectory data corresponding to a real-name file and second trajectory data corresponding to an anonymous file in a portrait file library, where the trajectory data of a file includes the acquisition time of the portrait image in the file, the longitude and latitude of the area corresponding to the portrait image, the age data and gender data of the portrait in the portrait image, and the number of days when the trajectory data appears;

[0173] An abnormal trajectory data deletion module 520, configured to respectively determine abnormal trajectory data in the first trajectory data and the second trajectory data, and delete the abnormal trajectory data in the first trajectory data and the second trajectory data;

[0174] A grid identifier determination module 530, configured to respectively perform geocoding on the longitude and latitude in the first trajectory data and the longitude and latitude in the second trajectory data to obtain a first grid identifier of the earth grid where the longitude and latitude in the first trajectory data are located, and a second grid identifier of the earth grid where the longitude and latitude in the second trajectory data are located;

[0175] A candidate file pair determination module 540, configured to match a real-name file and an anonymous file, and match different anonymous files based on the first grid identifier and the second grid identifier, the acquisition time of the portrait images in the first trajectory data and the second trajectory data, the age data and gender data in the first trajectory data and the second trajectory data, and the number of days of appearance corresponding to the first trajectory data and the second trajectory data respectively, to obtain a first candidate file pair composed of a real-name file and an anonymous file, and a second candidate file pair composed of different anonymous files;

[0176] A candidate file pair screening module 550, configured to screen the first candidate file pair and the second candidate file pair based on the similarity of portrait image features, to obtain a screened first candidate file pair and a screened second candidate file pair; for any candidate file pair in the screened first candidate file pair and the screened second candidate file pair, the similarity of the portrait image features of the two files included in the candidate file pair is greater than a preset similarity;

[0177] A file merging module 560, configured to merge the screened first candidate file pair and the screened second candidate file pair.

[0178] It can be seen that the present invention is based on big data analysis, conducts trajectory data matching analysis on all files in the portrait file library, improves the comprehensiveness of the analysis, can recall more and more accurate candidate file pairs. Compared with the prior art, the efficiency of anonymous file merging is higher, and the accuracy of anonymous file merging is relatively high, which can effectively solve the problem of multiple files for one person in the portrait file library, improve the actual combat effect of subsequent applications based on portrait file trajectories, and provide more powerful support for security work.

[0179] In a third aspect, an embodiment of the present invention provides an electronic device 600, as Figure 6 shown, including:

[0180] At least one processor 601;

[0181] A memory 602 for storing instructions executable by the at least one processor;

[0182] Wherein, the at least one processor is configured to execute the instructions to implement the method described in the first aspect.

[0183] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method described in the first aspect.

[0184] Fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program which, when executed by a processor, implements the method described in the first aspect.

[0185] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.

Claims

1. A method for merging anonymous files, characterized in that: The method comprises: Obtaining first trajectory data corresponding to the real-name archive and second trajectory data corresponding to the anonymous archive in the portrait archive library, wherein the trajectory data of one archive includes the collection time of the portrait image in the archive, the longitude and latitude of the area corresponding to the portrait image, the age data and gender data of the portrait in the portrait image, and the number of days the trajectory data appears; respectively determining abnormal trajectory data in the first trajectory data and the second trajectory data, and deleting the abnormal trajectory data in the first trajectory data and the second trajectory data; geocoding the longitude and latitude in the first trajectory data and the longitude and latitude in the second trajectory data respectively to obtain a first grid identifier of the earth grid where the longitude and latitude in the first trajectory data are located, and a second grid identifier of the earth grid where the longitude and latitude in the second trajectory data are located; Based on the first grid identifier and the second grid identifier, the collection time of the portrait image in the first trajectory data and the second trajectory data, the age data and the gender data in the first trajectory data and the second trajectory data, and the number of days corresponding to the first trajectory data and the second trajectory data, the real-name archive and the anonymous archive are matched, and different anonymous archives are matched to obtain a first candidate archive pair consisting of the real-name archive and the anonymous archive, and a second candidate archive pair consisting of different anonymous archives; The first candidate profile pair and the second candidate profile pair are screened based on similarity of portrait image features to obtain a screened first candidate profile pair and a screened second candidate profile pair; for any candidate profile pair of the screened first candidate profile pair and the screened second candidate profile pair, the similarity of portrait image features of the two profiles included in the candidate profile pair is greater than a preset similarity; The screened first candidate file pair and the screened second candidate file pair are merged.

2. The method according to claim 1, characterized in that The trajectory data also includes the storage time of the portrait image and the device identification of the portrait image acquisition device, and the determining of abnormal trajectory data in the first trajectory data and the second trajectory data respectively, and deleting the abnormal trajectory data in the first trajectory data and the second trajectory data includes: For any one of the first trajectory data and the second trajectory data, if the time difference between the collection time of the portrait image in the trajectory data and the storage time of the portrait image is greater than the first preset time difference, determining that the collection time of the portrait image is abnormal, and deleting the collection time of the portrait image from the trajectory data; The longitude and latitude in the trajectory data are compared with the longitude and latitude range of the city where the portrait image acquisition device is located. If the longitude and latitude in the trajectory data are not within the longitude and latitude range, it is determined that the longitude and latitude in the trajectory data are abnormal, and the abnormal longitude and latitude are deleted from the trajectory data.

3. The method according to claim 2, characterized in that The determining the abnormal trajectory data in the first trajectory data and the second trajectory data respectively, and deleting the abnormal trajectory data in the first trajectory data and the second trajectory data, comprises: For any one of the first trajectory data and the second trajectory data, comparing multiple collection times corresponding to the same device identifier in the trajectory data; If the time difference between the multiple collection times is less than the second preset time difference, the storage times corresponding to the multiple collection times are compared, and the target collection time corresponding to the latest storage time is retained, and the collection times other than the target collection time in the multiple collection times are deleted.

4. The method according to claim 1, characterized in that: The method further comprises: For each second trajectory data corresponding to the anonymous file in the portrait archive, counting the number of trajectories formed by the second trajectory data, the number of days on which the second trajectory data appears, and the number of device identifiers of the portrait image acquisition devices included in the second trajectory data; Determining whether the number of trajectories is within a first preset number range, whether the number of days is within a preset number of days, and whether the number of device identifiers is within a second preset number range; If the number of trajectories, the number of days and the number of target device identifiers do not meet the preset conditions, delete the second trajectory data, wherein the preset conditions are that the number of trajectories is within the first preset number range, the number of days is within the preset number of days, and the number of device identifiers is within the second preset number range.

5. The method according to any one of claims 1 to 4, characterized in that: The method of obtaining a first candidate profile pair consisting of a real-name profile and an anonymous profile based on the first grid identifier and the second grid identifier, the collection time of the portrait image in the first trajectory data and the second trajectory data, the age data and the gender data in the first trajectory data and the second trajectory data, and the appearance days corresponding to the first trajectory data and the second trajectory data, respectively, includes: For any real-name file and any anonymous file, matching a first grid identifier in the first trajectory data corresponding to the real-name file with a second grid identifier in the second trajectory data corresponding to the anonymous file, matching a collection time in the first trajectory data with a collection time in the second trajectory data, and matching gender data in the first trajectory data with gender data in the second trajectory data; If the matching condition is met, the real-name profile and the anonymous profile are determined as a pre-candidate profile pair; the matching condition is that the first grid identifier matches the second grid identifier, the collection time matches, and the gender data matches; For the real-name profile and the anonymous profile included in the pre-candidate profile pair, calculating an age difference between first age data and second age data, wherein the first age data is age data in first trajectory data corresponding to the real-name profile, and the second age data is age data in second trajectory data corresponding to the anonymous profile; If the age difference is less than the preset age difference, counting the number of days that the first trajectory data corresponding to the real-name file appears and the number of first grid identifiers in the first trajectory data, and the number of days that the second trajectory data corresponding to the anonymous file appears and the number of second grid identifiers in the second trajectory data; If the number of days on which the first trajectory data appears and the number of days on which the second trajectory data appears are both within a preset number of days, and the number of the first grid identifiers and the number of the second grid identifiers are both within a third preset number range, the pre-candidate archive pair is determined as a first candidate archive pair.

6. The method according to any one of claim 5, characterized in that: The method further comprises: After obtaining the first candidate profile pairs, determining whether there is a target anonymous profile matching multiple real-name profiles among the plurality of the first candidate profile pairs; If there is a case where a target anonymous profile matches multiple real-name profiles, the multiple real-name profiles are sorted in descending order of the number of first grid identifiers in the first trajectory data corresponding to the multiple real-name profiles and / or in descending order of the number of days on which the first trajectory data appear, and the real-name profile with a sorting number of 1 is determined as the target real-name profile; Among a plurality of first candidate profile pairs consisting of a target anonymous profile and a plurality of real-name profiles, first candidate profile pairs other than the target candidate profile pair are deleted, wherein the target candidate profile pair is a candidate profile pair consisting of the target anonymous profile and the target real-name profile.

7. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: After obtaining the first candidate profile pair and the second candidate profile pair, determining whether the first anonymous profile and the second anonymous profile included in any second candidate profile pair appear in different first candidate profile pairs; If the first anonymous profile and the second anonymous profile appear in different first candidate profile pairs, a second candidate profile pair consisting of the first anonymous profile and the second anonymous profile is deleted.

8. The method according to any one of claims 1 to 4, characterized in that: Merging the first candidate file pair after screening and the second candidate file pair after screening, including: For any first candidate profile pair after the screening, updating the profile identifier of the anonymous profile in the first candidate profile pair to the profile identifier of the real-name profile in the first candidate profile pair, updating the profile information of the anonymous profile to the profile information of the real-name profile, and deleting the profile information of the anonymous profile; For any second candidate profile pair after screening, the anonymous profile with a larger number of trajectories composed of trajectory data in the second candidate profile pair is used as the main profile, the other anonymous profile in the second candidate profile pair is used as the secondary profile, the profile identifier of the secondary profile is updated to the profile identifier of the main profile, the profile information of the secondary profile is updated to the profile information of the main profile, and the profile information of the secondary profile is deleted.

9. An anonymous file merging device, characterized in that: The device comprises: A trajectory data acquisition module, used to acquire first trajectory data corresponding to the real-name archive and second trajectory data corresponding to the anonymous archive in the portrait archive library, wherein the trajectory data of one archive includes the acquisition time of the portrait image in the archive, the longitude and latitude of the area corresponding to the portrait image, the age data and gender data of the portrait in the portrait image, and the number of days the trajectory data appears; an abnormal trajectory data deleting module, used to respectively determine the abnormal trajectory data in the first trajectory data and the second trajectory data, and delete the abnormal trajectory data in the first trajectory data and the second trajectory data; a grid identifier determination module, configured to geocode the longitude and latitude in the first trajectory data and the longitude and latitude in the second trajectory data, respectively, to obtain a first grid identifier of the earth grid where the longitude and latitude in the first trajectory data are located, and a second grid identifier of the earth grid where the longitude and latitude in the second trajectory data are located; a candidate profile pair determination module, configured to match the real-name profile and the anonymous profile, and to match different anonymous profiles based on the first grid identifier and the second grid identifier, the collection time of the portrait image in the first trajectory data and the second trajectory data, the age data and the gender data in the first trajectory data and the second trajectory data, and the number of days of appearance corresponding to the first trajectory data and the second trajectory data, respectively, to obtain a first candidate profile pair consisting of the real-name profile and the anonymous profile, and a second candidate profile pair consisting of different anonymous profiles; a candidate profile pair screening module, configured to screen the first candidate profile pair and the second candidate profile pair based on similarity of portrait image features to obtain a screened first candidate profile pair and a screened second candidate profile pair; for any candidate profile pair of the screened first candidate profile pair and the screened second candidate profile pair, the similarity of portrait image features of the two profiles included in the candidate profile pair is greater than a preset similarity; The file merging module is used to merge the first candidate file pair after screening and the second candidate file pair after screening.

10. An electronic device, characterized in that: include: at least one processor; a memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method according to any one of claims 1-8.