File merging method, device, electronic device and storage medium
By calculating the similarity, density and distance values of archived data of public places, and determining the clustering center point as the cover image, the problems of high randomness, low accuracy and representativeness of the cover image in the prior art are solved, and higher accuracy and representativeness are achieved.
Patent Information
- Application Number
- CN202111678208.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The prior art is more random when generating covers of personnel files in public places, resulting in lower accuracy and representativeness.
By calculating the similarity degree of any two archive data in the archive data set, a similarity set is obtained; determining the density value set based on the similarity set; calculating the distance value set; determining the archive data corresponding to the density value and distance value that meets the preset conditions is as the cluster center point; and using the archive data corresponding to the cluster center point as the cover image.
It effectively avoids the randomness of the cover image and improves the accuracy and representativeness of the cover image.
Smart Images

Figure CN114443875B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method, device, electronic device and storage medium for file merging. Background Art
[0002] At present, the combination of surveillance cameras and face recognition technology is used to monitor people in various public places, which improves the security of the city. While monitoring people in public places, a personnel file can be established for each person to facilitate personnel management. Currently, mainly the captured data is stratified by quality, and the pictures with better capture quality are compared with the existing file covers. If they are not similar, they are used as the covers of new files. However, the generated file covers have a large randomness, resulting in relatively low accuracy and representativeness. Summary of the Invention
[0003] In a first aspect, the main object of the present invention is to provide a method for file merging, including:
[0004] Calculating the similarity between any two archived data in the archived data set to obtain a similarity set, where the archived data set is an archived data set composed of multiple archived data of the same person;
[0005] Determining a density value set between the multiple archived data according to the similarity set;
[0006] Calculating a distance value set of the multiple archived data according to the density value set and the similarity set;
[0007] Determining multiple archived data corresponding to the density values and distance values that meet the preset conditions as multiple clustering centers;
[0008] Taking the archived data corresponding to the multiple clustering centers as the cover image.
[0009] Optionally, the determining a density value set between the multiple archived data according to the similarity set includes:
[0010] Determining the similarity threshold of each archived data in the similarity set;
[0011] Filtering the similarity set according to the similarity threshold of each archived data;
[0012] Calculating the filtered similarity set according to a predetermined algorithm to obtain the density value set.
[0013] Optionally, the calculating the filtered similarity set according to a predetermined algorithm to obtain the density value set includes:
[0014] Determine the corresponding similarity according to the filtered similarity set;
[0015] Sum up the similarities to obtain multiple density values corresponding to each archived data;
[0016] Sort the density values corresponding to multiple archived data to obtain the density value set.
[0017] Optionally, the calculating the distance value set of multiple archived data according to the density value set and the similarity set includes:
[0018] Sort the archived data set according to the density value set of each archived data, and filter the sorted archived data set;
[0019] Perform distance calculation on the filtered archived data set to obtain the distance value set.
[0020] Optionally, the sorting the archived data set according to the density value set of each archived data and filtering the sorted archived data set includes:
[0021] Sort the archived data set from large to small according to the density value of each archived data;
[0022] For each archived data,
[0023] Determine multiple reference archived data whose density values are greater than the density value of the archived data itself;
[0024] Determine the target archived data with the smallest similarity to the archived data among multiple reference archived data, so as to perform distance calculation according to the target archived data and the archived data.
[0025] Optionally, the method further includes:
[0026] Merge the archived data set according to the multiple cluster centers to obtain a cluster set; wherein, each cluster includes a corresponding cover image;
[0027] Determine the cover image of each cluster according to the cluster set;
[0028] When multiple clusters and the cover images satisfy a preset relationship, merge multiple clusters.
[0029] Optionally, the merging multiple clusters when the cluster set and the cover images satisfy a preset relationship includes:
[0030] Judge whether the number of multiple clusters is greater than the number of cover images;
[0031] When the number of multiple said clustering clusters is greater than the number of said cover images, determine the similarity between the clustering center points of pairwise clustering clusters;
[0032] In the case where the similarity between pairwise said clustering center points is greater than a preset similarity, merge the corresponding pairwise clustering clusters.
[0033] In a second aspect, an embodiment of the present invention provides an archive merging device, including:
[0034] A first calculation module, configured to calculate the similarity between any two archived data in the archived data set to obtain a similarity set, where the archived data set is an archived data set composed of multiple archived data of the same person;
[0035] A first determination module, configured to determine a density value set between multiple said archived data according to the similarity set;
[0036] A second calculation module, configured to calculate a distance value set of multiple said archived data according to the density value set and the similarity set;
[0037] A second determination module, configured to determine multiple archived data corresponding to density values and distance values that meet preset conditions as multiple clustering center points;
[0038] A third determination module, configured to use the archived data corresponding to the multiple clustering center points as cover images.
[0039] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned archive merging method are implemented.
[0040] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned archive merging method are implemented.
[0041] The above solution of the present invention has at least the following beneficial effects:
[0042] The file merging method provided by the present invention first calculates the similarity between any two archived data in the archived dataset, and obtains a similarity set. The archived dataset is an archived dataset composed of multiple archived data of the same person; according to the similarity set, a density value set between the multiple archived data is determined; according to the density value set and the similarity set, a distance value set of the multiple archived data is calculated; the multiple archived data corresponding to the density value and the distance value that meet the preset conditions are determined as multiple clustering center points; the archived data corresponding to the multiple clustering center points are used as the cover image. This avoids the large randomness of the generated cover image and improves the accuracy and representativeness of the cover image. Description of the Drawings
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the structures shown in these drawings.
[0044] Figure 1 It is a schematic diagram of the overall process of the file merging method provided by the embodiment of the present invention;
[0045] Figure 2 It is a schematic diagram of the process of step S20 provided by the embodiment of the present invention;
[0046] Figure 3 It is a specific process schematic diagram of step S23 provided by the embodiment of the present invention;
[0047] Figure 4 It is a specific process schematic diagram of step S30 provided by the embodiment of the present invention;
[0048] Figure 5 It is another process schematic diagram of step S32 provided by the embodiment of the present invention;
[0049] Figure 6 It is another process schematic diagram of the file merging method provided by the embodiment of the present invention;
[0050] Figure 7 It is another process schematic diagram of step 32 provided by the embodiment of the present invention;
[0051] Figure 8 It is a structural block diagram of the file merging device provided by the embodiment of the present invention;
[0052] Figure 9 It is a structural block diagram of the electronic device provided by the embodiment of the present invention.
[0053] The realization, functional features and advantages of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0055] The terms "first", "second", "third", etc. in the specification, claims and the above-mentioned drawings of the present invention are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0056] It can be understood that in the specific implementation manners of the present application, when it comes to relevant data such as archive data, image archives, and archived data sets, when the embodiments in the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use and processing of relevant data, as well as the construction, training and use of tools such as databases and police platforms need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0057] First, the solutions of the embodiments of the present application will be introduced by way of example in conjunction with the relevant drawings.
[0058] As Figure 1 shown, a specific embodiment of the present invention provides an archive merging method, including:
[0059] S10. Calculate the similarity between any two archived data in the archived data set to obtain a similarity set; the archived data set is an archived data set composed of multiple archived data of the same person.
[0060] In this embodiment, an archive data set is stored in the database. The archive data set includes an archived data set, and the archived data set may include multiple image archives. Of course, the archive data may also be archive data of other attributes; the archive data may be image archives collected in public places, such as those collected in urban streets, high-speed railway stations, bus stations, airports, etc. The database may be the database of the police platform. By collecting and clustering the image archives of various groups of people, the management of various groups of people can be more convenient.
[0061] Moreover, when calculating the similarity of the archived data sets, the cosine theorem can be used to calculate the similarity between pairwise archived data. The cosine theorem is also known as cosine similarity, which represents the cosine value of the angle between two vectors in a vector space as a measure of the difference between two individuals. The closer the cosine value is to 1, the closer the angle is to 0 degrees, that is, the more similar the two vectors are. In other words, when the cosine value between pairwise archived data is closer to 1, it indicates that the similarity between the pairwise archived data is greater.
[0062] Among them, the archive data set can be aidN{aid1, aid2, aid3...aidN}, and the archived data set in each archive data set aidN can be a n {a1, a2, a3,..a n}, by performing n×n calculations on the archived data set a n in the archive data set aidN, the similarity between pairwise archived data can be calculated. Therefore, the calculated similarity set can be sim ij {sim 12 , sim 13 ....sim ij}, where i, j represent the data a i , a j ; for example, the archived data set includes picture a, picture b, picture c, picture d. After calculating the similarity of picture a, picture b, picture c, picture d through the cosine theorem, the obtained similarity set is sim ab , sim ac , sim ad , sim bc , sim bd , sim cd ; thus, the picture with the greatest similarity in the archived data set can be determined.
[0063] S20. Determine the density value set among multiple archived data according to the similarity set.
[0064] In this embodiment, the density value is expressed as the similarity density of the archived data set, and the density value set contains multiple density values; each archived data can represent multiple corresponding image data, and each image data can be used as a data point. By calculating the density value between each data point corresponding to each cluster, the density value set can be determined; generally speaking, the similarity between each archived data is relatively small, but the similarity between each data point in each archived data is relatively large.
[0065] For example, after archiving the image file of A, the image file of A contains pictures a, b, c, and d. After archiving the image file of B, the image file of B contains pictures e, f, g, and h. After clustering and merging, the image file of A can be used as a cluster, and the image file of B can be used as a cluster. Therefore, multiple density values can be determined for the archived data corresponding to A through the similarities of pictures a, b, c, and d, and multiple density values can also be determined for the archived data corresponding to B through pictures e, f, g, and h.
[0066] As Figure 2 shown, the above method for determining the set of density values between multiple archived data according to the similarity set includes:
[0067] S21. Determine the similarity threshold for each archived data in the similarity set;
[0068] S22. Screen the similarity set according to the similarity threshold of each archived data;
[0069] S23. Calculate the screened similarity set according to a predetermined algorithm to obtain the set of density values.
[0070] In this embodiment, the similarity threshold can be preset. Through the similarity threshold of each archived data, multiple similarities of each archived data can be screened, and after screening, multiple similarities of each archived data can be calculated, thereby determining the set of density values for each archived data.
[0071] Among them, in the similarity set sim ij {sim 12 , sim 13 .... sim ij} of the archived data aidN obtained by the above calculation, the similarity threshold can be determined from the similarity set, and the similarity threshold is set as By the similarity threshold screen multiple similarities, and the screening condition can be That is to say, those with similarities less than or equal to can be filtered out, while those with similarities greater than can be retained. Thus, a screened similarity set can be obtained to calculate the set of density values.
[0072] As Figure 3 shown, the above method for calculating the screened similarity set according to a predetermined algorithm to obtain the set of density values includes:
[0073] S231. Determine the corresponding similarity according to the screened similarity set;
[0074] S232. Sum up the similarities to obtain multiple density values corresponding to each archived data;
[0075] S233. Sort the density values corresponding to multiple archived data to obtain a density value set.
[0076] In this embodiment, by summing up the determined similarities and using the summation result as the density value of the archived data, and after determining the density values of each archived data, sorting the multiple density values from largest to smallest, a density value set can be obtained. The density value set can be expressed as β(sim ij ) ∈ {β1, β2,.. β.β n}.
[0077] Specifically, the predetermined algorithm can be calculated using the following formula:
[0078]
[0079] Among them, β i represents the density value, sim ij is the similarity, is the minimum similarity, x represents the number of similarities. Therefore, the above formula means that when , then x = 1; when less than 0, then x = 0. By summing up the determined similarities, the density value set corresponding to each archived data can be determined.
[0080] For example, the archived data set includes the image files of three people, A, B, and C. The image file of A contains pictures a, b, c, and d. The image file of B contains pictures e, f, g, h, and i. After calculation, the similarity set corresponding to the image file of A is sim ab , sim ac , sim ad , sim bc , sim bd , sim cd ; the similarity set corresponding to the image file of B is sim ef , sim eg , sim eh , sim ei , sim fg , sim fh , sim fi , sim gh , sim gi , sim hi; When calculating the density value of A, when the density value obtained after summing the similarities of A is 4; when calculating the density value of B, when the density value obtained after summing the similarities of B is 9; after determining the density values corresponding to A and B, sort 9 and 4 from large to small, and the obtained density value set is β(sim AB ) ∈ {9, 4}.
[0081] S30. Calculate the distance value set of multiple archived data according to the density value set and the similarity set.
[0082] As Figure 4 shown, the specific implementation manner of the above step S30 includes:
[0083] S31. Sort the archived data set according to the density value set of each archived data, and filter the sorted archived data set;
[0084] S32. Calculate the distance of the filtered archived data set to obtain the distance value set.
[0085] In this embodiment, the distance value can represent the distance between archived data, that is, the distance value between each archived data; when sorting the archived data set, it can be sorted according to the density value set, and the density value set is sorted from large to small. Therefore, the archived data set can be sorted in sequence according to the density value. After sorting, the distance of each archived data can be calculated. By recursively comparing the density values of multiple archived data, and then determining the distance between multiple archived data; for example, in the four archived data A, B, C, and D, the obtained density value set is (600, 500, 400, 300). Therefore, B and A can be recursively compared, C and A, B can be recursively compared, and D and A, B, C can be recursively compared to determine the distance between multiple archived data.
[0086] As Figure 5 shown, the above sorting the archived data set according to the density value set of each archived data and filtering the sorted archived data set includes:
[0087] S311. Sort the archived data set from large to small according to the density value of each archived data;
[0088] S312. For each archived data, determine multiple reference archived data whose density value is greater than the density value of the archived data itself;
[0089] S313. Determine the target archived data with the smallest similarity to the archived data among the multiple reference archived data, so as to calculate the distance according to the target archived data and the archived data.
[0090] Among them, when recursively comparing the archived data, the reference archived data with each density value greater than itself can be recursively compared; optionally, for the archived data corresponding to the maximum density value in the density value set, the distance value between it and the target archived data with the maximum distance from it can be calculated; therefore, after determining the reference archived data, it can be determined whether the similarity between the archived data and its corresponding target archived data is the smallest. If the similarity between the two is the smallest, it can be determined as the target archived data, and the distance between the target archived data and the archived data can be calculated, and then the distance value can be determined. After calculating the distance values of all archived data, the distance value set can be determined; specifically, the distance value can be calculated using the following formula:
[0091]
[0092] Among them, γ i represents the distance value. After determining the density value set and the distance value set, the density value set and the distance value set can be represented by two-dimensional vector coordinates. For example, the x-axis is the density value and the y-axis is the distance value. Thus, in the two-dimensional vector coordinates, it can be determined whether the values of each density value and distance value are relatively large, and then the corresponding clustering center point can be determined; that is to say, when the density value and distance value corresponding to the archived data are both relatively large, the archived data can be determined as the clustering center point.
[0093] For example, for the four archived data A, B, C, and D, the obtained density value set is (600, 500, 400, 300). Therefore, B can be compared with A, C can be compared with A and B, D can be compared with A, B, and C, and the obtained distance value set is (0.9, 0.4, 0.7, 0.2), that is, the distance value with the largest distance from A is 0.9, the distance value between B and A is 0.4, the distance value with the smallest similarity after calculating C with A and B respectively is 0.7, and the distance value with the smallest similarity after calculating D with A, B, and C respectively is 0.2.
[0094] S40. Determine multiple archived data corresponding to density values and distance values that meet the preset conditions as multiple clustering center points.
[0095] In this embodiment, the preset condition means that the values of the density value and the distance value are relatively large. That is to say, the density value of each archive data is relatively large, but the distance value from other archive data is relatively large. Thus, the corresponding archived data is determined as the clustering center point. When re-clustering and archiving the archived data, the archived data can be more accurately archived through the clustering center point to improve the accuracy.
[0096] S50. Use the archived data corresponding to the multiple clustering center points as the cover image.
[0097] In this embodiment, the cover image may be a face image. Each archived data may contain multiple image files. After clustering and merging using the clustering centers obtained through the above calculations, the image files of each person may be merged to form one or more clustering clusters. Since the image files corresponding to each clustering center point have a high density and are farther away from other clustering center points, they can be used as representative cover images. When archiving subsequently, different image files can be more accurately clustered and merged through this cover image to improve accuracy.
[0098] For the file merging method provided by the present invention, first, the similarity between any two archived data in the archived data set is calculated to obtain a similarity set. The archived data set is an archived data set composed of multiple archived data of the same person; according to the similarity set, a density value set between multiple archived data is determined; according to the density value set and the similarity set, a distance value set of multiple archived data is calculated; multiple archived data corresponding to density values and distance values that meet the preset conditions are determined as multiple clustering center points; the archived data corresponding to the multiple clustering center points is used as the cover image. This avoids a large degree of randomness in the generated cover image and improves the accuracy and representativeness of the cover image.
[0099] As Figure 6 shown, the file merging method provided by the embodiment of the present invention further includes:
[0100] 30. According to multiple clustering center points, the archived data set is merged again to obtain a clustering cluster set; wherein, each clustering cluster includes the corresponding cover image;
[0101] 31. Determine the cover image of each clustering cluster according to the clustering cluster set;
[0102] 32. When multiple clustering clusters and the cover image meet the preset relationship, the multiple clustering clusters are merged.
[0103] In this embodiment, the preset relationship means that the number of clustering clusters is greater than the number of cover images. Each clustering cluster can represent the image files of a certain person, and the image files of one person can also form multiple clustering clusters. Multiple cover images can be determined among the multiple clustering clusters. To reduce the number of clustering clusters corresponding to the same file, multiple clustering clusters can be merged, and the corresponding cover image is re-determined after the merger; for example, the image files of A form 10 clustering clusters, and there are only 5 cover images in the image files of A. By merging the 10 clustering clusters into 5, the number of clustering clusters formed by the image files of A is reduced.
[0104] As Figure 7As shown above, when the cluster set and the cover image satisfy the preset relationship, merging multiple clusters to determine the corresponding cover image includes:
[0105] 321. Determine whether the number of multiple clusters is greater than the number of cover images;
[0106] 322. When the number of multiple clusters is greater than the number of cover images, determine the similarity between the cluster centers of pairwise clusters;
[0107] 323. In the case where the similarity between pairwise cluster centers is greater than the preset similarity, merge the corresponding pairwise clusters.
[0108] In this embodiment, the preset similarity can be the similarity preset by the user. When merging multiple clusters, by judging the similarity between pairwise clusters, when the similarity is greater than the preset similarity, the pairwise clusters can be merged; it can be understood that through the above-determined cluster centers, when judging whether the similarity between pairwise clusters is greater than the preset similarity, the similarity can be calculated through the cluster centers corresponding to the pairwise clusters. Among them, the preset similarity can be set to ο sim , and the similarity between pairwise cluster centers can be sim ij , therefore, when determining whether the similarity between pairwise clusters is greater than the preset similarity, through sim ij -ο sim > 0 for calculation, so as to determine whether the similarity between pairwise clusters is greater than the preset similarity to determine whether to perform cluster merging. It can be understood that when the number of clusters is less than the number of cover images, no merging operation needs to be performed between the clusters.
[0109] For example, in Party A's image file, there are clusters A, B, C, and D, and the cover images corresponding to Party A include face pictures e and f; in Party B's image file, there are clusters G, H, and I, and the cover images corresponding to Party B include face pictures 1, 2, 3, and 4; therefore, the number of clusters corresponding to Party A is greater than the number of its cover images, and clusters A, B, C, and D can be merged. By calculating the pairwise similarities between the cluster centers of each cluster, 6 similarities can be obtained. Calculating the 6 similarities and the preset similarity can determine the clusters that can be merged, making the merged clusters more representative; since the number of clusters corresponding to Party B is less than the number of its cover images, therefore, the clusters of Party B can not be merged. When there are more clusters formed subsequently, cluster merging can be performed again.
[0110] Such asFigure 8 As shown in the figure, an embodiment of the present invention provides an archive merging device 10, including:
[0111] A first calculation module 11, configured to calculate the similarity between any two archived data in the archived data set to obtain a similarity set, where the archived data set is an archived data set composed of multiple archived data of the same person;
[0112] A first determination module 12, configured to determine a density value set between multiple archived data according to the similarity set;
[0113] A second calculation module 13, configured to calculate a distance value set of multiple archived data according to the density value set and the similarity set;
[0114] A second determination module 14, configured to determine multiple archived data corresponding to density values and distance values that meet preset conditions as multiple clustering center points;
[0115] A third determination module 15, configured to use the archived data corresponding to multiple clustering center points as the cover image.
[0116] For the archive merging device 10 provided by the present invention, first, the similarity between any two archived data in the archived data set is calculated to obtain a similarity set, where the archived data set is an archived data set composed of multiple archived data of the same person; according to the similarity set, a density value set between multiple archived data is determined; according to the density value set and the similarity set, a distance value set of multiple archived data is calculated; multiple archived data corresponding to density values and distance values that meet preset conditions are determined as multiple clustering center points; the archived data corresponding to multiple clustering center points is used as the cover image. Thereby, the randomness of the generated cover image is avoided, and the accuracy and representativeness of the cover image are improved.
[0117] It should be noted that the archive merging device 10 provided by the specific embodiment of the present invention is a device corresponding to the above archive merging method. All embodiments of the above archive merging method are applicable to the archive merging device 10. There are corresponding modules in the embodiments of the above archive merging device 10 corresponding to the steps in the above archive merging method, and the same or similar beneficial effects can be achieved. To avoid excessive repetition, each module in the archive merging device 2 will not be described in detail here.
[0118] As Figure 9 shown, a specific embodiment of the present invention also provides an electronic device 20, including a memory 202, a processor 201, and a computer program stored in the memory 202 and executable on the processor 201. When the processor 201 executes the computer program, the steps of the above archive merging method are implemented.
[0119] Specifically, the processor 201 is configured to call a computer program stored in the memory 202 and execute the following steps:
[0120] Calculate the similarity between any two archived data in the archived data set to obtain a similarity set, where the archived data set is an archived data set composed of multiple archived data of the same person;
[0121] Determine a density value set among the multiple archived data according to the similarity set;
[0122] Calculate a distance value set of the multiple archived data according to the density value set and the similarity set;
[0123] Determine multiple archived data corresponding to the density values and distance values that meet the preset conditions as multiple clustering centers;
[0124] Use the archived data corresponding to the multiple clustering centers as the cover image.
[0125] Optionally, the determination of the density value set among the multiple archived data by the processor 201 includes:
[0126] Determine the similarity threshold of each archived data in the similarity set;
[0127] Filter the similarity set according to the similarity threshold of each archived data;
[0128] Calculate the filtered similarity set according to a predetermined algorithm to obtain a density value set.
[0129] Optionally, the calculation of the density value set by the processor 201 according to the predetermined algorithm for the filtered similarity set includes:
[0130] Determine the corresponding similarity according to the filtered similarity set;
[0131] Sum the similarities to obtain multiple density values corresponding to each archived data;
[0132] Sort the density values corresponding to the multiple archived data to obtain a density value set.
[0133] Optionally, the calculation of the distance value set of the multiple archived data by the processor 201 according to the density value set and the similarity set includes:
[0134] Sort the archived data set according to the density value set of each archived data, and filter the sorted archived data set;
[0135] Calculate the distance of the filtered archived data set to obtain a distance value set.
[0136] Optionally, the steps performed by the processor 201 to sort the archived data set according to the density value set of each archived data and screen the sorted archived data set include:
[0137] Sort the archived data set from largest to smallest according to the density value of each archived data;
[0138] For each archived data, determine multiple reference archived data whose density values are greater than the density value of the archived data itself;
[0139] Determine the target archived data with the smallest similarity to the archived data among the multiple reference archived data, so as to calculate the distance according to the target archived data and the archived data.
[0140] Optionally, the method performed by the processor 201 further includes:
[0141] According to multiple clustering center points, re-combine the archived data set to obtain a set of clustering clusters; wherein, each clustering cluster includes a corresponding cover image;
[0142] Determine the cover image of each clustering cluster according to the set of clustering clusters;
[0143] When the multiple clustering clusters and the cover images satisfy a preset relationship, merge the multiple clustering clusters.
[0144] Optionally, when the clustering cluster set and the cover image satisfy a preset relationship, the steps performed by the processor 201 to merge the multiple clustering clusters include:
[0145] Judge whether the number of multiple clustering clusters is greater than the number of cover images;
[0146] When the number of multiple clustering clusters is greater than the number of cover images, determine the similarity between the clustering center points of two-by-two clustering clusters;
[0147] In the case where the similarity between two-by-two clustering center points is greater than the preset similarity, merge the corresponding two-by-two clustering clusters.
[0148] That is, in the specific embodiment of the present invention, when the processor 201 of the electronic device 20 executes the computer program, the steps of the above-mentioned file merging method are implemented, thereby avoiding the relatively large randomness of the generated cover images and improving the accuracy and representativeness of the cover images.
[0149] It should be noted that since the steps of the above-mentioned file merging method are implemented when the processor 201 of the electronic device 20 executes the computer program, all embodiments of the above-mentioned file merging method are applicable to the electronic device 20 and can achieve the same or similar beneficial effects.
[0150] In the embodiments of the present invention, the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements each process of the file merging method or the application-side file merging method provided in the embodiments of the present invention, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0151] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0152] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0153] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural transformation made by using the description of the present invention and the content of the drawings under the concept of the present invention, or any direct / indirect application in other related technical fields is included in the patent protection scope of the present invention.
Claims
1. A file merging method, characterized in that, Including: Calculating the similarity between any two archived data in the archived dataset to obtain a similarity set, where the archived dataset is an archived dataset composed of multiple archived data of the same person; Determining a density value set among the multiple archived data according to the similarity set; Calculating a distance value set of the multiple archived data according to the density value set and the similarity set; Determining multiple archived data corresponding to the density value and distance value that meet the preset conditions as multiple clustering center points; Using the archived data corresponding to the multiple clustering center points as the cover image; The determining a density value set among the multiple archived data according to the similarity set includes: Determining the similarity threshold of each archived data in the similarity set; Filtering the similarity set according to the similarity threshold of each archived data; Calculating the filtered similarity set according to a predetermined algorithm to obtain the density value set; The calculating the filtered similarity set according to a predetermined algorithm to obtain the density value set includes: Determining the corresponding similarity according to the filtered similarity set; Summing the similarities to obtain multiple density values corresponding to the multiple archived data; Sorting the density values corresponding to the multiple archived data to obtain the density value set.
2. The file merging method according to claim 1, characterized in that, The calculating a distance value set of the multiple archived data according to the density value set and the similarity set includes: Sorting the archived dataset according to the density value set and filtering the sorted archived dataset; Performing distance calculation on the filtered archived dataset to obtain the distance value set.
3. The file merging method according to claim 2, characterized in that, The sorting the archived dataset according to the density value set and filtering the sorted archived dataset includes: Sorting the archived dataset from largest to smallest according to the density value of each archived data; For each archived data, determining multiple reference archived data whose density value is greater than the density value of the archived data itself; Determining the target archived data with the smallest similarity to the archived data among the multiple reference archived data, so as to perform distance calculation according to the target archived data and the archived data.
4. The file merging method according to claim 1, characterized in that, The method further includes: Merging the archived dataset according to the multiple clustering center points to obtain a clustering cluster set; where each clustering cluster includes the corresponding cover image; Determining the cover image of each clustering cluster according to the clustering cluster set; Merging the multiple clustering clusters when the multiple clustering clusters and the cover image satisfy a preset relationship.
5. The file merging method according to claim 4, characterized in that, The merging the multiple clustering clusters when the multiple clustering clusters and the cover image satisfy a preset relationship includes: Judging whether the number of the multiple clustering clusters is greater than the number of the cover images; When the number of the multiple clustering clusters is greater than the number of the cover images, determining the similarity between the clustering center points of two-by-two clustering clusters; in the case where the similarity between the two-by-two clustering center points is greater than the preset similarity, merging the corresponding two-by-two clustering clusters.
6. A file merging device, characterized in that, Including: A first calculation module, configured to calculate the similarity between any two archived data in the archived data set to obtain a similarity set, where the archived data set is an archived data set composed of multiple archived data of the same person; A first determination module, configured to determine a density value set among the multiple archived data according to the similarity set; A second calculation module, configured to calculate a distance value set of the multiple archived data according to the density value set and the similarity set; A second determination module, configured to determine multiple archived data corresponding to density values and distance values that meet preset conditions as multiple cluster centers; A third determination module, configured to use the archived data corresponding to the multiple cluster centers as the cover image; The first determination module is further configured to: Determine the similarity threshold of each archived data in the similarity set; Filter the similarity set according to the similarity threshold of each archived data; Calculate the filtered similarity set according to a predetermined algorithm to obtain the density value set; The first determination module is further configured to: Determine the corresponding similarity according to the filtered similarity set; Sum the similarities to obtain multiple density values corresponding to the multiple archived data; Sort the density values corresponding to the multiple archived data to obtain the density value set.
7. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the file merging method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the file merging method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Archive processing method and device, electronic equipment and computer readable storage medium
CN110059657A
Personnel archiving method and device and electronic equipment
CN112686141A