A method and apparatus for screening associated target profiles

By calculating the spatiotemporal trajectory similarity and other factors between the baseline archive and the archive to be associated, the problem of small archive screening scope and low efficiency in the existing technology is solved, and efficient archive data association is achieved.

CN116612309BActive Publication Date: 2026-05-01ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG DAHUA TECH CO LTD
Filing Date
2023-05-09
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing document screening methods are only applicable to identifying peer targets, with a limited scope and low screening efficiency, and cannot effectively achieve data association of documents.

Method used

By calculating the spatiotemporal trajectory similarity value, peer probability, time decay factor, and similarity value of the starting and ending points of the baseline archive and the archive to be associated, the association parameters are determined, thereby enabling the screening of the archive to be associated and the expansion of the known target archive database.

Benefits of technology

It improves the scope and efficiency of file screening, effectively filtering out files related to known targets and expanding the database of related target files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612309B_ABST
    Figure CN116612309B_ABST
Patent Text Reader

Abstract

The application provides a screening method and device for associated target archives, comprising: obtaining a space-time trajectory corresponding to a reference archive and a space-time trajectory corresponding to a to-be-associated archive, the reference archive being an archive of a known target; calculating a first similarity value of the space-time trajectory corresponding to the to-be-associated archive and the space-time trajectory corresponding to the reference archive; determining a same-trip probability according to the number of same position collection points in the space-time trajectory corresponding to the reference archive and the space-time trajectory corresponding to the to-be-associated archive, the number of position collection points included in the space-time trajectory corresponding to the reference archive, and the number of position collection points included in the space-time trajectory corresponding to the to-be-associated archive; determining an association parameter according to the first similarity value and the same-trip probability; and when the association parameter is greater than or equal to a first preset threshold, classifying the to-be-associated archive into an associated target archive library corresponding to the known target. Through the method, the problems that the current method is only applicable to judging same-trip targets, has a small range of action, and has low archive screening efficiency are solved.
Need to check novelty before this filing date? Find Prior Art

Description

A method and apparatus for filtering related target files Technical Field

[0001] This invention relates to the field of image clustering technology, and in particular to a method and apparatus for screening related target files. Background Technology

[0002] With the rapid development of computer vision, real-time computing, and hardware storage devices, target clustering has become a new research topic. At the same time, the rapid popularization and widespread application of image data acquisition devices such as surveillance cameras have led to a surge in target image samples, making target clustering a critical and challenging problem in the security field.

[0003] The follow-up steps of target clustering are file screening and identity archiving to achieve "data association" of files. However, existing file screening methods have problems such as being only applicable to judging peer targets, having a small scope of application, and low screening efficiency. Summary of the Invention

[0004] This invention provides a method and apparatus for filtering related target files, which solves the problems of current methods that are only applicable to judging peer targets, have a small scope of application, and have low file filtering efficiency.

[0005] In a first aspect, embodiments of the present invention provide a method for filtering associated target files, including:

[0006] Obtain the spatiotemporal trajectory corresponding to the baseline file and the spatiotemporal trajectory corresponding to the file to be associated, wherein the baseline file is the file of a known target;

[0007] Calculate the first similarity value between the spatiotemporal trajectory corresponding to the file to be associated and the spatiotemporal trajectory corresponding to the reference file;

[0008] The probability of being on the same track is determined based on the number of collection points at the same location in the spatiotemporal trajectory corresponding to the baseline file and the spatiotemporal trajectory corresponding to the file to be associated, the number of location collection points included in the spatiotemporal trajectory corresponding to the baseline file, and the number of location collection points included in the spatiotemporal trajectory corresponding to the file to be associated.

[0009] Based on the first similarity value and the peer probability, the association parameter is determined. When the association parameter is greater than or equal to the first preset threshold, the file to be associated is classified into the associated target file library corresponding to the known target.

[0010] According to the above method, a first similarity value is calculated between the spatiotemporal trajectory corresponding to the benchmark file and the spatiotemporal trajectory corresponding to the file to be associated. Simultaneously, the peer probability is obtained based on the number of collection points in the spatiotemporal trajectories of the benchmark file and the file to be associated. When the association parameter determined by the first similarity value and the peer probability is greater than or equal to a first preset threshold, it can be determined that the file to be associated is related to a known target. This method can filter out files to be associated with the benchmark file, and is not only suitable for filtering peer targets, but also has a wide scope and high filtering efficiency.

[0011] Optionally, a second similarity value is determined based on the spatial trajectory corresponding to the baseline file, the spatial trajectory corresponding to the file to be associated, and the time decay factor;

[0012] The time decay factor is determined based on the total time interval, which is determined based on the time interval of the timestamps corresponding to the same location collection points in the spatial trajectory corresponding to the reference file and the spatial trajectory corresponding to the file to be associated.

[0013] When the second similarity value is greater than or equal to the second preset threshold, the file to be associated is classified into the associated target file library corresponding to the known target.

[0014] Using the above method, a second similarity value can be calculated based on the spatial trajectories of the baseline file and the file to be associated, as well as the total time interval between the baseline file and the file to be associated at the same location collection points. Then, when the second similarity value is greater than a second preset threshold, it is determined that the file to be associated is related to the known target. This method, based on the spatial trajectories of the baseline file and the file to be associated, can achieve the screening of files to be associated that have a high degree of similarity to the baseline file in terms of spatial trajectory, with a wide scope and high screening efficiency.

[0015] Optionally, determining a second similarity value based on the spatial trajectory corresponding to the baseline file, the spatial trajectory corresponding to the file to be associated, and the time decay factor includes:

[0016] When the correlation parameter is less than the first preset threshold, the second similarity value is determined based on the spatial trajectory corresponding to the benchmark file, the spatial trajectory corresponding to the file to be correlated, and the time decay factor.

[0017] Optionally, the starting point similarity value is determined based on the starting point coordinates of the spatial trajectory corresponding to the reference file, the starting point coordinates of the spatial trajectory corresponding to the file to be associated, and a first time factor; the ending point similarity value is determined based on the ending point coordinates of the spatial trajectory corresponding to the reference file, the ending point coordinates of the spatial trajectory corresponding to the file to be associated, and a second time factor.

[0018] The first time factor is determined based on the start point time interval, which is determined based on the time interval between the timestamp of the start point coordinates of the spatial trajectory corresponding to the reference file and the timestamp of the start point coordinates of the spatial trajectory corresponding to the file to be associated; the second time factor is determined based on the end point time interval, which is determined based on the time interval between the timestamp of the end point coordinates of the spatial trajectory corresponding to the reference file and the timestamp of the end point coordinates of the spatial trajectory corresponding to the file to be associated.

[0019] When the similarity value at the starting point is greater than or equal to a third preset threshold and the similarity value at the ending point is greater than or equal to a fourth preset threshold, the file to be associated is assigned to the associated target file library corresponding to the known target.

[0020] Using the above method, the starting point similarity value can be calculated based on the starting point coordinates and corresponding time factors of the benchmark file and the file to be associated. Similarly, the ending point similarity value can be calculated based on the ending point coordinates and corresponding time factors of the benchmark file and the file to be associated. Then, when the starting point similarity value is greater than a third preset threshold and the ending point similarity value is greater than a fourth preset threshold, it is determined that the file to be associated is related to the known target. This method can effectively filter files to be associated that have the same or very similar starting and ending points as the benchmark file within a certain time interval, offering a wide scope and high filtering efficiency.

[0021] Optionally, the starting point similarity value is determined based on the starting point coordinates of the spatial trajectory corresponding to the reference file, the starting point coordinates of the spatial trajectory corresponding to the file to be associated, and a first time factor; the ending point similarity value is determined based on the ending point coordinates of the spatial trajectory corresponding to the reference file, the ending point coordinates of the spatial trajectory corresponding to the file to be associated, and a second time factor, including:

[0022] When the second similarity value is less than the second preset threshold, the starting point similarity value is determined based on the starting point coordinates of the spatial trajectory corresponding to the reference file, the starting point coordinates of the spatial trajectory corresponding to the file to be associated, and the first time factor; the ending point similarity value is determined based on the ending point coordinates of the spatial trajectory corresponding to the reference file, the ending point coordinates of the spatial trajectory corresponding to the file to be associated, and the second time factor.

[0023] Optionally, the associated target archive database is compared with the suspected target database to determine the target corresponding to the archive to be associated, wherein the suspected target database includes targets that are associated with the known target.

[0024] Optionally, the suspected target database is determined based on the target location collection points corresponding to the known targets, wherein the target location collection points corresponding to the known targets are the location collection points that the known targets have passed through with a repetition number greater than a preset threshold.

[0025] Secondly, embodiments of the present invention provide a filtering device for associated target files, comprising:

[0026] The transceiver unit is used to acquire the spatiotemporal trajectory corresponding to the reference file and the spatiotemporal trajectory corresponding to the file to be associated, wherein the reference file is a file of a known target;

[0027] The processing unit is configured to calculate a first similarity value between the spatiotemporal trajectory corresponding to the file to be associated and the spatiotemporal trajectory corresponding to the reference file; determine the peer probability based on the number of collection points at the same location in the spatiotemporal trajectory corresponding to the reference file and the spatiotemporal trajectory corresponding to the file to be associated, the number of location collection points included in the spatiotemporal trajectory corresponding to the reference file, and the number of location collection points included in the spatiotemporal trajectory corresponding to the file to be associated; determine the association parameter based on the first similarity value and the peer probability; and when the association parameter is greater than or equal to a first preset threshold, classify the file to be associated as part of the associated target file library corresponding to the known target.

[0028] Optionally, the processing unit is configured to determine a second similarity value based on the spatial trajectory corresponding to the reference file, the spatial trajectory corresponding to the file to be associated, and the time decay factor;

[0029] The time decay factor is determined based on the total time interval, which is determined based on the time interval of the timestamps corresponding to the same location collection points in the spatial trajectory corresponding to the reference file and the spatial trajectory corresponding to the file to be associated.

[0030] When the second similarity value is greater than or equal to the second preset threshold, the file to be associated is classified into the associated target file library corresponding to the known target.

[0031] Optionally, the processing unit is configured to determine the second similarity value based on the spatial trajectory corresponding to the baseline file, the spatial trajectory corresponding to the file to be associated, and the time decay factor when the association parameter is less than the first preset threshold.

[0032] Optionally, the processing unit is further configured to determine a starting point similarity value based on the starting point coordinates of the spatial trajectory corresponding to the reference file, the starting point coordinates of the spatial trajectory corresponding to the file to be associated, and a first time factor; and to determine a termination point similarity value based on the ending point coordinates of the spatial trajectory corresponding to the reference file, the ending point coordinates of the spatial trajectory corresponding to the file to be associated, and a second time factor.

[0033] The first time factor is determined based on the start point time interval, which is determined based on the time interval between the timestamp of the start point coordinates of the spatial trajectory corresponding to the reference file and the timestamp of the start point coordinates of the spatial trajectory corresponding to the file to be associated; the second time factor is determined based on the end point time interval, which is determined based on the time interval between the timestamp of the end point coordinates of the spatial trajectory corresponding to the reference file and the timestamp of the end point coordinates of the spatial trajectory corresponding to the file to be associated.

[0034] When the similarity value at the starting point is greater than or equal to a third preset threshold and the similarity value at the ending point is greater than or equal to a fourth preset threshold, the file to be associated is assigned to the associated target file library corresponding to the known target.

[0035] Optionally, the processing unit is configured to, when the second similarity value is less than the second preset threshold, determine the starting point similarity value based on the starting point coordinates of the spatial trajectory corresponding to the reference file, the starting point coordinates of the spatial trajectory corresponding to the file to be associated, and a first time factor; and determine the ending point similarity value based on the ending point coordinates of the spatial trajectory corresponding to the reference file, the ending point coordinates of the spatial trajectory corresponding to the file to be associated, and a second time factor.

[0036] Optionally, the processing unit is configured to compare the associated target archive database with the suspected target database to determine the target corresponding to the archive to be associated, wherein the suspected target database includes targets that are associated with the known target.

[0037] Optionally, the suspected target database is determined based on the target location collection points corresponding to the known targets, wherein the target location collection points corresponding to the known targets are the location collection points that the known targets have passed through with a repetition number greater than a preset threshold.

[0038] Thirdly, this application also provides an apparatus. This apparatus can perform the above-described method design. The apparatus may be a chip or circuit capable of performing the functions corresponding to the above-described method, or a device including the chip or circuit.

[0039] In one possible implementation, the device includes: a memory for storing computer-executable program code; and a processor coupled to the memory. The program code stored in the memory includes instructions that, when executed by the processor, cause the device or a device equipped with the device to perform any of the methods described above.

[0040] The device may also include a communication interface, which may be a transceiver, or, if the device is a chip or circuit, the communication interface may be the chip's input / output interface, such as input / output pins.

[0041] In one possible design, the device includes corresponding functional units, each used to implement the steps in the above method. The functions can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more units corresponding to the functions described above.

[0042] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on a device, executes the method described in any of the above possible designs.

[0043] Furthermore, the technical effects of any of the implementation methods in the second to fourth aspects can be found in the technical effects of different implementation methods in the first aspect, and will not be repeated here. Attached Figure Description

[0044] Figure 1 is a flowchart illustrating a method for filtering associated target files according to an embodiment of the present invention;

[0045] Figure 2 is a schematic diagram of a process for expanding the associated target archive of known targets according to an embodiment of the present invention;

[0046] Figure 3 is a schematic diagram of a process for further expanding the associated target archive of known targets according to an embodiment of the present invention;

[0047] Figure 4 shows a communication device 400 provided in an embodiment of the present invention;

[0048] Figure 5 shows another communication device 500 provided in this embodiment of the invention. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0050] The application scenarios described in the embodiments of this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. Those skilled in the art will understand that with the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems. In the description of this invention, unless otherwise stated, "multiple" means two or more.

[0051] With the rapid popularization and widespread application of image data acquisition equipment such as surveillance cameras, the number of target image samples has increased exponentially, leading to the rapid development of target clustering. Subsequent stages of target clustering include file screening and identity archiving to achieve "data association" among files. However, existing file screening methods have some problems, such as being only applicable to identifying peers, having a limited scope, and low screening efficiency.

[0052] Based on this, this application proposes a method for screening related target files to solve the problems of current methods, which are only applicable to judging peer targets, have a small scope of application, and have low file screening efficiency.

[0053] As shown in Figure 1, the specific process of the screening method for associated target files proposed in this application is as follows.

[0054] Step 101: Obtain the spatiotemporal trajectory corresponding to the baseline file and the spatiotemporal trajectory corresponding to the file to be associated.

[0055] Specifically, the spatiotemporal trajectory includes a time trajectory and a spatial trajectory. The time trajectory includes multiple timestamps, and the spatial trajectory includes information about the location collection points corresponding to the multiple timestamps.

[0056] For example, the steps of obtaining the spatiotemporal trajectory corresponding to each baseline file and obtaining the spatiotemporal trajectory corresponding to each file to be associated are as follows:

[0057] Step 1: First, specify the time period and geographical area, and acquire image data from location collection points within the specified time period and geographical area. Specifically, image data is obtained by capturing images of the target using cameras at the location collection points.

[0058] Step 2: After obtaining the image data, cluster the image data to obtain the clustering results.

[0059] Each clustering result is a file, and each file corresponds to a target.

[0060] Step 3: Based on the clustering results, select clustering results containing more than one target image and classify them as normal clustering results. Then, based on the normal clustering results, select normal clustering results where the target image was captured at different locations and classify them as usable clustering results.

[0061] Among them, the clustering result of normal files is the normal file archive, and the clustering result of usable files is the usable file archive.

[0062] Step 4: Based on the available file, arrange the images in the available file according to the timestamp of the capture time to obtain the time trajectory corresponding to the available file. Then, arrange the location collection points of the images in the available file according to the timestamp of the capture time to obtain the spatial trajectory corresponding to the available file. Combining the time trajectory and the spatial trajectory yields the spatiotemporal trajectory corresponding to the available file.

[0063] For example, the time trajectory of an available file can be represented as a set {timestamp 1, timestamp 2, timestamp 3, ..., timestamp n}, and the spatial trajectory of an available file can be represented as a set {location collection point 7, location collection point 2, location collection point 4, ..., location collection point n}, where n is a positive integer.

[0064] By repeating steps 1 to 4 above, you can obtain the spatiotemporal trajectories corresponding to multiple available files.

[0065] Step 5: After obtaining the spatiotemporal trajectories corresponding to multiple available files, the baseline files from these available files are placed in the baseline file library, and the files to be associated are placed in the file to be associated library. Available files can be either baseline files or files to be associated. Baseline files are files with known targets. Files to be associated are files with unknown targets, i.e., files whose targets are to be associated. For example, the baseline files may be files of registered users, and the files to be associated may be files of unregistered users.

[0066] Based on this, we can obtain the spatiotemporal trajectory corresponding to each benchmark file in the benchmark archive, as well as the spatiotemporal trajectory corresponding to each file to be associated in the archive to be associated.

[0067] Step 102: Calculate the first similarity value between the spatiotemporal trajectory of the file to be associated and the spatiotemporal trajectory of the reference file.

[0068] After obtaining the spatiotemporal trajectory corresponding to each benchmark file in the benchmark archive and the spatiotemporal trajectory corresponding to each file to be associated in the archive to be associated in step 101, a first similarity value is calculated between the spatiotemporal trajectory corresponding to the file to be associated in the archive to be associated and the spatiotemporal trajectory corresponding to the benchmark file in the benchmark archive. The first similarity value can be calculated using existing algorithms for calculating the similarity between two spatiotemporal trajectories, and this application does not impose any limitations on this.

[0069] Step 103: Determine the peer probability based on the number of identical location collection points in the spatiotemporal trajectory corresponding to the baseline file and the spatiotemporal trajectory corresponding to the file to be associated, the number of location collection points included in the spatiotemporal trajectory corresponding to the baseline file, and the number of location collection points included in the spatiotemporal trajectory corresponding to the file to be associated.

[0070] After calculating the first similarity value in step 102, it is also necessary to determine the peer probability based on the number of identical location collection points in the spatiotemporal trajectory corresponding to the benchmark file and the spatiotemporal trajectory corresponding to the file to be associated, the number of location collection points included in the spatiotemporal trajectory corresponding to the benchmark file, and the number of location collection points included in the spatiotemporal trajectory corresponding to the file to be associated.

[0071] Specifically, the formula for calculating the probability of being in the same row is:

[0072]

[0073] Where n is the number of collection points at the same location in the spatiotemporal trajectory corresponding to the baseline file and the spatiotemporal trajectory corresponding to the file to be associated, n1 is the number of location collection points included in the spatiotemporal trajectory corresponding to the baseline file, and n2 is the number of location collection points included in the spatiotemporal trajectory corresponding to the file to be associated.

[0074] Step 104: Determine the association parameters based on the first similarity value and peer probability. When the association parameters are greater than or equal to the first preset threshold, classify the file to be associated into the associated target file library corresponding to the known target.

[0075] After calculating the first similarity value in step 102 and the peer probability in step 103, when the association parameters determined by the first similarity value and the peer probability are greater than or equal to a first preset threshold, the files to be associated are classified into the associated target file library corresponding to the known target. The associated target file library determined at this time can also be called the first file library. The first preset threshold is determined based on empirical values, and the first file library may include multiple files to be associated.

[0076] Specifically, the formula for calculating the correlation parameter is as follows:

[0077] Association parameter = r1 * first similarity value + r2 * peer probability

[0078] The value of r1 can be between 0.7 and 0.9, which is determined based on empirical values, and the value of r2 can be between 1 and r1. In addition, the value of r1 can also be any other arbitrary value range, which is not limited in this application.

[0079] It is understandable that classifying a file to be associated into the associated target file library corresponding to a known target means that the unknown target corresponding to the file to be associated is regarded as an associated target of the known target, or a target that has a relationship with the known target. Files to be associated that are classified into the associated target file library corresponding to a known target can also be called associated target files.

[0080] For example, if the base file is the file of a registered user, and the file to be associated is the file of an unregistered user, then the file to be associated is placed in the associated user file database corresponding to the known user. In this case, the unregistered user is an associated user of the registered user. The file to be associated that is placed in the associated user file database corresponding to the known user can also be called an associated user file.

[0081] Furthermore, during the screening process of target files to be associated, the associated target file library corresponding to the known target can be expanded based on the spatial trajectory and time decay factor of the baseline file and the file to be associated. That is, the associated target file library corresponding to the known target is further expanded based on step 104 (i.e., the first file library). Alternatively, the files expanded based on the spatial trajectory and time decay factor of the baseline file and the file to be associated can also be called the second file library.

[0082] The process for expanding the associated target archive corresponding to known targets is shown in Figure 2. The expansion of the associated target archive corresponding to known targets can be further determined when the association parameter is less than a first threshold. In this case, for the files to be associated whose association parameter is less than the first threshold, the method shown in Figure 2 can be used to perform the next round of file filtering.

[0083] Step 201: Determine the second similarity value based on the spatial trajectory corresponding to the baseline file, the spatial trajectory corresponding to the file to be associated, and the time decay factor.

[0084] Specifically, the time decay factor corresponding to the file to be associated is determined based on the total time interval, which is determined based on the time interval of the timestamps corresponding to the same location collection points in the spatial trajectory corresponding to the baseline file and the spatial trajectory corresponding to the file to be associated. Among them, the time decay factor corresponding to the file to be associated is negatively correlated with the total time interval, and the value range of the time decay factor is [0, 1].

[0085] Specifically, the formula for calculating the second similarity value is as follows:

[0086]

[0087] Where t is the time decay factor, A i Let C be the coordinates of the data collection point at the i-th position of the baseline file. iLet be the coordinates of the data collection point at the i-th location of the file to be associated; l be the number of data collection points at the same location in the spatial trajectories of the reference file and the file to be associated; n be the total number of images captured at the same location in the spatial trajectories of the reference file and the file to be associated; and m be the total number of images captured at the same location in the spatial trajectories of the reference file and the file to be associated. Specifically, if only one image is captured at each data collection point, then l, n, and m are equal.

[0088] Specifically, the absolute value of the formula for calculating the second similarity value ranges from [0, 1]. Therefore, when the absolute value is the same, the larger the time decay factor, the larger the second similarity value, meaning the higher the similarity between the file to be associated and the reference file; the smaller the time decay factor, the smaller the second similarity value, meaning the lower the similarity between the file to be associated and the reference file.

[0089] For example, if the number of identical location collection points in the spatial trajectory corresponding to the baseline file and the spatial trajectory corresponding to the file to be associated is set to 3, i.e., l = n = m = 3, then the specific formula for calculating the second similarity is:

[0090]

[0091] The time decay factor t is determined by the total time interval consisting of the time interval between the timestamps at the first same location collection point of the reference file and the file to be associated, the time interval between the timestamps at the second same location collection point, and the time interval between the timestamps at the third same location collection point.

[0092] Step 202: When the second similarity value is greater than or equal to the second preset threshold, the file to be associated is classified into the associated target file library corresponding to the known target.

[0093] Specifically, after calculating the second similarity, the second similarity value is compared with a second preset threshold. When the second similarity value is greater than the second preset threshold, the file to be associated can be classified into the associated target file library corresponding to the known target. That is, based on step 104, the associated target file library corresponding to the known target can be further expanded (i.e., the first file library). Alternatively, the expanded associated target scheme library can be called the second file library, which is independent of the first file library. The second preset threshold is determined based on empirical values, and the second file library can include multiple files to be associated.

[0094] Furthermore, during the screening process of target files to be associated, the database of associated target files corresponding to known targets can be further expanded based on the starting point coordinates, ending point coordinates, first time factor, and second time factor of the baseline file and the file to be associated. That is, the database of associated target files corresponding to known targets can be further expanded based on step 104 or step 202. Alternatively, the files expanded based on the starting point coordinates, ending point coordinates, first time factor, and second time factor of the baseline file and the file to be associated can also be called a third database.

[0095] The process for further expanding the associated target archive of known targets is shown in Figure 3. The expansion of the associated target archive of known targets can be further determined when the association parameter is less than a first threshold and the second similarity value is less than a second preset threshold. In this case, for the files to be associated that have an association parameter less than the first threshold and a second similarity value less than the second preset threshold, the method shown in Figure 3 is used to perform the next round of file screening.

[0096] Step 301: Determine the starting point similarity value based on the starting point coordinates of the spatial trajectory corresponding to the reference file, the starting point coordinates of the spatial trajectory corresponding to the file to be associated, and the first time factor; determine the ending point similarity value based on the ending point coordinates of the spatial trajectory corresponding to the reference file, the ending point coordinates of the spatial trajectory corresponding to the file to be associated, and the second time factor.

[0097] Specifically, the first time factor is determined based on the start-point time interval, and the second time factor is determined based on the end-point time interval. The start-point time interval is determined by the time interval between the timestamps of the start-point coordinates of the spatial trajectory corresponding to the reference file and the start-point coordinates of the spatial trajectory corresponding to the file to be associated. Similarly, the end-point time interval is determined by the time interval between the timestamps of the end-point coordinates of the spatial trajectory corresponding to the reference file and the end-point coordinates of the spatial trajectory corresponding to the file to be associated.

[0098] Specifically, the first time factor corresponding to the file to be associated is negatively correlated with the time interval of the starting point, and the second time factor corresponding to the file to be associated is negatively correlated with the time interval of the ending point. Furthermore, the value range of the first time factor is [0, 1], and the value range of the second time factor is also [0, 1].

[0099] Specifically, the formula for calculating the starting point similarity value is as follows:

[0100]

[0101] Where T1 is the first time factor, a1 is the starting point coordinate of the spatial trajectory corresponding to the baseline file, and d1 is the starting point coordinate of the spatial trajectory corresponding to the file to be associated.

[0102] Specifically, the absolute value of the formula for calculating the starting point similarity value ranges from [0, 1]. Therefore, when the absolute value is the same, the larger the first time factor, the larger the starting point similarity value, meaning the higher the similarity between the file to be associated and the reference file; the smaller the first time factor, the smaller the starting point similarity value, meaning the lower the similarity between the file to be associated and the reference file.

[0103] Specifically, the formula for calculating the termination point similarity value is as follows:

[0104]

[0105] Where T2 is the second time factor, a2 is the coordinates of the endpoint of the spatial trajectory corresponding to the baseline file, and d2 is the coordinates of the endpoint of the spatial trajectory corresponding to the file to be associated.

[0106] Specifically, the absolute value of the formula for calculating the termination point similarity value ranges from [0, 1]. Therefore, when the absolute value is the same, the larger the second time factor, the larger the termination point similarity value, meaning the higher the similarity between the file to be associated and the reference file; the smaller the second time factor, the smaller the termination point similarity value, meaning the lower the similarity between the file to be associated and the reference file.

[0107] Step 302: When the similarity value of the starting point is greater than or equal to the third preset threshold and the similarity value of the ending point is greater than or equal to the fourth preset threshold, the file to be associated is classified into the associated target file library corresponding to the known target.

[0108] Specifically, after calculating the starting point similarity, the starting point similarity value is compared with a third preset threshold; after calculating the ending point similarity, the ending point similarity value is compared with a fourth preset threshold. When the starting point similarity value is greater than or equal to the third preset threshold, and the ending point similarity value is greater than or equal to the fourth preset threshold, the file to be associated can be classified into the associated target file library corresponding to the known target, that is, the associated target file library corresponding to the known target is further expanded based on step 104 or step 202. Alternatively, the expanded associated target scheme library at this time can also be called the third file library, that is, a third file library independent of the first file library and the second file library. The third preset threshold and the fourth preset threshold are both determined based on empirical values, and the third file library can include multiple files to be associated.

[0109] Furthermore, after step 104, step 202, or step 302, the target corresponding to the file to be associated is determined by comparing the associated target file database corresponding to the known target with the suspected target database. Before comparing the associated target file database corresponding to the known target with the suspected target database, the files to be associated in the associated target file database corresponding to the known target can be deduplicated.

[0110] For example, after determining the first archive database in step 104, the second archive database in step 202, and the third archive database in step 302, the first, second, and third archive databases are compared with the suspected target databases to determine the targets corresponding to each file to be associated in the first, second, and third archive databases, respectively, as the target results. Furthermore, after obtaining the target results, deduplication is required.

[0111] For example, the suspected target database is determined based on the target location collection points corresponding to known targets. These collection points are those where the known target's traversal frequency exceeds a preset threshold, determined empirically. Specifically, a known target can have multiple spatial trajectories, allowing it to repeatedly traverse the same collection point. After obtaining the target location collection points corresponding to known targets, information on Points of Interest (POIs) within one kilometer of these collection points is acquired, leading to the identification of the target associated with each POI and its inclusion in the suspected target database.

[0112] For example, let's take user A as a known target. User A passes by a certain snack shop on their way to and from get off work, to and from the supermarket, and while shopping, and the number of times they pass by the snack shop exceeds a preset threshold. Therefore, this snack shop is considered a target location collection point for user A. Furthermore, by acquiring POI information within 1 kilometer of this snack shop, we can directly obtain the information of the users included in the POI information. This information of all users included in the POI information is then added to the suspected target database.

[0113] For example, when the target is a user, the suspected target database can also be determined based on the known target's kinship and social relationships. The known target's social relationships may include, but are not limited to, close relationships, shared residence, and shared living / studying locations. Specifically, when obtaining the known target's kinship, if the known target's registered residence is not in the local area, the identity information of the known target's registered co-residents is also obtained.

[0114] Specifically, after initially determining a database of suspected targets based on known kinship relationships, known social relationships, and corresponding target location data collection points, the identity information in the database of suspected targets is deduplicated to form the final database of suspected targets. This database of suspected targets is then linked to an ID card database to form a database of suspected targets containing ID card photos.

[0115] For example, after identifying the suspected target database, a similarity comparison based on image features is performed between the known target's associated target database or any one of the three databases and the suspected target database. If the similarity comparison result is greater than the image feature similarity comparison threshold, the target corresponding to the file to be associated is determined, and the file to be associated is added to the baseline database. The image feature similarity comparison threshold is determined based on empirical values. The image feature similarity comparison threshold can also be adjusted according to the accuracy of target association.

[0116] The following describes, in conjunction with the embodiments shown in Figures 1 to 3, the following scenarios: taking the target as the user, the method shown in Figure 2 is used to expand the associated target archive of the known target when the association parameter is less than the first threshold; and the method shown in Figure 3 is used to further expand the associated target archive of the known target when the association parameter is less than the first threshold and the second similarity value is less than the second preset threshold.

[0117] Assume that the baseline archive includes the baseline file corresponding to Wang, and the archive to be associated includes 10 files to be associated, namely file 1, file 2, file 3, file 4, file 5, file 6, file 7, file 8, file 9, and file 10.

[0118] Wang's suspected target database includes Li and Zhang, who are related to Wang; Lin and Sun, who have social relationships with Wang; and Zhou, who was identified based on the target location collection points corresponding to Wang.

[0119] For example, firstly, assume that the filtering method for the associated target files corresponding to Figure 1 can determine that files 1 and 8 belong to the associated target file database corresponding to Wang. Next, for the remaining 8 files to be associated, according to the process shown in Figure 2, files 3, 4, and 10 can be determined to belong to the associated target file database corresponding to Wang, thus expanding the database. Then, for the remaining 5 files to be associated, according to the process shown in Figure 3, files 6 and 7 can be determined to belong to the associated target file database corresponding to Wang, further expanding the database. Finally, it can be determined that the associated target file database corresponding to Wang includes files 1, 3, 4, 6, 7, 8, and 10.

[0120] Furthermore, each file in the associated target archive can be compared with the suspected target base archive. Assuming that the comparison determines that file 1 belongs to Li, file 3 to Zhang, file 6 to Sun, file 7 to Zhou, and file 10 to Lin, this achieves the simultaneous identification of the targets corresponding to files 1, 3, 6, 7, and 10, and then adds the files corresponding to Li, Zhang, Sun, Zhou, and Lin to the baseline archive.

[0121] Furthermore, based on the processes shown in Figures 1 to 3 and the 10 files to be associated, a first archive database, a second archive database, and a third archive database can be determined respectively. The first archive database is defined as including files 1 and 8; the second archive database includes files 3, 4, 8, and 10; and the third archive database includes files 3, 6, and 7. These three archive databases are then compared with the suspected target database. Since the first and second archive databases contain duplicate files, and the second and third archive databases also contain duplicate files, after obtaining the comparison results of these three archive databases with the suspected target database, it is necessary to remove duplicates from the comparison results of the three archive databases. Finally, the targets corresponding to the same files 1, 3, 6, 7, and 10 as described above are obtained, and the files of Li, Zhang, Lin, Sun, and Zhou are all included in the baseline archive database.

[0122] This method determines a database of associated targets for known targets based on the first similarity value and peer probability between a baseline file and the file to be associated. It then expands this database based on the second similarity value between the baseline file and the file to be associated, and further expands it based on the starting and ending similarity values ​​between the baseline file and the file to be associated. Finally, it determines the target corresponding to the file to be associated based on a comparison between the database of associated targets and a base database of suspected targets, thus adding the file to be associated to the baseline database. This method is suitable for determining multiple associated targets based on known targets, has a wide scope, and is highly efficient.

[0123] The division of units in the embodiments of this invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this invention can be integrated into a single processor, exist as separate physical units, or be integrated into a single unit. The integrated units described above can be implemented in hardware or as software functional units.

[0124] This invention also provides a communication device 400, as shown in FIG4. The communication device 400 includes a processing module 410 and a transceiver module 420.

[0125] The transceiver module 420 may include a receiving module and a sending module. The processing module 410 is used to control and manage the operation of the communication device 400. The transceiver module 420 is used to support communication between the communication device 400 and other devices. Optionally, the communication device 400 may also include a storage module for storing the program code and data of the communication device 400.

[0126] Optionally, each module in the communication device 400 can be implemented by software.

[0127] Optionally, the processing module 410 may be a processor or a controller, and the transceiver module 420 may be a communication interface, transceiver, or transceiver circuit, etc. Here, the communication interface is a general term, and in a specific implementation, the communication interface may include multiple interfaces, and the storage module may be a memory.

[0128] In one possible implementation, the communication device 400 is adapted to a wireless access controller device or a wireless access point device;

[0129] The transceiver module 420 is used to acquire the spatiotemporal trajectory corresponding to the baseline file and the spatiotemporal trajectory corresponding to the file to be associated, wherein the baseline file is the file of a known target;

[0130] The processing module 410 is used to calculate a first similarity value between the spatiotemporal trajectory corresponding to the file to be associated and the spatiotemporal trajectory corresponding to the reference file; determine the peer probability based on the number of collection points at the same location in the spatiotemporal trajectory corresponding to the reference file and the spatiotemporal trajectory corresponding to the file to be associated, the number of location collection points included in the spatiotemporal trajectory corresponding to the reference file, and the number of location collection points included in the spatiotemporal trajectory corresponding to the file to be associated; determine the association parameter based on the first similarity value and the peer probability; and when the association parameter is greater than or equal to a first preset threshold, classify the file to be associated as an associated target file in the known target's associated target file library.

[0131] This invention also provides another communication device 500, which may be a terminal device or a chip system inside a terminal device, as shown in FIG5, including:

[0132] Communication interface 501, memory 502, and processor 503;

[0133] The communication device 500 communicates with other devices through the communication interface 501, such as sending and receiving messages; the memory 502 is used to store program instructions; and the processor 503 is used to call the program instructions stored in the memory 502 and execute them according to the obtained program.

[0134] Processor 503 executes program instructions stored in communication interface 501 and memory 502:

[0135] Obtain the spatiotemporal trajectory corresponding to the baseline file and the spatiotemporal trajectory corresponding to the file to be associated, wherein the baseline file is the file of a known target;

[0136] Calculate the first similarity value between the spatiotemporal trajectory corresponding to the file to be associated and the spatiotemporal trajectory corresponding to the reference file;

[0137] The probability of being on the same track is determined based on the number of collection points at the same location in the spatiotemporal trajectory corresponding to the baseline file and the spatiotemporal trajectory corresponding to the file to be associated, the number of location collection points included in the spatiotemporal trajectory corresponding to the baseline file, and the number of location collection points included in the spatiotemporal trajectory corresponding to the file to be associated.

[0138] Based on the first similarity value and the peer probability, the association parameter is determined. When the association parameter is greater than or equal to the first preset threshold, the file to be associated is classified into the associated target file library corresponding to the known target.

[0139] In this embodiment of the invention, the specific connection medium between the communication interface 501, the memory 502 and the processor 503 is not limited, such as a bus. A bus can be divided into an address bus, a data bus, a control bus, etc.

[0140] In this embodiment of the invention, the processor may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this embodiment of the invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in this embodiment of the invention can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0141] In embodiments of the present invention, the memory can be non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or it can be volatile memory, such as random-access memory (RAM). The memory can also be any other medium capable of carrying or storing desired program code having an instruction or data structure form and accessible by a computer, but is not limited thereto. The memory in embodiments of the present invention can also be a circuit or any other device capable of implementing a storage function for storing program instructions and / or data.

[0142] This invention also provides a computer-readable storage medium including program code. When the program code is run on a computer, the program code is used to cause the computer to perform the steps of the method provided in the above embodiments of this invention.

[0143] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0144] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0145] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0146] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0147] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0148] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for filtering associated target files, characterized in that, The method includes: acquiring the spatiotemporal trajectory corresponding to a baseline file and the spatiotemporal trajectory corresponding to a file to be associated, wherein the baseline file is a file of a known target; calculating a first similarity value between the spatiotemporal trajectory corresponding to the file to be associated and the spatiotemporal trajectory corresponding to the baseline file; determining a peer probability based on the number of collection points at the same location in the spatiotemporal trajectory corresponding to the baseline file and the spatiotemporal trajectory corresponding to the file to be associated, the number of location collection points included in the spatiotemporal trajectory corresponding to the baseline file, and the number of location collection points included in the spatiotemporal trajectory corresponding to the file to be associated; determining an association parameter based on the first similarity value and the peer probability; and when the association parameter is greater than or equal to a first preset threshold, classifying the file to be associated as part of the associated target file library corresponding to the known target.

2. The method as described in claim 1, characterized in that, Also includes: The second similarity value is determined based on the spatial trajectory corresponding to the baseline file, the spatial trajectory corresponding to the file to be associated, and the time decay factor. The time decay factor is determined based on the total time interval, which is determined based on the time interval of the timestamps corresponding to the same location collection points in the spatial trajectory corresponding to the baseline file and the spatial trajectory corresponding to the file to be associated; when the second similarity value is greater than or equal to the second preset threshold, the file to be associated is classified into the associated target file library corresponding to the known target.

3. The method as described in claim 2, characterized in that, Determining a second similarity value based on the spatial trajectory corresponding to the benchmark file, the spatial trajectory corresponding to the file to be associated, and the time decay factor includes: when the association parameter is less than the first preset threshold, determining the second similarity value based on the spatial trajectory corresponding to the benchmark file, the spatial trajectory corresponding to the file to be associated, and the time decay factor.

4. The method as described in claim 1, characterized in that, Also includes: The starting point similarity value is determined based on the starting point coordinates of the spatial trajectory corresponding to the benchmark file, the starting point coordinates of the spatial trajectory corresponding to the file to be associated, and the first time factor. The termination point similarity value is determined based on the termination point coordinates of the spatial trajectory corresponding to the baseline file, the termination point coordinates of the spatial trajectory corresponding to the file to be associated, and a second time factor. The first time factor is determined based on the start point time interval, which is determined based on the time interval between the timestamps of the start point coordinates of the spatial trajectory corresponding to the baseline file and the start point coordinates of the spatial trajectory corresponding to the file to be associated. The second time factor is determined based on the termination point time interval, which is determined based on the time interval between the timestamps of the termination point coordinates of the spatial trajectory corresponding to the baseline file and the termination point coordinates of the spatial trajectory corresponding to the file to be associated. When the start point similarity value is greater than or equal to a third preset threshold and the termination point similarity value is greater than or equal to a fourth preset threshold, the file to be associated is assigned to the associated target file library corresponding to the known target.

5. The method as described in claim 4, characterized in that, The starting point similarity value is determined based on the starting point coordinates of the spatial trajectory corresponding to the benchmark file, the starting point coordinates of the spatial trajectory corresponding to the file to be associated, and the first time factor. The similarity value of the termination point is determined based on the coordinates of the termination point of the spatial trajectory corresponding to the benchmark file, the coordinates of the termination point of the spatial trajectory corresponding to the file to be associated, and the second time factor, including: when the second similarity value is less than a second preset threshold, the similarity value of the starting point is determined based on the coordinates of the starting point of the spatial trajectory corresponding to the benchmark file, the coordinates of the starting point of the spatial trajectory corresponding to the file to be associated, and the first time factor. The similarity value of the termination point is determined based on the coordinates of the termination point of the spatial trajectory corresponding to the reference file, the coordinates of the termination point of the spatial trajectory corresponding to the file to be associated, and the second time factor.

6. The method according to any one of claims 1-5, characterized in that, Also includes: The associated target archive is compared with the suspected target database to determine the target corresponding to the archive to be associated. The suspected target database includes targets that are associated with the known target.

7. The method as described in claim 6, characterized in that, The suspected target database is determined based on the target location collection points corresponding to the known targets, wherein the target location collection points corresponding to the known targets are the location collection points that the known targets have passed through with a repetition number greater than a preset threshold.

8. A filtering device for associated target files, characterized in that, The device includes: a transceiver unit, configured to acquire the spatiotemporal trajectory corresponding to a reference file and the spatiotemporal trajectory corresponding to a file to be associated, wherein the reference file is a file of a known target; a processing unit, configured to calculate a first similarity value between the spatiotemporal trajectory corresponding to the file to be associated and the spatiotemporal trajectory corresponding to the reference file; determine a peer probability based on the number of identical location collection points in the spatiotemporal trajectory corresponding to the reference file and the spatiotemporal trajectory corresponding to the file to be associated, the number of location collection points included in the spatiotemporal trajectory corresponding to the reference file, and the number of location collection points included in the spatiotemporal trajectory corresponding to the file to be associated; determine an association parameter based on the first similarity value and the peer probability; and when the association parameter is greater than or equal to a first preset threshold, classify the file to be associated as part of the associated target file library corresponding to the known target.

9. The apparatus as claimed in claim 8, characterized in that, The processing unit is configured to determine a second similarity value based on the spatial trajectory corresponding to the baseline file, the spatial trajectory corresponding to the file to be associated, and a time decay factor; the time decay factor is determined based on the total time interval, which is determined based on the time interval of the timestamps corresponding to the same location collection points in the spatial trajectory corresponding to the baseline file and the spatial trajectory corresponding to the file to be associated; when the second similarity value is greater than or equal to a second preset threshold, the file to be associated is classified into the associated target file library corresponding to the known target.

10. The apparatus as claimed in claim 9, characterized in that, The processing unit is configured to determine the second similarity value based on the spatial trajectory corresponding to the baseline file, the spatial trajectory corresponding to the file to be associated, and the time decay factor when the association parameter is less than the first preset threshold.

11. The apparatus as claimed in claim 8, characterized in that, The processing unit is further configured to determine a starting point similarity value based on the starting point coordinates of the spatial trajectory corresponding to the benchmark file, the starting point coordinates of the spatial trajectory corresponding to the file to be associated, and a first time factor; and to determine an ending point similarity value based on the ending point coordinates of the spatial trajectory corresponding to the benchmark file, the ending point coordinates of the spatial trajectory corresponding to the file to be associated, and a second time factor; the first time factor is determined based on the starting point time interval, which is determined based on the time interval between the timestamps of the starting point coordinates of the spatial trajectory corresponding to the benchmark file and the timestamps of the starting point coordinates of the spatial trajectory corresponding to the file to be associated; the second time factor is determined based on the ending point time interval, which is determined based on the time interval between the timestamps of the ending point coordinates of the spatial trajectory corresponding to the benchmark file and the timestamps of the ending point coordinates of the spatial trajectory corresponding to the file to be associated; and when the starting point similarity value is greater than or equal to a third preset threshold and the ending point similarity value is greater than or equal to a fourth preset threshold, the file to be associated is assigned to the associated target file library corresponding to the known target.

12. The apparatus as claimed in claim 11, characterized in that, The processing unit is configured to determine the starting point similarity value based on the starting point coordinates of the spatial trajectory corresponding to the reference file, the starting point coordinates of the spatial trajectory corresponding to the file to be associated, and a first time factor when the second similarity value is less than a second preset threshold. The similarity value of the termination point is determined based on the coordinates of the termination point of the spatial trajectory corresponding to the reference file, the coordinates of the termination point of the spatial trajectory corresponding to the file to be associated, and the second time factor.

13. The apparatus according to any one of claims 8-12, characterized in that, The processing unit is used to compare the associated target archive database with the suspected target database to determine the target corresponding to the archive to be associated, wherein the suspected target database includes targets that are associated with the known targets.

14. The apparatus as claimed in claim 13, characterized in that, The suspected target database is determined based on the target location collection points corresponding to the known targets, wherein the target location collection points corresponding to the known targets are the location collection points that the known targets have passed through with a repetition number greater than a preset threshold.

15. A filtering device for associated target files, characterized in that, The device includes a processor and an interface circuit, the interface circuit being used to receive signals from other devices outside the device and transmit them to the processor or to send signals from the processor to other devices outside the device, the processor being used to implement the method as described in any one of claims 1 to 7 via logic circuits or execution code instructions.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on a computer, cause the computer to perform the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Monitoring method and device

    CN109800322A

  • Image file association method and device, electronic equipment and storage medium

    CN114549873A