A target object clustering method and device, electronic equipment and readable storage medium
By acquiring and managing the files of capture devices within the target area, statistically analyzing device relationships, and adjusting device latitude and longitude and clustering thresholds, the problem of low clustering accuracy and recall in existing technologies has been solved, achieving higher clustering accuracy and recall.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG DAHUA TECH CO LTD
- Filing Date
- 2023-02-20
- Publication Date
- 2026-05-12
AI Technical Summary
Existing target object clustering technologies may not achieve the required clustering conditions due to differences in capture distance, angle, and equipment, resulting in low clustering accuracy and recall.
By acquiring files of multiple capture devices within the target area, statistical analysis and management of the relationships between these devices are conducted, missing capture data is identified, and the latitude and longitude information of the devices and clustering thresholds are adjusted to improve clustering accuracy.
It improves the recall and accuracy of target object clustering and avoids omissions and incorrect clustering.
Smart Images

Figure CN116152535B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, electronic device, and readable storage medium for clustering target objects. Background Technology
[0002] With the continuous progress of society, target profiles have become an important means for relevant departments to identify targets.
[0003] Current target object clustering techniques primarily rely on images of target objects captured by cameras within a region over a period of time. These images are then analyzed using deep learning techniques to generate feature vectors. Similarity calculations are then performed based on these feature vectors to create target object profiles, forming a profile for each target object. Subsequently, target objects can be located based on the trajectory information within these profiles. However, due to variations in capture distance, angle, and equipment, the similarity between images of the same target object may not meet the clustering criteria. For example, the similarity between images of the same target object might be slightly below the clustering threshold, leading to missed clustering. Conversely, applying a uniform clustering standard to all images of the same target object can result in incorrect clustering. Therefore, the accuracy and recall rates of target object clustering are relatively low. Summary of the Invention
[0004] This invention provides a method, apparatus, electronic device, and readable storage medium for clustering target objects, which can improve the accuracy and recall of target object clustering.
[0005] In a first aspect, embodiments of the present invention provide a target object clustering method, including:
[0006] Acquire multiple target object files within a preset time period in a target area, wherein each of the multiple target object files includes multiple capture devices for capturing the corresponding target object, and capture time and latitude and longitude information of each of the multiple capture devices;
[0007] Based on the multiple target object files, the corresponding capture devices are processed to obtain processed files;
[0008] Based on the processed files, the correlation between two of the multiple capture devices is statistically analyzed.
[0009] Based on the aforementioned correlation, the missing snapshot data of the archives to be tested in the processed archives is extracted, and the corresponding archives are aggregated based on the snapshot data.
[0010] In one possible implementation, the step of processing the corresponding capture device based on the plurality of target object files to obtain the processed files includes:
[0011] The sequence of capture time differences between two capture devices is statistically analyzed for each of the multiple target object files. The capture time difference sequence includes multiple capture time differences, each of which is the time difference between the capture times of the two capture devices capturing the same target object.
[0012] The median of the capture time difference sequence is taken as the actual passage time between the two capture devices;
[0013] The actual passage time is compared with the preset passage time between the two capture devices obtained based on map crawling to obtain the comparison result;
[0014] If the comparison result indicates that the deviation between the actual passage time and the preset passage time is greater than the preset time threshold, then the two capture devices are rectified, and the latitude and longitude information of the two capture devices is re-recorded to obtain the rectified files.
[0015] In one possible implementation, the step of statistically analyzing the correlation between two of the multiple capture devices based on the processed archives includes:
[0016] The number of transitions between the two capture devices in each file of the processed archive is counted, and the number of transitions between the two capture devices in all files of the processed archive is summed to obtain the total number of transitions between the two capture devices;
[0017] The conversion rate between the two capture devices is determined based on the number of conversions and the total number of conversions.
[0018] If the conversion rate is greater than a preset conversion rate threshold, the correlation between the two capture devices indicates that the two capture devices are continuous capture devices.
[0019] In one possible implementation, the step of mining the missed snapshot data of the archive to be tested in the processed archive based on the association relationship includes:
[0020] If the relationship indicates that the first and second capture devices in the processed archives are continuous capture devices, and the archive to be tested in the processed archives is captured by the first capture device at the first moment and by the third capture device at the second moment after the first moment, then the archive to be tested must be captured by the second capture device at the third moment between the first moment and the second moment.
[0021] If the captured data of the file to be tested is missed at the third moment, the preset file aggregation threshold is lowered, and the captured data is aggregated with the file to be tested.
[0022] In one possible implementation, the step of statistically analyzing the correlation between two of the multiple capture devices based on the processed archives includes:
[0023] The time difference between the fourth and fifth capture devices is calculated for each file in the treated archives.
[0024] If the time difference is less than the preset passage time, it indicates that the corresponding file has passed through the fourth and fifth capture devices in one valid pass.
[0025] If the corresponding file passes through the sixth capture device while passing through the fourth and fifth capture devices, then the association between the fourth and fifth capture devices indicates the unique preset path that the corresponding file takes through between the fourth and fifth capture devices.
[0026] In one possible implementation, the step of mining the missed snapshot data of the archive to be tested in the processed archive based on the association relationship includes:
[0027] The percentage of the number of times the archives passed through the sixth capture device during the period when they effectively passed through the fourth and fifth capture devices, as well as the percentage of the total number of times they effectively passed through the fourth and fifth capture devices, is calculated.
[0028] If the ratio is greater than a preset number of times threshold, it indicates that the sixth capture device must have been passed between the fourth and fifth capture devices.
[0029] If the file to be tested in the processed archives effectively passes through the fourth and fifth capture devices, but no capture data is captured by the sixth capture device during the effective passage through the fourth and fifth capture devices, it indicates that there is a case of missing files in the archives to be tested.
[0030] The preset aggregation threshold is lowered, and the captured data that was captured by the sixth capture device during the effective passage through the fourth and fifth capture devices is aggregated again with the file to be tested.
[0031] Secondly, embodiments of the present invention provide a target object clustering device, comprising:
[0032] The acquisition unit is used to acquire multiple target object files within a preset time period in the target area. Each of the multiple target object files includes multiple capture devices for capturing the corresponding target object, and the capture time and latitude and longitude information of each of the multiple capture devices.
[0033] The processing unit is used to process the corresponding capture devices according to the multiple target object files to obtain the processed files;
[0034] The statistics unit is used to calculate the correlation between two capture devices among the plurality of capture devices based on the processed archives.
[0035] The file aggregation unit is used to mine the missed capture data of the files to be tested in the processed files according to the correlation relationship, and to aggregate the corresponding files according to the capture data.
[0036] In one possible implementation, the governance unit is specifically used for:
[0037] The sequence of capture time differences between two capture devices is statistically analyzed for each of the multiple target object files. The capture time difference sequence includes multiple capture time differences, which are the time differences between the capture times of the two capture devices capturing the same target object.
[0038] The median of the capture time difference sequence is taken as the actual passage time between the two capture devices;
[0039] The actual passage time is compared with the preset passage time between the two capture devices obtained based on map crawling to obtain the comparison result;
[0040] If the comparison result indicates that the deviation between the actual passage time and the preset passage time is greater than the preset time threshold, then the two capture devices are rectified, and the latitude and longitude information of the two capture devices is re-recorded to obtain the rectified files.
[0041] Thirdly, embodiments of the present invention also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and when the computer program is executed by the processor, it implements the method described in any of the above embodiments.
[0042] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method described in any of the above embodiments.
[0043] The beneficial effects of this invention are as follows:
[0044] This invention provides a method, apparatus, electronic device, and readable storage medium for target object clustering. First, multiple target object files within a preset time period are acquired for a target area. Each target object file includes multiple capture devices that captured images of the corresponding target object, along with the capture time and latitude / longitude information of each capture device. Then, the capture devices corresponding to the multiple target object files are processed to obtain processed files. Next, based on the processed files, the correlation between any two capture devices is statistically analyzed. Then, based on this correlation, missing capture data from the target files within the processed files is identified, and the corresponding files are clustered based on this missing capture data. In other words, by processing capture devices based on large volumes of target object files to obtain processed files, and statistically analyzing the correlation between any two capture devices within the processed files, the missing capture data from the target files within the processed files is identified, and the corresponding data is clustered based on this missing capture data, thereby improving the recall and accuracy of target object clustering. Attached Figure Description
[0045] Figure 1 This is a flowchart of a target object clustering method provided in an embodiment of the present invention;
[0046] Figure 2 for Figure 1 Flowchart of one method for step S102;
[0047] Figure 3 for Figure 1 Flowchart of one method of the first implementation of step S103;
[0048] Figure 4 for Figure 1 Flowchart of one method of the second implementation of step S103;
[0049] Figure 5 To adopt Figure 4 The implementation method shown Figure 1 Flowchart of one method for step S104;
[0050] Figure 6 This is a structural block diagram of a target object clustering device provided in an embodiment of the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Furthermore, the embodiments and features in the embodiments of the present invention can be combined with each other without conflict. Based on the described embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0052] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains. The terms "comprising" or "including," or similar terms as used in this invention, mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects.
[0053] In related technologies, deep learning is primarily used to extract features of target objects from images. Based on these features, images with similarity greater than a clustering threshold are then clustered. However, due to variations in capture distance, angle, and device, the similarity of images of the same target object may not meet the clustering criteria. Consequently, the accuracy and recall of target object clustering are relatively low.
[0054] In view of this, embodiments of the present invention provide a target object clustering method, apparatus, electronic device, and readable storage medium for improving the accuracy and recall of target object clustering.
[0055] like Figure 1 As shown, this embodiment of the invention provides a target object clustering method, which includes:
[0056] S101: Obtain multiple target object files within a preset time period in the target area, wherein each of the multiple target object files includes multiple capture devices for capturing the corresponding target object, and capture time and latitude and longitude information of each of the multiple capture devices;
[0057] In the specific implementation process, before clustering the target objects, multiple target object files within a preset time period of the target area are first acquired. Each target object file includes multiple capture devices that captured the corresponding target object, and the capture time and latitude / longitude information of each capture device. The target object can be a person or a vehicle, without limitation. The target area can be any area set by the user, without limitation. The preset time period can be any time period set by the user, such as one week. Of course, the target area and preset time period can be set according to actual application needs, without limitation. Furthermore, the capture devices involved in this embodiment of the invention can be checkpoints, image acquisition units, or other devices that collect images of target objects, without limitation.
[0058] In one exemplary embodiment, information on all portrait files within the target area for the previous N days is obtained, where N is a positive integer, for example, there exists a certain portrait file P. m The file P m The captured data is sorted by capture time, and the corresponding file P m The format is as follows:
[0059]
[0060] Among them, the archive P m Taking m1 as an example, T m1 Indicates file P m The capture time corresponding to the first captured image, C m1 For the corresponding capture device, G m1 This corresponds to the latitude and longitude of the captured image. Furthermore, the meanings related to m2, m3, m4, and m5 can be found in the description of the corresponding section for m1, and will not be detailed here. Thus, file P... m There are a total of 5 snapshots.
[0061] S102: Process the corresponding capture device according to the multiple target object files to obtain the processed files;
[0062] In the specific implementation process, after acquiring multiple target object files within a preset time period in the target area, the corresponding capture devices can be processed based on these multiple target object files to obtain processed files. Taking a capture device as a checkpoint as an example, if there is a significant deviation between the checkpoint's latitude and longitude and the actual latitude and longitude, the checkpoint's latitude and longitude can be re-recorded to achieve processing of the checkpoint. The specific processing procedure can be referred to the description in the relevant sections below, and will not be detailed here.
[0063] S103: Based on the processed files, calculate the correlation between two of the multiple capture devices;
[0064] In the specific implementation process, after obtaining the processed files, the correlation between two capture devices among multiple capture devices can be calculated based on these files. These two capture devices can be any two devices from the multiple capture devices. In one exemplary embodiment, the correlation between the two capture devices can be that they are consecutive checkpoints; for example, checkpoints A and B are consecutive checkpoints, and passing through checkpoint A necessarily leads to passing through checkpoint B. Of course, the correlation between the two checkpoints can also be other cases, which are not limited here.
[0065] S104: Based on the aforementioned correlation, extract the missing snapshot data of the archives to be tested in the processed archives, and aggregate the corresponding archives based on the snapshot data.
[0066] In the specific implementation process, after determining the correlation between two of the multiple capture devices, the missing capture data in the files to be tested within the processed archives can be extracted based on this correlation. Then, the corresponding files are aggregated based on this capture data. This can, to a certain extent, avoid the situation of missing files in the aggregation, thereby improving the recall rate and accuracy of the target object aggregation.
[0067] It should be noted that, in the specific implementation process, before acquiring multiple target object files within a preset time period in the target area, it is possible to acquire all capture devices within the target area and the latitude and longitude information of each capture device, as well as the map travel paths between pairs of capture devices and the corresponding map travel times based on map crawling. In one exemplary embodiment, if the target object is a person, the map travel path can be the optimal walking path between pairs of capture devices obtained from map crawling, and correspondingly, the map travel time can be the optimal walking time between pairs of capture devices obtained from map crawling.
[0068] In practice, based on the latitude and longitude information of checkpoints within the target area, the map can be crawled to retrieve the travel paths between any two checkpoints. For example, if there are three checkpoints C... i C j C k The distances corresponding to the map travel paths between these three checkpoints are D respectively. ij D ik D jk If D is satisfied ij +D jk <α*D ik Where α is a constant greater than 1, for example, α = 1.05, which indicates that the checkpoint C j To get from checkpoint C iWalk to checkpoint C k The optimal travel path passes through the checkpoints. It should be noted that the optimal travel path between two checkpoints can be one or more, and correspondingly, the optimal path can pass through one or more checkpoints, depending on the specific application. Furthermore, the time taken to crawl the optimal travel path between any two checkpoints on the map can also be determined, i.e., the optimal travel time. For example, if there are three checkpoints C... i C j C k Among them, the optimal passage time between each pair of checkpoints is Time. ij Time ik Time jk After obtaining the map travel paths and map travel times for each target object file within the target area, the same processing principle can be used to determine the map travel paths and map travel times for each file in the processed files, so as to subsequently determine the correlation between relevant capture devices. The specific process for determining the correlation between capture devices can be found in the descriptions below, and will not be detailed here.
[0069] In embodiments of the present invention, such as Figure 2 As shown, step S102: The corresponding capture device is processed according to the multiple target object files to obtain the processed files, including:
[0070] S201: Calculate the sequence of capture time differences between two capture devices for each of the multiple target object files. The sequence of capture time differences includes multiple capture time differences, and each capture time difference is the time difference between the capture times of the two capture devices capturing the same target object.
[0071] S202: The median of the capture time difference sequence is taken as the actual passage time between the two capture devices;
[0072] S203: Compare the actual passage time with the preset passage time between the two capture devices based on map crawling to obtain a comparison result;
[0073] S204: If the comparison result shows that the deviation between the actual passage time and the preset passage time is greater than the preset time threshold, then the two capture devices are rectified, the latitude and longitude information of the two capture devices is re-recorded, and the rectified files are obtained.
[0074] In the specific implementation process, steps S201 to S204 are implemented as follows:
[0075] First, we statistically analyze the time difference sequence between two capture devices for each target object file across multiple target object files. In other words, we calculate the time difference between two capture devices for each target object file. This time difference sequence includes multiple capture time differences, each representing the time difference between the capture times of the same target object by the two capture devices. For example, file P... m Passing through checkpoint C m2 and checkpoint C m1 The time difference is T m2 -T m1 Following the same processing method, the time difference between consecutive passages of all target object files through two different checkpoints is statistically calculated to obtain a time difference sequence between each pair of checkpoints. Then, the median of the capture time difference sequence is taken as the actual passage time between the two capture devices. Next, this actual passage time is compared with the preset passage time between the two capture devices obtained from map crawling, and the comparison result is obtained. Then, the corresponding capture devices are rectified based on the comparison result. For example, if the comparison result shows that the deviation between the actual passage time and the preset passage time is greater than a preset time threshold, the corresponding two capture devices are rectified, and the latitude and longitude information of the two capture devices is re-recorded to obtain the rectified files. The preset time threshold can be set according to actual application needs and is not limited here.
[0076] For example, checkpoint C i C j The map travel time between them is Time ij The corresponding actual travel time is S. ij If Time ij ≥γ, and S ij ≥γ, where γ is the longest time threshold, indicates that checkpoint C i and checkpoint C j If the distance is too far, travel time comparison is not performed; otherwise, if the condition is met... At that time, the map travel time is considered to be Time. ij And the actual travel time S ij The coordinates are basically consistent, and the latitude and longitude of the checkpoint are basically correct; if the above relationship is not met, then the coordinates are satisfied. This indicates the map travel time. ij And the actual travel time S ij There is a significant deviation; correspondingly, checkpoint C i The latitude and longitude of the checkpoint deviates significantly from its actual latitude and longitude. Checkpoint C jThe latitude and longitude of the checkpoints deviate significantly from their actual latitude and longitude. It should be noted that θ represents the time similarity threshold, and ∈ represents the regularization constant. Therefore, the latitude and longitude of the corresponding checkpoints can be re-recorded to address these issues. In this way, after addressing checkpoints with significant deviations from their actual latitude and longitude, the accuracy of subsequent target object clustering is effectively guaranteed. In this embodiment of the invention, there are two possible implementation methods to statistically analyze the correlation between two capture devices among multiple capture devices, but these are not the only methods available.
[0077] In the first implementation, such as Figure 3 As shown, step S103: Based on the processed files, statistically analyze the correlation between two capture devices among the multiple capture devices, including:
[0078] S301: Count the number of transitions between the two capture devices in each file of the treated archives, and sum up the number of transitions between the two capture devices for all files in the treated archives to obtain the total number of transitions between the two capture devices;
[0079] S302: Determine the conversion rate between the two capture devices based on the number of conversions and the total number of conversions;
[0080] S303: If the conversion rate is greater than the preset conversion rate threshold, the correlation between the two capture devices indicates that the two capture devices are continuous capture devices.
[0081] In the specific implementation process, steps S301 to S303 are implemented as follows:
[0082] First, the number of transitions between the two capture devices in each file of the processed archive is counted, and the number of transitions between the two capture devices in all files of the processed archive is summed to obtain the total number of transitions between the two capture devices. In one exemplary embodiment, file P mentioned above is still used. m For example, the archive P after treatment m The number of checkpoint conversions is R(C) m1 C m2 )=1,R(C m2 C m3 )=1,R(C m3 C m4 )=1,R(C m4 C m5 = 1. Accordingly, based on the same calculation principle, the number of checkpoint transitions for all files in the processed archives is counted. Then, the individual transition counts are summed to obtain the total number of transitions between checkpoints for all files.
[0083] Then, based on the number of conversions between the two capture devices for all files in the processed archives and the total number of conversions between the two capture devices, the conversion rate between the two capture devices is determined. In one exemplary embodiment, for example, there is a checkpoint C. s Their conversion ports are C and C respectively. t C u C v Among them, R(C) s C t ) = a st ,R(C s C u ) = a su ,R(C s C v ) = a sv Then the corresponding gate conversion rates for gate Cs are RT(C) s C t ) = a st / (a st +a su +a sv ),RT(C s C u ) = a su / (a st +a su +a sv ),RT(C s C v ) = a sv / (a st +a su +a sv Correspondingly, using the same calculation principle, the conversion rate between the two capture devices for all archives after treatment was statistically calculated.
[0084] Then, based on the conversion rate between the two capture devices, the correlation between the corresponding capture devices is determined. Taking the aforementioned checkpoint C as an example... s For example, if checkpoint C s Conversion rate Where RT is a preset conversion rate threshold, for example, if RT is 0.95, it indicates that the gate C s and checkpoint C t For continuous checkpoints, correspondingly, checkpoint C s and checkpoint C t The correlation between the two capture devices indicates that they are continuous capture devices, meaning they pass through checkpoint C. s It will inevitably pass through checkpoint C t .
[0085] It should be noted that, in the process of counting the number of times the two capture devices are switched in each file after statistical management, if the images captured in two consecutive files are from the same checkpoint, they will not be included in the statistics.
[0086] In this embodiment of the invention, the following is adopted: Figure 3 The implementation shown involves statistically analyzing the correlation between two capture devices among multiple capture devices. In step 104, based on this correlation, missing capture data from the processed archives is extracted, including:
[0087] If the relationship indicates that the first and second capture devices in the processed archives are continuous capture devices, and the archive to be tested in the processed archives is captured by the first capture device at the first moment and by the third capture device at the second moment after the first moment, then the archive to be tested must be captured by the second capture device at the third moment between the first moment and the second moment.
[0088] If the captured data of the file to be tested is missed at the third moment, the preset file aggregation threshold is lowered, and the captured data is aggregated with the file to be tested.
[0089] In the specific implementation process, based on the processed archives, after statistically analyzing the correlation between two capture devices among multiple capture devices, if the correlation indicates that the first and second capture devices in the processed archives are continuous capture devices, and the file to be tested in the processed archives is captured by the first capture device at the first moment and by the third capture device at the second moment after the first moment, then the file to be tested in the processed archives must have been captured by the second capture device at the third moment between the first and second moments; that is, the first and second capture devices are continuous capture devices. In this way, if the capture data of the file to be tested is missed at the third moment, the preset archive threshold can be lowered accordingly, and the captured data of the missed file to be tested can be clustered with the file to be tested, thereby improving the recall rate and accuracy of target object clustering.
[0090] Accordingly, in practical applications, there are two possible scenarios. In the first scenario, if there exists a file at a previous time T... cs C-shaped checkpoint s A snapshot, the next moment T cu C-shaped checkpoint u A snapshot will inevitably be taken in [T] cs T cu [Blocked by C within the time interval] t Capture the image. In the second scenario, there is a file aggregation error; the previous moment T... cs C-shaped checkpoint sThe captured image and the next moment T cu C-shaped checkpoint u The captured images are misfiled, meaning they are not images of the same target object.
[0091] In the first case, if [T] exists cs T cu [Not checked by checkpoint C within the time interval] t The captured images indicate that some image data has been missed in the current clustering process. In this case, the clustering threshold can be appropriately lowered to cluster the missed image data with the files to be tested, thereby ensuring the recall and accuracy of the target object clustering.
[0092] It should be noted that, in the first scenario, taking the above exemplary embodiment as an example, during the process of lowering the aggregation threshold and re-aggregating the missed image data with the file to be tested, if aggregation fails, the data can be re-aggregated using the checkpoint C. s The captured image data and the data from the checkpoint C t The captured data is compared for similarity. If the similarity is less than the cluster similarity threshold after the clustering threshold is reduced, it indicates that the clustering is incorrect. In this case, the data can be split, thus ensuring the accuracy of the target object clustering.
[0093] In the second implementation, such as Figure 4 As shown, step S103: Based on the processed files, statistically analyze the correlation between two capture devices among the multiple capture devices, including:
[0094] S401: Calculate the time difference between the fourth and fifth capture devices for each file in the treated archives;
[0095] S402: If the time difference is less than the preset passage time, it indicates that the corresponding file has passed through the fourth and fifth capture devices in one valid pass.
[0096] S403: If the corresponding file passes through the sixth capture device while passing through the fourth capture device and the fifth capture device, then the association between the fourth capture device and the fifth capture device indicates the unique preset path of the corresponding file passing through the fourth capture device and the fifth capture device.
[0097] In the specific implementation process, steps S401 to S403 are implemented as follows:
[0098] After obtaining the processed archives, the time difference between each archive passing through the fourth and fifth capture devices is calculated. If this time difference is less than the preset passage time, it indicates that the corresponding archive has validly passed through the fourth and fifth capture devices in one go; that is, the images captured by the fourth and fifth capture devices are valid data for the corresponding archive. Then, it is determined whether a sixth capture device exists during the period when the corresponding archive passes through the fourth and fifth capture devices. If it exists, it indicates that the correlation between the fourth and fifth capture devices shows a unique preset path for the corresponding archive passing through the fourth and fifth capture devices. Taking the aforementioned archive P as an example... m For example, if checkpoint C m2 For checkpoint C m1 and checkpoint C m3 The checkpoints passed through in the unique preset path. The checkpoints C passed through can be statistically analyzed based on all files in the processed archives. m1 and checkpoint C m3 In this case, calculate the time difference |T m3 -T m1 |, where T m3 For passing through checkpoint C m3 Time, T m1 For passing through checkpoint C m1 The time difference is less than the preset passage time, indicating that the corresponding file has passed through the fourth and fifth capture devices in one valid pass. Let's take file P as an example. m For example, if |T m3 -T m1 |<β*Time 13 Where β is the optimal passage path time threshold coefficient, that is, if the time difference is less than the preset time threshold, it can be considered as a valid passage through checkpoint C. m1 and checkpoint C m3 If passing through checkpoint C m1 and checkpoint C m3 During the period, passing through checkpoint C m2 Then it is considered that it has passed through checkpoint C m1 and checkpoint C m3 The only preset path.
[0099] In this embodiment of the invention, the following is adopted: Figure 4 The implementation shown involves calculating the correlation between two capture devices among multiple capture devices, such as... Figure 5 As shown, in step 104, based on the aforementioned correlation, the missing snapshot data of the files to be tested in the processed archives is extracted, including:
[0100] S501: Calculate the proportion of the number of times the archives effectively passed through the fourth and fifth capture devices during the period when they passed through the sixth capture device, relative to the total number of times they effectively passed through the fourth and fifth capture devices;
[0101] S502: If the ratio is greater than the preset number of times threshold, it indicates that the sixth capture device must be passed during the period between the fourth capture device and the fifth capture device;
[0102] S503: If the file to be tested in the processed file effectively passes through the fourth and fifth capture devices, but there is no capture data from the sixth capture device during the effective passage through the fourth and fifth capture devices, it indicates that there is a case of missing file aggregation in the file to be tested;
[0103] S504: Lower the preset aggregation threshold, and re-aggregate the captured data captured by the sixth capture device during the effective passage through the fourth and fifth capture devices with the file to be tested.
[0104] In the specific implementation process, steps S501 to S504 are as follows:
[0105] Based on the processed archives, after statistically analyzing the correlation between two of the multiple capture devices, the proportion of the number of times the device passed through the sixth capture device during the period when it effectively passed through the fourth and fifth capture devices, relative to the total number of times the device effectively passed through the fourth and fifth capture devices, is calculated. Using the aforementioned archive P... m For example, if checkpoint C m2 For checkpoint C m1 and checkpoint C m3 The checkpoints passed through in the unique preset path, and the valid checkpoints C in all files after statistical processing. m1 and checkpoint C m3 During the period, passing through checkpoint C m2 The number of times that passed through checkpoint C accounted for a significant portion of the valid passage. m1 and checkpoint C m3The proportion of total occurrences is then compared with a preset threshold. If the proportion is greater than the preset threshold, it indicates that the file passed through the fourth and fifth capture devices and must have passed through the sixth capture device. If the file under test in the processed archives is detected to have effectively passed through the fourth and fifth capture devices, but no capture data from the sixth capture device is found during the effective passage through the fourth and fifth capture devices, it indicates that the file under test has been missed in the aggregation process, or even that the aggregation process is incorrect. For cases of missed aggregation, the preset aggregation threshold can be lowered, and the capture data captured by the sixth capture device during the effective passage through the fourth and fifth capture devices can be aggregated again with the file under test, thereby improving the recall rate of the aggregation process.
[0106] Taking the aforementioned exemplary embodiment as an example, if file P m Valid through checkpoint C m1 and checkpoint C m3 During the period, passing through checkpoint C m2 The number of times that passed through checkpoint C accounted for a significant portion of the valid passage. m1 and C m3 If the proportion of total attempts is greater than the threshold after lowering the preset clustering threshold, it indicates that the signal has passed through checkpoint C. m1 and checkpoint C m3 It is necessary to pass through checkpoint C during the process. m2 Therefore, if file P is detected during the actual testing process... m Valid through checkpoint C m1 and checkpoint C m3 However, there was no checkpoint C during this period. m2 The captured data, then the file P m There will inevitably be instances of data aggregation, or even data aggregation errors. In this case, the data will pass through checkpoint C. m1 and checkpoint C m3 During the period, it was blocked at point C m2 Captured images and files P m By re-aggregating the files, the recall rate of the aggregated files is improved. In practical applications, if the files passing through checkpoint C are... m1 and checkpoint C m3 The device is held by C m2 Captured images and files P m If file aggregation fails again, the blocked port C can be used. m1 The captured image data and the data from the checkpoint C m3 The captured data is compared for similarity. If the similarity is less than the clustering similarity threshold, the clustering is considered incorrect and the data needs to be split, thereby further ensuring the accuracy of the clustering.
[0107] It should be noted that after determining the association between two of the multiple capture devices using the aforementioned first and second implementation methods, if there are any missing files in the file to be tested, the preset file aggregation threshold can be lowered. The degree of reduction in the preset file aggregation threshold can be a slight decrease by a certain percentage. For example, if the preset file aggregation threshold is X, the reduced threshold is aX, where a can be 0.01. Of course, the specific value of a can be set according to actual application needs and is not limited here. Furthermore, the degree of reduction in the preset file aggregation threshold can also be set according to actual application needs, and is not limited here.
[0108] Based on the same inventive concept, such as Figure 6 As shown, this embodiment of the invention also provides a target object clustering device, which includes:
[0109] The acquisition unit 10 is used to acquire multiple target object files within a preset time period in the target area. Each of the multiple target object files includes multiple capture devices for capturing the corresponding target object, and capture time and latitude and longitude information of each of the multiple capture devices.
[0110] The processing unit 20 is used to process the corresponding capture devices according to the multiple target object files to obtain the processed files;
[0111] The statistics unit 30 is used to calculate the correlation between two of the multiple capture devices based on the processed files.
[0112] The file aggregation unit 40 is used to mine the missing capture data of the files to be tested in the processed files according to the correlation relationship, and to aggregate the corresponding files according to the capture data.
[0113] In this embodiment of the invention, the governance unit 20 is specifically used for:
[0114] The sequence of capture time differences between two capture devices is statistically analyzed for each of the multiple target object files. The capture time difference sequence includes multiple capture time differences, which are the time differences between the capture times of the two capture devices capturing the same target object.
[0115] The median of the capture time difference sequence is taken as the actual passage time between the two capture devices;
[0116] The actual passage time is compared with the preset passage time between the two capture devices obtained based on map crawling to obtain the comparison result;
[0117] If the comparison result indicates that the deviation between the actual passage time and the preset passage time is greater than the preset time threshold, then the two capture devices are rectified, and the latitude and longitude information of the two capture devices is re-recorded to obtain the rectified files.
[0118] In this embodiment of the invention, the statistical unit 30 is specifically used for:
[0119] The number of transitions between the two capture devices in each file of the processed archive is counted, and the number of transitions between the two capture devices in all files of the processed archive is summed to obtain the total number of transitions between the two capture devices;
[0120] The conversion rate between the two capture devices is determined based on the number of conversions and the total number of conversions.
[0121] If the conversion rate is greater than a preset conversion rate threshold, the correlation between the two capture devices indicates that the two capture devices are continuous capture devices.
[0122] In this embodiment of the invention, the aggregation unit 40 is specifically used for:
[0123] If the relationship indicates that the first and second capture devices in the processed archives are continuous capture devices, and the archive to be tested in the processed archives is captured by the first capture device at the first moment and by the third capture device at the second moment after the first moment, then the archive to be tested must be captured by the second capture device at the third moment between the first moment and the second moment.
[0124] If the captured data of the file to be tested is missed at the third moment, the preset file aggregation threshold is lowered, and the captured data is aggregated with the file to be tested.
[0125] In this embodiment of the invention, the statistical unit 30 is specifically used for:
[0126] The time difference between the fourth and fifth capture devices is calculated for each file in the treated archives.
[0127] If the time difference is less than the preset passage time, it indicates that the corresponding file has passed through the fourth and fifth capture devices in one valid pass.
[0128] If the corresponding file passes through the sixth capture device while passing through the fourth and fifth capture devices, then the association between the fourth and fifth capture devices indicates the unique preset path that the corresponding file takes through between the fourth and fifth capture devices.
[0129] In this embodiment of the invention, the aggregation unit 40 is specifically used for:
[0130] The percentage of the number of times the archives passed through the sixth capture device during the period when they effectively passed through the fourth and fifth capture devices, as well as the percentage of the total number of times they effectively passed through the fourth and fifth capture devices, is calculated.
[0131] If the ratio is greater than a preset number of times threshold, it indicates that the sixth capture device must have been passed between the fourth and fifth capture devices.
[0132] If the file to be tested in the processed archives effectively passes through the fourth and fifth capture devices, but no capture data is captured by the sixth capture device during the effective passage through the fourth and fifth capture devices, it indicates that there is a case of missing files in the archives to be tested.
[0133] The preset aggregation threshold is lowered, and the captured data that was captured by the sixth capture device during the effective passage through the fourth and fifth capture devices is aggregated again with the file to be tested.
[0134] Based on the same inventive concept, embodiments of the present invention also provide an electronic device, which includes a memory and a processor. The memory stores a computer program that can run on the processor. When the computer program is executed by the processor, it implements the following method:
[0135] Acquire multiple target object files within a preset time period in a target area, wherein each of the multiple target object files includes multiple capture devices for capturing the corresponding target object, and capture time and latitude and longitude information of each of the multiple capture devices;
[0136] Based on the multiple target object files, the corresponding capture devices are processed to obtain processed files;
[0137] Based on the processed files, the correlation between two of the multiple capture devices is statistically analyzed.
[0138] Based on the aforementioned correlation, the missing snapshot data of the archives to be tested in the processed archives is extracted, and the corresponding archives are aggregated based on the snapshot data.
[0139] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the target object clustering method as described above.
[0140] This invention provides a method, apparatus, electronic device, and readable storage medium for target object clustering. First, multiple target object files within a preset time period are acquired for a target area. Each target object file includes multiple capture devices that captured images of the corresponding target object, along with the capture time and latitude / longitude information of each capture device. Then, the capture devices corresponding to the multiple target object files are processed to obtain processed files. Next, based on the processed files, the correlation between any two capture devices is statistically analyzed. Then, based on this correlation, missing capture data from the target files within the processed files is identified, and the corresponding files are clustered based on this missing capture data. In other words, by processing capture devices based on large volumes of target object files to obtain processed files, and statistically analyzing the correlation between any two capture devices within the processed files, the missing capture data from the target files within the processed files is identified, and the corresponding data is clustered based on this missing capture data, thereby improving the recall and accuracy of target object clustering.
[0141] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0142] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0143] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0144] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0145] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for clustering target objects, characterized in that, include: Acquire multiple target object files within a preset time period in a target area, wherein each of the multiple target object files includes multiple capture devices for capturing the corresponding target object, and capture time and latitude and longitude information of each of the multiple capture devices; Based on the multiple target object files, the corresponding capture devices are processed to obtain processed files; Based on the processed files, the correlation between two of the multiple capture devices is statistically analyzed. Based on the aforementioned correlation, the missing snapshot data of the archives to be tested in the processed archives is extracted, and the corresponding archives are aggregated based on the snapshot data; The step of processing the corresponding capture devices based on the multiple target object files to obtain processed files includes: The sequence of capture time differences between two capture devices is statistically analyzed for each of the multiple target object files. The capture time difference sequence includes multiple capture time differences, each of which is the time difference between the capture times of the two capture devices capturing the same target object. The median of the capture time difference sequence is taken as the actual passage time between the two capture devices; The actual passage time is compared with the preset passage time between the two capture devices obtained based on map crawling to obtain the comparison result; If the comparison result indicates that the deviation between the actual passage time and the preset passage time is greater than the preset time threshold, then the two capture devices are rectified, and the latitude and longitude information of the two capture devices is re-recorded to obtain the rectified files.
2. The method as described in claim 1, characterized in that, The step of calculating the correlation between two capture devices among the multiple capture devices based on the processed files includes: The number of transitions between the two capture devices in each file of the processed archive is counted, and the number of transitions between the two capture devices in all files of the processed archive is summed to obtain the total number of transitions between the two capture devices; The conversion rate between the two capture devices is determined based on the number of conversions and the total number of conversions. If the conversion rate is greater than a preset conversion rate threshold, the correlation between the two capture devices indicates that the two capture devices are continuous capture devices.
3. The method as described in claim 2, characterized in that, The step of mining the missed snapshot data of the archives to be tested in the processed archives based on the aforementioned correlation includes: If the relationship indicates that the first and second capture devices in the processed archives are continuous capture devices, and the archive to be tested in the processed archives is captured by the first capture device at the first moment and by the third capture device at the second moment after the first moment, then the archive to be tested must be captured by the second capture device at the third moment between the first moment and the second moment. If the captured data of the file to be tested is missed at the third moment, the preset file aggregation threshold is lowered, and the captured data is aggregated with the file to be tested.
4. The method as described in claim 1, characterized in that, The step of calculating the correlation between two capture devices among the multiple capture devices based on the processed files includes: The time difference between the fourth and fifth capture devices is calculated for each file in the treated archives. If the time difference is less than the preset passage time, it indicates that the corresponding file has passed through the fourth and fifth capture devices in one valid pass. If the corresponding file passes through the sixth capture device while passing through the fourth and fifth capture devices, then the association between the fourth and fifth capture devices indicates the unique preset path that the corresponding file takes through between the fourth and fifth capture devices.
5. The method as described in claim 4, characterized in that, The step of mining the missed snapshot data of the archives to be tested in the processed archives based on the aforementioned correlation includes: The percentage of the number of times the archives passed through the sixth capture device during the period when they effectively passed through the fourth and fifth capture devices, as well as the percentage of the total number of times they effectively passed through the fourth and fifth capture devices, is calculated. If the ratio is greater than a preset number of times threshold, it indicates that the sixth capture device must have been passed between the fourth and fifth capture devices. If the file to be tested in the processed archives effectively passes through the fourth and fifth capture devices, but no capture data is captured by the sixth capture device during the effective passage through the fourth and fifth capture devices, it indicates that there is a case of missing files in the archives to be tested. The preset aggregation threshold is lowered, and the captured data that was captured by the sixth capture device during the effective passage through the fourth and fifth capture devices is aggregated again with the file to be tested.
6. A target object clustering device, characterized in that, include: The acquisition unit is used to acquire multiple target object files within a preset time period in the target area. Each of the multiple target object files includes multiple capture devices for capturing the corresponding target object, and the capture time and latitude and longitude information of each of the multiple capture devices. The processing unit is used to process the corresponding capture devices according to the multiple target object files to obtain the processed files; The statistics unit is used to calculate the correlation between two capture devices among the plurality of capture devices based on the processed archives. The file aggregation unit is used to mine the missed capture data of the files to be tested in the processed files according to the correlation relationship, and to aggregate the corresponding files according to the capture data. Specifically, the governance unit is used for: The sequence of capture time differences between two capture devices is statistically analyzed for each of the multiple target object files. The capture time difference sequence includes multiple capture time differences, which are the time differences between the capture times of the two capture devices capturing the same target object. The median of the capture time difference sequence is taken as the actual passage time between the two capture devices; The actual passage time is compared with the preset passage time between the two capture devices obtained based on map crawling to obtain the comparison result; If the comparison result indicates that the deviation between the actual passage time and the preset passage time is greater than the preset time threshold, then the two capture devices are rectified, and the latitude and longitude information of the two capture devices is re-recorded to obtain the rectified files.
7. An electronic device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program that can run on the processor, and when the computer program is executed by the processor, it implements the method of any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.