An archive management method, device and computer readable storage medium
By sorting the captured data by time and analyzing the marked locations, and combining feature similarity and spatiotemporal rules, the problem of clustering errors in urban archive management was solved, achieving higher accuracy and data purity in archive management.
Patent Information
- Application Number
- CN202211088038.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-05
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2042-09-05
AI Technical Summary
In existing technologies, archive management based on spatiotemporal information has limited effectiveness in complex urban road environments, making it difficult to effectively distinguish between data of the same person and data of different people, leading to clustering errors.
By acquiring shooting data from multiple initial archives, sorting them according to shooting time, and setting marker point data, the archives are split using shooting location and time difference information. Combined with feature similarity and spatiotemporal rules, abnormal data is merged and excluded, thereby improving the accuracy and fault tolerance of archive management.
By deeply utilizing spatiotemporal information, we can improve the effectiveness of archive management, enhance its accuracy and error tolerance, ensure data purity, and make it suitable for archive management in complex urban environments.
Smart Images

Figure CN116304376B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision, and in particular to a method, apparatus, device, and computer-readable storage medium for document management. Background Technology
[0002] In existing technologies, feature clustering is used for file management. Feature clustering mainly refers to determining which images belong to the same person based on identity features and attributes such as faces and bodies, as well as spatiotemporal information. This allows data of the same person to be grouped together, while different people are assigned to different files, achieving one file per person. There is a lot of publicly available information on feature and attribute-based clustering, most of which compares the similarity of two data points and then sets a threshold; data exceeding the threshold is considered to belong to the same person.
[0003] However, due to the extensive network of roads in cities, the effectiveness of spatiotemporal information is often very limited. For example, two images with a long time interval can appear in almost any combination of spatial points, making it virtually impossible to apply spatiotemporal contradiction rules. This results in the limited role of spatiotemporal information in typical clustering systems. Summary of the Invention
[0004] This application provides at least one method, apparatus, and computer-readable storage medium for file management.
[0005] The document management method includes:
[0006] Multiple initial files are acquired, each initial file including several shooting data, the shooting data including shooting location information and shooting time information;
[0007] The shooting data in each initial file are sorted according to the shooting time information to obtain a shooting data list for each initial file;
[0008] In the list of captured data, at least one marker point data is set;
[0009] Obtain an adjacent shooting data in the shooting data list of the at least one marked point data, and obtain the shooting position difference information and shooting time difference information between the at least one marked point data and the adjacent shooting data;
[0010] When the shooting location difference information and the shooting time difference information meet the misfile condition, the first file is split into at least two first files based on the adjacent shooting data and the at least one marker point data.
[0011] Wherein, the marked point data is entrance and exit point data, and the first adjacent shooting data is the previous shooting data of the entrance and exit point data in the shooting data list;
[0012] the shooting position difference information and the shooting time difference information satisfy a wrong reel condition, splitting the initial reel into two first reels according to the adjacent shooting data and the at least one marked point data.
[0013] When the shooting position difference information and the shooting time difference information satisfy the wrong reel condition, the first adjacent shooting data and the shooting data before it are taken as a first reel, and the entrance and exit point data and the shooting data after it are taken as another first reel.
[0014] The marked point data is entrance and exit point data, and the second adjacent shooting data is the next shooting data of the entrance and exit point data in the shooting data list.
[0015] the shooting position difference information and the shooting time difference information satisfy a wrong reel condition, splitting the initial reel into two first reels according to the adjacent shooting data and the at least one marked point data.
[0016] When the shooting position difference information and the shooting time difference information satisfy the wrong reel condition, the second adjacent shooting data and the shooting data after it are taken as a first reel, and the entrance and exit point data and the shooting data before it are taken as another first reel.
[0017] The wrong reel condition is that the shooting positions of two shooting data do not belong to the same floor or the same floor entrance and exit, or the time interval of the shooting time is greater than or equal to a preset interval threshold.
[0018] After the initial reel is split into two first reels according to the adjacent shooting data and the at least one marked point data, the reel management method further includes:
[0019] Obtain a plurality of first reels split from the same initial reel.
[0020] Obtain data feature similarity information of two first reels in the plurality of first reels.
[0021] In the case where the data feature similarity information of two first reels satisfies a preset similarity condition, the two first reels are merged into a second reel.
[0022] The data feature similarity information of two first reels in the plurality of first reels includes:
[0023] Calculate a first feature similarity between the shooting data of one first reel and the shooting data of another first reel, and obtain a plurality of groups of first feature similarities of shooting data pairs.
[0024] The two first archives are merged into a second archive in a case where data feature similarity information of the two first archives satisfies a preset similarity condition, and the method comprises the following steps:
[0025] A first proportion in a total group number is obtained by counting a group number of the shooting data pairs whose first feature similarity is greater than or equal to a first preset similarity threshold value.
[0026] In a case where the first proportion is greater than or equal to a first preset proportion, the two first archives are merged into a second archive.
[0027] The first feature similarity comprises one or more of a face feature similarity, a human body feature similarity and a torso feature similarity.
[0028] After the first archive is split into at least two first archives according to the adjacent shooting data and the at least one marker point data, the archive management method further comprises the following steps:
[0029] For each first archive, a floor marker of all shooting data in the first archive is obtained, and a floor marker and a store marker of each shooting data are obtained.
[0030] When a floor marker of a previous shooting data is not the same as a floor marker of a current shooting data, or a store marker of the previous shooting data is not the same as a store marker of the current shooting data, the current shooting data is marked as an abnormal shooting data.
[0031] The abnormal shooting data is excluded from all the first archives to obtain a third archive.
[0032] The abnormal shooting data is excluded from all the first archives to obtain a third archive, which comprises the following steps:
[0033] A second feature similarity between the abnormal shooting data and all shooting data before the abnormal shooting data in each first archive is calculated.
[0034] A number of shooting data whose second feature similarity is greater than or equal to a second preset similarity threshold value is obtained, and a second proportion of the number of shooting data to a number of all shooting data before the abnormal shooting data is calculated.
[0035] In a case where the second proportion is greater than or equal to a second preset proportion, the abnormal shooting data is retained.
[0036] In a case where the second proportion is less than the second preset proportion, the abnormal shooting data is excluded.
[0037] The third archive is composed of the remaining shooting data of each first archive.
[0038] wherein, after the abnormal shooting data is excluded from all the first archives to obtain the third archives, the archive management method further comprises:
[0039] judging whether the shooting data after the abnormal shooting data excluded from the first archives is the same floor mark as the abnormal shooting data or is the same point group;
[0040] if yes, the abnormal shooting data is put back into the original first archives.
[0041] wherein, after the first archives are split into at least two first archives according to the adjacent shooting data and the at least one marked point data, the archive management method further comprises:
[0042] obtaining a time intersection archive group composed of any two first archives with time intersection in the at least two first archives;
[0043] calculating a third feature similarity between the shooting data of one first archive and the shooting data of another first archive in the time intersection archive group, to obtain third feature similarities of a plurality of shooting data pairs;
[0044] obtaining a third proportion in the total number of groups of the shooting data pairs with a first feature similarity greater than or equal to a first preset similarity threshold;
[0045] in the case that the third proportion is greater than or equal to a third preset proportion, the two first archives are merged into a fourth archive.
[0046] wherein, after the time intersection archive group composed of any two first archives with time intersection in the at least two first archives is obtained, the archive management method further comprises:
[0047] judging whether the earliest shooting data or the latest shooting data of the two first archives in the time intersection archive group are both building entrances, or whether the earliest shooting data of one first archive and the latest shooting data of another first archive in the two first archives are both building entrances;
[0048] if no, it is directly determined that the time intersection archive group cannot be merged.
[0049] wherein, after the time intersection archive group composed of any two first archives with time intersection in the at least two first archives is obtained, the archive management method further comprises:
[0050] the earliest shooting data in the first archive with later starting time in the time intersection archive group is inserted into the first archive with earlier starting time according to the shooting time, to form a fourth archive;
[0051] In the fourth profile, when the floor marker of the earliest photographed data is not the same floor marker as the previous photographed data, or is not the same point group, it is directly determined that the time intersection profile group cannot be merged.
[0052] In the fourth profile, when the floor marker of the earliest photographed data is not the same floor marker as the previous photographed data, or is not the same point group, it is directly determined that the time intersection profile group cannot be merged.
[0053] In the fourth profile, when the floor marker of the earliest photographed data is not the same floor marker as the previous photographed data, or is not the same point group, the fourth feature similarity between the earliest photographed data and all the photographed data before it in the fourth profile is calculated.
[0054] The number of photographed data with a fourth feature similarity greater than or equal to a fourth preset similarity threshold is obtained, and a fourth proportion of the number of photographed data to the number of all photographed data before the abnormal photographed data is calculated.
[0055] In the case where the fourth proportion is less than the fourth preset proportion, it is directly determined that the time intersection profile group cannot be merged.
[0056] In the fourth profile, when the floor marker of the earliest photographed data is not the same floor marker as the previous photographed data, or is not the same point group, it is directly determined that the time intersection profile group cannot be merged.
[0057] The latest photographed data in the first profile with an earlier end time in the time intersection profile group is inserted into the first profile with a later end time according to the photographed time, forming a fifth profile.
[0058] In the fourth profile, when the floor marker of the earliest photographed data is not the same floor marker as the previous photographed data, or is not the same point group, it is directly determined that the time intersection profile group cannot be merged.
[0059] In the fourth profile, when the floor marker of the earliest photographed data is not the same floor marker as the previous photographed data, or is not the same point group, it is directly determined that the time intersection profile group cannot be merged.
[0060] All the photographed data in the two first profiles in the time intersection profile group are merged according to the photographed time, forming a sixth profile.
[0061] When the floor marker of the previous photographing data is not the same as the floor marker of the current photographing data, or the shop marker of the previous photographing data is not the same as the shop marker of the current photographing data, it is directly determined that the time intersection file group cannot be merged.
[0062] The application further provides an archive management device, wherein the archive management device comprises a processor and a memory connected to the processor, wherein
[0063] The memory stores program instructions;
[0064] The processor is configured to execute the program instructions stored in the memory to implement the archive management method.
[0065] The application further provides a computer readable storage medium, wherein the storage medium stores program instructions, and the program instructions are executed to implement the archive management method.
[0066] Compared with the prior art, the application has the beneficial effects that: a plurality of initial archives are obtained, each initial archive comprises a plurality of photographing data, and the photographing data comprises photographing position information and photographing time information; the plurality of photographing data in each initial archive is sorted according to the photographing time information to obtain a photographing data list of each initial archive; at least one marker point data is set in the photographing data list; one adjacent photographing data of the at least one marker point data in the photographing data list is obtained, and photographing position difference information and photographing time difference information of the at least one marker point data and the adjacent photographing data are obtained; when the photographing position difference information and the photographing time difference information satisfy a misfile condition, the first archive is split into at least two first archives according to the adjacent photographing data and the at least one marker point data. In the foregoing manner, since the plurality of photographing data is sorted according to the time information, the photographing data is sorted according to the time sequence, the position difference and the time difference of the photographing data adjacent to the marker point data are investigated, the space-time information is deeply utilized to improve the archive management effect, and the effect and the fault tolerance of the archive management are improved. BRIEF DESCRIPTION OF DRAWINGS
[0067] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application.
[0068] Figure 1 A flowchart of a first embodiment of the archive management method provided by the application is shown in FIG. 1;
[0069] Figure 2 A flowchart of a second embodiment of the archive management method provided by the application is shown in FIG. 2;
[0070] Figure 3 is a sub-step flowchart of step S22 of the second embodiment provided by the present application;
[0071] Figure 4 is a flowchart of the third embodiment of the file management method provided by the present application;
[0072] Figure 5 is a flowchart of the fourth embodiment of the file management method provided by the present application;
[0073] Figure 6 is a framework diagram of an embodiment of the file management device provided by the present application;
[0074] Figure 7 is a structure diagram of an embodiment of the computer storage medium provided by the present application. DETAILED DESCRIPTION
[0075] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0076] Please refer to Figure 1 , Figure 1 is a flowchart of the first embodiment of the file management method provided by the present application.
[0077] Step S11: Obtain a plurality of initial files, each of which includes a plurality of shooting data, and the shooting data includes shooting position information and shooting time information.
[0078] It should be noted that the shooting data in the embodiments of the present application is anonymized scene feature data, which does not contain any biological features, facial information or privacy information that can identify personal identity, and does not involve personal privacy.
[0079] Among them, the initial file is obtained by the file management device through a clustering process, and the clustering method is not limited here. In an embodiment of the present application, the shooting data can be divided into different files based on DBSCAN, and clustering is performed based on features, attributes, etc. The files of the same person are preliminarily judged to belong to the same file, and different people are divided into different files. Each initial file includes a plurality of shooting data, and each initial file is a collection of a plurality of shooting data of the same person preliminarily judged, and the shooting data includes shooting position information and shooting time information.
[0080] Step S12: Sort the shooting data in each initial archive according to the shooting time information to obtain a shooting data list of each initial archive.
[0081] Specifically, the archive management device sorts the time information of the shooting data collection in chronological order, sorts the shooting data in the same initial archive, and further screens the spatial relationship between each data and the previous and next data through the time information.
[0082] Step S13: Set at least one marker point data in the shooting data list.
[0083] Specifically, the marker point is a reference point that can be recognized and obtained by the archive management device. The marker point data can be any data in the shooting list, including but not limited to building entrances, floor passageways, store entrances, and general points. Its setting method can be manual annotation of the shooting data in the shooting data list by the user, or scene detection of each frame of shooting data according to the intelligent detection algorithm, and then automatic annotation of the marker point according to the detection result.
[0084] Step S14: Obtain a neighboring shooting data of the at least one marker point data in the shooting data list, and obtain the shooting position difference information and the shooting time difference information of the at least one marker point data and the neighboring shooting data.
[0085] Among them, the neighboring shooting data can be the previous neighboring shooting data of the marker point data, or the next neighboring shooting data of the marker point data.
[0086] Specifically, the shooting position difference information includes but is not limited to any unreasonable position difference information, for example, a person appearing in a certain position without passing through the entrance of the floor where he is located or the entrance of the store where he is located; the shooting time difference information includes but is not limited to any unreasonable time difference information, for example, the data of two people who enter and exit the building without time intersection, which are combined together due to clustering error.
[0087] Step S15: When the shooting position difference information and the shooting time difference information meet the wrong archive condition, split the first archive into at least two first archives according to the neighboring shooting data and the at least one marker point data.
[0088] In the error condition judgment logic, the archive management device marks the data of the building entrance and exit point as a start and end point, and for each start and end point data, the archive management device checks the point of the data after the start and end point data, judges whether it is on the same floor or belongs to the building entrance and exit, and whether the time interval is less than the preset threshold t0 (such as half an hour), if not, the data after this data is removed as a whole, and temporarily listed as another new archive, that is, another first archive.
[0089] The error condition includes but is not limited to that the shooting positions of the two shooting data do not belong to the same floor or the same floor entrance, or the time interval of the shooting time is greater than or equal to the preset interval threshold.
[0090] Optionally, in an embodiment of the present application, the start and end point data is the entrance and exit point data, and the first adjacent shooting data is the previous shooting data of the start and end point data in the shooting data list.
[0091] Further, when the shooting position difference information and the shooting time difference information satisfy the error condition, the first adjacent shooting data and the shooting data before it are taken as a first archive, and the start and end point data and the shooting data after it are taken as another first archive.
[0092] For example, first, the building entrance and exit time is determined. The most common error in the archive is that the data of two people who do not have a time intersection in the building is merged due to clustering error, at which time it can be obviously found that the time is unreasonable. The archive management device marks the data of the building entrance and exit point as a start and end point, and for each start and end point data, the archive management device checks the point of the data before the start and end point data, whether it is on the same floor or belongs to the building entrance and exit, and whether the time interval is less than the preset threshold t0, for example, half an hour, if not, the data before this data is removed as a whole, and temporarily listed as another new archive, that is, another first archive.
[0093] Optionally, in another embodiment of the present application, the start and end point data is the entrance and exit point data, and the second adjacent shooting data is the next shooting data of the start and end point data in the shooting data list.
[0094] When the shooting position difference information and the shooting time difference information satisfy the error condition, the second adjacent shooting data and the shooting data after it are taken as a first archive, and the entrance and exit point data and the shooting data before it are taken as another first archive.
[0095] The application further provides an embodiment. After splitting the initial file into two first files according to the adjacent shooting data and the at least one start-end point data, data feature similarity information of the two first files is calculated, and the first files are combined according to the feature similarity information, so that the accuracy of file management and the purity of the file are further improved.
[0096] For details, see Figure 2 , Figure 2 The second embodiment provided by the application is shown in the flowchart, and the specific steps are as follows:
[0097] Step S20: obtaining a plurality of first files split from the same initial file.
[0098] Specifically, the plurality of first files split from the same initial file in the previous step are temporarily listed as the same person information not belonging to the original initial file, and are temporarily listed as suspicious files, waiting for the file management device to further determine whether they can be combined with the original initial file. Since the space-time relationship does not meet the requirements, it may be a phenomenon of missing shooting or recall failure, so further determination is needed in the next step.
[0099] Step S21: obtaining data feature similarity information of two first files in the plurality of first files.
[0100] The first feature similarity includes but is not limited to one or more of face feature similarity, body feature similarity and torso feature similarity.
[0101] Step S22: combining the two first files into a second file when the data feature similarity information of the two first files meets a preset similarity condition.
[0102] Specifically, whether the data feature similarity information of the two first files meets the preset similarity condition is determined by calculating the similarity between one first file and another first file, and further comparing the similarity with a similarity threshold. For details, see steps S221-S223.
[0103] In an embodiment of the application, a method for calculating feature similarity is provided. For details, see Figure 3 , Figure 3 FIG. 2 is a flowchart of a sub-step of step S22 of the second embodiment provided by the application, and the specific steps are as follows:
[0104] Step S221: calculating a first feature similarity between shooting data of one first file and shooting data of another first file, and obtaining a plurality of groups of first feature similarity of shooting data pairs.
[0105] Specifically, the initial archives excluding several first archives are subjected to similarity checking with the one or more first archives that are just temporarily listed as new archives, i.e., the several first archives split from the same initial archive are subjected to feature similarity checking in step S20.
[0106] Step S222: Obtain the number of groups of the photographed data whose first feature similarity is greater than or equal to the first preset similarity threshold, and the first proportion in the total number of groups.
[0107] Specifically, a threshold value is set, including but not limited to a preset threshold value, and when the similarity to be determined includes multiple features, multiple threshold values T1, T2, T3, etc. are set, and the proportion of the two data feature similarities of the archive and each new archive that are higher than the feature similarity threshold value T1 is counted, i.e., the archive contains N0 data, and the new archives being compared have M0 data, and there are N0*M0 combinations in total, of which K0 groups have a similarity higher than T1 (regardless of which feature, as long as it exceeds the threshold value T1 of the corresponding feature type). The proportion is K0 / (N0*M0). If the proportion is higher than the preset proportion R0, it is considered that the two archives are of the same person.
[0108] Step S223: In the case where the first proportion is greater than or equal to the first preset proportion, the two first archives are merged into a second archive.
[0109] Specifically, it is merged back into the initial archive before it was split out, and the comparison of the new archives that follow is continued, otherwise it is added as a new archive.
[0110] Through the above-mentioned manner, since the decision is based on the comparison result statistics of multiple features, and most errors mainly occur in the second step due to the high similarity of extremely rare data, the high score proportion will be very low under the N0*M0 group data statistics, so even if T1 is not necessarily higher than the threshold value of the second step, the error data can be screened out, and at the same time, only a small amount of false positives can be ensured.
[0111] Further, in order to implement the above-mentioned archive management method and enhance the purity of the archives, the present application further proposes an embodiment to compensate for the possible abnormal archives, including but not limited to the following ways:
[0112] In an embodiment of the present application, a similarity compensation method is proposed. For the possible abnormal data, the similarity of the data with the features of all the data in front of it in the sequence is calculated, such as face similarity, body similarity, etc. The proportion of the similarity of the data with the features of the data in front of it that is higher than T1 is calculated, i.e. there are N1 data in front of it, and M1 of them have a similarity with it, no matter what kind of feature, as long as it exceeds the corresponding type of feature similarity threshold T1, which is higher than T1, and the proportion is M1 / N1. If the proportion is higher than the preset proportion R0, it is considered that the data point is reliable, and the space-time relationship that does not meet the requirements may be only a missed shot or a recall failure of the way point, and the data point is checked and the next data point is checked.
[0113] Through the above two compensations, the data whose space-time relationship does not meet the requirements may be only a missed shot or a recall failure of the way point, and the next data point is checked, so that the purity of the data in the file is greatly improved.
[0114] In another embodiment of the present application, a post-compensation method is also proposed. After all the data in the file are checked, the temporarily removed data is reprocessed. According to the time sequence, it is judged whether the data point and the point of the next data (excluding the removed data point) meet the requirements of the same floor or the floor entrance; further, if it belongs to a certain point group, it also needs to meet the requirement that the point of the next data also belongs to the point group. If it does not meet the requirement, the data and all the data after it (excluding the removed point) are calculated in a similar manner as the similarity compensation method described above. If it does not meet the requirement, the data point is removed from the file. If at least one of the point judgment of the next data or the similarity compensation judgment is met, the data point is put back into the file.
[0115] Optionally, the similarity compensation method and the post-compensation method in the above embodiments can be applied to any step of the file in the present application, i.e. any file can be processed, and the position and sequence of the processed file are not limited.
[0116] After the first file is split into at least two first files according to the adjacent shooting data and the at least one marker point data in the present application, an embodiment is proposed, which is described in detail in Figure 4 , Figure 4 The flowchart of the third embodiment of the file management method provided by the present application is shown in the figure, and the specific steps are as follows:
[0117] Step S30: For each first file, mark the floor of all the shooting data in the first file, and obtain the floor mark and store mark of each shooting data.
[0118] The positions of the floor markers include, but are not limited to, floor entrances and connecting positions between floors, and the store markers include, but are not limited to, store entrances and point groups inside the stores.
[0119] Step S31: For each shooting data, if the floor marker of the previous shooting data is not the same as the floor marker of the current shooting data, or the store marker of the previous shooting data is not the same as the store marker of the current shooting data, the current shooting data is marked as abnormal shooting data.
[0120] Specifically, for each shooting data, according to the floor marker thereof, it is checked whether the point of the previous shooting data is on the same floor as the shooting data or is a floor entrance. If not, the shooting data is temporarily determined as possible abnormal data. Further, if the point of the shooting data belongs to a point group, it is required that the point of the previous shooting data also belongs to the point group. If not, the shooting data is temporarily determined as possible abnormal data. If a shooting data is not temporarily determined as possible abnormal data, it is considered that the checking is passed, and the checking of the next shooting data is continued. In particular, the first shooting data in the sequence does not need to be checked.
[0121] Step S32: The abnormal shooting data is excluded from all the first archives to obtain a third archive.
[0122] Specifically, in an embodiment, the application proposes a specific step of obtaining the third archive, and the steps are as follows:
[0123] The second feature similarity between the abnormal shooting data and all the previous shooting data in each first archive is calculated. The number of shooting data with a second feature similarity greater than or equal to a second preset similarity threshold is obtained, and a second proportion of the number of the shooting data to the number of all the previous shooting data of the abnormal shooting data is calculated. In the case where the second proportion is greater than or equal to a second preset proportion, the abnormal shooting data is retained. In the case where the second proportion is less than the second preset proportion, the abnormal shooting data is excluded.
[0124] The calculation method of the feature similarity in this embodiment is the same as that in steps S221-S223, which will not be described here.
[0125] The third archive is composed of the remaining shooting data of each first archive.
[0126] Through the above-mentioned archive management method based on the space-time rule, the technical problem of how to use space-time information to improve the archive management effect in buildings, such as office buildings, urban commercial complexes, and hotel apartments, is solved, and the accuracy of archive management is improved.
[0127] After obtaining the third file through step S32, the application further proposes an embodiment to further determine whether the abnormal shooting data excluded from the first file is the same floor mark as the shooting data after the first file, or whether it is the same point group;
[0128] If so, the abnormal shooting data is put back into the original first file.
[0129] Through the above method, it is further determined whether the abnormal shooting data is excluded due to the mistake in the exclusion process, thereby further improving the purity of the file.
[0130] In the application, any existing file can perform steps S30-S32. In an embodiment of the application, the second file obtained in step S23 can perform steps S30-S32. In another embodiment of the application, at least one first file obtained in step S15 can perform steps S30-S32. In other embodiments of the application, the initial file that does not meet step S15 can directly perform steps S30-S32.
[0131] Based on the above file management method, after excluding the wrong file, in order to further improve the accuracy of the file management method, the application further proposes a filing method, please refer to Figure 5 , Figure 5 is a flowchart of the fourth embodiment provided by the application. After splitting the first file into at least two first files according to adjacent shooting data and at least one mark point data, the specific steps are as follows:
[0132] Step S40: obtaining any two first files with time intersection in the at least two first files to form a time intersection file group.
[0133] Specifically, for each file, the earliest and latest data time is counted to form a time period. The files with time intersection are calculated to form a time intersection file group.
[0134] Further, after obtaining any two first files with time intersection in the at least two first files to form a time intersection file group, whether the time intersection file group can be merged includes but is not limited to the following ways, wherein the application does not limit the order and data of the following ways.
[0135] In an embodiment provided by the application, after obtaining any two first files with time intersection in the at least two first files to form a time intersection file group, it is determined that the earliest shooting data or the latest shooting data of the two first files in the time intersection file group is the building entrance, or the earliest shooting data of one of the two first files and the latest shooting data of the other first file is the building entrance.
[0136] If not, it is directly determined that the time intersection file group cannot be merged.
[0137] In an embodiment provided by the present application, when the earliest shooting data and the latest shooting data of the first file of any one or two of the time intersection file group are not the building entrance, the judgment logic of the above earliest shooting data and / or latest shooting data is directly skipped, and the subsequent spatiotemporal rule judgment logic is entered.
[0138] In the above manner, the split files that may meet the spatiotemporal rule with other files or meet the above screening rule after some multi-file integration into the original file are re-integrated, the purity of the file is improved, and the accuracy of the file management method is improved.
[0139] In an embodiment provided by the present application, after obtaining the time intersection file group composed of any two first files with time intersection in the at least two first files, the earliest shooting data in the first file with later start time in the time intersection file group is inserted into the first file with earlier start time according to the shooting time, to form a fourth file. When the floor mark of the earliest shooting data is not the same floor mark as the previous shooting data, or is not the building entrance, or is not the same point group in the fourth file, it is directly determined that the time intersection file group cannot be merged.
[0140] In the fourth file, when the floor mark of the earliest shooting data is not the same floor mark as the previous shooting data, or is not the same point group, the fourth feature similarity of the earliest shooting data and all the previous shooting data in the fourth file is calculated. The number of shooting data with fourth feature similarity greater than or equal to a fourth preset similarity threshold is obtained, and the fourth proportion of the number of shooting data to the number of all the previous shooting data of the abnormal shooting data is calculated. When the fourth proportion is less than the fourth preset proportion, it is directly determined that the time intersection file group cannot be merged.
[0141] For two archives with time intersection meeting the aforementioned multiple link troubleshooting, the similarity thereof is troubleshooted. The similarity here is required to be lower than steps S221-S223. A threshold T2 (if there are multiple features, there are multiple thresholds corresponding thereto) is set, which represents a weaker confidence, but is sufficient to distinguish whether two randomly selected persons are the same person. The similarity of randomly combined persons under a large amount of data will be very low, so the setting of T2 is not harsh. The proportion of two-by-two data features between the two archives is counted, that is, the two archives have N2 and M2 data respectively, and there are N2*M2 combinations in total, of which K2 are higher than T2 (regardless of which feature, as long as the threshold T2 of the corresponding feature is exceeded). The proportion is K2 / (N2*M2). If the proportion is higher than a preset proportion, it is considered that the two archives are the same person, and the two archives are combined.
[0142] In the above manner, the archives that have been split out and may meet the space-time rules with other archives, or can meet the above troubleshooting rules after some multi-archives are combined into the archives to which they originally belong, are recombined, the purity of the archives is improved, the accuracy of the archive management method is improved, and the accuracy of the archive management is further improved by setting a weaker confidence similarity threshold.
[0143] In an embodiment provided by the application, after any two first archives with time intersection in the at least two first archives are combined to form a time intersection archive group, the latest shooting data in the first archive with the earlier ending time in the time intersection archive group is inserted into the first archive with the later ending time according to the shooting time, to form a fifth archive. When the floor marker of the latest shooting data is not the same floor marker as the previous shooting data, or is not a floor entrance and exit, or is not the same point group as the previous shooting data in the fifth archive, it is directly determined that the time intersection archive group cannot be combined.
[0144] In an embodiment provided by the application, after any two first archives with time intersection in the at least two first archives are combined to form a time intersection archive group, all the shooting data in the two first archives in the time intersection archive group are combined according to the shooting time, to form a sixth archive.
[0145] When the floor marker of the previous shooting data is not the same floor marker as the current shooting data, or the store marker of the previous shooting data is not the same point group as the store marker of the current shooting data in the sixth archive, it is directly determined that the time intersection archive group cannot be combined.
[0146] Through the above manner, the split archives that may meet the space-time rules with other archives or meet the above checking rules after some multi-archives are combined into the archives to which the multi-archives originally belong are recombined, the purity of the archives is improved, and the precision of the archive management method is improved.
[0147] After the archives are determined to be combined according to any of the above embodiments, the following steps are further performed:
[0148] Step S41: Calculate a third feature similarity between shooting data of a first archive in the time intersection archive group and shooting data of another first archive, to obtain third feature similarities of a plurality of groups of shooting data pairs.
[0149] Step S42: Obtain a third proportion of the number of groups of shooting data pairs whose first feature similarity is greater than or equal to a first preset similarity threshold in the total number of groups.
[0150] Step S43: In a case where the third proportion is greater than or equal to a third preset proportion, combine the two first archives into a fourth archive.
[0151] Specifically, steps S41-S43 are calculated in the same manner as steps S221-S223, and details are not repeated here.
[0152] Through the above manner, the split archives that may meet the space-time rules with other archives or meet the above checking rules after some multi-archives are combined into the archives to which the multi-archives originally belong are recombined, the purity of the archives is improved, and the precision of the archive management method is improved.
[0153] To implement the image detection method in the above embodiments, the application further provides an archive management device, please refer to Figure 6 , Figure 6 is a frame schematic diagram of an embodiment of the archive management device provided by the application.
[0154] The archive management device 500 of the embodiment of the application comprises a processor 51, a memory 52, an input and output device 53, and a bus 54.
[0155] The processor 51, the memory 52, and the input and output device 53 are respectively connected to the bus 54, the memory 52 stores program data, and the processor 51 is used to execute the program data to implement the archive management method described in the above embodiments.
[0156] In the embodiments of the present application, the processor 51 can also be referred to as a CPU (Central Processing Unit). The processor 51 can be an integrated circuit chip having a processing capability of signals. The processor 51 can also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 51 can also be any conventional processor or the like.
[0157] The present application also provides a computer storage medium, please continue to refer to Figure 7 , Figure 7 FIG. 6 is a structural schematic diagram of an embodiment of a computer storage medium provided by the present application. The computer storage medium 600 stores a computer program 61. When the computer program 61 is executed by a processor, the computer program 61 is used to implement the archive management method of the above-mentioned embodiments.
[0158] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, the software functional units can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or all or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0159] The above only describes the embodiments of the present application, and does not limit the patent scope of the present application. The equivalent structure or equivalent flow transformation made by the content of the present application specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An archive management method characterized by, The archive management method comprises: obtaining a plurality of initial archives, each initial archive comprising a plurality of shooting data, the shooting data comprising shooting position information and shooting time information; sorting the plurality of shooting data in each initial archive according to the shooting time information to obtain a shooting data list of each initial archive; setting at least one marker point data in the shooting data list; obtaining a neighboring shooting data of the at least one marker point data in the shooting data list, and obtaining shooting position difference information and shooting time difference information of the at least one marker point data and the neighboring shooting data; when the shooting position difference information and the shooting time difference information satisfy a wrong archive condition, splitting the initial archive into at least two first archives according to the neighboring shooting data and the at least one marker point data.
2. The archive management method of claim 1, wherein: the marker point data is an entrance and exit point data, and a first neighboring shooting data is a previous shooting data of the entrance and exit point data in the shooting data list; when the shooting position difference information and the shooting time difference information satisfy the wrong archive condition, splitting the initial archive into two first archives according to the neighboring shooting data and the at least one marker point data comprises: when the shooting position difference information and the shooting time difference information satisfy the wrong archive condition, taking the first neighboring shooting data and the previous shooting data thereof as one first archive, and taking the entrance and exit point data and the subsequent shooting data thereof as another first archive.
3. The archive management method of claim 1 or 2, wherein: the marker point data is an entrance and exit point data, and a second neighboring shooting data is a subsequent shooting data of the entrance and exit point data in the shooting data list; when the shooting position difference information and the shooting time difference information satisfy the wrong archive condition, splitting the initial archive into two first archives according to the neighboring shooting data and the at least one marker point data comprises: when the shooting position difference information and the shooting time difference information satisfy the wrong archive condition, taking the second neighboring shooting data and the subsequent shooting data thereof as one first archive, and taking the entrance and exit point data and the previous shooting data thereof as another first archive.
4. The archive management method of claim 1, wherein: the wrong archive condition is that the shooting positions of two shooting data do not belong to the same floor or the same floor entrance and exit, or the time interval of the shooting time is greater than or equal to a preset interval threshold.
5. The archive management method of claim 1, wherein: after splitting the initial archive into two first archives according to the neighboring shooting data and the at least one marker point data, the archive management method further comprises: obtaining a plurality of first archives split from the same initial archive; obtaining data feature similarity information of two first archives in the plurality of first archives; In a case where the data feature similarity information of two first archives meets a preset similarity condition, the two first archives are merged into a second archive.
6. The archive management method of claim 5, wherein the data feature similarity information of each pair of first archives is obtained by: calculating a first feature similarity between the shooting data of one first archive and the shooting data of another first archive, and obtaining a first feature similarity of a plurality of groups of shooting data pairs; in a case where the first feature similarity of the two first archives meets a preset similarity condition, the two first archives are merged into a second archive.
7. The archive management method of claim 6, wherein the first feature similarity includes one or more of a face feature similarity, a human body feature similarity, and a torso feature similarity.
8. The archive management method of claim 1, wherein after the first archive is split into at least two first archives according to the adjacent shooting data and the at least one marker point data, the archive management method further comprises: performing floor marking on all shooting data in the first archive to obtain floor marking and store marking of each shooting data; iterating through each shooting data, and in a case where the floor marking of a previous shooting data is not the same as the floor marking of a current shooting data, or the store marking of the previous shooting data is not the same as the store marking of the current shooting data, marking the current shooting data as an abnormal shooting data; excluding the abnormal shooting data from all the first archives to obtain a third archive.
9. The archive management method of claim 8, wherein the excluding the abnormal shooting data from all the first archives to obtain a third archive comprises: calculating a second feature similarity between the abnormal shooting data and all shooting data before the abnormal shooting data in each first archive; obtaining a number of shooting data whose second feature similarity is greater than or equal to a second preset similarity threshold, and calculating a second proportion of the number of shooting data to a number of all shooting data before the abnormal shooting data; in a case where the second proportion is greater than or equal to a second preset proportion, retaining the abnormal shooting data; in a case where the second proportion is less than the second preset proportion, excluding the abnormal shooting data; composing the third archive from the remaining shooting data of each first archive.
10. The archive management method of claim 8 or 9, wherein after the excluding the abnormal shooting data from all the first archives to obtain a third archive, the archive management method further comprises: judging whether the shooting data after the first archive in which the abnormal shooting data is excluded is the same as the abnormal shooting data in floor marking or the same in a point site group. If yes, the abnormal shooting data is put back into the original first file.
11. The file management method of claim 1, further comprising: after the first file is split into at least two first files according to the adjacent shooting data and the at least one marked point data, acquiring a time intersection file group composed of any two first files having time intersection in the at least two first files; calculating a third feature similarity between shooting data of one first file and shooting data of another first file in the time intersection file group, and acquiring third feature similarities of a plurality of shooting data pairs; acquiring a number of shooting data pairs whose first feature similarity is greater than or equal to a first preset similarity threshold, and calculating a third proportion in the total number of groups; in a case where the third proportion is greater than or equal to a third preset proportion, merging the two first files into a fourth file.
12. The file management method of claim 11, further comprising: after the time intersection file group is acquired, judging whether the earliest shooting data or the latest shooting data of the two first files in the time intersection file group are both building entrances, or whether the earliest shooting data of one first file and the latest shooting data of another first file are both building entrances; if not, directly determining that the time intersection file group cannot be merged.
13. The file management method of claim 11 or 12, further comprising: after the time intersection file group is acquired, inserting the earliest shooting data in the first file with later starting time in the time intersection file group into the first file with earlier starting time according to shooting time, to form a fourth file; in a case where the floor mark of the earliest shooting data is not the same floor mark as the previous shooting data, or is not the same point group as the previous shooting data in the fourth file, directly determining that the time intersection file group cannot be merged.
14. The file management method of claim 13, wherein the case where the floor mark of the earliest shooting data is not the same floor mark as the previous shooting data, or is not the same point group as the previous shooting data in the fourth file, comprises: in the case where the floor mark of the earliest shooting data is not the same floor mark as the previous shooting data, or is not the same point group as the previous shooting data in the fourth file, calculating fourth feature similarities between the earliest shooting data and all previous shooting data in the fourth file; acquiring a number of shooting data whose fourth feature similarity is greater than or equal to a fourth preset similarity threshold, and calculating a fourth proportion between the number of shooting data and the number of all previous shooting data of the abnormal shooting data; in a case where the fourth proportion is less than a fourth preset proportion, directly determining that the time intersection file group cannot be merged. 15. The archive management method of claim 11 or 12, wherein, after the step of obtaining the time-intersected archive group by grouping any two of the at least two first archives having time intersection, the archive management method further comprises: inserting the latest shooting data in the first archive with earlier end time in the time-intersected archive group into the first archive with later end time according to shooting time, to form a fifth archive; when the floor marker of the latest shooting data is not the same as the floor marker of the previous shooting data, or is not the same point group as the previous shooting data, in the fifth archive, directly determining that the time-intersected archive group cannot be merged.
16. The archive management method of claim 11 or 12, wherein, after the step of obtaining the time-intersected archive group by grouping any two of the at least two first archives having time intersection, the archive management method further comprises: merging all shooting data in the two first archives in the time-intersected archive group according to shooting time, to form a sixth archive; traversing each shooting data in the sixth archive, when the floor marker of the previous shooting data is not the same as the floor marker of the current shooting data, or the shop marker of the previous shooting data is not the same point group as the current shooting data, directly determining that the time-intersected archive group cannot be merged. The archive management device comprises a processor and a memory connected to the processor, wherein: the memory stores program instructions; 17. An archival management apparatus characterized by comprising: the processor is configured to execute the program instructions stored in the memory to implement the archive management method of any one of claims 1-16. The computer readable storage medium has program instructions, and the program instructions are executed to implement the archive management method of any one of claims 1-16. 18. A computer-readable storage medium, characterized in that,
Citation Information
Patent Citations
Information processing method and device and storage medium
CN110348347A
Image archiving method and device, electronic equipment and storage medium
CN112487221A