Method and apparatus for data hierarchical clustering

By sorting the feature vector set by confidence and activity, this method solves the key vehicle management problem in the existing technology, realizes hierarchical filtering and rapid data aggregation, and improves the accuracy and efficiency of data aggregation.

CN116340261BActive Publication Date: 2026-03-24ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-13
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies are insufficient for effectively managing key vehicles, especially those from other areas and those transiting through the region. The large volume and inconsistent quality of video capture data make it difficult for traffic police departments to quickly establish effective vehicle files.

Method used

By identifying multiple feature vector sets and sorting them based on confidence and activity, hierarchical filtering and rapid data aggregation can be achieved.

Benefits of technology

It improves the accuracy and efficiency of data aggregation and archiving, and can more intuitively recommend important feature vector sets and target data, helping traffic police departments to quickly establish files for key vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116340261B_ABST
    Figure CN116340261B_ABST
Patent Text Reader

Abstract

The application provides a data hierarchical clustering method and device, which comprises the following steps: determining a plurality of feature vector sets, wherein the i-th feature vector set in the plurality of feature vector sets corresponds to the i-th confidence and the i-th activity; wherein the activity represents the number of occurrences of a target within a preset time length, and the confidence represents the confidence of the data corresponding to the target; the i-th feature vector set comprises at least one feature vector, the at least one feature vector corresponds to at least one target in a one-to-one manner, the confidence of each feature vector in the at least one feature vector is the i-th confidence, and the activity of each feature vector in the at least one feature vector is the i-th activity; wherein i is a positive integer. The plurality of feature vector sets are sorted according to the confidence and the activity corresponding to the plurality of feature vector sets. The above method is used to filter the data in layers, so that the data can be quickly clustered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to, but is not limited to, the field of data analysis technology, and particularly to a method and apparatus for hierarchical data aggregation. Background Technology

[0002] With the rapid development of urbanization, road traffic safety issues are becoming increasingly serious. Traffic violations by key vehicles, including passenger vehicles, dangerous goods transport vehicles, freight vehicles, and minivans, are serious and pose significant traffic safety hazards. Among them, passenger vehicles, dangerous goods transport vehicles, freight vehicles, and minivans refer to highway passenger vehicles, tourist passenger vehicles, dangerous goods transport vehicles, freight vehicles, and minivans. Currently, the supervision of key vehicles faces problems such as difficulty in supervision and high consumption of police resources.

[0003] Existing technologies for managing key vehicle records are immature, often relying on comparing real-time captured data with pre-entered data for identification and control. However, due to issues such as missing vehicle registration information, traffic police departments struggle to effectively identify and manage key vehicles, especially those from out of town or transiting through areas. Therefore, extracting relevant vehicle information from video capture images and quickly establishing records has become a viable approach. Given the massive volume and varying quality of images captured by electronic police and checkpoint cameras, finding effective methods to quickly extract more relevant data from this vast amount of information to help traffic police departments rapidly compile records is a crucial issue.

[0004] In summary, there is an urgent need for a method and device for hierarchical filtering of data to achieve rapid data aggregation. Summary of the Invention

[0005] This invention provides a method and apparatus for hierarchical data aggregation, which is used to perform hierarchical filtering on data and achieve rapid data aggregation.

[0006] Firstly, this application provides a method for hierarchical data archiving, the method comprising:

[0007] Multiple feature vector sets are defined, wherein the i-th feature vector set in the multiple feature vector sets corresponds to the i-th confidence level and the i-th activity level; wherein the activity level represents the number of times the target appears within a preset time period, and the confidence level represents the confidence level of the data corresponding to the target; the i-th feature vector set includes at least one feature vector, the at least one feature vector corresponds one-to-one with at least one target, the confidence level of each feature vector in the at least one feature vector is the i-th confidence level, and the activity level of each feature vector in the at least one feature vector is the i-th activity level; wherein i is a positive integer;

[0008] The multiple feature vector sets are sorted according to their respective confidence and activity levels.

[0009] The above design comprehensively considers two indicators: the confidence level of the target-corresponding data and the activity level of the target-corresponding data. Ranking the feature vector sets based on confidence level and activity level yields a more comprehensive and accurate ranking result. Furthermore, this design, by ranking multiple feature vector sets, can more intuitively recommend important feature vector sets. Since each feature vector set includes at least one feature vector, and at least one feature vector corresponds one-to-one with at least one target, this design can also recommend data for important targets corresponding to important feature vector sets.

[0010] In one possible design, the multiple feature vector sets are sorted according to their respective confidence and activity levels, including:

[0011] Based on the confidence and activity corresponding to the multiple feature vector sets, determine the total number of first vectors and the total number of second vectors corresponding to the multiple feature vector sets respectively;

[0012] The multiple feature vector sets are sorted according to the total number of first vectors and the total number of second vectors corresponding to the multiple feature vector sets respectively;

[0013] exist or When, the j-th eigenvector set is the dominated solution of the i-th eigenvector set, where P i For the i-th confidence level, a i For the i-th activity level, the j-th feature vector set is one of the multiple feature vector sets, where j is a positive integer not equal to i, and P j Let a be the confidence level corresponding to the j-th feature vector set. j The activity level corresponding to the j-th feature vector set;

[0014] The total number of first vectors corresponding to the i-th feature vector set is the total number of feature vectors included in the feature vector sets corresponding to each dominated solution of the i-th feature vector set;

[0015] exist or When the j-th eigenvector set is the dominant solution of the i-th eigenvector set;

[0016] The total number of second vectors corresponding to the i-th eigenvector set is the total number of eigenvectors included in the eigenvector sets corresponding to each dominating solution of the i-th eigenvector set.

[0017] In one possible design, the multiple feature vector sets are sorted according to the total number of first vectors and the total number of second vectors corresponding to the multiple feature vector sets, including:

[0018] The multiple feature vector sets are arranged in ascending order based on the total number of second vectors corresponding to the multiple feature vector sets.

[0019] In one possible design, for at least two feature vector sets in which the total number of second vectors is equal, the at least two feature vector sets are arranged in descending order according to the total number of first vectors corresponding to the at least two feature vector sets.

[0020] In one possible design, for at least two feature vector sets in which the total number of the first vector and the total number of the second vector are equal, the at least two feature vector sets are arranged in descending order according to the confidence levels corresponding to the at least two feature vector sets.

[0021] In one possible design, multiple sets of feature vectors are determined, including:

[0022] Acquire multiple data points, each including the identifier information of the target corresponding to the data, the confidence level of the data, and the acquisition time;

[0023] The plurality of feature vector sets are determined based on the plurality of data; wherein, the first feature vector is one of the i-th feature vector sets, the first feature vector corresponds to a first target, the first target is one of the at least one target, the confidence level corresponding to the first feature vector is the maximum confidence level in the data set corresponding to the first target, and the data set corresponding to the first target includes data with identification information corresponding to the first target;

[0024] The activity level corresponding to the first feature vector is determined based on the collection time of each data in the dataset corresponding to the first target and the preset duration;

[0025] In one possible design, each data also includes a target type. After acquiring the plurality of data, a target data set is selected from the plurality of data based on the target type included in each data. The target data set includes data of at least one preset target type.

[0026] Determining the multiple feature vector sets based on the multiple data includes:

[0027] The plurality of feature vector sets are determined based on the target data set.

[0028] Using the above design, data corresponding to key targets can be filtered from multiple acquired data based on the target type corresponding to the key targets. Furthermore, after determining multiple feature vector sets based on the target data set corresponding to the key targets using the above design, the sorting results of these feature vector sets can also be used to obtain the sorting results of the data corresponding to the key targets, thereby achieving the filtering of data corresponding to the key targets.

[0029] In one possible design, when the data containing the identification information corresponding to the first target includes at least two target types, the target type that appears most frequently is taken as the target type of the first target.

[0030] Secondly, this application provides a data hierarchical archiving device, which is a server or a chip within a server. The device includes a processing unit and a transceiver unit: the processing unit invokes the transceiver unit to execute:

[0031] Multiple feature vector sets are defined, wherein the i-th feature vector set in the multiple feature vector sets corresponds to the i-th confidence level and the i-th activity level; wherein the activity level represents the number of times the target appears within a preset time period, and the confidence level represents the confidence level of the data corresponding to the target; the i-th feature vector set includes at least one feature vector, the at least one feature vector corresponds one-to-one with at least one target, the confidence level of each feature vector in the at least one feature vector is the i-th confidence level, and the activity level of each feature vector in the at least one feature vector is the i-th activity level; wherein i is a positive integer;

[0032] The multiple feature vector sets are sorted according to their respective confidence and activity levels.

[0033] In one possible design, the processing unit is configured to, when sorting the plurality of feature vector sets according to the confidence and activity corresponding to the plurality of feature vector sets respectively, determine the total number of first vectors and the total number of second vectors corresponding to the plurality of feature vector sets respectively based on the confidence and activity corresponding to the plurality of feature vector sets respectively.

[0034] The processing unit is further configured to sort the plurality of feature vector sets according to the total number of first vectors and the total number of second vectors corresponding to the plurality of feature vector sets respectively;

[0035] exist or When, the j-th eigenvector set is the dominated solution of the i-th eigenvector set, where P i For the i-th confidence level, a iFor the i-th activity level, the j-th feature vector set is one of the multiple feature vector sets, where j is a positive integer not equal to i, and P j Let a be the confidence level corresponding to the j-th feature vector set. j The activity level corresponding to the j-th feature vector set;

[0036] The total number of first vectors corresponding to the i-th feature vector set is the total number of feature vectors included in the feature vector sets corresponding to each dominated solution of the i-th feature vector set;

[0037] exist or When the j-th eigenvector set is the dominant solution of the i-th eigenvector set;

[0038] The total number of second vectors corresponding to the i-th eigenvector set is the total number of eigenvectors included in the eigenvector sets corresponding to each dominating solution of the i-th eigenvector set.

[0039] In one possible design, the processing unit is used to sort the multiple feature vector sets according to the total number of first vectors and the total number of second vectors corresponding to the multiple feature vector sets, and to arrange the multiple feature vector sets in ascending order according to the total number of second vectors corresponding to the multiple feature vector sets.

[0040] In one possible design, for at least two feature vector sets in which the total number of second vectors is equal among the plurality of feature vector sets, the processing unit is configured to sort the at least two feature vector sets in descending order according to the total number of first vectors corresponding to the at least two feature vector sets.

[0041] In one possible design, for at least two feature vector sets in which the total number of first vectors and the total number of second vectors are equal, the processing unit is configured to sort the at least two feature vector sets in descending order according to the confidence levels corresponding to the at least two feature vector sets.

[0042] In one possible design, the transceiver unit is used to acquire multiple data when the processing unit determines multiple feature vector sets, each data including the identification information of the target corresponding to the data, the confidence level of the data, and the acquisition time;

[0043] The processing unit is configured to determine the plurality of feature vector sets based on the plurality of data; wherein, the first feature vector is one of the i-th feature vector sets, the first feature vector corresponds to a first target, the first target is one of the at least one target, the confidence level corresponding to the first feature vector is the maximum confidence level in the data set corresponding to the first target, and the data set corresponding to the first target includes data with identification information corresponding to the first target;

[0044] The activity level corresponding to the first feature vector is determined by the processing unit based on the collection time of each data in the dataset corresponding to the first target and the preset duration;

[0045] In one possible design, each data also includes a target type. After the transceiver unit acquires the plurality of data, the processing unit is configured to filter out a target data set from the plurality of data according to the target type included in each data, wherein the target data set includes data of at least one preset target type; the processing unit determines the plurality of feature vector sets according to the plurality of data, including: the processing unit determines the plurality of feature vector sets according to the target data set.

[0046] In one possible design, the processing unit is configured to, when the data containing the identification information corresponding to the first target includes at least two target types, select the target type that appears most frequently as the target type of the first target.

[0047] The technical effects of the device in the second aspect can be seen in the technical effects of the different implementation methods in the first aspect, and will not be repeated here.

[0048] Thirdly, this application also provides an apparatus. This apparatus can perform the above-described method design. The apparatus may be a chip or circuit capable of performing the functions corresponding to the above-described method, or a device including the chip or circuit.

[0049] In one possible implementation, the device includes: a memory for storing computer-executable program code; and a processor coupled to the memory. The program code stored in the memory includes instructions that, when executed by the processor, cause the device or a device equipped with the device to perform any of the methods described above.

[0050] The device may also include a communication interface, which may be a transceiver, or, if the device is a chip or circuit, the communication interface may be the chip's input / output interface, such as input / output pins.

[0051] In one possible design, the device includes corresponding functional units, each used to implement the steps in the above method. The functions can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more units corresponding to the functions described above.

[0052] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on a device, executes the method described in any of the above possible designs.

[0053] Furthermore, the technical effects of any of the implementation methods in the third to fourth aspects can be found in the technical effects of different implementation methods in the first aspect, and will not be repeated here. Attached Figure Description

[0054] Figure 1 A flowchart outlining a method for hierarchical data archiving provided in an embodiment of the present invention;

[0055] Figure 2 A schematic diagram of a feature vector set provided in an embodiment of the present invention;

[0056] Figure 3 A schematic diagram of another feature vector set provided in an embodiment of the present invention;

[0057] Figure 4 A schematic diagram of another feature vector set provided in an embodiment of the present invention;

[0058] Figure 5 This is a schematic diagram of the structure of a device provided in an embodiment of the present invention;

[0059] Figure 6 This is a schematic diagram of another device provided in an embodiment of the present invention. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0061] The application scenarios described in the embodiments of this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. Those skilled in the art will understand that with the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems. In the description of this invention, unless otherwise stated, "multiple" means two or more.

[0062] This invention provides a method and apparatus for hierarchical data aggregation, which is used to perform hierarchical filtering on data and achieve rapid data aggregation.

[0063] Figure 1 A flowchart outlining a data hierarchical archiving method provided in this embodiment of the invention includes the following steps:

[0064] S11. Determine multiple sets of feature vectors.

[0065] A set of feature vectors may include at least one feature vector, and at least one feature vector corresponds one-to-one with at least one target.

[0066] Each of the identified feature vector sets corresponds to a confidence level and an activity level. Specifically, the activity level represents the number of times the target appears within a preset time period, and the confidence level represents the confidence level of the data corresponding to the target.

[0067] For example, in the embodiments of the present invention, the target may refer to a vehicle, or other objects or operations, etc. The present invention does not limit this, and the following description only uses a vehicle as an example.

[0068] For example, if there exists an i-th feature vector set that is one of a plurality of determined feature vector sets, then the i-th feature vector set corresponds to the i-th confidence level and the i-th activity level, where i is a positive integer. The i-th feature vector set includes at least one feature vector, and each feature vector corresponds one-to-one with at least one target. The confidence level of each feature vector in the at least one feature vector included in the i-th feature vector set is the i-th confidence level, and the activity level of each feature vector in the at least one feature vector included in the i-th feature vector set is the i-th activity level.

[0069] In one possible implementation, determining multiple sets of feature vectors can be achieved, but is not limited to, the following methods:

[0070] First, multiple data sets need to be acquired. Each data set includes the target identification information, the confidence level of the data, and the collection time.

[0071] For example, when the target is a vehicle, the identification information can be license plate information. For instance, data on passing vehicles can be acquired using video detection equipment, such as a camera. Each data point includes license plate information. Specifically, when acquiring data on passing vehicles using video detection equipment, the license plate information of the passing vehicles is identified and obtained. The identified and obtained license plate information must meet license plate rules; for example, the license plate rule might be that the first character of the identified and obtained license plate is a Chinese character.

[0072] For example, data on past targets can be acquired through video detection equipment. When the same target appears multiple times, the video detection equipment can acquire data on the target multiple times. Each acquired data includes the corresponding acquisition time and the confidence level of the data.

[0073] For example, if the first target appears three times—on Monday, Thursday, and Sunday—within a preset one-week period, the video detection device can acquire data on the first target three times. This means there are three data sets containing the identification information corresponding to the first target. Specifically, the data sets containing the identification information corresponding to the first target can include: first data, second data, and third data. The confidence level of the first data is the first confidence level, and the confidence level of the first data was collected on Monday. The confidence level of the second data is the second confidence level, and the confidence level of the first data was collected on Thursday. The confidence level of the third data is the third confidence level, and the confidence level of the third data was collected on Sunday.

[0074] In addition, each data may also include a target type. When the data with identification information corresponding to the first target includes at least two target types, the target type that appears most frequently is taken as the target type of the first target.

[0075] For example, when the target is a vehicle and the identification information is license plate information, the target type is vehicle type. If the first data includes a vehicle type of van, the second data includes a vehicle type of van, and the third data includes a vehicle type of sedan, then the van is taken as the vehicle type of the first vehicle.

[0076] Furthermore, multiple sets of feature vectors are determined based on multiple data points.

[0077] In one possible implementation, each data also includes a target type. After acquiring multiple data, a target data set can be selected from the multiple data based on the target type included in each data. The target data set includes data of at least one preset target type. Furthermore, a set of feature vectors corresponding to the selected target data set can be determined based on the data included in the selected target data set.

[0078] For example, when the target is a vehicle, the identification information is license plate information, and the target type is vehicle type, each of the acquired data sets includes a vehicle type, such as a sedan, a hazardous materials transport vehicle, or a van. Based on the vehicle type included in each data set, a target data set is selected from the multiple data sets. For example, if the preset vehicle types are hazardous materials transport vehicles and vans, a target data set is selected from the multiple data sets based on these preset vehicle types. This target dataset includes data corresponding to the vehicle types of hazardous materials transport vehicles and vans.

[0079] For example, at least two feature vectors can be determined first based on multiple data or target datasets, along with the confidence level and activity level corresponding to each feature vector, and then feature vectors with the same confidence level and the same activity level can be combined into a feature vector set.

[0080] Each feature vector is determined by data with the same identification information. Taking the first feature vector as an example, the first feature vector corresponds to the first target. The first feature vector is determined based on the data set corresponding to the first target. The data set corresponding to the first target includes data with the identification information corresponding to the first target. The confidence level corresponding to the first feature vector is the maximum confidence level in the data set corresponding to the first target. The activity level corresponding to the first feature vector is determined based on the collection time and preset duration of each data in the data set corresponding to the first target.

[0081] For example, the data set corresponding to the first target includes data with identification information corresponding to the first target. Specifically, referring to the above example, assume that the data set corresponding to the first target includes first data, second data, and third data. The confidence level of the first data is the first confidence level, and the confidence level of the first data was collected on Monday. The confidence level of the second data is the second confidence level, and the confidence level of the first data was collected on Thursday. The confidence level of the third data is the third confidence level, and the confidence level of the third data was collected on Sunday. For example, assume that there exists a third confidence level greater than the second confidence level, and the second confidence level greater than the first confidence level.

[0082] The confidence level corresponding to the first feature vector is the highest confidence level in the dataset corresponding to the first target, which means the confidence level corresponding to the first feature vector is the third confidence level. The activity level corresponding to the first feature vector is determined based on the collection time and preset duration of each data point in the dataset corresponding to the first target. Specifically, assuming the preset duration is one week, the first data point is collected on Monday, the second data point is collected on Thursday, and the third data point is collected on Sunday, meaning the data corresponding to the first target is collected three times within a week. Therefore, the activity level corresponding to the first feature vector is three times.

[0083] Then, a feature vector set is formed by the feature vectors with an activity level of three and a confidence level of three, wherein the first feature vector is one of the feature vectors in the feature vector set.

[0084] Assuming the first feature vector set belongs to the i-th feature vector set, the i-th confidence level corresponding to the i-th feature vector set is the confidence level corresponding to the first feature vector, which is the third confidence level. The i-th activity level corresponding to the i-th feature vector set is the activity level corresponding to the first feature vector, which is the third activity level.

[0085] That is, the number of times each target corresponding to at least one feature vector in the i-th feature vector set appears within a preset time period is the same, and the activity level is i. The maximum confidence level in the data set corresponding to each target corresponding to at least one feature vector in the i-th feature vector set is also the same, and the confidence level is i.

[0086] S12. Sort the multiple feature vector sets according to their respective confidence and activity levels.

[0087] In one possible implementation, the total number of first vectors and the total number of second vectors corresponding to the multiple feature vector sets can be determined based on the confidence and activity corresponding to the multiple feature vector sets respectively. Then, the multiple feature vector sets are sorted according to the total number of first vectors and the total number of second vectors corresponding to the multiple feature vector sets respectively.

[0088] For example, there exists an i-th feature vector set that is one of the determined multiple feature vector sets, the i-th feature vector set corresponds to the i-th confidence level, and the i-th confidence level is P. i The i-th feature vector set corresponds to the i-th activity level, and the i-th confidence level is a. i .

[0089] There exists a j-th feature vector set that is one of the determined multiple feature vector sets, and the j-th feature vector set corresponds to the j-th confidence level, which is P. j The j-th feature vector set corresponds to the j-th activity level, and the j-th confidence level is a. j Where i and j are unequal positive integers.

[0090] Specifically, in the existence of or When the j-th eigenvector set is the dominated solution of the i-th eigenvector set, among the determined multiple eigenvector sets, the i-th eigenvector set may have multiple dominated solutions, and the total number of the first vectors of the i-th eigenvector set is the total number of eigenvectors included in the eigenvector sets corresponding to each dominated solution of the i-th eigenvector set.

[0091] In existence or When the j-th eigenvector set is the dominant solution of the i-th eigenvector set, among the determined multiple eigenvector sets, the i-th eigenvector set may have multiple dominant solutions, and the total number of second vectors in the i-th eigenvector set is the total number of eigenvectors included in the eigenvector sets corresponding to each dominant solution of the i-th eigenvector set.

[0092] For example, Figure 2 A schematic diagram of a feature vector set provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the horizontal axis P represents confidence level, and the vertical axis a represents activity level. There exists an i-th feature vector set X. i (p i ,a i ,s i ), where s i Let X be the number of eigenvectors included in the i-th eigenvector set, and let X be the number of eigenvectors in the j-th eigenvector set. j (p j ,a j ,s j ), where s j Let be the number of eigenvectors included in the j-th eigenvector set.

[0093] For example, such as Figure 2 As shown, there exists Therefore, the j-th eigenvector set X j (p j ,a j ,s j X is the set of the i-th feature vectors. i (p i ,a i ,s i A dominated solution to ). Assume the i-th eigenvector set X i (p i ,a i ,s i There is only one governed solution, which is the j-th eigenvector set X. j (p j ,a j ,s j At this point, the set X of the j-th eigenvectors...j (p j ,a j ,s j The number of eigenvectors included is s. j Therefore, the total number of the first vectors in the i-th eigenvector set is also s. j .

[0094] Similarly, there exists Therefore, the i-th eigenvector set X i (p i ,a i ,s i X is the set of the j-th eigenvectors. j (p j ,a j ,s j A dominant solution to ). Assume the set of eigenvectors X of the j-th eigenvector group. j (p j ,a j ,s j There is only one dominant solution, which is the set of eigenvectors X of the i-th eigenvector. i (p i ,a i ,s i At this point, the set of the i-th eigenvectors X i (p i ,a i ,s i The number of eigenvectors included is s. i Therefore, the total number of the second vectors in the j-th eigenvector set is also s. i .

[0095] In one possible implementation, the first sorting can be performed based on the total number of second vectors corresponding to the multiple feature vector sets. Specifically, the multiple feature vector sets can be arranged in ascending order based on the total number of second vectors corresponding to the multiple feature vector sets.

[0096] For example, the smaller the total number of second vectors corresponding to a feature vector set, the higher the ranking of that feature vector set among multiple feature vector sets. For instance, when the total number of second vectors corresponding to a feature vector set is 0, it indicates that the confidence level of that feature vector set is higher than that of other feature vectors, and the activity level of that feature vector set is also higher than that of other feature vector sets. Therefore, that feature vector set should be ranked first among multiple feature vector sets.

[0097] In one possible implementation, a second sorting can be performed based on the total number of first vectors corresponding to multiple feature vector sets. Specifically, for at least two feature vector sets with the same total number of second vectors among multiple feature vector sets, after the first sorting, the at least two feature vector sets are arranged in descending order based on the total number of first vectors corresponding to the at least two feature vector sets.

[0098] For example, when there exists an nth set of feature vectors X n With the m-th eigenvector set X m When the total number of the second vectors is equal, it indicates that the set of the nth eigenvectors X n With the m-th eigenvector set X m Since they do not dominate each other, they need to be sorted in descending order based on the total number of the first vectors in these two feature vector sets.

[0099] For example, Figure 3 A schematic diagram of another feature vector set provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the horizontal axis P represents confidence level, and the vertical axis a represents activity level. There exists a set of feature vectors X. n eigenvector set X m There also exists a set of eigenvectors X a eigenvector set X b eigenvector set X c .

[0100] like Figure 3 As shown, the feature vector set X n Corresponding activity level a n Greater than the eigenvector set X m Corresponding activity level a m However, the eigenvector set X n Corresponding activity level P n Less than the set of eigenvectors X m Corresponding activity level P m At this time, the feature vector set X n The total number of the second vectors is the eigenvector set X. a The number of eigenvectors included, the eigenvector set X m The total number of the second vector is also the eigenvector set X. a The number of eigenvectors included, that is, the set of eigenvectors X. n With the eigenvector set X m The total number of corresponding second vectors is equal.

[0101] At this point, it is necessary to determine the feature vector set X. n and the set of eigenvectors X m The corresponding first vectors are sorted in descending order of their total count. For example... Figure 3 As shown, the feature vector set X n The dominated solution set includes the eigenvector set X b eigenvector set X m The dominated solution set includes the eigenvector set X b and the set of eigenvectors X c That is, the set of eigenvectors X n The number of eigenvectors in the dominated solution set is less than the number of eigenvectors in the eigenvector set X. m The number of eigenvectors included, that is, the set of eigenvectors X. n The total number of the first vectors is less than the set of eigenvectors X m The total number of the first vector. Based on the total number of the first vector, the eigenvector set X... n and the set of eigenvectors X m Sort the feature vectors in descending order, i.e., the feature vector set X. m Arranged in the eigenvector set X n Before.

[0102] In one possible implementation, a third sorting can be performed based on the confidence levels corresponding to multiple feature vector sets. Specifically, for at least two feature vector sets in which the total number of the first vector and the total number of the second vector are equal, after the first and second sorting, the at least two feature vector sets are arranged in descending order based on the confidence levels corresponding to the at least two feature vector sets.

[0103] For example, such as Figure 4 As shown, the feature vector set X n The total number of the second vectors is the eigenvector set X. a The number of eigenvectors included, the eigenvector set X m The total number of the second vector is also the eigenvector set X. a The number of eigenvectors included, that is, the set of eigenvectors X. n With the eigenvector set X m The total number of the second vectors is equal.

[0104] Eigenvector set X n The total number of the first vectors is the eigenvector set X. b The number of eigenvectors included, the eigenvector set X m The total number of the first vectors is also the eigenvector set X. b The number of eigenvectors included, that is, the set of eigenvectors X. n With the eigenvector set X m The total number of the first vectors is equal.

[0105] like Figure 4 As shown, the feature vector set Xn The corresponding confidence level is less than that of the feature vector set X. m The corresponding confidence level, based on the confidence level, is applied to the feature vector set X. n and the set of eigenvectors X m Sort the feature vector set X in descending order. m Arranged in the eigenvector set X n Before.

[0106] For example, after sorting multiple feature vector sets for the first, second, and third time, a ranking result for these feature vector sets can be obtained. This ranking result comprehensively considers both the confidence level of the target-corresponding data and the activity level of the target, resulting in a more comprehensive and accurate ranking. In this ranking result, feature vector sets that rank higher are more important and require more attention. Since each feature vector set includes at least one feature vector, and each feature vector corresponds one-to-one with at least one target, obtaining the ranking results for multiple feature vector sets allows us to derive the ranking results for multiple data sets or target datasets corresponding to the feature vector sets. Data that ranks higher is more important and requires more attention, thus providing a clear indication of the most reliable data that needs focused attention.

[0107] For the data of multiple targets obtained by the present invention, the embodiments of the present invention also provide a method for hierarchical clustering of targets based on the confidence level of the data corresponding to the targets to obtain the clustered data of all targets.

[0108] Specifically, the acquired target data may contain multiple data sets corresponding to one target. For example, a first target may correspond to a set of data including first data, second data, and third data. Based on the confidence levels of the first, second, and third data, these data are sorted in descending order. The third data has the highest confidence level, so it is placed first. This first-ranked third data is then used as the target cluster data for the first target.

[0109] The confidence level of data indicates its reliability. Using the above method, the target archive data for the first target is the most reliable data among the multiple data corresponding to the first target. Using the above method, the target archive data corresponding to each target is obtained sequentially, and then aggregated to obtain the complete target archive data. This complete target archive data includes the most reliable data corresponding to each target.

[0110] For the data of multiple targets obtained by the present invention, the embodiments of the present invention also provide a method for hierarchical clustering of targets based on the activity level of the targets to obtain active target data.

[0111] For targets that frequently appear within the region, routine management and control are necessary. However, the occurrence patterns of different target types vary significantly. Therefore, different activity thresholds can be set for different target types to obtain active target data.

[0112] For example, when the target is a vehicle, the identification information is license plate information, and the target type is vehicle type, when the vehicle is within a preset time period [t]... m ,t n The total number of active users (∑a) is greater than the activity threshold (a). m When, it is considered that the vehicle is within the preset time period [t] m ,t n [Activity within the vehicle, activity thresholds a for different vehicle types] m The types of vehicles involved in this invention may differ. The types of vehicles involved may include hazardous materials transport vehicles, large trucks, large buses, and minivans. Preset duration [t] m ,t n The activity threshold is typically set to one week, one month, or one year. Different activity thresholds can be set for different vehicle types simultaneously. Specifically, after obtaining the active vehicles within the preset time period, the vehicle cluster data corresponding to the active vehicles can be filtered from the total vehicle cluster data. The active vehicle data includes the vehicle cluster data corresponding to all active vehicles.

[0113] The division of units in the embodiments of this invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this invention can be integrated into a single processor, exist as separate physical units, or be integrated into a single unit. The integrated units described above can be implemented in hardware or as software functional units.

[0114] This invention also provides a device 500, see [link to device 500]. Figure 5 As shown, it includes: a processing module 510 and a transceiver module 520.

[0115] The transceiver module 520 may include a receiving unit and a transmitting unit. The processing module 510 is used to control and manage the operation of the device 500. The transceiver module 520 is used to support communication between the device 500 and other devices. Optionally, the device 500 may also include a storage unit for storing the program code and data of the device 500.

[0116] Optionally, each module in the device 500 can be implemented by software.

[0117] Optionally, the processing module 510 may be a processor or controller, such as a general-purpose central processing unit (CPU), a general-purpose processor, a digital signal processing unit (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The transceiver module 520 may be a communication interface, a transceiver, or a transceiver circuit, etc., wherein the communication interface is a general term, and in a specific implementation, the communication interface may include multiple interfaces, and the storage unit may be a memory.

[0118] Processing module 510 calls transceiver module 520 to perform the following: Determine multiple feature vector sets, wherein the i-th feature vector set in the multiple feature vector sets corresponds to the i-th confidence level and the i-th activity level; wherein the activity level represents the number of times the target appears within a preset time period, and the confidence level represents the confidence level of the data corresponding to the target; the i-th feature vector set includes at least one feature vector, the at least one feature vector corresponds one-to-one with at least one target, the confidence level of each feature vector in the at least one feature vector is the i-th confidence level, and the activity level of each feature vector in the at least one feature vector is the i-th activity level; wherein i is a positive integer; sort the multiple feature vector sets according to the confidence level and activity level corresponding to the multiple feature vector sets respectively.

[0119] This invention also provides another device 600, see [link to previous document]. Figure 6 As shown, it includes:

[0120] Communication interface 601, memory 602 and processor 603;

[0121] The communication device 600 communicates with other devices through the communication interface 601, such as sending and receiving messages; the memory 602 is used to store program instructions; and the processor 603 is used to call the program instructions stored in the memory 602 and execute them according to the obtained program.

[0122] The processor 603 executes program instructions stored in the communication interface 601 and memory 602: It determines multiple feature vector sets, where the i-th feature vector set corresponds to the i-th confidence level and the i-th activity level; where the activity level represents the number of times the target appears within a preset time period, and the confidence level represents the confidence level of the data corresponding to the target; the i-th feature vector set includes at least one feature vector, each feature vector corresponding to at least one target, and the confidence level of each feature vector in the at least one feature vector is the i-th confidence level, and the activity level of each feature vector in the at least one feature vector is the i-th activity level; where i is a positive integer; and it sorts the multiple feature vector sets according to the confidence level and activity level corresponding to each feature vector set.

[0123] In this embodiment of the invention, the specific connection medium between the communication interface 601, the memory 602 and the processor 603 is not limited, such as a bus. A bus can be divided into an address bus, a data bus, a control bus, etc.

[0124] In this embodiment of the invention, the processor may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this embodiment of the invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in this embodiment of the invention can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0125] In embodiments of the present invention, the memory can be non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or it can be volatile memory, such as random-access memory (RAM). The memory can also be any other medium capable of carrying or storing desired program code having an instruction or data structure form and accessible by a computer, but is not limited thereto. The memory in embodiments of the present invention can also be a circuit or any other device capable of implementing a storage function for storing program instructions and / or data.

[0126] This invention also provides a computer-readable storage medium including program code. When the program code is run on a computer, the program code is used to cause the computer to perform the steps of the method provided in the above embodiments of this invention.

[0127] This invention also provides a computer program product comprising: computer program code, which, when run on a computer, causes the computer to perform the steps of the method provided in this invention.

[0128] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0129] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0130] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0131] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0132] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for hierarchical data archiving, characterized in that, The method includes: Multiple feature vector sets are defined, wherein the i-th feature vector set in the multiple feature vector sets corresponds to the i-th confidence level and the i-th activity level; wherein the activity level represents the number of times the target appears within a preset time period, and the confidence level represents the confidence level of the data corresponding to the target; the i-th feature vector set includes at least one feature vector, the at least one feature vector corresponds one-to-one with at least one target, the confidence level of each feature vector in the at least one feature vector is the i-th confidence level, and the activity level of each feature vector in the at least one feature vector is the i-th activity level; wherein i is a positive integer; The multiple feature vector sets are sorted according to their respective confidence and activity levels. The step of sorting the multiple feature vector sets according to their respective confidence and activity levels includes: Based on the confidence and activity corresponding to the multiple feature vector sets, determine the total number of first vectors and the total number of second vectors corresponding to the multiple feature vector sets respectively; The multiple feature vector sets are sorted according to the total number of first vectors and the total number of second vectors corresponding to the multiple feature vector sets respectively; The total number of first vectors corresponding to the i-th feature vector set is the total number of feature vectors included in the feature vector sets corresponding to each dominated solution of the i-th feature vector set; The total number of second vectors corresponding to the i-th eigenvector set is the total number of eigenvectors included in the eigenvector sets corresponding to each dominating solution of the i-th eigenvector set.

2. The method as described in claim 1, characterized in that, exist or When the j-th eigenvector set is the dominated solution of the i-th eigenvector set, where, For the i-th confidence level, For the i-th activity level, the j-th feature vector set is one of the multiple feature vector sets, where j is a positive integer not equal to i. The confidence level corresponding to the j-th feature vector set. The activity level corresponding to the j-th feature vector set; exist or When the j-th eigenvector set is the dominant solution of the i-th eigenvector set.

3. The method as described in claim 1, characterized in that, The multiple feature vector sets are sorted according to the total number of first vectors and the total number of second vectors corresponding to the multiple feature vector sets, including: The multiple feature vector sets are arranged in ascending order based on the total number of second vectors corresponding to the multiple feature vector sets.

4. The method as described in claim 3, characterized in that, Also includes: For at least two feature vector sets in which the total number of second vectors is equal, the at least two feature vector sets are arranged in descending order according to the total number of first vectors corresponding to the at least two feature vector sets.

5. The method as described in claim 3, characterized in that, Also includes: For at least two feature vector sets in which the total number of the first vector and the total number of the second vector are equal, the at least two feature vector sets are arranged in descending order according to the confidence levels corresponding to the at least two feature vector sets.

6. The method according to any one of claims 1-5, characterized in that, Determine multiple sets of feature vectors, including: Acquire multiple data points, each including the identifier information of the target corresponding to the data, the confidence level of the data, and the acquisition time; Determine the multiple feature vector sets based on the multiple data; Wherein, the first feature vector is one of the i-th feature vector sets, the first feature vector corresponds to the first target, the first target is one of the at least one target, the confidence level corresponding to the first feature vector is the maximum confidence level in the data set corresponding to the first target, the data set corresponding to the first target includes data with the identification information corresponding to the first target; the activity level corresponding to the first feature vector is determined based on the collection time of each data in the data set corresponding to the first target and the preset duration.

7. The method as described in claim 6, characterized in that, The method further includes: Each data also includes a target type. After acquiring the multiple data, a target data set is selected from the multiple data based on the target type included in each data. The target data set includes data of at least one preset target type. Determining the multiple feature vector sets based on the multiple data includes: The plurality of feature vector sets are determined based on the target data set.

8. The method as described in claim 7, characterized in that, When the data containing the identification information corresponding to the first target includes at least two target types, the target type that appears most frequently is taken as the target type of the first target.

9. A device for hierarchical data archiving, characterized in that, The device is a server or a chip within a server, and includes a processing unit and a transceiver unit: the processing unit invokes the transceiver unit to execute: Multiple feature vector sets are defined, wherein the i-th feature vector set in the multiple feature vector sets corresponds to the i-th confidence level and the i-th activity level; wherein the activity level represents the number of times the target appears within a preset time period, and the confidence level represents the confidence level of the data corresponding to the target; the i-th feature vector set includes at least one feature vector, the at least one feature vector corresponds one-to-one with at least one target, the confidence level of each feature vector in the at least one feature vector is the i-th confidence level, and the activity level of each feature vector in the at least one feature vector is the i-th activity level; wherein i is a positive integer; The multiple feature vector sets are sorted according to their respective confidence and activity levels. The processing unit is configured to, when sorting the multiple feature vector sets according to the confidence and activity corresponding to the multiple feature vector sets respectively, determine the total number of first vectors and the total number of second vectors corresponding to the multiple feature vector sets respectively based on the confidence and activity corresponding to the multiple feature vector sets respectively. The processing unit is further configured to sort the plurality of feature vector sets according to the total number of first vectors and the total number of second vectors corresponding to the plurality of feature vector sets respectively; The total number of first vectors corresponding to the i-th feature vector set is the total number of feature vectors included in the feature vector sets corresponding to each dominated solution of the i-th feature vector set; The total number of second vectors corresponding to the i-th eigenvector set is the total number of eigenvectors included in the eigenvector sets corresponding to each dominating solution of the i-th eigenvector set.

10. The apparatus as claimed in claim 9, characterized in that, exist or When the j-th eigenvector set is the dominated solution of the i-th eigenvector set, where, For the i-th confidence level, For the i-th activity level, the j-th feature vector set is one of the multiple feature vector sets, where j is a positive integer not equal to i. The confidence level corresponding to the j-th feature vector set. The activity level corresponding to the j-th feature vector set; exist or When the j-th eigenvector set is the dominant solution of the i-th eigenvector set.

11. The apparatus as claimed in claim 9, characterized in that, The processing unit is configured to sort the multiple feature vector sets in ascending order based on the second total number of vectors corresponding to the multiple feature vector sets when sorting the multiple feature vector sets according to the first total number of vectors and the second total number of vectors corresponding to the multiple feature vector sets respectively.

12. The apparatus as claimed in claim 11, characterized in that, For at least two feature vector sets in which the total number of second vectors is equal, the processing unit is configured to sort the at least two feature vector sets in descending order according to the total number of first vectors corresponding to the at least two feature vector sets.

13. The apparatus as claimed in claim 11, characterized in that, For at least two feature vector sets in which the total number of first vectors and the total number of second vectors are equal, the processing unit is configured to sort the at least two feature vector sets in descending order according to the confidence levels corresponding to the at least two feature vector sets.

14. The apparatus according to any one of claims 9-13, characterized in that, The transceiver unit is used to acquire multiple data when the processing unit determines multiple feature vector sets. Each data includes the identification information of the target corresponding to the data, the confidence level of the data, and the acquisition time. The processing unit is configured to determine the plurality of feature vector sets based on the plurality of data; wherein, the first feature vector is one of the i-th feature vector sets, the first feature vector corresponds to a first target, the first target is one of the at least one target, the confidence level corresponding to the first feature vector is the maximum confidence level in the data set corresponding to the first target, and the data set corresponding to the first target includes data with identification information corresponding to the first target; The activity level corresponding to the first feature vector is determined by the processing unit based on the collection time of each data in the dataset corresponding to the first target and the preset duration.

15. The apparatus as claimed in claim 14, characterized in that, Each data also includes a target type. After the transceiver unit acquires the plurality of data, the processing unit is used to filter out a target data set from the plurality of data according to the target type included in each data. The target data set includes data of at least one preset target type. The processing unit determines the multiple feature vector sets based on the multiple data, including: the processing unit determines the multiple feature vector sets based on the target data set.

16. The apparatus as claimed in claim 15, characterized in that, The processing unit is configured to, when the data containing the identification information corresponding to the first target includes at least two target types, select the target type that appears most frequently as the target type of the first target.

17. A communication device, characterized in that, The device includes a processor and an interface circuit, the interface circuit being used to receive signals from other devices outside the device and transmit them to the processor, or to send signals from the processor to other devices outside the device, the processor being used to implement the method as described in any one of claims 1 to 8 via logic circuits or execution code instructions.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on a computer, cause the computer to perform the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Content recommendation method based on deep learning model, related device and equipment

    CN115203568A

  • Power grid digital project Pareto optimization method and system

    CN115330201A