Stream data processing method, device, storage medium and computer equipment

By clustering and evaluating the flow data in road traffic scenarios in batches and combining optimization processing, the problem of difficulty in determining the clustering quality in multiple batches of clustering is solved, and the accuracy and efficiency of the processing are improved.

CN114416786BActive Publication Date: 2025-05-06ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111493756.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-08
Publication Date
2025-05-06
Estimated Expiration
2041-12-08

AI Technical Summary

Technical Problem

In road traffic scenarios, it is difficult to determine the quality of clustering after multiple batches of clustering, especially when category drifts and accuracy decreases.

Method used

By analyzing the stream data, analytical results are generated, and clustering them in batches, obtaining clustering results for multiple batches. The clustering results of each batch contain the matching probability between at least two records and any two records. Then, any clustering results are evaluated, the probability that any two or more records belong to the target record is determined, and optimization is performed based on this probability, including splitting, merging, deleting, and replacing.

Benefits of technology

The quality evaluation and optimization processing of multi-batch clustering results are realized, the accuracy and efficiency of clustering processing are improved, and the problems of category drift and accuracy reduction are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114416786B_ABST
    Figure CN114416786B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, storage medium and computer equipment for processing stream data. The method includes: parsing stream data to generate parsing results; clustering the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and the matching probability between any two records; evaluating one or more records included in any clustering result to determine the probability that any two or more records belong to the target record; optimizing any clustering result based on the probability that any two or more records belong to the target record, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing. The present invention solves the technical problem in the related art that it is difficult to determine the quality of clustering processing when performing multi-batch clustering processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular to a method, device, storage medium and computer equipment for processing stream data. Background Art

[0002] In the road traffic scenario, it is necessary to parse the streaming data of city cameras in real time to obtain the parsing results of people / motor vehicles / non-motor vehicles, such as feature vectors and attributes, and then batch cluster these parsing results and compare and archive them with the vehicle database / personnel database. However, in the related technology, the categories obtained after data circulation has been clustered for multiple batches, and when category drift and precision reduction occur, it is often difficult to detect.

[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0004] The embodiments of the present invention provide a method, apparatus, storage medium and computer device for processing stream data, so as to at least solve the technical problem in the related art that it is difficult to determine the quality of clustering processing when performing multi-batch clustering processing.

[0005] According to one aspect of an embodiment of the present invention, a method for processing stream data is provided, comprising: parsing the stream data to generate parsing results; clustering the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and a matching probability between any two records; evaluating one or more records included in any one of the clustering results to determine the probability that any two or more records belong to target records; and optimizing any one of the clustering results based on the probability that any two or more records belong to the target records, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0006] Optionally, one or more records contained in any one of the clustering results are evaluated to determine the probability that any two or more records belong to the target record, including: sampling from the clustering result to obtain records to be evaluated that meet a predetermined number of entries; segmenting the records to be evaluated with the predetermined number of entries according to a predetermined segmentation method to obtain multiple subclasses in the clustering result; matching the records contained in each subclass with the target record to obtain the matching probability of each record contained in each subclass; based on the matching probability of each record contained in each subclass, determining the probability of the record in each subclass belonging to the target record, wherein records with a probability higher than a threshold belong to the same type.

[0007] Optionally, the records to be evaluated that meet the predetermined number of entries are sampled from the clustering results by any one or more of the following methods: randomly extracting the predetermined number of entries from the clustering results as the records to be evaluated; extracting the predetermined number of entries from the clustering results as the records to be evaluated according to the spatiotemporal distribution; calculating at least one nearest neighbor record of each record according to the feature vector of each record in the clustering results, and selecting the predetermined number of entries from the records as the records to be evaluated, wherein the set of neighbor records of the predetermined number of entries selected exceeds a predetermined order of magnitude.

[0008] Optionally, after determining the probability of the records in each subclass belonging to the target record, the method further includes: counting the number of records in each subclass belonging to the target record, and determining the sample records in each subclass that do not belong to the target record; and fusing the sample records in each subclass that do not belong to the target record.

[0009] Optionally, the analysis results are clustered in batches to obtain multiple batches of clustering results, including: dividing the analysis results according to timestamps to obtain at least one batch of classification results; clustering the classification results of each batch separately to obtain the clustering results of the multiple batches.

[0010] According to another aspect of an embodiment of the present invention, a method for processing stream data is provided, comprising: parsing the stream data to generate parsing results; clustering the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and a matching probability between any two records; determining a class center from any one of the clustering results, wherein the class center is used to characterize the attributes of the class to be assigned; based on the class center, performing secondary clustering on records belonging to the same class in any one of the clustering results, wherein the attributes of the records in the clustering results after the secondary clustering are the same; and optimizing any one of the clustering results after the secondary clustering, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0011] According to another aspect of an embodiment of the present invention, a method for processing stream data is provided, comprising: if an assessment instruction is detected on an interactive interface of a city assessment system, parsing collected view flow data to generate a view parsing result, wherein the view stream data comprises: video stream data and / or image stream data; displaying on the interactive interface a plurality of batches of clustering results obtained after clustering the view parsing results in batches, wherein each batch of clustering results comprises: at least two view records with the same background, and a matching probability between any two view records; and displaying on the interactive interface a result of optimizing one or more batches of clustering results, wherein the optimization comprises any one or more of the following methods: splitting, merging, deleting and replacing.

[0012] Optionally, before displaying the results of optimizing one or more batches of clustering results on the interactive interface, the method further includes: evaluating one or more view records contained in any one of the clustering results to determine the probability that any two or more view records contain the same target object; and optimizing any one of the clustering results based on the probability that any two or more view records contain the same target object.

[0013] According to another aspect of an embodiment of the present invention, a stream data processing device is provided, comprising: a first parsing module, used to parse the stream data and generate a parsing result; a first clustering module, used to cluster the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and a matching probability between any two records; an evaluation module, used to evaluate one or more records included in any one of the clustering results to determine the probability that any two or more records belong to the target record; a first optimization module, used to optimize any one of the clustering results based on the probability that any two or more records belong to the target record, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0014] Optionally, the evaluation module includes: a sampling module, which is used to sample from the clustering results to obtain records to be evaluated that meet a predetermined number of entries; a segmentation module, which is used to segment the records to be evaluated with the predetermined number of entries according to a predetermined segmentation method, and obtain multiple subclasses in the clustering results; a matching module, which is used to match the records contained in each subclass with the target records to obtain the matching probability of each record contained in each subclass; and a first determination module, which is used to determine the probability of the records in each subclass belonging to the target record based on the matching probability of each record contained in each subclass, wherein records with probabilities higher than a threshold belong to the same type.

[0015] According to another aspect of an embodiment of the present invention, a stream data processing device is provided, comprising: a second parsing module, used to parse the stream data and generate parsing results; a second clustering module, used to cluster the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and a matching probability between any two records; a second determination module, used to determine a class center from any one of the clustering results, wherein the class center is used to characterize the attributes of the class to be assigned; a third clustering module, used to perform secondary clustering on records belonging to the same class in any one of the clustering results based on the class center, wherein the attributes of the records in the clustering results after the secondary clustering are the same; a second optimization module, used to optimize any one of the clustering results after the secondary clustering, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0016] According to another aspect of an embodiment of the present invention, a stream data processing device is provided, characterized in that it includes: a parsing module, which is used to parse the collected view flow data and generate a view parsing result if an assessment instruction is detected on the interactive interface of the city assessment system, wherein the view stream data includes: video stream data and / or image stream data; a first display module, which is used to display on the interactive interface a plurality of batches of clustering results obtained after clustering the view parsing results in batches, wherein each batch of clustering results includes: at least two view records with the same background, and a matching probability between any two view records; a second display module, which is used to display on the interactive interface a result of optimizing one or more batches of clustering results, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0017] According to another aspect of an embodiment of the present invention, a storage medium is provided, wherein the storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute any one of the above-mentioned methods for processing stream data.

[0018] According to one aspect of an embodiment of the present invention, a computer device is provided, comprising: a memory and a processor, wherein the memory stores a computer program; the processor is used to execute the computer program stored in the memory, wherein when the computer program is executed, the processor executes any one of the above-mentioned methods for processing stream data.

[0019] In an embodiment of the present invention, the parsing results generated by parsing the streaming data are clustered in batches. After obtaining multiple batches of clustering results, any record in the multiple batches of clustering results is evaluated, thereby achieving the purpose of determining the probability that each record belongs to the target record. Moreover, based on the probability that any two or more records belong to the target record, any clustering result can be optimized by splitting, merging, deleting and replacing, thereby solving the technical problem in the related art that it is difficult to determine the quality of clustering processing when performing multiple batches of clustering processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0021] Figure 1 A hardware structure block diagram of a computer terminal for implementing a method for processing stream data is shown;

[0022] Figure 2 is a flowchart of a first method for processing stream data according to embodiment 1 of the present invention;

[0023] Figure 3 is a flowchart of a second method for processing stream data according to embodiment 1 of the present invention;

[0024] Figure 4 is a flowchart of a third method for processing stream data according to embodiment 1 of the present invention;

[0025] Figure 5 is a schematic diagram of estimating target numbers under different situations provided by an optional embodiment of the present invention;

[0026] Figure 6 is a schematic diagram of edge matching probability and matching probability provided by an optional implementation manner of the present invention;

[0027] Figure 7 is a structural block diagram of a first device for processing stream data provided according to Embodiment 2 of the present invention;

[0028] Figure 8 is a structural block diagram of a second stream data processing device provided according to Embodiment 3 of the present invention;

[0029] Fig. 9 is a structural block diagram of a third stream data processing device provided in accordance with Embodiment 4 of the present invention;

[0030] Fig.10 It is a structural block diagram of a computer terminal according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0033] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:

[0034] Clustering: The process of grouping a number of structured records into multiple sets according to a certain category definition, that is, the process of dividing a set of physical or abstract objects into multiple classes composed of similar objects;

[0035] Multi-batch clustering: In streaming data scenarios, it is necessary to continuously cluster batches based on time and merge the clustering results of previous and next batches.

[0036] Probability sampling: Sampling is performed according to a certain probability density definition to obtain a sampling sequence result whose distribution is close to the true distribution;

[0037] Class evaluation: Quality evaluation of whether the elements in a class belong to the same class is performed based on the class definition. Usually, the number of real classes contained in a class (ideally 1) and the maximum proportion of records belonging to the same class are commonly used for intra-class quality evaluation.

[0038] View parsing: video stream / image stream parsing results, including feature vectors and attributes extracted from human / vehicle thumbnails;

[0039] Feature vector: a one-dimensional array calculated for an image. Usually, the similarity between two images can be obtained by calculating the Euclidean distance between their feature vectors.

[0040] Matching probability: the probability that any two records are of the same type. In the case of human body images, that is, the probability that two human images are of the same person when their feature vectors are d, which can be obtained through annotated data statistics or learning;

[0041] Indicator function: The letter I represents the indicator function, I(True)=1, I(False)=0.

[0042] Example 1

[0043] According to an embodiment of the present invention, an embodiment of a method for processing stream data is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0044] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for processing stream data. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0045] It should be noted that the one or more processors 102 and / or other processing circuits for stream data may generally be referred to herein as "processing circuits for stream data". The processing circuits for stream data may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the processing circuits for stream data may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the processing circuits for stream data are used as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0046] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the method for processing stream data in the embodiment of the present invention. The processor 102 executes various functional applications and stream data processing by running the software programs and modules stored in the memory 104, that is, the method for processing stream data of the above-mentioned application program is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0047] The transmission device is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device can be a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet wirelessly.

[0048] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0049] Under the above operating environment, this application provides Figure 2 The method for processing stream data shown. Figure 2 is a flowchart of a method for processing stream data according to Embodiment 1 of the present invention. Figure 2 As shown, the method comprises the following steps:

[0050] Step S202, parsing the stream data to generate parsing results;

[0051] Step S204, clustering the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and a matching probability between any two records;

[0052] Step S206, evaluating one or more records included in any clustering result to determine the probability that any two or more records belong to the target record;

[0053] Step S208, based on the probability that any two or more records belong to the target record, optimizing any clustering result, wherein the optimizing includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0054] Through the above steps, the parsing results generated by parsing the stream data are clustered in batches. After obtaining multiple batches of clustering results, any record in the multiple batches of clustering results is evaluated, thereby achieving the purpose of determining the probability that any two or more records belong to the target records. Moreover, based on the probability that any two or more records belong to the target records, it is possible to achieve the effect of optimizing the processing of splitting, merging, deleting and replacing any clustering result, thereby solving the technical problem in the related technology that it is difficult to determine the quality of clustering processing when performing multi-batch clustering processing.

[0055] It should be noted that the stream data may include multiple stream data, and the stream data included in different scenarios should not be the same. For example, in the scenario of urban traffic, the stream data obtained based on the camera may be video stream data; in the e-commerce scenario, it may also be commodity stream data, etc. The parsing results vary depending on the stream data. For example, in the scenario of road traffic, when the video stream data is obtained based on the camera, the parsing result obtained is the view parsing result.

[0056] As an optional embodiment, the stream data is parsed to generate parsing results. The parsing results include vector features and attributes of the target object obtained by parsing the stream data. In different scenarios, the parsed target object is different depending on the stream data. For example, in a road traffic scenario, the target object may be a vehicle, a human body, etc. By parsing the stream data and generating parsing results, batch clustering can be performed according to the vector features included in the parsing results, which can ensure accurate and effective batch clustering operations.

[0057] As an optional embodiment, taking the road traffic scenario as an example, streaming data is generated all the time. In order to better analyze the streaming data of each time period and to reduce the workload of processing a large number of parsing results, it is necessary to cluster the parsing results obtained by parsing the streaming data in batches, so as to efficiently analyze the clustering results based on the video parsing results. By clustering the parsing results in batches, multiple batches of clustering results are obtained, wherein each batch of clustering results includes: at least two records and the matching probability between any two records. That is, in different batches, clustering results including the matching probability between any two records can be obtained. The higher the matching probability, the more similar the two records are considered. By obtaining the matching probability between any two records, the similarity relationship between records can be more clearly understood, and then the quality of clustering processing can be understood, and the quality of multi-batch clustering can be scientifically evaluated.

[0058] As an optional embodiment, when performing batch clustering processing, the batches can be divided in a variety of ways, for example, the parsing results are divided according to timestamps to obtain at least one batch of classification results. Dividing the parsing results according to timestamps can not only better analyze the parsing results at each timestamp, but also reduce the workload of processing a large number of parsing results, and then cluster each batch of classification results separately to obtain multiple batches of clustering results.

[0059] As an optional embodiment, one or more records included in any clustering result are evaluated to determine the probability that any two or more records belong to the target record, that is, one or more records included in any clustering result are evaluated to determine the probability that at least two records included in any clustering result and the matching probability between any two records belong to the target record. That is, the similarity between one or more records included in the clustering result and the target record is judged. The higher the probability that the result of the record evaluation belongs to the target record, the more similar the record is to the target record. It should be noted that when determining the probability, it can be calculated based on the feature vector in the view feature result. When evaluating one or more records included in any clustering result, there is a situation where there are many records included in the clustering result. In this case, sampling can be performed from the clustering result to obtain records to be evaluated that meet the predetermined number of entries, so as to improve the efficiency of system operation and reduce the workload. After obtaining the records to be evaluated that meet the predetermined number of entries, the records to be evaluated with the predetermined number of entries are segmented according to a predetermined segmentation method to obtain multiple subclasses in the clustering result. That is, the records to be evaluated are divided into multiple possibilities through a predetermined segmentation method, and multiple subclasses represent multiple possible results that may exist between the record and the target record. Multiple subclasses in the clustering result are obtained through segmentation, and the records contained in each subclass are matched with the target record to obtain the matching probability of each record contained in each subclass. Then, based on the matching probability of each record contained in each subclass, the probability of the record belonging to the target record in each subclass can be determined, where the records with a probability higher than the threshold belong to the same type. The probability of each record belonging to the predetermined target can be judged more accurately, thereby more effectively reflecting the quality of multi-batch clustering.

[0060] As an optional embodiment, by evaluating one or more records included in any clustering result, the probability that any two or more records belong to the target record is determined, and the quality of the clustering result can be more accurately evaluated by continuous sampling. Through continuous sampling, a sampling result of which category one or more records included in the clustering result belong to can be obtained in each round. The sampling result can count the probability that one or more records included in the clustering result belong to the target record, where the probability can be expressed in the form of a percentage of the number of times, and the estimated distribution of different target numbers of multiple batches of clustering results can be obtained, thereby more effectively reflecting the quality of multiple batches of clustering.

[0061] As an optional embodiment, when obtaining records to be evaluated that meet the predetermined number of entries, any one or more of the following methods can be used, for example: randomly extracting records of a predetermined number of entries from the clustering results as records to be evaluated; extracting records of a predetermined number of entries from the clustering results according to the spatiotemporal distribution as records to be evaluated; calculating at least one record of the nearest neighbor of each record according to the characteristic vector of each record in the clustering results, and selecting the records of the predetermined number of entries as records to be evaluated, wherein the set of neighbor records of the selected records of the predetermined number of entries exceeds the predetermined order of magnitude. That is, assuming that m records are to be selected, the K records of the nearest neighbors of each record are calculated according to the characteristic vector of each record, and then m records are selected, so that the set composed of the m K nearest neighbor records is as large as possible. By setting different selection methods for records to be evaluated, the method of selecting the records of the predetermined number of entries to be evaluated can be selected accordingly in different applications and scenarios, so that the selection of records to be evaluated can be more flexible and more applicable.

[0062] As an optional embodiment, after determining the probability of the records belonging to the target record in each subclass, the following steps may also be included: counting the number of records belonging to the target record in each subclass, and determining the sample records that do not belong to the target record in each subclass; fusing the sample records that do not belong to the target record in each subclass. When performing the fusion process, the fusion calculation may be performed based on the matching probability of the selected records that appear in the nearest neighbor. For example, if record A is not selected into the sampling set, and there is record B in its nearest neighbor, and record C is selected into the sampling set, then the matching probability of record A and the target can take the maximum value from the following values: the direct matching probability of record A and the target record; the matching probability of record A and record B*the matching probability of record B and the target record; the matching probability of record A and record C*the matching probability of record C and the target record. By fusing the sample records that do not belong to the target record in each subclass, the matching probability can be obtained more accurately and the record loss can be reduced compared to the case where the sample records that do not belong to the target record of the target machine are not processed.

[0063] As an optional embodiment, based on the probability that any two or more records belong to the target record, any clustering result is optimized, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing. The probability that any two or more records belong to the target record can reflect the quality of multi-batch clustering, and based on the reflected quality of multi-batch clustering, the clustering results can be optimized by splitting, merging, deleting and replacing. Furthermore, when optimizing any clustering result, it can be processed in a human-computer collaborative manner, and the clustering result can be further processed quickly to achieve high precision and high classification rate of multi-batch clustering.

[0064] Figure 3 is a flow chart of a second method for processing stream data according to embodiment 1 of the present invention. Figure 3 As shown, the method comprises the following steps:

[0065] Step S302, parsing the stream data to generate parsing results;

[0066] Step S304, clustering the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and a matching probability between any two records;

[0067] Step S306, determining a cluster center from any clustering result, wherein the cluster center is used to characterize the attribute of the cluster to be assigned;

[0068] Step S308, based on the cluster center, performing secondary clustering on the records belonging to the same class in any clustering result, wherein the attributes of the records in the clustering result after the secondary clustering are the same;

[0069] Step S310, optimizing any clustering result after secondary clustering, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0070] Through the above steps, the parsing results generated by parsing the streaming data are clustered in batches. After obtaining multiple batches of clustering results, the class center representing the attributes of the class to be regressed is determined from any clustering result, so that based on the class center, the records belonging to the same class in any clustering result are clustered again. Because the attributes of the records in the clustering results after the secondary clustering are the same, the purpose of clustering records with the same attributes is achieved. Moreover, any clustering result after the secondary clustering can be split, merged, deleted and replaced for optimization processing, which solves the technical problem in the related technology that it is difficult to determine the quality of clustering processing when performing multi-batch clustering processing.

[0071] Figure 4 is a flowchart of a third method for processing stream data according to Embodiment 1 of the present invention. Figure 4 As shown, the method comprises the following steps:

[0072] Step S402: if an evaluation instruction is detected on the interactive interface of the city evaluation system, the collected view flow data is parsed to generate a view parsing result, wherein the view flow data includes: video flow data and / or picture flow data;

[0073] Step S404, displaying multiple batches of clustering results obtained by clustering the view parsing results in batches on the interactive interface, wherein each batch of clustering results includes: at least two view records with the same background, and a matching probability between any two view records;

[0074] Step S406, displaying the result of optimizing the clustering results of one or more batches on the interactive interface, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0075] Through the above steps, when an assessment instruction is detected on the interactive interface of the city assessment system, the collected view flow data including video stream data and / or image stream data is parsed to generate view analysis results, and the view analysis results are clustered in batches. The clustering results of multiple batches obtained after clustering the view analysis results in batches can be displayed on the interactive interface, and any clustering result can be optimized by splitting, merging, deleting and replacing, etc., thereby solving the technical problem in the related art that it is difficult to determine the quality of clustering processing when performing multi-batch clustering processing.

[0076] Based on the above embodiments and optional embodiments, an optional implementation is provided, which is described in detail below.

[0077] There is a technical problem in the related art that it is difficult to determine the quality of clustering processing when performing multi-batch clustering processing on streaming data.

[0078] Based on this, in an optional implementation of the present invention, a multi-batch clustering evaluation method based on probability sampling is provided, based on a human image scene, taking video stream data and view parsing results as examples. The optional implementation of the present invention is described in detail below.

[0079] S1, obtaining multi-batch / single-batch clustering results, wherein the clustering results are obtained by clustering the view parsing results generated by parsing the video stream data in batches;

[0080] S2, for each category, extract all records within a certain batch range;

[0081] S3, select m records within the category;

[0082] It should be noted that different methods can be used to select m records in a category according to the needs of actual applications and scenarios. For example, the following methods can be used: 1) If the number of records in a category is greater than m, then m records are randomly selected; 2) m records are selected according to the temporal and spatial feature distribution; 3) K records of the nearest neighbors of each record are calculated according to the feature vector; and then m records are selected so that the set of these m K nearest neighbor records is as large as possible.

[0083] S4, assuming that N records are finally selected in the category, there is a segmentation method π, which divides the records in the category into multiple subcategories, and the probability density of π is defined as:

[0084]

[0085] Where I is the indicator function, π={z i},z i represents the subcategory to which the i-th record belongs, and Ω is the normalization factor. At this time, the expected value of the number of subcategories within a category is: φ(X) = ∑ n P(X,π)×|π|.

[0086] It should be noted that, for π in the optional implementation mode of the present invention, π is initialized to That is, they are all of the same type, and can be different values ​​by continuously performing probability density sampling (Gibbs sampling). The probability density sampling formula is as follows:

[0087]

[0088]

[0089] in, The optional range is the category set to which the new category α / K neighbor records belong

[0090] S5, assuming that the sampling cycle is T times, based on the sampling results, the number of people in the category can be estimated as:

[0091]

[0092] Among them, |π t | indicates The number of different elements in is the number of different subclasses. Due to the limitation of computing resources, the number of sampling times cannot be infinite, and can be adjusted according to the size of N. The final number of people can be estimated by selecting the results of the second half of the stable sampling for statistics.

[0093] S6, based on the sampling results, the graph matching probability between each record and the target record can be estimated:

[0094]

[0095] Similarly, due to the limitation of computing resources, the number of sampling times cannot be infinite, but can be adjusted according to the size of N. Finally, the number of people can be estimated by selecting the results of the second half of the stable sampling for statistics.

[0096] S7, obtain the estimation of the number of targets contained in each category of the multi / single batch clustering results, and the graph matching probability that the records in the category belong to the target. The remaining samples that are not selected for sampling can be fused and calculated based on the matching probability of the selected records that appear in the nearest neighbors. For example, if record A is not selected into the sampling set, and there are records B and C in its nearest neighbors that are selected into the sampling set, then the matching probability of record A and the target can take the maximum value from the following values: the direct matching probability of record A and target record; the matching probability of record A and record B*the graph matching probability of record B and target record; the matching probability of record A and record C*the graph matching probability of record C and target record.

[0097] It should be noted that the optional implementation of the present invention mainly describes the intra-class evaluation. For the inter-class evaluation, the method is similar, and the quantitative data of class splitting can be performed based on specific indicators. That is, the classes suspected of being the same target are treated as the same class for probability density sampling. If the estimated number of different targets is close to 1, it can be determined that the two classes are the same class.

[0098] For example, Figure 5 is a schematic diagram of estimating target numbers under different situations provided by an optional embodiment of the present invention, such as Figure 5 As shown, Figure 5 The four cases of edge matching probability in the middle and upper edges indicate that a certain large category after multi-batch clustering has 2, 3, 3, and 3 records respectively, where each point represents a record and each edge represents the matching probability between two records. Through the above edge matching probabilities, we can get the results of estimating different target numbers.

[0099] by Figure 5 For example, in the first case of edge matching probability, there are 2 records, which means that there are 2 pictures in the first batch of this category. The probability that the two pictures are of the same person is 0.85. There is a segmentation method π. At this time, π has 2 segmentation methods, that is, there are two possibilities: the same person [0, 0], not the same person [0, 1]. The corresponding probability densities P(X, π) are 0.85 and 0.15 respectively. Therefore, the expected value of the number of subcategories in the category, here, refers to the estimated number of different targets in the first batch of this category = 0.85*1+0.15*2=1.15. According to sampling, it can be concluded that Figure 5 In the case of the middle edge matching probability, the estimated number of people in different batches of the same category is 1.149 / 1.382 / 2.141 / 1.099 respectively. That is, the estimated number of different targets in the first batch of the category obtained by sampling is 1.149, and the estimated number of different targets in the theoretical first batch calculated by the method provided by the optional embodiment of the present invention is 1.15. Therefore, the effectiveness and accuracy of the method provided by the optional embodiment of the present invention are verified.

[0100] Figure 5It also includes 4 cases of matching probability of 100 image records. When the matching probability between each pair is 0.5 / 0.51 / 0.60 / 0.70 respectively, according to sampling, a set of sampling results of the categories of images can be obtained in each round. The sampling results can count the proportion of times that each image and the target image are of the same category (that is, the matching probability). By counting the proportion of different categories in each round, it can be concluded that Figure 5 In the case of the matching probability of the picture pairs, the estimated number of people in different batches of the same category are 33.3 / 3.2 / 1.0 / 1.0 respectively, and the estimated distribution of different target numbers for multi-batch clustering results is obtained.

[0101] Figure 6 is a schematic diagram of edge matching probability and graph matching probability provided by an optional embodiment of the present invention, such as Figure 6 As shown in the figure, the above four situations indicate that a certain category after multi-batch clustering has 2, 3, 3, and 3 records respectively. Each point represents a record, and each edge represents the matching probability between two records. Based on the edge matching probability and the method provided by the optional embodiment of the present invention, the matching probability between different points in the background of the entire similarity graph can be calculated. The probability result is as follows: Figure 6 The following 4 situations are shown.

[0102] Through the above optional implementation, the following beneficial effects can be achieved:

[0103] (1) Rationally evaluate the number of different targets contained in each category in the multi-batch clustering results, as well as the probability that each record belongs to the predetermined target, and scientifically evaluate the quality of multi-batch clustering;

[0104] (2) It can directly identify possible garbage, noise, and noise records;

[0105] (3) Support optimization operations such as splitting, merging, deleting, and replacing categories.

[0106] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0107] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method for processing stream data according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of various embodiments of the present invention.

[0108] Example 2

[0109] According to an embodiment of the present invention, a device for implementing the first method for processing stream data is also provided. Figure 7 is a structural block diagram of a first device for processing stream data according to Embodiment 2 of the present invention. Figure 7 As shown, the device includes: a first parsing module 702, a first clustering module 704, an evaluation module 706 and a first optimization module 708. The device is described below.

[0110] The first parsing module 702 is used to parse the streaming data and generate parsing results; the first clustering module 704 is connected to the above-mentioned first parsing module 702, and is used to cluster the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and the matching probability between any two records; the evaluation module 706 is connected to the above-mentioned first clustering module 704, and is used to evaluate one or more records included in any clustering result to determine the probability that any two or more records belong to the target record; the first optimization module 708 is connected to the above-mentioned evaluation module 706, and is used to optimize any clustering result based on the probability that any two or more records belong to the target record, wherein the optimization processing includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0111] It should be noted that the first analysis module 702, the first clustering module 704, the evaluation module 706 and the first optimization module 708 correspond to steps S202 to S208 in Example 1, and the examples and application scenarios implemented by the multiple modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules as part of the device can be run in the computer terminal 10 provided in Example 1.

[0112] Example 3

[0113] According to an embodiment of the present invention, a device for implementing the second method for processing stream data is also provided. Figure 8 is a structural block diagram of a second stream data processing device according to Embodiment 3 of the present invention. Figure 8 As shown, the device includes: a second parsing module 802, a second clustering module 804, a second determination module 806, a third clustering module 808 and a second optimization module 810. The device is described below.

[0114] A second parsing module 802 is used to parse the stream data and generate parsing results; a second clustering module 804 is connected to the above-mentioned second parsing module 802, and is used to cluster the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and the matching probability between any two records; a second determination module 806 is connected to the above-mentioned second clustering module 804, and is used to determine the class center from any clustering result, wherein the class center is used to characterize the attributes of the class to be assigned; a third clustering module 808 is connected to the above-mentioned second determination module 806, and is used to perform secondary clustering on the records belonging to the same class in any clustering result based on the class center, wherein the attributes of the records in the clustering result after the secondary clustering are the same; a second optimization module 810 is connected to the above-mentioned third clustering module 808, and is used to optimize any clustering result after the secondary clustering, wherein the optimization processing includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0115] It should be noted that the second analysis module 802, the second clustering module 804, the second determination module 806, the third clustering module 808 and the second optimization module 810 correspond to steps S302 to S310 in Example 1, and the examples and application scenarios implemented by the multiple modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules as part of the device can be run in the computer terminal 10 provided in Example 1.

[0116] Example 4

[0117] According to an embodiment of the present invention, a device for implementing the third method for processing stream data is also provided. Fig. 9 is a structural block diagram of a third stream data processing device according to Embodiment 4 of the present invention. Fig. 9 As shown, the device includes: a parsing module 902, a first display module 904 and a second display module 906. The device is described below.

[0118] The parsing module 902 is used to parse the collected view flow data and generate view parsing results if an assessment instruction is detected on the interactive interface of the city assessment system, wherein the view flow data includes: video flow data and / or image flow data; the first display module 904 is connected to the above-mentioned parsing module 902, and is used to display multiple batches of clustering results obtained after clustering the view parsing results in batches on the interactive interface, wherein each batch of clustering results includes: at least two view records with the same background, and the matching probability between any two view records; the second display module 906 is connected to the above-mentioned first display module 904, and is used to display the results of optimizing one or more batches of clustering results on the interactive interface, wherein the optimization processing includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0119] It should be noted that the above-mentioned analysis module 902, the first display module 904 and the second display module 906 correspond to steps S402 to S406 in Example 1, and the examples and application scenarios implemented by the multiple modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules as part of the device can be run in the computer terminal 10 provided in Example 1.

[0120] Example 5

[0121] The embodiment of the present invention can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.

[0122] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.

[0123] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the method for processing stream data of the application: parsing the stream data to generate parsing results; clustering the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and the matching probability between any two records; evaluating one or more records included in any clustering result to determine the probability that any two or more records belong to the target record; optimizing any clustering result based on the probability that any two or more records belong to the target record, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0124] Optionally, Fig.10 is a structural block diagram of a computer terminal according to an embodiment of the present invention. Fig.10As shown, the computer terminal may include: one or more (only one is shown in the figure) processors 101 and a memory 102.

[0125] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the processing detection method and device of the stream data in the embodiment of the present invention. The processor executes various functional applications and the processing of stream data by running the software programs and modules stored in the memory, that is, the above-mentioned stream data processing method is realized. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0126] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: parse the streaming data and generate parsing results; cluster the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and the matching probability between any two records; evaluate one or more records included in any clustering result to determine the probability that any two or more records belong to the target record; optimize any clustering result based on the probability that any two or more records belong to the target record, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0127] Optionally, the processor may also execute program code for the following steps: evaluating one or more records contained in any clustering result to determine the probability that any two or more records belong to the target record, including: sampling from the clustering result to obtain records to be evaluated that meet a predetermined number of entries; segmenting the records to be evaluated with a predetermined number of entries according to a predetermined segmentation method to obtain multiple subclasses in the clustering result; matching the records contained in each subclass with the target record to obtain the matching probability of each record contained in each subclass; determining the probability of the record in each subclass belonging to the target record based on the matching probability of each record contained in each subclass, wherein records with a probability higher than a threshold belong to the same type.

[0128] Optionally, the processor may also execute program code for the following steps: sampling records to be evaluated that meet a predetermined number of entries from the clustering results by any one or more of the following methods: randomly extracting records with a predetermined number of entries from the clustering results as records to be evaluated; extracting records with a predetermined number of entries from the clustering results according to a spatiotemporal distribution as records to be evaluated; calculating at least one nearest neighbor record of each record according to a feature vector of each record in the clustering results, and selecting a predetermined number of records from them as records to be evaluated, wherein the set of neighbor records of the selected records with the predetermined number of entries exceeds a predetermined order of magnitude.

[0129] Optionally, the processor may also execute program code for the following steps: after determining the probability of records belonging to the target record in each subclass, the method further includes: counting the number of records belonging to the target record in each subclass, and determining sample records that do not belong to the target record in each subclass; and fusing the sample records that do not belong to the target record in each subclass.

[0130] Optionally, the processor may also execute the program code of the following steps: clustering the analysis results in batches to obtain multiple batches of clustering results, including: dividing the analysis results according to timestamps to obtain at least one batch of classification results; clustering the classification results of each batch separately to obtain multiple batches of clustering results.

[0131] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: parse the stream data and generate parsing results; cluster the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and the matching probability between any two records; determine the class center from any clustering result, wherein the class center is used to characterize the attributes of the class to be assigned; based on the class center, perform secondary clustering on the records belonging to the same class in any clustering result, wherein the attributes of the records in the clustering results after the secondary clustering are the same; optimize any clustering result after the secondary clustering, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0132] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: if an assessment instruction is detected on the interactive interface of the city assessment system, parse the collected view flow data and generate a view analysis result, wherein the view flow data includes: video flow data and / or image flow data; display multiple batches of clustering results obtained after clustering the view analysis results in batches on the interactive interface, wherein each batch of clustering results includes: at least two view records with the same background, and the matching probability between any two view records; display the results of optimizing one or more batches of clustering results on the interactive interface, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0133] Optionally, the processor may also execute program code for the following steps: before displaying the results of optimizing one or more batches of clustering results on the interactive interface, the method also includes: evaluating one or more view records contained in any clustering result to determine the probability that any two or more view records contain the same target object; optimizing any clustering result based on the probability that any two or more view records contain the same target object.

[0134] It can be understood by those skilled in the art that Fig.10 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Fig.10 It does not limit the structure of the above electronic device. For example, the computer terminal may also include Fig.10 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Fig.10 Different configurations are shown.

[0135] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0136] Example 4

[0137] The embodiment of the present invention further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the method for processing stream data provided in the first embodiment.

[0138] Optionally, in this embodiment, the above storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0139] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: parsing stream data to generate parsing results; clustering the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and the matching probability between any two records; evaluating one or more records included in any clustering result to determine the probability that any two or more records belong to the target record; optimizing any clustering result based on the probability that any two or more records belong to the target record, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0140] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: evaluating one or more records contained in any clustering result to determine the probability that any two or more records belong to the target record, including: sampling from the clustering result to obtain records to be evaluated that meet a predetermined number of entries; segmenting the records to be evaluated with a predetermined number of entries according to a predetermined segmentation method to obtain multiple subclasses in the clustering result; matching the records contained in each subclass with the target record to obtain the matching probability of each record contained in each subclass; based on the matching probability of each record contained in each subclass, determining the probability of the record in each subclass belonging to the target record, wherein records with a probability higher than a threshold belong to the same type.

[0141] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: sampling from the clustering results to obtain records to be evaluated that meet a predetermined number of entries by any one or more of the following methods: randomly extracting a predetermined number of records from the clustering results as records to be evaluated; extracting a predetermined number of records from the clustering results according to a spatiotemporal distribution as records to be evaluated; calculating at least one nearest neighbor record of each record according to a feature vector of each record in the clustering results, and selecting a predetermined number of records therefrom as records to be evaluated, wherein the set of neighbor records of the selected predetermined number of records exceeds a predetermined order of magnitude.

[0142] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: after determining the probability of records belonging to the target record in each subclass, the method also includes: counting the number of records belonging to the target record in each subclass, and determining sample records that do not belong to the target record in each subclass; and fusing the sample records that do not belong to the target record in each subclass.

[0143] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: clustering the analysis results in batches to obtain multiple batches of clustering results, including: dividing the analysis results according to timestamps to obtain at least one batch of classification results; clustering the classification results of each batch separately to obtain multiple batches of clustering results.

[0144] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: parsing stream data to generate parsing results; clustering the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and the matching probability between any two records; determining the class center from any clustering result, wherein the class center is used to characterize the attributes of the class to be assigned; based on the class center, performing secondary clustering on the records belonging to the same class in any clustering result, wherein the attributes of the records in the clustering results after the secondary clustering are the same; optimizing any clustering result after the secondary clustering, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0145] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: if an assessment instruction is detected on the interactive interface of the city assessment system, parsing the collected view flow data to generate a view analysis result, wherein the view flow data includes: video flow data and / or image flow data; displaying on the interactive interface multiple batches of clustering results obtained after clustering the view analysis results in batches, wherein each batch of clustering results includes: at least two view records with the same background, and a matching probability between any two view records; displaying on the interactive interface the results of optimizing one or more batches of clustering results, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing.

[0146] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: before displaying the results of optimizing one or more batches of clustering results on the interactive interface, the method also includes: evaluating one or more view records contained in any clustering result to determine the probability that any two or more view records contain the same target object; and optimizing any clustering result based on the probability that any two or more view records have the same target object.

[0147] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0148] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0149] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0150] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0151] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0152] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.

[0153] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for processing stream data, characterized in that: include: Parsing the stream data to generate a parsing result, wherein the stream data is video stream data and the parsing result is a view parsing result; Clustering the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and a matching probability between any two records, wherein the records are pictures in the video stream data; Evaluate one or more records included in any of the clustering results to determine the probability that any two or more records belong to the target record; Based on the probability that any two or more records belong to the target record, optimizing any one of the clustering results, wherein the optimizing process includes any one or more of the following methods: splitting, merging, deleting and replacing; The analysis results are clustered in batches to obtain multiple batches of clustering results, including: The parsing results are divided according to timestamps to obtain at least one batch of classification results; and the classification results of each batch are clustered to obtain the clustering results of the multiple batches.

2. The method according to claim 1, characterized in that Evaluating one or more records included in any of the clustering results to determine the probability that any two or more records belong to the target record includes: Sampling from the clustering results to obtain records to be evaluated that meet a predetermined number of entries; Segmenting the predetermined number of records to be evaluated according to a predetermined segmentation method to obtain a plurality of subclasses in the clustering result; Match the records contained in each subclass with the target records, and obtain the matching probability of each record contained in each subclass; Based on the matching probability of each record contained in each subclass, the probability of the record belonging to the target record in each subclass is determined, wherein the records with probabilities higher than a threshold belong to the same type.

3. The method according to claim 2, characterized in that Sampling the records to be evaluated that meet the predetermined number of entries from the clustering results is performed in any one or more of the following ways: Randomly extracting the predetermined number of records from the clustering results as the records to be evaluated; Extracting the predetermined number of records from the clustering results according to the temporal and spatial distribution as the records to be evaluated; At least one nearest neighbor record of each record is calculated according to the feature vector of each record in the clustering result, and the records of the predetermined number of entries are selected as the records to be evaluated, wherein the set of neighbor records of the predetermined number of entries selected exceeds a predetermined order of magnitude.

4. The method according to claim 2, characterized in that: After determining the probability of the record belonging to the target record in each subclass, the method further includes: Counting the number of records belonging to the target record in each subclass, and determining sample records that do not belong to the target record in each subclass; The sample records in each subclass that do not belong to the target record are fused.

5. A method for processing stream data, characterized in that: include: Parsing the stream data to generate a parsing result, wherein the stream data is video stream data and the parsing result is a view parsing result; Clustering the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and a matching probability between any two records, wherein the records are pictures in the video stream data; Determine a cluster center from any of the clustering results, wherein the cluster center is used to characterize the attributes of the cluster to be assigned; Based on the class center, performing secondary clustering on the records belonging to the same class in any of the clustering results, wherein the attributes of the records in the clustering results after the secondary clustering are the same; Optimizing any clustering result after secondary clustering, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing; The analysis results are clustered in batches to obtain multiple batches of clustering results, including: The parsing results are divided according to timestamps to obtain at least one batch of classification results; and the classification results of each batch are clustered to obtain the clustering results of the multiple batches.

6. A method for processing stream data, characterized in that: include: If an evaluation instruction is detected on the interactive interface of the city evaluation system, the collected view flow data is parsed to generate a view parsing result, wherein the view flow data includes: video flow data and / or picture flow data; Displaying, on the interactive interface, a plurality of batches of clustering results obtained by clustering the view parsing results in batches, wherein each batch of clustering results includes: at least two view records with the same background, and a matching probability between any two view records; Displaying the results of optimizing one or more batches of clustering results on the interactive interface, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing; Among them, multiple batches of clustering results obtained after clustering the view analysis results in batches are displayed on the interactive interface, including: dividing the analysis results according to timestamps to obtain at least one batch of classification results; clustering the classification results of each batch separately to obtain the clustering results of the multiple batches.

7. The method according to claim 6, characterized in that Before displaying the result of optimizing the clustering results of one or more batches on the interactive interface, the method further includes: Evaluate one or more view records included in any one of the clustering results to determine the probability that any two or more view records include the same target object; Based on the probability that any two or more view records have the same target object, any one of the clustering results is optimized.

8. A stream data processing device, characterized in that: include: A first parsing module, configured to parse stream data and generate a parsing result, wherein the stream data is video stream data and the parsing result is a view parsing result; A first clustering module is used to cluster the analysis results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and a matching probability between any two records, wherein the records are pictures in the video stream data; An evaluation module, used to evaluate one or more records included in any of the clustering results, and determine the probability that any two or more records belong to the target record; A first optimization module, configured to optimize any one of the clustering results based on the probability that any two or more records belong to the target record, wherein the optimization includes any one or more of the following methods: splitting, merging, deleting and replacing; The first clustering module is further used to segment the analysis results according to timestamps to obtain at least one batch of classification results; and cluster the classification results of each batch respectively to obtain the clustering results of the multiple batches.

9. The device according to claim 8, characterized in that The evaluation module includes: A sampling module, used to sample from the clustering results to obtain records to be evaluated that meet a predetermined number of entries; A segmentation module, used to segment the predetermined number of records to be evaluated according to a predetermined segmentation method, and segment to obtain multiple subclasses in the clustering result; A matching module, used to match the records contained in each subclass with the target records, and obtain the matching probability of each record contained in each subclass; The first determination module is used to determine the probability of the record belonging to the target record in each subclass based on the matching probability of each record contained in each subclass, wherein the records with probabilities higher than a threshold belong to the same type.

10. A stream data processing device, characterized in that: include: A second parsing module, configured to parse the stream data and generate a parsing result, wherein the stream data is video stream data and the parsing result is a view parsing result; A second clustering module is used to cluster the parsing results in batches to obtain multiple batches of clustering results, wherein each batch of clustering results includes: at least two records and a matching probability between any two records, wherein the records are pictures in the video stream data; A second determination module is used to determine a cluster center from any of the clustering results, wherein the cluster center is used to characterize the attributes of the cluster to be assigned; A third clustering module, configured to perform secondary clustering on the records belonging to the same class in any of the clustering results based on the class center, wherein the attributes of the records in the clustering results after the secondary clustering are the same; A second optimization module is used to optimize any clustering result after the secondary clustering, wherein the optimization process includes any one or more of the following methods: splitting, merging, deleting and replacing; The second clustering module is further used to segment the analysis results according to timestamps to obtain at least one batch of classification results; and cluster the classification results of each batch respectively to obtain the clustering results of the multiple batches.

11. A stream data processing device, characterized in that: include: A parsing module, configured to parse the collected view flow data and generate a view parsing result if an evaluation instruction is detected on the interactive interface of the city evaluation system, wherein the view flow data includes: video flow data and / or picture flow data; A first display module is used to display, on an interactive interface, a plurality of batches of clustering results obtained by clustering the view parsing results in batches, wherein each batch of clustering results includes: at least two view records with the same background, and a matching probability between any two view records; A second display module is used to display the optimization processing results of one or more batches of clustering results on the interactive interface, wherein the optimization processing includes any one or more of the following methods: splitting, merging, deleting and replacing; The first display module is further used to segment the analysis results according to timestamps to obtain at least one batch of classification results; and cluster the classification results of each batch to obtain the clustering results of the multiple batches.

12. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the method for processing stream data according to any one of claims 1 to 7.

13. A computer device, characterized in that: include: Memory and processor, The memory stores a computer program; The processor is used to execute the computer program stored in the memory, and when the computer program is run, the processor executes the method for processing stream data according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video identity management method and device

    CN108470195A

  • Image clustering method and device, electronic equipment and storage medium

    CN113052245A