A warning method for group behavior

By establishing a target group database and setting parameters, the early warning problem of group behavior in social governance data is solved, efficient and accurate group behavior discovery and data recommendation are achieved, and complex environments with multiple data sources are adapted to.

CN118394804BActive Publication Date: 2025-07-08NANJING TRANRUNS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410272420.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-11
Publication Date
2025-07-08
Estimated Expiration
2044-03-11

AI Technical Summary

Technical Problem

In the case of processing multiple data sources, many data impurities, and inconsistent data reliability, it is difficult to efficiently and accurately detect and warn of group behaviors, with low computing efficiency, lots of noise in the result, and data cannot be effectively aggregated and traced back.

Method used

By establishing a target group database, correlating and cleaning identity information, standardized trajectory data are formed, combining peer and peer scenario calculation methods, setting parameters, calculating group density, filtering out accurate candidate groups, and early warning and recommendation of group behaviors.

Benefits of technology

It has achieved efficient and accurate discovery of group behaviors in social governance data, improved the accuracy of data recommendations, ensured the credibility and transparency of data, adapted to data mining needs in different scenarios, and solved the problem of many data sources and many results noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118394804B_ABST
    Figure CN118394804B_ABST
Patent Text Reader

Abstract

The present invention relates to a warning method for group behavior, which includes: determining a group target; associating all identity information of the target; associating trajectory data related to the target identity information from social governance data; performing data preprocessing and normalization processing on the trajectory data, and organizing it into point data corresponding one-to-one to the location and time of the target, so as to complete the generation of trajectory data of the group target; combining the calculation method of the social governance business background and the same-route scenario, setting calculation parameters, inputting the group's same-route situation, and obtaining the same-route situation of the group target; through setting evaluation indicators and evaluation methods, cleaning, mining and recommending the same-route situation of the group target, tracing the origin of the recommended same-route situation, while improving the data processing efficiency, completing the precise mining and discovery of the group behavior of the group target, enhancing the reliability and credibility of the data discovery results, and empowering the social governance business system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for warning of group behavior, belonging to the technical field of big data processing. Background Art

[0002] With the rapid development of the Internet and the continuous improvement of computing power, a large amount of data has been generated and stored. The rich information contained in this data has become an important resource that cannot be ignored. Discovering hidden information from massive data, improving data utilization, and realizing data value have become popular research directions in recent years. Combining the characteristics of social governance data and the business background with big data technology provides convenience for special departments to discover specific group behaviors as early as possible during the performance of their duties and issue timely warnings. At the same time, it is also conducive to maintaining public group security. Furthermore, by means of discovery and warning, group behaviors can be located and prevented, and the cause investigation and root cause analysis can be carried out as early as possible, so as to have a clear understanding of the possible consequences and take preventive measures.

[0003] There are also some existing methods for discovering group behaviors, including methods such as spatio-temporal clustering, graph algorithms, and social network analysis. These methods mostly process and analyze under the conditions of a single data source, a single processing scenario, and unified data reliability, mainly reflecting the characteristics of the data itself. However, with the development of Internet information and the continuous strengthening of information collection technology, there is more adaptability to the types of collected devices, but the problem is that the data is numerous and miscellaneous. In the actual application process, it will lead to difficulties in locating the activity locations and activity personnel of activities that may cause mass incidents. When traditional analysis methods deal with this kind of data with many dimensions and high sparsity, they will show phenomena such as low calculation efficiency, many result noises, and unmet expectations. Moreover, due to the particularity of social governance data and being limited by the current data collection technology, there will also be a problem of omission rate in the actual operation process. In order to enrich the data, a method of complementary multi-party data is mostly adopted to reduce the sparsity of the data as much as possible and improve the probability of the result at the same time. The problem brought by this method is that it is very difficult for traditional methods to complete the analysis and description of a specific situation when facing multi-data sources and inconsistent reliability of each data source, and it is difficult to meet the actual business needs.

[0004] The method for warning of group behavior fully combines the characteristics of social governance data and the characteristics of group behavior. While ensuring the calculation efficiency, it accurately recommends groups that may have group behaviors, and can complete multiple characteristics such as data traceability, data credibility, and group behavior description, filtering impurities to the greatest extent, ensuring data rigor, and empowering social governance services. In addition to effectively warning of mass incidents, this method can also adapt to data mining in other scenarios by adjusting parameters, such as the association of multiple different device account information of the same person. Summary of the Invention

[0005] The purpose of the present invention is to provide a warning method for group behavior. By setting business parameters and standardizing the data processing flow, data aggregation, retrospective analysis, and recommendation of big data for social governance are carried out to solve problems such as multiple data sources, many data impurities, inability to aggregate data, and inability to trace back results encountered in the data research and judgment process, and to solve problems such as low efficiency and a lot of noise existing in the prior art.

[0006] The purpose of the present invention can be achieved through the following technical solutions:

[0007] A warning method for group behavior includes the following steps:

[0008] S100. Define the target group, establish a target group database, and store the initial identity information of the target group.

[0009] S200. Associate other identity information of the target group from the social governance data to supplement and improve all the identity information of the target group.

[0010] S300. According to the supplemented and improved identity information, associate all the data that can represent the target at a certain place at a certain time related to these identity information from the social governance data to form standardized trajectory data.

[0011] S400. According to the peer scenario calculation method, combined with the business background of the group, set parameters to complete the calculation of the peer situation in the target group.

[0012] S500. According to the same route scenario calculation method, combined with the business background of the group and the results of S400, set parameters to complete the calculation of the same route situation in the target group to obtain a candidate group.

[0013] S600. According to the set evaluation indicators and methods, complete the calculation of the group tightness of the candidate group, evaluate, rank, and recommend the candidate group.

[0014] S700. Set business parameters to re-judge the situation where the targets involved in the candidate group satisfy multiple same routes, further screen the candidate groups that meet the business background to obtain a refined result group, and complete the mining of the group and group behavior.

[0015] Furthermore, in S200, the method for supplementing the other identity information includes: from the big data of social governance, based on the initial identity information, associate all relevant other identity information, and combine factors such as data collection, data source, and data quality to preset a credibility value in the range of (0.0, 1.0] for the reliability of each type of identity information. For example: the credibility of the initial identity information mobile phone number is 1.0, and the credibility of the WeChat number associated with it through the social governance data is 0.9, etc.

[0016] Further, in S300, the method for forming the standardized trajectory - type data by associating all the data related to identity information in social governance data that can represent the target at a certain time and place includes:

[0017] S301. For each identity information (I p1 , I p2 ,..., I pn ) of each target p in the target group P, associate all the information related to the trajectory in the big data of the social governance platform [(p, I p1 , t, d),...], where p represents the target, I p1 represents the first identity information of target p, t represents time, and d represents the longitude and latitude of the location;

[0018] S302. Conduct data cleaning and verification on the trajectory information of each target. If different identity information of a target appears at relatively far - away places at close times, then according to the credibility values of different types of identity information, discard the trajectory information with relatively low credibility as dirty data;

[0019] S303. Form a list of cleaned target trajectory information [(p, t, d),...] in the manner of target, time, and location longitude and latitude.

[0020] Further, in S400, the method for calculating the situation of group - target traveling together by combining the calculation method of the peer - traveling scenario and the business background of the group target includes:

[0021] S401. According to the business background, set three parameters: time radius deltaT, space radius r, and clustering density n;

[0022] S402. Process the list of trajectory information, slice the data by day according to time, and sort the data in each slice in ascending order of time;

[0023] S403. Process the list of trajectory information of each slice in a sliding - window manner. Calculate that within the time difference deltaT, if the distance between the locations where any two target ids are located is within r, then all the target ids that meet the conditions are considered to form a cluster, and obtain cluster information such as C[(p 1 , t1, d1), (p 2 , t2, d2),...], where p 1 represents the first target, p 2 represents the second target, and t and d respectively represent the corresponding appearance time and the longitude and latitude of the location;

[0024] S404. Any two targets within each cluster after forming the clusters are considered to be in the same line relationship;

[0025] Further, in S500, the method for calculating the same - route situation of the group target by setting parameters in combination with the business background of the group target according to the same - route scenario calculation method includes:

[0026] S501. According to the business background, set the parameter q of the number of inter - cluster associations, the number η of the minimum number of clusters in the candidate group, and the threshold α of the proportion of the number of occurrences of each target in the candidate group;

[0027] S502. According to the different cluster - type target values C calculated based on the same - line situation, calculate the number of repeated target values between different clusters under the same time slice. If the number of target values is greater than or equal to q, establish an association relationship between the two clusters;

[0028] S503. Repeat the above steps of establishing inter - cluster relationships until the inter - cluster relationships are stable. All clusters that can establish association relationships together are recorded as the set of cluster groups CS;

[0029] S504. Calculate the number of clusters C involved in the cluster group CS. If the number satisfies being greater than or equal to η, retain it; otherwise, discard it;

[0030] S505. Calculate all the target values p involved in the cluster group CS, count the number of times each target appears. If the number of times is greater than or equal to α * the number of clusters, retain the target in this cluster group; otherwise, discard it;

[0031] S506. Calculate the number of remaining target values and the number of clusters in the cluster group CS. If the number of remaining target values is greater than or equal to q and the number of clusters is greater than or equal to η, retain this cluster group, which is recorded as the candidate group;

[0032] Further, in S600, the evaluation indicators include: the number of clusters included in the candidate group, the number of target values involved in the candidate group, the average number of times each target value appears, the average number of targets associated with each cluster, the average number of target values in each cluster, the earliest time and the latest time span in the candidate group, the farthest distance between locations in the candidate group, and the average value of the earliest time intervals between clusters.

[0033] Further, in S600, for the evaluation indicators and methods, according to the business background and data quality situation, preset weights for each evaluation indicator, calculate the total weight, and use it as the tightness value TN of each candidate group. The calculation formula is:

[0034]

[0035] Where M represents the index sequence and W represents the weight sequence.

[0036] Further, in S700, the set parameters are used to re-judge the targets involved in the candidate population. The method includes:

[0037] S701. Set the start time, the total number of time slices D, the number of time slices d that meet the conditions, and the compactness threshold β;

[0038] S702. Screen the trajectory data according to the start time and the total number of time slices;

[0039] S703. Calculate the performance of all candidate populations on each time slice, and obtain the compactness value sequence [tn1, tn2,... tn D of each candidate population on each time slice;

[0040] S704. Calculate the number of times the compactness value is greater than or equal to β in the compactness value sequence of each candidate population. When the number of times is greater than or equal to d, the candidate population is selected as the final target and sorted in descending order according to the number of times the compactness is greater than or equal to β.

[0041] Advantages of the present invention: Starting from the big data of social governance, the present invention first determines the target group through the business background, establishes a target group database, and stores the initial identity information of the target group; then, according to the initial identity information of the group, all other relevant identity information is associated and supplemented from the social governance data; and according to factors such as data collection, data source, and data quality, a credibility value is preset for the reliability of each type of identity information; the identity information is associated with all trajectory information, and the trajectory data is cleaned and sorted to form a trajectory information list for each target; according to the calculation method of the same-trip scenario, business parameters are set to obtain a target cluster; according to the calculation method of the same-route scenario, business parameters are set to obtain a target cluster group; finally, parameters are set to obtain a precise final group from the target cluster group and make recommendations. Specifically, by designing a method for early warning of group behavior based on social governance data, on the premise of ensuring the security of social governance data, by setting credibility, the problems of multi-data sources and inconsistent reliability of each data source are solved; by the calculation methods of the same-trip and the same-route, the situations of a lot of result noise and the effect not meeting the expectation are solved; by setting various parameters for tuning, the situation that this method adapts to the discovery of different types of group behaviors with different natures is solved; based on the statistical calculation method, the data can be traced back and checked, and the data can be efficiently and transparently transferred, solving the problem of opaque data that appears when other neural network algorithms solve similar problems. Finally, on the premise of meeting the characteristics of social governance data, useful information is analyzed and mined from the massive data, the data of different data sources is refined, the data results are refined and integrated, and while discovering group gathering behaviors from the clues of individual behaviors, the accuracy of data recommendation is improved, making the data results interpretable, reliable, and usable, so as to solve the actual business needs. In addition, a method for early warning of group behavior based on social governance data proposed by the present invention is also applicable to the discovery problem of the association relationship between different devices and different accounts of the same person encountered in the data research and judgment process. By clarifying one account information of this person, when other multiple account information and this account information meet the method proposed by the present invention by setting parameters, it can be determined that these account information may belong to the same person, or in other words, they have a very close association relationship. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a flowchart of the method for discovering group behavior of the present invention;

[0043] Figure 2 is a flowchart of the calculation process of group discovery. DETAILED DESCRIPTION OF THE INVENTION

[0044] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. The present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention. To solve the problems existing in the prior art, the present invention provides a method for warning of group behavior.

[0045] In this solution, all the relevant information collected is legally collected with the consent of the user.

[0046] Embodiment: Please refer to Figure 1 As shown, in one embodiment, a method for warning of group behavior is provided, including the following steps:

[0047] S100. Define the target group, establish a target group database, and store the initial identity information of the target group; in combination with the business objective, find the initial identity information of relevant personnel from the social governance data, such as: the basic personal information of a certain type of specific personnel;

[0048] S200. Associate other identity information of the target group from the social governance data to supplement and improve all the identity information of the target group; according to the basic personal information obtained in S100, associate other related identity information in the social governance data, such as: mobile phone number, WeChat account, QQ account, mac account, imsi account, to form the supplemented target identity information;

[0049] S300. According to the supplemented and improved identity information, associate all the data that can represent the target at a certain place and time from the social governance data to form standardized trajectory data; in combination with the target identity information obtained in S200, associate all the time and place data in the social governance data, and perform data cleaning and sorting to form a standardized trajectory information list.

[0050] In S300 of this embodiment, the specific method for forming the data from the target identity information to the standardized trajectory information list is as follows:

[0051] S301. For each identity information (I p1 , I p2 ,..., I pn ) of each target p in the target group P, associate all the information related to the trajectory in the big data of the social governance platform, and organize and form a trajectory list in the format of target, identity information, time, and place, such as [(p, I p1 , t, d),...], where p represents the target, Ip1 The first type of identity information of the target p is represented, t represents time, and d represents the longitude and latitude of the location;

[0052] S302. Perform data cleaning and verification on the trajectory information of each target. If different identity information of a target appears at two locations with a large difference in distance at similar times, then based on the credibility values of different types of identity information, the trajectory information with relatively low credibility is discarded as dirty data;

[0053] The following is the formula for calculating the distance between two locations based on longitude and latitude:

[0054]

[0055] Where R represents the radius of the earth, 6371.0 km, x represents the longitude of the point, and y represents the latitude of the point.

[0056] S303. Form a list of cleaned target trajectory information [(p, t, d),...] in the manner of target, time, and location. Clean and filter the data according to the method of S302 to obtain the complete trajectory information after integrating various data source identity information of the target.

[0057] S400. According to the standardized trajectory information list obtained in S300, set parameters according to the calculation method of the same-trip scenario and in combination with the business background of the group, and complete the calculation of the same-trip situation in the target group to obtain clusters that meet the conditions.

[0058] In S400 of this embodiment, from the trajectory information to the calculation of the same-trip situation to obtain clusters that meet the conditions, the specific method is as follows:

[0059] S401. According to the business background, set three parameters: the time radius deltaT, the space radius r, and the clustering density n;

[0060] Where the space radius r means that the distance between two locations is less than or equal to r, and the calculation formula of r refers to the formula for calculating the distance between two points based on longitude and latitude in S302. The clustering density n means that when the number of distinct targets within the time span delatT and the space range with a radius of r is greater than or equal to n, the targets within this time-space segment form a cluster.

[0061] S402. Process the trajectory information list in S400, sort the trajectory information list in ascending order of time, and perform data sharding in units of days to form multiple data shards;

[0062] S403. Each data slice obtained in S402 is processed separately, and the trajectory information list of each slice is processed by sliding window method. It is calculated that within the time difference deltaT, if the distance between any two target p locations is within r, then all targets p that meet the conditions can form a cluster, and the cluster information is obtained as C[(p 1 , t1, d1),(p 2 , t2,d2),...], where p 1 represents the first target, p 2 represents the second target, t and d represent the latitude and longitude of the time and place where the corresponding target appears;

[0063] The sliding window method is to lock the first track data (p 1 , t0, d0), and compare each piece of data one by one until such a piece of data (p k , t k , d k ) satisfies t k - t0≤deltaT and the next data t k+1 - t0>deltaT, then from the first data to (p k , t k , d k ) All the data of this data is divided into a window. Then lock the second data and continue the previous step to get all the window data.

[0064] For the data in each window, the distance between the location of each data in the window and the locations of other data is calculated according to the distance calculation formula in S302, and an association relationship is established between two targets whose distance is less than or equal to r.

[0065] Analyze the relationship network composed of association relationships and find the fully connected graph in the relationship network. When the number of targets in the connected graph is greater than or equal to the clustering density n, all targets in the connected graph form a cluster.

[0066] S404. According to the method in S403, any two targets in each cluster after clustering are formed are considered to have a peer relationship;

[0067] S500. According to the same-path scenario calculation method, combined with the business background of the group and the results of S400, set parameters, further complete the calculation of the same-path situation in the target group from the peer relationship, and obtain the candidate group;

[0068] In S500 of this embodiment, parameters are set in combination with the group business background, and based on the result cluster of S400, the calculation of the same-route situation in the target group is completed to obtain the candidate group. The specific method is:

[0069] S501. Set the parameter q for the number of inter-cluster associations, the number η of the smallest clusters in the candidate population, and the threshold α for the proportion of the number of occurrences of each target in the candidate population according to the business background.

[0070] Set the corresponding parameters according to the actual business background. The parameters can be the optimal parameter combination obtained after multiple adjustments based on historical data and corresponding results, or the optimal parameter combination finally obtained through machine learning methods by iteratively calculating the results of multiple parameter combinations. The number of inter-cluster associations refers to the number of the same targets between two clusters.

[0071] S502. Calculate the different cluster class target values C calculated according to the peer situation, and calculate the number of duplicate target values between different clusters under the same data shard. If the number of target values is greater than or equal to q, establish an association relationship between the two clusters.

[0072] For the clusters that meet the conditions obtained according to S400, there will be multiple windows [Wind1, Wind2,...] under each data shard [Slice1, Slice2,...], and there will be multiple clusters [C1, C2,...] that meet the conditions in each window. Calculate the number of inter-cluster associations between clusters under different windows in each data shard. When the number of associations is greater than or equal to q, establish an association relationship between the two clusters.

[0073] For example, window Wind a contains clusters [C a1 , C a2 , C a3 ,...], and window Wind b contains clusters [C b1 , C b2 , C b3 ,...]. In this step, for each cluster in Wind a , calculate the number of inter-cluster associations between it and each cluster in Wind b . If the number of the same targets between two clusters is greater than or equal to q, establish an association relationship between the two clusters.

[0074] In the graph network, establish an edge between two clusters with an association relationship, and record the weight of the edge as the number of inter-cluster associations between the two clusters. For example: (C a1 , k, C b3 ), where k represents the number of inter-cluster associations between the two clusters that is greater than or equal to q.

[0075] S503. Repeat the above steps for establishing inter-cluster relationships until the inter-cluster relationships are stable. All the clusters that can establish association relationships together are recorded as the set of cluster groups CS.

[0076] Repeat the steps of S502 to complete the establishment of all cluster association relationships. When there is a qualified association relationship between a cluster A and a cluster B, take A and B as a cluster group cs. For the remaining cluster C, compare C with each cluster in the cluster group cs. As long as there is one that satisfies the association relationship, add C to the cluster group cs. Perform iterative calculations in this way until the number of cluster groups and the number of clusters in each cluster group reach stability, then stop, forming a cluster group set with many cluster groups CS.

[0077] S504. For each cluster group CS in the cluster group set, calculate the number of clusters C involved therein. If the number meets the requirement of being greater than or equal to η, retain the cluster group; otherwise, discard it.

[0078] S505. For each cluster group CS in the cluster group set, calculate all the target values p involved therein, and count the number of times each target appears. If the number of times is greater than or equal to α * the number of clusters, retain the target in this cluster group; otherwise, discard it.

[0079] S506. After step S505, for each cluster group CS in the cluster group set, calculate the remaining number of target values and the number of clusters therein. If the number of remaining target values after deduplication is greater than or equal to q, and the number of clusters in the cluster group is greater than or equal to η, retain the cluster group, denoted as the candidate group.

[0080] S600. Combining the business background and data quality situation, set evaluation indicators and methods for the candidate group, complete the calculation of the group compactness of the candidate group, and evaluate, rank, and recommend the candidate group.

[0081] Among them, the evaluation indicators include the number of clusters included in the candidate group, the number of target values involved in the candidate group, the number of times each target value appears on average involved in the candidate group, the average number of targets associated with each cluster, the average number of target values in each cluster, the earliest time and the latest time span in the candidate group, the farthest distance between locations in the candidate group, and the average value of the earliest time intervals between clusters. Statistically calculate these indicators for each candidate group, and under the condition of preset indicator weights, calculate the compactness value TN of each candidate group according to the following formula:

[0082]

[0083] Among them, M represents the indicator sequence, with a total of 9 indicators, and W represents the preset weight sequence, and each indicator corresponds to a preset weight.

[0084] Up to this step, on each time slice, we can obtain some candidate groups containing the corresponding target values and their time and location data, as well as the evaluation results of the compactness values of the candidate groups.

[0085] S700. Set business parameters to re-judge the situation where the targets involved in the candidate group satisfy multiple same routes, further screen the candidate group that meets the business background, obtain the refined result group, and complete the mining of the group and group behavior;

[0086] In S700 of this embodiment, based on the calculation result of S600, re-judge the situation where the targets involved in the candidate group satisfy multiple same routes, further screen the candidate group that meets the business background, obtain the refined result group, and complete the mining of the group and group behavior. The specific method is as follows:

[0087] S701. First, set the start time of the calculation, the total number of time slices D, the number of time slices d that meet the conditions, and the tightness threshold β;

[0088] S702. Screen the trajectory data according to the start time and the total number of time slices;

[0089] According to the start time and the total number of time slices, with one slice representing one day, screen the trajectory data for this calculation in the trajectory information list. In the actual operation process, it may be necessary to regularly screen the trajectory data for a period of time for calculation to continuously and quantitatively update the group data, ensure the real-time nature of the data, and at the same time endow the system with an early warning function in a timely manner.

[0090] S703. Calculate the performance of all candidate groups on each time slice, and obtain the tightness value sequence [tn1, tn2,... tn D of each candidate group on each time slice;

[0091] According to the process of S600, calculate the candidate groups on each time slice, and de-duplicate and count the candidate groups on all slices.

[0092] For each candidate group, calculate its tightness value on each time slice. If the group does not appear on a certain time slice, the tightness value of the group on that time slice is 0.

[0093] Finally, each candidate group can obtain a tightness value sequence of length D [tn1, tn2,... tn D ;

[0094] S704. Calculate the number of times the tightness value in the tightness value sequence of each candidate group is greater than or equal to β. When the number of times is greater than or equal to d, the candidate group is screened out as the final target, and it is sorted in descending order according to the number of times the tightness is greater than or equal to β.

[0095] The performance of the candidate group with unclear characteristics at a certain time slice is evaluated by the compactness value of the candidate group being less than β, while the compactness value being greater than or equal to β is considered as the performance with obvious characteristics at the time slice. When such obvious performance appears d times or more in D time slices, it is considered that the group has formed a group aggregation behavior. Due to the particularity of the group itself, this aggregation behavior will be considered as a possible group behavior and be selected as the final target, so as to realize the early warning of group behavior based on social governance data. The record of the entire calculation process is regarded as data backtracking for checking the accuracy of the data.

[0096] Through the above technical solution, the group behavior discovery model also ensures avoiding the influence of data multi-source on the calculation result by setting parameters, and also ensures that the model can adapt to more business scenarios and handle more data situations through parameter adjustment. When the original input data is relatively sparse, the calculation parameters of the same route can be appropriately enlarged, and when the data density is relatively large, the calculation parameters of the same row can be appropriately reduced to ensure that the selected data is more suitable for the current data situation and business scenario.

[0097] Obviously, in the above detailed description process, in a single implementation, the description of each feature is only explained by simple data logic and process to simplify the present disclosure. This disclosure method should not be interpreted as reflecting the intention that the implementation of the claimed subject matter requires more features than those clearly stated in each claim. The above embodiment content is only an example and explanation of the concept of the present invention. Those skilled in the art of the present technology can make various modifications or supplements or use similar methods to replace the described specific embodiments, as long as they do not deviate from the concept of the invention or exceed the scope defined by the claims of the present invention, they should belong to the protection scope of the present invention.

Claims

1. A warning method for group behavior, characterized in that, It includes the following steps: S100. Define the target group, establish a target group database, and store the initial identity information of the target group; S200. Associate other identity information of the target group from social governance data, and supplement and improve all the identity information of the target group; S300. According to the supplemented and improved identity information, associate all the data that can represent the target of the target group at a certain time and place in a certain place from the social governance data, and form standardized trajectory data; S400. According to the peer scenario calculation method, combined with the business background of the group, set parameters to complete the calculation of the peer situation in the target group; S500. According to the same route scenario calculation method, combined with the business background of the group and the results of S400, set parameters to complete the calculation of the same route situation in the target group, and obtain the candidate group; S600. According to the set evaluation indicators and methods, complete the calculation of the group tightness of the candidate group, evaluate, rank, and recommend the candidate group; S700. Set business parameters to re-judge the situation where the targets involved in the candidate group satisfy the same route multiple times, further screen the candidate groups that meet the business background, obtain the refined result group, and complete the mining of the group and group behavior; In S200, the method for supplementing other identity information is: associate all identity information related to the initial identity information from the big data of the social governance platform, and combine data collection, data source, and data quality factors to preset a credibility value for the reliability of each type of identity information; In S300, the method of forming standardized trajectory data by associating all the data that can represent the target of the target group at a certain time and place in a certain place from the social governance data according to the supplemented and improved identity information is: S301. For each identity information of each target in the target group, associate all the information related to the trajectory in the big data of the social governance platform; S302. Perform data cleaning and verification on the trajectory information of each target. If different identity information of a target appears in places that are far away and cannot be reached by a conventional means of transportation within a short time, then according to the preset credibility value of different types of identity information, discard the trajectory information with relatively low credibility as dirty data; S303. Form the cleaned target trajectory information in the manner of target, time, and place; In S400, the calculation of the peer situation of the group target by setting parameters according to the peer scenario calculation method and combined with the business background of the group target includes: S401. According to the business background, set three parameters: time radius deltaT, space radius r, and clustering density n; S402. Process the trajectory information table and slice the data by time; S403. Process the data in each slice in a sliding window manner. If the distance between the locations of any two targets within the time difference deltaT is within the space radius r, it is considered to meet the condition for forming a cluster target; find out the values of each cluster target that meet the conditions from the trajectory data; S404. Each cluster that meets the conditions is considered that the target values within the cluster belong to the same-row relationship under this condition; In S500, according to the same-route scenario calculation method, combined with the business background of the group target, set parameters to complete the calculation of the same-route situation in the target group, and obtain the candidate group, including: S501. Set the parameter q for the number of inter-cluster associations, the minimum number of clusters η in the candidate group, and the proportion threshold α of the number of occurrences of each target in the candidate group; S502. According to the different cluster class target values calculated based on the same-row situation, calculate the number of repeated target values between clusters under the same time slice. If the number of target values is greater than or equal to q, establish an association relationship between the two clusters; S503. Repeat the above steps for establishing inter-cluster associations until the inter-cluster relationship is stable. All clusters that can establish an association relationship are grouped together and recorded as the set of cluster groups; S504. Calculate the number of clusters in the cluster group. If the number meets the requirement of being greater than or equal to η, keep it; otherwise, discard it; S505. Calculate the number of occurrences of the target value in each retained cluster group. If the number is greater than or equal to α * the number of clusters, keep it; otherwise, discard it; S506. Calculate the number of target values and the number of clusters in the remaining cluster groups. If the number of remaining target values is greater than or equal to q and the number of clusters is greater than or equal to η, keep this cluster group and record it as the candidate group; In S700, the above-set parameters are used to re-judge the targets involved in the candidate group, including: S701. Set the start time, the total number of time slices D, the number of time slices d that meet the conditions, and the compactness threshold β; S702. Filter the trajectory data according to the start time and the total number of time slices; S703. Calculate the performance of all candidate groups on each time slice to obtain the compactness value of each candidate group on each time slice; S704. Calculate the number of times that the compactness value is greater than or equal to β in the compactness values of each candidate group. When the number of times is greater than or equal to d, the candidate group is selected as the final target group and sorted in descending order according to the number of times that the compactness is greater than or equal to β.

2. The early warning method for group behavior according to claim 1, wherein In S600, the evaluation indicators include: the number of clusters included in the cluster group, the number of target values involved in the cluster group, the average number of occurrences of each target value involved in the cluster group, the average number of targets associated with each cluster group, the average number of target values in each cluster, the time span between the earliest time and the latest time in the cluster group, the maximum distance between locations in the cluster group, and the average value of the earliest time intervals between clusters.

3. The early warning method for group behavior according to claim 1, wherein In S600, for the above evaluation indicators and methods, preset weights for each evaluation indicator according to the business background and data quality conditions, calculate the total weight, use it as the compactness value of each candidate group, and sort them in reverse order according to the group compactness value.

Citation Information

Patent Citations

  • Method and system for realizing instant messaging among persons traveling together, travel together information sharing and content recommendation

    CN107615733A

  • Group discovery algorithm model based on big data mining and analysis module

    CN111191147A