Group mining method and system based on video image trajectory analysis

Through depth-first search and Louvain algorithm based on video image trajectory analysis, the identification problem of specific group members in large-scale data is solved, and efficient and accurate group structure mining is achieved.

CN120336740APending Publication Date: 2025-07-18VIDEO INVESTIGATION DETACHMENT OF WUHAN PUBLIC SECURITY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510220971.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately identify and mine specific group members, especially in large-scale social data, where manual analysis is inefficient and results are inaccurate.

Method used

The method based on video image trajectory analysis is adopted to identify sub-graphs through a depth-first search algorithm, and community division is carried out in combination with the Louvain algorithm to identify group graph structures.

Benefits of technology

It realizes efficient and accurate identification of specific group structures in large-scale data, reduces computational complexity and improves identification accuracy, and supports accurate analysis of specific groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336740A_ABST
    Figure CN120336740A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data mining, and particularly provides a group mining method and system based on video image trajectory analysis, and the method comprises the steps: obtaining a multi-dimensional attribute data set of a plurality of attention targets; focusing targets in the multi-dimensional attribute data set are defined as nodes, two nodes passing through the same attribute case are connected through edges, the edges connected end to end form a path, and a graph is obtained; a depth-first search algorithm is adopted to search the graph, a sub-graph is obtained through recognition, and any two nodes in the sub-graph are connected through a path; for each sub-graph, a Louvain algorithm is adopted to carry out more detailed community division, a group graph is obtained, and any two nodes in the group graph are in edge connection. The method is suitable for processing a large-scale and complex network structure, and can provide an accurate specific group structure mining result while ensuring high efficiency. The system can quickly screen out important data and deeply analyze the important data, so that identification and analysis work of a specific group is effectively supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data mining, and particularly to a group mining method and system based on video image trajectory analysis. Background Art

[0002] In current society, some behaviors show a development trend of being specific group (i.e., group)-oriented, structured, and purposeful. Compared with the behaviors of individuals in single events, due to the complex organizational structure of specific groups, their behaviors are more complex, and it is difficult to mine personnel correlations and obtain associated groups. Therefore, quickly and accurately identifying members of specific groups and mining specific groups is of great significance for improving the effect of specific group mining.

[0003] With the construction of informatization and the rapid development of the Internet, various types of social data are also increasing day by day. These data provide important support for the mining of specific groups, such as shopping mall transaction data, street travel video image data, traffic travel data, etc. How to associate these data to mine personnel groups with commonalities. However, due to the structural nature of specific groups, it is difficult and inefficient to rely solely on manual analysis of a large amount of social data to mine specific groups. Therefore, there is an urgent need for efficient and accurate big data analysis means to conduct specific group mining.

[0004] As an algorithm and method for discovering group, group, or community structures in a network or graph, specific group mining analysis usually has some limitations in the process of use. In particular, the problem of omission in specific group mining can lead to inaccurate or incomplete mining results. Summary of the Invention

[0005] The present disclosure aims to at least solve one of the technical problems existing in the prior art, and proposes a group mining method and system based on video image trajectory analysis.

[0006] In a first aspect, the present disclosure provides a group mining method based on video image trajectory analysis, including:

[0007] S1, obtaining a multi-dimensional attribute data set of multiple target objects of interest;

[0008] S2, defining the target objects of interest in the multi-dimensional attribute data set as nodes, connecting two nodes that have participated in the same attribute case with an edge, and the edges connected end to end form a path to obtain a graph;

[0009] S3, using a depth-first search algorithm to search the graph, and identifying a subgraph, where there is a path connection between any two nodes in the subgraph;

[0010] S4. For each of the subgraphs, use the Louvain algorithm for a more detailed community division to obtain a group graph, where there is an edge connection between any two nodes in the group graph.

[0011] Preferably, the multi-dimensional attribute data includes: specific personnel attributes and specific event behavior attributes.

[0012] Preferably, S3 specifically includes:

[0013] Let G=(V, E) be a graph, where V is the set of nodes, E is the set of edges, and s∈V is the starting node;

[0014] First step, initialization: Create a stack S and push the starting node s onto the stack. At the same time, create a set Visited to store the visited nodes, and add the starting node s to the set Visited.

[0015] Second step, explore nodes: When the stack S is not empty, pop a node u from the top of the stack, check all neighbor nodes v of u. If v has not been visited, push v onto the stack and mark it as visited, that is, add v to the set Visited.

[0016] Third step, recursion: Repeat the second step until the stack S is empty.

[0017] Fourth step, end: When the stack is empty, it means that all reachable nodes starting from the starting node s have been visited, and the search ends.

[0018] Preferably, S2 specifically includes:

[0019] According to the different numbers of edges connected to the nodes, different weights are given to the corresponding nodes, and the larger the number, the higher the weight.

[0020] Preferably, S4 specifically includes:

[0021] Let G=(V, E) be a graph, where V is the set of nodes, E is the set of edges, and the modularity Q is an index used to quantify the quality of community division, defined as:

[0022]

[0023] where A ij is the weight of the edge between nodes i and j, k i and k j are the degrees of nodes i and j, m is the sum of the weights of all edges in the graph, and δ(c i , c j ) is an indicator function, which is 1 when c i =c j and 0 otherwise;

[0024] Step 1, initialization: Each node is regarded as an independent community;

[0025] Step 2, local optimization: For each node, try to move it to the community of its neighbor. If this can increase the modularity Q, then make the move;

[0026] Step 3, community merging: Repeat Step 2 until the modularity Q cannot be improved by moving any node;

[0027] Step 4, community aggregation: Regard the formed communities as new nodes and reconstruct the network. The weight of the edge is equal to the sum of the weights of the original edges between the two communities;

[0028] Step 5, repeated optimization: Repeat Step 2 and Step 3 on the newly reconstructed network, continue to optimize and merge until the modularity Q reaches the maximum;

[0029] Preferably, the connection density of the nodes within the community finally obtained in the fifth step is greater than the connection density of the nodes between different communities.

[0030] Preferably, the S1 further includes data cleaning:

[0031] S10, data cleaning filters out dirty data and retains data with correct certificate information;

[0032] S11, uses SQL statements to filter and de-duplicate, de-duplicates the information of personnel in attribute cases and de-duplicates the attribute case information.

[0033] Preferably, the S10 specifically includes:

[0034] S101, for data with illegal characters in the ID field, remove the illegal characters and unify the case standard;

[0035] S102, if there is a unique identification code for the person, the data is retained;

[0036] S103, if the number of digits of the unique identification code for the person is correct but the check digit fails, resulting in the inability to directly verify the logic, then retain this part of the data;

[0037] S104, eliminate data that does not conform to normal rules, missing data, single-person attribute case data, and data with no clear meaning, and ensure that the processed data can construct relational data.

[0038] The present invention also provides a group mining system based on video image trajectory analysis. The system can be used to implement the above-mentioned group mining method based on video image trajectory analysis. The system includes:

[0039] A data acquisition module configured to acquire multi-dimensional attribute data sets of multiple attention targets;

[0040] A data processing module, configured to define the focus target in the multi-dimensional attribute dataset as nodes, connect two nodes that have participated in cases with the same attribute through edges, and the edges that are connected end to end form a path to obtain a graph;

[0041] A first screening module, configured to search the graph using a depth-first search algorithm to identify a subgraph, and there is a path connection between any two nodes in the subgraph;

[0042] A second screening module, configured to perform more detailed community division on each of the subgraphs using the Louvain algorithm to obtain a group graph, and there is an edge connection between any two nodes in the group graph.

[0043] The present invention also provides an electronic device, including:

[0044] One or more processors;

[0045] A memory for storing one or more programs;

[0046] When the one or more programs are executed by the one or more processors, the one or more processors implement the group mining method based on video image trajectory analysis. Description of the Drawings

[0047] Figure 1 It is a flowchart of a group mining method based on video image trajectory analysis provided by an embodiment of the present disclosure;

[0048] Figure 2 It is a data graph provided by an embodiment of the present disclosure;

[0049] Figure 3 It is a distribution graph of the number of street behaviors after filtering samples provided by an embodiment of the present disclosure;

[0050] Figure 4 It is a structural block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed Embodiment

[0051] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the present disclosure will be further described in detail below in conjunction with the drawings and specific embodiments.

[0052] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the field to which this disclosure belongs. The "first", "second" and similar terms used in this disclosure are not for any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "a", "an" or "the" are not limited by quantity, but mean that there is at least one. Words such as "including" or "comprising" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used for relative position relationships, and when the absolute position of the object being described changes, the relative position relationship may also change accordingly.

[0053] In the respective drawings, like elements are denoted by like reference numerals. For the sake of clarity, not all parts in the drawings are drawn to scale. In addition, some well-known parts may not be shown in the figures.

[0054] Many specific details of this disclosure are described below, such as the structure, materials, dimensions, processing techniques and technologies of components, in order to understand this disclosure more clearly. However, as those skilled in the art can understand, this disclosure can be implemented without these specific details.

[0055] As Figure 1 shown, an embodiment of the present invention provides a group mining method based on video image trajectory analysis, including:

[0056] S1, obtaining a multi-dimensional attribute data set of multiple target objects of interest;

[0057] S2, defining the target objects of interest in the multi-dimensional attribute data set as nodes, connecting two nodes that have participated in the same attribute case with an edge, and the edges connected end to end form a path to obtain a graph;

[0058] S3, using a depth-first search algorithm to search the graph to identify subgraphs, and there is a path connection between any two nodes in the subgraph;

[0059] S4, for each of the subgraphs, using the Louvain algorithm for more detailed community division to obtain a group graph, and there is an edge connection between any two nodes in the group graph.

[0060] It should be particularly noted that the order of the above steps S1-S4 is not limited.

[0061] This method identifies and mines specific groups by combining connectivity analysis and community relationship algorithms (Louvain algorithm). The key parameters include:

[0062] Node: Represents a single entity in the dataset (such as people walking together in the same space on the street, or people with overlapping spatio-temporal relationships at the same counter in a mall).

[0063] Edge: Represents the association relationship between entities (such as the association of multiple points in the same space on the street, or the overlapping spatio-temporal relationships of multiple different counters).

[0064] Connected Component: A subgraph in which there is a path between any two nodes.

[0065] Community: A group of nodes with close internal connections and relatively few connections to external nodes.

[0066] Modularity: An index to measure the quality of community division, indicating the difference between the density of internal connections in the community and the corresponding connections in a random graph. A high modularity indicates dense internal connections in the community and sparse connections to the outside.

[0067] Generally speaking, people with spatio-temporal overlaps in different locations tend to have type convergence. For example, often being in the same street area (such as a mall, park, etc.) often indicates the same hobbies (such as entertainment, exercise, etc.). Therefore, such people can be associated. This facilitates the platform to analyze relevant data.

[0068] Based on the above method, the idea can be further elaborated as follows:

[0069] 1. Connectivity analysis

[0070] Initial division: First, perform connectivity analysis on the graph to identify the connected components in the graph. Each connected component is a subgraph, and the nodes of the subgraph are connected by paths. This step helps to identify the initial group structure and exclude nodes that have no relationship with other parts at all.

[0071] Data filtering: Through connectivity analysis, those independent nodes or subgraphs outside the scope of interest can be pre-filtered, thereby reducing the data volume and concentrating computing resources on more meaningful parts.

[0072] Specifically, it includes the following steps:

[0073] 1. Graph construction: Construct the dataset as a graph, where the nodes represent person entities and the edges represent the relationships of people participating in the same attribute cases.

[0074] 2. Connectivity Analysis: Apply graph algorithms (Depth-First Search DFS) to identify connected components.

[0075] 3. Subgraph Partitioning: Divide the large graph into multiple subgraphs, each subgraph being a connected component for further analysis.

[0076] Depth-First Search (DFS) is an algorithm used to traverse and search trees or graphs. The characteristic of DFS is to search each branch path as deeply as possible, usually implemented through recursive or stack data structures. It starts from the starting node, searches each branch as deeply as possible until the end of the path is reached, then backtracks to the previous node and continues to explore the next unvisited path.

[0077] Mathematical Representation:

[0078] Let G=(V, E) be a graph, where V is the set of nodes (such as the set of people often appearing in street park areas), and E is the set of edges (that is, the set of relationships of people often appearing in the park participating in cases of the same attribute, such as having overlapping time and space, or participating in square dancing together). s∈V is the starting node.

[0079] Algorithm Steps:

[0080] 1. Initialization: Create a stack S and push the starting node s onto the stack. At the same time, create a set Visited to store the visited nodes and add s to Visited.

[0081] 2. Explore Nodes:

[0082] When the stack S is not empty, pop a node u from the top of the stack.

[0083] Check all neighbor nodes v of u:

[0084] If v has not been visited (i.e., ), push v onto the stack and mark it as visited (i.e., add v to the Visited set).

[0085] 3. Recursion: Repeat step 2 until the stack S is empty.

[0086] 4. End: When the stack is empty, it means that all reachable nodes starting from the starting node s have been visited and the search ends.

[0087] DFS can effectively identify all connected components in the graph. Starting a DFS on an unvisited node in the graph can traverse the entire connected component where the node is located, thus quickly identifying each connected subgraph in the data and providing a basis for further specific group analysis. Especially when dealing with large-scale data, through preliminary connectivity analysis, dividing the entire large graph into multiple subgraphs can effectively reduce the complexity and computational amount of subsequent processing.

[0088] 2. Community Division

[0089] Refined Division: On each connected component generated by the connectivity analysis, the Louvain algorithm is applied for community detection. The Louvain algorithm determines the community structure by locally optimizing modularity. The algorithm first treats each node as an independent community and then gradually merges nodes to maximize modularity, finally forming multiple communities.

[0090] Iterative Optimization: The Louvain algorithm undergoes multiple rounds of iterative optimization to gradually merge and reduce communities until the modularity cannot be further improved.

[0091] The Louvain algorithm can effectively discover community structures in networks of different scales and is particularly suitable for large-scale networks. The hierarchical approximation and fast modularity optimization of the algorithm make it a common choice in community detection. By gradually optimizing and merging communities, the Louvain algorithm not only improves the efficiency of community detection but also maintains high accuracy and resolution.

[0092] For the business scenario of specific group mining, the Louvain algorithm provides an accurate and computationally efficient method to analyze and identify group structures in large-scale personnel relationship networks.

[0093] 3. Joint Application:

[0094] Phase 1: First, perform connectivity analysis to preliminarily divide the connected components in the large-scale data, and identify and isolate the significantly independent nodes or small groups. This reduces the amount of data that requires complex community detection.

[0095] Phase 2: Apply the Louvain algorithm to the larger connected components for more detailed community division. By optimizing modularity, accurately identify the community structures with close internal connections.

[0096] Dynamic Adjustment: In practical applications, the thresholds of connectivity analysis and community division can be adjusted according to the group structure and business requirements. For example, the processing of small-scale connected components can be directly marked or further refined.

[0097] This method combines connectivity analysis with the community relationship algorithm, combining efficiency and accuracy. First, connectivity analysis can quickly identify and separate different group structures in the data, reducing the number of nodes and complexity to be processed. Then, the Louvain algorithm performs fine-grained community division within each connected component, further improving the accuracy of community identification.

[0098] This combined method is particularly suitable for dealing with large-scale and complex network structures, and can provide accurate mining results of specific group structures while ensuring high efficiency. By combining these two algorithms, the system can quickly screen out important data and conduct in-depth analysis, thus effectively supporting the identification and analysis of specific groups.

[0099] In a preferred solution, the multi-dimensional attribute dataset includes: specific personnel attributes and specific event behavior attributes. The specific personnel attributes include: a unique identification code for personnel, and the entities involved in the data are mainly: attribute cases and targets of concern. There are many personnel attributes, and the multi-dimensional attribute data of each target of concern contains information related to all event behavior attributes of the specific personnel attributes. There are many specific event behavior attributes of the attribute case entity, including any events that may be associated, such as traveling together in the same time and space, practicing Tai Chi together; participating in square dancing together, participating in competitions together, etc.

[0100] In a preferred solution, S2 specifically includes: according to the different numbers of edges connected to the nodes, different weights are given to the corresponding nodes, and the larger the number, the higher the weight. The higher the weight indicates that the target of concern corresponding to the node is more critical and is one of the group members that need to be mined.

[0101] In addition, the case personnel within the same attribute case can be sorted first, and the case personnel are arranged in order according to the severity. Then, they are combined in pairs to form personnel relationship combination data, and weights are assigned according to the attribute case category or other criteria, and the aggregated weighted sum processing is performed on the data of the same combination, so as to obtain the total weight of each combination, and this data is the interpersonal network data of the members.

[0102] In a further solution, S4 specifically includes:

[0103] Let G=(V, E) be a graph, where V is the set of nodes, E is the set of edges, and the modularity Q is an index used to quantify the quality of community partitioning, defined as:

[0104]

[0105] where A ij is the weight of the edge between nodes i and j, k i and k j are the degrees of nodes i and j, m is the sum of the weights of all edges in the graph, and δ(c i , c j ) is an indicator function, which is 1 when c i =c j and 0 otherwise;

[0106] The first step is initialization: each node is used as an independent community;

[0107] Step 2, Local Optimization: For each node, attempt to move it to the community of its neighbor. If this can increase the modularity Q, then make the move.

[0108] Step 3, Community Merging: Repeat Step 2 until the modularity Q cannot be improved by moving any node.

[0109] Step 4, Community Aggregation: Treat the communities that have been formed as new nodes, and reconstruct the network. The weight of an edge is equal to the sum of the weights of the original edges between the two communities.

[0110] Step 5, Repeated Optimization: Repeat Step 2 and Step 3 on the newly reconstructed network, continue to optimize and merge until the modularity Q reaches the maximum.

[0111] After the groups are mined, according to the weights of each focus target in the previous step S2, the risk of each node corresponding to the focus target within the group can be further confirmed, which is beneficial to more accurately locate individuals, improve the detection efficiency, and save human resources.

[0112] In a specific implementation scenario, the specific group mining model mainly targets specific groups entering and leaving the park on the street surface. The specific group mining model is based on the park social activity data, focusing on two parts: mining specific group members and verifying event relevance feedback. Taking park exercise personnel as the core, by deeply studying the characteristics and spatio-temporal relationships of park access personnel, summarizing the interpersonal network of park access personnel, and using technical means such as statistics and community division to achieve accurate identification of specific groups and screening of core members, so as to improve the accuracy of data push with computing power and greatly enhance the connectivity of relevant group members.

[0113] In terms of data: The data accessed by this model are the attribute cases and personnel information data of the whole city in recent years. The park activity data provides data to us in the way of providing database connection information, which can be obtained incrementally or in full. The main data tables and corresponding primary keys are shown in Table 1 below:

[0114] Table 1 Data Source Information

[0115]

[0116] Attribute Case Information: It mainly includes the basic information of the attribute case, 60 fields such as the attribute case number, input time, attribute case name, attribute case category, behavior occurrence time (such as square dancing time), and behavior occurrence location (such as park).

[0117] Person-Case Association: It mainly includes the basic information of the attribute case, 14 fields in total such as the attribute case number, personnel number, input time, attribute case name, and personnel type.

[0118] Case personnel information: mainly includes the basic information of attribute cases, such as personnel number, input time, ID number, personnel type, age, etc., a total of 94 fields.

[0119] Entity data processing: as Figure 2 shown, the entities involved in the data are mainly: attribute cases, objects of interest. Among them, there are many personnel attributes, including all relevant information of the objects of interest such as age, occupation, household register, etc. There are many case types of attribute case entities, including any events that may be associated, such as square dancing, tai chi; participating in a basketball game, walking a pet, etc., with the start and end time being 20200421 - 20231217, and the model analysis mainly expands for these two entities.

[0120] According to the analysis business requirements, filter out the personnel who enter and leave the park more than 3 times a month and have at least 1 same spatio - temporal relationship overlap, as well as the attribute case data associated with these personnel as the analysis objects for specific group mining applications. As Figure 3 shown, the personnel who enter and leave the park more than 3 times a month are in the blue area, and the personnel with at least one same spatio - temporal relationship overlap in a month are in the yellow area, that is, the personnel who have met at least once in the park in a month. That is to say, the yellow area means that in the park within a month, at least two people have had the same behavior together, such as square dancing together, etc.

[0121] After data exploration, it is found that the overall quality of the original data is good, but there are still problems such as duplicate data and dirty data. And the data for model analysis should be ensured to be pure enough. Therefore, corresponding pre - processing operations are performed on the data, including cleaning of dirty data and invalid data, and unification of data formats, only retaining the data that we need and is valid:

[0122] S101, for the data with illegal characters in the ID field, remove the illegal characters and unify the case standard; the personnel unique identification code is composed of pure numbers and can be set artificially. And there is a check digit among them.

[0123] S102, if there is a personnel unique identification code, retain this part of the data;

[0124] S103, if the number of digits of the personnel unique identification code is correct but the check digit fails, resulting in the inability to directly verify the logic, then retain this part of the data;

[0125] S104, eliminate the data that does not conform to the normal rules, missing data, single - person attribute case data, and data without clear meaning, ensuring that the processed data can construct relational data.

[0126] Use SQL statements for filtering and deduplication, while screening out the data of interest for the business, also ensuring the purity of each table (deduplication of personnel information, deduplication of attribute case information);

[0127] After preprocessing, the data becomes relatively pure. There are 60,164 remaining target samples, among which 15,068 samples are related to square dancing. After cleaning the attribute cases, there are 15,470 remaining attribute cases. There are 4,828 overall square dancing cases. Among them, the attribute case with the largest number of cases is 111 people, and the smallest is 2 people. Single-person attribute cases have been filtered out. The distribution of the cleaning completion results is shown in Table 2 below:

[0128] Table 2 Distribution Table of Attribute Cases after Filtering Samples

[0129]

[0130]

[0131] The distribution of the personnel cases after filtering samples is shown in Table 3 below:

[0132] Table 3 Distribution of the Number of Personnel Cases after Filtering Samples

[0133] Number of personnel cases Number of people Number of people doing square dancing 1 11897 6776 2 15968 2783 3 3097 911 4 893 305 5 316 131 6 124 35 7 56 23 8 25 9 9 9 4 10 1 0 11 3 2 12 3 3 13 3 1 14 1 1 15 6 0 16 5 0 17 3 0 18 1 1 19 1 1 20 3 2 21 1 1

[0134] The distribution of the number of personnel cases after filtering samples is as Figure 3 shown.

[0135] Relationship data processing: Based on all the data related to this application, sort out the basic information of all case personnel and their members under each attribute case. First, sort the case personnel within the same attribute case, and then combine them in pairs to form personnel relationship combination data. Assign weights according to the attribute case category or other criteria, and perform aggregation and weighted summation processing on the data of the same combination to obtain the total weight of each combination. This data is the interpersonal relationship network data of the members.

[0136] For example, in the construction of the relationship among Zhang San, Li Si, and Wang Wu, if the square dancing attribute case information of Zhang San, Li Si, and Wang Wu is as shown in Table 4 below:

[0137] Table 4 Member Information of Attribute Cases

[0138]

[0139]

[0140] Among them, the different attribute case numbers can be people who do square dancing in different places, or people who have appeared in the same place and done square dancing or tai chi or other behaviors.

[0141] Then, through pairwise combination of the attribute case personnel, the following data can be obtained, which is the input data of the algorithm model, as shown in Table 5 below:

[0142] Table 5 Attention Target Relationship Weight Table

[0143] Personnel 1 Personnel 2 Weight Zhang San Li Si 2 Zhang San Wang Wu 1 Li Si Wang Wu 1

[0144] Construct a graph: First, connect to the database to load entity and relationship data, and then use the networkx program to construct a graph with the personnel entity data as nodes and the same-attribute case relationship data between members as edges.

[0145] An embodiment of the present invention also provides a group mining system based on video image trajectory analysis. The system can be used to implement the group mining method based on video image trajectory analysis described above. The system includes:

[0146] A data acquisition module configured to acquire multi-dimensional attribute data sets of multiple attention targets;

[0147] A data processing module configured to define the attention targets in the multi-dimensional attribute data set as nodes, connect the two nodes that have participated in the same-attribute case through edges, and the edges that are connected end to end form a path to obtain a graph;

[0148] A first screening module configured to search the graph using a depth-first search algorithm to identify subgraphs, and there is a path connection between any two nodes in the subgraph;

[0149] A second screening module configured to perform more detailed community division on each of the subgraphs using the Louvain algorithm to obtain a group graph, and there is an edge connection between any two nodes in the group graph.

[0150] Please refer to Figure 4 which is a schematic diagram of an embodiment of an electronic device provided by an embodiment of the present disclosure. As Figure 4 shown, an embodiment of the present disclosure provides an electronic device 1300, including a memory 1310, a processor 1320, and a computer program 1311 stored in the memory 1310 and executable on the processor 1320. When the processor 1320 executes the computer program 1311, the following steps are implemented: S1, acquire multi-dimensional attribute data sets of multiple attention targets;

[0151] S2, define the attention targets in the multi-dimensional attribute data set as nodes, connect the two nodes that have participated in the same-attribute case through edges, and the edges that are connected end to end form a path to obtain a graph;

[0152] S3, search the graph using a depth-first search algorithm to identify subgraphs, and there is a path connection between any two nodes in the subgraph;

[0153] S4. For each of the subgraphs, the Louvain algorithm is used for more detailed community division to obtain a group graph, and there is an edge connection between any two nodes in the group graph.

[0154] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principle of the present disclosure. However, the present disclosure is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present disclosure, and these modifications and improvements are also regarded as the protection scope of the present disclosure.

Claims

1. A group mining method based on video image trajectory analysis, characterized in that Including: S1. Obtain multi-dimensional attribute datasets of multiple focus targets; S2. Define the focus targets in the multi-dimensional attribute datasets as nodes, connect two nodes that have participated in the same attribute cases with edges, and the edges that are connected end to end form a path to obtain a graph; S3. Use the depth-first search algorithm to search the graph, identify subgraphs, and there is a path connection between any two nodes within the subgraphs; S4. For each of the subgraphs, use the Louvain algorithm for more detailed community division to obtain a community graph, and there is an edge connection between any two nodes within the community graph.

2. The group mining method based on video image trajectory analysis according to claim 1, characterized in that The multi-dimensional attribute data includes: specific personnel attributes and specific event behavior attributes.

3. The group mining method based on video image trajectory analysis according to claim 1, characterized in that, The specific content of S3 includes: Let G=(V, E) be a graph, where V is the node set, E is the edge set, and s∈V is the starting node; The first step, initialization: create a stack S and push the starting node s onto the stack, and at the same time create a set Visited to store the visited nodes, and add the starting node s to the set Visited; The second step, explore nodes: when the stack S is not empty, pop a node u from the top of the stack, check all neighbor nodes v of u, if v has not been visited, push v onto the stack and mark it as visited, that is, add v to the set Visited; The third step, recursion: repeat the second step until the stack S is empty; The fourth step, end: when the stack is empty, it means that all reachable nodes starting from the starting node s have been visited and the search ends.

4. The group mining method based on video image trajectory analysis according to claim 1, characterized in that The specific content of S2 includes: According to the different numbers of edges connected to the nodes, assign different weights to the corresponding nodes, and the larger the number, the higher the weight.

5. The group mining method based on video image trajectory analysis according to claim 4, wherein The specific content of S4 includes: Let G=(V, E) be a graph, where V is the node set, E is the edge set, and the modularity Q is an index used to quantify the quality of community division, defined as: where A ij is the weight of the edge between nodes i and j, k i and k j are the degrees of nodes i and j, m is the sum of the weights of all edges in the graph, δ(c i , c j ) is an indicator function that is 1 when c i = c j and 0 otherwise; The first step, initialization: each node is an independent community; The second step, local optimization: for each node, try to move it to the community of its neighbor. If doing so can increase the modularity Q, then make the move; The third step, community merging: repeat the second step until the modularity Q cannot be improved by moving any node; The fourth step, community aggregation: regard the formed communities as new nodes and reconstruct the network, and the weight of the edge is equal to the sum of the weights of the original edges between the two communities; The fifth step, repeated optimization: repeat the second step and the third step on the newly reconstructed network, continue to optimize and merge until the modularity Q reaches the maximum.

6. The group mining method based on video image trajectory analysis according to claim 5, wherein The connection density of the nodes within the community finally obtained in the fifth step is greater than the connection density of the nodes between different communities.

7. The method for group mining based on video image trajectory analysis according to claim 1, wherein S1 further includes data cleaning: S10. Filter out dirty data during data cleaning and retain the data with correct certificate information; S11. Use SQL statements to filter and remove duplicates, remove duplicates from the personnel information of attribute cases and remove duplicates from the information of attribute cases.

8. The group mining method based on video image trajectory analysis according to claim 7, wherein The specific content of S10 includes: S101. For the data with illegal characters in the ID field, remove the illegal characters and unify the case standard; S102. If there is a unique identifier for the person, the data is retained. S103. If the number of digits of the unique personal identification code is correct but the check digit fails, resulting in the inability to directly verify the logic, then the data of this part is retained; S104. Eliminate the data that does not conform to the normal rules, missing data, single-person attribute case data, and data without clear meaning, so as to ensure that the processed data can construct relational data.

9. A group mining system based on video image trajectory analysis, characterized in that, The system can be used to implement the group mining method based on video image trajectory analysis described in any one of claims 1 to 8 above. The system includes: A data acquisition module configured to acquire multi-dimensional attribute data sets of multiple targets of interest; A data processing module configured to define the targets of interest in the multi-dimensional attribute data set as nodes, connect two nodes that have participated in the same attribute case through edges, and the edges that are connected end to end form a path to obtain a graph; A first screening module configured to search the graph using a depth-first search algorithm to identify subgraphs, and there is a path connection between any two nodes within the subgraph; A second screening module configured to perform more detailed community division on each of the subgraphs using the Louvain algorithm to obtain a group graph, and there is an edge connection between any two nodes within the group graph.

10. An electronic device, characterized in that, Including: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the group mining method based on video image trajectory analysis described in any one of claims 1 to 8.