A method and system for identifying potential risk personnel based on graph-data fusion
Through graph-number fusion technology, combining view and business data, a relationship network diagram is built, impact index is calculated, and potential risk personnel are determined. The problem of failing to make full use of video image big data in the existing technology is solved, and more accurate and intuitive risk personnel identification is achieved.
Patent Information
- Application Number
- CN202510265371.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-07
AI Technical Summary
In the early warning or identification of risk personnel, the existing technology fails to fully combine the data collected by the video image big data platform, resulting in the analysis results being not intuitive enough and require manual review.
A method for identifying potential risk personnel based on graph-number fusion is proposed. By collecting view data and business data, a fusion information database is built, risk personnel and related personnel are extracted as nodes, a relationship network diagram is constructed, the impact index between nodes is calculated, and whether the candidate node is potential risk personnel is determined.
The multi-dimensional data fusion analysis of risk personnel has been achieved, the accuracy and intuitiveness of risk personnel identification have been improved, the need for manual review has been reduced, and the efficiency of risk warning has been improved.
Smart Images

Figure CN119782753B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for identifying potential risk personnel based on graph-data fusion, belonging to the technical field of risk personnel identification. Background Art
[0002] For the early warning or identification of risk personnel, currently, it is mainly through obtaining multi-dimensional data such as the call records, fund transactions, travel records, accommodation records, and Internet cafe access of risk personnel, and combining with the actual business logic to construct an analysis and early warning model, so as to apply it to business such as intelligence analysis related to risks, risk control, and personnel early warning. It is mainly based on the big data analysis application of business-related data.
[0003] Currently, there are various application systems and methods for analyzing criminal intelligence through big data analysis technologies such as machine learning and data mining, including risk transactions, personnel mobility, fund flows, etc. For example, a drug control intelligence research and judgment system based on big data disclosed in a Chinese invention patent with the patent number "CN201910917858.7", and a drug situation monitoring, early warning and evaluation method and system disclosed in the patent number "CN202210726928.2", etc.
[0004] The above solutions are mainly based on business-related data and do not combine the large amount of vivid view data collected by the video image big data platform. It is impossible to comprehensively grasp the latest status characteristics, life trajectories, contact personnel, etc. of the target object, and the analysis results are not intuitive enough. Often, manual review in combination with the video image big data platform is required. Summary of the Invention
[0005] In order to solve the problems existing in the above-mentioned prior art, the present invention proposes a method and system for identifying potential risk personnel based on graph-data fusion.
[0006] The technical solution of the present invention is as follows:
[0007] On the one hand, the present invention proposes a method for identifying potential risk personnel based on graph-data fusion, including the following steps:
[0008] Collect view data and business data of risk personnel with a criminal record and fuse them to form graph-data fusion data, and create a fusion information database based on the graph-data fusion data;
[0009] Extract risk personnel and associated personnel in the fusion information database as nodes to construct a relationship network diagram. At the same time, collect associated personnel information to supplement the nodes in the relationship network diagram, create an attribute list and a relationship list of the nodes according to the fusion information database and the associated personnel information, and define the node corresponding to the risk personnel as a risk node, the node corresponding to the associated personnel as a risk neighbor node, and other nodes as ordinary nodes;
[0010] Calculate the influence index between two adjacent nodes through the relationships of each node. Select any risk neighbor node or ordinary node as a candidate node. Taking the selected candidate node as the center, construct the influence relationship network of the candidate node. The nodes inserted into the influence relationship network of the candidate node are determined by the influence index and a set threshold.
[0011] Confirm whether the candidate node is a potential risk person through the influence relationship network of the candidate node.
[0012] As a preferred embodiment, the step of calculating the influence index between two adjacent nodes through the relationships of each node includes:
[0013] Select two adjacent nodes, and obtain the interaction behavior data of the two nodes through the attribute list and relationship list of the two nodes.
[0014] According to the behavior category, assign different behavior weights to the interaction behaviors. Based on the number of different interaction behaviors of the two nodes and the total number of interaction behaviors, combined with the frequency of risk vocabulary words in the interaction behavior data of the two nodes, calculate the influence index between the two adjacent nodes.
[0015] As a preferred embodiment, in the step of taking the selected candidate node as the center, constructing the influence relationship network of the candidate node, and determining the nodes inserted into the influence relationship network of the candidate node by the influence index and a set threshold:
[0016] According to the set first threshold, second threshold,..., k-th threshold, in the neighbor node set, two-hop neighbor node set,..., k-hop neighbor node set of the candidate node, select the neighbor nodes whose comprehensive influence index with the candidate node is greater than the corresponding threshold and insert them into the influence relationship network of the candidate node;
[0017] Among them, in the neighbor node set of the candidate node, the comprehensive influence index between the neighbor node and the candidate node is the influence index between the two nodes; in the k-hop neighbor node set, the comprehensive influence index between the k-hop neighbor node and the candidate node is calculated through the influence indexes between all nodes on the relationship path from the k-hop neighbor node to the candidate node.
[0018] As a preferred embodiment, the method for confirming whether the candidate node is a potential risk person through the influence relationship network of the candidate node is specifically:
[0019] Confirm whether the influence relationship network of each candidate node contains risk nodes, and exclude the candidate nodes that do not contain risk nodes;
[0020] For the candidate nodes whose influence relationship network contains risk nodes, if the candidate node and the risk node are multi-hop neighbor nodes, then calculate the indirect influence index between the candidate node and the risk node through the relationship path between the candidate node and the risk node.
[0021] Calculate the influence of candidate nodes in the influence relationship network through the influence relationship network of candidate nodes;
[0022] Combine the influence index or indirect influence index between the candidate node and the risk node, and the influence of the candidate node in the influence relationship network, calculate the risk factor of the corresponding candidate node, and compare the risk factor with the risk threshold to determine whether the corresponding candidate node is a potential risk person.
[0023] As a preferred embodiment, the view class data includes personnel image data and vehicle image data; the service class data includes identity data, call record data, fund data, and registration data.
[0024] As a preferred embodiment, the method further includes:
[0025] Preprocess the service class data, including data cleaning and data standardization;
[0026] Preprocess the view class data, including view data filtering, target extraction, and target matching.
[0027] As a preferred embodiment, in the steps of taking the risk persons and associated persons in the extraction and fusion information library as nodes, and collecting the information of associated persons to supplement the nodes in the relationship network diagram:
[0028] Determine the closely related persons as associated persons through the graph-data fusion data of risk persons;
[0029] After determining the associated persons, collect the graph-data fusion data of the associated persons and determine the closely related persons as supplementary nodes.
[0030] On the other hand, the present invention also proposes a potential risk person identification system based on graph-data fusion, including:
[0031] A graph-data fusion information library construction module, configured to collect view class data and service class data of risk persons with a criminal record and perform fusion to form graph-data fusion data, and create a fusion information library based on the graph-data fusion data;
[0032] A relationship network diagram construction module, configured to extract the risk persons and associated persons in the fusion information library as nodes, construct a relationship network diagram, and at the same time collect the information of associated persons to supplement the nodes in the relationship network diagram, create an attribute list and a relationship list of the nodes according to the fusion information library and the information of associated persons, and at the same time define the node corresponding to the risk person as a risk node, the node corresponding to the associated person as a risk neighbor node, and other nodes as ordinary nodes;
[0033] An influence network construction module is used to calculate the influence index between two adjacent nodes through the relationships of each node, select any risk neighbor node or ordinary node as a candidate node, and construct an influence relationship network of the candidate nodes with the selected candidate node as the center. The nodes inserted into the influence relationship network of the candidate nodes are determined by the influence index and a set threshold;
[0034] A potential risk personnel identification module is used to confirm whether the candidate node is a potential risk personnel through the influence relationship network of the candidate nodes.
[0035] In another aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for identifying potential risk personnel based on graph-data fusion as described in any embodiment of the present invention.
[0036] In another aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for identifying potential risk personnel based on graph-data fusion as described in any embodiment of the present invention.
[0037] The additional aspects and advantages of the present invention will be set forth in the following description, and some of them will be obvious from the description, or can be understood by practicing the present invention. Description of the Drawings
[0038] Figure 1 It is a schematic flowchart of the method according to Embodiment 1 of the present invention;
[0039] Figure 2 It is a schematic diagram of the relationship network diagram shown in the embodiment of the present invention;
[0040] Figure 3 It is a schematic diagram of the influence relationship network shown in the embodiment of the present invention. Detailed Embodiments
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0042] It should be understood that the step numbers used in the text are only for convenient description and do not limit the execution order of the steps.
[0043] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0044] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0045] The term "and / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0046] Embodiment 1:
[0047] See Figure 1 , this embodiment provides a method for identifying potential risk personnel based on graph-data fusion, which is applied to the public security big data system. The method specifically includes the following steps:
[0048] S100. Collect view-class data and business-class data of risk personnel with a criminal record as the basic information database, including multi-dimensional data information such as call data, fund data, virtual identity data, video image data, registration data, etc. Since view-class data and business-class data are applied in different systems, it is necessary to fuse these two types of data and apply them uniformly in one system. According to the public security big data standard specifications, access various business-class data and view-class data that need to be analyzed, and perform operations such as data processing, association, analysis, comparison, and fusion on the data to form graph-data fusion data, and create a fusion information database based on the graph-data fusion data.
[0049] S200. Extract the risk personnel and associated personnel in the fusion information database as nodes to construct a relationship network diagram. At the same time, collect the information of the associated personnel to supplement the nodes in the relationship network diagram, create an attribute list and a relationship list of the nodes according to the fusion information database and the information of the associated personnel. At the same time, define the node corresponding to the risk personnel as a risk node, the node corresponding to the associated personnel as a risk neighbor node, and other nodes as ordinary nodes (other nodes are the nodes supplemented after collecting the information of the associated personnel).
[0050] In this embodiment, by analyzing the graph-data fusion data of the risk personnel, those personnel who have close connections with the risk personnel are identified, and these personnel are determined as associated personnel. After successfully determining the associated personnel, continue to collect the graph-data fusion data of these associated personnel, and further determine other personnel who have close connections with the associated personnel based on these data, and these personnel are added to the relationship network diagram as supplementary nodes.
[0051] S300. Through the relationships between nodes, the influence index between any two adjacent nodes can be calculated; based on the influence index, any risk neighbor node or ordinary node can be selected as a candidate node, and with the selected candidate node as the center, an influence relationship network corresponding to the selected candidate node can be constructed. In this influence relationship network, some nodes will be inserted, and the inserted nodes can be risk nodes, risk neighbor nodes or other nodes. Whether a node is inserted into the inserted influence relationship network is determined according to the influence index between nodes and a set threshold.
[0052] S400. Confirm whether the candidate node is a potential risk person through the influence relationship network of the candidate node.
[0053] S500. If the candidate node is confirmed as a potential risk person, then upgrade the attribute of the candidate node from a risk neighbor node or other node to a potential risk node, and update the relationship network diagram. At the same time, re-join the graph number fusion data of this potential risk person into the fusion information database for subsequent analysis and monitoring.
[0054] The public security big data system applying this method will continuously monitor the activities of all risk nodes and potential risk nodes. If abnormal behaviors or new associated persons are found, the re-evaluation process will be automatically triggered to ensure the accuracy and timeliness of risk person identification.
[0055] In one embodiment, in step S300, the steps of calculating the influence index between two adjacent nodes through the relationships between nodes include:
[0056] S301. Select two adjacent nodes in the network diagram. These two nodes can be any two directly connected nodes. Once these two nodes are selected, the next step is to obtain the interaction behavior data between them by referring to the relationship list between the two nodes. The interaction behavior data includes the behavior list between the two nodes, such as fund transactions, social media interactions, social relationships (such as relatives, friends, classmates, colleagues), communication behaviors, peer behaviors, etc., and interaction characteristics, such as the number of transactions, transaction amount, social media interaction frequency, communication frequency, number of peer interactions, etc.
[0057] S302. Assign different behavior weights to the interaction behaviors according to the behavior categories. According to the number of different interaction behaviors between the two nodes and the total number of interaction behaviors, combined with the frequency of risk words in the interaction behavior data between the two nodes, calculate the influence index between the two adjacent nodes.
[0058] In one embodiment, the specific calculation steps of the influence index are as follows:
[0059] First, calculate the interaction behavior weights of the two nodes:
[0060] ;
[0061] Among them, is the interaction behavior weight of two nodes, represents the behavior category, is the total number of behavior categories, represents the th behavior weight of the behavior category, represents the th number of times of the behavior category, represents the total number of interaction behaviors.
[0062] Then, according to the interaction behavior weights of two nodes and the risk vocabulary word frequency, combined with the TF-IDF mining technology, calculate the influence index between adjacent two nodes:
[0063] ;
[0064] Among them, represents the influence index between node and node , represents the interaction behavior weight between node and node , represents the risk vocabulary word frequency that appears in the th interaction behavior data, represents the inverse document frequency index of the risk vocabulary.
[0065] In one embodiment, the behavior categories can be directly distinguished according to specific categories. For example, fund transactions, social media interactions, communication behaviors, peer behaviors, etc. are respectively regarded as one behavior category, and different behavior weights are assigned to different behavior categories. For example, the behavior weight of fund transactions is 0.3, the behavior weight of social media interactions is 0.05, the behavior weight of communication behaviors is 0.1, the behavior weight of peer behaviors is 0.15, etc.
[0066] In another embodiment, the behavior categories can also be divided according to the risk level. For example, they are divided into serious risk behavior categories, general risk behavior categories, and ordinary behavior categories. Public security personnel can determine which risk level behavior category each interaction behavior data belongs to according to case handling experience or logical reasoning. For example: classify fund transactions into serious risk behavior categories, classify communication behaviors and peer behaviors into general risk behavior categories, and classify social media interactions into ordinary behavior categories. Different behavior weights are assigned to different behavior categories. For example, the behavior weight of serious risk behavior categories is 0.5, the behavior weight of general risk behavior categories is 0.35, and the behavior weight of ordinary behavior categories is 0.15.
[0067] In one embodiment, in the step of constructing a candidate node influence relationship network centered on the selected candidate node in step S300, in the step of determining the nodes inserted into the candidate node influence relationship network through the influence index and a set threshold:
[0068] According to the set first threshold, second threshold, …, k-th threshold, in the neighbor node set, two-hop neighbor node set, …, k-hop neighbor node set of the candidate node, select the neighbor nodes whose comprehensive influence index with the candidate node is greater than the corresponding threshold and insert them into the candidate node influence relationship network;
[0069] Among them, in the neighbor node set of the candidate node, the comprehensive influence index between the neighbor node and the candidate node is the influence index between the two nodes; in the k-hop neighbor node set, the comprehensive influence index between the k-hop neighbor node and the candidate node is calculated through the influence indexes between all the nodes on the relationship path from the k-hop neighbor node to the candidate node.
[0070] For specific reference Figure 2 , Figure 2 in, the selected candidate node is Zhang San, the neighbor node set of Zhang San includes [Li Si, Li Wu, Li Liu, Li Qi, Li Ba], the two-hop neighbor node set of Zhang San includes [Zhao San, Zhao Si, Zhao Wu], and the three-hop neighbor node set of Zhang San includes [Chen San].
[0071] In this embodiment, the first threshold, second threshold, and third threshold are respectively set to 0.7, 0.65, and 0.5. In the neighbor node set of Zhang San, the influence indexes between Li Wu, Li Qi, and Li Ba and Zhang San are greater than 0.7, so Li Wu, Li Qi, and Li Ba are inserted into Zhang San's influence relationship network, and the three nodes are still neighbor nodes, removing Li Si and Li Liu.
[0072] In one embodiment, the comprehensive influence index between the k-hop neighbor node and the candidate node is calculated by the average value of the influence indexes between all the nodes on the relationship path. For example, among the two-hop neighbor nodes of Zhang San, the relationship path between Zhao San and Zhang San is Zhao San - Li Qi - Zhang San, and the average value of the influence indexes of this relationship path is (0.79 + 0.78) / 2 = 0.785, which is greater than the second threshold, so Zhao San is inserted into Zhang San's influence relationship network, and the relationship path between Zhao San and Zhang San remains the same. Among the three-hop neighbor nodes of Zhang San, the relationship path between Chen San and Zhang San is Chen San - Zhao San - Li Qi - Zhang San, and the average value of the influence indexes of this relationship path is (0.79 + 0.78 + 0.95) / 3 = 0.84, which is greater than the third threshold, so Chen San is inserted into Zhang San's influence relationship network, and the relationship path between Chen San and Zhang San remains the same. The influence relationship network of Zhang San formed according to this embodiment is as Figure 3 shown.
[0073] In another embodiment, the comprehensive influence index between the k-hop neighbor node and the candidate node is calculated by weighted summation of the influence indexes between all nodes on the relationship path. This calculation method takes into account the contribution degree of each node on the path to the comprehensive influence index, where the influence index of each node is assigned different weights according to its position and importance in the path. The determination of these weights can be based on the expert's experience judgment or can be learned and optimized through a machine learning model. In practical applications, the setting of weights is a key step, which directly affects the accuracy of the final comprehensive influence index. Although there are various methods for setting weights, the specific implementation details thereof will not be elaborated in detail in this embodiment.
[0074] In one embodiment, in step S400, the method for confirming whether the candidate node is a potential risk person through the influence relationship network of the candidate node is specifically as follows:
[0075] S401. Confirm whether the influence relationship network of each candidate node contains a risk node, and exclude the candidate nodes that do not contain a risk node.
[0076] S402. For the candidate nodes whose influence relationship network contains a risk node, if the candidate node and the risk node are multi-hop neighbor nodes, calculate the indirect influence index between the candidate node and the risk node through the relationship path between them; in this embodiment, the calculation formula of the indirect influence index is:
[0077] ;
[0078] Among them, represents the indirect influence index between the candidate node and the risk node , represents the influence index between the candidate node and the 1-hop neighbor node of the candidate node on the relationship path, represents the average influence index between the candidate node and its neighbor nodes in the influence relationship network, represents the influence index between the 1-hop neighbor node and the 2-hop neighbor node of the candidate node on the relationship path, represents the average influence index between the 1-hop neighbor node of the candidate node and its neighbor nodes in the influence relationship network; represents the influence index between the -hop neighbor node of the candidate node and the risk node , represents the -hop neighbor node of the candidate node The average influence index between a hopping neighbor node and its neighbor nodes;
[0079] S403. Calculate the influence of the candidate node in the influence relationship network through the influence relationship network of the candidate node; specifically, obtain all interaction behavior data in the influence relationship network, determine the total number of interaction behaviors, the number of active interaction behaviors initiated by the candidate node, and the number of passive behaviors received by the candidate node passively, and calculate the influence of the candidate node in the influence relationship network according to the following formula:
[0080] ;
[0081] Where, is the influence of the candidate node in the influence relationship network, is the preset behavior weight, is the candidate node the number of active interaction behaviors initiated actively, is the candidate node the number of passive behaviors received passively, is the total number of interaction behaviors;
[0082] Based on the above solution, this embodiment comprehensively considers the active interaction behavior and passive interaction behavior of the candidate node, and balances the importance of the two through the preset behavior weight . If has a large value, it indicates that the system focuses more on the active behavior of the candidate node; conversely, if has a small value, it values more the passive behavior of the candidate node. Through such a calculation method, the actual influence of the candidate node in the influence relationship network can be evaluated more comprehensively.
[0083] S404. Combine the influence index or indirect influence index between the candidate node and the risk node, and the influence of the candidate node in the influence relationship network, calculate the risk factor of the corresponding candidate node, and compare the risk factor with the risk threshold to determine whether the corresponding candidate node is a potential risk person;
[0084] The calculation formula of the risk factor is:
[0085] .
[0086] Where, represents the risk factor of the candidate node ; is the dynamically adjusted weight, adjusted according to the number of risk nodes in the influence relationship network of the candidate node , represents the total number of risk nodes in the influence relationship network of the candidate node .
[0087] After calculating the risk factors of the candidate nodes, the risk factors will be compared with a preset risk threshold. If the risk factor is greater than or equal to the risk threshold, the candidate node is determined to be a potential risk person; otherwise, it is determined to be a non-potential risk person. In this embodiment, the weight value is dynamically adjusted and is positively correlated with the number of risk nodes in the influence relationship network of the candidate node , because the more the number of risk nodes in the influence relationship network of the candidate node , the greater the probability that the candidate node can come into contact with or be affected by the risk nodes.
[0088] In one embodiment, the view class data includes personnel image data and vehicle image data; the business class data includes identity data, call record data, fund data, and registration data.
[0089] This embodiment also includes preprocessing the business class data, specifically:
[0090] For the business class data, the data mainly consists of structured data such as form data and transaction data. The processing of this type of data includes data cleaning, filtering, conversion, etc.
[0091] For example, for the fund transaction data of a bank, the data processing process is as follows:
[0092] 1) Key data acquisition: Obtain the fund transaction data from the database, check the key data such as the fund transaction time, transaction amount, transaction account, counterparty account, transaction type, and transaction location, and obtain complete resources.
[0093] 2) Data cleaning:
[0094] Data deduplication: Check and delete duplicate transaction records to avoid interfering with the analysis results.
[0095] Format conversion: Ensure that the formats of all date and time fields are consistent, such as unified to the format of "YYYY-MM-DD HH:MM:SS". For the amount field, ensure that the number of digits after the decimal point is consistent.
[0096] Remove abnormal characters: In text fields (such as transaction remarks), remove any illegal characters or garbled codes that may exist.
[0097] 3) Data conversion:
[0098] Unify the amount unit: If there are amounts in different currency units in the data, they need to be unified and converted to the same currency unit (such as RMB).
[0099] Timestamp conversion: Convert timestamps into a format that is convenient for analysis, such as extracting information like the hour when a transaction occurred, the day of the week, etc., for subsequent analysis of time-related patterns.
[0100] Data screening: Screen transfer transactions: Filter out records of the "transfer" type from all transaction types.
[0101] Filter small-value transactions: Based on business experience, small-value transactions are usually less likely to be abnormal transactions. Filter data with a transaction amount below 500 yuan.
[0102] Preprocess view-class data, specifically:
[0103] The data processing process is as follows:
[0104] 1) Data collection:
[0105] Access various types of target image data of personnel, vehicles, etc. collected by the road surface area capture cameras.
[0106] 2) Image modeling:
[0107] Use the view analysis algorithm deployed locally by the public security department to extract and model target features, such as the facial contours, eyes, mouths, noses of personnel, and license plates, vehicle brands, license plate numbers, license plate colors, vehicle colors, etc. of vehicles.
[0108] 3) Data filtering:
[0109] According to the data results after modeling analysis, filter out invalid data, such as pictures without targets, blurred pictures, pictures without license plates, pictures with incomplete target captures, etc., to ensure the validity of the data.
[0110] 4) Model matching:
[0111] Match the captured picture data after modeling analysis with the local personnel list library that has been collected. If the model similarity between the captured data and the local data reaches the set threshold, add a label of the identity information of the local personnel to this piece of data. If multiple pieces of data match, use the identity information of the list with the higher similarity as the standard.
[0112] 5) Data archiving:
[0113] The data after modeling processing is stored and archived in the locally deployed distributed storage system for storage and provides efficient data retrieval and query.
[0114] After completing data preprocessing, data association is also required. For business data, data is associated based on the key fields of various types of business data. For example, for fund flow data, the transfer account is used as the key field for data association to form a penetrable fund data chain; for call record data, the communication number is used as the key field for data association to form call association relationship data; for accommodation data, the registered identity information of the accommodation personnel is used as the key field to form personnel accommodation record association data.
[0115] For view data, it is mainly associated based on the identity information of personnel and the license plate information of vehicles. For example, for personnel picture data, the identity information tag added by matching the local list is used as the personnel identity ID for association, and the personnel pictures are aggregated according to the identity ID to form personnel profile data with individuals as units; for vehicle picture data, the analyzed license plate number is used as the ID for association to form vehicle profile data with license plate numbers as units.
[0116] After preprocessing and data association of business data and view data, graph-data fusion is performed. Using the identity ID of personnel as the key field, the business data and view data are fused based on the identity ID to form a graph-data fusion personnel thematic database, including various call record data, accommodation data, fund data, Internet access data, capture data, vehicle data, etc. of personnel, and comprehensive data queries can be performed through one identity ID.
[0117] In one embodiment, the method for further determining the personnel who have close contact with risk personnel or associated personnel based on graph-data fusion data is as follows:
[0118] The real-name identity information, associated fellow traveler information, travel location information, travel time information, and fellow traveler capture data information of risk personnel or associated personnel can be input, and analysis conditions such as the number of fellow travels, the number of days of fellow travels, the number of fellow travelers, and the travel location can be set to output the identity information of the fellow travelers of risk personnel or associated personnel who meet the analysis conditions.
[0119] The mobile communication data of risk personnel or associated personnel can be input, including call time, call duration, call recipient number, affiliated operator, and call location area, and analysis conditions such as call duration, call days, call times, and call location can be set to judge the intimacy of communication with risk personnel or associated personnel, and output the information of the close communicators with risk personnel or associated personnel who meet the conditions.
[0120] It is possible to input the train ticket booking data, air ticket booking data, and online car-hailing travel data of risk personnel or associated personnel, including travel time, travel type, departure location, arrival location, co-travelers on the same train, and co-travelers on the same plane. Set analysis conditions such as travel time period, travel location, and co-travelers, and output the information of co-travelers with risk personnel or associated personnel who meet the conditions.
[0121] It is possible to input the vehicle travel data of risk personnel or associated personnel, including vehicle capture records, license plate data, vehicle co-rider data, and vehicle travel time. Set analysis conditions such as vehicle travel time, vehicle travel location, vehicle travel route, number of vehicle co-rides, and number of days of vehicle co-rides, and output the information of co-travelers in the same vehicle with risk personnel or associated personnel who meet the conditions.
[0122] It is possible to input the accommodation data of risk personnel or associated personnel, including accommodation time, accommodation hotel, accommodation floor, accommodation room, and co-accommodation personnel. Set analysis conditions such as accommodation in the same hotel, accommodation in the same room, number of accommodation times, number of accommodation days, and accommodation location, and output the information of co-accommodation personnel in the same hotel or the same room with risk personnel or associated personnel who meet the conditions.
[0123] Based on the above implementation solutions, this embodiment can screen out multi-dimensional associated information of risk personnel from a large amount of business data and view data, and through the methods of constructing relationship graphs and influence graphs, use multi-dimensional data to warn of the associated personnel of risk personnel, discover potential risk personnel, and can analyze potential risk personnel more accurately from more dimensions.
[0124] Embodiment 2:
[0125] This embodiment provides a potential risk personnel identification system based on graph-data fusion, which is characterized by including:
[0126] A graph-data fusion information library construction module, which is used to collect view data and business data of risk personnel with a criminal record and fuse them to form graph-data fusion data, and create a fusion information library based on the graph-data fusion data; this module is used to implement the function of step S100 in Embodiment 1, and will not be elaborated here;
[0127] A relationship network graph construction module, which is used to extract risk personnel and associated personnel in the fusion information library as nodes to construct a relationship network graph, and at the same time collect associated personnel information to supplement the nodes in the relationship network graph, create an attribute list and a relationship list of the nodes according to the fusion information library and the associated personnel information, and at the same time define the node corresponding to the risk personnel as a risk node, the node corresponding to the associated personnel as a risk neighbor node, and other nodes as ordinary nodes; this module is used to implement the function of step S200 in Embodiment 1, and will not be elaborated here;
[0128] The influence network construction module is used to calculate the influence index between two adjacent nodes through the relationships of each node, select any risk neighbor node or ordinary node as a candidate node, and construct an influence relationship network of the candidate nodes with the selected candidate node as the center. The nodes inserted into the influence relationship network of the candidate nodes are determined by the influence index and a set threshold; this module is used to implement the function of step S300 in the first embodiment and will not be elaborated here;
[0129] The potential risk personnel identification module is used to confirm whether the candidate node is a potential risk personnel through the influence relationship network of the candidate nodes; this module is used to implement the function of step S400 in the first embodiment and will not be elaborated here.
[0130] Embodiment 3:
[0131] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the potential risk personnel identification method based on graph-data fusion as described in any embodiment of the present invention.
[0132] Embodiment 4:
[0133] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the potential risk personnel identification method based on graph-data fusion as described in any embodiment of the present invention.
[0134] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent the situation where A exists alone, A and B exist simultaneously, or B exists alone. Where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c may represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c may be single or multiple.
[0135] Those of ordinary skill in the art can realize that the units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0136] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0137] In several embodiments provided in the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (hereinafter referred to as ROM), random access memory (hereinafter referred to as RAM), magnetic disks, or optical discs that can store program codes.
[0138] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for identifying potential risk personnel based on image-data fusion, characterized in that: The following steps are involved: Collect the visual data and business data of risk personnel with criminal records and fuse them into graph-number fusion data, and create a fusion information database based on the graph-number fusion data; Extract risk personnel and related personnel in the fusion information database as nodes, build a relationship network diagram, and collect related personnel information to supplement the nodes in the relationship network diagram. Create a node attribute list and relationship list based on the fusion information database and related personnel information, and define the node corresponding to the risk personnel as a risk node, the node corresponding to the related personnel as a risk neighbor node, and other nodes as ordinary nodes; Through the relationship between each node, the influence index between two adjacent nodes is calculated, and any risk neighbor node or ordinary node is selected as a candidate node. With the selected candidate node as the center, an influence relationship network of the candidate node is constructed. The nodes inserted into the influence relationship network of the candidate node are determined by the influence index and the set threshold; Confirm whether the candidate node is a potential risk person through the influence relationship network of the candidate node; The step of calculating the influence index between two adjacent nodes according to the relationship between the nodes includes: Select two adjacent nodes, and obtain the interaction behavior data of the two nodes through the attribute lists and relationship lists of the two nodes; Different behavior weights are assigned to interaction behaviors according to behavior categories. The impact index between two adjacent nodes is calculated based on the number of different interaction behaviors and the total number of interaction behaviors between the two nodes, combined with the frequency of risk words in the interaction behavior data of the two nodes. Among them, in the step of constructing a candidate node influence relationship network with the selected candidate node as the center, and the node inserted in the candidate node influence relationship network is determined by the influence index and the set threshold: According to the set first threshold, second threshold, ..., kth threshold, from the neighbor node set, two-hop neighbor node set, ..., k-hop neighbor node set of the candidate node, select the neighbor node whose comprehensive influence index with the candidate node is greater than the corresponding threshold and insert it into the candidate node influence relationship network; Among them, in the neighbor node set of the candidate node, the comprehensive influence index between the neighbor node and the candidate node is the influence index between the two nodes; in the k-hop neighbor node set, the comprehensive influence index between the k-hop neighbor node and the candidate node is calculated by the influence index between all nodes on the relationship path from the k-hop neighbor node to the candidate node.
2. According to claim 1, a method for identifying potential risk personnel based on image-data fusion is characterized in that: The method for confirming whether a candidate node is a potential risk person through the influence relationship network of the candidate node is specifically as follows: Confirm whether the influence relationship network of each candidate node contains risk nodes, and exclude candidate nodes that do not contain risk nodes; For candidate nodes whose influence relationship network includes risk nodes, if the candidate node and the risk node are multi-hop neighbor nodes, the indirect influence index between the candidate node and the risk node is calculated through the relationship path between the candidate node and the risk node; And calculate the influence of the candidate node in the influence relationship network through the influence relationship network of the candidate node; Combined with the influence index or indirect influence index between the candidate node and the risk node, as well as the influence of the candidate node in the influence relationship network, the risk factor of the corresponding candidate node is calculated, and the risk factor and risk threshold are compared to determine whether the corresponding candidate node is a potential risk person.
3. The method for identifying potential risk personnel based on image-data fusion according to claim 1 is characterized by: The view data includes personnel image data and vehicle image data; the business data includes identity data, call record data, capital data and registration data.
4. The method for identifying potential risk personnel based on image-data fusion according to claim 3 is characterized in that: The method further comprises: Pre-process business data, including data cleaning and data standardization; Preprocess view data, including view data filtering, target extraction, and target matching.
5. The method for identifying potential risk personnel based on image-data fusion according to claim 1 is characterized in that: In the step of extracting risk personnel and related personnel in the fusion information database as nodes, and collecting related personnel information to supplement the nodes in the relationship network diagram: By fusing the graph and number data of risk personnel, we can identify the close contact personnel as related personnel; After determining the related persons, collect the graph fusion data of the related persons and determine the closely connected persons as supplementary nodes.
6. A potential risk personnel identification system based on image and data fusion, characterized in that: include: A graph-data fusion information base construction module is used to collect the view data and business data of risk personnel with criminal records and fuse them to form graph-data fusion data, and create a fusion information base based on the graph-data fusion data; The relationship network diagram construction module is used to extract risk personnel and related personnel in the fusion information database as nodes, construct a relationship network diagram, and collect related personnel information to supplement the nodes in the relationship network diagram. According to the fusion information database and related personnel information, the attribute list and relationship list of the nodes are created, and the nodes corresponding to risk personnel are defined as risk nodes, the nodes corresponding to related personnel are defined as risk neighbor nodes, and other nodes are defined as ordinary nodes. The influence network construction module is used to calculate the influence index between two adjacent nodes through the relationship between each node, select any risk neighbor node or ordinary node as a candidate node, and build an influence relationship network of the candidate node with the selected candidate node as the center. The node inserted in the influence relationship network of the candidate node is determined by the influence index and the set threshold; A potential risk personnel identification module is used to confirm whether a candidate node is a potential risk person through the influence relationship network of the candidate node; The step of calculating the influence index between two adjacent nodes according to the relationship between the nodes includes: Select two adjacent nodes, and obtain the interaction behavior data of the two nodes through the attribute lists and relationship lists of the two nodes; Different behavior weights are assigned to interaction behaviors according to behavior categories. The impact index between two adjacent nodes is calculated based on the number of different interaction behaviors and the total number of interaction behaviors between the two nodes, combined with the frequency of risk words in the interaction behavior data of the two nodes. Among them, in the step of constructing a candidate node influence relationship network with the selected candidate node as the center, and the node inserted in the candidate node influence relationship network is determined by the influence index and the set threshold: According to the set first threshold, second threshold, ..., kth threshold, from the neighbor node set, two-hop neighbor node set, ..., k-hop neighbor node set of the candidate node, select the neighbor node whose comprehensive influence index with the candidate node is greater than the corresponding threshold and insert it into the candidate node influence relationship network; Among them, in the neighbor node set of the candidate node, the comprehensive influence index between the neighbor node and the candidate node is the influence index between the two nodes; in the k-hop neighbor node set, the comprehensive influence index between the k-hop neighbor node and the candidate node is calculated by the influence index between all nodes on the relationship path from the k-hop neighbor node to the candidate node.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for identifying persons with potential risks based on image and data fusion as described in any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for identifying persons with potential risks based on image-data fusion as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Drug-forbidden intelligence research and judgment system based on big data
CN110674238A
Poison condition monitoring, early warning and evaluating method and system
CN115049284A
Method and system for calculating user influence in social network
CN101770487A
Method for selecting the initial node with maximum influence in online social network
CN106355506A