A method and system for analyzing personnel relationships based on multi-source heterogeneous data

CN122547862APending Publication Date: 2026-08-11GUIZHOU YUNTENGZHIYUAN TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]然而,上述方法存在以下缺陷:现有技术仅关注待筛查人员与每一个目标人员之间的独立关联强度,而未考虑该待筛查人员同时与多个目标人员存在关联这一整体性特征

Benefits of technology

[0015] The beneficial effects of this invention include: First, existing technologies only output independent correlation scores between the person to be screened and the target person to determine their relationship, without considering the overall characteristic that the person to be screened is simultaneously associated with multiple target persons. Based on this, this application statistically analyzes the number of effectively associated target persons and introduces a connection breadth gain coefficient that monotonically increases with this number, enabling individuals with extensive connections to multiple target persons to receive non-linear amplification in overall correlation. This prioritizes organizers, liaisons, and other similar individuals within a group in the scoring ranking, solving the problem of systematic underestimation of such high-risk individuals in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122547862A_ABST
    Figure CN122547862A_ABST
Patent Text Reader

Abstract

This invention relates to a method and system for analyzing personnel relationships based on multi-source heterogeneous data, belonging to the field of data processing technology. The method includes: acquiring multi-source heterogeneous association data for each target person in a set of personnel to be screened and a set of target personnel; calculating a comprehensive association score between the personnel to be screened and each target person based on the multi-source heterogeneous association data; counting the number of target persons with effective associations to the personnel to be screened based on multiple comprehensive association scores; determining a connection breadth gain coefficient based on the number of target persons; wherein the connection breadth gain coefficient monotonically increases with the number of target persons; and determining the overall association degree between the personnel to be screened and the set of target persons based on multiple comprehensive association scores and the connection breadth gain coefficient. This method enables individuals with extensive connections to multiple target persons to achieve non-linear amplification in the overall association degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for analyzing personnel relationships based on multi-source heterogeneous data. Background Technology

[0002] In the field of personnel relationship analysis, existing technologies typically collect individual data such as call records and travel records to calculate the strength of association between the person to be screened and known target persons (such as call frequency, co-occurrence frequency, etc.), and use this as the basis for determining whether they are related persons. For example, statistical methods are used to calculate the association score of personnel pairs, or machine learning models are used to output the association probability, and then the scores are sorted and truncated.

[0003] However, the above methods have the following drawbacks: existing technologies only focus on the individual association strength between the person to be screened and each target person, without considering the overall characteristic that the person to be screened may have associations with multiple target persons simultaneously. In real-world business scenarios, an individual may have moderate-strength connections with multiple target persons (e.g., having phone calls or cohabitation records with multiple members of a group), and their group embedding degree is much higher than that of an individual with a high-strength connection only with a single target person. Because the scoring system of existing technologies lacks quantification of the breadth of connections, the overall association degree of the former type of person is systematically underestimated, making it difficult to highlight their actual risk level in the ranking process, thus affecting the accuracy of personnel analysis and the effectiveness of business decisions. Summary of the Invention

[0004] To address the aforementioned problems in the prior art, this invention provides a method and system for analyzing personnel relationships based on multi-source heterogeneous data.

[0005] In a first aspect, this application provides a method for analyzing personnel relationships based on multi-source heterogeneous data, comprising: acquiring multi-source heterogeneous association data for each target person in a set of persons to be screened and a set of target persons; wherein the data types of the multi-source heterogeneous association data include at least two of call data, accommodation data, flight data, and train data; calculating a comprehensive association score between the persons to be screened and each target person based on the multi-source heterogeneous association data; counting the number of target persons with effective associations with the persons to be screened based on multiple comprehensive association scores; determining a connection breadth gain coefficient based on the number of target persons; wherein the connection breadth gain coefficient monotonically increases with the number of target persons; and determining the overall association degree between the persons to be screened and the set of target persons based on multiple comprehensive association scores and the connection breadth gain coefficient.

[0006] Optionally, the comprehensive correlation score between the person to be screened and any target person is determined through the following steps: dividing the multi-source heterogeneous correlation data into multiple correlation dimensions according to data type; calculating the time difference in days between the occurrence time of each correlation record and the current time for each correlation record in each correlation dimension; assigning a time decay weight to each correlation record based on the time decay weight; wherein, the correlation record closer to the current time has a higher weight; within the same correlation dimension, accumulating all correlation records between the person to be screened and the target person by combining the assigned time decay weights to obtain the equivalent interaction quantity of the correlation dimension, and mapping the equivalent interaction quantity to a unified scoring interval through logarithmic normalization to obtain the single-dimensional score of the correlation dimension; and weighting and fusing the single-dimensional scores of each correlation dimension to obtain the comprehensive correlation score between the person to be screened and the target person.

[0007] Optionally, the time decay weight is determined by the following steps: processing the difference in the number of days of validity using an exponential decay function, so that the weight decreases exponentially over time according to a preset half-life; processing the difference in the number of days of validity using a linear zeroing function, so that the weight of associated records exceeding a preset deadline is zeroed; and multiplying the value of the exponential decay function by the value of the linear zeroing function to determine the time decay weight.

[0008] Optionally, the step of weightedly fusing the single-dimensional scores of each correlation dimension to obtain the comprehensive correlation score between the person to be screened and the target person includes: when there is only one effective correlation dimension, using the single-dimensional score of that correlation dimension as the comprehensive correlation score; when there are two or more effective dimensions, weighting the single-dimensional scores of each correlation dimension according to a preset dimension weight, and then multiplying by a multi-path evidence enhancement coefficient to obtain the comprehensive correlation score; wherein, the multi-path evidence enhancement coefficient increases with the number of effective correlation dimensions.

[0009] Optionally, the step of mapping the equivalent interaction quantity to a unified scoring interval through logarithmic normalization to obtain a single-dimensional score for the correlation dimension includes: mapping the equivalent interaction quantity to a unified scoring interval through logarithmic normalization to obtain an initial score for the correlation dimension; determining a phased enhancement coefficient based on the mapping relationship between the equivalent interaction quantity and a preset tier interval; and obtaining a single-dimensional score for the correlation dimension based on the phased enhancement coefficient and the initial score for the correlation dimension.

[0010] Optionally, the method further includes: retrieving judgment dimensions; wherein the judgment dimensions include multiple association dimensions that divide the multi-source heterogeneous association data according to data type and a target personnel quantity dimension; setting a threshold parameter for each judgment dimension; determining the number of judgment dimensions triggered by the person to be screened through the threshold parameter set for each judgment dimension; when the number of judgment dimensions triggered by the person to be screened reaches a preset minimum number of triggered dimensions, marking the person to be screened as a hidden associated person; and / or, if there is a co-occurrence event between the person to be screened and any target personnel, then marking the person to be screened as a hidden associated person; wherein the co-occurrence event includes any one of the following: having a record of sharing a room, having a record of traveling in the same carriage, or having a record of being in the same flight cabin.

[0011] Optionally, the preset minimum number of trigger dimensions is a first value, or the preset minimum number of trigger dimensions is a second value; the first value is 1; the second value is greater than or equal to 2; by setting the first value or the second value, the screening mode for hidden associated persons can be distinguished.

[0012] Optionally, the method further includes: extracting the original records of the hidden associated personnel and each target personnel in each association dimension; wherein the multi-source heterogeneous association data is divided into multiple association dimensions according to data type; calculating the equivalent interaction volume of each association dimension; calculating the recent activity proportion index; wherein the recent activity proportion index represents the contribution ratio of records within 30 days to the equivalent interaction volume; and outputting the recent activity proportion index.

[0013] Optionally, the method further includes: using all hidden related personnel and all target personnel in the target personnel set as nodes; the attributes of the nodes include at least personnel identifier, personnel category label, and overall correlation degree; using the relationship between each pair of personnel as a directed weighted edge; the attributes of the directed weighted edge include at least one of the following: set of relationship dimension types, comprehensive correlation score, equivalent interaction volume, and recent activity ratio index; and outputting the structured data of the nodes and the directed weighted edges to form a personnel relationship network graph.

[0014] Secondly, this application provides a personnel relationship analysis system based on multi-source heterogeneous data, comprising: an acquisition module for acquiring multi-source heterogeneous association data of each target person in the set of personnel to be screened and the set of target personnel; wherein the data types of the multi-source heterogeneous association data include at least two of call data, accommodation data, flight data, and train data; a calculation module for calculating a comprehensive association score between the personnel to be screened and each target person based on the multi-source heterogeneous association data; a statistics module for counting the number of target persons with effective associations with the personnel to be screened based on multiple comprehensive association scores; a determination module for determining a connection breadth gain coefficient based on the number of target persons; wherein the connection breadth gain coefficient monotonically increases with the number of target persons; and an analysis module for determining the overall association degree between the personnel to be screened and the set of target personnel based on multiple comprehensive association scores and the connection breadth gain coefficient.

[0015] The beneficial effects of this invention include: First, existing technologies only output independent correlation scores between the person to be screened and the target person to determine their relationship, without considering the overall characteristic that the person to be screened is simultaneously associated with multiple target persons. Based on this, this application statistically analyzes the number of effectively associated target persons and introduces a connection breadth gain coefficient that monotonically increases with this number, enabling individuals with extensive connections to multiple target persons to receive non-linear amplification in overall correlation. This prioritizes organizers, liaisons, and other similar individuals within a group in the scoring ranking, solving the problem of systematic underestimation of such high-risk individuals in existing technologies.

[0016] Second, existing technologies that rely solely on a single data source (such as call records) are prone to missing valid association determinations due to parties deliberately avoiding that data source, resulting in a significantly underestimated number of target individuals and rendering the connection breadth gain coefficient ineffective. Therefore, this application employs at least two types of heterogeneous data, allowing association records from different sources to independently trigger valid association determinations. For example, a person to be screened may have no call records with target A, but co-occurrence records of accommodation; call records with target B, but no accommodation records; and a flight co-travel record with target C. In a single data source approach, these associations may be partially or completely omitted, resulting in a significantly smaller number of target individuals with valid associations compared to the actual number. In this approach, by integrating multi-source heterogeneous data, it is possible to fully identify valid associations between the person to be screened and multiple target individuals across different dimensions, thereby accurately calculating the number of target individuals and ensuring that the overall association degree truly reflects the degree of group embedding. Attached Figure Description

[0017] Figure 1A flowchart illustrating the steps of a personnel relationship analysis method based on multi-source heterogeneous data provided in this embodiment of the invention; Figure 2 A flowchart illustrating the steps of another method for analyzing personnel relationships based on multi-source heterogeneous data provided in this embodiment of the invention; Figure 3 This is a block diagram of a personnel relationship analysis system based on multi-source heterogeneous data provided in an embodiment of the present invention. Figure 4 This is a module block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0018] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application can also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0019] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0020] Research has revealed that existing technologies only focus on the individual association strength between the person being screened and each target person, neglecting the overall characteristic of the person being screened simultaneously having associations with multiple target persons. In real-world business scenarios, an individual may have moderately strong connections with multiple target persons (e.g., having phone calls or cohabitation records with multiple members of a group), resulting in a much higher level of group embedding than an individual with only a high-strength connection to a single target person. Because the scoring system of existing technologies lacks quantification of the breadth of connections, the overall association strength of the former type of individual is systematically underestimated, making it difficult to highlight their actual risk level in ranking, thus affecting the accuracy of personnel analysis and the effectiveness of business decisions.

[0021] In view of the above problems, this application proposes the following embodiments to solve the above technical problems.

[0022] Please see Figure 1 This application provides a method for analyzing personnel relationships based on multi-source heterogeneous data, including steps 101 to 105.

[0023] Step 101: Obtain multi-source heterogeneous association data for each target person in the set of people to be screened and the set of target people.

[0024] Among them, the data types of multi-source heterogeneous correlation data include at least two of the following: call data, accommodation data, flight data, and train data.

[0025] First, obtain the information of the people to be screened. Gather with the target personnel Multi-source heterogeneous correlation data for each target individual; among which, Represents the target personnel set The m-th target person in the list.

[0026] It should be noted that in this application, the resident identity card number is used as the unique natural person index to construct a cross-data source personnel identification mapping table, ensuring that records of the same natural person (including those to be screened and those targeted) in various data sources such as calls, accommodations, flights, and trains can be accurately merged. Based on this, composite data such as flight numbers and dates, train numbers, carriage numbers and dates, and hotel names and check-in date ranges are extracted to identify related events of personnel at specific times and locations.

[0027] After identity merging is completed, associated data is constructed based on the business characteristics of each data source to identify associated events between two individuals at specific times and locations.

[0028] For example, for flight data, the concatenated string of flight number and departure date is used as the association key, and two people with the same association key are considered to be traveling on the same flight.

[0029] For train data, the concatenated string of train number, departure date and carriage number is used as the association key. If the train number, carriage number and departure date are the same, it is considered a valid ride-sharing, or if the train number and departure date are the same, it is considered a valid ride-sharing.

[0030] For accommodation data, it can be determined that two conditions must be met simultaneously: identical hotel identifiers and overlapping check-in times. Specifically, if person A (the person to be screened) has an accommodation period of... The accommodation area for Personnel B (target personnel) is as follows: Then when and When both are established, they are determined to be stays during the same period, thus avoiding misjudging non-overlapping check-in records as related based solely on the same hotel name.

[0031] Step 102: Based on multi-source heterogeneous association data, calculate the comprehensive association score between the person to be screened and each target person.

[0032] Based on the multi-source heterogeneous association data obtained in step 101, the individuals to be screened are calculated respectively. With each target person Comprehensive correlation score between For each pair of people ( , ), Represents the target personnel set For the j-th target person in the dataset, after normalization, the output is a value located in [0, ..., ... Values ​​within the range , The larger the value, the stronger the connection between the individuals. To set a preset scoring limit, this embodiment takes... =88. Regarding For specific calculations, please refer to the detailed calculation process of steps 201 to 205 in the subsequent embodiments. For example, scores can be assigned to call data, accommodation data, flight data, and train data respectively, and then weighted and fused to obtain the personnel to be screened. With target personnel Comprehensive correlation score between .

[0033] The specific normalization process will be explained in subsequent embodiments.

[0034] Step 103: Based on multiple comprehensive correlation scores, count the number of target individuals who have a valid correlation with the individuals to be screened.

[0035] After obtaining the individuals to be screened With each target person Comprehensive correlation score Afterwards, statistics were compiled on the individuals to be screened. The number of target personnel with valid associations is denoted as .

[0036] For example, if If the value is greater than 0, then the target personnel are identified. People to be screened The number of effectively associated target individuals is obtained by counting the number of individuals with a comprehensive association score greater than 0. .

[0037] Of course, any threshold can be set, for example, if If the value is greater than 1, then the target personnel are identified. People to be screened One of the effective associated target persons, which is not limited in this application.

[0038] Step 104: Determine the contact breadth gain coefficient based on the number of target personnel.

[0039] Among them, the breadth gain coefficient of the connection As the number of target personnel Monotonically increasing.

[0040] Specifically, in the embodiments of this application, when At 1 o'clock, ;when hour, ,when hour ,when hour, .

[0041] It should be noted that the above values ​​are just examples. In actual applications, different base values ​​and increments can be configured according to business needs.

[0042] Step 105: Based on multiple comprehensive correlation scores and the connection breadth gain coefficient, determine the overall correlation between the individuals to be screened and the target set of individuals.

[0043] Finally, based on multiple comprehensive correlation scores And the breadth gain coefficient Identify the individuals to be screened With the entire target group T Overall correlation .

[0044] The specific calculation process is as follows: First, calculate the number of people to be screened. The sum of the overall relevance scores with all target individuals is then multiplied by the connection breadth gain coefficient. If the product exceeds the preset score limit, the score will be multiplied. Then it is cut off to the upper limit.

[0045] Expressed as a formula: ; in, This represents the total number of target personnel in the target personnel set.

[0046] Through the above five steps, this embodiment of the application outputs the personnel to be screened. Overall correlation This score combines the sum of the correlation strength with each target individual and the breadth of the number of associated target individuals, providing a more comprehensive reflection of the individuals to be screened. The degree to which the target audience is embedded.

[0047] The following describes some application scenarios of the personnel relationship analysis method based on multi-source heterogeneous data provided in the embodiments of this application: (1) Applied to financial crime investigations, the target group consists of confirmed criminal members, and the individuals to be screened are a large number of customers of banks or payment platforms. In addition to those shown in the above embodiments, the multi-source heterogeneous data may also include bank transfer records, etc.

[0048] (2) Internal personnel survey: The target personnel are members of the departments or organizations already surveyed. The personnel to be screened are all employees of the company. In addition to those shown in the above embodiments, multi-source heterogeneous data may also include WeChat records / email records, etc.

[0049] (3) Epidemiological tracing investigation, the target group is the group of confirmed patients. The people to be screened are the residents of the area.

[0050] In summary, the personnel relationship analysis method based on multi-source heterogeneous data provided in this application has the following beneficial effects: First, existing technologies only output independent association scores between the person to be screened and the target individuals to determine their relationship, without considering the overall characteristic that the person to be screened may be associated with multiple target individuals simultaneously. Based on this, this application introduces a connection breadth gain coefficient that monotonically increases with the number of effectively associated target individuals, allowing individuals with extensive connections to multiple target individuals to receive a non-linear amplification in overall association score. This prioritizes organizers, liaisons, and other high-risk individuals within a group in the scoring ranking, addressing the systematic underestimation of such high-risk individuals in existing technologies.

[0051] Second, existing technologies rely solely on a single data source (such as call records). This makes it highly susceptible to errors in determining effective associations, as related parties may deliberately circumvent the data source. This results in a significantly underestimated number of target individuals, rendering the connection breadth gain coefficient ineffective. This application employs at least two types of heterogeneous data, allowing association records from different sources to independently trigger effective association determinations. For example, a person to be screened may have no call records with target A, but co-occurrence records of accommodation; call records with target B, but no accommodation records; and a shared flight record with target C. In a single data source approach, these associations may be partially or completely omitted, leading to a significantly smaller number of target individuals with effective associations compared to the actual number. However, this approach, by integrating multi-source heterogeneous data, can fully identify effective associations between the person to be screened and multiple target individuals across different dimensions. This allows for accurate calculation of the number of target individuals, ensuring that the overall association degree truly reflects the group embedding level.

[0052] Optionally, please refer to Figure 2 The overall correlation score between the person to be screened and any target person is determined through the following steps, including: steps 201 to 205.

[0053] Step 201: Divide the multi-source heterogeneous related data into multiple related dimensions according to data type.

[0054] For example, call data is divided into a call dimension, accommodation data into an accommodation dimension, flight data into a flight dimension, and train data into a train dimension. Each dimension is processed independently before being merged.

[0055] Step 202: For each associated record in each associated dimension, calculate the difference in the number of days between the occurrence time of the associated record and the current time.

[0056] For each association record between the person to be screened and a target person (such as a phone call, a co-occurrence of a one-night stay, a shared flight, or a shared train ride), extract the occurrence time of the association record (accurate to the day), and calculate the difference in days between that time and the analysis baseline time (i.e., the current analysis time, which remains constant within the same batch of tasks): This can be expressed by the formula: ; in, This represents the difference in the number of days the item is invalid, in days. It represents the current time and is a fixed timestamp that remains unique and unchanged within the same batch of analysis tasks to determine the reproducibility of results within the batch; This represents the i-th associated record. When... =0 indicates that the associated record occurred on the day of the analysis. The larger the value, the older the associated record.

[0057] Step 203: Assign a time decay weight to each associated record based on the difference in the number of days of time elapsed.

[0058] Among them, the closer the related record is to the current time, the higher its weight.

[0059] In this embodiment of the application, the time decay weight ranges from [0,1] and decreases monotonically with the difference in the number of days of time elapsed.

[0060] Step 204: Within the same association dimension, accumulate all association records between the person to be screened and the target person by combining them with the assigned time decay weights to obtain the equivalent interaction quantity of the association dimension. Then, map the equivalent interaction quantity to a unified scoring interval through logarithmic normalization to obtain the single-dimensional score of the association dimension.

[0061] In terms of call dimensions, personnel Collection of all call records between ,in, This represents the nth call record, and each record... (Call record i) Contribution to the number of calls made (1 for a single call) and call duration (Unit: minutes) Calculate the attenuation-weighted equivalent number of calls. With equivalent call duration : ; ; Among them, the number of equivalent calls With equivalent call duration This can be used as the equivalent interaction quantity in the call dimension. Summation and traversal of personnel... All call records between them To record the total number of entries. Then... and Substituting each value into the logarithmic normalization formula defined in step one, we obtain the degree sub-quantities. and duration subdivision Then according to the weight parameters (Number of times weight, default 0.6) and (Duration weight, default 0.4) Weighted merging yields the call dimension score. : ; It should be noted that the design principle of setting the weight of the number of calls higher than that of the duration in the embodiments of this application is as follows: the number of calls can reflect the frequency of the willingness to take the initiative to contact, and has a stronger associated initiative signal; the call duration may be inflated in noisy calls (such as sales calls), and using the duration alone will introduce misjudgment, but completely ignoring the duration will lose the strong associated signal of long-term in-depth communication. Therefore, in the embodiments of this application, the number of calls is the main factor and the duration is the auxiliary factor for fusion.

[0062] Regarding accommodation, related records are divided into two categories, with fundamentally different levels of confidence in their association. Sharing a room implies a very high probability of direct contact between two individuals, constituting high-confidence, strongly related evidence; staying in different rooms of the same hotel at overlapping times only indicates the possibility of two individuals being present at the same time, representing medium-confidence evidence. Assigning equal weight to both types of events would significantly underestimate the strength of evidence related to sharing a room. Therefore, in this application's embodiments, an event confidence coefficient is introduced. Records of co-occupancy in the same room Records of different rooms in the same hotel with overlapping time periods are retrieved. In other cases, take . Multiplying by the time decay weight, a two-dimensional modulation of the event confidence score × time-sensitivity weight is achieved for each accommodation record: ; in, Used to represent the equivalent amount of interaction under the accommodation dimension; for Total number of accommodation records between the two locations. For the first The event confidence coefficient of each record. This corresponds to the time decay weight. Substituting into the logarithmic normalization formula, we obtain the accommodation dimension score. .

[0063] For the flight and train dimensions, the processing logic is the same as for the accommodation dimension, assigning different values ​​based on the confidence level of co-occurring events. Coefficient: Records of the same flight and the same cabin class (First Class / Business Class / Economy Class) within the flight dimension are taken. Records for the same flight, regardless of cabin class, are retrieved. Records from the same carriage within the train dimension are retrieved. Records from different carriages on the same train The coefficient for the same train carriage is set to 2.0, consistent with the coefficient for the same room in accommodation, because both have a nearly certain direct contact nature. Then, the equivalent interaction quantity under the flight dimension is obtained. Log-normalization was performed to obtain the flight dimension score. And obtain the equivalent interaction quantity in the train dimension. Logarithmic normalization was performed to obtain the train dimension score. .

[0064] This step outputs the personnel's... Independent rating vectors in four dimensions Each score incorporates the modulation effect of time decay weights and event confidence coefficients, and is uniformly placed within a certain range. Within the comparable interval.

[0065] The normalization method used in this application will be explained below: By performing normalization, raw feature values ​​(such as number of calls, total call duration, number of times living together, etc.) from different data sources and different units of measurement are mapped to a unified standard. The interval, where, In real-world scenarios, behavioral data such as phone calls, accommodations, and travel all follow a heavy-tailed distribution: the vast majority of people interact with each other infrequently, while a few extremely high-frequency association pairs (such as members of the same group) generate observations far exceeding the mean. If the maximum value of the entire dataset is directly used as the normalization benchmark, the feature scores of most normal association pairs will be compressed to near-zero ranges by extreme values, severely losing discriminative power. Therefore, this application's embodiment employs a normalization scheme combining percentile truncation and logarithmic transformation: the 99th percentile of the feature in the entire dataset is taken as a robust upper bound. All exceeding the robust upper bound The original value is truncated to Then, the truncated values Perform logarithmic transformation normalization: ; in, The original observation value of a certain feature. This is the value after truncation at the 99th percentile. This is the 99th percentile of this feature in the entire dataset. It is the natural logarithm. This represents the upper limit of the scoring interval. In this embodiment, logarithmic transformation is chosen over linear transformation because the diminishing marginal returns of the logarithmic function closely match the actual pattern of correlation strength in this scenario: the amount of correlation information carried by increasing the number of calls from 0 to 10 is far greater than that from 90 to 100. Linear normalization would overestimate the incremental value of the high-frequency interval, while logarithmic transformation can naturally compress the high-frequency interval and amplify the resolution of the low-frequency interval, making the scoring curve reasonably interpretable in business terms. Using the 99th percentile instead of the maximum value of the entire dataset as the normalization benchmark is to prevent individual extreme abnormal records from raising the benchmark and causing the overall scores of all other records to be compressed, thereby maintaining the stability of the scoring while preserving the strong correlation signals of high-frequency behaviors.

[0066] Step 205: Weight and fuse the single-dimensional scores of each correlation dimension to obtain the comprehensive correlation score between the person to be screened and the target person.

[0067] When multiple related dimensions exist (e.g., both call and accommodation dimensions have valid records), the individual dimension scores of each dimension are weighted and merged according to preset dimension weights to obtain the individuals to be screened. With target personnel Comprehensive correlation score between If only one valid dimension exists, then the score of that single dimension is used directly as the score. .

[0068] In this embodiment of the application, the initial weights are configured as follows: call dimension 0.40, accommodation dimension 0.30, flight dimension 0.15, and train dimension 0.15.

[0069] In summary, this embodiment unifies heterogeneous data from multiple sources, such as calls, accommodations, flights, and trains, into independent dimensions. It calculates the difference in the number of days elapsed for each record and applies a time decay weight, ensuring that recent interactions dominate the equivalent interaction volume while effectively suppressing the contribution of older historical records. This fundamentally solves the problem of historical noise interference caused by the equal-weighted processing of all-time data in existing technologies. Simultaneously, logarithmic normalization maps equivalent interaction volumes with different dimensions and heavy-tailed distributions to a unified scoring interval. This suppresses the stretching of the scoring benchmark by extremely high-frequency values ​​while maintaining good differentiation between low-frequency and mid-frequency ranges. This allows for direct weighted fusion of scores from different dimensions, such as call counts and accommodation co-occurrence, achieving horizontal comparability of multi-source data. Finally, the comprehensive correlation score output through dimension-weighted fusion objectively and quantitatively reflects the true correlation strength between any pair of individuals, providing reliable and interpretable foundational data for subsequent overall correlation calculations and the discovery of hidden individuals.

[0070] Optionally, the time decay weight is determined through the following steps: processing the difference in the number of days of validity using an exponential decay function, so that the weight decreases exponentially over time according to a preset half-life; processing the difference in the number of days of validity using a linear zeroing function, so that the weight of associated records exceeding the preset deadline is zeroed; and multiplying the value of the exponential decay function by the value of the linear zeroing function to determine the time decay weight.

[0071] Specifically, it can be expressed by the following formula: ; in, This is the difference in the number of days of validity for this record (calculated in step one). The exponential decay coefficient is... The linear zeroing time limit is set to 365 days by default, meaning that the weight of records older than one year is forcibly reset to zero. In practice, it can be configured according to business needs. From half-life parameter The only certainty: Order ,exist Under the condition of approximately obtaining Default value Records older than 30 days decay to half their initial weight. hour, ;when The linear term returns to zero, therefore This achieves a hard cutoff. It should be noted that in the above formula, the exponential term is responsible for continuously adjusting the weights based on recent activity within the validity period, while the linear term is responsible for... The forced zeroing at the point and the multiplication of the two results in the weight curve decaying rapidly in the near term and smoothly returning to zero at the hard cutoff boundary, thus eliminating the problem of residual weights at the tail of the pure exponential function.

[0072] Considering that equal-weighting all-time data can cause older records to interfere with the determination of current association strength, a time decay weight is introduced for each associated record. While pure exponential decay achieves a monotonically decreasing weight over time, its asymptotic approach to zero and never returning to zero means that any historical record, even those occurring ten years ago, still contributes a non-zero weight to the current score. This contradicts the business requirement that records older than a certain number of years should be completely invalidated. Pure linear cutoff (returning to zero upon expiration) results in an overly uniform decay, failing to highlight the additional weight of recent activities relative to medium-term activities. Therefore, in this embodiment, the exponential decay term is multiplied by the linear zeroing term to form the aforementioned dual time decay function.

[0073] Optionally, the above steps map the equivalent interaction quantity to a unified scoring interval through logarithmic normalization to obtain the single-dimensional score of the correlation dimension, including: mapping the equivalent interaction quantity to a unified scoring interval through logarithmic normalization to obtain the initial score of the correlation dimension; determining the phased enhancement coefficient based on the mapping relationship between the equivalent interaction quantity and the preset level interval; and obtaining the single-dimensional score of the correlation dimension based on the phased enhancement coefficient and the initial score of the correlation dimension.

[0074] The four-dimensional scoring vector received in this step First, segmented step enhancements are applied to the scores of each dimension.

[0075] Because logarithmic normalization is used to map the equivalent interaction quantity to the scoring interval, the diminishing marginal property of the logarithmic function, while suppressing extreme value interference, inevitably compresses the resolution of the high-frequency interval: the score increment brought about by increasing the equivalent interaction quantity from 10 to 20 times is very small on the logarithmic scale compared to the score increment brought about by increasing it from 200 to 400 times, but they are fundamentally different in the actual meaning of the association strength. The former belongs to ordinary frequency of interaction, while the latter belongs to extremely abnormal high-frequency contact. The strong association signal between the core members of the group corresponding to the latter should be significantly reflected in the fused score. Therefore, before the scores of each dimension are included in the fused calculation, a piecewise step enhancement coefficient is introduced. It is amplified compensatorily.

[0076] Set up association dimensions The output score is The equivalent interaction volume corresponding to this dimension is (i.e., the equivalent number of times after time decay weighting in the aforementioned steps, such as...) 、 (etc.). Enhancement coefficient in accordance with The gear range to be entered is determined, and the gear threshold and coefficient values ​​are as follows: When hour, ,when hour, ,when hour, ,when hour, ,when hour, Each dimension's threshold level is configured independently and does not interfere with others (for example, the threshold level for the accommodation dimension can be lowered overall because the confidence level of a single co-stay is high). The default parameters given here are for the call dimension. The enhanced score is: ;in, This represents the upper limit of the rating range. The operation ensures that the enhanced score does not exceed the interval boundary. This is the enhanced score for that dimension, which is directly used in subsequent fusion calculations, i.e., the single-dimensional score for that related dimension.

[0077] Optionally, the above steps involve weighted fusion of the individual scores of each correlation dimension to obtain a comprehensive correlation score between the person to be screened and the target person. This includes: when there is only one effective correlation dimension, the individual score of that correlation dimension is used as the comprehensive correlation score; when there are two or more effective dimensions, the individual scores of each correlation dimension are weighted according to a preset dimension weight, and then multiplied by a multi-path evidence enhancement coefficient to obtain the comprehensive correlation score; wherein, the multi-path evidence enhancement coefficient increases with the number of effective correlation dimensions.

[0078] After completing the dimension enhancements in the aforementioned steps, determine the effective set of dimensions. If no related records satisfying the identity merging requirements are found for a certain dimension, it is considered an invalid dimension, not included in the fusion calculation, and its enhancement score is set to zero. Let the number of valid dimensions be... The fusion strategy is executed according to the following logic: when At that time, personnel If only a single dimension of correlation evidence exists among the four types of data sources, the premise of multi-source cross-validation is not valid. Therefore, the enhancement score of this single valid dimension is directly used as the comprehensive correlation score. Output (for any related dimension): .

[0079] when At that time, the enhancement scores of each effective dimension are weighted and averaged, and then a multi-path evidence enhancement coefficient is applied. Let the preset weights for each dimension be... (The original configuration was 0.40 for calls, 0.30 for accommodations, 0.15 for flights, and 0.15 for trains. If a dimension is invalid, its weight is proportionally distributed to the remaining valid dimensions and then normalized.) Multi-channel evidence enhancement coefficient in accordance with The value of is determined when: hour, ,when hour, ,when hour, Overall Relevance Score The calculation is as follows: Among them, the multi-path evidence enhancement coefficient The design rationale for increasing the number of effective dimensions is that cross-validation of two types of data sources has significantly reduced the probability of false alarms in statistics compared to a single dimension. Simultaneous validation from three or more types of data sources means that the associated person is associated with the known target person in multiple independent life dimensions such as communication, travel, and accommodation. This has an extremely low probability in random situations and should be rewarded non-linearly rather than accumulated non-linearly. The specific value of is such that the maximum reward cap when all four dimensions are effective is about 30%, which is considered sufficient in business to distinguish between strong multidimensional correlation and strong single-dimensional correlation, while preventing the reward coefficient from dominating the final score and obscuring the information of the scores of each dimension themselves.

[0080] The above calculations yielded That is, this application is for the personnel. The quantitative expression of comprehensive correlation strength, with a range of values. .

[0081] Optionally, in one embodiment of this application, a method for determining hidden associated persons is also provided. The method further includes: retrieving determination dimensions; wherein, the determination dimensions include multiple association dimensions that divide multi-source heterogeneous association data according to data type and a target person quantity dimension; setting a threshold parameter for each determination dimension; determining the number of determination dimensions triggered by the person to be screened through the threshold parameter set for each determination dimension; when the number of determination dimensions triggered by the person to be screened reaches a preset minimum number of trigger dimensions, the person to be screened is marked as a hidden associated person; and / or, if there is a co-occurrence event between the person to be screened and any target person, the person to be screened is marked as a hidden associated person; wherein, the co-occurrence event includes any one of the following: having a record of sharing a room, having a record of traveling in the same carriage, or having a record of being in the same flight cabin.

[0082] Configure threshold parameters independently for each dimension to form a threshold parameter group. The parameter settings are as follows: Threshold for the number of target personnel to be contacted in the call. (Right now ), the threshold for equivalent interaction volume in a call (correspond (Meeting at least 5 equivalent weighted call counts) and equivalent interaction thresholds in the accommodation dimension. (correspond (Meeting at least two equivalent weighted cohabitation times). The threshold conditions for each of the above dimensions are triggered by the cumulative amount reaching the target.

[0083] In addition, a separate direct trigger condition is set for high-risk co-occurrence events with extremely high confidence: if there is any record of co-occupancy in the same room (not requiring reaching the cumulative number threshold), or any record of traveling in the same train carriage, or any record of traveling in the same cabin class on the same flight, then regardless of the scores of other dimensions, the person to be screened will be directly included. These individuals are marked as having hidden connections. The reason for this setting is that the probability of such co-occurrence events happening randomly in reality is extremely low. Once such records exist with a known target person, they constitute sufficiently strong evidence of a connection, without needing to wait for the accumulation of evidence.

[0084] Define the condition satisfaction function for each dimension. If the person to be screened In dimensions If the evidence reaches the corresponding threshold, then Otherwise, it is 0. If the above co-occurrence event is triggered, then Set to 1 directly. Total number of dimensions that meet the condition. Recorded as: ; in, This is a set of all dimensions for determining correlation (including the number of callers, call volume, accommodation volume, and various high-risk co-occurrence event dimensions). Those to be screened The number of dimensions of the triggering conditions.

[0085] In summary, by independently setting configurable absolute business thresholds for each related dimension such as calls, accommodation, flights, and trains, as well as the target number of people, and by counting the number of dimensions triggered by the people to be screened, the system marks the data when the preset minimum number of triggered dimensions is reached. In addition, it adopts a direct triggering rule that rejects high-risk co-occurrence events such as sharing a room, traveling in the same carriage, or being in the same cabin class on the same flight, thus completely abandoning the uninterpretable Top-K truncation strategy that relies on ranking the entire dataset in existing technologies. This solution ensures that each labeled conclusion corresponds to a clear and readable business trigger condition (such as ≥2 target people associated with a call or ≥2 equivalent accommodation interactions). Analysts can understand the reasons for the labeling without having to backtrack the original data, fundamentally solving the problem of untraceable black-box probability scores. At the same time, the judgment based on absolute thresholds does not depend on data distribution, and there is no need to recalibrate parameters when the target personnel set changes dynamically or the data volume continues to grow, significantly reducing operation and maintenance costs. The rejection mechanism ensures that a single high-risk co-occurrence event can be labeled immediately even if it does not meet the cumulative counting conditions, avoiding the omission of key evidence due to insufficient accumulation, thus providing clear business interpretability while ensuring a high recall rate.

[0086] Optionally, the preset minimum number of trigger dimensions is a first value, or the preset minimum number of trigger dimensions is a second value.

[0087] The first value is 1; the second value is greater than or equal to 2.

[0088] By setting a first or second value, the screening mode for hidden related persons can be distinguished.

[0089] Specifically, this step receives the output of each person to be screened from the preceding steps. Quadruple And an index of original evidence records across all dimensions. First, the final labeling decision is executed. For individuals that do not trigger the veto condition, the final labeling decision is determined by configurable parameters. control: ; in Indicates the individuals to be screened Marked as an invisible associated person, Indicates no marking; This is an indicator function; it takes the value 1 if the condition within the square brackets is true, and 0 otherwise. This is the minimum number of trigger dimensions that business personnel can configure; its value determines the judgment mode. When... When any dimension reaches its corresponding threshold, a flag is triggered. This is a lenient mode, suitable for large-scale preliminary screening scenarios, aiming to minimize missed detections. At that time, at least A flag is triggered only if all dimensions are met simultaneously; this is the strict mode, suitable for precise targeted analysis scenarios, aiming to reduce the false flag rate. Both modes share the same scoring and threshold system, differing only in their adjustment... This parameter enables switching without reconfiguring any parameters in the preceding steps, ensuring the comparability of the output results between the two modes.

[0090] Optionally, the method further includes: extracting the original records of hidden associated personnel and each target personnel in each association dimension; wherein, the multi-source heterogeneous association data is divided into multiple association dimensions according to the data type; calculating the equivalent interaction volume of each association dimension; calculating the recent activity ratio index; wherein, the recent activity ratio index represents the contribution ratio of records within 30 days to the equivalent interaction volume; and outputting the recent activity ratio index.

[0091] After the labeling decision is completed, a structured and interpretable evidence summary is automatically generated for each hidden associated individual. Each numerical field in the summary is directly linked to the original record index, rather than simply outputting summary statistics, ensuring that each conclusion is fully traceable and verifiable at the data level. The evidence summary includes the following fields: individual identification information (basic attributes such as ID number and name); overall correlation. and its corresponding connection breadth gain coefficient The trigger flag's mode type (negation trigger or...) (Triggered upon reaching the target dimension); for each target person with a valid connection. List the following separately: Dimension type (call / accommodation / flight / train), equivalent interaction volume (Including the weighted result with time decay weight), number of original record entries, and time of the most recent associated record occurrence; for veto-type high-risk co-occurrence events, additionally note the specific event type (shared room / same carriage / same flight cabin) and the corresponding original data record index.

[0092] Efficient interaction volume in the abstract The temporal distribution characteristics are characterized by the recent proportion index to intuitively reflect the timeliness of related activities. (Regarding dimensions...) Define the recent percentage of all related records under this category. Within the last 30 days (i.e.) The ratio of the contribution of the records (days) to the total equivalent interaction volume in this dimension: ; in, The dual time decay weights defined in the preceding steps have a numerator of the sum of decay weights recorded within the last 30 days and a denominator of the sum of decay weights for all records in that dimension. ,when When the value is close to 1, it indicates that the related activities in this dimension are concentrated in the recent period and have high timeliness; when When the value is low, it indicates that the association mainly comes from historical records, and analysts can use this to determine the current activity level of the association. Embedded as an independent field in the summary, it allows analysts to directly obtain timeliness judgments when reading evidence summaries, without having to trace back to the original time series data.

[0093] Optionally, the method further includes: using all hidden related personnel and all target personnel in the target personnel set as nodes; the attributes of the nodes include at least personnel identifier, personnel category label, and overall correlation degree; using the relationship between each pair of personnel as a directed weighted edge; the attributes of the directed weighted edge include at least one of the following: set of relationship dimension types, comprehensive correlation score, equivalent interaction volume, and recent activity ratio index; and outputting the structured data of the nodes and directed weighted edges to form a personnel relationship network diagram.

[0094] Finally, using all personnel (known target personnel and implicitly related personnel) as nodes and the relationships as directed weighted edges, a structured relationship graph is constructed and output. Node attributes include: ID number index, personnel category label (known target personnel or implicitly related personnel), and overall relationship degree. Trigger mode type; edge attributes include: set of related dimension types, comprehensive relatedness score. Equivalent interaction volume Recent percentage The number of original record entries. The graph data is output in a standardized structure, with node sets and edge sets stored separately. It can be directly connected to the front-end visualization engine to render an interactive personnel relationship network graph. It supports filtering and coloring of nodes and edges according to attributes such as association dimension, score range, and timeliness, enabling analysts to evaluate the mining results from both the macro-relationship network and micro-evidence details levels.

[0095] Please see Figure 3 Based on the same inventive concept, embodiments of this application also provide a personnel relationship analysis system 30 based on multi-source heterogeneous data, including: The acquisition module 301 is used to acquire multi-source heterogeneous association data of each target person in the set of people to be screened and the set of target people; wherein, the data types of the multi-source heterogeneous association data include at least two of the following: call data, accommodation data, flight data and train data.

[0096] The calculation module 302 is used to calculate the comprehensive correlation score between the person to be screened and each target person based on the multi-source heterogeneous correlation data.

[0097] The statistics module 303 is used to count the number of target personnel who have a valid association with the person to be screened based on multiple comprehensive correlation scores.

[0098] The determination module 304 is used to determine the contact breadth gain coefficient based on the number of target personnel; wherein the contact breadth gain coefficient increases monotonically with the number of target personnel.

[0099] Analysis module 305 is used to determine the overall correlation between the person to be screened and the target person set based on multiple comprehensive correlation scores and the connection breadth gain coefficient.

[0100] Please see Figure 4 Based on the same inventive concept, this application provides a module block diagram of an electronic device 400 applying the above-described method. The electronic device 400 includes: at least one processor 401 ( Figure 4 (Only one is shown in the image), memory 402, computer program 403 stored in memory 402 and executable on at least one processor 401, processor 401 executing computer program 403 to implement the steps of the method in any of the foregoing embodiments.

[0101] The electronic device 400 can be a server, a personal computer, etc.

[0102] Those skilled in the art will understand that Figure 4 This is merely an example of electronic device 400 and does not constitute a limitation on electronic device 400. It may include more or fewer components than shown, or combine certain components, or use different components.

[0103] The processor 401 may be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0104] In some embodiments, the memory 402 may be an internal storage unit of the electronic device 400, such as a hard disk or memory of the electronic device 400. In other embodiments, the memory 402 may also be an external storage device of the electronic device 400, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 400. Furthermore, the memory 402 may include both internal storage units and external storage devices of the electronic device 400.

[0105] It should be noted that the above-mentioned systems, devices, etc. are based on the same concept as the method embodiments of this application. The modules designed in the system, as well as the steps performed by the device and the resulting technical effects, can all be found in the method embodiments section, and will not be repeated here.

[0106] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the above-described method embodiments.

[0107] This application provides a computer program product that, when run on a mobile terminal, causes the mobile terminal to execute the steps described in the various method embodiments above.

[0108] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographic device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0109] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0110] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0111] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for analyzing personnel relationships based on multi-source heterogeneous data, characterized in that, include: Obtain multi-source heterogeneous correlation data for each target person in the set of people to be screened and the set of target people; wherein, the data types of the multi-source heterogeneous correlation data include at least two of the following: call data, accommodation data, flight data, and train data; Based on the multi-source heterogeneous association data, the comprehensive association score between the person to be screened and each target person is calculated respectively; Based on the comprehensive correlation scores, the number of target individuals who have a valid correlation with the individuals to be screened is counted. Based on the number of target personnel, a connection breadth gain coefficient is determined; wherein, the connection breadth gain coefficient increases monotonically with the number of target personnel. Based on multiple comprehensive correlation scores and the connection breadth gain coefficient, the overall correlation between the person to be screened and the target personnel set is determined.

2. The method for analyzing personnel relationships based on multi-source heterogeneous data according to claim 1, characterized in that, The comprehensive correlation score between the person to be screened and any target person is determined through the following steps: The multi-source heterogeneous associated data is divided into multiple association dimensions according to data type; For each associated record in each associated dimension, calculate the difference in the number of days between the occurrence time of the associated record and the current time; Based on the difference in the number of days of validity, a time decay weight is assigned to each associated record; among them, the associated record that is closer to the current time has a higher weight; Within the same association dimension, all association records between the person to be screened and the target person are combined and accumulated with assigned time decay weights to obtain the equivalent interaction quantity of the association dimension. The equivalent interaction quantity is then mapped to a unified scoring range through logarithmic normalization to obtain the single-dimensional score of the association dimension. The scores of each correlation dimension are weighted and fused to obtain the comprehensive correlation score between the person to be screened and the target person.

3. The method for analyzing personnel relationships based on multi-source heterogeneous data according to claim 2, characterized in that, The time decay weight is determined through the following steps: The difference in the number of days of time is processed by an exponential decay function, so that the weight decreases exponentially over time according to a preset half-life. The difference in the number of days of validity is processed using a linear zeroing function, so that the weight of the associated records that exceed the preset deadline is zeroed. The time decay weight is determined by multiplying the value of the exponential decay function by the value of the linear zero-return function.

4. The method for analyzing personnel relationships based on multi-source heterogeneous data according to claim 2, characterized in that, The step of weighting and fusing the single-dimensional scores of each correlation dimension to obtain the comprehensive correlation score between the person to be screened and the target person includes: When there is only one valid correlation dimension, the single-dimensional score of that correlation dimension shall be used as the comprehensive correlation score. When there are two or more valid dimensions, the single-dimensional scores of each related dimension are weighted and averaged according to the preset dimension weights, and then multiplied by the multi-way evidence enhancement coefficient to obtain the comprehensive correlation score. The multi-path evidence enhancement coefficient increases with the number of effective association dimensions.

5. The method for analyzing personnel relationships based on multi-source heterogeneous data according to claim 2, characterized in that, The step of mapping the equivalent interaction quantity to a unified scoring interval through logarithmic normalization to obtain a single-dimensional score for the association dimension includes: The equivalent interaction quantity is mapped to a unified scoring range through logarithmic normalization to obtain the initial score of this association dimension; Based on the mapping relationship between the equivalent interaction amount and the preset gear range, the phased enhancement coefficient is determined; Based on the phased enhancement coefficient and the initial score of the correlation dimension, the single-dimensional score of the correlation dimension is obtained.

6. The method for analyzing personnel relationships based on multi-source heterogeneous data according to claim 1, characterized in that, The method further includes: Retrieve the judgment dimensions; wherein, the judgment dimensions include multiple association dimensions that divide the multi-source heterogeneous association data according to data type and the target personnel quantity dimension; a threshold parameter is set for each judgment dimension; The number of judgment dimensions triggered by the person to be screened is determined by setting the threshold parameter for each judgment dimension; When the number of judgment dimensions triggered by the person to be screened reaches the preset minimum number of trigger dimensions, the person to be screened is marked as a hidden associated person; and / or, If there is a co-occurrence event between the person to be screened and any target person, then the person to be screened is marked as a hidden associated person; The co-occurrence events include any one of the following: having a record of sharing a room, having a record of traveling in the same carriage, or having a record of being in the same flight cabin.

7. The method for analyzing personnel relationships based on multi-source heterogeneous data according to claim 6, characterized in that, The preset minimum number of trigger dimensions is a first value, or the preset minimum number of trigger dimensions is a second value; The first value is 1; The second value is greater than or equal to 2; By setting the first value or the second value, the screening mode for hidden related persons can be distinguished.

8. The method for analyzing personnel relationships based on multi-source heterogeneous data according to claim 6, characterized in that, The method further includes: Extract the original records of the hidden associated personnel and each target personnel in each association dimension; wherein, the multi-source heterogeneous association data is divided into multiple association dimensions according to data type; Calculate the equivalent interaction volume for each related dimension; Calculate the recent activity percentage index; wherein, the recent activity percentage index represents the proportion of records within 30 days that contribute to the equivalent interaction volume; Output the recent activity percentage metric.

9. The method for analyzing personnel relationships based on multi-source heterogeneous data according to claim 6, characterized in that, The method further includes: All hidden related personnel and all target personnel in the target personnel set are used as nodes; the attributes of the nodes include at least personnel identifier, personnel category label, and overall correlation degree; The relationships between each pair of people are used as directed weighted edges; The attributes of the directed weighted edge include at least one of the following: set of association dimension types, comprehensive association score, equivalent interaction volume, and recent activity ratio. The structured data of the nodes and the directed weighted edges are output to form a personnel relationship network graph.

10. A personnel relationship analysis system based on multi-source heterogeneous data, characterized in that, include: The acquisition module is used to acquire multi-source heterogeneous association data for each target person in the set of people to be screened and the set of target people; wherein, the data types of the multi-source heterogeneous association data include at least two of the following: call data, accommodation data, flight data, and train data; The calculation module is used to calculate the comprehensive correlation score between the person to be screened and each target person based on the multi-source heterogeneous correlation data. The statistics module is used to count the number of target individuals who have a valid association with the individuals to be screened, based on multiple comprehensive correlation scores. The determination module is used to determine the contact breadth gain coefficient based on the number of target personnel; wherein the contact breadth gain coefficient increases monotonically with the number of target personnel. The analysis module is used to determine the overall correlation between the person to be screened and the target set of persons based on multiple comprehensive correlation scores and the connection breadth gain coefficient.