Model for calculating relationship intimacy
By using the relationship intimacy calculation model in the public security system, combining random forest, logistic regression and naive Bayes algorithm, and combining principal component analysis, the problem of being unable to accurately evaluate relationship in the existing technology is solved, and higher prediction accuracy and more accurate relationship evaluation are achieved.
Patent Information
- Application Number
- CN202510270899.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
The relationship intimacy calculation methods used by existing public security systems rely on simple rules or statistical methods, cannot effectively process complex multi-dimensional data, and it is difficult to capture the mutual influence between variables, resulting in the inability to accurately evaluate the relationship intimacy and affecting the public security department to identify crimes.
A model is provided for calculating relationship intimacy. The data set is obtained through internal channels of the public security, divided into multiple sets of data and cleaned and deduplicated. The model is trained using random forest, logistic regression and naive Bayes algorithms, and the intimacy of two people who know each other is calculated and output.
By integrating random forests of multiple decision trees, parameter intuitiveness of logistic regression and naive Bayes mathematical foundation, the prediction accuracy of relationship intimacy can be improved, and relationship strength can be evaluated more accurately, helping the public security department more effectively identify crimes and prevent potential threats.
Smart Images

Figure CN120217178A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of calculating intimacy, and specifically relates to a model for calculating relationship intimacy. Background Art
[0002] A model for calculating relationship intimacy is a technical tool that quantitatively evaluates the strength of relationships between individuals by analyzing information such as interaction data, behavior patterns, and social network structures among them. Such models usually combine technologies such as data analysis, machine learning, graph theory, and natural language processing, and are widely used in fields such as social network analysis, public security systems, and recommendation systems. Using a model for calculating relationship intimacy in the public security system can bring many benefits. By analyzing suspects and their social relationships, it is possible to quickly locate key individuals and core members of criminal gangs, identify individuals closely related to criminal activities, and help discover hidden criminal networks. By analyzing the relationship intimacy in different cases, it is possible to identify whether there are related cases and improve the efficiency of case linking. During crime prevention, by analyzing an individual's social relationships and behavior patterns, potential high-risk groups or individuals with criminal tendencies can be identified. By analyzing the relationship intimacy between groups, possible social conflicts or mass incidents can be predicted. In social governance and public security, by analyzing the relationship intimacy among community residents, potential contradictions or conflicts can be identified, the relationship network in areas with complex public security can be analyzed, and targeted control measures can be formulated. And during emergencies, key individuals can be quickly located through the relationship network to assist in emergency response.
[0003] The existing public security system uses traditional methods for calculating relationship intimacy. This algorithm relies on simple rules or statistical methods, cannot effectively process complex multi-dimensional data, is difficult to capture the mutual influence between variables, and cannot accurately evaluate relationship intimacy, thus affecting the public security department's ability to identify crimes. Summary of the Invention
[0004] The purpose of the present invention is to provide a model for calculating relationship intimacy to solve the problem that the algorithm used in the existing public security system relies on simple rules or statistical methods and cannot accurately evaluate relationship intimacy, thus affecting the public security department's ability to identify crimes.
[0005] To achieve the above objective, the basic solution provided by the present invention is: A model for calculating relationship intimacy, comprising the following steps: S1. Obtain a data set through internal channels of the public security department, and divide the data set into multiple groups of data according to characteristic variables; S2. Divide the multiple groups of data into positive samples and negative samples according to user requirements; S3. Clean and deduplicate the data of the positive samples and negative samples; S4. Use the random forest, logistic regression, and Naive Bayes algorithms to train the model on 80% of the samples, and classify multiple pairs of people whose relationship needs to be judged as knowing each other or not knowing each other. S5. Use 20% of the samples for verification, obtain the optimal algorithm based on the verification results, and output the classifier. S6. Input the information of the two people whose relationship needs to be judged into the trained model, repeat steps S1 to S3, use the classifier in step S5 to judge whether the two people know each other. If the two people do not know each other, end; if the two people know each other, then proceed to the next step. S7. Calculate and output the intimacy of the two acquaintance personnel through principal component analysis.
[0006] The principle and beneficial effects of the present invention are as follows: When calculating the intimacy relationship between two personnel, first divide multiple groups of data into positive samples and negative samples according to user needs, clean and remove duplicates, and then perform model training. Use the random forest, logistic regression, and Naive Bayes algorithms to train the model on 80% of the samples, and then use 20% of the samples for verification. Obtain the optimal algorithm suitable for user needs based on the results. Finally, input the processed information of the two people whose intimacy needs to be judged into the optimal algorithm to obtain whether the two people know each other. If they do not know each other, end; if they know each other, then calculate the specific intimacy value of the two acquaintance personnel through principal component analysis. The random forest improves the prediction accuracy by integrating multiple decision trees. The parameters of logistic regression are intuitive and can explain the influence of each feature on the result, which is suitable for scenarios that require interpretability. Naive Bayes is based on Bayes' theorem and has a solid mathematical foundation, and has advantages in high-dimensional data processing and other aspects. Principal component analysis can extract the main features, help understand the data structure, and increase the accuracy.
[0007] Solution 2, which is the optimization of the basic solution. In step S1, the feature variables include: registration relationships, trajectory relationships, communication relationships, and other relationships. The data of the feature variables are subjected to data type conversion and keyword field de-duplication; the registration relationships include: same family, classmates in the same class, and population registration, etc.; the trajectory relationships include: same passenger transportation, same flight, same train, and same location perception, etc.; the communication relationships include: telephone connection and LBS - location-based service, etc.; the other relationships include: same clan, same case, same prison area, and same unit, etc.
[0008] Solution 3, which is the optimization of the basic solution. In step S3, through missing value processing, outlier processing, and normalization processing, correct and fill the outliers, extreme values, and missing values in the positive samples and negative samples; the missing value processing is to fill the missing data in the samples with 0 during the data analysis and modeling process. The outlier processing is to use the capping method to remove the extreme values higher than the 85% quantile value. The normalization processing is that the normalized data eliminates the dimensional difference between features, and the model only focuses on the distribution characteristics of features at the statistical level, meeting the modeling standards.
[0009] Solution 4, which is the optimization of the basic solution. In step S4, Random forest training model: Randomly select features from the feature variables in S1 to construct decision trees, set the number of feature variables in the random forest, the amount of data for each feature variable, and the minimum number of samples required for each feature variable to split. Finally, use the majority voting mechanism to determine whether two people know each other; Logistic regression training model: According to the feature variables in S1, map the linear output to between 0 and 1 through a logistic function to determine whether two people know each other; Naive Bayes training model: According to the feature variables in S1, calculate the prior probability, conditional probability, and posterior probability of each feature variable, and then select the category with the highest posterior probability as the prediction result to determine whether two people know each other.
[0010] By integrating multiple decision trees, the random forest improves the prediction accuracy, especially performs well on complex data sets, effectively reduces the risk of overfitting, enhances the generalization ability of the model, can handle a large number of features, and automatically selects important features during the training process to reduce the dimension; Logistic regression is simple to implement, has high computational efficiency, is suitable for large-scale data sets, the model parameters are intuitive, can explain the impact of each feature on the result, is suitable for scenarios that require interpretability. Compared with complex models, logistic regression has lower requirements for computing resources, is suitable for environments with limited resources, can update the model gradually, is suitable for data stream or online learning scenarios, is not sensitive to small fluctuations in the data, and the model performance is stable; Naive Bayes is simple to implement, performs well in high-dimensional data, can handle a large number of features, has fast training and prediction speeds, is suitable for large-scale data sets, and has lower requirements for computing resources, is suitable for environments with limited resources, and has a solid mathematical foundation based on Bayes' theorem.
[0011] Solution 5, which is the optimization of the basic solution. In step S5, based on the test results of 20% of the samples, if the accuracy rate ≥ 80%, precision rate ≥ 80%, and f1 ≥ 80% of the test results are all satisfied, output the classifier according to the optimal algorithm; if the accuracy rate ≥ 80%, precision rate ≥ 80%, and f1 ≥ 80% of the test results are all not satisfied, loop back to S2 to adjust the samples and feature values until the classifier is output.
[0012] Solution 6, which is the optimization of the basic solution. In step S7, the steps for calculating the intimacy include: a. Extract the positive samples of the original data of the feature variables that two people know; b. Calculate the values of each continuous feature variable; c. Fill in the null values in b, perform positive transformation on the reverse variables, and then perform data normalization processing; d. Output the original variables, principal component coefficients, and the contribution rate of each principal component; e. Calculate the intimacy score by substituting the original variables, principal component coefficients, and the contribution rate of each principal component into the formula; Principal component analysis is based on linear algebra and statistics, with a reliable theoretical foundation. It can effectively reduce the data dimension, retain the main information, reduce the computational complexity, reduce the amount of data after dimensionality reduction, speed up the model training speed, improve the efficiency and performance of the model by eliminating redundant features, and can also remove noise and improve the data quality.
[0013] Solution VII, which is an optimization of Solution VI. In step e, the intimacy calculation formula is: , where F n represents the variance contribution rate of the nth principal component, U n represents the original variable and the principal component coefficient, and X n represents the value of the original variable after preprocessing. Brief Description of the Drawings
[0014] Figure 1 is a flowchart of the model training for calculating relationship intimacy of the present invention; Figure 2 is a flowchart of the model usage for calculating relationship intimacy of the present invention. Detailed Description of the Invention
[0015] The present invention will be further described in detail below through specific embodiments: Embodiment As Figure 1 and Figure 2 shown: A model for calculating relationship intimacy includes the following steps: S1. Obtain a data set through the internal channels of the public security department, and divide the data set into multiple groups of data according to characteristic variables, where the characteristic variables include: registration relationships, trajectory relationships, communication relationships, and other relationships. The data of the characteristic variables are subjected to data type conversion and duplicate removal of key fields; S2. Divide the multiple groups of data into positive samples and negative samples according to user requirements; S3. Clean and remove duplicates from the data of the positive samples and negative samples, and correct and fill the outliers, extreme values, and missing values in the positive samples and negative samples through missing value processing, outlier processing, and normalization processing; S4. Use the random forest, logistic regression, and naive Bayes algorithms to train the model on 80% of the samples respectively, Random forest training model: Randomly select features from the feature variables in S1 to construct decision trees. Set the number of feature variables in the random forest, the amount of data for each feature variable, and the minimum number of samples required for each feature variable to split. Finally, use the majority voting mechanism to determine whether two people know each other; Logistic regression training model: Based on the feature variables in S1, map the linear output to between 0 and 1 through a logistic function to determine whether two people know each other; Naive Bayes training model: Based on the feature variables in S1, calculate the prior probability, conditional probability, and posterior probability of each feature variable, and then select the category with the highest posterior probability as the prediction result to determine whether two people know each other; S5. Use 20% of the samples for verification. Through the test results of the 20% samples, if the accuracy rate of the test results ≥ 80%, the precision rate ≥ 80%, and f1 ≥ 80% are all satisfied, then output the classifier according to the optimal algorithm; if the accuracy rate of the test results ≥ 80%, the precision rate ≥ 80%, and f1 ≥ 80% are all not satisfied, then loop back to S2 to adjust the samples and feature values until the classifier is output; S6. Input the information of two people whose relationship needs to be judged into the trained model. Repeat steps S1 to S3, and use the classifier obtained in step S5 to determine whether the two people know each other. If the two people do not know each other, end; if the two people know each other, then proceed to the next step; S7. Calculate and output the intimacy of two people who know each other through principal component analysis. The steps for calculating intimacy include: a. Extract the positive samples of the original data of the feature variables that two people know; b. Calculate the values of each continuous feature variable; c. Fill in the null values in b, perform forward processing on the reverse variables, and then perform data normalization processing; d. Output the original variables, principal component coefficients, and contribution rates of each principal component; e. Substitute the output original variables, principal component coefficients, and contribution rates of each principal component into the formula to calculate the intimacy score. The intimacy calculation formula is: , where F n represents the variance contribution rate of the nth principal component, U n represents the original variable and the principal component coefficient, and X n represents the value of the original variable after preprocessing.
[0016] The implementation method of this embodiment is as follows: When calculating the intimacy relationship between two persons, model training is first performed. The data set is obtained through the internal channels of the public security department and divided into multiple groups of data according to feature variables, where the feature variables include: registration relationships, trajectory relationships, communication relationships, and other relationships. Among them, the registration relationships include: same family, classmates in the same class, population registration, etc.; the trajectory relationships include: same passenger transportation, same flight, same train, same sense of place, etc.; the communication relationships include: telephone connection and LBS - Location - Based Service, etc.; the other relationships include: same clan, same case, same prison area, same unit, etc. The data of the feature variables are subjected to data type conversion and duplicate removal of key fields. Then, according to the user's needs, the multiple groups of data are divided into positive samples and negative samples. Next, the data of the positive samples and negative samples are cleaned and de - duplicated. Through missing value processing, outlier processing, and normalization processing, the outliers, extreme values, and missing values in the positive samples and negative samples are corrected and filled. The random forest, logistic regression, and naive Bayes algorithms are respectively used to train the model with 80% of the samples. Finally, 20% of the samples are used for verification. If the accuracy rate ≥ 80%, precision rate ≥ 80%, and f1 ≥ 80% of the test results are all satisfied, the classifier is output according to the optimal algorithm. If the accuracy rate ≥ 80%, precision rate ≥ 80%, and f1 ≥ 80% of the test results are not all satisfied, it loops back to S2 to adjust the samples and feature values until the classifier is output. After cleaning, de - duplicating, and pre - processing the information of the two persons whose intimacy relationship needs to be judged, the information is input into the classifier obtained from the above - mentioned model training. It is judged whether the two persons know each other through the classifier. If they don't know each other, the process ends. If they know each other, the intimacy of the two acquaintance persons is calculated by principal component analysis. The steps for calculating the intimacy are as follows: First, extract the positive samples of the original data of the feature variables of the two persons' acquaintance. Secondly, calculate the values of each continuous feature variable, then fill in the null values of the feature variables, perform forward - transformation on the reverse variables, then perform data normalization processing, and then output the original variables, principal component coefficients, and the contribution rates of each principal component. Finally, substitute the output original variables, principal component coefficients, and the contribution rates of each principal component into Calculate the intimacy score.
[0017] The above are only the embodiments of the present invention. Common knowledge such as specific structures and characteristics in the prior art are not described in detail here. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be subject to the content of its claims, and the specific implementation manners in the specification can be used to interpret the content of the claims.
Claims
1. A model for calculating relationship intimacy, characterized in that: The following steps are involved: S1. Obtain the data set through internal channels of the public security department and divide the data set into multiple groups according to characteristic variables; S2. Divide multiple groups of data into positive samples and negative samples according to user needs; S3. Clean and remove duplicate data of positive and negative samples; S4. Use random forest, logistic regression and naive Bayes algorithms to train models on 80% of the samples, and divide multiple groups of two people whose relationships need to be judged into acquaintances and non-acquaintances; S5. Use 20% of the samples for verification, derive the optimal algorithm based on the verification results, and output the classifier; S6. Input the information of the two people whose relationship needs to be determined into the trained model, repeat steps S1 to S3, and use the classifier in step S5 to determine whether the two people know each other. If the two people do not know each other, the process ends; if the two people know each other, proceed to the next step; S7. Calculate and output the intimacy between two acquaintances through principal component analysis.
2. A model for calculating relationship intimacy according to claim 1, characterized in that ,In step S1, the feature variables include: registration class relations, ,trajectory class relations, communication relations and other relations, and the data of the ,feature variables are converted into data types and the key fields are ,deduplicated.
3. A model for calculating relationship intimacy according to claim 1, characterized in that: In step S3, the abnormal values, outliers and missing values in the positive samples and negative samples are corrected and filled through missing value processing, outlier processing and normalization processing.
4. A model for calculating relationship intimacy according to claim 1, characterized in that: In step S4, Random forest training model: randomly select features from the feature variables of S1, build a decision tree, set the number of feature variables in the random forest, the amount of data for each feature variable, and the minimum number of samples required for each feature variable to split, and finally use the majority voting mechanism to determine whether two people know each other; Logistic regression training model: Based on the feature variables in S1, the linear output is mapped to between 0 and 1 through a logical function to determine whether two people know each other; Naive Bayes training model: Based on the feature variables in S1, the prior probability, conditional probability, and posterior probability of each feature variable are calculated, and then the category with the highest posterior probability is selected as the prediction result to determine whether two people know each other.
5. A model for calculating relationship intimacy according to claim 1, characterized in that: In step S5, through the test results of 20% samples, if the accuracy of the test results is ≥80%, the precision is ≥80% and f1 is ≥80%, the classifier is output according to the optimal algorithm; if the accuracy of the test results is ≥80%, the precision is ≥80% and f1 is ≥80%, the loop goes to S2 to adjust the samples and feature values until the classifier is output.
6. A model for calculating relationship intimacy according to claim 1, characterized in that: In step S7, the step of calculating the intimacy includes: a. Extract the positive samples of the original data of the characteristic variables of two people’s acquaintance; b. Calculate the value of each continuous characteristic variable; c. Fill the empty values in b, perform positive processing on the inverse variables, and then normalize the data; d. Output the original variables, principal component coefficients and the contribution rate of each principal component; e. Calculate the intimacy score by substituting the original variables, principal component coefficients and the contribution rate of each principal component into the formula.
7. A model for calculating relationship intimacy according to claim 6, characterized in that: In step e, the intimacy calculation formula is: , where F n Indicates the variance contribution rate of the nth principal component, U n Represents the original variables and the principal component coefficients, X n Represents the value of the original variable after preprocessing.