A method and system for identifying economically disadvantaged college students based on frustrated random walk and feature weighted clustering

By applying the frustrated random walk algorithm and feature weighted clustering method in the identification of students with financial difficulties in colleges and universities, the problem of inaccurate data diversity and feature indicator calculations in traditional methods is solved, and more efficient identification and clustering accuracy of students with financial difficulties is achieved.

CN116522120BActive Publication Date: 2025-05-09SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211425243.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-05-09
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

The existing technology has problems such as data diversity, inaccurate calculation of characteristic indicators and overfitting in the identification of students with economic difficulties in colleges and universities, making it difficult to effectively identify students with economic difficulties.

Method used

The frustrated random walk algorithm and feature weighting clustering method are used to construct a fully connected student information graph, calculate the weight of feature indicators, filter and compress features, and effectively identify the generation of economic difficulties.

Benefits of technology

The data universality and clustering accuracy of economic difficulties identification are improved, adapt to changes in the influence of characteristics during different time periods, without retraining parameters, and ensuring business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116522120B_ABST
    Figure CN116522120B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for identifying college students with financial difficulties based on frustrated random walk and feature weighted clustering, including: obtaining and preprocessing the historical data of the one-card flow; obtaining student work data and integrating it with the historical data of the one-card flow to obtain a data set that is ultimately used to identify college students with financial difficulties; constructing a fully connected graph of student information on each feature; executing a frustrated random walk algorithm on the fully connected graph; calculating the weights of feature indicators using the random walk results; and performing feature optimization on the weights of feature indicators to achieve identification of college students with financial difficulties. The present invention solves the differences in data from various schools, solves the problem of the diversity of one-card consumption data in different canteens of different colleges and universities, and generates universal feature indicators; the present invention solves the problem that a simple random walk algorithm cannot find key graph node features that affect the current feature from the feature association graph, thereby comprehensively improving the feature weight calculation performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning and data mining of information technology, and in particular to a method and system for identifying college students with financial difficulties by using frustrated random walk and feature weighted clustering. Background Art

[0002] Financial difficulties will bring inconvenience to students' campus life and have a certain negative impact on their physical and mental development. However, some students from poor families are unwilling to apply for identification as students with financial difficulties for various reasons. Therefore, identification of students with financial difficulties is an important task in college student management. Using big data to "invisibly subsidize" students with financial difficulties is to use big data technology analysis to find a balance between financial aid fairness and student dignity. It is both humane and accurate and efficient. The accuracy of funding lies in finding students who have the same behavior as students who have been identified as students with financial difficulties, and it is necessary to identify students with financial difficulties. At present, traditional data analysis methods for students with financial difficulties mainly focus on the identification of students with financial difficulties based on the consumption data of the one-card pass. The consumption data of the one-card pass can directly reflect the economic status of students and is the core data for analyzing whether students are in financial difficulties. However, the single-source one-card consumption data has certain limitations for analyzing students with financial difficulties, which are mainly reflected in the following three aspects:

[0003] First, the proportion of students with financial difficulties in colleges and universities is less than 1 / 4. If a single classification model is used, the model training on the original data set will be statistically biased towards ordinary students, which leads to poor classification results and overfitting problems. Second, the one-card data is diverse, and traditional research methods do not provide a method to calculate common indicators for different forms of consumption records. Third, in addition to the one-card consumption data, some other data of students in school (such as the number of work-study programs and per capita family income) can also reflect the students' economic status and need to be considered comprehensively. Therefore, how to use multi-dimensional student behavior time series data to identify students with financial difficulties is a very challenging research task, that is, how to deal with the differences in the degree of impact of different data factors on students' financial difficulties in different periods. For example, during the epidemic school closure, students can only eat at school, then the increase in the impact factor of the one-card consumption data may not be the only criterion for identifying students with financial difficulties; during the holidays, the decrease in the impact factor of the one-card data also cannot reflect the students' true consumption capacity; in addition, different schools have different data completeness and data accuracy, which will also affect the identification of students with financial difficulties, such as the authenticity of the family status filled in by students when they enter the school, etc. These factors will affect the final identification of students with financial difficulties. Summary of the invention

[0004] In order to overcome the shortcomings of the prior art, the present invention first proposes common indicator features for identifying financially disadvantaged college students, including consumption times, average consumption amount, total consumption amount, canteen consumption amount, Engel coefficient, gender, scholarship information, family population, per capita family income information, student aid amount, work-study times, loan amount, and minimum living security line of the student's place of origin; then, based on the frustrated random walk algorithm, the weight proportion of the common feature indicators in different periods is calculated; secondly, the feature optimization module is used to screen out the indicator features with higher influence, and the indicator features with lower influence are effectively deleted; thirdly, the simplified indicator feature dimensions are effectively compressed, and finally, on the basis of indicator feature screening and streamlining, the effective identification of financially disadvantaged students is achieved.

[0005] The present invention also proposes a system for identifying college students with financial difficulties based on frustrated random walk and feature weighted clustering.

[0006] Terminology explanation:

[0007] LDA model, Linear Discriminant Analysis (LDA) is to project high-dimensional pattern samples into low-dimensional vector space, so that after the samples are projected in the new space, the distance between samples of the same category is minimized, and the distance between samples of different categories is maximized, thereby achieving the effect of extracting classification information and compressing the dimension of feature space. Therefore, it is an effective feature extraction method. Let S w is the intra-class dispersion matrix, S N is the inter-class discreteness matrix, then the goal of LDA is to find a transformation V so that the Fisher criterion is w Maximum when non-singular: Get V i Solution for (i=1,2,…,m): V i =λS w V i Among them, V i (i=1,2,…,m) is the matrix composed of the eigenvectors corresponding to the first m larger eigenvalues ​​λ.

[0008] The technical solution of the present invention is:

[0009] A method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering includes the following steps:

[0010] Obtain the historical data of the card transaction and perform preprocessing;

[0011] The student work data is obtained and integrated with the pre-processed one-card flow history data to obtain the final data set used to identify students with financial difficulties in colleges and universities; the student work data includes students' basic information, family situation and school activity information;

[0012] Construct a fully connected graph of student information on each feature;

[0013] Execute the frustrated random walk algorithm on the constructed fully connected graph to obtain the random walk result;

[0014] The weights of characteristic indicators are calculated using random walk results;

[0015] The weights of characteristic indicators are optimized to identify students with financial difficulties in colleges and universities.

[0016] According to a preferred embodiment of the present invention, obtaining the historical data of the card flow and preprocessing it includes:

[0017] Collect the historical data of the card flow in the previous period, including: student number S o , Consumption time C t 、Consumption location C p 、Consumption amount C m Four attributes, let the card transaction history data be y, and there are n data in total, then

[0018] Consumption time C of the historical data of the card flow t Processing, converting the specific time point into the consumption date + consumption interval mode, let sub(C t ) is C t The specific conversion rules are as follows: C′ t Date+sub(C t ); Convert the card transaction history data into

[0019] According to student number S o , C′ t Aggregate y′ and calculate Sum(C m )|group by S o ,C′ t , then the amount of each student's consumption is obtained, thus obtaining n1 data sets Then y″ is calculated according to S o Aggregate and get the total amount of consumption of each student Sum(C′ m )|group by S o , average consumption amount Avg(C′ m )|group by S o ,Consumption times count(1)|group by S o , select the original one-card flow history data y as C p For the canteen, and according to S oAggregate and calculate the total amount of cafeteria consumption for each student, Sum(C m )|group byS o ; Then we get the Engel coefficient

[0020] The final processed card data set is n2 in total, which is: Among them, A i represents the number of consumptions, B i is the average consumption amount, C i is the total amount of consumption, D i is the total amount of canteen consumption, E i is the Engel coefficient.

[0021] According to the preferred embodiment of the present invention, the basic information includes: student gender information F, student scholarship information G; family situation includes: student family population H, student family per capita income information J; school activity information includes: student aid amount K, number of work-study programs L, student loan amount O, and minimum living allowance line P of the place of origin;

[0022] D y Integrate with the student affairs data to obtain the final dataset d for identifying students with financial difficulties in colleges and universities s ,

[0023] Preferably, according to the present invention, a fully connected graph of student information on each feature is constructed, including:

[0024] For the dataset d s The information of n2 card data is extracted by features, and each student information Info i =(A i ,B i ,C i ,D i ,E i ,F i ,G i ,H i ,J i ,K i ,L i ,O i ,P i ,S oi ) constructs the sub-feature data of each feature; X is used to represent any feature, and each sub-feature data of each student information is That is, the collection of student ID and certain feature data; each student's sub-feature data subInfo X Considered as the fully connected graph of the sub-feature X X A node in the sub-feature dataset subInfos of all studentsX The node set V that constitutes the fully connected graph of feature X X ;

[0025] In feature X, traverse all nodes i, and for any node j that is not i itself, calculate the absolute value diff of the difference between the eigenvalues ​​of nodes i and j X (i,j)=|X i -X j |, and diff X (i, j) is used as the edge weight between node i and node j, and an undirected edge E is added to the fully connected graph that does not repeat the previous edge. x (i,j), represents the edge between nodes i and j, the set of all edges E x As feature X connected graph subgraph X The edge set of

[0026] Complete the above construction of subgraph for each sub-feature X After the operation, we get the feature fully connected graph set subgraphs = (subgraph A ,subgraph B ,...,subgraph P ).

[0027] Preferably, according to the present invention, a frustrated random walk algorithm is executed on the constructed fully connected graph to obtain a random walk result; comprising:

[0028] Sampling a small amount of sample data Sample that has been marked as students with financial difficulties, construct sub-feature data Sample:subInfo X , respectively as each sub-feature fully connected graph subgraph X The end point of the random walk;

[0029] Select Sample:subInfo X As a subgraph x The walking target of all nodes in the random walk is the end point of the fully connected graph node. Sample:subInfo X Hitting time value; Hitting time is: from subgraph X Any point in the first walk to Sample:subInfo X The number of walks required;

[0030] In the frustrated random walk algorithm, the nodes and the transition probability are optimized, that is, in a transition from node i to node j, the degree of a node is defined as Degree i = ∑ k wik , where k is a node adjacent to node i, and w ik is the weight of the edge connecting two nodes i and k, and the transition probability P in the frustrated random walk algorithm transaction is the probability P of i transferring to j i→j The probability P that j accepts i's walk j→i The product of P transaction =P i→j ×P j→i ; W ij is the weight of the edge connecting two nodes i and j, degree i and degree j refer to the degree of node i and node j respectively;

[0031] Using the transfer probability of the frustrated random walk algorithm, we calculate the transfer probability of each node to all other nodes. Assuming that there are M points in the graph, we get an M×M transfer matrix B, B ij represents the frustrated walk probability P of transferring from node i to node j transaction , that is, the probability P of i transferring to j i→j The probability P that j accepts i's walk j→i The product of; let N(t,s) be the first hit time from node s to t, t is Sample:subInfo X , then Nt is the total vector of the first hit time of all nodes transferred to t:

[0032] use Directly find N t The exact result is the random walk result on the graph; after executing the random walk on each subgraph, the Hitting time result list HTlist = (N At ,N Bt ,...,N Pt ).

[0033] Preferably, according to the present invention, the weight of the characteristic index is calculated using the random walk result, including:

[0034] Using HTlist=(N At ,N Bt ,...,N Pt ) Calculate the Hitting time variance V of all nodes a ,V b ...V p ,in, Perform normalization on it The weight of each feature is obtained, that is, the weight vector of the feature index I = (w A ,wB ,w C ,w D ,w E ,w F ,w G ,w H ,w J ,w K ,w L ,w o ,w P ).

[0035] Preferably, according to the present invention, the weights of the feature indicators are subjected to feature optimization, including:

[0036] First, sort I in descending order to obtain the corresponding relationship vector I between I′ and the sorted vector subscript and the original vector. d , I d The columns representing I′ and the source dataset d s The correspondence between the columns;

[0037] Define the sum of the first j elements of I′ as: Let j take values ​​from 0 to 12 in a loop, if both It is believed that w j is an irrelevant feature; since the elements after j are larger than w j The weight of is lower than that of , so the weight vector with the subscript interval [j,12] is irrelevant features, so we jump out of the loop and select the 0th to j-1th bits of I′, thus obtaining the feature vector I″, which is a vector composed of the first j-1 elements of I′, representing the dimension related to students with financial difficulties;

[0038] Find the important and unimportant elements in I″. In I″, find the slope of any two adjacent points. The position with the largest slope, that is, the fastest decline, is marked as Index. p ;

[0039] From vector I d In the example, find the value corresponding to the element from 0 to j-1, which represents the selection of the source data set d s Which columns of elements in the new data set R are composed of? Let R′=(I″ T ×R T ) T , then R′ is the new matrix after the weights are applied, and the matrix has j columns, where Index p It is the critical point between important elements and unimportant elements, that is, 0 to Index p -1 columns are important elements, Index p j-1 is a non-important element. After being compressed by the LDA model, the non-important elements become a column, and this column is placed in the Index. p-1 column, the new dimension is Index p The matrix R″ is clustered using FCM to obtain the final result, that is, a classification result vector. Each student is labeled with a digital label representing the cluster to which the student belongs. If k clusters are finally divided, the value range of the label is 0 to k-1. The classification result vector is automatically matched with the real label of each student to obtain the clustering accuracy.

[0040] Further preferably, the classification result vector is automatically matched with the real label of each student, and the implementation method is:

[0041] First, set up a map to store cluster labels and the number of members in the label, denoted as cmap. cmap includes two values, key and value. The key represents the cluster label and the value represents the number of members of the label. Let the classification result vector generated by FCM clustering be Vec, that is, Vec = FCM (R ″), and the length of Vec be vl, that is, vl = range (Vec). Loop through Vec, let i be the current cursor, and determine whether Vec [i] exists in cmap. If not, it means that the category appears for the first time and the number of members is 1, then execute cmap. .put(Vec[i],1), cmap.put() means putting the number of members owned by the cluster label into cmap; otherwise, get the number of members of the category and add 1 to it; let temp = cmap.get(Vec[i]), cmap.get means getting the number of members corresponding to Vec[i], execute cmap.put(Vec[i],temp++), temp represents the number of members of the cluster, temp++ means adding 1 to the number of members, after the loop is executed, cmap stores the number of members in each cluster;

[0042] Then, calculate the proportion of the number of each type of cluster members to the total number; the method is: first convert cmap into an iterator Ite, set the cursor j and initialize it to 0, set the vector VecCluster to store the class cluster objects; if Ite still has a value, then key j =Ite.next(),key j Represents the cluster label represented by the jth element in the set. Ite.next() means taking out the next element; using tempc = cmap.get(key j ), get the number of members corresponding to the cluster label, that is, key j The number of students with the label tempc, then let rate j is the proportion of the cluster label members represented by the jth element. Since vl is the total number of students, the proportion is Set the cluster object cluster j =(key j ,rate j ), cluster the object j Add VecCluster and set j to j+1 until the loop ends; sort the objects in VecCluster by rate value;

[0043] In the database, the categories of students are sorted according to their proportions, namely:

[0044] VecClusterReal=clusterkey,count(*) / length|group by clusterkeyorderbycount(*);

[0045] Among them, group by is a group aggregation operation, which groups students according to clusterkey, that is, the cluster to which they belong. count(*)

[0046] represents the number of members in the group, orderby is to sort by the number of members, and | represents taking the result on its left and putting it into VecClusterReal. Given the cluster of students with financial difficulties and the location of the cluster in VecClusterReal, the corresponding category in the clustering result vector is: poorkey = VecCluster[cindex].key;

[0047] In the initial Vec, we get students of the poorkey category. Among these students, poorcount students are marked as students with financial difficulties in the database. The accuracy acc is Here, count is the number of students marked as having financial difficulties in the database.

[0048] Further preferably, the formula for obtaining the slope is as shown in formula (I):

[0049]

[0050] In formula (I), the importance values ​​of each feature obtained by the frustrated random walk algorithm are mapped to the coordinate axis. Since the feature importance value is a vector and has been sorted in reverse order, for the jth value in the vector, the horizontal axis is set to x j , the vertical coordinate is y j , and since the ordinate represents the feature weight, y j =w j , and the horizontal axis represents the subscript of the vector, that is, x j =j, then the slope of point j is Tdj is the difference between the ordinate of point j and its previous point divided by the difference of the abscissa, that is,

[0051] A system for identifying college students with financial difficulties based on frustrated random walk and feature weighted clustering, including:

[0052] The one-card transaction history data acquisition and preprocessing module is configured to: acquire the one-card transaction history data and perform preprocessing;

[0053] The student affairs data acquisition and integration module is configured to: acquire student affairs data and integrate it with the pre-processed one-card flow history data to obtain a data set that is ultimately used to identify college students with financial difficulties;

[0054] The fully connected graph construction module is configured to: construct a fully connected graph of student information on each feature;

[0055] The frustrated random walk algorithm execution module is configured to: execute the frustrated random walk algorithm on the constructed fully connected graph to obtain a random walk result;

[0056] The weight calculation module of the characteristic index is configured to: calculate the weight of the characteristic index using the random walk result;

[0057] The feature optimization module is configured to: perform feature optimization on the weights of feature indicators to identify students with financial difficulties in colleges and universities.

[0058] The beneficial effects of the present invention are:

[0059] 1. The present invention solves the differences in data of various schools and the diversity of consumption data of one-card passes in different canteens of different universities through a one-card data calculation model, and generates universal feature indicators; the present invention uses a frustrated random walk algorithm to replace a simple random walk algorithm, and calculates the influence weight of each feature on the identification of students with financial difficulties in different universities, at different times, and in different scenarios, and solves the problem that a simple random walk algorithm cannot find the key graph node features that affect the current features from the feature association graph, thereby comprehensively improving the feature weight calculation performance.

[0060] 2. The present invention uses a feature optimization module to screen and compress features, thereby solving the noise impact of irrelevant features and features with low influence on clustering, and also solving the problem of poor clustering results of high-dimensional data.

[0061] 3. The present invention can adapt to the influence of changes in the influence of features in different periods on clustering by adjusting the feature weights, without the need to retrain parameters like classification algorithms, thereby ensuring business continuity. In short, the invention improves the data versatility, method scalability, business continuity and clustering accuracy in the process of identifying students with financial difficulties. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 The whole framework diagram of the method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering of the present invention;

[0063] Figure 2 Schematic diagram for the execution of the frustrated random walk algorithm;

[0064] Figure 3 It is a schematic diagram for feature optimization of the weights of feature indicators;

[0065] Figure 4 This is a comparison chart of the effects of the method of the present invention and the traditional clustering method for identifying students with financial difficulties. DETAILED DESCRIPTION

[0066] The present invention is further defined below in conjunction with the accompanying drawings and embodiments.

[0067] Example 1

[0068] A method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering includes the following steps:

[0069] Obtain the historical data of the card transaction and perform preprocessing;

[0070] The student work data is obtained and integrated with the pre-processed one-card flow history data to obtain the final data set used to identify students with financial difficulties in colleges and universities; the student work data includes students' basic information, family situation and school activity information;

[0071] Construct a fully connected graph of student information on each feature;

[0072] Execute the frustrated random walk algorithm on the constructed fully connected graph to obtain the random walk result;

[0073] The weights of characteristic indicators are calculated using random walk results;

[0074] The weights of characteristic indicators are optimized to identify students with financial difficulties in colleges and universities.

[0075] Example 2

[0076] The difference between the method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering described in Example 1 is that:

[0077] Obtain the historical data of the card transaction and perform preprocessing, including:

[0078] At the beginning of each month, the card transaction history data of all students in the previous month is collected from the card system, including: o , Consumption time C t 、Consumption location C p 、Consumption amount C m Four attributes, let the card transaction history data be y, and there are n data in total, then

[0079] Since there are two consumption modes of the one-card system, one is to swipe the card once at the total settlement counter, and the other is to swipe the card multiple times at different windows for one consumption, it is necessary to solve the problem of calculating the same consumption behavior under different consumption modes. t Processing, converting the specific time point into the consumption date + consumption interval mode, let sub(C t ) is C t The specific conversion rules are as follows: C′ t Date+sub(C t ); For example, "2022-08-05 08:00" is converted to "2022-08-05 AM", and the historical data of the card flow is converted to

[0080] According to student number S o , C′ t Aggregate y′ and calculate Sum(C m )|group by S o ,C′ t , then the amount of each student's consumption is obtained, thus obtaining n1 data sets Then y″ is calculated according to S o Aggregate and get the total amount of consumption of each student Sum(C′ m )|group by S o , average consumption amount avg(C′ m )|group by S o ,Consumption times count(1)|group by S o , select the original one-card flow history data y as C p For the canteen, and according to S o Aggregate and calculate the total amount of cafeteria consumption for each student, Sum(C m )|group byS o ; Then we get the Engel coefficient

[0081] The final processed card data set is n2 in total, which is: Among them, A i represents the number of consumptions, B i is the average consumption amount, C i is the total amount of consumption, D i is the total amount of canteen consumption, E i is the Engel coefficient.

[0082] Obtain the basic information of students from the student affairs system, including: student gender information F, student scholarship information G; family situation includes: number of family members H, student family per capita income information J; school activity information includes: scholarship amount K, number of work-study sessions L, student loan amount O, and minimum living allowance P for the place of origin;

[0083] D y Integrate with the student affairs data to obtain the final dataset d for identifying students with financial difficulties in colleges and universities s ,

[0084] Construct a fully connected graph of student information on each feature, including:

[0085] It is necessary to select the range of students to be identified as students with financial difficulties from the n2 data sets, such as all students, students of a certain grade, students of a certain campus, or students of a certain college, to obtain a subset of the n2 data sets, a total of n3, and then sample from the subset to construct a full connectivity graph, where the ratio of students with financial difficulties to students without financial difficulties in the sample is 1:4. If n3 is greater than 2000, only 2000 are sampled as samples, and if it is less than or equal to 2000, all data are used as samples.

[0086] For the dataset d s The information of n2 card data is extracted by features, and each student information Construct the sub-feature data of each feature; use X to represent any feature, and each sub-feature data of each student information is That is, the collection of student ID and certain feature data; each student's sub-feature data subInfo X Considered as the fully connected graph of the sub-feature X X A node in the sub-feature dataset subInfos of all students X The node set V that constitutes the fully connected graph of feature X X ;

[0087] In feature X, traverse all nodes i, and for any node j that is not i itself, calculate the absolute value diff of the difference between the eigenvalues ​​of nodes i and j X (i,j)=|X i -X j |, and diff X (i, j) is used as the edge weight between node i and node j, and an undirected edge E is added to the fully connected graph that does not repeat the previous edge. x (i,j), represents the edge between nodes i and j, the set of all edges E x As feature X connected graph subgraph X The edge set of

[0088] Complete the above construction of subgraph for each sub-feature X After the operation, we get the feature fully connected graph set subgraphs = (subgraph A ,subgraph B ,...,subgraph P ).

[0089] Execute the frustrated random walk algorithm on the constructed fully connected graph to obtain the random walk results; including:

[0090] In the student work system, a small amount of sample data Sample (from the data set d s Extracted from the source), construct sub-feature data Sample:subInfo X , respectively as each sub-feature fully connected graph subgraph X The end point of the random walk;

[0091] Select Sample:subInfo X As a subgraph x The walking target of all nodes in the random walk is the end point of the fully connected graph node. Sample:subInfo X Hitting time value; Sample: subInfo X That is, a data set consisting of a small number of sub-features X that have been marked as students with financial difficulties. Hitting time means: Assuming that a person P starts from any node and walks to its neighboring nodes with a certain probability each time, after a long enough time, P can always reach any target node connected to its starting point. The number of transfers required for P to reach the specified target node for the first time is defined as Hitting time; in this situation, Hitting time is: from subgraph X Any point in the first walk to Sample:subInfoX The number of walks required; Hitting time is a random variable. By calculating its probability distribution, we can get the expected value and variance of the Hitting time for each node to reach the target point. For a given connected graph, the Hitting time only depends on the structural relationship between the initial point S and the target point T of the random walk.

[0092] In the frustrated random walk algorithm, in order to solve the problem that only the follower nodes on the graph but not the key nodes can be found in the simple random walk, the nodes and the transition probabilities are optimized, that is, in a transition from node i to node j, the degree of a node is defined as Degreei = ∑ k w ik , where k is a node adjacent to node i, and w ik is the weight of the edge connecting two nodes i and k, and the transition probability P in the frustrated random walk algorithm transaction is the probability P of i transferring to j i→j The probability P that j accepts i's walk j→i The product of P transaction =P i→j ×P j→i ; W ij is the weight of the edge connecting two nodes i and j, degreei and degreej refer to the degree of node i and node j respectively; this mechanism can increase the transfer probability of nodes with large weight but slightly far distance on the graph, increase the difficulty of transferring to nodes with close distance on the graph but small weight, and complete the frustrated design.

[0093] Using the transfer probability of the frustrated random walk algorithm, we calculate the transfer probability of each node to all other nodes. Assuming that there are M points in the graph, we get an M×M transfer matrix B, B ij represents the frustrated walk probability P of transferring from node i to node j transaction , that is, the probability P of i transferring to j i→j The probability P that j accepts i's walk j→i The product of (B ii represents the probability of staying at this node without transferring); N(t,s) is the first hit time of transferring from node s to t, t is Sample:subInfo X , then Nt is the total vector of the first hit time of all nodes transferred to t:

[0094] use Directly find N t The exact result is the random walk result on the graph; after executing the random walk on each subgraph, the Hitting time result list HTlist = (NAt ,N Bt ,...,N Pt ).

[0095] The random walk results are used to calculate the weights of the characteristic indicators, including:

[0096] According to the properties of random walk, subgraph X A node in the sample arrives at the node Sample:subInfo X The lower the Hittingtime value, the higher the similarity between the node and the end point of the walk. The higher the importance of a feature, the greater the discrimination of the feature. The variance of the similarity between the node and the end point of the walk can just represent the discrimination of the feature. At ,N Bt ,...,N Pt ) Calculate the Hitting time variance V of all nodes a ,V b ...V p ,in, Perform normalization on it The weight of each feature is obtained, that is, the weight vector of the feature index I = (w A ,w B ,w C ,w D ,w E ,w F ,w G ,w H ,w J ,w K ,w L ,w O ,w P ).

[0097] The weights of feature indicators are optimized, including:

[0098] First, sort I in descending order to obtain the corresponding relationship vector I between I′ and the sorted vector subscript and the original vector. d , I d The columns representing I′ and the source dataset d s The correspondence between the columns;

[0099] Define the sum of the first j elements of I′ as: Let j take values ​​from 0 to 12 in a loop, if both It is believed that w j is an irrelevant feature; where α means: the sum of the weights of the first j elements needs to exceed a certain threshold, otherwise even if w jSmall, can not be considered as irrelevant features, to avoid the feature is too many, and the feature weight is relatively average resulting in w j is misidentified as an irrelevant feature; and β means: when the sum of the previous feature weights exceeds the threshold, w j It must be less than a certain threshold to be considered as an irrelevant feature. The setting of β is related to the feature dimension. The more dimensions, the smaller β. Users can set the values ​​of α and β, or use the default values. The value of β is related to the dimension. If the dimension is high, the value of β is small. If there is W d dimensions, the average value is β defaults to half the mean, which is α can be set to 0.9 by default. The meaning of using these two rules is that for the descending weight vector, if the weight before j has exceeded α, that is, the proportion is high, and the current weight of j is small, since the elements after j are larger than w j The weight of is lower than that of , so the weight vector with the subscript interval [j,12] is irrelevant features, so we jump out of the loop and select the 0th to j-1th bits of I′, thus obtaining the feature vector I″, which is a vector composed of the first j-1 elements of I′, representing the dimension related to students with financial difficulties;

[0100] Find the important and unimportant elements in I″. Since the weights of the elements in I″ are arranged in descending order, and there is a sharp drop from important elements to unimportant elements, that is, the slope between two points will increase significantly, so we only need to find the slope of any two adjacent points in I″, and the position with the largest slope, that is, the fastest drop, is marked as Index. p ; Used to distinguish the importance of elements;

[0101] From vector I d In the example, find the value corresponding to the element from 0 to j-1, which represents the selection of the source data set d s Which columns of elements in the new data set R are composed of? Let R′=(I″ T ×R T ) T , then R′ is the new matrix after the weights are applied, and the matrix has j columns, where Index p It is the critical point between important elements and unimportant elements, that is, 0 to Index p -1 columns are important elements, Index p j-1 is a non-important element. After being compressed by the LDA model, the non-important elements become a column, and this column is placed in the Index. p -1 column, the new dimension is Index pMatrix R″; this not only reduces the dimension of the matrix, retains the relevant features, but also distinguishes the importance of the features. The matrix R″ is clustered using FCM. FCM is a clustering algorithm based on fuzzy theory proposed by JcBezdek, which is mainly used for data clustering analysis. The process of clustering using FCM is to solve the cluster centers of each category after initializing the parameters, and then iteratively update the cluster centers and the membership matrix to optimize the objective function, and finally divide the data according to the membership matrix. The clustering result is the membership of each data point to the cluster center, expressed as a numerical value. The final result of clustering using FCM is a classification result vector, which gives each student a digital label representing the cluster to which the student belongs. If it is finally divided into k clusters, the label value range is 0 to k-1; the classification result vector is automatically matched with the true label of each student to obtain the accuracy of clustering. Since the clustering process is an unsupervised learning algorithm, the labels obtained by clustering can only represent that students with the same label belong to the same category, but have no actual meaning. Therefore, the value of the clustering label may be different each time. For example, in the first run, students A and B are clustered into one category and labeled as 0. In the next run, they may be labeled as 1. The students’ actual economic difficulties are labels with real meanings provided by the student affairs department. If the clustering is not improved, it must be manually identified, which will cause great trouble to subsequent work.

[0102] Check the number of students marked as students with financial difficulties in the final results, and compare them with the total number to get the clustering accuracy. Use kmeans, FCM, simple random walk + FCM, frustrated random walk + FCM, frustrated random walk + LDA + FCM to test on the October, November, and December data sets. The experiment shows that frustrated random walk + LDA + FCM (the algorithm of the present invention) has the highest accuracy.

[0103] The present invention also makes improvements in the comparison between cluster labels and real labels, and realizes automatic matching. Since the cluster label is a random value from 0 to k-1, the value assigned to a certain cluster each time may be different, resulting in the need for manual judgment. The algorithm calculates the proportion of each label and compares it with the proportion of the real label to realize the automatic correspondence between cluster labels and real labels and automatically calculate the clustering accuracy. The classification result vector is automatically matched with the real label of each student. The implementation method is:

[0104] First, set up a map to store cluster labels and the number of members in the labels, denoted as cmap. cmap includes two values, key and value. The key represents the cluster label and the value represents the number of members of the label. cmap contains two common operations. One is cmap.put(key,value), which puts the number of members owned by the cluster label into cmap. The other operation is cmap.get(key), which means to get the number of members corresponding to the label key. Suppose the classification result vector generated by FCM clustering is Vec, that is, Vec = FCM(R″), the length of Vec is vl, that is, vl = range(Vec), loop over Vec, let i be the current cursor, and determine whether Vec[i] exists in cmap. If not, it means that the category appears for the first time and the number of members is 1, then execute cmap.put(Vec[i],1), cmap.put() means putting the number of members owned by the clustering label into cmap; otherwise, get the number of members of the category and add 1 to it; let temp = cmap.get(Vec[i]), cmap.get means getting the number of members corresponding to Vec[i], execute cmap.Eut(Vec[i],temp++), temp represents the number of members of the cluster, and temp++ means adding 1 to the number of members. After the loop is completed, cmap stores the number of members in each cluster;

[0105] Then, calculate the proportion of the number of each type of cluster members to the total number; the method is: first convert cmap into an iterator Ite. Ite is a collection that stores data. There are two common operations: Ite.hasnext() is used to determine whether the iterator has the next element, and Ite.next() is to take out the next element, set the cursor j, and initialize it to 0, set the vector VecCluster to store the class cluster object; if Ite still has a value, then key j =Ite.next(),key j Represents the cluster label represented by the jth element in the set. Ite.next() means taking out the next element; using tempc = cmap.get(key j ), get the number of members corresponding to the cluster label, that is, key j The number of students with the label tempc, then let rate j is the proportion of the cluster label members represented by the jth element. Since vl is the total number of students, the proportion is Set the cluster object cluster j =(key j ,rate j ), cluster the object jAdd VecCluster and set j to j+1 until the loop ends; sort the objects in VecCluster by rate value;

[0106] In the database, the categories of students are sorted according to their proportions, namely:

[0107] VecClusterReal=clusterkey,count(*) / length|group by clusterkeyorderbycount(*);

[0108] Among them, group by is a grouping aggregation operation, which groups students according to clusterkey, that is, the cluster to which they belong. Count(*) represents the number of members in the group. Order by is sorting according to the number of members, and | represents taking the result on its left and putting it into VecClusterReal. Given the cluster to which students with financial difficulties belong and the location of the cluster in VecClusterReal, cindex, the corresponding category in the clustering result vector is: Eoorkey = VecCluster[cindex].key;

[0109] In the initial Vec, we get students of the poorkey category. Among these students, poorcount students are marked as students with financial difficulties in the database. The accuracy acc is Among them, count is the number of students marked as having financial difficulties in the database. For students marked as poorkey in Vec, but not on the list of students with financial difficulties given by the student affairs department, they can be listed as "invisible aid" pending review list to provide a basis for funding decisions, which is also one of the advantages of this invention. In addition, this invention calculates the proportion of each cluster instead of the number, which can also intuitively show the difference between the distribution of clusters and the actual distribution.

[0110] The formula for obtaining the slope is shown in formula (I):

[0111]

[0112] In formula (I), the importance values ​​of each feature obtained by the frustrated random walk algorithm are mapped to the coordinate axis. Since the feature importance value is a vector and has been sorted in reverse order, for the jth value in the vector, the horizontal axis is set to x j , the vertical coordinate is y j , and since the ordinate represents the feature weight, y j =w j , and the horizontal axis represents the subscript of the vector, that is, x j =j, then the slope of point j is Td jis the difference between the ordinate of point j and its previous point divided by the difference of the abscissa, that is,

[0113] Example 3

[0114] The difference between the method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering described in Example 2 and that described in Example 1 is that:

[0115] like Figure 1 As shown in the figure, firstly, the card flow data is processed, and the consumption records in the same period are merged through the data integration module of the same period, and a consumption record is calculated, thereby solving the difference between one consumption and multiple consumption, and generating a universal flow y′. According to the universal flow y′, the total consumption amount, average consumption amount, and consumption times can be calculated, and the total consumption amount in the cafeteria can be calculated through the original flow data y, and then the Engel coefficient is calculated according to the total consumption amount and the total consumption amount in the cafeteria, thus forming 5 card data feature dimensions d y .

[0116] D y The original data is generated by associating the student data with the student ID. s . s The weight vector I of each feature is generated by random walk, such as Figure 2 As shown in the figure, the blue nodes in the single-feature fully connected graph are the end points of the random walk. I is sorted in descending order according to the weight to generate I′, and d s and I′ feature correspondence I d ,like Figure 3 shown.

[0117] According to the rules of feature optimization module, if Discard irrelevant dimensions after j. The values ​​of α and β can be set by the user. If not set, α defaults to 0.9 and β defaults to W d is the number of dimensions, which is 13, then β is 0.038, and the first j-1 dimensions are intercepted to generate I″. d Find the first j-1 elements in order, starting from d s Find the relevant columns in the dataset to form a new dataset R. Perform a dot multiplication of R and I″ to get R′. From I″, find the position where the gradient descent is fastest according to the slope of the element. p , Index p The following elements are compressed by the LDA model and are consistent with the previous index of R′. p The columns are fused to obtain a new matrix R″. FCM is used to cluster R″ to obtain the final result, that is, the category of each student.

[0118] Check the students marked with the label of students with financial difficulties, and divide the number of students with financial difficulties into the total number to get the clustering accuracy. Taking the data from October 2021 as an example, two common clustering methods, kmeans and FCM, were used to perform clustering on the original data set in 13 dimensions, with accuracies of 0.627 and 0.674 respectively, indicating that the FCM algorithm is superior to kmeans in the data set for identifying students with financial difficulties. Then, a simple random walk algorithm was used to calculate the weights of relevant features, filter out irrelevant feature dimensions, and then the weights were applied to the data set for clustering. The accuracy was increased to 0.721, and when frustrated random walk was used for weighted clustering, the accuracy was further improved to 0.787. Finally, after the data set was weighted by frustrated random walk, irrelevant feature dimensions were filtered out, and the weights were applied to the data set. After the weights with a lower proportion were compressed, the accuracy was the highest, which was 0.801. In order to prove the usability of the algorithm,

[0119] The present invention uses a frustrated random walk and feature-weighted clustering algorithm, which has three advantages over the classification algorithm. First, it solves the problem in the classification algorithm that some students from poor families do not apply for financial aid due to their strong self-esteem and are labeled as ordinary students; second, while clustering the data marked as students with financial difficulties into the same cluster, the ordinary students in the cluster are also marked as students to be identified, providing a basis for "invisible aid"; third: when the influencing factors of each feature change, there is no need to re-train, thereby ensuring business continuity.

[0120] Example 4

[0121] A system for identifying college students with financial difficulties based on frustrated random walk and feature weighted clustering, including:

[0122] The one-card transaction history data acquisition and preprocessing module is configured to: acquire the one-card transaction history data and perform preprocessing;

[0123] The student affairs data acquisition and integration module is configured to: acquire student affairs data and integrate it with the pre-processed one-card flow history data to obtain a data set that is ultimately used to identify college students with financial difficulties;

[0124] The fully connected graph construction module is configured to: construct a fully connected graph of student information on each feature;

[0125] The frustrated random walk algorithm execution module is configured to: execute the frustrated random walk algorithm on the constructed fully connected graph to obtain a random walk result;

[0126] The weight calculation module of the characteristic index is configured to: calculate the weight of the characteristic index using the random walk result;

[0127] The feature optimization module is configured to: perform feature optimization on the weights of feature indicators to identify students with financial difficulties in colleges and universities.

Claims

1. A method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering, characterized in that: The steps include: Obtain the historical data of the card transaction and perform preprocessing; The student work data is obtained and integrated with the pre-processed one-card flow history data to obtain the final data set used to identify students with financial difficulties in colleges and universities; the student work data includes students' basic information, family situation and school activity information; Construct a fully connected graph of student information on each feature; Execute the frustrated random walk algorithm on the constructed fully connected graph to obtain the random walk result; The weights of characteristic indicators are calculated using random walk results; Perform feature selection on the weights of feature indicators to identify college students with financial difficulties; In the frustrated random walk algorithm, the nodes and the transition probability are optimized, that is, in a transition from node i to node j, the degree of a node is defined as Degree i = ∑ k w ik , where k is a node adjacent to node i, and w ik is the weight of the edge connecting two nodes i and k, and the transition probability P in the frustrated random walk algorithm transaction is the probability P of i transferring to j i→j The probability P that j accepts i's walk j→i The product of W ij is the weight of the edge connecting two nodes i and j. Degree i and degree j refer to the degree of node i and node j respectively.

2. A method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering according to claim 1, characterized in that: Obtain the historical data of the card transaction and perform preprocessing, including: Collect the historical data of the card flow in the previous period, including: student number S o , Consumption time C t 、Consumption location C p 、Consumption amount C m Four attributes, let the card transaction history data be y, and there are n data in total, then Consumption time C of the historical data of the card flow t Processing, converting the specific time point into the consumption date + consumption interval mode, let sub(C t ) is C t The specific conversion rules are as follows: C′ t Date+sub(C t ); Convert the historical data of the card flow into According to student number S o , C′ t Aggregate y′ and calculate Sum(C m )|group dy S o ,C′ t , then we get the consumption amount of each student at one time, and thus get n + Dataset Then y″ is calculated according to S o Aggregate and get the total amount of consumption of each student Sum(C′ m )|group by S o , average consumption amount Avg(C′ m )|group by S o ,Consumption times count(1)|group by S o , select the original one-card flow history data y as C p For the canteen, and according to S o Aggregate and calculate the total amount of cafeteria consumption for each student, Sum(C m )|group byS o ; Then we get the Engel coefficient The final processed card data set is n , Article, for: Among them, A i represents the number of consumptions, B i is the average consumption amount, C i is the total amount of consumption, D i is the total amount of canteen consumption, E i is the Engel coefficient.

3. A method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering according to claim 1, characterized in that: Basic information includes: student gender information F, student scholarship information G; family situation includes: number of family members H, student family per capita income information J; school activity information includes: scholarship amount K, number of work-study programs L, student loan amount O, and minimum living allowance for the student’s place of origin P; D P Integrate with the student affairs data to obtain the final dataset d for identifying students with financial difficulties in colleges and universities s , 4. The method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering according to claim 1, characterized in that: Construct a fully connected graph of student information on each feature, including: For data set d s Medium , The information of each piece of card data is extracted by features. Construct the sub-feature data of each feature; use X to refer to any feature, and each sub-feature data of each student information is That is, the collection of student ID and certain feature data; each student's sub-feature data subInfo X Considered as the fully connected graph of the sub-feature X X A node in the sub-feature dataset subInfos of all students X The node set V that constitutes the fully connected graph of feature X X ; In feature X, traverse all nodes i, and for any node j that is not i itself, calculate the absolute value diff of the difference between the eigenvalues ​​of nodes i and j X (i,j)=|X i -X j |, and diff X (i, j) is used as the edge weight between node i and node j, and an undirected edge E is added to the fully connected graph that does not repeat the previous edge. x (i,j), represents the edge between nodes i and j, the set of all edges E x As feature X connected graph subgraph X The edge set of Complete the above construction of subgraph for each sub-feature X After the operation, we get the feature fully connected graph set subgraphs = (subgraph A ,subgraph B ,...,subgraph P ).

5. The method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering according to claim 1, characterized in that: Execute the frustrated random walk algorithm on the constructed fully connected graph to obtain the random walk results; including: Sampling a small amount of sample data Sample that has been marked as students with financial difficulties, construct sub-feature data Sample:subInfo X , respectively as each sub-feature fully connected graph subgraph X The end point of the random walk; Select Sample:subInfo X As a subgraph X The walking target of all nodes in the random walk is the end point of the fully connected graph node. Sample:subInfo X Hitting time value; Hitting time is: from subgraph X Any point in the first walk to Sample:subInfo X The number of walks required; Using the transfer probability of the frustrated random walk algorithm, we calculate the transfer probability of each node to all other nodes. Assuming that there are M points in the graph, we get an M×M transfer matrix B, B ij represents the frustrated walk probability P of transferring from node i to node j transaction , that is, the probability P of i transferring to j i→j The probability P that j accepts i's walk j→i The product of; let N(t,s) be the first hit time from node s to t, t is Sample:subInfo X , then Nt is the total vector of the first hit time of all nodes transferred to t: use Directly find N t The exact result is the random walk result on the graph; after executing the random walk on each subgraph, the Hiting time result list HTlist = (N At ,N Bt ,...,N Pt ).

6. A method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering according to claim 1, characterized in that: The random walk results are used to calculate the weights of the characteristic indicators, including: Using HTlist=(N At ,N Bt ,...,N Pt ) Calculate the Hitting time variance V of all nodes a ,V b ...V p ,in, Perform normalization on it The weight of each feature is obtained, that is, the weight vector of the feature index I = (w A ,w B ,w C ,w D ,w E ,w F ,w G ,w H ,w J ,w K ,w L ,w O ,w P ).

7. The method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering according to claim 1, characterized in that: Perform feature selection on the weights of feature indicators, including: First, sort I in descending order to obtain the corresponding relationship vector I between I′ and the sorted vector subscript and the original vector. d , I d The columns representing I′ and the source dataset d s The correspondence between the columns; Define the sum of the first j elements of I′ as: Let j take values ​​from 0 to 12 in a loop, if both It is believed that w j is an irrelevant feature; since the elements after j are larger than w j The weight of is lower than that of , so the weight vector with the subscript interval [j,12] is irrelevant features, so we jump out of the loop and select the 0th to j-1th bits of I′, thus obtaining the feature vector I″, which is a vector composed of the first j-1 elements of I′, representing the dimension related to students with financial difficulties; Find the important and unimportant elements in I″. In I″, find the slope of any two adjacent points. The position with the largest slope, that is, the fastest decline, is marked as Index. p ; From vector I d In the example, find the value corresponding to the element from 0 to j-1, which represents the selection of the source data set d s Which columns of elements in the new data set R are composed of? Let R′=(I″ T ×R T ) T , then R′ is the new matrix after the weights are applied, and the matrix has j columns, where Index p It is the critical point between important elements and unimportant elements, that is, 0 to Index p -1 columns are important elements, Index p j-1 is a non-important element. After being compressed by the LDA model, the non-important elements become a column, and this column is placed in the Index. p -1 column, the new dimension is Index p The matrix R″ is clustered using FCM to obtain the final result, that is, a classification result vector. Each student is labeled with a digital label representing the cluster to which the student belongs. If k clusters are finally divided, the value range of the label is 0 to k-1. The classification result vector is automatically matched with the real label of each student to obtain the clustering accuracy.

8. A method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering according to claim 7, characterized in that: The classification result vector is automatically matched with the true label of each student by: First, set up a map to store cluster labels and the number of members in the label, denoted as cmap. cmap includes two values, key and value. The key represents the cluster label and the value represents the number of members of the label. Let the classification result vector generated by FCM clustering be Vec, that is, Vec = FCM (R ″), and the length of Vec be vl, that is, vl = range (Vec). Loop through Vec, let i be the current cursor, and determine whether Vec [i] exists in cmap. If not, it means that the category appears for the first time and the number of members is 1, then execute cmap. .put(Vec[i],1), cmap.put() means putting the number of members owned by the cluster label into cmap; otherwise, get the number of members of the category and add 1 to it; let temp = cmap.get(Vec[i]), cmap.get means getting the number of members corresponding to Vec[i], execute cmap.put(Vec[i],temp++), temp represents the number of members of the cluster, temp++ means adding 1 to the number of members, after the loop is executed, cmap stores the number of members in each cluster; Then, calculate the proportion of the number of each type of cluster members to the total number; the method is: first convert cmap into an iterator Ite, set the cursor j and initialize it to 0, set the vector VecCluster to store the class cluster objects; if Ite still has a value, then key j =Ite.next(),key j Represents the cluster label represented by the jth element in the set. Ite.next() means taking out the next element; using tempc = cmap.get(key j ), get the number of members corresponding to the cluster label, that is, key j The number of students with the label tempc, then let rate j is the proportion of the cluster label members represented by the jth element. Since vI is the total number of students, the proportion is Set the cluster object cluster j =(key j ,rate j ), cluster the object j Add VecCluster and set j to j+1 until the loop ends; sort the objects in VecCluster by rate value; In the database, the categories of students are sorted according to their proportions, namely: VecClusterReal=clusterkey,count(*) / length|group byclusterkey order bycount(*); Among them, group by is a grouping aggregation operation, which groups students according to clusterkey, that is, the cluster to which they belong. Count(*) represents the number of members in the group. Order by is sorting according to the number of members, and | represents taking the result on its left and putting it into VecClusterReal. Given the cluster to which students with financial difficulties belong and the location of the cluster in VecClusterReal, cindex, the corresponding category in the clustering result vector is: poorkey = VecCluster[cindex].key; In the initial Vec, we get students of the porky category, among which poorcount students are marked as students with financial difficulties in the database, so the accuracy acc is Here, count is the number of students marked as having financial difficulties in the database.

9. The method for identifying college students with financial difficulties using frustrated random walk and feature weighted clustering according to claim 7, characterized in that: The formula for obtaining the slope is shown in formula (I): In formula (I), the importance values ​​of each feature obtained by the frustrated random walk algorithm are mapped to the coordinate axis. Since the feature importance value is a vector and has been sorted in reverse order, for the jth value in the vector, the horizontal axis is set to x j , the vertical coordinate is y j , and since the ordinate represents the feature weight, y j =w j , and the horizontal axis represents the subscript of the vector, that is, x j =j, then the slope of point j is Td j is the difference between the ordinate of point j and its previous point divided by the difference of the abscissa, that is, 10. A system for identifying college students with financial difficulties based on frustrated random walk and feature weighted clustering, characterized in that: include: The one-card transaction history data acquisition and preprocessing module is configured to: acquire the one-card transaction history data and perform preprocessing; The student affairs data acquisition and integration module is configured to: acquire student affairs data and integrate it with the pre-processed one-card flow history data to obtain a data set that is ultimately used to identify college students with financial difficulties; The fully connected graph construction module is configured to: construct a fully connected graph of student information on each feature; The frustrated random walk algorithm execution module is configured to: execute the frustrated random walk algorithm on the constructed fully connected graph to obtain a random walk result; The weight calculation module of the characteristic index is configured to: calculate the weight of the characteristic index using the random walk result; The feature optimization module is configured to: perform feature selection on the weights of feature indicators to identify students with financial difficulties in colleges and universities.

Citation Information

Patent Citations

  • Method and device for identifying financial difficulty students on basis of smart card consumption behavior analysis

    CN103632238A

  • Semantically sensitive knowledge graph random walk sampling method

    CN111444317A