A pre-examination community discovery algorithm-based abnormal enrollment group detection method

By constructing a candidate graph structure in the Neo4j database, using community detection algorithms and graph neural networks to generate risk scores, and marking and propagating high-risk labels, the problem of difficulty in identifying pre-exam exam risks in existing technologies is solved, achieving efficient and accurate risk prediction.

CN122114659APending Publication Date: 2026-05-29人力资源和社会保障部人事考试中心
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
人力资源和社会保障部人事考试中心
Filing Date
2026-04-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to identify potential exam risks before the actual exam using big data, resulting in low identification efficiency and accuracy.

Method used

We use the Neo4j database to construct a graph structure among candidates, generate embedded representations of candidates through community detection algorithms and graph neural networks, determine risk scores, and perform high-risk labeling and label propagation to identify potential abnormal registration groups.

Benefits of technology

It improves the efficiency and accuracy of pre-exam risk prediction for test takers, and can identify potential test risk groups in advance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114659A_ABST
    Figure CN122114659A_ABST
Patent Text Reader

Abstract

The application discloses a pre-examination abnormal registration group detection method based on a community discovery algorithm. A graph structure for representing the association relationship between examinees is constructed, then the examinees are divided into communities or non-community individuals based on the graph structure, and then the graph structure is subjected to message passing to generate the embedded representation corresponding to the examinees, and whether there is a risk is predicted according to the embedded representation corresponding to the examinees, so as to determine the confidence degree corresponding to the examinees and the preset label, and then generate a risk score for the examinees. Then, the high-risk examinee characteristics are counted based on the risk score, so that the high-risk label marking and dyeing of the high-risk examinee nodes are performed, and the label propagation of the high-risk label in the graph structure is performed. Through the method, the risk characteristics can be counted, the risk label can be marked and propagated based on the graph structure, so that the examinee group that may have risks can be determined in a big data manner before the examination, and the efficiency and accuracy of the risk prediction of the examinees before the examination are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of pre-exam risk identification technology, and in particular to a method for detecting abnormal registration groups based on a community detection algorithm before exams. Background Technology

[0002] With the continuous advancement of information technology, various examinations currently face certain risks. For example, there is the possibility of organized cheating by certain groups. Therefore, relevant technical personnel are constantly researching how to better identify or mitigate these examination risks.

[0003] In existing technologies, monitoring for risky behavior during examinations is typically done manually. For example, personnel are organized to manage each stage of the examination process.

[0004] In existing technologies, since exam risks are mainly identified through human supervision, it is difficult to identify hidden exam risks in advance using big data, thus reducing the efficiency and accuracy of exam risk identification.

[0005] There is currently no effective solution to the technical problem that the existing technologies mentioned above make it difficult to identify hidden exam risks in advance through big data, thus reducing the efficiency and accuracy of exam risk identification. Summary of the Invention

[0006] The embodiments of this disclosure provide a method for detecting abnormal registration groups based on a community detection algorithm before an exam, which at least solves the technical problem in the prior art that it is difficult to identify hidden exam risks in advance through big data, thereby reducing the efficiency and accuracy of exam risk identification.

[0007] According to one aspect of the present disclosure, a method for detecting abnormal registration groups based on a community detection algorithm before an exam is provided, comprising: constructing a graph structure using a Neo4j database to represent the relationships between candidates; classifying candidates into community or non-community individuals based on the graph structure, wherein nodes in the graph structure represent candidates, edges in the graph structure are weighted undirected edges created based on the similarity of candidate features, and communities are used to indicate potentially organized abnormal registration groups of candidates; generating embedded representations of each node as embedded representations corresponding to the corresponding candidates by message passing through a preset graph neural network based on the graph structure; determining the confidence level of the corresponding candidate and a preset label based on the embedded representation of the corresponding candidate, and generating a risk score for the corresponding candidate based on the confidence level, wherein the preset label is used to indicate that the candidate is at risk; determining high-risk candidate features based on the risk score of the candidate, and marking the nodes corresponding to candidates with high-risk candidate features as high-risk candidate nodes with high-risk labels, and performing a coloring operation in the Neo4j database; and performing label propagation based on the graph structure to determine risk candidate groups associated with high-risk candidate nodes.

[0008] According to another aspect of the present disclosure, a storage medium is also provided, the storage medium including a stored program, wherein, when the program is executed, a processor performs any of the methods described above.

[0009] According to another aspect of the present disclosure, an abnormal registration group detection device based on a community detection algorithm before an exam is also provided, comprising: a construction module, configured to construct a graph structure representing the relationships between candidates using a Neo4j database, and to classify candidates into community or non-community individuals based on the graph structure, wherein nodes in the graph structure refer to candidates, edges in the graph structure are weighted undirected edges created based on the similarity of candidate features, and communities are used to indicate potentially organized abnormal registration groups of candidates; and a message passing module, configured to perform message passing based on the graph structure using a preset graph neural network, generating embedded representations of each node as corresponding to the respective candidates. The system includes an embedded representation; a scoring determination module, used to determine the confidence level of the corresponding candidate and the preset label based on the embedded representation of the candidate, and generate a risk score for the corresponding candidate based on the confidence level, where the preset label is used to indicate that the candidate has a risk; a coloring module, used to determine the characteristics of high-risk candidates based on the risk score of the candidate, and to mark the nodes corresponding to candidates with high-risk candidate characteristics as high-risk candidate nodes and perform coloring operations in the Neo4j database; and a propagation module, used to propagate the label of high-risk label based on the graph structure to determine the risk candidate group associated with the high-risk candidate node.

[0010] According to another aspect of the present disclosure, an abnormal registration group detection device based on a community detection algorithm before an exam is also provided, comprising: a processor; and a memory connected to the processor, for providing the processor with instructions to perform the following processing steps: constructing a graph structure representing the relationships between candidates using a Neo4j database, and classifying candidates into community or non-community individuals based on the graph structure, wherein nodes in the graph structure refer to candidates, edges in the graph structure are weighted undirected edges created based on the similarity of candidate features, and communities are used to indicate potentially organized abnormal registration groups of candidates; and performing message passing based on the graph structure using a preset graph neural network. The process involves generating embedded representations for each node, which serve as the embedded representations corresponding to the respective candidates. Based on these embedded representations, the confidence level of each candidate against a preset label is determined, and a risk score is generated for each candidate based on the confidence level. The preset label indicates that the candidate is at risk. Based on the risk score, high-risk candidate characteristics are identified, and the nodes corresponding to candidates with high-risk characteristics are labeled as high-risk candidate nodes and colored in the Neo4j database. Finally, based on the graph structure, high-risk label propagation is performed to identify the risk candidate groups associated with the high-risk candidate nodes.

[0011] In this embodiment, a graph structure representing the relationships between candidates is first constructed. Based on this graph structure, candidates are then divided into community or non-community individuals. This allows for subsequent risk group analysis based on the graph structure representing candidate relationships, focusing only on candidates within their respective communities. The graph structure is then used for message passing to generate embedded representations for each candidate. Based on these embedded representations, a risk prediction is made, determining the confidence level of each candidate against a preset label, and generating a risk score. High-risk candidate characteristics are then statistically analyzed based on the risk score, and high-risk candidate nodes are labeled and marked with high-risk tags. These high-risk tags are then propagated within the graph structure. This method enables the statistical analysis of risk characteristics and the labeling and propagation of risk tags based on the graph structure. This allows for the identification of potentially risky candidate groups before the exam using big data, improving the efficiency and accuracy of pre-exam risk prediction. Attached Figure Description

[0012] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this application, illustrate exemplary embodiments of this disclosure and are used to explain this disclosure, but do not constitute an undue limitation of this disclosure. In the drawings: Figure 1 This is a hardware structure block diagram of a computing device for implementing the method described in Embodiment 1 of this disclosure; Figure 2This is a flowchart illustrating the method for detecting abnormal registration groups based on a community detection algorithm before an exam, according to the first aspect of Embodiment 1 of this disclosure. Figure 3 This is a schematic diagram of a dyeing process provided in Embodiment 1 of this disclosure; Figure 4 This is a schematic diagram of an abnormal registration group detection device based on a community detection algorithm before an exam, according to the first aspect of Embodiment 2 of this disclosure; and Figure 5 This is a schematic diagram of an abnormal registration group detection device based on a community discovery algorithm before the exam, as described in the first aspect of Embodiment 3 of this disclosure. Detailed Implementation

[0013] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this disclosure.

[0014] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0015] Example 1

[0016] According to this embodiment, an embodiment of an abnormal registration group detection method based on community detection algorithm before the exam is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than that shown here.

[0017] The method embodiments provided in this example can be executed on mobile terminals, computer terminals, servers, or similar computing devices. Figure 1A hardware block diagram of a computing device for implementing a pre-exam anomaly registration group detection method based on a community detection algorithm is shown. Figure 1 As shown, a computing device may include one or more processors (processors may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include a display, keyboard, and cursor control device connected to the input / output interface. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, a computing device may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0018] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element in a computing device. As involved in the embodiments of this disclosure, the data processing circuits serve as processor control (e.g., selection of a variable resistor termination path connected to an interface).

[0019] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the abnormal registration group detection method based on the community detection algorithm before the exam in this embodiment of the disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the above-mentioned method for detecting abnormal registration groups based on the community detection algorithm before the exam. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the computing device via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0020] The transmission device is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the computing device's communication provider. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0021] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows users to interact with the user interface of the computing device.

[0022] It should be noted here that, in some optional embodiments, the above... Figure 1 The computing device shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computing devices.

[0023] Under the aforementioned operating environment, according to the first aspect of this embodiment, a method for detecting abnormal registration groups based on a community detection algorithm before examination is provided. This method can be implemented by... Figure 1 The computing device implementation is shown. Figure 2 A flowchart illustrating the method is shown below. (Refer to...) Figure 2 As shown, the method includes: S202: Construct a graph structure using the Neo4j database to represent the relationships between candidates, and classify candidates into community or non-community individuals based on the graph structure. In the graph structure, nodes refer to candidates, and the edges are weighted undirected edges created based on the similarity of candidate features. The community is used to indicate groups of candidates with potentially organized and abnormal registration. S204: By using a pre-set graph neural network to pass messages based on the graph structure, embedded representations of each node are generated as embedded representations corresponding to the corresponding candidates. S206: Based on the embedded representation of the corresponding candidate, determine the confidence level of the corresponding candidate and the preset label, and generate a risk score for the corresponding candidate based on the confidence level, wherein the preset label is used to indicate that the candidate has a risk; S208: Based on the risk scores of the examinees, identify the characteristics of high-risk examinees, and mark the nodes corresponding to examinees with high-risk characteristics as high-risk examinee nodes, then perform a coloring operation in the Neo4j database; and S210: Based on the graph structure, perform label propagation of high-risk labels to identify risk candidate groups associated with high-risk candidate nodes.

[0024] The computing device can utilize the Neo4j database to construct a graph structure representing the relationships between candidates, and based on this graph structure, classify candidates into community or non-community individuals. In the graph structure, nodes represent candidates, and edges are weighted undirected edges created based on the similarity of candidate features. Communities are used to indicate potentially organized groups of candidates who have registered absurdly (S202).

[0025] Each node possesses several attribute information (i.e., the attribute information of the corresponding candidate), such as the registration IP address, password, mailing address, work unit, and security questions. This attribute information forms the basis for analyzing the relationships between candidates. Connections between nodes, i.e., edges in the graph structure, can be established by finding correlations between candidates based on this attribute information.

[0026] Then, test takers can be divided into community or non-community individuals. This can be achieved using a community partitioning algorithm, such as Louvain's algorithm, to divide the nodes in the graph structure, thus classifying the nodes corresponding to test takers as either nodes in the community or nodes not in the community (i.e., non-community individuals). A community represents a group with pre-exam risk.

[0027] Graph structure construction and community partitioning play three key roles in identifying high-risk candidate groups: First, they provide a more focused path for subsequent in-depth analysis, excluding individual candidate samples outside the group, significantly reducing computational complexity, and thus greatly improving the overall efficiency of the detection process. Second, they can more effectively capture the characteristics of candidates who violate regulations, more accurately identify cheating behavior and its participants, and form a common feature identification of high-risk groups. Third, the application of the Louvain community partitioning algorithm further improves the accuracy of community detection.

[0028] After the graph structure is constructed and the community is divided, the computing device can use a preset graph neural network to pass messages based on the graph structure and generate embedded representations of each node as embedded representations corresponding to the corresponding candidates (S204).

[0029] Then, the computing device can determine the confidence level of the corresponding candidate and the preset label based on the embedded representation of the candidate, and generate a risk score for the corresponding candidate based on the confidence level, where the preset label is used to indicate that the candidate has a risk (S206). How to generate a risk score for the candidate will be described in detail below.

[0030] The computing device can determine the characteristics of high-risk candidates based on the risk scores of the candidates, and mark the nodes corresponding to candidates with high-risk candidate characteristics as high-risk candidate nodes, and perform coloring operations in the Neo4j database (S208).

[0031] The high-risk candidate characteristics mentioned here refer to the common features of candidates identified through statistical methods as having a relatively high risk profile. For example, if an IP address is associated with multiple high-risk candidates, then that IP address may be a centralized registration point for a risky organization related to the exam.

[0032] Therefore, by identifying the characteristics of high-risk candidates, the nodes corresponding to candidates with risks (high-risk candidate nodes) can be determined. These nodes can then be marked, and coloring operations can be performed in the Neo4j database to clearly and obviously visualize the candidates with risks.

[0033] Since risky behaviors related to exams in reality are usually organized group behaviors, after identifying high-risk candidate nodes, the associated groups can be determined based on these nodes in a graph structure. Therefore, the computing device can perform high-risk label propagation based on the graph structure to identify the risk candidate groups associated with the high-risk candidate nodes (S210). The specific method of high-risk label propagation will be described below.

[0034] As described in the background section, existing technologies typically rely on human intervention to monitor for risky behaviors during examinations. For example, this involves organizing personnel to manage various stages of the examination process. Because current technologies primarily identify examination risks through human oversight, it is difficult to identify hidden risks in advance using big data, thus reducing the efficiency and accuracy of risk identification.

[0035] In view of this, according to the technical solution of this embodiment, a graph structure representing the relationships between candidates is first constructed. Then, based on this graph structure, candidates are divided into community or non-community individuals. This allows for subsequent risk group analysis based on the graph structure representing candidate relationships, focusing only on candidates classified as belonging to a community. Next, message passing is performed on the graph structure to generate embedded representations corresponding to candidates. Based on these embedded representations, a prediction of risk is made, determining the confidence level of each candidate against a preset label (i.e., the label corresponding to risk), and generating a risk score. Then, based on the risk score, characteristics of high-risk candidates are statistically analyzed, and high-risk candidate nodes are labeled and colored with high-risk tags. High-risk tags are then propagated within the graph structure. This method enables the statistical analysis of risk characteristics and the labeling and propagation of risk tags based on the graph structure, thereby identifying potentially risky candidate groups before the exam using big data, improving the efficiency and accuracy of pre-exam risk prediction.

[0036] Optionally, the operation of constructing a graph structure representing the relationships between candidates using the Neo4j database specifically includes: determining the candidate's attribute information, including registration IP address, password, mailing address, work unit, and security questions; constructing the nodes corresponding to the candidates; for two candidates, determining the common attribute information between them, and determining the preset weights corresponding to the common attribute information; determining the edge weights between the two candidates based on the preset weights corresponding to the common attribute information; and constructing the edges between the two candidates based on the edge weights, thus obtaining the graph structure.

[0037] In other words, in a graph structure, the edge weight between any two nodes can be determined by the shared attribute information and corresponding weights among the corresponding candidates. The edge weight can be determined by the following formula:

[0038] Among them, W ij Let w be the edge weight between the i-th node and the j-th node, and k be the number of attribute information. k The preset weight for the k-th attribute can be set based on experience. For example, passwords and security questions have relatively high weights, while IP addresses have relatively low weights. This is because, under normal circumstances, it's almost impossible for candidates to have the same password or security question. However, due to the widespread use of NAT technology, people in the same school or organization may share a single IP address. k (i, j) is an indicator function used to represent the condition where the k-th attribute information of the i-th node and the j-th node are the same. k (i, j) is 1, otherwise it is 0.

[0039] Of course, the computing device can first construct a multigraph, with each graph corresponding to a type of attribute information. When two nodes share the same attribute information, an edge can be constructed between them. Then, after weighted summation using the above method, the original graph is transformed from a multigraph into a simple graph, where nodes are connected by at most one edge. This edge represents the overall similarity between the two candidates; the more shared features, the greater the edge weight, meaning a closer connection between the two candidates. The construction of the simple graph also facilitates subsequent community partitioning algorithms, reducing complexity. For example, if two candidates have used the same IP address and the same communication address, two edges will be constructed between the two nodes. After constructing the relationship edges for various features, the edges between all candidate nodes are weighted and summarized to form a total relationship edge. Secondly, regarding community partitioning, the Louvain community partitioning algorithm automatically identifies closely related subgraphs, i.e., communities, by calculating the comprehensive similarity between each pair of candidates. The Louvain algorithm is a greedy algorithm based on maximizing modularity. For a given modular partitioning scheme in a network, its modularity is defined as:

[0040] Among them, e c Let represent the sum of weights of community c, and m represent the sum of weights of all edges in the graph. This represents the sum of the weights of the edges connecting to nodes within community c. The modularity ranges from -1 / 2 to 1. A higher modularity indicates a more reasonable community division, while a negative modularity indicates a low degree of rationality in the current community division, suggesting that each node should be considered as a separate community.

[0041] For nodes in the network, by trying to place them into different candidate groups and evaluating the improvement in overall network modularity, the most ideal community for each node can be found. After multiple iterations, if the overall modularity no longer improves, the current partitioning is output as the community discovery result. After community partitioning, the modularity is as high as 0.7, indicating that there is a very high probability of high-risk groups among the test takers.

[0042] To further identify high-risk candidates within the defined communities, nodes not belonging to a community can be considered as nodes unrelated to other nodes. Alternatively, during subsequent message passing via graph neural networks, feature extraction can be performed only on nodes within the communities, thus allowing for risk prediction only on those nodes.

[0043] Optionally, the operation of generating embedded representations of each node by passing messages based on the graph structure using a preset graph neural network specifically includes: dividing the candidate's relevant information into different intervals according to the degree of repetition of the candidate's features and assigning corresponding encoding values; passing messages on the graph structure using a semi-supervised graph neural network with non-random missing values, and generating embedded representations of each node by combining node features and the topological information of the graph structure.

[0044] The goal of this stage is to further consolidate the characteristics of high-risk candidates within the communities segmented in the previous step. Since it's impossible to capture all violations during actual invigilation, whether a candidate has violated regulations should be considered a missing label field. This method leverages the advantages of graph neural network models, combining node features with graph structure information for deep analysis to handle complex relationships between candidates. Specifically, this method employs a semi-supervised graph neural network (Graph-based joint model with Nonignorable Missingness, GNM). In the GNM model, for each observed node, the GNM weights it according to the probability of the node being observed, thus inversely weighting the observed data to reduce bias caused by non-random missing values. For missing nodes, the GNM predicts the latent labels of these nodes, and then predicts the missing labels based on these representations.

[0045] Specifically, it is necessary to perform detailed segmented coding of the candidates' relevant information. This information can include attribute information. For example, based on the number of times an IP address is repeated, or the degree of repetition of a password, a high coding value may indicate proxy registration behavior; based on the number of times an address is repeated, a high coding value may indicate concentrated registration; based on the degree of repetition of an organization, a high coding value may mean unified organization of registration. This coding method divides features into different intervals based on their degree of repetition and assigns corresponding coding values. Through this segmented coding method based on repetition, all features are transformed into multi-level coding values, which not only allows for a more flexible representation of feature repetition but also helps the model better capture the similarities between candidates, thereby enhancing its ability to identify potential organized proxy registration behavior.

[0046] Thus, the encoded value corresponding to each node can be obtained through the above method, thereby obtaining the initial node features. Then, through message passing via the GNM model, the node features of each node can be updated and iterated continuously, ultimately obtaining the embedded representation of each node.

[0047] The GNM model captures the features of each node within its network structure. In this process, each node gathers information from its neighbors, aggregates and updates this information with its own features, forming a new node representation. In this mechanism, node features are continuously updated with each iteration, eventually incorporating information from neighbors, thus capturing both global and local relationships within the graph. After receiving messages from neighboring nodes, a weighted sum of features within the node's neighborhood is performed using graph convolution operations. Let h be the feature of node i in the l-th iteration. i (l) Then, after one convolution operation, its new features can be obtained by the following calculation:

[0048] Where, N (i) Let k be the set of neighbors of node i. i k j W represents the degree of nodes i and j. (l) It is a weight matrix. ( ) represents the activation function, typically the ReLU function. The GNM model iteratively updates the features of each node through message passing and graph convolution operations, ultimately obtaining the embedded representation of each node.

[0049] After determining the embedded representation of each examinee, the computing device can determine the confidence level corresponding to the pre-defined label for each examinee, and generate a risk score for each examinee based on the confidence level. The pre-defined label indicates that the examinee is at risk. In other words, GNM can predict whether an examinee is at risk based on the examinee's embedded representation. GNM can output the predicted confidence level during prediction, where only the confidence level corresponding to the pre-defined label can be obtained; that is, the probability that GNM considers a particular examinee to be at risk.

[0050] This stage focuses only on candidates with a predicted label of 0 (i.e., the preset label), collecting their probabilities and calculating their mean and standard deviation. The second step, based on the concept of z-scores, calculates a risk score between 0 and 100 for each candidate, effectively ranking and normalizing their risk probabilities. Candidates with lower z-scores are considered to have higher risk. For example, candidates can be categorized as high-risk, medium-risk, and low-risk. High-risk candidates (0-40 points): These candidates are considered to have a higher risk of cheating and should be given priority for attention and monitoring.

[0051] Medium-risk candidates (41-69 points): These candidates are considered to have a potential risk of cheating and should be closely monitored to prevent possible cheating.

[0052] Low-risk candidates (70-100 points): These candidates are considered low-risk and generally do not require special attention.

[0053] Optionally, the operation of determining the characteristics of high-risk candidates based on their corresponding risk scores includes: identifying candidates whose risk scores are below a first preset threshold, thus obtaining a set of high-risk candidates, where the lower the risk score, the higher the risk; and determining the frequency of occurrence of each characteristic in the characteristic set based on the high-risk candidate set. These characteristics include: registration IP address, mailing address, workplace, password, security question answers, reviewer, and whether the candidate is from a different location. The characteristics of high-risk candidates are determined based on the frequency of occurrence of each characteristic in the characteristic set.

[0054] In other words, we can statistically identify features that appear frequently among high-risk candidates and observe which features are repeated among these high-risk candidates. For example, if an IP address is associated with multiple high-risk candidates, then this IP may be a centralized registration point for cheating organizations. Specifically, we set a scoring threshold and select the set V of nodes representing all high-risk candidates with scores below a certain threshold. high For all v i ∈V high Examine the frequency of occurrence of each feature in the feature set F:

[0055] Wherein, I(v) i F k ) is the indicator function, node v i With characteristic F k At that time, I(v) i F k ) = 1, otherwise I(v) i F k =0. Set the threshold for the frequency of feature occurrence (i.e., the first preset threshold mentioned above) θ. f , with frequencies higher than θ f The features are labeled as "high-risk candidate features", and the high-risk feature set F is obtained. high .

[0056] In the above, a GNM model was used to generate a predicted probability of cheating risk for each examinee, and these predicted probabilities were quantified into easily understandable scores using a scoring card. Based on this, feature pattern recognition was used to analyze examinees with lower scores and identify patterns of certain high-frequency features, such as frequently shared IP addresses or communication addresses. The purpose of feature analysis and label propagation is to identify which features are predominantly present among high-risk examinees, thereby identifying key characteristics of organized cheating.

[0057] After identifying the characteristics of high-risk candidates, the nodes corresponding to these candidates can be designated as high-risk candidate nodes. These nodes are then labeled as high-risk, and a coloring operation is used to highlight them, facilitating subsequent label propagation. The label propagation algorithm identifies potential groups posing a test-taking risk (e.g., cheating rings) by propagating risk labels through the graph structure.

[0058] Further, the Neo4j graph database is used for coloring operations, specifically for high-risk candidate nodes v. i ∈V high Initialize its coloring in the graph and mark it as "high risk", that is, let v i .color=red. For any node v j ∈V, if it possesses feature F k ∈F high Then update the node's attributes, making v j .color=red.

[0059] Optionally, based on the risk score corresponding to the examinee and a graph structure, the operation of high-risk label propagation is performed, including: setting an initial propagation intensity for each node in the graph structure; in each iteration, starting from the current high-risk examinee node, determining the propagation intensity corresponding to the neighboring nodes of the current high-risk examinee node; if the propagation intensity is higher than a second preset threshold, the neighboring node is regarded as a high-risk examinee node, and the next iteration is entered after the label propagation of the current iteration is completed; and the iteration ends when the propagation intensity of each node tends to converge.

[0060] In other words, after initially identifying the high-risk candidate nodes, the high-risk label can be propagated to a subset of nodes within a certain neighborhood of each high-risk candidate node through the relationships between nodes in the graph structure. The method for propagating the high-risk label can be determined iteratively. In each iteration, the propagation strength of each node is calculated to determine whether the high-risk label can be propagated to the corresponding node. After continuous iteration, if the label propagation to each node no longer changes (i.e., the propagation strength of each node tends to converge), the iteration can end. Therefore, each high-risk candidate node can potentially propagate its high-risk label to nodes within a certain range. Thus, this method can identify groups at risk, rather than just discrete high-risk candidate nodes.

[0061] Optionally, the operation of determining the propagation strength corresponding to the neighboring nodes of the current high-risk candidate node includes: determining the propagation strength corresponding to the neighboring nodes of the current high-risk candidate node based on the propagation strength of the current high-risk candidate node, the normalized weight between the current high-risk candidate node and its neighboring nodes, and the propagation strength of the neighboring nodes after the previous iteration.

[0062] The label propagation algorithm uses a greedy strategy, iteratively labeling each colored node with its directly associated nodes. A propagation strength *s* is introduced to describe the propagation effect on the examinee. First, the initial strength of each node can be set:

[0063] Set a threshold γ for the propagation intensity; if the propagation intensity v of a certain node... j If s > γ, then v j It will also be considered as being stained, i.e., marked v. j .color=red, and v j Add to set V high In the middle, each iteration starts from v i ∈V high , for v i neighbor node v n Propagation occurs, satisfying the following condition in the l-th iteration: v n .s=v n .s+v i .s·α l ·w' in .

[0064] Where α∈(0,1) is the update rate, w' in For node v i With node v n After multiple iterations of propagation, the normalized weights (i.e., edge weights) between nodes will converge, and the number of colored nodes will stabilize, at which point the label propagation process ends. The overall coloring and label propagation process is as follows: Figure 3 As shown.

[0065] Figure 3 This is a schematic diagram of a dyeing process provided in Embodiment 1 of this disclosure.

[0066] exist Figure 3 The diagram illustrates the original graph structure and the three-round label propagation process based on this original graph structure. First, in the original graph structure, each node represents a candidate, and the edges between nodes represent the edge weights determined in step S202. Since nodes 1 and 3 are the initially determined candidate nodes, the initial coloring sets these two nodes to 1, and the propagation strength of the remaining nodes to 0. In the first round of propagation, the propagation strengths of the other three nodes are updated, but do not exceed the set second preset threshold (the second preset threshold is...). Figure 3(The corresponding value in the example is set to 0.7). In the second round of propagation, the propagation strength of node 2 after the update exceeds the second preset threshold, so node 2 is colored. In the third round of propagation, the update magnitude of the propagation strength of the uncolored nodes is relatively small and does not exceed the second preset threshold. Therefore, no new nodes are colored in the third round of propagation, and the propagation strength has converged, stopping the iteration of label propagation.

[0067] Finally, the computing device can visualize the groups of high-risk test-takers and the identification criteria. Through Neo4j's graph database, the relationships between test-takers can be visualized, allowing administrators to more intuitively understand the structure and scale of the cheating network. This visualization helps administrators better understand the complexity of cheating behavior, thereby developing more effective preventative measures.

[0068] In addition, refer to Figure 1 As shown, according to a second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein, when the program is executed, a processor performs any of the methods described above.

[0069] Therefore, according to this embodiment, (1) based on the characteristics of candidate groups with potential organized abnormal registration, a unified encoding method of feature aggregation is adopted to express candidate features. For example, according to the number of times IP is repeated or the degree of password repetition, a high encoding value may point to the registration behavior on behalf of others; according to the number of times address is repeated, a high encoding value often indicates the phenomenon of concentrated registration before the exam; according to the degree of unit repetition, a high encoding value may mean the phenomenon of unified organized registration before the exam. According to the degree of repetition of features, they are divided into different intervals and assigned corresponding encoding values. Through this segmented encoding method based on the degree of repetition, all features are transformed into multi-level encoding values, which can not only more flexibly represent the degree of repetition of features, but also help the model to better capture the similarity between candidates, thereby enhancing the ability to identify potential organized registration behavior on behalf of others.

[0070] (2) High-risk candidates are statistically analyzed to identify patterns of certain high-frequency characteristics as "high-risk candidate characteristics". If an IP address is associated with multiple high-risk candidates, then this IP may be a centralized registration point for cheating organizations. Therefore, high-risk candidate nodes are identified through "high-risk candidate characteristics", making the identification of high-risk candidate nodes more representative. (3) Coloring operations are performed in the Neo4j graph database, followed by label propagation, to identify risk candidate groups associated with high-risk candidate nodes. Through coloring operations, nodes with high-risk candidate characteristics can be marked prominently to facilitate subsequent label propagation operations; the label propagation algorithm can identify potential cheating groups by propagating risk labels in the graph structure.

[0071] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0072] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0073] Example 2

[0074] Figure 4 An abnormal registration group detection device based on a community detection algorithm for pre-exam registration is shown according to the first aspect of this embodiment, which corresponds to the method described according to the first aspect of Embodiment 1. (Reference) Figure 4 As shown, the device includes: a construction module 410, used to construct a graph structure representing the relationships between candidates using the Neo4j database, and to classify candidates into community or non-community individuals based on the graph structure, wherein nodes in the graph structure refer to candidates, edges in the graph structure are weighted undirected edges created based on the similarity of candidate features, and communities are used to indicate groups of candidates with potentially organized and abnormal registration; a message passing module 420, used to pass messages based on the graph structure using a preset graph neural network, generating embedded representations of each node as embedded representations corresponding to the corresponding candidates; and a scoring determination module 430, used to... Based on the embedded representation of the corresponding candidates, the confidence level of each candidate and the preset label is determined, and a risk score is generated for each candidate based on the confidence level. The preset label is used to indicate that the candidate is at risk. The coloring module 440 is used to determine the characteristics of high-risk candidates based on the risk score of the candidates, and to mark the nodes corresponding to candidates with high-risk candidate characteristics as high-risk candidate nodes and perform the coloring operation in the Neo4j database. The propagation module 450 is used to propagate the high-risk labels based on the graph structure to determine the risk candidate group associated with the high-risk candidate nodes.

[0075] Optionally, the construction module 410 is specifically used to: determine the candidate's attribute information, including the registration IP address, password, mailing address, work unit, and security questions; construct the node corresponding to the candidate; for two candidates, determine the common attribute information between the two candidates, and determine the preset weights corresponding to the common attribute information between the two candidates; determine the edge weights between the two candidates based on the preset weights corresponding to the common attribute information between the two candidates; and construct the edges between the two candidates based on the edge weights between the two candidates, thus obtaining a graph structure.

[0076] Optionally, the message passing module 420 is specifically used to divide the candidate's relevant information into different intervals according to the degree of repetition of the candidate's features and assign corresponding encoding values; to perform message passing on the graph structure through a semi-supervised graph neural network with non-random missing values, and to generate embedded representations of each node by combining node features and graph topology information.

[0077] Optionally, the coloring module 440 is specifically used to: identify candidates whose risk scores are lower than a first preset threshold to obtain a set of high-risk candidates, wherein the lower the risk score of the corresponding candidate, the higher the risk of the candidate; determine the frequency of occurrence of each feature in the feature set based on the set of high-risk candidates, the features include: registration IP, mailing address, work unit, password, security question answer, reviewer, and whether the candidate is from another location, etc.; and determine the characteristics of high-risk candidates based on the frequency of occurrence of each feature in the feature set.

[0078] Optionally, the propagation module 450 is specifically used to: set an initial propagation intensity for each node in the graph structure; in each iteration, starting from the current high-risk candidate node, determine the propagation intensity corresponding to the neighboring nodes of the current high-risk candidate node; if the propagation intensity is higher than a second preset threshold, the neighboring node is regarded as a high-risk candidate node, and the next iteration is entered after the label propagation of the current iteration is completed; and the iteration ends when the propagation intensity of each node tends to converge.

[0079] Optionally, the propagation module 450 is specifically used to determine the propagation intensity corresponding to the neighboring nodes of the current high-risk candidate node based on the propagation intensity of the current high-risk candidate node, the normalized weight between the current high-risk candidate node and its neighboring nodes, and the propagation intensity of the neighboring nodes after the previous iteration.

[0080] Therefore, according to this embodiment, (1) based on the characteristics of candidate groups with potential organized abnormal registration, a unified encoding method of feature aggregation is adopted to express candidate features. For example, according to the number of times IP is repeated or the degree of password repetition, a high encoding value may point to the registration behavior on behalf of others; according to the number of times address is repeated, a high encoding value often indicates the phenomenon of concentrated registration before the exam; according to the degree of unit repetition, a high encoding value may mean the phenomenon of unified organized registration before the exam. According to the degree of repetition of features, they are divided into different intervals and assigned corresponding encoding values. Through this segmented encoding method based on the degree of repetition, all features are transformed into multi-level encoding values, which can not only more flexibly represent the degree of repetition of features, but also help the model to better capture the similarity between candidates, thereby enhancing the ability to identify potential organized registration behavior on behalf of others.

[0081] (2) High-risk candidates are statistically analyzed to identify patterns of certain high-frequency characteristics as "high-risk candidate characteristics". If an IP address is associated with multiple high-risk candidates, then this IP may be a centralized registration point for cheating organizations. Therefore, high-risk candidate nodes are identified through "high-risk candidate characteristics", making the identification of high-risk candidate nodes more representative. (3) Coloring operations are performed in the Neo4j graph database, followed by label propagation, to identify risk candidate groups associated with high-risk candidate nodes. Through coloring operations, nodes with high-risk candidate characteristics can be marked prominently to facilitate subsequent label propagation operations; the label propagation algorithm can identify potential cheating groups by propagating risk labels in the graph structure.

[0082] Example 3

[0083] Figure 5 An abnormal registration group detection device based on a community detection algorithm for pre-exam registration is shown according to the first aspect of this embodiment, which corresponds to the method described according to the first aspect of Embodiment 1. (Reference) Figure 5As shown, the device includes: a processor 510; and a memory 520 connected to the processor 510, for providing the processor with instructions to process the following steps: constructing a graph structure representing the relationships between candidates using the Neo4j database, and classifying candidates into community or non-community individuals based on the graph structure, wherein nodes in the graph structure represent candidates, edges in the graph structure are weighted undirected edges created based on candidate feature similarity, and communities are used to indicate groups of candidates with potentially organized abnormal registration; generating embedded representations of each node as embedded representations corresponding to the corresponding candidates by message passing based on the graph structure through a preset graph neural network; determining the confidence level of the corresponding candidate and the preset label based on the embedded representation of the corresponding candidate, and generating a risk score for the corresponding candidate based on the confidence level, wherein the preset label is used to indicate that the candidate is at risk; determining the characteristics of high-risk candidates based on the risk score of the candidates, and marking the nodes corresponding to candidates with high-risk candidate characteristics as high-risk candidate nodes with high-risk labels, and performing a coloring operation in the Neo4j database; and performing label propagation of high-risk labels based on the graph structure to determine the risk candidate groups associated with the high-risk candidate nodes.

[0084] Optionally, the operation of constructing a graph structure representing the relationships between candidates using the Neo4j database specifically includes: determining the candidate's attribute information, including registration IP address, password, mailing address, work unit, and security questions; constructing the nodes corresponding to the candidates; for two candidates, determining the common attribute information between them, and determining the preset weights corresponding to the common attribute information; determining the edge weights between the two candidates based on the preset weights corresponding to the common attribute information; and constructing the edges between the two candidates based on the edge weights, thus obtaining the graph structure.

[0085] Optionally, the operation of generating embedded representations of each node by passing messages based on the graph structure using a preset graph neural network specifically includes: dividing the candidate's relevant information into different intervals according to the degree of repetition of the candidate's features and assigning corresponding encoding values; passing messages on the graph structure using a semi-supervised graph neural network with non-random missing values, and generating embedded representations of each node by combining node features and the topological information of the graph structure.

[0086] Optionally, the operation of determining the characteristics of high-risk candidates based on their risk scores includes: identifying candidates whose risk scores are below a first preset threshold to obtain a set of high-risk candidates, wherein the lower the risk score of a candidate, the higher the risk of the candidate; determining the frequency of occurrence of each feature in the feature set based on the set of high-risk candidates, where features include: registration IP address, mailing address, work unit, password, security question answer, reviewer, and whether the candidate is from a different location; and determining the characteristics of high-risk candidates based on the frequency of occurrence of each feature in the feature set.

[0087] Optionally, based on the risk score corresponding to the examinee and a graph structure, the operation of high-risk label propagation is performed, including: setting an initial propagation intensity for each node in the graph structure; in each iteration, starting from the current high-risk examinee node, determining the propagation intensity corresponding to the neighboring nodes of the current high-risk examinee node; if the propagation intensity is higher than a second preset threshold, the neighboring node is regarded as a high-risk examinee node, and the next iteration is entered after the label propagation of the current iteration is completed; and the iteration ends when the propagation intensity of each node tends to converge.

[0088] Optionally, the operation of determining the propagation strength corresponding to the neighboring nodes of the current high-risk candidate node includes: determining the propagation strength corresponding to the neighboring nodes of the current high-risk candidate node based on the propagation strength of the current high-risk candidate node, the normalized weight between the current high-risk candidate node and its neighboring nodes, and the propagation strength of the neighboring nodes after the previous iteration.

[0089] Therefore, according to this embodiment, (1) based on the characteristics of candidate groups with potential organized abnormal registration, a unified encoding method of feature aggregation is adopted to express candidate features. For example, according to the number of times IP is repeated or the degree of password repetition, a high encoding value may point to the registration behavior on behalf of others; according to the number of times address is repeated, a high encoding value often indicates the phenomenon of concentrated registration before the exam; according to the degree of unit repetition, a high encoding value may mean the phenomenon of unified organized registration before the exam. According to the degree of repetition of features, they are divided into different intervals and assigned corresponding encoding values. Through this segmented encoding method based on the degree of repetition, all features are transformed into multi-level encoding values, which can not only more flexibly represent the degree of repetition of features, but also help the model to better capture the similarity between candidates, thereby enhancing the ability to identify potential organized registration behavior on behalf of others.

[0090] High-risk candidates are statistically analyzed to identify patterns of certain high-frequency characteristics as "high-risk candidate characteristics". If an IP address is associated with multiple high-risk candidates, then this IP may be a centralized registration point for cheating organizations. Therefore, high-risk candidate nodes are identified through "high-risk candidate characteristics", making the identification of high-risk candidate nodes more representative. (3) Coloring operations are performed in the Neo4j graph database, followed by label propagation, to identify risk candidate groups associated with high-risk candidate nodes. Through coloring operations, nodes with high-risk candidate characteristics can be marked prominently to facilitate subsequent label propagation operations; the label propagation algorithm can identify potential cheating groups by propagating risk labels in the graph structure.

[0091] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0092] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0093] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0094] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0095] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0096] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0097] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for detecting abnormal registration groups before an exam based on a community detection algorithm, characterized in that, include: A graph structure representing the relationships between candidates is constructed using the Neo4j database, and candidates are divided into community or non-community individuals based on the graph structure. In the graph structure, nodes refer to candidates, and the edges are weighted undirected edges created based on the similarity of candidate features. The community is used to indicate groups of candidates with potentially organized and abnormal registration. The preset graph neural network performs message passing based on the graph structure to generate embedded representations of each node, which serve as embedded representations corresponding to the respective candidates. Based on the embedded representation of the corresponding candidate, the confidence level of the corresponding candidate and the preset label is determined, and a risk score is generated for the corresponding candidate based on the confidence level, wherein the preset label is used to indicate that the candidate has a risk. Based on the risk score of the examinee, the characteristics of high-risk examinees are determined, and the nodes corresponding to examinees with high-risk examinee characteristics are marked as high-risk examinee nodes and then colored in the Neo4j database. as well as Based on the graph structure, the high-risk labels are propagated to identify risk candidate groups associated with high-risk candidate nodes.

2. The method according to claim 1, characterized in that, The operations for constructing a graph structure representing the relationships between candidates using the Neo4j database include: Determine the candidate's attribute information, which includes at least one of the following: registration IP address, password, mailing address, work unit, and security questions; Construct the nodes corresponding to the candidates; For two candidates, determine the common attribute information that exists between the two candidates, and determine the preset weights corresponding to the common attribute information that exists between the two candidates; Based on the preset weights corresponding to the shared attribute information between the two candidates, the edge weights between them are determined; and Based on the edge weights between the two candidates, the edges between the two candidates are constructed to obtain the graph structure.

3. The method according to claim 1, characterized in that, The operation of generating embedded representations of each node through message passing based on the graph structure using a preset graph neural network specifically includes: Based on the degree of repetition of candidate characteristics, the relevant information of the candidates is divided into different intervals and assigned corresponding coding values; By using a semi-supervised graph neural network with non-random missing values, message passing is performed on the graph structure. By combining node features and topological information of the graph structure, embedded representations of each node are generated.

4. The method according to claim 1, characterized in that, Based on the candidate's risk score, the procedures for identifying high-risk candidates include: Candidates whose risk scores are below a first preset threshold are identified, resulting in a set of high-risk candidates. The lower the risk score of a candidate, the higher the risk of that candidate. Based on the set of high-risk candidates, the frequency of occurrence of each feature in the feature set is determined. The features include at least one of the following: registration IP address, mailing address, work unit, password, security question answer, reviewer, and whether the candidate is from a different location. The characteristics of high-risk candidates are determined based on the frequency of occurrence of each feature in the feature set.

5. The method according to claim 1, characterized in that, Based on the graph structure, the operation of propagating the high-risk label includes: Set the initial propagation strength for each node in the graph structure; In each iteration, starting from the current high-risk candidate node, the propagation strength corresponding to the neighboring nodes of the current high-risk candidate node is determined. If the propagation strength is higher than a second preset threshold, the neighboring node is designated as a high-risk candidate node, and the next iteration begins after the label propagation of the current iteration is completed; and The iteration ends when the propagation intensity of each node tends to converge.

6. The method according to claim 5, characterized in that, The operation of determining the propagation strength corresponding to the neighboring nodes of the current high-risk candidate node includes: Based on the propagation strength of the current high-risk candidate node, the normalized weight between the current high-risk candidate node and its neighboring nodes, and the propagation strength of the neighboring nodes after the previous iteration, the propagation strength corresponding to the neighboring nodes of the current high-risk candidate node is determined.

7. A storage medium, characterized in that, The storage medium includes a stored program, wherein the method described in any one of claims 1 to 6 is generated and executed by a processor when the program is run.

8. A device for detecting abnormal registration groups before exams based on a community detection algorithm, characterized in that, include: The module is used to construct a graph structure representing the relationships between candidates using the Neo4j database, and to classify candidates into community or non-community individuals based on the graph structure. In the graph structure, nodes refer to candidates, and the edges are weighted undirected edges created based on the similarity of candidate features. The community is used to indicate groups of candidates with potentially organized and abnormal registration. The message passing module is used to pass messages based on the graph structure through a preset graph neural network, and generate embedded representations of each node as embedded representations corresponding to the corresponding candidates. The scoring determination module is used to determine the confidence level of the corresponding candidate and the preset label based on the embedded representation of the corresponding candidate, and generate a risk score for the corresponding candidate based on the confidence level, wherein the preset label is used to indicate that the candidate has a risk. The coloring module is used to determine the characteristics of high-risk candidates based on the risk scores of the candidates, and to mark the nodes corresponding to candidates with high-risk candidate characteristics as high-risk candidate nodes and perform coloring operations in the Neo4j database. as well as The propagation module is used to propagate the high-risk tags based on the graph structure in order to identify the risk candidate groups associated with the high-risk candidate nodes.

9. A device for detecting abnormal registration groups before exams based on a community detection algorithm, characterized in that, include: processor; as well as A memory, connected to the processor, for providing the processor with instructions to perform the following processing steps: A graph structure representing the relationships between candidates is constructed using the Neo4j database, and candidates are divided into community or non-community individuals based on the graph structure. In the graph structure, nodes refer to candidates, and the edges are weighted undirected edges created based on the similarity of candidate features. The community is used to indicate groups of candidates with potentially organized and abnormal registration. The preset graph neural network performs message passing based on the graph structure to generate embedded representations of each node, which serve as embedded representations corresponding to the respective candidates. Based on the embedded representation of the corresponding candidate, the confidence level of the corresponding candidate and the preset label is determined, and a risk score is generated for the corresponding candidate based on the confidence level, wherein the preset label is used to indicate that the candidate has a risk. Based on the risk score of the examinee, the characteristics of high-risk examinees are determined, and the nodes corresponding to examinees with high-risk examinee characteristics are marked as high-risk examinee nodes and then colored in the Neo4j database. as well as Based on the graph structure, the high-risk labels are propagated to identify risk candidate groups associated with high-risk candidate nodes.