A community discovery method and system based on Word2vec-FL
By improving the loss function of word2vec and using focal loss embedded in word2vec's loss function, the problem of poor expression of low-frequency word vectors is solved, and more accurate community discovery is achieved.
Patent Information
- Application Number
- CN202210662377.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-13
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-06-13
AI Technical Summary
The existing word2vec method has poor expression of vectors for low-frequency words in the community discovery, resulting in inaccurate clustering and data omissions.
By improving the loss function of word2vec, the loss function of word2vec is embedded by focal loss, and the vector expression quality of low-frequency words is improved.
The vector expression quality of low-frequency words is improved, the clustering inaccurate problem in community discovery is solved, and the reasonable vector representation is obtained for each IP.
Smart Images

Figure CN115062699B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of community discovery, and in particular relates to a community discovery method and system based on Word2vec-FL. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Word embedding methods based on word2vec technology (including doc2vec, node2vec, etc.) are applied to the field of community discovery. They can convert various logs into vectors using word embedding methods and perform clustering, classification, and other calculations in the vector space.
[0004] Before training, word2vec calculates the frequency of each word and removes low-frequency words based on a manually set minimum frequency. Training sample sets often have a significant long-tail problem, where some words are ignored due to their low frequency and thus fail to obtain corresponding vector representations. Furthermore, the richness of the remaining words in the training sample set is significantly uneven. While the vectors learned for frequently occurring words are relatively accurate, low-frequency words are discarded or treated as random vectors.
[0005] Regarding network access relationships in social networks, when a source IP initiates access to a target IP, it cannot be said that such access relationships are frequent and therefore important; on the contrary, if the access relationship frequency is very low, it can be ignored. In the access relationship, we regard IPs as words, so the access relationship can be expressed as a causal relationship between the source IP and the destination IP. In the access relationship, we hope that each IP will receive its own vector after passing through word2vec, rather than being ignored because it appears less frequently. However, the original word2vec loss function focuses on whether the vectors generated by words with higher frequency are reasonable. Even if low-frequency words have vectors, their vectors are close to random vectors, and there is no explainable vector distance. This will lead to inaccurate community clustering and data omission problems.
[0006] Therefore, we must focus on the vector representation of low-frequency words. As the input of a series of subsequent neural networks, the quality of the word vector has a great impact on the final judgment result. Summary of the Invention
[0007] In order to solve the problems of poor low-frequency word vector expression and missing long-tail word vector expression in the original word embedding method, the present invention provides a community discovery method and system based on Word2vec-FL, which improves the focal loss of the word2vec loss function to obtain more reasonable vector expression for low-frequency words.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] A first aspect of the present invention provides a community discovery method based on Word2vec-FL.
[0010] A community discovery method based on Word2vec-FL, including:
[0011] Get traffic data;
[0012] Extract the source IP field and destination IP field in the traffic data;
[0013] Construct sample data based on the source IP field and the destination IP field;
[0014] Based on the sample data, the word2vec-FL model is input for training to obtain the vector corresponding to each IP; wherein the word2vec-FL model includes: embedding focal loss into the word2vec loss function;
[0015] The cosine similarity of the vectors is calculated, and each IP is clustered based on the cosine similarity of the vectors to obtain a community corresponding to each IP.
[0016] Furthermore, after extracting the source IP field and the destination IP field in the traffic data, the method further includes: deleting the default values in both the source IP field and the destination IP field.
[0017] Furthermore, constructing the sample data based on the source IP field and the destination IP field specifically includes: writing the source IP field and the destination IP field in pairs into a txt file with spaces as separators, with each traffic record in a separate line, to construct the sample data.
[0018] Furthermore, embedding the focal loss into the loss function of word2vec includes: embedding the focal loss into the loss function of Hierarchical softmax and the loss function of Negative Sampling respectively.
[0019] Furthermore, the loss function of embedding focal loss into Hierarchical softmax is specifically:
[0020]
[0021] Among them, α and γ are balance factors, γ>0, is the probability of being classified as positive or negative, w is the word, and β is the probability of the node being classified as positive.
[0022] Furthermore, the loss function of embedding focal loss into Negative Sampling is specifically:
[0023] L(w,u)=α(1-β) γ L w (u)×log(β)+(1-α)β γ (1-L w (u))×log(1-β)
[0024] Among them, α and γ are balance factors, γ>0, L w (u) is the probability of classification into positive and negative classes, w is the word, and β is the probability of the node being classified as positive.
[0025] Furthermore, the calculating of the vector cosine similarity between IPs based on the frequency of occurrence of the vector corresponding to each IP specifically includes: judging whether the frequency of occurrence of each vector is greater than a set threshold based on the frequency of occurrence of the vector corresponding to each IP; if so, it is a high-frequency IP; otherwise, it is a low-frequency IP; and calculating the vector cosine similarity between the high-frequency IP and the low-frequency IP.
[0026] A second aspect of the present invention provides a community discovery system based on Word2vec-FL.
[0027] A community discovery system based on Word2vec-FL, including:
[0028] A data acquisition module is configured to: acquire traffic data;
[0029] An extraction module is configured to: extract a source IP field and a destination IP field from the traffic data;
[0030] A sample construction module is configured to: construct sample data based on a source IP field and a destination IP field;
[0031] A computing module is configured to: input a word2vec-FL model for training based on sample data to obtain a vector corresponding to each IP; wherein the word2vec-FL model includes: embedding focal loss into the word2vec loss function;
[0032] The clustering module is configured to calculate the cosine similarity of the vectors, cluster each IP based on the cosine similarity of the vectors, and obtain the community corresponding to each IP.
[0033] A third aspect of the present invention provides a computer-readable storage medium.
[0034] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the community discovery method based on Word2vec-FL as described in the first aspect above.
[0035] A fourth aspect of the present invention provides a computer device.
[0036] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the Word2vec-FL-based community discovery method described in the first aspect are implemented.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] The present invention improves the focal loss of the word2vec loss function to obtain a more reasonable vector expression for low-frequency words, solving the problems of poor low-frequency word vector expression and missing long-tail word vector expression in the original word embedding method. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0040] Figure 1 4 is a flowchart of a community discovery method based on Word2vec-FL shown in an embodiment of the present invention. DETAILED DESCRIPTION
[0041] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0042] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0043] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0044] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods and systems according to the various embodiments of the present disclosure. It should be noted that each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code can include one or more executable instructions for implementing the logical functions specified in the various embodiments. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the flowchart and / or block diagram, and the combination of the boxes in the flowchart and / or block diagram, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0045] Example 1
[0046] like Figure 1 As shown, this embodiment provides a community discovery method based on Word2vec-FL. This embodiment uses the method applied to the server as an example for illustration. It is understandable that the method can also be applied to terminals, and can also be applied to a system including terminals, servers, and is implemented through the interaction between terminals and servers. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected by wired or wireless communication, which is not limited in this application. In this embodiment, the method includes the following steps:
[0047] Get traffic data;
[0048] Extract the source IP field and destination IP field in the traffic data;
[0049] Construct sample data based on the source IP field and the destination IP field;
[0050] Based on the sample data, the word2vec-FL model is input for training to obtain the vector corresponding to each IP; wherein the word2vec-FL model includes: embedding focal loss into the word2vec loss function;
[0051] The cosine similarity of the vectors is calculated, and each IP is clustered based on the cosine similarity of the vectors to obtain a community corresponding to each IP.
[0052] First, word2vec is a very simple neural network. Originally developed in the field of natural language processing, it represents the contextual relationships between words using two-dimensional vectors. This abstract, discrete contextual relationship is represented as a vector matrix in word2vec space. Therefore, if two words never have a contextual relationship, the cosine distance between their corresponding vectors is large.
[0053] At the same time, word2vec itself has a problem: it ignores low-frequency words. Word2vec defines a threshold for word frequency. If a word appears less than the threshold, it is ignored. Specifically, a text may contain 40,000 words, but word2vec may only generate vectors for 10,000 of them. This is because the frequency of the remaining 30,000 words in the text is less than the threshold. Although there are many words, their frequency is very low, which has little impact on subsequent classification and clustering.
[0054] This embodiment draws on the idea of focal loss in the field of computer vision to improve the loss function of word embedding methods such as word2vec, aiming to solve the problems of poor expression of low-frequency word vectors and missing expression of long-tail word vectors in the original word embedding methods.
[0055] 1. Algorithm Introduction
[0056] (1) word2vec
[0057] Word2vec is a language model that learns semantic knowledge from a large amount of text corpus in an unsupervised manner and is widely used in natural language processing.
[0058] Word2vec is a lightweight neural network consisting of only an input layer, hidden layers, and an output layer. Depending on the input and output, the model framework mainly includes CBOW and Skip-gram models. CBOW predicts the current word given the context of the word, while Skip-gram predicts the context of the word given the word itself.
[0059] Typically, neural network language models output the probability of predicting the target word. This means that each prediction must be calculated based on the entire dataset, which undoubtedly incurs a significant time overhead. Unlike other neural networks, Word2vec proposes two methods to accelerate training: Hierarchical softmax and negative sampling.
[0060] A.Hierarchical softmax
[0061] Hierarchical softmax records the positive and negative classification of samples in the form of a Huffman tree, where the left subtree 0 is the positive class and the right subtree 1 is the negative class.
[0062]
[0063] in, is the probability of being classified as the left and right subtrees of the Huffman tree, is the probability of classification into positive and negative classes, and w is the word.
[0064] When performing binary classification, the Sigmoid function is selected here. Then, the probability of a node being classified as a positive class is:
[0065]
[0066] For any word w in the dictionary, there must be a path p in the Huffman tree from the root node to the leaf node corresponding to word w. w , and this path is unique. Consider each branch as a binary classification, and the objective function is p(w|Context(w)).
[0067]
[0068] in:
[0069]
[0070] in, is the probability of each branch of the Huffman tree, and the loss function L(w,j) is:
[0071]
[0072] B. Negative Sampling
[0073] The essence of Negative Sampling is to use a known probability density function to estimate an unknown probability density function. Compared with Hierarchical Softmax, it no longer uses the complex Huffman tree, but instead uses relatively simple random negative sampling, which can significantly improve performance.
[0074] For a given positive sample (Context(w), w), the objective function we want to maximize is:
[0075]
[0076] Among them, L w (u) is the probability of classification into positive and negative classes.
[0077]
[0078] The loss function L(w,u) is:
[0079] L(w,u)=L w (u)×log(β)+[1-L w (u)]×log(1-β) (8)
[0080] (2) focal loss
[0081] Focal loss is primarily designed to address the severe imbalance in the ratio of positive and negative samples in one-stage object detection. This loss function reduces the weight of a large number of simple negative samples in training and can also be understood as a form of difficult sample mining.
[0082] Focal loss is a modification of the cross-entropy loss function. It adds a tuning factor, γ > 0, to reduce the loss of easily classified samples and focus more on difficult, misclassified samples. A balancing factor, α, is also added to balance the uneven ratio of positive and negative samples. When γ is 0 and α is 1, this is the cross-entropy loss function. As γ increases, the influence of the tuning factor increases.
[0083]
[0084] Among them, α and γ are balance factors, y is the classification probability, Lf l is the loss function.
[0085] 2. Algorithm Improvement
[0086] This example draws on the idea of focal loss to improve the word2vec loss function and proposes the word2vec-fl algorithm. The word2vec-fl algorithm supports the skip- and CBOW models of the original word2vec algorithm and improves the loss functions of two optimization methods, hierarchical softmax and negative sampling, respectively.
[0087] (1)Hierarchical softmax-focal loss
[0088] The objective function p(w|Context(w)) is
[0089]
[0090] in,
[0091]
[0092] The loss function is
[0093]
[0094] The gradient is calculated as
[0095]
[0096] (2)Negative Sampling-focal loss
[0097] The objective function g(w) is
[0098]
[0099] in,
[0100]
[0101] The loss function is
[0102] L(w,u)=α(1-β) γ L w (u)×log(β)+(1-α)β γ (1-L w (u))×log(1-β) (16)
[0103] The gradient is calculated as
[0104]
[0105] Algorithm Flow
[0106] Taking traffic access relations as an example, the following contents can be considered for implementation:
[0107] Step 1: First download some traffic data.
[0108] Step 2: Extract the source IP field and destination IP field from the traffic data and delete the default values.
[0109] Step 3: Write the source IP and destination IP into a txt file in pairs with spaces as separators, with each traffic record on a separate line to construct training samples.
[0110] Step 4: Input the training samples into Word2Vec-FL for training.
[0111] Step 5: Based on the training results, extract the vector corresponding to each IP.
[0112] Step 6: Based on the frequency of each IP address, select high-frequency and low-frequency IP addresses, calculate the vector cosine similarity between the two IP addresses, and compare it with the cosine similarity of the vectors generated by the original algorithm. Focus on comparing whether the vector cosine similarity between low-frequency IP addresses with close or no access relationship decreases or increases compared to the original algorithm. Using the vector cosine similarity, cluster the IP addresses using DBSCAN to determine the community corresponding to each IP address.
[0113] Practical verification has shown that word2vec-fl performs better than word2vec in generating vectors for low-frequency words. It can obtain more reasonable vector representations for low-frequency words based on limited training samples. The concept of focal loss effectively addresses the problems of poor vector representation for low-frequency words and missing vector representation for long-tail words caused by sample imbalance and long-tail words.
[0114] Example 2
[0115] This embodiment provides a community discovery system based on Word2vec-FL.
[0116] A community discovery system based on Word2vec-FL, including:
[0117] A data acquisition module is configured to: acquire traffic data;
[0118] An extraction module is configured to: extract a source IP field and a destination IP field from the traffic data;
[0119] A sample construction module is configured to: construct sample data based on a source IP field and a destination IP field;
[0120] A computing module is configured to: input a word2vec-FL model for training based on sample data to obtain a vector corresponding to each IP; wherein the word2vec-FL model includes: embedding focal loss into the word2vec loss function;
[0121] The clustering module is configured to calculate the cosine similarity of the vectors, cluster each IP based on the cosine similarity of the vectors, and obtain the community corresponding to each IP.
[0122] It should be noted that the examples and application scenarios implemented by the above-mentioned data acquisition module, extraction module, sample construction module, model extraction module, calculation module, and clustering module are the same as those in the steps of Example 1, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules, as part of the system, can be executed in a computer system, such as a set of computer-executable instructions.
[0123] Example 3
[0124] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the community discovery method based on Word2vec-FL as described in the first embodiment are implemented.
[0125] Example 4
[0126] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the Word2vec-FL-based community discovery method described in the first embodiment are implemented.
[0127] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0128] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0129] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0131] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0132] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A community discovery method based on Word2vec-FL, characterized by: include: Get traffic data; Extract the source IP field and destination IP field in the traffic data; Construct sample data based on the source IP field and the destination IP field; Based on the sample data, the word2vec-FL model is input for training to obtain the vector corresponding to each IP; wherein the word2vec-FL model includes: embedding focal loss into the word2vec loss function; Word2vec is a lightweight neural network whose model includes an input layer, a hidden layer, and an output layer. The model framework mainly includes CBOW and Skip-gram models, depending on the input and output. The word2vec-fl algorithm improves the loss functions of two optimization methods that speed up training: Hierarchical softmax and Negative Sampling. Embedding focal loss into the word2vec loss function involves embedding focal loss into the Hierarchical softmax loss function and the Negative Sampling loss function, respectively. The loss function of embedding focal loss into Hierarchical softmax is specifically: Among them, α and γ are balance factors, γ>0, is the probability of classification into positive and negative classes, w is a word, and β is the probability of a node being classified as a positive class; the loss function of embedding focal loss into Negative Sampling is specifically: L(w,u)=α(1-β) γ L w (u)×log(β)+(1-α)β γ (1-L w (u))×log(1-β); where α and γ are balance factors, γ>0, L w (u) is the probability of classification into positive and negative classes, w is the word, and β is the probability of the node being classified into the positive class; The cosine similarity of the vectors is calculated, and each IP is clustered based on the cosine similarity of the vectors to obtain a community corresponding to each IP.
2. The community discovery method based on Word2vec-FL according to claim 1, characterized in that After extracting the source IP field and the destination IP field in the traffic data, the method further includes: deleting the default values in both the source IP field and the destination IP field.
3. The community discovery method based on Word2vec-FL according to claim 1, characterized in that The construction of sample data based on the source IP field and the destination IP field specifically includes: writing the source IP field and the destination IP field into a txt file in pairs with spaces as separators, with each traffic record in a separate line, to construct the sample data.
4. The community discovery method based on Word2vec-FL according to claim 1, characterized in that The calculating of the vector cosine similarity between IPs based on the frequency of occurrence of the vector corresponding to each IP specifically includes: judging whether the frequency of occurrence of each vector is greater than a set threshold based on the frequency of occurrence of the vector corresponding to each IP; if so, it is a high-frequency IP; otherwise, it is a low-frequency IP; and calculating the vector cosine similarity between the high-frequency IP and the low-frequency IP.
5. A community discovery system based on Word2vec-FL, characterized by: include: A data acquisition module is configured to: acquire traffic data; An extraction module is configured to: extract a source IP field and a destination IP field from the traffic data; A sample construction module is configured to: construct sample data based on a source IP field and a destination IP field; The computing module is configured to: input a word2vec-FL model for training based on sample data to obtain a vector corresponding to each IP; wherein the Word2Vec-FL model includes: embedding focal loss into the word2vec loss function; word2vec is a lightweight neural network, and its model includes an input layer, a hidden layer, and an output layer. The model framework mainly includes CBOW and Skip-gram models depending on the input and output; the word2vec-fl algorithm makes corresponding improvements to the loss functions of two optimization methods for accelerating training speed, Hierarchical softmax and Negative Sampling, and embedding focal loss into the word2vec loss function includes: embedding focal loss into the Hierarchical softmax loss function and the Negative Sampling loss function respectively; The loss function of embedding focal loss into Hierarchical softmax is specifically: Among them, α and γ are balance factors, γ>0, is the probability of classification into positive and negative classes, w is a word, and β is the probability of a node being classified as a positive class; the loss function of embedding focal loss into Negative Sampling is specifically: L(w,u)=α(1-β) γ L w (u)×log(β)+(1-α)β γ (1-L w (u))×log(1-β); where α and γ are balance factors, γ>0, L w (u) is the probability of classification into positive and negative classes, w is the word, and β is the probability of the node being classified into the positive class; The clustering module is configured to calculate the cosine similarity of the vectors, cluster each IP based on the cosine similarity of the vectors, and obtain the community corresponding to each IP.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the community discovery method based on Word2vec-FL are implemented as described in any one of claims 1 to 4.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the Word2vec-FL-based community discovery method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Water consumption pattern mining and matching method, system and equipment
CN110597880A
Address matching algorithm based on deep learning model
CN111881677A