A label confirmation method, device, apparatus, storage medium and program product

By clustering and optimizing IP addresses, the problem of low IP address labeling accuracy was solved, achieving higher labeling accuracy and consistency.

CN122132859APending Publication Date: 2026-06-02中国移动通信集团江西有限公司 +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
中国移动通信集团江西有限公司
Filing Date
2026-02-13
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in labeling IP addresses, making it difficult to accurately associate labels with anonymized IP addresses.

Method used

Multiple IP addresses are clustered to generate multiple clusters, and a label is set for each cluster. The reward value is calculated by combining the node state vector and action parameters. The action parameter with the highest reward value is selected to send the label, and the label content is optimized among multiple nodes to improve accuracy.

Benefits of technology

It improves the accuracy of IP address labeling by ensuring the accuracy and consistency of label content through clustering and label optimization between nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132859A_ABST
    Figure CN122132859A_ABST
Patent Text Reader

Abstract

This invention provides a tag confirmation method, apparatus, device, storage medium, and program product, relating to the field of data processing technology. The method includes: clustering multiple Internet Protocol (IP) addresses to obtain multiple first clusters and a first tag for each first cluster; determining a node state vector of a first node; setting multiple node action parameters for the first node; calculating a reward value for each node action parameter based on the node state vector and the multiple node action parameters; sending the first tags of the multiple IP addresses to a second node based on the target action parameters; receiving the first IP address and a second tag of the first IP address, wherein the first tag content of the first IP address is different in at least two of the multiple nodes, and the tag content of the second tag is the tag content with the highest score among at least two contents, and the at least two contents are the tag content of the first IP address in the first tags of the multiple nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and specifically to a tag verification method, apparatus, device, storage medium, and program product. Background Technology

[0002] Internet Protocol (IP) addresses are typically associated with users. By labeling IP addresses, user behavior and interests can be directly identified, facilitating targeted services. However, in related technologies, IP addresses are often anonymized, making it difficult to accurately associate them with labels, resulting in low label accuracy.

[0003] It is evident that the relevant technologies suffer from low accuracy in IP address labeling. Summary of the Invention

[0004] This invention provides a tag verification method, apparatus, device, storage medium, and program product to solve the problem of low accuracy of IP address-based tags in related technologies.

[0005] To solve the above problems, the present invention is implemented as follows: In a first aspect, embodiments of the present invention provide a tag confirmation method, applied to a first node, the method comprising: Multiple Internet Protocol (IP) addresses are clustered to obtain multiple first clusters and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster. Determine the node state vector of the first node, wherein the node state vector is used to characterize at least one of the clustering accuracy of the plurality of first clusters, the communication latency of the first node, and the computational load; Set multiple node action parameters for the first node, where different node action parameters are used to characterize different strategies for the first node to send tags; Based on the node state vector and the multiple node action parameters, calculate the reward value for each node action parameter. Based on the target action parameters, a first tag of the plurality of IP addresses is sent to the second node, wherein the target action parameters are the node action parameters with the largest reward value among the plurality of node action parameters; The system receives a first IP address and a second tag for the first IP address. The first IP address has different tag content in at least two nodes across multiple nodes. The tag content of the second tag is the tag content with the highest score among at least two contents. The at least two tag contents are the tag content of the first IP address within the first tag across the multiple nodes. The score is calculated by the confidence level of each tag content within the node.

[0006] Secondly, embodiments of the present invention provide a tag confirmation method applied to a second node, the method comprising: The system receives first labels for multiple IP addresses sent by each of multiple nodes, wherein the first labels for the multiple IP addresses sent by each node are obtained by clustering the multiple IP addresses by the nodes. In the case where a first IP address exists among the plurality of IP addresses, at least two tag contents are obtained, as well as the confidence level of each tag content within a node. The tag contents of the first IP address are different in the first tags of at least two nodes among the plurality of nodes. The at least two tag contents are the tag contents of the first IP address within the first tags of the plurality of nodes. A score for each tag content is calculated based on the confidence level of each tag content; Set the tag content with the highest score among the at least two tag contents as the tag content of the second tag of the first IP address; Send the first IP address and the second tag of the first IP address to the plurality of nodes.

[0007] Thirdly, embodiments of the present invention also provide a tag verification device applied to a first node, the device comprising: The first clustering module is used to cluster multiple Internet Protocol (IP) addresses to obtain multiple first clusters and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster. The determination module is used to determine the node state vector of the first node, wherein the node state vector is used to characterize at least one of the clustering accuracy of the plurality of first clusters, the communication latency of the first node, and the computational load. The setting module is used to set multiple node action parameters of the first node, wherein different node action parameters are used to characterize different strategies for the first node to send tags. The first calculation module is used to calculate the reward value of each node action parameter based on the node state vector and the multiple node action parameters. The sending module is used to send a first label of the plurality of IP addresses to the second node based on the target action parameter, wherein the target action parameter is the node action parameter with the largest reward value among the plurality of node action parameters; A receiving module is used to receive a first IP address and a second tag of the first IP address. The first IP address has different tag content in at least two of the multiple nodes. The tag content of the second tag is the tag content with the highest score among at least two contents. The at least two tag contents are the tag contents of the first IP address in the first tags of the multiple nodes. The score is calculated by the confidence level of each tag content in the node.

[0008] Fourthly, embodiments of the present invention also provide a tag verification device applied to a first node, the device comprising: The receiving module is used to receive the first labels of multiple IP addresses sent by each of the multiple nodes, wherein the first labels of the multiple IP addresses sent by each node are obtained by clustering the multiple IP addresses by the nodes. The acquisition module is used to acquire at least two tag contents and the confidence level of each tag content within a node when a first IP address exists among the plurality of IP addresses. The tag contents of the first tag of the first IP address are different in at least two of the plurality of nodes. The at least two tag contents are the tag contents of the first IP address within the first tag of the plurality of nodes. The calculation module is used to calculate the score of each tag content based on the confidence level of each tag content; The setting module is used to set the tag content with the highest score among the at least two tag contents as the tag content of the second tag of the first IP address; The sending module is used to send the first IP address and the second tag of the first IP address to the plurality of nodes.

[0009] Fifthly, embodiments of the present invention also provide an electronic device applied to a first node, including a transceiver and a processor. The processor is configured to cluster multiple Internet Protocol (IP) addresses to obtain multiple first clusters and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster. The processor is further configured to determine the node state vector of the first node, the node state vector being used to characterize at least one of the clustering accuracy of the plurality of first clusters, the communication latency of the first node, and the computational load. The processor is further configured to set multiple node action parameters for the first node, wherein different node action parameters are used to characterize different strategies for the first node to send tags. The processor is further configured to calculate the reward value for each node action parameter based on the node state vector and the plurality of node action parameters. The transceiver is used to send a first tag of the plurality of IP addresses to the second node based on a target action parameter, wherein the target action parameter is the node action parameter with the largest reward value among the plurality of node action parameters. The transceiver is further configured to receive a first IP address and a second tag for the first IP address, wherein the first IP address has different tag content in at least two of the multiple nodes, and the tag content of the second tag is the tag content with the highest score among at least two contents, wherein the at least two tag contents are the tag contents of the first IP address in the first tags of the multiple nodes, and the score is calculated by the confidence level of each tag content in the node.

[0010] Sixthly, embodiments of the present invention also provide an electronic device applied to a second node, including a transceiver and a processor. The processor is configured to cluster multiple Internet Protocol (IP) addresses to obtain multiple first clusters and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster. The processor is further configured to determine the node state vector of the first node, the node state vector being used to characterize at least one of the clustering accuracy of the plurality of first clusters, the communication latency of the first node, and the computational load. The processor is further configured to set multiple node action parameters for the first node, wherein different node action parameters are used to characterize different strategies for the first node to send tags. The processor is further configured to calculate the reward value for each node action parameter based on the node state vector and the plurality of node action parameters. The transceiver is used to send a first tag of the plurality of IP addresses to the second node based on a target action parameter, wherein the target action parameter is the node action parameter with the largest reward value among the plurality of node action parameters. The transceiver is further configured to receive a first IP address and a second tag for the first IP address, wherein the first IP address has different tag content in at least two of the multiple nodes, and the tag content of the second tag is the tag content with the highest score among at least two contents, wherein the at least two tag contents are the tag contents of the first IP address in the first tags of the multiple nodes, and the score is calculated by the confidence level of each tag content in the node.

[0011] In a seventh aspect, embodiments of the present invention provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the tag verification method described in the first aspect.

[0012] Eighthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the tag verification method described in the first aspect.

[0013] In a ninth aspect, the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the tag verification method described in the first aspect.

[0014] In this embodiment of the invention, multiple Internet Protocol (IP) addresses are clustered to obtain multiple first clusters, and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster. A node state vector of the first node is determined, which characterizes at least one of the clustering accuracy of the multiple first clusters, the communication latency of the first node, and the computational load. Multiple node action parameters of the first node are set, where different node action parameters characterize different strategies for the first node to send labels. Based on the above... The system uses a point state vector and multiple node action parameters to calculate the reward value for each node action parameter. Based on the target action parameter, it sends a first label for each of the multiple IP addresses to a second node. The target action parameter is the node action parameter with the highest reward value among the multiple node action parameters. It receives a first IP address and a second label for that first IP address. The first IP address has different label content in at least two nodes among the multiple nodes. The second label's label content is the label content with the highest score among at least two contents. The at least two label contents are the label content of the first IP address within the first labels of the multiple nodes. The score is calculated based on the confidence level of each label content within the node. In this way, the first labels for multiple IP addresses are obtained through clustering. The target action parameter for the first node to send the label to the second node is determined based on the reward value, improving the accuracy of the first labels received by the second node. Furthermore, the second node optimizes the first label of the first IP address to obtain a second label for the first IP address, which is then sent to multiple nodes to update the label of the first IP address, further improving the label accuracy of the IP address. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart of a tag confirmation method applied to a first node provided by an embodiment of the present invention; Figure 2 This is a flowchart of a tag confirmation method applied to a second node provided by an embodiment of the present invention; Figure 3 This is a structural diagram of a tag verification device applied to a first node according to an embodiment of the present invention; Figure 4 This is a structural diagram of a tag verification device applied to a second node according to an embodiment of the present invention; Figure 5 This is a structural diagram of an electronic device applied to a first node according to an embodiment of the present invention; Figure 6 This is a structural diagram of an electronic device applied to a second node, provided by an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Please see Figure 1 , Figure 1 This is a flowchart of a tag confirmation method applied to a first node provided by an embodiment of the present invention, such as... Figure 1 As shown, it includes the following steps: Step 101: Cluster multiple Internet Protocol (IP) addresses to obtain multiple first clusters and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster.

[0019] The first node mentioned above is the node that directly receives requests sent by user equipment. The first node will receive requests from multiple IP addresses, and can then cluster different IP addresses and generate labels.

[0020] The above clustering of multiple IP addresses yields multiple first clusters. Each IP address is in only one first cluster, and each cluster includes at least one IP address. The first labels of different first clusters are different.

[0021] This can be achieved by clustering multiple IP addresses based on their characteristics to obtain multiple first clusters. Since the IP addresses in each first cluster are similar, they can be considered to be of the same type. Therefore, the first label of the first cluster is set as the first label of the IP address, thereby obtaining the first label of each IP address through clustering.

[0022] The above-mentioned clustering of multiple IP addresses can be performed according to a pre-configured clustering algorithm, such as hierarchical clustering algorithm and density-based spatial clustering of applications with noise (DBSCAN) algorithm.

[0023] Step 102: Determine the node state vector of the first node. The node state vector is used to characterize at least one of the clustering accuracy of the plurality of first clusters, the communication latency of the first node, and the computational load.

[0024] The aforementioned node state vector is obtained by statistically analyzing the data of the first node, and is specifically used to characterize the clustering situation and / or network situation of the first node.

[0025] The clustering status of the first node can be represented by the clustering accuracy, and the network status of the first node can be represented by parameters such as communication latency and node computing load.

[0026] In some implementations, the clustering accuracy, communication latency, and computational load of the first node can be collected, and the node state vector of the first node can be constructed using the clustering accuracy, communication latency, and computational load.

[0027] It should be understood that the node state vector of the first node may change. Therefore, when constructing the node state vector of the first node through clustering accuracy, communication latency and computational load, the clustering accuracy, communication latency and computational load can be statistically analyzed after multiple IP addresses are clustered to obtain multiple first clusters, and the node state vector can be constructed based on the statistical data.

[0028] Step 103: Set multiple node action parameters for the first node. Different node action parameters are used to characterize different strategies for the first node to send tags.

[0029] It should be understood that after clustering multiple IP addresses, the first node can directly provide business services based on the first labels of the clustered IPs, or it can upload the clustering results to the master node or server for optimization of the first labels. Therefore, the first node has multiple node action parameters. Among these parameters, one node action parameter is selected to determine whether the first node sends the clustering results (i.e., the labels of the multiple IP addresses).

[0030] Among the multiple node action parameters, each node's action parameters are different, and these different node action parameters can represent different strategies adopted by the first node when sending tags. For example, the strategy could be to send tags at a different set rate (such as 1 MB / s), or not to send tags at all.

[0031] Step 104: Based on the node state vector and the multiple node action parameters, calculate the reward value for each node action parameter.

[0032] The aforementioned reward value represents the difference between label accuracy and label upload cost. A higher reward value indicates a higher clustering accuracy for the first node and a lower cost for uploading the first labels from multiple IP addresses within the first node to the master node or server. In this case, the first node is more inclined to send labels. Conversely, a lower reward value indicates a lower distance accuracy for the first node and a higher cost for uploading the first labels from multiple IP addresses within the first node to the master node or server. In this case, the first node is more inclined not to send labels. In this embodiment, the reward value can be used to determine the label sending strategy.

[0033] The reward value of each node action parameter is calculated based on the node state vector and multiple node action parameters. Alternatively, the cost of each node action parameter can be calculated, and the difference between the node state vector and the cost of each node action parameter can be set as the reward value of each node action parameter.

[0034] Step 105: Send the first tag of the multiple IP addresses to the second node based on the target action parameter, wherein the target action parameter is the node action parameter with the largest reward value among the multiple node action parameters.

[0035] The target action parameter mentioned above is one of the node action parameters among multiple node action parameters. Different node action parameters correspond to different sending strategies. Sending the first label of multiple IP addresses to the second node based on the target action parameter can be understood as sending the first label of multiple IP addresses based on the sending strategy corresponding to the target action parameter.

[0036] In this embodiment, sending the first tags of multiple IP addresses to the second node through the target action parameters can make the first tags of multiple IP addresses received by the second node more accurate.

[0037] In some implementations, an action threshold can be set to determine whether to send a tag. If the maximum reward value corresponding to the action parameters of multiple nodes is less than the set action threshold, then no tag will be sent.

[0038] The second node mentioned above is the node corresponding to the master node or server device. The second node does not directly receive requests, but receives the first labels of multiple IP addresses sent by the first node, and optimizes the first labels of multiple IP addresses to improve the accuracy of IP address labels.

[0039] Step 106: Receive a first IP address and a second tag for the first IP address. The first IP address has different tag content in at least two of the multiple nodes. The tag content of the second tag is the tag content with the highest score among at least two contents. The at least two tag contents are the tag contents of the first IP address in the first tags of the multiple nodes. The score is calculated by the confidence level of each tag content in the node.

[0040] The aforementioned multiple nodes include a first node. Each of the multiple nodes is used to receive requests from multiple IP addresses and cluster the multiple IP addresses to generate first labels for the multiple IP addresses. After obtaining the first labels for the multiple IP addresses, the multiple nodes send the first labels for the multiple IP addresses to a second node, so that the second node can receive the first labels for the multiple IP addresses from different nodes, and thus optimize the first labels for the multiple IP addresses.

[0041] It should be understood that since the first label of multiple IP addresses can be obtained from clustering at different nodes, and the clustering results of different nodes are different, there may be cases where the label content corresponding to the first label of the same IP address is different after the second node receives the label. In this case, the second node needs to determine the second label of the IP address. The second label is more accurate than the first label, thereby improving the overall accuracy of multiple IP addresses.

[0042] Taking the first IP address as an example, if the first IP address has different label content for at least two nodes across multiple addresses, then the first IP address needs optimization. For instance, the first IP address's first label includes two different label contents. In the second node, a score is calculated for each label content based on the confidence level of the node containing the different label contents. The label content with the highest score is then selected as the second label content, resulting in a more accurate second label. The second label is then sent to multiple nodes (including the first node) so that all nodes can update the first label of the first IP address to the second label, improving label accuracy.

[0043] In some embodiments, an agent can be set up at the first node to execute the steps of the tag verification method applied to the first node in this invention.

[0044] In this embodiment of the invention, multiple Internet Protocol (IP) addresses are clustered to obtain multiple first clusters, and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster. A node state vector of the first node is determined, which characterizes at least one of the clustering accuracy of the multiple first clusters, the communication latency of the first node, and the computational load. Multiple node action parameters of the first node are set, where different node action parameters characterize different strategies for the first node to send labels. Based on the above... The system uses a point state vector and multiple node action parameters to calculate the reward value for each node action parameter. Based on the target action parameter, it sends a first label for each of the multiple IP addresses to a second node. The target action parameter is the node action parameter with the highest reward value among the multiple node action parameters. It receives a first IP address and a second label for that first IP address. The first IP address has different label content in at least two nodes among the multiple nodes. The second label's label content is the label content with the highest score among at least two contents. The at least two label contents are the label content of the first IP address within the first labels of the multiple nodes. The score is calculated based on the confidence level of each label content within the node. In this way, the first labels for multiple IP addresses are obtained through clustering. The target action parameter for the first node to send the label to the second node is determined based on the reward value, improving the accuracy of the first labels received by the second node. Furthermore, the second node optimizes the first label of the first IP address to obtain a second label for the first IP address, which is then sent to multiple nodes to update the label of the first IP address, further improving the label accuracy of the IP address.

[0045] In one embodiment, the reward value of each node action parameter is calculated using the following formula: ; in the formula This represents the reward value for each node's action parameter, 0. T Represents the coefficient matrix. Let R() represent the discount factor, R() represent the calculation formula, st represent the node state vector of the first node, and a() represent the discount factor. t This represents the node action parameters.

[0046] Furthermore, the reward value for the target action parameter can be expressed as: ; The above formula can be used to determine the reward value for the target action parameters.

[0047] In one embodiment, before clustering the multiple IP addresses to obtain multiple first clusters, the method further includes: Obtain the plurality of IP addresses, each of which includes a network prefix and a host identifier; A random salt value is generated for each IP address based on the timestamp obtained for each IP address; The host identifier of the corresponding IP address is encrypted based on the random salt value of each IP address to obtain multiple encrypted IP addresses; The process of clustering multiple IP addresses yields multiple first clusters, including: The multiple encrypted IP addresses are clustered to obtain multiple first clusters.

[0048] It should be noted that IP addresses are 32-bit binary, segmented into a network prefix (first 24 bits) and a host identifier (last 8 bits). The network prefix is ​​retained to maintain group coherence, while the host identifier is encrypted and obfuscated to encrypt the IP address. Specifically, the host identifier is converted to an unsigned integer before encryption, as shown by the following formula: ; IP in the formula int The host is represented by a, b, c, and d, which are the four bytes of the IP address.

[0049] The above method generates a random salt value for each IP address based on the timestamp of each IP address. The timestamp of each IP address can be the timestamp when the IP address is received or the timestamp when the IP address is encrypted.

[0050] Furthermore, generating a random salt value for each IP address can ensure the unpredictability of the salt value using the HMAC-SHA256 algorithm, which can be expressed by the following formula: ; In the formula, salt represents the salt value, HMAC-SHA256() represents the HMAC-SHA256 algorithm, timestamp represents the timestamp, device_id represents the identifier of the device corresponding to the IP address, and K master This represents the root key.

[0051] One approach is to set the salt value to be updated periodically (e.g., with a 1-hour cycle) to prevent rainbow table attacks. Experiments show that this mechanism increases the difficulty of brute-force attacks to 2. 28 .

[0052] The above method encrypts the host identifier of the corresponding IP address based on the random salt value of each IP address to obtain multiple encrypted IP addresses. This encryption of IP addresses can be achieved using the Crypto-PAn algorithm.

[0053] Specifically, the host identifier of each IP address is encrypted based on a random salt value, which can be expressed by the following formula: ; masked represents the encrypted host identifier, and K represents the encryption key generated by the salt value.

[0054] The encrypted host identifier is then combined with the first 24 bits of the reserved IP address (i.e., the network prefix) to generate an encrypted IP address, specifically represented as follows: ; In the formula, masked_IP is the encrypted IP address, IP is the address before encryption, and 0xFFFFFF00 and 0x000000FF represent the specific values ​​of the IP.

[0055] In this embodiment of the invention, a plurality of IP addresses are obtained, each IP address including a network prefix and a host identifier; a random salt value is generated for each IP address based on the timestamp of each IP address acquisition; the host identifier of the corresponding IP address is encrypted based on the random salt value of each IP address, resulting in a plurality of encrypted IP addresses. Thus, by encrypting the IP addresses, anonymization of the IP addresses can be achieved, reducing the risk of IP address leakage and improving IP address security.

[0056] For example, an attack can be launched against an encrypted IP address to verify the security of the encrypted IP.

[0057] Specifically, assume there exists a multinomial-time algorithm A that can achieve a multinomial-time result with probability. By cracking anonymized IPs, algorithm B can be constructed to solve the pseudo-random function problem of the Rijndael algorithm: 1. Algorithm B receives the challenger's (i.e., Algorithm A) proposal. ; in, This refers to a pseudo-random function oracle, a common construct in cryptographic security proofs. Indicates a key is The pseudo-random function; An oracle is an abstract computational model that a challenger can query for input and receive output, but the queryer is unaware of the internal implementation. In this proof, the challenger provides Algorithm B with a pseudo-random function oracle. Algorithm B can be given arbitrary input and output F as if it were a black box. k (x), without needing to know its internal structure.

[0058] 2. B simulates the IP anonymization process, forwarding A's query to Oracle; 3. If A successfully cracks the code, then B can distinguish between PRF and random functions with a probability ≥ 1. ; neg(n) is an abbreviation for negligible function, which represents a function that tends to be minimal as the security parameter $n$ increases. It is often used to formalize cryptographic security. This represents the probability that attack algorithm A will be successfully cracked; neg(n) represents a negligible function, that is, for any polynomial poly(n), when When it is large enough, neg(n) < 1 / poly(n).

[0059] Since the Rijndael algorithm has been proven to be a safe PRF, this algorithm satisfies computational safety.

[0060] Furthermore, collision resistance analysis can be performed on encrypted IP addresses.

[0061] Specifically, for any two different IP addresses (IP1 ≠ IP2), the collision probability is calculated using the following formula:

[0062] in the formula This indicates the encrypted IP1 (i.e. ) and encrypted IP2 (i.e. The collision probability of two IPs is calculated using a formula based on the prefix-preserving obfuscation property. The first 24 bits of the network prefix remain unchanged, while only the last 8 bits of the host identifier are obfuscated. Therefore, two IPs can only be considered identical after encryption if their first 24 bits are the same and their last 8 bits collide. Assuming the encryption algorithm is an ideal random function, the probability of a collision in the last 8 bits of two different IPs is... ; Understandably, given that the first 24 bits are the same, the probability of two different host identifiers (the last 8 bits) colliding under ideal encryption is... That is, the collision probability is <0.4%, which meets the needs of practical applications.

[0063] In some implementations, parallel encryption processing can be used, such as GPU acceleration (NVIDIA A100), which uses CUDA kernel functions to achieve parallel encryption of IP addresses, with a single card throughput of 1 million IPs / second, which is 20 times higher than that achieved by CPU.

[0064] In some implementations, a cache is established for frequently accessed IP addresses (such as an LRU policy), with a cache hit rate of up to 65%, reducing the average encryption time from 1μs to 0.3μs, effectively improving encryption efficiency.

[0065] In one embodiment, clustering multiple IP addresses to obtain multiple first clusters includes: Extract the feature indicators for each IP address, wherein the feature indicators are used to represent at least one of the time characteristics, geographical characteristics and behavioral characteristics of the IP address; Based on the characteristic indicators of each IP address, the multiple IP addresses are clustered to obtain the multiple first clusters; Generate a first label for each of the plurality of first clusters.

[0066] The aforementioned feature indicators are used to represent at least one of the time, geographic, and behavioral characteristics of an IP address. These feature indicators can be used to determine the characteristics of each IP address, and subsequently, multiple IP addresses can be clustered and labeled. The feature indicators can be represented by a 20-dimensional feature vector.

[0067] Among them, time features can be 5-dimensional parameters, such as hourly activity, weekly access cycle, access interval entropy, concentration of active time periods, and day-night access ratio; regional features can be 3-dimensional parameters, such as city clustering, regional density, and IP location stability; and behavioral features can be 12-dimensional parameters.

[0068] Behavioral characteristics can include 3D parameters of click preferences, such as category click share, price range preference, and brand preference entropy; 3D parameters of conversion characteristics, such as add-to-cart to purchase conversion rate, browse-to-add-to-cart conversion rate, and average order value; 3D parameters of interaction depth, such as page dwell time, visit depth, and interaction frequency; and 3D parameters of social characteristics, such as sharing rate, social fission coefficient, and community activity.

[0069] In some implementations, characteristic indicators for each IP address are extracted, operations on the IP address are collected by setting a sliding window, and the collected characteristic data is processed to obtain the characteristic indicators.

[0070] For example, using stream processing SQL as an example, a sliding window is configured through FlinkSQL, and parameters such as the number of clicks (COUNT()), average dwell time (AVG(duration)), and preferred category (MODE(category)) are collected. The sliding window is configured to last 5 minutes, with a sliding step size of 1 minute and a parallelism of 100.

[0071] Furthermore, the collected data is processed to obtain feature indicators, which can be specifically represented as: ; In the formula Feature index, where x is the collected feature data. The mean of the collected feature data. The standard deviation of the collected feature data is denoted as .

[0072] Furthermore, based on the characteristic indicators of each IP address, the multiple IP addresses are clustered to obtain the multiple first clusters. The specific process is as follows: 1. Mesh generation: The feature space is meshed (mesh size = ...). This reduces distance calculations by 60%.

[0073] 2. Core grid selection: Count the number of samples in each grid and select core grids with a density higher than the threshold.

[0074] 3. Cluster merging: Adjacent core grids are merged based on connectivity to form initial clusters.

[0075] 4. Boundary assignment: Non-core grid samples are assigned to neighboring clusters according to the nearest neighbor principle.

[0076] Clustering multiple IP addresses can be done using a pre-configured algorithm. This pre-configured algorithm is a multi-IP address clustering algorithm, such as K-Means, hierarchical clustering, or Density-Based Spatial Clustering of Applications with Noise (DBSCAN).

[0077] In some implementations, the parameters of the DBSCAN algorithm can be dynamically optimized to improve the BDSCAN algorithm, thereby further improving the clustering accuracy.

[0078] Specifically, the radius of the DBSCAN algorithm can be... Assuming automatic adjustment based on sample density, the formula is: =0.5×mean(dmin), where mean(dmin) is the mean nearest neighbor distance.

[0079] Furthermore, the minimum number of clusters after clustering can be adjusted. The minimum number can be adjusted according to the data characteristics of IP addresses. For example, it can be set to 5 when the data characteristics of IP addresses are low-dimensional (such as three-dimensional features) and 10 when the data characteristics of IP addresses are high-dimensional (such as ten-dimensional features).

[0080] For example, the clustering effects of different clustering algorithms were compared through experiments, and the clustering results of different clustering algorithms are shown in Table 1 below.

[0081] Table 1. Clustering results of different clustering algorithms

[0082] As shown in the table above, the improved DBSCAN algorithm exhibits higher accuracy, higher speed, and better noise resistance. Therefore, the improved DBSCAN algorithm can be used to cluster multiple IP addresses in this invention.

[0083] The above-mentioned generation of the first label of each of the plurality of first clusters can be achieved by calculating the central feature of the cluster based on the feature index of at least one IP address included in each first cluster, and then generating the first label of each first cluster based on the central feature.

[0084] In this embodiment of the invention, feature indicators are extracted for each IP address. These feature indicators represent at least one of the IP address's temporal, geographical, and behavioral characteristics. Based on the feature indicators of each IP address, the multiple IP addresses are clustered to obtain multiple first clusters. A first label is generated for each of the multiple first clusters. Thus, clustering of multiple IP addresses and generation of a first label for each first cluster are achieved using the feature indicators of IP addresses.

[0085] In one embodiment, the method further includes: Acquire multiple sample data sets and sample labels for each sample data set, wherein the multiple sample data sets and the multiple IP addresses belong to the same domain; Calculate the distribution divergence based on the first label of the multiple IP addresses and the sample label of the multiple sample data; If the distribution divergence is greater than a set divergence threshold, the multiple IP addresses are re-clustered based on the feature indicators of each IP address to obtain multiple second clusters, each of which includes at least one IP address. Generate a third label for each of the plurality of second clusters, and set the third label of each second cluster as the third label of the second cluster including the IP address.

[0086] The aforementioned multiple sample data and multiple IP addresses belong to the same domain. This can be achieved by extracting a portion of the data from multiple IP addresses as multiple sample data, and manually assigning sample labels to each sample data.

[0087] It should be noted that if multiple sample data points and multiple IP addresses belong to the same domain, the distribution of the first labels of the multiple IP addresses should be similar to the distribution of the multiple sample data points. Therefore, in this embodiment, the distribution divergence is calculated based on the first labels of the multiple IP addresses and the sample labels of the multiple sample data points. The distribution divergence is used to determine whether the accuracy of the first labels of the multiple IP addresses meets the requirements. If the distribution divergence is greater than a set divergence threshold, it indicates that the distribution of the first labels of the multiple IP addresses should differ from the distribution of the multiple sample data points. In this case, the first labels of the multiple IP addresses are not accurate enough, and a third label needs to be regenerated.

[0088] Furthermore, after obtaining the third label, the distribution divergence can be calculated by calculating the third labels of multiple IP addresses and the sample labels of the multiple sample data, and then a divergence threshold can be set to determine whether the labels need to be regenerated.

[0089] The distribution divergence can be the JS divergence.

[0090] This invention also provides a specific example of generating labels for IP addresses to demonstrate that the method of this invention can effectively improve label accuracy.

[0091] Specifically, a five-layer modular architecture, consisting of a data acquisition layer, an anonymization processing layer, a feature calculation layer, a tag generation layer, and a federated collaboration layer, is used to generate tags for IP addresses.

[0092] The aforementioned data acquisition layer includes a Kafka cluster with a 3-replica architecture, 100 partitions, a throughput of 100,000 msg / s, and a message loss rate of <0.1%. The data acquisition layer also includes a traffic cleaning module, which is used to filter abnormal IPs (requests per second >100), identify crawlers (based on User-Agent + behavioral characteristics), and complete data (fill in missing values ​​with mean).

[0093] The above anonymization layer includes a dynamic IP obfuscation service. The dynamic IP obfuscation service includes 20 nodes deployed distributively, with each node having an 8-core CPU + 32GB of memory, and supporting 1 million IP processing per second. The anonymization layer also includes a key management center, which is based on HashiCorp Vault, with a key update period of 1 hour, and supports key version control and auditing.

[0094] The above feature calculation layer includes a Flink cluster. The Flink cluster includes 20 nodes, with each node having an 8-core 64GB of memory, a parallelism of 100, and a state backend of RocksDB. The feature calculation layer is also used for feature storage, which can be implemented through a Redis cluster (RedisCluster), and supports 500,000 reads and writes per second.

[0095] The above label generation layer includes a clustering engine. The clustering engine is used to improve the DBSCAN algorithm for clustering, and supports 100,000 feature vector processing per second. The label generation layer also includes a drift detection service, which calculates the distribution difference every 5 minutes, and a JS divergence threshold > 0.3 triggers reclustering.

[0096] The above federated collaboration layer includes a parameter server. The parameter server can be a TensorFlow Parameter Server, including 10 parameter nodes, and supports parallel aggregation of 100+ edge nodes. The federated collaboration layer uses a lightweight aggregation protocol, based on gRPC + TLS encryption, with a compression ratio of 3:1 and an aggregation delay < 100ms.

[0097] Among them, the above different cluster services can adopt the parameter configurations shown in Table 2.

[0098] Table 2 Hardware Configuration Parameters

[0099] Furthermore, the label generation process for IP addresses includes: 1. Data collection process: User behavior occurs → SDK collection → Kafka transmission (delay < 10ms) → Traffic cleaning (filtering abnormal data).

[0100] 2. Anonymization process: Original IP → Dynamic obfuscation service (1 million IP / second) → Anonymous IP → Feature calculation layer.

[0101] 3. Feature calculation process: Anonymous IP + Behavioral data → Flink window aggregation → 20-dimensional feature vector → Feature storage.

[0102] 4. Label generation process: Feature vector → Clustering engine → Initial label → Drift detection → Dynamic update.

[0103] 5. Federal Collaboration Process: Local tagging → Parameter compression → Encrypted transmission → Global aggregation → Model update.

[0104] Among them, the divergence threshold can be the JS divergence threshold: obtained through testing with 100 sets of samples (covering e-commerce, social and financial scenarios), when the JS divergence > 0.3, it is considered that the group behavior pattern has changed significantly (accuracy decreases by > 15%), at which point the false positive rate is < 5% and the false negative rate is < 3%; therefore, 0.3 is set as the divergence threshold to trigger re-clustering.

[0105] Furthermore, the results of clustering using different clustering parameters are shown in Table 3.

[0106] Table 3 Clustering results with different clustering parameters

[0107] Table 3 `minPts` represents the radius of the cluster, and `minPts` represents the minimum number of IP addresses required for each cluster. As shown in Table 3, using... Clustering parameters of 0.5 and minPts=5 can achieve better clustering accuracy.

[0108] The method of this invention was applied to IP addresses in different fields for compatibility testing, and different compatibility results were obtained, as shown in Table 4.

[0109] Table 4: Data Adaptation Results for Different Industries

[0110] Furthermore, clustering constraints can be set to ensure that the clustering process consumes fewer resources and achieves better clustering results. Specifically, these can include: Throughput test parameters: Single-machine feature extraction parameter is 100,000 feature vectors / second; clustering engine parameter is 100,000 samples / second; anonymization service parameter is 1 million IPs / second.

[0111] Latency test parameters: end-to-end latency <200ms (from data acquisition to label generation); feature extraction latency <50ms; clustering calculation latency <100ms; federated synchronization latency <100ms.

[0112] Furthermore, the communication protocol stack parameters can be configured as follows: the transport layer uses gRPC+HTTP / 2 (supporting multiplexing); the application layer uses Protocol Buffers serialization (compression ratio 3:1); and the security layer uses TLS1.3 encryption (key exchange uses ECDHE, and authentication uses RSA).

[0113] Furthermore, key parameters were retained through model distillation, including 100 core cluster centers and 50-dimensional key feature weights; parameters removed included intermediate layer calculation parameters and historical iteration parameters; the compression effect was reduced from the original 10MB to 1MB, with a compression rate of 90%.

[0114] Please see Figure 2 , Figure 2 This is a flowchart of a tag verification method applied to a second node provided by an embodiment of the present invention, such as... Figure 2 As shown, it includes the following steps: Step 201: Receive the first labels of multiple IP addresses sent by each of the multiple nodes, wherein the first labels of the multiple IP addresses sent by each node are obtained by clustering the multiple IP addresses by the nodes; Step 202: If a first IP address exists among the multiple IP addresses, obtain at least two tag contents and the confidence level of each tag content within the node. The tag contents of the first IP address are different in the first tags of at least two nodes among the multiple nodes. The at least two tag contents are the tag contents of the first IP address within the first tags of the multiple nodes. Step 203: Calculate the score for each tag content based on the confidence level of each tag content; Step 204: Set the tag content with the highest score among the at least two tag contents as the tag content of the second tag of the first IP address; Step 205: Send the first IP address and the second tag of the first IP address to the plurality of nodes.

[0115] In this embodiment of the invention, first tags for multiple IP addresses sent by each of multiple nodes are received. The first tags for the multiple IP addresses sent by each node are obtained through node clustering of the multiple IP addresses. If a first IP address exists among the multiple IP addresses, at least two tag contents and the confidence score of each tag content within a node are obtained. The tag contents of the first IP address are different in the first tags of at least two nodes among the multiple nodes; the at least two tag contents are the tag contents of the first IP address within the first tags of the multiple nodes. A score is calculated for each tag content based on its confidence score. The tag content with the highest score among the at least two tag contents is set as the tag content of the second tag for the first IP address. The first IP address and its second tag are then sent to the multiple nodes. Thus, even when a first IP address exists, the scores of different tag contents are calculated based on confidence scores, and the tag content of the second tag is determined based on these scores, thereby effectively improving the accuracy of the second tag.

[0116] It should be understood that the first IP address has only one tag content within a node, and there is also a confidence level corresponding to that tag content. In this invention, the score of each tag content is calculated through the confidence level, so that the tag content of the second tag can be determined through the score.

[0117] Taking the first IP address as an example, if the first IP address has different label content for at least two nodes across multiple addresses, then the first IP address needs optimization. For instance, the first IP address's first label includes two different label contents. In the second node, a score is calculated for each label content based on the confidence level of the node containing the different label contents. The label content with the highest score is then selected as the second label content, resulting in a more accurate second label. The second label is then sent to multiple nodes (including the first node) so that all nodes can update the first label of the first IP address to the second label, improving label accuracy.

[0118] In some implementations, the score for each tag content can be calculated by averaging the confidence level of each tag content within a node and using the average value as the score for the tag content.

[0119] In one embodiment, the label content of the second label is calculated using the following formula: ; In the formula, label represents the label content of the second label. This represents the rating of tag content k, where n represents the number of tags, and w iThe label represents the confidence level of the first IP address at the i-th node, I() represents the indicator function, and label i =k indicates that the label content of the i-th node is k.

[0120] In this embodiment of the invention, the scores of different tag contents are calculated by the above formula, so that the tag content with the highest score among the tag contents can be set as the tag content of the second tag.

[0121] In this way, the above method is used to cluster and generate labels for multiple IP addresses of multiple nodes, achieving cross-domain label consistency. For example, experimental tests show that when 10 nodes collaborate, the label consistency reaches 92%, which is 15% higher than independent clustering. Communication costs are reduced, with each round of parameter transmission being 1MB and bandwidth usage being <1Mbps, making it suitable for edge scenarios. Aggregation latency is also reduced, with aggregation latency of <100ms when 100 nodes participate, meeting real-time requirements.

[0122] Please see Figure 3 , Figure 3 This is a structural diagram of a tag verification device applied to a first node according to an embodiment of the present invention, as shown below. Figure 3 As shown, the label verification device 300 includes: The first clustering module 301 is used to cluster multiple Internet Protocol (IP) addresses to obtain multiple first clusters and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster. The determination module 302 is used to determine the node state vector of the first node, wherein the node state vector is used to characterize at least one of the clustering accuracy of the plurality of first clusters, the communication latency of the first node, and the computational load. The setting module 303 is used to set multiple node action parameters of the first node, wherein different node action parameters are used to characterize different strategies for the first node to send tags. The first calculation module 304 is used to calculate the reward value of each node action parameter based on the node state vector and the multiple node action parameters. Sending module 305 is used to send a first label of the plurality of IP addresses to the second node based on the target action parameter, wherein the target action parameter is the node action parameter with the largest reward value among the plurality of node action parameters; The receiving module 306 is used to receive a first IP address and a second tag of the first IP address. The first IP address has different tag content in at least two nodes among multiple nodes. The tag content of the second tag is the tag content with the highest score among at least two contents. The at least two tag contents are the tag contents of the first IP address in the first tags of the multiple nodes. The score is calculated by the confidence level of each tag content in the node.

[0123] In one embodiment, the reward value of each node action parameter is calculated using the following formula: ; in the formula This represents the reward value for each node's action parameter, 0. T Represents the coefficient matrix. Let R() represent the discount factor, R() represent the calculation formula, st represent the node state vector of the first node, and a() represent the discount factor. t This represents the node action parameters.

[0124] In one embodiment, the label verification device 300 further includes: The first acquisition module is used to acquire the plurality of IP addresses, wherein each IP address includes a network prefix and a host identifier. The first generation module is used to generate a random salt value for each IP address based on the timestamp of each IP address being acquired. An encryption module is used to encrypt the host identifier of the corresponding IP address based on a random salt value for each IP address, thereby obtaining multiple encrypted IP addresses; The first clustering module 301 includes: The first clustering unit is used to cluster the multiple encrypted IP addresses to obtain multiple first clusters.

[0125] In one embodiment, the first clustering module 301 includes: An extraction unit is used to extract feature indicators for each IP address, wherein the feature indicators represent at least one of the IP address’s time characteristics, geographic characteristics, and behavioral characteristics. The second clustering unit is used to cluster the multiple IP addresses based on the feature indicators of each IP address to obtain the multiple first clusters; A generation unit is used to generate a first label for each of the plurality of first clusters.

[0126] In one embodiment, the label verification device 300 further includes: The second acquisition module is used to acquire multiple sample data and a sample label for each sample data in the multiple sample data, wherein the multiple sample data and the multiple IP addresses are data in the same domain; The second calculation module is used to calculate the distribution divergence based on the first label of the multiple IP addresses and the sample label of the multiple sample data. The second clustering module is used to re-cluster the multiple IP addresses based on the feature indicators of each IP address when the distribution divergence is greater than a set divergence threshold, to obtain multiple second clusters, each second cluster including at least one IP address. The second generation module is used to generate a third label for each of the plurality of second clusters, and set the third label of each second cluster as a third label including the IP address of the second cluster.

[0127] The tag verification device for the first node provided in this embodiment of the invention can implement the various processes of the various embodiments of the tag verification method for the first node described above. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0128] It should be noted that the tag verification device applied to the first node in the embodiments of the present invention can be a device, or it can be a component, integrated circuit, or chip in an electronic device.

[0129] Please see Figure 4 , Figure 4 This is a structural diagram of a tag verification device applied to a second node according to an embodiment of the present invention, as shown below. Figure 4 As shown, the label verification device 400 includes: The receiving module 401 is used to receive the first labels of multiple IP addresses sent by each of the multiple nodes, wherein the first labels of the multiple IP addresses sent by each node are obtained by clustering the multiple IP addresses by the nodes. The acquisition module 402 is used to acquire at least two tag contents and the confidence level of each tag content within a node when a first IP address exists among the plurality of IP addresses. The tag contents of the first tag of the first IP address are different in at least two of the plurality of nodes. The at least two tag contents are the tag contents of the first IP address within the first tag of the plurality of nodes. The calculation module 403 is used to calculate the score of each tag content based on the confidence level of each tag content; Setting module 404 is used to set the tag content with the highest score among the at least two tag contents as the tag content of the second tag of the first IP address; The sending module 405 is used to send the first IP address and the second tag of the first IP address to the plurality of nodes.

[0130] In one embodiment, the label content of the second label is calculated using the following formula: ; In the formula, label represents the label content of the second label. This represents the rating of tag content k, where n represents the number of tags, and w i The label represents the confidence level of the first IP address at the i-th node, I() represents the indicator function, and label i =k indicates that the label content of the i-th node is k.

[0131] The tag verification device for the second node provided in this embodiment of the invention can implement the various processes of the various embodiments of the tag verification method for the second node described above. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0132] It should be noted that the tag verification device applied to the second node in the embodiments of the present invention can be a device, or it can be a component, integrated circuit, or chip in an electronic device.

[0133] This invention also provides an electronic device applied to a first node, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the above-described functionality. Figure 1 The various processes of the label confirmation method embodiment shown herein can achieve the same technical effect, and will not be described again here to avoid repetition.

[0134] For details, see Figure 5 As shown, this embodiment of the invention also provides an electronic device, including a bus 501, a transceiver 502, an antenna 503, a bus interface 504, a processor 505, and a memory 506.

[0135] The processor 505 is configured to cluster multiple Internet Protocol (IP) addresses to obtain multiple first clusters and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster. The processor 505 is further configured to determine the node state vector of the first node, the node state vector being used to characterize at least one of the clustering accuracy of the plurality of first clusters, the communication latency of the first node, and the computational load. The processor 505 is also configured to set multiple node action parameters for the first node, wherein different node action parameters are used to characterize different strategies for the first node to send tags. The processor 505 is further configured to calculate the reward value of each node action parameter based on the node state vector and the plurality of node action parameters. The transceiver 502 is used to send a first label of the plurality of IP addresses to the second node based on the target action parameter, wherein the target action parameter is the node action parameter with the largest reward value among the plurality of node action parameters; The transceiver 502 is further configured to receive a first IP address and a second tag of the first IP address, wherein the first IP address has different tag content in at least two of the multiple nodes, and the tag content of the second tag is the tag content with the highest score among at least two contents, wherein the at least two tag contents are the tag contents of the first IP address in the first tags of the multiple nodes, and the score is calculated by the confidence level of each tag content in the node.

[0136] In one embodiment, the reward value of each node action parameter is calculated using the following formula: ; in the formula This represents the reward value for each node's action parameter, 0. T Represents the coefficient matrix. Let R() represent the discount factor, R() represent the calculation formula, st represent the node state vector of the first node, and a() represent the discount factor. t This represents the node action parameters.

[0137] In one embodiment, the transceiver 502 is further configured to obtain the plurality of IP addresses, each of the plurality of IP addresses including a network prefix and a host identifier; The processor 505 is further configured to generate a random salt value for each IP address based on the timestamp of each IP address being acquired; The processor 505 is further configured to encrypt the host identifier of the corresponding IP address based on the random salt value of each IP address, thereby obtaining multiple encrypted IP addresses; The process of clustering multiple IP addresses yields multiple first clusters, including: The multiple encrypted IP addresses are clustered to obtain multiple first clusters.

[0138] In one embodiment, clustering multiple IP addresses to obtain multiple first clusters includes: Extract the feature indicators for each IP address, wherein the feature indicators are used to represent at least one of the time characteristics, geographical characteristics and behavioral characteristics of the IP address; Based on the characteristic indicators of each IP address, the multiple IP addresses are clustered to obtain the multiple first clusters; Generate a first label for each of the plurality of first clusters.

[0139] In one embodiment, the transceiver 502 is further configured to acquire multiple sample data and a sample label for each sample data in the multiple sample data, wherein the multiple sample data and the multiple IP addresses are data in the same domain; The processor 505 is also used to calculate the distribution divergence based on the first label of the plurality of IP addresses and the sample label of the plurality of sample data; The processor 505 is further configured to, when the distribution divergence is greater than a set divergence threshold, re-cluster the multiple IP addresses based on the feature indicators of each IP address to obtain multiple second clusters, each second cluster including at least one IP address. The processor 505 is further configured to generate a third label for each of the plurality of second clusters, and to set the third label of each second cluster as a third label including the IP address of the second cluster.

[0140] exist Figure 5 In this document, a bus architecture (represented by bus 501) is used. Bus 501 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 505 and memory represented by memory 506. Bus 501 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 504 provides an interface between bus 501 and transceiver 502. Transceiver 502 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 505 is transmitted over a wireless medium via antenna 503, which further receives data and transmits it to processor 505.

[0141] Processor 505 manages bus 501 and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory 506 can be used to store data used by processor 505 during operation.

[0142] Optionally, the processor 505 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a graphics processing unit (GPU).

[0143] This invention also provides an electronic device applied to a second node, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the above-described functionality. Figure 2 The various processes of the label confirmation method embodiment shown herein can achieve the same technical effect, and will not be described again here to avoid repetition.

[0144] For details, see Figure 6 As shown, this embodiment of the invention also provides an electronic device, including a bus 601, a transceiver 602, an antenna 603, a bus interface 604, a processor 605, and a memory 606.

[0145] The transceiver 602 is used to receive first tags of multiple IP addresses sent by each of the multiple nodes, wherein the first tags of the multiple IP addresses sent by each node are obtained by clustering the multiple IP addresses by the nodes. The processor 605 is configured to, when a first IP address exists among the plurality of IP addresses, obtain at least two tag contents and the confidence level of each tag content within a node, wherein the tag contents of the first IP address are different in the first tags of at least two of the plurality of nodes, and the at least two tag contents are the tag contents of the first IP address within the first tags of the plurality of nodes. The processor 605 is also configured to calculate a score for each tag content based on the confidence level of each tag content; The processor 605 is further configured to set the tag content with the highest score among the at least two tag contents as the tag content of the second tag of the first IP address; The transceiver 602 is also used to send the first IP address and the second tag of the first IP address to the plurality of nodes.

[0146] In one embodiment, the label content of the second label is calculated using the following formula: ; In the formula, label represents the label content of the second label. This represents the rating of tag content k, where n represents the number of tags, and w i The label represents the confidence level of the first IP address at the i-th node, I() represents the indicator function, and label i =k indicates that the label content of the i-th node is k.

[0147] exist Figure 6 In this document, a bus architecture (represented by bus 601) is used. Bus 601 may include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 605 and memory represented by memory 506. Bus 601 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 604 provides an interface between bus 601 and transceiver 602. Transceiver 602 may be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 605 is transmitted over a wireless medium via antenna 603, which further receives data and transmits data to processor 605.

[0148] Processor 605 manages bus 601 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 606 can be used to store data used by processor 605 during operation.

[0149] Optionally, the processor 605 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a graphics processing unit (GPU).

[0150] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the above-described functions. Figure 1 or Figure 2The various processes of the corresponding label verification method embodiments achieve the same technical effect, and will not be described again here to avoid repetition. The computer-readable storage medium mentioned includes, for example, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0151] The present invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described... Figure 1 or Figure 2 The various processes of the corresponding label confirmation method embodiments can achieve the same technical effect, and will not be described again here to avoid repetition.

[0152] In the embodiments of this invention, the terms "first," "second," etc., are used to distinguish similar object parameters and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected object parameters, such as A and / or B and / or C, representing seven possibilities: A alone, B alone, C alone, both A and B present, both B and C present, both A and C present, and A, B, and C present.

[0153] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0154] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of the various embodiments of this application.

[0155] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A tag confirmation method, applied to a first node, characterized in that, include: Multiple Internet Protocol (IP) addresses are clustered to obtain multiple first clusters and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster. Determine the node state vector of the first node, wherein the node state vector is used to characterize at least one of the clustering accuracy of the plurality of first clusters, the communication latency of the first node, and the computational load; Multiple node action parameters are set for the first node, and different node action parameters are used to characterize different strategies for the first node to send tags. Based on the node state vector and the multiple node action parameters, calculate the reward value for each node action parameter. Based on the target action parameter, a first tag of the plurality of IP addresses is sent to the second node, wherein the target action parameter is the node action parameter with the largest reward value among the plurality of node action parameters; The system receives a first IP address and a second tag for the first IP address. The first IP address has different tag content in at least two nodes across multiple nodes. The tag content of the second tag is the tag content with the highest score among at least two contents. The at least two tag contents are the tag content of the first IP address within the first tag across the multiple nodes. The score is calculated by the confidence level of each tag content within the node.

2. The method as described in claim 1, characterized in that, The reward value for each node's action parameter is calculated using the following formula: ; in the formula This represents the reward value for each node's action parameter, 0. T Represents the coefficient matrix. Let R() represent the discount factor, R() represent the calculation formula, st represent the node state vector of the first node, and a() represent the discount factor. t This represents the node action parameters.

3. The method as described in claim 1 or 2, characterized in that, Before clustering multiple IP addresses to obtain multiple first clusters, the method further includes: Obtain the plurality of IP addresses, each of which includes a network prefix and a host identifier; A random salt value is generated for each IP address based on the timestamp obtained for each IP address; The host identifier of the corresponding IP address is encrypted based on the random salt value of each IP address to obtain multiple encrypted IP addresses; The process of clustering multiple IP addresses yields multiple first clusters, including: The multiple encrypted IP addresses are clustered to obtain multiple first clusters.

4. The method as described in claim 1 or 2, characterized in that, The process of clustering multiple IP addresses yields multiple first clusters, including: Extract the feature indicators for each IP address, wherein the feature indicators are used to represent at least one of the time characteristics, geographical characteristics and behavioral characteristics of the IP address; Based on the characteristic indicators of each IP address, the multiple IP addresses are clustered to obtain the multiple first clusters; Generate a first label for each of the plurality of first clusters.

5. The method as described in claim 1 or 2, characterized in that, The method further includes: Acquire multiple sample data sets and sample labels for each sample data set, wherein the multiple sample data sets and the multiple IP addresses belong to the same domain; Calculate the distribution divergence based on the first label of the multiple IP addresses and the sample label of the multiple sample data; If the distribution divergence is greater than a set divergence threshold, the multiple IP addresses are re-clustered based on the feature indicators of each IP address to obtain multiple second clusters, each of which includes at least one IP address. Generate a third label for each of the plurality of second clusters, and set the third label of each second cluster as the third label of the second cluster including the IP address.

6. A tag confirmation method, applied to a second node, characterized in that, include: The system receives first labels for multiple IP addresses sent by each of multiple nodes, wherein the first labels for the multiple IP addresses sent by each node are obtained by clustering the multiple IP addresses by the nodes. In the case where a first IP address exists among the plurality of IP addresses, at least two tag contents are obtained, as well as the confidence level of each tag content within a node. The tag contents of the first IP address are different in the first tags of at least two nodes among the plurality of nodes. The at least two tag contents are the tag contents of the first IP address within the first tags of the plurality of nodes. A score for each tag content is calculated based on the confidence level of each tag content; Set the tag content with the highest score among the at least two tag contents as the tag content of the second tag of the first IP address; Send the first IP address and the second tag of the first IP address to the plurality of nodes.

7. The method as described in claim 6, characterized in that, The content of the second tag is calculated using the following formula: ; In the formula, label represents the label content of the second label. This represents the rating of tag content k, where n represents the number of tags, and w i The label represents the confidence level of the first IP address at the i-th node, I() represents the indicator function, and label i =k indicates that the label content of the i-th node is k.

8. A tag verification device, applied to a first node, characterized in that, include: The first clustering module is used to cluster multiple Internet Protocol (IP) addresses to obtain multiple first clusters and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster. The determination module is used to determine the node state vector of the first node, wherein the node state vector is used to characterize at least one of the clustering accuracy of the plurality of first clusters, the communication latency of the first node, and the computational load. The setting module is used to set multiple node action parameters of the first node, wherein different node action parameters are used to characterize different strategies for the first node to send tags. The first calculation module is used to calculate the reward value of each node action parameter based on the node state vector and the multiple node action parameters. The sending module is used to send a first label of the plurality of IP addresses to the second node based on the target action parameter, wherein the target action parameter is the node action parameter with the largest reward value among the plurality of node action parameters; A receiving module is used to receive a first IP address and a second tag of the first IP address. The first IP address has different tag content in at least two of the multiple nodes. The tag content of the second tag is the tag content with the highest score among at least two contents. The at least two tag contents are the tag contents of the first IP address in the first tags of the multiple nodes. The score is calculated by the confidence level of each tag content in the node.

9. A tag verification device, applied to a second node, characterized in that, include: The receiving module is used to receive the first labels of multiple IP addresses sent by each of the multiple nodes, wherein the first labels of the multiple IP addresses sent by each node are obtained by clustering the multiple IP addresses by the nodes. The acquisition module is used to acquire at least two tag contents and the confidence level of each tag content within a node when a first IP address exists among the plurality of IP addresses. The tag contents of the first tag of the first IP address are different in at least two of the plurality of nodes. The at least two tag contents are the tag contents of the first IP address within the first tag of the plurality of nodes. The calculation module is used to calculate the score of each tag content based on the confidence level of each tag content; The setting module is used to set the tag content with the highest score among the at least two tag contents as the tag content of the second tag of the first IP address; The sending module is used to send the first IP address and the second tag of the first IP address to the plurality of nodes.

10. An electronic device applied to a first node, characterized in that, Including transceivers and processors, The processor is configured to cluster multiple Internet Protocol (IP) addresses to obtain multiple first clusters and a first label for each first cluster. Each first cluster includes at least one IP address, and the first label for each first cluster is the first label of the IP address included in the first cluster. The processor is further configured to determine the node state vector of the first node, the node state vector being used to characterize at least one of the clustering accuracy of the plurality of first clusters, the communication latency of the first node, and the computational load. The processor is further configured to set multiple node action parameters for the first node, wherein different node action parameters are used to characterize different strategies for the first node to send tags. The processor is further configured to calculate the reward value for each node action parameter based on the node state vector and the plurality of node action parameters. The transceiver is used to send a first tag of the plurality of IP addresses to the second node based on a target action parameter, wherein the target action parameter is the node action parameter with the largest reward value among the plurality of node action parameters. The transceiver is further configured to receive a first IP address and a second tag for the first IP address, wherein the first IP address has different tag content in at least two of the multiple nodes, and the tag content of the second tag is the tag content with the highest score among at least two contents, wherein the at least two tag contents are the tag contents of the first IP address in the first tags of the multiple nodes, and the score is calculated by the confidence level of each tag content in the node.

11. An electronic device applied to a second node, characterized in that, Including transceivers and processors, The transceiver is used to receive first tags of multiple IP addresses sent by each of the multiple nodes, wherein the first tags of the multiple IP addresses sent by each node are obtained by clustering the multiple IP addresses by the nodes. The processor is configured to, when a first IP address exists among the plurality of IP addresses, acquire at least two tag contents and the confidence level of each tag content within a node, wherein the tag contents of the first IP address are different in the first tags of at least two of the plurality of nodes, and the at least two tag contents are the tag contents of the first IP address within the first tags of the plurality of nodes. The processor is also configured to calculate a score for each tag content based on the confidence level of each tag content; The processor is further configured to set the tag content with the highest score among the at least two tag contents as the tag content of the second tag of the first IP address; The transceiver is also used to send the first IP address and a second tag of the first IP address to the plurality of nodes.

12. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the tag verification method as described in any one of claims 1 to 7.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the tag verification method as described in any one of claims 1 to 7.

14. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the tag verification method as described in any one of claims 1 to 7.