Method and apparatus for identifying mobile network assets, and device and storage medium

By clustering and joint analysis of network data flows in multiple historical time periods, network addresses that meet the conditions of mobile network assets are identified, which solves the problem of low identification accuracy in the prior art and achieves higher identification accuracy.

WO2025130547A1PCT designated stage expired Publication Date: 2025-06-26CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/135390
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-11-28
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

When identifying mobile network assets in the prior art, the identification accuracy is low, and it is easy to identify invalid assets as mobile network assets, misidentify non-mobile network assets as mobile network assets, and miss mobile network assets.

Method used

By obtaining network data flows in multiple historical time periods, clustering processes are performed, and the clustering cluster contains network addresses with the same address type. Combined with the joint analysis of multiple historical time periods, the network address that meets the conditions for mobile network assets is determined, and the results of mobile network assets are generated.

Benefits of technology

Improve the accuracy of identifying mobile network assets, avoid the situation of misidentifying invalid or non-mobile network assets as mobile network assets, and enhance the ability to identify mobile network assets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135390_26062025_PF_FP_ABST
    Figure CN2024135390_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for identifying mobile network assets, and a device and a storage medium, which can be applied to fields such as communications, and are used for solving the problem of low identification accuracy in the identification of mobile network assets. The method at least comprises: acquiring a plurality of network data flows generated within each of a plurality of historical time periods, wherein each network data flow comprises a network address; for a plurality of network data flows generated within each historical time period, executing the following operation: performing clustering processing on the plurality of network data flows on the basis of network addresses comprised in the plurality of respective network data flows, so as to obtain a plurality of clusters, wherein the address types of network addresses comprised in each of the respective network data flows contained in each cluster are the same; and on the basis of a plurality of clusters obtained for each of the plurality of historical time periods, determining at least one network address, which meets a mobile network asset condition, among network addresses comprised in each of the respective network data flows, and generating a mobile network asset identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment and storage medium for identifying mobile network assets

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 22, 2023, with application number 202311779255.8 and application name "A method, device, equipment and storage medium for identifying mobile network assets", the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for identifying mobile network assets. Background Art

[0004] With the continuous advancement of technology, more and more devices can serve as network addresses, such as source or destination addresses, to enable the sending and receiving of network data. The source or destination address represented by a device can be a mobile network asset, allowing the device to move from one network resource system to another, while maintaining the same source or destination address. This allows the device to maintain connectivity while moving, enabling roaming across different network segments. By identifying mobile network assets, a network resource whitelist can be provided to network threat intelligence systems, preventing them from mistakenly identifying mobile network assets as abnormal network addresses.

[0005] In related technologies, mobile network assets are typically identified based on the network address's home organization or the network address's Autonomous System Number (ASN). For example, if the network address's home organization is a carrier, then the network address is identified as a mobile network asset. Alternatively, if the network address's ASN is associated with a carrier, then the network address is identified as a mobile network asset.

[0006] However, due to the large number of network addresses, some may not be used in a standardized manner or managed in a timely manner. For example, some network addresses with a carrier as their home institution may have been abandoned or expired long ago but have not been cleared. Another example is that some network addresses have ASNs that are not associated with a carrier but are allocated for use as mobile network assets. Another example is that some network addresses with a carrier as their home institution may be allocated for use as non-mobile network assets.

[0007] It can be seen that when identifying mobile network assets in related technologies, it is easy to identify expired assets as mobile network assets; it is also easy to mistakenly identify non-mobile network assets as mobile network assets; it is also easy to miss mobile network assets, etc., that is, the accuracy of identifying mobile network assets is low. Summary of the Invention

[0008] Embodiments of the present application provide a method, apparatus, device, and storage medium for identifying mobile network assets, which are used to solve the problem of low accuracy in identifying mobile network assets.

[0009] In a first aspect, a method for identifying mobile network assets is provided, comprising:

[0010] Acquire multiple network data flows generated in multiple historical time periods; wherein the network data flows include network addresses;

[0011] For each of the plurality of network data flows generated within the historical time period, the following operations are performed: clustering the plurality of network data flows based on the network addresses included in each of the plurality of network data flows to obtain a plurality of clusters; wherein the network addresses included in each of the network data flows included in each cluster have the same address type;

[0012] Based on the multiple clusters obtained for the multiple historical time periods, at least one network address that meets the mobile network asset condition among the network addresses included in each network data flow is determined to generate a mobile network asset identification result.

[0013] Optionally, obtaining multiple network data flows generated in multiple historical time periods includes:

[0014] Obtain multiple initial data streams generated in multiple historical time periods;

[0015] For each of the multiple initial data streams generated within the historical time period, perform the following operations:

[0016] Based on a pre-stored data filtering strategy, the multiple initial data streams are filtered to obtain multiple filtered data streams;

[0017] Data deduplication is performed on at least two filtered data flows including the same network address among the multiple filtered data flows to obtain the multiple network data flows.

[0018] Optionally, the initial data stream further includes a data communication protocol;

[0019] The filtering of the multiple initial data streams based on the pre-stored data filtering strategy to obtain multiple filtered data streams includes:

[0020] When determining that a first data flow including a specified network address exists among the multiple initial data flows, deleting the first data flow from the multiple initial data flows;

[0021] When determining that a second data stream including a data communication protocol other than the preset communication protocol exists in the plurality of initial data streams, deleting the second data stream from the plurality of initial data streams;

[0022] The multiple filtered data streams are obtained based on the multiple initial data streams after deleting the first data stream and / or the second data stream.

[0023] Optionally, clustering the multiple network data flows based on the network addresses respectively included in the multiple network data flows to obtain multiple clusters includes:

[0024] Determining the flow similarity between the corresponding two network data flows based on the network addresses respectively included in each of the two network data flows;

[0025] Based on the obtained similarities of each flow, multiple rounds of clustering processing are performed on the multiple network data flows to obtain multiple clusters; wherein each round of clustering processing includes:

[0026] Acquire multiple intermediate clusters from a previous round; wherein, if a previous round of clustering processing exists, the multiple intermediate clusters from the previous round are multiple current intermediate clusters obtained after the previous round of clustering processing; if no previous round of clustering processing exists, the multiple intermediate clusters from the previous round are the multiple network data flows;

[0027] Based on the flow similarity between each network data flow included in each two intermediate clusters of the previous round, clustering processing is performed on the multiple intermediate clusters of the previous round to obtain multiple current intermediate clusters.

[0028] Optionally, clustering the multiple intermediate clusters of the previous round based on the flow similarity between the network data flows respectively included in every two intermediate clusters of the previous round to obtain multiple current intermediate clusters includes:

[0029] For every two intermediate clusters in the previous round, perform the following operations:

[0030] Determining a maximum flow similarity based on flow similarities between each network data flow included in one intermediate cluster of the previous round and each network data flow included in another intermediate cluster of the previous round;

[0031] The maximum flow similarity is used as the cluster similarity between the corresponding two intermediate clusters of the previous round;

[0032] Based on the cluster similarity between every two intermediate clusters in the previous round, clustering processing is performed on the multiple intermediate clusters in the previous round to obtain multiple current intermediate clusters.

[0033] Optionally, the determining, based on the multiple clusters obtained for the multiple historical time periods, at least one network address that meets the mobile network asset condition among the network addresses included in each network data flow, and generating the mobile network asset identification result includes:

[0034] For the multiple clusters obtained in each historical time period, perform the following operations:

[0035] The plurality of clusters obtained in historical time periods other than the historical time period in which the cluster is located are respectively used as the plurality of other clusters obtained for each other historical time period; when it is determined that there are other clusters matching the cluster among the plurality of other clusters obtained for each other historical time period, the number of historical time periods in which the other clusters matching the cluster are located is counted;

[0036] When it is determined that the number of the obtained historical time periods reaches a first number threshold, the network address included in each network data flow included in the cluster is used as a mobile network asset to obtain a mobile network asset identification result.

[0037] Optionally, taking the network address included in each network data flow included in the cluster as a mobile network asset to obtain a mobile network asset identification result includes:

[0038] Determining a geographical location corresponding to a network address included in each network data flow included in the cluster;

[0039] When the number of identical geographical locations among the obtained geographical locations is determined to be lower than a second number threshold, the network addresses included in the network data flows included in the clusters are used as mobile network assets to obtain a mobile network asset identification result.

[0040] Optionally, after obtaining the mobile network asset identification result, the method further includes:

[0041] Setting a mobile asset validity period for each network address included in the mobile network asset identification result; wherein the mobile asset validity period represents: a period during which the corresponding network address is used as a mobile network asset;

[0042] An expiration status flag is added to each of the network addresses included in each of the multiple network data flows generated in the multiple historical time periods, except for the network addresses included in the mobile network asset identification result; wherein the expiration status flag is used to indicate that the corresponding network address has not been used as a mobile network asset.

[0043] In a second aspect, a device for identifying mobile network assets is provided, comprising:

[0044] Acquisition module: used to acquire multiple network data flows generated in multiple historical time periods; wherein the network data flows are used to include network addresses, and the network addresses include: source addresses and destination addresses;

[0045] a processing module configured to perform the following operations on the plurality of network data flows generated in each of the historical time periods: clustering the plurality of network data flows based on the network addresses included in each of the plurality of network data flows to obtain a plurality of clusters; wherein the network addresses included in each of the network data flows included in each cluster have the same address type;

[0046] The processing module is further configured to determine, based on the plurality of clusters obtained for the plurality of historical time periods, at least one network address satisfying a mobile network asset condition among the network addresses included in each network data flow, and generate a mobile network asset identification result.

[0047] Optionally, the acquisition module is specifically configured to:

[0048] Obtain multiple initial data streams generated in multiple historical time periods;

[0049] For each of the multiple initial data streams generated within the historical time period, perform the following operations:

[0050] Based on a pre-stored data filtering strategy, the multiple initial data streams are filtered to obtain multiple filtered data streams;

[0051] Data deduplication is performed on at least two filtered data flows including the same network address among the multiple filtered data flows to obtain the multiple network data flows.

[0052] Optionally, the initial data stream further includes a data communication protocol;

[0053] The acquisition module is specifically used for:

[0054] When determining that a first data flow including a specified network address exists among the multiple initial data flows, deleting the first data flow from the multiple initial data flows;

[0055] When determining that a second data stream including a data communication protocol other than the preset communication protocol exists in the plurality of initial data streams, deleting the second data stream from the plurality of initial data streams;

[0056] The multiple filtered data streams are obtained based on the multiple initial data streams after deleting the first data stream and / or the second data stream.

[0057] Optionally, the processing module is specifically configured to:

[0058] Determining the flow similarity between the corresponding two network data flows based on the network addresses respectively included in each of the two network data flows;

[0059] Based on the obtained similarities of each flow, multiple rounds of clustering processing are performed on the multiple network data flows to obtain multiple clusters; wherein each round of clustering processing includes:

[0060] Acquire multiple intermediate clusters from a previous round; wherein, if a previous round of clustering processing exists, the multiple intermediate clusters from the previous round are multiple current intermediate clusters obtained after the previous round of clustering processing; if no previous round of clustering processing exists, the multiple intermediate clusters from the previous round are the multiple network data flows;

[0061] Based on the flow similarity between each network data flow included in each two intermediate clusters of the previous round, clustering processing is performed on the multiple intermediate clusters of the previous round to obtain multiple current intermediate clusters.

[0062] Optionally, the processing module is specifically configured to:

[0063] For every two intermediate clusters in the previous round, perform the following operations:

[0064] Determining a maximum flow similarity based on flow similarities between each network data flow included in one intermediate cluster of the previous round and each network data flow included in another intermediate cluster of the previous round;

[0065] The maximum flow similarity is used as the cluster similarity between the corresponding two intermediate clusters of the previous round;

[0066] Based on the cluster similarity between every two intermediate clusters in the previous round, clustering processing is performed on the multiple intermediate clusters in the previous round to obtain multiple current intermediate clusters.

[0067] Optionally, the processing module is specifically configured to:

[0068] For the multiple clusters obtained in each historical time period, perform the following operations:

[0069] The plurality of clusters obtained in historical time periods other than the historical time period in which the cluster is located are respectively used as the plurality of other clusters obtained for each other historical time period; when it is determined that there are other clusters matching the cluster among the plurality of other clusters obtained for each other historical time period, the number of historical time periods in which the other clusters matching the cluster are located is counted;

[0070] When it is determined that the number of the obtained historical time periods reaches a first number threshold, the network address included in each network data flow included in the cluster is used as a mobile network asset to obtain a mobile network asset identification result.

[0071] Optionally, the processing module is specifically configured to:

[0072] Determining a geographical location corresponding to a network address included in each network data flow included in the cluster;

[0073] When the number of identical geographical locations among the obtained geographical locations is determined to be lower than a second number threshold, the network addresses included in the network data flows included in the clusters are used as mobile network assets to obtain a mobile network asset identification result.

[0074] Optionally, the processing module is further configured to:

[0075] After obtaining the mobile network asset identification result, respectively setting a mobile asset validity period for each network address included in the mobile network asset identification result; wherein the mobile asset validity period represents: the period during which the corresponding network address is used as a mobile network asset;

[0076] An expiration status flag is added to each of the network addresses included in each of the multiple network data flows generated in the multiple historical time periods, except for the network addresses included in the mobile network asset identification result; wherein the expiration status flag is used to indicate that the corresponding network address has not been used as a mobile network asset.

[0077] According to a third aspect, a computer program product is provided, comprising a computer program, which implements the method according to the first aspect when executed by a processor.

[0078] According to a fourth aspect, a computer device is provided, comprising:

[0079] a memory for storing program instructions;

[0080] The processor is configured to call the program instructions stored in the memory and execute the method described in the first aspect according to the obtained program instructions.

[0081] In a fifth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method described in the first aspect.

[0082] In an embodiment of the present application, mobile network assets are identified by analyzing the network data streams generated in real time during a historical time period, thereby avoiding the situation where abandoned or invalid network addresses are identified as mobile network assets, and to a certain extent improving the accuracy of identifying mobile network assets.

[0083] Furthermore, by analyzing multiple network data streams generated in multiple historical time periods, it is possible to analyze from a macro perspective whether a network address has the characteristics of a mobile network asset, thereby identifying at least one network address that meets the mobile network asset conditions, further improving the accuracy of identifying mobile network assets.

[0084] Furthermore, clustering processing is performed on multiple network data streams generated in each historical time period, and network data streams including network addresses of the same address type can be aggregated into a cluster cluster, so that network addresses that may be used as mobile network assets can be aggregated into a cluster cluster. Combined with the joint analysis of multiple historical time periods, mobile network assets can be accurately identified. For example, mobile network assets that have no association with operator organizations can be identified, and those that belong to operator organizations but are allocated as non-mobile network assets will not be identified as mobile network assets, thereby improving identification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] FIG1 is an application scenario of the method for identifying mobile network assets provided by an embodiment of the present application;

[0086] FIG2 is a flow chart of a method for identifying mobile network assets provided in an embodiment of the present application;

[0087] FIG3 is a schematic diagram showing a first principle of a method for identifying mobile network assets provided by an embodiment of the present application;

[0088] FIG4 is a second schematic diagram of a method for identifying mobile network assets provided by an embodiment of the present application;

[0089] FIG5A is a third schematic diagram of a method for identifying mobile network assets provided by an embodiment of the present application;

[0090] FIG5B is a fourth schematic diagram of a method for identifying mobile network assets provided in an embodiment of the present application;

[0091] FIG6A is a fifth schematic diagram of a method for identifying mobile network assets provided by an embodiment of the present application;

[0092] FIG6B is a sixth schematic diagram of a method for identifying mobile network assets provided by an embodiment of the present application;

[0093] FIG7A is a seventh schematic diagram of a method for identifying mobile network assets provided by an embodiment of the present application;

[0094] FIG7B is a schematic diagram showing a method for identifying mobile network assets according to an embodiment of the present application; FIG.

[0095] FIG8 is a first structural diagram of an apparatus for identifying mobile network assets according to an embodiment of the present application;

[0096] FIG9 is a second structural diagram of the apparatus for identifying mobile network assets provided in an embodiment of the present application. DETAILED DESCRIPTION

[0097] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0098] Some of the terms used in the embodiments of the present application are explained below to facilitate understanding by those skilled in the art.

[0099] (1) Threat Intelligence

[0100] Threat intelligence, also known as security intelligence or security threat intelligence, generally refers to information related to cyberspace threats extracted from security data. This can include information on threat sources, attack intent, attack methods, attack targets, and knowledge that can be used to resolve threats or address hazards.

[0101] (2) Mobile Internet Protocol (IP) address (Mobile IP, or IP mobility):

[0102] Mobile IP, also known as "mobile IP," is an Internet transmission protocol standard developed by the Internet Engineering Task Force (IETF). Mobile IP is designed to allow mobile device users to move from one network system to another while maintaining the same IP address. This allows mobile nodes to maintain connectivity while on the move, enabling roaming across different network segments.

[0103] (3) Hierarchical clustering and Agglomerative clustering algorithms:

[0104] Hierarchical clustering creates a hierarchical, nested cluster tree by calculating the similarity between data points of different categories. The advantage of hierarchical clustering is that it does not require a specific number of categories. The resulting tree is a single tree, and after clustering is complete, it can be cut at any level to obtain a specified number of clusters.

[0105] The Agglomerative clustering algorithm is a hierarchical clustering algorithm, a bottom-up clustering method. First, each point in all samples is considered a cluster, then the two clusters with the smallest distance are merged, and the process is repeated until the desired cluster is found or other termination conditions are met.

[0106] (4) Quintuple:

[0107] Quintuple is a computer communication term, usually referring to the source Internet Protocol (IP) address, source port, destination IP address, destination port, and data communication protocol.

[0108] (5) Reserved address:

[0109] Reserved IP addresses refer to IP address ranges used for specific purposes or network environments, such as 0.0.0.0, 127.0.0.0 to 127.255.255.255, 10.0.0.0 to 10.255.255.255, 172.16.0.0 to 172.31.255.255, and 192.168.0.0 to 192.168.255.255.

[0110] (6) Autonomous System Number (ASN)

[0111] On the Internet, an autonomous system (AS) is a small unit that has the authority to independently determine which routing protocol to use within the system. An AS is the collection of all IP networks and routers under the jurisdiction of one or more entities, which implement a common routing policy for the Internet. An AS can be a simple network structure or a group of networks controlled by one or more common network administrators. An AS is a single, manageable network unit (such as a university, an enterprise, or an individual company). An AS is sometimes also referred to as a routing domain. An AS is assigned a globally unique number, the Autonomous System Number (ASN).

[0112] It should be noted that the embodiments of the present application involve operations involving obtaining network data streams and other data. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0113] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0114] The following is a brief introduction to the application areas of the method for identifying mobile network assets provided in the embodiments of the present application.

[0115] With the continuous advancement of technology, more and more devices can serve as network addresses, such as source or destination addresses, to enable the sending and receiving of network data. The source or destination address represented by a device can be a mobile network asset, allowing the device to move from one network resource system to another, while maintaining the same source or destination address. This allows the device to maintain connectivity while moving, enabling roaming across different network segments. By identifying mobile network assets, a network resource whitelist can be provided to network threat intelligence systems, preventing them from mistakenly identifying mobile network assets as abnormal network addresses.

[0116] In related technologies, mobile network assets are typically identified based on the network address's home organization or the network address's Autonomous System Number (ASN). For example, if the network address's home organization is a carrier, then the network address is identified as a mobile network asset. Alternatively, if the network address's ASN is associated with a carrier, then the network address is identified as a mobile network asset.

[0117] However, due to the large number of network addresses, some may not be used in a standardized manner or managed in a timely manner. For example, some network addresses with a carrier as their home institution may have been abandoned or expired long ago but have not been cleared. Another example is that some network addresses have ASNs that are not associated with a carrier but are allocated for use as mobile network assets. Another example is that some network addresses with a carrier as their home institution may be allocated for use as non-mobile network assets.

[0118] It can be seen that when identifying mobile network assets in related technologies, it is easy to identify expired assets as mobile network assets; it is also easy to mistakenly identify non-mobile network assets as mobile network assets; it is also easy to miss mobile network assets, etc., that is, the accuracy of identifying mobile network assets is low.

[0119] In order to solve the problem of low accuracy in identifying mobile network assets, the present application proposes a method for identifying mobile network assets. In this method, multiple network data streams generated in multiple historical time periods are obtained, and the network data streams include network addresses. For the multiple network data streams generated in each historical time period, the following operations are performed: based on the network addresses included in each of the multiple network data streams, the multiple network data streams are clustered to obtain multiple clusters, and the address types of the network addresses included in each of the network data streams included in each cluster are the same. Based on the multiple clusters obtained for the multiple historical time periods, at least one network address that meets the mobile network asset conditions is determined among the network addresses included in each of the network data streams, and a mobile network asset identification result is generated.

[0120] In an embodiment of the present application, mobile network assets are identified by analyzing the network data streams generated in real time during a historical time period, thereby avoiding the situation where abandoned or invalid network addresses are identified as mobile network assets, and to a certain extent improving the accuracy of identifying mobile network assets.

[0121] Furthermore, by analyzing multiple network data streams generated in multiple historical time periods, it is possible to analyze from a macro perspective whether a network address has the characteristics of a mobile network asset, thereby identifying at least one network address that meets the mobile network asset conditions, further improving the accuracy of identifying mobile network assets.

[0122] Furthermore, clustering processing is performed on multiple network data streams generated in each historical time period, and network data streams including network addresses of the same address type can be aggregated into a cluster cluster, so that network addresses that may be used as mobile network assets can be aggregated into a cluster cluster. Combined with the joint analysis of multiple historical time periods, mobile network assets can be accurately identified. For example, mobile network assets that have no association with operator organizations can be identified, and those that belong to operator organizations but are allocated as non-mobile network assets will not be identified as mobile network assets, thereby improving identification accuracy.

[0123] The following describes the application scenarios of the method for identifying mobile network assets provided by this application.

[0124] Please refer to Figure 1, which is a schematic diagram of an application scenario of the method for identifying mobile network assets provided by this application. The application scenario includes a client 101 and a server 102. The client 101 and the server 102 can communicate with each other. The communication method can be to use wired communication technology, for example, by connecting a network cable or a serial port cable; or to use wireless communication technology, for example, by using Bluetooth or wireless fidelity (WIFI) technology, without specific limitation.

[0125] The client 101 generally refers to a device that can send and receive network data streams, such as a terminal device, a third-party application that the terminal device can access, or a web page that the terminal device can access. Terminal devices include but are not limited to mobile phones, computers, smart medical devices, smart home appliances, vehicle-mounted terminals or aircraft, etc. The server 102 generally refers to a device that can identify mobile network assets, such as a terminal device or a server, etc. The server includes but is not limited to a cloud server, a local server or an associated third-party server, etc. Both the client 101 and the server 102 can use cloud computing to reduce the occupation of local computing resources; cloud storage can also be used to reduce the occupation of local storage resources.

[0126] As an embodiment, the client 101 and the server 102 may be the same device or different devices, without any specific limitation.

[0127] The following is a detailed introduction to the method for identifying mobile network assets provided by an embodiment of the present application based on Figure 1. Please refer to Figure 2, which is a flow chart of the method for identifying mobile network assets provided by an embodiment of the present application.

[0128] S201, obtaining a plurality of network data flows generated in a plurality of historical time periods.

[0129] The historical time period is a time period that includes historical times with the current time as a reference. For example, if the current time is 1 p.m., then the historical time period can be the time period from 10 a.m. to 1 p.m. on the same day; for another example, the historical time period can be the entire day of the previous day, etc., and there is no specific restriction.

[0130] Multiple historical time periods can be continuous or discontinuous. For example, multiple historical time periods are two historical time periods, one historical time period is the whole day of September 5, and the other historical time period is the whole day of September 6; for another example, multiple historical time periods are two historical time periods, one historical time period is the whole day of the previous day, and the other time period is from 10:00 am to 1:00 pm today.

[0131] A network data stream includes a network address. The network address may include one or more of a source IP address, a destination IP address, a source port address, and a destination port address. The network data stream may include a quintuple, etc., without limitation. The quintuple may refer to the above description. In the embodiments of the present application, the network data stream may include a quintuple as an example.

[0132] Network data flow can be net flow data. A net flow can be a data packet stream transmitted in one direction between a source IP address and a destination IP address, and all data packets in the data packet stream have a common transport layer source and destination port address.

[0133] Each historical time period corresponds to multiple network data flows. You can obtain multiple network data flows generated during multiple historical time periods from devices such as routers or switches, or from data flow monitoring software, without limitation.

[0134] As an embodiment, since the data streams generated in the network environment may include multiple data streams with the same five-tuple, or multiple data streams with the same source IP address, destination IP address, source port address, and destination port address in the five-tuple; and may also include invalid data streams, such as junk data streams, when obtaining network data streams, multiple initial data streams generated in multiple historical time periods may be obtained first. For the multiple initial data streams generated in each historical time period, the following operations are performed:

[0135] Based on a pre-stored data filtering strategy, multiple initial data streams are filtered to obtain multiple filtered data streams. Among the multiple filtered data streams, at least two filtered data streams with the same network address are deduplicated to obtain multiple network data streams.

[0136] Through data filtering and data deduplication, the obtained initial data streams are preprocessed to ensure the data quality during subsequent clustering processing, thereby improving the final recognition accuracy.

[0137] As an embodiment, since the reserved addresses in the IP addresses are IP address ranges used for specific purposes or specific network environments, the reserved addresses can be used as designated network addresses. When a first data flow including the designated network address is determined to be present among multiple initial data flows, the first data flow is deleted from the multiple initial data flows. The initial data flows including the reserved addresses are filtered out to prevent them from interfering with the accuracy of subsequent clustering processing.

[0138] If the initial data stream also includes a data communication protocol, then when it is determined that a second data stream including a data communication protocol other than the preset communication protocol exists among the multiple initial data streams, the second data stream is deleted from the multiple initial data streams. The preset communication protocol includes the Transmission Control Protocol (TCP) and the User Datagram Protocol (UDP), etc., and can be specifically set according to the usage scenario and is not limited here.

[0139] Thus, multiple filtered data streams can be obtained based on the multiple initial data streams after deleting the first data stream and / or the second data stream.

[0140] S202 , for the multiple network data flows generated in each historical time period, perform the following operations: cluster the multiple network data flows based on the network addresses included in each of the multiple network data flows to obtain multiple clusters.

[0141] Clustering is performed on multiple network data flows generated within each historical time period. This allows for microscopic analysis of the characteristics of network data flows. This, combined with the characteristics analyzed for multiple historical time periods, allows for more accurate identification of mobile network assets.

[0142] The network data flows contained in each cluster have network addresses of the same address type. The same address type of the network addresses can mean the same IP addresses, or the C segment of the IP addresses is within the same address range. For example, in two network data flows, the network address in one network data flow includes a first source IP address and a first destination IP address, and the network address in the other network data flow includes a second source IP address and a second destination IP address. If the first source IP address and the second source IP address are the same, then the network addresses included in the two network data flows can be considered to have the same address type; if the first destination IP address and the second destination IP address are the same, then the network addresses included in the two network data flows can be considered to have the same address type; if the first source IP address and the second source IP address are the same except for the C segment, and the corresponding C segments belong to the same address range, then the network addresses included in the two network data flows can be considered to have the same address type; if the first destination IP address and the second destination IP address are the same except for the C segment, and the corresponding C segments belong to the same address range, then the network addresses included in the two network data flows can be considered to have the same address type, etc., without further limitation.

[0143] Therefore, the obtained clusters can be network data flows representing a source IP address and its corresponding multiple destination IP addresses; or network data flows representing a destination IP address and its corresponding multiple source IP addresses; or network data flows representing the C segment of a source IP address within an address range and its corresponding multiple destination IP addresses; or network data flows representing the C segment of a destination IP address within an address range and its corresponding multiple source IP addresses, etc.

[0144] Since mobile network assets tend to cluster in segment C and, in some cases, also cluster in segment B, network addresses used as mobile network assets can be aggregated into clusters by clustering the network addresses by address type.

[0145] As an embodiment, a specific clustering process is introduced below as an example.

[0146] Based on the network addresses included in each of the two network data flows, the flow similarity between the two corresponding network data flows is determined. Based on the obtained flow similarities, multiple rounds of clustering processing are performed on the multiple network data flows to obtain multiple clusters. If the network addresses include source IP addresses and destination IP addresses, the flow similarity between the two network data flows can be determined based on the weighted sum of the errors between the IP addresses included in the two network data flows and the errors between the destination IP addresses.

[0147] Multiple rounds of clustering processing can be performed until only one cluster is included after the clustering processing, and then based on the flow similarity between the network data flows contained in the current intermediate clusters obtained after each round of clustering processing, the current intermediate clusters obtained after the number of clustering processing rounds are selected as the final multiple clusters. Alternatively, a cluster number can be pre-set, and when the number of current intermediate clusters obtained after a certain round of clustering processing reaches this number of clusters, the obtained current intermediate clusters can be used as the final multiple clusters; alternatively, a flow number can be pre-set, and when the number of network data flows contained in each current intermediate cluster obtained after a certain round of clustering processing reaches this number of flows, the obtained current intermediate clusters can be used as the final multiple clusters, etc., without specific limitations.

[0148] Please refer to Figure 3. The horizontal axis represents each network data flow (not shown in the figure), and the vertical axis represents the number of network data flows contained in a cluster. Two connected network data flows form an intermediate cluster; a network data flow and a connected intermediate cluster form a new intermediate cluster; two connected intermediate clusters form a new intermediate cluster, and so on. Finally, all network data flows form a cluster.

[0149] Please refer to Figure 4, the horizontal axis represents each network data flow (not shown in the figure), and the vertical axis represents the number of network data flows contained in a cluster. The thickest horizontal line in Figure 4 represents the preset number of flows, such as 30, then all the intermediate clusters covered by the horizontal line are regarded as the multiple clusters obtained.

[0150] The following is an introduction to each round of clustering processing:

[0151] First, multiple intermediate clusters from the previous round are obtained. Based on the flow similarity between the network data flows contained in each pair of intermediate clusters from the previous round, these clusters are clustered to obtain multiple current intermediate clusters. If a previous round of clustering exists, these multiple intermediate clusters are the multiple current intermediate clusters obtained after the previous round of clustering. If no previous round of clustering exists, these multiple intermediate clusters are multiple network data flows.

[0152] For example, in the first round of clustering processing, multiple network data flows are respectively used as multiple intermediate clusters of the previous round. Thus, based on the flow similarity between each two network data flows, the cluster similarity between each two intermediate clusters of the previous round can be determined. Then, based on the obtained cluster similarity, the multiple intermediate clusters of the previous round can be clustered to obtain multiple current intermediate clusters.

[0153] In the second round of clustering processing, the multiple current intermediate clusters obtained after the first round of clustering processing are respectively used as multiple previous round intermediate clusters, so that the cluster similarity between each two previous round network data flows can be determined, and then based on the obtained cluster similarity, the multiple previous round intermediate clusters can be clustered to obtain multiple current intermediate clusters, and so on.

[0154] As an embodiment, when clustering multiple intermediate clusters from the previous round based on the flow similarity between each network data flow contained in each of the two intermediate clusters from the previous round to obtain multiple current intermediate clusters, the following operations are performed for each of the two intermediate clusters from the previous round:

[0155] Based on the flow similarities between each network data flow contained in a previous round intermediate cluster and each network data flow contained in another previous round intermediate cluster, the maximum flow similarity is determined. The maximum flow similarity is used as the cluster similarity between the corresponding two previous round intermediate clusters. Based on the cluster similarity between each pair of previous round intermediate clusters, multiple previous round intermediate clusters are clustered to obtain multiple current intermediate clusters. In other words, the Manhattan distance between the two points with the greatest distance between two clusters is used as the cluster similarity between the two clusters.

[0156] S203 : Based on the multiple clusters obtained for the multiple historical time periods, determine at least one network address that meets the mobile network asset condition among the network addresses included in each network data flow, and generate a mobile network asset identification result.

[0157] After obtaining multiple clusters for each historical time period, at least one network address that meets the mobile network asset condition among the network addresses included in each network data flow can be determined based on the multiple clusters obtained for the multiple historical time periods, so that a mobile network asset identification result can be generated based on the at least one network address that meets the mobile network asset condition.

[0158] As an embodiment, when determining whether the mobile network asset condition is met, the following operations are performed for each of the multiple clusters obtained in each historical time period:

[0159] Multiple clusters obtained in historical time periods other than the historical time period in which the cluster cluster exists are respectively used as multiple other clusters obtained for each other historical time period. When it is determined that there are other clusters matching the cluster cluster among the multiple other clusters obtained for each other historical time period, the number of historical time periods in the other historical time periods in which the other clusters matching the cluster cluster exist is counted. When it is determined that the number of obtained historical time periods reaches a first threshold, the network address included in each network data flow included in the cluster cluster is used as a mobile network asset to obtain a mobile network asset identification result.

[0160] When counting the number of historical time periods of other clusters matching the cluster, the number of consecutive historical time periods within the other historical time periods of other clusters matching the cluster can be further counted to further improve recognition accuracy. The first number threshold can be any value, specifically set according to the usage scenario, and is not limited here.

[0161] In one embodiment, when the network addresses included in each network data flow within a cluster are used as mobile network assets to obtain a mobile network asset identification result, in order to further confirm whether the network addresses included in each network data flow within the cluster are mobile network assets, the geographic locations corresponding to the network addresses included in each network data flow within the cluster can be determined. When the number of identical geographic locations among the obtained geographic locations is determined to be below a second threshold, the network addresses included in each network data flow within the cluster are used as mobile network assets to obtain a mobile network asset identification result. If the number of identical geographic locations among the obtained geographic locations is below the second threshold, the geographic locations are relatively dispersed and not clustered, which is consistent with the characteristics of a mobile network asset. Therefore, the geographic locations can be used for secondary confirmation to further improve identification accuracy. The second threshold can be determined based on the number of geographic locations corresponding to the network addresses included in each network data flow within the cluster. The number of locations is positively correlated with the second threshold: a larger number of locations corresponds to a larger second threshold, and a smaller number of locations corresponds to a smaller second threshold. The second threshold can also be a preset value, which is not limited herein.

[0162] As an embodiment, after identifying a mobile network asset, the mobile asset validity period can be set for each network address included in the mobile network asset identification result. The mobile asset validity period represents the period during which the corresponding network address is used as a mobile network asset. Therefore, after the mobile asset validity period expires, it is possible to re-determine whether the corresponding network address is a mobile asset, avoiding the identification of abandoned or expired network addresses as mobile network assets, further improving identification accuracy. At the same time, compared to traditional fixed lifecycle determinations, it improves flexibility and interpretability, and reduces the risk of false positives.

[0163] Furthermore, an expired status flag can be added to each of the network addresses included in multiple network data flows generated during multiple historical time periods, excluding the network addresses included in the mobile network asset identification results. The expired status flag indicates that the corresponding network address is no longer used as a mobile network asset. Thus, the next time a mobile network asset is identified, only network addresses associated with the expired status flag can be identified, and network addresses within the mobile asset's validity period will no longer be identified. This reduces the number of network addresses required for each mobile network asset identification, reduces computational complexity, and improves identification efficiency.

[0164] Below, taking the net flow network data flow generated by a certain operator every day from September 1, 2023 to September 15, 2023 as an example, the method for identifying mobile network assets provided in the embodiment of the present application is introduced by way of example. 950 network data flows can be randomly selected from each network data flow as a test, and 50 network data flows that are confirmed to be mobile network assets are added as verification.

[0165] Please refer to Figure 5A, which includes six modules: data collection module, data filtering module, data extraction module, data clustering module, floating date determination module, and data life cycle assessment module. Please refer to Figure 5B, which is a schematic diagram of the principle of identifying mobile network assets based on these six modules.

[0166] Preparation stage:

[0167] The data collection module is used to collect multiple initial data streams generated in multiple historical time periods and record each quintuple. Please refer to Figure 6A, which is a format diagram of the quintuple of some initial data streams, in which some addresses are not fully shown.

[0168] The data filtering module is used to filter the initial data stream collected by the data collection module, such as filtering the reserved addresses in the source IP address and the destination IP address, and filtering the data communication protocols that do not belong to TCP and UDP.

[0169] The data extraction module is used to deduplicate the initial data stream after the data filtering module has performed data filtering, and obtain multiple network data streams, such as deduplicating data based on the granularity of source IP address, destination IP address, source port address, and destination port address. For multiple net flows with the same source IP address, destination IP address, source port address, and destination port address in the net flow, only one record is retained as the network data flow.

[0170] Clustering stage:

[0171] The data clustering module is used to cluster multiple network data flows generated within each historical time period. For example, it uses the source IP address and destination IP address as a mapping relationship, adds the C segment of the source IP address and the C segment of the destination IP address as clustering features, and uses an agglomerative algorithm. Within the agglomerative algorithm, a full-link clustering strategy is selected for clustering. Please refer to Figure 6B for a schematic diagram of a clustering feature format.

[0172] The generated clusters are divided by the number of network data flows included in them. Clusters containing more than a preset number of network data flows are extracted, and features such as the corresponding source IP address, destination IP address, and the C segment of the source IP address and the C segment of the destination IP address are recorded to obtain multiple clusters. Please refer to FIG7A , which is a schematic diagram of a cluster format, wherein each cluster has a cluster number that uniquely identifies the corresponding cluster, and network data flows associated with the same cluster number belong to the same cluster.

[0173] Second confirmation stage:

[0174] The geographical locations corresponding to the source IP addresses and destination IP addresses in the network data flows contained in the multiple clusters are determined in batches and recorded in the database. Please refer to FIG7B , which is a schematic diagram of a format of the recorded geographical locations.

[0175] As shown in Figure 7B, we can preliminarily determine that the geographical distribution of the destination IP addresses in cluster number 1, whose source IP addresses are mobile network assets, is not regular, but rather multi-regional, which is consistent with the characteristics of mobile network asset communication. Therefore, we can re-confirm that cluster number 1 is a mobile network asset cluster, and that the source IP addresses it contains are mobile network assets.

[0176] The floating date determination module is used to enter threat intelligence for the source IP address of cluster number 1 in the cluster cluster. Based on the cluster cluster characteristics and geographic location characteristics, as well as the characteristics that the data communication time is a floating date, it is comprehensively determined that the source IP address of cluster number 1 in the cluster cluster has mobile network asset characteristics and can be entered as mobile network threat intelligence as a whitelist of mobile network threats.

[0177] Intelligence maintenance phase:

[0178] The data lifecycle assessment module is used to designate the cluster numbered 1 as the mobile network asset cluster based on the characteristics of its changeable geographical location. The validity period of the mobile asset intelligence containing the source IP address of the mobile network asset is 3 months, and the net flow data is extracted again after 3 months for re-identification.

[0179] The data lifecycle assessment module sets an expired status flag for network data flows that do not have mobile network asset characteristics; for network data flows that still have mobile network asset characteristics, the mobile asset validity period of the network data flows is extended, and a new intelligence mobile asset validity period is specified based on the cluster characteristics and geographic location characteristics.

[0180] In the embodiment of the present application, the C-segment features of the IP address selected based on the mobile network asset features can effectively deduce a large number of mobile network assets, avoiding the overly cumbersome and inaccurate judgment of too many IP addresses when judging mobile network assets by a single IP address. Using the Agglomerative algorithm for clustering processing can avoid the disadvantages of other algorithms in performance and principle, and at the same time support classification based on the number of flows or clusters set based on the results, with better results. The floating date judgment is performed on the IP addresses contained in the well-clustered clusters, and according to the connection features of the destination IP addresses corresponding to the source IP addresses contained in the multi-day clusters, it is found that the mobile network asset features with less intersection and lower overlap are present, thereby increasing the accuracy of judging mobile network assets. The geographical location features of the destination IP corresponding to the source IP addresses contained in the clusters are queried, and it is found that the geographical location is distributed in multiple places and has no specific regular characteristics, which is more in line with the geographical location features of mobile network assets, further increasing the accuracy of judging mobile network assets. The mobile network asset intelligence data maintenance mechanism based on the validity period can effectively record the life cycle of mobile network assets and avoid false alarms caused by expired assets.

[0181] Based on the same inventive concept, the present embodiment provides a device for identifying mobile network assets, which can implement the functions corresponding to the aforementioned method for identifying mobile network assets. Referring to FIG8 , the device includes an acquisition module 801 and a processing module 802, wherein:

[0182] Acquisition module 801: used to acquire multiple network data flows generated in multiple historical time periods; wherein the network data flows are used to include network addresses, and the network addresses include: source addresses and destination addresses;

[0183] Processing module 802 is configured to perform the following operations on the multiple network data flows generated in each historical time period: clustering the multiple network data flows based on the network addresses included in each of the multiple network data flows to obtain multiple clusters; wherein the network addresses included in each of the network data flows included in each cluster have the same address type;

[0184] The processing module 802 is further configured to: determine, based on the multiple clusters obtained for the multiple historical time periods, at least one network address that meets the mobile network asset condition among the network addresses included in each network data flow, and generate a mobile network asset identification result.

[0185] In a possible embodiment, the acquisition module 801 is specifically configured to:

[0186] Obtain multiple initial data streams generated in multiple historical time periods;

[0187] For multiple initial data streams generated in each historical time period, perform the following operations:

[0188] Based on the pre-stored data filtering strategy, multiple initial data streams are filtered to obtain multiple filtered data streams;

[0189] Data deduplication is performed on at least two filtered data flows including the same network address in the multiple filtered data flows to obtain multiple network data flows.

[0190] In one possible embodiment, the initial data stream further includes a data communication protocol;

[0191] The acquisition module 801 is specifically used for:

[0192] When determining that a first data flow including the specified network address exists among the multiple initial data flows, deleting the first data flow from the multiple initial data flows;

[0193] When determining that a second data stream including a data communication protocol other than the preset communication protocol exists in the plurality of initial data streams, deleting the second data stream from the plurality of initial data streams;

[0194] Based on the multiple initial data streams after deleting the first data stream and / or the second data stream, multiple filtered data streams are obtained.

[0195] In a possible embodiment, the processing module 802 is specifically configured to:

[0196] Determining the flow similarity between the corresponding two network data flows based on the network addresses respectively included in each of the two network data flows;

[0197] Based on the obtained similarities of each flow, multiple rounds of clustering processing are performed on multiple network data flows to obtain multiple clusters. Each round of clustering processing includes:

[0198] Obtaining multiple previous-round intermediate clusters; wherein, when a previous-round clustering process exists, the multiple previous-round intermediate clusters are multiple current intermediate clusters obtained after the previous-round clustering process; when no previous-round clustering process exists, the multiple previous-round intermediate clusters are multiple network data flows;

[0199] Based on the flow similarity between each network data flow contained in each two intermediate clusters in the previous round, a clustering process is performed on multiple intermediate clusters in the previous round to obtain multiple current intermediate clusters.

[0200] In a possible embodiment, the processing module 802 is specifically configured to:

[0201] For every two intermediate clusters from the previous round, perform the following operations:

[0202] Determine the maximum flow similarity based on the flow similarities between each network data flow included in an intermediate cluster in the previous round and each network data flow included in another intermediate cluster in the previous round;

[0203] The maximum flow similarity is used as the cluster similarity between the corresponding two intermediate clusters in the previous round;

[0204] Based on the cluster similarity between every two intermediate clusters in the previous round, multiple intermediate clusters in the previous round are clustered to obtain multiple current intermediate clusters.

[0205] In a possible embodiment, the processing module 802 is specifically configured to:

[0206] For the multiple clusters obtained in each historical time period, perform the following operations:

[0207] The multiple clusters obtained in the historical time periods other than the historical time period where the cluster cluster is located are respectively used as the multiple other clusters obtained for each other historical time period. When it is determined that there are other clusters matching the cluster cluster among the multiple other clusters obtained for each other historical time period, the number of historical time periods in which the other clusters matching the cluster cluster are located is counted.

[0208] When it is determined that the number of the obtained historical time periods reaches a first number threshold, the network addresses included in each network data flow included in the cluster are used as mobile network assets to obtain a mobile network asset identification result.

[0209] In a possible embodiment, the processing module 802 is specifically configured to:

[0210] Determine the geographical location corresponding to the network address included in each network data flow contained in the cluster;

[0211] When the number of identical geographical locations among the obtained geographical locations is determined to be lower than a second number threshold, the network addresses included in the network data flows included in the clusters are used as mobile network assets to obtain a mobile network asset identification result.

[0212] In a possible embodiment, the processing module 802 is further configured to:

[0213] After obtaining the mobile network asset identification result, respectively set the mobile asset validity period of each network address included in the mobile network asset identification result; wherein the mobile asset validity period represents: the period during which the corresponding network address is used as a mobile network asset;

[0214] An expiration status flag is added to each of the network addresses included in each of the multiple network data flows generated in the multiple historical time periods, except for the network addresses included in the mobile network asset identification result; wherein the expiration status flag is used to indicate that the corresponding network address has not been used as a mobile network asset.

[0215] Please refer to Figure 9, which shows a computer device 900 provided in an embodiment of the present application. The computer device 900 may be, for example, the terminal device 101 or the server 102 in Figure 1. The current version and historical versions of the data storage program and the application software corresponding to the data storage program may be installed on the computer device 900. The computer device 900 includes a processor 980 and a memory 920. In some embodiments, the computer device 900 may include a display unit 940, which includes a display panel 941 for displaying a user interactive operation interface, etc.

[0216] In a possible embodiment, the display panel 941 may be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).

[0217] The processor 980 is configured to read a computer program and then execute the method defined by the computer program. For example, the processor 980 reads a data storage program or file, thereby running the data storage program on the computer device 900 and displaying a corresponding interface on the display unit 940. The processor 980 may include one or more general-purpose processors and may also include one or more DSPs (Digital Signal Processors) to perform related operations to implement the technical solutions provided in the embodiments of the present application.

[0218] The memory 920 generally includes internal memory and external memory. The internal memory may be a random access memory (RAM), a read-only memory (ROM), and a cache (CACHE), etc. The external memory may be a hard disk, an optical disk, a USB disk, a floppy disk, or a tape drive, etc. The memory 920 is used to store computer programs and other data. The computer program includes an application corresponding to each client, etc. Other data may include data generated after the operating system or application is run, and the data includes system data (such as configuration parameters of the operating system) and user data. In the embodiment of the present application, the computer program is stored in the memory 920, and the processor 980 executes the computer program in the memory 920 to implement any of the methods discussed in the previous figure.

[0219] The display unit 940 is used to receive input digital information, character information, or contact touch operations / contactless gestures, and to generate signal input related to user settings and function control of the computer device 900. Specifically, in the embodiment of the present application, the display unit 940 may include a display panel 941. The display panel 941, such as a touch screen, can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any other suitable object or accessory on or on the display panel 941) and drive corresponding connected devices according to a pre-set program.

[0220] In one possible embodiment, the display panel 941 may include a touch detection device and a touch controller. The touch detection device detects the player's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller. The touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 980. The touch controller can also receive and execute commands from the processor 980.

[0221] The display panel 941 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 940, in some embodiments, the computer device 900 may further include an input unit 930. The input unit 930 may include an image input device 931 and other input devices 932. The other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, a joystick, and the like.

[0222] In addition to the above, the computer device 900 may also include a power supply 990 for powering other modules, an audio circuit 960, a near-field communication module 970, and an RF circuit 910. The computer device 900 may also include one or more sensors 950, such as an accelerometer, a light sensor, a pressure sensor, etc. The audio circuit 960 specifically includes a speaker 961 and a microphone 962. For example, the computer device 900 can use the microphone 962 to collect the user's voice and perform corresponding operations.

[0223] As an embodiment, the number of the processors 980 may be one or more, and the processor 980 and the memory 920 may be coupled or relatively independently configured.

[0224] As an embodiment, the processor 980 in FIG. 9 may be used to implement the functions of the acquisition module 801 and the processing module 802 in FIG. 8 .

[0225] As an embodiment, the processor 980 in FIG. 9 may be used to implement the functions corresponding to the server or terminal device discussed above.

[0226] Those skilled in the art will appreciate that all or part of the steps of implementing the above-mentioned method embodiments may be accomplished by a computer program. The aforementioned computer program may be stored in a computer-readable storage medium. When the computer program is executed, it executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0227] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, for example, through a computer program product, which is stored in a storage medium and includes a computer program for enabling a computer device to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program code, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0228] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for identifying mobile network assets, characterized in that: include: Acquire multiple network data flows generated in multiple historical time periods respectively; wherein the network data flows include network addresses; For each of the plurality of network data flows generated in the historical time period, the following operations are performed: based on the network addresses respectively included in the plurality of network data flows, the plurality of network data flows are clustered to obtain a plurality of clusters; wherein the network addresses respectively included in the network data flows included in each of the clusters have the same address type; Based on the multiple clusters respectively obtained for the multiple historical time periods, at least one network address that meets the mobile network asset condition among the network addresses included in each network data flow is determined, and a mobile network asset identification result is generated.

2. The method according to claim 1, characterized in that The obtaining of multiple network data flows generated in multiple historical time periods respectively includes: Obtain multiple initial data streams generated in multiple historical time periods; For each of the multiple initial data streams generated in the historical time period, the following operations are performed: Based on a pre-stored data filtering strategy, the multiple initial data streams are filtered to obtain multiple filtered data streams; Data deduplication is performed on at least two filtered data flows including the same network address among the multiple filtered data flows to obtain the multiple network data flows.

3. The method according to claim 2, characterized in that The initial data stream also includes a data communication protocol; The method of filtering the multiple initial data streams based on the pre-stored data filtering strategy to obtain multiple filtered data streams includes: When determining that a first data flow including a specified network address exists among the multiple initial data flows, deleting the first data flow among the multiple initial data flows; When determining that a second data stream including a data communication protocol other than the preset communication protocol exists in the plurality of initial data streams, deleting the second data stream in the plurality of initial data streams; The multiple filtered data streams are obtained based on the multiple initial data streams after deleting the first data stream and / or the second data stream.

4. The method according to any one of claims 1 to 3, characterized in that: The clustering of the multiple network data flows based on the network addresses respectively included in the multiple network data flows to obtain multiple clusters includes: Determining the flow similarity between the corresponding two network data flows based on the network addresses respectively included in each of the two network data flows; Based on the obtained similarities of each flow, multiple rounds of clustering processing are performed on the multiple network data flows to obtain multiple cluster clusters; wherein each round of clustering processing includes: Acquire multiple previous round intermediate clusters; wherein, when there is a previous round of clustering processing, the multiple previous round intermediate clusters are multiple current intermediate clusters obtained after the previous round of clustering processing, and when there is no previous round of clustering processing, the multiple previous round intermediate clusters are the multiple network data flows; Based on the flow similarity between each network data flow included in each two intermediate clusters of the previous round, clustering processing is performed on the multiple intermediate clusters of the previous round to obtain multiple current intermediate clusters.

5. The method according to claim 4, characterized in that The clustering process is performed on the multiple intermediate clusters of the previous round based on the flow similarity between each network data flow respectively included in each two intermediate clusters of the previous round to obtain multiple current intermediate clusters, including: For each of the two intermediate clusters in the previous round, perform the following operations: Determine the maximum flow similarity based on the flow similarities between each network data flow included in one intermediate cluster of the previous round and each network data flow included in another intermediate cluster of the previous round; The maximum flow similarity is used as the cluster similarity between the corresponding two intermediate clusters in the previous round; Based on the cluster similarity between every two intermediate clusters of the previous round, clustering processing is performed on the multiple intermediate clusters of the previous round to obtain multiple current intermediate clusters.

6. The method according to any one of claims 1 to 3, characterized in that: The step of determining, based on the plurality of clusters obtained for the plurality of historical time periods, at least one network address satisfying the mobile network asset condition among the network addresses included in each of the network data flows, and generating a mobile network asset identification result includes: For the multiple clusters obtained in each historical time period, perform the following operations: Using multiple clustering clusters obtained in historical time periods other than the historical time period in which the clustering cluster is located as multiple other clustering clusters obtained for each other historical time period, and when determining that there are other clustering clusters matching the clustering cluster among the multiple other clustering clusters obtained for each other historical time period, counting the number of historical time periods in other historical time periods in which the other clustering clusters matching the clustering cluster are located; When it is determined that the number of historical time periods obtained reaches a first number threshold, a network address included in each network data flow included in the cluster is used as a mobile network asset to obtain a mobile network asset identification result.

7. The method according to claim 6, characterized in that The step of using the network address included in each network data flow included in the cluster as a mobile network asset to obtain a mobile network asset identification result includes: Determine the geographical location corresponding to the network address included in each network data flow included in the cluster; When the number of locations of the same geographical location among the obtained geographical locations is determined to be lower than a second number threshold, the network address included in each network data flow included in the cluster is used as a mobile network asset to obtain a mobile network asset identification result.

8. The method according to claim 7, characterized in that After obtaining the mobile network asset identification result, the method further includes: The mobile asset validity period of each network address included in the mobile network asset identification result is set respectively; wherein the mobile asset validity period represents: the period during which the corresponding network address is used as a mobile network asset; An expiration status flag is added to the network addresses included in each of the multiple network data flows generated in the multiple historical time periods, except for the network addresses included in the mobile network asset identification result; wherein the expiration status flag is used to indicate that the corresponding network address has not been used as a mobile network asset.

9. A device for identifying mobile network assets, characterized in that: include: Acquisition module: used to acquire multiple network data flows generated in multiple historical time periods; wherein the network data flows are used to include network addresses, and the network addresses include: source addresses and destination addresses; Processing module: for performing the following operations for the multiple network data flows generated in each of the historical time periods: clustering the multiple network data flows based on the network addresses respectively included in the multiple network data flows to obtain multiple clusters; wherein the network addresses respectively included in the network data flows included in each of the clusters have the same address type; The processing module is further configured to: determine at least one network address satisfying a mobile network asset condition among network addresses included in each network data flow based on the multiple clusters obtained for the multiple historical time periods, and generate a mobile network asset identification result.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

11. A computer device, characterized in that: include: A memory for storing program instructions; The processor is used to call the program instructions stored in the memory, and execute the method according to any one of claims 1 to 8 according to the obtained program instructions.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Network data flow detection method and device

    CN110213227A

  • Asset identification method, device and equipment and computer readable storage medium

    CN114972827A

  • Illegal external connection detection method and device

    CN117176456A

  • Method, device and equipment for identifying mobile network assets and storage medium

    CN117768183A

  • Machine learning techniques for associating network addresses with information object access locations

    US20220230078A1