Data processing method and device

By processing multi-source data, extracting key fields and generating indexes, the problems of insufficient information timeliness and data integration capabilities under the traditional IP management method are solved, and accurate association and rapid response of risk identities are achieved.

CN121864467APending Publication Date: 2026-04-14BEIJING BITE YIPAI INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-05
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Under traditional IP management methods, enterprises face difficulties in locating cybersecurity incidents, lack timely information, have insufficient data integration capabilities, and experience a disconnect between risk identification and risk management.

Method used

By acquiring multi-source data (Dynamic Host Configuration Protocol data, Virtual Private Network data, Wireless LAN data, and Network Security Management data), key fields in a specified format are extracted, an index is generated and stored in the search engine database, and user identifiers are queried based on risk query requests to generate alarm information.

Benefits of technology

It improves information timeliness and data integration capabilities, enables accurate association of risk identities, and ensures rapid response to cybersecurity incidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121864467A_ABST
    Figure CN121864467A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device. The method comprises the following steps: acquiring network data which comprises one or more of dynamic host configuration protocol data, virtual private network data, wireless local area network data and network security management data, extracting a key field in a specified format from the network data, generating an index according to the key field, and transmitting the index to a server according to the index. The method comprises the steps of obtaining a risk terminal identifier, storing the risk terminal identifier in a search engine database, obtaining a risk query request which comprises the risk terminal identifier, querying a corresponding user identifier in the search engine database according to the risk terminal identifier, and generating alarm information according to the user identifier. Therefore, through normalization processing of the multi-source data, the information timeliness and the data integration capability can be improved, and risk identity association is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer information processing, and more particularly to a data processing method and apparatus. Background Technology

[0002] Driven by the wave of enterprise digital transformation, the scale of enterprise internal networks continues to expand, the types of connected devices are becoming increasingly diverse, and mobile office and remote access have become the new normal for employees. This makes the dynamic binding relationship between internal IP (Internet Protocol) addresses and users intricate and frequently changing. In the event of network security incidents such as virus outbreaks, abnormal access, or data breaches, information security personnel urgently need to quickly and accurately locate the specific employees and devices corresponding to problematic IP addresses for timely response and handling. However, traditional IP management methods rely on periodically collecting and analyzing DHCP (Dynamic Host Configuration Protocol) server allocation logs to build a database of IP address-user correspondences, thereby indirectly locating employees. But the information obtained in this way is often delayed, resulting in insufficient timeliness. Another approach is to use network access control systems to record authentication information, forming a real-time online user list. However, this method results in fragmented IP user information, making it difficult to effectively integrate and share data between different systems, leading to insufficient data integration capabilities and a disconnect between risk identification and data sharing. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a data processing method and apparatus that can improve information timeliness and data integration capabilities and achieve risk identity association through the normalization processing of multi-source data.

[0004] In a first aspect, embodiments of the present invention provide a data processing method, the method comprising: Acquire network data, which includes one or more of the following: Dynamic Host Configuration Protocol data, Virtual Private Network data, Wireless Local Area Network data, and Network Security Management data; Extract key fields in a specified format from the network data; An index is generated based on the key fields and stored in the search engine database; Obtain a risk query request, wherein the risk query request includes a risk terminal identifier; Based on the risk terminal identifier, query the corresponding user identifier in the search engine database; Alarm information is generated based on the user identifier.

[0005] In some embodiments, acquiring network data includes: Collect Dynamic Host Configuration Protocol (DHCP) data using a data acquisition tool; Collect virtual private network (VPN) data using data collection tools; The wireless LAN data of the wireless LAN system is obtained by calling the network interface through the data acquisition script. The network security management data of the access control system is obtained by calling the network interface through the data acquisition script.

[0006] In some embodiments, extracting key fields in a specified format from the network data includes: The Dynamic Host Configuration Protocol (DHCP) data and Virtual Private Network (VPN) data are extracted using a filter to obtain a first field information. The first field information includes a first extraction timestamp, a first terminal protocol address, a terminal hardware identifier, and a first event type. The first event type includes DHCP allocation, DHCP release, VPN connection, and VPN disconnection. The second field information is obtained by extracting the wireless LAN data and network security management data through the data parsing library. The second field information includes the second terminal protocol address, user identifier, second extraction timestamp, and second event type. The second event type includes wireless LAN online, wireless LAN offline, access granted, and access denied. The first field information and the second field information are formatted to obtain the key field.

[0007] In some embodiments, generating an index based on the key field and storing it in the search engine database includes: A first index and a second index are generated based on the first key field. The first index is an index of Dynamic Host Configuration Protocol data, and the second index is an index of Virtual Private Network data. A third index and a fourth index are generated based on the second key field. The third index is an index for wireless LAN data, and the fourth index is an index for network security management data. The first index, second index, third index, and fourth index are stored in the search engine database.

[0008] In some embodiments, the risk query request includes: Collect operational and network data from terminal devices; Risk query requests are obtained based on the operational data and network data using a pre-set risk detection model.

[0009] In some embodiments, querying the corresponding user identifier in the search engine database based on the risk terminal identifier includes: Obtain the risk terminal identifier and risk event timestamp based on the risk query request; The search engine database is used to query the risk terminal protocol address corresponding to the risk terminal identifier based on the risk terminal identifier. The user identifier corresponding to the risk terminal protocol address can be queried using the third index and / or the fourth index.

[0010] In some embodiments, querying the risk terminal protocol address corresponding to the risk terminal identifier through the search engine database includes: In response to the retrieval of multiple risky terminal protocol addresses, the terminal protocol address that meets the time matching rule is determined as the risky terminal protocol address based on the risk event timestamp.

[0011] In some embodiments, generating alarm information based on the user identifier includes: Based on the user identifier, obtain the corresponding risk terminal protocol address, risk event description information, risk event timestamp, and risk terminal information; An alarm message is generated based on the user identifier, risk terminal protocol address, risk event description information, risk event timestamp, and risk terminal information, and then sent to the management terminal.

[0012] In some embodiments, the method further includes: A fifth index is generated based on the alarm information and stored in the search engine database.

[0013] In a second aspect, embodiments of the present invention provide a data processing apparatus, the apparatus comprising: The data acquisition module is used to acquire network data, which includes one or more of the following: Dynamic Host Configuration Protocol data, Virtual Private Network data, Wireless Local Area Network data, and Network Security Management data. A formatting module is used to extract key fields in a specified format from the network data; An index building module is used to generate an index based on the key fields and store it in the search engine database; The request acquisition module is used to acquire a risk query request, wherein the risk query request includes a risk terminal identifier; The query module is used to query the corresponding user identifier in the search engine database based on the risk terminal identifier; The alarm generation module is used to generate alarm information based on the user identifier.

[0014] The technical solution of this invention acquires network data, including one or more of Dynamic Host Configuration Protocol (DHCP) data, Virtual Private Network (VPN) data, Wireless Local Area Network (WLAN) data, and Network Security Management (WLAN) data. Key fields of a specified format are extracted from the network data, an index is generated based on the key fields, and stored in a search engine database. A risk query request is then obtained, including a risk terminal identifier. The corresponding user identifier is queried from the search engine database based on the risk terminal identifier, and an alarm message is generated based on the user identifier. Thus, through the normalization processing of multi-source data, information timeliness and data integration capabilities can be improved, and risk identity association can be achieved. Attached Figure Description

[0015] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which: Figure 1 This is a schematic diagram of the data processing system according to an embodiment of the present invention; Figure 2 This is a flowchart of the data processing method according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating the acquisition of network data according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating the extraction of key fields according to an embodiment of the present invention; Figure 5 This is a flowchart of index generation according to an embodiment of the present invention; Figure 6 This is a flowchart illustrating the process of obtaining a risk query request according to an embodiment of the present invention; Figure 7 This is a flowchart illustrating the process of querying a user identifier according to an embodiment of the present invention; Figure 8 This is a flowchart illustrating the generation of alarm information according to an embodiment of the present invention; Figure 9 This is a schematic diagram of a data processing device according to an embodiment of the present invention; Figure 10 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0016] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.

[0017] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0018] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".

[0019] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0020] Figure 1 This is a schematic diagram of a data processing system according to an embodiment of the present invention. Figure 1 In the illustrated embodiment, the data processing system includes a terminal device 11, a log server 12, a wireless local area network system 13, an access control system 14, and a network security management server 15.

[0021] The terminal device 11 serves as the direct carrier of user operations, generating operation data and network data, and simultaneously receiving and displaying alarm information for the security team to review and handle. The log server 12 stores log data such as DHCP (Dynamic Host Configuration Protocol) data and VPN (Virtual Private Network) data. The wireless LAN system 13 manages terminal wireless access behavior and generates Wi-Fi (Wireless Fidelity) access data. The access control system 14 manages terminal intranet access permissions and generates network security management data. The network security management server 15 acquires one or more types of network data, including Dynamic Host Configuration Protocol data, Virtual Private Network data, wireless LAN data, and network security management data; constructs an index based on the network data; obtains risk query requests; queries the corresponding user identifier in the search engine database based on the risk terminal identifier; and generates and pushes alarms.

[0022] When user terminal device 11 performs operations, such as powering on and connecting to the network, remotely logging in via VPN, or accessing business systems, the network security management server 15 deploys a lightweight data collection tool, Filebeat, or a log processing tool, Logstash on the log server 12 to monitor log file changes in real time. When terminal device 11 obtains an IP address via DHCP or accesses the intranet via VPN, the log server 12 generates corresponding Dynamic Host Configuration Protocol (DHCP) data and Virtual Private Network (VPN) data. The data collection tool acquires this dynamic data and pushes it to the network security management server 15. Simultaneously, the network security management server 15 deploys Python scripts on the API (Application Programming Interface) data collection server to periodically call the RESTful API of the wireless LAN system 13 to retrieve wireless LAN data from terminal device 11 in a pull-based manner. Through the Python scripts deployed on the API (Application Programming Interface) data collection server, the server periodically calls the RESTful API (Representational State Transfer API) to retrieve access system data in a pull-based manner.

[0023] After the network security management server 15 acquires the network data, it uses Grok (a parsing filter) to extract fields from the Dynamic Host Configuration Protocol (DHCP) data and Virtual Private Network (VPN) data collected by the log server 12 to obtain the first field information. Then, it uses a Python data parsing library to extract the second field information from the wireless LAN data and network security management data. The first and second field information are standardized; for example, the terminal identifier is standardized to a MAC address, the timestamp to ISO format (e.g., 2025-01-08T09:05:00Z), the data format to JSON, and supplementary data source tags (e.g., log server-DHCP, wireless LAN system-WIFI) to ensure consistent data formats across different sources. Based on the standardized first and second field information, four types of structured indexes are constructed according to data source and purpose. The first index is a dedicated index for DHCP data, which associates the storage terminal's MAC address with the DHCP-assigned IP address; the second index is a dedicated index for VPN data, which associates the storage terminal's MAC address with the VPN access IP address and access account; the third index is a dedicated index for WIFI data, which associates the storage terminal's MAC address, the WIFI-assigned IP address, and the user identifier; and the fourth index is a dedicated index for access control data, which associates the storage terminal's MAC address, the access IP address, the user identifier, and the compliance status.

[0024] When terminal device 11 performs a dangerous operation (such as starting a mining process, accessing a malicious IP, or unauthorized login), network security management server 15 identifies the dangerous behavior through a risk detection model and, based on the four types of indexes it has built, locks down the corresponding user information. Specifically, the risk detection model extracts the risk terminal identifier corresponding to the dangerous operation from the operation data and network data of terminal device 11, and simultaneously determines the risk event timestamp. Network security management server 15 sends a search request to the search engine database, using the risk terminal identifier as the search key, to query all terminal protocol addresses used by the terminal within the risk event timestamp in the first to fourth indexes. If multiple terminal protocol addresses are found, the final risk terminal protocol address is determined through time matching rules. Using the determined risk terminal protocol address as the search key, the user identifier in the third and / or fourth indexes is queried. Because these two types of indexes store the binding relationship between IP and user identifier, the user identifier corresponding to the access control system 14 is extracted first; if not found, the user identifier corresponding to the WIFI system is extracted. If no match is found, it is marked as an unbound user terminal and subsequent verification is triggered.

[0025] After locking user information, the network security management server 15 completes alarm generation, push notifications, and data archiving. Specifically, based on the locked user identifier, it traces back the associated risk terminal protocol address, risk event description information, risk event timestamp, and risk terminal information, and generates alarm information in a standardized format. The network security management server 15 sends the alarm information to the corresponding terminal device 11. After receiving the alarm, the terminal device 11 displays the alarm content on its operation interface, ensuring that the security team is aware of the risks in real time. The network security management server 15 constructs a fifth index based on the alarm information and stores it in the search engine database. The index includes fields such as alarm identifier, user identifier, risk terminal information, alarm time, and handling status, for subsequent risk tracing.

[0026] This invention, through the normalization of multi-source data, improves information timeliness and data integration capabilities, and enables risk identity association. The network data includes one or more of Dynamic Host Configuration Protocol (DHCP) data, Virtual Private Network (VPN) data, Wireless Local Area Network (WLAN) data, and network security management data. Key fields of a specified format are extracted from this network data. An index is generated based on these key fields and stored in a search engine database. A risk query request, including a risk terminal identifier, is then retrieved from the search engine database. The corresponding user identifier is then searched for in the search engine database based on the risk terminal identifier. Finally, an alarm message is generated based on the user identifier. Therefore, through the normalization processing of multi-source data, risk identity association can be achieved.

[0027] Figure 2 This is a flowchart of a data processing method according to an embodiment of the present invention. Figure 2 The data processing method shown includes the following steps: Step S110: Obtain network data, which includes one or more of the following: Dynamic Host Configuration Protocol data, Virtual Private Network data, Wireless Local Area Network data, and Network Security Management data.

[0028] Specifically, the network security management server uses data collection tools and data collection scripts to call the network to obtain one or more of the following: Dynamic Host Configuration Protocol data, Virtual Private Network data, Wireless LAN data, and network security management data, in order to achieve multi-source data collection.

[0029] Figure 3 This is a flowchart of the process of acquiring network data according to an embodiment of the present invention. Figure 3 The process of obtaining network data, as shown, includes the following steps: Step S111: Collect Dynamic Host Configuration Protocol (DHCP) data from the DHCP server using a data acquisition tool.

[0030] Specifically, configure the log collection path for the DHCP (Dynamic Host Configuration Protocol) server in Filebeat or Logstash Agent. Use this collection tool to monitor log file changes in real time. When the DHCP server allocates or releases an IP address, a new record is added to the log file. The Agent captures this raw log entry and obtains DHCP data based on the captured raw log.

[0031] Step S112: Collect virtual private network data of the virtual private network using a data collection tool.

[0032] Specifically, the Filebeat or Logstash Agent is deployed on the server hosting the VPN (Virtual Private Network) or a data collection node accessible from the intranet. Log paths are configured, and log matching rules for connection and disconnection events are set. When a terminal device disconnects from the intranet via the VPN, the Agent captures new logs and obtains VPN data.

[0033] Step S113: Obtain wireless LAN data from the wireless LAN system by calling the network interface through the data acquisition script.

[0034] Specifically, the data acquisition script calls the network interface to automatically send query requests to the wireless LAN system at preset intervals, and receives the wireless LAN data returned by the wireless LAN system after receiving the request.

[0035] Step S114: Obtain network security management data of the access control system by calling the network interface through the data acquisition script.

[0036] Specifically, the data acquisition script calls the network interface to send a query request to the access control system, and receives network security management data returned by the access control system after receiving the request.

[0037] Step S120: Extract key fields in a specified format from the network data.

[0038] Specifically, after receiving the network data, the network security management server extracts key fields from the network data through a parsing filter and a data parsing library, and performs format conversion on the key fields to ensure that the network data has a consistent format.

[0039] Figure 4 This is a flowchart illustrating the extraction of key fields according to an embodiment of the present invention. Figure 4 The extraction of key fields shown includes the following steps: Step S121: Extract the Dynamic Host Configuration Protocol (DHCP) data and Virtual Private Network (VPN) data through a filter to obtain the first field information. The first field information includes a first extraction timestamp, a first terminal protocol address, a terminal hardware identifier, and a first event type. The first event type includes DHCP allocation, DHCP release, VPN connection, and VPN disconnection.

[0040] Specifically, the core logic for extracting Dynamic Host Configuration Protocol (DHCP) data and Virtual Private Network (VPN) data through filters to obtain the first field information utilizes Logstash's Grok filter and pre-sets dedicated extraction templates for the two types of logs to match and extract target fields to determine the first field information. Specifically, Grok regular expressions are used to match the time field in the DHCP logs and automatically standardize it to a standard format (e.g., 2025-01-08T09:05:30Z) as the first extraction timestamp. The IP address assigned to the terminal in the DHCP logs is matched as the first terminal protocol address. The terminal's physical address (MAC address) is matched as a unique terminal hardware identifier. Operational behaviors recorded in the DHCP logs are matched to generate corresponding first event types: DHCP allocation and DHCP release. DHCP allocation is generated when DHCP assigns an IP address to the terminal, and DHCP release is generated when the terminal releases the DHCP-assigned IP address. Pre-set VPN log Grok regular expressions are used to match the time field in the VPN logs, standardizing it to a standard format as the first extraction timestamp. The internal network IP address assigned to the terminal during VPN access is matched as the first terminal protocol address. The MAC address of the terminal accessing the VPN is matched and used as the terminal hardware identifier. Operational behaviors recorded in the VPN logs are matched to generate corresponding first event types: VPN connection and VPN disconnection. The VPN connection is generated when the terminal successfully accesses the VPN, and the VPN disconnection is generated when the terminal disconnects from the VPN.

[0041] Step S122: Extract the second field information from the wireless LAN data and network security management data through the data parsing library. The second field information includes the second terminal protocol address, user identifier, second extraction timestamp, and second event type. The second event type includes wireless LAN online, wireless LAN offline, access granted, and access denied.

[0042] Specifically, considering the characteristics of Wi-Fi data and network security management data, two types of dedicated extraction templates are preset. After accurately extracting target information through these templates, the second field information is obtained. For the aforementioned Wi-Fi data template, the IP address is extracted, corresponding to the second terminal protocol address in the second field information. This IP address is the access IP assigned to the terminal by the Wi-Fi system and is the core network identifier for the terminal in a wireless access scenario; it is directly used as the second terminal protocol address. The template extracts the employee ID, corresponding to the user identifier in the second field information. The employee ID is a unique user identifier within the enterprise and is directly used as the user identifier in the second field information, achieving terminal-user binding. The template extracts the time, corresponding to the second extraction timestamp in the second field information. The extracted time information is standardized using a data parsing library and used as the second extraction timestamp to ensure consistency in the time dimension. The template extracts the Wi-Fi operation type, corresponding to the second event type in the second field information; the extracted operation type is directly mapped to the second event type. When the terminal successfully accesses the Wi-Fi, it corresponds to Wi-Fi online; when the terminal disconnects from the Wi-Fi, it corresponds to Wi-Fi offline.

[0043] The template extracts the IP address, corresponding to the second terminal protocol address in the second field information. This IP address is the one assigned by the access control system when allowing the terminal to access the intranet, or the IP address used by the terminal when accessing the network. The template also extracts the employee ID, corresponding to the user identifier in the second field information. Consistent with the logic of extracting data from the wireless LAN, the extracted employee ID is directly used as the user identifier, enabling the binding of the terminal and user in the access control scenario. The template extracts the time, corresponding to the second extraction timestamp in the second field information. After being standardized to ISO format by the data parsing library, it is used as the second extraction timestamp to ensure consistency with the time format of the wireless LAN data extraction. The template extracts the access control system operation type, corresponding to the second event type in the second field information. Based on the operation result of the access control system, it is mapped to the second event type. When the terminal passes the access control verification, access is granted; when the terminal fails the access control verification, access is denied.

[0044] Step S123: Format the first field information and the second field information to obtain the key field.

[0045] Specifically, the two types of information extracted in the first two steps are standardized in format, merged, and organized into key fields that are universally applicable across the entire system. For example, the first and second extraction timestamps are all converted to a standard format. The first and second terminal protocol addresses are uniformly free of spaces and lowercase letters, and labeled with IP type. MAC addresses are all converted to uppercase, with standardized separators. Dispersed event types are grouped into common tags; for example, DHCP allocation, VPN connection, wireless LAN online and access approval are uniformly labeled as "Network Access," while DHCP release, VPN disconnection, wireless LAN offline and access denial are uniformly labeled as "Network Outbound." Furthermore, the first and second terminal protocol addresses can be uniformly named `ip_address`, the first and second extraction timestamps can be uniformly named `collect_time`, the terminal hardware identifier can be named `mac_address`, and the user identifier can be named `employee_id`, ensuring consistency of field names across the entire system.

[0046] Step S130: Generate an index based on the key fields and store it in the search engine database.

[0047] Specifically, the obtained key fields are indexed and stored in a search engine database, which can also be an Elasticsearch (ES) cluster. The ES cluster is an open-source search and analytics engine cluster built on a distributed architecture, consisting of multiple cooperating nodes. The index is a distributed data collection composed of multiple shards, serving as the core unit for data organization and access in ES, and is used to categorize and store data according to data type and / or purpose. After receiving an index creation request, the master node of the ES cluster allocates the primary shard and replica shards to different data nodes based on the cluster node load and sharding strategy. When a data node receives a data write request, it first writes the document to the primary shard and then synchronizes it to the replica shards to ensure data consistency. When a retrieval request is initiated, the coordinating node routes the request to the data node containing the relevant shard, aggregates the retrieval results from all shards, and returns them to the terminal device.

[0048] Figure 5 This is a flowchart of index generation according to an embodiment of the present invention. Figure 5 The index generation process shown includes the following steps: Step S131: Generate a first index and a second index based on the first key field. The first index is an index of dynamic host configuration protocol data, and the second index is an index of virtual private network data.

[0049] Specifically, based on the common characteristics of DHCP and VPN data, a general index template is created, defining field mapping rules, fragmentation strategies, and lifecycle strategies. Simultaneously, a first index for Dynamic Host Configuration Protocol (DHCP) data is generated based on the first key field. This first index is a structured data set specifically storing DHCP server IP allocation or release records. For example, following a preset naming convention, the first index can be named ip-mapping-dhcp, and each record stored within it is a standardized DHCP event record, containing fixed fields: The IP address acquired or released by the terminal (e.g., 192.168.1.105). The terminal's hardware MAC address (e.g., AA:BB:CC:DD:EE:FF). Standard timestamp for data collection (e.g., 2025-01-08T09:15:30.000Z); Specific event type (DHCP allocation or DHCP release); A fixed data source identifier (such as a DHCP server (IDC data center area A)).

[0050] Simultaneously, a second index for the virtual private network (VPN) data is generated based on the first key field. This second index is a structured data set specifically storing VPN gateway remote access or disconnection records. For example, following a preset naming convention, the second index can be named ip-mapping-vpn, where each data entry is a standardized record of a VPN remote access event, with core fields aligned with the DHCP index. The internal network IP address (e.g., 192.168.2.203) obtained by the terminal through VPN. The terminal's hardware MAC address (format consistent with the DHCP index); The collected standard timestamps (format consistent with the DHCP index); Specific event type (VPN connection or VPN disconnection); A fixed data source identifier (such as a VPN gateway); Additional business field: terminal public IP address (e.g., 117.136.8.25, recording the public network location for remote access).

[0051] Step S132: Generate a third index and a fourth index based on the second key field. The third index is an index for wireless local area network data, and the fourth index is an index for network security management data.

[0052] Specifically, based on the general index templates for DHCP and VPN, an additional `employee_id` field is added for precise matching of employee identifiers, forming dedicated templates for WIFI (Wireless Fidelity) and NAC data. Format validation rules for `employee_id` are defined to prevent the writing of illegally formatted data. Simultaneously, a third index for WIFI data is generated based on the second key field. This third index is a structured data set specifically storing WIFI controller terminal online / offline records. For example, following a preset naming convention, the third index can be named `ip-mapping-wifi`, where each data entry is a standardized record of a WIFI access event. Fields are extended based on DHCP / VPN to adapt to WIFI scenarios. The IP address obtained by the terminal via WIFI; The terminal's hardware MAC address; The standard timestamp collected; The newly added core field is employee ID (e.g., 00123). Specific event type (WLAN online or WLAN offline). A fixed data source identifier (e.g., a Wi-Fi controller (xx AC6805)); The name of the connected Wi-Fi network (e.g., Corp-WIFI-Office). The access point (AP) number (e.g., AP-Floor3-08).

[0053] Simultaneously, a fourth index for network security management data is generated based on the second key field. This fourth index is a structured data set specifically storing records of compliant access or denial of terminals in the NAC access control system. For example, following a preset naming convention, the fourth index can be named ip-mapping-nac. Each data entry within it is a standardized record of a terminal access event, with fields similar to those in the WIFI index. The terminal's IP address; The terminal's hardware MAC address; The standard timestamp collected; The core field is the employee ID number; Specific event types (admission granted or admission denied); Fixed data source identifiers (such as NAC access control systems); Terminal compliance level (e.g., high, medium, low); The access switch port (e.g., Switch-Floor2-03, record the physical port of wired access).

[0054] Step S133: Store the first index, second index, third index and fourth index in the search engine database.

[0055] Specifically, the master node of the Elasticsearch cluster distributes the primary and replica shards of the first, second, third, and fourth indexes to different data nodes in the search engine database according to the sharding load balancing algorithm. Priority is given to assigning different primary shards of the same index to nodes on different racks to avoid index unavailability due to failure. Replica shards are assigned to different nodes than their corresponding primary shards to ensure that replica shards can be promoted to primary shards if a primary shard node fails.

[0056] Step S140: Obtain a risk query request, wherein the risk query request includes a risk terminal identifier.

[0057] Specifically, collecting behavioral data from terminal devices provides the initial basis for risk assessment. This behavioral data can be divided into operational data and network data. After the system automatically identifies a security risk in a terminal device, it generates a risk query request corresponding to that risky terminal.

[0058] Figure 6 This is a flowchart illustrating the process of obtaining a risk query request according to an embodiment of the present invention. Figure 6 The risk query request shown includes the following steps: Step S141: Collect the operation data and network data of the terminal device.

[0059] Specifically, a terminal detection and response client is pre-installed on all terminal devices of the enterprise to collect operation data at a fixed frequency. The operation data includes process running records (e.g., whether mining or virus processes are started), file operations (e.g., accessing sensitive directories), login behavior (e.g., unauthorized account login), peripheral device access (e.g., USB flash drive insertion), and system vulnerability status (e.g., high-risk vulnerabilities not patched).

[0060] The network data of the terminal is collected in real time through the mirror port or API interface of the network device. The network data includes the terminal MAC address, accessed IP or domain name (e.g., malicious domain name), port communication records (e.g., abnormal external connection port), traffic characteristics (e.g., large volume of data transmission), and access method (wired, WIFI, and VPN).

[0061] Step S142: Obtain a risk query request based on the operation data and network data using a pre-set risk detection model.

[0062] Specifically, the collected operational and network data are standardized and transformed into a structured format recognizable by the risk detection model. The risk detection model has a pre-built multi-dimensional risk trigger condition library, which includes operational and network risk conditions. Operational risk trigger conditions include the presence of pre-set malicious processes in process execution records, access to sensitive system directories, modification of system configuration files, unauthorized account logins, multiple regional logins, and unpatched high-risk system vulnerabilities on the terminal. Network risk conditions include accessed IPs or domains that match malicious IP or domain blacklists, abnormal external ports, large upload volumes in a short period, unusual traffic fluctuations, and access to core business systems via VPN outside of office hours. The model compares the operational and network data in the input data against the pre-set risk trigger condition library item by item, marking matched risk conditions and obtaining risk query requests. Simultaneously, the model can calculate a comprehensive risk score for the terminal based on the risk weight of the matched conditions and set trigger thresholds. If the terminal meets the trigger threshold, it is determined to be a risky terminal; otherwise, it is determined to be a normal terminal, and the process terminates.

[0063] Step S150: Query the corresponding user identifier in the search engine database based on the risk terminal identifier.

[0064] Specifically, the risk terminal identifier and risk event time range are obtained from the risk query request. Using the risk terminal identifier as the search key, all IP addresses used by the terminal during the risk period are retrieved from the first, second, third, and fourth indexes of the ES cluster. After deduplication and sorting of the results, the most recently used IP address is taken as the risk terminal protocol address. Using the risk terminal protocol address as the search key, the user identifier is accurately matched in the third and fourth indexes containing user information.

[0065] Figure 7 This is a flowchart illustrating the process of querying user identifiers according to an embodiment of the present invention. Figure 7 The query user identifier shown specifically includes the following steps: Step S151: Obtain the risk terminal identifier and risk event timestamp according to the risk query request.

[0066] Specifically, a JSON parser extracts the first field indicating the risk terminal identifier and the second field indicating the risk event timestamp from the risk query request. To avoid invalid queries, the fields are validated after parsing. The first field is checked to see if it conforms to the MAC address format (e.g., XX:XX:XX:XX:XX:XX), and identifiers with invalid formats are removed. The timestamp is checked to see if it is in ISO standard format and if the end time is later than the start time. If the start time is later than the end time, the query window is taken by default as the 24 hours before the risk request was generated. Simultaneously, the validated risk terminal identifier and risk event timestamp are encapsulated into an Elasticsearch (ES) retrieval condition object for convenient subsequent query operations.

[0067] Step S152: Query the risk terminal protocol address corresponding to the risk terminal identifier through the search engine database based on the risk terminal identifier.

[0068] Specifically, using the terminal risk identifier as the search key, the system matches all IP addresses used by the terminal during the risky period within the full index of the Elasticsearch (ES) cluster. The core is precise cross-index matching to obtain the terminal's IP trajectory. The terminal's IP address may originate from DHCP, VPN, Wi-Fi, or NAC access control; therefore, all four types of indexes in the ES cluster must be queried. A Boolean query is used to construct the search statement, filtered by the terminal's MAC address and time range. The search statement is submitted to the ES cluster's query API, and the cluster's coordinating node routes the query request to all relevant shards of the four types of indexes, executing the search in parallel. The ES cluster returns the terminal's IP usage records during the risky period, determining the risky terminal protocol address corresponding to the risky terminal identifier.

[0069] Simultaneously, in response to the retrieval of multiple risky terminal protocol addresses, the terminal protocol addresses that meet the time matching rules are determined as risky terminal protocol addresses based on the risk event timestamps. Specifically, the time range of risk event occurrences is obtained from the risk query request, and the precise occurrence time of the risk event is determined by the risk detection model. If the time difference between the collection timestamp of the terminal protocol address and the precise occurrence time of the risk event is less than or equal to a preset time, the terminal protocol address is directly determined as a risky terminal protocol address. If no terminal protocol address meets the precise matching requirement, all terminal protocol addresses whose collection timestamps fall within the time range of the risk event are filtered out, and then sorted by the time difference between the collection timestamp and the precise occurrence time of the risk event from smallest to largest. The terminal protocol address with the smallest time difference is selected as the risky terminal protocol address.

[0070] Step S153: Query the user identifier corresponding to the risk terminal protocol address based on the risk terminal protocol address using the third index and / or the fourth index.

[0071] Specifically, the third index stores Wi-Fi access records, including the binding relationship between IP addresses and employee IDs, and the fourth index stores NAC access records, also including the binding relationship between IP addresses and employee IDs. Search conditions for the third and fourth indices are constructed based on the determined risky terminal protocol addresses, and search requests for both indices are submitted to the ES cluster. If both the third and fourth indices return results, the user identifier from the fourth index is prioritized. This is because NAC access is the final barrier for terminal access to the intranet, and its IP-user binding relationship has the highest accuracy. If only one of the third and fourth indices returns a result, the user identifier corresponding to the risky terminal protocol address is directly extracted.

[0072] Step S160: Generate alarm information based on the user identifier.

[0073] Specifically, based on the retrieved user identifiers, the system aggregates full-link data including the corresponding risky terminal IP, risk event description, risk event time, and terminal access information. The integrated information is then used to generate alarms in a standardized format and pushed to management terminals and operations personnel through multiple channels such as security operations platforms, instant messaging tools, and email. Simultaneously, the alarm information is structured and stored in the fifth index of the Elasticsearch cluster, enabling long-term retention and traceability of alarm data.

[0074] Figure 8 This is a flowchart illustrating the generation of alarm information according to an embodiment of the present invention. Figure 8 The generation of alarm information shown includes the following steps: Step S161: Obtain the corresponding risk terminal protocol address, risk event description information, risk event timestamp, and risk terminal information based on the user identifier.

[0075] Specifically, by using the user identifier in the third and / or fourth index, all terminal information bound to that user identifier is queried, and terminal protocol addresses and device information matching the risky IP are filtered out. The description information and timestamp of the risk event are extracted from the risk query request log. Thus, the risky terminal protocol address, risk event description information, risk event timestamp, and risky terminal information corresponding to the user identifier are obtained. Simultaneously, detailed information such as the terminal's access method and access location can be supplemented from the four types of ES service indexes.

[0076] Step S162: Generate alarm information based on the user identifier, risk terminal protocol address, risk event description information, risk event timestamp, and risk terminal information, and send the alarm information to the management terminal.

[0077] Specifically, based on a preset alarm template or format, the collected user identifier, risk terminal protocol address, risk event description information, risk event timestamp, and risk terminal information are integrated into structured alarm information.

[0078] Step S163: Generate a fifth index based on the alarm information and store it in the search engine database.

[0079] Specifically, the configuration rules for the fifth index are standardized according to a pre-created index template. Field mappings are defined, specifying the storage types for core fields such as alarm ID, user identifier, risk IP, and timestamp. The fifth index is generated based on the alarm information using the index template and field mappings and stored in the search engine database. This fifth index is a dedicated index within the search engine database specifically for carrying alarm information.

[0080] This invention, through the normalization of multi-source data, improves information timeliness and data integration capabilities, and enables risk identity association. The network data includes one or more of Dynamic Host Configuration Protocol (DHCP) data, Virtual Private Network (VPN) data, Wireless Local Area Network (WLAN) data, and network security management data. Key fields of a specified format are extracted from this network data. An index is generated based on these key fields and stored in a search engine database. A risk query request, including a risk terminal identifier, is then retrieved from the search engine database. The corresponding user identifier is then searched for in the search engine database based on the risk terminal identifier. Finally, an alarm message is generated based on the user identifier. Therefore, through the normalization processing of multi-source data, risk identity association can be achieved.

[0081] Figure 9 This is a schematic diagram of a data processing apparatus according to an embodiment of the present invention. Figure 9 As shown, the data processing device of this embodiment includes a data acquisition module 81 for acquiring network data, which includes one or more of Dynamic Host Configuration Protocol (DHCP) data, Virtual Private Network (VPN) data, Wireless Local Area Network (WLAN) data, and Network Security Management (WLAN) data; a formatting module 82 for extracting key fields of a specified format from the network data; an index building module 83 for generating an index based on the key fields and storing it in a search engine database; a request acquisition module 84 for acquiring risk query requests, which include risk terminal identifiers; a query module 85 for querying the corresponding user identifier in the search engine database based on the risk terminal identifier; and an alarm generation module 86 for generating alarm information based on the user identifier.

[0082] This invention, through the normalization of multi-source data, improves information timeliness and data integration capabilities, and enables risk identity association. The network data includes one or more of Dynamic Host Configuration Protocol (DHCP) data, Virtual Private Network (VPN) data, Wireless Local Area Network (WLAN) data, and network security management data. Key fields of a specified format are extracted from this network data. An index is generated based on these key fields and stored in a search engine database. A risk query request, including a risk terminal identifier, is then retrieved from the search engine database. The corresponding user identifier is then searched for in the search engine database based on the risk terminal identifier. Finally, an alarm message is generated based on the user identifier. Therefore, through the normalization processing of multi-source data, risk identity association can be achieved.

[0083] Figure 10 This is a schematic diagram of an electronic device according to an embodiment of the present invention. (For example...) Figure 10 As shown, Figure 10 The illustrated electronic device is a data processing apparatus, comprising a general computer hardware architecture, including at least a processor 91 and a memory 92. The processor 91 and memory 92 are connected via a bus 93. The memory 92 is adapted to store instructions or programs executable by the processor 91. The processor 91 can be a standalone microprocessor or a collection of one or more microprocessors. Thus, the processor 91 executes the instructions stored in the memory 92, thereby performing the method flow of the embodiments of the present invention as described above to process data and control other devices. The bus 93 connects the aforementioned components together, and also connects these components to a display controller 94, a display device, and an input / output (I / O) device 95. The input / output (I / O) device 95 can be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, the input / output device 95 is connected to the system via an input / output (I / O) controller 96.

[0084] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus (devices), or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0085] This application is described with reference to flowchart illustrations of methods, apparatus (devices), and computer program products according to embodiments of this application. It should be understood that each step in the flowchart can be implemented by computer program instructions.

[0086] These computer program instructions may be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction means, the implementation process of which is described in the instruction means. Figure 1 The function specified in one or more processes.

[0087] These computer program instructions may also be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, produce instructions for implementing processes. Figure 1 A device for a function specified in one or more processes.

[0088] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program for use by a computer to execute some or all of the above-described method embodiments.

[0089] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program specifying the relevant hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0090] The solutions described in this specification and embodiments, if involving the processing of personal information, will be processed only on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be processed within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0091] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A data processing method, characterized in that, The method includes: Acquire network data, which includes one or more of the following: Dynamic Host Configuration Protocol data, Virtual Private Network data, Wireless Local Area Network data, and Network Security Management data; Extract key fields in a specified format from the network data; An index is generated based on the key fields and stored in the search engine database; Obtain a risk query request, wherein the risk query request includes a risk terminal identifier; Based on the risk terminal identifier, query the corresponding user identifier in the search engine database; Alarm information is generated based on the user identifier.

2. The method according to claim 1, characterized in that, The acquisition of network data includes: Collect Dynamic Host Configuration Protocol (DHCP) data using a data acquisition tool; Collect virtual private network (VPN) data using data collection tools; The wireless LAN data of the wireless LAN system is obtained by calling the network interface through the data acquisition script. The network security management data of the access control system is obtained by calling the network interface through the data acquisition script.

3. The method according to claim 1, characterized in that, The key fields in a specified format extracted from the network data include: The Dynamic Host Configuration Protocol (DHCP) data and Virtual Private Network (VPN) data are extracted using a filter to obtain a first field information. The first field information includes a first extraction timestamp, a first terminal protocol address, a terminal hardware identifier, and a first event type. The first event type includes DHCP allocation, DHCP release, VPN connection, and VPN disconnection. The second field information is obtained by extracting the wireless LAN data and network security management data through the data parsing library. The second field information includes the second terminal protocol address, user identifier, second extraction timestamp, and second event type. The second event type includes wireless LAN online, wireless LAN offline, access granted, and access denied. The first field information and the second field information are formatted to obtain the key field.

4. The method according to claim 3, characterized in that, The process of generating an index based on the key fields and storing it in the search engine database includes: A first index and a second index are generated based on the first key field. The first index is an index of Dynamic Host Configuration Protocol data, and the second index is an index of Virtual Private Network data. A third index and a fourth index are generated based on the second key field. The third index is an index for wireless LAN data, and the fourth index is an index for network security management data. The first index, second index, third index, and fourth index are stored in the search engine database.

5. The method according to claim 1, characterized in that, The risk query request includes: Collect operational and network data from terminal devices; Risk query requests are obtained based on the operational data and network data using a pre-set risk detection model.

6. The method according to claim 4, characterized in that, The step of querying the corresponding user identifier in the search engine database based on the risk terminal identifier includes: Obtain the risk terminal identifier and risk event timestamp based on the risk query request; The search engine database is used to query the risk terminal protocol address corresponding to the risk terminal identifier based on the risk terminal identifier. The user identifier corresponding to the risk terminal protocol address can be queried using the third index and / or the fourth index.

7. The method according to claim 6, characterized in that, The step of querying the risk terminal protocol address corresponding to the risk terminal identifier through the search engine database includes: In response to the retrieval of multiple risky terminal protocol addresses, the terminal protocol address that meets the time matching rule is determined as the risky terminal protocol address based on the risk event timestamp.

8. The method according to claim 1, characterized in that, The step of generating alarm information based on the user identifier includes: Based on the user identifier, obtain the corresponding risk terminal protocol address, risk event description information, risk event timestamp, and risk terminal information; An alarm message is generated based on the user identifier, risk terminal protocol address, risk event description information, risk event timestamp, and risk terminal information, and then sent to the management terminal.

9. The method according to claim 8, characterized in that, The method further includes: A fifth index is generated based on the alarm information and stored in the search engine database.

10. A data processing apparatus, characterized in that, The device includes: The data acquisition module is used to acquire network data, which includes one or more of the following: Dynamic Host Configuration Protocol data, Virtual Private Network data, Wireless Local Area Network data, and Network Security Management data. A formatting module is used to extract key fields in a specified format from the network data; An index building module is used to generate an index based on the key fields and store it in the search engine database; The request acquisition module is used to acquire a risk query request, wherein the risk query request includes a risk terminal identifier; The query module is used to query the corresponding user identifier in the search engine database based on the risk terminal identifier; The alarm generation module is used to generate alarm information based on the user identifier.