Knowledge mining and application method and device for network access log data, computer equipment and readable storage medium

By dividing and building knowledge graphs of network access log data, combining node clustering and large language model query, the problem of difficulty in digging deep into network access log data knowledge in the existing technology is solved, and accurate and intelligent analysis of massive log data is achieved.

CN119940498AActive Publication Date: 2025-05-06DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510004470.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-06
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

The existing technology is difficult to deeply explore the knowledge and relevance in network access log data, especially in terms of integrating time and space dimension information and using knowledge graph technology, which leads to the inability to accurately and intelligently obtain valuable information from massive log data.

Method used

By dividing the original network access log data, a probability syntax model based on the knowledge graph is constructed, this step is repeated and the graph merging is completed, and node clustering is performed to obtain the target probability knowledge graph. Then, the knowledge graph query language corresponding to the user query request is obtained, the query is executed in the target map, the context information is constructed and the large language model is called to get the query answer.

Benefits of technology

It realizes effective knowledge mining of network access log data, can meet the complex query needs of users, and accurately extract valuable information from massive log data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940498A_ABST
    Figure CN119940498A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge mining and application method and device for network access log data, computer equipment and a readable storage medium, and the method comprises the steps: firstly dividing original network access log data into subsets, constructing a probability grammar model based on a knowledge graph, repeating the step and completing graph combination, and obtaining a knowledge graph set; and performing node clustering to obtain a target probability knowledge graph. And obtaining a knowledge graph query language corresponding to a user query request, executing query in the target graph to obtain a node related result, constructing context information, constructing a cue word, and calling large language model processing to obtain a query answer. According to the method, network access log knowledge is effectively mined, and user query requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a method, device, computer equipment and readable storage medium for knowledge mining and application of network access log data. Background Art

[0002] With the development of network technology, the amount of network access log data has increased dramatically. Traditional methods of analyzing network access log data are mostly simple statistics, which makes it difficult to deeply explore the knowledge and relevance behind the data. Existing technologies lack the means to effectively integrate time and space dimension information and use knowledge graph technology for in-depth mining. When faced with complex query requirements, it is impossible to accurately and intelligently obtain valuable information from massive log data. Summary of the invention

[0003] The object of the present invention is to provide a method, device, computer equipment and readable storage medium for knowledge mining and application of network access log data.

[0004] In a first aspect, an embodiment of the present invention provides a knowledge mining and application method for network access log data, comprising:

[0005] Divide the original network access log data according to preset rules to obtain multiple network access log data subsets;

[0006] For each subset of the network access log data, a probabilistic grammar model based on a knowledge graph is constructed, wherein the knowledge graph is composed of a triple including a source IP node, a target node, and an initiated connection as a relationship, wherein the source IP node and the target node include a unique identifier, a probability parameter, a raw count, a start time, and an end time;

[0007] Repeat the step of constructing a probabilistic grammar model based on a knowledge graph for each subset of the network access log data, and complete graph merging based on the original count, the start time, and the end time to obtain a knowledge graph merged based on time and space;

[0008] Performing node clustering on the merged knowledge graph to obtain a target probability knowledge graph corresponding to each of the network access log data subsets;

[0009] Obtain the knowledge graph query language corresponding to the query request input by the user;

[0010] Performing a query operation in each of the target probabilistic knowledge graphs based on the knowledge graph query language to obtain node-related results corresponding to each of the target probabilistic knowledge graphs;

[0011] Context construction is performed according to the knowledge graph query language and the node-related results to obtain context information for large language model query;

[0012] A prompt word is constructed according to the context information and the query request, and a preset large language model is called to perform processing to obtain an answer content for the query request.

[0013] In a second aspect, an embodiment of the present invention provides a knowledge mining and application device for network access log data, comprising:

[0014] A mining module is used to divide the original network access log data according to preset rules to obtain multiple network access log data subsets; construct a probabilistic grammar model based on a knowledge graph for each of the network access log data subsets, wherein the knowledge graph is composed of a triple including a source IP node, a target node, and an initiated connection as a relationship, and the source IP node and the target node include a unique identifier, a probability parameter, an original count, a start time, and an end time; repeatedly execute the steps of constructing a probabilistic grammar model based on a knowledge graph for each of the network access log data subsets, and complete graph merging based on the original count, the start time, and the end time to obtain a knowledge graph merged based on time and space; perform node clustering on the merged knowledge graph to obtain a target probabilistic knowledge graph corresponding to each of the network access log data subsets;

[0015] The application module is used to obtain the knowledge graph query language corresponding to the query request input by the user; perform a query operation in each of the target probabilistic knowledge graphs based on the knowledge graph query language to obtain node-related results corresponding to each of the target probabilistic knowledge graphs; perform context construction based on the knowledge graph query language and the node-related results to obtain context information for large language model query; construct prompt words based on the context information and the query request, and call a preset large language model to perform processing to obtain the answer content for the query request.

[0016] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor and a non-volatile memory storing computer instructions, wherein when the computer instructions are executed by the processor, the computer device executes the method described in the first aspect.

[0017] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, wherein the readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method described in the first aspect.

[0018] Compared with the prior art, the beneficial effects provided by the present invention include: using a knowledge mining and application method, device, computer equipment and readable storage medium for network access log data disclosed by the present invention, including: first dividing the original network access log data into subsets, building a probabilistic grammar model based on the knowledge graph, repeating this step and completing the graph merging, performing node clustering to obtain the target probabilistic knowledge graph. Obtain the knowledge graph query language corresponding to the user query request, execute the query in the target graph to obtain node-related results, construct context information, construct prompt words to call the large language model for processing, and obtain the query answer. This method effectively mines network access log knowledge and meets user query needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative work.

[0020] Figure 1 A schematic diagram of the steps of the knowledge mining and application method for network access log data provided by an embodiment of the present invention;

[0021] Figure 2 A schematic block diagram of the structure of a knowledge mining and application device for network access log data provided by an embodiment of the present invention;

[0022] Figure 3 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0024] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.

[0025] In order to solve the technical problems in the aforementioned background technology, Figure 1 This is a flow chart of a method for knowledge mining and application of network access log data provided by an embodiment of the present disclosure. The method for knowledge mining and application of network access log data is introduced in detail below.

[0026] Step S201, dividing the original network access log data according to a preset rule to obtain multiple network access log data subsets;

[0027] Step S202, constructing a probabilistic grammar model based on a knowledge graph for each subset of the network access log data, wherein the knowledge graph is composed of a triple including a source IP node, a target node, and an initiated connection as a relationship, and the source IP node and the target node include a unique identifier, a probability parameter, a raw count, a start time, and an end time;

[0028] Step S203, repeatedly executing the step of constructing a probabilistic grammar model based on a knowledge graph for each subset of the network access log data, and completing graph merging based on the original count, the start time, and the end time to obtain a knowledge graph merged based on time and space;

[0029] Step S204, performing node clustering on the merged knowledge graph to obtain a target probability knowledge graph corresponding to each of the network access log data subsets;

[0030] Step S205, obtaining the knowledge graph query language corresponding to the query request input by the user;

[0031] Step S206, performing a query operation in each of the target probabilistic knowledge graphs based on the knowledge graph query language to obtain node-related results corresponding to each of the target probabilistic knowledge graphs;

[0032] Step S207, constructing a context according to the knowledge graph query language and the node-related results to obtain context information for large language model query;

[0033] Step S208: construct prompt words according to the context information and the query request, and call a preset large language model to perform processing to obtain answer content for the query request.

[0034] In an embodiment of the present invention, illustratively, assuming that there is one month of network access log data, in order to facilitate processing and analysis, it is necessary to divide the data according to preset rules. First, according to the node splitting method, the source IP address and the destination IP address are used as the basis. For example, it is found that there are a large number of access records from the source IP address "192.168.1.100" and many accesses to the destination IP address "10.0.0.50". Then, the log data related to these two IP addresses are divided out separately to form two subsets.

[0035] At the same time, the data is also split according to the time window. Assuming that every 24 hours is set as a time window, the log data from 0:00 to 24:00 on the first day will be divided into one subset, and the log data from 0:00 to 24:00 on the second day will be divided into another subset, and so on.

[0036] For example, if "192.168.1.100" frequently accesses "10.0.0.50" during the time period from 12:00 to 13:00 on a certain day, then this part of the data will be divided into the time window subset corresponding to that day.

[0037] Through such a combination of node splitting and time window splitting, numerous subsets of network access log data are obtained, each of which has specific characteristics, which facilitates subsequent processing.

[0038] For each data subset, we start to build a probabilistic grammar model based on the knowledge graph. Take a specific subset as an example, which contains a series of access records with a source IP address of "192.168.1.100" and a destination IP address of "10.0.0.50".

[0039] First, create the nodes of the knowledge graph. The source IP node "192.168.1.100" and the target IP node "10.0.0.50" are established, and each node has attributes such as unique identifier, probability parameter, raw count, start time and end time.

[0040] Then, we build relationships by scanning the log records line by line. For example, if we find a record showing "192.168.1.100 initiated a connection to 10.0.0.50 using protocol type 80 and a packet size of 500 bytes", we will add a triple "<192.168.1.100, 10.0.0.50, initiated connection>" and increase the relationship strength between the two nodes by 1.

[0041] Next, we model the probability. For the conditional probability distribution of "protocol type-packet size", we find that the source IP node "192.168.1.100" has a packet size of 400-600 bytes when using protocol type 80 in multiple visits. Then, this probability distribution will be adjusted and recorded accordingly.

[0042] For the joint probability distribution of "source node port-destination node port", assuming that the port used by the source IP node "192.168.1.100" is usually 8080, and the port received by the destination IP node "10.0.0.50" is usually 80, then this two-dimensional probability distribution will also be recorded.

[0043] Similarly, conditional probability distributions such as “source node-protocol type”, “target node-protocol type”, “source node-target node”, “target node-source node”, and the prior probability distributions of the occurrence of source nodes and target nodes are constructed by analyzing and counting a large number of access records.

[0044] In this way, a rich and accurate probabilistic grammar model is constructed for each data subset.

[0045] After constructing the probabilistic grammar models of multiple data subsets, we start to merge the graphs. Assume that we first divide the data into 1-hour time intervals and obtain a probabilistic grammar model for 24 hours.

[0046] First, the nodes are grouped according to their start and end times. For example, all nodes from 0 to 1 are grouped into one group, all nodes from 1 to 2 are grouped into another group, and so on.

[0047] Then, for nodes in the same time period, they are merged spatially. For example, in the time period from 0:00 to 1:00, there are multiple source IP nodes and target IP nodes. They are merged into the same knowledge graph to ensure that the node relationship and probability distribution in this time period can be fully presented.

[0048] Next, we use the original count information to perform a temporal merge. Assuming that the source IP node "192.168.1.100" appears 100 times from 0:00 to 1:00 and 80 times from 1:00 to 2:00, then we can calculate its probability distribution in the longer time period from 0:00 to 2:00 through the original count, thereby completing the temporal merge and obtaining the probability model of the node at different time scales.

[0049] Through such graph merging, a more comprehensive and accurate knowledge graph can be obtained, providing a more valuable foundation for subsequent node clustering and application.

[0050] For the merged knowledge graph, node clustering is performed. Taking a knowledge graph containing many nodes as an example, each node has a series of probability distribution characteristics.

[0051] First, the feature vector of each node is calculated. Assume that the probability distribution of node A includes the probability distribution of "protocol type-packet size" as P1, the probability distribution of "source node port-destination node port" as P2, and so on.

[0052] Then, by setting the distance threshold, unsupervised clustering methods such as k-NN are used for clustering. For example, if the distance threshold is set to 0.5, if the distance between node A and node B is less than 0.5, then they are classified into one category.

[0053] After clustering is completed, set the relevant clustering information for each node. For example, node A is classified as category 1, its distance from the class center is 0.2, and the maximum distance between classes is 0.8.

[0054] Through node clustering, nodes with similar characteristics can be grouped into one category, which is convenient for subsequent analysis and application.

[0055] When a user enters a query request, such as "query the target nodes accessed between 10 and 11 am yesterday with a source IP of 192.168.1.100 and the type of protocol used", the server uses natural language processing technology to parse the request.

[0056] First, the server identifies the key entity information, namely the source IP address "192.168.1.100" and the time range "10:00 to 11:00 yesterday morning".

[0057] The server then converts these entity and relationship information into a knowledge graph query language, such as the SPARQL query statement: "SELECT DISTINCT?destination_ip?protocol_type WHERE{<192.168.1.100><Initiate connection>?destination_ip.?destination_ip<Use protocol>?protocol_type.FILTER(TIMESTAMP>='Yesterday 10 am'ANDTIMESTAMP<='Yesterday 11 am')}".

[0058] Through such semantic understanding and conversion, the server can accurately understand the user's query intention and convert it into a query language that can be executed in the knowledge graph.

[0059] Based on the query language generated by semantic understanding, the server performs query operations in each target probabilistic knowledge graph.

[0060] Assuming there are multiple target probabilistic knowledge graphs, corresponding to different time periods and node sets, the server will search for information that meets the query conditions in these graphs in turn.

[0061] Taking the query just now as an example, the server will filter out the access records with the source IP address of "192.168.1.100" and the timestamp between 10:00 and 11:00 yesterday morning from all knowledge graphs, and then extract the corresponding target node and the type of protocol used.

[0062] For example, the server finds a record such as "192.168.1.100 accessed 10.0.0.50 at 10:15 yesterday morning, using protocol type 80" and returns it as a query result.

[0063] Based on the results returned by the knowledge query module and the user's query statement, the server generates a context that can be used for large language model queries.

[0064] Assume that the result returned by the knowledge query is "192.168.1.100 accessed 10.0.0.50 at 10:15 yesterday morning, and the protocol type used was 80", and the user's query statement is "Query the access target nodes and the protocol type used between 10 and 11 am yesterday morning with a source IP of 192.168.1.100".

[0065] The server converts this information into a string, for example: "(192.168.1.100,10.0.0.50,initiate connection), protocol type: 80".

[0066] At the same time, the server will also convert the node's related attribute list into a string, such as "node name: 192.168.1.100, - conditional probability distribution of "source node-protocol type": the value range is [0,1000], represented by a Gaussian mixture model, with parameters (k=1, =100, =1), - category ID: 1".

[0067] In this way, the server builds complete and clear context information to prepare for the processing of large language models.

[0068] The server uses the constructed context information as a character string STR1 and the user's inquiry content as a character string STR2 to construct a prompt word.

[0069] Assume that STR1 is "(192.168.1.100, 10.0.0.50, initiate connection), protocol type: 80, node name: 192.168.1.100, - conditional probability distribution of "source node-protocol type": the value range is [0, 1000], represented by Gaussian mixture model, parameters are (k=1, =100, =1), - category ID: 1", STR2 is "query the access target node and the protocol type used between 10 and 11 am yesterday with source IP 192.168.1.100".

[0070] The prompt words constructed by the server may be: "Known: (192.168.1.100, 10.0.0.50, initiate connection), protocol type: 80, node name: 192.168.1.100, - conditional probability distribution of "source node-protocol type": the value range is [0, 1000], represented by Gaussian mixture model, parameters are (k=1, =100, =1), - category ID: 1; the user's inquiry is to query the target node accessed between 10 and 11 am yesterday with source IP 192.168.1.100 and the protocol type used; as a professional knowledge question-answering system, what is your answer?"

[0071] Then, the server calls the preset large language model, such as Tongyi Qianwen or Wenxin Yiyan, and inputs the prompt word. The large language model processes and analyzes the input information and finally generates the answer content for the user's query request, such as "Between 10 and 11 am yesterday, the source IP address 192.168.1.100 accessed the target node 10.0.0.50, and the protocol type used was 80."

[0072] Through such a series of steps and processing, the server can effectively mine the knowledge in the network access log data and provide accurate and useful answers to users.

[0073] In the embodiment of the present invention, the data division of the original network access log data according to the preset rules to obtain multiple network access log data subsets can be implemented through the following examples.

[0074] Divide the original network access log data by nodes according to the source IP address and the destination IP address to obtain the plurality of network access log data subsets; or

[0075] Dividing the original network access log data according to a preset time interval to obtain the plurality of network access log data subsets; or,

[0076] The original network access log data is divided according to the source IP address, the destination IP address and the preset time interval to obtain the multiple network access log data subsets.

[0077] In the embodiment of the present invention, for example, the server receives a large amount of original network access log data, and first starts to divide the nodes according to the source IP address and the destination IP address. Assume that there are the following records in the original data:

[0078] Record 1: The source IP address is "10.10.10.1" and the destination IP address is "20.20.20.2".

[0079] Record 2: The source IP address is "10.10.10.1" and the destination IP address is "30.30.30.3".

[0080] Record 3: The source IP address is "20.20.20.2" and the destination IP address is "10.10.10.1".

[0081] The server will group records with the same source IP address and destination IP address combination together. For example, record 1 and record 2 above will be grouped into one subset because their source IP addresses are both "10.10.10.1". Record 3, however, will be grouped into another subset because its source IP address is "20.20.20.2" and its destination IP address is "10.10.10.1", which is different from the previous two records in terms of source IP and destination IP combination.

[0082] Assume that the preset time interval is every 2 hours. The original network access log data contains the following records, each with an accurate timestamp:

[0083] Record 4: timestamp is 08:00:00;

[0084] Record 5: timestamp is 08:30:00;

[0085] Record 6: timestamp is 09:50:00;

[0086] Record 7: timestamp is 10:10:00;

[0087] The server divides the data based on the timestamp. Records 4, 5, and 6 between 08:00:00 and 09:59:59 are divided into one subset, while record 7 between 10:00:00 and 11:59:59 is divided into another subset.

[0088] For example, the preset time interval is still every 2 hours, and both the source IP address and the destination IP address are considered.

[0089] The original data contains the following records:

[0090] Record 8: The source IP address is "30.30.30.3", the destination IP address is "40.40.40.4", and the timestamp is 12:00:00;

[0091] Record 9: The source IP address is "30.30.30.3", the destination IP address is "40.40.40.4", and the timestamp is 13:30:00;

[0092] Record 10: The source IP address is "50.50.50.5", the destination IP address is "60.60.60.6", and the timestamp is 13:00:00;

[0093] The server first performs a preliminary classification based on the source IP address and the destination IP address. Records 8 and 9 are classified into one category because they have the same source IP address and destination IP address. Then, they are further classified based on the time interval. Records 8, 9, and 10 between 12:00:00 and 13:59:59 are classified into one subset.

[0094] Through the above different data division methods, the server can effectively decompose the original network access log data into multiple subsets with specific characteristics and rules, providing a clear and organized data foundation for subsequent knowledge mining and application work, and helping to analyze and process data more accurately and efficiently.

[0095] In the embodiment of the present invention, the construction of a probabilistic grammar model based on a knowledge graph for each subset of the network access log data may be implemented through the following examples.

[0096] Obtaining an initial knowledge image corresponding to a target network access log data subset; the target network access log data subset is any data subset among multiple network access log data subsets;

[0097] Configuring parameterized probability attributes for the nodes and edges of the initial knowledge graph, and limiting the nodes of the initial knowledge graph to the source IP node and the target node, and limiting the relationship between the source IP node and the target node to initiating a connection;

[0098] The target network access log data subset is scanned line by line to obtain multiple triplets consisting of the source IP node, the target node, and the initiated connection as relations in turn, until a probabilistic grammar model based on the knowledge graph is completed, wherein for each newly added triple, the relationship strength between the corresponding nodes increases by one unit strength.

[0099] In the embodiment of the present invention, illustratively, the server first obtains a subset of target network access log data, assuming that this subset includes the following records:

[0100] Record 1: The source IP address is "192.168.1.10", the destination IP address is "10.0.0.50", the protocol type is "TCP", and the data packet size is "500 bytes";

[0101] Record 2: The source IP address is "192.168.1.10", the destination IP address is "10.0.0.50", the protocol type is "UDP", and the data packet size is "800 bytes";

[0102] Record 3: The source IP address is "192.168.1.20", the destination IP address is "10.0.0.50", the protocol type is "TCP", and the data packet size is "600 bytes";

[0103] The server obtains the initial knowledge graph corresponding to the target network access log data subset. The probability attribute has not yet been configured in the initial knowledge graph.

[0104] Next, the server configures parameterized probabilistic attributes for the nodes and edges of the initial knowledge graph. In this example, the source IP nodes "192.168.1.10" and "192.168.1.20", and the target node "10.0.0.50" are defined as nodes of the knowledge graph, and the relationship between them is defined as "initiating a connection".

[0105] Then, the server scans the target network access log data subset line by line. When scanning record 1, a new triple "<192.168.1.10,10.0.0.50,initiate connection>" is added, and the relationship strength between the two nodes increases by one unit strength. Then scan record 2. Since the source IP node and the destination IP node are the same as the previous triple, the relationship strength between the two nodes is increased again.

[0106] When record 3 is scanned, a new triple “<192.168.1.20,10.0.0.50,initiate connection>” is added, and the strength of the relationship between the two nodes is increased accordingly.

[0107] In this process, the server not only records the connection relationship and strength between nodes, but also models the probability. For example, for the conditional probability distribution of "protocol type-packet size", the server finds that when the source IP node "192.168.1.10" connects to "10.0.0.50" multiple times, the probability of using the "TCP" protocol and the packet size between 400-600 bytes is higher; for the joint probability distribution of "source node port-destination node port", the server counts the common ports used by the source IP node "192.168.1.10" and the common ports received by the destination node "10.0.0.50".

[0108] By scanning and analyzing the target network access log data subset line by line, the server continuously improves the probabilistic grammar model until all records are scanned and the construction of the probabilistic grammar model based on the knowledge graph is completed. This model contains rich information such as the connection relationship between nodes, relationship strength, various probability distributions, etc., providing strong support for subsequent analysis and application.

[0109] For example, this model can be used to infer information such as the protocol type, packet size, and port that may be used when "192.168.1.10" initiates a connection to "10.0.0.50" again, which is helpful for network optimization, security monitoring, etc.

[0110] In an embodiment of the present invention, the probabilistic grammar model includes:

[0111] The conditional probability distribution of protocol type-packet size is: P1=P(c|S i ), is about the source node S i A one-dimensional probability distribution, where S i represents the i-th source node, c represents the size of the data packet;

[0112] The joint probability distribution of source node port and target node port is: P2 = P(Dp|S i .p), is about the source node S i A two-dimensional probability distribution of , where S i .p represents node S i All ports of the target node, Dp represents all ports of the target node;

[0113] The conditional probability distribution of source node-protocol type is: P3=P(r|S i ), is the source node S i A one-dimensional probability distribution of , where r represents the protocol type;

[0114] The conditional probability distribution of target node-protocol type is: P4=P(r|D i ), is the target node D i A one-dimensional probability distribution of, where r represents the protocol type, where D i represents the i-th target node;

[0115] The conditional probability distribution of source node-target node is: P5=P(D|S i ), is the source node S i A one-dimensional probability distribution of ;

[0116] The conditional probability distribution of the target node-source node is: P6 = P(S|D i ), is the target node D i A one-dimensional probability distribution of ;

[0117] The prior probability distribution of source node i is:

[0118] In an embodiment of the present invention, the step of repeatedly executing the step of constructing a probabilistic grammar model based on a knowledge graph for each subset of the network access log data, and completing graph merging based on the original count, the starting time, and the ending time to obtain a knowledge graph merged based on time and space can be implemented through the following examples.

[0119] Repeat the step of constructing a probabilistic grammar model based on a knowledge graph for each subset of the network access log data, and obtain a probabilistic knowledge graph of different nodes and different time periods corresponding to each subset of the network access log data;

[0120] The nodes are grouped by the starting time and the ending time to obtain multiple candidate knowledge graphs, and all nodes in the same time period are placed in the same knowledge graph to complete the spatial merging;

[0121] Based on the original count, the probability distribution of each node in a preset time range is obtained to complete time merging;

[0122] Based on the spatial merging and the temporal merging, a knowledge graph based on temporal and spatial merging is obtained.

[0123] In the embodiment of the present invention, illustratively, the server first repeatedly executes the steps of constructing a probabilistic grammar model based on a knowledge graph on multiple network access log data subsets. Assume that there are three data subsets: subset A, subset B, and subset C.

[0124] Subset A covers the access logs from 8:00 a.m. to 9:00 a.m., including source IP nodes "10.10.10.1" and "10.10.10.2", and destination IP nodes "20.20.20.1" and "20.20.20.2".

[0125] Subset B covers the access logs from 9 am to 10 am, including source IP nodes "10.10.10.1" and "10.10.10.3", and destination IP nodes "20.20.20.1" and "20.20.20.3".

[0126] Subset C covers the access logs from 10 a.m. to 11 a.m., including source IP nodes "10.10.10.2" and "10.10.10.3", and destination IP nodes "20.20.20.2" and "20.20.20.3".

[0127] By constructing a probabilistic grammar model, the server obtains the probabilistic knowledge graph corresponding to different nodes and different time periods of each subset.

[0128] Next, the server groups the nodes by start time and end time. For the nodes in subset A, their start time is 8 o'clock and their end time is 9 o'clock. Similarly, the nodes in subset B have a start time of 9 o'clock and an end time of 10 o'clock, and the nodes in subset C have a start time of 10 o'clock and an end time of 11 o'clock.

[0129] The server obtains multiple candidate knowledge graphs based on this time information. For example, the nodes from 8 to 9 o'clock are placed in one candidate knowledge graph, the nodes from 9 to 10 o'clock are placed in another candidate knowledge graph, and the nodes from 10 to 11 o'clock are placed in another candidate knowledge graph. Then, all nodes in the same time period are placed in the same knowledge graph to complete the spatial merge. For example, in the time period from 9 to 10 o'clock, the node "10.10.10.1" appears in both subset A and subset B. Then, when merging in space, its information in the two subsets is integrated into the same knowledge graph.

[0130] After completing the spatial merging, the server obtains the probability distribution of each node in the preset time range based on the original count to complete the time merging. For example, the original count of node "10.10.10.1" in subset A is 50 times, and the original count in subset B is 30 times. Then the server calculates and obtains the probability distribution of node "10.10.10.1" in the longer time range of 8 o'clock to 10 o'clock.

[0131] Finally, through spatial merging and temporal merging, the server obtains a knowledge graph based on time and space merging. This knowledge graph integrates the node information and probability distribution of different subsets in different time periods, and can more comprehensively and accurately reflect the laws and characteristics of network access. For example, through this merged knowledge graph, you can clearly see which source IP nodes and destination IP nodes are more frequently connected throughout the morning time period, as well as the probability change trend of such connections. This has important guiding significance for network monitoring, optimization, and security analysis.

[0132] In the embodiment of the present invention, the node clustering is performed on the merged knowledge graph to obtain the target probability knowledge graph corresponding to each subset of the network access log data, which can be implemented through the following examples.

[0133] The nodes of the merged knowledge graph are clustered by unsupervised clustering to obtain the target probability knowledge graph corresponding to each subset of the network access log data. After clustering is completed, the category, distance from the class center, and maximum distance between classes are set in the node information to record the clustering information of the node.

[0134] In an embodiment of the present invention, exemplarily, the server obtains a merged knowledge graph, assuming that the knowledge graph contains a large number of nodes, and each node has a series of probability distribution characteristics.

[0135] First, the server selects the k-NN algorithm in the unsupervised clustering method to cluster nodes. The server calculates the feature vector of each node, which contains various probability distribution information of the node, such as the probability distribution of "protocol type-packet size", the joint probability distribution of "source node port-destination node port", etc.

[0136] Assume that the feature vector of node A contains the following probability distribution: the probability distribution of "protocol type-packet size" is that the packet size under a specific protocol is mainly concentrated in 500-800 bytes; the joint probability distribution of "source node port-destination node port" shows that the commonly used source port is 8080 and the destination port is 80.

[0137] The server calculates the distance between node A and other nodes according to the set distance calculation method. For example, when calculating the distance between node A and node B, the distance calculation comprehensively considers the difference in probability distribution and the difference in absolute value of probability value.

[0138] If the distance threshold set by the server is 10, and after calculation it is found that the distance between node A and node B is 8, which is less than the distance threshold, then node A and node B are considered to be similar and may be classified into the same category.

[0139] By continuously calculating and comparing the distances between nodes, the server completes the clustering process. Assume that three cluster categories are finally obtained, namely category 1, category 2 and category 3.

[0140] For nodes that have been clustered, the server sets the relevant clustering information in the node information. For example, if node A is classified as category 1, the server calculates the distance between node A and the center of category 1 and records it. At the same time, the server also calculates the maximum distance between nodes in category 1 and records it in the information of node A.

[0141] Assuming that the distance between node A and the center of category 1 is 5, and the maximum distance between nodes in category 1 is 15, then the information of node A will be set as "category: 1", "distance from category center: 5", and "maximum distance between categories: 15".

[0142] In this way, the server completes node clustering for the merged knowledge graph and records detailed clustering information for each node, thereby obtaining the target probabilistic knowledge graph corresponding to each subset of network access log data. The clustering information in these target probabilistic knowledge graphs can help better understand and analyze the patterns and characteristics of network access. For example, it can be found that certain types of nodes have similarities in access behavior, or that there are large differences between nodes in certain cluster categories and other categories, providing valuable references for further network optimization, security monitoring and other work.

[0143] In an embodiment of the present invention, the knowledge graph query language corresponding to the query request input by the user can be obtained and implemented through the following examples.

[0144] Get the query request entered by the user;

[0145] Convert the query request into the knowledge graph query language containing entity and relationship query information;

[0146] The context construction based on the knowledge graph query language and the node-related results to obtain context information for large language model query can be implemented through the following examples.

[0147] The knowledge graph query language and the node-related results are converted into a character string, and the character string is used as the context information.

[0148] In the embodiment of the present invention, the server first obtains the query request input by the user. Assume that the query request input by the user is: "Find the destination IP address and the protocol type used by the source IP address 192.168.0.10 between 9:00 and 10:00 yesterday morning."

[0149] After receiving this query request, the server begins to convert it into a knowledge graph query language containing entity and relationship query information. The server identifies the key entities, namely the source IP address "192.168.0.10" and the time range "9:00 to 10:00 yesterday morning", as well as the relationships to be queried "destination IP address visited" and "type of protocol used".

[0150] The server converts this information into a knowledge graph query language. For example, a SPARQL query statement may be: "SELECT DISTINCT?destination_ip?protocol_type WHERE{<192.168.0.10><Initiate connection>?destination_ip.?destination_ip<Use protocol>?protocol_type.FILTER(TIMESTAMP>='Yesterday's 9 am' AND TIMESTAMP<='Yesterday's 10 am')}".

[0151] Next, the server performs a query operation based on the knowledge graph query language in each target probabilistic knowledge graph to obtain node-related results related to the query request. Assume that the query result is "the source IP address 192.168.0.10 accessed the destination IP address 10.0.0.50 at 9:30 am yesterday, and the protocol type used was TCP".

[0152] Then, the server constructs the context according to the knowledge graph query language and the node-related results. The server converts the knowledge graph query language "SELECT DISTINCT?destination_ip?protocol_type WHERE{<192.168.0.10><Initiate connection>?destination_ip.?destination_ip<Use protocol>?protocol_type.FILTER(TIMESTAMP>='Yesterday's 9 am'AND TIMESTAMP<='Yesterday's 10 am')}" and the node-related results "The source IP address 192.168.0.10 accessed the destination IP address 10.0.0.50 at 9:30 am yesterday, and the protocol type used was TCP" into a string.

[0153] The converted string may be: "Knowledge graph query language: SELECT DISTINCT? destination_ip? protocol_type WHERE{<192.168.0.10><Initiate connection>? destination_ip.? destination_ip<Use protocol>? protocol_type.FILTER(TIMESTAMP>='Yesterday 9 am' ANDTIMESTAMP<='Yesterday 10 am')} Node related results: The source IP address 192.168.0.10 accessed the destination IP address 10.0.0.50 at 9:30 am yesterday, and the protocol type used was TCP".

[0154] The server uses this string as context information for large language model query, so as to subsequently construct prompt words based on this context information and the user's query request, and call the preset large language model for processing, thereby obtaining an accurate answer to the user's query request.

[0155] Through such detailed and precise processing flow, the server can effectively understand the user's needs, extract valuable information from complex network access log data, and respond to user queries in a clear and useful manner.

[0156] In order to more clearly describe the solution provided by the implementation of the present invention, a relatively complete implementation method is provided below.

[0157] This embodiment uses more than 30,000 access data records of a company's intranet node as input to demonstrate a specific implementation example of the present invention.

[0158] Knowledge mining and expression system

[0159] The data used in this example are as follows Figure 1 As shown in the figure, the collection time is 1 hour, and the total number of intranet nodes involved does not exceed 200. The extranet nodes are divided into about 200 countries and regions according to the IP location.

[0160] Data partitioning: This module sets the slice time to 1 hour, so data with a duration of 1 hour is not partitioned;

[0161] Probabilistic grammar model construction:

[0162] For each row of data, assuming its content is "Node 1 initiates a connection to Node 2", a triple of the form "<Node 1, Node 2, initiate connection>" is added to the knowledge graph.

[0163] For example, in the currently scanned row, the source node is "192.168.28.31" and the destination node is "219.151.145.151" (the location is China and the code is 86). Then a new triple <"192.168.28.31", "86", initiate connection> is added, and the strength value of the relationship between the two nodes is increased by 1.

[0164] For each node, the following probabilities are calculated:

[0165] The conditional probability distribution of “source node-protocol type” P3=P(r|S i ), is the current node as the source node S i is the one-dimensional probability distribution of the protocol type, where r represents the protocol type and its value range is between 0 and 255;

[0166] The conditional probability distribution of “target node-protocol type” is P4=P(r|D i ), the current node is the target node D i is the one-dimensional probability distribution of the protocol type, where r represents the protocol type, and D i represents the i-th target node;

[0167] The conditional probability distribution of "source node - target node" P5 = P(D|S i ), is the current node as the source node S i , the one-dimensional probability distribution when selecting the target node D;

[0168] Prior probability distribution of the current node appearing as source node i

[0169] The prior probability distribution of the current node appearing as the target node i

[0170] Considering that the sample spaces of the above five probabilities are not very large, this embodiment uses one-dimensional vectors to represent them. Specifically, for each node, probability distributions P3 and P4 are represented by 1x256-dimensional normalized (element sum is 1) vectors; since the total number of nodes does not exceed 512, probability distribution P5 is represented by a 1x512-dimensional normalized vector.

[0171] For all nodes, probability distributions P7 and P8 are also represented by 1x512 dimensional normalized vectors.

[0172] Graph merging:

[0173] Since this embodiment does not divide the data, this module only performs spatial merging of the knowledge graph. Specifically, in step c (probabilistic grammar model construction), triples have been obtained by scanning the log records line by line. In this module, all triples need to be stored in the same database (this embodiment uses Triple Store to store triples).

[0174] Node clustering:

[0175] In this embodiment, for each node i, a 1x1026-dimensional (256+256+512+2=1026) feature vector about the node can be established. Where P7(i) and P8(i) are the specific values ​​of the node taken from the probability distribution P7 and P8. The distance between nodes i and j is defined as

[0176]

[0177] After obtaining the distance between two nodes, the K-NN algorithm is used to complete node clustering.

[0178] Knowledge application system:

[0179] Semantic understanding:

[0180] The user question types supported by this embodiment include but are not limited to

[0181] Ask questions about a node mentioned in the log record, such as "Do you think the node 192.168.10.23 is a server or a user terminal?", "Is the node a an active node?", "Which regions' external networks does the node b most frequently access?"

[0182] Ask about the behavioral similarity between any two nodes mentioned in the log records, such as "Are nodes a and b similar in behavior?", "Find all nodes that behave similarly to node a"

[0183] Ask the strength of the association between any two nodes mentioned in the log record, for example, "Which external network region does node a visit the most?", "Which internal network node visits node c the most?"

[0184] Ask for statistics in the log records, such as "What are the top five most active IP addresses in this hour?", "Which category contains the most nodes among all categories?"

[0185] Knowledge query:

[0186] This embodiment uses a large language model to implement this module. The main principle is to set up a semantic classifier and a problem solver. Among them, the semantic classifier decides which problem solver to call; and the problem solver sets solutions for typical problems. The specific classification of semantic classifiers includes "node information query", "node attribute question and answer", "relationship information query", "relationship attribute question and answer", "statistical information query", "statistical information question and answer", etc.

[0187] Taking the problem of "node information query" as an example, this module builds a solver in the following way. First, define the string variables STR_TABLE (used to save the specific description of the data table), STR_NODE (used to save the attribute description and data table fields of the source node and the target node), and STR_QUERY (used to save the user's query text), and then use the prompt words shown in Table 1 to call a large language model with strong command generation capabilities (such as qwen2-72B).

[0188] Table 1

[0189]

[0190] Based on the query language generated by semantic understanding, the generated code is executed in all probabilistic knowledge graphs to extract the most relevant information fragments to solve this type of problem.

[0191] Context construction:

[0192] Define the return result string of the knowledge query module as STR_RES (for example, query the information of a certain node), define the semantic classification of the user query (for example, "node information query") as the string variable STR_Q_TYPE, and convert the node attribute list and relationship list into the string STR_RES_NODE. Then the context string variable is

[0193] STR_Context = "[Known user intent is STR_Q_TYPE, knowledge graph query result is STR_RES, and information about related nodes and relationships is STR_RES_NODE]".

[0194] Note that string variables should be replaced with corresponding texts in actual use.

[0195] Content Generator:

[0196] This module is implemented with the help of the content generation capability of the large language model qwen2-72B. The specific method is to call the large model using the following prompt words:

[0197] "Given STR_Context, the user query is STR_QUERY. You are a knowledge application assistant for log data. Please give an appropriate response. Only a response is required, no other content is required."

[0198] Please refer to Figure 2 , Figure 2 A knowledge mining and application device 110 for network access log data provided by an embodiment of the present invention includes:

[0199] The mining module 1101 is used to divide the original network access log data according to preset rules to obtain multiple network access log data subsets; construct a probabilistic grammar model based on a knowledge graph for each of the network access log data subsets, wherein the knowledge graph is composed of a triple including a source IP node, a target node, and an initiated connection as a relationship, and the source IP node and the target node include a unique identifier, a probability parameter, an original count, a start time, and an end time; repeatedly execute the steps of constructing a probabilistic grammar model based on a knowledge graph for each of the network access log data subsets, and complete graph merging based on the original count, the start time, and the end time to obtain a knowledge graph merged based on time and space; perform node clustering on the merged knowledge graph to obtain a target probabilistic knowledge graph corresponding to each of the network access log data subsets;

[0200] Application module 1102 is used to obtain the knowledge graph query language corresponding to the query request input by the user; perform a query operation in each of the target probabilistic knowledge graphs based on the knowledge graph query language to obtain node-related results corresponding to each of the target probabilistic knowledge graphs; perform context construction based on the knowledge graph query language and the node-related results to obtain context information for large language model query; construct prompt words based on the context information and the query request, and call a preset large language model to perform processing to obtain the answer content for the query request.

[0201] It should be noted that the implementation principle of the aforementioned knowledge mining and application device 110 for network access log data can refer to the implementation principle of the aforementioned knowledge mining and application method for network access log data, and will not be repeated here. It should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also be fully implemented in the form of hardware; some modules can also be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the knowledge mining and application device 110 for network access log data can be a separately established processing element, or it can be integrated in a chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a processing element of the above device. The functions of the knowledge mining and application device 110 for network access log data. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in a processor element or an instruction in software form.

[0202] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more microprocessors (digital signal processors, DSP), or one or more field programmable gate arrays (FPGA), etc. For another example, when a module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0203] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned knowledge mining and application device 110 for network access log data. Figure 3 As shown, Figure 3The computer device 100 provided in the embodiment of the present invention is a structural block diagram. The computer device 100 includes a knowledge mining and application device 110 for network access log data, a memory 111, a processor 112 and a communication unit 113.

[0204] To achieve data transmission or interaction, the memory 111, processor 112 and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The knowledge mining and application device 110 for network access log data includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the knowledge mining and application device 110 for network access log data stored in the memory 111, such as the software function modules and computer programs included in the knowledge mining and application device 110 for network access log data.

[0205] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the aforementioned knowledge mining and application device 110 for network access log data.

[0206] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.

Claims

1. A knowledge mining and application method for network access log data, characterized in that: include: Divide the original network access log data according to preset rules to obtain multiple network access log data subsets; For each subset of the network access log data, a probabilistic grammar model based on a knowledge graph is constructed, wherein the knowledge graph is composed of a triple including a source IP node, a target node, and an initiated connection as a relationship, wherein the source IP node and the target node include a unique identifier, a probability parameter, a raw count, a start time, and an end time; Repeat the step of constructing a probabilistic grammar model based on a knowledge graph for each subset of the network access log data, and complete graph merging based on the original count, the start time, and the end time to obtain a knowledge graph merged based on time and space; Performing node clustering on the merged knowledge graph to obtain a target probability knowledge graph corresponding to each of the network access log data subsets; Obtain the knowledge graph query language corresponding to the query request input by the user; Performing a query operation in each of the target probabilistic knowledge graphs based on the knowledge graph query language to obtain node-related results corresponding to each of the target probabilistic knowledge graphs; Context construction is performed according to the knowledge graph query language and the node-related results to obtain context information for large language model query; A prompt word is constructed according to the context information and the query request, and a preset large language model is called to perform processing to obtain an answer content for the query request.

2. The method according to claim 1, characterized in that The original network access log data is divided according to the preset rules to obtain multiple network access log data subsets, including: Divide the original network access log data by nodes according to the source IP address and the destination IP address to obtain the plurality of network access log data subsets; or Dividing the original network access log data according to a preset time interval to obtain the plurality of network access log data subsets; or, The original network access log data is divided according to the source IP address, the destination IP address and the preset time interval to obtain the multiple network access log data subsets.

3. The method according to claim 1, characterized in that The constructing of a probabilistic grammar model based on a knowledge graph for each subset of the network access log data includes: Obtaining an initial knowledge image corresponding to a target network access log data subset; the target network access log data subset is any data subset among multiple network access log data subsets; Configuring parameterized probability attributes for the nodes and edges of the initial knowledge graph, and limiting the nodes of the initial knowledge graph to the source IP node and the target node, and limiting the relationship between the source IP node and the target node to initiating a connection; The target network access log data subset is scanned line by line to obtain multiple triplets consisting of the source IP node, the target node, and the initiated connection as relations in turn, until a probabilistic grammar model based on the knowledge graph is completed, wherein for each newly added triple, the relationship strength between the corresponding nodes increases by one unit strength.

4. The method according to claim 3, characterized in that The probabilistic grammar model includes: The conditional probability distribution of protocol type-packet size is: P1=P(c|S i ), is about the source node S i A one-dimensional probability distribution, where S i represents the i-th source node, c represents the size of the data packet; The joint probability distribution of source node port and target node port is: P2 = P(Dp|S i .p), is about the source node S i A two-dimensional probability distribution of , where S i .p represents node S i All ports of the target node, Dp represents all ports of the target node; The conditional probability distribution of source node-protocol type is: P3=P(r|S i ), is the source node S i A one-dimensional probability distribution of , where r represents the protocol type; The conditional probability distribution of target node-protocol type is: P4=P(r|D i ), is the target node D i A one-dimensional probability distribution of, where r represents the protocol type, where D i represents the i-th target node; The conditional probability distribution of source node-target node is: P5=P(D|S i ), is the source node S i A one-dimensional probability distribution of ; The conditional probability distribution of the target node-source node is: P6 = P(S|D i ), is the target node D i A one-dimensional probability distribution of ; The prior probability distribution of source node i is: The prior probability distribution of the target node i is:

5. The method according to claim 1, characterized in that: The step of repeatedly executing the step of constructing a probabilistic grammar model based on a knowledge graph for each subset of the network access log data, and completing graph merging based on the original count, the start time, and the end time to obtain a knowledge graph merged based on time and space, includes: Repeat the step of constructing a probabilistic grammar model based on a knowledge graph for each subset of the network access log data, and obtain a probabilistic knowledge graph of different nodes and different time periods corresponding to each subset of the network access log data; The nodes are grouped by the starting time and the ending time to obtain multiple candidate knowledge graphs, and all nodes in the same time period are placed in the same knowledge graph to complete the spatial merging; Based on the original count, the probability distribution of each node in a preset time range is obtained to complete time merging; Based on the spatial merging and the temporal merging, a knowledge graph based on temporal and spatial merging is obtained.

6. The method according to claim 1, characterized in that The node clustering is performed on the merged knowledge graph to obtain a target probability knowledge graph corresponding to each of the network access log data subsets, including: The nodes of the merged knowledge graph are clustered by unsupervised clustering to obtain the target probability knowledge graph corresponding to each subset of the network access log data. After clustering is completed, the category, distance from the class center, and maximum distance between classes are set in the node information to record the clustering information of the node.

7. The method according to claim 1, characterized in that The step of obtaining the knowledge graph query language corresponding to the query request input by the user includes: Get the query request entered by the user; Convert the query request into the knowledge graph query language containing entity and relationship query information; The context construction is performed according to the knowledge graph query language and the node related results to obtain context information for large language model query, including: The knowledge graph query language and the node-related results are converted into a character string, and the character string is used as the context information.

8. A knowledge mining and application device for network access log data, characterized in that: include: A mining module is used to divide the original network access log data according to preset rules to obtain multiple network access log data subsets; A probabilistic grammar model based on a knowledge graph is constructed for each subset of the network access log data, wherein the knowledge graph is composed of a triple including a source IP node, a target node, and an initiated connection as a relationship, and the source IP node and the target node include a unique identifier, a probability parameter, an original count, a start time, and an end time; the steps of constructing a probabilistic grammar model based on a knowledge graph for each subset of the network access log data are repeated, and graph merging is completed based on the original count, the start time, and the end time to obtain a knowledge graph merged based on time and space; Performing node clustering on the merged knowledge graph to obtain a target probability knowledge graph corresponding to each of the network access log data subsets; An application module is used to obtain a knowledge graph query language corresponding to a query request input by a user; perform a query operation based on the knowledge graph query language in each of the target probabilistic knowledge graphs to obtain node-related results corresponding to each of the target probabilistic knowledge graphs; Context construction is performed according to the knowledge graph query language and the node-related results to obtain context information for large language model query; A prompt word is constructed according to the context information and the query request, and a preset large language model is called to perform processing to obtain an answer content for the query request.

9. A computer device, characterized in that: The computer device comprises a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Distributed computing platform based multisource vertical knowledge graph classified integration query method

    CN107341215A

  • Information analysis method and device based on knowledge graph, equipment and storage medium

    CN111368096A

  • Network routing mechanism vulnerability analysis method based on knowledge graph

    CN117834508A

  • Generative language model knowledge editing method and device based on time sequence knowledge graph and medium

    CN118484513A

  • Data management method and system for internet of things API docking platform

    CN118838783A