Unsupervised network flow detection method and device based on large language model reasoning, equipment and medium

Through the unsupervised network traffic detection method of large language model reasoning, the problem of low detection efficiency and accuracy in the existing technology is solved, and efficient and accurate abnormal traffic detection is achieved, which is suitable for large-scale and complex network environments.

CN120474744APending Publication Date: 2025-08-12WUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510533599.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing abnormal network traffic detection methods are difficult to effectively deal with unknown threats and new types of cyber attacks, and the detection efficiency and accuracy are low, especially in high-dimensional and complex cyber attacks.

Method used

Unsupervised network traffic detection method based on large language model reasoning is adopted, network traffic data is processed textually, and labeled network traffic text with high similarity is determined using text retrieval tools, vectorized processing and similarity total score calculation are performed, and labels of network traffic text are obtained by combining inference prompt word templates and large language models.

Benefits of technology

It improves the efficiency and accuracy of network traffic detection, can accurately detect abnormal traffic in large-scale or complex network environments, and reduces computing resource requirements and model fine-tuning costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120474744A_ABST
    Figure CN120474744A_ABST
Patent Text Reader

Abstract

The invention relates to an unsupervised network flow detection method and device based on large language model reasoning, equipment and a medium. The method comprises the following steps: determining a first number of first annotated network traffic texts in a text retrieval tool, wherein the first similarity between the first annotated network traffic texts and obtained test network traffic texts meets a certain condition; determining a vector similarity between the test network traffic text and each first labeled network traffic text; based on the first similarity and the vector similarity, determining a total similarity score between the test network traffic text and each first labeled network traffic text, so as to determine a second number of second labeled network traffic texts in the first number of first labeled network traffic texts; and based on the second number of second labeled network traffic texts and the labels thereof, the test network traffic text and the reasoning prompt word template, obtaining a first label of the test network traffic text through large language model reasoning. By adopting the method, the efficiency and accuracy of network flow detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of network traffic detection, and in particular to an unsupervised network traffic detection method, apparatus, device and medium based on large language model reasoning. Background Art

[0002] With the rapid development of information technology, the means of network attacks have become more and more complex. Traditional rule-based or model-based abnormal network traffic detection methods have become difficult to cope with unknown threats or new network attack methods.

[0003] Currently, existing methods for detecting abnormal network traffic, including rule-based detection methods, statistical model-based detection methods, and machine learning-based detection methods, have made some progress in detection efficiency and accuracy, but they still face several challenges. Specifically, existing methods for detecting abnormal network traffic often rely on manual rule design or manual feature extraction, making them difficult to effectively respond to unknown network attack patterns. Furthermore, these existing methods are susceptible to performance limitations when dealing with high-dimensional and complex network attacks, making it difficult to improve detection efficiency and accuracy. Summary of the Invention

[0004] Based on this, it is necessary to provide an unsupervised network traffic detection method, device, computer equipment, computer-readable storage medium and computer program product based on large language model reasoning, which can improve the efficiency and accuracy of network traffic detection in response to the above technical problems.

[0005] In a first aspect, the present application provides an unsupervised network traffic detection method based on large language model reasoning, comprising:

[0006] Acquire test network traffic text; the test network traffic text is obtained by textually processing the network traffic data to be tested;

[0007] Determine, in a text search tool, a first number of first annotated network flow texts whose first similarity with the test network flow text satisfies a first preset condition; each first annotated network flow text is respectively associated with a label;

[0008] Performing vectorization processing on the test network traffic text and each first annotated network traffic text to obtain a test network traffic text vector and a first number of annotated network traffic text vectors, and determining the similarity between the test network traffic text vector and each annotated network traffic text vector as the vector similarity between the test network traffic text and each first annotated network traffic text;

[0009] Determine a total similarity score between the test network traffic text and each first annotated network traffic text based on the first similarity and vector similarity between the test network traffic text and each first annotated network traffic text;

[0010] Based on the total similarity scores between the test network traffic text and each of the first annotated network traffic texts, determining a second number of second annotated network traffic texts whose corresponding total similarity scores meet a second preset condition from the first number of first annotated network traffic texts;

[0011] A target inference prompt word is obtained based on a second number of second annotated network traffic texts, a label of each second annotated network traffic text, a test network traffic text, and an inference prompt word template, and the target inference prompt word is input into the large language model to obtain a first label of the test network traffic text.

[0012] In a second aspect, the present application further provides an unsupervised network traffic detection device based on large language model reasoning, comprising:

[0013] An acquisition module is used to acquire a test network traffic text; the test network traffic text is obtained by textually processing the network traffic data to be tested;

[0014] A retrieval module is used to determine, in a text retrieval tool, a first number of first annotated network flow texts whose first similarity with the test network flow text satisfies a first preset condition; each first annotated network flow text is respectively associated with a label;

[0015] a vectorization module, configured to perform vectorization processing on the test network traffic text and each first annotated network traffic text to obtain a test network traffic text vector and a first number of annotated network traffic text vectors, and determine the similarity between the test network traffic text vector and each annotated network traffic text vector as the vector similarity between the test network traffic text and each first annotated network traffic text;

[0016] A similarity calculation module is used to determine a total similarity score between the test network traffic text and each first annotated network traffic text based on the first similarity and vector similarity between the test network traffic text and each first annotated network traffic text;

[0017] a reordering module, configured to determine, from the first number of first annotated network traffic texts, a second number of second annotated network traffic texts whose corresponding total similarity scores satisfy a second preset condition, based on the total similarity scores between the test network traffic text and each of the first annotated network traffic texts;

[0018] The large language model inference module is used to obtain a target inference prompt word based on a second number of second annotated network traffic texts, a label of each second annotated network traffic text, a test network traffic text, and an inference prompt word template, and input the target inference prompt word into the large language model to obtain a first label for the test network traffic text.

[0019] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements some or all of the steps described in any method of the first aspect of the present application.

[0020] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps described in any method of the first aspect of the embodiments of the present application.

[0021] In a fifth aspect, the present application further provides a computer program product, comprising a computer program that, when executed by a processor, implements some or all of the steps described in any method of the first aspect of the present application.

[0022] The above-mentioned unsupervised network traffic detection method, device, computer equipment, computer-readable storage medium and computer program product based on large language model reasoning obtain a test network traffic text obtained by textual processing of the network traffic data to be detected; determine in a text retrieval tool a first number of first annotated network traffic texts whose first similarity with the test network traffic text meets a first preset condition; each first annotated network traffic text corresponds to a label; vectorize the test network traffic text and each first annotated network traffic text to obtain a test network traffic text vector and a first number of annotated network traffic text vectors, and determine the similarity between the test network traffic text vector and each annotated network traffic text vector as the similarity between the test network traffic text and each first annotated network traffic text vector. Vector similarity between traffic texts; based on the first similarity and vector similarity between the test network traffic text and each first labeled network traffic text, determining the total similarity score between the test network traffic text and each first labeled network traffic text; based on the total similarity score between the test network traffic text and each first labeled network traffic text, determining a second number of second labeled network traffic texts whose corresponding total similarity scores meet a second preset condition among the first number of first labeled network traffic texts; obtaining a target inference prompt word based on the second number of second labeled network traffic texts, the labels of each second labeled network traffic text, the test network traffic text and the inference prompt word template, and inputting the target inference prompt word into the large language model to obtain a first label for the test network traffic text. The unsupervised network traffic detection method based on large language model reasoning provided in the present application is firstly determined by using a text retrieval tool to determine a first number of first annotated network traffic texts having a high first similarity with the test network traffic text. Secondly, the vector similarity between the test network traffic text and each first annotated network traffic text is determined. Thirdly, the total similarity score between the test network traffic text and each first annotated network traffic text is determined respectively using the first similarity and the vector similarity, so as to determine a second number of second annotated network traffic texts having a high total similarity score from the first number of first annotated network traffic texts. Finally, the target inference prompt word is obtained by inputting the second number of second annotated network traffic texts and their labels and the test network traffic text into an inference prompt word template, so that the large language model obtains a first label for the test network traffic text based on the target inference prompt word. Obviously, the process of obtaining the first label of the test network traffic text in this embodiment is highly efficient, and the obtained first label has a high accuracy. Based on the first label with a high accuracy, it is possible to accurately determine whether the network traffic data to be detected is abnormal. That is, this embodiment can improve the efficiency and accuracy of network traffic detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0024] Figure 1 FIG1 is an application environment diagram of an unsupervised network traffic detection method based on large language model reasoning in one embodiment;

[0025] Figure 2 1 is a flow chart of an unsupervised network traffic detection method based on large language model reasoning in one embodiment;

[0026] Figure 3 1 is a structural block diagram of an unsupervised network traffic detection device based on large language model reasoning in one embodiment;

[0027] Figure 4 is a structural block diagram of an unsupervised network traffic detection device based on large language model reasoning in another embodiment;

[0028] Figure 5 is a diagram of the internal structure of a computer device in one embodiment;

[0029] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0031] The unsupervised network traffic detection method based on large language model reasoning provided by the embodiment of the present application can be applied to Figure 1In the application environment shown. The terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, projection devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0032] In an exemplary embodiment, Figure 2 As shown in the figure, an unsupervised network traffic detection method based on large language model reasoning is provided. Figure 1 The terminal in FIG is taken as an example to illustrate, including the following steps 202 to 212. Among them:

[0033] Step 202, obtaining a test network traffic text; the test network traffic text is obtained by textually processing the network traffic data to be tested.

[0034] Network traffic data refers to data generated based on the network behavior of the network requester and various network communication protocols and transmitted over the network. Network behavior may include user behavior, device interaction behavior, application behavior, and network attack behavior. Accordingly, network traffic data to be tested refers to network traffic data that needs to be tested to determine whether the network behavior that caused it is abnormal.

[0035] The test network traffic text, in this embodiment, is the network traffic text whose first tag needs to be determined.

[0036] Optionally, the network traffic data to be detected includes multiple data features related to the network traffic; the multiple data features related to the network traffic may include basic connection features, abnormal behavior features, statistical features, and target host features of the network traffic.

[0037] Basic connection characteristics refer to the basic attribute characteristics used to characterize the network connection corresponding to the network traffic data. Optionally, basic connection characteristics may include at least one of connection duration, network communication protocol type, target port service type, connection status flag, data transmission volume, and data reception volume.

[0038] Abnormal behavior characteristics refer to behavioral characteristics that indicate that the network connection corresponding to the network traffic data may contain attack behavior or illegal operation behavior. Optionally, abnormal behavior characteristics may include at least one of the following: the number of failed logins, the number of files created, the number of files accessed, the number of incorrect fragments, and the number of user switch command executions. Furthermore, in order to determine whether the network traffic to be tested is abnormal, the network traffic data to be tested should at least include abnormal behavior characteristics.

[0039] Statistical features refer to features used to characterize the statistical state of network traffic data within a specific time period. Optionally, the statistical features may include at least one of the number of connections within a preset time period, the number of connections of a selected service type within a preset time period, the proportion of connections of the same service type, the proportion of connections of different service types, the error rate of the source host in the connection, and the error rate of the source host in the selected service type.

[0040] The target host characteristics refer to characteristics used to characterize the network connection information of the target host selected in the network traffic data. Optionally, the target host characteristics may include at least one of the number of source hosts connected to the target host, the number of service types connected to the target host, and the error rate of the selected service type connected to the target host.

[0041] Specifically, since the network traffic data to be tested includes multiple data features related to the network traffic data, the test network traffic text obtained after textual processing of the network traffic data to be tested includes multiple text features. It is easy to understand that each text feature is obtained after the textual processing of each data feature, and each text feature corresponds to each data feature.

[0042] Optionally, the network traffic data to be detected is generated in an industrial Internet network environment or a large-scale network environment.

[0043] Optionally, the test network traffic text can be obtained by textually processing the network traffic data to be tested based on a text encoding method, a templated natural language generation method, or other textual processing methods.

[0044] Step 204 : determining in the text search tool a first number of first annotated network traffic texts whose first similarity with the test network traffic text satisfies a first preset condition; each first annotated network traffic text has a corresponding label.

[0045] Among them, the text retrieval tool includes multiple annotated network traffic texts and labels corresponding to each annotated network traffic text. The annotated network traffic text is a text with a predetermined label; the annotated network traffic text includes multiple text features, and the multiple text features included in the annotated network traffic text correspond to the multiple text features included in the test network traffic text; the text retrieval tool is used to search based on the multiple annotated network traffic texts included in the test network traffic text, and can determine the first similarity between each annotated network traffic text and the test network traffic text based on the retrieval results. Based on this, the text retrieval tool is a tool that can determine the first annotated network traffic text with a higher similarity to the test network traffic text among multiple annotated network traffic texts.

[0046] Optionally, the text retrieval tool may include an ES (Elasticsearch) index or other indexing tools that can be used to store, manage, and retrieve annotated network traffic text.

[0047] Optionally, satisfying the first preset condition may be that the first similarity is greater than or equal to the first preset similarity, or may be that the first similarities are ranked first in descending order. Exemplarily, when satisfying the first preset condition is that the first similarities are ranked first in descending order, the first number of first annotated network traffic texts satisfying the first preset condition refer to the first number of annotated network traffic texts having the highest first similarities among the multiple annotated network traffic texts.

[0048] The label of the annotated network traffic text refers to a label used to characterize the classification result of the annotated network traffic data corresponding to the annotated network traffic text. Optionally, there may be at least two classification results for the annotated network traffic data. Exemplarily, when there are two classification results, the classification result of the annotated network traffic data may be used to characterize the annotated network traffic data as a normal type of network traffic data or an abnormal type of network traffic data. Based on this, optionally, the label of the annotated network traffic text may be expressed as normal or abnormal. When the label of the annotated network traffic text is normal, it characterizes the annotated network traffic data as a normal type of network traffic data. When the label of the annotated network traffic text is abnormal, it characterizes the annotated network traffic data as an abnormal type of network traffic data. Exemplarily, when there are more than three classification results, the classification result of the annotated network traffic data may also be used to characterize the annotated network traffic data as a normal type of network traffic data, a first abnormal type of network traffic data, a second abnormal type of network traffic data...

[0049] Optionally, when there are two classification results for a label, the label may be a first result label or a second result label, and the first result label and the second result label may be labels indicating normal or abnormal, respectively. Optionally, when the first result label of the annotated network traffic text indicates normal, the annotated network traffic data corresponding to the annotated network traffic text is a normal type of network traffic data; when the second result label of the annotated network traffic text indicates abnormal, the annotated network traffic data corresponding to the annotated network traffic text is an abnormal type of network traffic data, that is, the annotated network traffic data is network attack traffic data. Abnormal network traffic data may be caused by network attacks, malware activities, configuration errors, or system failures.

[0050] The label of the first annotated network traffic text refers to a label used to represent the classification result of the annotated network traffic data corresponding to the first annotated network traffic text.

[0051] In an exemplary embodiment, the above-mentioned determining in a text retrieval tool a first number of first annotated network traffic texts whose first similarity with the test network traffic text satisfies a first preset condition, includes: in the text retrieval tool, respectively determining the first similarity between each annotated network traffic text in a plurality of annotated network traffic texts and the test network traffic text; and determining the first number of annotated network traffic texts whose first similarity satisfies the first preset condition as the first annotated network traffic text.

[0052] Specifically, when the text retrieval tool is the ES index, the ES index converts multiple data features of multiple network traffic data into text, thereby obtaining annotated network traffic text including multiple text features, and then saves the multiple annotated network traffic texts as multiple index documents in a preset storage format. Each index document contains multiple text features of the corresponding annotated network traffic text and its corresponding label category. In order to achieve efficient retrieval, the ES index adopts an inverted index structure, that is, the position of each text feature (or word) in the document collection will be recorded, so that the ES index can quickly find documents containing the text feature, that is, the annotated network traffic text containing the text feature.

[0053] Optionally, assuming that there are labeled network traffic text 1, labeled network traffic text 2, etc. in the ES index, and labeled network traffic text 1 corresponds to label 1, labeled network traffic text 2 corresponds to label 2, etc., then the preset storage format of the labeled network traffic text in the ES index can be: index: [{labeled network traffic text 1, label 1}, {labeled network traffic text 2, label 2}...], wherein labeled network traffic text 1 and labeled network traffic text 2 both include multiple text features, that is, each data item in the ES index includes multiple text features in text form and their corresponding labels.

[0054] Specifically, when the ES index performs text retrieval on the annotated network traffic text based on the test network traffic text, it first preprocesses the test network traffic text. The preprocessing includes word segmentation and stop word removal. The ES index then converts the test network traffic text into a query format suitable for retrieval within the ES index. Optionally, the query format can represent the test network traffic text as a query vector containing multiple text features. Furthermore, the query format can correspond to a preset storage format for the annotated network traffic text within the ES index. That is, the query format can be a query vector containing multiple text features of the test network traffic text.

[0055] Specifically, the ES index calculates the first similarity between the test network traffic text and each annotated network traffic text based on text similarity calculation methods such as the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm and the BM25 algorithm.

[0056] Specifically, when the ES index calculates the first similarity based on the TF-IDF algorithm, the calculation of the first similarity is measured based on the following factors: Term Frequency (TF), that is, the frequency of a text feature in the test network traffic text appearing in each annotated network traffic text. The higher the term frequency of a text feature, the more important the text feature is in the test network traffic text; Inverse Document Frequency (IDF), that is, the frequency of a text feature in the test network traffic text appearing in the test network traffic text. Frequently appearing text features (for example, common words, stop words, etc.) will be assigned a lower IDF value, and rare text features will obtain a higher IDF value. The higher the IDF value of a text feature, the better the text feature can help distinguish each annotated network traffic text. In the process of calculating the first similarity based on the TF-IDF algorithm, in order to avoid the text length of the test network traffic text and each annotated network traffic text being too long or too short to affect the calculation of the first similarity, the length of the test network traffic text and each annotated network traffic text needs to be normalized.

[0057] Specifically, when the ES index calculates the first similarity based on the BM25 algorithm, in addition to considering the TF value and IDF value, the calculation of the first similarity also needs to consider a first adjustment coefficient for controlling the nonlinearity of the word frequency of the text feature and a second adjustment coefficient for controlling the influence of text length on word frequency. Optionally, there is a positive correlation between the word frequency of the text feature and the degree to which the text feature improves the first similarity, and a negative correlation between the text length and the influence of word frequency. Compared with the TF-IDF algorithm, the BM25 algorithm can more accurately determine the first similarity between the test network traffic text and each annotated network traffic text.

[0058] Specifically, after the ES index calculates the first similarity between the test network traffic text and each annotated network traffic text, the ES index will sort the annotated network traffic texts according to each first similarity. The annotated network traffic texts with higher first similarities will be ranked closer to the front, and the annotated network traffic texts ranked closer to the front have higher first similarities with the test network traffic text. Ultimately, the ES index returns to the terminal the first number of first annotated network traffic texts that are most similar to the test network traffic text and the first similarity corresponding to each first annotated network traffic text.

[0059] Specifically, when the retrieval tool is ES index, in order to avoid the dimensionality problem of the first similarity, ES index will normalize the calculated first similarity to ensure that all first similarities are on a unified scale (for example, 0 to 1). The normalized first similarity represents the strength of the correlation between the test network traffic text and each first annotated network traffic text (or each annotated network traffic text).

[0060] In another exemplary embodiment, when the text retrieval tool is an ES index, the above-mentioned first similarity between each of the multiple annotated network traffic texts and the test network traffic text is determined in the text retrieval tool, including: based on the word frequency and inverse document frequency of each text feature in the test network traffic text, the first similarity between each of the multiple annotated network traffic texts and the test network traffic text is determined; or, based on the word frequency, inverse document frequency, first adjustment coefficient and second adjustment coefficient of each text feature in the test network traffic text, the first similarity between each of the multiple annotated network traffic texts and the test network traffic text is determined; the first adjustment coefficient is used to control the nonlinearity of the word frequency of the text feature, and the second adjustment coefficient is used to control the influence of text length on word frequency.

[0061] Optionally, the first adjustment coefficient may be 1.2 to 2 or other values, and the second adjustment coefficient may be 0.75 or other values.

[0062] Step 206, vectorize the test network traffic text and each first annotated network traffic text to obtain a test network traffic text vector and a first number of annotated network traffic text vectors, and determine the similarity between the test network traffic text vector and each annotated network traffic text vector as the vector similarity between the test network traffic text and each first annotated network traffic text.

[0063] The test network traffic text and each first labeled network traffic text may be vectorized using a large language model and / or label encoding to obtain a test network traffic text vector and a first number of labeled network traffic text vectors. Specifically, the test network traffic text vector corresponds to the test network traffic text, and the labeled network traffic text vector corresponds to the first labeled network traffic text.

[0064] Optionally, the similarity between the test network traffic text vector and each annotated network traffic text vector can be determined by cosine similarity and / or cosine similarity with feature weight coefficients between the test network traffic text vector and each annotated network traffic text vector.

[0065] Specifically, by vectorizing the test network traffic text and each first annotated network traffic text, the multiple text features included in the test network traffic text and the multiple text features included in each first annotated network traffic text can be quantified to accurately quantify the vector similarity between the test network traffic text and each first annotated network traffic text.

[0066] Step 208 : Determine a total similarity score between the test network traffic text and each first annotated network traffic text based on the first similarity and vector similarity between the test network traffic text and each first annotated network traffic text.

[0067] Among them, the total similarity score between the test network traffic text and each first-annotated network traffic text can be the sum of the first similarity and the vector similarity between the test network traffic text and each first-annotated network traffic text, that is, the total similarity score = first similarity + vector similarity; when the first similarity and vector similarity respectively correspond to the first similarity weight coefficient and the vector similarity weight coefficient, the total similarity score between the test network traffic text and each first-annotated network traffic text can also be the weighted summation result of the first similarity, the first similarity weight coefficient, the vector similarity and the vector similarity weight coefficient between the test network traffic text and each first-annotated network traffic text, that is, the total similarity score = first similarity × first similarity weight coefficient + vector similarity × vector similarity weight coefficient.

[0068] Step 210: Based on the total similarity scores between the test network traffic text and each first annotated network traffic text, determine a second number of second annotated network traffic texts whose corresponding total similarity scores meet a second preset condition from the first number of first annotated network traffic texts.

[0069] Here, satisfying the second preset condition may be that the total similarity score is greater than or equal to the second preset similarity, or that the second similarity is ranked first in a second number of items in descending order. For example, when satisfying the second preset condition is that the total similarity score is ranked first in a second number of items in descending order, the second number of second annotated network traffic texts satisfying the second preset condition refers to the first number of first annotated network traffic texts having the highest total similarity scores among the first number of first annotated network traffic texts.

[0070] It is easy to understand that the first number is at least one, the second number is at least one, and the second number is smaller than the first number.

[0071] Specifically, since each first annotated network flow text corresponds to a label, and the second number of second annotated network flow texts are determined from the first number of first annotated network flow texts, accordingly, each second annotated network flow text also corresponds to a label.

[0072] The label of the second annotated network traffic text refers to a label used to represent the classification result of the annotated network traffic data corresponding to the second annotated network traffic text.

[0073] Step 212: Obtain a target inference prompt word based on the second number of second annotated network traffic texts, the labels of each second annotated network traffic text, the test network traffic text, and the inference prompt word template, and input the target inference prompt word into the large language model to obtain a first label for the test network traffic text.

[0074] The inference prompt word template refers to an input instruction template used to guide the large language model to infer the first label of the test network traffic text through a structured logical framework. The input instruction template has not yet been filled with the second number of second annotated network traffic texts, the labels of each second annotated network traffic text, and the test network traffic text. In other words, the inference prompt word template is an input instruction framework that is universal for different second annotated network traffic texts and different test network traffic texts.

[0075] It is easy to understand that after the second number of second annotated network flow texts, the labels of the second annotated network flow texts, and the test network flow text are filled into the inference prompt word template, the target inference prompt word is obtained.

[0076] The target inference prompt word refers to the input instruction generated after filling in the inference prompt word template with the second number of second annotated network traffic texts, the labels of each second annotated network traffic text, and the test network traffic text, and is used to guide the large language model to infer the first label of the test network traffic text through a structured logical framework.

[0077] For example, assuming that the second number is k, the multiple text features included in each second annotated network flow text are represented as ti1, ti2... (i=1, 2... k), and the labels of each second annotated network flow text are represented as lab1, lab2... lab k , the multiple text features included in the test network traffic text are represented as t1, t2, ..., then the target reasoning prompt word can be "Please act as a network traffic anomaly detection expert, based on the k feature-label pairs provided, classify the test feature as one of the two categories: attack / normal.

[0078] Rank1 features:t 11 ,t 12 ...,label:lab1

[0079] Rank2 features:t 21 ,t 22 ...,label:lab2 ...

[0081] Rankk features:t k1 ,t k2 ...,label:lab k

[0082] test features:t1,t2...,label:?"; In another exemplary embodiment, the target reasoning prompt word can also be "Please play the role of a network traffic anomaly detection expert, and according to the k text feature-label pairs provided below, determine that the test network traffic text including multiple text features is one of the following two categories: attack (attack) or normal (normal).

[0083] The first second annotated network traffic text text feature-label pair: t 11 , t 12 ...label:lab1

[0084] The second second annotated network traffic text text feature-label pair: t 21 , t 22 ...label:lab2 ...

[0086] The third second annotated network traffic text feature-label pair: t k1 , t k2 ...label:lab k

[0087] Test the text features of network traffic text - the first label pair: t1, t2...label:?"

[0088] Optionally, the “label:?” in the above two target reasoning prompts can be lab1, lab2…lab k One of the label results after removing repeated labels from a total of k labels.

[0089] It should be noted that the target reasoning prompt words provided in this example are merely examples of target reasoning prompt words. In actual implementation, target reasoning prompt words can also be designed in other styles, and this application does not impose specific restrictions here. Furthermore, target reasoning prompt words can be written in languages including, but not limited to, Chinese, English, and other languages.

[0090] A Large Language Model (LLM) is a natural language processing model based on deep learning. It can understand and reason about natural language and generate corresponding responses. In this embodiment, the LLM can infer the first label of the test network traffic text using the target inference prompt word.

[0091] Optionally, the large language model can be ChatGPT, DeepSeek, Kimi, Wenxin Yiyan, Zhipu Qingyan or other large language models, as long as it can realize the function of obtaining the first label of the test network traffic text through target inference prompt words. This application does not limit the specific type and version number of the large language model.

[0092] Specifically, in this embodiment, a large language model is used to infer the first label of the test network traffic text. However, unlike traditional methods, in this embodiment, there is no need to fine-tune the large language model. Instead, an efficient inference prompt word template is designed to generate target inference prompt words for different textual network traffic. Thus, the large language model is driven to infer the first label of the test network traffic text based on the efficient target inference prompt words, thereby reducing the complexity of the large language model inference process and improving the efficiency and accuracy of network traffic detection.

[0093] The first label of the test network traffic text refers to a label used to characterize the classification result of the network traffic data to be detected corresponding to the test network traffic text. Optionally, there can be at least two classification results of the test network traffic text. Exemplarily, when there are two classification results, the classification result of the network traffic data to be detected can be used to characterize that the network traffic data to be detected is a normal type of network traffic data or an abnormal type of network traffic data. Based on this, optionally, the first label of the test network traffic text can be expressed as normal or abnormal. When the first label of the test network traffic text is normal, it characterizes that the network traffic data to be detected is a normal type of network traffic data. When the first label of the test network traffic text is abnormal, it characterizes that the network traffic data to be detected is an abnormal type of network traffic data. Exemplarily, when there are more than three classification results, the classification result of the network traffic data to be detected can also be used to characterize that the network traffic data to be detected is a normal type of network traffic data, a first abnormal type of network traffic data, a second abnormal type of network traffic data...

[0094] Optionally, when there are two label classification results, similarly to the label for annotating network traffic text, the first label of the test network traffic text can be a first result label or a second result label. Optionally, when the first result label of the test network traffic text indicates normal, it indicates that the network traffic data to be detected corresponding to the test network traffic text is normal type of network traffic data; when the second result label of the test network traffic text indicates abnormal, it indicates that the network traffic data to be detected corresponding to the test network traffic text is abnormal type of network traffic data, that is, the network traffic data to be detected is network attack traffic data.

[0095] In an exemplary embodiment, when the second result label indicates an anomaly, after inputting the target inference prompt word into the large language model to obtain the first label of the test network traffic text, the above method also includes: when the first label of the test network traffic text is the second result label, generating an alarm signal for the network traffic data to be detected and / or taking security defense measures for the network traffic to be detected.

[0096] In the above-mentioned unsupervised network traffic detection method based on large language model reasoning, first, a first number of first annotated network traffic texts having a high first similarity with the test network traffic text are determined through a text retrieval tool. Second, the vector similarity between the test network traffic text and each first annotated network traffic text is determined. Third, the total similarity score between the test network traffic text and each first annotated network traffic text is determined respectively through the first similarity and the vector similarity, so as to determine a second number of second annotated network traffic texts having a high total similarity score from the first number of first annotated network traffic texts. Finally, the target inference prompt word is obtained by inputting the second number of second annotated network traffic texts and their labels and the test network traffic text into the inference prompt word template, so that the large language model obtains the first label of the test network traffic text based on the target inference prompt word. Obviously, the process of obtaining the first label of the test network traffic text in this embodiment has high efficiency, and the obtained first label has high accuracy. Based on the first label with high accuracy, it is possible to accurately determine whether the network traffic data to be detected is abnormal. That is, this embodiment can improve the efficiency and accuracy of network traffic detection.

[0097] Specifically, the unsupervised network traffic detection method based on large language model reasoning provided by this application can effectively improve the efficiency and accuracy of network traffic detection. Therefore, this embodiment is particularly suitable for detecting network traffic in large-scale or complex network traffic data environments. In large-scale or complex network traffic data environments, this embodiment can significantly enhance network security protection capabilities because it can accurately detect potential abnormal network traffic.

[0098] Specifically, the unsupervised network traffic detection method based on large language model reasoning provided by this application adopts an unsupervised learning method, that is, no fine-tuning training is required, thereby avoiding the need for traditional network traffic detection methods to rely on a large amount of labeled data for fine-tuning training, effectively reducing the computational cost and time overhead of model fine-tuning using the method provided by this application. At the same time, the prompt word driven mechanism provided by this embodiment can also enable more efficient detection of network traffic during the reasoning process, significantly reducing the demand for computing resources.

[0099] In an exemplary embodiment, the method further includes:

[0100] A plurality of labeled network flow data are obtained; the labeled network flow data include a plurality of data features, and each labeled network flow data has a corresponding label.

[0101] Determine the correlation coefficient between each data feature and the label.

[0102] The correlation coefficient between each data feature and the label is determined as the feature weight coefficient of each text feature.

[0103] The multiple data features included in the annotated network traffic data correspond to the multiple data features included in the network traffic data to be detected.

[0104] Labels for annotating network traffic data refer to labels used to represent the classification results of the annotated network traffic data. Similar to labels for annotating network traffic text, labels for annotating network traffic data can be expressed as a first result label or a second result label, or as normal, abnormal, or other representations.

[0105] In an exemplary embodiment, the determining of the correlation coefficient between each data feature and the label includes: determining the correlation coefficient between each data feature and the label based on a Pearson correlation analysis method.

[0106] Exemplarily, the number of multiple data features of the labeled network traffic data is represented as t, and the process of determining the correlation coefficient between the j-th data feature (j=1~t) and the label is described. The number of multiple labeled network traffic data is represented as n, the data feature j of the i-th labeled network traffic data is represented as x, the feature mean of the data feature j of the n labeled network traffic data is represented as x', the label of the i-th labeled network traffic data is represented as y, the label mean of the n labels is represented as y', and the correlation coefficient between the data feature j and the label is represented as w j , then w j The calculation formula can be expressed as:

[0107]

[0108] Specifically, since there is a one-to-one correspondence between each text feature and each data feature, the correlation coefficient between each data feature and the label can be determined as the feature weight coefficient of each text feature.

[0109] In an exemplary embodiment, the above-mentioned obtaining of multiple annotated network traffic data includes: obtaining multiple initial network traffic data; performing data cleaning, missing value filling and denoising on the multiple initial network traffic data to obtain multiple annotated network traffic data.

[0110] Specifically, after data cleaning, missing value filling and denoising, the multiple labeled network traffic data do not contain redundant information, irrelevant data, missing values, and do not contain abnormal points or noise points.

[0111] In this embodiment, by determining the correlation coefficient between each data feature and the label, and determining the correlation coefficient between each data feature and the label as the feature weight coefficient of each text feature, thereby identifying the closeness of the relationship between each data feature and the label, the degree of influence of each data feature on the label is quantified, and further, based on the determined feature weight coefficient of each text feature, the third similarity between the test network traffic text and each first labeled network traffic text can be accurately determined to improve the efficiency and accuracy of network traffic detection.

[0112] In an exemplary embodiment, the test network traffic text and each first annotated network traffic text respectively correspond to a plurality of text features.

[0113] The above-mentioned vectorization processing of the test network traffic text and each first labeled network traffic text to obtain a test network traffic text vector and a first number of labeled network traffic text vectors, and determining the similarity between the test network traffic text vector and each labeled network traffic text vector as the vector similarity between the test network traffic text and each first labeled network traffic text, includes:

[0114] The test network traffic text and each first annotated network traffic text are vectorized based on a large language model to obtain a test network traffic text semantic vector and a first number of annotated network traffic text semantic vectors, and the similarity between the test network traffic text semantic vector and each annotated network traffic text semantic vector is determined as the second similarity between the test network traffic text and each first annotated network traffic text.

[0115] The test network traffic text and each first annotated network traffic text are vectorized based on the label encoding method to obtain a test network traffic text feature vector and a first number of annotated network traffic text feature vectors, and the similarity between the test network traffic text feature vector and each annotated network traffic text feature vector is determined based on the feature weight coefficient of each text feature, and the similarity between the test network traffic text feature vector and each annotated network traffic text feature vector is determined as the third similarity between the test network traffic text and each first annotated network traffic text.

[0116] The above-mentioned determining the total similarity score between the test network traffic text and each first annotated network traffic text based on the first similarity and vector similarity between the test network traffic text and each first annotated network traffic text includes:

[0117] Based on the first similarity, the second similarity and the third similarity between the test network traffic text and each first annotated network traffic text, a total similarity score between the test network traffic text and each first annotated network traffic text is determined.

[0118] Among them, the test network traffic text and each first labeled network traffic text are vectorized based on the large language model, and the test network traffic text and each first labeled network traffic text are word embedded based on the large language model to respectively obtain a test network traffic text semantic vector represented by a vector for the test network traffic text and a labeled network traffic text semantic vector represented by a vector for each first labeled network traffic text.

[0119] Specifically, word embedding processing is a vectorization processing technology that converts text features into numerical vectors. Word embedding processing is performed through a large language model. Since the large language model can capture the deep semantic information contained in the test network traffic text and each first labeled network traffic text, the test network traffic text and each first labeled network traffic text are vectorized based on the large language model to obtain a test network traffic text semantic vector and a labeled network traffic text semantic vector that are highly accurate and can fully reflect the semantic information. Based on this, the similarity between the test network traffic text semantic vector and the semantic vectors of each labeled network traffic text can reflect the semantic similarity between the first test network traffic text and each first labeled network traffic text, that is, the second similarity is the semantic similarity.

[0120] Optionally, the test network traffic text and each first labeled network traffic text are vectorized based on the label encoding method, and the test network traffic text and each first labeled network traffic text are vectorized based on LabelEncoder to respectively obtain a test network traffic text feature vector represented by a high-dimensional sparse vector and a labeled network traffic text feature vector represented by a high-dimensional sparse vector for each first labeled network traffic text.

[0121] Specifically, LabelEncoder is a tool for converting classification labels into numerical form. Since LabelEncoder can vectorize the test network traffic text and each first labeled network traffic text to obtain an intuitive test network traffic text feature vector and each labeled network traffic text feature vector expressed in numerical form, by converting the text features into an intuitive numerical form, the text similarity between the test network traffic text and each first labeled network traffic text can be reflected, that is, the third similarity is the text similarity.

[0122] It should be noted that the label encoding method refers to encoding the text content vector label in the test network traffic text feature vector and the text content vector label in the labeled network traffic text feature vector after the test network traffic text and each first labeled network traffic text are vectorized. That is to say, the "label" in the label encoding method refers to the text content vector label, rather than the label that represents the classification results of the test network traffic text and each first labeled network traffic text.

[0123] Specifically, the test network traffic text semantic vector and the test network traffic text feature vector both correspond to the test network traffic text, and the annotated network traffic text semantic vector and the annotated network traffic text feature vector both correspond to the first annotated network traffic text.

[0124] Optionally, the total similarity score between the test network traffic text and each first-annotated network traffic text may be the sum of the first similarity, the second similarity and the third similarity between the test network traffic text and each first-annotated network traffic text, that is, the total similarity score = the first similarity + the second similarity + the third similarity; when the first similarity, the second similarity and the third similarity respectively correspond to the first similarity weight coefficient, the second similarity weight coefficient and the third similarity weight coefficient, the total similarity score between the test network traffic text and each first-annotated network traffic text may also be the weighted sum of the first similarity, the first similarity weight coefficient, the second similarity, the second similarity weight coefficient, the third similarity and the third similarity weight coefficient between the test network traffic text and each first-annotated network traffic text, that is, the total similarity score = the first similarity × the first similarity weight coefficient + the second similarity × the second similarity weight coefficient + the third similarity × the third similarity weight coefficient.

[0125] Specifically, when the first similarity is calculated based on the ES index, the second similarity is calculated based on the cosine similarity, and the third similarity is calculated based on the cosine similarity with the feature weight coefficient, based on the properties of the trigonometric function, it can be seen that the value range of the second similarity and the third similarity is 0 to 1. Therefore, at this time, the first similarity is the similarity after normalization processing, that is, the value range of the first similarity is also 0 to 1.

[0126] Exemplarily, when the second similarity is calculated based on cosine similarity, the process of determining the similarity between the test network traffic text semantic vector and each annotated network traffic text semantic vector is described. The test network traffic text semantic vector is represented as A, the annotated network traffic text semantic vector is represented as B, and the similarity between the test network traffic text semantic vector and each annotated network traffic text semantic vector is represented as cosine_similarity(A, B). The calculation formula of cosine_similarity(A, B) can be expressed as:

[0127]

[0128] Optionally, the second similarity may be calculated based on Euclidean distance or other methods in addition to cosine similarity. Specifically, the closer the cosine similarity is to 1, the higher the second similarity is, and the smaller the Euclidean distance is, the higher the second similarity is.

[0129] For example, when the third similarity is calculated based on the cosine similarity with a feature weight coefficient, the process of determining the similarity between the test network traffic text feature vector and each labeled network traffic text feature vector is described. Since the correlation coefficient between each data feature and the label is the feature weight coefficient of each text feature, the multiple text feature vectors included in the test network traffic text feature vector are represented as a1, a2...a t , the multiple text feature vectors included in the labeled network traffic text feature vector are represented as b1, b2...b t , the feature weight coefficients of each text feature are expressed as w1, w2...w t , the similarity between the test network traffic text feature vector and each labeled network traffic text feature vector is expressed as Weighted Cosine Similarity, and the calculation formula of Weighted Cosine Similarity can be expressed as:

[0130]

[0131] Optionally, in addition to being calculated based on the cosine similarity with a feature weight coefficient, the third similarity may also be calculated based on the Euclidean distance with a feature weight coefficient or in other ways.

[0132] In this embodiment, a large language model is used to vectorize the test network traffic text and each first-labeled network traffic text to determine the semantic similarity between the test network traffic text and each first-labeled network traffic text, and a label encoding method is used to vectorize the test network traffic text and each first-labeled network traffic text to determine the text similarity between the test network traffic text and each first-labeled network traffic text. Thus, on the one hand, this embodiment combines the semantic understanding ability of the large language model with the text feature extraction ability of the label encoding method, and can more accurately and comprehensively capture the similarity between the test network traffic text and each first-labeled network traffic text, thereby obtaining a similarity total score that more accurately reflects the similarity between the texts. On the other hand, this embodiment designs a re-ranking mechanism for the first similarity, the second similarity, and the third similarity to obtain the total similarity score. By re-ranking the scores of similar texts, it is possible to quickly and accurately locate and filter out the first-labeled network traffic text with the highest similarity to the test network traffic text. Obviously, this embodiment can significantly improve the efficiency and accuracy of network traffic detection.

[0133] In an exemplary embodiment, the method further includes:

[0134] Determine a second label of a second annotated network traffic text having the highest first similarity to the test network traffic text among the second number of second annotated network traffic texts.

[0135] Determine, among the second number of second annotated network traffic texts, a third label of the second annotated network traffic text corresponding to the annotated network traffic text semantic vector having the second highest similarity with the test network traffic text semantic vector.

[0136] Determine, among the second number of second annotated network traffic texts, a fourth label of the second annotated network traffic text corresponding to the annotated network traffic text feature vector having the third highest similarity to the test network traffic text feature vector.

[0137] A target label of the test network traffic text is determined based on the first label, the second label, the third label, and the fourth label.

[0138] The test network traffic text, in this embodiment, is a network traffic text whose target tag needs to be determined.

[0139] Specifically, since the second annotated network traffic text is determined in the first annotated network traffic text, the second label, third label and fourth label of the second annotated network traffic text all correspond to the label of the first annotated network traffic text, that is, the second label, third label and fourth label can all be expressed as the first result label or the second result label. In other words, they can all be expressed as normal or abnormal or in other ways.

[0140] Optionally, the target label of the test network traffic text is determined based on the first label, the second label, the third label, and the fourth label, and the result with the largest number of occurrences can be directly used as the target label of the test network traffic text. For example, if the first label, the second label, and the third label all indicate normal, and the fourth label indicates abnormal, then the result indicating normal, which accounts for the majority, can be directly determined as the target label of the test network traffic text. In other words, the target label at this time indicates normal.

[0141] For example, assuming that the label indicates normal or abnormal, there are second-labeled network traffic text 1, second-labeled network traffic text 2 and second-labeled network traffic text 3 corresponding to labels indicating normal, normal and abnormal respectively; the first similarity, second similarity and third similarity corresponding to the second-labeled network traffic text 1 are 0.9, 0.8 and 0.7 respectively, the first similarity, second similarity and third similarity corresponding to the second-labeled network traffic text 2 are 0.8, 0.9 and 0.6 respectively, and the first similarity, second similarity and third similarity corresponding to the second-labeled network traffic text 3 are 0.6, 07 and 0.9 respectively. At this time, the first similarity corresponding to the second-labeled network traffic text 1 is the highest, the second similarity of the second-labeled network traffic text 2 is the highest and the third similarity corresponding to the second-labeled network traffic text 3 is the highest. Therefore, at this time, the label of the second-labeled network traffic text 1 is the second label, that is, the second label indicates normal, the label of the second-labeled network traffic text 2 is the third label, that is, the third label indicates normal, and the label of the second-labeled network traffic text 3 is the fourth label, that is, the fourth label indicates abnormal.

[0142] In an exemplary embodiment, after determining the target label of the test network traffic text based on the first label, the second label, the third label and the fourth label, the above method also includes: when the target label is the second result label, generating an alarm signal for the network traffic data to be tested and / or taking security defense measures for the network traffic to be tested.

[0143] In this embodiment, by respectively determining the second annotated network traffic text with the highest first similarity, the highest second similarity, and the highest third similarity with the test network traffic text among the second number of second annotated network traffic texts, the corresponding second label, third label, and fourth label are obtained respectively, so as to finally determine the target label of the test network traffic text by jointly determining the first label, the second label, the third label, and the fourth label. Obviously, since different similarities are obtained through different calculation methods, this embodiment realizes the comprehensive consideration of the target label of the test network traffic text from multiple angles. On the one hand, it can more comprehensively capture the similarity between the test network traffic text and different second annotated network traffic texts. On the other hand, by dynamically adjusting the weight coefficients corresponding to each similarity to select the result label with the largest weight coefficient as the target label, it can effectively balance the contribution of each text feature to the target label. Therefore, this embodiment can significantly improve the efficiency and accuracy of network traffic detection while improving the generalization ability.

[0144] In an exemplary embodiment, determining the target tag of the test network traffic text based on the first tag, the second tag, the third tag, and the fourth tag includes:

[0145] A first weighted result between the first weight coefficient and the first label, a second weighted result between the second weight coefficient and the second label, a third weighted result between the third weight coefficient and the third label, and a fourth weighted result between the fourth weight coefficient and the fourth label are determined respectively.

[0146] Based on the first weighted result, the second weighted result, the third weighted result, and the fourth weighted result, a target label of the test network traffic text is determined.

[0147] The first weight coefficient is respectively greater than the second weight coefficient, the third weight coefficient and the fourth weight coefficient.

[0148] Optionally, the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient may be pre-set in size and stored in the terminal.

[0149] Specifically, the first weighted result is the product of the first weight coefficient and the first label, the second weighted result is the product of the second weight coefficient and the second label, the third weighted result is the product of the third weight coefficient and the third label, and the fourth weighted result is the product of the fourth weight coefficient and the fourth label.

[0150] Specifically, the first weight coefficient is respectively greater than the second weight coefficient, the third weight coefficient and the fourth weight coefficient, and the first weight coefficient corresponds to the first label. Therefore, the target label of the test network traffic text will depend more on the inference result made by the large language model based on the target inference prompt word. In other words, the target label of the test network traffic text will be more inclined to refer to the inference result of the large language model.

[0151] Exemplarily, the first weight coefficient, the second weight coefficient, the third weight coefficient and the fourth weight coefficient are represented as λ1, λ2, λ3 and λ4 respectively, and the first label, the second label, the third label and the fourth label are represented as lab1, lab2, lab3 and lab4 respectively, then the first weighted result = λ1×lab1, the second weighted result = λ2×lab2, the third weighted result = λ3×lab3, and the fourth weighted result = λ4×lab4.

[0152] In an exemplary embodiment, determining a target label for the test network traffic text based on the first weighted result, the second weighted result, the third weighted result, and the fourth weighted result includes:

[0153] Performing similar item merging processing on the first weighted result, the second weighted result, the third weighted result, and the fourth weighted result, and accumulating the weighted results of the same result label to obtain the total weight coefficient corresponding to each result label;

[0154] The total weight coefficients of the result labels are compared, and the result label with the largest total weight coefficient is determined as the target label of the test network traffic text.

[0155] Among them, since there may be some identical labels in the first label, the second label, the third label and the fourth label, accordingly, there may also be some identical result labels in the first weighted result, the second weighted result, the third weighted result and the fourth weighted result. Therefore, it is necessary to accumulate the weighted results of the same result labels to obtain the total weight coefficient corresponding to each result label.

[0156] Exemplarily, assuming that λ1=0.7, λ2=0.1, λ3=0.1, λ4=0.1, lab1 is the first result label, lab2 is the first result label, lab3 is the second result label, and lab4 is the second result label, then the first weighted result = λ1×lab1=0.7×first result label, the second weighted result = λ2×lab2=0.1×first result label, the third weighted result = λ3×lab3=0.1×second result label, and the fourth weighted result = λ4×lab4=0.1×second result label. At this time, the first weighted result, the second weighted result, the third weighted result and the fourth weighted result are merged for similar items, wherein the first weighted result and the second weighted result are similar items, and the third weighted result and the fourth weighted result are similar items. Thus, the first weighted result + the second weighted result = 0.8×first result label, and the third weighted result + the fourth weighted result = 0.2×second result label. It can be seen that the label weight coefficient of the first result label is 0.8, and the label weight coefficient of the second result label is 0.2, that is, the label weight coefficient of the first result label is greater than the label weight coefficient of the second result label. Therefore, the first result label is determined as the target label of the test network traffic text. It should be noted that this exemplary embodiment uses the case where there are two result labels in lab1, lab2, lab3 and lab4 for illustration, but in the specific implementation, lab1, lab2, lab3 and lab4 also include the case where there are one, three or more result labels. The implementation ideas for determining the target label of the test network traffic text in these different cases are the same as those of this exemplary embodiment, so they will not be repeated here.

[0157] In this embodiment, after determining the first label, second label, third label and fourth label obtained by different similarity calculation methods, a classification reordering mechanism of a weighted summation strategy is implemented. Specifically, the label weight coefficients of different label categories are calculated through the weighted summation strategy, and the optimal label is selected as the target label of the test network traffic text based on the label weight coefficients of different label categories. Thus, the second labeled network traffic text with the highest similarity to the test network traffic text can be accurately screened out based on the weighted summation result. Furthermore, this embodiment further improves the efficiency and accuracy of network traffic detection, that is, makes the target label more robust.

[0158] In an exemplary embodiment, the reasoning prompt word template includes a role setting element, a reasoning task content element, a guidance text element for inputting a second number of second annotated network traffic texts and a label for each second annotated network traffic text, a reasoning text element for inputting a test network traffic text, and an output format element.

[0159] The method of obtaining a target inference prompt word based on the second number of second annotated network traffic texts, the labels of each second annotated network traffic text, the test network traffic text, and the inference prompt word template, and inputting the target inference prompt word into the large language model to obtain a first label for the test network traffic text, includes:

[0160] Determine the role setting element, the reasoning task content element and the output format element, input the second number of second annotated network traffic texts and the label of each second annotated network traffic text into the guidance text element, and input the test network traffic text into the reasoning text element to obtain the target reasoning prompt word.

[0161] The target reasoning prompt word is input into the large language model so that the large language model can infer the first label of the test network traffic text based on the role setting elements, the reasoning task content elements, the second number of second annotated network traffic texts, the labels of each second annotated network traffic text and the test network traffic text.

[0162] The role-setting elements are used to make the large language model clear about the role it is set to play during the reasoning process. For example, the target reasoning prompts "Please act as a network traffic anomaly detection expert" and "Please play the role of a network traffic anomaly detection expert" are role-setting elements.

[0163] Reasoning task content elements refer to the elements used to describe the specific tasks that a large language model needs to complete. For example, the target reasoning prompts "Based on the following feature-label pairs provided, classify the test feature as one of the two categories: normal or attack" and "Based on the following k text feature-label pairs provided, classify the test network traffic text including multiple text features as one of the following two categories: attack or normal" are reasoning task content elements.

[0164] For example, the target reasoning prompt word "Rank1 features:t 11 ,t 12 ...,label:lab1

[0165] Rank2 features:t 21 ,t 22 ...,label:lab2 ...

[0167] Rankk features:t k1 ,t k2 ...,label:lab k ” and “The first and second labeled network traffic text feature-label pairs: t 11 , t 12 ...label:lab1

[0168] The second second annotated network traffic text text feature-label pair: t 21 , t 22 ...label:lab2 ...

[0170] The third second annotated network traffic text feature-label pair: t k1 , t k2 ...label:lab k ”, which is the guiding text element.

[0171] For example, the target reasoning prompt words “test features: t1, t2..., label:?” and “text features of test network traffic text - first label pair: t1, t2...label:?” are reasoning text elements.

[0172] The output format element is an element used to define the specific format of the output first label. For example, the target inference prompt words "attack / normal" and "attack (attack) or normal (normal)" are output format elements.

[0173] In an exemplary embodiment, the above-mentioned determination of the role setting elements, the reasoning task content elements and the output format elements includes: determining the role setting elements based on the role played by the large language model in the reasoning process of the first label; determining the reasoning task content elements based on the reasoning task for the first label that the large language model needs to complete; and determining the output format elements based on the expected format of the first label output by the large language model.

[0174] Among them, since this embodiment provides an unsupervised network traffic detection method based on large language model reasoning, the role setting element can be to specify the large language model to play the role of a network traffic detection expert, and further, it can be to specify the large language model to play the role of a network traffic detection expert who can detect abnormal network traffic in the industrial Internet; the reasoning task content element can be to specify the large language model to infer the first label of the test network traffic text based on the second number of second annotated network traffic texts, the labels of each second annotated network traffic text and the test network traffic text; the output format element can be to specify the large language model to output the first label of the test network traffic text as one of the label results after removing the repeated labels of the labels of each second annotated network traffic text, or to specify the large language model to output the first label of the test network traffic text as the first result label or the second result label, or to output it as "normal" or "abnormal".

[0175] Specifically, by setting the role setting element in the target reasoning prompt word, the large language model can understand the specific role it needs to play and its specific perspective in the reasoning process, thereby being able to infer a first label that is more in line with expectations.

[0176] Specifically, by setting the reasoning task content elements in the target reasoning prompt words, the large language model can understand the specific tasks it needs to perform during the reasoning process, thereby ensuring that the reasoning direction and reasoning goals of the large language model are consistent with expectations.

[0177] Specifically, by setting the guiding text element in the target reasoning prompt word, the large language model can understand the relationship between different text features and labels based on the provided second number of second annotated network traffic texts, labels of each second annotated network traffic text and other annotation data.

[0178] Specifically, by setting the inference text element in the target inference prompt word, the large language model can combine the relationship between different text features and labels and multiple features of the test network traffic text to accurately infer the first label of the test network traffic text.

[0179] Specifically, by setting the output format element in the target inference prompt word, the large language model can output the first label that conforms to the expected format, so as to facilitate unified subsequent processing based on the first label that conforms to the expected format.

[0180] In this embodiment, a complete inference prompt word template is constructed by defining a role setting element, an inference task content element, a guidance text element, an inference text element, and an output format element. By inputting a second number of second-annotated network traffic texts and the labels of each second-annotated network traffic text into the guidance text element, and inputting the test network traffic text into the inference text element, a target inference prompt word that enables the large language model to perform efficient inference can be constructed. Thus, by designing a target inference prompt word that includes contextual features such as similar text features and labels, this embodiment enables the large language model to simulate the role of a network traffic monitoring expert to infer the first label of the test network traffic text, ultimately generating an expected inference result. Obviously, this embodiment can improve the accuracy, intelligence, and robustness of network traffic monitoring through an efficient prompt word inference driving mechanism.

[0181] The following describes the application process of the above-mentioned unsupervised network traffic detection method based on large language model reasoning in conjunction with a detailed embodiment, as follows:

[0182] (1) Determination process of feature weight coefficients for each text feature

[0183] Acquire multiple annotated network traffic data; the annotated network traffic data includes multiple data features, and each annotated network traffic data corresponds to a label; determine the correlation coefficient between each data feature and the label; determine the correlation coefficient between each data feature and the label as a feature weight coefficient of each text feature.

[0184] (2) Test the process of obtaining network traffic text

[0185] Acquire test network traffic text; the test network traffic text is obtained by textually processing the network traffic data to be tested.

[0186] (3) Determination process of the first number of first annotated network traffic texts

[0187] A first number of first annotated network traffic texts whose first similarity with the test network traffic text meets a first preset condition are determined in a text retrieval tool; each first annotated network traffic text corresponds to a label.

[0188] (4) Determination process of the second similarity and the third similarity between the test network traffic text and each first annotated network traffic text

[0189] Performing vectorization processing on the test network traffic text and each first annotated network traffic text based on the large language model to obtain a test network traffic text semantic vector and a first number of annotated network traffic text semantic vectors, and determining the similarity between the test network traffic text semantic vector and the semantic vectors of each annotated network traffic text as a second similarity between the test network traffic text and each first annotated network traffic text;

[0190] The test network traffic text and each first annotated network traffic text are vectorized based on the label encoding method to obtain a test network traffic text feature vector and a first number of annotated network traffic text feature vectors, and the similarity between the test network traffic text feature vector and each annotated network traffic text feature vector is determined based on the feature weight coefficient of each text feature, and the similarity between the test network traffic text feature vector and each annotated network traffic text feature vector is determined as the third similarity between the test network traffic text and each first annotated network traffic text.

[0191] (5) Determination of the total similarity score between the test network traffic text and each first annotated network traffic text

[0192] Based on the first similarity, the second similarity and the third similarity between the test network traffic text and each first annotated network traffic text, a total similarity score between the test network traffic text and each first annotated network traffic text is determined.

[0193] (6) Determination process of the second number of second annotated network traffic texts

[0194] Based on the total similarity score between the test network traffic text and each first annotated network traffic text, a second number of second annotated network traffic texts whose corresponding total similarity scores meet a second preset condition are determined from the first number of first annotated network traffic texts.

[0195] (7) Test the reasoning process of the first label of network traffic text

[0196] Determine a role setting element, a reasoning task content element, and an output format element, input a second number of second annotated network flow texts and a label of each second annotated network flow text into a guidance text element, and input the test network flow text into a reasoning text element to obtain a target reasoning prompt word;

[0197] The target reasoning prompt word is input into the large language model so that the large language model can infer the first label of the test network traffic text based on the role setting elements, the reasoning task content elements, the second number of second annotated network traffic texts, the labels of each second annotated network traffic text and the test network traffic text.

[0198] (8) The process of determining the target label of the test network traffic text

[0199] Determine, from a second number of second annotated network traffic texts, a second label of a second annotated network traffic text having the highest first similarity to the test network traffic text;

[0200] Determine, from the second number of second annotated network traffic texts, a third label of the second annotated network traffic text corresponding to the annotated network traffic text semantic vector having the second highest similarity to the test network traffic text semantic vector;

[0201] Determine, among the second number of second annotated network traffic texts, a fourth label of the second annotated network traffic text corresponding to the annotated network traffic text feature vector having the highest third similarity to the test network traffic text feature vector;

[0202] Determining a first weighted result between the first weight coefficient and the first label, a second weighted result between the second weight coefficient and the second label, a third weighted result between the third weight coefficient and the third label, and a fourth weighted result between the fourth weight coefficient and the fourth label, respectively; wherein the first weight coefficient is greater than the second weight coefficient, the third weight coefficient, and the fourth weight coefficient, respectively;

[0203] Performing similar item merging processing on the first weighted result, the second weighted result, the third weighted result, and the fourth weighted result, and accumulating the weighted results of the same result label to obtain the total weight coefficient corresponding to each result label;

[0204] The total weight coefficients of the result labels are compared, and the result label with the largest total weight coefficient is determined as the target label of the test network traffic text.

[0205] In this embodiment, on the one hand, since the present application adopts an unsupervised learning strategy to detect network traffic, there is no need to fine-tune the model using this method. Therefore, the present embodiment does not rely on a large amount of labeled data, which solves the problem of high dependence on large-scale labeled data in the prior art; on the other hand, the present application designs a double reordering mechanism, and the first reordering is a reordering mechanism that combines the first similarity, the second similarity and the third similarity to obtain a total similarity score. By reordering the first labeled network traffic text with a high similarity to the test network traffic text, the second number of second labeled network traffic texts that are most relevant to the test network traffic text can be accurately further screened out, which can ensure the subsequent network The accuracy of the network traffic detection results is improved. The second reordering is a classification reordering mechanism based on the weighted summation strategy. The label weight coefficients of different result labels are calculated through weighted summation measurement, and the most preferred result label is selected as the target label of the test network traffic text according to the label weight coefficients of different result labels, which further improves the robustness of the network traffic detection results. Thirdly, the present application adopts a large language model for reasoning. However, unlike traditional methods, the present application does not need to fine-tune the large language model. Instead, it drives the reasoning process of the first label of the test network traffic text by designing efficient target reasoning prompt words, thereby reducing the complexity of the large language model reasoning process and improving the efficiency and accuracy of network traffic detection.

[0206] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0207] Based on the same inventive concept, an embodiment of the present application also provides an unsupervised network traffic detection device based on large language model reasoning for implementing the above-mentioned unsupervised network traffic detection method based on large language model reasoning. The implementation solution provided by this device is similar to the implementation solution described in the above-mentioned method. Therefore, the specific limitations of one or more embodiments of the unsupervised network traffic detection device based on large language model reasoning provided below can be found in the above-mentioned limitations of the unsupervised network traffic detection method based on large language model reasoning, and will not be repeated here.

[0208] In an exemplary embodiment, Figure 3 As shown, an unsupervised network traffic detection device based on large language model reasoning is provided, including: an acquisition module 302, a retrieval module 304, a vectorization module 306, a similarity calculation module 308, a reordering module 310 and a large language model reasoning module 312, wherein:

[0209] The acquisition module 302 is used to acquire the test network traffic text; the test network traffic text is obtained by textually processing the network traffic data to be tested.

[0210] The retrieval module 304 is used to determine a first number of first annotated network traffic texts whose first similarity with the test network traffic text meets a first preset condition in a text retrieval tool; each first annotated network traffic text corresponds to a label.

[0211] The vectorization module 306 is used to perform vectorization processing on the test network traffic text and each first annotated network traffic text to obtain a test network traffic text vector and a first number of annotated network traffic text vectors, and determine the similarity between the test network traffic text vector and each annotated network traffic text vector as the vector similarity between the test network traffic text and each first annotated network traffic text.

[0212] The similarity calculation module 308 is used to determine the total similarity score between the test network traffic text and each first annotated network traffic text based on the first similarity and vector similarity between the test network traffic text and each first annotated network traffic text.

[0213] The reordering module 310 is used to determine a second number of second annotated network traffic texts whose corresponding total similarity scores meet a second preset condition from the first number of first annotated network traffic texts based on the total similarity scores between the test network traffic text and each first annotated network traffic text.

[0214] The large language model inference module 312 is used to obtain a target inference prompt word based on a second number of second annotated network traffic texts, labels of each second annotated network traffic text, a test network traffic text, and an inference prompt word template, and input the target inference prompt word into the large language model to obtain a first label for the test network traffic text.

[0215] In an exemplary embodiment, the vectorization module 306 is also used to perform vectorization processing on the test network traffic text and each first labeled network traffic text based on a large language model to obtain a test network traffic text semantic vector and a first number of labeled network traffic text semantic vectors, and determine the similarity between the test network traffic text semantic vector and the semantic vectors of each labeled network traffic text as the second similarity between the test network traffic text and each first labeled network traffic text; perform vectorization processing on the test network traffic text and each first labeled network traffic text based on a label encoding method to obtain a test network traffic text feature vector and a first number of labeled network traffic text feature vectors, and determine the similarity between the test network traffic text feature vector and each labeled network traffic text feature vector based on the feature weight coefficient of each text feature, and determine the similarity between the test network traffic text feature vector and each labeled network traffic text feature vector as the third similarity between the test network traffic text and each first labeled network traffic text.

[0216] The similarity calculation module 308 is further configured to determine a total similarity score between the test network traffic text and each first annotated network traffic text based on the first similarity, the second similarity and the third similarity between the test network traffic text and each first annotated network traffic text.

[0217] In an exemplary embodiment, Figure 4 As shown, the above-mentioned device also includes a weight coefficient determination module 314, which is used to obtain multiple annotated network traffic data; the annotated network traffic data includes multiple data features, and each annotated network traffic data corresponds to a label; determine the correlation coefficient between each data feature and the label; and determine the correlation coefficient between each data feature and the label as the feature weight coefficient of each text feature.

[0218] In an exemplary embodiment, Figure 4 As shown, the above-mentioned device also includes a label determination module 316, which is used to determine the second label of the second labeled network traffic text with the highest first similarity to the test network traffic text among the second number of second labeled network traffic texts; determine the third label of the second labeled network traffic text corresponding to the labeled network traffic text semantic vector with the second highest similarity to the test network traffic text semantic vector among the second number of second labeled network traffic texts; determine the fourth label of the second labeled network traffic text corresponding to the labeled network traffic text feature vector with the third highest similarity to the test network traffic text feature vector among the second number of second labeled network traffic texts; and determine the target label of the test network traffic text based on the first label, the second label, the third label and the fourth label.

[0219] In an exemplary embodiment, the above-mentioned label determination module 316 is also used to respectively determine a first weighted result between the first weight coefficient and the first label, a second weighted result between the second weight coefficient and the second label, a third weighted result between the third weight coefficient and the third label, and a fourth weighted result between the fourth weight coefficient and the fourth label; based on the first weighted result, the second weighted result, the third weighted result and the fourth weighted result, determine the target label of the test network traffic text; wherein, the first weight coefficient is respectively greater than the second weight coefficient, the third weight coefficient and the fourth weight coefficient.

[0220] In an exemplary embodiment, the above-mentioned label determination module 316 is also used to merge similar items of the first weighted result, the second weighted result, the third weighted result and the fourth weighted result, and accumulate the weighted results of the same result label to obtain the total weight coefficient corresponding to each result label; compare the total weight coefficients of each result label, and determine the result label with the largest total weight coefficient as the target label for the test network traffic text.

[0221] In an exemplary embodiment, the above-mentioned large language model reasoning module 312 is also used to determine the role setting elements, the reasoning task content elements and the output format elements, input the second number of second annotated network traffic texts and the labels of each second annotated network traffic text into the guidance text element, and input the test network traffic text into the reasoning text element to obtain the target reasoning prompt word; input the target reasoning prompt word into the large language model, so that the large language model infers the first label of the test network traffic text based on the role setting elements, the reasoning task content elements, the second number of second annotated network traffic texts, the labels of each second annotated network traffic text and the test network traffic text.

[0222] Each module in the aforementioned unsupervised network traffic detection device based on large language model reasoning can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0223] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store network traffic data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an unsupervised network traffic detection method based on large language model reasoning is implemented.

[0224] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 6 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, an unsupervised network traffic detection method based on large language model reasoning is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0225] Those skilled in the art will understand that Figure 6The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0226] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0227] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0228] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0229] It should be noted that the data involved in this application (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0230] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile memory and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, an artificial intelligence (AI) processor, and the like.

[0231] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0232] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. An unsupervised network traffic detection method based on large language model reasoning, characterized in that: The method comprises: Acquire a test network traffic text; the test network traffic text is obtained by textually processing the network traffic data to be tested; Determining in a text search tool a first number of first annotated network traffic texts whose first similarity with the test network traffic text satisfies a first preset condition; each of the first annotated network traffic texts has a corresponding label; Performing vectorization processing on the test network traffic text and each of the first annotated network traffic texts to obtain a test network traffic text vector and a first number of annotated network traffic text vectors, and determining the similarity between the test network traffic text vector and each of the annotated network traffic text vectors as the vector similarity between the test network traffic text and each of the first annotated network traffic texts; Determining a total similarity score between the test network traffic text and each of the first annotated network traffic texts based on a first similarity and a vector similarity between the test network traffic text and each of the first annotated network traffic texts; Based on the total similarity score between the test network traffic text and each of the first annotated network traffic texts, determining a second number of second annotated network traffic texts whose corresponding total similarity scores meet a second preset condition from the first number of first annotated network traffic texts; A target inference prompt word is obtained based on a second number of the second annotated network traffic texts, a label of each of the second annotated network traffic texts, a test network traffic text, and an inference prompt word template, and the target inference prompt word is input into a large language model to obtain a first label for the test network traffic text.

2. The method according to claim 1, characterized in that The test network traffic text and each of the first annotated network traffic texts respectively correspond to a plurality of text features; The vectorizing process of the test network traffic text and each of the first annotated network traffic texts to obtain a test network traffic text vector and a first number of annotated network traffic text vectors, and determining the similarity between the test network traffic text vector and each of the annotated network traffic text vectors as the vector similarity between the test network traffic text and each of the first annotated network traffic texts, includes: Performing vectorization processing on the test network traffic text and each of the first annotated network traffic texts based on a large language model to obtain a test network traffic text semantic vector and a first number of annotated network traffic text semantic vectors, and determining the similarity between the test network traffic text semantic vector and each of the annotated network traffic text semantic vectors as a second similarity between the test network traffic text and each of the first annotated network traffic texts; The test network traffic text and each of the first annotated network traffic texts are vectorized based on a label encoding method to obtain a test network traffic text feature vector and a first number of annotated network traffic text feature vectors, and the similarity between the test network traffic text feature vector and each of the annotated network traffic text feature vectors is determined based on a feature weight coefficient of each of the text features, and the similarity between the test network traffic text feature vector and each of the annotated network traffic text feature vectors is determined as a third similarity between the test network traffic text and each of the first annotated network traffic texts; The determining, based on the first similarity and vector similarity between the test network traffic text and each of the first annotated network traffic texts, a total similarity score between the test network traffic text and each of the first annotated network traffic texts includes: Based on the first similarity, the second similarity and the third similarity between the test network traffic text and each of the first annotated network traffic texts, a total similarity score between the test network traffic text and each of the first annotated network traffic texts is determined.

3. The method according to claim 2, characterized in that The method further comprises: Acquire a plurality of annotated network flow data; the annotated network flow data includes a plurality of data features, and each annotated network flow data has a corresponding label; Determining a correlation coefficient between each of the data features and the label; The correlation coefficient between each of the data features and the label is determined as the feature weight coefficient of each text feature.

4. The method according to claim 2, characterized in that The method further comprises: Determine, from a second number of second annotated network traffic texts, a second label of a second annotated network traffic text having the highest first similarity to the test network traffic text; Determine, from a second number of second annotated network traffic texts, a third label of the second annotated network traffic text corresponding to the annotated network traffic text semantic vector having the second highest similarity with the test network traffic text semantic vector; Determine, from a second number of second annotated network traffic texts, a fourth label of the second annotated network traffic text corresponding to the annotated network traffic text feature vector having the highest third similarity to the test network traffic text feature vector; A target label of the test network traffic text is determined based on the first label, the second label, the third label, and the fourth label.

5. The method according to claim 4, characterized in that The determining a target label of the test network traffic text based on the first label, the second label, the third label, and the fourth label includes: respectively determining a first weighted result between a first weight coefficient and the first label, a second weighted result between a second weight coefficient and the second label, a third weighted result between a third weight coefficient and the third label, and a fourth weighted result between a fourth weight coefficient and the fourth label; Determining a target label for the test network traffic text based on the first weighted result, the second weighted result, the third weighted result, and the fourth weighted result; The first weight coefficient is respectively greater than the second weight coefficient, the third weight coefficient and the fourth weight coefficient.

6. The method according to claim 5, characterized in that The determining the target label of the test network traffic text based on the first weighted result, the second weighted result, the third weighted result, and the fourth weighted result includes: Performing similar item merging processing on the first weighted result, the second weighted result, the third weighted result, and the fourth weighted result, and accumulating the weighted results of the same result label to obtain a total weight coefficient corresponding to each result label; The total weight coefficients of the result labels are compared, and the result label with the largest total weight coefficient is determined as the target label of the test network traffic text.

7. The method according to any one of claims 1 to 6, characterized in that The inference prompt word template includes a role setting element, an inference task content element, a guide text element for inputting a second number of the second annotated network flow texts and a label of each of the second annotated network flow texts, an inference text element for inputting the test network flow text, and an output format element; The step of obtaining a target inference prompt word based on the second number of the second annotated network traffic texts, the label of each of the second annotated network traffic texts, the test network traffic text, and the inference prompt word template, and inputting the target inference prompt word into the large language model to obtain a first label for the test network traffic text includes: Determining a role setting element, a reasoning task content element, and an output format element, inputting a second number of the second annotated network flow texts and a label of each of the second annotated network flow texts into the guidance text element, and inputting a test network flow text into the reasoning text element to obtain the target reasoning prompt word; The target inference prompt word is input into the large language model so that the large language model infers the first label of the test network traffic text based on the role setting elements, the inference task content elements, the second number of the second annotated network traffic texts, the labels of each of the second annotated network traffic texts, and the test network traffic text.

8. An unsupervised network traffic detection device based on large language model reasoning, characterized in that: The device comprises: An acquisition module is used to acquire a test network traffic text; the test network traffic text is obtained by textually processing the network traffic data to be tested; A retrieval module is configured to determine, in a text retrieval tool, a first number of first annotated network flow texts whose first similarity with the test network flow text satisfies a first preset condition; each of the first annotated network flow texts is respectively associated with a label; a vectorization module, configured to perform vectorization processing on the test network traffic text and each of the first annotated network traffic texts to obtain a test network traffic text vector and a first number of annotated network traffic text vectors, and determine the similarity between the test network traffic text vector and each of the annotated network traffic text vectors as the vector similarity between the test network traffic text and each of the first annotated network traffic texts; a similarity calculation module, configured to determine a total similarity score between the test network traffic text and each of the first annotated network traffic texts based on a first similarity and a vector similarity between the test network traffic text and each of the first annotated network traffic texts; a reordering module, configured to determine, from the first number of first annotated network traffic texts, a second number of second annotated network traffic texts whose corresponding total similarity scores satisfy a second preset condition, based on the total similarity scores between the test network traffic text and each of the first annotated network traffic texts; The large language model inference module is used to obtain a target inference prompt word based on a second number of the second annotated network traffic texts, a label of each of the second annotated network traffic texts, a test network traffic text, and an inference prompt word template, and input the target inference prompt word into the large language model to obtain a first label for the test network traffic text.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.