Malicious URL detection method based on character-level language model and structural feature fusion

Through the method of integrating character-level language model and structural features, the problem of performance degradation of malicious URL detection methods in the prior art when dealing with unlogined words and new URLs is solved, and more efficient malicious URL detection is achieved.

CN120277500AActive Publication Date: 2025-07-08CHANGSHU INSTITUTE OF TECHNOLOGY

Patent Information

Application Number
CN202510764171.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing malicious URL detection methods have deteriorated performance when processing unlogined words (OOV) or new malicious URLs, and fail to effectively integrate the structural characteristics and semantic information of the URL, resulting in poor detection results.

Method used

The method of fusion of character-level language model and structural features is adopted, and the substitute word embedding module is used for character-level convolutional network (CNN), combined with the multi-scale attention mechanism, the semantic features of the URL are extracted and the structural features of the URL are fused, and the gating mechanism is used for dynamic weighted fusion.

Benefits of technology

It improves the detection ability of malicious URLs, enhances the adaptability and accuracy of the model, and effectively alleviates the misjudgment problems caused by incomplete semantic features or lack of structural information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277500A_ABST
    Figure CN120277500A_ABST
Patent Text Reader

Abstract

The invention discloses a malicious URL detection method based on character-level language model and structural feature fusion. The method comprises the steps that character-level URL semantic features are obtained through a character-level language model; the URL semantic features are sent to expansion pyramid attention, semantic enhancement is carried out, and enhanced character-level URL semantic features are obtained; extracting structural features of the URL character string to obtain URL structural features; and carrying out dynamic weighted fusion on the URL structural features and the enhanced character-level URL semantic features, and outputting a malicious URL judgment result through a classifier. Through character-level semantic understanding, a predefined word list library is separated, the adaptive capacity of random characters appearing in the URL is improved, meanwhile, the structural features of the URL are fused, and the recognition performance of the model for the malicious URL is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of malicious URL detection, and the present invention relates to a malicious URL detection method based on the fusion of character-level language model and structural features. Background Art

[0002] Malicious URL detection is an important part of network security, aiming to identify and intercept potential network threats, especially playing a key role in protecting user privacy. Malicious URLs are often used to implement phishing, information theft and other attacks, so their detection technology is of great significance for building a secure and trustworthy network environment.

[0003] At present, there are three main methods for detecting malicious websites: blacklist and whitelist databases, machine learning (ML) algorithms, and deep learning (DL) technology. Early methods mainly relied on blacklist mechanisms and rule-based feature engineering. Although they were simple to implement and had high detection efficiency, they were often unable to cope with the rapidly evolving characteristics of new or variant malicious URLs, and the detection effect was limited.

[0004] Detection methods based on machine learning have improved the level of intelligence to a certain extent, but most of them still rely on manual extraction of URL features, which is not only time-consuming and labor-intensive, but also performs poorly when faced with malicious URLs with complex structures and frequent changes.

[0005] In recent years, with the rapid development of pre-trained language models (such as BERT, GPT, etc.), they have demonstrated strong capabilities in semantic understanding and feature expression, opening up a new direction for malicious URL detection. Researchers have begun to try to apply such models to URL detection tasks to mine the deep semantic information contained in URLs, and have achieved initial results. However, existing methods generally focus on semantic modeling, ignoring the structural characteristics of the URL itself, and fail to achieve effective fusion of structural information and semantic information.

[0006] In addition, the pre-trained language model is highly dependent on predefined vocabulary, and URLs often contain a large number of random characters. Traditional word segmentation methods are difficult to capture their highly fine-grained semantics, resulting in a decrease in the model's performance when dealing with out-of-date (OOV) words or new malicious URLs. This limitation significantly restricts the model's detection effect in practical applications. Summary of the invention

[0007] The purpose of the present invention is to provide a malicious URL detection method based on the fusion of character-level language model and structural features. Through character-level semantic understanding, it is separated from the predefined vocabulary library, improves the adaptability to random characters appearing in URLs, and at the same time integrates the structural features of URLs to further improve the model's recognition performance for malicious URLs.

[0008] The technical solution for achieving the object of the present invention is as follows: A malicious URL detection method based on the fusion of character-level language models and structural features, comprising the following steps: S01: Obtain character-level URL semantic features using a character-level language model; S02: Feed the URL semantic features into dilated pyramid attention for semantic enhancement to obtain enhanced character-level URL semantic features; S03: Extract the structural features of the URL string to obtain URL structural features; S04: Dynamically weight and fuse the URL structural features with the enhanced character-level URL semantic features, and output the malicious URL judgment result through a classifier.

[0009] In a preferred technical solution, step S01 further includes segmenting each URL string according to special symbols and alphanumeric characters, adding special markers at the beginning and end respectively, converting each segmented word and special marker into a fixed-length character-level index representation, padding the insufficient part, and the overall character-level index of the URL consists of all segmented indexes, which is used as the model input.

[0010] In a preferred technical solution, step S01 further includes: Using Character CNN to replace the word embedding layer to generate a unified representation at the character level. Character CNN extracts character-level features through a character CNN and a Highway network to generate character-level embeddings; Character CNN includes a character embedding layer, a convolutional feature extractor, a Highway network, and a projection layer; The character embedding layer is used to map characters into fixed-length vectors; The convolutional feature extractor includes multiple one-dimensional convolutional kernels for extracting sequences composed of n consecutive characters at the character level; multiple convolutional kernels of different sizes are applied in parallel to the character sequence to obtain different fine-grained information in short texts. The output of each CNN is max-pooled in the character sequence and connected to the outputs of other CNNs to generate a single representation; The Highway network is used to screen the information of the CNN representation; The projection layer is used for dimension conversion.

[0011] In a preferred technical solution, it further includes: Combining character-level word embeddings with position embeddings and segment embeddings, and generating a final text representation that combines BERT and character-level embeddings after normalization and regularization; After being processed by a multi-layer Transformer encoder and a pooling layer, the embedding finally outputs 12 hidden layers. The hidden states in these layers include both the representation of each token and the sentence-level pooling result.

[0012] In the preferred technical solution, step S02 of dilated pyramid attention includes depthwise separable convolution and spatial pyramid attention. For the input features Apply a single depthwise separable convolution to extract the common information of each branch : ;

[0013] Among them, denotes the depthwise separable convolution operation. Perform dilated depthwise separable convolution operations with different dilation rates on different branches to obtain the output features of the th branch : ;

[0014] Among them, denotes the dilated depthwise separable convolution operation at branch . is the number of branches. Pass the context information from different scales through the use of residual connections, that is: ;

[0015] Adopt the element-wise summation strategy for optimization and refine the aggregated features through 1x1 convolution operations: ;

[0016] denotes the standard convolution operation for fusing context information at different scales. The original feature map is integrated through the residual connection mechanism and different dilation rates are used to capture multi-scale context information; The spatial pyramid attention consists of pointwise convolution, a spatial pyramid structure, and a multi-layer perceptron. Pointwise convolution is used to integrate information and align channels; the spatial pyramid structure fuses adaptive average pooling of multiple scales to enhance feature integration and regularization; the multi-layer perceptron generates an attention map to provide key spatial information; The attention mechanism extracts attention weights from the input data and multiplies them with each channel through learnable weights to generate the output. The output expression of the spatial pyramid structure is as follows: ;

[0017] is the concatenation operation, For tensor reshaping, For adaptive average pooling, is the feature map output by depthwise separable convolution; Omitting the batch normalization and activation layers, the basic transformation is represented as follows: ;

[0018] is the fully connected layer, is the Sigmoid activation function; At the end of the network, average pooling is performed on the weighted feature map along the fixed sequence length dimension to extract the representative feature vector.

[0019] In the preferred technical solution, the method for extracting the structural features of the URL string in step S03 includes: Using the tld library to parse the URL, extracting its structural components, and counting the structural feature information. The structural features include the total length of the URL, the number of special characters, whether it contains sensitive words, the number of top-level domains, and the number of subdomains. These features are combined into a URL structure feature list.

[0020] In the preferred technical solution, the method for dynamic weighted fusion in step S04 includes: S41: Project the two feature vectors into the same dimensional space respectively: ;

[0021]

[0022] Among them, and are the projection matrices, and are the bias terms, is the semantic feature vector, is the URL structure feature vector.

[0023] S42: Calculate the gating weight : ;

[0024] Among them, is the Sigmoid activation function; S43: Use the gating weight to perform weighted fusion on the two feature vectors: ;

[0025] Among them, represents element-wise multiplication; S44: Generate the fused feature vector through a fully connected layer: ;

[0026] Among them, and are the weights and biases of the fully connected layer.

[0027] In a preferred technical solution, after obtaining the URL structure feature in step S03, it further includes performing dimensionality increase processing on the obtained URL structure feature through a linear layer to make its dimension consistent with the dimension of the semantic feature.

[0028] The present invention also discloses a malicious URL detection system based on the fusion of character-level language model and structural features, including: A semantic feature extraction module that uses a character-level language model to obtain character-level URL semantic features; A high-fine-grained semantic module that sends the URL semantic features into a dilated pyramid attention for semantic enhancement to obtain enhanced character-level URL semantic features; A structural feature extraction module that extracts the structural features of the URL string to obtain URL structure features; A fusion judgment module that dynamically weights and fuses the URL structure features and the enhanced character-level URL semantic features, and outputs a malicious URL judgment result through a classifier.

[0029] The present invention also discloses a computer storage medium, on which a computer program is stored, and when the computer program is executed, it implements the above-mentioned malicious URL detection method based on the fusion of character-level language model and structural features.

[0030] Compared with the prior art, the present invention has the following significant advantages: 1. It does not need to rely on a predefined vocabulary, and enhances the character-level semantic perception ability Compared with traditional models such as BERT that rely on WordPiece, the present invention uses a character-level convolutional network (CharacterCNN) to replace the word embedding module, completely breaking away from the vocabulary limit, and can directly process the original URL string containing random characters, and introduces a multi-scale attention mechanism to enhance semantic modeling, greatly improving the detection ability of malicious URLs.

[0031] 2. It fully explores the internal semantics and structural patterns of URLs and improves the detection accuracy The present invention fuses the structural feature information of URLs (such as length, number of subdomains, whether it contains sensitive words, etc.) with character-level semantic vectors, and dynamically learns the fusion strategy through an innovative gating mechanism to achieve joint representation modeling of "structure + semantics", effectively alleviating the misjudgment problem caused by incomplete semantic features or missing structural information. Description of the Drawings

[0032] Figure 1 Flowchart of the malicious URL detection method based on the fusion of character-level language model and structural features in this embodiment; Figure 2 Overall structure diagram of the model of the malicious URL detection method based on the fusion of character-level language model and structural features in this embodiment; Figure 3 Model flowchart of the malicious URL detection method based on the fusion of character-level language model and structural features in this embodiment; Figure 4 Differences between CharacterBERT and BERT in this embodiment; Figure 5 Flowchart of CharacterBERT in this embodiment. Detailed implementation mode

[0033] Principle of the invention: Construct a multi-feature malicious URL detection framework with a character-level language model as the semantic backbone, structural URL features as the supplementary information source, and a fusion mechanism as the connection bridge; The said architecture is applicable to any character-level encoder, context modeling module, and multi-scale feature enhancement module that can output embedding vectors.

[0034] Embodiment 1:

[0035] As Figure 1 shown, a malicious URL detection method based on the fusion of character-level language model and structural features includes the following steps: S01: Obtain character-level URL semantic features using a character-level language model; S02: Send the URL semantic features into dilated pyramid attention for semantic enhancement to obtain enhanced character-level URL semantic features; S03: Extract the structural features of the URL string to obtain URL structural features; S04: Dynamically weight and fuse the URL structural features with the enhanced character-level URL semantic features, and output the malicious URL judgment result through a classifier.

[0036] Another embodiment, a malicious URL detection system based on the fusion of character-level language model and structural features, includes: A semantic feature extraction module that obtains character-level URL semantic features using a character-level language model; A high-fine-grained semantic module that sends the URL semantic features into dilated pyramid attention for semantic enhancement to obtain enhanced character-level URL semantic features; A structural feature extraction module that extracts the structural features of the URL string to obtain URL structural features; The fusion judgment module dynamically weights and fuses the URL structure features and the enhanced character-level URL semantic features, and outputs the malicious URL judgment result through a classifier.

[0037] Specifically, the character-level language model is illustrated by taking CharacterBERT as an example. The overall model composition is as Figure 2 shown, including: 1) Data preprocessing module It is used to perform preliminary word segmentation and structure information extraction on the original URL input, including: Symbol splitting and token splitting (such as splitting www.baidu.com into ['www', '.', 'baidu', '.', 'com']); Extract structural indicators such as URL length, number of subdomains, and whether it contains sensitive words to form a structure feature vector.

[0038] 2) Character-level encoding module (CharacterCNN) Perform character ID mapping on each token, and extract character n-gram patterns through multiple groups of convolutional kernels to obtain token-level character embedding representations.

[0039] 3) CharacterBERT module The difference from the traditional BERT is that CharacterBERT uses CharacterCNN to replace the word embedding part of the original BERT. As Figure 4 shown.

[0040] Combine the character-level word embeddings generated by CharacterCNN with positional embeddings and segment embeddings, and finally use them as the input of the BERT encoder. Obtain character-level context representations through multi-layer Transformer encoding to form a semantic vector sequence of the entire URL.

[0041] 4) Dilated pyramid attention (as shown in the Figure 2 attention part) Send the output of the above 3) into the DPAM module containing dilated convolution and spatial pyramid pooling to expand the receptive field without changing the resolution, and realize multi-scale context information extraction and feature enhancement.

[0042] 5) Structure feature mapping module Dimensionality increase of the structure feature vector obtained by data preprocessing through linear mapping to make its dimension consistent with the semantic vector, preparing for the subsequent fusion module.

[0043] 6) Gated fusion unit Utilize the gating mechanism to dynamically learn the fusion strategy between semantic information and structural information, and obtain the joint representation.

[0044] 7) Classification and discrimination module Feed the fused joint representation into Dropout and the fully connected layer, and output the prediction result of whether the URL is malicious or not.

[0045] The principle working process of the model is as Figure 3 shown: URL input → Through the CharacterBERT module and the structural feature extraction model, obtain semantic information and structural features; The primary semantics obtained through the CharacterBERT module enters the dilated pyramid attention to obtain semantic information with a higher level of granularity.

[0046] The structural features are dimensionally increased through the dimension processing module to make their dimensions consistent with those of CharacterBERT.

[0047] Input the outputs obtained in step 2 and step 3 into the gating fusion unit to obtain the output vector that fuses the URL structural features and semantic information; The fused vector passes through the loss layer and the fully connected layer to obtain the prediction result - malicious / benign.

[0048] The following describes the specific implementation steps: Step S1: Obtain the dataset of URL samples, where the dataset includes the URL text and the corresponding labels for the URL - benign and malicious; Step S2: Use CharacterCNN to convert the URL string into a URL character vector, and feed the URL character vector into CharacterBERT for context semantic modeling to obtain the character-level semantic features of the URL; Step S3: Feed the URL semantic features obtained in step S2 into the dilated pyramid attention for semantic enhancement to obtain the enhanced URL character-level semantic feature F c ; Step S4: Use the tld library to parse the URL string, and perform structural feature extraction on the parsed URL string to construct the URL structural features; Step S5: Dimensionally increase the URL structural features obtained in step S4 through a linear layer to align them with the dimensions of CharacterBERT, obtaining the dimensionally increased URL structural feature F s ; Step S6: The F s obtained in step S5 and the F cSend it into the door control fusion unit, so that the two features can be dynamically weighted and fused to obtain the fused URL feature F url ; Step S7: Process the F obtained in step S6 url through a standard Dropout layer and a fully connected layer, convert the URL feature into a binary class representation for prediction, that is, obtain the final URL detection model.

[0049] The specific method for step S2 to convert the URL string into a URL character vector and send it into CharacterBERT for semantic modeling is as follows: (a) Data preprocessing: Each URL string is tokenized according to special symbols and alphanumeric characters. For example, "www.baidu.com" becomes: [ ’ www’, ’ . ’ , ‘baidu’ , ‘ . ’ , ‘com’ ] after tokenization. The maximum length of this tokenization is n, and if it exceeds, it is truncated; (b) Character-level word embedding: CharacterBERT is similar to traditional BERT, but the method is different when constructing the initial representation. Traditional BERT relies on a built-in vocabulary to disassemble unknown tokens into word pieces for independent embedding, while CharacterBERT uses CharacterCNN to replace the word embedding layer and generates a unified representation at the character level. This method improves the flexibility of text processing and eliminates the dependence on the vocabulary. As Figure 4 shown.

[0050] CharacterCNN consists of a character embedding layer, a convolutional feature extractor, a Highway network, and a projection layer.

[0051] Character embedding layer: In CharacterCNN, a character is mapped to a vector with a fixed length of 50. The dimension of the embedding matrix is usually 262×16, (262 represents the supported character types, including the utf-8 character set and 6 custom special tokens, and 16 is the embedding dimension). The character embedding layer aims to convert discrete character indices into trainable continuous vector representations. These continuous vectors are then used as the input for the subsequent convolutional layer to extract character-level features.

[0052] Convolutional feature extractor: The core of this module is multiple one-dimensional convolutional kernels (1D-CNN) used to extract character-level n-gram (referring to a sequence composed of n consecutive characters) patterns, enhancing the ability to capture local features, thereby capturing higher fine-grained semantic information. Multiple convolutional kernels of different sizes are applied in parallel to the character sequence to obtain different fine-grained information in short texts. Then, the output of each CNN is max-pooled in the character sequence and connected with the outputs of other CNNs to generate a single representation.

[0053] Highway Network: To enhance the non-linear representation ability of the model and enable it to selectively retain or discard features, CharacterCNN uses the Highway network to screen the representations of CNNs.

[0054] Projection Layer: After convolution and the Highway network, the output feature vector has a large dimension and needs to be reduced in dimension through the projection layer to ensure compatibility with models such as BERT. The projection layer uses a linear transformation to map the output features to the 768-dimensional BERT input format.

[0055] (c) CharacterBERT Semantic Modeling: Combine character-level word embeddings with position embeddings and segment embeddings. After normalization and the Dropout layer (regularization), generate the final BertCharacterEmbeddings. Subsequently, after processing by multiple layers of Transformer encoders and a pooling layer, finally output 12 hidden layers. The hidden states in these layers include both the representations of each token and the sentence-level pooling results - the character-level semantic information of the URL string. As Figure 5 shown.

[0056] The detailed method of the dilated pyramid attention in step S3 is as follows: The dilated pyramid attention consists of depthwise separable convolution (DSConv) and spatial pyramid attention, and its specific implementation is as follows: Depthwise Separable Convolution: Set the high-dimensional output feature representation of CharacteBERT as the input as , where C, H, and W represent the number of channels, height, and width of the feature map respectively. We first apply a single depthwise separable conv3 x 3 (abbreviated as DSConv3x3) to the input to extract the common information of each branch , and the specific formula is as follows: ;

[0057] Among them, represents the depthwise separable conv3 x 3 operation; on different branches, the dilated DSConv3 x 3 with different dilation rates is applied to to obtain the output feature of the th branch, that is: ;

[0058] Among them, represents the dilated DSConv3x3 operation at branch ​ is the number of branches, and context information from different scales is connected using residual connections, namely: ;

[0059] Since the concatenation operation significantly increases the number of channels, aggravates the computational complexity and parameter expansion, we turn to the element summation strategy optimization. Finally, the 1x1 convolution operation is used to refine the aggregated features, aiming to achieve compact and efficient feature representation, namely: ;

[0060] Represents a standard conv1x1 operation that fuses context information of different scales. Original feature map Integration is performed through a residual connection mechanism, which aims to facilitate gradient flow and improve training efficiency. We adopt the expansion rate configuration of [1, 2, 4, 8] to capture multi-scale contextual information, thereby enhancing the richness and diversity of URL feature representation.

[0061] Spatial Pyramid Attention The spatial pyramid attention mechanism consists of point-by-point convolution, spatial pyramid structure and multi-layer perceptron. Point-by-point convolution integrates information and aligns channels; the spatial pyramid structure integrates adaptive average pooling of three scales to enhance feature integration and regularization; the multi-layer perceptron generates an attention map to provide key spatial information.

[0062] For precise description, we mark adaptive average pooling, fully connected layer, connection operation, Sigmoid activation function and tensor reshaping as , , , , The feature map output by the multi-scale learning module can be expressed as In this framework, the attention mechanism extracts attention weights from the input data and multiplies them with the learnable weights of each channel to generate the output. Spatial pyramid structure The output expression is as follows: ;

[0063] For clarity, batch normalization and activation layers are omitted. The basic transformation It can be expressed as follows: ;

[0064] The number of channels is 12, which is used to align the 12 hidden layers output by the CharacterBERT model. At the end of the network, the weighted feature map is mean-pooled along the fixed sequence length dimension to extract representative feature vectors. The result is passed to the gating unit and dynamically weighted and fused with the URL structure features.

[0065] In step S4, the tld library is used to parse the URL structure and extract the structural features. The detailed method is as follows: Use the tld library to parse the URL, separate each structure of the URL, and count the required structural feature information, such as: the total length of the URL, whether the URL contains sensitive words, the number of top-level domains in the URL, the number of subdomains of the URL, etc. The judgment type uses 0 (no) and 1 (yes), and the statistical type uses numerical values. The URL structure feature list is composed of the above information.

[0066] The specific method of step S6 is as follows: Project the two feature vectors into the same dimensional space: ;

[0067] ;

[0068] Among them, and are the projection matrices, and are the bias terms. represents the semantic feature vector from BERT, represents the feature vector from the URL structure features.

[0069] Calculate the gating weights : ;

[0070] Among them, is the Sigmoid activation function to ensure that the gating weights are in the range of [0,1].[[]]

[0071] Use the gating weights to perform weighted fusion on the two feature vectors: ;

[0072] Among them, represents element-wise multiplication.

[0073] Finally, generate the fused feature vector through a fully connected layer: ;

[0074] Among them, and are the weights and biases of the fully connected layer.

[0075] Finally, input the output of the gating fusion unit into the dropout layer and the fully connected layer to obtain the prediction result.

[0076] In another embodiment, a computer storage medium stores a computer program, and when the computer program is executed, the malicious URL detection method based on the fusion of character-level language models and structural features described in any one of the above is performed.

[0077] To verify the effectiveness of the model, this paper selects two URL classification datasets, Grambeddings and kaggle_1. Grambeddings contains a total of 800,000 real phishing samples collected from PhishTank and OpenPhish, with 400,000 malicious and 400,000 benign samples each; kaggle_1 is from kaggle and contains approximately 316,000 malicious and benign URL samples each. To reduce the consumption of computing resources and improve the experimental efficiency, we randomly selected half of the samples from each category of each dataset for the experiment. The specific information is shown in Table 1.

[0078] Table 1 URL type and quantity statistics of three datasets

[0079] The experimental evaluation metrics for the model are Precision (P), Recall (R), F1 (comprehensive evaluation metric), and Accuracy, as shown in formulas (12) - (15).

[0080] (12) (13) (14) (15) In the formulas: TP represents the number of URLs that are actually malicious among the URLs predicted by the model as malicious; FP represents the number of URLs that are actually benign among the URLs predicted by the model as malicious; TN represents the number of URLs that are actually benign among the URLs predicted by the model as benign; FN represents the number of URLs that are actually malicious among the URLs predicted by the model as benign.

[0081] Comparative experiment analysis The experiment used python3.8 and the pytorch-2.0 deep learning framework. The specific hardware configuration was GPU A800 with 48GB of video memory. Using the AdamW optimizer, the number of training epochs was 3, and the extracted dataset was divided into a training set and a test set in a ratio of 8:2.

[0082] The experiment selected PMANet and TransURL as comparative models, and the performance comparison is shown in Table 2.

[0083] Table 2 Results of comparative experiments

[0084] The results show that the performance of the model in this paper is better than the comparison methods on both datasets. Especially, it has obvious advantages in terms of recall rate and F1 metric, can effectively reduce misjudgment and missed judgment, and has high application value in actual security scenarios.

[0085] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A malicious URL detection method based on the fusion of character-level language models and structural features, characterized in that, It includes the following steps: S01: Obtain character-level URL semantic features using a character-level language model; S02: Feed the URL semantic features into dilated pyramid attention for semantic enhancement to obtain enhanced character-level URL semantic features; S03: Extract the structural features of the URL string to obtain URL structural features; S04: Dynamically weight and fuse the URL structural features with the enhanced character-level URL semantic features, and output the malicious URL judgment result through a classifier.

2. The malicious URL detection method based on the fusion of character-level language model and structural features according to claim 1, wherein Step S01 further includes segmenting each URL string by special symbols and alphanumeric characters, adding special markers at the beginning and end respectively, converting each segment and special marker into a fixed-length character-level index representation, padding the insufficient part, and the overall character-level index of the URL consists of all segment indexes, serving as the model input.

3. The malicious URL detection method based on the fusion of character-level language model and structural features according to claim 1, wherein Step S01 further includes: Using Character CNN to replace the word embedding layer to generate a unified representation at the character level. Character CNN extracts character-level features through a character CNN and a Highway network to generate character-level embeddings; Character CNN includes a character embedding layer, a convolutional feature extractor, a Highway network, and a projection layer; The character embedding layer is used to map characters into fixed-length vectors; The convolutional feature extractor includes multiple one-dimensional convolutional kernels for extracting sequences composed of n consecutive characters at the character level; multiple convolutional kernels of different sizes are applied in parallel to the character sequence to obtain different fine-grained information in the short text. The output of each CNN is max-pooled in the character sequence and connected with the outputs of other CNNs to generate a single representation; The Highway network is used to screen the information of the CNN representation; The projection layer is used for dimension transformation.

4. The malicious URL detection method based on the fusion of character-level language model and structural features according to claim 3, wherein It also includes: Combining character-level word embeddings with position embeddings and segment embeddings, and generating the final text representation that combines BERT and character-level embeddings after normalization and regularization; After being processed by a multi-layer Transformer encoder and a pooling layer, the embedding finally outputs 12 hidden layers. The hidden states in these layers include both the representation of each token and the pooling result at the sentence level.

5. The malicious URL detection method based on the fusion of character-level language model and structural features according to claim 1, characterized in that The dilated pyramid attention in step S02 includes depthwise separable convolution and spatial pyramid attention; For the input features Apply a single depthwise separable convolution to extract the common information of each branch : , Among them, represents a depthwise separable convolution operation. A dilated depthwise separable convolution operation with different dilation rates is performed on different branches to obtain the output feature of the th branch : , Among them, represents a depthwise separable convolution operation for expansion at the branch where, is the number of branches. The context information from different scales is passed through the use of residual connections, that is: , Optimized by the element-wise summation strategy, and refining the aggregated features through 1x1 convolution operations: , Denotes a standard convolution operation that fuses context information at different scales, the original feature map Is integrated through a residual connection mechanism, and different dilation rates are adopted to capture multi-scale context information; The spatial pyramid attention consists of pointwise convolution, a spatial pyramid structure, and a multi-layer perceptron. The pointwise convolution is used to integrate information and align channels; the spatial pyramid structure fuses adaptive average pooling of multiple scales to enhance feature integration and regularization; the multi-layer perceptron generates an attention map to provide key spatial information; The attention mechanism extracts attention weights from the input data, multiplies them with each channel through learnable weights, and generates the output. The output of the spatial pyramid structure is expressed as follows: , is a connection operation, is a tensor reshaping, is an adaptive average pooling, is a feature map output by depthwise separable convolution; Omit the batch normalization and activation layers, basic transformation It is expressed as follows: , is a fully connected layer, is a Sigmoid activation function; At the end of the network, mean pooling is performed on the weighted feature map along the fixed sequence length dimension to extract representative feature vectors.

6. The malicious URL detection method based on the fusion of character-level language model and structural features according to claim 1, wherein The method for extracting the structural features of the URL string in step S03 includes: Parse the URL using the tld library, extract its structural components, and count the structural feature information. The structural features include the total length of the URL, the number of special characters, whether it contains sensitive words, the number of top-level domains, and the number of subdomains. Combine these features into a URL structural feature list.

7. The malicious URL detection method based on the fusion of character-level language model and structural features according to claim 1, wherein The method of dynamic weighted fusion in step S04 includes: S41: Project the two feature vectors into the same dimensional space respectively: , , Among them, and are projection matrices, and are bias terms, is the semantic feature vector, is the URL structure feature vector; S42: Calculate the gating weight : , Among them, is the Sigmoid activation function; S43: Use gated weights to perform weighted fusion on the two feature vectors: , Among them, represents element-wise multiplication; S44: Generate a fused feature vector through a fully connected layer: , Among them, and are the weights and biases of the fully connected layer.

8. The malicious URL detection method based on the fusion of character-level language model and structural features according to claim 1, characterized in that, After obtaining the URL structural features in step S03, it further includes performing dimensionality increase processing on the obtained URL structural features through a linear layer to make its dimension consistent with the dimension of the semantic features.

9. A malicious URL detection system based on the fusion of character-level language models and structural features, characterized in that, It includes: A semantic feature extraction module that uses a character-level language model to obtain character-level URL semantic features; A high-fine-grained semantic module that sends the URL semantic features into dilated pyramid attention for semantic enhancement to obtain enhanced character-level URL semantic features; A structural feature extraction module that extracts the structural features of the URL string to obtain URL structural features; A fusion judgment module that dynamically weights and fuses the URL structural features and the enhanced character-level URL semantic features, and outputs a malicious URL judgment result through a classifier.

10. A computer storage medium, on which a computer program is stored, characterized in that, When the computer program is executed, it implements the malicious URL detection method based on the fusion of the character-level language model and the structural features according to any one of claims 1-8.

Citation Information

Patent Citations

  • Novel coupling sharing hyperspectral and LiDAR data collaborative classification method

    CN115841599A

  • Malicious domain name detection method based on large language model

    CN118413402A

Cited By

  • Multi-task URL (Uniform Resource Locator) detection method and system fusing structure and semantic features

    CN120856460A

  • Multi-task url detection method and system fusing structural and semantic features

    CN120856460B