Self-driven multi-element malicious URL (Uniform Resource Locator) detection method and system fusing width-depth model

By integrating wide-depth model and self-driven multi-learning strategy and combining time-sequent convolutional visual network, the problems of insufficient feature fusion and sample diversity in malicious URL detection are solved, and higher detection accuracy and robustness are achieved.

CN120372104APending Publication Date: 2025-07-25HUNAN UNIV OF SCI & ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510334069.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing malicious URL detection methods are difficult to effectively utilize the structural characteristics of URLs, and their feature fusion effects are poor, and they cannot cope with sample diversity and dynamic changes, resulting in insufficient detection accuracy and robustness.

Method used

The self-driven multi-learning strategy of fusion wide-deep models is adopted, and the deep structural features are extracted through the wide model and the deep model are extracted. The self-driven multi-learning strategy is used to dynamically adjust the sample weight, and the URL structural features are captured in combination with the time-sequential convolutional visual network, and the joint loss function and regularized objective function are constructed to optimize feature fusion and sample selection.

Benefits of technology

It improves the accuracy and robustness of malicious URL detection, can better deal with diverse and dynamically changing malicious URLs, avoid overfitting, and enhance model generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372104A_ABST
    Figure CN120372104A_ABST
Patent Text Reader

Abstract

The invention discloses a self-driven multivariate malicious URL detection method and system fused with a wide-deep model, the wide-deep model is fused through a self-driven multivariate learning method, a factorization machine is utilized to learn potential interaction of statistics and vocabulary features in the wide model, a time sequence convolution visual network with position embedding is utilized to capture structural features in the deep model, and a malicious URL detection result is obtained. And meanwhile, sample weights are dynamically adjusted through diversified self-paced learning with automatic adjustment, samples with different difficulties are adaptively selected to participate in training, and a more comprehensive basis is provided for malicious URL detection. On the basis of a traditional self-paced learning framework, self-driven multivariate learning has the advantages that on one hand, sample learning difficulty is dynamically adjusted through a complexity perception function, and the model is prevented from falling into local optimum in the initial training stage; on the other hand, diversity evaluation items are constructed, a sample diversity compensation mechanism is established, balanced learning of different feature partition samples is ensured, and the robustness of resisting sample confusion is remarkably improved while the generalization ability of the model is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular, to a self-driven multi-source malicious URL detection method integrating wide and deep models. Background Art

[0002] Malicious URLs can lead users to access malicious websites, download malicious software, or disclose users' sensitive information, posing a huge security threat to individuals, enterprises, and society. Traditional malicious URL detection methods, such as blacklist-based methods, need to maintain a huge URL blacklist. However, in the face of the rapid growth of malicious URLs and constantly changing obfuscation strategies, it is difficult to ensure timeliness and accuracy.

[0003] Machine learning methods have improved the detection effect to a certain extent, but often rely on complex manual feature engineering and have insufficient generalization ability when facing new types of malicious URLs. In recent years, deep learning technology has made remarkable progress and demonstrated powerful capabilities in fields such as image recognition and natural language processing. In the aspect of malicious URL detection, there have also been some deep learning-based methods tried, but there are still problems such as insufficient utilization of URL structural characteristics, poor feature fusion effect, and inability to effectively handle sample diversity and dynamic changes. For example, some methods do not fully consider the position sensitivity and long-distance dependencies of URL tokens and lack effective strategies when fusing lexical, statistical, and structural features. Summary of the Invention

[0004] The self-driven multi-source malicious URL detection method integrating wide and deep models provided by the present invention improves on the deficiencies of the prior art. In terms of feature extraction, the wide and deep models are used to extract the lexical statistics and deep structural features of URLs respectively to achieve complementary advantages. When fusing features, a unique joint loss function is constructed to reasonably fuse features according to weight coefficients. The training strategy adopts a self-driven multi-source learning strategy, screening samples according to dynamic thresholds and regularizing the objective function, considering sample diversity, preventing overfitting, and improving the accuracy and robustness of the model to meet the requirements of network security protection.

[0005] In a first aspect, a self-driven multi-source malicious URL detection method integrating wide and deep models includes:

[0006] Step 1: Collect URL data and perform preprocessing;

[0007] Step 2: Perform positional embedding on the preprocessed URL data;

[0008] Step 3: Use the wide model in the wide and deep model to extract lexical features and statistical features of the preprocessed URL data;

[0009] Step 4: Use the deep model in the wide and deep model to extract the deep structural features of the URL data after positional embedding;

[0010] Step 5: Adopt a self-driven multi-source learning strategy to fuse the features obtained in Steps 3 and 4, and calculate the feature losses generated in Steps 3 and 4 based on the fused features for training sample screening. When the feature losses generated in Steps 3 and 4 are recalculated with the screened samples and reach a set value, update the parameters of the wide and deep model to obtain a self-driven multi-source malicious URL detection model that fuses the wide and deep model;

[0011] Step 6: Use the self-driven multi-source malicious URL detection model that fuses the wide and deep model to detect the real-time collected URL data after processing according to Steps 1 - 2.

[0012] Further, the specific steps of adopting a self-driven multi-source learning strategy to fuse the features extracted in Steps 3 and 4 are as follows:

[0013] Step 5.1: Construct the joint loss of the wide and deep model;

[0014]

[0015] where L WD represents the joint loss of the training sample set in the wide and deep model fusion, n represents the number of samples in the training dataset; d represents the total dimension of the fused features, j represents the dimension index, and the value range is from 1 to d, which is used to traverse each dimension of the feature; u j represents the weight coefficient of the j-th dimension feature of the fused feature, and satisfies which is used to measure the relative importance of each dimension in the overall loss calculation;

[0016] and respectively represent the losses generated by the wide model and the deep model in the wide and deep model fusion when processing the j-th dimension feature of the i-th sample, where y wij is the true label corresponding to the j-th dimension feature of the i-th sample of the wide model, y dij is the true label corresponding to the j-th dimension feature of the i-th sample of the deep model, and are the predicted labels of the outputs of the wide model and the deep model in the wide and deep model fusion for the j-th dimension feature of the i-th sample respectively; the predicted labels of the outputs of the wide model and the deep model in the wide and deep model fusion for the j-th dimension feature of the i-th sample;

[0017] and are respectively used to quantify the difference degrees between the predicted values and the true values of the wide model and the deep model in the wide and deep model fusion;

[0018] Step 5.2: After the joint loss based on the fusion wide and deep model, train and update the wide and deep model according to the objective function corresponding to the following formula;

[0019]

[0020] Wherein, represents the objective function, which is used to train the wide and deep model and update the parameters. v i represents the parameter for controlling the selection and weight assignment of training sample i. During the training process, v is determined through iterative optimization; i value; represents the basic coefficient of the regularization term, which is a constant parameter; η represents the decay rate for controlling the change with the training process in the regularization term, which is a hyperparameter; v represents the set of v i ;

[0021] Specifically, in each round of iteration, calculate the loss based on the prediction result of the wide and deep model for the sample and the true label, and continuously adjust v i according to the loss situation, so that v i can reflect the importance of the sample in the current training stage;

[0022] affects the change rhythm of the regularization term with the number of training rounds, making the regularization term show different change trends in the early, middle, and late stages of training to adapt to the learning needs of the model for samples at different training stages;

[0023] Step 5.3: Based on the wide and deep model with updated parameters, recalculate the feature losses of each sample extracted by the wide model and the deep model, and screen the training samples according to the decision variable formula;

[0024]

[0025] Wherein, S i represents the joint loss of the i-th sample in the fusion wide and deep model; δ represents the set loss threshold; represents whether the sample is selected. When the value is 0, it means that sample i is excluded. When the value is 1, it means that sample i is selected and retained in the sample set; t represents the number of training rounds or time steps; T represents a preset constant related to the training rounds or time;

[0026] is an intermediate result in the optimization process of v i and is a way to determine the assignment of v i under given conditions;

[0027] After the wide and deep model extracts lexical features, statistical features, and deep structural features, it will initially fuse the features extracted by the wide model and the deep model. The fusion method is to add the features output by the two element by element, and then use the decision variable formula in the following text to determine whether to select the sample into the training set, thereby completing the sample screening work to ensure that the selected samples are more valuable for model training and help the model learn and generalize better;

[0028] Step 5.4: Use the screened samples to train and update the wide and deep model again. Based on the wide and deep model updated again, extract the features of the samples again and perform element-wise addition to obtain the fused features.

[0029] The Self-Driven Multivariate Learning (SDM-SPLS) strategy is one of the core innovations of this solution. It adds a diversity term on the basis of self-paced learning to balance sample diversity, and automatically adjusts sample complexity and diversity by introducing specific parameters. This dynamic and adaptive learning strategy design requires in-depth research and unique thinking on the sample distribution and learning dynamics during model training. Different from traditional fixed sample selection or simple difficulty progressive learning methods, the SDM-SPLS strategy can more intelligently guide model learning, avoid overfitting and bias towards a certain type of sample. This innovative idea is relatively rare in the field of malicious URL detection and requires crossing the limitations of traditional thinking patterns.

[0030] Furthermore, perform regularization processing on the objective function to optimize the objective function, and use the optimized objective function and the screened sample set to train and update the wide and deep model;

[0031]

[0032] Among them, represents the optimized objective function, which is used to control sample diversity and complexity, and v j is used to control the selection and weight assignment of the screened training sample j. j is an index variable, and v j is the j-th element in the updated parameter v; ‖v j ‖2 represents the L2 norm of v j . Introduce the L2 norm in the regularization term to constrain the value of the parameter v j . μ represents the basic coefficient of the adjustment term and is a constant parameter.

[0033] In malicious URL detection, since samples may exhibit complex distributions in different dimensions and partitions, it is difficult to fully exploit their potential value through only the initial screening. Therefore, an additional regularization term is introduced to further consider the diversity of samples, and then the samples after the initial screening are screened again. In this process, the regularization term is used to constrain the parameter values to prevent overfitting. At the same time, samples are rescreened according to the new loss and adjusted conditions to better balance the diversity of samples in different partitions. For the samples selected after the second screening, their feature fusion methods are further optimized, so that the model can comprehensively learn the features of various samples, improving the generalization ability of the model and the accuracy of malicious URL detection.

[0034] Further, the loss function L of the wide model in the wide and deep model CFL includes the minimum cross-entropy loss L MCE and the contrastive loss L Cont ;

[0035] L CFL = αL MCE + βL Cont

[0036]

[0037] where α and β represent the weights used to balance the cross-entropy loss and the contrastive loss in the total loss, and their values are hyperparameters; n represents the number of training samples; y i represents the true label of sample i, y i = 1 indicates that sample i is a malicious URL, y i = 0 indicates that sample i is a normal URL; p i is the probability that the URL detection model predicts sample i as a malicious URL; d(a i , a j ) represents the distance metric between the feature representations of sample i and sample j, y ij is the label of the sample pair (i, j), used to indicate the class relationship of the sample pair. If sample pair i and j belong to the same malicious or normal URL, then y ij = 1. If one is a malicious URL and the other is a normal URL, then y ij = 0; m is a preset boundary value used to control the range of the contrastive loss, and its value range is generally [0.1, 5].

[0038] Further, the distance metric between the feature representations of sample i and sample j is calculated using the Euclidean distance:

[0039]

[0040] where, x ik and x jkrepresents the kth eigenvalue of samples i and j respectively; k represents an index variable used to traverse the dimension of the feature vector to comprehensively consider the two feature vectors a i and a j Differences in all dimensions.

[0041] Furthermore, the deep model in the wide-depth model is a TCVN network, and the TCVN network includes a CNN network and a TCN network connected in sequence. The input data first passes through the CNN network for local feature extraction, and then enters the TCN network to learn long-distance dependencies, and outputs a feature vector for classification or prediction;

[0042] Among them, the CNN network contains a one-dimensional convolution layer and a ReLU activation function. The input data is passed through the one-dimensional convolution layer to extract local features, and then the nonlinear transformation is performed through the ReLU activation function;

[0043] The TCN network is implemented by the TemporalConvNet class, which is a class used to build a temporal convolutional network.

[0044] Its main function is to integrate multiple TemporalBlocks to form a complete network structure for processing data with time series characteristics, including multiple TemporalBlocks. TemporalBlock is the basic component of TemporalConvNet. Each TemporalBlock has two one-dimensional convolutional layers, which perform feature extraction operations on the input data in sequence. At the same time, a residual connection is used to add part of the original input data to the result after processing by two convolutional layers.

[0045] The proposed Temporal Convolutional Visual Network (TCVN) model is specially designed for the characteristics of URL structure, which requires in-depth insights into the complexity of URL structure and the limitations of traditional models. The TCVN model cleverly combines the local feature processing capabilities of CNN and the medium- and long-sequence dependency processing capabilities of TCN. This combination is not a conventional technology combination. General research work may only try to make simple improvements on traditional TCN or CNN models, but it is difficult to think of this new architecture design for URL data characteristics to achieve more comprehensive feature learning and more powerful representation capabilities for URL structure.

[0046] Furthermore, the URL is preprocessed, specifically including:

[0047] Step A1: Clean the URL data to remove duplicates and correct obvious spelling errors, and remove records containing missing items;

[0048] Step A2: Perform data reduction on the cleaned URL data;

[0049] (1) Delete the string after the question mark "?" in the URL.

[0050] In the URL structure, the question mark is often used as a separator, and the content after it is usually a query string. Deleting this part of the content helps improve the accuracy of subsequent analysis and reduces interference from unnecessary information.

[0051] (2) Delete the string after the hash mark "#" in the URL;

[0052] In the URL structure, "#" represents a part of the web page and is not sent to the server side.

[0053] Step A3: Perform a copy operation on the URL data after data reduction processing to generate two identical sample sets, one for retention for subsequent steps and one for data obfuscation;

[0054] Data obfuscation includes at least five types of malicious URL obfuscations:

[0055] (1) Disguise the URL using a fake domain name;

[0056] (2) Obfuscate the URL through specific keywords;

[0057] (3) Obfuscate the URL using a domain name with similar spelling or an extremely long domain name;

[0058] (4) Replace the domain name with an IP address to avoid detection;

[0059] (5) Use URL short link services for obfuscation;

[0060] Step A4: After encoding the obfuscated data, classify all URL data;

[0061] For samples that disguise the URL using a fake domain name, the encoding is set to 1; for samples that obfuscate the URL through specific keywords, the encoding is set to 2; for samples that obfuscate using a domain name with similar spelling or an extremely long domain name, the encoding is set to 3; if a sample has the situation of replacing the domain name with an IP address to avoid detection, the encoding is set to 4; for samples that use URL short link services for obfuscation, the encoding is set to 5.

[0062] Divide all URL data into two major categories, the positive sample set and the negative sample set. The positive sample set is the sample set retained in Step A3, and the negative sample set is the URL data encoded in Step A4.

[0063] Furthermore, perform positional embedding on the preprocessed URL data, and the specific process is as follows:

[0064] Step B1: Cut the positive sample set and the negative sample set into four blocks according to the position of " / ": protocol, domain, path, and file;

[0065] (1) The block before the first " / " is defined as the protocol part;

[0066] (2) The string before the second " / " is defined as the domain part;

[0067] (3) The string after the last " / " is defined as the file part;

[0068] (4) The remaining string is defined as the path part;

[0069] Step B2: Split the consecutive strings trimmed in Step B1 with a set of delimiters defined in A = { / , -, :,., #, @,?, =}, and introduce special tokens to replace the special delimiters defined in the middle. The special tokens include slash, dash, colon, dot, hash, at, mark, equal. Specifically:

[0070] To minimize information loss, introduce special tokens to replace the special delimiters. This avoids the confusion between these special delimiters and their grammatical meanings in the URL itself during subsequent processing, and also makes it easier for the model to process them like ordinary text tokens; to overcome the problem of token ambiguity, the alignment strategy will use different types of parentheses to locate the tokens in different blocks;

[0071] (1) Each token in the protocol part is placed in curly braces {};

[0072] (2) Each token in the domain part is placed in parentheses ();

[0073] (3) The tokens in the path part are placed in angle brackets <>;

[0074] (4) The tokens in the file part are modified with square brackets [];

[0075] In a second aspect, a system for a self-driven multi-source malicious URL detection method using the above-mentioned fusion wide and deep model includes:

[0076] Data collection and preprocessing module: Collect URL data and perform preprocessing;

[0077] Position embedding module: Perform position embedding on the preprocessed URL data;

[0078] Wide model feature extraction module: Use the wide model in the wide and deep model to extract lexical features and statistical features from the preprocessed URL data;

[0079] Deep module feature extraction module: Extract the deep structural features of the URL data after position embedding by using the deep model in the wide and deep model;

[0080] Fusion model learning module: Adopt a self-driven multi-learning strategy to fuse the features extracted by the wide model and the deep model modules, calculate the feature loss generated when the wide model and the deep model modules extract features based on the fused features, and perform training sample screening. When the feature loss generated when the wide model and the deep model modules extract features is calculated again with the screened samples and reaches the set value, update the parameters of the wide and deep model to obtain a self-driven multi-malicious URL detection model that fuses the wide and deep model;

[0081] URL detection module: Use the self-driven multi-malicious URL detection model that fuses the wide and deep model to detect the real-time collected URL data after being processed by the data collection and preprocessing module and the position embedding module.

[0082] Thirdly, a computer storage medium stores a computer program, and the computer program is called by a processor to implement:

[0083] The steps of the above-mentioned self-driven multi-malicious URL detection method that fuses the wide and deep model.

[0084] Beneficial effects

[0085] Compared with the existing methods, the advantages of the present invention are:

[0086] The technical solution of the present invention fuses the wide and deep model through a self-driven multi-method, fully considering the different characteristics of URLs. In the wide model, the factorization machine is used to learn the potential interactions of statistical and lexical features, and in the deep model, the temporal convolutional vision network (TCVN) with position embedding is used to capture structural features. At the same time, the sample weights are dynamically adjusted through the diverse self-regulated self-paced learning, and samples of different difficulties are adaptively selected to participate in the training, providing a more comprehensive basis for malicious URL detection.

[0087] The present invention proposes a self-driven multi-source learning strategy integrating wide and deep models. Considering the complexity of URL data itself, including structural diversity, diversified obfuscation means, etc., traditional detection methods often have difficulty in comprehensively capturing its features. It pays attention to the impact of the difficulty and diversity of samples on the model performance during the model training process, and avoids the model from overfitting simple samples or being biased towards a certain type of samples. A Temporal Convolutional Vision Network (TCVN) model is designed. By organically integrating the local feature extraction ability of CNN and the temporal modeling advantage of TCN, it breaks through the limitations of traditional deep learning models in processing multi-dimensional structural features of URLs and can better extract the multi-dimensional structural features of URLs. A self-driven multi-source learning (SDM-SPLS) strategy is proposed, which is the core innovation point of this research. Based on the traditional self-paced learning framework, this strategy innovatively introduces a dual-parameter regulation mechanism: on the one hand, it dynamically adjusts the learning difficulty of samples through a complexity perception function to avoid the model falling into local optimum in the initial stage of training; on the other hand, it constructs a diversity evaluation term and establishes a sample diversity compensation mechanism to ensure the balanced learning of samples in different feature partitions. This two-dimensional regulation enables the model to autonomously adapt to complex and changeable sample distributions, significantly enhancing the generalization ability of the model while improving the robustness against adversarial sample obfuscation. Description of the Drawings

[0088] Figure 1 It is a flowchart for detecting malicious URL models of the technical solution of the present invention;

[0089] Figure 2 It is a display diagram of URL position embedding tags of the technical solution of the present invention;

[0090] Figure 3 It is a graph showing the changing trend of the accuracy of the malicious URL detection model of the technical solution of the present invention;

[0091] Figure 4 It is a graph showing the changing trend of the loss value of the malicious URL detection model of the technical solution of the present invention;

[0092] Figure 5 It is a comparison schematic diagram of the accuracies of five different obfuscation types of malicious URLs and positive samples in the training set;

[0093] Figure 6 It is a comparison schematic diagram of the accuracies of five different obfuscation types of malicious URLs and positive samples in the validation set;

[0094] Figure 7 It is a comparison schematic diagram of the accuracies of five different obfuscation types of malicious URLs and positive samples in the test set. Detailed Embodiment

[0095] The following will further illustrate the technical solution of the present invention in combination with the drawings and embodiments.

[0096] Embodiment 1

[0097] like Figure 1 As shown in the figure, a self-driven multi-dimensional malicious URL detection method integrating wide and deep models includes:

[0098] Step 1: Collect URL data and pre-process it;

[0099] Preprocess the URL, including:

[0100] Step A1: Clean the URL data to remove duplicates and correct obvious spelling errors, and remove records containing missing items;

[0101] Step A2: simplifying the cleaned URL data;

[0102] (1) Delete the string after the question mark "?" in the URL.

[0103] In the URL structure, the question mark is often used as a separator, and the content after it is usually a query string. Deleting this part of the content can help improve the accuracy of subsequent analysis and reduce the interference of unnecessary information;

[0104] (2) Delete the string after the # sign in the URL;

[0105] In the URL structure, "#" represents a part of the web page that will not be sent to the server.

[0106] Step A3: Duplicate the URL data after data simplification to generate two identical sample sets, one for subsequent steps and one for data obfuscation;

[0107] Data obfuscation includes at least five types of malicious URL obfuscation:

[0108] (1) Disguise the URL using a fake domain name;

[0109] (2) URL obfuscation through specific keywords;

[0110] (3) Using similarly spelled domain names or very long domain names to obfuscate URLs;

[0111] (4) Using IP addresses to replace domain names to evade detection;

[0112] (5) Obfuscation through URL shortening services;

[0113] Step A4: After encoding the obfuscated data, classify all URL data;

[0114] For samples that use a fake domain name to disguise the URL, the encoding is set to 1; for samples that obfuscate the URL with specific keywords, the encoding is set to 2; for samples that use similar spelling or extremely long domain names for obfuscation, the encoding is set to 3; if a sample uses an IP address to replace the domain name to avoid detection, the encoding is set to 4; for samples that use URL short link services for obfuscation, the encoding is set to 5;

[0115] All URL data is divided into two categories: the positive sample set and the negative sample set. The positive sample set is the sample set retained in step A3, and the negative sample set is the URL data after encoding in step A4.

[0116] Step 2: Perform positional embedding on the preprocessed URL data;

[0117] Perform positional embedding on the preprocessed URL data, and the specific process is as follows:

[0118] Step B1: Cut the positive and negative sample sets into four blocks according to the position of " / ": protocol, domain, path, and file;

[0119] (1) The block before the first " / " is defined as the protocol part;

[0120] (2) The string before the second " / " is defined as the domain part;

[0121] (3) The string after the last " / " is defined as the file part;

[0122] (4) The remaining string is defined as the path part;

[0123] Step B2: Split the continuous string trimmed in step B1 with a set of delimiters defined in A = { / , -, :,., #, @,?, =}, and introduce special tokens to replace the special delimiters defined in it. The special tokens include slash, dash, colon, dot, hash, at, mark, equal. The processing process is as Figure 2 shown. Specifically:

[0124] To minimize information loss, introduce special tokens to replace the special delimiters. This avoids confusion between these special delimiters and their grammatical meanings in the URL itself in subsequent processing, and also makes it easier for the model to process them like ordinary text tokens; to overcome the problem of token ambiguity, the alignment strategy will use different types of brackets to locate the tokens in different blocks;

[0125] (1) Each token in the protocol part is placed in curly braces {};

[0126] (2) Each tag in the domain part is placed in parentheses ().

[0127] (3) The tags in the path part are placed in angle brackets <>.

[0128] (4) The tags in the file part are modified with square brackets [].

[0129] Step 3: Use the wide model in the wide and deep model to extract lexical features and statistical features from the preprocessed URL data;

[0130] The loss function L of the wide model in the wide and deep model CFL includes the minimum cross-entropy loss L MCE and the contrastive loss L Cont ;

[0131] L CFL = αL MCE + βL Cont

[0132]

[0133] where α and β represent the weights used to balance the cross-entropy loss and the contrastive loss in the total loss, and their values are hyperparameters; n represents the number of training samples; y i represents the true label of sample i, y i = 1 indicates that sample i is a malicious URL, y i = 0 indicates that sample i is a normal URL; p i is the probability that the URL detection model predicts sample i as a malicious URL; d(a i , a j ) represents the distance metric between the feature representations of sample i and sample j, y ij is the label of the sample pair (i, j), used to indicate the class relationship of the sample pair. If sample pair i and j both belong to malicious or normal URLs, then y ij = 1. If one is a malicious URL and the other is a normal URL, then y ij = 0; m is a preset boundary value used to control the range of the contrastive loss, and its value range is generally [0.1, 5].

[0134] In the article "Supervised Contrastive Loss in Python", the author mentioned that when training the model, the value of m needs to be selected within a certain range according to experience;

[0135] The distance metric between the feature representations of sample i and sample j is calculated using the Euclidean distance:

[0136]

[0137] where, xik and x jk represent the k-th eigenvalue of samples i and j, respectively; k represents an index variable used to traverse the dimensions of the feature vector to comprehensively consider the two feature vectors a i and a j The differences in each dimension.

[0138] Step 4: Use the deep model in the wide and deep model to extract the deep structural features of the URL data after position embedding;

[0139] The deep model in the wide and deep model is a TCVN network. The TCVN network includes a CNN network and a TCN network connected in sequence. The input data first undergoes local feature extraction by the CNN network and then enters the TCN network to learn long-range dependency relationships, and outputs a feature vector for classification or prediction;

[0140] Among them, the CNN network contains a one-dimensional convolutional layer and a ReLU activation function. The input data extracts local features through the one-dimensional convolutional layer and then undergoes a non-linear transformation through the ReLU activation function;

[0141] The TCN network is implemented by the TemporalConvNet class. TemporalConvNet is a class for constructing a temporal convolutional network.

[0142] Its main role is to integrate multiple TemporalBlocks to form a complete network structure for processing data with temporal characteristics, including multiple TemporalBlocks. TemporalBlock is the basic building block of TemporalConvNet. Each TemporalBlock has two layers of one-dimensional convolutional layers, and these two convolutional layers sequentially perform feature extraction operations on the input data. At the same time, a residual connection method is adopted to add a part of the original input data to the result after being processed by the two convolutional layers.

[0143] The proposed Temporal Convolutional Vision Network (TCVN) model is specifically designed for the URL structure characteristics, which requires an in-depth understanding of the complexity of the URL structure and the limitations of traditional models. The TCVN model cleverly combines the local feature processing ability of CNN and the long and medium sequence dependency processing ability of TCN. This combination method is not a conventional technical combination. General research work may only try to make simple improvements on traditional TCN or CNN models, and it is difficult to come up with this brand-new architecture design for URL data characteristics to achieve more comprehensive feature learning and more powerful representation ability for the URL structure.

[0144] Step 5: Adopt a self-driven multi-modal learning strategy to fuse the features obtained in Steps 3 and 4, and calculate the feature losses generated in Steps 3 and 4 based on the fused features to screen the training samples. When the feature losses generated in Steps 3 and 4 calculated with the screened samples reach a set value, update the parameters of the wide and deep model to obtain a self-driven multi-modal malicious URL detection model that fuses the wide and deep model;

[0145] The specific steps to fuse the features extracted in Steps 3 and 4 by adopting a self-driven multi-modal learning strategy are as follows:

[0146] Step 5.1: Construct the joint loss of the wide and deep model;

[0147]

[0148] Among them, L WD represents the joint loss of the training sample set in the fused wide and deep model, n represents the number of samples in the training dataset; d represents the total dimension of the fused features, j represents the dimension index, and the value range is from 1 to d, which is used to traverse each dimension of the features; u j represents the weight coefficient of the j-th dimension feature of the fused feature, and satisfies which is used to measure the relative importance of each dimension in the overall loss calculation;

[0149] and respectively represent the losses generated by the wide model and the deep model in the fused wide and deep model when processing the j-th dimension feature of the i-th sample, where y wij is the true label corresponding to the j-th dimension feature of the i-th sample of the wide model, y dij is the true label corresponding to the j-th dimension feature of the i-th sample of the deep model, and are the predicted labels of the outputs of the wide model and the deep model in the fused wide and deep model for the j-th dimension feature of the i-th sample respectively; the predicted labels of the outputs of the wide model and the deep model in the fused wide and deep model for the j-th dimension feature of the i-th sample respectively;

[0150] and are respectively used to quantify the degree of difference between the predicted values and the true values of the wide model and the deep model in the fused wide and deep model;

[0151] Step 5.2: After the joint loss of the fused wide and deep model, train and update the wide and deep model according to the objective function corresponding to the following formula;

[0152]

[0153] Among them, represents the objective function, which is used to train the wide and deep model and update the parameters, v i represents the parameter for controlling the selection of training sample i and the weight assignment. During the training process, v is determined through iterative optimization i value; represents the base coefficient of the regularization term, which is a constant parameter; η represents the decay rate for controlling the change of the regularization term with the training process, which is a hyperparameter; v represents the set of v i set.

[0154] Specifically, in each round of iteration, the loss is calculated based on the prediction result of the wide and deep model for the sample and the true label, and v is continuously adjusted according to the loss situation i such that v i can reflect the importance of the sample at the current training stage.

[0155] affects the change rhythm of the regularization term with the number of training rounds, making the regularization term show different change trends in the early, middle, and late stages of training to adapt to the sample learning needs of the model at different training stages.

[0156] Step 5.3: Based on the wide and deep model with updated parameters, recalculate the feature losses of each sample extracted by the wide model and the deep model, and screen the training samples according to the decision variable formula;

[0157]

[0158]

[0159] Among them, S i represents the joint loss of the i-th sample in the integrated wide and deep model; δ represents the set loss threshold; represents whether the sample is selected. When the value is 0, it means that the sample i is excluded. When the value is 1, it means that the sample i is selected and retained in the sample set; t represents the number of training rounds or time steps; T represents a preset constant related to the training rounds or time;

[0160] is an intermediate result in the optimization process of v i and is a way to determine the assignment of v under given conditions i ;

[0161] After the wide and deep model extracts lexical features, statistical features, and deep structural features, the features extracted by the wide model and the deep model will be preliminarily integrated. The integration method is to add the features output by the two element by element, and then use the decision variable formula in the following text to determine whether to select the sample into the training set, thereby completing the sample screening work to ensure that the selected samples are more valuable for model training and help the model better learn and generalize;

[0162] Step 5.4: Use the filtered samples to train and update the wide and deep model again. Based on the wide and deep model updated again, extract the features of the samples again and perform pixel-by-pixel addition to obtain the fused features.

[0163] The Self-Driven Multi-View Learning (SDM-SPLS) strategy is one of the core innovations of this solution. It adds a diversity term on the basis of self-paced learning to balance sample diversity, and automatically adjusts sample complexity and diversity by introducing specific parameters. This design of dynamic adaptive learning strategy requires in-depth research and unique thinking on the sample distribution and learning dynamics during the model training process. Different from traditional fixed sample selection or simple difficulty progressive learning methods, the SDM-SPLS strategy can more intelligently guide the model to learn, avoid overfitting and bias towards a certain type of sample. This innovative idea is relatively rare in the field of malicious URL detection and requires transcending the limitations of traditional thinking patterns.

[0164] Regularize the objective function, optimize the objective function, and use the optimized objective function and the filtered sample set to train and update the wide and deep model;

[0165]

[0166] where represents the optimized objective function, which is used to control sample diversity and complexity, and v j is used to control the selection and weight assignment of the filtered training sample j. j is an index variable, and v j is the j-th element in the updated parameter v; |v j ||2 represents the L2 norm of v j . Introduce the L2 norm in the regularization term to constrain the value of the parameter v j . μ represents the basic coefficient of the adjustment term, which is a constant parameter.

[0167] In malicious URL detection, since the samples may present complex distributions in different dimensions and partitions, it is difficult to fully exploit their potential value with only the initial screening. Therefore, an additional regularization term is introduced to further consider sample diversity, and then the samples after the initial screening are screened again. During this process, use the regularization term to constrain the parameter values to prevent overfitting, and at the same time re-screen the samples according to the new loss and adjusted conditions to better balance the diversity of samples in different partitions. For the samples selected after the re-screening, further optimize their feature fusion method, so that the model can comprehensively learn the features of various samples and improve the generalization ability of the model and the accuracy of malicious URL detection.

[0168] Step 6: Use the self-driven multi-source malicious URL detection model that integrates the wide and deep model to detect the real-time collected URL data after processing according to Steps 1 - 2.

[0169] Obtain a total sample set containing a positive sample set and a negative sample set, perform a replication operation on the total sample set to obtain two identical total sample sets. Perform position embedding processing (the position embedding processing method in this solution) on one of the total sample sets to obtain a position-embedded sample set, and perform Word2Vec word vector processing on the other total sample set to obtain a non-position-embedded sample set.

[0170] To more intuitively evaluate the performance of the self-driven multi-source malicious URL detection method that integrates the wide and deep model, a series of experiments were conducted and charts were drawn. In the entire model process, through Figure 1 The key steps and process directions of the self-driven multi-source malicious URL detection method that integrates the wide and deep model can be clearly shown. After position embedding processing, through Figure 2 The conversion of the URL format can be clearly seen, which lays a foundation for the effective learning of the subsequent model. During the model training and testing process, Figure 3 and Figure 4 Respectively show the overall learning trend of the model from the perspectives of accuracy and loss value, and it can be seen that the model is gradually optimized during continuous iterative training. Further, Figure 5 、 Figure 6 and Figure 7 Detailed accuracy analysis was carried out for positive and negative samples, especially different types of negative samples, on the training set, validation set, and test set. This is crucial for deeply exploring the detection ability of the model for various malicious URLs, helping us comprehensively grasp the advantages and disadvantages of the model, and thus providing a strong basis for subsequent improvements.

[0171] In Figure 2 https: / / www.newexample.com / docs / info / main.doc represents a complete URL. After being processed by a specific step B1, it presents the form of https www.newexample.com docs / info main.doc. Among them, https represents the Hypertext Transfer Protocol Secure, www.newexample.com represents the domain name, docs / info represents the file part, main.doc represents the path part, protocol means protocol, domain means domain name, subdirectory means subdirectory, file means file; {https} indicates marking this part as the protocol content, www.newexample.com indicates marking this part as the domain name, <docs>slash <info>It indicates that this part is a table of contents. "slash" represents the slash symbol. "[main]dot[doc]" indicates that this part is a document, and "dot" represents the dot symbol.

[0172] In Figure 3 Training, Validation and Test Accuracy represent the training, validation, and test accuracies, which are used to show the performance of the model in terms of accuracy on different datasets. Among them, the x-axis represents the training rounds (Training Round), and the y-axis represents the accuracy (Accuracy). The blue dashed line represents Train Accuracy (training set accuracy), which is the prediction accuracy of the model on the training dataset; the green dashed line represents Validation Accuracy (validation set accuracy), which is the prediction accuracy of the model on the validation dataset and is used to adjust the model parameters; the red dashed line represents Test Accuracy (test set accuracy), which is the prediction accuracy of the model on the test dataset and reflects the final generalization ability of the model.

[0173] In Figure 4 Training, Validation and Test Loss represent the training, validation, and test loss values, which show the changes in training loss, validation loss, and test loss with the number of training rounds. Among them, the x-axis represents the training rounds (Training Round), and the y-axis represents the loss value (Loss). The blue dashed line represents Train Loss (training set loss), which is the degree of difference between the prediction result and the true value of the model on the training data; the green dashed line represents Validation Loss (validation set loss), which is used to adjust the model parameters during training to prevent overfitting; the red dashed line represents Test Loss (test set loss), which reflects the final adaptability of the model to new data.

[0174] In Figure 5 Among them, Training Accuracy for neg0-neg4 represents the training accuracy from neg0 to neg4, showing the variation of the training accuracy of different categories (neg0, neg1, neg2, neg3, neg4, and pos) with the number of training rounds. Among them, the x-axis represents the number of training rounds (Training Round); the y-axis represents the accuracy value (Accuracy Value); neg0-neg4 corresponds to the five malicious URL obfuscation types in the text, representing different types of negative samples; pos represents positive samples. The blue dashed line represents neg0 train acc (the training set accuracy of the neg0 category); the orange dashed line represents neg1 train acc (the training set accuracy of the neg1 category); the green dashed line represents neg2 train acc (the training set accuracy of the neg2 category); the red dashed line represents neg3 train acc (the training set accuracy of the neg3 category); the purple dashed line represents neg4 train acc (the training set accuracy of the neg4 category); the brown dashed line represents pos train acc (the training set accuracy of the pos category).

[0175] In Figure 6 Among them, Valid Accuracy for neg0-neg4 represents the validation accuracy from neg0 to neg4, showing the variation of the validation accuracy of different categories (neg0, neg1, neg2, neg3, neg4, and pos) with the number of training rounds. Among them, the x-axis represents the number of training rounds (Training Round); the y-axis represents the accuracy value (Accuracy Value); neg0-neg4 corresponds to the five malicious URL obfuscation types in the text, representing different types of negative samples; pos represents positive samples. The blue dashed line represents neg0 valid acc (the validation set accuracy of the neg0 category); the orange dashed line represents neg1 valid acc (the validation set accuracy of the neg1 category); the green dashed line represents neg2 valid acc (the validation set accuracy of the neg2 category); the red dashed line represents neg3 valid acc (the validation set accuracy of the neg3 category); the purple dashed line represents neg4 valid acc (the validation set accuracy of the neg4 category); the brown dashed line represents pos valid acc (the validation set accuracy of the pos category).

[0176] In Figure 7 Among them, Test Accuracy for neg0-neg4 represents the test accuracy from neg0 to neg4, showing the variation of the test accuracy of different categories (neg0, neg1, neg2, neg3, neg4, and pos) with the number of training rounds. Among them, the x-axis represents the number of training rounds (Training Round); the y-axis represents the accuracy value (Accuracy Value); neg0-neg4 correspond to five malicious URL obfuscation types in the text, representing different types of negative samples; pos represents positive samples. The blue dashed line represents neg0 test acc (the test set accuracy of the neg0 category); the orange dashed line represents neg1 test acc (the test set accuracy of the neg1 category); the green dashed line represents neg2 test acc (the test set accuracy of the neg2 category); the red dashed line represents neg3 test acc (the test set accuracy of the neg3 category); the purple dashed line represents neg4 test acc (the test set accuracy of the neg4 category); the brown dashed line represents pos test acc (the test set accuracy of the pos category).

[0177] Example 2

[0178] A system for a self-driven multi-source malicious URL detection method using the above-mentioned integrated wide and deep model, including:

[0179] Data collection and preprocessing module: Collect URL data and perform preprocessing;

[0180] Position embedding module: Perform position embedding on the preprocessed URL data;

[0181] Wide model feature extraction module: Use the wide model in the wide and deep model to extract lexical features and statistical features from the preprocessed URL data;

[0182] Deep module feature extraction module: Use the deep model in the wide and deep model to extract the deep structural features of the URL data after position embedding;

[0183] Integrated model learning module: Adopt a self-driven multi-source learning strategy to fuse the features extracted by the wide model and the deep model modules, and calculate the feature loss generated when the wide model and the deep model modules extract features based on the fused features, perform training sample screening. When the feature loss generated when the wide model and the deep model modules extract features with the screened samples reaches a set value, update the parameters of the wide and deep model to obtain a self-driven multi-source malicious URL detection model of the integrated wide and deep model;

[0184] URL Detection Module: Using a self-driven multi-source malicious URL detection model that integrates wide and deep models, the URL data collected in real time is detected after being processed by the data collection and preprocessing module and the location embedding module.

[0185] It should be understood that the implementation process of each setting can refer to the content description of the foregoing method.

[0186] Embodiment 3

[0187] In a third aspect, a computer storage medium stores a computer program, and the computer program is called by a processor to implement:

[0188] The steps of the foregoing self-driven multi-source malicious URL detection method that integrates wide and deep models.

[0189] For the specific implementation process of each step, please refer to the description of the foregoing method.

[0190] It should be understood that in the embodiments of the present invention, the so-called processor may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0191] The readable storage medium is a computer-readable storage medium, which may be an internal storage unit of the software and hardware device described in any of the foregoing embodiments, such as the hard disk or memory of the controller. The readable storage medium may also be an external storage device of the controller, such as a plug-in hard disk equipped on the controller, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the readable storage medium may also include both the internal storage unit of the controller and the external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0192] Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing readable storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0193] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The present application is a device that generates, according to the flowchart of the method, device (system), and computer program product of the embodiments of the present application and the instructions executed by the processor, for implementing the functions specified in one process or multiple processes of the flowchart and / or one block or multiple blocks of the block diagram. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in one process or multiple processes of the flowchart and / or one block or multiple blocks of the block diagram. These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process or multiple processes of the flowchart and / or one block or multiple blocks of the block diagram.

[0194] It should be emphasized that the examples described in the present invention are illustrative rather than restrictive. Therefore, the present invention is not limited to the examples described in the specific embodiments. Any other embodiments obtained by those skilled in the art based on the technical solutions of the present invention, whether modified or replaced, as long as they do not depart from the spirit and scope of the present invention, also belong to the protection scope of the present invention.< / info> < / docs>

Claims

1. A self-driven multi-source malicious URL detection method integrating wide and deep models, characterized in that include: Step 1: Collect URL data and pre-process it; Step 2: Perform location embedding on the preprocessed URL data; Step 3: Use the wide model in the wide-deep model to extract lexical features and statistical features from the preprocessed URL data; Step 4: Use the deep model in the wide-deep model to extract the deep structural features of the URL data after the location is embedded; Step 5: Adopt a self-driven multi-learning strategy to fuse the features obtained in step 3 and step 4, and calculate the feature loss generated in step 3 and step 4 based on the fused features, and screen the training samples. When the feature loss generated in step 3 and step 4 is calculated again with the screened samples and reaches the set value, update the parameters of the wide and deep model to obtain a self-driven multi-malicious URL detection model that fuses the wide and deep model; Step 6: Use the self-driven multi-dimensional malicious URL detection model that integrates the wide and deep models to process the URL data collected in real time according to steps 1 and 2, and then perform detection.

2. The method according to claim 1, wherein The specific steps of using the self-driven multi-learning strategy to fuse the features extracted in step 3 and step 4 are as follows: Step 5.1: Construct the joint loss of wide and deep models; Among them, L WD represents the joint loss of the training sample set in the fusion wide and deep model, n represents the number of samples in the training dataset; d represents the total dimension of the fusion features, j represents the dimension index, and its value range is from 1 to d, which is used to traverse each dimension of the features; u j represents the weight coefficient of the j-th dimension feature of the fusion features, and satisfies which is used to measure the relative importance of each dimension in the overall loss calculation; and respectively represent the losses generated by the wide model and the deep model in the fused wide-deep model when processing the j-th dimensional feature of the i-th sample, where y wij is the true label corresponding to the j-th dimensional feature of the i-th sample of the wide model, and y dij is the true label corresponding to the j-th dimensional feature of the i-th sample of the deep model. and are the predicted labels of the outputs made by the wide model and the deep model in the fused wide-deep model for the k-th dimensional feature of the i-th sample; the predicted labels of the outputs made by the wide model and the deep model in the fused wide-deep model for the j-th dimensional feature of the i-th sample. Step 5.2: Based on the joint loss of the fused wide and deep models, the wide and deep models are trained and updated according to the objective function corresponding to the following formula; Among them, represents the objective function, which is used to train the wide and deep model and update parameters, and v i represents the parameter for controlling the selection of training sample i and weight allocation. During the training process, v is determined by iterative optimization i value; represents the basic coefficient of the regularization term, which is a constant parameter; η represents the decay rate that varies with the training process in the regularization term and is a hyperparameter; v represents the set of v i values. Step 5.3: Based on the wide and deep models with updated parameters, recalculate the feature loss of each sample extracted by the wide model and the deep model, and screen the training samples according to the decision variable formula; Among them, S i represents the combined loss of the i-th sample in the fused wide and deep model; δ represents the set loss threshold; indicates whether sample i is selected. When the value is 0, it means sample i is excluded. When the value is 1, it means sample i is selected and retained in the sample set; t represents the number of training rounds or time steps; T represents a preset constant related to the training rounds or time; Step 5.4: Use the screened samples to train the wide-depth model again and update its parameters. Based on the updated wide-depth model, extract the features of the samples again and add them pixel by pixel to obtain the fused features.

3. The method according to claim 2, wherein Regularize and optimize the objective function, and use the optimized objective function and the screened sample set to train and update the wide and deep model; Among them, represents the optimization objective function, which is used to control the sample diversity and complexity, v j is used to control the selection and weight assignment of the filtered training sample j, where j is an index variable, v j is the j-th element in the updated parameter v; ||v j ||2 represents the L2 norm of v j and the introduction of the L2 norm in the regularization term constrains the value of the parameter v j ; μ represents the basic coefficient of the adjustment term and is a constant parameter.

4. The method according to claim 1, wherein The loss function L of the wide model in the wide and deep model CFL includes the minimum cross-entropy loss L MCE and the contrastive loss L Cont ; L CFL = αL MCE + βL Cont Among them, α and β represent the weights used to balance the cross-entropy loss and the contrastive loss in the total loss, and their values are hyperparameters; n represents the number of training samples; y i represents the true label of sample i, and y i = 1 indicates that sample i is a malicious URL, and y i = 0 indicates that sample i is a normal URL; p i is the probability that the URL detection model predicts sample i as a malicious URL; d(a i , a j ) represents the distance metric between the feature representations of sample i and sample j, and y ij is the label of the sample pair (i, j), which is used to indicate the class relationship of the sample pair. If sample pair i and j both belong to malicious or normal URLs, then y ij = 1. If one is a malicious URL and the other is a normal URL, then y ij = 0; m is a preset boundary value used to control the range of the contrastive loss, and its value range is [0.1, 5].

5. The method according to claim 4, characterized in that, The distance metric between the feature representations of sample i and sample j is calculated using the Euclidean distance: where x ik and x jk represent the k-th eigenvalue of samples j and j, respectively; k represents an index variable used to traverse the dimensions of the eigenvector to comprehensively consider the differences between the two eigenvectors a i and a j in each dimension.

6. The method according to claim 1, wherein The deep model in the wide-depth model is a TCVN network, and the TCVN network includes a CNN network and a TCN network connected in sequence. The input data first passes through the CNN network for local feature extraction, and then enters the TCN network to learn long-distance dependencies, and outputs a feature vector for classification or prediction; Among them, the CNN network contains a one-dimensional convolution layer and a ReLU activation function. The input data is passed through the one-dimensional convolution layer to extract local features, and then the nonlinear transformation is performed through the ReLU activation function; The TCN network is implemented by the TemporalConvNet class, which is a class used to build a temporal convolutional network.

7. The method according to claim 1, wherein Preprocess the URL, including: Step A1: Clean the URL data to remove duplicates and correct obvious spelling errors, and remove records containing missing items; Step A2: simplifying the cleaned URL data; (1) Delete the string after the question mark "?" in the URL. (2) Delete the string after the pound sign "#" in the URL; Step A3: Copy the URL data after data reduction processing to generate two identical sample sets. One is retained for subsequent steps, and the other is used for data obfuscation; Data obfuscation includes at least five types of malicious URL obfuscations: (1) Disguise the URL using a false domain name; (2) Obfuscate the URL through specific keywords; (3) Obfuscate the URL using a domain name with similar spelling or an extremely long domain name; (4) Replace the domain name with an IP address to avoid detection; (5) Use a URL short link service for obfuscation; Step A4: After encoding the obfuscated data, classify all URL data; For samples that disguise the URL using a false domain name, the encoding is set to 1; for samples that obfuscate the URL through specific keywords, the encoding is set to 2; for samples that obfuscate using a domain name with similar spelling or an extremely long domain name, the encoding is set to 3; if a sample replaces the domain name with an IP address to avoid detection, the encoding is set to 4; for samples that use a URL short link service for obfuscation, the encoding is set to 5; All URL data is divided into two major categories, the positive sample set and the negative sample set. The positive sample set is the sample set retained in Step A3, and the negative sample set is the URL data encoded in Step A4.

8. The method according to claim 7, wherein Perform position embedding on the preprocessed URL data. The specific process is as follows: Step B1: Cut the URL into four blocks according to the position of " / ": protocol, domain, path, and file for both the positive sample set and the negative sample set; (1) The block before the first " / " is defined as the protocol part; (2) The string before the second " / " is defined as the domain part; (3) The string after the last " / " is defined as the file part; (4) The remaining string is defined as the path part; Step B2: Split the continuous string trimmed in Step B1 using a set of delimiters defined in A = { / , -, :., #, @,?, =}, and introduce special tokens to replace the special delimiters defined. The special tokens include slash, dash, colon, dot, hash, at, mark, equal. Specifically: (1) Each token in the protocol part is placed in curly braces {}; (2) Each token in the domain part is placed in parentheses (); (3) The tokens in the path part are placed in angle brackets <>; (4) The tokens in the file part are modified with square brackets []; 9. A system for a self-driven multi-source malicious URL detection method using the fusion wide and deep model according to any one of claims 1-8, characterized in that, Include: Data collection and preprocessing module: Collect URL data and perform preprocessing; Position embedding module: Perform position embedding on the preprocessed URL data; Wide model feature extraction module: Extract lexical features and statistical features of the preprocessed URL data using the wide model in the wide and deep model; Deep module feature extraction module: Extract the deep structure features of the URL data after position embedding using the deep model in the wide and deep model; Fusion model learning module: Adopt a self-driven multi-faceted learning strategy to fuse the features extracted by the wide model and the deep model modules, calculate the feature loss generated when the wide model and the deep model modules extract features based on the fused features, and perform training sample screening. When the feature loss generated when the wide model and the deep model modules extract features is recalculated with the screened samples and reaches a set value, update the parameters of the wide and deep models to obtain a self-driven multi-faceted malicious URL detection model that fuses the wide and deep models; URL detection module: Use the self-driven multi-faceted malicious URL detection model that fuses the wide and deep models to detect the real-time collected URL data after processing by the data collection and preprocessing module and the position embedding module.

10. A computer storage medium, characterized in that: Stores a computer program, and the computer program is called by a processor to implement: The steps of the method according to any one of claims 1-8.