Script detection method and device, electronic equipment and storage medium

By building a script detection model based on a large language model, using syntax trees and semantic information to be processed, the problem of not identifying phishing websites that are not on the blacklist in the prior art is solved, and higher detection accuracy and generalization capabilities are achieved.

CN120238318APending Publication Date: 2025-07-01SANGFOR TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311790891.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing blacklist-based script detection methods cannot accurately identify phishing websites that are not on the blacklist, resulting in insufficient accuracy of phishing attack detection.

Method used

The script detection model is built based on a large language model. By obtaining the syntax tree of the script information to be detected, converting it into a token and adding semantic information, it uses the feature extraction module, the encoder module and the full connection layer for processing, which supports segmentation and prediction of longer script strings.

Benefits of technology

It improves the accuracy and generalization ability of detection scripts, can more effectively identify phishing websites, and enhances the detection accuracy of phishing attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238318A_ABST
    Figure CN120238318A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a script detection method and device, electronic equipment and a storage medium. The script detection method comprises the following steps: acquiring to-be-detected script information, wherein the to-be-detected script information comprises semantic information between scripts; the to-be-detected script information is input into a pre-trained script detection model, a prediction result corresponding to the to-be-detected script information output by the script detection model is obtained, and the script detection model is constructed based on a large language model. According to the method, the big language model has good generalization ability, so that the to-be-detected script information is processed through the script detection model constructed based on the big language model, and the to-be-detected script information can be predicted more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computers, and particularly relates to a script detection method, device, electronic device, and readable storage medium. Background Art

[0002] Phishing attacks have become one of the most serious threats faced by Internet users and service providers. Related script detection methods generally maintain a list of information about known phishing websites based on blacklist detection methods, so as to detect whether the currently accessed website is a phishing website according to the list. This method is simple in design and easy to implement. However, since this method is limited by the phishing websites recorded in the blacklist, when there is a phishing website not in the blacklist, the electronic device cannot accurately identify the phishing website. Therefore, the accuracy of detecting phishing websites needs to be improved. Summary of the Invention

[0003] In view of the above problems, this application proposes a script detection method, device, electronic device, and storage medium to improve the above problems.

[0004] In a first aspect, an embodiment of this application provides a script detection method, and the method includes: obtaining information of a script to be detected, where the information of the script to be detected includes a script and semantic information between scripts; inputting the information of the script to be detected into a pre-trained script detection model, and obtaining a prediction result corresponding to the string of the script to be detected output by the script detection model, where the script detection model is constructed based on a large language model.

[0005] Further, the step of inputting the information of the script to be detected into a pre-trained script detection model includes: converting the script to be detected into tokens; adding the semantic information to the tokens, and inputting the tokens after addition into a pre-trained script detection model.

[0006] Further, the step of obtaining the string of the script to be detected includes: obtaining a syntax tree corresponding to the script to be detected, where the syntax tree includes multiple nodes; determining a target node from the multiple nodes, where the target node is a node that includes at least one child node among the multiple nodes; adding a preset symbol to the character corresponding to the target node to obtain an updated target node; and determining the string of the script to be detected based on the character corresponding to the updated target node and the characters corresponding to the child nodes included in the updated target node.

[0007] Further, the script detection model includes a feature extraction module, an encoder module, and a fully connected layer; the step of inputting the script string to be detected into the pre-trained script detection model to obtain the prediction result corresponding to the script string to be detected output by the script detection model includes: determining whether the length of the script string to be detected is less than or equal to a preset length threshold; if it is determined that the length of the script string to be detected is less than or equal to the preset length threshold, inputting the script string to be detected into the feature extraction module to obtain the embedding vector of the script string to be detected output by the feature extraction module; inputting the embedding vector into the encoder module to obtain the encoded vector of the script string to be detected output by the encoder module; and inputting the encoded vector into the fully connected layer to obtain the prediction result output by the fully connected layer.

[0008] Further, the method further includes: if it is determined that the length of the script string to be detected is greater than the preset length threshold, splitting the script string to be detected based on the preset length threshold to obtain a plurality of split strings to be detected; inputting the plurality of split strings to be detected into the feature extraction module to obtain a plurality of embedded split vectors output by the feature extraction module, where one embedded split vector corresponds to one split string to be detected; inputting the plurality of embedded split vectors into the encoder module to obtain a plurality of encoded split vectors output by the encoder, where one encoded split vector corresponds to one embedded split vector; inputting the plurality of encoded split vectors into the fully connected layer to obtain a plurality of reference prediction results output by the fully connected layer, where one reference prediction result corresponds to one encoded split vector; and selecting the reference prediction result with the largest value among the plurality of reference prediction results as the prediction result corresponding to the script string to be detected. Through the above method, when it is determined that the length of the script string to be detected is greater than the preset length threshold, the script string to be detected is split to obtain a plurality of split strings to be detected, so that the processing of the script string to be processed can be completed by processing the plurality of split strings to be detected, thereby supporting the processing of longer strings and expanding the detection range of website scripts.

[0009] In a second aspect, an embodiment of the present application provides a model training method, the method includes: obtaining a training script data set, where the training data in the training data set includes scripts and semantic information between the scripts; training a script detection model according to the training script data set to obtain a trained script detection model, where the script detection model is constructed based on a large language model.

[0010] In a third aspect, an embodiment of the present application provides a script detection device, which includes: a string acquisition unit for acquiring information of a script to be detected, where the information of the script to be detected includes the script and the semantic information between scripts; a prediction result acquisition unit for inputting the information of the script to be detected into a pre-trained script detection model, and acquiring a prediction result corresponding to the string of the script to be detected output by the script detection model, where the script detection model is constructed based on a large language model.

[0011] In a fourth aspect, an embodiment of the present application provides a model training device, which includes: a data set acquisition unit for acquiring a training script data set, where the training script data in the training script data set includes the script and the semantic information between scripts; training the script detection model according to the training script data set to obtain a trained script detection model, where the script detection model is constructed based on a large language model.

[0012] In a fifth aspect, an embodiment of the present application provides an electronic device, including one or more processors and a memory; one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the above-mentioned method.

[0013] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored, where the above-mentioned method is executed when the program code runs.

[0014] An embodiment of the present application provides a script detection method, device, electronic device and storage medium. This script detection method detects the string of the script to be detected corresponding to the script to be detected through a script detection model constructed based on a large language model. Since the large language model itself has good generalization ability, processing the string of the script to be detected through the script detection model constructed based on the large language model can more accurately predict the string of the script to be detected. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 Shows a flowchart of a script detection method proposed in an embodiment of the present application;

[0017] Figure 2The flowchart of a script detection method proposed in another embodiment of the present application is shown;

[0018] Figure 3 The flowchart of a script detection method proposed in yet another embodiment of the present application is shown;

[0019] Figure 4 The schematic diagram of the abstract syntax tree of a script detection method proposed in an embodiment of the present application is shown;

[0020] Figure 5 The schematic diagram of the model of a script detection method proposed in an embodiment of the present application is shown;

[0021] Figure 6 The flowchart of a script detection method proposed in yet another embodiment of the present application is shown;

[0022] Figure 7 The flowchart of a model training method proposed in an embodiment of the present application is shown;

[0023] Figure 8 The structural block diagram of a script detection method proposed in an embodiment of the present application is shown;

[0024] Figure 9 The structural block diagram of a model training method proposed in an embodiment of the present application is shown;

[0025] Figure 10 The structural block diagram of the vehicle for executing the script detection method of the embodiments of the present application in real time of the present application is shown;

[0026] Figure 11 The storage unit for storing or carrying the program code for implementing the script detection method according to the embodiments of the present application in real time of the present application is shown. Detailed implementation manners

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0028] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0029] When a user uses an electronic device, they often accidentally log in to some phishing websites, which may lead to the user being scammed online and suffering economic losses. Therefore, it is necessary to detect whether a website is a phishing website.

[0030] The inventor found in the research on the detection methods of relevant phishing websites that the detection of relevant phishing websites generally maintains a list of information of known phishing websites based on the blacklist detection method, so as to detect whether the currently accessed website is a phishing website according to the list. This method is simple in design and easy to implement. However, since this method is limited by the phishing websites recorded in the blacklist, when there is a phishing website not in the blacklist, the electronic device cannot accurately identify the phishing website not in the blacklist.

[0031] Therefore, the inventor proposed the script detection method, device, electronic device and readable storage medium in the embodiments of this application. This script detection method includes: obtaining a script string to be detected, where the script string to be detected is obtained based on the syntax tree of the script to be detected; inputting the script string to be detected into a pre-trained script detection model, and obtaining the prediction result corresponding to the script string to be detected output by the script detection model, where the script detection model is constructed based on a large language model. Through the above method, since the large language model itself has good generalization ability, by processing the script string to be detected through the script detection model constructed based on the large language model, the script string to be detected can be predicted more accurately.

[0032] The following will specifically describe the embodiments of this application with reference to the drawings.

[0033] Please refer to Figure 1 , the embodiments of this application provide a script detection method, and the method includes:

[0034] Step S110: Obtain script information to be detected, where the script information to be detected includes semantic information between scripts.

[0035] In an embodiment of the present application, the form of the script information to be detected is a string, which is obtained based on the syntax tree of the script to be detected. When a user browses a website using an electronic device, the script detection application obtains the script to be detected corresponding to the website and sends the script to be detected to the background. In the background, the syntax tree corresponding to the script to be detected is determined, and the string corresponding to the script to be detected is obtained from the syntax tree, and this string is used as the script information to be detected. Herein, the syntax tree refers to a tree structure representing the syntax structure of the script to be detected; the script to be detected refers to the code that executes and controls the website to be detected; the string is obtained by processing the script to be detected through a preset algorithm, and the preset algorithm is an algorithm for processing the characters corresponding to the nodes in the syntax tree corresponding to the script to be detected.

[0036] Step S120: Input the script information to be detected into a pre-trained script detection model, and obtain the prediction result corresponding to the script information to be detected output by the script detection model, wherein the script detection model is constructed based on a large language model.

[0037] In an embodiment of the present application, the script detection model is a model for detecting the script to be detected to determine whether the script to be detected is a phishing script. The script detection model is constructed based on a large language model. When constructing the script detection model based on the large language model, the encoders in multiple Transformer layers in the large language model are selected and added to the script detection model, that is, the script detection model includes multiple encoders, and the output of the previous encoder layer in the multiple encoder layers is the input of the next encoder layer, and the output of the last encoder layer is used as the input of the fully connected layer. The script information to be detected is input into the pre-trained script detection model, and the multiple encoder layers and the fully connected layer in the script detection model process the script information to be detected to obtain the prediction result corresponding to the script information to be detected output by the script detection model.

[0038] A script detection method provided by an embodiment of the present application first obtains the script information to be detected, where the script information to be detected is obtained based on the syntax tree of the script to be detected, and then inputs the script information to be detected into a pre-trained script detection model to obtain the prediction result corresponding to the script information to be detected output by the script detection model, wherein the script detection model is constructed based on a large language model. Through the above method, since the large language model itself has good generalization ability, the script detection model constructed based on the large language model processes the script information to be detected, and can more accurately predict the script information to be detected.

[0039] Please refer to Figure 2 , an embodiment of the present application provides a script detection method, and the method includes:

[0040] Step S210: Obtain the script information to be detected, where the script information to be detected includes the semantics between scripts.

[0041] The specific implementation of step S210 can refer to the detailed explanation in the above embodiments, so it will not be elaborated in this embodiment.

[0042] Step S220: Convert the script information to be detected into tokens.

[0043] In the embodiment of the present application, since the script information to be detected is a string, the string is split to obtain multiple characters, and one character is used as a token, thereby completing the conversion of the script information to be detected.

[0044] Step S230: Add the semantic information to the token, and input the token with the added semantic information into a pre-trained script detection model.

[0045] In the embodiment of the present application, after obtaining the tokens, since each token corresponds to a character, the semantic information corresponding to the character is determined, and the semantic information corresponding to the character is added to the token, thereby adding semantic information to each token, and inputting the token with the added semantic information into a pre-trained script detection model.

[0046] Step S240: Obtain the prediction result corresponding to the script information to be detected output by the script detection model, where the script detection model is constructed based on a large language model.

[0047] The specific implementation of step S240 can refer to the detailed explanation in the above embodiments, so it will not be elaborated in this embodiment.

[0048] A script detection method provided by an embodiment of the present application first obtains the script information to be detected, where the script information to be detected includes the semantics between scripts, then converts the script information to be detected into tokens, then adds the semantic information to the tokens, inputs the tokens with the added semantic information into a pre-trained script detection model, and finally obtains the prediction result corresponding to the script information to be detected output by the script detection model, where the script detection model is constructed based on a large language model. Through the above method, since the large language model itself has good generalization ability, the script information to be detected is processed by the script detection model constructed based on the large language model, and the script information to be detected can be predicted more accurately.

[0049] Please refer to Figure 3 , an embodiment of the present application provides a script detection method, and the method includes:

[0050] Step S310: Obtain the syntax tree corresponding to the script to be detected, where the syntax tree includes multiple nodes.

[0051] In the embodiment of the present application, the syntax tree includes an Abstract Syntax Tree (AST) and a Concrete Syntax Tree (CST). For the same piece of code, when generating the concrete syntax tree of the code, the number of nodes included in the concrete syntax tree is more than the number of nodes in the abstract syntax tree. However, most of the information contained in these nodes is useless information, while the abstract syntax tree omits the intermediate redundant information and directly generates the semantics corresponding to the code. Therefore, the redundant information contained in the abstract syntax tree is much less than that in the concrete syntax tree. Therefore, when using the abstract syntax tree for subsequent detection, the detection effect is better than using the concrete syntax tree. Therefore, the abstract syntax tree is used in this solution. After obtaining the script to be detected, convert the script to be detected into an abstract syntax tree. In the abstract syntax tree, there are multiple nodes, and each node represents a structure in the code of the script to be tested.

[0052] Exemplarily, the abstract syntax tree of a piece of Python code can be as Figure 4 shown. For "def mean(data);" in this piece of code, this piece of code includes seven nodes in the abstract syntax tree, and the characters corresponding to these seven nodes are: "def", "mean", "parameters", ";", "(", "data", and ")". Among them, the nodes corresponding to "(", "data", and ")" are the child nodes of the node corresponding to "parameters".

[0053] Step S320: Determine a target node from the multiple nodes, where the target node is a node that includes at least one child node among the multiple nodes.

[0054] In the embodiment of the present application, after obtaining the abstract syntax tree corresponding to the script to be detected, first, the root node in the abstract syntax tree corresponding to the script to be detected can be obtained by calling the tree_sitter third-party library, and then the abstract syntax tree is traversed through a preset algorithm to determine the nodes that have at least one child node among the multiple nodes in the abstract syntax tree, and the nodes that have at least one child node are used as the target nodes of the abstract syntax tree.

[0055] Among them, for the preset algorithm, its working principle is to traverse each node in the abstract syntax tree corresponding to the script to be detected, find a node including at least one child node from multiple nodes in the abstract syntax tree, use the node with at least one child node as the target node, and add a preset symbol to the character corresponding to the target node.

[0056] Step S330: Add a preset symbol to the character corresponding to the target node to obtain an updated target node.

[0057] In the embodiment of the present application, after determining the target node in the abstract syntax tree, a preset symbol is added to the character corresponding to the target node according to the foregoing preset algorithm. In this solution, the added preset characters are left and right. After adding the preset symbol to the character of the target node, an updated target node is obtained.

[0058] Step S340: Determine the script information to be detected based on the character corresponding to the updated target node and the characters corresponding to the child nodes included in the updated target node.

[0059] In the embodiment of the present application, since multiple encoder layers in the script detection model cannot be well compatible with the tree structure and the syntax tree cannot be directly input into the script detection model for processing, it is necessary to convert the corresponding abstract syntax tree into a string without losing semantic information. Each updated target node included in the abstract syntax tree and each child node included in each updated target node are used as reference nodes, and the nodes other than the reference nodes in the abstract syntax tree are used as non-reference nodes. The positions of the reference nodes and non-reference nodes in the abstract syntax tree are determined respectively. According to the positions of the reference nodes and non-reference nodes in the abstract syntax tree, the characters corresponding to the reference nodes and the characters corresponding to the non-reference nodes are sorted in sequence, so as to obtain the string corresponding to the script to be detected, and this string is used as the script information to be detected. For example, in the abstract syntax tree, the character corresponding to the target node is argument-list, and the characters corresponding to the child nodes under this target node are (, data, and ), respectively. After adding the preset character to the character corresponding to the target node, the character corresponding to the target node with the preset character added, and the characters corresponding to the child nodes included in the target node are displayed in the script information to be detected as: <argument_list.left>(data)<argument_list.right>.

[0060] Step S350: Determine whether the length of the script information to be detected is less than or equal to a preset length threshold.

[0061] In an embodiment of the present application, after obtaining the script information to be detected, the length of the string corresponding to the script information to be detected is determined. Since the maximum length of the string supported by the Transformer layer is 1024, generally, the preset length can be default set to 1024, and it is determined whether the length of the string to be detected is less than or equal to 1024.

[0062] As a way, the preset length threshold can also be set by the user to a value smaller than 1024. However, due to the situation where the two pieces of attacking code are far apart, if the preset length is divided into smaller values, it is possible that these two pieces of attacking code are in different strings, resulting in an inaccurate prediction result for the script information to be detected. Therefore, generally, setting the preset length threshold to a smaller value is not considered.

[0063] Step S360: If it is determined that the length of the script information to be detected is less than or equal to the preset length threshold, input the script information to be detected into the feature extraction module, and obtain the embedding vector of the script information to be detected output by the feature extraction module.

[0064] In an embodiment of the present application, if it is determined that the length of the string corresponding to the script information to be detected is less than or equal to the preset length threshold, it can be determined that the Transformer can support the processing of the script information to be detected, that is, it indicates that multiple encoder layers support the processing of the script information to be detected. Therefore, the script information to be detected is input into the feature extraction module, and in the feature extraction module, the embedding vector corresponding to the script information to be detected is output by the feature extraction module. Among them, the embedding vector is composed of character vectors corresponding to each script character in the script information to be detected, and the character vector corresponding to each script character is composed of the word vector and position vector corresponding to the script character.

[0065] Step S370: Input the embedding vector into the encoder module, and obtain the encoded vector of the script information to be detected output by the encoder module.

[0066] In the embodiments of the present application, in order to increase the complexity of the model, improve the understanding ability and generalization ability of the model, and improve the interpretation effect of the model on complex scripts, the encoder module of this solution is composed of multiple encoder layers in multiple Transformer layers. Each encoder layer is connected end to end in sequence. The first encoder layer takes the embedding vector output by the feature extraction module as input, the output of the previous encoder layer is the input of the next encoder layer, and the output of the last encoder layer is used as the encoded vector of the script information to be detected. Each encoder layer is mainly composed of two parts of networks: a network layer that implements the multi-head self-attention mechanism and a two-layer feed-forward neural network. At the same time, for these two parts of networks, residual connections are added, and layer normalization operations are also performed after the residual connections. The embedding vector output by the feature extraction module is input into the encoder layer module. First, the first encoder layer in the encoder layer module processes the embedding vector, calculates the embedding vector through the multi-head self-attention mechanism, and processes the calculated embedding vector through the feed-forward neural network, residual connection, and layer normalization operations included in the first encoder layer, so as to obtain the first reference encoded vector output by the first encoder layer, and input the first reference encoded vector into the second encoder layer for processing. The processing method is as above, and the above process is repeated until the encoded vector corresponding to the script information to be detected is output by the last encoder layer in the encoder layer module.

[0067] Exemplarily, for the encoder layer i in the encoder layer module, if the vector output by the encoder layer i-1 before the encoder layer i is the reference encoded vector H i-1 , since the encoder layer i includes multiple head self-attention mechanisms, for each head, first based on the encoded vector H i-1 extract the three elements of the attention mechanism: Q (query), K (Key), and V (Value). The specific calculation formulas are: Q = H i-1 W Q , K = H i-1 W K , V = H i-1 W V . Then, through the formula V calculation, the calculation result of one head is obtained. Repeating the above calculation, the calculation results of multiple heads can be obtained. Through the formula Multihead(Q, K, V) = Concat(head1, head2,... head n )W o calculation, the multi-head attention vector output by the multi-head self-attention mechanism included in the encoder layer i can be obtained, and by processing the multi-head attention vector through residual connection, layer normalization, and feed-forward neural network, the reference encoded vector output by the encoder layer i can be output.

[0068] Step S380: Input the encoded vector into the fully connected layer, and obtain the prediction result output by the fully connected layer.

[0069] In the embodiment of the present application, the encoded vector is input into the fully connected layer through the encoder module. When the fully connected layer receives the encoded vector, the encoded vector is processed through a normalization function to obtain the probability distribution corresponding to the encoded vector, and the cross-entropy loss function is used to calculate this probability distribution to obtain the prediction result corresponding to the script information to be detected, and the fully connected layer outputs this prediction result.

[0070] As a way, if it is determined according to the prediction result that the script to be detected is a phishing script, it can be determined that the website currently browsed by the user is a phishing website, and a pop-up window can be issued to remind the user.

[0071] Exemplarily, steps S310 - S380 can be as Figure 5 shown. After obtaining the abstract syntax tree corresponding to the script to be detected, convert the abstract syntax tree into the script information to be detected. When the length of the script information to be detected is less than or equal to the preset length threshold, in the feature extraction module, determine the embedding vector corresponding to the script information to be detected through word embedding and position embedding, and input the embedding vector into the encoder module. The multiple encoder layers in the encoder module process the embedding vector to obtain the encoded vector output by the encoder module, and input the encoded vector into the fully connected layer. The softmax function and cross-entropy loss function in the fully connected layer process the encoded vector, and output to obtain the prediction result corresponding to the script to be detected.

[0072] A script detection method provided by an embodiment of the present application, by converting the script to be detected into the script information to be detected. When it is determined that the length of the script information to be detected is less than or equal to the preset length threshold, input the script information to be detected into the feature extraction module in the script detection model to obtain the embedding vector of the script information to be detected output by the feature extraction module, then input the embedding vector into the encoder module to obtain the encoded vector of the script information to be detected output by the encoder module, and finally input the encoded vector into the fully connected layer to obtain the prediction result output by the fully connected layer, so as to determine whether the script to be detected is a phishing script. Since the large language model itself has good generalization ability, by processing the script information to be detected through the large language model, the script information to be detected can be predicted more accurately.

[0073] Please refer to Figure 6 , an embodiment of the present application provides a script detection method, the method includes:

[0074] Step S410: Obtain the syntax tree corresponding to the script to be detected, and the syntax tree includes multiple nodes.

[0075] Determine a target node from the multiple nodes, where the target node is a node among the multiple nodes that includes at least one child node.

[0076] Step S420: Add a preset symbol to the character corresponding to the target node to obtain an updated target node.

[0077] Step S430: Determine the script information to be detected based on the character corresponding to the updated target node and the characters corresponding to the child nodes included in the updated target node.

[0078] The specific details of steps S410 - S430 can refer to the detailed explanations in the above embodiments, so they will not be elaborated in this embodiment.

[0079] Step S440: If it is determined that the length of the script information to be detected is greater than the preset length threshold, segment the script information to be detected based on the preset length threshold to obtain a plurality of segmented strings to be detected.

[0080] In the embodiment of the present application, if it is determined that the length of the script information to be detected is greater than the preset length threshold, it means that multiple encoder layers do not support processing the script information to be detected with the existing length. Therefore, it is necessary to segment the string corresponding to the script information to be detected, and the multiple encoder layers process the segmented script information to be detected, so as to complete the processing of the script information to be detected. According to step S250, the length of the segmentation of the script information to be detected cannot be too short, otherwise it will affect the accuracy of the prediction result of the script information to be detected. Therefore, in this solution, the script information to be detected is segmented based on the preset length threshold to obtain a plurality of segmented strings to be detected, and the multiple encoder layers process the plurality of segmented strings to be detected respectively, so as to complete the processing of the script information to be detected. Among them, since the maximum length of the string supported by Transformer is 1024, the maximum length of the string supported by the encoder layer is also 1024, that is, the preset length threshold is 1024. Therefore, the script information to be detected is segmented according to 1024.

[0081] Step S450: Input the plurality of segmented strings to be detected into the feature extraction module, and obtain a plurality of embedded segmented vectors output by the feature extraction module, where one embedded segmented vector corresponds to one segmented string to be detected.

[0082] In an embodiment of the present application, a plurality of to-be-detected segmented strings are input into a feature extraction module, and the feature extraction module outputs a plurality of embedded segmented vectors corresponding to the plurality of to-be-detected segmented strings. Among them, one embedded segmented vector corresponds to one to-be-detected segmented string, and each embedded segmented vector is composed of character vectors corresponding to each script character in the corresponding to-be-detected segmented string, and the character vector corresponding to each script character is composed of the word vector and the position vector corresponding to the script character.

[0083] Specifically, the processing process of the feature extraction module for a plurality of to-be-detected segmented strings is as follows: The plurality of to-be-detected segmented strings are input into the feature extraction module. For each to-be-detected segmented string, the feature extraction module converts the plurality of script characters included in the to-be-detected segmented string into their respective corresponding word vectors through token embedding (word embedding). The process of converting script characters into word vectors is generally completed through a pre-trained word vector model (such as Word2Vec, GloVe, etc.). In this solution, the word vector model is placed in the feature extraction module. The feature extraction module converts the plurality of script characters included in the to-be-detected segmented string into their respective corresponding position vectors through position embedding (position embedding). The position vector has the same dimension as the word vector. Add the word vector and the position vector corresponding to each script character to obtain the character vector corresponding to each script character, and stack the character vectors corresponding to each script character to obtain the embedded segmented vector corresponding to each to-be-detected segmented string. The feature extraction module outputs a plurality of embedded segmented vectors corresponding to the plurality of to-be-detected segmented strings.

[0084] Step S460: Input the plurality of embedded segmented vectors into the encoder module, and obtain a plurality of encoded segmented vectors output by the encoder, where one encoded segmented vector corresponds to one embedded segmented vector.

[0085] In an embodiment of the present application, the plurality of embedded segmented vectors output by the feature extraction module are input into the encoder module. First, the first encoder layer in the encoder module processes each embedded segmented vector, calculates each embedded segmented vector through a multi-head self-attention mechanism, and processes each calculated embedded segmented vector through the feed-forward neural network, residual connection, and layer normalization operations included in the first encoder layer to obtain a plurality of first reference encoded segmented vectors output by the first encoder layer, and input the plurality of first reference encoded segmented vectors into the second encoder layer for processing. The processing method is as above. Repeat the above process until the last encoder layer in the encoder module outputs a plurality of encoded segmented vectors corresponding to the plurality of embedded segmented vectors.

[0086] Step S470: Input the multiple encoded split vectors into the fully connected layer, and obtain multiple reference prediction results output by the fully connected layer, where one reference prediction result corresponds to one encoded split vector.

[0087] In an embodiment of the present application, the encoder module inputs multiple encoded split vectors into the fully connected layer. When the fully connected layer receives the multiple encoded split vectors, the normalization function is used to process the multiple encoded split vectors respectively to obtain multiple normalized split vectors, and the cross-entropy loss function is used to calculate the multiple normalized split vectors respectively to obtain multiple reference prediction results corresponding to the multiple to-be-detected split strings, and the fully connected layer outputs the multiple reference prediction results.

[0088] Step S480: Among the multiple reference prediction results, select the reference prediction result with the largest value as the prediction result corresponding to the to-be-detected script information.

[0089] In an embodiment of the present application, after the fully connected layer outputs multiple reference prediction results, compare the numerical values of each reference prediction result, and select the reference prediction result with the largest value among the multiple reference prediction results as the prediction result corresponding to the to-be-detected script information, and judge the probability that the to-be-detected script is a phishing script according to this prediction result.

[0090] A script detection method provided by an embodiment of the present application first converts the to-be-detected script into to-be-detected script information. When it is determined that the length of the to-be-detected script information is greater than a preset length threshold, the to-be-detected script information is segmented to obtain multiple to-be-detected split strings, the multiple to-be-detected split strings are input into the feature extraction module in the script detection model to obtain multiple embedded split vectors of the multiple to-be-detected split strings output by the feature extraction module, the multiple embedded split vectors are input into the encoder module to obtain multiple encoded split vectors corresponding to the multiple to-be-detected split strings, and finally the multiple encoded split vectors are input into the fully connected layer to obtain multiple reference prediction results output by the fully connected layer, and select the reference prediction result with the largest value among the multiple reference prediction results as the prediction result corresponding to the to-be-detected script information, so as to determine whether the to-be-detected script is a phishing script. Since the large language model itself has good generalization ability, by using the large language model to process the to-be-detected script information, the to-be-detected script information can be predicted more accurately.

[0091] Please refer to Figure 7 , an embodiment of the present application provides a model training method, and the method includes:

[0092] Step S510: Obtain a training script data set, and the training data in the training script data set includes scripts and semantic information between the scripts.

[0093] In the application embodiment, the engineer uses multiple marked script data as multiple training data, and uses the set of the multiple training data as a training script data set. Among the multiple training data included in the training script data set, there are positive sample script data and negative sample script data. The positive sample script data is the script related to a normal website and the semantic information between scripts, and the negative sample script data is the script related to a phishing website and the semantic information between scripts.

[0094] Step S520: Train a script detection model according to the training script data set to obtain a trained script detection model, where the script detection model is constructed based on a large language model.

[0095] In the application embodiment, the training script data set is input into the script detection model. In the script detection model, the embedding vector corresponding to the positive sample script data is determined through token embedding and position embedding, and the embedding vector corresponding to the positive sample script data is processed through multiple encoder layers included in multiple Transformers, so as to obtain the encoded vector corresponding to the positive sample script data. Then, the encoded vector corresponding to the positive sample script data is processed through a normalization function to obtain the probability distribution corresponding to the positive sample script data. Similarly, in the script detection model, the embedding vector corresponding to the negative sample script data is determined through token embedding and position embedding, and the embedding vector corresponding to the negative sample script data is processed through multiple encoder layers included in multiple Transformers, so as to obtain the encoded vector corresponding to the negative sample script data. Then, the encoded vector corresponding to the negative sample script data is processed through a normalization function to obtain the probability distribution corresponding to the negative sample script data. After obtaining the probability distributions corresponding to the positive sample script data and the negative sample script data respectively, the loss function value is obtained through the cross-entropy loss function. In the process of iteratively training the script detection model, the loss function value will gradually decrease. When the loss function value drops to a certain extent, it can be determined that the model training is completed, and a trained script detection model is obtained. Exemplarily, when the loss function value drops to 0.05, it is determined that the model training is completed, and a trained script detection model is obtained. Among them, in the process of training the script detection model, the parameters included in the reference hidden layer state of the script detection model are continuously optimized. When the script detection model training is completed, the parameters included in the reference hidden layer state are optimized, and the reference hidden layer state with the optimized parameters is used as the hidden layer state.

[0096] As a way, the end condition of model training can also be that the number of iterative trainings of the model reaches a preset number. For example, the preset number of times for the script detection model to end training is set to 10 times, that is, when the script detection model performs 10 iterative trainings, it can be determined that the model training is completed.

[0097] A model training method provided by an embodiment of this application first obtains a training script data set. The training data in the training data set includes scripts and semantic information between the scripts. Then, the script detection model is trained according to the training script data set to obtain a trained script detection model, where the script detection model is constructed based on a large language model.

[0098] Please refer to Figure 8 , an embodiment of this application provides a script detection device 600. The device 600 includes:

[0099] A string acquisition unit 610, configured to acquire a script string to be detected, where the script string to be detected is obtained based on the syntax tree of the script to be detected.

[0100] As a way, the string acquisition unit 410 is further configured to acquire the syntax tree corresponding to the script to be detected. The syntax tree includes multiple nodes; determine a target node from the multiple nodes, where the target node is a node among the multiple nodes that includes at least one child node; add a preset symbol to the character corresponding to the target node to obtain an updated target node; and determine the script string to be detected based on the character corresponding to the updated target node and the characters corresponding to the child nodes included in the updated target node.

[0101] A prediction result acquisition unit 420, configured to input the script string to be detected into a pre-trained script detection model, and acquire a prediction result corresponding to the script string to be detected output by the script detection model, where the script detection model is constructed based on a large language model.

[0102] As a way, the prediction result acquisition unit 420 is further configured to determine whether the length of the script string to be detected is less than or equal to a preset length threshold; if it is determined that the length of the script string to be detected is less than or equal to the preset length threshold, input the script string to be detected into a feature extraction module to acquire an embedding vector of the script string to be detected output by the feature extraction module; input the embedding vector into the encoder module to acquire a coding vector of the script string to be detected output by the encoder module; and input the coding vector into the fully connected layer to acquire a prediction result output by the fully connected layer.

[0103] Optionally, the prediction result obtaining unit 420 is further configured to, if it is determined that the length of the to-be-detected script string is greater than the preset length threshold, segment the to-be-detected script string based on the preset length threshold to obtain a plurality of to-be-detected segmented strings; input the plurality of to-be-detected segmented strings into the feature extraction module, and obtain a plurality of embedded segmented vectors output by the feature extraction module, where one embedded segmented vector corresponds to one to-be-detected segmented string; input the plurality of embedded segmented vectors into the encoder module, and obtain a plurality of encoded segmented vectors output by the encoder, where one encoded segmented vector corresponds to one embedded segmented vector; input the plurality of encoded segmented vectors into the fully connected layer, and obtain a plurality of reference prediction results output by the fully connected layer, where one reference prediction result corresponds to one encoded segmented vector.

[0104] Optionally, the prediction result obtaining unit 420 is further configured to input the plurality of to-be-detected segmented strings into the feature extraction module, and obtain a plurality of embedded segmented vectors output by the feature extraction module, where each embedded segmented vector is a vector obtained by the feature extraction module by superimposing character vectors corresponding to each script character based on the position of each script character included in each to-be-detected segmented string in the to-be-detected script, and each character vector is a vector obtained by adding the word vector and the position vector corresponding to each script character included in each to-be-detected segmented string.

[0105] Optionally, the prediction result obtaining unit 420 is further configured to select the reference prediction result with the largest value among the plurality of reference prediction results as the prediction result corresponding to the to-be-detected script string.

[0106] Please refer to Figure 9 , an embodiment of the present application provides a model training device 700, and the device 700 includes:

[0107] A dataset obtaining unit 710, configured to obtain a training script dataset, where the training data in the training script dataset includes scripts and semantic information between the scripts;

[0108] A model obtaining unit 720, configured to train a script detection model according to the training script dataset to obtain a trained script detection model, where the script detection model is constructed based on a large language model.

[0109] Next, an electronic device provided by the present application will be described in conjunction with Figure 10 an electronic device provided by the present application will be described.

[0110] Please refer to Figure 10, based on the above-mentioned emotion generation method and device, an embodiment of the present application further provides an electronic device 500 that can execute the foregoing emotion generation method. The electronic device 500 includes one or more (only one is shown in the figure) processors 502, a memory 504, and a network module 506 that are coupled to each other. Among them, the memory 504 stores a program that can execute the content in the foregoing embodiments, and the processor 502 can execute the program stored in the memory 504.

[0111] Among them, the processor 502 may include one or more processing cores. The processor 502 uses various interfaces and lines to connect various parts within the entire electronic device 500, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 504, and by calling data stored in the memory 504, it executes various functions of the server 500 and processes data. Optionally, the processor 502 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 502 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 502 and may be implemented separately by a communication chip.

[0112] The memory 504 may include random access memory (RAM), and may also include read-only memory. The memory 504 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 504 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the electronic device 500 (such as phone books, audio and video data, chat record data, etc.).

[0113] The network module 506 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices, such as communicating with an audio playback device. The network module 506 may include various existing circuit elements for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The network module 506 can communicate with various networks such as the Internet, enterprise intranets, wireless networks or communicate with other devices through a wireless network. The above-mentioned wireless network may include a cellular phone network, a wireless local area network or a metropolitan area network. For example, the network module 506 can interact with a base station.

[0114] Please refer to Figure 11 , which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable medium 600, and the program code can be called by a processor to execute the method described in the above method embodiment.

[0115] The computer-readable storage medium 600 may be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk or a ROM. Optionally, the computer-readable storage medium 600 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 600 has a storage space for the program code 610 for executing any method step in the above method. These program codes can be read out from or written into one or more computer program products. The program code 610 can be compressed in an appropriate form, for example.

[0116] An embodiment of the present application provides a script detection method, device, electronic device and storage medium. The script detection method includes: obtaining script information to be detected, where the script information to be detected includes semantic information between scripts; inputting the script information to be detected into a pre-trained script detection model, and obtaining a prediction result corresponding to the script information to be detected output by the script detection model, where the script detection model is constructed based on a large language model. Through the above method, since the large language model itself has good generalization ability, the script information to be detected is processed by the script detection model constructed based on the large language model, and the script information to be detected can be predicted more accurately.

[0117] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims, and all of them belong to the protection scope of the present invention.

Claims

1. A script detection method, characterized in that, The method includes: Obtaining the script information to be detected, where the script information to be detected includes the semantics between scripts; Inputting the script information to be detected into a pre-trained script detection model, and obtaining a prediction result corresponding to the script information to be detected output by the script detection model, where the script detection model is constructed based on a large language model.

2. The method according to claim 1, wherein The step of inputting the script information to be detected into a pre-trained script detection model includes: Converting the script information to be detected into tokens; Adding the semantic information to the tokens, and inputting the tokens with added semantic information into a pre-trained script detection model.

3. The method according to claim 1, wherein The step of obtaining the script information to be detected includes: Obtaining a syntax tree corresponding to the script to be detected, where the syntax tree includes multiple nodes; Determining a target node from the multiple nodes, where the target node is a node that includes at least one child node among the multiple nodes; Adding a preset symbol to the characters corresponding to the target node to obtain an updated target node; Determining the script information to be detected based on the characters corresponding to the updated target node and the characters corresponding to the child nodes included in the updated target node.

4. The method according to claim 1, wherein The script detection model includes a feature extraction module, an encoder module, and a fully connected layer; the step of inputting the script information to be detected into a pre-trained script detection model and obtaining a prediction result corresponding to the script information to be detected output by the script detection model includes: judging whether the length of the script information to be detected is less than or equal to a preset length threshold; If it is determined that the length of the script information to be detected is less than or equal to the preset length threshold, inputting the script information to be detected into the feature extraction module, and obtaining an embedding vector of the script information to be detected output by the feature extraction module; Inputting the embedding vector into the encoder module, and obtaining an encoded vector of the script string to be detected output by the encoder module; Inputting the encoded vector into the fully connected layer, and obtaining the prediction result output by the fully connected layer.

5. The method according to claim 4, wherein The method further includes: If it is determined that the length of the script information to be detected is greater than the preset length threshold, splitting the script information to be detected based on the preset length threshold to obtain multiple split strings to be detected; Inputting the multiple split strings to be detected into the feature extraction module, and obtaining multiple embedding split vectors output by the feature extraction module, where one embedding split vector corresponds to one split string to be detected; Inputting the multiple embedding split vectors into the encoder module, and obtaining multiple encoded split vectors output by the encoder, where one encoded split vector corresponds to one embedding split vector; Inputting the multiple encoded split vectors into the fully connected layer, and obtaining multiple reference prediction results output by the fully connected layer, where one reference prediction result corresponds to one encoded split vector; Among the multiple reference prediction results, selecting the reference prediction result with the largest value as the prediction result corresponding to the script string to be detected.

6. A model training method, characterized in that, The method includes: Obtain a training script dataset, where the training data in the training script dataset includes scripts and semantic information between the scripts; Train a script detection model according to the training script dataset to obtain a trained script detection model, where the script detection model is constructed based on a large language model.

7. A script detection device, characterized in that, The device includes: A string acquisition unit, configured to acquire script information to be detected, where the script information to be detected includes scripts and semantic information between the scripts; A prediction result acquisition unit, configured to input the script information to be detected into a pre-trained script detection model, and obtain a prediction result corresponding to the script string to be detected output by the script detection model, where the script detection model is constructed based on a large language model.

8. A model training device, characterized in that, The device includes: A dataset acquisition unit, configured to acquire a training script dataset, where the training data in the training script dataset includes scripts and semantic information between the scripts; A model acquisition unit, configured to train a script detection model according to the training script dataset to obtain a trained script detection model, where the script detection model is constructed based on a large language model.

9. An electronic device, characterized in that, Includes one or more processors and a memory, and one or more programs are stored in the memory and configured to be executed by one or more processors to perform the method according to any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code, and the program code includes instructions for executing the method according to any one of claims 1-6.