Virus protection method based on host security, computer device

By implementing a host security-based virus protection method in the training stage of neural network model, generating risk description information and distributing training tasks to distributed nodes, the malicious tampering and operation problems of super-large parameter models in the training stage are solved, and the improvement of model quality and resource utilization is achieved.

CN119442237BActive Publication Date: 2025-05-30BEIJING EASYNETWORKS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510024952.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-30
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Neural network models with super large parameters are susceptible to malicious sample tampering and malicious operations during the training stage, resulting in slowing down the convergence speed of the model, decreasing quality, and wasting computing resources.

Method used

Through a virus protection method based on host security, operational risk description information and sample risk description information are generated, model training tasks are distributed to distributed nodes, and training samples are subjected to tamper-proof verification to ensure the safety and efficiency of model training.

Benefits of technology

It effectively protects the neural network model in the model training stage, ensures the training quality and effective utilization of computing resources, and avoids model pollution and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119442237B_ABST
    Figure CN119442237B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a virus protection method and a computer device based on host security. A specific implementation manner of the method includes: distributing a model training task corresponding to training task information to a first node; generating operation risk description information for the task modification operation in response to monitoring a task modification operation for the first node; performing anti-tampering verification of training sample information on the training sample information according to the training sample verification information in response to receiving training sample information and training sample verification information sent by a second node respectively; and performing protection security for the model training task corresponding to the training task information according to the operation risk description information and / or sample risk description information. This implementation manner effectively protects the neural network model in the training stage, thereby ensuring the training quality and avoiding waste of computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technology, and more particularly to a virus protection method based on host security and a computer device. Background Art

[0002] A neural network model is a machine learning model that mimics the processing and learning mechanisms of the human nervous system, especially the way neurons work. It can analyze and process data to achieve tasks such as prediction, classification, and recognition. Currently, with the rise of models such as GPT (Generative Pre-trained Transformer), neural network models with extremely large numbers of parameters have become the mainstream research trend. Currently, the supervised model training method, as one of the important training methods for neural network models, is still widely used.

[0003] However, due to the characteristic of extremely large numbers of parameters in neural network models with extremely large numbers of parameters, it is necessary to rely on a large amount of computer resources for model training during the model training stage. During the training stage, when malicious sample tampering or malicious operations occur, it will, at the least, affect the model convergence speed, and at the worst, affect the quality of the trained model, and at the same time, it will also cause a waste of a large amount of computer resources. Summary of the Invention

[0004] This section of the present application is used to briefly introduce concepts that will be described in detail in the subsequent Detailed Description section. This section of the present application is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] Some embodiments of the present application propose a virus protection method based on host security and a computer device to solve one or more of the technical problems mentioned in the above Background Art section.

[0006] In a first aspect, some embodiments of the present application provide a virus protection method based on host security. The method includes: creating a model training task for the target neural network model according to the model description file corresponding to the target neural network model to obtain a training task information set, where the target neural network model is a neural network model to be trained, the model description file includes a model structure description file and a model training parameter description file, and the training task information in the training task information set includes task description information, node address, training mode, and training sample address; for each training task information in the training task information set, perform the following processing steps: distributing the model training task corresponding to the training task information to a first node according to the training task information, where the first node is a distributed node for executing the model training task corresponding to the training task information; in response to monitoring a task modification operation for the first node, generating operation risk description information for the task modification operation, where the operation risk description information includes operation type, operation risk confidence level, operation anomaly confidence level, operation initiation source credibility, and virus confidence level, and the virus confidence level represents the confidence level that the task modification operation is a host virus; in response to receiving training sample information and training sample verification information sent by a second node respectively, performing a training sample anti-tampering check on the training sample information according to the training sample verification information to generate sample risk description information, where the second node is the distributed node corresponding to the training sample address included in the training task information, the training sample information includes a training sample and a sample label, and the training sample is an image type training sample; performing security protection for the model training task corresponding to the training task information according to the operation risk description information and / or the sample risk description information.

[0007] In a second aspect, the present application further provides a computer device. The computer device includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the method described in any implementation manner of the first aspect is implemented.

[0008] In a third aspect, the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0009] The above embodiments of the present application have the following beneficial effects: Through the virus protection method based on host security in some embodiments of the present application, effective protection during the model training stage is achieved, avoiding problems such as the model being contaminated, which affects the quality of the trained model, and waste of computing resources. Specifically, the reasons for the above problems are as follows: Due to the characteristic of a neural network model with an extremely large number of parameters, a large amount of computer resources are required for model training during the model training stage. During the training stage, when malicious sample tampering or malicious operations occur, it will, at the least, affect the model convergence speed, and at the most, affect the quality of the trained model. At the same time, it will also cause a waste of a large amount of computer resources. Based on this, in some embodiments of the virus protection method based on host security of the present application, first, according to the model description file corresponding to the target neural network model, a model training task for the above target neural network model is created to obtain a training task information set. Among them, the above target neural network model is the neural network model to be trained, and the above model description file includes: a model structure description file and a model training parameter description file. The training task information in the above training task information set includes: task description information, node address, training mode, and training sample address. For a neural network model with an extremely large number of parameters, it is difficult to match the corresponding training requirements using a single-machine training method. Therefore, a distributed method is usually used for model training. Therefore, through task decomposition, the entire model training task can be effectively disassembled into training tasks processed by nodes in different distributions, thereby improving the training efficiency. Secondly, for each training task information in the above training task information set, the following processing steps are performed: The first step is to distribute the model training task corresponding to the above training task information to the first node according to the above training task information, where the above first node is a distributed node used to execute the model training task corresponding to the above training task information. The second step is to generate operation risk description information for the above task modification operation in response to monitoring the task modification operation for the above first node. Among them, the above operation risk description information includes: operation type, operation risk confidence level, operation anomaly confidence level, operation initiation source credibility, virus confidence level, where the virus confidence level represents the confidence level that the above task modification operation is a host virus. In order to improve the training speed, it is often necessary to adaptively adjust the task according to the current training state. During this period, there may be invasive malicious task modifications or non-invasive, internally initiated malicious task modifications. Therefore, by generating corresponding operation risk description information for each generated task modification operation, the risk of the task modification operation can be evaluated.In the third step, in response to receiving the training sample information and the training sample verification information sent by the second node respectively, according to the above training sample verification information, perform anti-tampering verification on the above training sample information to generate sample risk description information, where the second node is the distributed node corresponding to the training sample address included in the above training task information, and the above training sample information includes: training samples and sample labels, and the training samples are image-type training samples. In practice, for a supervised training method, training samples and sample labels are crucial for the training accuracy of the model. However, during the training process, through the way of sample contamination, it is extremely easy to cause the training accuracy of the model to be damaged, and even the situation of non-convergence may occur. Therefore, it is necessary to perform corresponding verification on whether the training samples are tampered with to obtain corresponding sample risk description information. Finally, perform security protection for the model training task corresponding to the above training task information according to the above operation risk description information and / or the above sample risk description information. In this way, the neural network model in the training stage is effectively protected, thereby ensuring the training quality and avoiding the waste of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present application will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0011] Figure 1 is a flowchart of some embodiments of the virus protection method based on host security according to the present application;

[0012] Figure 2 is a schematic diagram of a model structure;

[0013] Figure 3 is a schematic diagram of the distribution process of distributing the model training task corresponding to the above training task information to the first node;

[0014] Figure 4 is a schematic diagram of the structure of a computer device suitable for implementing some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.

[0016] In addition, it should be noted that for ease of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0017] It should be noted that the concepts such as "first" and "second" mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0018] It should be noted that the modifiers "one" and "multiple" mentioned in the present application are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".

[0019] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0020] The present application will be described in detail below with reference to the drawings and in conjunction with embodiments.

[0021] Reference Figure 1 , a flow 100 of some embodiments of a virus protection method based on host security according to the present application is shown. The virus protection method based on host security includes the following steps:

[0022] Step 101, create a model training task for the target neural network model according to the model description file corresponding to the target neural network model, and obtain a training task information set.

[0023] In some embodiments, the execution subject (for example, a computer device) of the virus protection method based on host security can create a model training task for the above-mentioned target neural network model according to the model description file corresponding to the target neural network model, and obtain a training task information set.

[0024] Among them, the above-mentioned target neural network model is a neural network model to be trained. In practice, the above-mentioned target neural network model is trained using a supervised model training method. The model description file includes: a model structure description file and a model training parameter description file.

[0025] The model structure description file is a file used to describe the model structure of the target neural network model. In practice, the model structure description file may include: the type of layer, the number of layers, the connection order of layers, the parameters of layers, and the input and output sizes of layers. For example, the type of layer may include: convolutional layer, fully connected layer, input layer, output layer, flattening layer, etc. The parameters of layers may include: the input size of the layer, the output size of the layer, the convolutional kernel size, the stride, the activation function, the padding method, etc.

[0026] The model training parameter description information is used to describe the training parameters of the target neural network model during the model training phase. For example, the model training parameter description information may include: learning rate, batch size, number of iterations, regularization parameter, loss function, and optimizer. Among them, the learning rate is a parameter used to describe the step size during each weight update. In practice, a smaller learning rate may increase the model training time, and a larger learning rate may cause oscillations, resulting in the model being unable to converge. The batch size refers to the number of training samples used in each gradient update process. In practice, a large batch can use more training samples in a single gradient update process, but it will occupy more memory. The number of iterations refers to the number of times the entire dataset corresponding to the training samples is completely traversed. In practice, when the number of iterations is set too large, it may cause overfitting of the model, and when the number of iterations is set too small, it may cause underfitting of the model. The regularization parameter is used to prevent overfitting and to a certain extent improve the generalization ability of the model. For example, L1 regularization coefficient, L2 regularization coefficient. The loss function measures the difference between the predicted value and the true value. For example, the loss function may include: mean squared error loss function, root mean squared error loss function, mean absolute error loss function, etc. The optimizer represents the method of minimizing the loss function. For example, the optimizer may include: gradient descent optimizer, Adam optimizer, etc.

[0027] The training task information is a task description of the model training task. Among them, the training task information in the training task information set includes: task description information, node address, training mode, and training sample address. The task description information is a task description of the model training task assigned to a single distributed node. The node address represents the communication address of the distributed node for the model training task corresponding to the task description information to be processed. The training mode represents the training mode of the distributed node for the model training task corresponding to the task description information to be processed when processing the distributed model training task corresponding to the task description information to be processed. In practice, the training mode may include: data parallel mode, model parallel mode, etc. Among them, the data parallel mode means that multiple distributed nodes perform model training in parallel independently. The model parallel mode means that the model is split into multiple parts and multiple distributed nodes perform model training in parallel. The training sample address refers to the storage address of the training samples required for the model training task corresponding to the training task information.

[0028] In practice, when the model structure is complex and there are numerous model parameters, it is impossible to efficiently train the target neural network model on a single machine. Therefore, a distributed approach is often required. The model training task corresponding to the target neural network model is decomposed into multiple sub-training tasks to be executed by multiple distributed nodes (computers or virtual computer nodes). Through various optimization methods such as parallel training, the training efficiency of the target neural network model can be improved. In practice, the above execution entity can use frameworks such as TensorFlow, PyTorch, and PaddlePaddle to create a model training task for the target neural network model according to the model description file corresponding to the target neural network model, and obtain a set of training task information.

[0029] It should be noted that the above computing device can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is embodied as software, it can be installed in the above-listed hardware devices. It can be implemented as, for example, multiple software or software modules for providing distributed services, or as a single software or software module. Specific limitations are not made here. In practice, the above computing execution entity can be a distributed framework for training the target neural network model.

[0030] In some optional implementation manners of some embodiments, the above execution entity creates a model training task for the target neural network model according to the model description file corresponding to the target neural network model, and obtains a set of training task information, including:

[0031] In the first step, according to the model structure description file included in the above model description file, determine the model structure of the target neural network model to generate a model structure diagram.

[0032] Among them, the model structure diagram is a graph structure used to describe the model structure of the target neural network model. Among them, the graph nodes in the model structure tree correspond to the layers included in the target neural network model.

[0033] As an example, in practice, the file formats of different model structure description files may vary. Therefore, by uniformly converting to obtain a model structure diagram, a unified description of the model structure corresponding to the model structure description file in the file format can be realized. See Figure 2 A schematic diagram of the model structure diagram shown, where Figure 2 The model structure diagram shown includes: an input layer, a convolutional layer A, a convolutional layer B, a convolutional layer C, a splicing layer, and a fully connected layer. Taking Python code as an example, the code corresponding to the model structure description file can be (which is the same asFigure 2 corresponding to the model structure diagram shown below:

[0034] # Define the input layer

[0035] input_layer = Input(shape=(64, 64, 3))

[0036] # Define three parallel convolutional layers A, B, and C

[0037] conv_layer_A = Conv2D(filters = 32, kernel_size=(3, 3), padding='same', activation='relu')(input_layer)

[0038] conv_layer_B = Conv2D(filters = 64, kernel_size=(5, 5), padding='same', activation='relu')(input_layer)

[0039] conv_layer_C = Conv2D(filters = 128, kernel_size=(7, 7), padding='same', activation='relu')(input_layer)

[0040] # Define the concatenation layer

[0041] concatenated = Concatenate(axis=-1)([conv_layer_A, conv_layer_B, conv_layer_C])

[0042] flattened = Flatten()(concatenated)

[0043] # Define the fully connected layer

[0044] dense_layer = Dense(128, activation='relu')(flattened)

[0045] output_layer = Dense(10, activation='softmax')(dense_layer)

[0046] # Build the model

[0047] model = Model(inputs=input_layer, outputs=output_layer)

[0048] Step 2: According to the above model structure diagram, decompose the above target neural network model to generate a set of sub-models.

[0049] In practice, each graph node in the model structure diagram can be a sub-model, and thus a set of sub-models is obtained. In addition, in order to improve the utilization efficiency of distributed nodes, multiple connected graph nodes can also be used as a sub-model.

[0050] Step 3: For each sub-model in the above set of sub-models, perform the following steps for creating a model training task:

[0051] The first sub-step: According to the number of model parameters and operator types corresponding to the above sub-model, allocate distributed nodes that match the above sub-model to obtain the node address included in the training task information corresponding to the above sub-model.

[0052] In practice, since the model structure of the sub-model is known and the size of the feature map corresponding to the sub-model is known, therefore, the number of model parameters corresponding to the sub-model can be quantified according to the sub-model combined with the size of the feature map processed by the sub-model. Operator types can include: linear operators (matrix multiplication, convolution, transposed convolution), normalization operators, pooling operators, etc. Therefore, the above execution entity can quantify the computing power requirements corresponding to the sub-model in combination with the number of model parameters and operator types corresponding to the sub-model, and thus allocate distributed nodes that match the sub-model. After the allocation is completed, the node address included in the training task information corresponding to the above sub-model can be obtained.

[0053] The second sub-step: In response to the completion of the allocation, generate a model training task for the above sub-model to obtain the task description information included in the training task information corresponding to the above sub-model.

[0054] In practice, a sub-model can be understood as a neural network model to be trained with a relatively small model size. Therefore, the above execution entity can use frameworks such as TensorFlow, PyTorch, and PaddlePaddle to create a model training task for the sub-model to obtain the task description information included in the training task information corresponding to the above sub-model.

[0055] The third sub-step: Predict the training time complexity according to the number of model parameters and operator types corresponding to the above sub-model.

[0056] In practice, since the number of model parameters is known and the operator type is known, the above execution entity can predict the training time complexity corresponding to the sub-model according to historical training tasks.

[0057] The fourth sub-step is to determine the training mode included in the training task information corresponding to the sub-model according to the above training time complexity and the above sub-model.

[0058] In practice, on the premise of knowing the training time complexity corresponding to the sub-model and the model structure corresponding to the sub-model, the above-mentioned execution entity can match the training mode that is optimally matched with the sub-model as the training mode included in the training task information corresponding to the sub-model.

[0059] The fifth sub-step is to determine the storage address of the training samples used for model training of the above sub-model as the training sample address included in the training task information corresponding to the above sub-model.

[0060] In practice, the above-mentioned execution entity can dynamically allocate the training samples used for model training of the above sub-model and use the node address (storage address) of the distributed node storing the training samples used for model training of the above sub-model as the training sample address included in the training task information corresponding to the above sub-model.

[0061] As an example, the model training task corresponding to the sub-model is executed on distributed node A, and the training samples used for model training of the above sub-model are stored on distributed node B. Therefore, the training sample address can be the node address of distributed node B.

[0062] Step 102: For each piece of training task information in the training task information set, perform the following processing steps:

[0063] Step 1021: Distribute the model training task corresponding to the training task information to the first node according to the training task information.

[0064] In some embodiments, the above-mentioned execution entity can distribute the model training task corresponding to the training task information to the first node according to the training task information. Among them, the above-mentioned first node is a distributed node used to execute the model training task corresponding to the above training task information. In practice, the above-mentioned execution entity can distribute the training task information to the first node in the form of broadcasting.

[0065] As an example, see Figure 3Schematic diagram of the distribution process of distributing the model training task corresponding to the above training task information to the first node. Among them, the distributed node A can be elected as the scheduling node, and the distributed node A can publish the training task information B in the broadcast channel in the form of broadcasting. Since the training task information B is the first node corresponding to the distributed node B, when the distributed node B receives the training task information B, it indicates successful distribution. For other distributed nodes, such as the distributed node C, the distributed node D,..., the distributed node Z, there is no corresponding relationship with the training task information B. Therefore, the distributed node C, the distributed node D,..., the distributed node Z will ignore the training task information B broadcast by the distributed node A.

[0066] In some optional implementation manners of some embodiments, according to the above training task information, distributing the model training task corresponding to the above training task information to the first node includes:

[0067] First step, generate a first temporary key.

[0068] In practice, the first temporary key can be a temporary key created by the scheduling node (for example, Figure 3 the distributed node A shown). Among them, the first temporary key includes: a public key and a private key.

[0069] Second step, obtain the identity key, signature key, and second temporary key corresponding to the first node from the key server.

[0070] Among them, the signature key corresponding to the above first node is updated regularly. The key server is a server for storing the identity key, signature key, and second temporary key corresponding to each distributed node. The identity key represents the node identity for the first node. The signature key represents the node signature for the first node. The second temporary key is a temporarily generated temporary key for the first node. Among them, the identity key, signature key, and second temporary key all include a private key and a public key.

[0071] Third step, determine the initial key according to the private key included in the above first temporary key, the public key included in the above identity key, the public key of the above signature key, and the public key of the above second temporary key.

[0072] Among them, the key length of the above initial key is 32 bytes. Initial key = KDF(DH(private key included in the first temporary key, public key included in the identity key) | DH(private key included in the first temporary key, public key of the signature key) | DH(private key included in the first temporary key, public key of the second temporary key)). Among them, DH() represents the ECDHE elliptic curve function. "|" represents concatenation. KDF() represents the key derivation function.

[0073] Step 4: Send a start message to the first node.

[0074] Among them, the start message may include the public key included in the identity key of the scheduling node (executing entity) and the public key included in the first temporary key.

[0075] Step 5: In response to receiving the verification message sent by the first node for the above start message, encrypt the above training task information according to the above initial key to obtain the encrypted training task information.

[0076] In practice, during the model training process, tampering with the training task information is one of the most direct ways to affect the execution of the model training task. When the tampering is initiated externally, it can be identified and intercepted through interception technologies such as firewalls. However, when the tampering is initiated internally, it is difficult to intercept through technologies such as firewalls. Therefore, on this basis, this application defaults that the distributed nodes are not trusted. When the distributed nodes communicate with each other, especially when transmitting information and data related to model training, the above encryption method is used for encryption to ensure the security of the information transmission process.

[0077] Step 6: Send the above encrypted training task information to the above first node to distribute the model training task corresponding to the above training task information to the first node.

[0078] In practice, since the above encryption method is essentially asymmetric encryption, after receiving the encrypted training task information, the first node can decrypt the encrypted training task information by combining the corresponding public key and private key.

[0079] Step 1022: In response to monitoring a task modification operation for the first node, generate operation risk description information for the task modification operation.

[0080] In some embodiments, the above executing entity may generate operation risk description information for the task modification operation in response to monitoring a task modification operation for the first node. In practice, the task modification operation may be a modification operation initiated externally to the distributed system or a modification operation initiated internally to the distributed system. Among them, the above operation risk description information includes: operation type, operation risk confidence level, operation anomaly confidence level, operation source credibility, virus confidence level, where the virus confidence level represents the confidence level that the above task modification operation is a host virus. The operation type represents the specific operation type of the task modification operation. The operation risk confidence level represents the risk level of executing the task modification operation. The operation anomaly confidence level represents the behavioral anomaly degree of the task modification operation. The operation source credibility represents the credibility of the node that initiates the task modification operation.

[0081] In some embodiments and some optional implementation manners, the above-mentioned execution entity generates operation risk description information for the above-mentioned task modification operation in response to monitoring a task modification operation for the above-mentioned first node, including:

[0082] In the first step, an operation feature extraction model is used to extract operation features from the above-mentioned task modification operation to generate operation features.

[0083] In practice, first, the instruction corresponding to the task modification operation can be converted into a two-dimensional grayscale image. Then, taking the two-dimensional grayscale image as the input, through a convolutional neural network model with 7 convolutional layers, as the operation feature extraction model, the above-mentioned task modification operation is subjected to operation feature extraction to generate operation features. Specifically, when converting the instruction corresponding to the task modification operation into a two-dimensional grayscale image, the instruction corresponding to the task modification operation can be converted into a binary string, which is cut into substrings with a length of 8 bits. Among them, a substring with a length of 8 bits corresponds to 1 pixel point, so that the instruction corresponding to the task modification operation can be converted into a two-dimensional grayscale image. Considering that the lengths of the instructions corresponding to the task modification operations are often different, resulting in inconsistent sizes of the obtained two-dimensional grayscale images and unable to be directly input into the convolutional neural network model, the operation feature extraction model in the convolutional neural network model with 7 convolutional layers further includes: 3 downsampling layers arranged in parallel for different grayscale image scales. The input size of each downsampling layer is preset. When the size of the two-dimensional grayscale image is consistent with the input size of the downsampling layer, it can be directly input. When they are inconsistent, the input size of the downsampling layer closest to the size of the two-dimensional grayscale image is determined as the constraint, and the two-dimensional grayscale image is padded with 0 to ensure that the padded two-dimensional grayscale image can be input into the downsampling layer. Among them, the output sizes of the 3 downsampling layers are all consistent with the input size of the convolutional neural network model with 7 convolutional layers.

[0084] In the second step, according to the operation type classifier and the above-mentioned operation features, the operation type included in the above-mentioned operation risk information is determined.

[0085] Among them, the operation type classifier is a multi-classifier, which is used to output the operation type included in the operation risk information.

[0086] In the third step, according to the operation risk confidence predictor and the above-mentioned operation features, the operation risk confidence included in the above-mentioned operation risk information is determined.

[0087] In practice, the operation risk confidence predictor can be a single-classifier, and the confidence corresponding to its classification result can be the operation risk confidence.

[0088] In the fourth step, the operation initiation source, operation initiation time, operation initiation period, and operation initiation frequency of the above-mentioned task modification operation are determined.

[0089] In practice, the operation initiation source of the above task modification operation can be located through reverse tracing, and according to the operation log, the operation initiation time, operation initiation period, and operation initiation frequency corresponding to the task modification operation can be obtained.

[0090] Step 5: Determine the operation anomaly confidence level included in the above operation risk information according to the above operation characteristics, the above operation initiation source, the above operation initiation time, the above operation initiation period, and the above operation initiation frequency.

[0091] In practice, the above operation initiation source, the above operation initiation time, the above operation initiation period, and the above operation initiation frequency are static characteristics. The above execution entity can perform feature encoding on the above operation initiation source, the above operation initiation time, the above operation initiation period, and the above operation initiation frequency through feature encoding. Then, the encoded features and operation characteristics are concatenated as the input of the operation anomaly classifier to obtain the operation anomaly confidence level. Among them, the operation anomaly classifier can also be a binary classifier (the classification results include operation anomaly and operation normal), and the confidence level of the corresponding classification results can be the operation anomaly confidence level.

[0092] In practice, the operation feature extraction model, operation type classifier, risk confidence predictor, and operation anomaly classifier can be taken as a whole and the model can be trained in a supervised manner.

[0093] Step 6: Determine whether the above operation initiation source is located in the list of trusted operation sources.

[0094] Among them, the list of trusted operation sources is a list used to maintain the identities of trusted operation sources. The above execution entity can regularly perform trusted verification on distributed nodes and store the node addresses of the operation sources (distributed nodes) that meet the trusted verification in the list of trusted operation sources.

[0095] Step 7: In response to determining that the above operation initiation source is not located in the list of trusted operation sources, determine whether there is a historical operation portrait corresponding to the above operation initiation source.

[0096] In practice, the historical operation portrait can be a behavior portrait constructed for the historical operation behavior of the operation initiation source.

[0097] Step 8: In response to the existence of a historical operation portrait corresponding to the above operation initiation source, determine the credibility of the operation initiation source included in the above operation risk information according to the above historical operation portrait, the above operation characteristics, the above operation initiation source, the above operation initiation time, the above operation initiation period, and the above operation initiation frequency.

[0098] Among them, the credibility of the operation initiation source can be combined with a pre-constructed decision tree model, taking the above-mentioned historical operation portraits, the above-mentioned operation characteristics, the above-mentioned operation initiation source, the above-mentioned operation initiation time, the above-mentioned operation initiation period, and the above-mentioned operation initiation frequency as inputs to obtain the credibility of the operation initiation source.

[0099] The ninth step, in response to the non-existence of the historical operation portrait corresponding to the above-mentioned operation initiation source, according to the above-mentioned operation characteristics, the above-mentioned operation initiation source, the above-mentioned operation initiation time, the above-mentioned operation initiation period, and the above-mentioned operation initiation frequency, determine the credibility of the operation initiation source included in the above-mentioned operation risk information.

[0100] In practice, the above-mentioned execution entity can still combine the above-mentioned decision tree model to obtain the credibility of the operation initiation source.

[0101] The tenth step, compare the above-mentioned operation characteristics with the host virus characteristics in the host virus feature library to generate a virus similarity, which is used as the virus confidence included in the above-mentioned operation risk description information.

[0102] In practice, the above-mentioned execution entity can adopt a similarity calculation method to determine the virus similarity, which is used as the virus confidence included in the above-mentioned operation risk description information.

[0103] Quantify the operation risk of the task modification operation from multiple perspectives including operation type, operation risk confidence, operation anomaly confidence, credibility of the operation initiation source, and virus confidence.

[0104] Step 1023, in response to receiving the training sample information and the training sample verification information sent by the second node respectively, according to the training sample verification information, perform a training sample anti-tampering verification on the training sample information to generate sample risk description information.

[0105] In some embodiments, the above-mentioned execution entity can respond to receiving the training sample information and the training sample verification information sent by the second node respectively, and according to the training sample verification information, perform a training sample anti-tampering verification on the training sample information to generate sample risk description information. Among them, the above-mentioned second node is the distributed node corresponding to the training sample address included in the above-mentioned training task information. The above-mentioned training sample information includes: training samples and sample labels, and the training samples are image type training samples.

[0106] In some optional implementation manners of some embodiments, the above-mentioned execution entity responds to receiving the training sample information and the training sample verification information sent by the second node respectively, and according to the above-mentioned training sample verification information, performs a training sample anti-tampering verification on the above-mentioned training sample information to generate sample risk description information, including:

[0107] In the first step, decrypt the above training sample verification information to obtain the decrypted training sample verification information.

[0108] Among them, the above decrypted training sample verification information includes: training sample key point features, trusted sample labels, feature extraction parameters, and check codes. The check code is an anti-tampering check code for the training sample key point features, the above trusted sample labels, and the above feature extraction parameters. In practice, the encryption and decryption methods for transmitting information and data between the first node and the second node can refer to the content in step 1021. The training sample key point features are key point features collected in advance for the training samples. The trusted sample labels can be trusted sample labels. The feature extraction parameters are extraction parameters for sample features in the training samples used for anti-tampering verification. The check code can be a check code for the training sample key point features, trusted sample labels, and feature extraction parameters. Specifically, the check code can be a CRC (Cyclic Redundancy Check) cyclic check code.

[0109] In the second step, perform anti-tampering verification on the above training sample key point features, the above trusted sample labels, and the above feature extraction parameters according to the above check code.

[0110] In practice, the above execution entity can perform anti-tampering verification on the above training sample key point features, the above trusted sample labels, and the above feature extraction parameters according to the above check code through the CRC verification method.

[0111] In the third step, in response to passing the verification information anti-tampering verification, perform the following verification steps:

[0112] The first sub-step is to extract at least M local key point features from the above training sample key point features.

[0113] Among them, 32 ≤ M ≤ 64. The M local key point features are M non-connected local features among the training sample key point features.

[0114] The second sub-step is to extract features from the training samples included in the above training sample information according to the above feature extraction parameters to generate the training sample features to be verified.

[0115] In practice, since the training sample is an image type training sample, the feature extraction parameters can be parameters related to the convolution kernel for image features constructed in advance. The training sample key point features are also extracted according to the feature extraction parameters.

[0116] The third sub-step is to perform feature comparison on the corresponding positions of the above training sample features to be verified according to the above M local key point features to generate the first comparison result.

[0117] Among them, the first comparison result characterizes whether the corresponding positions of the M local key-point features and the features of the training sample to be verified are consistent.

[0118] The fourth sub-step is to determine the label consistency between the above-mentioned reliable sample label and the sample label included in the above-mentioned training sample information, so as to generate a second comparison result.

[0119] Among them, the second comparison result characterizes whether the reliable sample label is consistent with the sample label included in the above-mentioned training sample information.

[0120] The fifth sub-step is to generate the above-mentioned sample risk description information according to the above-mentioned first comparison result and the above-mentioned second comparison result.

[0121] Step 1024 is to perform the protection security for the model training task corresponding to the training task information according to the operation risk description information and / or the sample risk description information.

[0122] In some embodiments, the above-mentioned execution subject may perform the protection security for the model training task corresponding to the training task information according to the operation risk description information and / or the sample risk description information.

[0123] In some optional implementation manners of some embodiments, the above-mentioned execution subject performs the protection security for the model training task corresponding to the above-mentioned training task information according to the above-mentioned operation risk description information and / or the above-mentioned sample risk description information, including:

[0124] The first step is to, in response to the above-mentioned operation risk description information indicating that there is an operation risk in the above-mentioned task modification operation, perform the following first protection step:

[0125] The first sub-step is to determine the operation protection for the above-mentioned task modification operation according to the above-mentioned operation risk description information and the operation protection decision tree.

[0126] Among them, the operation protection decision tree is a pre-constructed decision tree used to combine risk description information for operation protection type decision-making.

[0127] The second sub-step is to initiate a task risk warning for the model training task corresponding to the above-mentioned training task information.

[0128] In practice, the above-mentioned execution subject may initiate a task risk warning to the object that initiates the model training task for the target neural network model.

[0129] The second step is to, in response to the above-mentioned sample risk description information indicating that the training sample has a sample risk, freeze the above-mentioned training sample information and perform the following second protection step:

[0130] The first sub-step is to initiate a node lock for the above-mentioned second node.

[0131] In practice, the above-mentioned execution entity can initiate node locking for the second node through the scheduling node to avoid communication, data transmission, etc. between the second node and other distributed nodes.

[0132] The second sub-step is to switch the training sample sending node for the model training task corresponding to the above training task information.

[0133] In practice, to ensure the stability of the training process, data often adopts a multi-node backup method. Therefore, when the second node is abnormal, it can be switched to the backup node corresponding to the second node as the training sample sending node for the model training task corresponding to the above training task information.

[0134] Broadcast the node untrusted broadcast for the above-mentioned second node to other distributed nodes.

[0135] The third sub-step is to broadcast the node untrusted broadcast for the above-mentioned second node to other distributed nodes.

[0136] In practice, through the method of node untrusted broadcast, this can avoid other distributed nodes from initiating invalid communication and data exchange with the second node.

[0137] The above embodiments of the present application have the following beneficial effects: Through the virus protection method based on host security in some embodiments of the present application, effective protection during the model training stage is achieved, avoiding problems such as the model being contaminated, which affects the quality of the trained model, and the waste of computing resources. Specifically, the reasons for the above problems are as follows: Due to the characteristics of a neural network model with an extremely large number of parameters, a large amount of computer resources are required for model training during the model training stage. During the training stage, when malicious sample tampering or malicious operations occur, it will, at the very least, affect the model convergence speed, and at the worst, affect the quality of the trained model. At the same time, it will also cause a waste of a large amount of computer resources. Based on this, in some embodiments of the virus protection method based on host security of the present application, first, according to the model description file corresponding to the target neural network model, a model training task for the above target neural network model is created to obtain a training task information set. Among them, the above target neural network model is a neural network model to be trained, and the above model description file includes: a model structure description file and a model training parameter description file. The training task information in the above training task information set includes: task description information, node address, training mode, and training sample address. For a neural network model with an extremely large number of parameters, it is difficult to match the corresponding training requirements using a single-machine training method. Therefore, a distributed method is usually used for model training. Therefore, through task decomposition, the entire model training task can be effectively disassembled into training tasks processed by nodes in different distributions, thereby improving the training efficiency. Secondly, for each training task information in the above training task information set, the following processing steps are executed: The first step is to distribute the model training task corresponding to the above training task information to the first node according to the above training task information, where the above first node is a distributed node used to execute the model training task corresponding to the above training task information. The second step is to generate operation risk description information for the above task modification operation in response to monitoring the task modification operation for the above first node. Among them, the above operation risk description information includes: operation type, operation risk confidence level, operation anomaly confidence level, operation origin credibility, and virus confidence level, where the virus confidence level represents the confidence level that the above task modification operation is a host virus. In order to improve the training speed, it is often necessary to adaptively adjust the task according to the current training state. During this period, there may be invasive malicious task modifications or non-invasive, internally initiated malicious task modifications. Therefore, for each generated task modification operation, the corresponding operation risk description information is generated to evaluate the risk of the task modification operation.In the third step, in response to receiving the training sample information and the training sample verification information sent by the second node respectively, according to the above-mentioned training sample verification information, perform anti-tampering verification on the above-mentioned training sample information to generate sample risk description information, where the second node is the distributed node corresponding to the training sample address included in the above-mentioned training task information, and the above-mentioned training sample information includes: training samples and sample labels, and the training samples are image-type training samples. In practice, for a supervised training method, training samples and sample labels are crucial for the training accuracy of the model. However, during the training process, through the method of sample contamination, it is extremely easy to cause damage to the training accuracy of the model, and even the situation of non-convergence may occur. Therefore, it is necessary to perform corresponding verification on whether the training samples are tampered with to obtain corresponding sample risk description information. Finally, perform the protection security for the model training task corresponding to the above-mentioned training task information according to the above-mentioned operation risk description information and / or the above-mentioned sample risk description information. In this way, the neural network model in the training stage is effectively protected, thereby ensuring the training quality and avoiding the waste of computing resources.

[0138] This application also provides a computer device 400. As Figure 4 shown, the computer device 400 includes: a bus 401, a processor 402, a memory 403, and a communication interface 404. The processor 402, the memory 403, and the communication interface 404 communicate with each other through the bus 401. The computer device 400 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computer device 400.

[0139] The bus 401 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 4 only one line is shown here, but it does not mean that there is only one bus or one type of bus. The bus 401 can include a path for transmitting information between various components of the computer device 400 (for example, the memory 403, the processor 402, the communication interface 404).

[0140] The processor 402 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0141] The memory 403 may include a volatile memory, such as a random access memory (RAM). The memory 403 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0142] The executable program code is stored in the memory 403, and the processor 402 executes the executable program code to respectively implement the functions of the foregoing acquisition module, sampling module, determination module, and mixing module, thereby implementing the foregoing virus protection method based on host security. That is, instructions for executing the foregoing virus protection method based on host security are stored on the memory 403.

[0143] The communication interface 404 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computer device 400 and other devices or a communication network.

[0144] An embodiment of this application also provides a chip, which includes a processor and a data interface. The processor reads the instructions stored on the memory through the data interface to execute the foregoing virus protection method based on host security.

[0145] An embodiment of this application also provides a computer-readable storage medium. The foregoing computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center that includes one or more available media. The foregoing available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state drive), etc. The computer-readable storage medium includes instructions, and the foregoing instructions instruct the computing device to execute the foregoing virus protection method based on host security.

[0146] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered that the scope described in this specification is covered.

[0147] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A virus protection method based on host security, comprising: According to the model description file corresponding to the target neural network model, a model training task for the target neural network model is created to obtain a training task information set, wherein the target neural network model is a neural network model to be trained, the model description file includes: a model structure description file and a model training parameter description file, and the training task information in the training task information set includes: task description information, node address, training mode and training sample address; For each training task information in the training task information set, the following processing steps are performed: Distribute the model training task corresponding to the training task information to a first node according to the training task information, wherein the first node is a distributed node for executing the model training task corresponding to the training task information; In response to monitoring a task modification operation for the first node, generating operation risk description information for the task modification operation, wherein the operation risk description information includes: operation type, operation risk confidence, operation abnormality confidence, operation origin credibility, and virus confidence, wherein the virus confidence represents the confidence that the task modification operation is a host virus; In response to receiving the training sample information and the training sample verification information sent by the second node respectively, performing a training sample tamper-proof check on the training sample information according to the training sample verification information to generate sample risk description information, wherein the second node is a distributed node corresponding to the training sample address included in the training task information, and the training sample information includes: a training sample and a sample label, and the training sample is an image type training sample; Perform security protection for the model training task corresponding to the training task information according to the operation risk description information and / or the sample risk description information.

2. The method according to claim 1, wherein: The step of creating a model training task for the target neural network model according to the model description file corresponding to the target neural network model, and obtaining a training task information set includes: Determining the model structure of the target neural network model according to the model structure description file included in the model description file to generate a model structure diagram; According to the model structure diagram, the target neural network model is deconstructed to generate a sub-model set; For each sub-model in the sub-model set, perform the following model training task creation steps: According to the model parameter quantity and operator type corresponding to the sub-model, a distributed node matching the sub-model is allocated to obtain a node address included in the training task information corresponding to the sub-model; In response to the allocation being completed, generating a model training task for the sub-model, and obtaining task description information included in the training task information corresponding to the sub-model; Predicting the training time complexity according to the model parameter quantity and operator type corresponding to the sub-model; Determine, according to the training time complexity and the sub-model, a training mode included in the training task information corresponding to the sub-model; Determine a storage address of a training sample used for model training of the sub-model as the training sample address included in the training task information corresponding to the sub-model.

3. The method according to claim 2, wherein: The step of distributing the model training task corresponding to the training task information to the first node according to the training task information includes: generating a first temporary key; Obtaining an identity key, a signature key, and a second temporary key corresponding to the first node from a key server, wherein the signature key corresponding to the first node is updated regularly; Determine an initial key according to the private key included in the first temporary key, the public key included in the identity key, the public key of the signature key, and the public key of the second temporary key, wherein the key length of the initial key is 32 bytes; Sending an initiation message to the first node; In response to receiving a verification message for the start message sent by the first node, encrypting the training task information according to the initial key to obtain encrypted training task information; The encrypted training task information is sent to the first node to distribute the model training task corresponding to the training task information to the first node.

4. The method according to claim 3, wherein: In response to receiving the training sample information and the training sample verification information sent by the second node respectively, performing a training sample tamper-proof check on the training sample information according to the training sample verification information to generate sample risk description information, including: Decrypting the training sample verification information to obtain decrypted training sample verification information, wherein the decrypted training sample verification information includes: key point features of the training sample, a trusted sample label, feature extraction parameters and a check code, wherein the check code is a tamper-proof check code for the key point features of the training sample, the trusted sample label and the feature extraction parameters; According to the verification code, anti-tampering verification is performed on the key point features of the training sample, the trusted sample label and the feature extraction parameter; In response to passing the verification information tamper-proof check, the following verification steps are performed: Extracting at least M local key point features from the key point features of the training samples; According to the feature extraction parameters, feature extraction is performed on the training samples included in the training sample information to generate features of the training samples to be verified; According to the M local key point features, feature comparison is performed on the points corresponding to the features of the training sample to be verified to generate a first comparison result; Determining label consistency between the credible sample label and the sample label included in the training sample information to generate a second comparison result; The sample risk description information is generated according to the first comparison result and the second comparison result.

5. The method according to claim 4, wherein: The performing of security protection for the model training task corresponding to the training task information according to the operation risk description information and / or the sample risk description information includes: In response to the operational risk description information indicating that the task modification operation has an operational risk, the following first protection step is performed: Determining the operation protection for the task modification operation according to the operation risk description information and the operation protection decision tree; Initiate a task risk warning for the model training task corresponding to the training task information; In response to the sample risk description information indicating that the training sample has sample risk, freezing the training sample information, and performing the following second protection step: Initiating node locking for the second node; Switching a training sample sending node for the model training task corresponding to the training task information; The node broadcast for the second node is not credible and is broadcast to other distribution nodes.

6. A computer device, wherein: The computer device comprises a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

7. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Intelligent model construction method based on meta-features

    CN111857680A

  • Data splitting and computing power processing method and device, electronic equipment and storage medium

    CN118606061A