File security detection method and device, electronic equipment and storage medium

By calling the multi-agent system in file security detection, extracting and understanding the multimodal features of the file, the problem that the existing technology is difficult to accurately detect file security is solved, and flexible response and efficient detection of complex network attacks are achieved.

CN120017307APending Publication Date: 2025-05-16安全智能有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411976667.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-07
Filing Date
2024-12-30
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing file security detection technologies are difficult to flexibly deal with increasingly complex new cyber attacks, and it is difficult to accurately detect file security.

Method used

By calling the multi-agent system to work together, the multi-modal features in the file to be detected are first extracted, and the multi-modal features are understood information is generated, and the detection results of the file to be detected are determined based on the multi-modal features understanding information. The specific steps include: pre-training multiple initial models using a pre-training data set, post-training multiple pre-training models using a post-training data set, and then building a multi-agent system.

Benefits of technology

It realizes the security of accurate file detection, can flexibly respond to complex new cyber attacks, and improves the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017307A_ABST
    Figure CN120017307A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a file security detection method and device, electronic equipment and a storage medium, and relates to the technical field of network security. The file security detection method comprises the steps of obtaining a detection instruction for a to-be-detected file, and calling a multi-agent system; the multi-agent system comprises a first agent, a second agent and a third agent, adopting the first intelligent agent to generate a plurality of tasks according to the detection instruction; wherein the plurality of tasks comprise a feature extraction task, a feature understanding task and a security judgment task; processing the feature extraction task by adopting a second intelligent agent, and extracting multi-modal features in the to-be-detected file; processing the feature understanding task by adopting a third agent, and generating multi-modal feature understanding information; and processing the security judgment task by adopting a second intelligent agent, and determining a detection result of the to-be-detected file according to the understanding information. According to the embodiment of the invention, the technical effect of accurately detecting the security of the file can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and in particular to a file security detection method, device, electronic device and storage medium. Background Art

[0002] With the rapid growth of network information, network security issues are becoming more and more serious, and file security detection technology has emerged. The goal of file security detection technology is to detect and block malicious software that may pose a threat to network security by monitoring and analyzing files, such as viruses, worms and Trojans carried by files.

[0003] However, existing file security detection technologies are often unable to flexibly respond to increasingly complex new network attack methods, and it is difficult to accurately detect the security of files. Therefore, it is urgent to improve file security detection technology. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a file security detection method, device, electronic device and storage medium to achieve the technical effect of accurately detecting the security of files.

[0005] In a first aspect, an embodiment of the present application provides a file security detection method, comprising: Obtaining a detection instruction for a file to be detected, and calling a multi-agent system; wherein the multi-agent system includes a first agent, a second agent, and a third agent; Using the first agent, generating a plurality of tasks according to the detection instruction; wherein the plurality of tasks include a feature extraction task, a feature understanding task and a safety judgment task; Using the second agent to process the feature extraction task to extract multimodal features from the file to be detected; Using the third agent to process the feature understanding task to generate understanding information of the multimodal feature; The second intelligent agent is used to process the security judgment task, and the detection result of the file to be detected is determined according to the understanding information.

[0006] In the above implementation process, by calling the multi-agent system to work collaboratively, the multimodal features in the file to be detected are first extracted, the understanding information of the multimodal features is generated, and then the detection results of the file to be detected are determined based on the understanding information of the multimodal features, so that the security of the file can be accurately detected.

[0007] Furthermore, before calling the multi-agent system, the method further includes: Pre-training multiple initial models using pre-training data sets to obtain multiple pre-trained models; Post-training the multiple pre-trained models using a post-training data set to obtain multiple target models; The multi-agent system is constructed according to the multiple target models.

[0008] In the above implementation process, by using a pre-training data set to pre-train multiple initial models, and using a post-training data set to post-train the multiple pre-trained models, and then building a multi-agent system based on the obtained multiple target models, it is possible to ensure that the multi-agent system works collaboratively and accurately detect the security of files.

[0009] Further, the multi-agent system further includes a fourth agent; After the first agent is used to generate a plurality of tasks according to the detection instruction, the method further includes: Using the fourth agent to monitor whether the processing flow of the first agent is abnormal, and obtaining a first monitoring result; In the case where the first monitoring result is abnormal, sending the first monitoring result to the first agent to re-adopt the first agent and generate a plurality of new tasks according to the detection instruction; and / or, After the second agent is used to process the feature extraction task and extract the multimodal features in the to-be-detected file, the method further includes: Using the fourth agent to monitor whether the processing flow of the second agent is abnormal, and obtaining a second monitoring result; In the case where the second monitoring result is abnormal, the second monitoring result is sent to the second agent so as to re-adopt the second agent to process the feature extraction task.

[0010] In the above implementation process, by adding a fourth agent in the multi-agent system, the fourth agent is used to monitor whether the processing flow of the first agent and / or the second agent is abnormal, and when the processing flow of the first agent and / or the second agent is abnormal, the first agent and / or the second agent is re-used for corresponding processing, which can further ensure the security of accurate detection of files.

[0011] Further, the multimodal features include text features and image features; The extracting of multimodal features from the file to be detected includes: Parsing the file to be detected to obtain text data and binary data in the file to be detected; Extracting features of the text data to obtain the text features; generating a target image according to the text data and the binary data; Extract the features of the target image to obtain the image features.

[0012] In the above implementation process, by adopting a second intelligent agent to take the features of the text data in the file to be detected as text features, and generating a target image based on the text data and binary data in the file to be detected, and taking the features of the target image as image features, the multimodal features of the file to be detected can be fully extracted.

[0013] Further, generating a target image according to the text data and the binary data includes: Generating a text image according to the text data; wherein the text image includes a statistical graph for representing the distribution of the text data; Generate a binary image according to the binary data; wherein the binary image includes at least one of a statistical graph and a binary grayscale graph for representing the distribution of the binary data; The text image and the binary image are reconstructed to obtain the target image.

[0014] In the above implementation process, by adopting a second intelligent agent to generate a text image according to the text data in the file to be detected, and to generate a binary image according to the binary data in the file to be detected, and then reconstructing the text image and the binary image to obtain the target image, the target image can be accurately generated, ensuring comprehensive and accurate extraction of multimodal features in the file to be detected.

[0015] Further, the multimodal features include text features and image features; The generating the understanding information of the multimodal feature includes: Fusing the text feature and the image feature to obtain a target feature; The understanding information is generated according to the target feature.

[0016] In the above implementation process, by adopting a third intelligent agent to first fuse text features and image features into target features, and then generate understanding information based on the target features, it is possible to ensure accurate generation of understanding information of multimodal features.

[0017] Further, the multi-agent system further includes a fifth agent; The method further comprises: Using the fifth agent to generate process explanation information according to the processing flow of the multi-agent system; and / or, The fifth agent is used to generate result explanation information according to the detection result.

[0018] In the above implementation process, by adding a fifth agent in the multi-agent system, using the fifth agent to generate process explanation information according to the processing flow of the multi-agent system, and / or using the fifth agent to generate result explanation information according to the detection results, it is possible to provide users with explanatory information on the security detection process of the file to be detected.

[0019] Furthermore, the method further comprises: According to the detection results obtained within a preset period and user feedback information corresponding to the detection results, statistical evaluation index values ​​are calculated; When the evaluation index value is less than a preset index threshold, the multi-agent system is updated.

[0020] In the above implementation process, by statistically evaluating the evaluation index value based on the detection results obtained within a preset period and the user feedback information corresponding to the detection results, and updating the multi-agent system when the evaluation index value is less than the preset index threshold, the security of the multi-agent system can be effectively guaranteed to accurately detect files.

[0021] In a second aspect, an embodiment of the present application provides a file security detection device, comprising: An instruction acquisition module, used to acquire a detection instruction for a file to be detected and call a multi-agent system; wherein the multi-agent system includes a first agent, a second agent and a third agent; A task planning module, configured to use the first agent to generate a plurality of tasks according to the detection instruction; wherein the plurality of tasks include a feature extraction task, a feature understanding task and a safety judgment task; A feature extraction module, used to use the second agent to process the feature extraction task and extract multimodal features from the file to be detected; A feature understanding module, configured to process the feature understanding task using the third agent to generate understanding information of the multimodal feature; A security judgment module is used to use the second agent to process the security judgment task and determine the detection result of the file to be detected according to the understanding information.

[0022] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; the memory is coupled to the processor, and the processor implements the method described above when executing the computer program.

[0023] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method described above.

[0024] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes instructions, and when the instructions are executed by a computer, the computer implements the method as described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0026] Figure 1 A flowchart of a file security detection method provided in the first embodiment of the present application; Figure 2 This is a schematic diagram of the structure of a multi-agent system illustrated in the first embodiment of the present application; Figure 3 A flowchart of a file security detection method exemplified in the first embodiment of the present application; Figure 4 A schematic diagram of the structure of a file security detection device provided in the second embodiment of the present application; Figure 5 A schematic diagram of the structure of an electronic device provided in the third embodiment of the present application. DETAILED DESCRIPTION

[0027] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0028] It should be noted that in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance. At the same time, the step numbers in the text are only for the convenience of explanation of the embodiments of this application and do not serve to limit the order of execution of the steps. The method provided in the embodiments of this application can be executed by a related terminal device, and the following description will be taken as an example of a user terminal as the execution subject.

[0029] It should be noted that user terminals include terminal devices such as mobile phones, tablets and computers held by users.

[0030] Please see Figure 1 , Figure 1The first embodiment of the present application provides a method for detecting file security, including steps S101 to S105: S101, obtaining a detection instruction for a file to be detected, and calling a multi-agent system; wherein the multi-agent system includes a first agent, a second agent, and a third agent; S102, using a first agent to generate multiple tasks according to the detection instruction; wherein the multiple tasks include a feature extraction task, a feature understanding task and a safety judgment task; S103, using a second agent to process a feature extraction task to extract multimodal features from the file to be detected; S104, using a third agent to process a feature understanding task to generate understanding information of multimodal features; S105. Use the second intelligent agent to process the security judgment task and determine the detection result of the file to be detected based on the understanding information.

[0031] As an exemplary embodiment, according to actual application requirements, a multi-agent system is pre-constructed, wherein the multi-agent system includes a first agent, a second agent, and a third agent.

[0032] It should be noted that an intelligent agent refers to a computing entity with intelligence. A neural network model can be selected, such as a neural network model based on a self-attention mechanism, namely the Transformer model, a time recurrent neural network model, namely the LSTM (Long Short-Term Memory) model, and a linearized Transformer model, namely the RWKV model.

[0033] In practical applications, the neural network models selected by multiple agents such as the first agent, the second agent and the third agent in the multi-agent system can be all the same, partially the same, or completely different.

[0034] When a user needs to detect the security of any executable file, the user can use the executable file as a file to be detected, upload the file to be detected to the user terminal, and input a detection instruction for the file to be detected to the user terminal.

[0035] The user terminal obtains the detection instruction for the file to be detected and calls the pre-built multi-agent system.

[0036] After calling the multi-agent system, the first agent can be used to generate multiple tasks according to the detection instructions, where the multiple tasks include feature extraction tasks, feature understanding tasks and safety judgment tasks, and send the feature extraction tasks and safety judgment tasks to the second agent, and send the feature understanding tasks to the third agent.

[0037] After sending the feature extraction task to the second agent, the second agent is continued to process the feature extraction task to extract the multimodal features in the file to be detected so as to send the multimodal features to the third agent.

[0038] It should be noted that multimodal features refer to the features of multiple heterogeneous modal data, such as text and image data.

[0039] After sending the feature understanding task and the multimodal features to the third agent, the third agent is continued to process the feature understanding task to generate understanding information of the multimodal features to send the understanding information to the second agent.

[0040] It should be noted that understanding information refers to the information obtained by analyzing and understanding the data characteristics described by multimodal features.

[0041] In practical applications, understanding information may include the correlation between multimodal features, such as the correlation between features of different modalities, such as the correlation between text feature A1 and image feature B1, or the correlation between features of the same modality, such as text feature A1 and text feature A2. It may also include detailed software information represented by text features, such as software type and software function.

[0042] After sending the security judgment task and the understanding information to the second agent, the second agent continues to process the security judgment task, and determines the detection result of the file to be detected based on the understanding information.

[0043] In actual applications, the detection result may include a conclusion on whether the file to be detected carries malware, as well as detailed information about the malware, such as the type of malware, the virus family to which the malware belongs, and the paragraph location of the malware in the file to be detected.

[0044] It is understandable that there are interactive operations between multiple agents such as the first agent, the second agent and the third agent in the multi-agent system. By calling the multi-agent system to work together, the multimodal features in the file to be detected are first extracted, the understanding information of the multimodal features is generated, and then the detection result of the file to be detected is determined based on the understanding information of the multimodal features. The powerful learning ability of the agent can be used to analyze and understand the multimodal features in the file to be detected, comprehensively identify various network attacks, and accurately determine whether the file to be detected carries malware, so as to flexibly respond to the increasingly complex new network attack methods and accurately detect the security of the file.

[0045] The embodiment of the present application calls on a multi-agent system to work collaboratively, first extracts multimodal features from the file to be detected, generates understanding information of the multimodal features, and then determines the detection result of the file to be detected based on the understanding information of the multimodal features, thereby accurately detecting the security of the file.

[0046] In an optional embodiment, before calling the multi-agent system, it also includes: using a pre-training data set to pre-train multiple initial models to obtain multiple pre-trained models; using a post-training data set to post-train the multiple pre-trained models to obtain multiple target models; and constructing the multi-agent system based on the multiple target models.

[0047] As an example, according to actual application requirements, a pre-training data set is obtained, and the pre-training data set includes text data and attack knowledge data collected from the network, wherein the text data collected from the network includes Chinese text data, English text data and Blog (network log) text data, and the attack knowledge data collected from the network includes attack event data, attack code, TTPs (tactics, techniques and procedures) knowledge data, vulnerability knowledge data, network security general knowledge data, virus family knowledge data and feature data of virus family attack codes.

[0048] In order to ensure that the pre-training data in the pre-training data set effectively participates in model pre-training, after obtaining the pre-training data set, the pre-training data set can be pre-processed first, and then the pre-processed pre-training data set can be used to pre-train multiple initial models to obtain multiple pre-trained models.

[0049] In an optional implementation of the present embodiment, before the pre-training of multiple initial models using a pre-training data set to obtain multiple pre-trained models, it also includes: pre-processing the pre-training data set, wherein the pre-processing includes at least one of data filling processing, data deduplication processing, data ratio adjustment processing, data normalization processing and data weighting processing.

[0050] As an example, for a pre-training data set, if a certain pre-training data is incomplete, the pre-training data is filled in. For example, if the attack code lacks a necessary code field, such as a variable, the code field of the attack code is filled with the field value corresponding to the code field, such as an empty value; if a certain pre-training data is repeated with other pre-training data, one of these repeated pre-training data is selected and retained to avoid data redundancy in the pre-training data set; after dividing multiple categories of pre-training data according to classification criteria such as language category, data type, and knowledge field, the data volume ratio of the multiple categories of pre-training data can be adjusted, and some pre-training data can be randomly eliminated to meet the data volume ratio, further avoiding data redundancy in the pre-training data set; if the data format of a certain pre-training data is not a standard format that can be recognized and processed by the user terminal, the pre-training data is converted into a standard format that can be recognized and processed by the user terminal; if a certain pre-training data is important data, the importance of the pre-training data can be considered and the pre-training data can be given a corresponding weight so that the initial model focuses on the important data.

[0051] After obtaining the pre-processed pre-training data set, multiple initial models are selected according to actual application requirements, such as the Transformer model, LSTM model, and RWKV model, and the pre-processed pre-training data set is used to pre-train multiple initial models to obtain multiple pre-trained models. The pre-training process of multiple initial models can refer to the pre-training process of the existing neural network model, which will not be described in detail here.

[0052] In practical applications, the preprocessed pre-training data set can be used to pre-train multiple initial models at the same time to obtain multiple pre-trained models, or the preprocessed pre-training data set can be used to pre-train each of the multiple initial models separately to obtain multiple pre-trained models.

[0053] And obtaining a post-training data set, the post-training data set includes text data in the file to be trained, binary data in the file to be trained, text images generated according to the text data in the file to be trained, binary images generated according to the binary data in the file to be trained, and detection results of the file to be trained, wherein the text data in the file to be trained includes a file import and export table, a file printable string and a file N-gram sequence (referring to a sequence of any consecutive N characters in a text), the binary data in the file to be trained refers to data stored in binary codes 0 and 1, the text images generated according to the text data in the file to be trained include statistical graphs for representing the distribution of these text data, such as a text character distribution histogram and a text word cloud graph, and the binary images generated according to the binary data in the file to be trained include binary grayscale images and statistical graphs for representing the distribution of these binary data, such as a binary byte distribution histogram, a binary byte entropy graph and a binary byte Markov graph.

[0054] In order to ensure that the post-training data in the post-training dataset effectively participates in the model post-training, after obtaining the post-training dataset, the post-training dataset can be preprocessed first, and then the preprocessed post-training dataset can be used to post-train multiple pre-trained models to obtain multiple target models.

[0055] In another optional implementation of the present embodiment, before the post-training of multiple pre-trained models using a post-training data set to obtain multiple target models, it also includes: pre-processing the post-training data set; wherein the pre-processing includes at least one of data filling processing, data deduplication processing, data ratio adjustment processing, data normalization processing and data weighting processing.

[0056] As an example, for a post-training data set, if a certain post-training data is incomplete, the post-training data is filled in. For example, if the number of characters in the file N-gram sequence is less than N, the file N-gram sequence is filled with characters to make the number of characters reach N; if a certain post-training data is repeated with other post-training data, one of these repeated post-training data is selected and retained to avoid data redundancy in the post-training data set; after dividing multiple categories of post-training data according to classification criteria such as data type, the data volume ratio of the multiple categories of post-training data can be adjusted, and some post-training data can be randomly eliminated to meet the data volume ratio, further avoiding data redundancy in the post-training data set; if the data format of a certain post-training data is not a standard format that can be recognized and processed by the user terminal, the post-training data is converted into a standard format that can be recognized and processed by the user terminal; if a certain post-training data is important data, the importance of the post-training data can be considered and the post-training data can be given a corresponding weight so that the pre-trained model focuses on the important data.

[0057] After obtaining the preprocessed post-training data set, the preprocessed post-training data set is used to post-train multiple pre-trained models to obtain multiple target models to achieve model alignment. The post-training process of multiple pre-trained models can refer to the post-training process of the existing neural network model, which will not be described in detail here.

[0058] In practical applications, a variety of post-training methods can be selected, such as fine-tuning methods and reinforcement learning methods. For each post-training method, a post-training dataset is obtained, and each post-training dataset is used to post-train multiple pre-trained models to obtain multiple target models to achieve model alignment.

[0059] For example, assuming that a first post-training data set is obtained for the fine-tuning method, and a second post-training data set is obtained for the reinforcement learning method, the first post-training data set can be first used to perform multiple fine-tuning on multiple pre-trained models to obtain multiple fine-tuned models, and then the second post-training data set can be used to continue multiple reinforcement learning training on the multiple fine-tuned models to obtain multiple final reinforcement learning trained models, that is, multiple target models; the second post-training data set can be first used to perform multiple reinforcement learning training on multiple pre-trained models to obtain multiple reinforcement learning trained models, and then the first post-training data set can be used to continue multiple fine-tuning on the multiple reinforcement learning trained models to obtain multiple final fine-tuned models, that is, multiple target models; the first post-training data set can also be first used to perform single fine-tuning on multiple pre-trained models to obtain multiple one-round fine-tuned models, and then the second post-training data set can be used to continue single reinforcement learning training on multiple fine-tuned models to obtain multiple two-round reinforcement learning trained models, and this alternating operation is performed until the cumulative number of post-training times reaches a preset post-training times threshold, and multiple final post-trained models, that is, multiple target models, are obtained.

[0060] After obtaining multiple target models, a multi-agent system is constructed according to the multiple target models to complete the construction of the multi-agent system.

[0061] The embodiment of the present application uses a pre-training data set to pre-train multiple initial models, and uses a post-training data set to post-train the multiple pre-trained models obtained, and then builds a multi-agent system based on the obtained multiple target models, which can ensure the collaborative operation of the multi-agent system and accurately detect the security of files.

[0062] In an optional embodiment, the multi-agent system also includes a fourth agent; after the first agent is used to generate multiple tasks according to the detection instruction, it also includes: using the fourth agent to monitor whether the processing flow of the first agent is abnormal, and obtaining a first monitoring result; when the first monitoring result is abnormal, sending the first monitoring result to the first agent to re-adopt the first agent to generate multiple new tasks according to the detection instruction; and / or, after the second agent is used to process the feature extraction task and extract the multimodal features in the file to be detected, it also includes: using the fourth agent to monitor whether the processing flow of the second agent is abnormal, and obtaining a second monitoring result; when the second monitoring result is abnormal, sending the second monitoring result to the second agent to re-adopt the second agent to process the feature extraction task.

[0063] As an example, according to actual application requirements, in order to ensure that the first agent can correctly plan multiple tasks for the files to be detected according to expected requirements, a fourth agent can be added to the multi-agent system.

[0064] After the first intelligent agent generates multiple tasks according to the detection instruction, the fourth intelligent agent is used to monitor whether the processing flow of the first intelligent agent is abnormal to obtain a first monitoring result. If the first monitoring result shows that the processing flow of the first intelligent agent is abnormal, it is considered that the multiple tasks generated by the first intelligent agent do not meet the expected requirements, and subsequent operations are temporarily not allowed at this time. The first monitoring result is sent to the first intelligent agent to re-adopt the first intelligent agent to generate multiple new tasks according to the detection instruction; if the first monitoring result shows that the processing flow of the first intelligent agent is normal, it is considered that the multiple tasks generated by the first intelligent agent meet the expected requirements, and subsequent operations are allowed at this time.

[0065] In actual applications, multiple reference tasks can be pre-configured, and a fourth agent can be used to monitor the processing flow of the first agent. The multiple tasks generated by the first agent can be compared with the multiple reference tasks. If the multiple tasks generated by the first agent are different from the multiple reference tasks, the first monitoring result is determined to be that the processing flow of the first agent is abnormal; otherwise, the first monitoring result is determined to be that the processing flow of the first agent is normal.

[0066] According to actual application requirements, in order to ensure that the second agent can correctly extract the multimodal features in the file to be detected as expected, a fourth agent can also be added to the multi-agent system.

[0067] After the second agent extracts the multimodal features in the file to be detected, the fourth agent is used to monitor whether the processing flow of the second agent is abnormal, and a second monitoring result is obtained. If the second monitoring result shows that the processing flow of the second agent is abnormal, it is considered that the multimodal features extracted by the second agent do not meet the expected requirements. At this time, subsequent operations are temporarily not allowed to be performed, and the second monitoring result is sent to the second agent to re-use the second agent to process the feature extraction task; if the second monitoring result shows that the processing flow of the first agent is normal, it is considered that the multimodal features extracted by the second agent meet the expected requirements, and subsequent operations are allowed to be performed.

[0068] In practical applications, reference multimodal features can be pre-configured, and a fourth agent can be used to detect the processing flow of the second agent. The multimodal features extracted by the second agent are compared with the reference multimodal features. If the multimodal features extracted by the second agent are different from the reference multimodal features, the second monitoring result is determined to be that the processing flow of the second agent is abnormal; otherwise, the second monitoring result is determined to be that the processing flow of the second agent is normal.

[0069] It should be noted that the use of the fourth intelligent agent to monitor whether the processing flow of the first intelligent agent is abnormal, and / or the use of the fourth intelligent agent to monitor whether the processing flow of the second intelligent agent is abnormal, includes the following three situations: only using the fourth intelligent agent to monitor whether the processing flow of the first intelligent agent is abnormal; only using the fourth intelligent agent to monitor whether the processing flow of the second intelligent agent is abnormal, and using the fourth intelligent agent to monitor whether the processing flow of the first intelligent agent and the processing flow of the second intelligent agent are abnormal.

[0070] In practical applications, in order to avoid interrupting the file security detection process for a long time and improve the efficiency of file security detection, an end condition can be set in advance, such as the end condition that the cumulative number of processing times reaches a preset processing number threshold. If the cumulative number of processing times of the first agent and / or the second agent reaches the preset processing number threshold, but the result of the last processing still does not meet the expected requirements, subsequent operations are allowed to be performed.

[0071] The embodiment of the present application can further ensure the security of accurate detection of files by adding a fourth agent in the multi-agent system, using the fourth agent to monitor whether the processing flow of the first agent and / or the second agent is abnormal, and re-using the first agent and / or the second agent to perform corresponding processing when the processing flow of the first agent and / or the second agent is abnormal.

[0072] In an optional embodiment, the multimodal features include text features and image features; the extraction of multimodal features in the file to be detected includes: parsing the file to be detected to obtain text data and binary data in the file to be detected; extracting features of the text data to obtain text features; generating a target image based on the text data and binary data; extracting features of the target image to obtain image features.

[0073] As an example, a second intelligent agent is used to process feature extraction tasks, parse the file to be detected, and obtain text data and binary data in the file to be detected, wherein the text data in the file to be detected includes file import and export tables, file printable strings and file N-gram sequences, and the binary data in the file to be detected refers to data stored in binary codes 0 and 1.

[0074] After obtaining the text data and binary data in the file to be detected, the features of the text data are extracted to obtain text features, and a target image is generated based on the text data and the binary data, and the features of the target image are extracted to obtain image features, thereby obtaining multimodal features, wherein the multimodal features include text features and image features.

[0075] The embodiment of the present application uses a second intelligent agent to take the features of the text data in the file to be detected as text features, generates a target image based on the text data and binary data in the file to be detected, and takes the features of the target image as image features, thereby being able to comprehensively extract multimodal features in the file to be detected.

[0076] In an optional embodiment, the generating of a target image based on text data and binary data includes: generating a text image based on the text data; wherein the text image includes a statistical graph for representing the distribution of the text data; generating a binary image based on the binary data; wherein the binary image includes at least one of a statistical graph and a binary grayscale image for representing the distribution of the binary data; and reconstructing the text image and the binary image to obtain the target image.

[0077] As an exemplary embodiment, a second intelligent agent is used to generate a text image based on the text data after obtaining the text data and binary data in the file to be detected, wherein the text image includes a statistical graph for representing the distribution of text data, such as a text character distribution histogram and a text word cloud graph.

[0078] And a binary image is generated according to the binary data, wherein the binary image includes a binary grayscale image and a statistical graph for representing the distribution of the binary data, such as a binary byte distribution histogram, a binary byte entropy graph and a binary byte Markov graph.

[0079] After obtaining the text image and the binary image, the text image and the binary image are reconstructed to obtain the target image.

[0080] In practical applications, a super-resolution image reconstruction method or other image reconstruction methods can be used to reconstruct text images and binary images to obtain a target image.

[0081] Since existing methods for converting binary data into images rely on a uniform and fixed image width, this conversion can easily lead to loss or distortion of feature information. Therefore, by reconstructing the text image and the binary image, the influence of the image width can be removed without ensuring that the image width is uniformly fixed.

[0082] The embodiment of the present application uses a second intelligent agent to generate a text image based on the text data in the file to be detected, generates a binary image based on the binary data in the file to be detected, and then reconstructs the text image and the binary image to obtain a target image. The target image can be accurately generated to ensure comprehensive and accurate extraction of multimodal features in the file to be detected.

[0083] In an optional embodiment, the multimodal features include text features and image features; generating understanding information of the multimodal features includes: fusing the text features and the image features to obtain target features; and generating understanding information according to the target features.

[0084] As an exemplary embodiment, a third intelligent agent is used to process a feature understanding task, and text features and image features are fused to obtain target features, and understanding information is generated according to the target features.

[0085] By first fusing text features and image features into target features and then generating understanding information based on the target features, we can comprehensively analyze and understand the data characteristics described by multimodal features and accurately generate understanding information.

[0086] The embodiment of the present application uses a third intelligent agent to first fuse text features and image features into target features, and then generates understanding information based on the target features, thereby ensuring accurate generation of understanding information of multimodal features.

[0087] In an optional embodiment, the multi-agent system further includes a fifth agent; the method further includes step S106: S106. Use the fifth agent to generate process explanation information according to the processing flow of the multi-agent system; and / or use the fifth agent to generate result explanation information according to the detection results.

[0088] As an example, according to actual application requirements, in order to provide explanatory information for the file security detection process to be detected and to provide the file security detection process and results transparently to users, a fifth agent may be added to the multi-agent system.

[0089] After obtaining the detection result of the file to be detected, the fifth agent is used to generate process explanation information according to the processing flow of multiple agents in the multi-agent system, wherein the process explanation information includes input data, output data, various steps and flow chart of the processing flow.

[0090] After obtaining the detection result of the file to be detected, the fifth agent may also be used to generate result explanation information according to the detection result of the file to be detected, wherein the result explanation information is used to explain the reason for determining whether the file to be detected carries malware.

[0091] It should be noted that the use of the fifth agent to generate process explanation information according to the processing flow of the multi-agent system, and / or the use of the fifth agent to generate result explanation information according to the detection results, includes the following three situations: only using the fifth agent to generate process explanation information according to the processing flow of the multi-agent system; only using the fifth agent to generate result explanation information according to the detection results; both using the fifth agent to generate process explanation information according to the processing flow of the multi-agent system and using the fifth agent to generate result explanation information according to the detection results.

[0092] The embodiments of the present application add a fifth agent to the multi-agent system, use the fifth agent to generate process explanation information according to the processing flow of the multi-agent system, and / or use the fifth agent to generate result explanation information according to the detection results, thereby providing users with explanatory information on the security detection process of the file to be detected.

[0093] In an optional embodiment, the method further includes steps S107-S108: S107, according to the detection results obtained within a preset period, and the user feedback information corresponding to the detection results, statistically analyzing the evaluation index values; S108. When the evaluation index value is less than the preset index threshold, update the multi-agent system.

[0094] As an example, according to actual application requirements, a cycle is preset and set. The user terminal waits to obtain the detection results of multiple files to be detected within the preset cycle, and calculates the evaluation index value according to the detection results obtained within the preset cycle and user feedback information corresponding to the detection results.

[0095] After obtaining the evaluation index value, the evaluation index value is compared with the preset index threshold. If the evaluation index value is less than the preset index threshold, it is considered that the detection accuracy of the multi-agent system does not meet the standard, and the multi-agent system should be updated at this time. If the evaluation index value is greater than or equal to the preset index threshold, it is considered that the detection accuracy of the multi-agent system meets the standard, and the multi-agent system may not be updated at this time.

[0096] In practical applications, the detection accuracy can be used as an evaluation indicator. For example, assuming that 10 detection results are obtained within a preset period, 5 detection results correspond to user feedback information that the detection is correct, and 5 detection results correspond to user feedback information that the detection is wrong, then the detection accuracy of the multi-agent system is 50%, which is less than the preset indicator threshold of 80%. At this time, the multi-agent system should be updated.

[0097] The embodiment of the present application can effectively ensure the security of the multi-agent system by accurately detecting files by statistically evaluating the value of an evaluation index based on the detection results obtained within a preset period and the user feedback information corresponding to the detection results, and updating the multi-agent system when the evaluation index value is less than a preset index threshold.

[0098] In order to more clearly illustrate a file security detection device provided by the first embodiment of the present application, a multi-agent system is pre-constructed, and the multi-agent system includes a first agent, a second agent, a third agent, a fourth agent, and a fifth agent. The structural diagram of the multi-agent system is as follows: Figure 2 As shown, based on the multi-agent system, the flowchart of applying the file security detection method is as follows Figure 3 shown.

[0099] Please see Figure 4 , Figure 4 A schematic diagram of the structure of a file security detection device provided for the second embodiment of the present application. The second embodiment of the present application provides a file security detection device, including: an instruction acquisition module 201, used to acquire detection instructions for a file to be detected and call a multi-agent system; wherein the multi-agent system includes a first agent, a second agent and a third agent; a task planning module 202, used to use the first agent to generate multiple tasks according to the detection instructions; wherein the multiple tasks include feature extraction tasks, feature understanding tasks and security judgment tasks; a feature extraction module 203, used to use the second agent to process the feature extraction task, and extract multimodal features from the file to be detected; a feature understanding module 204, used to use the third agent to process the feature understanding task, and generate understanding information of the multimodal features; a security judgment module 205, used to use the second agent to process the security judgment task, and determine the detection result of the file to be detected according to the understanding information.

[0100] In an optional embodiment, the device also includes: a system construction module, which is used to pre-train multiple initial models using a pre-training data set to obtain multiple pre-trained models before calling the multi-agent system; post-train the multiple pre-trained models using a post-training data set to obtain multiple target models; and construct the multi-agent system based on the multiple target models.

[0101] In an optional embodiment, the multi-agent system also includes a fourth agent; the device also includes: a process monitoring module, which is used to use the fourth agent to monitor whether the processing flow of the first agent is abnormal after the first agent is used to generate multiple tasks according to the detection instruction, and obtain a first monitoring result; if the first monitoring result is abnormal, send the first monitoring result to the first agent to re-adopt the first agent to generate multiple new tasks according to the detection instruction; and / or, after the second agent is used to process the feature extraction task and extract the multimodal features in the file to be detected, use the fourth agent to monitor whether the processing flow of the second agent is abnormal to obtain a second monitoring result; if the second monitoring result is abnormal, send the second monitoring result to the second agent to re-adopt the second agent to process the feature extraction task.

[0102] In an optional embodiment, the multimodal features include text features and image features; the extraction of multimodal features in the file to be detected includes: parsing the file to be detected to obtain text data and binary data in the file to be detected; extracting features of the text data to obtain text features; generating a target image based on the text data and binary data; extracting features of the target image to obtain image features.

[0103] In an optional embodiment, the generating of a target image based on text data and binary data includes: generating a text image based on the text data; wherein the text image includes a statistical graph for representing the distribution of the text data; generating a binary image based on the binary data; wherein the binary image includes at least one of a statistical graph and a binary grayscale image for representing the distribution of the binary data; and reconstructing the text image and the binary image to obtain the target image.

[0104] In an optional embodiment, the multimodal features include text features and image features; generating understanding information of the multimodal features includes: fusing the text features and the image features to obtain target features; and generating understanding information according to the target features.

[0105] In an optional embodiment, the multi-agent system also includes a fifth agent; the device also includes: an overall interpretation module, which uses the fifth agent to generate process interpretation information according to the processing flow of the multi-agent system; and / or uses the fifth agent to generate result interpretation information according to the detection results.

[0106] In an optional embodiment, the device also includes: a system update module, which is used to statistically calculate the evaluation index value based on the detection results obtained within a preset period and user feedback information corresponding to the detection results; when the evaluation index value is less than a preset index threshold, the multi-agent system is updated.

[0107] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.

[0108] Please see Figure 5 , Figure 5 The third embodiment of the present application provides an electronic device 30, comprising a processor 301, a memory 302, and a computer program stored in the memory 302 and configured to be executed by the processor 301; the memory 302 is coupled to the processor 301, and when the processor 301 executes the computer program, the method described in the first embodiment of the present application is implemented and the same beneficial effects can be achieved.

[0109] When the processor 301 reads the computer program from the memory 302 through the bus 303 and executes the computer program, the method of any embodiment included in the method described in the first embodiment of the present application can be implemented.

[0110] Processor 301 can process digital signals and can include various computing structures, such as complex instruction set computer structure, reduced instruction set computer structure, or a structure that implements a combination of multiple instruction sets. In some examples, processor 301 can be a microprocessor.

[0111] The memory 302 may be used to store instructions executed by the processor 301 or data related to the execution of instructions. These instructions and / or data may include codes for implementing some or all functions of one or more modules described in the embodiments of the present application. The processor 301 of the disclosed embodiment may be used to execute instructions in the memory 302 to implement the method described in the first embodiment of the present application. The memory 302 includes a dynamic random access memory, a static random access memory, a flash memory, an optical memory, or other memory known to those skilled in the art.

[0112] The fourth embodiment of the present application provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method described in the first embodiment of the present application, and can achieve the same beneficial effects as the method.

[0113] The fifth embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed by a computer, the computer implements the method described in the first embodiment of the present application and can achieve the same beneficial effects.

[0114] The method described in the first embodiment of the present application can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed on a computer, the process or function described in each embodiment of the present application is executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM (Open Application Model) or other programmable device.

[0115] The computer program or instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program or instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a digital video disk; it may also be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.

[0116] In summary, the embodiments of the present application provide a file security detection method, device, electronic device and storage medium, the file security detection method comprising: obtaining a detection instruction for a file to be detected, and calling a multi-agent system; wherein the multi-agent system comprises a first agent, a second agent and a third agent; using the first agent to generate multiple tasks according to the detection instruction; wherein the multiple tasks comprise a feature extraction task, a feature understanding task and a security judgment task; using the second agent to process the feature extraction task, extracting multimodal features from the file to be detected; using the third agent to process the feature understanding task, generating understanding information of the multimodal features; using the second agent to process the security judgment task, determining the detection result of the file to be detected according to the understanding information. The embodiments of the present application can accurately detect the security of the file by calling the multi-agent system to work collaboratively, first extracting the multimodal features from the file to be detected, generating understanding information of the multimodal features, and then determining the detection result of the file to be detected according to the understanding information of the multimodal features.

[0117] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0118] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0119] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.

[0120] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A file security detection method, characterized in that: include: Obtaining a detection instruction for a file to be detected, and calling a multi-agent system; wherein the multi-agent system includes a first agent, a second agent, and a third agent; Using the first agent, generating a plurality of tasks according to the detection instruction; wherein the plurality of tasks include a feature extraction task, a feature understanding task and a safety judgment task; Using the second agent to process the feature extraction task to extract multimodal features from the file to be detected; Using the third agent to process the feature understanding task to generate understanding information of the multimodal feature; The second intelligent agent is used to process the security judgment task, and the detection result of the file to be detected is determined according to the understanding information.

2. The method according to claim 1, characterized in that Before calling the multi-agent system, it also includes: Pre-training multiple initial models using pre-training data sets to obtain multiple pre-trained models; Post-training the multiple pre-trained models using a post-training data set to obtain multiple target models; The multi-agent system is constructed according to the multiple target models.

3. The method according to claim 1, characterized in that: The multi-agent system further includes a fourth agent; After the first agent is used to generate a plurality of tasks according to the detection instruction, the method further includes: Using the fourth agent to monitor whether the processing flow of the first agent is abnormal, and obtaining a first monitoring result; In the case where the first monitoring result is abnormal, sending the first monitoring result to the first agent to re-adopt the first agent and generate a plurality of new tasks according to the detection instruction; and / or, After the second agent is used to process the feature extraction task and extract the multimodal features in the file to be detected, the method further includes: Using the fourth agent to monitor whether the processing flow of the second agent is abnormal, and obtaining a second monitoring result; In the case where the second monitoring result is abnormal, the second monitoring result is sent to the second agent so as to re-adopt the second agent to process the feature extraction task.

4. The method according to claim 1, characterized in that: The multimodal features include text features and image features; The extracting of multimodal features from the file to be detected includes: Parsing the file to be detected to obtain text data and binary data in the file to be detected; Extracting features of the text data to obtain the text features; generating a target image according to the text data and the binary data; Extract the features of the target image to obtain the image features.

5. The method according to claim 4, characterized in that The step of generating a target image according to the text data and the binary data comprises: Generating a text image according to the text data; wherein the text image includes a statistical graph for representing the distribution of the text data; Generate a binary image according to the binary data; wherein the binary image includes at least one of a statistical graph and a binary grayscale graph for representing the distribution of the binary data; The text image and the binary image are reconstructed to obtain the target image.

6. The method according to claim 1, characterized in that The multimodal features include text features and image features; The generating the understanding information of the multimodal feature includes: Fusing the text feature and the image feature to obtain a target feature; The understanding information is generated according to the target feature.

7. The method according to claim 1, characterized in that The multi-agent system further includes a fifth agent; The method further comprises: Using the fifth agent to generate process explanation information according to the processing flow of the multi-agent system; and / or, The fifth agent is used to generate result explanation information based on the detection result.

8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: According to the detection results obtained within a preset period and user feedback information corresponding to the detection results, statistical evaluation index values ​​are calculated; When the evaluation index value is less than a preset index threshold, the multi-agent system is updated.

9. A file security detection device, characterized in that: include: An instruction acquisition module, used to acquire a detection instruction for a file to be detected and call a multi-agent system; wherein the multi-agent system includes a first agent, a second agent and a third agent; A task planning module, configured to use the first agent to generate a plurality of tasks according to the detection instruction; wherein the plurality of tasks include a feature extraction task, a feature understanding task and a safety judgment task; A feature extraction module, used to use the second agent to process the feature extraction task and extract multimodal features from the file to be detected; A feature understanding module, configured to process the feature understanding task using the third agent to generate understanding information of the multimodal feature; A security judgment module is used to use the second agent to process the security judgment task and determine the detection result of the file to be detected according to the understanding information.

10. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; the memory is coupled to the processor, and the processor implements the method according to any one of claims 1 to 8 when executing the computer program.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 8.