Risk document security processing method and device, equipment and storage medium

By combining artificial intelligence algorithms and sandbox technology in document security processing, analyzing and processing the risk levels and dynamic behavior of target documents, the problem that existing technologies cannot efficiently handle complex dynamic content and potential threats is solved, and a more efficient level of document security processing and automation is achieved.

CN120068055APending Publication Date: 2025-05-30BEIJING THUNDERSTONE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510133000.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing document security processing solutions cannot efficiently handle complex dynamic content and potential threats, and mainly rely on static analysis and manual inspection, and cannot deeply identify and handle dynamic behaviors.

Method used

By analyzing the risk level of the target document using an artificial intelligence algorithm before running it, and creating a sandbox environment based on the risk level. In a sandbox environment, the properties of dynamic behavior are identified through artificial intelligence algorithms and the corresponding processing strategies are determined for safe processing.

Benefits of technology

It improves the ability to identify and process threats in documents, improves the automation level of document isolation processing, and is suitable for handling complex dynamic documents and high security needs scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068055A_ABST
    Figure CN120068055A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides a safety processing method for a risk document, and the method comprises the steps: analyzing a target document through an artificial intelligence algorithm before the target document is operated, and determining the risk level of the target document; creating a corresponding sandbox environment for the target document according to the risk level of the target document; based on the risk level of the target document and the dynamic behavior in the sandbox environment of the target document, identifying the property of the dynamic behavior through an artificial intelligence algorithm; according to the property of the dynamic behavior of the target document in the sandbox environment of the target document, determining a corresponding processing strategy to perform security processing; and outputting the target document subjected to security processing. According to the technical scheme, the capability of recognizing and processing the threats in the document can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, equipment and storage medium for secure processing of risk documents. Background Art

[0002] With the diversification of document types (such as PDF, Word, Excel, etc.) and the introduction of various complex functions (such as embedded macros, forms, buttons, scripts, etc.), the security processing solutions for documents have also received attention. These security processing solutions include inspection, isolation and conversion of risk documents. However, the existing document security processing solutions mainly rely on static analysis and manual inspection by humans, and cannot efficiently handle complex dynamic content and potential threats. Summary of the Invention

[0003] The present application provides a method, device, equipment and storage medium for secure processing of risk documents, which can improve the ability to identify and process threats in documents.

[0004] On the one hand, the present application provides a method for secure processing of risk documents, the method comprising:

[0005] Before running a target document, analyzing the target document through an artificial intelligence algorithm to determine the risk level of the target document;

[0006] Creating a corresponding sandbox environment for the target document according to the risk level of the target document;

[0007] Identifying the nature of the dynamic behavior based on the risk level of the target document and its dynamic behavior in the sandbox environment through an artificial intelligence algorithm;

[0008] Determining a corresponding processing strategy for secure processing according to the nature of the dynamic behavior of the target document in its sandbox environment;

[0009] Outputting the target document that has undergone secure processing.

[0010] On the other hand, the present application provides a device for secure processing of risk documents, the device comprising:

[0011] A determination module, configured to analyze the target document through an artificial intelligence algorithm before running the target document to determine the risk level of the target document;

[0012] A creation module, configured to create a corresponding sandbox environment for the target document according to the risk level of the target document;

[0013] An identification module, configured to identify the nature of the dynamic behavior through an artificial intelligence algorithm based on the risk level of the target document and its dynamic behavior in its sandbox environment;

[0014] A processing module, configured to determine a corresponding processing strategy for security processing according to the nature of the dynamic behavior of the target document in its sandbox environment;

[0015] An output module, configured to output the target document that has undergone security processing.

[0016] In a third aspect, the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the technical solution of the security processing method of the risk document as described above are implemented.

[0017] In a fourth aspect, the present application provides a storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the technical solution of the security processing method of the risk document as described above are implemented.

[0018] As can be seen from the technical solutions provided by the present application above, after determining the risk level of the target document, according to the risk level of the target document, a corresponding sandbox environment is created for the target document. Based on the risk level of the target document and its dynamic behavior in its sandbox environment, the nature of the dynamic behavior is identified through an artificial intelligence algorithm. According to the nature of the dynamic behavior of the target document in its sandbox environment, a corresponding processing strategy is determined for security processing. Since the sandbox technology has deficiencies or limitations in the in-depth identification and processing of dynamic content and internal threats, its advantage is that it can run the target document in an isolated environment, thereby preventing potential threats from directly affecting the main system or network. And the artificial intelligence technology can just make up for the deficiencies or limitations of the sandbox technology. Therefore, the technical solution of the present application combines the sandbox technology and the artificial intelligence algorithm, which not only improves the ability to identify and process threats in the document, but also improves the automation level of document isolation processing, and is applicable to scenarios of processing complex dynamic documents and high-security requirements. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a flowchart of the security processing method of the risk document provided by the embodiment of the present application;

[0021] Figure 2 It is a schematic structural diagram of a security processing device for risk documents provided by an embodiment of the present application;

[0022] Figure 3 It is a schematic structure of an electronic device provided by an embodiment of the present application. Specific embodiments

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present application.

[0024] In this specification, adjectives such as first and second can only be used to distinguish one element or action from another element or action, and do not necessarily require or imply any actual such relationship or order. Where circumstances permit, reference elements or components or steps (etc.) should not be construed as limited to only one of the elements, components, or steps, but can be one or more of the elements, components, or steps, etc.

[0025] In this specification, for ease of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0026] With the diversification of document types (such as PDF, Word, Excel, etc.) and the introduction of various complex functions (such as embedded macros, forms, buttons, scripts, etc.), security processing solutions for documents have also received attention. These security processing solutions include inspection, isolation, and conversion of risk documents, etc. However, existing document security processing solutions mainly rely on static analysis and manual inspection by humans, and cannot efficiently process complex dynamic content and potential threats. Even for sandbox technology, its isolation mechanism usually only observes and restricts the overall environmental behavior of the file, and there are limitations in the in-depth identification and processing of dynamic content and internal threats. For example, a sandbox can monitor whether there is a request to execute code, but may not be able to analyze the complex dynamic content inside the document line by line or module by module, especially when malicious code is embedded in macros, scripts, or compressed packages and requires specific triggering conditions to be activated; again, the dynamic content inside the document (such as macros, scripts, embedded objects) may require specific triggering conditions (such as user clicks or specific operations) to exhibit malicious behavior, and a sandbox usually detects the file within a short period of time and it is difficult to capture these complex triggering conditions and corresponding threat behaviors, and so on.

[0027] In view of the above problems of the prior art, the present application proposes a method for securely processing risk documents, and its flowchart is as shown in Appendix Figure 1 as follows, mainly including steps S101 to S105, which are described in detail as follows:

[0028] Step S101: Before running the target document, analyze the target document through an artificial intelligence algorithm to determine the risk level of the target document.

[0029] Since the target document is analyzed through an artificial intelligence algorithm before running the target document, this analysis essentially belongs to the static analysis of the target document, that is, the content of the target document is analyzed without running the target document, mainly focusing on features such as the structure, grammar, keywords, and embedded objects of the document, rather than relying on the actual behavior and runtime state of the target document, so that potential risks and problems can be quickly identified in the preliminary screening. As an embodiment of the present application, analyzing the target document through an artificial intelligence algorithm to determine the risk level of the target document can be achieved through steps S1011 to S1013, and the detailed description is as follows:

[0030] Step S1011: Extract the features of each object in the target document.

[0031] In the embodiment of the present application, the objects in the target document include the content itself of the document, metadata, and embedded code, etc. Among them, metadata is information such as the creation time, author, and modification history of the document, and the embedded code includes dynamic elements such as macros, scripts, and embedded links in the target document. For the object of the content itself of the document, keyword extraction, sentiment analysis, topic model extraction, etc. can be performed on it to generate a content feature vector; for the object of metadata, classification and coding can be performed on it to generate a metadata feature vector; and for the embedded code, by analyzing information such as the corresponding event type, interaction frequency, behavior time, and event correlation, these information are converted into numerical features, thereby generating a behavior feature vector; these vectors can be used as the features of the corresponding objects in this embodiment.

[0032] Step S1012: Input the features of each object into the trained artificial intelligence model, and the trained artificial intelligence model evaluates the risk score values of each object in the target document.

[0033] In the embodiment of the present application, the trained artificial intelligence model can be obtained by training a model based on machine learning algorithms such as decision trees, support vector machines, or random forests using historical behavior data. The specific training scheme is as follows:

[0034] First, divide the historical behavior data, i.e., the dataset, into a training set and a test set. For example, divide the training set and the test set in a ratio of 7:3 or 8:2, i.e., Training set:Test set = 7:3 or 8:2.

[0035] Secondly, start training the model. Taking the decision tree-based model as an example, the information gain IG can be used to select the splitting point for training. The information gain formula is as follows:

[0036]

[0037] where H(D) is the entropy of the dataset D, D v is the subset with value v on feature A, and D and D v are the sample numbers of the dataset D and the subset D v respectively.

[0038] Thirdly, evaluate the performance of the model through the test set, and calculate metrics such as accuracy, precision, recall, and F1-score. For example, use the Confusion Matrix for evaluation. The Confusion Matrix is an important tool for evaluating the performance of a classification model. By showing the correspondence between the model's prediction results and the actual situation, it is more intuitive to understand the performance of the model. The core of the Confusion Matrix is to compare the prediction results with the actual results and count the number of different types of predictions. The Confusion Matrix is a 2×2 matrix, and an example is as follows:

[0039]

[0040] where TP (True Positive) is the true positive example, representing the number of examples that are actually positive and are predicted as positive; FN (False Negative) is the false negative example, representing the number of examples that are actually positive but are predicted as negative; FP (False Positive) is the false positive example, representing the number of examples that are actually negative but are predicted as positive; and TN (True Negative) is the true negative example, representing the number of examples that are actually negative and are predicted as negative.

[0041] Finally, further improve the model performance by adjusting the hyperparameters or using more complex models (such as grid search, random search, etc.) until the performance of the model reaches the expectation or the number of iterations of the hyperparameters reaches the preset number, and then stop training the model to obtain the trained artificial intelligence model.

[0042] Step S1013: Weight the risk score values of each object in the target document based on the weight coefficients of the features of each object, and determine the risk level corresponding to the weighted sum value as the risk level of the target document.

[0043] Suppose the risk score values of the content feature vector, metadata feature vector, and behavior feature vector evaluated by the trained artificial intelligence model are C, M, and D respectively, and the weight coefficients of C, M, and D are w C , w M and w D respectively. Then, the risk score values of each object in the target document are weighted based on the weight coefficients of the features of each object, and the weighted sum value S = w C *C + w M *M + w D *D. Then, according to the correspondence between the risk score value of the document and the risk level, the risk level corresponding to S is determined as the risk level of the target document. As for the above weight coefficients w C , w M and w D , they can be implemented through steps S10131 to S10133, and the detailed description is as follows:

[0044] Step S10131: Set the initial values of the weight coefficients w C , w M and w D according to historical data and / or expert opinions.

[0045] Step S10132: In practical applications, continuously collect the performance data of the risk scores of various feature vectors, and use the collected performance data to train a machine learning model (such as linear regression, neural network, etc.) to identify the impact of each factor on the final risk assessment.

[0046] In the embodiments of the present application, the performance data of the risk scores of various feature vectors can be actual results, prediction accuracies, feedback data, time series data, etc. Among them, the actual results can be the actual risk events and their severities that occur, the prediction accuracy is the difference between the risk score predicted by the machine learning model and the actual result, the feedback data is the feedback obtained from users or systems using the machine learning model, and these feedbacks can help evaluate the accuracy and practicality of the model, while the time series data is the change situation of different risk factors over time. After training the machine learning model (such as linear regression, neural network, etc.) with the collected performance data, an optimized machine learning model is obtained. This model can evaluate the influence of each risk factor (that is, identify and quantify each input factor, such as the factors corresponding to w C , w M and w D on the final risk score) and the improvement of the prediction ability (that is, by continuous learning and adjustment, improve the prediction accuracy of the model for future risks). Specifically, the machine learning model will output the weight coefficients of each factor, and these coefficients will tell the user the importance of each risk factor to the overall risk assessment.

[0047] Step S10133: Regularly adjust the weight coefficients according to the results output by the machine learning model to reflect the changes in the importance of different factors in risk assessment.

[0048] The data of various risk factors at different time points, the occurrence of actual risk events corresponding to these time points and their impacts, etc. are input into the machine learning model. The machine learning model will output the weight values corresponding to each risk factor, and these values represent the contribution degrees of each factor to the final risk assessment. Adjust the current weight coefficients according to the results output by the machine learning model. For example, if the output of the machine learning model shows that the importance of w C increases while the importance of w M and w D decreases, then the value of w C can be increased accordingly, and the values of w M and w D can be decreased. Use the adjusted weight coefficients for risk assessment, and continue to collect new performance data to verify whether the adjusted model has improved the prediction accuracy. If the effect is not ideal, iterative adjustment can be continued. Since the risk environment is dynamically changing, it is necessary to regularly retrain the machine learning model, adjust the weight coefficients according to the latest data, and ensure that the machine learning model can always reflect the current risk situation.

[0049] Step S102: Create a corresponding sandbox environment for the target document according to the risk level of the target document.

[0050] When selecting a sandbox environment, different configuration and resource allocation strategies need to be adopted for target documents with different risk levels. The differences in the requirements of target documents with different risk levels for the sandbox environment are mainly reflected in aspects such as security, performance, resource consumption, and monitoring requirements. Therefore, creating a corresponding sandbox environment for the target document according to the risk level of the target document can be to create a sandbox environment by configuring resources and security mechanisms corresponding to its risk level for the target document. In principle, for target documents with a low risk level (i.e., the risk level is lower than a preset first threshold), a basic sandbox environment is configured, emphasizing low resource consumption and basic security protection, mainly focusing on performance and efficiency, suitable for large-scale and rapid document processing. For target documents with a high risk level (i.e., the risk level is higher than a preset second threshold, where the second threshold is higher than the first threshold), a high-security sandbox environment is configured, paying attention to resource investment and in-depth monitoring mechanisms to ensure that potential malicious behaviors can be comprehensively identified and prevented, applicable to scenarios with extremely high security requirements. For target documents with a medium risk level (i.e., the risk level is higher than the preset first threshold and lower than the preset second threshold), it is necessary to balance security and performance, avoid overprotection in high-risk documents, and also avoid simplified strategies in low-risk documents, and configure a moderately secure sandbox environment through reasonable resource allocation and security mechanisms.

[0051] Step S103: Based on the risk level of the target document and its dynamic behavior in the sandbox environment, identify the nature of the dynamic behavior through an artificial intelligence algorithm.

[0052] In the embodiment of the present application, the dynamic behavior of the target document in the sandbox environment refers to the behaviors or activities triggered by the program or document when the target document runs in the sandbox environment. These behaviors may include file operations, network communications, system calls, process creation, and dynamic code execution behaviors, etc. Correspondingly, the analysis of the dynamic behavior of the target document in its sandbox environment is relative to the aforementioned static analysis and essentially belongs to dynamic analysis, that is, the analysis is carried out when the target document is running, and can capture the actual behavior and runtime state of the target document. As an embodiment of the present application, identifying the nature of the dynamic behavior through an artificial intelligence algorithm based on the risk level of the target document and its dynamic behavior in the sandbox environment can be achieved through steps S1031 to S1033, which are described in detail as follows:

[0053] Step S1031: Based on real-time monitoring, extract the dynamic behavior characteristics of the target document in its sandbox environment.

[0054] In the embodiments of the present application, the dynamic behavior characteristics of the target document in its sandbox environment include network behavior characteristics, endogenous behavior characteristics, resource usage characteristics, etc. That is, based on real-time monitoring, the dynamic behavior characteristics of the target document in its sandbox environment are extracted, including: by capturing and analyzing all network traffic when the target document runs in its sandbox environment, calculating the network behavior characteristics of the target document in its sandbox environment; matching the endogenous behavior of the target document when it runs in its sandbox environment with malicious behaviors to obtain the endogenous behavior characteristics of the target document when it runs in its sandbox environment; and by real-time monitoring the resource consumption data when the target document runs in its sandbox environment, calculating the resource usage characteristics of the target document in its sandbox environment. The following will be described in detail respectively.

[0055] 1) By capturing and analyzing all network traffic when the target document runs in its sandbox environment, calculate the network behavior characteristics of the target document in its sandbox environment. Specifically, when the target document runs in its sandbox environment, use a network traffic monitoring tool to capture all outgoing network packets, including the following traffic data: source IP address, destination IP address, transport protocol, destination port number, and the amount of transmitted data, etc.; according to the captured traffic data, calculate network behavior characteristics such as connection frequency (i.e., the number of network connections initiated per unit time) and data transmission rate (i.e., the amount of data transmitted per unit time (e.g., bytes / second)); by calculating the network behavior characteristics, if the network behavior of a certain process is significantly different from the normal behavior, it can be determined as a malicious behavior. For example, when the data transmission rate suddenly surges, it may be that a ransomware is encrypting or uploading data.

[0056] 2) Match the endogenous behavior of the target document when it runs in its sandbox environment with malicious behaviors to obtain the endogenous behavior characteristics of the target document when it runs in its sandbox environment. Specifically, some open-source sandbox tools can be used to monitor the endogenous behaviors such as file system activities, process behaviors, and system calls when the target document runs in its sandbox environment, that is, record the files created, modified, and deleted during the execution of the target document and their quantities, the processes started during the execution of the target document, including process names, process trees, parent-child process relationships, etc., and the system calls generated during the execution process, especially the system calls involving sensitive operations such as privilege escalation, network access, and file operations. Then, extract the characteristics of the target document when performing the above endogenous behaviors, and match the characteristics of the target document when performing the above endogenous behaviors with the characteristics of malicious documents. If the match is successful, determine the characteristics corresponding to the endogenous behavior as the endogenous behavior characteristics of the target document when it runs in its sandbox environment.

[0057] 3) By monitoring the resource consumption data of the target document in real time when it is running in its sandbox environment, calculate the resource usage characteristics of the target document in its sandbox environment. Specifically, according to the monitored data, calculate the CPU peak value (i.e., the maximum value of the CPU usage rate per unit time), memory consumption (i.e., the peak value of the memory usage per unit time), and disk I / O frequency (i.e., the number of disk read and write operations per unit time) and other resource usage characteristics of the target document when it is running in its sandbox environment. If the resource consumption of the target document exceeds a certain threshold, it is marked as an abnormal behavior. For example, when the CPU usage rate is higher than the set threshold, it is determined to be abnormal; another example is that if the memory consumption or disk read and write volume is too high, it is regarded as a potential malicious behavior (especially in large-scale ransomware attacks, there is usually extremely high resource consumption).

[0058] Step S1032: Integrate the risk level of the target document and its dynamic behavior characteristics in its sandbox environment to obtain the multi-modal behavior characteristics of the target document.

[0059] Specifically, integrating the risk level of the target document and its dynamic behavior characteristics in its sandbox environment can be in the way of weighted sum, that is, multiply the weight corresponding to the risk level of the target document by its corresponding feature vector, add the multiplication of the weight corresponding to the dynamic behavior characteristics of the target document in its sandbox environment by its corresponding feature vector, and construct a comprehensive feature vector - the multi-modal behavior characteristics of the target document.

[0060] Step S1033: Input the multi-modal behavior characteristics of the target document into a preset machine learning classification model, and let the machine learning classification model learn the multi-modal behavior characteristics to obtain the nature of the dynamic behavior of the target document in its sandbox environment.

[0061] In the embodiments of the present application, the machine learning classification model can be a trained decision tree, support vector machine SVM, random forest and other classification models. These classification models can identify the nature of the dynamic behavior of the target document when it is running in its sandbox environment according to the input behavior characteristics, including the execution of malicious code, the execution of macros, scripts or embedded objects, and low-risk automated operations, and so on.

[0062] It should be noted that, as can be seen from the technical solutions of the above embodiments, creating a corresponding sandbox environment for the target document is based on the risk level of the target document, and the risk level of the target document is determined by analyzing the target document through an artificial intelligence algorithm before running the target document. As mentioned above, analyzing the target document through an artificial intelligence algorithm before running the target document essentially belongs to static analysis of the target document. As the target document runs in its sandbox environment, the sandbox environment configured based on static analysis in the early stage may no longer be suitable. Therefore, after obtaining the multi-modal behavior characteristics of the target document by integrating the risk level of the target document and its dynamic behavior characteristics in its sandbox environment, the following may also be included: dynamically adjusting the resource and security mechanism configurations of the sandbox environment based on the multi-modal behavior characteristics of the target document.

[0063] Step S104: Determine a corresponding processing strategy for security processing according to the nature of the dynamic behavior of the target document in its sandbox environment.

[0064] Specifically, determining a corresponding processing strategy for security processing according to the nature of the dynamic behavior of the target document in its sandbox environment may be: analyzing the nature of the dynamic behavior of the target document in its sandbox environment; if there are activities corresponding to malicious code when the target document runs in its sandbox environment, abort the activities corresponding to the malicious code; if there are activities corresponding to macros, scripts, or embedded objects when the target document runs in its sandbox environment, isolate the macros, scripts, or embedded objects; if there are low-risk automated operations when the target document runs in its sandbox environment, perform a security conversion on the objects corresponding to the low-risk automated operations. In the above embodiments, the so-called low-risk automated operations may be activities such as simple text replacement without external calls, format adjustment, and automatic embedding of non-executable objects. For these low-risk automated operations, the objects corresponding to these low-risk automated operations, such as simple text, scripts, etc., can be converted into comments, or pictures can be converted into embedded static images to eliminate risks without changing the content of the target document.

[0065] Step S105: Output the target document after security processing.

[0066] From the above appendix Figure 1According to the method for securely processing a risk document in the example, after determining the risk level of the target document, a corresponding sandbox environment is created for the target document according to the risk level of the target document. Based on the risk level of the target document and its dynamic behavior in the sandbox environment, the nature of the dynamic behavior is identified through an artificial intelligence algorithm. According to the nature of the dynamic behavior of the target document in its sandbox environment, a corresponding processing strategy is determined for secure processing. Since the sandbox technology has deficiencies or limitations in the in-depth identification and processing of dynamic content and internal threats, its advantage lies in being able to run the target document in an isolated environment, thereby preventing potential threats from directly affecting the main system or network. And the artificial intelligence technology can just make up for the deficiencies or limitations of the sandbox technology. Therefore, the technical solution of this application combines the sandbox technology and the artificial intelligence algorithm, which not only improves the ability to identify and process threats in the document, but also improves the automation level of document isolation processing, and is applicable to scenarios of processing complex dynamic documents and high security requirements.

[0067] Please refer to the attached Figure 2 , which is a security processing device for risk documents provided by an embodiment of this application. The device may include a determination module 201, a creation module 202, an identification module 203, a processing module 204, and an output module 205, which are described in detail as follows:

[0068] The determination module 201 is configured to analyze the target document through an artificial intelligence algorithm before running the target document to determine the risk level of the target document;

[0069] The creation module 202 is configured to create a corresponding sandbox environment for the target document according to the risk level of the target document;

[0070] The identification module 203 is configured to identify the nature of the dynamic behavior through an artificial intelligence algorithm based on the risk level of the target document and its dynamic behavior in the sandbox environment;

[0071] The processing module 204 is configured to determine a corresponding processing strategy for secure processing according to the nature of the dynamic behavior of the target document in its sandbox environment;

[0072] The output module 205 is configured to output the target document that has been securely processed.

[0073] From the above attached Figure 2As can be seen from the security processing device for the exemplary risk document, after determining the risk level of the target document, a corresponding sandbox environment is created for the target document according to the risk level of the target document. Based on the risk level of the target document and its dynamic behavior in the sandbox environment, the nature of the dynamic behavior is identified through an artificial intelligence algorithm. According to the nature of the dynamic behavior of the target document in its sandbox environment, a corresponding processing strategy is determined for security processing. Since the sandbox technology has deficiencies or limitations in the in-depth identification and processing of dynamic content and internal threats, its advantage lies in being able to run the target document in an isolated environment, thereby preventing potential threats from directly affecting the main system or network. And the artificial intelligence technology can just make up for the deficiencies or limitations of the sandbox technology. Therefore, the technical solution of this application combines the sandbox technology and the artificial intelligence algorithm, which not only improves the ability to identify and process threats in the document, but also improves the automation level of document isolation processing, and is applicable to scenarios of processing complex dynamic documents and high-security requirements.

[0074] Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 3 shown, the electronic device 3 in this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program for the security processing method of the risk document. When the processor 30 executes the computer program 32, the steps in the above-mentioned embodiment of the security processing method of the risk document are implemented, such as Figure 1 the steps S101 to S105 shown. Alternatively, when the processor 30 executes the computer program 32, the functions of each module / unit in the above-mentioned device embodiments are implemented, such as Figure 2 the functions of the determination module 201, the creation module 202, the identification module 203, the processing module 204, and the output module 205 shown.

[0075] Exemplarily, the computer program 32 for the secure processing method of risk documents mainly includes: before running the target document, analyzing the target document through an artificial intelligence algorithm to determine the risk level of the target document; creating a corresponding sandbox environment for the target document according to the risk level of the target document; identifying the nature of the dynamic behavior through an artificial intelligence algorithm based on the risk level of the target document and its dynamic behavior in the sandbox environment; determining a corresponding processing strategy for secure processing according to the nature of the dynamic behavior of the target document in its sandbox environment; and outputting the target document after secure processing. The computer program 32 can be divided into one or more modules / units, and one or more modules / units are stored in the memory 31 and executed by the processor 30 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into the functions of a determination module 201, a creation module 202, an identification module 203, a processing module 204, and an output module 205 (modules in the virtual device). The specific functions of each module are as follows: The determination module 201 is used to analyze the target document through an artificial intelligence algorithm before running the target document to determine the risk level of the target document; the creation module 202 is used to create a corresponding sandbox environment for the target document according to the risk level of the target document; the identification module 203 is used to identify the nature of the dynamic behavior through an artificial intelligence algorithm based on the risk level of the target document and its dynamic behavior in the sandbox environment; the processing module 204 is used to determine a corresponding processing strategy for secure processing according to the nature of the dynamic behavior of the target document in its sandbox environment; the output module 205 is used to output the target document after secure processing.

[0076] The electronic device 3 may include but is not limited to the processor 30 and the memory 31. Those skilled in the art can understand that Figure 3 merely examples of the electronic device 3 do not constitute a limitation on the electronic device 3, and it may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0077] The so-called processor 30 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0078] The memory 31 may be an internal storage unit of the electronic device 3, such as the hard disk or memory of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk equipped on the electronic device 3, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 31 may also include both the internal storage unit of the electronic device 3 and the external storage device. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 may also be used to temporarily store data that has been output or is to be output.

[0079] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above-mentioned device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0080] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0081] Those of ordinary skill in the art will recognize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0082] In the embodiments provided in this application, it should be understood that the disclosed devices / apparatuses and methods can be implemented in other ways. For example, the device / apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0083] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0084] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0085] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by a computer program instructing relevant hardware. The computer program of the security processing method for risk documents can be stored in a storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented, that is, splitting the target document by page, and allocating agents to each split page; the agent accepting the page allocation detects the embedded resources of the page it is allocated; if the agent accepting the page allocation detects that there are risk elements on the page it is allocated, then remove the risk elements and convert the page with risk elements into a page in a secure format; if the agent accepting the page allocation detects that there is an error on the page it is allocated, then notify the error to other agents, and each agent that learns of the error generates an error correction plan for the error page through a shared knowledge base; use the error correction plan to correct the page with the error, and convert the corrected page into a page in a secure format; recombine the pages in the secure format into a complete document for output. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The storage medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.

[0086] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application. The specific implementation manners described above further elaborate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above is only the specific implementation manners of the present application, and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should all be included in the protection scope of the present invention.

Claims

1. A method for securely processing risk documents, characterized in that: The method comprises: Before running the target document, the target document is analyzed by an artificial intelligence algorithm to determine the risk level of the target document; Creating a corresponding sandbox environment for the target document according to the risk level of the target document; Based on the risk level of the target document and its dynamic behavior in its sandbox environment, identifying the nature of the dynamic behavior through an artificial intelligence algorithm; Determine a corresponding processing strategy for security processing according to the nature of the dynamic behavior of the target document in its sandbox environment; Output the target document after security processing.

2. The method for securely processing risk documents according to claim 1, characterized in that: The step of analyzing the target document by an artificial intelligence algorithm to determine the risk level of the target document includes: Extracting features of each object in the target document; Inputting the characteristics of each object into a trained artificial intelligence model, and using the trained artificial intelligence model to evaluate the risk score value of each object; The risk score values ​​of the objects in the target document are weighted based on the weight coefficients of the features of the objects, and the risk level corresponding to the weighted sum is determined as the risk level of the target document.

3. The method for securely processing risk documents according to claim 1, characterized in that: The step of creating a corresponding sandbox environment for the target document according to the risk level of the target document includes: According to the risk level of the target document, a sandbox environment is created by configuring resources and security mechanisms corresponding to the risk level for the target document.

4. The method for securely processing risk documents according to claim 1, characterized in that: The method of identifying the nature of the dynamic behavior based on the risk level of the target document and the dynamic behavior in the sandbox environment thereof by an artificial intelligence algorithm includes: Based on real-time monitoring, extract the dynamic behavior characteristics of the target document in its sandbox environment; The risk level of the target document and the dynamic behavior characteristics in the sandbox environment are integrated to obtain the multimodal behavior characteristics of the target document; The multimodal behavior features of the target document are input into a preset machine learning classification model, and the multimodal behavior features are learned by the machine learning classification model to obtain the properties of the dynamic behavior of the target document in its sandbox environment.

5. The method for securely processing risk documents according to claim 4, characterized in that: The dynamic behavior characteristics include network behavior characteristics, endogenous behavior characteristics and resource usage characteristics; the extraction of the dynamic behavior characteristics of the target document in its sandbox environment based on real-time monitoring includes: By capturing and analyzing all network traffic when the target document is running in its sandbox environment, the network behavior characteristics of the target document in its sandbox environment are calculated; Matching the intrinsic behavior of the target document when it is running in its sandbox environment with the malicious behavior to obtain the intrinsic behavior characteristics of the target document when it is running in its sandbox environment; and By real-time monitoring of resource consumption data of the target document when it is running in its sandbox environment, resource usage characteristics of the target document in its sandbox environment are calculated.

6. The method for securely processing risk documents according to claim 4, characterized in that: After fusing the risk level of the target document and the dynamic behavior characteristics in the sandbox environment to obtain the multimodal behavior characteristics of the target document, the method further includes: Based on the multimodal behavior characteristics of the target document, the resource and security mechanism configuration of the sandbox environment is dynamically adjusted.

7. The method for securely processing risk documents according to any one of claims 1 to 6, characterized in that: Determining a corresponding processing strategy for security processing according to the nature of the dynamic behavior of the target document in its sandbox environment includes: Analyzing the nature of the dynamic behavior of the target document in its sandbox environment; If the analysis shows that the target document has activities corresponding to malicious code when running in its sandbox environment, the activities corresponding to the malicious code are terminated; If the analysis shows that the target document has activities corresponding to macros, scripts or embedded objects when running in its sandbox environment, the macros, scripts or embedded objects are isolated; If the analysis shows that the target document has a preset low-risk automated operation when running in its sandbox environment, the object corresponding to the preset low-risk automated operation is securely converted.

8. A device for securely processing risk documents, characterized in that: The device comprises: A determination module, used to analyze the target document by artificial intelligence algorithm before running the target document to determine the risk level of the target document; A creation module, used to create a corresponding sandbox environment for the target document according to the risk level of the target document; An identification module, configured to identify the nature of the dynamic behavior through an artificial intelligence algorithm based on the risk level of the target document and the dynamic behavior in its sandbox environment; A processing module, used to determine a corresponding processing strategy for security processing according to the nature of the dynamic behavior of the target document in its sandbox environment; The output module is used to output the target document after security processing.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.