Method and device for extracting macro code from composite document

By analyzing the structure of the composite document, the target offset information in the structure that records the macro code offset information is extracted, and the problem of difficult to extract malicious macro codes that do not use known feature string identification in the prior art is solved, and efficient macro code extraction and detection are achieved.

CN120180431APending Publication Date: 2025-06-20QI AN XIN TECHNOLOGY GROUP INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510179824.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to extract malicious macro code from composite documents that does not use known feature string identification.

Method used

By analyzing the structure of the composite document, the target offset information in the structure that records the macro code offset information is extracted, so that the macro code is directly extracted from the data stream without relying on the feature string of the macro code.

Benefits of technology

It realizes the effective extraction of macro code from composite documents without relying on feature strings, improving the detection ability of malicious macro code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180431A_ABST
    Figure CN120180431A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for extracting a macro code from a composite document, relates to the technical field of network security, and mainly aims to extract the macro code from the composite document without depending on a feature character string for identifying the macro code. According to the main technical scheme, the method comprises the following steps: recording offset information of a macro code in a data stream of the composite document by a structural body of the composite document; obtaining a target composite document of a to-be-extracted macro code; analyzing a structural body of the target composite document, and extracting target offset information; and based on the target offset information, extracting a macro code from a data stream of the target composite document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technologies, and particularly to a method and device for extracting macro code from a compound document. Background Art

[0002] A macro virus is a malicious macro code that can be stored in a compound document in a format such as OLE (Object Linking and Embedding). When the compound document carrying the macro virus is opened, the malicious behavior of the macro virus will be executed, thereby affecting the security of the device where the compound document is located. Therefore, usually, the macro code in the compound document is extracted to determine whether the compound document contains a macro virus by detecting malicious features of the macro code.

[0003] Currently, macro code is usually extracted from a compound document based on well-known macro code feature strings in the industry (such as Attribut). Specifically, the feature string is searched for in the data stream of the compound document, and the data after the feature string in the data stream is extracted as the macro code. However, some malicious parties may not use well-known feature strings to identify the macro code when storing malicious macro code in the compound document, resulting in difficulty in extracting the macro code from the compound document based on well-known feature strings. Summary of the Invention

[0004] In view of this, this application provides a method and device for extracting macro code from a compound document, and the main purpose is to extract macro code from a compound document without relying on a feature string for identifying the macro code.

[0005] To achieve the above objective, this application mainly provides the following technical solutions:

[0006] In a first aspect, this application provides a method for extracting macro code from a compound document. The compound document records the offset information of the macro code in the data stream of the compound document in a structure. The method for extracting macro code from the compound document includes: obtaining a target compound document from which macro code is to be extracted; parsing the structure of the target compound document to extract target offset information; and extracting the macro code from the data stream of the target compound document based on the target offset information.

[0007] In some embodiments of the present application, the offset information includes data stream information for indicating the data stream where the macro code is located and a byte offset for indicating the macro code relative to the start byte of the data stream. Then, based on the target offset information, extracting the macro code from the data stream of the target compound document includes: determining the target data stream where the macro code is located in the target compound document based on the data stream information included in the target offset information; and extracting the macro code from the target data stream based on the byte offset included in the target offset information.

[0008] In some embodiments of the present application, the method for extracting the macro code from the compound document further includes: allocating a storage area for the target data stream in a preset memory based on the data volume of the target data stream; copying the target data stream to the storage area, and then extracting the macro code from the target data stream based on the byte offset included in the target offset information according to the target data stream in the storage area.

[0009] In some embodiments of the present application, the compound document includes a first data stream for recording a structure directory, and each structure included in the compound document has a corresponding first field in the structure directory; each type of compound document has a corresponding preset first field, and the structure corresponding to the preset first field is used to record the offset information of the macro code in the data stream of the compound document.

[0010] Then, the method for extracting the macro code from the compound document further includes: determining the preset first field corresponding to the target compound document.

[0011] Then, parsing the structure of the target compound document to extract the target offset information includes: parsing the first data stream of the target compound document to obtain the structure directory of the target compound document; determining the first structure corresponding to the preset first field in the target compound document based on the structure directory; and extracting the target offset information from the first structure.

[0012] In some embodiments of the present application, the structure for recording the offset information of the macro code in the data stream of the compound document includes at least one second structure corresponding to a second field, the second field is used to indicate the macro code, and the second structure records the data stream information for indicating the data stream where the macro code indicated by the second field is located and the byte offset for indicating the macro code relative to the start byte of the data stream. Then, extracting the target offset information from the first structure includes: determining the second fields included in the first structure; for each of the second fields, extracting the data stream information and the byte offset recorded in the second structure corresponding to the second field as the target offset information.

[0013] In some embodiments of the present application, the structure for recording the offset information of the macro code in the data stream of the compound document in the compound document further includes the target quantity, and the target quantity is used to indicate the total number of macro codes in the compound document. Then, the method for extracting the macro code from the compound document further includes: determining whether the target quantity included in the first structure is zero; if not zero, performing the step of determining the second field included in the first structure; if zero, prompting that the target compound document does not contain macro code.

[0014] In some embodiments of the present application, the method for extracting the macro code from the compound document further includes: detecting whether the target compound document includes all specified data streams, where the specified data streams are the data streams that a compound document containing macro code needs to have; if it includes, performing the step of parsing the structure of the target compound document to extract the target offset information; if it does not include, prompting that the target compound document does not contain macro code.

[0015] In some embodiments of the present application, each data stream of the compound document has a corresponding data stream name, and the number of the specified data streams is at least one. Then, detecting whether the target compound document includes all specified data streams includes: matching the data stream name of each specified data stream with the data stream name of the data stream of the target compound document; if the data stream names of all specified data streams are successfully matched, it is detected that the target compound document includes all specified data streams; if the data stream name of any one of the specified data streams fails to be matched successfully, it is detected that the target compound document does not include all specified data streams.

[0016] In some embodiments of the present application, each data stream of the compound document is used to record data of a corresponding data category, and the number of the specified data streams is at least one. Then, detecting whether the target compound document includes all specified data streams includes: detecting whether there are matching data streams for all the specified data streams in the target compound document, where the specified data stream and the matching data stream record data of the same data category; if there are matching data streams for all specified data streams, it is detected that the target compound document includes all specified data streams; if there is no matching data stream for any one of the specified data streams, it is detected that the target compound document does not include all specified data streams.

[0017] In some embodiments of the present application, the specified data stream includes a second data stream, and the second data stream is used to record the code information of the code included in the compound document. Then, the method for extracting macro code from the compound document further includes: if it is detected that the target compound document includes all the specified data streams, detecting whether there is any keyword indicating macro code in the code information recorded in the second data stream of the target compound document, and the number of keywords indicating macro code is at least one; if there is, performing the step of parsing the structure of the target compound document and extracting the target offset information; if not, prompting that the target compound document does not contain macro code.

[0018] In some embodiments of the present application, the compound document is a compound document in the Object Linking and Embedding (OLE) format.

[0019] In a second aspect, the present application provides an apparatus for extracting macro code from a compound document. The compound document records the offset information of the macro code in the data stream of the compound document in a structure. The apparatus for extracting macro code from the compound document includes:

[0020] An acquisition module, configured to acquire a target compound document from which macro code is to be extracted;

[0021] A parsing module, configured to parse the structure of the target compound document and extract target offset information;

[0022] An extraction module, configured to extract macro code from the data stream of the target compound document based on the target offset information.

[0023] In a third aspect, the present application provides a computer-readable storage medium. The storage medium includes a stored program. When the program runs, it controls the device where the storage medium is located to execute the method for extracting macro code from a compound document described in the first aspect.

[0024] In a fourth aspect, the present application provides an electronic device. The storage management device includes: a memory, configured to store a program; a processor, coupled to the memory, configured to run the program to execute the method for extracting macro code from a compound document described in the first aspect.

[0025] The method and device for extracting macro code from a compound document provided by this application, when obtaining a target compound document from which macro code is to be extracted, based on the feature that the compound document records the offset information of the macro code in the data stream of the compound document in a structure, parses the structure of the target compound document to extract the target offset information corresponding to the target compound document. Then, based on the target offset information, the macro code is extracted from the data stream of the target compound document. It can be seen that the solution provided by this application, based on the feature that compound documents with macro code will record the offset information of the macro code in the data stream of the compound document in a structure, extracts the offset information of the macro code in the data stream of the compound document from the structure of the compound document from which macro code is to be extracted. In this way, it is not necessary to rely on the feature string of the macro code, and based on the offset information, the macro code can be extracted from the data stream of the compound document.

[0026] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically gives the specific implementation manners of this application. Brief Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0028] Figure 1 Shows a flowchart of a method for extracting macro code from a compound document provided by an embodiment of this application;

[0029] Figure 2 Shows a schematic structural diagram of a device for extracting macro code from a compound document provided by an embodiment of this application;

[0030] Figure 3 Shows a schematic structural diagram of a device for extracting macro code from a compound document provided by another embodiment of this application. Detailed Description of the Preferred Embodiments

[0031] The following will describe the exemplary embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully communicated to those skilled in the art.

[0032] Currently, when some malicious parties deposit malicious macro code in a compound document, any of the following situations may occur: One is to use a non-public characteristic field string to identify the macro code; the other is not to use any characteristic string to identify the macro code. When a malicious party deposits malicious macro code in a compound document, no matter which of the above situations is adopted, the existing method of extracting macro code from the compound document based on the publicly known macro code characteristic string in the industry (such as, Attribut) is difficult to extract the macro code from the compound document.

[0033] Through research, it is found that when a malicious party deposits malicious macro code into a compound document, whether or not a publicly known characteristic string is used to identify the macro code, the compound document will record the offset information of the macro code in the data stream of the compound document in a structure, so that when the compound document is opened, the malicious macro code can be executed based on the offset information in the structure. Based on this, this embodiment proposes to extract the offset information from the structure of the compound document, so that it is not necessary to rely on the characteristic string of the macro code, and the macro code can be extracted from the data stream of the compound document based on the offset information.

[0034] Based on the above discovery, this embodiment provides a technical solution for extracting macro code from a compound document. Specifically, it is to obtain the target compound document from which the macro code is to be extracted; parse the structure of the target compound document and extract the target offset information; based on the target offset information, extract the macro code from the data stream of the target compound document.

[0035] The technical solution for extracting macro code from a compound document provided by this embodiment can extract macro code from office documents such as Office and WPS, so as to identify whether there is a macro virus in the office document based on the extracted macro code.

[0036] Based on the above technical solution for extracting macro code from a compound document, this embodiment specifically provides a method and device for extracting macro code from a compound document. The method and device for extracting macro code from a compound document provided by this embodiment will be specifically described below.

[0037] As Figure 1 shown, an embodiment of the present application provides a method for extracting macro code from a compound document. The method for extracting macro code from a compound document may at least include the following steps 101 to 103:

[0038] 101. Obtain the target compound document from which the macro code is to be extracted.

[0039] The compound document in this embodiment is a compound document in the Object Linking and Embedding (OLE) format. In practical applications, the compound document includes at least, but is not limited to, the following two types: First, the compound document is a compound document in the OLE format in office documents such as Office or WPS. Exemplarily, the compound document is a document in a version of Office before 2007. For example, Office documents with suffixes such as doc, dot, xls, and xlt are all compound documents in the OLE format. Second, the compound document is an OLE-format compound document obtained by converting a non-OLE-format document in an office document such as Office or WPS. Exemplarily, documents in versions of Office after 2007 are in the OpenXML format. For example, Office documents with suffixes such as docm, dotm, and xlsm are all in the OpenXML format. In order to extract macro code from non-OLE-format office documents, these documents are converted from the OpenXML format to the OLE format.

[0040] In this embodiment, the compound document records the offset information of the macro code in the data stream of the compound document in a structure. In this way, when the compound document is opened, the execution of malicious macro code is realized based on the offset information in the structure. Based on this, this embodiment mainly extracts the offset information from the structure of the compound file, and extracts the macro code from the data stream of the compound document based on the offset information.

[0041] In this embodiment, there are three methods to obtain the target compound document from which the macro code is to be extracted: First, the compound document specified by the user to extract the macro code is obtained as the target compound document from which the macro code is to be extracted. This method can flexibly obtain the target compound document based on the user's needs. Second, it is monitored whether the electronic device receives a compound document transmitted externally. If it is monitored, the externally transmitted compound document is obtained as the target compound document from which the macro code is to be extracted. This method can timely verify whether there is a macro virus in the externally transmitted compound document by extracting the macro code. Third, it is monitored whether there is a compound document about to be opened. If there is, the opening of the compound document is intercepted, and the compound document is obtained as the target compound document from which the macro code is to be extracted. This method takes into account that only when the compound document is opened, the macro virus will be run. Therefore, in order to reduce both the workload of extracting the macro code and the probability of the macro virus being run, only the compound document about to be opened is obtained as the target compound document from which the macro code is to be extracted. At least one of the above three methods for obtaining the target compound document from which the macro code is to be extracted can be selected based on business needs. It should be noted that no matter which method for obtaining the target compound document from which the macro code is to be extracted is selected, the storage path of the target compound document will be obtained first, so as to read the target compound document based on the storage path to execute the subsequent steps.

[0042] Further, considering that not all compound documents will have macro code, after obtaining the target compound document from which the macro code is to be extracted, it is also necessary to detect whether the target compound document includes macro code, so as to terminate the subsequent steps related to macro code extraction in time when the target compound document does not include macro code, thereby avoiding unnecessary computing power consumption.

[0043] The specific process of detecting whether the target compound document includes macro code may include the following steps A to B:

[0044] A. Detect whether the target compound document includes all specified data streams; if it does, execute step 102 to parse the structure of the target compound document and extract the target offset information; if it does not, execute step B.

[0045] The specified data stream is a data stream that a compound document containing macro code needs to have. That is, for any compound document with macro code, the data streams in the compound document must include all specified data streams, that is, the compound document needs to include all specified data streams at the same time. The number of specified data streams is set based on specific services, and this embodiment does not limit this.

[0046] There are two methods for detecting whether the target compound document includes all specified data streams:

[0047] The first one is that the compound document stores data in at least one data stream, each data stream of the compound document has a corresponding data stream name, and the number of specified data streams is at least one. Then, the specific process of detecting whether the target compound document includes all specified data streams may include the following steps: match the data stream name of each specified data stream with the data stream name of the data streams of the target compound document; if the data stream names of all specified data streams match successfully, it is detected that the target compound document includes all specified data streams; if the data stream name of any specified data stream does not match successfully, it is detected that the target compound document does not include all specified data streams.

[0048] The process of matching the data stream name of each specified data stream with the data stream name of the data streams of the target compound document is essentially: for each data stream name of the specified data stream, execute to detect whether there is a data stream name in the data stream names of the data streams of the target compound document that is the same as the data stream name of this specified data stream. If there is, it is determined that the data stream name of this specified data stream matches successfully; if there is not, it is determined that the data stream name of this specified data does not match successfully.

[0049] If the data stream names of all specified data streams match successfully, it means that all specified data streams exist in the target compound document at the same time. Therefore, it is detected that the target compound document includes all specified data streams.

[0050] If the data stream name of any specified data stream fails to match successfully, it indicates that some specified data streams do not exist in the target compound document. Therefore, it is detected that the target compound document does not include all specified data streams.

[0051] Exemplarily, the number of specified data streams is three, and the names of the three specified data streams are: PROJECT, VBA\\_VBA_PROJECT, and VBA\\dir. Then, when it is detected that the data stream names of the data streams in the target compound document simultaneously include PROJECT, VBA\\_VBA_PROJECT, and VBA\\dir, it is determined that the target compound document includes all specified data streams.

[0052] Second, each data stream of the compound document is used to record data of a corresponding data category, and the number of specified data streams is at least one. Then, the specific process of detecting whether the target compound document includes all specified data streams may include the following steps: detecting whether there is a matching data stream for each specified data stream in the target compound document, and the specified data stream and the matching data stream record data of the same data category; if there is a matching data stream for all specified data streams, it is detected that the target compound document includes all specified data streams; if there is no matching data stream for any specified data stream, it is detected that the target compound document does not include all specified data streams.

[0053] Considering that the naming rules of the data streams in the compound document may be different in different versions or enterprises and institutions, it is difficult to accurately detect whether the target compound document includes all specified data streams based on the data stream names of the data streams. The data streams in the compound document are used to record data of corresponding data categories, and the data categories of the data recorded by the specified data streams do not change regardless of the naming rules. Therefore, in order to be able to detect whether the target compound document includes all specified data streams under different naming rules in this embodiment, the following technical solution is adopted: for each specified data stream, it is executed to detect whether there is a matching data stream for the specified data stream in the target compound document, and the specified data stream and the matching data stream record data of the same data category.

[0054] In this embodiment, each data stream of the compound document has corresponding configuration information, and the configuration information is used to describe the data category of the data recorded by the data stream. Therefore, the data category corresponding to each data stream of the target compound document can be determined through the configuration information of the data stream. After detection, if there is a matching data stream for all specified data streams in the target compound document, it indicates that all specified data streams exist in the target compound document at the same time, and it is detected that the target compound document includes all specified data streams. If there is no matching data stream for any specified data stream, it indicates that some specified data streams do not exist in the target compound document, and therefore it is detected that the target compound document does not include all specified data streams.

[0055] Exemplarily, the number of specified data streams is three, and the data stream names of the three specified data streams are: PROJECT, VBA\\_VBA_PROJECT, and VBA\\dir. For the specified data stream with the data stream name PROJECT, its corresponding data category is code information, that is, it is used to record the code information of the code included in the compound document. For the specified data stream with the data stream name VBA\\_VBA_PROJECT, its corresponding data category is version-related project information, that is, it is used to record the version-related project information of the compound document. For the specified data stream with the data stream name VBA\\dir, its corresponding data category is the structure directory, that is, it is used for the structure directory. When it is detected that the target compound document includes a data stream for recording code information, a data stream for recording the structure directory, and a data stream for recording version-related project information at the same time, it is determined that the target compound document includes all specified data streams.

[0056] The above two methods for detecting whether the target compound document includes all specified data streams can be selected based on business needs, and this embodiment does not make specific limitations. For example, one can be selected for use or two can be selected for use. When two are selected for use, one method is used to verify the detection result of the other method to improve the accuracy of the detection result. When the detection results of the two methods are different, a prompt can be issued so that the user can perform manual verification based on the prompt.

[0057] Further, in order to more accurately detect whether the target compound document includes macro code, after obtaining the result that the target compound document includes all specified data streams through the above method for detecting whether the target compound document includes all specified data streams, continue to perform a secondary detection on whether the target compound document includes macro code.

[0058] The specified data stream includes a second data stream, and the second data stream is used to record the code information of the code included in the compound document. That is, once there is macro code in the compound document, the corresponding information of the macro code will necessarily be recorded in the code information. Therefore, the specific steps for the secondary detection of whether the target compound document includes macro code can include: if it is detected that the target compound document includes all specified data streams, then detect whether there is any keyword indicating macro code in the code information recorded by the second data stream of the target compound document, and the number of keywords indicating macro code is at least one. If it exists, execute step 102 to parse the structure of the target compound document and extract the target offset information; if it does not exist, execute step B.

[0059] Keywords for indicating macro code can be determined based on business requirements, and this embodiment does not limit this. Exemplarily, the keyword may include at least one of the following: Document, Module, Class, BaseClass.

[0060] If one or more keywords for indicating macro code are found in the code information of the second data stream record of the detected target compound document, it indicates that the target compound document probably includes macro code. Therefore, step 102 can be continued to extract the macro code from the target compound document.

[0061] If no keyword for indicating macro code is found in the code information of the second data stream record of the detected target compound document, it indicates that the target compound document probably does not include macro code. Therefore, step B is continued to let the user know the fact that the target compound document does not include macro code.

[0062] B. Prompt that the target compound document does not contain macro code.

[0063] If it is detected that the target compound document does not include all specified data streams, it indicates that the target compound document probably does not include macro code. Then, prompt that the target compound document does not contain macro code to let the user know this fact. The way to prompt that the target compound document does not contain macro code can be determined based on business requirements, and this embodiment does not limit this. Exemplarily, a prompt in document format is sent to the specified terminal used by the user.

[0064] 102. Analyze the structure of the target compound document and extract the target offset information.

[0065] The compound document includes a first data stream for recording the structure directory. Each structure included in the compound document has a corresponding first field in the structure directory. That is to say, for any structure in the compound document, there is a first field in the structure directory corresponding to the structure for identifying the existence of the structure in the compound document. There are different types of compound documents, and each type of compound document has a corresponding preset first field. The structure corresponding to the preset first field is used to record the offset information of the macro code in the data stream of the compound document. That is, the preset first field is the first field corresponding to the structure for recording the offset information. The structure corresponding to the non-preset first field is not a structure for recording offset information.

[0066] In order to accurately locate the structure for recording offset information in the target compound document, the method for extracting macro code from the compound document further includes the following steps: determining the corresponding preset first field of the target compound document. Each type of compound document has a corresponding preset first field, and the preset first field corresponding to each type of compound document is obtained by summarizing a large number of compound documents in advance. Therefore, the corresponding preset first field of the target compound document can be directly determined based on the type of the target compound document.

[0067] After determining the corresponding preset first field of the target compound document, the steps of parsing the structure of the target compound document and extracting the target offset information can be executed. The specific execution process of this step can include the following steps 102A to 102C:

[0068] 102A. Parse the first data stream of the target compound document to obtain the structure directory of the target compound document.

[0069] Data streams are usually stored in a compressed format. Therefore, it is necessary to decompress the first data stream to be able to read the structure directory in the first data stream. Exemplarily, the first data stream can be a data stream with the data stream name VBA\\dir.

[0070] 102B. Based on the structure directory, determine the first structure corresponding to the preset first field in the target compound document.

[0071] Each first field in the structure directory has a corresponding structure in the first data stream. The structure corresponding to the preset first field records the offset information. Therefore, based on the structure directory, find the first structure corresponding to the preset first field to extract the offset information from the first structure.

[0072] Exemplarily, if the preset first field is ModulesRecord, then based on the structure directory, obtain the first structure corresponding to ModulesRecord.

[0073] 102C. Extract the target offset information from the first structure.

[0074] The structure for recording the offset information of the macro code in the data stream of the compound document in the compound document includes: at least one second structure corresponding to a second field, where the second field is used to indicate the macro code, and the second structure records the data stream information of the data stream where the macro code indicated by the second field is located and the byte offset relative to the start byte of the data stream. Therefore, the specific process of extracting the target offset information from the first structure can include the following steps: determining the second fields included in the first structure; for each second field, respectively execute, and extract the data stream information and byte offset recorded in the second structure corresponding to the second field as the target offset information.

[0075] Further, the structure for recording the offset information of the macro code in the data stream of the compound document in the compound document further includes the number of targets, and the number of targets is used to indicate the total number of macro codes in the compound document. Based on this, it is determined whether to continue to execute the subsequent steps based on the number of targets. Therefore, the method for extracting macro codes from a compound document provided in this embodiment may further include the following steps: determining whether the number of targets included in the first structure is zero; if not zero, then executing the step of determining the second field included in the first structure; if zero, then prompting that the target compound document does not contain macro codes.

[0076] Extract the number of targets from the first structure. If the number of targets is not zero, it indicates that there are macro codes in the target compound document. Therefore, in order to be able to extract the macro codes, the step of determining the second field included in the first structure is continued. If the number of targets is zero, it indicates that the probability of the existence of macro codes in the target compound document is relatively low. Therefore, it is prompted that the target compound document does not contain macro codes, so that the user can be aware of the fact that the target compound document does not contain macro codes based on the prompt.

[0077] Further, the second field is used to indicate the macro code, so the number of targets is the same as the number of second fields. In order to avoid missing extraction of the target offset information, when the number of targets is not zero, the number of targets can be used as the number of loops. Each time a loop is performed, a second field for extracting the target offset information is obtained, and for this second field, the steps of extracting the data stream information and byte offset recorded in the second structure corresponding to the second field as the target offset information are performed. After the target offset information is extracted, this second segment is removed from the second field of the target offset to be extracted. In this way, based on the loop of the target data, all the target offset information is extracted in an orderly manner, avoiding missing extraction of the target offset information.

[0078] The following uses a specific example to detail the specific process of parsing the structure of the target compound document in step 102 to extract the target offset information. First, the following characteristics of the compound document are clarified: The compound document includes a first data stream (the first data stream can be represented as a dir data stream), and the first data stream is used to record the structure directory. Each structure included in the compound document has a corresponding first field in the structure directory. If there are macro codes in the compound document, there is a structure for recording the offset information of the macro code in the data stream of the compound document, and the structure for recording the offset information of the macro code in the data stream of the compound document includes at least one second structure corresponding to the second field. The second field is used to indicate the macro code. The second structure includes a first information field for indicating the data stream information and a second information field for indicating the offset information. The first information field has a first structure sub-body for recording the data stream information, and the second information field has a second structure sub-body for recording the target offset information.

[0079] When the target compound document for which macro code is to be extracted is obtained, first determine the first data stream of the target compound document, and decompress the first data stream to obtain the structure directory of the target compound document. Each structure included in the target compound document has a corresponding first field in the structure directory.

[0080] Then determine that the corresponding preset first field of the target compound document is "ModulesRecord", and the structure corresponding to the preset first field "ModulesRecord" is used to record the offset information of the macro code in the data stream of the compound document. After determining the preset first field "ModulesRecord", based on the structure directory of the target compound document, determine that the first structure corresponding to the preset first field "ModulesRecord" in the target compound document is the "PROJECTMODULES structure".

[0081] The first structure "PROJECTMODULES structure" includes at least one second structure "MODULE structure" corresponding to the second field "Modules". The "PROJECTMODULES structure" also includes a target quantity "Count", and the target quantity "Count" is used to indicate the total number of macro codes in the target compound document. Each second field "Modules" is used to indicate a macro code, and each second structure "MODULE structure" corresponding to the second field "Modules" includes a first information field, a second information field, a first structural sub-body, and a second structural sub-body.

[0082] After detection, if the target quantity "Count" is not zero, then use Count as the number of loop iterations. And in each loop, obtain a second field "Modules" for which the target offset information is to be extracted, and perform the following for the extracted second field "Modules": determine the first information field and the second information field corresponding to the document type of the target compound document, and then extract the data stream information in the first structural sub-body corresponding to the determined first information field and the byte offset in the second structural sub-body corresponding to the determined second information field as the target offset information. When the target offset information is extracted, remove the obtained second field "Modules" from the second fields for which the target offset is to be extracted. Such a loop based on the target quantity "Count" can orderly extract all the target offset information and avoid missing extraction of the target offset information.

[0083] It should be noted that the first information field and the second information field correspond to the document type of the target compound document. Exemplarily, the first information field can be any one of NameRecord, NameUnicodeRecord, StreamNameRecord, DocStringRecord. The second information field is OffsetRecord. For example, if the first information field is NameRecord, the first structural sub-body corresponding to the first information field is the MODULENAME structure body, and the data stream information is extracted from the MODULENAME structure body. If the second information field is OffsetRecord, the second structural sub-body corresponding to the second information field is the MODULEOFFSET structure body, and the byte offset is extracted from the MODULEOFFSET structure body.

[0084] 103. Extract the macro code from the data stream of the target compound document based on the target offset information.

[0085] The offset information recorded by the structure body can clarify the following two key pieces of information: First, the data stream where the macro code exists; second, the byte offset of the macro code in the corresponding data stream. Therefore, the target offset information includes: the data stream information used to indicate the data stream where the macro code is located and the byte offset used to indicate the macro code relative to the start byte of the data stream. Here, the data stream information can be the data stream name. There can be two situations for the byte offset: One is that when the last byte of the macro code is the last byte of the data stream where the macro code is located, the byte offset is only the byte offset of the start byte of the macro code relative to the start byte of the data stream. The other is that when the last byte of the macro code is not the last byte of the data stream where the macro code is located, the byte offset includes the byte offset of the start byte of the macro code relative to the start byte of the data stream and the byte offset of the last byte of the macro code relative to the start byte node of the data stream.

[0086] After extracting the target offset information from the target compound document, perform the step of extracting the macro code from the data stream of the target compound document based on the target offset information. The specific execution process of this step can include the following steps 103A to 103B:

[0087] 103A. Determine the target data stream where the macro code is located in the target compound document based on the data stream information included in the target offset information.

[0088] The data stream information is used to indicate the data stream where the macro code is located. When extracting the macro code from the target compound document, it is necessary to determine the target data stream where the macro code is located in the target compound document based on the data stream information included in the target offset information, so as to extract the macro code from the target data stream in a targeted manner.

[0089] 103B. Extract macro code from the target data stream based on the byte offset included in the target offset information.

[0090] After determining the target data stream where the macro code is located in the target compound document, extract the macro code from the target data stream based on the byte offset included in the target offset information. The process of extracting the macro code is related to the specific situation of the byte offset, so it includes the following two cases:

[0091] The first case is that when the last character of the macro code is the last byte of the data stream where the macro code is located, and the byte offset is only the byte offset of the start byte of the macro code relative to the start byte of the data stream, then the start byte of the macro code can be located based on the byte offset, and the data from the start byte of the macro code to the last byte of the target data stream is extracted as the macro code.

[0092] The second case is that the last byte of the macro code is not the last byte of the data stream where the macro code is located, and the byte offset includes the byte offset of the start byte of the macro code relative to the start byte of the data stream and the byte offset of the last byte of the macro code relative to the start byte of the data stream node. Then, based on the byte offset, the start byte and the last byte of the macro code are located in the target data stream, and the data from the start byte of the macro code to the last byte is extracted as the macro code.

[0093] It should be noted that when extracting the macro code from the target data stream, if the macro code exists in the target data stream in plain text form, the macro code can be directly extracted based on the byte offset. If the macro code exists in the target data stream in cipher text form, the cipher text is first converted to plain text, and then the macro code is extracted based on the byte offset. Exemplarily, if the macro code is stored in the data stream in a compressed format, the data stream is first decompressed, and after the decompressed plain text is obtained, the macro code is extracted based on the byte offset.

[0094] Furthermore, there may be more than one macro code in some compound documents. Therefore, the number of target offset information extracted from the target compound document may be multiple. Each of the target offset information here has a corresponding macro code. Therefore, it is necessary to extract the macro code from the target compound document for each target offset information.

[0095] Furthermore, considering that directly extracting the macro code from the target compound document will inevitably invade the target compound document, and the invasion may damage some data in the target compound document. Therefore, in order to reduce the invasion of the target compound file, the following technical solution is proposed in this embodiment. After the above step 103A, allocate a storage area for the target data stream in the preset memory based on the data volume of the target data stream; copy the target data stream to the storage area to execute step 103B based on the target data stream in the storage area.

[0096] The method for obtaining the data volume of the target data stream may include: obtaining the corresponding pointer of the target data stream, and using a preset data volume obtaining method according to the pointer to obtain the data volume of the target data stream. Exemplarily, when the compound document is an OLE format office document, the IStorage::OpenStream method can be used to obtain the pointer of the target data stream, and the IStream::Start method is used according to the pointer to obtain the data volume of the target data stream. After obtaining the data volume of the target data stream, allocate a storage area in a preset memory (such as memory) that can sufficiently store the target data stream, and then read the data in the target data stream through the IStream::Read method and write it into the storage area, thereby completing the copy of the target data stream.

[0097] The target data stream copied to the storage area has been separated from the target compound document. Therefore, performing step 103B based on the target data stream in the storage area will neither invade the target compound document nor prevent the macro code from being extracted from the target data stream.

[0098] After extracting the macro code from the target compound document, the macro code can be provided to a specified security device so that the security device can detect whether the macro code is a macro virus and obtain a detection result indicating whether the macro code is a macro virus. Then, corresponding operations are performed on the target compound document based on the detection result. For example, when the detection result indicates that none of the macro codes extracted from the target compound document are macro viruses, it means that the target compound document is safe, and no opening and usage restrictions need to be imposed on the target compound document. For example, when the detection result indicates that there are macro viruses in the macro codes extracted from the target compound document, it means that the target compound document is unsafe. Therefore, the target compound document can be restricted from being opened and isolated to prevent the macro virus from infecting the device where the target compound document is located.

[0099] In the method for extracting macro code from a compound document provided by the embodiments of the present application, when obtaining the target compound document from which the macro code is to be extracted, based on the feature that the compound document records the offset information of the macro code in the data stream of the compound document in a structure, the structure of the target compound document is parsed to extract the target offset information corresponding to the target compound document. Then, based on the target offset information, the macro code is extracted from the data stream of the target compound document. It can be seen that the solution provided by the present application extracts the offset information of the macro code in the data stream of the compound document from the structure of the compound document from which the macro code is to be extracted based on the feature that all compound documents with macro codes record the offset information of the macro code in the data stream of the compound document in a structure. In this way, it is not necessary to rely on the feature string of the macro code, and the macro code can be extracted from the data stream of the compound document based on the offset information.

[0100] Further, another embodiment of the present application further provides an apparatus for extracting macro code from a composite document. The composite document records the offset information of the macro code in the data stream of the composite document in a structure, as Figure 2 shown. The apparatus for extracting macro code from a composite document may include:

[0101] An acquisition module 21, configured to acquire a target composite document from which macro code is to be extracted;

[0102] A parsing module 22, configured to parse the structure of the target composite document and extract target offset information;

[0103] An extraction module 23, configured to extract macro code from the data stream of the target composite document based on the target offset information.

[0104] For the apparatus for extracting macro code from a composite document provided by the embodiment of the present application, when acquiring a target composite document from which macro code is to be extracted, according to the feature that the composite document records the offset information of the macro code in the data stream of the composite document in a structure, the structure of the target composite document is parsed, and the target offset information corresponding to the target composite document is extracted. Then, based on the target offset information, macro code is extracted from the data stream of the target composite document. It can be seen that the solution provided by the present application extracts the offset information of the macro code in the data stream of the composite document from the structure of the composite document to be extracted for macro code according to the feature that any composite document with macro code will record the offset information of the macro code in the data stream of the composite document in a structure. In this way, it is not necessary to rely on the feature string of the macro code, and the macro code can be extracted from the data stream of the composite document based on the offset information.

[0105] In some embodiments of the present application, as Figure 3 shown, the offset information includes data stream information for indicating the data stream where the macro code is located and a byte offset for indicating the offset of the macro code relative to the start byte of the data stream. Then, the extraction module 23 includes:

[0106] A first determination unit 231, configured to determine the target data stream where the macro code is located in the target composite document based on the data stream information included in the target offset information;

[0107] A first extraction unit 232, configured to extract macro code from the target data stream based on the byte offset included in the target offset information.

[0108] In some embodiments of the present application, as Figure 3 shown, the extraction module 23 further includes:

[0109] An allocation unit 233, configured to allocate a storage area for the target data stream in a preset memory based on the data volume of the target data stream determined by the first determination unit 231;

[0110] A copy unit 234, configured to copy the target data stream to the storage area, and extract macro code from the target data stream based on the byte offset included in the target offset information according to the target data stream in the storage area.

[0111] In some embodiments of the present application, as Figure 3 shown, the compound document includes a first data stream for recording a structure directory, and each structure included in the compound document has a corresponding first field in the structure directory; each type of compound document has a corresponding preset first field, and the structure corresponding to the preset first field is used to record the offset information of the macro code in the data stream of the compound document. Then, the apparatus for extracting macro code from the compound document further includes:

[0112] A determination module 24, configured to determine the preset first field corresponding to the target compound document;

[0113] Then, the parsing module 22 includes:

[0114] A parsing unit 221, configured to parse the first data stream of the target compound document to obtain the structure directory of the target compound document;

[0115] A second determination unit 222, configured to determine the first structure corresponding to the preset first field in the target compound document based on the structure directory;

[0116] A second extraction unit 223, configured to extract the target offset information from the first structure.

[0117] In some embodiments of the present application, as Figure 3 shown, the structure for recording the offset information of the macro code in the data stream of the compound document in the compound document includes at least one second structure corresponding to a second field, the second field is used to indicate the macro code, and the second structure records the data stream information for indicating the data stream where the macro code indicated by the second field is located and the byte offset relative to the start byte of the data stream. Then, the second extraction unit 223 includes:

[0118] A determination subunit 2231, configured to determine the second field included in the first structure;

[0119] An extraction subunit 2232, configured to, for each second field, extract the data stream information and byte offset recorded by the second structure corresponding to the second field as the target offset information.

[0120] In some embodiments of the present application, as Figure 3As shown in the figure, the structure for recording the offset information of macro code in the data stream of a compound document in the compound document further includes the number of targets, and the number of targets is used to indicate the total number of macro codes in the compound document. Then, the second extraction unit 223 includes:

[0121] A judgment subunit 2233, configured to judge whether the number of targets included in the first structure is zero; if not zero, trigger the determination subunit 2231 to execute the step of determining the second field included in the first structure; if zero, trigger the prompt subunit 2234 to prompt that the target compound document does not contain macro code.

[0122] In some embodiments of the present application, as Figure 3 shown, the apparatus for extracting macro code from a compound document further includes:

[0123] A detection module 25, configured to detect whether the target compound document includes all specified data streams, where the specified data streams are the data streams that a compound document containing macro code needs to have; if included, trigger the parsing module 22 to execute the step of parsing the structure of the target compound document and extracting the target offset information; if not included, trigger the prompt module 26 to prompt that the target compound document does not contain macro code.

[0124] In some embodiments of the present application, as Figure 3 shown, each data stream of the compound document has a corresponding data stream name, and the number of the specified data streams is at least one. Then, the detection module 25 includes:

[0125] A first detection unit 251, configured to match the data stream name of each specified data stream with the data stream name of the data stream of the target compound document; if the data stream names of all specified data streams are successfully matched, it is detected that the target compound document includes all specified data streams; if the data stream name of any one specified data stream fails to be matched, it is detected that the target compound document does not include all specified data streams.

[0126] In some embodiments of the present application, as Figure 3 shown, each data stream of the compound document is used to record data of a corresponding data category, and the number of the specified data streams is at least one. Then, the detection module 25 includes:

[0127] A second detection unit 252, configured to detect whether there is a matching data stream for each specified data stream in the target compound document, where the specified data stream and the matching data stream record data of the same data category; if there is a matching data stream for all specified data streams, it is detected that the target compound document includes all specified data streams; if there is no matching data stream for any one specified data stream, it is detected that the target compound document does not include all specified data streams.

[0128] In some embodiments of the present application, as Figure 3 shown, the specified data stream includes a second data stream, and the second data stream is used to record the code information of the code included in the composite document. Then, the detection module 25 is further configured to, if it is detected that the target composite document includes all the specified data streams, detect whether there is any keyword for indicating macro code in the code information recorded by the second data stream of the target composite document, and the number of the keywords for indicating macro code is at least one; if so, trigger the parsing module 22 to execute the step of parsing the structure of the target composite document and extracting the target offset information; if not, trigger the prompting module 26 to prompt that the target composite document does not include macro code.

[0129] In some embodiments of the present application, the composite document is a composite document in the Object Linking and Embedding (OLE) format.

[0130] In the apparatus for extracting macro code from a composite document provided by the embodiments of the present application, the detailed explanations adopted during the operation of each functional module can refer to the corresponding explanations in the above method embodiments for extracting macro code from a composite document, and will not be elaborated here.

[0131] Further, an embodiment of the present application further provides a computer-readable storage medium. The storage medium includes a stored program. When the program runs, it controls the device where the storage medium is located to execute the above method for extracting macro code from a composite document.

[0132] Further, an embodiment of the present application further provides an electronic device. The storage management device includes: a memory for storing a program; a processor coupled to the memory for running the program to execute the above method for extracting macro code from a composite document.

[0133] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not elaborated in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0134] It can be understood that the relevant features in the above methods and apparatuses can be referred to each other. In addition, the "first", "second", etc. in the above embodiments are used to distinguish the respective embodiments, and do not represent the superiority or inferiority of the respective embodiments.

[0135] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.

[0136] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. A variety of general-purpose systems may also be used in conjunction with the teachings based herein. The structure required to construct such systems will be apparent from the above description. In addition, the present application is not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using a variety of programming languages, and the description of a particular language above is for the purpose of disclosing the preferred embodiments of the present application.

[0137] In addition, the memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0138] Those skilled in the art will appreciate that the embodiments of the present application may be provided as a method, system, or computer program product. Accordingly, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0139] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0140] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.

[0142] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0143] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0144] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0145] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.

[0146] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0147] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for extracting macro codes from a compound document, characterized in that: The compound document records the offset information of the macro code in the data stream of the compound document in a structure, and the method includes: Obtain a target compound document to be extracted macro codes; Parsing the structure of the target compound document to extract target offset information; Based on the target offset information, macro codes are extracted from the data stream of the target compound document.

2. The method according to claim 1, characterized in that The offset information includes data stream information for indicating the data stream where the macro code is located and a byte offset for indicating the macro code relative to the starting byte of the data stream. Then, based on the target offset information, extracting the macro code from the data stream of the target compound document includes: Determine the target data stream where the macro code is located in the target compound document based on the data stream information included in the target offset information; A macro code is extracted from the target data stream based on the byte offset included in the target offset information.

3. The method according to claim 2, characterized in that The method further comprises: Allocating a storage area for the target data stream in a preset memory based on the data volume of the target data stream; The target data stream is copied to the storage area, so as to extract macro codes from the target data stream based on the byte offset included in the target offset information according to the target data stream in the storage area.

4. The method according to claim 1, characterized in that: The compound document includes a first data stream, the first data stream is used to record a structure directory, each structure included in the compound document has a corresponding first field in the structure directory; each compound document has a corresponding preset first field, and the structure corresponding to the preset first field is used to record the offset information of the macro code in the data stream of the compound document, Then, the method further comprises: determining a preset first field corresponding to the target compound document; Then, parsing the structure of the target compound document and extracting the target offset information includes: Parsing the first data stream of the target compound document to obtain a structure directory of the target compound document; Based on the structure directory, determining a first structure corresponding to the preset first field in the target compound document; The target offset information is extracted from the first structure.

5. The method according to claim 4, characterized in that The structure for recording the offset information of the macro code in the data stream of the compound document in the compound document includes at least one second structure corresponding to the second field, the second field is used to indicate the macro code, the second structure records the data stream information indicating the data stream where the macro code indicated by the second field is located and the byte offset of the macro code relative to the starting byte of the data stream. Then, extracting the target offset information from the first structure includes: Determine a second field included in the first structure; For each of the second fields, the data stream information and the byte offset recorded in the second structure corresponding to the second field are extracted as target offset information.

6. The method according to claim 5, characterized in that The structure for recording the offset information of the macro code in the data stream of the compound document also includes a target number, and the target number is used to indicate the total number of macro codes in the compound document. Then, the method further includes: Determining whether the target quantity included in the first structure is zero; If it is not zero, executing the step of determining the second field included in the first structure; If it is zero, it indicates that the target compound document does not contain macro code.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Detecting whether the target compound document includes all specified data streams, where the specified data streams are data streams that the compound document containing the macro code needs to have; If included, the step of parsing the structure of the target compound document and extracting target offset information is performed; If not included, it is prompted that the target compound document does not contain macro code.

8. The method according to claim 7, characterized in that Each data stream of the compound document has a corresponding data stream name, and the number of the designated data streams is at least one. Then, detecting whether the target compound document includes all designated data streams includes: Matching the data stream name of each of the specified data streams with the data stream name of the data stream of the target compound document; If the data stream names of all designated data streams are matched successfully, it is detected that the target compound document includes all designated data streams; If the data stream name of any designated data stream fails to match successfully, it is detected that the target compound document does not include all designated data streams.

9. The method according to claim 7, characterized in that: Each data stream of the compound document is used to record data of a corresponding data category, and the number of the specified data streams is at least one. Then, detecting whether the target compound document includes all the specified data streams includes: Detecting whether each of the designated data streams has a matching data stream in the target compound document, wherein the designated data stream and the matching data stream record data of the same data category; If there are matching data streams for all designated data streams, it is detected that the target compound document includes all designated data streams; If there is no matching data stream for any designated data stream, it is detected that the target compound document does not include all designated data streams.

10. The method according to claim 7, characterized in that The designated data stream includes a second data stream, and the second data stream is used to record code information of the code included in the compound document. Then, the method further includes: If it is detected that the target compound document includes all designated data streams, detecting whether there is any keyword for indicating a macro code in the code information of the second data stream record of the target compound document, wherein the number of the keyword for indicating a macro code is at least one; If it exists, executing the step of parsing the structure of the target compound document and extracting the target offset information; If it does not exist, it indicates that the target compound document does not contain macro code.

11. The method according to any one of claims 1 to 6, characterized in that: The compound document is a compound document in an object-linked and embedded OLE format.

12. A device for extracting macro codes from a compound document, characterized in that: The compound document records the offset information of the macro code in the data stream of the compound document in a structure, and the device comprises: An acquisition module is used to acquire a target compound document to be extracted macro codes; A parsing module, used for parsing the structure of the target compound document and extracting target offset information; The extraction module is used to extract macro codes from the data stream of the target compound document based on the target offset information.

13. A computer-readable storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the method for extracting macro codes from a compound document as claimed in any one of claims 1 to 11.

14. An electronic device, characterized in that: The storage management device comprises: Memory, used to store programs; A processor is coupled to the memory and is used to run the program to execute the method for extracting macro codes from a compound document as claimed in any one of claims 1 to 11.