Determination method, determination device, and program

By extracting and comparing functional unit category information across languages, the method enhances the determination of malicious software, addressing language-dependent limitations and improving detection accuracy.

WO2025154322A1PCT designated stage expired Publication Date: 2025-07-24PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/033059
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-30
Filing Date
2024-09-17
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing methods for determining malicious software using machine learning are limited by language-dependent features, leading to difficulties in making language-independent determinations and potential overlooks of functional features like API usage, which affects determination performance and practicality.

Method used

A method that extracts functional unit category information from software in different languages, using a malicious determination model trained on functional unit category information to determine software malice, with adjustments for covariate shift and domain differences.

Benefits of technology

Enables language-independent determination of malicious software by comparing functional unit category information across languages, improving inference accuracy and identifying malicious components within the software.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024033059_24072025_PF_FP_ABST
    Figure JP2024033059_24072025_PF_FP_ABST
Patent Text Reader

Abstract

This determination method is executed by a computer and determines whether or not the target software is malignant. The determination method includes: a step for extracting functional unit category information which is information in functional units constituting the target software, with each individual piece of software constituted by one or more functional units, and which classifies functional units into a common category in a first language for writing software and in a second language for writing software and different from the first language; and a step for determining whether or not the target software is malignant on the basis of the extracted functional unit category information.
Need to check novelty before this filing date? Find Prior Art

Description

Determination method, determination device, and program

[0001] The present disclosure relates to a determination method, a determination device, and a program.

[0002] In recent years, the use of open source software (OSS) has become widespread, and the number of attack methods targeting it has also increased and diversified. To address this wide variety of attack methods, a method has been proposed for determining whether software contains an attack (or anomaly) by inference using machine learning that has learned known attack methods (see, for example, Non-Patent Document 1).

[0003] Sejfia et al., “Practical automated detection of malicious npm packages”, ICSE 2022

[0004] On the other hand, there are cases where the above-mentioned method of determining by inference is difficult to apply.

[0005] In view of the above, an object of the present disclosure is to provide a determination method and the like that allows for expanded application of methods for determining by inference.

[0006] A determination method according to one aspect of the present disclosure is a determination method executed by a computer to determine whether target software is malicious, wherein each piece of software is composed of one or more functional units, and includes the steps of extracting functional unit category information of the functional units that constitute the target software, the functional unit category information classifying the functional units into common categories in a first language for writing the software and a second language for writing the software that is different from the first language, and determining whether the target software is malicious based on the extracted functional unit category information.

[0007] In addition, a determination device according to one aspect of the present disclosure is a determination device that determines whether target software is malicious, wherein each piece of software is composed of one or more functional units, and the determination device includes: an extraction unit that extracts functional unit category information of the functional units that constitute the target software, the functional unit category information classifying the functional units into common categories in a first language for writing the software and a second language for writing the software that is different from the first language; and a determination unit that determines whether the target software is malicious based on the extracted functional unit category information.

[0008] Furthermore, one aspect of the present disclosure can also be realized as a program for causing a computer to execute the above-described determination method.

[0009] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0010] According to the present disclosure, it is possible to expand the application of methods for making judgments by inference.

[0011] FIG. 1 is a block diagram showing a functional configuration of a malicious software detection device according to an embodiment. FIG. 2 is a flowchart showing an example of operation of the malicious software detection device according to an embodiment. FIG. 3 is a flowchart showing part of operation of the malicious software detection device according to an embodiment. FIG. 4 is a diagram for explaining API categorization according to an embodiment. FIG. 5 is a diagram for explaining API categorization according to an embodiment. FIG. 6 is a flowchart showing part of operation of the malicious software detection device according to an embodiment. FIG. 7 is a diagram showing an example of extracted feature quantities according to an embodiment. FIG. 8 is a diagram showing an example of domain application according to an embodiment. FIG. 9 is a flowchart showing part of operation of the malicious software detection device according to an embodiment. FIG. 10 is a flowchart showing part of operation of the malicious software detection device according to an embodiment. FIG. 11 is a diagram showing an example of a malicious software detection result according to an embodiment. FIG. 12 is a flowchart showing part of operation of the malicious software detection device according to an embodiment. FIG. 13 is a diagram for explaining information presented by the malicious software detection device according to an embodiment.

[0012] (Knowledge that formed the basis of the disclosure) In recent years, the use of open source software (OSS) has become widespread, and the means of attacking it have also increased and diversified. In response to the wide variety of attack means, a method has been proposed for determining by inference whether or not software contains an attack (or anomaly) using machine learning that has learned known attack means.

[0013] On the other hand, various programming languages ​​(hereinafter simply referred to as languages) are used to write software, and the writing styles differ for each language. Therefore, it is difficult to perform inference-based judgment (hereinafter sometimes referred to as cross-language judgment) on software written in a language different from the language in which the software used for training was written. For example, there have been attempts to perform cross-language judgment using structural features of software. This method utilizes structural features such as the number of characters in the program and the number of files in the software, but does not utilize functional features such as the called API (Application Programming Interface). As a result, there is a risk of overlooking automatic code generation and abuse of network communication functions, and the method has poor judgment performance and is not practical. Therefore, the present disclosure aims to provide a judgment method, etc., that can expand the application of inference-based judgment methods while maintaining judgment performance by performing cross-language judgment using functional features of software.

[0014] More specifically, the determination method according to the first aspect of the present disclosure is a determination method executed by a computer for determining whether or not target software is malicious, wherein each piece of software is composed of one or more functional units, and includes the steps of extracting functional unit category information of the functional units that constitute the target software, the functional unit category information classifying the functional units into common categories in a first language for writing the software and a second language for writing the software that is different from the first language, and determining whether or not the target software is malicious based on the extracted functional unit category information.

[0015] According to this determination method, by extracting functional unit category information that classifies functional units that make up software into common categories, it is possible to compare functional unit category information even for software written in different languages. As a result, a model trained to determine whether software is malicious using functional unit category information can infer whether the target software is malicious based on the extracted functional unit category information, regardless of the language in which the software is written. Therefore, this method of determining whether software is malicious based on inference can be applied across languages ​​(expanding its application).

[0016] Furthermore, for example, a determination method according to a second aspect of the present disclosure is a determination method according to the first aspect, and in the determination step, the extracted functional unit category information is input into a malignancy determination model trained using at least one of functional unit category information of functional units that constitute benign software and functional unit category information of functional units that constitute malicious software, thereby determining whether the target software is malicious or not.

[0017] According to this, it is possible to determine whether or not the target software is malicious by inputting the extracted functional unit category information into the maliciousness determination model.

[0018] Furthermore, for example, a determination method according to a third aspect of the present disclosure is a determination method according to the second aspect, further including a step of constructing a malignancy determination model using at least one of functional unit category information of functional units that constitute benign software and functional unit category information of functional units that constitute malignant software.

[0019] This allows the construction and use of a malignancy determination model.

[0020] Furthermore, for example, a determination method according to a fourth aspect of the present disclosure is a determination method according to the third aspect, and in the construction step, an adjustment is made to the malignancy determination model to reduce the covariate shift between the distribution of functional unit category information of the functional units that constitute the target software and the distribution of functional unit category information used to train the malignancy determination model.

[0021] This reduces the covariate shift between the two distributions, making it possible to construct a malignancy determination model with improved inference accuracy.

[0022] Furthermore, for example, a determination method according to a fifth aspect of the present disclosure is a determination method according to any one of the second to fourth aspects, in which the target software is written in a language different from a language that describes at least one of malicious software and benign software that are composed of functional units related to functional unit category information used to train the maliciousness determination model.

[0023] This makes it possible to determine whether or not target software written in a language different from the language in which the software used to train the maliciousness determination model is written is malicious.

[0024] Furthermore, for example, a determination method relating to a sixth aspect of the present disclosure is a determination method relating to any one of the first to fifth aspects, and in the determination step, a score indicating the likelihood that the target software is malicious is calculated based on the extracted functional unit category information, and whether the target software is malicious or not is determined based on the calculated score.

[0025] According to this, a score is calculated for the target software, and it is possible to determine whether the target software is malicious or not from the score.

[0026] Furthermore, for example, a determination method according to a seventh aspect of the present disclosure is the determination method according to any one of the first to sixth aspects, in which the functional unit is an API (Application Programming Interface).

[0027] This allows functional unit category information of APIs that constitute the target software to be extracted and used.

[0028] Furthermore, for example, a determination method according to an eighth aspect of the present disclosure is a determination method according to any one of the first to seventh aspects, and further includes a step of calculating the similarity between functional unit category information of functional units constituting target software determined to be malicious and functional unit category information of functional units constituting reference software known to be malicious, and a step of identifying, based on the calculation result, reference software composed of functional units associated with functional unit category information similar to the functional unit category information of the functional units constituting the target software determined to be malicious.

[0029] With this, for target software determined to be malicious, it is possible to identify reference software that is similar to the target software and is known to be malicious, based on the similarity of the functional unit category information.

[0030] Furthermore, for example, a determination method relating to a ninth aspect of the present disclosure is a determination method relating to any one of the first to eighth aspects, and further includes a step of estimating the contribution of each individual functional unit to the maliciousness determination in target software determined to be malicious, and determining, based on the estimation result, the location among the functional units where the malicious functional unit that contributes to the maliciousness determination is written.

[0031] This makes it possible to determine, for each functional unit of the target software that has been determined to be malicious, the location where the malicious functional unit that contributes to the maliciousness determination is written.

[0032] Furthermore, for example, a determination method relating to a tenth aspect of the present disclosure is a determination method relating to any one of the first to ninth aspects, and further includes a step of estimating the contribution of each individual functional unit to the maliciousness determination in target software that has been determined to be malicious, and generating information regarding abnormalities in the target software that are expected from malicious functional units that contribute to the maliciousness determination among the functional units based on the estimation results.

[0033] This makes it possible to generate, for each functional unit of target software that has been determined to be malicious, information regarding anomalies in the target software that are predicted from the malicious functional units that contribute to the determination of maliciousness.

[0034] Furthermore, for example, a determination method relating to an eleventh aspect of the present disclosure is a determination method relating to any one of the first to tenth aspects, and further includes a step of estimating the contribution of each individual functional unit to the maliciousness determination in target software that has been determined to be malicious, and based on the estimation result, deleting from the target software any malicious functional unit among the functional units that contributes to the maliciousness determination.

[0035] According to this, for each functional unit of the target software that has been determined to be malicious, the malicious functional unit that contributed to the determination of maliciousness can be deleted from the target software.

[0036] Furthermore, for example, a program according to a twelfth aspect of the present disclosure is a program for causing a computer to execute the determination methods according to the first to eleventh aspects.

[0037] According to this, by executing the program using a computer, the same effects as those of the determination method described above can be achieved.

[0038] Furthermore, for example, a determination device relating to a thirteenth aspect of the present disclosure is a determination device that determines whether target software is malicious, wherein each piece of software is composed of one or more functional units, and the determination device includes: an extraction unit that extracts functional unit category information of the functional units that constitute the target software, the functional unit category information classifying the functional units into common categories in a first language for writing the software and a second language, different from the first language, for writing the software; and a determination unit that determines whether the target software is malicious based on the extracted functional unit category information.

[0039] This provides the same effects as the above-described determination method.

[0040] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0041] Hereinafter, embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not recited in independent claims will be described as optional components. Note that each figure is a schematic diagram and is not necessarily an exact illustration. Furthermore, in each figure, substantially identical components are assigned the same reference numerals, and duplicated descriptions may be omitted or simplified.

[0042] (Embodiment) [Configuration] First, the configuration of a malicious software determination device according to an embodiment will be described. FIG. 1 is a block diagram showing the functional configuration of the malicious software determination device according to an embodiment. FIG. 1 also shows the malicious software determination device 100, as well as various pieces of information input and output to and from the malicious software determination device 100. The malicious software determination device 100 according to this embodiment is an example of a determination device, and as shown in FIG. 1, includes feature extraction units 11 and 11a, an API categorizer storage unit 12, a domain application setting storage unit 13, a learning unit 14, a maliciousness determination model storage unit 15, a scoring unit 16, a visualization unit 17, an anomaly removal determination unit 18, and a category model learning unit 19. The malicious software determination device 100 is realized by sequentially executing predetermined programs using a processor and a memory on a computer such as an information processing server, for example.

[0043] In the malicious software determination device 100 of this embodiment, a maliciousness determination model is constructed by machine learning, and whether or not the target software to be inspected is malicious is determined by inference using the maliciousness determination model. Note that the determination of whether or not the software is malicious is performed by at least one of determining that the software is malicious, determining that the software is benign (not malicious), determining that the software is not malicious, or determining that the software is not benign.

[0044] To build a maliciousness determination model, training is performed using a benign software package (hereinafter, the terms "package" and "software" are used synonymously) written in language X and a malicious package written in language X, such as malware. The training package 21 includes at least one of such benign packages and malicious packages written in a first language. The training package may also include a package written in a second language (e.g., language Y) different from the first language. In this case, benign packages written in language Y are easily available, making this preferable. In this embodiment, packages written in two or more languages ​​can be used to train the maliciousness determination model. This is because the feature extraction unit 11 is provided at the destination where the training package 21 is input, and feature category information, which is information on features in a common category classification regardless of language, can be obtained.

[0045] A package, i.e., software, is composed of a combination of one or more functional units. Categorization by these functional units allows one or more functional units contained in software to be classified into a common category in any language. In other words, by extracting functional unit category information from software indicating which functional units are classified into which categories the software is constructed, it is possible to compare the functional characteristics of different software programs, even if they are written in different languages. Therefore, by using functional unit category information for training, the trained maliciousness detection model can be used across languages, even if the languages ​​of the training software and the software to be tested are different.

[0046] An example of a functional unit is an API, but other functional units such as functions, functional unit codes, etc. may be defined in any way as long as they are parts of a program that correspond to one function.

[0047] Furthermore, the feature category information as described above may include any other information as long as it includes at least functional unit category information. In this embodiment, the feature extraction unit 11 extracts feature category information including functional unit category information and structural feature quantities from the learning package 21.

[0048] In order to extract the function unit category information, this embodiment uses an API categorizer stored in the API categorizer storage unit 12. The API categorizer is constructed in advance by the category model learning unit 19 and stored therein.

[0049] The category model learning unit 19 can construct an API categorizer that associates API descriptions in each language with categories by using a large-scale language model or the like, and inputs the APIs defined in each target language, API descriptions from their official pages or definition documents, etc., and predefined category information (all of which are included in the category model learning information 24). Because the categories here are common across languages, APIs classified into the same category, even if they are in different languages, correspond to the same functional unit.

[0050] When using a large-scale language model, it is possible to use API names, descriptions, category information, etc. as data for fine-tuning. Furthermore, when categorizing a target API, the API name and description are also used as input. The input of the API name and description is converted into a vector by the natural language model, and categorization is performed by calculating the proximity of the API to each predefined category.

[0051] The domain application setting storage unit 13 stores information for domain application processing used when training a maliciousness determination model. Differences in the way functional units are included between the software used for training and the target software may occur due to differences in the languages ​​used, etc. In other words, a covariate shift may occur between the feature vectors of the features of the software used for training and the feature vectors of the features of the target software. In other words, the feature vectors of the features of the software used for training and the feature vectors of the features of the target software may be in different domains. By taking such domain differences into account and making adjustments to reduce the covariate shift, it is possible to improve the determination accuracy.

[0052] Here, the difference in domains refers to, for example, differences in distributions caused by differences in the intended use of packages between languages, such as Python (registered trademark) and JavaScript (registered trademark), which arise when the training language and the testing language are different. Furthermore, the difference in domains refers to the difference in the distribution of data between training and testing, which occurs when the set of packages used in training is biased toward those used in software development in a specific field, and the testing package is frequently used in a different field from the set of packages used in training. In such cases, adjusting the model using a so-called transfer learning model can absorb the difference (reducing covariate shift) and improve the accuracy of the assessment.

[0053] The learning unit 14 is a processing unit that constructs a malignancy determination model by learning. The learning unit 14 is an example of a construction unit that constructs a malignancy determination model by learning. The learning unit 14 stores the constructed malignancy determination model in the malignancy determination model storage unit 15.

[0054] In constructing a maliciousness determination model, a machine learning model is trained to determine whether an input feature vector is malicious or not, based on feature vectors of benign features and feature vectors of malicious features. A classifier such as a random forest can be used for the model. Since functional unit category information related to functional unit categories common across languages ​​can be used as the feature, the constructed maliciousness determination model can inspect feature vectors whose features are functional unit category information extracted from the target package 22, which is the package to be inspected. The feature extraction unit 11a has the same functions as the feature extraction unit 11, except for whether the input package is the training package 21 or the target package 22, and therefore a description thereof will be omitted here.

[0055] The scoring unit 16 calculates a score for the target package 22 being inspected, representing the likelihood of the package 22 being malicious, based on the proximity to a feature vector of a malicious feature or the distance from a feature vector of a benign feature. If this score is higher than a preset threshold, the target package 22 is determined to be a malicious package (malicious software), and if the score is equal to or lower than the preset threshold, the target package 22 is determined to be a benign package (benign software). In this way, the scoring unit 16 functions as a determination unit that determines whether the target software is malicious or not based on the calculated score. Note that the determination unit may also include a determination unit that only determines whether the software is malicious or not without calculating a score.

[0056] The visualization unit 17 visualizes the determination result in a form that is easy for the user to understand (for example, as an image with information written on it). The visualization unit 17 outputs the visualized image or the like as output data 23.

[0057] The anomaly removal determination unit 18 is a functional unit that removes the attack site (the anomaly) that caused the maliciousness of the target package 22 that has been determined to be malicious. The anomaly removal determination unit 18 outputs the package corrected by the removal as output data 23.

[0058] [Operation] Next, an example of the operation of the malicious software determination device 100 configured as described above will be described with further reference to FIG. 2 and subsequent figures.

[0059] 2 and 3 are flowcharts showing an example of the operation of the malicious software determination device according to the embodiment, respectively.

[0060] As shown in Fig. 2, the malicious software determination device 100 first constructs an API categorizer (S10). Here, Fig. 3 shows detailed operations of step S10 in Fig. 2. As shown in Fig. 3, the category model learning unit 19 acquires API details (i.e., category model learning information 24) from the official website of the organization that standardizes the programming language (S11). Then, it prepares API categories common to all languages ​​to be categorized, either manually or by machine learning (S12). This API category is also acquired (read) by the category model learning unit 19 as the category model learning information 24. The category model learning unit 19 then constructs an API categorizer, which is then stored in the API categorizer storage unit 12 (S13).

[0061] 4 and 5 are diagrams illustrating categorization of APIs according to an embodiment. To categorize APIs, a list of APIs in the target language is first prepared using a group of APIs called standard APIs, and APIs that appear on the list are then identified as being used within the package. The extracted APIs are then classified into predefined categories of API functions frequently used by malware. In other words, the prepared API categories are not comprehensive, but can be narrowed down to APIs that are likely to be used in attacks. However, in case an API does not fit into any category, it can be classified into a category such as "other" to prevent it from being used as a feature. Furthermore, API categories may be hierarchically structured, with main categories and subcategories, such as an OS Information category under the SystemInfo category.

[0062] 2, the learning unit 14 constructs a malicious software determination model by learning (S20). Fig. 6 is a flowchart showing part of the operation of the malicious software determination device according to the embodiment. Fig. 6 shows the detailed operation of step S20 in Fig. 2.

[0063] As shown in Figure 6, the feature extraction unit 11 acquires one or more learning packages 21 (S21). If package meta information such as update frequency, number of pull requests, and package name is available, the feature extraction unit 11 uses these to extract them as part of the feature (S22). These feature values ​​are neither structural nor functional features, but are different features. Since such feature values ​​may also reveal signs of maliciousness, they may be used in combination in this way.

[0064] The feature extraction unit 11 extracts structural information of the program, such as the number of lines, the number of characters, and the presence or absence of encoded character strings, as part of the feature (S23). This feature corresponds to the structural feature described above.

[0065] The feature extraction unit 11 further extracts APIs used in the learning package 21 by program analysis or the like, categorizes them using an API categorizer, and extracts them as feature amounts (S24). These feature amounts correspond to the functional feature amounts described above.

[0066] The features extracted as described above are organized for each package, as shown in Fig. 7. Fig. 7 is a diagram showing an example of extracted features according to an embodiment. As shown in Fig. 7, the features used for learning can be a combination of functional features common to all languages, extracted by categorizing APIs used in a package, and structural features of programs and the like in the package.

[0067] As mentioned above, examples of structural features that can be used include the number of lines in a program, the number of encoded characters, and hard-coded IP addresses. Functional features may include not only features indicating whether APIs related to each category (or classified into categories) are used, but also features indicating whether functions from multiple categories are used simultaneously. For example, the Network & System info shown in the figure corresponds to this.

[0068] When constructing a malignancy detection model by learning using features using a random forest or the like, it is possible to quantify the importance of each feature in discrimination. Similar quantification is also possible using an expressible method such as SHAP. Hereinafter, this numerical value is referred to as feature importance. Based on the feature importance information, APIs used in a package determined to be malignant and associated with an API category with a relatively high feature importance (e.g., above a threshold or within the top few percent) can be presented to the user, allowing the user to use this information as a reference for determining the anomaly of a portion determined to be anomaly by the user (i.e., whether the user considers it to be anomaly).

[0069] Returning to Fig. 6, the learning unit 14 constructs a model in which domain application is performed so as to absorb the difference in distribution between the feature values ​​extracted in the target package 22 and the feature values ​​extracted in the learning package 21 (reducing the covariate shift) (S25). Fig. 8 is a diagram showing an example of domain application according to an embodiment. Fig. 8 shows an image diagram of domain application, where (a) shows the distribution before application and (b) shows the distribution after application.

[0070] For example, when different languages ​​such as JavaScript (registered trademark) and Python (registered trademark) are used as the description languages ​​for the learning package 21 and the target package 22, respectively, there is a possibility that the purpose of the created packages will be biased and the distribution of data will be different. In this case, there is a possibility that the malignancy determination model constructed from the learning package 21 cannot be effectively applied to determining whether a package is malignant or not based on the feature quantities of the target package 22.

[0071] Therefore, it is possible to use domain adaptation techniques. For example, importance-weighted learning or feature transfer learning can be performed so that the malignancy determination model functions even when the distribution of features extracted from the learning package 21 and the distribution of features extracted from the target package 22 are different (different domains).

[0072] As already mentioned, it is believed that similar domain application processing will be effective not only in cases of differences between languages, but also in cases where the domains are different and accuracy is reduced because the learning package 21 is a package that is commonly used in a specific field, such as embedded systems, while the target package 22 is a package that is commonly used in another field, such as IoT systems.

[0073] Returning to FIG. 6, the learning unit 14 uses the information obtained as described above to learn and construct a malignancy determination model (S26).

[0074] Returning to Fig. 2, the scoring unit 16 determines whether the target package 22 is malicious or not using the constructed maliciousness determination model (S30). Fig. 9 is a flowchart showing part of the operation of the malicious software determination device according to the embodiment. Fig. 9 shows the detailed operation of step S30 in Fig. 2.

[0075] 9, the feature extraction unit 11 acquires the target package 22 and extracts feature quantities similar to those shown in steps S22 to S24 (extracts feature category information, S31). Next, the scoring unit 16 scores the extracted feature quantities using the constructed malignancy determination model (S32). The scoring unit 16 detects anomalies in the target package 22 based on threshold determination using the calculated score (S33). If an anomaly is detected, the target package is determined to be malignant.

[0076] The visualization unit 17 visualizes the determination result so that the user can confirm the abnormality of the part where the abnormality was detected (S40). Here, Fig. 10 is a flowchart showing part of the operation of the malicious software determination device according to the embodiment. Fig. 10 shows the detailed operation of step S40 in Fig. 9.

[0077] The visualization unit 17 quantifies the features determined to be functionally abnormal and ranks them according to the likelihood of abnormality (S41). The visualization unit 17 presents information about the target package 22 by including it in the visualized information (S42). The visualization unit 17 also estimates attack details highly related to features such as the APIs used (S43). The visualized information then includes the estimated attack details and the API functions related to the attack (S44). FIG. 11 is a diagram illustrating an example of a malicious software determination result according to an embodiment. As shown in FIG. 11, the system presents a program location within the package that is suspected to be abnormal, a possible attack, and an overview of the package, allowing the user to confirm the abnormality of the portion of the target package 22 where an abnormality was detected from this information.

[0078] The package overview can be, for example, the one related to the target package 22 obtained from a website or the like. Possible attacks can be presented according to rules that are created and stored in advance for presenting information linked to the contents of the APIs used. Possible abnormalities can be presented by determining and presenting the portions that contain many APIs that contributed to the maliciousness determination.

[0079] Returning to Fig. 9, the anomaly deletion determination unit 18 determines whether or not to delete the portion in which the anomaly was detected (S50). Fig. 12 is a flowchart showing part of the operation of the malicious software determination device according to the embodiment. Fig. 12 shows the detailed operation of step S50 in Fig. 9.

[0080] As shown in FIG. 12 , first, a set of expected inputs and outputs, such as test code, for the program is prepared (S51). Then, the change in output for the prepared inputs is verified when the part where the abnormality was detected is deleted and when it is not deleted (S52). If the output changes (Yes in S53), the visualized information is presented to the user by including a deletion suggestion (S54). For example, as shown in FIG. 11 , the user is presented with the option of whether or not to delete the abnormal part. If the user then allows the deletion (Yes in S55), the program is corrected by deleting the part (S56). On the other hand, if the output does not change (No in S53) or if the user does not allow the deletion (No in S55), step S56 is not performed and the process ends.

[0081] 13 is a diagram illustrating information presented by the malicious software determination device according to the embodiment. As shown in FIG. 13, for example, assuming that package A, package X, and package Y, which have been determined to be malicious, exist, the system checks whether package A is similar to package X or package Y (reference software). This similarity information can be used to infer an attack by package A based on whether package A is more similar to package X or package Y, the details of which are known about the attack, than package A.

[0082] First, as shown in the figure, the API category features are limited to those expressed as main category > subcategory, and the API category features with the top M feature importance for package A are extracted, where M is a constant, and the APIs that are actually used are extracted (hereinafter referred to as an important API list).

[0083] A possible method for checking whether packages A are similar is to extract important API lists that contribute to anomaly detection from each of two packages, X and Y, and calculate the sum of the feature importance of the APIs included in the set (important API list of package X) ∩ (important API list of package A) or the set (important API list of package Y) ∩ (important API list of package A) and the sum of the feature importance of the APIs included in the set (important API list of package X) ∪ (important API list of package A) or the set (important API list of package Y) ∪ (important API list of package A), and use the sum of the former divided by the sum of the latter as the similarity.

[0084] Furthermore, when indicating the API used by the malicious package in the important API list, the feature importance of the API category feature to which the API belongs may be used as a similar factor of feature importance.

[0085] Other Embodiments Although the embodiments have been described above, the present disclosure is not limited to the above-described embodiments.

[0086] For example, in the above-described embodiment, a process executed by a specific processing unit may be executed by another processing unit, the order of multiple processes may be changed, or multiple processes may be executed in parallel.

[0087] In the above-described embodiments, each component may be realized by executing a software program suitable for that component, or by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0088] Furthermore, each component may be realized by hardware. For example, each component may be a circuit (or integrated circuit). These circuits may form a single circuit as a whole, or each may be a separate circuit. Furthermore, each of these circuits may be a general-purpose circuit or a dedicated circuit.

[0089] Furthermore, the general or specific aspects of the present disclosure may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, etc. Furthermore, the general or specific aspects of the present disclosure may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0090] For example, the present disclosure may be realized as a determination method executed by a computer, or as a program for causing a computer to execute the determination method. The present disclosure may also be realized as a computer-readable non-transitory recording medium on which such a program is recorded.

[0091] In addition, this disclosure also includes forms obtained by applying various modifications to each embodiment that a person skilled in the art would think of, or forms realized by arbitrarily combining the components and functions of each embodiment within the scope that does not deviate from the intent of this disclosure.

[0092] This disclosure is useful in determining whether software is malicious.

[0093] REFERENCE SIGNS LIST 11, 11a Feature extraction unit 12 API categorizer storage unit 13 Domain application setting storage unit 14 Learning unit 15 Maliciousness determination model storage unit 16 Scoring unit 17 Visualization unit 18 Anomaly removal determination unit 19 Category model learning unit 21 Learning package 22 Target package 23 Category model learning information 24 Output data 100 Malicious software determination device

Claims

1. A determination method executed by a computer to determine whether target software is malicious, wherein each software is composed of one or more functional units, and the method includes: extracting functional unit category information for classifying the functional units into a common category in a first language for describing software and a second language different from the first language for describing software, which is the functional unit category information of the functional units constituting the target software; and determining whether the target software is malicious based on the extracted functional unit category information.

2. The determination method according to claim 1, wherein in the determining step, the extracted functional unit category information is input to a malicious determination model learned using at least one of the functional unit category information of the functional units constituting benign software and the functional unit category information of the functional units constituting malicious software, so as to determine whether the target software is malicious.

3. The determination method according to claim 2, further including a step of constructing the malicious determination model using at least one of the functional unit category information of the functional units constituting the benign software and the functional unit category information of the functional units constituting the malicious software.

4. The determination method according to claim 3, wherein in the constructing step, an adjustment for reducing the co-variance shift between the distribution of the functional unit category information of the functional units constituting the target software and the distribution of the functional unit category information used for learning the malicious determination model is performed on the malicious determination model.

5. The determination method according to claim 2, wherein the target software is described in a language different from the language describing at least one of the malicious software and the benign software composed of the functional units related to the functional unit category information used for learning the malicious determination model.

6. The determination method according to claim 1, wherein in the determining step, based on the extracted functional unit category information, a score indicating the likelihood that the target software is malicious is calculated, and it is determined whether the target software is malicious based on the calculated score.

7. The determination method according to claim 1, wherein the functional unit is an API (Application Programming Interface).

8. The determination method according to claim 1, further comprising: calculating a similarity between the functional unit category information of the functional units constituting the target software determined to be malicious and the functional unit category information of the functional units constituting reference software known to be malicious; and identifying the reference software constituted by the functional units related to the functional unit category information similar to the functional unit category information of the functional units constituting the target software determined to be malicious based on the calculation result.

9. The determination method according to any one of claims 1 to 8, further comprising: estimating a contribution degree of each of the functional units to the malicious determination in the target software determined to be malicious, and determining a location where a malicious functional unit contributing to the malicious determination among the functional units is described based on the estimation result.

10. The determination method according to any one of claims 1 to 8, further comprising: estimating a contribution degree of each of the functional units to the malicious determination in the target software determined to be malicious, and generating information regarding an abnormality in the target software assumed from the malicious functional units contributing to the malicious determination among the functional units based on the estimation result.

11. The determination method according to any one of claims 1 to 8, further comprising: estimating a contribution degree of each of the functional units to the malicious determination in the target software determined to be malicious, and deleting the malicious functional units contributing to the malicious determination among the functional units from the target software based on the estimation result.

12. A program for causing the computer to execute the determination method according to claim 1.

13. A determination device that determines whether target software is malicious, wherein each software is composed of one or more functional units, and functional unit category information of the functional units constituting the target software, which classifies the functional units into a common category in a first language for describing software and a second language different from the first language for describing software, an extraction unit that extracts the functional unit category information, and a determination unit that determines whether the target software is malicious based on the extracted functional unit category information. The determination device is provided with the above components.

Citation Information

Patent Citations

  • Illegal application category identification method and device

    CN111143833A

  • Android malicious software detection and malicious code positioning system and method

    CN111523117A

  • System, apparatus and method for classifying a file as malicious using static scanning

    US10192052B1

  • Extending dynamic detection of malware using static and dynamic malware analyses

    US20200026851A1

  • Method for machine learning of malicious code detecting model and method for detecting malicious code using the same

    US20210133323A1