Program identification method and program identification device

By generating a machine learning model, the feature vectors of programming languages with more label information are transformed into programming language forms with less label information, which solves the accuracy of malicious program detection in programming languages with less label information, and realizes high-precision malicious program recognition.

CN120359512APending Publication Date: 2025-07-22PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380084935.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-06
Filing Date
2023-11-13
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to detect malicious programs with high precision in programming languages with less label information.

Method used

By generating a machine learning model, a programming language with more label information is used to learn feature vectors, and transform it into a programming language with less label information to identify malicious programs.

Benefits of technology

Even in programming languages with little label information, malicious programs can be identified with high accuracy, improving the accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359512A_ABST
    Figure CN120359512A_ABST
Patent Text Reader

Abstract

A program recognition method includes: acquiring a machine learning model (S21) obtained by machine learning, the machine learning model being a model generated by learning learning data including a plurality of first feature vectors and a plurality of pieces of recognition information indicating whether or not a first program is a rogue program, the first feature vector is represented in a first form representing whether each of a plurality of first functions of the program in the first language is used by the first program; generating a second feature vector (S22), the second feature vector being expressed in a second form representing whether each of a plurality of second functions of the program in the second language is used by the second program; converting the second feature vector into a first form (S23); and outputting an identification result (S24) indicating whether the second program is a rogue program, the identification result being obtained by inputting the converted second feature vector into the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a program recognition method and a program recognition apparatus. Background Art

[0002] Malicious (abnormal) code may sometimes be mixed into open-source software. In Non-Patent Document 1, a technique of using machine learning to detect maliciously mixed software is disclosed.

[0003] Prior Art Documents

[0004] Non-Patent Documents

[0005] Non-Patent Document 1: Ke Xu et al., "DroidEvolver: Self-Evolving Android Malware Detection System", 2019 IEEE European Symposium on Security and Privacy (EuroS&P)

[0006] Non-Patent Document 2: Trong Duc Nguyen et al., "Exploring API Embedding for API Usages and Applications", 2017 IEEE / ACM 39th International Conference on Software Engineering Summary of the Invention

[0007] Problems to be Solved by the Invention

[0008] The present disclosure provides a program recognition method and the like capable of detecting malicious programs with high accuracy.

[0009] Technical Solutions for Solving the Problems

[0010] A program recognition method according to an aspect of the present disclosure includes: obtaining a machine learning model generated by learning using training data indicating whether each of a plurality of first programs is a malicious program (whether there is malice), each of the plurality of first programs being expressed (written) in a first language, the machine learning model being a model generated by learning learning data generated based on the training data, the learning data including a plurality of first feature vectors and a plurality of identification information, the plurality of first feature vectors being obtained by extracting features from each of the plurality of first programs, the plurality of identification information respectively indicating whether a corresponding one of the plurality of first programs is malicious, each of the plurality of first feature vectors being expressed in a first form, the first form indicating whether each of a plurality of first functions of a program expressed in the first language is used by a corresponding one of the first programs; generating a second feature vector by extracting features of a second program expressed in a second language different from the first language, the second feature vector being expressed in a second form, the second form indicating whether each of a plurality of second functions of a program expressed in the second language is used by the second program; transforming the generated second feature vector into the first form; outputting an identification result indicating whether the second program is a malicious program, which is obtained by inputting the second feature vector transformed into the first form into the machine learning model.

[0011] A program recognition device according to an aspect of the present disclosure includes a processor and a memory, and the processor executes, using the memory: generating a machine learning model by learning using training data indicating whether each of a plurality of first programs is a malicious program, each of the plurality of first programs being expressed in a first language, the machine learning model being a model generated by learning learning data generated based on the training data, the learning data including a plurality of first feature vectors and a plurality of identification information, the plurality of first feature vectors being obtained by extracting features from each of the plurality of first programs, the plurality of identification information respectively indicating whether a corresponding one of the plurality of first programs is malicious, each of the plurality of first feature vectors being expressed in a first form, the first form indicating whether each of a plurality of first functions of a program expressed in the first language is used by a corresponding one of the first programs; transforming a second feature vector obtained by extracting features of a second program expressed in a second language different from the first language into the first form, the second feature vector being expressed in a second form, the second form indicating whether each of a plurality of second functions of a program expressed in the second language is used by the second program; and outputting an identification result indicating whether the second program is a malicious program, which is obtained by inputting the second feature vector transformed into the first form into the machine learning model.

[0012] In addition, these general or specific technical solutions can be implemented by a system, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a method, a device, a system, an integrated circuit, a computer program, and a non-transitory recording medium.

[0013] Advantageous Effects of the Invention

[0014] According to the program recognition method and the like according to the present disclosure, malicious programs can be detected with high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a block diagram showing an example of the configuration of a program recognition device according to an embodiment.

[0016] Figure 2 is a diagram for explaining a first feature vector.

[0017] Figure 3 is a diagram for explaining a part of the correspondence relationship.

[0018] Figure 4 is a diagram for explaining another part of the correspondence relationship.

[0019] Figure 5It is a flowchart showing an example of the learning process performed by the program recognition device.

[0020] Figure 6 It is a flowchart showing an example of the recognition process performed by the program recognition device.

[0021] Figure 7 It is information showing an example of the first malicious information.

[0022] Figure 8 It is a diagram for explaining a method of calculating the similarity between the first program and the second program.

[0023] Figure 9 It is a flowchart showing an example of the process of determining the first program similar to the second program. Detailed implementation

[0024] (Insight underlying the present disclosure)

[0025] In the prior art as described above, for programming languages in which a lot of source codes are given label information indicating benign or malicious, by using a machine learning model generated through supervised learning, it is possible to detect with high accuracy a program (hereinafter referred to as a malicious program) including an abnormal source code.

[0026] However, in the case of a programming language with less label information, supervised learning cannot be sufficiently performed, so it is difficult to detect malicious programs with high accuracy.

[0027] Therefore, the present disclosure provides a program recognition method and the like that can detect malicious programs with high accuracy.

[0028] The program recognition method according to the first aspect of the present disclosure includes: obtaining a machine learning model generated by learning using training data indicating whether each of a plurality of first programs is a malicious program, each of the plurality of first programs being expressed in a first language, the machine learning model being a model generated by learning learning data generated based on the training data, the learning data including a plurality of first feature vectors and a plurality of identification information, the plurality of first feature vectors being obtained by extracting features from each of the plurality of first programs, the plurality of identification information respectively indicating whether a corresponding one of the plurality of first programs is malicious, each of the plurality of first feature vectors being expressed in a first form, the first form indicating whether each of a plurality of first functions of a program expressed in the first language is used by a corresponding one of the first programs; generating a second feature vector by extracting features of a second program expressed in a second language different from the first language, the second feature vector being expressed in a second form, the second form indicating whether each of a plurality of second functions of a program expressed in the second language is used by the second program; transforming the generated second feature vector into the first form; and outputting an identification result indicating whether the second program is a malicious program obtained by inputting the second feature vector transformed into the first form into the machine learning model.

[0029] Accordingly, in order to identify whether a second program expressed in a second language is a malicious program, a machine learning model obtained by machine learning regarding whether a plurality of first programs expressed in a first language are malicious programs can be used. Thus, for example, even a program expressed in a programming language with less label information can be identified as malicious using a machine learning model learned from programs expressed in a programming language with more label information, and thus malicious programs can be identified with high accuracy.

[0030] The program recognition method according to the second aspect of the present disclosure is based on the program recognition method according to the first aspect. In the transformation, the second feature vector is transformed into the first form by using the correspondence between the plurality of first functions and the plurality of second functions.

[0031] Therefore, the form of the second feature vector can be easily transformed into the form of the first feature vector. Thus, even when using a machine learning model learned from programs in other programming languages, it is possible to accurately identify whether the second program is malicious.

[0032] The program recognition method according to the third aspect of the present disclosure is based on the program recognition method according to the second aspect, and the correspondence indicates that one of the plurality of first functions corresponds to one of the plurality of second functions.

[0033] Therefore, the form of the second feature vector can be easily transformed into the form of the first feature vector.

[0034] Based on the program recognition method according to the third aspect of the present disclosure, the correspondence relationship indicates that two or more other first functions among the multiple first functions are in correspondence with one other second function among the multiple second functions.

[0035] Therefore, even when in the relationship where two or more other first functions are in correspondence with one other second function, the form of the second feature vector can be easily transformed into the form of the first feature vector.

[0036] Based on the program recognition method according to the fourth aspect of the present disclosure, the correspondence relationship has weights among the two or more other first functions for the one other second function.

[0037] Therefore, according to the relationship between each of the two or more other first functions and the one other second function, the form of the second feature vector can be easily transformed into the form of the first feature vector.

[0038] Based on the program recognition method according to the sixth aspect of the present disclosure, the correspondence relationship indicates the similarity between the vector representations of each of the multiple first functions and the vector representations of each of the multiple second functions.

[0039] Therefore, the form of the second feature vector can be easily transformed into the form of the first feature vector.

[0040] The program recognition method according to the seventh aspect of the present disclosure, based on the program recognition method according to any one of the first to sixth aspects, further includes: for each of one or more first programs represented as malicious programs by the training data, obtaining first malicious information, the first malicious information including one or more first malicious contribution degrees, and the one or more first malicious contribution degrees respectively corresponding to one or more first functions represented as being used in the first feature vector corresponding to the first program; obtaining second malicious information, the second malicious information including one or more second malicious contribution degrees, and the one or more second malicious contribution degrees respectively corresponding to one or more first functions represented as being used in the second feature vector corresponding to the second program whose recognition result indicates maliciousness and transformed into the first form; by comparing each of the one or more first malicious information corresponding to each of the one or more first programs obtained and the second malicious information, determining the first program corresponding to the similar malicious information similar to the second malicious information; outputting information indicating the determined first program.

[0041] Therefore, it is possible to determine the first program having characteristics similar to those of the second program.

[0042] The program recognition method according to the eighth aspect of the present disclosure, based on the program recognition method according to the seventh aspect, in the determination, uses M first malicious contribution degrees selected in descending order of contribution degree from the one or more first malicious contribution degrees and M second malicious contribution degrees selected in descending order of contribution degree from the one or more second malicious contribution degrees to calculate the similarity between the second malicious information and the first malicious information to be a comparison object, and determines the similar malicious information according to the similarity, where M is an integer of 2 or more.

[0043] Therefore, it is possible to simplify the calculation of the similarity for determining the similar malicious information.

[0044] The program recognition device according to the ninth aspect of the present disclosure includes a processor and a memory, and the processor executes, using the memory: generating a machine learning model by learning using training data indicating whether each of a plurality of first programs is a malicious program, each of the plurality of first programs being expressed in a first language, the machine learning model being a model generated by learning learning data generated based on the training data, the learning data including a plurality of first feature vectors and a plurality of identification information, the plurality of first feature vectors being obtained by extracting features from each of the plurality of first programs, the plurality of identification information respectively indicating whether a corresponding one of the plurality of first programs is malicious, each of the plurality of first feature vectors being expressed in a first form, the first form indicating whether each of a plurality of first functions of a program expressed in the first language is used by a corresponding one of the first programs, transforming a second feature vector obtained by extracting features of a second program expressed in a second language different from the first language into the first form, the second feature vector being expressed in a second form, the second form indicating whether each of a plurality of second functions of a program expressed in the second language is used by the second program, and outputting an identification result indicating whether the second program is a malicious program obtained by inputting the second feature vector transformed into the first form into the machine learning model.

[0045] Accordingly, in order to identify whether a second program expressed in a second language is a malicious program, a machine learning model obtained by machine learning regarding whether a plurality of first programs expressed in a first language are malicious programs can be used. Thus, for example, even a program expressed in a programming language with less label information can be identified as malicious using a machine learning model learned from programs expressed in a programming language with more label information, and thus malicious programs can be identified with high accuracy.

[0046] The program according to the tenth aspect of the present disclosure is a program for causing a computer to execute the program recognition method according to any one of the first aspect to the eighth aspect.

[0047] Hereinafter, with reference to the accompanying drawings, an anomaly detection device according to an embodiment of the present disclosure will be described. In addition, all the embodiments to be described below represent a preferred specific example of the present disclosure. That is to say, the numerical values, shapes, materials, constituent elements, the arrangement and connection manners of the constituent elements, steps, the order of steps, etc. shown in the following embodiments are only examples of the present disclosure and are not intended to limit the present disclosure. The present disclosure is determined based on the description of the technical solution. Therefore, among the constituent elements of the following embodiments, the constituent elements not described in the independent claims representing the most general concept of the present disclosure are not necessary to complete the subject matter of the present disclosure, but are described as constituent elements constituting a more preferred mode.

[0048] (Embodiment)

[0049] [1. Configuration]

[0050] The program recognition device according to the embodiment is a device that uses a machine learning model generated through supervised learning to recognize whether a program expressed in a programming language different from the language of the program used in the supervised learning is a malicious program.

[0051] Figure 1 It is a block diagram showing an example of the configuration of the program recognition device according to the embodiment.

[0052] The program recognition device 100 uses a machine learning model obtained by learning using the following training data in order to recognize whether a program expressed in a second language is a malicious program. The training data includes label information indicating whether each of a plurality of programs expressed in a first language is a malicious program. The first language and the second language are different programming languages. A program expressed in the first language is called a first program. A program expressed in the second language is called a second program. For example, the number of programs included in the first program is larger than the number of programs included in the second program to which the label information is given.

[0053] The program recognition device 100 includes an acquisition unit 101, a generation unit 102, a transformation unit 103, an identification unit 104, an acquisition unit 105, a generation unit 106, a learning unit 107, and a storage unit 108. In addition, if the program recognition device 100 has a function of acquiring the following machine learning model, it may not include the acquisition unit 105, the generation unit 106, the learning unit 107, and the storage unit 108. The machine learning model is obtained by learning using training data including label information indicating whether each of a plurality of programs expressed in a first language is a malicious program.

[0054] Here, the training data that is the object of machine learning will be described.

[0055] The training data includes multiple first feature vectors and multiple identification information. The multiple first feature vectors are obtained by extracting features from each of the multiple first programs. The multiple identification information respectively indicates whether the corresponding one of the multiple first programs is malicious. Each of the multiple identification information is an example of label information. The multiple identification information respectively corresponds to the multiple first programs. The multiple first feature vectors respectively correspond to the multiple first programs. Each of the multiple first feature vectors is represented in a first form, and this first form indicates whether each of the multiple first functions of the program represented in the first language is used by the corresponding one of the first programs. The first form is common to the multiple different first programs. The first programs that are the objects for extracting the first feature vectors can be programs represented in source code or programs represented in binary code.

[0056] Figure 2 It is a figure for explaining the first feature vector.

[0057] For example, the multiple first functions of the program represented in the first language are represented by a list of multiple APIs (Application Programming Interface) that can be used by the program represented in the first language. As Figure 2 shown, the first feature vector of the first program is information indicating whether each of the multiple APIs that can be used by the program represented in the first language is used by the first program. The APIs used by the first program are, for example, the APIs called within the first program. This list can also be a group of APIs called standard APIs corresponding to the program represented in the first language. In Figure 2 the example shown, the unused APIs are represented by "0", and the used APIs are represented by "1". In this way, the first feature vector is a binary vector. In addition, the multiple APIs that can be used by the program represented in the first language are called multiple first APIs.

[0058] Return to Figure 1 and explain the configuration of the program recognition device 100.

[0059] The acquisition unit 101 acquires a second program that is an object to be recognized as a malicious or benign program. The acquisition unit 101 can acquire one second program or multiple second programs. The second programs that are the objects to be recognized can be programs represented in source code or programs represented in binary code.

[0060] The generation unit 102 generates a second feature vector by extracting the features of the second program. The second feature vector is represented in a second form, and this second form indicates whether each of a plurality of second functions of the program expressed in the second language is used by the second program. The second form is common to more than one different second program.

[0061] The plurality of second functions is a list of a plurality of APIs (Application Programming Interface) that can be used by a program expressed in the second language. That is to say, similar to the first feature vector described in Figure 2 , the second feature vector of the second program is information indicating whether each of a plurality of APIs that can be used by a program expressed in the second language is used by the second program. An API used by the second program refers to, for example, an API that is called within the second program. This list can also be a group of APIs called standard APIs corresponding to the program expressed in the second language. In addition, the plurality of APIs that can be used by a program expressed in the second language is called a plurality of second APIs.

[0062] The transformation unit 103 transforms the generated second feature vector into a first form. Specifically, the transformation unit 103 uses the correspondence between the plurality of first APIs and the plurality of second APIs to transform the second feature vector into a first form. Transforming into the first form means transforming the second feature vector into information indicating whether each of a plurality of first APIs that can be used by a program expressed in the first language is used by the second program. Transforming into the first form can also be said to be a process of mapping the feature quantity extracted from the second program in the second language to the same space as the feature quantity extracted from the first program in the first language.

[0063] Figure 3 It is a diagram for explaining a part of the relationship of the correspondence.

[0064] As Figure 3As shown, a part of the correspondence relationship, for example, indicates that a one-to-one correspondence is established between a plurality of first APIs and a plurality of second APIs. That is to say, the correspondence relationship includes a first correspondence relationship, which indicates that one first API among the plurality of first APIs and one second API among the plurality of second APIs establish a correspondence relationship. If a second API is used in the second feature vector, the transformation unit 103 regards that a first API corresponding to the one second API in the first correspondence relationship is used, and transforms the second feature vector into a first form. For example, if it is represented by "1" whether a second API is used, the transformation unit 103 sets the information indicating whether it is used corresponding to the first API corresponding to the one second API in the first correspondence relationship to "1". For example, if it is represented by "0" whether a second API is used, the transformation unit 103 sets the information indicating whether it is used corresponding to the first API corresponding to the one second API in the first correspondence relationship to "0". In this way, the second feature vector is a binary vector.

[0065] Figure 4 It is a diagram for explaining another part of the correspondence relationship.

[0066] As Figure 4 shown, another part of the correspondence relationship, for example, indicates that one second API and a plurality of first APIs establish a correspondence relationship. That is to say, the correspondence relationship includes a second correspondence relationship, which indicates that two or more other first APIs among the plurality of first APIs and one other second API among the plurality of second APIs establish a correspondence relationship. In the second correspondence relationship, utilization degrees (weights) of two or more first APIs for the one second API corresponding to them can be assigned. That is to say, the second correspondence relationship has weights among two or more first APIs for one second API.

[0067] For example, in the second correspondence relationship, it is also possible that 4 first APIs are associated with 1 second API. When 1 second API is used, it means that each of the 4 first APIs is used at 0.25. In this way, the second correspondence relationship can be the following correspondence relationship: when N first APIs are associated with 1 second API, when 1 second API is used, it means that each of the N first APIs is used at 1 / N. In addition, Figure 4 shows an example of equally distributing weights to 4 first APIs. However, different weights can also be distributed according to the degree of similarity of each of the 4 first APIs to 1 second API. In this case, the greater the degree of similarity to 1 second API, the greater the value of the weight assigned.

[0068] For example, if it is indicated by "1" whether another second API is used, the transformation unit 103 sets the information indicating whether each of the four first APIs corresponding to one second API in the first correspondence relationship is used to "0.25". For example, if it is indicated by "0" whether one second API is used, the transformation unit 103 sets the information indicating whether each of the four first APIs corresponding to one second API in the first correspondence relationship is used to "0".

[0069] In addition, Figure 4 shows that the correspondence relationship includes a plurality of first correspondence relationships and a second correspondence relationship.

[0070] As Figure 3 and Figure 4 shown, the correspondence relationship can also be an API transformation table indicating which API in the second language corresponds to the API in the first language.

[0071] In addition, the correspondence relationship can also include a third correspondence relationship, which indicates that one first API among a plurality of first APIs corresponds to two or more second APIs among a plurality of second APIs.

[0072] The correspondence relationship can also be composed of a combination of the first correspondence relationship, the second correspondence relationship, and the third correspondence relationship.

[0073] In Non-Patent Document 2, the API2vec method is used to transform a plurality of APIs that can be used in Java JDK and C#.NET respectively into vector representations 10 for visualization.

[0074] The transformation unit 103 can transform APIs that are similar in the vector representation 10 obtained by using API2vec into similar vector representations 10. Therefore, the transformation unit 103 can determine, for example, using the nearest neighbor method, the API in the vector representation 10 of the APIs in the target language that is closest to the vector representation 10 of the APIs in the source language. In addition, regarding the method of determining APIs similar to the vector representation 10 of the APIs in the source language, not only the nearest neighbor method can be used, but also a method using the K-nearest neighbor method or a method of extracting features by dispersing to a plurality of APIs according to the distance from the source API can be used.

[0075] In this way, the correspondence relationship can also represent the similarity between the vector representation 10 of each of the plurality of first APIs and the vector representation 10 of each of the plurality of second APIs.

[0076] The recognition unit 104 outputs a recognition result indicating whether the second program is a malicious program, which is obtained by inputting the second feature vector transformed into the first form into the machine learning model. The recognition unit 104 obtains the machine learning model from the storage unit 108.

[0077] The acquisition unit 105 acquires training data.

[0078] The generation unit 106 generates a plurality of first feature vectors by extracting features from each of the plurality of first programs. The generation unit 106 outputs learning information including the plurality of first feature vectors and a plurality of recognition information to the learning unit 107, and the plurality of recognition information respectively indicates whether the corresponding one of the plurality of first programs is malicious.

[0079] The learning unit 107 generates a machine learning model by learning using the learning information. The learning unit 107 stores the generated machine learning model in the storage unit 108. In the learning performed by the learning unit 107, any method can be used as long as it is a method used in supervised learning. For example, logistic regression can also be used.

[0080] The storage unit 108 stores the machine learning model generated by the learning unit 107.

[0081] In addition, as described above, the acquisition unit 105, the generation unit 106, the learning unit 107, and the storage unit 108 may not be provided in the program recognition device 100, or may be provided in an external device communicably connected to the program recognition device 100. In this case, the program recognition device 100 obtains the machine learning model from the external device.

[0082] [2. Operation]

[0083] Next, the operation of the program recognition device 100 according to the embodiment will be described.

[0084] First, the learning process of the program recognition device 100 will be described. Figure 5 It is a flowchart showing an example of the learning process performed by the program recognition device.

[0085] The program recognition device 100 acquires training data (S11). Step S11 is the process performed by the acquisition unit 105.

[0086] Next, the program recognition device 100 extracts features from each of the plurality of first programs to generate a plurality of first feature vectors (S12). Thus, learning information including the plurality of first feature vectors and a plurality of recognition information respectively indicating whether the corresponding one of the plurality of first programs is a malicious program is generated. Step S12 is the process performed by the generation unit 106.

[0087] Next, the program recognition device 100 generates a machine learning model by learning using the learning information (S13). Step S13 is a process performed by the learning unit 107.

[0088] Next, the program recognition device 100 stores the generated machine learning model in the storage unit 108 (S14). Step S14 is a process that uses the storage unit 108 and is performed by the learning unit 107.

[0089] Next, the recognition process of the second program of the program recognition device 100 will be described. Figure 6 It is a flowchart showing an example of the recognition process performed by the program recognition device.

[0090] The program recognition device 100 acquires a second program that is an object of recognition as to whether the recognized program is malicious or a benign program (S21). Step S21 is a process performed by the acquisition unit 101.

[0091] Next, the program recognition device 100 generates a second feature vector by extracting the features of the second program (S22). Step S22 is a process performed by the generation unit 102.

[0092] Next, the program recognition device 100 transforms the form of the generated second feature vector from the second form to the first form (S23). Step S23 is a process performed by the transformation unit 103.

[0093] Next, the program recognition device 100 outputs an identification result indicating whether the second program is a malicious program, which is obtained by inputting the second feature vector transformed into the first form into the machine learning model (S24). Step S24 is a process performed by the identification unit 104.

[0094] [3. Effects, etc.]

[0095] The program recognition device 100 according to this embodiment executes a program recognition method. In the program recognition method, first, a machine learning model generated by learning using training data indicating whether each of a plurality of first programs is a malicious program is obtained (S21). Here, each of the plurality of first programs is expressed in a first language. The training data includes a plurality of first feature vectors and a plurality of recognition information. The plurality of first feature vectors are obtained by extracting features from each of the plurality of first programs, and the plurality of recognition information respectively indicates whether a corresponding one of the plurality of first programs is malicious. Each of the plurality of first feature vectors is expressed in a first form, and the first form indicates whether each of a plurality of first functions of the program expressed in the first language is used by a corresponding one of the first programs. Further, in the program recognition method, a second feature vector is then generated by extracting features of a second program expressed in a second language different from the first language (S22). Here, the second feature vector is expressed in a second form, and the second form indicates whether each of a plurality of second functions of the program expressed in the second language is used by the second program. Next, in the program recognition method, the generated second feature vector is transformed into the first form (S23), and an identification result indicating whether the second program is a malicious program obtained by inputting the second feature vector transformed into the first form into the machine learning model is output (S24).

[0096] Accordingly, in order to identify whether a second program expressed in a second language is a malicious program, a machine learning model obtained by machine learning regarding whether a plurality of first programs expressed in a first language are malicious programs can be used. Thus, for example, even for a program expressed in a programming language with less label information, a machine learning model learned through programs expressed in a programming language with more label information can be used to identify whether it is malicious, and thus malicious programs can be identified with high accuracy.

[0097] In addition, in the program recognition device 100 according to this embodiment, in the transformation, the correspondence between a plurality of first APIs (first functions) and a plurality of second APIs (second functions) is used to transform the second feature vector into the first form.

[0098] Therefore, the form of the second feature vector can be easily transformed into the form of the first feature vector. Thus, even when using a machine learning model learned for programs in other programming languages, it is possible to accurately identify whether the second program is malicious.

[0099] In addition, in the program recognition device 100 according to this embodiment, the correspondence indicates that one first API among a plurality of first APIs (first functions) is in correspondence with one second API among a plurality of second APIs (second functions).

[0100] Therefore, the form of the second feature vector can be easily transformed into the form of the first feature vector.

[0101] In addition, in the program recognition device 100 according to the present embodiment, the correspondence relationship indicates that two or more other first APIs among the plurality of first APIs (first functions) are in correspondence with one other second API among the plurality of second APIs (second functions).

[0102] Therefore, even when two or more other first APIs are in correspondence with one other second API, the form of the second feature vector can be easily transformed into the form of the first feature vector.

[0103] In addition, in the program recognition device 100 according to the present embodiment, the correspondence relationship has weights among two or more other first APIs (first functions) for one other second API (second function).

[0104] Therefore, according to the relationship between each of the two or more other first APIs and one other second API, the form of the second feature vector can be easily transformed into the form of the first feature vector.

[0105] In addition, in the program recognition device 100 according to the present embodiment, the correspondence relationship represents the similarity between the vector representations of each of the plurality of first APIs (first functions) and the vector representations of each of the plurality of second APIs (second functions).

[0106] Therefore, the form of the second feature vector can be easily transformed into the form of the first feature vector.

[0107] [4. Modification Example]

[0108] The program recognition device 100 according to the above-described embodiment may further determine, among the plurality of first programs that have been learned and are given malicious label information, a first program that is similar to the second program identified as a malicious program. Here, the configuration for performing the process of determining the first program that is similar to the second program identified as a malicious program will be described. In addition, this configuration is implemented in the configuration included in the program recognition device 100.

[0109] The learning unit 107 generates, for each of one or more first programs represented as malicious programs by the training data through machine learning described in the embodiment, first malicious information that includes one or more first malicious contribution degrees, and the one or more first malicious contribution degrees respectively correspond to one or more first APIs represented as being used in the first feature vector corresponding to the first program. Thus, the first malicious information is obtained.

[0110] Figure 7 is information showing an example of the first malicious information.

[0111] As Figure 7 shown, the first malicious information includes a plurality of first malicious contribution degrees respectively corresponding to a plurality of first APIs that can be used by a program expressed in the first language. The plurality of first malicious contribution degrees are calculated when generating a machine learning model.

[0112] The identification unit 104 generates second malicious information, which includes one or more second malicious contribution degrees, and the one or more second malicious contribution degrees respectively correspond to one or more first APIs that are represented as being used in a second feature vector corresponding to a second program whose output identification result indicates maliciousness and transformed into a first form. Thus, the second malicious information is obtained.

[0113] Since the second malicious information is malicious information corresponding to the second feature vector transformed into the first form, it is the same as the first malicious information described in Figure 7 and includes a plurality of second malicious contribution degrees respectively corresponding to a plurality of first APIs that can be used by a program expressed in the first language. The plurality of second malicious contribution degrees are calculated when the second feature vector transformed into the first form is input into the machine learning model.

[0114] The identification unit 104 determines the first program corresponding to the similar malicious information similar to the second malicious information by comparing each of the one or more first malicious information corresponding to each of the one or more first programs obtained and the second malicious information. Specifically, when making the determination, the identification unit 104 uses M (M is an integer of 2 or more) first malicious contribution degrees selected in descending order of contribution degree from the plurality of first malicious contribution degrees and M second malicious contribution degrees selected in descending order of contribution degree from the plurality of second malicious contribution degrees, calculates the similarity between the second malicious information and the first malicious information to be the comparison object, and determines the similar malicious information based on the similarity.

[0115] Using Figure 8 , a specific example of a method for calculating the similarity between the first program and the second program is described.

[0116] Figure 8 is a diagram for explaining a method for calculating the similarity between the first program and the second program.

[0117] Figure 8 The (a) of Figure 8 represents the malicious contribution degrees corresponding to the respective elements (each API) of the transformed second feature vector, Figure 8The (c) represents the malicious contribution degree corresponding to each element (each API) of the first feature vector of another (learning program 2) in the malicious first program. That is to say, Figure 8 The (a) of Figure 8 is an example of the second malicious information, Figure 8 The (b) and (c) of Figure 8 are examples of the first malicious information. The first malicious information and the second malicious information include identification information for identifying the first API, usage information indicating whether the corresponding first API is used, and the malicious contribution degree of the corresponding first API.

[0118] Based on the second malicious information, the identification unit 104 determines three first APIs corresponding to three second malicious contribution degrees selected in descending order of contribution degree from the malicious contribution degrees of a plurality of first APIs used by the second program to be identified. As the three first APIs, the identification unit 104 determines the first APIs with indexes "1", "3", and "5".

[0119] Based on the first malicious information of the learning program 1, the identification unit 104 determines three first APIs corresponding to three first malicious contribution degrees selected in descending order of contribution degree from the malicious contribution degrees of a plurality of first APIs used by the learning program 1. As the three first APIs, the identification unit 104 determines the first APIs with indexes "1", "3", and "7".

[0120] The identification unit 104 calculates the sum of the malicious contribution degrees corresponding to the indexes "1" and "3" included in the intersection of the set of the first APIs with indexes "1", "3", and "5" determined using the second malicious information and the set of the first APIs with indexes "1", "3", and "7" determined using the first malicious information of the learning program 1 (that is, 0.3 + 0.8). Next, the identification unit 104 calculates the sum of the malicious contribution degrees corresponding to the indexes "1", "3", "5", and "7" included in the union of the set of the first APIs with indexes "1", "3", and "5" determined using the second malicious information and the set of the first APIs with indexes "1", "3", and "7" determined using the first malicious information of the learning program 1 (that is, 0.3 + 0.8 + 0.1 + 0.8). Moreover, the identification unit 104 calculates the value obtained by dividing the sum of the malicious contribution degrees of the intersection by the sum of the malicious contribution degrees of the union as the similarity between the second malicious information and the first malicious information to be compared. In this case, the similarity is calculated as 0.55.

[0121] In addition, the recognition unit 104 determines three first APIs corresponding to three first malicious contribution degrees selected in descending order of contribution degree from the malicious contribution degrees of a plurality of first APIs used by the learning program 1 based on the first malicious information of the learning program 2. As the three first APIs, the recognition unit 104 determines the first APIs with indexes "3", "8", and "10".

[0122] The recognition unit 104 calculates the sum of the malicious contribution degrees corresponding to the index "3" included in the intersection of the set of the first APIs with indexes "1", "3", and "5" determined using the second malicious information and the set of the first APIs with indexes "3", "8", and "10" determined using the first malicious information of the learning program 1 (that is, 0.8). Next, the recognition unit 104 calculates the sum of the malicious contribution degrees corresponding to the indexes "1", "3", "5", "8", and "10" included in the union of the set of the first APIs with indexes "1", "3", and "5" determined using the second malicious information and the set of the first APIs with indexes "3", "8", and "10" determined using the first malicious information of the learning program 1 (that is, 0.3 + 0.8 + 0.1 + 0.1 + 0.8). Moreover, the recognition unit 104 calculates the value obtained by dividing the sum of the malicious contribution degrees of the intersection by the sum of the malicious contribution degrees of the union as the similarity between the second malicious information and the first malicious information to be the comparison object. In this case, the similarity is calculated as 0.38.

[0123] In addition, the higher the degree of similarity, the larger the calculated similarity value.

[0124] When the calculated similarity is greater than a specified threshold, the recognition unit 104 may determine the first program that is the calculation object of the similarity as a similar program similar to the second program, or may determine the first program with the largest similarity or the top N (N is an integer greater than or equal to 1) as a similar program similar to the second program.

[0125] Moreover, the recognition unit 104 outputs information indicating the determined first program.

[0126] Figure 9 It is a flowchart showing an example of the process of determining the first program similar to the second program.

[0127] In addition to the processing performed by the program recognition device 100 according to the embodiment, the program recognition device 100 according to Modification 1 further obtains, for each of one or more first programs represented as malicious programs by the training data, first malicious information including one or more first malicious contribution degrees, and the one or more first malicious contribution degrees respectively correspond to one or more first functions represented as being used in a first feature vector corresponding to the first program (S31).

[0128] Next, the program recognition device 100 obtains second malicious information including one or more second malicious contribution degrees, and the one or more second malicious contribution degrees respectively correspond to one or more first functions represented as being used in a second feature vector corresponding to a second program whose recognition result indicates maliciousness and transformed into a first form (S32).

[0129] Next, the program recognition device 100 determines a first program corresponding to similar malicious information similar to the second malicious information by comparing each of the one or more first malicious information corresponding to each of the one or more first programs obtained and the second malicious information (S33).

[0130] Next, the program recognition device 100 outputs information indicating the determined first program (S34).

[0131] Therefore, it is possible to output a first program having characteristics similar to those of the second program.

[0132] In the program recognition device 100 according to the modification, a plurality of first malicious contribution degrees are calculated when generating a machine learning model. A plurality of second malicious contribution degrees are calculated when inputting a second feature vector transformed into a first form into the machine learning model.

[0133] Therefore, it is possible to output a first program having characteristics similar to those of the second program by using the malicious contribution degrees corresponding to the elements (each API) of the feature vectors of one or more first programs and the second program respectively.

[0134] In the program recognition device 100 according to the modification, in the determination process, M (M is an integer of 2 or more) first malicious contribution degrees selected in descending order of contribution degree from a plurality of first malicious contribution degrees and M second malicious contribution degrees selected in descending order of contribution degree from a plurality of second malicious contribution degrees are used to calculate the similarity between the second malicious information and the first malicious information to be a comparison object, and similar malicious information is determined based on the similarity.

[0135] Therefore, it is possible to simplify the calculation of the similarity for determining similar malicious information.

[0136] (Other Embodiments)

[0137] In the above embodiments, each component may be constituted by dedicated hardware, or may also be implemented by executing a software program suitable for each component. Each component may be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.

[0138] In addition, each component may be a circuit (or an integrated circuit). These circuits may form a single circuit as a whole, or may be separate circuits respectively. In addition, these circuits may be general-purpose circuits respectively, or may be dedicated circuits.

[0139] In addition, the general or specific technical solutions of the present disclosure may be implemented by a system, a device, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM. In addition, it may also be implemented by any combination of a system, a device, a method, an integrated circuit, a computer program, and a computer-readable non-transitory recording medium.

[0140] For example, the present disclosure may also be implemented as a program recognition method executed by a program recognition device (a computer or a DSP), and may also be implemented as a program for causing a computer or a DSP to execute the above program recognition method.

[0141] In addition, in the above embodiments, the processing executed by a specific processing unit may be executed by another processing unit. In addition, the order of multiple processes in the operation of the program recognition device described in the above embodiments may also be changed, and multiple processes may also be executed in parallel.

[0142] In addition, modes obtained by applying various modifications conceivable by those skilled in the art to each embodiment, or modes implemented by arbitrarily combining the components and functions in each embodiment without departing from the gist of the present disclosure are also included in the present disclosure.

[0143] Industrial Applicability

[0144] The present disclosure is useful as a program recognition method or the like that can detect malicious programs with high accuracy.

[0145] Description of Reference Numerals

[0146] 100 Program recognition device

[0147] 101, 105 Acquisition unit

[0148] 102, 106 Generation unit

[0149] 103 Transformation unit

[0150] 104 Recognition unit

[0151] 107 Learning Department

[0152] 108 Storage Department

Claims

1. A program recognition method, comprising: Obtaining a machine learning model generated by learning using training data indicating whether each of a plurality of first programs is a malicious program, Each of the plurality of first programs is expressed in a first language, The machine learning model is a model generated by learning learning data generated based on the training data, the learning data includes a plurality of first feature vectors and a plurality of recognition information, the plurality of first feature vectors are obtained by extracting features from each of the plurality of first programs, and the plurality of recognition information respectively indicates whether a corresponding one of the plurality of first programs is malicious, Each of the plurality of first feature vectors is expressed in a first form, and the first form indicates whether each of a plurality of first functions of a program expressed in the first language is used by a corresponding one of the first programs; Generating a second feature vector by extracting features of a second program expressed in a second language different from the first language, The second feature vector is expressed in a second form, and the second form indicates whether each of a plurality of second functions of a program expressed in the second language is used by the second program; Transforming the generated second feature vector into the first form; Outputting an identification result indicating whether the second program is a malicious program obtained by inputting the second feature vector transformed into the first form into the machine learning model.

2. The program recognition method according to claim 1, In the transformation, using the correspondence between the plurality of first functions and the plurality of second functions, the second feature vector is transformed into the first form.

3. The program recognition method according to claim 2, The correspondence indicates that one first function among the plurality of first functions corresponds to one second function among the plurality of second functions.

4. The program recognition method according to claim 3, The correspondence indicates that two or more other first functions among the plurality of first functions correspond to one other second function among the plurality of second functions.

5. The program recognition method according to claim 4, The correspondence has weights among the two or more other first functions for the one other second function.

6. The program recognition method according to claim 2, The correspondence indicates the similarity between the vector representation of each of the plurality of first functions and the vector representation of each of the plurality of second functions.

7. The program recognition method according to any one of claims 1 to 6, further comprising: For each of one or more first programs represented as malicious programs by the training data, obtaining first malicious information, the first malicious information includes one or more first malicious contribution degrees, and the one or more first malicious contribution degrees respectively correspond to one or more first functions represented as being used in the first feature vector corresponding to the first program; Obtain second malicious information, where the second malicious information includes more than one second malicious contribution degree, and each of the more than one second malicious contribution degrees corresponds to one or more first functions represented as being used in a second feature vector that corresponds to a second program indicated as malicious by the recognition result and is transformed into the first form; By comparing each of the more than one first malicious information corresponding to each of the more than one first programs obtained and the second malicious information, determine the first program corresponding to the similar malicious information similar to the second malicious information; Output information indicating the determined first program.

8. The program recognition method according to claim 7, In the determination, use M first malicious contribution degrees selected from the more than one first malicious contribution degrees in descending order of contribution degree and M second malicious contribution degrees selected from the more than one second malicious contribution degrees in descending order of contribution degree, calculate the similarity between the second malicious information and the first malicious information to be the comparison object, and determine the similar malicious information according to the similarity, where M is an integer greater than or equal to 2.

9. A program recognition device includes a processor and a memory, The processor uses the memory to execute: Generate a machine learning model by learning using training data indicating whether each of a plurality of first programs is a malicious program, Each of the plurality of first programs is expressed in a first language, The machine learning model is a model generated by learning learning data generated based on the training data. The learning data includes a plurality of first feature vectors and a plurality of recognition information. The plurality of first feature vectors are obtained by extracting features from each of the plurality of first programs, and the plurality of recognition information respectively indicate whether a corresponding one of the plurality of first programs is malicious, Each of the plurality of first feature vectors is expressed in a first form, and the first form indicates whether each of a plurality of first functions of a program expressed in the first language is used by a corresponding first program; Transform a second feature vector obtained by extracting features of a second program expressed in a second language different from the first language into the first form, The second feature vector is expressed in a second form, and the second form indicates whether each of a plurality of second functions of a program expressed in the second language is used by the second program; Output a recognition result indicating whether the second program is a malicious program obtained by inputting the second feature vector transformed into the first form into the machine learning model.

10. A program for causing a computer to execute the program recognition method according to any one of claims 1 to 6.