Collection device, collection method, and collection program

The system addresses the limitations of conventional phishing attack detection by using security and co-occurrence keywords to collect and analyze tweets, enabling comprehensive threat intelligence through text and image extraction.

JP7838668B2Active Publication Date: 2026-04-01NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-10-27
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Conventional technologies are limited in the scope of information collection for phishing attacks, primarily focusing on specific keywords and text formats, failing to extract a wide range of security threat information, including images and reports from various users.

Method used

A system comprising a collection device and classification device that uses security keywords and co-occurrence keywords to collect and analyze tweets, extracting both text and image data from social networking service posts to identify phishing attacks.

Benefits of technology

The system effectively collects and classifies a wide range of phishing attack reports from various users, extracting valuable information from both text and images, enhancing the accuracy and scale of threat intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007838668000001
    Figure 0007838668000001
  • Figure 0007838668000002
    Figure 0007838668000002
  • Figure 0007838668000003
    Figure 0007838668000003
Patent Text Reader

Abstract

This collection device uses security keywords to collect tweets related to reports of phishing attacks from tweets of each user. Then, the collection device extracts, from the collected tweets, co-occurrence keywords that co-occur more frequently than a predetermined frequency. Subsequently, the collection device collects, from the tweets of each user, tweets containing the co-occurrence keywords and data (for example, text and image) associated with the tweets. Subsequently, the collection device screens posts which are highly possibly related to reports of phishing attacks on the basis of URLs or domain names extracted from the text and images of the collected tweets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a collection device, a collection method, and a collection program for collecting posts related to security threat information.

Background Art

[0002] On social platforms, in addition to security experts, ordinary users with good intentions also share many cases of suspicious phishing attacks they have observed themselves as warnings, such as images (e.g., screenshots). If these information can be collected, analyzed, and extracted as early and accurately as possible, it is useful for countermeasures against phishing attacks.

[0003] Targets for extracting security threat information such as phishing attacks include security blogs, security reports, social platforms, etc.

[0004] For example, as in Non-Patent Documents 3 and 4, by applying natural language processing technology to blogs and reports summarizing threat information analyzed by security experts and extracting it as formalized data, it can be mechanically utilized.

[0005] Also, in Non-Patent Document 5, as collection targets for threat information, Twitter (registered trademark), Facebook (registered trademark), news sites, security blogs, security forums, etc. were compared and evaluated, and it was reported that Twitter is the most excellent in both the quantity and quality of collectable information.

[0006] In Non-Patent Documents 6, 7, and 8, technologies for extracting URLs, domain names, hash values, IP addresses, vulnerability information, etc. related to threats from the Tweets of each user by focusing on specific users and keywords on Twitter have been proposed. According to this technology, it has been reported that a large number of useful threat information can be obtained.

Prior Art Documents

Non-Patent Documents

[0007] [Non-Patent Document 1] Phishing attacks continue to rage - approximately 270 unique URLs per day on average, Security NEXT, [online], [searched October 13, 2022], Internet<URL:https: / / www.security-next.com / 134607> [Non-Patent Document 2] Phishing Report Status, February 2022, [online], Council of Anti-Phishing Japan, [Retrieved October 13, 2022], Internet<URL:https: / / www.antiphishing.jp / report / monthly / 202202.html> [Non-Patent Document 3] Zhu, Ziyun and Dumitras, Tudor, “ChainSmith: Automatically Learning the Semantics of Malicious Campaigns by Mining Threat Intelligence Reports”, 2018 IEEE European Symposium on Security and Privacy [Non-Patent Document 4] Satvat, Kiavash and Gjomemo, Rigel and Venkatakrishnan, VN, “EXTRACTOR: Extracting Attack Behavior from Threat Reports”, IEEE EuroS&P 2021. [Non-Patent Document 5] Shin, Hyejin and Shim, WooChul and Moon, Jiin and Seo, Jae Woo and Lee, Sol and Hwang, Yong Ho, “Cybersecurity Event Detection with New and Re-emerging Words,” ASIA CCS 2020. [Non-Patent Document 6] Alves, Fernando and Andongabo, Ambrose and Gashi, Ilir and Ferreira, Pedro M. and Bessani, Alysson, “Follow the Blue Bird: A Study on threat data published on Twitter”, ESORICS 2020. [Non-Patent Document 7] Shin, Hyejin and Shim, WooChul and Kim, Saebom and Lee, Sol and Kang, Yong Goo and Hwang, Yong Ho, “#Twiti: Social Listening for Threat Intelligence”, WWW 2021. [Non-Patent Document 8] Roy, Sayak Saha and Karanjit, Unique and Nilizadeh, Shirin, “Evaluating the Effectiveness of Phishing Reports on Twitter”, eCrime 2021. [Overview of the project] [Problems that the invention aims to solve]

[0008] However, the above-mentioned conventional technologies have the following problems.

[0009] (1) The number of Tweets collected is limited. Conventional technologies limit information collection to specific user accounts, making it impossible to collect information on phishing attacks reported by various users. Furthermore, conventional technologies only collect tweets within a limited range, as they target specific keywords such as "#phishing" and "#warning."

[0010] (2) The information to be extracted is limited to text of a certain format contained in the Tweet. Reports of phishing attacks via Tweet often include images such as screenshots, but conventional techniques only extract information from the text within the Tweet. Therefore, conventional techniques cannot extract information contained within images. Furthermore, since users post information in various formats, conventional techniques that are specialized for a certain format can only extract limited information.

[0011] As a result, conventional technologies have the problem of not being able to extract a wide range of security threat information. Therefore, the present invention aims to solve the above problem and extract a wide range of security threat information. [Means for solving the problem]

[0012] To solve the aforementioned problems, the present invention is characterized by comprising: a first collection unit that collects posts related to security threats from SNS (Social Networking Service) posts using security keywords, which are keywords related to security threats; a keyword extraction unit that extracts co-occurring keywords, which are keywords that co-occur more than a predetermined frequency, from the collected posts related to security threats; and a second collection unit that collects posts containing the co-occurring keywords and images associated with those posts from SNS posts. [Effects of the Invention]

[0013] According to the present invention, a wide range of security threat information can be extracted. [Brief explanation of the drawing]

[0014] [Figure 1] Figure 1 shows an example of the system configuration. [Figure 2A] Figure 2A shows an example of the configuration of the collection device. [Figure 2B] Figure 2B is a flowchart showing an example of the processing procedure performed by the collection device. [Figure 3] Figure 3 is a diagram illustrating a specific example of the processing procedure performed by the collection device. [Figure 4] Figure 4 is a diagram showing an example of security keywords. [Figure 5] Figure 5 is a diagram for explaining an example of generating Co-occurrence Keywords. [Figure 6] Figure 6 is a diagram showing an example of a Tweet targeted for data collection. [Figure 7] Figure 7 is a diagram for explaining the process of extracting URLs and domain names from Tweet text and images. [Figure 8A] Figure 8A is a diagram showing a configuration example of a classification device. [Figure 8B] Figure 8B is a flowchart showing an example of a processing procedure executed by the classification device. [Figure 9] Figure 9 is a diagram for explaining a specific example of a processing procedure executed by the classification device. [Figure 10] Figure 10 is a diagram showing an example of features generated from Tweets. [Figure 11] Figure 11 is a diagram showing an example of Tweet Account Feature. [Figure 12] Figure 12 is a diagram showing an example of Tweet Content Feature. [Figure 13] Figure 13 is a diagram showing an example of Tweet URL Feature. [Figure 14] Figure 14 is a diagram showing an example of Tweet OCR Feature. [Figure 15] Figure 15 is a diagram showing an example of Tweet Visual Feature. [Figure 16] Figure 16 is a diagram showing an example of Tweet Context Feature. [Figure 17] Figure 17 is a diagram showing an example of features selected by the selection unit in Figure 8A. [Figure 18] Figure 18 is a diagram showing the evaluation result of the classification accuracy of the system. [Figure 19]Figure 19 shows the number of phishing attack reports and URLs associated with phishing attacks extracted by the system over a predetermined period. [Figure 20] Figure 20 shows a comparison of the system and OpenPhish. [Figure 21] Figure 21 shows the comparison results between the system and PhishTank. [Figure 22] Figure 22 shows the results of a survey comparing the number of user reports and the number of phishing URLs. [Figure 23] Figure 23 shows the effects of dynamically selecting keywords. [Figure 24] Figure 24 shows a computer running a program. [Modes for carrying out the invention]

[0015] The following describes embodiments for carrying out the present invention with reference to the drawings. The present invention is not limited to these embodiments.

[0016] [overview] First, using Figure 1, we will describe the overview of the system comprising the collection device and classification device of this embodiment.

[0017] The explanation will use Twitter posts (Tweets) as an example of the SNS (Social Networking Service) posts handled by the system, but it is not limited to this. Also, SNS posts can be in Japanese or English.

[0018] Furthermore, in this embodiment, the system is described as collecting posts reporting phishing attacks from SNS posts as an example, but it may also collect posts reporting security threats other than phishing attacks.

[0019] The system, for example, extracts tweets reporting phishing attacks from each user's tweets early and with high accuracy. For example, the system consists of a collection device 10 and a classification device 20. The collection device 10 and the classification device 20 may be connected via a network such as the Internet, or they may be installed in the same device.

[0020] (1) Collection device 10: Collects a wide range of tweets that may be reports of phishing attacks. For example, collection device 10 extracts co-occurrence keywords that appear in reports of phishing attacks. Then, collection device 10 uses security keywords and the above co-occurrence keywords to collect a wide range of tweets that may be reports of phishing attacks (Screened Tweets in Figure 1).

[0021] (2) Classification device 20: Classifies phishing attack reports from among the Tweets collected by the collection device 10. For example, the classification device 20 extracts text and image features of phishing attack reports from Tweets using machine learning, and uses these extracted features to classify whether each Tweet is a phishing attack report or another Tweet.

[0022] Furthermore, after the classification of Tweets by the classification device 20, the collection device 10 may extract Co-occurrence Keywords from the group of Tweets classified as phishing attack reports. The collection device 10 may then use the extracted Co-occurrence Keywords to collect Tweets that are likely to be phishing attack reports. In this way, the system can dynamically expand / narrow the keywords used to collect Tweets that are likely to be phishing attack reports, and collect Tweets that should be collected at the appropriate time.

[0023] Such a system can collect phishing attack reports via Tweets not only from security experts but also from well-meaning ordinary users. Furthermore, because the system collects Tweets using a wide range of keywords, it can analyze phishing attack reports on a large scale.

[0024] Furthermore, the system can accurately extract reports of phishing attacks from the large volume of tweets it collects. In addition, because the system extracts information about phishing attacks from both the text and images contained in the tweets, it can extract useful information that would not be obtainable by analyzing only the text of the tweets.

[0025] This system offers the following benefits in combating phishing attacks: (1) It becomes possible to collect threat information from a wider range of areas than the limited monitoring scope of conventional technologies, and to provide threat information from a new perspective.

[0026] (2) In particular, it will be possible to quickly provide threat information that can be used to counter phishing attacks targeting Japanese people, which has been lacking until now.

[0027] (3) Applying the data obtained by this system to the filtering rules of telecommunications carriers will help reduce the number of victims of phishing attacks and other similar attacks.

[0028] [Multiple devices] [Example Configuration] Next, the data collection device 10 will be described in detail. First, an example of the configuration of the data collection device 10 will be described using Figure 2A. The data collection device 10 includes, for example, an input / output unit 11, a storage unit 12, and a control unit 13.

[0029] The input / output unit 11 is an interface that handles the input and output of various types of data. For example, the input / output unit 11 accepts Tweets collected from Twitter as input. The input / output unit 11 also outputs Tweets that may be reports of phishing attacks, extracted by the control unit 13 (Screened Tweets in Figure 1).

[0030] The memory unit 12 stores data, programs, etc., that are referenced when the control unit 13 performs various processes. The memory unit 12 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or storage devices such as hard disks and optical discs. The memory unit 12 stores, for example, security keywords, co-occurrence keywords, etc., extracted by the control unit 13.

[0031] The control unit 13 is responsible for controlling the entire data collection device 10. The functions of the control unit 13 are realized, for example, by the CPU (Central Processing Unit) executing a program stored in the memory unit 12.

[0032] The control unit 13 includes, for example, a first collection unit 131, a keyword extraction unit 132, a second collection unit 133, and a data collection unit 134. Note that the URL / domain name extraction unit 135 and the sorting unit 136, shown by dashed lines, may or may not be included; the cases in which they are included will be described later.

[0033] The first collection unit 131 uses Security Keywords, which are keywords related to security threats, to collect tweets reporting phishing attacks from each user's tweets.

[0034] The keyword extraction unit 132 extracts co-occurrence keywords, which are keywords that co-occur more than a predetermined frequency, from the phishing attack report tweets collected by the first collection unit 131. These co-occurrence keywords may also be extracted from tweets classified as phishing attack report tweets by the classification device 20.

[0035] The second collection unit 133 uses Co-occurrence Keywords to collect tweets from each user that may be reports of phishing attacks. For example, the second collection unit 133 collects tweets from each user that contain Security Keywords and Co-occurrence Keywords in the text of the tweet or in the image associated with the tweet. The collected tweets are stored, for example, in the storage unit 12.

[0036] The data collection unit 134 collects data necessary for input to the classification device 20. For example, the data collection unit 134 collects the following data from Tweets collected by the second collection unit 133: (1) the Tweet string (e.g., hashtags, number of characters, etc.), (2) metadata associated with the Tweet (e.g., application information, presence or absence of defang, etc.), (3) information about the Tweet account (e.g., number of followers of the account, account registration period, etc.), and (4) images included in the Tweet (e.g., up to 4 images associated with the Tweet). The collected data is stored in the storage unit 12, for example.

[0037] [Example of processing procedure] Next, an example of the processing procedure performed by the collection device 10 will be explained using Figure 2B. First, the first collection unit 131 of the collection device 10 collects Tweets reporting phishing attacks using, for example, Security Keywords (S1: Collection of Tweets using Security Keywords). Then, the keyword extraction unit 132 extracts Co-occurrence Keywords, which are keywords that co-occur more than a predetermined frequency, from the Tweets reporting phishing attacks collected in S1 (S2: Extraction of Co-occurrence Keywords).

[0038] After S2, the second collection unit 133 collects tweets from each user that may be reports of phishing attacks using Security Keywords and Co-occurrence Keywords (S3). Subsequently, the data collection unit 134 collects data necessary for input to the classification device 20 from the tweets collected in S3 (S4).

[0039] By performing the above process, the collection device 10 can collect Tweets that may be reports of phishing attacks.

[0040] The collection device 10 may also include a URL / domain name extraction unit 135 and a sorting unit 136, as shown in Figure 2A.

[0041] The URL / domain name extraction unit 135 extracts URLs and domain names from the text and images of Tweets collected by the second collection unit 133. The selection unit 136 selects Tweets that are likely to be phishing attack reports from the Tweets collected by the second collection unit 133 based on the URLs or domain names extracted by the URL / domain name extraction unit 135.

[0042] For example, the selection unit 136 selects a Tweet that is likely to be a phishing attack report if the URL or domain included in the Tweet collected by the second collection unit 133 is not included in the list of legitimate website URLs or domain names. The selection unit 136 also selects a Tweet that is likely to be a phishing attack report if the domain name of the URL included in the Tweet has been in use for less than a predetermined period. For example, the selection unit 136 selects a domain name that has been registered in WHOIS for less than a predetermined number of days as a Tweet that is likely to be a phishing attack report.

[0043] Subsequently, the data collection unit 134 collects data necessary for input to the classification device 20 (for example, the string of the Tweet) from the Tweets selected by the sorting unit 136.

[0044] In this way, the collection device 10 can collect Tweets and their data that are more likely to be reports of phishing attacks from the collected Tweets.

[0045] [Specific example of processing procedure] Next, using Figure 3, a specific example of the processing procedure performed by the collection device 10 will be explained. The explanation will be based on the case where the collection device 10 is equipped with a URL / domain name extraction unit 135 and a sorting unit 136.

[0046] (1) Generating Keywords The collection device 10 generates two types of keywords (Security Keywords and Co-occurrence Keywords) for searching for Tweets that contain reports of phishing attacks.

[0047] (1-1) Security Keywords First, let's explain Security Keywords. For example, the collection device 10 generates Security Keywords such as keywords related to security threats and the media through which they are spread, such as "SMS" and "fake website," and keywords for sharing security threat information such as "#phishing" and "#fraud" (see Figure 4). Note that these Security Keywords may also use existing keywords related to security threats.

[0048] (1-2) Security Keywords Next, we will explain Co-occurrence Keywords. For example, the collection device 10 extracts keywords that co-occur more frequently than a predetermined value (Co-occurrence Keywords) only in phishing attack reports collected using Security Keywords as the key.

[0049] For example, the first collection unit 131 of the collection device 10 collects tweets reporting phishing attacks from each user's tweets using Security Keywords. Subsequently, the keyword extraction unit 132 extracts Co-occurrence Keywords from the collected tweets. For example, the keyword extraction unit 132 extracts new Co-occurrence Keywords from the tweets collected during a predetermined period.

[0050] For example, the keyword extraction unit 132 extracts proper nouns from the strings of Tweets over a predetermined period and calculates PMI (Pointwise Mutual Information) using the following formula (1). In formula (1), X and Y are proper nouns contained in the Tweets.

[0051] PMI(X,Y)=log(P(X,Y) / P(X)P(Y))…Equation (1)

[0052] Next, the keyword extraction unit 132 calculates the SoA using equation (2). In equation (2), W is a proper noun contained in the Tweet, and L is a label (security threat information or other).

[0053] SoA(W,L)=PMI(W,L)-PMI(W,¬L)…Equation (2)

[0054] The keyword extraction unit 132 then extracts proper nouns whose SoA exceeds a predetermined threshold. For example, Tweets containing the Security Keyword "fraud" include Tweets related to phishing reports shown in Figure 5 (1) and Tweets unrelated to phishing reports shown in Figure 5 (2). The keyword extraction unit 132 extracts "d company" and "SMS" as Co-occurrence Keywords, as these are proper nouns that frequently appear (where SoA exceeds a predetermined threshold) only in Tweets related to phishing reports that contain "fraud" ((1)).

[0055] (2) Searching Tweets Next, the collection device 10 collects data from Twitter that is necessary for input to the classification device 20. For example, the second collection unit 133 uses the Co-occurrence Keywords extracted by the keyword extraction unit 132 to collect tweets from each user that may be reports of phishing attacks. As a result, the second collection unit 133 can collect tweets that include URLs and domains of Potentially Phishing Sites, for example, as shown in Figure 3.

[0056] In other words, the second collection unit 133 can collect Tweets (Screened Tweets) from each user's Tweets, excluding Tweets related to Legitimate Sites (Unrelated Tweets). The data collection unit 134 collects the following data regarding the Tweets collected by the second collection unit 133 (see Figure 6).

[0057] The tweet's text (e.g., hashtags, character count, etc.), associated metadata (e.g., application information, whether or not it's defanged, etc.), information about the account associated with the tweet (e.g., number of followers, account registration period, etc.), and images included in the tweet (e.g., up to 4 images associated with the tweet).

[0058] (3)Extracting URLs and Domain Names Next, the URL / domain name extraction unit 135 of the collection device 10 extracts URLs and domain names from the text and images of the Tweets (Screened Tweets) collected by the second collection unit 133.

[0059] For example, the URL / domain name extraction unit 135 applies optical character recognition to the image in the Tweet to extract the string. The URL / domain name extraction unit 135 also reverts any defangs (e.g., https -> ttps) in the Tweet string. Then, the URL / domain name extraction unit 135 extracts the URL and domain name from the text and image strings of the Tweet using regular expressions. After that, the URL / domain name extraction unit 135 checks whether the extracted domain name is possible using a Public Suffix List (see Reference 1), etc.

[0060] ·Reference 1: “Public Suffix List”, https: / / publicsuffix.org /

[0061] Then, the URL / domain name extraction unit 135, upon confirming the existence of the extracted domain name, extracts the domain name and the URL containing that domain name. For example, the URL / domain name extraction unit 135 extracts the following URL and domain name from the Tweet shown in Figure 7.

[0062] ·URL: https: / / tinyurl.com / yph6pswp, https: / / atavollwei.duckdns.org / • Domain names: tinyurl.com, atavolwei.duckdns.org

[0063] (4)Screening Phishing-related URLs and Domain Names Next, the selection unit 136 screens for URLs and domain names related to phishing from the URLs and domain names extracted by the URL / domain name extraction unit 135.

[0064] For example, if the extracted URL or domain name does not match the Allowlist (e.g., a list of legitimate website URLs or domain names) and is not a Long-lived Domain Name (e.g., a domain name that has been registered in WHOIS for more than a specified number of days), the selection unit 136 determines the extracted URL and domain name to be Potentially Phishing Sites. The selection unit 136 then selects Tweets containing URLs or domain names determined to be Potentially Phishing Sites as Tweets that are highly likely to be reports of phishing attacks.

[0065] On the other hand, if the extracted URL and domain name match the Allowlist, or if they are Long-lived Domain Names, the selection unit 136 designates the URL and domain name as Legitimate Sites.

[0066] For example, the selection unit 136 allows an extracted domain name to pass if it matches a predefined URL shortening service domain name. Furthermore, if an extracted domain name matches the Tranco List (see Reference 2), the selection unit 136 excludes it as a domain name not related to phishing attacks.

[0067] ·Reference 2: “A research-oriented top sites ranking hardened against manipulation - Tranco”, https: / / tranco-list.eu /

[0068] Furthermore, the selection unit 136 queries WHOIS for the extracted domain names, and if information cannot be obtained, it allows the domain names to pass through. In addition, based on the WHOIS information, the selection unit 136 excludes domain names that have been registered for more than 365 days, and allows them to pass through if they have not been registered for more than 365 days. Then, for example, the selection unit 136 selects Tweets that contain at least one type of URL or domain name that passed through the above process as Tweets that are highly likely to be reports of phishing attacks.

[0069] In this way, the collection device 10 can extract tweets from each user that are highly likely to be reports of phishing attacks.

[0070] [Classification device] [Example Configuration] Next, the classification device 20 will be described in detail. First, an example of the configuration of the classification device 20 will be described using Figure 8A. The classification device 20 includes, for example, an input / output unit 21, a storage unit 22, and a control unit 23.

[0071] The input / output unit 21 is an interface that handles the input and output of various types of data. For example, the input / output unit 21 accepts Tweets and their data that may be reports of phishing attacks collected by the collection device 10. The input / output unit 21 also outputs the classification results from the control unit 23.

[0072] The memory unit 22 stores data, programs, etc., that are referenced when the control unit 23 performs various processes. The memory unit 22 is implemented by semiconductor memory elements such as RAM and flash memory, or by storage devices such as hard disks and optical discs. For example, the memory unit 22 stores Tweets and their data (collected data) that are highly likely to be phishing attack reports received by the input / output unit 21. The memory unit 22 also stores the parameters of the classification model after the control unit 23 has trained the classification model.

[0073] The control unit 23 is responsible for controlling the entire classification device 20. The functions of the control unit 23 are realized, for example, by the CPU executing a program stored in the memory unit 22.

[0074] The control unit 23 includes, for example, a data acquisition unit 231, a feature extraction unit 232, a feature selection unit 233, a learning unit 234, a classification unit 235, and an output processing unit 236.

[0075] The data acquisition unit 231 acquires Tweets and their data that are highly likely to be reports of phishing attacks from the collection device 10.

[0076] The feature extraction unit 232 extracts features from the Tweets and data acquired by the data acquisition unit 231. For example, the feature extraction unit 232 extracts features from both the text and images of the Tweets acquired by the data acquisition unit 231.

[0077] For example, the feature extraction unit 232 extracts features from the Tweet acquired by the data acquisition unit 231, such as the features of the Tweet's account, the features of the Tweet's content, the features of the URL or domain name included in the Tweet, the features of the string obtained by optical character recognition of the image included in the post, the features of the image included in the Tweet, and the contextual features of the text included in the Tweet. Details of the feature extraction unit 232's extraction of Tweet features will be described later using a specific example.

[0078] The feature selection unit 233 selects features from the features extracted by the feature extraction unit 232 that are effective in classifying whether or not a tweet is related to a phishing attack report. For example, Boruta-SHAP (see references 3 and 4) can be used as a feature selection method.

[0079] ·Reference 3: Kursa, Miron B. and Rudnicki, Witold R., “Feature Selection with the Boruta Package,” Journal of Statistical Software 2010. ·Reference 4: “BorutaShap : A wrapper feature selection method which combines the Boruta feature selection algorithm with Shapley values,” https: / / zenodo.org / badge / latestdoi / 255354538

[0080] For example, the feature selection unit 233 selects features from the features extracted by the feature extraction unit 232 that are effective in classifying whether or not a tweet is related to a phishing attack, according to the following procedure.

[0081] (1) First, the feature selection unit 233 generates spurious features that include random values ​​in addition to the features to be selected. (2) Next, the feature selection unit 233 classifies the selected features and fake features using a decision tree-based algorithm and calculates the variable importance of each feature. (3) Next, the feature selection unit 233 counts any feature whose variable importance is greater than the variable importance of the false feature calculated in (2). (4) The feature selection unit 233 repeats the processes in (1) to (3) multiple times and selects the features that it deems statistically significant as features that are effective for classification.

[0082] The learning unit 234 trains a machine learning model (classification model) to classify whether an input Tweet is a phishing attack report or not, using supervised learning with the features selected by the feature selection unit 233. For example, the learning unit 234 trains the classification model using supervised learning with the features selected by the feature selection unit 233 on training data related to phishing attacks (data to which each Tweet is assigned a correct label indicating whether it is a phishing attack or not).

[0083] The classification unit 235 uses the classification model learned by the learning unit 234 to classify whether the input Tweet is a report of a phishing attack or not. The output processing unit 236 outputs the result of the classification of the Tweet by the classification unit 235.

[0084] [Example of processing procedure] Next, using Figure 8B, an example of the processing procedure performed by the classification device 20 will be explained. First, the data acquisition unit 231 of the classification device 20 acquires Tweets and their data that are highly likely to be reports of phishing attacks collected by the collection device 10 (S11: Acquisition of collected data). Subsequently, the feature extraction unit 232 extracts features from the Tweets and their data acquired by the data acquisition unit 231 (S12: Extraction of Tweet features).

[0085] After S12, the feature selection unit 233 selects features from the features extracted in S12 that are effective in classifying whether or not a tweet is related to a phishing attack report (S13). Then, the learning unit 234 uses the features selected in S13 to train a classification model on training data related to phishing attacks to classify whether or not an input tweet is a phishing attack report (S14).

[0086] After S14, the classification unit 235 uses the classification model learned in S14 to classify whether the input Tweet is a phishing attack report or not (S15). Then, the output processing unit 236 outputs the classification result from S16 (S16).

[0087] [Specific example of processing procedure] Next, using Figure 9, we will explain a specific example of the processing procedure performed by the classification device 20.

[0088] (5) Feature Engineering First, the data acquisition unit 231 of the classification device 20 acquires the Tweets (Screened Tweets) and their data collected by the collection device 10. Then, the feature extraction unit 232 extracts features from the Tweets and their data acquired by the data acquisition unit 231.

[0089] For example, as shown in Figure 10, the feature extraction unit 232 generates a total of 27 features across six types: Account Feature (1) from the Tweet account, Content Feature (2) from the information associated with the Tweet, URL Feature (3) from the extracted URL, OCR Feature (5) from the string extracted by OCR, Visual Feature (6) from the appearance of the image, and Context Feature (4) from the context of the Tweet. Each feature will be explained in detail below.

[0090] (5-1) Account Feature The feature extraction unit 232 generates an Account Feature for each Tweet from the user's account information (e.g., number of followers, number of following, number of tweets, number of media, number of lists, account registration date, etc.) to capture the characteristics of Twitter users, for example as shown in Figure 11.

[0091] (5-2) Content Feature The feature extraction unit 232 generates a Content Feature for each Tweet from information associated with the Tweet itself (e.g., string, mentioned users, hashtags, images, URLs or domain names, applications used for tweeting, defang type, etc.) in order to capture the characteristics of content that frequently appears in Tweets reporting phishing attacks, as shown in Figure 12.

[0092] (5-3) URL Feature The feature extraction unit 232 generates a URL feature for each tweet from URLs (or domain names) extracted from both the Tweet string and image, as shown in Figure 13, in order to capture features related to the misuse of subdomains specific to phishing URLs and the misuse of specific top-level domains. URL features include, for example, the URL string, domain name, path, numbers included in the URL, top-level domain, etc.

[0093] (5-4) OCR Feature The feature extraction unit 232 generates OCR features for each tweet from strings extracted by optical character recognition (OCR), for example, as shown in Figure 14, in order to capture the characteristics of similar strings in tweets related to phishing attacks. OCR features can be strings, words, symbols, numbers, URLs, or domain names, etc.

[0094] (5-5) Visual Feature The feature extraction unit 232 generates a Visual Feature for each Tweet from the images associated with the Tweet in order to capture the visual commonalities of the images included in Tweets reporting phishing attacks.

[0095] The feature extraction unit 232 uses the EfficientNet model (see Reference 5), which has shown excellent results in image classification, to generate a fixed-dimensional vector of images associated with the Tweet. Subsequently, the feature extraction unit 232 compresses the dimensionality of the vector using Truncated SV (see Reference 6) to convert sparse vectors into dense vectors. Finally, the feature extraction unit 232 uses the compressed vector as the Visual Features of the images contained in the Tweet.

[0096] ·Reference 5: Tan, Mingxing and Le, Quoc., “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks”, ICML 2019. ·Reference 6: “The truncatedsvd as a method for regularization”, BIT Numerical Mathematics.

[0097] The feature extraction unit 232, for example as shown in Figure 15, uses an EfficientNet model pre-trained on a large number of images from ImageNet to convert images associated with Tweets into vectors of eigendimensionality. Then, the feature extraction unit 232 compresses the converted vectors using Truncated SV to a cumulative contribution rate of 99% in the training data.

[0098] (5-6) Context Feature The feature extraction unit 232 generates a Context Feature for each tweet from the strings within the tweets in order to capture the common context in tweets reporting phishing attacks.

[0099] The feature extraction unit 232 generates a fixed-dimensional vector from the strings in the Tweet, for example, using the BERT model, which has shown excellent results in text classification. Then, the feature extraction unit 232 compresses the dimensionality of the vector using Truncated SV. Finally, the feature extraction unit 232 uses the compressed vector as the Context Feature of the Tweet.

[0100] The feature extraction unit 232 converts the strings in the Tweet into eigendimensional vectors using a BERT model pre-trained on a large amount of strings from English and Japanese Wikipedia, for example, as shown in Figure 16. Then, the feature extraction unit 232 compresses the converted vectors to a cumulative contribution rate of 99% in the training data using Truncated SV.

[0101] (6) Feature Selection The feature selection unit 233 selects (important) features from the set of features generated by the feature extraction unit 232 in (5) that are effective in classifying phishing attack reports from other tweets.

[0102] Figure 17 shows examples of features that were determined to be important for classification based on the Feature Selection results.

[0103] Account Features: 6 English versions (6 dimensions), 5 Japanese versions (5 dimensions) Content Feature: 6 English versions (9 dimensions), 4 Japanese versions (7 dimensions) URL Feature: Two types in English (2D), three types in Japanese (3D) OCR Features: 3 types of English (3D), 3 types of Japanese (3D) Visual Feature: English 9 dimensions, Japanese 5 dimensions Context Feature: English 58 dimensions, Japanese 33 dimensions

[0104] Of the Context Features shown in Figure 17, for App source (14), Twitter Web App, Twitter for iPhone®, and Twitter for Android® are important in both languages, while PhishingPicker is important only in English. Furthermore, for Defined type (15), example[.]com is important in both languages, while hxxp is important only in Japanese. Additionally, for Top-level domain (20) shown in Figure 17, .xyz is important only in Japanese.

[0105] Ultimately, we confirmed that 87-dimensional features for English and 56-dimensional features for Japanese are important for classifying phishing attack reports from other tweets.

[0106] (7) Offline Training The learning unit 234 learns a classification model (Machine Learning Model) using the features (feature vectors) selected by the feature selection unit 233 in (6) and the training data (Ground-Truth Dataset) to which the correct labels indicating whether or not it is a phishing attack are assigned.

[0107] The algorithms used to train the classification model include, for example, Random Forest, Neural Network, Decision Tree, Support Vector Machine, Logistic Regression, Naive Bayes, Gradient Boosting, and Stochastic Gradient Descent. After evaluating these algorithms on training data, it was confirmed that Random Forest is preferable for the following three reasons.

[0108] Random Forest demonstrated superior classification accuracy compared to all other algorithms. Random Forest performed at a stable speed during both the learning and estimation (classification) phases. • In Random Forest, the feature importance was distributed across all six types of features.

[0109] (8) Online Classification The classification unit 235 uses the Machine Learning Model (classification model) learned in (7) to classify whether the Tweets collected by the collection device 10 are positive or negative reports of phishing attacks. The output processing unit 236 then outputs the result of this classification.

[0110] Furthermore, the classification device 20 may extract proper nouns that appear in the phishing attack reports and classified Tweets, and the collection device 10 may use these proper nouns when extracting Co-occurrence Keywords.

[0111] [Evaluation Results] Next, we will explain the evaluation results of the system of this embodiment. For example, by using the features selected by the system, it was confirmed that it can classify whether a tweet is a phishing attack report or not with an accuracy of approximately 95% for both English and Japanese (see Figure 18).

[0112] Furthermore, during the experimental period (August 1, 2021 to September 30, 2021), the system of this embodiment was able to extract 77,004 phishing attack reports (User Reports) and 85,027 phishing URLs, as shown in Figure 19.

[0113] Furthermore, when comparing phishing URLs collected by the existing data feed OpenPhish (see Reference 7) with those collected by the system of this embodiment (see Figure 20), it was found that of the 4,802 phishing URLs common to both, 2,686 (55.9% of the total) were collected faster by the system of this embodiment.

[0114] ·Reference 7: “OpenPhish - Phishing Intelligence”, https: / / openphish.com

[0115] Furthermore, when comparing phishing URLs collected by the existing data feed PhishTank (see Reference 8) with those collected by the system of this embodiment (see Figure 21), the system of this embodiment was able to collect 3,183 of the 5,323 phishing URLs common to both (59.8% of the total) faster.

[0116] ·Reference 8: “PhishTank | Join the fight against phishing”, https: / / www.phishtank.com / .

[0117] Furthermore, an investigation into the number of phishing attack reports by users and the number of phishing URLs revealed that 49.8% of all phishing URLs were reported only once by users (see Figure 22). In other words, it was confirmed that phishing attack reports from a wide range of users are likely to contain highly unique phishing URLs. From this, it was confirmed that collecting phishing attack reports from a wide range of users, as in the system of this embodiment, is extremely effective.

[0118] Furthermore, we confirmed the effectiveness of using not only fixed keywords (Security Keywords) but also dynamic keywords (Co-occurrence Keywords) in collecting tweets reporting phishing attacks (see Figure 23). As a result, we found that using dynamic keywords (Co-occurrence Keywords) in addition to fixed keywords (Security Keywords) allowed us to extract 23.3% more User Reports (Tweets reporting phishing attacks) than using only fixed keywords (Security Keywords). We also found that using dynamic keywords (Co-occurrence Keywords) in addition to fixed keywords allowed us to extract 24.1% more phishing URLs.

[0119] From this, it was confirmed that collecting Tweets using not only fixed keywords (Security Keywords) but also dynamic keywords (Co-occurrence Keywords), as in the system of this embodiment, is extremely effective for gathering information on phishing attacks.

[0120] [System configuration, etc.] Furthermore, the components of each part shown in the diagram are functional concepts and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown in the diagram, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions. Moreover, all or any part of the processing functions performed by each device can be realized by a CPU and the program executed on that CPU, or by hardware using wired logic.

[0121] Furthermore, among the processes described in the embodiments described above, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above document and drawings can be arbitrarily changed unless otherwise specified.

[0122] [program] The aforementioned system can be implemented by installing the program as packaged software or online software on a desired computer. For example, by having the above program run on an information processing device, the information processing device can be made to function as the aforementioned system. The information processing device referred to here includes mobile communication terminals such as smartphones, mobile phones and PHS (Personal Handyphone System), as well as terminals such as PDA (Personal Digital Assistant).

[0123] Figure 24 shows an example of a computer running a program. Computer 1000 has, for example, memory 1010 and a CPU 1020. Computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0124] Memory 1010 includes ROM (Read Only Memory) 1011 and RAM (Random Access Memory) 1012. ROM 1011 stores, for example, a boot program such as BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. For example, a removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.

[0125] The hard disk drive 1090 stores, for example, the OS 1091, application programs 1092, program modules 1093, and program data 1094. That is, the programs that define each process executed by the system are implemented as program modules 1093, in which executable code is written. The program modules 1093 are stored, for example, on the hard disk drive 1090. For example, a program module 1093 for performing processes similar to those in the system's functional configuration is stored on the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0126] Furthermore, the data used in the processing of the above-described embodiment is stored as program data 1094 in, for example, memory 1010 or hard disk drive 1090. The CPU 1020 then reads the program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as needed and executes them.

[0127] Furthermore, the program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090; for example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via a network interface 1070. [Explanation of symbols]

[0128] 10 Collection device 11,21 Input / output section 12,22 Storage section 13,23 Control Unit 20 Classifier 131 First Collection Section 132 Keyword Extraction Unit 133 Second Collection Section 134 Data Collection Department 135 URL / Domain Name Extraction Section 136 Sorting Department 231 Data Acquisition Unit 232 Feature Extraction Unit 233 Feature Selection Unit 234 Learning Department 235 Classification Department 236 Output Processing Unit

Claims

1. The first collection unit collects posts related to security threats from SNS (Social Networking Service) posts using security keywords, which are keywords related to security threats. A keyword extraction unit extracts co-occurring keywords, which are keywords that co-occur more frequently than a predetermined frequency, from the collected posts related to the security threats. A second collection unit collects posts containing the aforementioned co-occurring keywords and images associated with those posts from SNS posts, A selection unit selects and outputs posts that are likely to be related to security threats, based on the URL or domain name extracted from the text and images of the posts collected by the second collection unit. Equipped with, The sorting unit is, A data collection device characterized in that, if the usage period of a domain name extracted from the text and images of the post collected by the second collection unit is less than a predetermined period, the post is selected as a post that may be related to a security threat.

2. The first collection unit is, Collect the aforementioned posts at predetermined intervals, The keyword extraction unit, The co-occurring keywords are extracted from posts collected during the predetermined period. The collection device according to feature 1.

3. A collection method performed by a collection device, The process involves collecting posts related to security threats from SNS (Social Networking Service) posts using security keywords, which are keywords related to security threats. A step of extracting co-occurring keywords, which are keywords that co-occur more frequently than a predetermined frequency, from the collected posts related to the security threats, The process involves collecting the text of posts containing the aforementioned co-occurring keywords and images associated with those posts from social media posts, The process involves selecting and outputting posts that are likely to be related to security threats, based on URLs or domain names extracted from the text and images of the collected posts. Includes, The output process described above is: A collection method characterized by selecting a post as potentially related to a security threat if the usage period of a domain name extracted from the text and images of the collected post is less than a predetermined period.

4. The process involves collecting posts related to security threats from SNS (Social Networking Service) posts using security keywords, which are keywords related to security threats. A step of extracting co-occurring keywords, which are keywords that co-occur more frequently than a predetermined frequency, from the collected posts related to the security threats, The process involves collecting posts containing the aforementioned co-occurring keywords and images associated with those posts from social media posts, The process involves selecting and outputting posts that are likely to be related to security threats, based on URLs or domain names extracted from the text and images of the collected posts. Have the computer run it, The output process described above is: A collection program that identifies posts as potentially related to security threats if the domain name extracted from the text and images of the collected posts is less than a predetermined period of use.

Citation Information

Patent Citations

  • Method for detecting expression capable of becoming dangerous expression by relying on specific theme and electronic device and program for electronic device for detecting the same expression

    JP2015072614A

  • Search device, control method thereof, and program, as well as search system, control method thereof, and program

    JP2019066979A

  • Automated Extraction and Classification of Malicious Indicators

    JP2024512266A

  • Automated extraction and classification of malicious indicators

    WO2022182568A1