Estimation device, estimation method, and estimation program
The estimation device uses a large language model to categorize OSS based on its name, addressing the challenge of expert-dependent category estimation and enhancing efficiency and accuracy in software analysis.
Patent Information
- Application Number
- PCT/JP2023/044711
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
Existing technologies struggle to easily estimate the category of Open Source Software (OSS) from its name, relying on expert analysis which is costly and subjective.
An estimation device that acquires OSS information using its name, inputs this information into a large language model, and estimates the category based on the output, allowing for easy categorization without expert analysis.
Enables the easy estimation of OSS categories from their names, reducing the need for expert analysis and improving efficiency and accuracy in software configuration analysis.
Smart Images

Figure JP2023044711_19062025_PF_FP_ABST
Abstract
Description
Estimation device, estimation method, and estimation program
[0001] The present invention relates to an estimation device, an estimation method, and an estimation program.
[0002] In recent software development, third-party libraries and source code from OSS (Open Source Software) are often reused to improve development efficiency. However, since reused source code may contain vulnerabilities, it is important to understand the software configuration and information about the source code used to mitigate security risks.
[0003] Known technologies for understanding software configuration include dependency detection technology that identifies the OSS from which software is reused (see, for example, Non-Patent Document 1) and vulnerability detection technology that detects vulnerabilities caused by the OSS from which software is reused (see, for example, Non-Patent Document 2). These technologies detect OSS based on lexical string similarity in source code, and when a target software is input, a list of names of OSS reused in the software is output as a result.
[0004] Yang Kanemoto, Tomoki Yamanaka, Eitaro Shioji, Kazufumi Aoki, and Mitsuaki Akiyama, "The Actual State of Implicit Relationships in Software as Seen from Code Clone Investigation," SCIS 2023 2023 Symposium on Cryptography and Information Security, Fukuoka, Japan, January 24-27, 2023. Reika Arakawa, Yang Kanemoto, Eitaro Shioji, and Mitsuaki Akiyama, "A Method for Detecting Vulnerable Code Clone Inherent in Software Using a Set Similarity Calculation Algorithm," IEICE Technical Journal, vol. 123, no. 86, ICSS2023‐7, pp. 32-39, June 2023.
[0005] However, with the above-mentioned conventional technology, it is not possible to easily estimate the OSS category from the OSS name. For example, with the above-mentioned conventional technology, it is possible to detect a list of reused OSSs present in the software being inspected, but the output content is limited to the OSS name, so the analysis content, such as the OSS category, must be understood through analysis by an expert.
[0006] In order to solve the above-mentioned problems and achieve the object, the estimation device of the present invention is characterized by having an acquisition unit that acquires information about an OSS using the name of the OSS, an estimation unit that inputs the information acquired by the acquisition unit into a large-scale language model and estimates the category of the OSS based on the output result of the large-scale language model, and an output unit that outputs the category estimated by the estimation unit.
[0007] According to the present invention, it is possible to easily estimate the category of an OSS from the name of the OSS.
[0008] FIG. 1 is a diagram showing an estimation system according to an embodiment. FIG. 2 is a diagram showing an example of the configuration of an estimation device according to an embodiment. FIG. 3 is a diagram showing an example of data stored in the estimation device according to an embodiment. FIG. 4 is a diagram showing an example of data stored in the estimation device according to an embodiment. FIG. 5 is a diagram showing an example of data stored in the estimation device according to an embodiment. FIG. 6 is a diagram showing an example of data stored in the estimation device according to an embodiment. FIG. 7 is a diagram showing a specific example of processing by the estimation device according to an embodiment. FIG. 8 is a diagram showing a specific example of processing by the estimation device according to an embodiment. FIG. 9 is a diagram showing a specific example of processing by the estimation device according to an embodiment. FIG. 10 is a flowchart showing an example of the flow of estimation processing according to an embodiment. FIG. 11 is a diagram showing an example of a computer that executes an estimation program.
[0009] Hereinafter, embodiments of an estimation device, an estimation method, and an estimation program according to the present application will be described in detail with reference to the accompanying drawings. Note that the estimation device, the estimation method, and the estimation program according to the present application are not limited to these embodiments.
[0010] [1. Introduction] (1-1. Overview of Prior Art) First, an overview of the prior art in the OSS analysis technology according to this embodiment will be described. In recent software development, it has become common to reuse OSS to improve development efficiency, and research reports have shown that reused OSS source code accounts for approximately 90% of the software developed by companies. Reused OSS may contain code that leaves vulnerabilities, or may contain OSS that is reused in a nested manner (such as a reused OSS reusing another OSS) that is unknown to the developer.
[0011] Therefore, if an older version of OSS or an OSS with reported vulnerabilities is reused, it may pose a security risk, so it is important to identify and analyze the existence of reused OSS itself in order to reduce security risks.
[0012] In order to understand the OSS of a reused software, there are known prior art technologies for identifying the OSS of a reused software and for detecting vulnerabilities caused by the OSS of a reused software. These technologies output a list of OSSs reused in software based on the lexical similarity of the source code, but the output content is limited to the OSS name. Therefore, analysis of the functionality of the reused OSSs is left to experts.
[0013] Additionally, there are currently over 500,000 OSSs in the world, with over 80 categories. However, only a few hundred representative OSSs are categorized. Given this current situation, manual analysis is costly and subjective.
[0014] Therefore, analysis of the functions of the reused OSS is left to experts, and there are problems with the analysis by these experts, such as high work costs and the possibility of subjectivity being introduced into the analysis.
[0015] (1-2. Overview of the estimation device according to this embodiment) The estimation device according to this embodiment that performs the process of estimating the OSS category was invented for the purpose of solving the above-mentioned problems, and has the effect of being able to easily estimate the OSS category from the OSS name.
[0016] Next, an estimation device according to the present embodiment will be described. Fig. 1 is a diagram showing an estimation system according to the present embodiment. The estimation device 100 is, for example, a server device that receives the name of an OSS to be estimated from an external information processing device and acquires information about the OSS to be estimated via a network, and is realized by a PC (Personal Computer), a cloud system, or the like.
[0017] The external device 200 is an information processing device that transmits and receives information to and from the estimation device 100 via a network, and is realized by a PC or the like. The external device 200 is, for example, a terminal possessed by a client that transmits the name of the OSS to be estimated to the estimation device 100 and receives the name and category of the OSS output by the estimation device 100.
[0018] The estimation device 100 acquires information about the OSS using the name of the OSS, inputs the acquired information into a large-scale language model, estimates the category of the OSS based on the output result of the large-scale language model, and outputs the estimated category.
[0019] For example, the estimation device 100 creates a search query for the OSS to be estimated based on the name of the OSS received from the external device 200, and obtains the top few search results when searching for the search query on a search site.
[0020] The estimation device 100 then inputs the acquired search result data and preset category information into a large-scale language model, estimates the category output by the large-scale language model as the category of the OSS, and outputs the estimated category together with the name of the OSS to the external device 200.
[0021] As a result, the estimation device 100 can estimate the OSS category from the OSS name, making it possible to easily grasp the OSS category without requiring analysis by an expert.
[0022] Here, the estimation process of the estimation device 100 described above has been described as a process of estimating an OSS category using a large-scale language model. However, utilizing a large-scale language model rather than an existing machine learning model has the following advantages.
[0023] First, large-scale language models perform processing using information other than the data entered as the question, preventing the data used for category estimation from depending on the input data. Second, even when expanding the types of categories to be estimated, no additional training is required. Furthermore, because large-scale language models have a high affinity with natural language processing, search results data from search sites can be input without the need for preprocessing.
[0024] 2. Configuration of the Estimation Device 100 Next, the configuration of the estimation device 100 shown in Fig. 1 will be described with reference to Fig. 2. Fig. 2 is a block diagram showing an example configuration of the estimation device according to an embodiment. The estimation device 100 has a communication unit 110, a control unit 120, and a storage unit 130, and is connected to an external device 200 via a network N so that the devices can communicate with each other.
[0025] The communication unit 110 is realized by, for example, a network interface card (NIC). The communication unit 110 is connected to a network N and transmits and receives information to and from the external device 200 and an external search application programming interface (API). The communication unit 110 mediates, for example, the acquisition of information related to the OSS from the external search API and the output of an estimated category to the external device 200.
[0026] The storage unit 130 is realized by a storage device such as a RAM (Random Access Memory) or a hard disk, for example. The storage unit 130 stores data and programs necessary for various processes by the control unit 120. The storage unit 130 includes a category information storage unit 131, an OSS information storage unit 132, a search result information storage unit 133, and an estimation result information storage unit 134, which are closely related to the present invention.
[0027] The category information storage unit 131 stores a category list indicating the relationship between an OSS and its category. Here, the information stored in the category information storage unit 131 will be described with reference to FIG. 3 . FIG. 3 is a diagram showing an example of data stored in the estimation device according to the embodiment. The category list may be a category list published by an external organization such as "Openstandia NRI."
[0028] 3, the category information storage unit 131 stores, for example, a "category number" and a "category name." The "category number" stores a number assigned to identify each category, and the "category name" stores the name of the classified category. For example, the category information storage unit 131 stores information such as "category name: OS" for "category number: 1."
[0029] The OSS information storage unit 132 stores information about the name of the OSS to be estimated. Here, the information stored in the OSS information storage unit 132 will be described with reference to Fig. 4. Fig. 4 is a diagram illustrating an example of data stored in the estimation device according to the embodiment.
[0030] 4, the OSS information storage unit 132 stores, for example, an "index number" and an "OSS name." The "index number" stores a number assigned to identify each OSS, and the "OSS name" stores the name of the OSS to be estimated. For example, the OSS information storage unit 132 stores information such as "OSS name: Keras" for "index number: 1."
[0031] The search result information storage unit 133 stores information about the OSS to be estimated, acquired by the acquisition unit 121 (described later). Here, the information stored in the search result information storage unit 133 will be described with reference to Fig. 5. Fig. 5 is a diagram illustrating an example of data stored in the estimation device according to the embodiment.
[0032] As shown in FIG. 5 , the search result information storage unit 133 stores, for example, an "index number," a "rank," a "title," a "url," and a "snippet." The "rank" stores the ranking order of the site in the search results, and the "title" stores the title of the site listed in each rank. The "url" stores the uniform resource locator (URL) of the site listed in each rank, and the "snippet" stores a summary description of the site listed in each rank. Here, the search result information storage unit 133 may store each of the above-mentioned information in, for example, a CSV (Comma Separated Values) file format or in a database format.
[0033] The estimation result information storage unit 134 stores information about the estimation results of categories output by the large-scale language model. Here, the information stored in the estimation result information storage unit 134 will be described with reference to Fig. 6. Fig. 6 is a diagram illustrating an example of data stored in the estimation device according to the embodiment.
[0034] 6, the estimation result information storage unit 134 stores, for example, an "index number," an "estimated category name," an "OSS name," and a "snippet." The "estimated category name" stores the name of the estimated category of the OSS output by the large-scale language model. The "snippet" stores information linking snippets of each rank in the OSS search results.
[0035] Returning to the explanation of Fig. 2, the control unit 120 is realized by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or the like executing various programs stored in a storage device within the device using RAM as a work area. The control unit 120 is also realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 120 has an acquisition unit 121, an estimation unit 122, an output unit 123, and a reconstruction unit 124.
[0036] The acquisition unit 121 acquires information about the OSS by using the name of the OSS. For example, the acquisition unit 121 creates a predetermined search query from the name of the OSS and acquires a predetermined number of top search results when the search query is entered into a search site.
[0037] For example, the acquisition unit 121 creates a search query character string such as "What is Keras (an example of an OSS name)" from the name of the OSS to be estimated stored in the OSS information storage unit 132. After that, the acquisition unit 121 stores, in CSV file format, information on "rank," "title," "URL," and "snippet" for the top 10 search results obtained by searching a search site for the search query, for example, in the search result information storage unit 133.
[0038] Here, the number of search results to be stored in the above-described process is, for example, an arbitrary number that is set in advance in consideration of the reliability of the search results. In addition to the above-described information, the acquisition unit 121 can also acquire information about the source code of the OSS as information about the OSS.
[0039] In addition to the name of the OSS itself, the name of the OSS mentioned above can include information such as the project name searched for on platforms such as Github and Gitlab regarding the OSS to be estimated.
[0040] The estimation unit 122 inputs the information acquired by the acquisition unit 121 into the large-scale language model, and estimates the category of the OSS based on the output result of the large-scale language model. For example, the estimation unit 122 refers to the information stored in the category information storage unit 131, the OSS information storage unit 132, and the information stored in the search result information storage unit 133, connects snippets of sites searched based on the name of the OSS to be estimated, and inputs the connected snippets and a category list into the large-scale language model to perform an estimation process of the OSS category.
[0041] The estimation unit 122 inputs, for example, a question such as "You are an excellent software engineer. Please estimate the main functions of the following open source software as categories," along with the snippet linked to the name of the OSS to be estimated and the category list into the large-scale language model. After that, the estimation unit 122 stores, for example, information output from the large-scale language model in the estimation result information storage unit 134.
[0042] The estimation unit 122 also calculates the certainty of the estimated category. For example, the estimation unit 122 calculates the certainty of the estimated category by inputting a question sentence that outputs the certainty (accuracy) of the estimated category along with the OSS category estimation to the large-scale language model. Here, the estimation unit 122 stores information such as "encryption (certainty 0.7)," "authentication (certainty 0.2)," and "communication (certainty 0.1)" in the estimation result information storage unit 134, for example, by one estimation process.
[0043] Furthermore, the estimation unit 122 inputs the category list reconstructed by the reconstruction unit 124 into the large-scale language model to estimate a category. For example, the estimation unit 122 inputs a category list stored in the category information storage unit 131 and reconstructed by the reconstruction unit 124 (described later) into the large-scale language model, thereby estimating the category of the OSS under conditions where the reconstructed category list is specified as an estimation candidate.
[0044] The output unit 123 outputs the category estimated by the estimation unit 122. For example, the output unit 123 references the information stored in the estimation result information storage unit 134 and outputs the category estimated for the main function of the OSS together with the name of the OSS to be estimated to the external device 200, etc. Here, the output unit 123 may output, for example, one category for one OSS, or when multiple categories are estimated for one OSS, may output all of the estimated categories.
[0045] Furthermore, the output unit 123 outputs the category that has been estimated most frequently for the OSS based on the estimation results obtained by the estimation unit multiple times. For example, the estimation unit 122 outputs the category that has been output most frequently as the category of the OSS from among the multiple categories output by the estimation process multiple times.
[0046] For example, if the first estimated result for a certain OSS is "encryption," the second estimated result is "authentication," and the third estimated result is "encryption," the output unit 123 outputs "encryption," which has been output the most times, as the category of the OSS.
[0047] Furthermore, the output unit 123 outputs the certainty factor calculated by the estimation unit 122 together with the category. For example, the output unit 123 outputs the certainty factor of the estimated category stored in the estimation result information storage unit 134 as information that reinforces the estimation information.
[0048] The reconstructing unit 124 performs a predetermined process on the category for which the number of erroneous predictions is equal to or greater than a predetermined threshold, and reconstructs the category list. For example, the reconstructing unit 124 uses external information in which OSS names and their corresponding categories are shown in a one-to-one correspondence as correct answer data, and determines that the category predicted by the above-mentioned estimation unit 122 is an incorrect prediction when it differs from the correct answer data.
[0049] Then, for example, for categories described in the category list stored in the category information storage unit 131, the reconstructing unit 124 extracts categories with an erroneous estimation rate of 80% or more from estimation results for each OSS above a certain level, and performs processing such as converting category words, adding new categories, deleting categories, etc. After that, the reconstructing unit 124 stores, for example, the category list in which the processed categories have been re-registered as an updated category list in the category information storage unit 131.
[0050] Here, the processing performed by the reconfiguration unit 124 on categories will be described below. For example, if the reconfiguration unit 124 recognizes an inclusion relationship in main functions between categories such as "storage" and "distributed system," it converts the category into a category word that does not create an inclusion relationship. Also, if the main function of a category itself is broad, such as "framework" or "infrastructure construction," it breaks down the main function and adds a category. Furthermore, if there are categories with overlapping main functions, it deletes one of the categories.
[0051] 7 to 9, a specific example of the estimation process performed by the estimation device 100 will be described. FIGS. 7 to 9 are diagrams showing a specific example of the process performed by the estimation device according to the embodiment. First, the overall flow of the process performed by the estimation device 100 will be described with reference to FIG. 7.
[0052] First, for example, an OSS list for which a category is to be estimated is input to the estimation device 100 and stored in the OSS information storage unit 132. The OSS list lists OSS names as shown in FIG. 7 . Next, the acquisition unit 121 acquires the top 10 search results for each OSS listed in the OSS list from an external search API. Here, the search query created is "What is [OSS name] OSS" and is used for the search. Next, the acquisition unit 121 stores the acquired information in the search result information storage unit 133.
[0053] Here, referring to FIG. 8 , a specific example of information stored in the search result information storage unit 133 in this specific example will be described. In FIG. 8 , information on the first six of the top ten search results is shown, and similar information is stored from the sixth search result onwards up to the tenth search result. In FIG. 8 , the first line indicates that the "rank, title, URL, snippet" of the search site is stored in this order, and the second line and onwards indicate that specific information of the search results is stored. For example, the second line stores the information "1, Pricing, Table Elements-iDempiere ADempiere OSS ERP SES, https: / / www.as-link.com / en / portfolio-items / pricing-table-elements / , 'Mar 11, 2017...Carefully crafted...'."
[0054] 7 , the description will be continued. The estimation unit 122 then implements "gpt-4 Function Calling" and defines "estimate_function_categorize" to obtain the category estimation results of the OSS using the large-scale language model, and stores the results in the estimation result information storage unit 134. Note that the estimated category candidates refer to categories listed in a pre-stored category list.
[0055] Here, referring to FIG. 9 , a specific example of information stored in the inference result information storage unit 134 in this specific example will be described. FIG. 9 shows that the inference result information storage unit 134 stores "index number, inferred category, OSS name, linked snippet" in this order. For example, for the OSS with "index number: 26" written on the first line in FIG. 9 , the following information is stored: "26, language, Perl, 'Perl is a highly capable, feature-rich programming language with...'"
[0056] Thereafter, the output unit 123 outputs a list of OSSs and estimated categories, which are category results for each input OSS estimated by the estimation unit 122. For example, for "OSS_name: Keras", the output unit 123 outputs a list listing information such as "estimated_category: machine learning AI" for each OSS.
[0057] Through the above-described series of processes, the estimating device 100 can easily estimate the category of an OSS from the name of the input OSS. Regarding the accuracy of the estimation process by the estimating device 100, when the category estimation process was performed on a total of 227 OSSs using the category list of "Openstandia NRI" as the correct answer data, the number of matching OSSs was 172 and the number of mismatching OSSs was 55, resulting in a category precision rate of 0.758.
[0058] 4. Example of Estimation Process Next, the flow of the estimation process performed by the estimation device 100 will be described with reference to Fig. 10. Fig. 10 is a flowchart showing an example of the flow of the estimation process according to the embodiment. First, the estimation device 100 receives the name of the OSS to be estimated from the outside (S101).
[0059] When the name of the OSS to be estimated is received from the outside (S101; Yes), the acquisition unit 121 creates a search query from the OSS name (S102). On the other hand, when the name of the OSS to be estimated is not received from the outside (S101; No), the acquisition unit 121 waits until the name of the OSS to be estimated is received from the outside. After the process of S102, the acquisition unit 121 then acquires the top few search results when the search query is input to a search site (S103).
[0060] Next, the acquisition unit 121 stores the acquired search results in the search result information storage unit 133 (S104). After that, the estimation unit 122 inputs the stored search results into the large-scale language model (S105). Next, the estimation unit 122 estimates the category of the OSS based on the output result of the large-scale language model (S106). Then, the output unit 123 outputs the name of the OSS and the estimated category (S107), and the estimation device 100 ends the process.
[0061] 5. Effects of the Embodiment As described above, the estimation device 100 according to the present embodiment includes an acquisition unit 121, an estimation unit 122, and an output unit 123. The acquisition unit 121 acquires information about an OSS using the name of the OSS. The estimation unit 122 inputs the information acquired by the acquisition unit 121 into a large-scale language model and estimates the category of the OSS based on the output result of the large-scale language model. The output unit 123 outputs the category estimated by the estimation unit 122.
[0062] This allows the estimation device 100 to input OSS information searched from the OSS name into a large-scale language model and estimate the category from the output result, making it possible to easily estimate the OSS category from the OSS name. As a result, even non-experts can grasp the main function indicated by the OSS category from the OSS name, thereby enabling quick software configuration analysis.
[0063] The acquisition unit 121 also creates a predetermined search query from the name of the OSS and acquires a predetermined number of top search results when the search query is entered into a search site. This allows the estimation device 100 to easily acquire information used for estimation from the name of the OSS using an external search API.
[0064] Furthermore, the output unit 123 outputs the category that has been estimated most frequently for the OSS based on the estimation results obtained multiple times by the estimation unit 122. As a result, the estimation device 100 determines the estimated category based on the results of multiple estimation processes for one OSS, and can therefore output highly accurate estimation results.
[0065] Furthermore, the estimation unit 122 calculates the certainty of the estimated category, and the output unit 123 outputs the certainty calculated by the estimation unit 122 together with the category. In this way, the estimation device 100 outputs the certainty indicating the accuracy of the estimation together with the category, allowing the accuracy of the category estimation to be grasped as supplementary information.
[0066] The estimation device 100 also includes a reconstructing unit 124. The reconstructing unit 124 performs a predetermined process on categories for which the number of erroneous estimations is equal to or greater than a predetermined threshold, and reconstructs a category list indicating categories that are candidates for estimation. In this case, the estimation unit 122 inputs the category list reconstructed by the reconstructing unit 124 into a large-scale language model and estimates a category.
[0067] As a result, the estimation device 100 reconstructs categories with clearer main functions for categories in which a certain number of incorrect estimations have occurred in the estimation process when an external category list is used as correct answer data, and performs the estimation process using the reconstructed categories, thereby enabling estimation of categories with clearer main functions.
[0068] [6. System Configuration, etc.] Of the processes described in the above embodiments, some of the processes described as being performed automatically can also be performed manually. Alternatively, all or some of the processes described as being performed manually can be performed automatically using known methods. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.
[0069] Furthermore, the components of each device shown in the figure are conceptual functional units and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of the devices can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions performed by each device can be realized by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic.
[0070] 2 may be stored in a storage server or the like, rather than being stored in the estimation device 100. In this case, the estimation device 100 acquires various pieces of information by accessing the storage server.
[0071] 7. Hardware Configuration Fig. 11 is a diagram showing an example of a hardware configuration. The estimation device 100 according to the embodiment described above is realized by a computer 1000 having the configuration shown in Fig. 11, for example.
[0072] 11 is a diagram showing an example of a computer that executes an estimation program. The computer 1000 includes, for example, a memory 1010 and a CPU 1020. The computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0073] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1041. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1041. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.
[0074] The hard disk drive 1090 stores, for example, an operating system (OS) 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the estimation device 100 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing processes similar to those of the functional configuration of the estimation device 100 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced with an SSD (Solid State Drive).
[0075] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1090. Then, CPU 1020 reads out program module 1093 or program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as necessary and executes them.
[0076] The program module 1093 and program data 1094 may not necessarily be stored in the hard disk drive 1090, but may instead be stored in a removable storage medium and read by the CPU 1020 via the disk drive 1041 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (LAN, WAN, etc.). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.
[0077] REFERENCE SIGNS LIST 100 Estimation device 110 Communication unit 120 Control unit 121 Acquisition unit 122 Estimation unit 123 Output unit 124 Reconstruction unit 130 Storage unit 131 Category information storage unit 132 OSS information storage unit 133 Search result information storage unit 134 Estimation result information storage unit
Claims
1. An estimation device, comprising: an acquisition unit that acquires information about the OSS using the name of the OSS; an estimation unit that inputs the information acquired by the acquisition unit into a large language model and estimates the category of the OSS based on the output result of the large language model; and an output unit that outputs the category estimated by the estimation unit.
2. The estimation device according to claim 1, wherein the acquisition unit creates a predetermined search query from the name of the OSS and acquires search results for a predetermined number of top search results when the search query is input to a search site.
3. The estimation device according to claim 1, wherein the output unit outputs the category with the most estimated times based on the estimation results by the estimation unit for the OSS multiple times.
4. The estimation device according to claim 1, wherein the estimation unit calculates the confidence level of the estimated category, and the output unit outputs the confidence level calculated by the estimation unit together with the category.
5. The estimation device according to claim 1, further comprising a reconstruction unit that performs a predetermined process on the category for which the misestimation is equal to or greater than a predetermined threshold and reconstructs a category list indicating the category as an estimation candidate, and the estimation unit inputs the category list reconstructed by the reconstruction unit into the large language model and estimates the category.
6. An estimation method executed by an estimation device, comprising: an acquisition step of acquiring information about the OSS using the name of the OSS; an estimation step of inputting the information acquired by the acquisition step into a large language model and estimating the category of the OSS based on the output result of the large language model; and an output step of outputting the category estimated by the estimation step.
7. An estimation program for causing a computer to execute an acquisition procedure for acquiring information about the OSS using the name of the OSS, an estimation procedure for inputting the information acquired by the acquisition procedure into a large language model and estimating the category of the OSS based on the output result of the large language model, and an output procedure for outputting the category estimated by the estimation procedure.
Citation Information
Patent Citations
Systems and methods to analyze open source components in software products
US20190005206A1