Credit Information Generation Device and Credit Information Generation Method
The credit information generation device accurately evaluates machine learning programs by using development element information to identify similar projects and calculate uncertainty indices, addressing the challenge of assessing commercial value and feasibility in financial decisions.
Patent Information
- Application Number
- JP2021129829
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-08-06
AI Technical Summary
Existing methods fail to accurately evaluate the development of machine learning programs due to the lack of suitable risk assessment techniques, making it difficult to determine the commercial value and feasibility during financing decisions.
A credit information generation device and method that utilizes development element information from past and new machine learning programs to identify similar projects, calculate learning difficulty, and generate credit information based on technical uncertainty indices, enabling more precise evaluation of future machine learning program development.
Enables accurate evaluation of machine learning programs by quantifying technical uncertainty, allowing for better financial decision-making and risk assessment during program development.
Smart Images

Figure 0007702726000001 
Figure 0007702726000002 
Figure 0007702726000003
Abstract
Description
Technical Field
[0001] The present invention relates to a credit information generation device and a credit information generation method, and is particularly suitable for use in a device and a method for generating credit information related to the development of a machine learning program.
Background Art
[0002] When requesting a development vendor to develop a machine learning program to be used in-house, it is assumed that a financial institution will provide financing for the development funds. When providing financing, it is required to evaluate the development of the machine learning program. On the other hand, in the development of a machine learning program, the commercial value of the machine learning program (the profit that the machine learning program can generate in the future) may be clarified through PoC (Proof of Concept).
[0003] PoC is an experimental verification process for a newly developed machine learning program to confirm whether it is feasible and whether the intended effects and functions can be obtained. PoC is not a verification by theory or calculation, but a specific verification is performed by creating a simplified version of the product or system and actually using it. Therefore, it is common that PoC has not been started at the time of determining financing, and it has been difficult to perform an evaluation related to financing by PoC.
[0004] Note that prior art documents that disclose risk management in program development and the like are known (see, for example, Patent Documents 1 and 2). Patent Document 1 discloses that the "difficulty of problems" is quantified as a feature amount by referring to the performance information of past projects, and a risk evaluation of the quantified risk value is performed. Patent Document 2 discloses that similar past cases are extracted, and a risk determination of a new case is performed based on the risk information of the similar cases.
[0005] However, in the systems described in Patent Documents 1 and 2, the program development subject to risk assessment is not the development of machine learning programs. Therefore, even if the technologies described in Patent Documents 1 and 2 are used, an accurate assessment related to the development of machine learning programs cannot be performed.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0007] The present invention has been made to solve such problems, and an object thereof is to enable more accurate evaluation of a machine learning program scheduled for development.
Means for Solving the Problems
[0008] In order to solve the above-described problems, in the present invention, based on the development element information of a program scheduled for development, which is a machine learning program to be developed, and the development element information of a developed program, which is a machine learning program developed in the past, a developed program having a similarity to the program scheduled for development equal to or greater than a predetermined value is specified, and using the past development element information of the specified developed program and the new development element information of the program scheduled for development, credit information related to the development of the program scheduled for development is generated.
Effects of the Invention
[0009] According to the present invention configured as described above, since credit information related to the development of a machine learning program is generated using development element information specific to the development of a machine learning program, it is possible to more accurately evaluate a machine learning program scheduled for development.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Modes for Carrying Out the Invention
[0011] (First Embodiment) Hereinafter, a first embodiment of the present invention will be described with reference to the drawings. FIG. 1 is a block diagram showing a functional configuration example of a credit information generation device 1 according to the first embodiment. The credit information generation device 1 according to the first embodiment is a device that generates credit information related to the development of a machine learning program, and as a functional configuration, includes a new development information acquisition unit 11, a past development information acquisition unit 12, a past similar program identification unit 13, and a credit information generation unit 14. Further, a past development information DB storage unit 101 as a storage medium is connected to the credit information generation device 1.
[0012] Each of the above functional blocks 11 to 14 can be configured by any of hardware, DSP (Digital Signal Processor), and software. For example, when configured by software, each of the above functional blocks 11 to 14 is actually configured with a computer's CPU, RAM, ROM, etc., and is realized by the operation of a program stored in a storage medium such as RAM, ROM, hard disk, or semiconductor memory. The past development information DB storage unit 101 may be connected to the credit information generation device 1 by wire or wirelessly, or may be provided in a server that can be connected via a communication network such as the Internet.
[0013] In this embodiment, the machine learning program for which credit information is generated is a program developed through machine learning, which discovers certain rules from the input data and realizes inference, prediction, etc. for unknown data based on those rules. Machine learning may be based on any of regression models, tree models, neural network models, Bayesian models, clustering models, etc. Also, the machine learning program targeted in this embodiment is a program for any business utilization of BtoB, CtoC, or BtoC. The credit information generated in this embodiment is information that can be used as a reference when a financial institution provides financing for the development funds of a machine learning program scheduled for development, and is information that evaluates the uncertainty of the development of the machine learning program.
[0014] The new development information acquisition unit 11 acquires development element information (hereinafter referred to as new development element information) of a machine learning program scheduled for development (hereinafter referred to as the development-scheduled program). For example, the new development information acquisition unit 11 acquires the new development element information by reading a file 102 in which the new development element information is recorded. Alternatively, the new development information acquisition unit 11 may provide an input screen for the user to manually input the new development element information to the display, and acquire the new development element information input by the user operating an input device such as a mouse or keyboard through the input screen.
[0015] The new development element information acquired by the new development information acquisition unit 11 includes information regarding data (hereinafter sometimes referred to as learning data) to be learned by the development schedule program. For example, the new development element information includes at least one of information indicating the type and amount of data to be learned by the development schedule program, information indicating the type of machine learning to be performed on the development schedule program, information indicating the type of one or more algorithms used in the development schedule program, information indicating the quality of annotation, information indicating the number of labels identified by the development schedule program, and information indicating the accuracy index required for the development schedule program.
[0016] The type of learning data is the type of data format such as image / audio / text. The amount of learning data is the quantity of data. Note that the amount of learning data may include additional information regarding the data (such as metadata such as size in an image). For example, when the amount of data used for the development of the development schedule program is less than the amount of data used for the development of a machine learning program developed in the past (hereinafter referred to as a developed program), the technical uncertainty related to the development of the development schedule program increases.
[0017] The type of machine learning is a type specified by the combination of the type of data used in machine learning (such as images / audio / natural language / tabular format / ··· etc.) and the method used in machine learning (such as mining / classification / detection / tracking / preference extraction / ··· etc.). The type of algorithm is of types such as linear regression, logistic regression, support vector machine, decision tree, regression tree, random forest, gradient boosting tree, perceptron, convolutional neural network, recurrent neural network, residual network, RBF network, probabilistic neural network, spiking neural network, complex neural network, Bayesian inference, k-nearest neighbor method, hierarchical clustering, non-hierarchical clustering, topic model, etc. For example, if the type of machine learning of the program scheduled for development is not similar to any of the types of machine learning of the developed programs, the technical uncertainty related to the development of the program scheduled for development will increase. The same applies to the type of algorithm.
[0018] The quality of annotation is the quality of the correct data given as teacher data to a machine learning program. Since the annotation to be given varies depending on the learning data and the type of machine learning, the information on what kind of annotation was used may be important in evaluating the development of a machine learning program. The information indicating the quality of annotation can be, for example, information related to a relative index preset by an engineer or the like from the viewpoint of whether it is correctly annotated. The quality of annotation may also be determined considering the number of identification labels described below. For example, if the quality of the annotation used for the development of the program scheduled for development is lower than the quality of the annotation used during the development of the developed program, the technical uncertainty related to the development of the program scheduled for development will increase.
[0019] The number of identification labels is the number of correct data used for machine learning. For example, in the case of a machine learning program that simply determines whether a person is present or not, it is a single label. However, in the case of a machine learning program that simultaneously identifies multiple targets, it is necessary to prepare multiple correct data, and the number of identification labels becomes multiple. Generally, when the number of identification labels increases, the quality of annotation may decrease. Therefore, it is possible to standardize the quality of annotation in consideration of the number of identification labels. Note that depending on the type of machine learning, such as unsupervised learning (clustering), annotation may not be necessary. In this case, information indicating the quality of annotation and information indicating the number of identification labels are not required.
[0020] The required accuracy index is the accuracy to be achieved required for the machine learning program. For example, in the case of a machine learning program that handles identification problems, it is possible to use indices such as Confusion Matrix, F-value, accuracy rate, precision rate, and recall rate. Also, in the case of a machine learning program that handles regression problems, it is possible to use indices such as RMSE (Root Mean Square Error), MSE (Mean Square Error), and MAE (Mean Absolute Error). Additionally, indices such as AIC as a relative evaluation index of the model itself and the value of the loss function (Loss) during learning may be used. For example, when the accuracy required for the program scheduled for development is higher than the accuracy required during the development of the already developed program, the technical uncertainty related to the development of the program scheduled for development increases.
[0021] The past development information acquisition unit 12 acquires development element information of developed programs (hereinafter referred to as past development element information) from the past development information DB storage unit 101. The past development information DB storage unit 101 stores past development element information regarding a plurality of machine learning programs developed in the past. This past development element information includes information regarding data (learning data) learned by the developed program. Similar to the new development element information, this past development element information also includes at least one of information indicating the type and amount of data learned by the developed program, information indicating the type of machine learning performed on the developed program, information indicating the type of one or more algorithms used in the developed program, information indicating the quality of the annotation, information indicating the number of labels identified by the developed program, and information indicating the accuracy index required for the developed program.
[0022] Based on the new development element information of the planned development program acquired by the new development information acquisition unit 11 and the past development element information of the developed program acquired by the past development information acquisition unit 12, the past similar program identification unit 13 identifies a developed program whose similarity to the planned development program is equal to or greater than a predetermined value. Hereinafter, the developed program identified in this way is referred to as a past similar program. Here, the past similar program identification unit 13 calculates the similarity using all or part of the development element information. For example, it is possible to calculate the similarity between the planned development program and the developed program based on the information indicating the type of machine learning.
[0023] When calculating similarity using machine learning types, for example, for each of the planned development program and the developed program, for each of the types of data used in machine learning (such as images / audio / natural language / tabular form / ··· etc.) and the methods used in machine learning (such as mining / classification / detection / tracking / preference extraction / ··· etc.), the corresponding items are extracted and vectorized, and it is possible to use a method of calculating the similarity of the vectors. Specifically, the distance between two vectors is calculated, and the similarity of the n nearest neighbors is calculated. Cosine similarity or the like may be used as an index for evaluating the degree of proximity. When calculating similarity, weighting may be performed for each item.
[0024] The credit information generation unit 14 generates credit information related to the development of the planned development program using the new development element information of the planned development program and the past development element information of the past similar program specified by the past similar program specifying unit 13. In this embodiment, as an example, the credit information generation unit 14 uses the new development element information of the planned development program and the past development element information of the past similar program to evaluate the technical uncertainty related to the development of the planned development program, and generates credit information using this evaluation result.
[0025] FIG. 2 is a block diagram showing a specific functional configuration example of the credit information generation unit 14 according to the first embodiment. As shown in FIG. 2, the credit information generation unit 14 according to the first embodiment includes a learning difficulty evaluation unit 14A, a technical uncertainty index value calculation unit 14B, and a credit information generation unit 14C as functional configurations.
[0026] The learning difficulty evaluation unit 14A evaluates the learning difficulty of the planned development program based on the new development element information of the planned development program and the past development element information of the past similar program. Here, the learning difficulty evaluation unit 14A first calculates the learning difficulty LD -P of the past similar program based on the past development element information of the past similar program. Then, based on the calculated learning difficulty LD -P and the new development element information of the planned development program, the learning difficulty LD -B of the planned development program is calculated.
[0027] For example, the learning difficulty evaluation unit 14A uses known machine learning methods such as multiple regression analysis, with information indicating the type of machine learning, the amount of learning data, and the quality of annotation (which may also be the number of identification labels) among the past development element information of past similar programs as input, to calculate the learning difficulty LD of the past similar program -P It is possible to do so. Note that the information to be input is not limited to this. For example, information indicating the type of learning data, the type of algorithm, and the accuracy index required for the program to be developed may be further used
[0028] In addition, the learning difficulty evaluation unit 14A may calculate the learning difficulty LD of the past similar program by the following different method -P First, calculate the approximate learning difficulty LD of the machine learning program from the order of the amount of learning data required for learning. Here, the learning difficulty LD of the machine learning program is obtained by logarithmic transformation of the amount of data required for learning. The number of digits obtained by this logarithmic transformation is defined as the learning difficulty LD in the amount of learning data of the machine learning program. Also, in the case of a machine learning program where the number of identification labels is determined (excluding natural language systems, etc.), the number of identification labels is logarithmically transformed. The number of digits obtained by this logarithmic transformation is defined as the learning difficulty LD in the number of identification labels of the machine learning program. Then, the learning difficulty evaluation unit 14A calculates the learning difficulty LD of the past similar program according to the calculation formula LD = LD × LD -X Calculate the learning difficulty LD of the machine learning program -X Here, the learning difficulty LD of the machine learning program is obtained by logarithmic transformation of the amount of data required for learning. The number of digits obtained by this logarithmic transformation is defined as the learning difficulty LD in the amount of learning data of the machine learning program -X Let it be so. Also, in the case of a machine learning program where the number of identification labels is determined (excluding natural language systems, etc.), the number of identification labels is logarithmically transformed. The number of digits obtained by this logarithmic transformation is defined as the learning difficulty LD in the number of identification labels of the machine learning program -Z Let it be so. And the learning difficulty evaluation unit 14A -P = LD -X × LD -Z Calculate the learning difficulty LD of the past similar program according to the calculation formula -P Let it be so
[0029] Next, the learning difficulty evaluation unit 14A, based on the learning difficulty LD of the past similar program calculated as above and the new development element information of the program to be developed, calculates the learning difficulty LD of the program to be developed -P And the learning difficulty LD of the program to be developed -BCalculate it. Here, as an example, among the newly developed element information, the amount of learning data is used. For example, let the number of learning data available for machine learning of the planned development program be n, and perform a logarithmic transformation on the number n of the learning data (n' = log 10 n). At this time, in the case of a machine learning program for a multi-class classification problem, the logarithmic transformation value n' may be calculated in consideration of the variation in classes included in the learning data. For example, the standard deviation of the data amount between classes may be multiplied by the logarithmic transformation value n'. Then, the learning difficulty evaluation unit 14A calculates the learning difficulty LD -B =LD -P / n' according to the calculation formula, and calculates the learning difficulty LD -B of the planned development program. As is clear from this calculation formula, when the learning difficulty LD -P of the past similar program is high or the number of learning data available for machine learning is small, the learning difficulty LD -B of the planned development program becomes high. Note that standardization may be performed using a known normalization method such as the sigmoid function.
[0030] Based on the learning difficulty LD -B of the planned development program evaluated by the learning difficulty evaluation unit 14A and the similarity between the planned development program obtained when the past similar program specifying unit 13 specifies the past similar program, the technical uncertainty index value calculation unit 14B calculates an index value indicating the technical uncertainty related to the development of the planned development program. When the similarity is DS and the technical uncertainty index value is UI -T , the technical uncertainty index value calculation unit 14B calculates the technical uncertainty index value UI -T =LD -B / DS according to the calculation formula, and calculates the technical uncertainty index value UI -T related to the development of the planned development program. As is clear from this calculation formula, the larger the similarity DS with the past similar program, the smaller the technical uncertainty index value UI -T of the planned development program, and the larger the learning difficulty LD -B of the planned development program, the larger the technical uncertainty index value UI -T .
[0031] Note that there may be cases where the past similar program identification unit 13 fails to identify a developed program with a similarity to the planned development program equal to or greater than a predetermined value. In such a case, the technical uncertainty index value calculation unit 14B may calculate the technical uncertainty index value UI -T to be maximized.
[0032] The credit information generation unit 14C generates credit information based on the technical uncertainty index value UI -T evaluated by the technical uncertainty index value calculation unit 14B. The credit information generated here can be, for example, the volatility of the expected profit that the planned development program may generate in the future. Alternatively, the credit information may be the development success rate or the expected development period of the planned development program. Alternatively, the credit information may be flag information indicating whether the planned development program has a prospect of generating profit in the future. The credit information generation unit 14C uses a predetermined function with the technical uncertainty index value UI -T as an explanatory variable and any one of volatility, development success rate, expected development period, or the presence or absence of a prospect of generating profit as an objective variable to generate credit information.
[0033] FIG. 3 is a flowchart showing an operation example of the credit information generation apparatus 1 according to the first embodiment configured as described above. In FIG. 3, first, the new development information acquisition unit 11 acquires new development element information of the planned development program (step S1). Also, the past development information acquisition unit 12 acquires past development element information of the developed program from the past development information DB storage unit 101 (step S2).
[0034] Then, the past similar program identification unit 13 identifies a past similar program with a similarity to the planned development program equal to or greater than a predetermined value based on the new development element information of the planned development program acquired by the new development information acquisition unit 11 and the past development element information of the developed program acquired by the past development information acquisition unit 12 (step S3).
[0035] Next, the learning difficulty evaluation unit 14A of the credit information generation unit 14 calculates the learning difficulty LD of a past similar program based on the past development element information of the past similar program (for example, information indicating the type of machine learning, the amount of learning data, the quality of annotation, or the number of identification labels). -P (Step S4). Further, the learning difficulty evaluation unit 14A calculates the learning difficulty LD' of the planned development program based on the calculated learning difficulty LD of the past similar program and the new development element information of the planned development program (for example, information indicating the amount of learning data) using the calculation formula LD' = LD / n' (Step S5). -P -B -P -B
[0036] Next, the technical uncertainty index value calculation unit 14B calculates the technical uncertainty index value UI of the planned development program based on the learning difficulty LD of the planned development program evaluated by the learning difficulty evaluation unit 14A and the similarity DS between the planned development program obtained when the past similar program identification unit 13 identifies the past similar program using the calculation formula UI = LD / DS (Step S6). Then, the credit information generation unit 14C generates credit information based on the technical uncertainty index value UI evaluated by the technical uncertainty index value calculation unit 14B (Step S7). Thus, the processing of the flowchart shown in FIG. 3 is completed. -B -T -B -T -T
[0037] As described in detail above, in the first embodiment, based on the new development element information of the planned development program (including information related to the data to be learned by the planned development program) and the past development element information of the developed program (including information related to the data learned by the developed program), a past similar program with a similarity to the planned development program equal to or greater than a predetermined value is identified, and the past development element information of the identified past similar program and the new development element information of the planned development program are used to generate credit information related to the development of the planned development program.
[0038] According to the first embodiment configured as described above, since the credit information related to the development of the machine learning program is generated using the development element information including information on the learning data specific to the development of the machine learning program, compared with the techniques described in Patent Documents 1 and 2 that perform risk assessment of program development without using learning data, the evaluation of the machine learning program scheduled for development can be performed more accurately.
[0039] Further, according to the first embodiment, it is possible to quantitatively present the technical uncertainty related to the development of the machine learning program. Here, in the first embodiment, the learning difficulty LD -P of the past similar program and the new development element information of the program scheduled for development are used to calculate the learning difficulty LD -B of the program scheduled for development, and then the technical uncertainty index value UI -T related to the development of the machine learning program is calculated. As a result, different from the techniques described in Patent Documents 1 and 2 that regard the risks of past similar cases as the risks of new cases as they are and perform risk assessment by presuming that the development difficulty of new cases similar to those with high development difficulty in the past cases is also high, it is possible to perform an accurate evaluation related to development according to the actual development situation of the program scheduled for development.
[0040] In the above first embodiment, an example of generating credit information based on the technical uncertainty index value UI -T related to the development of the program scheduled for development has been described, but the present invention is not limited to this. For example, credit information may be generated based on the learning difficulty LD -B of the program scheduled for development evaluated by the learning difficulty evaluation unit 14A. In this case, the technical uncertainty index value calculation unit 14B is omitted.
[0041] In addition, in the above-described first embodiment, an example has been described in which the past development element information of the developed program includes at least one of information indicating the type and amount of learning data, the type of machine learning, the type of algorithm, the quality of annotation, the number of identification labels, and the required accuracy index. However, information indicating the number of times of implementation of past cases and their respective success or failure may be included. This success or failure information may be classified for each type of machine learning. The success or failure information may be the total of the number of cases in which the machine learning program could be delivered as successful cases among the past implemented cases. In this case, the learning difficulty evaluation unit 14A calculates the learning difficulty LD -P of the past similar programs including this success or failure information. For example, the learning difficulty evaluation unit 14A may calculate LD -P =LD -X ×LD -Z × success rate, and calculate the learning difficulty LD -P of the past similar programs according to such a calculation formula.
[0042] In addition, in the above-described first embodiment, an example has been described in which the learning difficulty LD -P of the past similar programs is calculated from the past development element information by calculation. However, the present invention is not limited to this. For example, the past development element information stored in the past development information DB storage unit 101 may include information indicating the learning difficulty of the developed program. The past learning difficulty stored in the past development information DB storage unit 101 in advance may be the result calculated as described above, or may be an index determined in advance by an engineer or the like. In addition, the past learning difficulty stored in the past development information DB storage unit 101 may be classified for each type of machine learning. When storing the information on the past learning difficulty in the past development information DB storage unit 101 in this way, the learning difficulty evaluation unit 14A acquires the information indicating the past learning difficulty of the past similar programs from the past development information DB storage unit 101, and based on this and the new development element information of the program to be developed, calculates the learning difficulty LD -B of the program to be developed.
[0043] In the first embodiment described above, an example of generating trust information for development has been described for a machine learning program developed using supervised machine learning and unsupervised machine learning that require learning data. However, the present invention is not limited to this. For example, it is also possible to generate trust information for development for a machine learning program developed using reinforcement learning that does not require learning data, using the same method as in the first embodiment.
[0044] Reinforcement learning is a machine learning technique that explores a method of maximizing value under given conditions through trial and error. In reinforcement learning, the size of the state space in which an actor (agent) in a machine learning program can exist and the size of the action space that the agent can take are important elements in calculating the technical uncertainty index value. Therefore, it is possible to use information regarding the size of the state space and the size of the action space of reinforcement learning as development element information of the machine learning program.
[0045] In this case, the past similar program identification unit 13 identifies a developed program whose similarity to the planned development program is equal to or greater than a predetermined value as a past similar program based on the new development element information regarding the size of the state space and the size of the action space of the planned development program and the past development element information regarding the size of the state space and the size of the action space of the developed program. Further, the trust information generation unit 14 generates trust information related to the development of the planned development program using the new development element information of the planned development program and the past development element information of the past similar program. Information regarding the size of the state space and the size of the action space of the developed program may be stored in advance in the past development information DB storage unit 101.
[0046] Generally, the state space is updated every time the agent takes an action. For example, when the machine learning program is for a tic-tac-toe game, since each of the 3×3 grids has three states of "〇", "×", and "blank", 3 9=19683 is the size of the discrete state space. Also, when the machine learning program targets the control of a robotic arm, the number of possible combinations regarding the range of motion of each joint of the robotic arm, the angular velocity of each joint, etc. becomes the size of the state space.
[0047] The action space in a machine learning program can also be said to be a virtual space corresponding to the options that an agent can take. For example, when the machine learning program targets a tic-tac-toe game, the number of patterns that can be taken as the state of the grid (where "〇" or "×" is in the grid) in each game board position becomes the size of the action space. Also, when the machine learning program targets the control of a robotic arm, the range of motion of each joint of the robotic arm that can be taken in the next frame, etc. becomes the size of the action space.
[0048] (Second Embodiment) Next, a second embodiment of the present invention will be described with reference to the drawings. FIG. 4 is a block diagram showing a functional configuration example of the credit information generation device 2 according to the second embodiment. In this FIG. 4, those denoted by the same reference numerals as those shown in FIG. 1 have the same functions, so duplicate explanations are omitted here.
[0049] The credit information generation device 2 according to the second embodiment further includes a new utilization scenario information acquisition unit 21, a past utilization scenario information acquisition unit 22, and a past similar utilization scenario identification unit 23. Also, the credit information generation device 2 according to the second embodiment includes a credit information generation unit 24 instead of the credit information generation unit 14. Further, a past business information DB storage unit 201 is connected to the credit information generation device 2 as a storage medium.
[0050] Each of the above functional blocks 21 to 24 can also be configured by any of hardware, DSP, and software. For example, when configured by software, each of the above functional blocks 21 to 24 is actually configured with a computer's CPU, RAM, ROM, etc., and is realized by the operation of a program stored in a storage medium such as RAM, ROM, hard disk, or semiconductor memory. The past business information DB storage unit 201 may be connected to the credit information generation device 2 by wire or wirelessly, or may be provided in a server that can be connected via a communication network such as the Internet.
[0051] The new utilization scenario information acquisition unit 21 acquires utilization scenario information (hereinafter referred to as new utilization scenario information) indicating the business utilization scenario of the development-scheduled program by users who utilize the development-scheduled program in business (such as requesters who request a development vendor to develop a machine learning program). The new utilization scenario information acquired by the new utilization scenario information acquisition unit 21 includes information related to at least one of the company situation, business situation, financial status of the user of the development-scheduled program, and the utilization scenario of the machine learning program by the user.
[0052] The company situation refers to the situation of officers, employees, the existence and situation of affiliated companies, the situation of shareholders, an overview of corporate governance, the situation of audits, etc. The business situation refers to the industry type, occupation type, business format, business environment, major risks affecting the business, management policy, business model, the position of the user in the industry value chain, the trend of major management indicators and the management's recognition of them. The industry type is classified with reference to, for example, the "Japanese Standard Industrial Classification" provided by the Statistics Bureau of the Ministry of Internal Affairs and Communications. The occupation type is classified with reference to, for example, the value chain that classifies business activities by function. The business format is classified with reference to, for example, the "Business Format Classification Table" provided by the Ministry of Economy, Trade and Industry.
[0053] The financial condition refers to information related to financial statements, sources of capital and liquidity of funds, methods and situations of fundraising, status of stocks, etc. The utilization scenarios of the machine learning program refer to information on the value chain where the program scheduled for development will be introduced, issues to be solved in the value chain, means for solution, expected effects, etc.
[0054] The new utilization scenario information acquisition unit 21 acquires new utilization scenario information, for example, by reading a file 202 in which new development element information is recorded. Alternatively, the new utilization scenario information acquisition unit 21 may provide a display with an input screen for the user to manually input new utilization scenario information, and acquire the new utilization scenario information input by the user operating an input device such as a mouse or keyboard through the input screen. Note that the new utilization scenario information may be acquired by referring to securities reports, etc. from EDINET provided by the Financial Services Agency, obtaining it from a research company, conducting on-site inspections, or other methods of obtaining public information.
[0055] The past utilization scenario information acquisition unit 22 acquires utilization scenario information (hereinafter referred to as past utilization scenario information) indicating past business utilization scenarios of developed programs from the past business information DB storage unit 201. The past business information DB storage unit 201 stores past utilization scenario information regarding a plurality of machine learning programs developed in the past. This past utilization scenario information includes at least one of the business situation (industry type, occupation type, business form, etc.) and the utilization scenario of the machine learning program by the user (information on the value chain where the developed program is introduced, issues solved in the value chain, means for solution, introduction effect, etc.). Note that the past utilization scenario information may also include information on the company situation and financial condition of the user of the program scheduled for development, similar to the new utilization scenario information.
[0056] The past similar usage scene specifying unit 23 specifies past usage scene information with a similarity to the new usage scene information of a predetermined value or more based on the new usage scene information acquired by the new usage scene information acquisition unit 21 and the past usage scene information acquired by the past usage scene information acquisition unit 22. Hereinafter, the past usage scene information specified in this way is referred to as a past similar scene. Here, the past similar usage scene specifying unit 23 calculates the similarity using all or part of the usage scene information. For example, it is possible to calculate the similarity between the new usage scene information and the past usage scene information based on information indicating the business situation (industry type, occupation type, business form). The similarity may be calculated using further information indicating the usage scene of the machine learning program (for example, the problem to be solved in the value chain). The number of past similar scenes to be specified may be one or more.
[0057] The similarity between the new usage scene information and the past usage scene information can be evaluated, for example, by the proximity degree DN of the feature amounts included in the various information used for calculating the similarity. This proximity degree DN can be obtained by, for example, extracting keywords from the feature amounts of the various information and using a well-known algorithm such as the k-nearest neighbor method. Here, when the proximity degree DN is large, it is determined that the similarity is large, and when the proximity degree DN is small, it is determined that the similarity is small. Note that the determination of the similarity may be automated using a natural language processing algorithm.
[0058] The credit information generation unit 24 generates credit information related to the development of the program scheduled for development using the new development element information of the program scheduled for development, the past development element information of the past similar program specified by the past similar program specifying unit 13, and the past usage scene information of the past similar scene specified by the past similar usage scene specifying unit 23. When generating this credit information, the credit information generation unit 24, in addition to the technical uncertainty index value UI -T calculates an index value UI -B indicating the business uncertainty related to the development of the program scheduled for development, and generates credit information based on the technical uncertainty index value UI -T and the business uncertainty index value UI -B .
[0059] FIG. 5 is a block diagram showing a specific functional configuration example of the credit information generation unit 24 according to the second embodiment. In FIG. 5, components denoted by the same reference numerals as those in FIG. 2 have the same functions, and thus redundant descriptions are omitted here. As shown in FIG. 5, the credit information generation unit 24 according to the second embodiment further includes a utilization risk evaluation unit 24A and a business uncertainty index value calculation unit 24B as functional components. Further, the credit information generation unit 24 according to the second embodiment includes a credit information generation unit 24C instead of the credit information generation unit 14C.
[0060] The utilization risk evaluation unit 24A evaluates the degree of risk related to the business utilization of the developed program based on the past utilization scene information of the past similar scenes acquired by the past utilization scene information acquisition unit 22. Here, the utilization risk evaluation unit 24A evaluates the risk of machine learning program development and calculates a risk score RS -B For example, the utilization risk evaluation unit 24A evaluates the risk of machine learning program development as follows based on the information on the introduction effect on the value chain of the developed program. That is, when the introduction effect extends to a wide range of the value chain, the risk is large, and a large risk score RS -B is output. The case where the introduction effect extends to a wide range of the value chain is, for example, the case of developing a machine learning program related to the launch of a new business. On the other hand, when the introduction effect extends to a narrow range of the value chain, the risk is small, and a small risk score RS -B is output. The case where the introduction effect extends to a narrow range of the value chain is, for example, the case of developing a machine learning program for the purpose of improving the efficiency of a certain one operation.
[0061] Here, an example of calculating the risk score RS -B based on the information on the introduction effect among the past utilization scene information is shown, but the risk score RS -Bmay be calculated. Also, here, the risk score RS is calculated from the past usage scenario information. -B Although an example of calculating the risk score RS has been shown, the risk score RS determined in advance considering the past usage scenario information -B may be stored in advance in the past business information DB storage unit 201.
[0062] The business uncertainty index value calculation unit 24B is based on the risk score RS calculated by the usage risk evaluation unit 24A -B (the degree of usage risk evaluated by the usage risk evaluation unit 24A) and the similarity (proximity DN) with the new usage scenario information obtained when the past similar usage scenario identification unit 23 identifies the past similar scenario, and calculates the business uncertainty index value UI -B =RS -B / DN according to the calculation formula, and calculates the business uncertainty index value UI -B relating to the development of the planned development program. As is clear from this calculation formula, the greater the proximity DN to the past usage scenario information, the smaller the business uncertainty index value UI -B of the planned development program, and the greater the risk score RS -B , the greater the business uncertainty index value UI -B .
[0063] The credit information generation unit 24C generates credit information based on the technical uncertainty index value UI -T evaluated by the technical uncertainty index value calculation unit 14B and the business uncertainty index value UI -B evaluated by the business uncertainty index value calculation unit 24B.
[0064] The types of business are diverse. However, considering the application of the machine learning program, whether there has been a past performance in the same business may be important in measuring the uncertainty related to the development of the machine learning program. That is, when the target business of the planned development program is similar to the target business of the developed program, it is conceivable that the business uncertainty will be low. In the second embodiment, this business uncertainty index value UI -BSince the credit information is generated taking this into account as well, it is possible to more accurately evaluate the development of the development schedule program.
[0065] (Third Embodiment) Next, a third embodiment of the present invention will be described with reference to the drawings. FIG. 6 is a block diagram showing a functional configuration example of the credit information generation device 3 according to the third embodiment. In this FIG. 6, those denoted by the same reference numerals as those shown in FIG. 4 have the same functions, and thus redundant explanations are omitted here.
[0066] The credit information generation device 3 according to the third embodiment further includes a development process information acquisition unit 31. Also, the credit information generation device 3 according to the third embodiment includes a credit information generation unit 34 instead of the credit information generation unit 24. These functional blocks 31, 34 can also be configured by any of hardware, DSP, and software.
[0067] The development process information acquisition unit 31 acquires the development process information of the development schedule program. The development process information acquired by the development process information acquisition unit 31 includes, for example, at least one of the presence or absence of a collaborative vendor, the past development achievements of the collaborative vendor, the capital status of the collaborative vendor, the development schedule, and experts on the client side. The development process information acquisition unit 31 acquires the development process information, for example, by reading a file 302 in which the development process information is recorded. Alternatively, the development process information acquisition unit 31 may provide an input screen for the user to manually input the development process information to the display, and acquire the development process information input by the user operating an input device such as a mouse or a keyboard through the input screen.
[0068] The credit information generation unit 34 generates credit information related to the development of the planned development program using the new development element information of the planned development program, the past development element information of the past similar program specified by the past similar program specifying unit 13, the past utilization scene information of the past similar scene specified by the past similar utilization scene specifying unit 23, and the development process information acquired by the development process information acquisition unit 31. When generating this credit information, the credit information generation unit 34 uses the technical uncertainty index value UI -T and the business uncertainty index value UI -B In addition, it calculates the development process uncertainty index value UI -P and generates credit information based on the technical uncertainty index value UI -T , the business uncertainty index value UI -B and the development process uncertainty index value UI -P .
[0069] In the development of a machine learning program, it is possible to consider developing it alone by a single company, developing it including collaborating vendors, or developing it within the framework of a development consortium. Generally, as the number of stakeholders involved in the development increases, the need to align progress may arise, which may be a factor increasing the uncertainty of the development. Therefore, the credit information generation unit 34 calculates the development process uncertainty index value UI -P according to a function such that the value increases as the number of collaborating vendors increases, the value increases as the past development record of the collaborating vendors is poorer, and the value increases as the development schedule is shorter.
[0070] According to the third embodiment configured as described above, in addition to the technical uncertainty index value UI -T and the business uncertainty index value UI -B , the development process uncertainty index value UI -P is also taken into account to generate credit information, so it is possible to more accurately evaluate the development of the planned development program.
[0071] In the third embodiment described above, the technical uncertainty index value UI -T , the business uncertainty index value UI -Band the development process uncertainty index value UI -P Although an example of generating credit information based on these three index values has been described, the present invention is not limited thereto. For example, the technical uncertainty index value UI -T and the development process uncertainty index value UI -P may be used to generate credit information based on these two index values.
[0072] Furthermore, each of the above-described first to third embodiments merely shows an example of implementation of the present invention, and the technical scope of the present invention should not be construed as being limited thereby. That is, the present invention can be implemented in various forms without departing from the gist or main features thereof.
Explanation of Reference Numerals
[0073] 1, 2, 3 Credit information generation device 11 New development information acquisition unit 12 Past development information acquisition unit 13 Past similar program identification unit 14 Credit information generation unit 14A Learning difficulty evaluation unit 14B Technical uncertainty index value calculation unit 14C Credit information generation unit 21 New utilization scenario information acquisition unit 22 Past utilization scenario information acquisition unit 23 Past similar utilization scenario identification unit 24 Credit information generation unit 24A Utilization risk evaluation unit 24B Business uncertainty index value calculation unit 24C Credit information generation unit 31 Development process information acquisition unit 34 Credit information generation unit
Claims
1. An apparatus for generating credit information related to the development of a machine learning program, comprising: a new development information acquisition unit that acquires development element information of a program to be developed, which is a machine learning program to be developed; a past development information acquisition unit that acquires development element information of a developed program, which is a machine learning program developed in the past; a past similar program identification unit that identifies a developed program having a similarity with the program to be developed equal to or greater than a predetermined value based on new development element information, which is the development element information of the program to be developed acquired by the new development information acquisition unit, and past development element information, which is the development element information of the developed program acquired by the past development information acquisition unit; a credit information generation unit that generates credit information for evaluating the uncertainty of the development of the program to be developed by using the new development element information of the program to be developed and the past development element information of the developed program identified by the past similar program identification unit; wherein the credit information generation unit comprises a learning difficulty evaluation unit that evaluates the learning difficulty of the program to be developed based on the new development element information and the past development element information; and a technical uncertainty index value calculation unit that calculates an index value indicating the technical uncertainty related to the development of the program to be developed based on the learning difficulty evaluated by the learning difficulty evaluation unit and the similarity obtained by the past similar program identification unit; and generates the credit information based on the technical uncertainty index value evaluated by the technical uncertainty index value calculation unit. A credit information generation apparatus characterized by the above.
2. The new development element information acquired by the new development information acquisition unit includes information related to data for training the program to be developed. The past development element information acquired by the past development information acquisition unit includes information related to data for training the developed program. The credit information generation apparatus according to claim 1, characterized by the above.
3. The new development element information and the past development element information each include information indicating the type and amount of data for training the machine learning program. The credit information generation apparatus according to claim 2, characterized by the above.
4. The credit information generation device according to claim 2 or 3, wherein the new development element information and the past development element information each include information indicating the number of labels identified by the machine learning program.
5. The credit information generation device according to any one of claims 2 to 4, wherein the new development element information and the past development element information each include information indicating the quality of annotation given as teacher data to the machine learning program.
6. The credit information generation device according to any one of claims 2 to 5, wherein the new development element information and the past development element information each include information indicating the type of machine learning to be performed on the machine learning program.
7. The credit information generation device according to claim 6, wherein the information indicating the type of machine learning is information indicating the type of one or more algorithms used in the machine learning program.
8. The credit information generation device according to any one of claims 2 to 7, wherein the new development element information and the past development element information each include information indicating an accuracy index required for the machine learning program.
9. The new development element information acquired by the new development information acquisition unit includes information regarding the size of the state space and the size of the action space in the development plan program. The past development element information acquired by the past development information acquisition unit includes information regarding the size of the state space and the size of the action space in the developed program. The credit information generation device according to claim 1, characterized in that.
10. The learning difficulty evaluation unit calculates the learning difficulty of the developed program based on the past development element information, and evaluates the learning difficulty of the development plan program based on the calculated learning difficulty of the developed program and the new development element information. The credit information generation device according to claim 1, characterized in that.
11. The past development element information further includes information indicating the learning difficulty of the machine learning program. The learning difficulty evaluation unit evaluates the learning difficulty of the development plan program based on the new development element information and the information indicating the learning difficulty included in the past development element information. The credit information generation device according to claim 1, characterized in that.
12. A new usage scene information acquisition unit that acquires usage scene information indicating a business usage scene of the development schedule program by a user who uses the development schedule program for business; A past usage scene information acquisition unit that acquires usage scene information indicating past business usage scenes of the developed program; Based on the new usage scene information, which is the usage scene information acquired by the new usage scene information acquisition unit, and the past usage scene information, which is the usage scene information acquired by the past usage scene information acquisition unit, a past similar usage scene identification unit that identifies the past usage scene information whose similarity to the new usage scene information is equal to or greater than a predetermined value; The credit information generation unit further uses the past usage scene information identified by the past similar usage scene identification unit to generate the credit information The credit information generation device according to any one of claims 1 to 11, characterized in that.
13. The credit information generation unit A usage risk evaluation unit that evaluates the degree of risk related to the business usage of the developed program based on the past usage scene information; Based on the usage risk degree evaluated by the usage risk evaluation unit and the similarity obtained by the past similar usage scene identification unit, a business uncertainty index value calculation unit that calculates an index value indicating the business uncertainty related to the development of the development schedule program; The credit information is generated using the business uncertainty index value evaluated by the business uncertainty index value calculation unit. The credit information generation device according to claim 12, characterized in that.
14. Further comprising a development process information acquisition unit that acquires development process information of the development schedule program; The credit information generation unit further uses the development process information to generate the credit information The credit information generation device according to any one of claims 1 to 13, characterized in that.
15. A method for generating credit information related to the development of a machine learning program, A first step in which a new development information acquisition unit of a computer acquires development element information of a development schedule program, which is a machine learning program scheduled for development; A second step in which a past development information acquisition unit of the computer acquires development element information of a developed program, which is a machine learning program developed in the past; A third step in which a past similar program specifying unit of the computer specifies a developed program whose similarity to the planned development program is equal to or greater than a predetermined value based on new development element information, which is development element information of the planned development program acquired by the new development information acquisition unit, and past development element information, which is development element information of the developed program acquired by the past development information acquisition unit; A fourth step in which a credit information generation unit of the computer generates credit information for evaluating the uncertainty of development of the planned development program using the new development element information of the planned development program and the past development element information of the developed program specified by the past similar program specifying unit; In the fourth step, the credit information generation unit A process in which a learning difficulty evaluation unit of the computer evaluates the learning difficulty of the planned development program based on the new development element information and the past development element information; A process in which a technical uncertainty index value calculation unit of the computer calculates an index value indicating the technical uncertainty related to the development of the planned development program based on the learning difficulty evaluated by the learning difficulty evaluation unit and the similarity obtained by the past similar program specifying unit; generates the credit information based on the technical uncertainty index value evaluated by the technical uncertainty index value calculation unit A credit information generation method characterized by the above.
Citation Information
Patent Citations
Project management system, program of the system, and recording medium with the program recorded thereon
JP2007058409A
Device, method and program for managing project, and recording medium
JP2011175593A
Artificial intelligence module development method, artificial intelligence module development system, and artificial intelligence module development integration system
JP2019003603A
Automatic identification of relevant software project for cross project learning
JP2019153270A
Estimation support apparatus, estimation support program and estimation support method
JP2020017148A