Evaluation device, evaluation method, and computer program

By generating qualitative assessment results through a question-and-answer system and assessment equipment, the problems of traditional qualitative analysis being time-consuming, labor-intensive, and susceptible to subjective bias are solved, enabling rapid and objective large-scale qualitative analysis.

JP7799312B2Active Publication Date: 2026-01-15NAT INST OF INFORMATION & COMM TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022013852
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-01
Publication Date
2026-01-15
Estimated Expiration
2042-02-01

AI Technical Summary

Technical Problem

Traditional quantitative analysis methods are easy to confirm based on objective information, but the effectiveness of qualitative analysis is difficult to verify. Furthermore, qualitative analysis is time-consuming, labor-intensive, and easily affected by subjective biases, making it difficult to apply on a large scale.

Method used

An evaluation device and computer program are used to obtain responses to multiple questions through a question-and-answer system and perform qualitative analysis based on predefined evaluation criteria. The qualitative evaluation results are generated using the text generation and evaluation model of the question-and-answer system, including the question-and-answer system, question-and-answer table, domain evaluation model and evaluation device.

Benefits of technology

It enables rapid and objective large-scale qualitative analysis, reduces subjective bias, improves analytical efficiency and accuracy, and reduces the workload of manual investigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799312000013
    Figure 0007799312000013
  • Figure 0007799312000014
    Figure 0007799312000014
  • Figure 0007799312000015
    Figure 0007799312000015
Patent Text Reader

Abstract

To perform evaluation through qualitative analysis with respect to an object, in short time, easily, and on a large scale according to objective standards.SOLUTION: An evaluation device includes: a question generation unit 64 for acquiring answers to each of a plurality of questions from a question answering system 52, about each of a plurality of objects to be evaluated; and an evaluation unit 68 for evaluating each of the plurality of objects to be evaluated in accordance with evaluation standards defined on the basis of a predetermined viewpoint, on the basis of the answers acquired by the question generation unit 64 about each of the plurality of objects to be evaluated.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an evaluation system, an evaluation device, and a computer program for evaluating an object. [Background technology]

[0002] There are two methods for evaluating business activities: quantitative analysis and qualitative analysis. With quantitative analysis, it is easy to confirm that it is based on objective information and is conducted using objective methods. However, with qualitative analysis, there is a problem in that it is not easy to confirm the validity of the analysis.

[0003] When attempting to conduct a qualitative analysis of corporate activities, it is necessary to investigate the content of corporate activities that cannot be expressed numerically. However, such investigations pose challenges different from those faced with quantitative analysis. Surveys are often used by those investigating corporate activities. However, conducting a large-scale survey using a survey requires a great deal of work, including creating the survey, responding to the survey, compiling the survey results, and analyzing them. Each of these stages requires a great deal of effort and expense. In particular, when conducting a detailed survey, the number of survey items increases, and the number of items requiring written responses also increases. As a result, the burden on survey respondents increases, leading to a lower survey response rate. Furthermore, the amount of work required to analyze the survey results is extensive, difficult, and time-consuming.

[0004] Furthermore, in the case of qualitative analysis, even experts have a limit to the amount of data they can grasp. Therefore, manual qualitative analysis can be prone to oversights and subjective bias in the analysis results. In other words, it is difficult to eliminate subjectivity and maintain consistency in evaluations when conducting manual qualitative analysis. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Machida, Kayoko, "Advantages and Points to Note when Using Text Mining in Qualitative Research," Sapporo City University Research Papers, Vol. 13, No. 1, pp. 47-53 (2019) Summary of the Invention [Problem to be solved by the invention]

[0006] Due to these factors, conventional qualitative analysis of corporate activities usually requires a huge amount of effort and expense, which results in a long time and often results that are not obtained in a timely manner. These problems arise not only in the analysis of corporate activities, but also in many other qualitative analyses, such as the analysis of university research activities and the analysis of the quality of life of ordinary citizens.

[0007] Therefore, an object of the present invention is to provide an evaluation system, an evaluation method, and a computer program therefor that can evaluate objects through qualitative analysis easily and on a large scale in accordance with objective criteria in a short time. [Means for solving the problem]

[0008] An evaluation device according to a first aspect of the present invention includes a response acquisition means for acquiring a response to each of a plurality of questions from a question-answering system for each of a plurality of evaluation objects, and an evaluation means for evaluating each of the plurality of evaluation objects in accordance with evaluation criteria defined based on a predetermined viewpoint, based on the response acquired by the response acquisition means for each of the plurality of evaluation objects.

[0009] Preferably, the evaluation device further includes a question expansion table storage means for storing a question expansion table, a domain evaluation model storage means for storing a domain evaluation model, and a question generation means for generating a plurality of questions for each of a plurality of evaluation objects using the domain evaluation model and the question expansion table and providing the questions to the response acquisition means.

[0010] More preferably, the question expansion table includes a plurality of question templates, each of the plurality of question templates having one or more slots each assigned with one of a predetermined plurality of tags, the domain evaluation model storing a plurality of character strings each associated with one of the predetermined plurality of tags, and the question generation means including, for each of the plurality of question templates, slot insertion means for generating a plurality of questions by inserting, into each of the one or more slots of the question template, one of the plurality of character strings stored in the domain evaluation model that is associated with the tag of the slot.

[0011] More preferably, the plurality of question templates are classified into a plurality of question types, and the evaluation means includes a type-specific evaluation means for evaluating the plurality of evaluation objects according to a plurality of question types in accordance with evaluation criteria based on the responses acquired by the response acquisition means for each of the plurality of evaluation objects, and an overall evaluation means for calculating an overall evaluation of the plurality of evaluation objects using the evaluations for the plurality of question types acquired by the type-specific evaluation means.

[0012] Preferably, the type-specific evaluation means includes a summary generation means for merging responses acquired by the response acquisition means by question type for each of the plurality of evaluation objects to generate a summary for each question type, and a type-specific summary evaluation means for evaluating each of the plurality of evaluation objects based on the reliability, amount of information, and similarity between the summary and specified reference data, or any combination of these.

[0013] More preferably, the type-specific summary evaluation means includes a multiple-type summary evaluation means for evaluating each of the multiple evaluation objects based on any combination of at least two of the reliability, information volume, and similarity generated for each question type, and the evaluation device further includes a type-specific ranking calculation means for calculating a ranking of each of the multiple evaluation objects for each question type using the evaluation by the multiple-type summary evaluation means.

[0014] More preferably, the overall evaluation means includes an extended reverse rank integration means for calculating an overall evaluation by applying an extended reverse rank integration method to each rank of the plurality of evaluation objects calculated by the type-specific rank calculation means.

[0015] Preferably, the evaluation device further includes a variant storage means for storing variants of the name of each of the plurality of evaluation objects in association with the official name of the object, a variant search means for searching for variants stored in the variant storage means for each of the plurality of evaluation objects, and an additional response acquisition means for acquiring additional responses to each of the plurality of questions about the variant from the question-answering system in response to the variant search means detecting the existence of a variant, and the evaluation means includes means for evaluating the plurality of evaluation objects in accordance with evaluation criteria based on the responses acquired by the response acquisition means and the additional response acquisition means for each of the plurality of evaluation objects.

[0016] More preferably, the objects of the multiple evaluations are multiple companies, and the companies that serve as the evaluation standards include companies that are evaluated as being above a certain rank in the rankings of existing companies.

[0017] A computer program according to a second aspect of the present invention causes a computer to function as a response acquisition means for acquiring a response to each of a plurality of questions from a question-answering system for each of a plurality of evaluation objects, and an evaluation means for evaluating each of the plurality of evaluation objects in accordance with predetermined criteria established based on a predetermined viewpoint, based on the response acquired by the response acquisition means for each of the plurality of evaluation objects.

[0018] An evaluation method according to a third aspect of the present invention includes a step in which a computer acquires a response to each of a plurality of questions for each of a plurality of evaluation objects from a question-answering system, and a step in which the computer evaluates each of the plurality of evaluation objects in accordance with predetermined criteria defined based on a predetermined viewpoint, based on the responses acquired in the acquiring step for each of the plurality of evaluation objects. [Brief explanation of the drawings]

[0019] [Figure 1] FIG. 1 is a block diagram of a business activity evaluation system according to one embodiment of the present invention. [Figure 2] FIG. 2 is a schematic diagram showing an example of the contents of the question expansion table shown in FIG. [Figure 3] FIG. 3 is a schematic diagram showing an example of the contents of the domain evaluation model shown in FIG. [Figure 4] FIG. 4 is a flowchart showing a control structure of a computer program (hereinafter simply referred to as "the program") for implementing the question generation unit shown in FIG. 1 by a computer. [Figure 5] FIG. 5 is a flowchart showing a control structure of a program for realizing the abstract generation process performed for each company in the flowchart shown in FIG. [Figure 6] FIG. 6 is a flowchart showing a control structure of a program for realizing the abstract generation process in the flowchart shown in FIG. [Figure 7] FIG. 7 is a flowchart showing a control structure of a program for realizing the ranking process in the flowchart shown in FIG. [Figure 8] FIG. 8 is a block diagram showing the configuration of a coupling coefficient estimation system for estimating coupling coefficients for a multi-query ensemble used by the business activity evaluation system shown in FIG. [Figure 9] FIG. 9 is a graph showing an example of a Precision-Recall curve used to estimate the combining coefficient for a multi-query ensemble. [Figure 10] FIG. 10 is a diagram showing, in table form, the score functions used when analyzing corporate activities according to one embodiment of the present invention and the AUPR (Area Under the Precision-Recall Curve) obtained for each question type. [Figure 11] FIG. 11 is a diagram showing the appearance of a computer system for realizing the business activity evaluation system shown in FIG. [Figure 12] FIG. 12 is a hardware block diagram of the computer system shown in FIG. DETAILED DESCRIPTION OF THE INVENTION

[0020] In the following description and drawings, identical parts are assigned the same reference numbers. Therefore, detailed descriptions thereof will not be repeated. The following embodiments relate to a system for evaluating a company's DX (Digital Transformation) activities. However, the present invention is not limited to such embodiments. The present invention can be applied to any system for performing qualitative analysis of any evaluation target. Even when limited to corporate activities, the present invention can be applied to surveys related to, for example, security management, carbon neutrality, work style reform, and SDGs (Sustainable Development Goals).

[0021] 1. First embodiment 1. Configuration A. Overall configuration 1 shows a block diagram of a corporate activity evaluation system 50 according to a first embodiment of the present invention. The corporate activity evaluation system 50 is realized as a dedicated system for corporate evaluation by programming a device including a processor, appropriate input / output devices, and a storage device, such as a computer, with an appropriate program, as will be described later.

[0022] This corporate activity evaluation system 50 receives a company list 54, which is a list of company names, and uses a question-answering system 52 to evaluate the DX activities of each company listed in the company list 54, and creates and outputs a corporate ranking 56 from the perspective of DX activities.

[0023] In recent years, it has become clear that digital transformation (DX) across all socioeconomic activities is an urgent issue for Japan. In Japan, guidelines for promoting DX were published in 2018. DX is defined as "a process in which companies respond to rapid changes in the business environment, expand data and digital technology, and transform products, services, and business models based on the needs of customers and society, while also transforming operations, organizations, processes, and corporate culture and climate to establish a competitive advantage" (Ministry of Economy, Trade and Industry DX Promotion Guidelines, https: / / www.meti.go.jp / press / 2018 / 12 / 20181212004 / 20181212004.html, 2018 / 12 / 12).

[0024] Many companies are currently facing changes in their environment, putting their business continuity at risk. In this situation, the gap between companies that can adapt flexibly to environmental changes and those that cannot is widening. This situation is not limited to Japan, but is similar in many countries that have been considered developed until now. And just like with companies, the gap is widening between countries that can adapt flexibly to environmental changes and those that cannot. Therefore, it is no exaggeration to say that the success or failure of DX will determine Japan's future.

[0025] To advance DX in Japan, it is necessary to grasp the DX status of companies and other organizations in a timely and appropriate manner and widely share effective cases. To achieve this, it is necessary to conduct large-scale surveys in a timely manner and obtain results early. This first embodiment is used to evaluate companies' DX activities for such purposes.

[0026] Referring to FIG. 1, the question answering system 52 has a function of outputting text that is appropriate as a response to a question when it receives the question. In particular, the question answering system 52 has a function of outputting text that is considered to be the most appropriate response to a given question based on the text of many web pages on the Internet. The question answering system 52 also outputs a reliability score that indicates how appropriate the output response is as a response to the question. As will be described later, this score is used in evaluating DX activities. For example, if the question answering system includes a neural network, the score here is the probability that each candidate response output from the output layer indicates that it is an answer to the question.

[0027] In this embodiment, the question answering system 52 is assumed to have the function of accepting both so-called factoid questions and non-factoid questions and providing appropriate answers to each of them.

[0028] Factoid questions are questions that expect a noun as a response, such as "What does global warming cause?" Non-factoid questions include why questions such as "Why does global warming occur?", how questions such as "How do we prevent global warming?", what questions such as "What will happen if global warming progresses?", and definition questions such as "What is global warming?" that ask for a definition of something rather than just a noun.

[0029] Recently, natural language information retrieval techniques have been developed to find answers to questions from information on the Web. The volume of information on the Web is vast, and the variety of information contained therein is greater than ever. These factors have made it possible to find the right answer to a question with a high probability.

[0030] B. Company List 54 Company List 54 lists the names of the companies that are the subject of the survey.

[0031] The business activity evaluation system 50 includes a question expansion table 62 for storing a plurality of question templates, which are templates of questions to be given to the question answering system 52 .

[0032] Referring to FIG. 2, in this embodiment, each question template is classified into six question types as shown in the question expansion table 62 in FIG. 2. These six question types correspond to the so-called 5W1H. Each question template has a plurality of slots, which are positions where words can be inserted. The "slot" here indicates a position where a word can be inserted in a character string representing a question, and may be anything that can be identified by a computer. In this embodiment, a tag indicating the attribute of the word to be inserted at that position is inserted in advance into each slot. This tag indicates the slot position. In this embodiment, a tag indicating that the slot is a slot into which the subject of the question is inserted is inserted. タグ、述語が入るスロットであることを示す <pred> タグ、及び述語の目的語が入るスロットであることを示す <obj> タグを使用する。

[0033] C.ドメイン評価モデル60 図3を参照して、ドメイン評価モデル60には質問テンプレートにおいて使用されている3種類のタグに応じて3種類の文字列が格納されている。

[0034] a. <obj> タグ <obj> タグを持つスロット( <obj> スロット)に挿入される文字列としてドメイン評価モデル60に記憶されているのは、この実施形態においては「デジタルトランスフォーメーション」(又はこれとあわせて「DX」)である。 <obj> スロットに挿入される文字列は、質問テンプレートにより生成される質問において、 <pred> タグが示すスロット(「 <pred> スロット」に挿入される述語の目的語となる。

[0035] b. <pred> タグ <pred> スロットに挿入される文字列としてドメイン評価モデル60に記憶されているのは、「した」「成し遂げた」「達成した」などである。これら文字列は、 <obj> スロットに挿入された単語とともに、DX活動として高く評価されていることを示す文章(質問文)を生成するためのものである。

[0036] c. タグタグを持つスロット(スロット)に挿入される文字列は、質問の主語となる文字列である。この実施形態においては、企業リスト54に記載された各企業の名称がスロットに挿入される。それらの企業名は企業リスト54に記載されているのでドメイン評価モデル60には格納されない。

[0037] しかし、企業の中には正式名称に加えて異表記を持つものがある。ここでいう異表記とは、同じ企業を指す文字列であって互いに異なる文字からなるものをいう。例えば「株式会社」などという表記は省かれることがある。「株式会社」という文字列に代えて「(株)」という省略形が用いられることもある。さらに、社名の一部が省略されたり短縮されたりした文字列が使用されることもある。カタカナ表記の社名の場合には、「ソフトウェア」と「ソフトウエア」、「ウェブ」と「ウエブ」のように文字の大きさが異なる場合もある。アルファベットの社名を持つ企業の場合には、大文字と小文字の違いがある場合もあるし、一部又は全部がカタカナ表示される場合もある。ときには文字列中に「·」が入っていたり入っていなかったりする場合もある。特に質問応答システム52が質問応答のために情報を収集する場であるウェブにおいては、企業の正式名称ではなく様々な異表記が使用されることが想定される。企業のDX活動を適切に評価するためには、そのような異表記とともに記載された情報であっても収集する必要がある。

[0038] そうした理由から、ドメイン評価モデル60は、スロットに入力される文字列として、各企業の正式名称にそれぞれ関連付けられた各企業の異表記のリストを記憶する。したがって、図3に示す例においては、例えば正式名称が「企業A」である場合に、その異表記を探索し、異表記A1、異表記A2などが検出されたことに応答して、異表記A1を含む記事、異表記A2を含む記事なども全て応答探索の対象にできる。

[0039] なお、企業リスト54に正式名称によらず異表記により記載された企業がある場合でも、このドメイン評価モデル60を用いてその正式名称を調べたり、別の異表記を調べたりできる。さらに、異表記を記憶していなくても、ある規則に従って異表記を生成することもできる。それらの方法については自明であるためここではその詳細については記載しない。本実施形態においては、後述するように企業の正式名からその異表記を探索するためにこのドメイン評価モデルを使用する。

[0040] D.質問生成部64質問生成部64は、質問展開テーブル62に格納された各質問テンプレートと、ドメイン評価モデル60に格納された各文字列、及び企業リスト54に記載された企業名及びその異表記の任意の組み合わせから質問を生成するためのものである。質問生成部64の詳細については図5を参照して後述する。

[0041] E.DX優良事例66DX優良事例66は、2015年―2020年のDX銘柄、DX注目企業、攻めのIT経営銘柄に関する経済産業省発行のDX関連レポートを基に作成した。

[0042] F.学習データ70学習データ70は、後述する評価部68が使用するマルチクエリアンサンブルのための結合係数の推定に用いるためのものである。学習データ70についてもDX優良事例66と同じレポートを基に、企業活動評価システム50が評価した結果とDX関連レポートによる評価とを用いて作成した。

[0043] G.評価部68評価部68は、企業ごとに質問生成部64が作成した質問群に対して質問応答システム52から得られた応答から各企業のDX活動に関するスコアを算出し、ランキングして出力するためのものである。

[0044] より具体的には、評価部68は、企業ごとに質問生成部64が作成した各質問タイプの複数の質問群について、摘要を生成する。

[0045] この摘要を得るために評価部68は、まず同一の質問タイプの質問に対する質問応答システム52からの応答から以下の3つ組からなる応答要素を抽出する。各応答要素は、(1)質問応答システム52が応答の情報源として用いたウェブページのスニペット、(2)そのURL(Uniform Resource Locator)、及び(3)その応答要素に対して質問応答システム52が推定する信頼度、の3つ組である。

[0046] 評価部68はさらに、これら応答要素から重複したものを削除して重複のない応答要素の集合を作成する。この作業を応答要素の「マージ」と呼び、この応答要素の集合をその企業に関するDX摘要と呼ぶ。

[0047] この方法により、1つの企業に対し、6つの質問タイプの各々について1つずつのDX摘要が得られる。すなわち、1つの企業に対して質問タイプに対応した6種類のDX摘要が生成される。

[0048] 評価部68はさらに、質問タイプごとに所定のスコア関数により企業ごとのスコアを算出する。評価部68はさらに、このスコアを用いて質問タイプごとに企業をランキングする。この結果、各企業には、質問タイプの種類に対応した6種類の順位が付される。評価部68は、最後に、各企業の6種類の順位を後述する拡張版逆順位統合法により統合して総合スコアを算出する。評価部68は企業をこの総合スコアに基づいてランキングし企業ランキング56として出力する。

[0049] 評価部68を実現するプログラムの構成については後述する。

[0050] H.プログラム構成H1.全体のプログラム構成図4を参照して、企業活動評価システム50を実現するプログラムは以下のような制御構造を有する。このプログラムは、記憶領域の確保及び初期値による初期化、ファイルオープン又はデータベース(以下「DB(Database)」という。)への接続などの初期設定を行うステップ100を含む。

[0051] このプログラムはさらに、ステップ100に続いて図1に示す企業リスト54を読んでメモリに格納するステップ102を含む。この処理においては、このプログラムは、例えば企業リスト54がファイル形式であれば外部記憶装置からファイルを先頭から最後まで読み出してメモリに格納する。企業リスト54がDBのテーブルに格納されていれば、このプログラムが適切なクエリをDBに発行することにより、企業名の配列をDBから受け取ってメモリに格納する。

[0052] このプログラムはさらに、ステップ102に続き、ステップ102において読み込んだ各企業の名称に対して企業ごとの処理を実行して、ランキングに必要な情報である、前述したDX摘要を企業ごと及び質問タイプごとに作成するステップ106を実行するステップ104を含む。ステップ106の処理の詳細については図5を参照して後述する。

[0053] このプログラムはさらに、ステップ104に続き、ステップ104において企業ごと及び質問タイプごとに作成されたDX摘要を用いて各企業をランキングし企業ランキング56を出力するステップ108を含む。

[0054] H2.企業ごとの処理(ステップ106)図5を参照して、図4に示すステップ106は、質問展開テーブル62から質問テンプレートを読み出すステップ150と、処理対象の企業名の異表記を図1に示すドメイン評価モデル60の内部において探索するステップ152とを含む。

[0055] ステップ106はさらに、処理対象の正式企業名について、その企業のDX活動に関する情報を質問応答システム52により収集するための質問群を作成するステップ156を実行するステップ154を含む。ステップ154においては、ステップ152において企業名の異表記が探索されたことに応答して、探索されたその異表記についてもその企業のDX活動に関する情報を質問応答システム52により収集するための質問群を追加する。

[0056] ステップ156は、ドメイン評価モデル60に記憶されている <obj> タグに対応する文字列の各々と、 <pred> タグに対応する文字列の各々との組み合わせ(対)の各々に基づいて、質問テンプレートの各スロットに、そのスロットのタグに対応する文字列を挿入することにより質問を生成するステップ172を実行するステップ170を含む。

[0057] ステップ172は、質問展開テーブル62に記憶されている質問テンプレートの各々について、 <obj> スロットに処理対象の <obj> タグに対応する文字列を挿入し、 <pred> スロットに処理対象の <pred> タグに対応する文字列を挿入し、 スロットには処理対象の企業の正式名又は異表記を挿入することにより質問を生成するステップ180を含む。ステップ172はさらに、ステップ180において生成された質問を正式企業名と関連付けて記憶装置に記憶してステップ172を終了するステップ182を含む。

[0058] 以上のようにステップ154により、各企業についての質問群が作成され記憶装置に記憶される。これら各企業についての質問群は、その企業の正式名称を含む質問だけではなく、もしもその企業の名称に異表記があれば、その企業の異表記を含む質問も含む。

[0059] ステップ106はさらに、ステップ154において作成された質問を質問応答システム52に与えて各質問に対する応答を得ることにより、企業ごと、及び質問タイプごとにDX摘要を生成しステップ106を終了するステップ158を含む。

[0060] H3.摘要生成処理(ステップ158)図6を参照して、ステップ158は、DX摘要を生成するステップ192を各質問タイプについて実行するステップ190を含む。

[0061] ステップ192は、各質問について以下のステップ202を実行して質問応答システム52から応答を取得するステップ200を含む。

[0062] ステップ200は、処理対象の質問を質問応答システム52に入力するステップ230と、ステップ230において質問応答システム52に与えられた質問に対して質問応答システム52が出力する応答を取得し記憶するステップ232とを含む。

[0063] ステップ192はさらに、ステップ200において各質問に対して取得された応答をマージしてDX摘要を生成するステップ204と、ステップ204において生成されたDX摘要を記憶装置に記憶してステップ192を終了するステップ206とを含む。

[0064] H4.ランキング(ステップ108)図7を参照して、図4に示すステップ108を実現するプログラムは、各質問タイプについて企業別にランキングするステップ252を実行するステップ250と、ステップ250によるランキングにより各企業について質問タイプごとに得られた順位から、各企業の総合スコアを算出するステップ254と、ステップ254において算出された総合スコアにより各社をランキングするステップ256と、ステップ256によるランキング結果を出力してランキングのステップ108を終了するステップ258とを含む。

[0065] ステップ252は、処理中の質問タイプに関するスコアを算出するステップ282を各企業について実行するステップ280と、ステップ280により各企業について処理中の質問タイプに関して得られたスコアに基づいて各企業をランキングしてステップ252を終了するステップ284とを含む。

[0066] ステップ282は、処理中の質問タイプに関して、処理中の企業について得られたDX摘要を記憶装置から読み出すステップ310と、ステップ310において読み出されたDX摘要に対して後述する8種類のスコアのいずれか又はこれらの任意の組み合わせによるスコア計算を行い、その結果を保存してステップ282を終了するステップ312とを含む。

[0067] I.スコアリングI1.DX摘要の生成図1に示す質問生成部64は、各企業に対して、質問タイプごとに複数の質問文を生成する。その結果、質問応答システム52からは、同一の企業、同一の質問タイプに関して複数の応答が得られる。

[0068] この応答の集合は、ドメイン評価モデル60に登録された異表記及び同義語に由来するものを含む。そのため、前述したように、まず、各応答から(1)質問応答システムが応答の情報源として用いたWebページのスニペット、(2)URL、(3)各応答要素に対して質問応答システム52が推定する信頼度、の3つ組である応答要素を抽出する。続いて、抽出された全ての応答要素から、重複のない応答要素の集合を生成する。この集合がすなわちDX摘要を形成する。この方法により、異表記や同義語に由来する複数の応答は1つのDX摘要に統合される。その結果、企業ごとに、質問タイプに対応し6種類のDX摘要が生成される。

[0069] I2.スコア関数の定式化この実施形態に係る企業活動評価システム50は、企業のDXに対する取り組みの優良さ及び活動の活発さなどを、質問応答システム52から得られる情報の「掲載情報量」、DX優良事例との「類似度」、情報の「信頼度」の3要素を軸にスコアリングする。以降、このスコアリングの定式化を説明する。なお、以下においては8種類のスコアリング方法を説明するが、後述するように上記実施形態においてはいずれのスコアリング方法を用いてもよい。

[0070] また「情報量」については情報理論におけるような厳密な定義もあり得る。しかしここでは、以下に述べるように簡易な指標を情報量とする。

[0071] まず、スコアリングにおける各入力要素の表記を導入する。以降では、質問タイプtに対するDX摘要をDt、さらにDtの要素をdtにより表す。DX摘要の要素dtは、応答要素由来のスニペット、URL、信頼度の組を持つ。このうち要素dtの信頼度をconf(dt)、要素dtのスニペットに含まれる単語集合の要素をwt∈dtにより表す。また、DX優良事例のテキストをdh、同テキストに含まれる単語集合を{wh}により表す。

[0072] この実施形態においては、スコアリングを構成する3要素のうち、「信頼度」については、要素dtについて質問応答システム52から出力される信頼度conf(dt)、「掲載情報量」についてはDX摘要Dtに含まれる要素dtの要素数を用いることができる。一方、「類似度」については、DX摘要Dtに含まれる要素dtと、DX優良事例のテキストdhの間の類似度を用いることができる。ここでは、DX摘要Dtに含まれる単語wtとDX優良事例のテキストdhに含まれる単語集合{wh}の間に定義される次式のような語彙的類似度sim(wt,{wh})を「類似度」の具体的な実装として用いる。

[0073]

[0074] ここまでの議論により導入された「掲載情報量」、「類似度」、「信頼度」の3要素を軸とするスコアリングの基本設計に沿って、DX摘要Dtに対するスコア関数Score(Dt)の定式化を行う。まず、「掲載情報量」の1軸のみを用いて、次式のようなスコア関数を定式化できる。

[0075] 「掲載情報量」と「信頼度」の2軸からは、次式のような定式化が可能である。

[0076] 「掲載情報量」と「類似度」の2軸からは、まず語彙的類似度sim(wt,{wh})のみを「類似度」の評価値として用いれば、次式のような定式化が可能である。

[0077] 「類似度」の評価に、語彙的類似度sim(wt,{wh})を単語の重みidf(wt)により補正した定式化を用いれば、次式のようなスコア関数が得られる。

[0078] 単語頻度tf(w)も併用すれば、次式のようなスコア関数となる。

[0079] 「掲載情報量」、「類似度」、「信頼度」の3軸を用いる場合、語彙的類似度sim(wt,{wh})を「類似度」の評価に用いれば、次式のスコア関数が得られる。

[0080] 「類似度」の評価値に単語の重みidf(wt)により補正した語彙的類似度を用いた場合、次のスコア関数を得る。

[0081] 「類似度」に、さらに単語頻度tf(w)の補正を加える場合は次式となる。

[0082] この実施形態においては、以上によって導入された計8つのスコア関数のいずれかをDX摘要のスコアリングに用いる。これら8つのスコア関数の任意の組み合わせを用いてもよい。例えばこれら8つのスコア関数の任意の組み合わせの加重平均などを用いることも考えられる。

[0083] I3.マルチクエリアンサンブルこの実施形態は、6種類の質問タイプに関して、それぞれ6つのDX摘要を生成し、それぞれに対してスコアリングを行う。このため、これらスコアを用いて企業のランキングを得るためには、6種類の質問タイプに対して得られた6種類のスコアを1つに統合する必要がある。

[0084] ここでは、このように複数のランキング結果を統合する手法として教師無しの統合手法であるRRF(Reciprocal Rank Fusion)を使用する。複数のランキング結果を統合する手法としては、これ以外にもCondorcet Fuse、CombMNZなどがある。

[0085] RRFは、順位の逆数(Reciprocal Rank、ただし定数補正項付き)を重みなしで足し合わせた単純な定式化である。それにもかかわらず、RRFは、NIST TRECにおける複数の関連文書ランキングの統合では、標準的な統合手法であるCondorcetや他の学習ベースの統合手法と比較しても高い性能を示すことが報告されている。この実施形態が扱う6つのDX摘要に対するスコアの統合でもRRFを適用できる。ただし、6種類の質問タイプの中で、どの質問タイプが有効かを事前に知ることはできない。そのため、この実施形態ではRRFに結合パラメータを組み込んだ拡張版RRFを導入する。この統合手法は、以下の手順によってランキングの統合と、結合係数の推定を行う。なお、以下では拡張版RRFを拡張版逆順位統合法と呼ぶ。

[0086] 第1ステップ6種類の質問タイプに対してそれぞれ得られた各DX摘要Dtに対し、前節のスコア関数を用いてスコアを求める。

[0087] 第2ステップ第1ステップにおいて求められたDX摘要のスコアを、質問文の種類が同じグループに分けてソートする。この結果、質問文の種類ごとに各企業の企業集合における順位rank(Dt)が求められる。

[0088] 第3ステップ質問文の種類ごとに求めた企業の順位rank(Dt)から総合スコアScoreensembleを次式により求める。

[0089] ここで、{t}は6種の質問タイプの集合、rank(Dt)-1はある質問タイプtに関するDX摘要のスコアから生成される企業のランキングの逆数(Reciprocal Rank)、^ctは質問タイプtの結合係数を表す。すなわち、この実施形態においては、ある企業の総合スコアは、その企業の各質問タイプに関するランキング(順位)の逆数にそれぞれの結合係数を乗算して和をとることにより算出される。結合係数は、ランキングの評価尺度を目的関数として、これを最大化するように推定する。AUPRをランキングの評価尺度とする場合には、以下の式を用いて結合係数が推定される。

[0090] ただしytrueは、学習データ中の二値分類の教師ラベルを表す。この実施形態においては、2020年のDX銘柄、注目企業、及び、2015-2019年の攻めのIT経営銘柄を正例、それ以外の銘柄を負例とする教師データを用い、グリッドサーチにより結合係数を推定する。つまり、既存のランキングにおいて一定の順位より上にランクされた企業を正例、それ以外を負例とする。ただしここでいう一定の順位とは、全てのランキングにおいて同一の順位というわけではなく、各ランキングにおいて優れている企業の最低順位とされた順位のことをいう。すなわちこの実施形態においては、既存のランキングが各企業の最終的なランキングのための基準として使用される。なお、「既存のランキング」としては、その調査方法及び評価方法などが客観的であると考えられるもの、及び評価対象が今回の発明を適用するものと共通するものであることが望ましい。既存のランキングとしては特に大規模であることは必須ではない。しかし、客観的な基準で行われた大規模な調査結果に基づくランキングであればより好ましい。

[0091] I4.結合係数の推定図8を参照して、この結合係数を推定するための構成についてより具体的に説明する。結合係数を推定するためのシステム340は、2020年のDX銘柄及び注目企業に関するリストを記憶するための記憶装置350と、2015年から2019年の攻めのIT経営銘柄に関するリストを記憶するための記憶装置352と、それ以外の銘柄に関するリストを記憶するための記憶装置354とを含む。記憶装置350及び記憶装置352に記憶されたリストの銘柄は図1に示すDX優良事例66に相当し、結合係数の推定処理における正例となる。記憶装置354に記憶されたリストの銘柄は負例となる。

[0092] システム340はさらに、記憶装置350、352及び354に記憶されたリストに基づいて、質問応答システム52を使用して学習データを生成するための学習データ生成部356と、学習データ生成部356により生成された学習データを記憶するための学習データ記憶装置358とを含む。

[0093] 学習データ生成部356は、実質的には企業活動評価システム50と同じ構成である。ただし、結合係数を用いた総合ランキングの算出は行わない。すなわち、学習データ生成部356は、企業活動評価システム50とほぼ同じプログラムにより実現される。ただし学習データ生成部356は、図7に示すステップ254、256及び258を行わない点において企業活動評価システム50と異なる。学習データ記憶装置358に記憶される学習データは、各企業について、上記した6つの質問タイプ別に算出された順位であるランキング結果と、その企業が正例か負例かを示す教師ラベルとを含む。

[0094] システム340はさらに、学習データ記憶装置358に記憶された各企業の質問タイプ別のランキング結果と各企業に関する教師ラベルとに基づいて、ランキングの評価尺度を目的関数とするグリッドサーチにより、この目的関数の値を最大化するような結合係数の組を推定するための結合係数推定処理部362を含む。この実施形態においては、ランキングの評価尺度として上記したようにAUPRを用いる。したがってシステム340はさらに、結合係数推定処理部362の指示に従って、学習データ記憶装置358に記憶された学習データと結合係数の組の候補とを用いて、候補ごとにAUPRを算出する処理を行うためのAUPR算出部364を含む。

[0095] 図9にAUPRの算出に使用されるPrecision-Recall曲線400の1例を示す。AUPRとは、Precision-Recall曲線下の面積を計算したものである。一般的に、分類システムの適合率と再現率はトレードオフの関係にある。そのため、Precision-Recall曲線は図9にも示されるように右下がりの曲線を描く。AUPRが高いほどシステムの予測精度は高い。グリッドサーチにおいては、結合係数の組み合わせを様々に変化させてAUPRを算出し、最も高いAUPRが得られた結合係数の組み合わせを採用する。

[0096] 2.動作上記した企業活動評価システム50は以下のように動作する。まず質問別のランキングを生成するところまでの企業活動評価システム50の動作について説明する。その後、教師データを用いて結合係数を推定する際のシステム340の動作について説明する。最後に、システム340により決定された結合係数を用いて企業の総合ランキングを生成する際の企業活動評価システム50の動作を説明する。

[0097] A.質問別のランキング生成A1.質問の生成図1に示す質問生成部64は、まず記憶領域の確保及び初期値による初期化、ファイルオープン又はDBへの接続などの初期設定を行う(図4のステップ100)。質問生成部64はさらに、図1に示す企業リスト54を読んでメモリに格納する(ステップ102)。

[0098] 続いて質問生成部64は、読み込んだ各企業に対して図4に示すステップ106において各企業についての処理を実行する。より具体的には、図5を参照して、質問生成部64は各企業に対して、図1に示す質問展開テーブル62を読む(ステップ150)。質問生成部64は、ドメイン評価モデル60を探索して処理対象の企業名の異表記を取り出す(ステップ152)。さらに質問生成部64は、企業名とその異表記との各々に対してステップ156を繰り返す(ステップ154)。

[0099] ステップ156においては質問生成部64は、ドメイン評価モデル60に格納されている <obj> タグに対応する文字列と <pred> タグに対応する文字列との組み合わせの各々について、質問展開テーブルに記憶されている各質問テンプレートに挿入することにより質問を生成する(ステップ180)。このとき、 タグには企業名及び探索されたその異表記との各々が挿入される。さらに質問生成部64は、生成された質問を正式企業名と関連付けて記憶装置に記憶する(ステップ182)。

[0100] こうした処理を全ての質問テンプレートについて実行することにより、処理対象である企業の企業名がスロットに挿入された質問と、その企業の異表記が挿入された質問とが生成される。各質問と <obj> タグに挿入される文字列及び <pred> タグに挿入される文字列は、この実施形態においては、企業に関するDXの肯定的な評価を含む文章が検索されるように予め選ばれている。

[0101] 質問は予め質問タイプ別に分類されている。したがって、図5に示すステップ106の処理が実行されることにより、処理対象の企業について、ランキングに必要な情報であるDX摘要を質問タイプごとに作成するために必要な質問が生成される。

[0102] A2.DX摘要の生成 図5を参照して、ステップ154までの処理により処理対象の企業に関する質問が生成されると、図1に示す評価部68が以下のようにしてその企業に関するDX摘要を生成する。

[0103] より具体的には、図6を参照して、評価部68は、各質問タイプについて処理対象の企業に関するDX摘要を生成する処理を実行する(ステップ190)。

[0104] この処理において評価部68は、まず処理中の質問タイプの各質問を質問応答システム52に入力し(ステップ230)、質問応答システム52からの応答結果(回答)を取得して記憶装置に記憶する(ステップ232)。この処理により、処理対象の企業について、処理中の質問タイプの質問の各々に対する質問応答システム52の応答が集積される。

[0105] 続いて評価部68は、このようにして処理中の質問タイプについて集積された応答をマージしてDX摘要を生成し(ステップ204)、記憶装置に記憶する(ステップ206)。

[0106] 以上説明した処理を各質問タイプに対して実行することにより、6つの質問タイプの全てについてタイプ別のDX摘要が生成される。

[0107] A3.質問タイプ別のランキング 図7を参照して、図1に示す評価部68は以下のようにして質問タイプ別の企業のランキングを生成する。すなわち、各質問タイプについて、評価部68は、その質問タイプに関する各企業のDX摘要を読み(ステップ310)、その質問タイプのDX摘要に関するスコア計算を行って記憶装置に保存する(ステップ312)。この処理を各企業について実行することにより、質問タイプごとに、企業の評価を示すスコアが算出される。評価部68はさらに、このように質問タイプごとに算出されたスコアを用いて、質問タイプごとに企業をランキングする(ステップ284)。ここで使用されるスコアは、前述した8種類のスコアのいずれでもよい。

[0108] この後、評価部68はステップ254、256及び258を実行して企業のランキングを生成する。しかし、そのためには予めステップ254において用いられるRRFのための結合係数を算出しておく必要がある。そこで、ステップ254以降における評価部68の動作を説明する前に結合係数の推定のためのシステム340の動作を説明する。

[0109] B.結合係数の推定 結合係数の推定処理は大きく分けて学習データの生成と、グリッドサーチによる結合係数の生成とに分けられる。

[0110] B1.学習データの生成 学習データの生成時には、図8に示すシステム340の学習データ生成部356は以下のようにして学習データを生成する。まず学習データ生成部356は記憶装置350及び記憶装置352に記憶された銘柄を読み、これら優良銘柄の各々について上記した方法を用いて質問タイプ別のDX摘要を生成する。学習データ生成部356はさらに、これらDX摘要の各々についてスコアリングを行う。この際、学習データ生成部356は、各企業のスコアにその企業が正例であることを示す教師ラベルを付す。学習データ生成部356は同様に、記憶装置354に記憶された、優良銘柄以外の企業のリストを読み、各企業についてDX摘要を生成する。学習データ生成部356はさらに、これらDX摘要の各々についてスコアリングを行う。この際、学習データ生成部356は、各企業のスコアにその企業が負例であることを示す教師ラベルを付す。

[0111] このようにして、各企業についてスコアが算出されたことにより、これを学習データとしてグリッドサーチによる結合係数を推定できるようになる。

[0112] B2.グリッドサーチによる結合係数の推定 結合係数推定処理部362は、各企業について質問タイプ別のスコアに対するグリッドサーチにより最も高いAUPRを与える結合係数の組み合わせを探索する。

[0113] より具体的には結合係数推定処理部362は、予め定められた探索方法により結合係数の組み合わせを決定する。結合係数推定処理部362は、各企業について計算された質問タイプ別のスコアに結合係数を乗じて和をとることにより、その企業の総合スコアを算出する。結合係数推定処理部362は、こうして算出された総合スコアに基づいて企業をランキングする。AUPR算出部364がそのランキング結果と教師データとを比較してその結合係数の組み合わせに対するAUPRを算出する。結合係数推定処理部362はその値を記憶する。

[0114] 結合係数推定処理部362がこの処理を結合係数の全ての組み合わせに対して算出する。結合係数推定処理部362は、記憶されたAUPRのうちで最も高いものを特定し、そのAUPRが得られたときの結合係数の組み合わせを最終的な結合係数に決定し記憶装置に記憶する。このように結合係数が決定されると、企業活動評価システム50は各企業に関する総合スコアを算出できるようになる。

[0115] C.総合ランキングの生成 図7を参照して、企業活動評価システム50の評価部68(図1を参照)は、総合ランキングのための結合係数の組み合わせを記憶装置から読み出す。評価部68はさらに、各企業の質問タイプ別のスコアにそれぞれの結合係数を乗算してその和を算出することにより各企業の総合スコアを算出する(ステップ254)。評価部68はこの総合スコアを用いて企業をランキングする(ステップ256)。評価部68は最終的にそのランキングを出力する(ステップ258)。

[0116] 3.実施形態の効果 以上のようにこの実施形態によれば、短時間しか要さずに各企業のDX活動の評価をランキングの形により生成できる。アンケートを作成する必要も、アンケートを評価する必要もない。質問テンプレートは基本的な5W1Hに従って生成するだけであるため、特に難しいこともない。評価に複数の人手を要することもないため、従来のように評価を行う人によって結果にばらつきが生ずることもない。

[0117] 以下に説明する実験によれば、このようにして作成した企業のDX活動のランキングは、一般的な検索エンジンによる検索結果を用いるアプローチよりも精度がよいことが分かる。この一因は、実験に用いた質問応答システムが必要な情報を精緻に収集できることにもよると思われる。しかし、こうした効果は、実験に用いたものと異なる質問応答システムであっても一定程度は得られると考えられる。質問応答システムは、一般的な検索エンジンによる検索結果と比較して、質問に対する応答として適切なものを出力することが期待できるためである。もちろん、そのためには質問応答システムに適切な質問を与えることが必要である。そうした観点から、上記実施形態のドメイン評価モデル60と質問展開テーブル62とを用いて質問文を生成する手法も適切と考えられる。質問展開テーブル62に格納される質問テンプレートも適切なものにする必要があるが、上記実施形態のように基本的なテンプレートでも十分に有効と考えられる。

[0118] 第2 実験 1.実験方法 上記実施形態に係る企業活動評価システム50により、DX先進企業をどの程度の精度により判定可能かを確認するための実験を行った。実験においては、客観的な観点から大規模な企業の分析と評価の自動化を実現する見通しを調べることも目的の一つとした。評価対象はDX銘柄2021にエントリした企業464社(33業種)であった。評価対象の464社に含まれる正例は、DX銘柄2021(28社)とDX注目企業2021(20社)の合計48社であった。実験においてはこの約10%の正例を企業活動評価システム50によるランキングによってどの程度識別可能かを評価した。

[0119] なお、企業活動評価システム50においてはどのような質問応答システムを使用するかによって結果も異なってくると考えられる。以下に述べる実験においては、出願人が提供している質問応答サービス(https: / / www.wisdom-nict.jp / )において使用されているものと同等の質問応答システムを使用した。この質問応答システムは、Web60億ページから抽出した情報を基にして、ファクトイド型、なぜ型、どうやって型、どうなる型、定義型といった多様なタイプの質問に応答できる。この質問応答システムは2015年より公開されていたが、近年の改良により精度が向上し、応答できる質問のバリエーションも増大した。またこの質問応答サービスは高速かつ効率的に質問を処理できる。特にこの質問応答システムは、質問に対して名詞句、動詞句、文など必要な範囲に応じピンポイントな応答を出力できる。

[0120] 企業活動評価システム50のスコアリングを実現するDX優良事例、及び学習データは、2015-2020年のDX銘柄、DX注目企業、攻めのIT経営銘柄に関する経済産業省発行のDX関連レポートを基に作成した。専門家による実際のDX先進企業の選定においては、464社が提出した選択式項目の回答及び3年平均ROE(Return On Equity:自己資本利益率)が選考の中において考慮されている。しかし、この実験においては、あくまでも企業活動評価システム50の性能と課題の把握に主眼を置いているため、これらの情報は用いなかった。すなわち、この実験においては、DX関連レポート、及び、質問応答システムを通じて得られるWebページのテキスト情報のみを用いて評価を実施した。企業活動評価システム50は概ね実施形態に記載した手法を用いて464社のランキングを生成した。

[0121] さらに、過去年度のDX銘柄、注目企業、攻めのIT経営銘柄の分析の結果、専門家による企業の選定において、特定業種に選定企業が集中するのを避けるため、業種毎の選定数が一定数を超えない様な調整が行われていることが推察された。そこで、この実験においては、同一業種内の企業の順位r segに対する、以下の業種内ランキングに関するヒンジ関数型のコストを導入してランキングの調整を行った。

[0122] ここでr segは同一業種内の企業の順位である。またN=464であり、n maxとaはコスト関数のパラメータである。実験においては、学習データによる予備実験においてAUPRが最大となったn max=3、a=0.5を用いた。

[0123] また、上記企業活動評価システム50の有効性を検証するために、代表的なウェブ検索サービスを用いたベースライン評価を作成した。具体的には、上記464社について、企業名と"デジタルトランスフォーメーション”の2つのキーワードのAND検索をウェブ検索サービスを用いて実行し、検索件数の大きい順に464社のランキングを生成した。評価尺度としては、正例が少ないインバランスデータであることを考慮し、ブレークイーブンポイントにおける適合率(上位48件に正例が含まれる件数の割合)と、AUPRとを用いた。

[0124] 2.評価結果 評価対象である464社に対して、企業活動評価システム50、及び代表的ウェブ検索サービスによるベースラインを用いてランキングを生成した。企業活動評価システム50に関しては、上記した6種類の質問タイプを用いて、同じく上記した8種類のスコア関数を用いてそれぞれ評価を行った。企業活動評価システム50によって生成されたランキングのAUPRを図10に表形式により示す。

[0125] 図10を参照して、いずれのスコア関数を用いた場合も、6種類の質問タイプをそれぞれ独立に用いたときよりも、マルチクエリアンサンブルを用いたときの方がAUPRの値は高かった。スコア関数による結果の相違としては、マルチクエリアンサンブルの結果においては、sim_confを用いた場合のAUPRが0.543と最も高い値となった。同じスコア関数sim_confを用いてマルチクエリアンサンブルを適用した場合のPrecision-Recall曲線が図9に示したものである。図9において企業活動評価システム50によるランキングのブレークイーブンポイントにおける適合率(=再現率)は56.3%(上位48件中27件が正解)であった。一方、代表的ウェブ検索サービスによるベースラインのAUPRは0.202、ブレークイーブンポイントにおける適合率は22.9%(上位48件中11件が正解)であった。その結果、AUPRの値、ブレークイーブンポイントにおける適合率、上位48件中に入る正解企業名のいずれにおいても、企業活動評価システム50による結果がベースラインを大幅に上回った。

[0126] 3.効果 以上のように、企業活動評価システム50によれば、代表的なウェブ検索サービスを用いて収集されたデータに基づいて企業のDX活動を評価しランキングした場合よりも、専門家が大規模に行った本格的な調査の結果に近い結果を得ることができる。企業活動評価システム50においては、アンケートを作る必要も、アンケートに回答する必要もない。アンケートを人手により集計する必要もない。その結果、企業活動の評価などの大規模な調査を、短時間しか要さず、低コストに、高い精度をもって行えるという効果がある。

[0127] ただしこの企業活動評価システム50は、マルチクエリアンサンブルにより最終的な評価をする。そのためには結合係数の推定をする必要がある。しかし、上記実験においても利用したように、先行した同種の調査結果があればそれを利用できる。また先行する調査結果において結合係数の推定をした場合、その結合係数を他の調査にも転用可能である。

[0128] 第3 コンピュータによる実現 図11は、上記各実施形態を実現するコンピュータシステムの1例の外観図である。図12は、図11に示すコンピュータシステムのハードウェア構成の1例を示すハードウェアブロック図である。

[0129] 図11を参照して、このコンピュータシステム950は、DVD(Digital Versatile Disc)ドライブ1002を有するコンピュータ970と、いずれもコンピュータ970に接続された、ユーザと対話するためのキーボード974、マウス976、及びモニタ972とを含む。もちろんこれはユーザ対話のための構成の一例であって、ユーザ対話に利用できる一般のハードウェア及びソフトウェア(例えばタッチパネル、音声入力、ポインティングデバイス一般)であればどのようなものも利用できる。

[0130] 図12を参照して、コンピュータ970は、DVDドライブ1002に加えて、CPU(Central Processing Unit)990と、GPU(Graphics Processing Unit)992と、CPU990、GPU992、DVDドライブ1002に接続されたバス1010と、バス1010に接続され、コンピュータ970のブートアッププログラムなどを記憶するROM(Read-Only Memory)996と、バス1010に接続され、上記実施形態に係る企業活動評価システム50などを実現するプログラムを構成する命令、システムプログラム、学習データ、及び作業データなどを記憶するRAM(Random Access Memory)998と、バス1010に接続された不揮発性メモリであるHDD(Hard Disk Drive)1000とを含む。HDD1000は、CPU990及びGPU992が実行するプログラム、並びにCPU990及びGPU992が実行するプログラムが使用するデータなどを記憶するためのものである。コンピュータ970はさらに、他端末との通信を可能とするネットワーク986への接続を提供するネットワークI / F(Interface)1008と、USB(Universal Serial Bus)メモリ984が着脱可能で、USBメモリ984とコンピュータ970内の各部との通信を提供するUSBポート1006とを含む。

[0131] コンピュータ970はさらに、マイクロフォン982及びスピーカ980とバス1010とに接続され、CPU990により生成されRAM998又はHDD1000に保存された音声信号をCPU990の指示に従って読み出し、アナログ変換及び増幅処理をしてスピーカ980を駆動したり、マイクロフォン982からのアナログの音声信号をデジタル化し、RAM998又はHDD1000の、CPU990により指定される任意のアドレスに保存したりするための音声I / F1004を含む。

[0132] 上記実施形態においては、図1に示す企業活動評価システム50、企業リスト54、ドメイン評価モデル60、質問展開テーブル62、DX優良事例66、学習データ70などのデータ及びパラメータなどは、いずれも例えば図12に示すHDD1000、RAM998、DVD978又はUSBメモリ984、若しくはネットワークI / F1008及びネットワーク986を介して接続された図示しない外部装置の記憶媒体などに格納される。典型的には、これらのデータ及びパラメータなどは、例えば外部からHDD1000に書込まれコンピュータ970の実行時にはRAM998にロードされる。

[0133] このコンピュータシステムを図1に示す企業活動評価システム50及びその各部、並びに図8に示す学習のためのシステム340、及びその各構成要素の機能を実現するよう動作させるためのコンピュータプログラムは、DVDドライブ1002に装着されるDVD978に記憶され、DVDドライブ1002からHDD1000に転送される。又は、これらのプログラムはUSBメモリ984に記憶される。このUSBメモリ984をUSBポート1006に装着し、これらプログラムをハードディスク1000に転送できる。又は、これらのプログラムはネットワーク986を通じてコンピュータ970に送信されHDD1000に記憶されてもよい。プログラムは実行のときにRAM998にロードされる。もちろん、キーボード974、モニタ972及びマウス976を用いてソースプログラムを入力し、コンパイルした後のオブジェクトプログラムをHDD1000に格納してもよい。スクリプト言語の場合には、キーボード974などを用いて入力したスクリプトをHDD1000に格納してもよい。仮想マシンにおいて動作するプログラムの場合には、仮想マシンとして機能するプログラムを予めコンピュータ970にインストールしておく必要がある。

[0134] CPU990は、その内部のプログラムカウンタと呼ばれるレジスタ(図示せず)により示されるアドレスに従ってRAM998からプログラムを読み出す。CPU990はさらに、命令を解釈し、命令の実行に必要なデータを命令により指定されるアドレスに従ってRAM998、ハードディスク1000又はそれ以外の機器から読み出して命令により指定される処理を実行する。CPU990は、実行結果のデータを、RAM998、ハードディスク1000、CPU990内のレジスタなど、プログラムにより指定されるアドレスに格納する。このとき、プログラムカウンタの値もプログラムによって更新される。コンピュータプログラムは、DVD978から、USBメモリ984から、又はネットワークを介して、RAM998に直接にロードしてもよい。なお、CPU990が実行するプログラムの中の、一部のタスク(主として数値計算)については、プログラムに含まれる命令により、又はCPU990による命令実行時の解析結果に従って、GPU992にディスパッチされる。企業活動評価システム50などを実現するプログラムのうち、並列して実行可能な処理は、ソースプログラムの並列構造として、コンパイル時にコンパイラにより見いだされたオブジェクトプログラムの並列構造として、又はプログラムの実行時に実行系システムが見出した並列実行可能なスレッドとして、例えばCPU990が複数のコアを持てばそれらに分散されて実行される。並列処理はGPU992により並列処理として実行されてもよい。

[0135] コンピュータ970により上記した実施形態に係る各部の機能を実現するプログラムは、それら機能を実現するようコンピュータ970を動作させるように記述され配列された複数の命令を含む。この命令を実行するのに必要な基本的機能のいくつかはコンピュータ970において動作するオペレーティングシステム(OS)若しくはサードパーティのプログラム、又はコンピュータ970にインストールされる各種ツールキットのモジュールにより提供されることがある。そうした場合、このプログラムはこの実施形態のシステム及び方法を実現するのに必要な機能全てを必ずしも含まなくてよい。このプログラムは、命令の中で、所望の結果が得られるように制御されたやり方に従って、OSが提供する機能のうち適切なものを読み出すように作成できる。又はこのプログラムは、サードパーティが提供するダイナミックプログラミングライブラリに含まれる機能の中の適切なものを呼出すことにより、上記した各装置及びその構成要素としての動作を実行する命令を含んでいてもよい。コンピュータ970の動作方法は周知なので、ここでは繰返さない。なお、GPU992は並列処理を行うことが可能である。またCPU990も複数のコアを含む場合には完全な並列処理を行うことが可能である。例えばプログラムのコンパイル時にプログラム中において発見された並列的計算要素、又はプログラムの実行時に発見された並列的計算要素は、随時、CPU990の各コア又はGPU992にディスパッチされ、実行され、その結果が直接に、又はRAM998の所定アドレスを介してCPU990により実行されているプログラムに戻され、プログラム中の所定の変数に代入される。もちろん、ソースプログラムの段階からそうした並列的要素を加味してプログラミングを行ってもよい。

[0136] 第4 変形例 上記実施形態においては、企業のDX活動に関する評価を行った。しかしこの発明はそのような実施形態には限定されない。何らかの評価対象に関する質的分析の全般にこの発明を適用できる。企業活動を例に採ると、例えばセキュリティ管理、カーボンニュートラル、働き方改革、SDGsなどに関する調査にこの発明を適用できる。政府及びその各省庁などの官庁、地方自治体、独立行政法人、教育機関、社団法人など、ウェブ上に豊富に情報が存在する団体の活動に関する、多様な調査にこの発明を適用できる。そのためには、図3に示すドメイン評価モデル60の内容を調査対象に応じて変更すればよい。

[0137] 上記実施形態においては、全ての企業についてDX摘要を生成した後、企業ごとに各質問タイプについてスコアを算出している。しかしこの発明はそのような実施形態には限定されない。各企業について、DX摘要を生成した後、続けて質問タイプごとにその企業のスコアを算出するようにしてもよい。

[0138] また上記実施形態においては、5W1Hに対応した6種類の質問タイプを採用している。しかしこの発明はそのような実施形態には限定されない。これ以外の質問タイプがあればそれを採用してもよい。また上記実施形態においては、各質問タイプについて質問テンプレートを1種類だけ使用している。しかし、質問タイプにより質問テンプレートの数を変えてもよいし、また質問タイプごとの質問テンプレートの数も1つに限定されるわけではない。さらに上記実施形態において示した質問テンプレートはあくまで1例である。別の形の質問テンプレートを採用してもよい。

[0139] また上記実施形態においては、各企業に対する処理を順番に行っている。しかし、本発明はそのような実施形態には限定されない。各企業の質問タイプごとのスコアを算出する処理は、本質的に並列に行うことができる。したがって、企業ごとの処理を並列に行ってもよい。さらに企業ごとに、質問タイプごとの処理を並列に行ってもよい。

[0140] また上記実施形態においては、質問応答システムは自然言語による質問を受け付けて自然言語による応答を出力している。しかしこの発明はそのような実施形態には限定されない。質問応答システムへの入力及び出力のいずれも、自然言語ではない形式でもよい。

[0141] 上記実施形態の図5においては、異表記を正式企業名と対応付けておき、正式企業名による質問生成及びその応答の取得と、異表記による質問生成及びその応答の取得とを別々に行っている。そして摘要生成時に正式企業名によりこれら情報を総合して摘要を作成している。しかしこの発明はそのような実施形態には限定されない。例えば正式企業名ごとにフォルダを生成しておき、その正式名と異表記とによる質問を実行して応答をいずれも同じ正式企業名のフォルダに蓄積するようにしてもよい。この場合、摘要生成時に改めて両者を総合する必要はない。結果をDBに記憶する場合には、正式名称による質問と異表記による質問との双方を同じ正式名称をキーとするレコードとしてテーブルに格納しておいてもよい。この場合、応答を取得するときには、正式名称をキーとして質問を検索し、検索された質問が正式名称によるものか異表記によるものかを問わず質問応答システム52に与えてその応答を正式名称と関連付けて記憶するようにすればよい。

[0142] 今回開示された実施形態は単に例示であって、本発明が上記した実施形態のみに制限されるわけではない。本発明の範囲は、発明の詳細な説明の記載を参酌した上で、特許請求の範囲の各請求項によって示され、そこに記載された文言と均等の意味及び範囲内での全ての変更を含む。

符号の説明

[0143] 50 企業活動評価システム 52 質問応答システム 54 企業リスト 56 企業ランキング 60 ドメイン評価モデル 62 質問展開テーブル 64 質問生成部 66 DX優良事例 68 評価部 70 学習データ 340 システム 350、352、354 記憶装置 356 学習データ生成部 358 学習データ記憶装置 362 結合係数推定処理部 364 AUPR算出部 400 Precision-Recall曲線 < / pred> < / obj> < / pred> < / obj> < / pred> < / pred> < / obj> < / obj> < / pred> < / obj> < / obj> < / pred> < / pred> < / pred> < / pred> < / obj> < / obj> < / obj> < / obj> < / obj> < / pred>

Claims

1. a response acquisition means for acquiring a response to each of a plurality of questions from the question answering system for each of a plurality of evaluation objects; evaluation means for evaluating each of the plurality of evaluation objects in accordance with evaluation criteria defined based on a predetermined viewpoint, based on the responses acquired by the response acquisition means for each of the plurality of evaluation objects; a question expansion table storage means for storing a question expansion table; a domain evaluation model storage means for storing a domain evaluation model; a question generation means for generating the plurality of questions for each of the plurality of evaluation objects using the domain evaluation model and the question expansion table, and providing the questions to the response acquisition means; the question expansion table includes a plurality of question templates; each of the plurality of question templates has one or more slots each having one of a plurality of predetermined tags attached thereto; the domain reputation model stores a plurality of character strings, each of which is associated with one of the predetermined plurality of tags; the question generation means includes slot insertion means for generating the plurality of questions by inserting, into each of the one or more slots of the question template, any of the plurality of character strings stored in the domain evaluation model that are associated with the tag of the slot; the plurality of question templates are classified into a plurality of question types; The evaluation means a type-based evaluation means for evaluating each of the plurality of evaluation objects for each of the plurality of question types in accordance with the evaluation criteria based on the responses acquired by the response acquisition means for each of the plurality of evaluation objects; and a comprehensive evaluation means for calculating a comprehensive evaluation of the objects of the plurality of evaluations using the evaluations for each of the plurality of question types obtained by the type-specific evaluation means.

2. The type-specific evaluation means a summary generating means for merging the responses acquired by the response acquiring means for each of the plurality of evaluation objects by question type to generate a summary for each question type; 2. The evaluation device according to claim 1, further comprising: a type-specific summary evaluation means for evaluating each of the plurality of evaluation objects based on the reliability, information amount, and similarity between the summary generated for each question type and predetermined reference data, or any combination thereof.

3. Computer, a response acquisition means for acquiring a response to each of a plurality of questions from the question answering system for each of a plurality of evaluation objects; evaluation means for evaluating each of the plurality of evaluation objects in accordance with predetermined criteria determined based on predetermined viewpoints, based on the responses acquired by the response acquisition means for each of the plurality of evaluation objects; a question expansion table storage means for storing a question expansion table; a domain evaluation model storage means for storing a domain evaluation model; a computer program that functions as a question generation means for generating the plurality of questions for each of the plurality of evaluation objects using the domain evaluation model and the question expansion table and providing the questions to the response acquisition means, the question expansion table includes a plurality of question templates; each of the plurality of question templates has one or more slots each having one of a plurality of predetermined tags attached thereto; the domain reputation model stores a plurality of character strings, each of which is associated with one of the predetermined plurality of tags; the question generation means includes slot insertion means for generating the plurality of questions by inserting, into each of the one or more slots of the question template, any of the plurality of character strings stored in the domain evaluation model that are associated with the tag of the slot; the plurality of question templates are classified into a plurality of question types; The evaluation means a type-based evaluation means for evaluating each of the plurality of evaluation objects for each of the plurality of question types in accordance with the evaluation criteria based on the responses acquired by the response acquisition means for each of the plurality of evaluation objects; and comprehensive evaluation means for calculating a comprehensive evaluation of the object of the plurality of evaluations using the evaluations for each of the plurality of question types obtained by the type-specific evaluation means.

4. The type-specific evaluation means a summary generating means for merging the responses acquired by the response acquiring means for each of the plurality of evaluation objects by question type to generate a summary for each question type; 4. The computer program according to claim 3, further comprising: a type-specific summary evaluation means for evaluating each of the plurality of evaluation objects based on the reliability, information content, and similarity between the summary generated for each question type and predetermined reference data, or any combination thereof.

5. A step in which a computer obtains an answer to each of a plurality of questions for each of a plurality of evaluation objects from a question answering system; A step in which a computer evaluates each of the plurality of evaluation objects according to predetermined criteria defined based on a predetermined viewpoint, based on the responses acquired by the acquiring step for each of the plurality of evaluation objects; a step of causing a computer to store the question expansion table in a question expansion table storage means; A step in which the computer stores the domain evaluation model in a domain evaluation model storage means; a question generating step in which the computer generates the plurality of questions for each of the plurality of evaluation objects using the domain evaluation model and the question expansion table and provides the questions to the response acquisition means; the question expansion table includes a plurality of question templates; each of the plurality of question templates has one or more slots each having one of a plurality of predetermined tags attached thereto; the domain reputation model stores a plurality of character strings, each of which is associated with one of the predetermined plurality of tags; the question generation step includes a slot insertion step of generating the plurality of questions by inserting, into each of the one or more slots of each of the plurality of question templates, any of the plurality of character strings stored in the domain evaluation model that are associated with the tag of the slot; the plurality of question templates are classified into a plurality of question types; The evaluating step includes: a type-based evaluation step of evaluating each of the plurality of evaluation objects for each of the plurality of question types in accordance with the evaluation criteria based on the responses acquired in the acquisition step for each of the plurality of evaluation objects; and a comprehensive evaluation step in which a computer uses the evaluations for each of the plurality of question types obtained in the evaluation step for each type to calculate a comprehensive evaluation of the object of the plurality of evaluations.

6. The type-specific evaluation step includes: a summary generating step in which the computer merges the responses acquired in the acquiring step by question type for each of the plurality of evaluation objects to generate a summary by question type; 6. The evaluation method according to claim 5, further comprising a type-by-type summary evaluation step in which a computer evaluates each of the plurality of evaluation objects based on the reliability, information content, and similarity between the summary generated for each question type and predetermined reference data, or any combination thereof.

Citation Information

Patent Citations

  • Question answering system, data retrieval method, and computer program

    JP2006293830A

  • Query generating device, method, and program

    JP2017027233A

  • Information processing system, information processing method, and program

    JP2021157775A