Dashboard system and program
By estimating topics and similarity calculations on health management questionnaire data and enterprise integrated report data, a dashboard system with relevant scores is solved, and the problem of difficult to quickly find and compare enterprise human capital management and health management information in the existing technology is solved, and rapid and objective information access and comparison are achieved.
Patent Information
- Application Number
- JP2023188046
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-15
- Estimated Expiration
- 2043-11-01
AI Technical Summary
It is difficult for the existing technology to quickly and objectively extract human capital management and health management related information from enterprise integration reports, resulting in difficulty in finding and finding information.
A dashboard system has been developed to estimate topics and similarity calculations on health management questionnaire data and enterprise integrated report data, and generate relevant scores, so that users can quickly find and compare corporate efforts in human capital management and health management.
It realizes rapid and objective access and comparison of enterprise human capital management and health management information, and improves the efficiency and accuracy of information search.
Smart Images

Figure 2025076211000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a dashboard system and program configured by a computer that presents information on human capital management, including health management, and can be used, for example, to improve the efficiency and sophistication of the human resources department or corporate planning department of a health insurance association's employer company, or the work of consultants at management consulting firms. [Background technology]
[0002] Generally, modern corporate management places emphasis on corporate efforts regarding human capital management, including health management (in this application, the term human capital management includes a company's non-financial situation). The Health Management Survey (see Figure 2) published by the Ministry of Economy, Trade and Industry is one survey conducted as part of this effort. This survey on health management by the Ministry of Economy, Trade and Industry is conducted every year, and the 2022 Health Management Survey data includes response data from approximately 2,000 companies.
[0003] On the other hand, companies that are not making sufficient efforts in human capital management, including health management, or that want to further enhance their efforts may want to refer to the efforts of other companies. Also, management consulting firms may introduce the efforts of excellent companies that are making sufficient efforts to companies that are not making sufficient efforts. In such cases, companies that want to refer to the efforts of other companies or management consulting firms that want to introduce the efforts of excellent companies often refer to corporate information published in integrated reports and other documents issued by each company.
[0004] In this invention, a dashboard system is constructed using integrated report data, and in this respect, a related technology is an integrated report evaluation device that can evaluate an integrated report quickly and objectively while making it easy to understand the basis for the evaluation (see Patent Document 1). The "Solution" in the "Abstract" of Patent Document 1 states that "the integrated report evaluation device 1 includes a concept vector acquisition unit 12 that calculates an average value of each word vector of multiple concept words set for an evaluation item to acquire a concept vector, a sentence vector acquisition unit 13 that calculates an average value of each word vector of multiple words included in each sentence of the integrated report to acquire a sentence vector, a feature calculation unit 14 that calculates a feature value indicating the similarity between the concept vector and the sentence vector for each sentence of the integrated report, and a score calculation unit 15 that calculates a score for the evaluation item based on the feature value calculated for each sentence of the integrated report." [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2023-43426 A (Abstract) Summary of the Invention [Problem to be solved by the invention]
[0006] As mentioned above, integrated reports can be used to learn about the efforts of each company regarding human capital management, including health and productivity management. However, because integrated reports are corporate documents written in a format by each company, information on human capital management, including health and productivity management, is not always organized. For this reason, companies must find out where the information they (their companies) want to refer to or use in the integrated report, i.e., the information they need regarding human capital management, including health and productivity management, is, which takes time and effort.
[0007] An object of the present invention is to provide a dashboard system and program that enable easy reference of information about each company regarding human capital management, including health management. [Means for solving the problem]
[0008] The present invention is a dashboard system configured by a computer that presents information on human capital management, including health management, a sub-theme information storage means for performing a topic vector calculation process in which, for health and productivity management level questionnaire data including response data from each company to themes raised as issues in a survey on human capital management including health and productivity management or other questionnaire data, each of the response data from each company is divided into sentences, and using the multiple response sentence data obtained by dividing the response data, for each theme, a topic estimation process is performed to estimate multiple topics for the theme by soft clustering or a neural language model, thereby obtaining a topic value indicating the occurrence probability of each topic in each of the response sentence data, selecting and extracting a predetermined number of top answer sentence data having large topic values from the multiple answer sentence data for each of the multiple topics, and using the multiple answer sentence vectors obtained by vectorizing each of the multiple selected and extracted top answer sentence data, a topic vector calculation process is performed to obtain a topic vector expressing each topic, and storing the obtained topic vector in association with sub-theme identification information that identifies a sub-theme indicating a topic within the theme; A corporate document data storage means for dividing each company's integrated report data or other corporate document data, which is different from the questionnaire data and in which each company describes information about its own human capital management, including health management, in a free-form format for each company, into sentences, and storing the multiple corporate document data obtained by dividing them in association with corporate identification information that identifies the company and corporate document identification information that identifies the corporate document data; a similarity calculation means for calculating a similarity between a plurality of company sentence vectors obtained by vectorizing each of the plurality of company sentence data by a vectorization process using the same method as that used in the vectorization process of the answer sentence data for each sub-theme and a topic vector of each sub-theme stored in the sub-theme information storage means, thereby executing a process of obtaining a relevance score for each of the plurality of company sentence data with respect to each of the sub-themes; an association score storage means for storing the association score for each sub-theme calculated by the similarity calculation means in association with the sub-theme identification information, the company identification information, and the company sentence identification information; an output means for executing a process of displaying the company sentence data stored in the company sentence data storage means on a screen by using the related scores for each sub-theme for the company sentence data of each company stored in the related score storage means; The present invention is characterized by comprising:
[0009] Here, the "sub-theme identification information" may be a single identification information capable of identifying which topic belongs to which theme, or it may be a combination of theme identification information and topic identification information limited to a theme capable of identifying topics only within each theme.
[0010] Furthermore, the process of calculating the relevance score by the "similarity calculation means" may be executed as a pre-processing before the user uses the system, or may be executed as a real-time processing while the user is using the system. In the former case of pre-processing, the "relevance score storage means" is composed of a non-volatile memory, and in the latter case of real-time processing, the "relevance score storage means" may be either a volatile memory or a non-volatile memory.
[0011] Furthermore, "health and productivity management questionnaire data or other questionnaire data" refers to questionnaire data in which response data on a set topic is organized, and the survey body is not limited to a national or local government agency such as the Ministry of Economy, Trade and Industry, but may be a private research organization. In short, it is sufficient that the data includes responses from each company according to a set topic, and that the response data functions as training data for building this system. Therefore, even if the name of the questionnaire currently named the Health and Productivity Management Questionnaire by the Ministry of Economy, Trade and Industry is changed later, the data after the name change may be used as long as it includes response data on a similar topic. Also, the data may be in accordance with the guidelines of ISO30414 (Guidelines for disclosure of information on human capital).
[0012] Also, unlike the questionnaire data above, "integrated report data or other corporate document data" means information about human capital management, including health and productivity management, written in a free format by each company, and includes, in addition to integrated reports, for example, the section of a securities report that describes human capital management, including health and productivity management, and sustainability reports. Therefore, it does not include organized information such as the questionnaire data above, that is, document data in which each company answers questions in a standard form according to a theme. This is because such information that is organized from the beginning does not require the user to spend time and effort finding the information they need.
[0013] In the dashboard system of the present invention, a topic estimation process is performed for each theme using the "health management survey data or other survey data" to determine sub-themes that represent each topic, and topic vectors that express these sub-themes are obtained. The similarity between the corporate text vectors of the corporate text data created from the "integrated report data or other corporate document data" and the topic vectors is calculated, and the calculated similarity is used as a relevance score for each sub-theme for each corporate text data. Output processing is then performed to display each corporate text data on the screen using this relevance score.
[0014] Therefore, because the themes of issues in the "Health and Productivity Management Survey Data or Other Survey Data" are automatically subdivided by the topic estimation process, it becomes possible to grasp issues in terms of sub-themes (each topic of each theme) with finer granularity than the theme, and to refer to or compare company document data. As a result, users of this system can refer to the details of each company's efforts regarding human capital management (including non-financial situations), including health and productivity management, without spending time or effort, thereby achieving the above-mentioned objective.
[0015] <Configuration including representative related score calculation means and output means executing overhead view display process>
[0016] In the dashboard system of the present invention, a representative related score calculation means for executing a process of extracting, for each sub-theme, a maximum related score from among a plurality of related scores associated with the same company identification information, using the related scores for each sub-theme stored in the related score storage means, and setting the extracted related score as a representative related score for the company of the company identification information, or extracting a plurality of related scores with the highest values, and setting an average value or other calculated value calculated using the extracted plurality of top related scores as a representative related score for the company of the company identification information, The output means is It is desirable to configure the device to execute an overhead view display process that displays an overhead view on the screen in which one axis of a matrix-like display formed by arranging cells vertically and horizontally represents each sub-theme and the other axis represents each company, and in which the intensity of the display color of each cell corresponds to the magnitude of the representative related score calculated by the representative related score calculation means.
[0017] Here, the calculation process of the representative related score by the "representative related score calculation means" may be performed as a pre-processing before the user uses the system, or may be performed as a real-time processing while the user is using the system.
[0018] In this manner, when the representative related score calculation means is provided and the output means is configured to execute the bird's-eye view display process, the status of each company's efforts with respect to each sub-theme can be seen at a glance by the shade of the color of the cell, making it possible to easily grasp, for example, the status of efforts in the industry as a whole and differences in the status of efforts between companies.
[0019] <Configuration with sub-theme name acquisition means using large-scale language models>
[0020] Furthermore, in the dashboard system of the present invention, A topic estimation means for executing a topic estimation process; a topic model storage means for storing, for each theme, the occurrence probability of each word in each topic among the themes obtained by executing a topic estimation process for each theme by the topic estimation means; a sub-theme name acquisition means for extracting a predetermined number of top words having a high occurrence probability for each theme and for each topic from the occurrence probability of each word in each topic for each theme stored in the topic model storage means, inputting the extracted multiple words and sub-theme name creation request data for each theme and for each topic, which contains a request to create a sub-theme name using these multiple words, into a chat generative pretrained transformer (ChatGPT) or other large-scale language model, and storing the sub-theme name output from the large-scale language model in the sub-theme information storage means in association with the sub-theme identification information; The output means is It is desirable that when a user's request to display information for each sub-theme is received, or when information for each sub-theme is displayed, a process is executed using the sub-theme name stored in the sub-theme information storage means.
[0021] Here, the "request to create a sub-theme name" constituting the "sub-theme name creation request data" does not mean directly inputting the words "sub-theme name" into the large-scale language model. Since the large-scale language model does not recognize that the input words are a request to obtain a name to be used as the sub-theme name, the words input as the request can be "theme" or an equivalent word such as "subject" rather than a sub-theme.
[0022] In the case of a configuration equipped with a sub-theme name acquisition means using a large-scale language model in this way, in the past, when topic names (names equivalent to the sub-theme names of the present invention) had to be determined in the topic estimation process, the topic names were determined manually, but by determining the topic names using a large-scale language model, the system developer's work can be saved and more objective and appropriate names can be adopted. Note that since the number of topics when performing the topic estimation process is specified by the system developer, topic names are not always necessary when performing topic estimation, but in the present invention, topic names (names equivalent to the sub-theme names of the present invention) are necessary because there are cases where a person who refers to or searches corporate text data specifies a sub-theme (i.e., each topic within each theme) or understands the company's issues in units of sub-themes.
[0023] <Configuration in which output means executes processing for displaying a list of comparisons between companies>
[0024] In the dashboard system of the present invention, The output means is It is desirable to have a configuration in which a user's selection of a sub-theme and a selection of multiple companies are accepted, and for each of the multiple selected companies, the top company sentence data having the highest related score for the selected sub-theme is selected and extracted from the multiple company sentence data, and the selected and extracted company sentence data is compiled and compared by company in a list, and an inter-company initiative comparison list display process is executed to display the selected and extracted company sentence data on the screen together with company identification information or together with company identification information and related scores.
[0025] Here, the "multiple companies" selected and compared by the user are preferably two companies, from the viewpoint of preventing the font size displayed on the screen from becoming too small, but the corporate document data of three or more companies may be displayed in a list for comparison, grouped by company.
[0026] In this manner, when the output means is configured to execute processing for displaying a list for comparing initiatives between companies, it becomes possible to easily compare the status of initiatives by a plurality of companies regarding a certain sub-theme.
[0027] <Configuration equipped with corporate document acquisition means using large-scale language models>
[0028] In addition, in the case where the output means is configured to execute a process for displaying a list of comparisons of business ventures, The apparatus is provided with a company commitment statement acquisition means for inputting the top multiple company statement data for each selected and extracted company and company commitment statement creation request data, which includes a request to create a list of multiple commitment statements showing the company's efforts using the multiple company statement data, into a chat generative pretrained transformer (ChatGPT) or other large-scale language model when executing a process for displaying a list of company commitment comparisons by the output means, and outputting the multiple commitment statements from the large-scale language model in list form; The output means is As a process for displaying a list of inter-company initiative comparisons, it is desirable to configure the system to also execute a process for displaying on the screen a number of initiative statements acquired by the company initiative statement acquisition means in list form in correspondence with a number of higher-ranking company statement data for each company.
[0029] When the system is configured to include a corporate commitment statement acquisition means that utilizes a large-scale language model in this manner, the large-scale language model uses the top corporate statement data for each company to output multiple commitment statements showing the company's efforts in list format, and these are displayed on the screen, making it easier for users of the system to understand the details of the company's efforts and also allowing them to use these commitment statements as search strings.
[0030] <Configuration in which output means executes similar sentence search process>
[0031] Furthermore, in the dashboard system of the present invention, The output means is It is desirable to have a configuration in which a similar sentence search process is executed in which an input of an arbitrary search string by a user is accepted, the input search string data is vectorized by a vectorization process using the same method as that used in the vectorization process of the answer sentence data, a similarity between the obtained search string vector and each of a plurality of company sentence vectors obtained by vectorizing a plurality of company sentence data stored in a company sentence data storage means is calculated, an input string related score for the search string data for each of the plurality of company sentence data is obtained, and the company sentence data is displayed in order of the highest obtained input string related score.
[0032] Here, the "search string" may be a sentence or a word, and in the case of a sentence, it may be one sentence or multiple sentences.
[0033] When the output means is configured to execute a similar sentence search process in this manner, a relevance score (this relevance score is called an input string relevance score to distinguish it from the relevance score for the sub-theme) is calculated for each of the multiple corporate sentence data with respect to any search string data input by the user, and the corporate sentence data is displayed using this input string relevance score, thereby enabling users of the system to more easily display the information they wish to refer to.
[0034] <Configuration in which output means executes similar sub-theme display processing during similar sentence search processing>
[0035] In addition, in the case where the output means is configured to execute a similar sentence search process, The output means is A similarity between the search character string vector and topic vectors for a plurality of subthemes stored in the subtheme information storage means may be calculated to extract subthemes of topic vectors whose similarity to the search character string vector is greater than or equal to a predetermined threshold, or to extract subthemes of topic vectors whose similarity is highest among a predetermined number of cases, and to display the extracted subthemes on a screen to present them to the user.
[0036] In this manner, when the output means is configured to execute a similar subtheme display process during the similar sentence search process, subthemes similar to any search string data entered by the user of the system can be presented to the user, and the display method can be changed using the subthemes (such as changing the order in which corporate sentence data is displayed).
[0037] <Configuration in which the output means executes a similar subtheme display process during a similar sentence search process, and then executes an integrated relevance score order display process>
[0038] Furthermore, as described above, in the case where the output means is configured to execute a similar subtheme display process during the similar sentence search process, The output means is an input string association score; using the related score stored in the related score storage means in association with the subtheme identification information of at least one subtheme selected by the user from the at least one subtheme presented by the similar subtheme display process, For each of the multiple corporate text data, the input string relevance score and at least one relevance score are used to calculate an average value, a weighted average value, or other calculated value, which is used as an integrated relevance score; The configuration may be such that an integrated association score order display process is executed to sort and display the company statement data in the order of highest integrated association score.
[0039] In this manner, when the output means is configured to execute the similar subtheme display process and then the integrated relevance score order display process during the similar sentence search process, the user of the present system can refer to the company sentence data sorted in descending order of the input string relevance score (relevance score for any search string data entered by the user) in a state sorted in descending order of the integrated relevance score. In other words, the user can refer to the data sorted in a state taking into account the relevance score for the subthemes presented in the similar subtheme display process.
[0040] <Configuration in which the output means executes a similar subtheme display process during a similar sentence search process, and then executes an adjusted related score order display process>
[0041] As described above, in the case where the output means is configured to execute a similar subtheme display process during the similar sentence search process, The output means is Using the search character string vector and at least one topic vector associated with the subtheme identification information of at least one subtheme selected by the user from at least one subtheme presented by the similar subtheme display process and stored in the subtheme information storage means, a composite vector is created by calculating an average value or a weighted average value or other calculated value of the elements of each vector; By calculating the similarity between the composite vector and each of a plurality of company sentence vectors obtained by vectorizing the plurality of company sentence data stored in the company sentence data storage means, an adjusted association score is obtained for each of the plurality of company sentence data; The system may be configured to execute an adjusted related score order display process in which the company statement data is sorted and displayed in descending order of the adjusted related score.
[0042] In this manner, when the output means is configured to execute the similar subtheme display process and then the adjusted relevance score order display process during the similar sentence search process, the user of the present system can refer to the corporate sentence data sorted in descending order of the input string relevance score (relevance score for any search string data entered by the user) in a state sorted in descending order of the adjusted relevance score. In other words, the user can refer to the data sorted using the topic vectors of the subthemes presented in the similar subtheme display process.
[0043] <Configuration including representative related score calculation means and output means executing process for extracting initiative recommendation sub-themes, process for displaying initiative texts of focused companies, and process for displaying recommended initiative texts>
[0044] Furthermore, in the dashboard system of the present invention, a representative related score calculation means for executing a process of extracting, for each sub-theme, a maximum related score from among a plurality of related scores associated with the same company identification information, using the related scores for each sub-theme stored in the related score storage means, and setting the extracted related score as a representative related score for the company of the company identification information, or extracting a plurality of related scores with the highest values, and setting an average value or other calculated value calculated using the extracted plurality of top related scores as a representative related score for the company of the company identification information, The output means is accepts a user's selection of one target company, and executes an initiative recommendation sub-theme extraction process, which calculates a deviation of the target company's representative related score from the representative related scores of companies other than the target company by calculating, for each sub-theme, a difference between the representative related score of the selected target company and the largest representative related score among the representative related scores of companies other than the target company, or a difference between the representative related score of the selected target company and an average or other calculated value of a plurality of top representative related scores among the representative related scores of companies other than the target company, and extracts a plurality of top sub-themes with the largest calculated deviation as initiative recommendation sub-themes and displays them on the screen; Accepting a user's selection of at least one initiative recommendation sub-theme from among the multiple initiative recommendation sub-themes extracted by the initiative recommendation sub-theme extraction process; A target company initiative display process for displaying the target company's initiative recommendation sub-theme related to the user's selection of the target company's initiative recommendation sub-theme on the screen, the target company initiative display process being performed by displaying the target company ... The system may be configured to execute a recommended initiative display process that displays on a screen the corporate text data of companies other than the focal company stored in the corporate text data storage means together with company identification information, or together with company identification information and the associated score, in order of the highest associated score for the initiative recommendation sub-theme selected by the user for the corporate text data of companies other than the focal company stored in the associated score storage means.
[0045] In this way, when the representative related score calculation means is provided and the output means is configured to execute the process of extracting the recommended effort sub-theme, the process of displaying the focused company's effort, and the process of displaying the recommended effort, it is possible to display the recommended effort for the focused company that needs to improve its efforts regarding human capital management, including health management. Therefore, when a company that is not fully committed to human capital management, including health management, or a company that wants to further enhance its efforts, is a user of this system, it can easily refer to the efforts of other companies that should be used as reference, and can efficiently collect information. In addition, when a management consulting firm is a user of this system, it can efficiently introduce the efforts of excellent companies that are making sufficient efforts to a company with insufficient efforts.
[0046] <Configuration with extended sub-theme creation means>
[0047] In the dashboard system of the present invention, An extended subtheme creating means for creating an extended subtheme to be added to the subtheme obtained by the topic estimation process, This method of creating an extended sub-theme is as follows: an extended sub-theme keyword input reception process for receiving an input of an extended sub-theme keyword; an extended sub-theme challenge statement creation request data, which describes a request to create multiple challenge statements showing efforts related to the extended sub-theme keyword using the received extended sub-theme keyword, is input to a chat generative pretrained transformer (ChatGPT) or other large-scale language model, and multiple challenge statement data is output from the large-scale language model; A process of creating a plurality of approach sentence vectors by performing a vectorization process for each of the plurality of output approach sentence data by the same method as that used in the vectorization process for the answer sentence data; It is desirable to have a configuration in which an extended subtheme vector creation process is executed in which an average of the multiple created initiative text vectors is taken to create an extended subtheme vector, and the obtained extended subtheme vector is associated with the additionally added subtheme identification information and stored in the subtheme information storage means.
[0048] In this way, when the system is configured with an extended sub-theme creation means, any sub-theme can be constructed and used as an extended sub-theme. In other words, in addition to the sub-themes (each topic within each theme) obtained by subdividing the theme of the assignment, at least one extended sub-theme can be prepared and provided to the user. This makes it possible to build a system that meets the needs of the user, thereby improving the convenience of the system.
[0049] <Program invention>
[0050] The program of the present invention is for causing a computer to function as the above-mentioned dashboard system.
[0051] The above-mentioned program or a part of it can be recorded on a recording medium such as a magneto-optical disk (MO), a compact disk (CD), a digital versatile disk (DVD), a flexible disk (FD), a magnetic tape, a read-only memory (ROM), an electrically erasable and programmable read-only memory (EEPROM), a flash memory, a random access memory (RAM), a hard disk drive (HDD), a solid-state drive (SSD), a flash disk, etc., and can be stored or distributed, and can be transmitted using a transmission medium such as a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a wired network such as the Internet, an intranet, an extranet, a wireless communication network, or a combination of these, and can also be carried on a carrier wave. Furthermore, the above-mentioned program may be a part of another program, or may be recorded on a recording medium together with a separate program. Effect of the Invention
[0052] As described above, according to the present invention, a topic estimation process is performed for each theme using health management survey data or other survey data to determine sub-themes that represent each topic, topic vectors that express the sub-themes are obtained, and the similarity between the topic vector and the corporate text vector of corporate text data created from integrated report data or other corporate document data is calculated. The calculated similarity is used as an association score for each sub-theme for each corporate text data, and each corporate text data is displayed on the screen using this association score, so that users of the system can refer to the details of each company's efforts regarding human capital management, including health management, without spending time and effort. [Brief description of the drawings]
[0053] [Figure 1] 1 is a diagram showing the overall configuration of a dashboard system according to an embodiment of the present invention; [Diagram 2] FIG. 13 is an example diagram of health management questionnaire data including response data according to the embodiment. [Diagram 3] FIG. 4 is an explanatory diagram of a topic inference process and its preparation process according to the embodiment. [Figure 4] FIG. 4 is an explanatory diagram of a topic vector calculation process according to the embodiment. [Diagram 5] 6 is an explanatory diagram of a calculation process of an association score and a representative association score in the embodiment. [Figure 6] 5A to 5C are diagrams illustrating screen transitions in the embodiment. [Figure 7] FIG. 4 is a diagram showing an example of an overhead view display screen in the embodiment. [Figure 8] FIG. 13 is a diagram showing an example of an inter-company deal comparison list display screen according to the embodiment. [Figure 9] FIG. 13 is an example of a screen displaying each company's commitment documents according to the embodiment. [Figure 10] FIG. 4 is a view showing an example of a similar sentence search screen in the embodiment. [Figure 11] FIG. 13 is an explanatory diagram of the merging or adjustment of related scores by similar sub-themes in the embodiment. [Figure 12] FIG. 11 is a view showing an example of a recommendation screen according to the embodiment. [Figure 13] FIG. 11 is an explanatory diagram of a sub-theme name acquisition process according to the embodiment. [Figure 14] FIG. 4 is an explanatory diagram of the process of acquiring a company initiative statement according to the embodiment. [Figure 15] FIG. 4 is a flowchart showing a process flow of the dashboard system according to the embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0054] An embodiment of the present invention will be described below with reference to the drawings. FIG. 1 shows the overall configuration of a dashboard system 10 of this embodiment, and FIG. 2 shows an example of health management survey data including response data. FIG. 3 is an explanatory diagram of topic estimation processing by the topic estimation means 34 and preparation processing by the topic estimation preparation means 33, FIG. 4 is an explanatory diagram of topic vector calculation processing by the topic vector calculation means 36, and FIG. 5 is an explanatory diagram of relevance score calculation processing by the similarity calculation means 39 and representative relevance score calculation processing by the representative relevance score calculation means 40. FIG. 6 shows an example of screen transition by the output means 41, FIG. 7 shows an example of a bird's-eye view display screen, FIG. 8 shows an example of an inter-company initiative comparison list display screen, FIG. 9 shows an example of an initiative statement display screen for each company, FIG. 10 shows an example of a similar sentence search screen, FIG. 11 is an explanatory diagram of integration or adjustment of relevance scores by similar sub-themes, and FIG. 12 shows an example of a recommendation screen. Furthermore, Figure 13 is an explanatory diagram of the process of acquiring sub-theme names by the sub-theme name acquisition means 35, Figure 14 is an explanatory diagram of the process of acquiring company commitment statements by the company commitment statement acquisition means 42, and Figure 15 is a flowchart showing the processing flow by the dashboard system 10.
[0055] <Overall configuration of dashboard system 10>
[0056] 1, a dashboard system 10 includes a main body 20 configured with one or more computers, a display means 80 such as a liquid crystal display, and an input means 81 such as a mouse or a keyboard. A service providing system 70 using large language models (LLMs) is connected to the main body 20 via a network 1.
[0057] Here, the service provision system 70 using a large-scale language model (LLM) is a system that provides a cloud API (Application Programming Interface) service. In this embodiment, Chat Generative Pre-trained Transformer (ChatGPT: Chat Generative Pre-trained Transformer, Azure OpenAI ChatCompletion API GPT-4) is used, but is not limited to this. For example, OpenAI's GPT-3.5, Google's Palm, Palm2, Amazon Web Services' (AWS) Titan, Meta Platforms' (Meta) Llama, etc. may be adopted.
[0058] Furthermore, network 1 is an external network mainly composed of the Internet, but it may also be a combination of the Internet with an internal network such as a LAN or an intranet, and it does not matter whether it is wired or wireless, or even a combination of wired and wireless; in short, it is sufficient if it is capable of transmitting information at a certain degree of speed between multiple points (regardless of the distance).
[0059] Although not shown in the figure, external service providing systems using cloud APIs connected via network 1 include, in addition to the service providing system 70 using a large-scale language model (LLM), for example, a service providing system that performs vectorization processing of text data, and a service providing system that performs OCR processing to convert PDF data, etc., into text data.
[0060] The main body 20 is equipped with a processing means 30 that executes various processes for presenting information regarding human capital management, including health management, and a storage means 50 that stores various data necessary for the processing by this processing means 30.
[0061] The processing means 30 is composed of an answer sentence data creation means 31, an answer sentence vector creation means 32, a topic estimation preparation means 33, a topic estimation means 34, a sub-theme name acquisition means 35, a topic vector calculation means 36, a company sentence data creation means 37, a company sentence vector creation means 38, a similarity calculation means 39, a representative related score calculation means 40, an output means 41, and a company initiative sentence acquisition means 42.
[0062] Here, each of the means 31 to 42 included in the processing means 30 is realized by a central processing unit (CPU) provided inside the main body 20, one or more programs that define the operation procedure of this CPU, and a working memory such as a main memory, a cache memory, etc. The details of each of these means 31 to 42 will be described later.
[0063] The storage means 50 is composed of a questionnaire data storage means 51, a response sentence data storage means 52, a response sentence vector storage means 53, a topic model storage means 54, a sub-theme information storage means 55, a company document data storage means 56, a company sentence data storage means 57, a company sentence vector storage means 58, an association score storage means 59, and a representative association score storage means 60.
[0064] Here, each of the storage means 51-60 included in the storage means 50 may be a non-volatile memory such as a hard disk drive (HDD) or a solid state drive (SSD). The data is preferably stored in a database format, but the questionnaire data storage means 51 and the company document data storage means 56 may be stored in a file format. However, the relevance score storage means 59 and the representative relevance score storage means 60 may be main memory (volatile memory) in the case where the relevance scores and the representative relevance scores are calculated in real time processing in accordance with the operation of the user of the present system, rather than calculated in advance by processing. The details of these storage means 51-60 will be described later.
[0065] <Configuration of response sentence data creation means 31>
[0066] The answer sentence data creation means 31 creates answer sentence data for each company by dividing the answer data (see Figure 2) for each theme (task theme) contained in the health management questionnaire data or other questionnaire data of each company stored in the questionnaire data storage means 51 into sentences for each period, and executes a process of storing the answer sentence data for each company in the answer sentence data storage means 52 (see Figure 3) in association with the theme identification information, company identification information (stock code, stock name, etc.), and answer sentence identification information.
[0067] At this time, if the questionnaire data is not text data but is PDF data or the like, data conversion processing, OCR processing, etc. are performed to convert it to text data and then split it. Note that in the topic estimation process, since it is not distinguished which company's answer sentence data belongs to, the answer sentence identification information is identification information (serial number) within the theme, not identification information within the company.
[0068] 2, the health management level survey data includes (a) themes of issues, (b) contents of issues, (c) implementation results of measures, and (d) results of effect verification, and among these, the entered data of (b), (c), and (d) are used as response data in the present invention. Therefore, the response data creation means 31 divides the response data of (b) into response data b1, b2, b3, ..., the response data of (c) into response data c1, c2, c3, ..., and the response data of (d) into response data d1, d2, d3, ....
[0069] Some companies may respond by misunderstanding the division into (b), (c), and (d), but this is not a problem because the response data is collected from many companies to create the response sentence data, and the response sentence data created from the response data of (b), (c), and (d) is used in the topic inference process without distinguishing between them. In addition, to ensure appropriate topic inference, it is preferable to use all three response data of (b), (c), and (d), but topic inference processing can be performed without necessarily using all three of (b), (c), and (d). In this sense, if the division into (b), (c), and (d) of the questionnaire changes in the future and the number of categories increases or decreases, it is preferable to use the response data of all categories, but it is acceptable for some categories not to be used.
[0070] Furthermore, questionnaire data such as health management questionnaire data is basically created every year, but the topic estimation process may be performed each time, or topic vectors obtained by the topic estimation process performed in the previous year or earlier may be used. When the topic estimation process is performed again, past response data and the latest response data may be mixed and used, or only the latest response data may be used, as long as the content of the theme matches. Note that when the year changes and the theme no longer matches (for example, when the number of themes changes from 10 to 9, or from 10 to 11, etc.), in principle, it is better not to mix past response data and the latest response data. This is because the topic estimation process is performed for each theme, and if the number of themes changes (i.e., if the content of each theme changes), the response content of each company for each theme will differ.
[0071] <Configuration of the response text vector creation means 32>
[0072] The answer message vector creation means 32 executes a process of vectorizing the answer message data stored in the answer message data storage means 52 (see Figure 3) and storing the obtained answer message vector in the answer message vector storage means 53 (see Figure 4) in association with theme identification information, company identification information (stock code, stock name, etc.), and answer message identification information.
[0073] In this embodiment, as an example, the answer sentence vector creation means 32 uses an Azure Open AI Embeddings API service provision system (Azure OpenAI Embeddings API text-embedding-ada-002 version2) to vectorize the answer sentence data and obtain a 1,536-dimensional answer sentence vector as shown in Fig. 4. Note that the embedding method may be Doc2Vec, BERT, Transformer, etc. The number of dimensions of the answer sentence vector is not limited to 1,536 dimensions.
[0074] <Configuration of topic estimation preparation means 33>
[0075] As shown in FIG. 3, the topic estimation preparation means 33 removes unnecessary symbols and tags (for example, ☆, (1) (environment-dependent “maruichi”), The process includes a process of removing unnecessary words, a process of decomposing (dividing) the text into words by morphological analysis and extracting only nouns, and a process of removing unnecessary words. Note that these preparatory processes are similar to the processes described in Japanese Patent Application Laid-Open No. 2021-26413 and Japanese Patent Application Laid-Open No. 2022-190557 by the applicant of the present application.
[0076] At this time, the topic estimation preparation means 33 can execute morphological analysis by using an existing analysis tool.
[0077] Furthermore, the topic estimation preparation means 33 narrows down the words in the unnecessary word removal process. That is, first, the words are filtered based on the part of speech and the number of occurrences of the words. In the set of all answer sentence data, words that have appeared less than three times, for example, are discarded. The relationship between each word and the number of occurrences in the set of all answer sentence data is stored in a word occurrence count storage means (not shown). Next, the topic estimation preparation means 33 eliminates unnecessary words (noise words) stored in an unnecessary word dictionary storage means (not shown). Specifically, for example, words that are thought to appear regardless of the theme, such as "company" and "our company", are eliminated as unnecessary words.
[0078] In the example of FIG. 3, the words remaining after the above-mentioned process of removing unnecessary symbols and tags by the topic estimation preparation means 33, the process of breaking down into words and extracting only nouns by morphological analysis, and the process of removing unnecessary words are "sleepiness", "at work", .... Therefore, as shown in FIG. 3, the relationship between each remaining word and its occurrence frequency is obtained, and this relationship becomes information necessary for topic estimation. That is, it is the occurrence frequency of each word in one answer sentence data (i=00001234) treated as one document data in the topic estimation process (the document data here does not mean the corporate document data in the present invention, but the document data generally referred to when performing (explaining) the topic estimation process). i=00001234 is the answer sentence identification information (for example, the identification number of the answer sentence b1 of X company). Then, the occurrence frequency of each word is obtained for all answer sentence data (i=1 to n: n is the number of answer sentence data). Furthermore, although the example of FIG. 3 is the process for theme 4, the above process is performed for each theme and for all themes. In addition, since the topic estimation process is performed for each theme, in this embodiment, the answer sentence identification information is identification information within a theme (serial number within a theme) that can identify an answer sentence within each theme. However, the answer sentence identification information may be identification information across all themes.
[0079] <Configuration of topic estimation means 34>
[0080] The topic estimation means 34 executes a topic estimation process for estimating multiple topics for each theme by soft clustering or a neural language model using the response sentence data of each company stored in the response sentence data storage means 52 (see FIG. 3) for each theme. This means that the process for estimating multiple topics for one theme is executed for all themes.
[0081] More specifically, in this embodiment, as an example of the soft clustering or neural language model used when performing the topic estimation process, Latent Dirichlet Allocation (LDA) is adopted. That is, in this embodiment, as shown in FIG. 3, the topic estimation means 34 executes a topic estimation process for obtaining a topic value (column vector π(i)) indicating the occurrence probability of each topic in the answer sentence data (i) and the occurrence probability (matrix β) of each word in each topic by Latent Dirichlet Allocation (LDA) using the occurrence frequency of each word in each answer sentence data (i=1 to n: n is the number of answer sentence data in the theme) obtained by the process by the topic estimation preparation means 33 for each theme, and executes a process of storing the topic distribution indicated by the column vector π(i) and the matrix β obtained by this topic estimation process as a topic model in the topic model storage means 54 for each theme. As shown in FIG. 3, the matrix β is a matrix of K rows and p columns (K is the number of topics, and p is the number of words). The number of topics K in a theme is specified by the system developer, but it is not necessary to set the same number for each theme. Specifically, for example, the number of topics K for themes 1 to 10 is 5, 5, 7, 5, 5, 6, 5, 6, 4, 4, in order from theme 1.
[0082] In addition to Latent Dirichlet Allocation (LDA), other algorithms that can be used include Fuzzy c-means, mixture distribution model, Non-negative Matrix Factorization (NMF), probabilistic Latent Semantic Indexing (pLSI), Doc2Vec, Sparse Compose Document Vectors (SCDV), etc. For example, when implementing Doc2Vec, an existing library called Gensim can be used.
[0083] <Configuration of the sub-theme name acquisition means 35>
[0084] The sub-theme name acquisition means 35 executes a process of extracting a predetermined number of top (e.g., top 10) words with a high occurrence probability for each theme and for each topic from the occurrence probability (matrix β) of each word in each topic for each theme stored in the topic model storage means 54, that is, extracting, for example, the top 10 words with a high occurrence probability in each row of the matrix β, creating sub-theme name creation request data for each theme and for each topic, which describes the extracted multiple (e.g., 10) words and a request to create a sub-theme name using these multiple words, inputting the created sub-theme name creation request data into a chat generative pretrained transformer (ChatGPT) or other large-scale language model (LLM), and storing the sub-theme name output from this large-scale language model (LLM) in the sub-theme information storage means 55 in association with the sub-theme identification information.
[0085] Specifically, as shown in FIG. 13, the sub-theme name acquisition means 35 creates sub-theme name creation request data by adding 10 words, "0 working hours...9 report" to text data, "Please think of one theme with about 10 characters common to the following words," and transmits the created sub-theme name creation request data to the service provision system 70 using a large-scale language model (LLM) via the network 1, receives data returned from the large-scale language model, and stores the received data in the sub-theme information storage means 55 as a name indicating a topic in the theme, that is, a sub-theme name, in association with the sub-theme identification information. In the example of FIG. 13, "Report on the implementation status of improvement of working hours" is returned from the service provision system 70 using a large-scale language model, and this becomes the sub-theme name. Note that the reason why "Think of a theme..." is input into the request to the large-scale language model (LLM) instead of "Think of a sub-theme..." is because the large-scale language model does not recognize that its reply data will be used as a sub-theme name.
[0086] <Configuration of topic vector calculation means 36>
[0087] The topic vector calculation means 36 executes a topic vector calculation process in which, for each theme, for each of a plurality of topics in the theme, from among the plurality of answer sentence data, the top (e.g., top 10) answer sentence data having the largest topic value is selected and extracted, and each of the selected and extracted top multiple (e.g., U=10) answer sentence data is vectorized to obtain a plurality of answer sentence vectors (in this embodiment, 1,536-dimensional answer sentence vectors created by the answer sentence vector creation means 32 and stored in the answer sentence vector storage means 53 (see FIG. 4 )), to obtain a topic vector expressing each topic (in this embodiment, a 1,536-dimensional vector), and the obtained topic vector is stored in the sub-theme information storage means 55 in association with the sub-theme identification information.
[0088] At this time, the topic vector calculation means 36 obtains an average vector for the answer sentence vectors of the top multiple (e.g., U=10) answer sentence data selected and extracted, and sets it as the topic vector, as shown in Fig. 4. However, it is not limited to an average vector obtained by simply averaging the values of the elements of each dimension of the vector, and it may be, for example, a weighted average vector or a harmonic average vector, and in short, it is sufficient to perform some kind of calculation using the top multiple (e.g., U=10) answer sentence vectors to calculate the topic vector. Even when calculating a weighted average vector to use as a topic vector, the weights can be determined arbitrarily. For example, if you want to increase the influence of the answer sentence vector of the answer sentence data with the largest topic value, then you can set the weight of the answer sentence vector of the answer sentence data with the largest topic value to 10, the weight of the answer sentence vector of the answer sentence data with the second largest topic value to 9, the weight of the answer sentence vector of the answer sentence data with the third largest topic value to 8, ..., the weight of the answer sentence vector of the answer sentence data with the tenth largest topic value to 1, and then add these up and divide by (10 + 9 + 8 + ... + 1) = 55. In addition, the weighted average can be calculated using the length of the answer sentence data (number of characters) as the weight.
[0089] More specifically, the topic vector calculation means 36, as shown in Fig. 4, first focuses on the topic with topic number = 1 among multiple topics (topic numbers = 1 to K: K is the number of topics, for example, K = 5) in theme 4. That is, it focuses on the topic value π(i, 1) of the top element (first-dimension element) of the vertical vector π(i) indicating the topic distribution stored in the topic model storage means 54, and extracts answer sentence identification information of multiple answer sentence data items (for example, U = 10 items) with the largest topic value. In the example of Fig. 4, i = 00000264, i = 00000967, i = 00001764, ... are extracted. Then, U answer sentence vectors (for example, U=10) of i=00000264, i=00000967, i=00001764, ... (1,536-dimensional vectors) are obtained from the answer sentence vector storage means 53, and the average vector of these U answer sentence vectors is calculated. This is treated as the topic vector for topic number 1 in theme 4, and is stored in the sub-theme information storage means 55 in association with the sub-theme identification information "4-1".
[0090] Next, the topic vector calculation means 36 focuses on the topic with topic number=2. That is, it focuses on the topic value π(i,2) of the second element from the top (second-dimensional element) of the vertical vector π(i) indicating the topic distribution stored in the topic model storage means 54, and extracts answer sentence identification information of the top multiple answer sentence data items (e.g., U=10 items) with the largest topic value. In the example of Fig. 4, i=00000134, i=00000777, i=00001348, ... are extracted. Then, U answer sentence vectors (for example, U=10) of i=00000134, i=00000777, i=00001348, ... (1,536-dimensional vectors) are obtained from the answer sentence vector storage means 53, and the average vector of these U answer sentence vectors is calculated. This is treated as the topic vector for topic number = 2 in theme 4, and is stored in the sub-theme information storage means 55 in association with the sub-theme identification information = "4-2".
[0091] Then, when the number of topics K for theme 4 is 5, the topic vector calculation means 36 executes the above process for topic numbers 3, 4, and 5. Furthermore, this process is executed not only for theme 4, but also for the other themes 1 to 3 and 5 to 10.
[0092] <Configuration of Corporate Document Data Creation Method 37>
[0093] As shown in Figure 5, the corporate document data creation means 37 executes a process of dividing each of the integrated report data or other corporate document data of each company stored in the corporate document data storage means 56 (unlike questionnaire data, which is information about human capital management, including health management, related to the company, written in a free format for each company) into sentences at each period, and storing the multiple corporate document data obtained by the division in the corporate document data storage means 57 in association with company identification information (stock code, stock name, etc.) and company document identification information that identifies the corporate document data.
[0094] In this case, if the corporate document data such as the integrated report data is PDF data, the corporate document data creation means 37 must convert the corporate document data such as the integrated report data, which has a variety of formats depending on the company, into text before dividing the corporate document data into corporate document data. However, in a general method of converting PDF data into text data, problems occur such as garbled characters and character strings written in different blocks on the same page being connected in an unnatural manner across blocks (such as the inconvenience of each character string in different blocks (blocks aligned horizontally) that should not be connected being connected because they are located at the same line position on the same page). Therefore, the corporate document data creation means 37 applies OCR technology to the corporate document data such as the integrated report data, which is PDF data, to largely solve the problem of garbled characters and extract character strings for each block.
[0095] Here, it is preferable to adopt, for example, Azure Form Recognizer as the OCR technology. This is a service provision system using a cloud API. In addition, for converting English text, Textract from Amazon Web Services (AWS) or paddleOCR, which is an open source software (OSS) from Google, may be adopted.
[0096] <Configuration of the corporate document vector creation means 38>
[0097] 5, the company text vector creation means 38 executes a process of vectorizing each of the multiple company text data stored in the company text data storage means 57 to create a company text vector, and storing the multiple created company text vectors in the company text vector storage means 58 in association with company identification information (stock code, stock name, etc.) and company text identification information. The company text identification information corresponding to the company text vector is the same as the company text identification information corresponding to the original company text data.
[0098] At this time, the company statement vector creation means 38 vectorizes each of the multiple company statement data by a vectorization process using the same method (for example, Azure OpenAI Embeddings API text-embedding-ada-002 version2 in this embodiment) as the method used in the vectorization process of the answer statement data by the answer statement vector creation means 32. Therefore, if the answer statement vector is a 1,536-dimensional vector, the company statement vector will also be a 1,536-dimensional vector. In addition, since the topic vector is created using U answer statement vectors (for example, U=10) as shown in FIG. 4, the number of dimensions is the same as that of the answer statement vector, and it is a 1,536-dimensional vector. Therefore, the company statement vector and the topic vector have the same number of dimensions, and similarity such as cosine similarity can be calculated as shown in FIG. 5.
[0099] <Configuration of Similarity Calculation Means 39>
[0100] As shown in FIG. 5, the similarity calculation means 39 calculates the similarity (in this embodiment, cosine similarity) between the multiple company sentence vectors stored in the company sentence vector storage means 58 and the topic vector of each subtheme stored in the subtheme information storage means 55 for each subtheme, thereby obtaining an association score for each of the multiple company sentence data, and performs a process of storing the obtained association score in the association score storage means 59 in association with the subtheme identification information, company identification information (stock code, stock name, etc.), and company sentence identification information.
[0101] Therefore, any one piece of corporate text data has a relevance score for each of a plurality of sub-themes (each topic of each theme). Note that this relevance score may be calculated in advance and stored in a non-volatile memory, or may be calculated when the user refers to or searches the corporate text data. In the latter case, the relevance score storage means 59 may be the main memory.
[0102] <Configuration of representative association score calculation means 40>
[0103] As shown in Figure 5, the representative related score calculation means 40 extracts, for each sub-theme, the maximum related score from among the multiple related scores associated with the same company identification information using the related scores for each sub-theme stored in the related score storage means 59, and sets the extracted related score as the representative related score for the company of the company identification information, or extracts the top multiple (V numbers, for example, V=5) related scores with the highest values, and sets the average or weighted average or other calculated value calculated using the extracted top multiple (V numbers) related scores as the representative related score for the company of the company identification information, and stores the representative related score in the representative related score storage means 60 in association with the sub-theme identification information and company identification information (stock code, stock name, etc.).
[0104] Therefore, since any one company has multiple company sentence data, there is a relevance score for each sub-theme for each company sentence data. In other words, any one company has as many relevance scores as the number of company sentence data multiplied by the number of sub-themes. In contrast, any one company has one representative relevance score for each sub-theme. In other words, any one company has as many representative relevance scores as the number of sub-themes. For this reason, the representative relevance scores are used in the bird's-eye view display process and the initiative recommendation sub-theme extraction process by the output means 41.
[0105] <Configuration of output means 41>
[0106] The output means 41 executes various output processes, including the process of displaying the company sentence data on a screen, using the relevance scores for each sub-theme for the company sentence data of each company stored in the relevance score storage means 59.
[0107] More specifically, the output means 41 executes a bird's-eye view display process (see FIG. 7), a process for displaying a list of comparisons of company initiatives (see FIG. 8), a process for displaying each company's initiative statement (see FIG. 9), a process for searching for similar sentences (see FIG. 10), a process for displaying similar sub-themes (see FIG. 10), a process for displaying in order of integrated related scores (see FIG. 10 and FIG. 11), a process for displaying in order of adjusted related scores (see FIG. 10 and FIG. 11), a recommendation screen display process (see FIG. 12), a process for extracting initiative recommendation sub-themes (see FIG. 12), a process for displaying initiative statements of companies of interest (see FIG. 12), and a process for displaying recommended initiative statements (see FIG. 12), and also executes a process for displaying a factoid-type question screen 300 and a keyword search screen 310 (see FIG. 6).
[0108] A user can start a reference / search operation from any of the screens 100, 130, 160, 200, 230, 300, and 310 shown in Fig. 6. And, as shown by the dotted arrow in Fig. 6, a user can move from any screen to another screen. Also, as shown by the solid arrow in Fig. 6, a screen transition can be made by passing transition information from the previous screen to the next screen.
[0109] <Output means 41 / bird's-eye view display processing: Fig. 7, Fig. 6>
[0110] 7, an overhead view display screen 100 displayed by the output means 41 is provided with a display section 101 showing the status of efforts of each company with respect to each sub-theme using a plurality of cells arranged vertically and horizontally in a matrix. The vertical axis of this display section 101 is provided with a display section 102 of company identification information (such as stock name and stock code) of each company and a company selection section 103 for selecting a company displayed on this display section 102, and the horizontal axis is provided with a display section 104 of the sub-theme name of each sub-theme and a sub-theme selection section 105 for selecting a sub-theme displayed on this display section 104.
[0111] Each cell constituting the display unit 101 is colored, and the shade of the color indicates the magnitude of the value of the representative related score stored in the representative related score storage means 60, that is, the status (degree of the effort) of each company with respect to each sub-theme. The relationship between the shade of the color and the magnitude of the value of the representative related score is shown in the shade-score relationship display unit 106. The darker the color, the higher the value of the representative related score, and the better the effort status. Therefore, in the example of the display unit 101 in FIG. 7, when looking at the shade of the color for each sub-theme belonging to theme 10 on the right side (the same as theme 10 in FIG. 2), it can be seen that the overall shade is light, and therefore it can be seen that each company's effort with respect to each sub-theme of theme 10 is insufficient. However, among them, Company Y is darker in color compared to other companies, and therefore it can be seen that sufficient efforts are being made.
[0112] The sub-theme name displayed on the display unit 104 is the name stored in the sub-theme information storage means 55 in association with the sub-theme identification information.
[0113] 7, Company Y is highlighted 107. This highlighted display 107 indicates the company specified by the user in the company selection section 205 of the similar sentence search screen 200 (see FIGS. 6 and 10).
[0114] In the example of Figure 7, due to space constraints on the drawing, the relationships between all companies and all sub-themes are not shown, but in an actual overview diagram, the cells omitted by dotted lines are also filled in.
[0115] Furthermore, the bird's-eye view display screen 100 in FIG. 7 is provided with a "Display inter-company initiative comparison list" button 110 for moving to the inter-company initiative comparison list display screen 130 (see FIG. 8), a "Display each company's initiative statement" button 111 for moving to the each company initiative statement display screen 160 (see FIG. 9), a "Search for similar sentences" button 112 for moving to the similar sentence search screen 200 (see FIG. 10), a "Recommend" button 113 for moving to the recommendation screen 230 (see FIG. 12), a "Factoid question" button 114 for moving to the factoid question screen 300 (see FIG. 6), and a "Keyword search" button 115 for moving to the keyword search screen 310 (see FIG. 6).
[0116] On the bird's-eye view display screen 100 in Fig. 7, when the user selects two companies in the company selection section 103 and one sub-theme in the sub-theme selection section 105, and then presses the "Display inter-company initiative comparison list" button 110, the company identification information of the two selected companies and the sub-theme identification information of one sub-theme are used in the inter-company initiative comparison list display process (see Fig. 8) (see Fig. 6). Note that even if the "Display inter-company initiative comparison list" button 110 is pressed without selecting anything, it is possible to move to the inter-company initiative comparison list display screen 130 (see Fig. 8), but the user will perform a selection operation on that screen 130 for the desired display.
[0117] In addition, when the user selects one sub-theme in the sub-theme selection section 105 on the bird's-eye view display screen 100 in Fig. 7 and presses the "Display each company's commitment statement" button 111, the sub-theme identification information of the selected one sub-theme is used in the each company's commitment statement display process (see Fig. 9) (see Fig. 6). Note that even if the "Display each company's commitment statement" button 111 is pressed without selecting anything, it is possible to move to the each company's commitment statement display screen 160 (see Fig. 9), but the user will perform a selection operation on that screen 160 for the desired display.
[0118] Furthermore, when the user selects one company in the company selection section 103 on the bird's-eye view display screen 100 in Fig. 7 and then presses the "recommend" button 113, the company identification information of the selected company (the company of interest that needs improvement efforts) is used in the recommendation screen display process (see Fig. 12) (see Fig. 6). Note that even if the "recommend" button 113 is pressed without selecting anything, it is possible to move to the recommendation screen 230 (see Fig. 12), but the user will perform a selection operation on that screen 230 to display the desired content.
[0119] Although not shown in the figure, there is an industry-specific storage means for storing the association between company identification information (stock code, stock name, etc.) and industry identification information, so that the user can specify the industry for which he / she wants to refer to and search for information, and thus an overview of the desired industry, for example, an overview of each company in the securities and financial industry, can be displayed on the screen. In addition, company document data such as integrated report data, company document data created from this company document data, company document vectors, association scores, and representative association scores may be prepared for each industry (associated with industry identification information) or for only a certain industry. In that case, the company document data storage means 56, company document data storage means 57, company document vector storage means 58, association score storage means 59, and representative association score storage means 60 may be prepared for each industry (associated with industry identification information) or for only a certain industry.
[0120] <Output means 41 / Company-to-company initiative comparison list display process: Fig. 8, Fig. 6>
[0121] In Figure 8, the inter-company initiative comparison list display screen 130 displayed by the output means 41 includes a theme selection section 131, a sub-theme selection section 132, a company selection section 133 for selecting a company (company name), a comparison company selection section 134 for selecting a company (company name) to compare, a threshold setting section 135 for setting a threshold for the relevance score of the text (company text data) to be displayed, a company text data display section 140 and an initiative list display section 141 of the selected company (Company B in the example of Figure 8), and a company text data display section 142 and an initiative list display section 143 of the selected comparison company (Company X in the example of Figure 8).
[0122] In the theme selection section 131 and the subtheme selection section 132, if the user specifies a subtheme in the subtheme selection section 105 of the bird's-eye view display screen 100 in Fig. 7, the subtheme and the theme to which it belongs are automatically selected. In the company selection section 133 and the comparison company selection section 134, if the user specifies two companies in the company selection section 103 of the bird's-eye view display screen 100 in Fig. 7, those two companies are automatically selected.
[0123] Company sentence data having an association score equal to or exceeding the threshold value specified by the threshold setting unit 135 is displayed in the company sentence data display units 140, 142. In this case, when displaying all of the company sentence data having an association score equal to or exceeding the threshold value, the part that does not fit within the display frame can be viewed by scrolling. On the other hand, among the company sentence data having an association score equal to or exceeding the threshold value, only the company sentence data having the highest association score (for example, the top 10) of a predetermined number may be displayed.
[0124] The output means 41 displays on the company sentence data display unit 140 the related scores that satisfy the conditions set by the threshold setting unit 135 among the related scores stored in the related score storage means 59 (see FIG. 5) in association with the company identification information of the company selected by the user in the company selection unit 133 and the subtheme identification information of the subtheme selected by the user in the subtheme selection unit 132, in descending order of related scores. In addition, the output means 41 acquires the company identification information and company sentence identification information associated with the related scores from the related score storage means 59, and displays the company sentence data and year associated with the company identification information and the company sentence identification information stored in the company sentence data storage means 57 (see FIG. 5) in association with the related scores in the same row of the company sentence data display unit 140. Note that the display of the year may be omitted.
[0125] Similarly, the output means 41 displays on the company sentence data display section 142 the related scores that satisfy the conditions set by the threshold setting section 135 among the related scores stored in the related score storage means 59 (see FIG. 5) in association with the company identification information of the comparison company selected by the user in the comparison company selection section 134 and the subtheme identification information of the subtheme selected by the user in the subtheme selection section 132, in descending order of related scores. In addition, the output means 41 acquires the company identification information and company sentence identification information associated with the related scores from the related score storage means 59, and displays the company sentence data and year associated with the company identification information and the company sentence identification information stored in the company sentence data storage means 57 (see FIG. 5) in the same row of the company sentence data display section 142 in correspondence with the related scores. Note that the display of the year may be omitted.
[0126] In addition, the output means 41 passes the company statement data with the highest relevance scores (for example, the top 10) for each of the two companies displayed in the company statement data display sections 140, 142 to the company statement acquisition means 42, and then receives the multiple statement data in list format from the company statement acquisition means 42, and displays the received multiple statement data in list format in the statement list display sections 141, 143.
[0127] Furthermore, the inter-company initiative comparison list display screen 130 in FIG. 8 is provided with a "bird's-eye view display" button 150 for moving to the bird's-eye view display screen 100 (see FIG. 7), a "display each company's initiative statement" button 151 for moving to the each company's initiative statement display screen 160 (see FIG. 9), a "similar sentence search" button 152 for moving to the similar sentence search screen 200 (see FIG. 10), a "recommend" button 153 for moving to the recommendation screen 230 (see FIG. 12), a "factoid question" button 154 for moving to the factoid question screen 300 (see FIG. 6), and a "keyword search" button 155 for moving to the keyword search screen 310 (see FIG. 6).
[0128] When one of the multiple attempted sentences displayed in list format in the attempted sentence list display areas 141, 143 is selected and clicked, the attempted sentence data is used in the similar sentence search process (see FIG. 10) (see FIG. 6).
[0129] <Output means 41 / Company approach text display processing: Figure 9, Figure 6>
[0130] In FIG. 9, a company initiative document display screen 160 displayed by the output means 41 includes a theme selection section 161, a sub-theme selection section 162, a company initiative document data display section 170, and an initiative list display section 171.
[0131] Furthermore, if the user specifies a sub-theme to be selected in the sub-theme selection section 105 of the overhead view display screen 100 in FIG. 7, the theme selection section 161 and the sub-theme selection section 162 are automatically selected for that sub-theme and the theme to which it belongs.
[0132] The company statement data display section 170 displays the fiscal year, company identification information (stock name, stock code, etc.), the overall evaluation score and overall deviation value of the company's efforts by the Ministry of Economy, Trade and Industry, the related scores for the sub-theme selected in the sub-theme selection section 162, and the text (company statement data) in a corresponding state. In addition to or instead of the overall evaluation score and overall deviation value, evaluation scores and deviation values by other evaluators / evaluation organizations (whether they are the government, local government, or private) for the company's efforts may be displayed. Evaluation scores and deviation values by other evaluators / evaluation organizations include, for example, well-being reports in which each company is scored in multiple items.
[0133] Unlike the inter-company initiative comparison list display screen 130 in FIG. 8, the company-by-company initiative document display screen 160 in FIG. 9 does not specify a company, so the company document data display section 170 displays a mixture of initiative details (company document data) for multiple companies for the sub-theme selected in the sub-theme selection section 162.
[0134] Specifically, the output means 41 displays the related scores stored in the related score storage means 59 (see FIG. 5) in association with the subtheme identification information of the subtheme selected by the user in the subtheme selection unit 162 in descending order of value on the company sentence data display unit 170. The output means 41 also acquires the company identification information and company sentence identification information associated with the related scores from the related score storage means 59, and displays the company sentence data and year associated with the company identification information and company sentence identification information stored in the company sentence data storage means 57 (see FIG. 5) in the same row on the company sentence data display unit 170 in correspondence with the related scores. Note that the display of the year may be omitted.
[0135] Furthermore, the output means 41 acquires the overall evaluation score, the overall deviation value, or other evaluation scores or deviation values stored in the questionnaire data storage means 51 or the company document data storage means 56 in association with the company identification information, and displays them on the company document data display section 170. Note that for companies that have not submitted a questionnaire, there is no data for the overall evaluation score or the overall deviation value, so these are left blank.
[0136] In addition, the output means 41 passes the top multiple (e.g., top 10) corporate text data items with the highest related scores displayed in the corporate text data display unit 170 to the corporate commitment text acquisition means 42, and then receives the multiple commitment text data items in list format from the corporate commitment text acquisition means 42, and displays the received multiple commitment text data items in list format in the commitment list display unit 171.
[0137] Furthermore, the company engagement statement display screen 160 in Figure 9 is provided with a "bird's-eye view display" button 180 for moving to the bird's-eye view display screen 100 (see Figure 7), a "company engagement comparison list display" button 181 for moving to the company engagement comparison list display screen 130 (see Figure 8), a "similar sentence search" button 182 for moving to the similar sentence search screen 200 (see Figure 10), a "recommend" button 183 for moving to the recommendation screen 230 (see Figure 12), a "factoid question" button 184 for moving to the factoid question screen 300 (see Figure 6), and a "keyword search" button 185 for moving to the keyword search screen 310 (see Figure 6).
[0138] When any one of the multiple attempted sentences displayed in list format in the attempted sentence list display section 171 is selected and clicked, the attempted sentence data is used for the similar sentence search process (see FIG. 10) (see FIG. 6).
[0139] <Output means 41 / Similar sentence search process, Similar sub-theme display process, Integrated related score order display process, Adjusted related score order display process: Figs. 10, 11, 6>
[0140] In FIG. 10, the similar sentence search screen 200 displayed by the output means 41 includes a search string input section 201 that accepts input of any search string data (text data) by the user, a "Search execution" button 202 that starts a search for company sentence data using the search string data inputted in the search string input section 201, a display section 203 for the input search string (the search string actually used in the search by pressing the "Search execution" button 202), a company sentence data display section 204, a company selection section 205 in the company sentence data display section 204, and an initiative list display section 206.
[0141] In addition, if the user selects and specifies initiative sentence data from the initiative list display sections 141, 143, 171, 206, 242, 243 on the inter-company initiative comparison list display screen 130 in FIG. 8, the individual company initiative sentence display screen 160 in FIG. 9, the recommendation screen 230 in FIG. 12, or the similar sentence search screen 200 itself in FIG. 10, the search string input section 201 will be automatically input with the selected initiative sentence data as search string data.
[0142] In the example of Fig. 10, the search string input section 201 states "Please enter the string (sentence or word) you want to search for similarity from the corporate document (integrated report, etc.)", but the search string actually used for the search by pressing the "Search" button 202 ("Wearable device" in the example of Fig. 10) is displayed in the search string display section 203, so the display prompting the input of this search string is a display prompting the input of the next search string data. If the user selects and specifies the target sentence data in the target sentence list display section 206 of the similar sentence search screen 200 of Fig. 10, the selected and specified target sentence data is automatically input into the search string input section 201, and when the "Search" button 202 is pressed in this state, the search process is executed with the selected and specified target sentence data as new search string data, so that repeated search processes are possible only within the similar sentence search screen 200 of Fig. 10 (see Fig. 6).
[0143] In the company sentence data display section 204 in Fig. 10, the year, company identification information (stock name, stock code, etc.), relevance score, and text (company sentence data) are displayed in a corresponding state. However, unlike the company sentence data display sections 140, 142 in Fig. 8 and the company sentence data display section 170 in Fig. 9, the relevance score in the initial display (display before sorting) in the company sentence data display section 204 in Fig. 10 is not the relevance score for the subtheme for each company sentence data, but the relevance score for the search string data for each company sentence data, and in this application (particularly, in the claims of this application), this is called the input string relevance score and is distinguished from the relevance score for the subtheme.
[0144] As shown in FIG. 11, the output means 41 executes a similar sentence search process in the initial display (display before sorting). In this similar sentence search process, the search string input unit 201 accepts input of any search string data by the user (including cases where the action sentence data selected and specified by the user in the action list display units 141, 143, 171, 206, 242, and 243 in Figures 8, 9, 10, and 12 is automatically input as search string data), and the input search string data is vectorized by a vectorization process using the same method (in this embodiment, for example, Azure OpenAI Embeddings API text-embedding-ada-002 version2) as the method used in the vectorization process of the answer sentence data, and the similarity (in this embodiment, cosine similarity) between the obtained search string vector and each of the multiple company sentence vectors (i.e., the multiple company sentence vectors stored in the company sentence vector storage means 58) obtained by vectorizing the multiple company sentence data stored in the company sentence data storage means 57 is calculated, thereby obtaining an input string related score for the search string data for each of the multiple company sentence data, and displaying the company sentence data in the company sentence data display unit 204 in order of the obtained input string related score. The number of items displayed in the company sentence data display unit 204 may be a predetermined number (e.g., the top 10 items), or the top 10 company sentence data may be displayed first, and then company sentence data with high input string related scores may be displayed one after another, for example, 10 items at a time, or all company sentence data may be displayed from highest to lowest input string related scores by scrolling the screen.
[0145] More specifically, the output means 41 displays the obtained input string related scores in the "related score" column of the company sentence data display section 204 in descending order of value, and displays in the same row of the company sentence data display section 204 the company sentence data, year, and company identification information (stock name, stock code, etc.) associated with the company sentence identification information corresponding to the company sentence vector used to obtain the input string related score and stored in the company sentence data storage means 57 (see Figure 5).
[0146] In addition, the output means 41 passes the top multiple (e.g., top 10) corporate text data items with the highest input string related scores displayed in the corporate text data display unit 204 to the corporate commitment text acquisition means 42, and then receives the multiple commitment text data items in list format from the corporate commitment text acquisition means 42, and displays the received multiple commitment text data items in list format in the commitment list display unit 206.
[0147] Furthermore, the similar sentence search screen 200 in FIG. 10 is provided with a similar subtheme display section 210 for displaying subthemes similar to the search string data displayed in the search string display section 203, i.e., the search string data actually used for the search by pressing the “Execute search” button 202 (in the example of FIG. 10, the weights are 0%, 10%, 40%, and 20%, but the input does not necessarily have to be in the form of a percentage value), a similar subtheme weight input section 211 for inputting the weights of the search string data displayed in the search string display section 203 (in the example of FIG. 10, the weights are 30%, but the input does not necessarily have to be in the form of a percentage value), a display order selection section 213 for sorting the display order of the company sentence data in the company sentence data display section 204 in various ways, and a “Execute sort” button 214 for changing the display order of the company sentence data.
[0148] After performing the similar sentence search process for the first display (display before sorting), the output means 41 performs the similar subtheme display process as shown in Fig. 11. In this similar subtheme display process, the similarity (in this embodiment, cosine similarity) between the search character string vector and each of the topic vectors for a plurality of subthemes stored in the subtheme information storage means 55 is calculated, and subthemes of topic vectors whose similarity to the search character string vector is greater than or equal to a predetermined threshold value are extracted, or subthemes of topic vectors whose similarity is the top of a predetermined number (for example, the top five) are extracted, and the extracted subthemes are displayed on the screen of the similar subtheme display unit 210 to present to the user. In the example of Figs. 10 and 11, four subthemes, subthemes [1-2], [1-1], [4-1], and [4-4], are presented.
[0149] After executing the similar subtheme display process, the output means 41 changes the display order of the company sentence data in the company sentence data display section 204 according to the method selected by the user in the display order selection section 213 of Fig. 10. In the display order selection section 213, the user can select from a method of sorting by taking into account the relevance to the subtheme according to the input designated ratio (weight) by the user in the similar subtheme weight input section 211 and the search string weight input section 212, a method of sorting only by the similarity to each similar subtheme displayed in the similar subtheme display section 210 (in the example of Fig. 10, there are methods of sorting only by the similarity to the subtheme "1-2 Health Promotion", only by the similarity to the subtheme "1-1 Strengthening Health Management", only by the similarity to the subtheme "4-1 Mental Health Measures", and only by the similarity to the subtheme "4-4 Stress Relief Initiatives"). Alternatively, the user can select a display order based on the similarity to the input search string data (initial display order).
[0150] In the display order selection section 213, when a method of sorting is selected that takes into account the relevance to the subtheme at a ratio (weight) input by the user and the "sort execution" button 214 is pressed, the user inputs a ratio (weight) into the similar subtheme weight input section 211 and the search string weight input section 212. In this case, the output means 41 executes an integrated related score order display process or an adjusted related score order display process as shown in FIG. 11. Here, the ratio (weight) input into the search string weight input section 212 is α (30% in the example of FIG. 10), and the ratios (weights) of the subthemes [1-2], [1-1], [4-1], and [4-4] input into the similar subtheme weight input section 211 are β (0% in the example of FIG. 10), γ (10% in the example of FIG. 10), δ (40% in the example of FIG. 10), and ε (20% in the example of FIG. 10). If the user selects a method of sorting by taking into account the relevance to the sub-theme based on the ratios (weights) input by the user without inputting these ratios (weights), and then presses the "Sort" button 214, each ratio (weight) is considered to be equal. In other words, a=β=γ=δ=ε=20%.
[0151] In FIG. 11, when sorting is performed by executing the integrated related score order display process by the output means 41, the input string related score obtained by the similar sentence search process for the initial display (display before sorting) and the related score stored in the related score storage means 59 (see FIG. 5) in association with the subtheme identification information of at least one subtheme selected by the user from at least one subtheme (four subthemes in the example of FIG. 10) presented by the similar subtheme display process (in the example of FIG. 10, the proportion (weight) of the subtheme [1-2] is input as 0%, so the remaining three subthemes [1-1], [4-1], and [4-4] are selected.) are used to obtain an average value or a weighted average value or other calculated value (in this embodiment, a weighted average value) for each of the multiple company sentence data using the input string related score and at least one related score, and this is used as the integrated related score. The company sentence data is sorted in descending order of the integrated related score and displayed on the company sentence data display unit 204.
[0152] That is, in FIG. 11, if four sub-themes [1-2], [1-1], [4-1], and [4-4] are selected, the relevance score storage means 59 (see FIG. 5) stores the relevance score for the sub-theme [1-2] for each of the multiple company sentence data (similarity between each of the multiple company sentence vectors and the topic vector of the sub-theme [1-2]), and similarly, the relevance score for each of the multiple company sentence data for each of the sub-themes [1-1], [4-1], and [4-4] (similarity between each of the multiple company sentence vectors and each of the topic vectors of the sub-themes [1-1], [4-1], and [4-4]) is stored. Therefore, the calculation formula for the integrated relevance score for any one company sentence data is as follows.
[0153] Integrated relevance score = {Input string relevance score × α + relevance score for subtheme [1-2] × β + relevance score for subtheme [1-1] × γ + relevance score for subtheme [4-1] × δ + relevance score for subtheme [4-4] × ε} / (α + β + γ + δ + ε)
[0154] In addition, in FIG. 11 , when sorting is performed by executing the display process in order of adjusted related scores by the output means 41, the search character string vector obtained by the similar sentence search process for the initial display (display before sorting) and at least one topic vector (in the example of FIG. 10 , the proportion (weight) of the subtheme [1-2] is input as 0%, so the remaining three subthemes [1-1], [4-1], and [4-4] are selected) associated with the subtheme identification information of at least one subtheme selected by the user from at least one subtheme (four subthemes in the example of FIG. 10 ) presented by the similar subtheme display process are stored in the subtheme information storage means 55. Then, using the topic vectors of the remaining three sub-themes [1-1], [4-1], and [4-4]), the average value or weighted average value or other calculated value (weighted average value in this embodiment) of the elements of each vector is calculated to create a composite vector, and the similarity between this composite vector and each of the multiple company sentence vectors (i.e., multiple company sentence vectors stored in the company sentence vector storage means 58) obtained by vectorizing multiple company sentence data stored in the company sentence data storage means 57 (see FIG. 5) is calculated to obtain an adjusted related score for each of the multiple company sentence data, and the company sentence data is sorted in descending order of the adjusted related score and displayed on the company sentence data display unit 204. The calculation formula for the composite vector is as follows.
[0155] Composite vector = {search string vector × α + topic vector of subtheme [1-2] × β + topic vector of subtheme [1-1] × γ + topic vector of subtheme [4-1] × δ + topic vector of subtheme [4-4] × ε} / (α + β + γ + δ + ε)
[0156] In the display order selection unit 213, when the user selects a method of sorting only by the similarity to each similar sub-theme displayed in the similar sub-theme display unit 210 and presses the "sort execution" button 214, the user does not need to input a ratio (weight) in the similar sub-theme weight input unit 211 and the search string weight input unit 212. In this case, if the user selects a method of sorting only by the similarity to the sub-theme "1-2 Health Promotion", for example, the output means 41 sorts the company sentence data in descending order of the relevance scores (similarity between each of the multiple company sentence vectors and the topic vector of the sub-theme [1-2]) for each of the multiple company sentence data, and displays them in the company sentence data display unit 204. The same applies when the user selects a method of sorting only by the similarity to the sub-themes [1-1], [4-1], and [4-4].
[0157] Furthermore, the similar sentence search screen 200 in FIG. 10 is provided with a "bird's-eye view display" button 220 for moving to the bird's-eye view display screen 100 (see FIG. 7), a "display inter-company initiative comparison list" button 221 for moving to the inter-company initiative comparison list display screen 130 (see FIG. 8), a "display each company's initiative statement" button 222 for moving to the each company initiative statement display screen 160 (see FIG. 9), a "recommend" button 223 for moving to the recommendation screen 230 (see FIG. 12), a "factoid question" button 224 for moving to the factoid question screen 300 (see FIG. 6), and a "keyword search" button 225 for moving to the keyword search screen 310 (see FIG. 6).
[0158] When the user selects two companies in the company selection section 205 and presses the "Company-to-Company Partnership Comparison List Display" button 221, the company identification information of the selected companies is used for the company-to-company partnership comparison list display process (see FIG. 8) (see FIG. 6). When the user selects three or more companies in the company selection section 205 and presses the "Bird's-eye view display" button 220, the company identification information of the selected companies is used for the highlight display 107 (see FIG. 7) in the bird's-eye view display process (see FIG. 6).
[0159] <Output means 41 / recommendation screen display process, initiative recommendation sub-theme extraction process, focus company initiative text display process, recommended initiative text display process: FIG. 12>
[0160] In FIG. 12, a recommendation screen 230 displayed by the output means 41 includes a company selection section 231 for selecting a company that needs to improve its efforts (a company of interest), an "Execute" button 232 for deciding a recommended initiative sub-theme to be recommended as a sub-theme that the company that needs to improve its efforts (a company of interest) selected in the company selection section 231 should work on, an initiative recommendation sub-theme display section 233 for displaying the initiative recommendation sub-theme decided by pressing down the "Execute" button 232, a sub-theme selection section 234 for accepting two or more selections by the user from the initiative recommendation sub-themes displayed in the initiative recommendation sub-theme display section 233, and various methods for displaying the current initiative details (company statement data) of the company that needs to improve its efforts (a company of interest) and the recommendation initiative statement (company statement data) for the company that needs to improve its efforts (a company of interest). The display unit 200 is provided with a selection unit 235, a selection unit 236 for selecting to limit the display of recommendation initiative letters (company initiative data) to texts (company initiative data) of companies having an overall evaluation score, a "display" button 237 for displaying the current initiative details (company initiative data) of the company in need of improvement (focus company) and the recommendation initiative letters (company initiative data) for the company in need of improvement (focus company), a company initiative data display unit 240 for displaying the current initiative details (company initiative data) of the company in need of improvement (focus company) after pressing the "display" button 237, a recommendation initiative display unit 241 for displaying the recommendation initiative letters (company initiative data) for the company in need of improvement (focus company), an initiative list display unit 242 for the company in need of improvement (focus company) and an initiative list display unit 243 for the recommendation initiative letters (company initiative data).
[0161] The recommendation screen display process for displaying the recommendation screen 230 in FIG. 12 executed by the output means 41 includes the following initiative recommendation sub-theme extraction process, focus company initiative statement display process, and recommendation initiative statement display process.
[0162] In the company selection section 231, if the user selects a company requiring improvement efforts (a company of interest) in the company selection section 103 of the bird's-eye view display screen 100 in FIG.
[0163] When the user presses the “Execute” button 232, the recommended initiative sub-theme extraction process is executed, and the recommended initiative sub-themes recommended as sub-themes that the company in need of improvement (the company of interest) should tackle are displayed in the recommended initiative sub-theme display section 233.
[0164] When executing the recommended initiative sub-theme extraction process, the output means 41 accepts the user's selection of a company requiring improvement (focus company) in the company selection section 231 (including the case of automatic selection by the user's selection specification in the company selection section 103 of the overhead view display screen 100 of Figure 7), and uses the representative related score stored in the representative related score storage means 60 (see Figure 5) to calculate, for each sub-theme, the difference between the representative related score of the selected focus company and the largest representative related score among the representative related scores of companies other than the focus company, or the difference between the representative related score of the selected focus company and the average or other calculated value of the top representative related scores (for example, a specified number such as the top 5) of the representative related scores of companies other than the focus company, thereby calculating the deviation of the representative related score of the focus company from the representative related score of companies other than the focus company, and extracts the top multiple sub-themes (for example, a specified number such as the top 5) with the largest calculated deviation as recommended initiative sub-themes and displays them on the screen in the recommended initiative sub-theme display section 233.
[0165] The sub-theme selection section 234 is used when selecting two or more sub-themes from the initiative recommendation sub-themes displayed in the initiative recommendation sub-theme display section 233 and displaying the company statement data display section 240 and the recommended initiative display section 241. Therefore, when selecting one sub-theme and displaying the company statement data display section 240 and the recommended initiative display section 241, it is not necessary to make a selection in the sub-theme selection section 234.
[0166] In the selection unit 235, the user can select a display method for the company text data display unit 240 and the recommended initiative text display unit 241 from among a method in which two or more subthemes selected in the subtheme selection unit 234 are mixed and displayed in descending order of relevance score, and a method in which the subthemes displayed in the initiative recommendation subtheme display unit 233 are displayed in descending order of relevance score for one of the subthemes (in the example of Figure 12, there are a method of displaying in descending order of relevance score for the subtheme ``[1-2] Health promotion'', a method of displaying in descending order of relevance score for the subtheme ``1-1] Strengthening health management'', a method of displaying in descending order of relevance score for the subtheme ``[4-1] Mental health measures'', and a method of displaying in descending order of relevance score for the subtheme ``[4-4] Stress relief efforts'').
[0167] In addition, since there are companies that have an overall evaluation score or an overall standard deviation (companies that have submitted a questionnaire) and companies that do not have these (companies that have not submitted a questionnaire), by checking the selection section 236, it is possible to narrow down the display to the corporate document data of companies that have an overall evaluation score or an overall standard deviation.
[0168] After the user selects the display method in the selection unit 235, when the user presses the "Display" button 237, the output means 41 executes a process for displaying the target company's initiative statement, which causes it to be displayed in the company statement data display unit 240 and the initiative list display unit 242, and also executes a process for displaying a recommended initiative statement, which causes it to be displayed in the recommended initiative display unit 241 and the initiative list display unit 243.
[0169] When the output means 41 executes the target company initiative statement display process according to the display method selected by the user in the selection unit 235, the output means 41 displays the target company's company statement data and year stored in the company statement data storage means 57 (see FIG. 5) together with the associated scores on the screen of the company statement data display unit 240 in order of the highest associated scores for the initiative recommendation sub-theme selected by the user for the target company's company statement data stored in the associated score storage means 59 (see FIG. 5). At this time, if the initiative recommendation sub-theme selected by the user is two or more sub-themes, the display order is determined with the associated scores for each sub-theme mixed.
[0170] In addition, when the output means 41 executes the recommended initiative display process according to the display method selected by the user in the selection unit 235, the company statement data and years of companies other than the focused company stored in the company statement data storage means 57 (see FIG. 5) are displayed on the recommended initiative display unit 241 in order of the highest related score for the initiative recommendation sub-theme selected by the user for the company statement data of companies other than the focused company stored in the related score storage means 59 (see FIG. 5) together with company identification information (stock name, stock code, etc.) and related score. At this time, if the initiative recommendation sub-theme selected by the user is two or more sub-themes, the display order is determined in a state where the related scores for each sub-theme are mixed. In addition, since there are multiple companies other than the focused company, the display order is determined in a state where the related scores for each sub-theme for the company statement data of multiple companies are mixed.
[0171] Then, the output means 41 passes the company statement data of the top multiple (for example, top 10) in the magnitude of the related score displayed in the company statement data display unit 240 to the company commitment statement acquisition means 42, and then receives the multiple commitment statement data in list form from the company commitment statement acquisition means 42, and displays the received multiple commitment statement data in list form on the commitment list display unit 242. Similarly, the output means 41 passes the company statement data of the top multiple (for example, top 10) in the magnitude of the related score displayed in the recommended commitment statement display unit 241 to the company commitment statement acquisition means 42, and then receives the multiple commitment statement data in list form from the company commitment statement acquisition means 42, and displays the received multiple commitment statement data in list form on the commitment list display unit 243. Note that, due to the space constraints of the drawing, only one commitment statement data is displayed on each of the commitment list display units 242 and 243, but multiple commitment statement data are displayed in list form in each, as in the case of the commitment list display units 141 and 143 of FIG. 8.
[0172] Furthermore, the recommendation screen 230 in FIG. 12 is provided with a "bird's-eye view display" button 250 for moving to the bird's-eye view display screen 100 (see FIG. 7), a "display inter-company initiative comparison list" button 251 for moving to the inter-company initiative comparison list display screen 130 (see FIG. 8), a "display each company's initiative statement" button 252 for moving to the each company initiative statement display screen 160 (see FIG. 9), a "search for similar sentences" button 253 for moving to the similar sentence search screen 200 (see FIG. 10), a "factoid question" button 254 for moving to the factoid question screen 300 (see FIG. 6), and a "keyword search" button 255 for moving to the keyword search screen 310 (see FIG. 6).
[0173] When one of the multiple attempted sentences displayed in list format in the attempted sentence list display areas 242, 243 is selected and clicked, the attempted sentence data is used in the similar sentence search process (see FIG. 10) (see FIG. 6).
[0174] <Display process of output means 41 / factoid question screen 300: FIG. 6>
[0175] The factoid question screen 300 displayed by the output means 41 is provided with an input section for the user to input factoid question data (text data asking who, what, when, where, how much, etc.). The output means 41 transmits the factoid question data inputted into this input section to a system that provides an external service (not shown) via the network 1, and receives reply data from the system that provides the external service. The factoid question screen 300 is provided with a reply data display section that displays the received reply data on screen.
[0176] <Display process of output means 41 / keyword search screen 310: FIG. 6>
[0177] The keyword search screen 310 displayed by the output means 41 is provided with an input section for keyword search data (text data) by the user. The output means 41 transmits the keyword search data inputted into this input section to a system that provides an external service (not shown) via the network 1, and receives response data from the system that provides the external service. The keyword search screen 310 is provided with a response data display section that displays the received response data on the screen.
[0178] <Composition of Corporate Initiative Acquisition Method 42>
[0179] When the output means 41 executes the inter-company initiative comparison list display process (see FIG. 8), the individual company initiative display process (see FIG. 9), the similar sentence search process (see FIG. 10), and the recommendation screen display process (see FIG. 12), the company initiative acquisition means 42 receives from the output means 41 the top multiple (e.g., top 10) company sentence data with the highest related score displayed on the screen or the input string related score in the case of FIG. 10, creates company initiative creation request data containing a request to create multiple initiative sentences showing the company initiatives in list format using the received top multiple (e.g., top 10) company sentence data and these multiple company sentence data, inputs the created company initiative creation request data into a chat generative pretrained transformer (ChatGPT) or other large-scale language model (LLM), causes the large-scale language model (LLM) to output multiple initiative sentence data in list format, and passes the output list-format multiple initiative sentence data to the output means 41.
[0180] Specifically, the company commitment statement acquisition means 42 creates company commitment statement creation request data (text data) such as, for example, "Please output company commitment statements from the text below in a list format separated by commas, with no more than 20 characters. {message}", as shown in Figure 14, transmits the created company commitment statement creation request data to a service provision system 70 using a large scale language model (LLM) via network 1, receives data returned from the large scale language model, and passes the received data to the output means 41.
[0181] Here, the {message} in the company commitment statement creation request data is passed to the large-scale language model (LLM) with the top 10 text data (top 10 company sentence data) for each company displayed on the inter-company commitment comparison list display screen 130 (see FIG. 8) input. In other words, the company commitment statement creation request data is created for each company. The same is true for the company commitment statement display screen 160 (see FIG. 9) and the similar sentence search screen 200 (see FIG. 10). However, in the case of these screens 160 and 200, a predetermined number of top text data (for example, the top 10 company sentence data) for a mixture of multiple companies is input, rather than for each company, and passed to the large-scale language model (LLM).
[0182] The same is true for the recommendation screen 230 (see FIG. 12). However, in the case of this recommendation screen 230, when displaying on the initiative list display unit 242 as the target company initiative display process executed by the output means 41, the top 10 text data (top 10 company text data) for one company that is the target company is input and passed to the large-scale language model (LLM), while when displaying on the initiative list display unit 243 as the recommended initiative display process, the top predetermined number of text data (for example, top 10 company text data) in a state where multiple companies are mixed is input and passed to the large-scale language model (LLM).
[0183] <Configuration of the questionnaire data storage means 51>
[0184] The questionnaire data storage means 51 stores the health management questionnaire data (see FIG. 2) or other questionnaire data in association with company identification information (stock code, stock name, etc.). In addition to the questionnaire data, it also stores information on the fiscal year, the overall evaluation score, the overall deviation value, or other evaluation scores and deviation values.
[0185] <Configuration of the reply sentence data storage means 52>
[0186] As shown in FIG. 3, the answer message data storage means 52 stores answer message data for each theme of each company in association with theme identification information, company identification information (stock code, stock name, etc.), and answer message identification information.
[0187] <Configuration of reply sentence vector storage means 53>
[0188] The answer message vector storage means 53 stores the answer message vector (see FIG. 4) obtained by vectorizing the answer message data of each theme of each company (in this embodiment, 1,536 dimensions as an example), in association with theme identification information, company identification information (stock code, stock name, etc.), and answer message identification information.
[0189] <Configuration of topic model storage means 54>
[0190] The topic model storage means 54 stores, as a topic model for each theme, the topic values (topic distribution indicated by vertical vector π(i)) indicating the occurrence probability of each topic in the answer sentence data (i=1 to n: n is the number of answer sentence data) obtained by executing a topic estimation process (LDA in this embodiment) for each theme by the topic estimation means 34, and the occurrence probability of each word in each topic (matrix β).
[0191] <Configuration of the sub-theme information storage means 55>
[0192] The subtheme information storage means 55 stores topic vectors expressing each topic in each theme, and subtheme names, which are the names of each topic in each theme, in association with subtheme identification information.
[0193] <Configuration of the enterprise document data storage means 56>
[0194] The corporate document data storage means 56 stores integrated report data or other corporate document data of each company in association with corporate identification information (stock code, stock name, etc.) In addition to the corporate document data, it also stores information on the fiscal year, overall evaluation score, overall deviation value, or other evaluation score or deviation value.
[0195] In addition, the companies that have entered corporate document data such as integrated report data (the companies that have collected integrated report data, etc. to build the dashboard system 10) and the companies that have entered responses to questionnaire data such as the health management questionnaire data (see Figure 2) (the companies that used the response data to perform topic estimation, approximately 2,000 companies in 2022) do not have to match, and the number of companies that have entered corporate document data such as integrated report data may be greater, or vice versa. Since the questionnaire data is used for topic estimation processing, the number of companies that have entered responses to the questionnaire data needs to be a number that can obtain a large number of response sentence data to perform appropriate topic estimation. On the other hand, the number of companies that are the targets of collection of corporate document data needs to be a number that satisfies users who refer to and search for information, but in order for this system to function, corporate document data from a minimum of two companies is sufficient.
[0196] <Configuration of the enterprise document data storage means 57>
[0197] As shown in FIG. 5, the company text data storage means 57 stores the company text data created by the company text data creation means 37 in association with company identification information (stock code, stock name, etc.) and company text identification information.
[0198] <Configuration of the corporate text vector storage means 58>
[0199] As shown in FIG. 5, the company text vector storage means 58 stores the company text vector created by the company text vector creation means 38 in association with company identification information (stock code, stock name, etc.) and company text identification information.
[0200] <Configuration of the relevance score storage means 59>
[0201] As shown in Figure 5, the related score storage means 59 stores the related score for each sub-theme calculated by the similarity calculation means 39 in association with the sub-theme identification information, company identification information (stock code, stock name, etc.), and company statement identification information.
[0202] <Configuration of representative association score storage means 60>
[0203] As shown in FIG. 5, the representative related score storage means 60 stores the representative related scores calculated by the representative related score calculation means 40 in association with the subtheme identification information and company identification information (stock code, stock name, etc.).
[0204] <Processing flow by dashboard system 10>
[0205] In this embodiment, the dashboard system 10 performs a display process of information related to human capital management including health management as shown in FIG.
[0206] In Figure 15, first, the answer sentence data creation means 31 creates answer sentence data for each period using answer data contained in questionnaire data such as the health management questionnaire data of each company stored in the questionnaire data storage means 51, and stores the answer sentence data in the answer sentence data storage means 52 (see Figure 3) (step S1).
[0207] Next, in this embodiment, the answer sentence data stored in the answer sentence data storage means 52 is vectorized by the answer sentence vector creation means 32 using, for example, the Azure Open AI Embeddings API service provision system (Azure Open AI Embeddings API text-embedding-ada-002 version 2), and the obtained answer sentence vector (in this embodiment, a 1,536-dimensional vector) is stored in the answer sentence vector storage means 53 (step S2).
[0208] Next, the topic estimation preparation means 33 performs the step of removing unnecessary symbols and tags (for example, ☆, (1) (environment-dependent "maruichi"), etc.) for each theme (theme 4 in the example of FIG. 3) from the response sentence data of each company stored in the response sentence data storage means 52 (see FIG. 3). The process performs a process of removing unnecessary words (i=1 to n: n is the number of answer sentence data in a theme) and a process of breaking down the sentences into words using morphological analysis and extracting only nouns, and a process of removing unnecessary words, and obtains the relationship between each word and its frequency of occurrence in each answer sentence data (i=1 to n: n is the number of answer sentence data in a theme) as information necessary for topic estimation (topic estimation is performed for each theme) (step S3 in FIG. 15).
[0209] Then, for each theme, the topic estimation means 34 executes a topic estimation process (LDA in this embodiment) for each theme using the relationship between each word in all answer sentence data (i = 1 to n) within the theme and the number of times they appear, and determines a topic value (column vector π(i)) indicating the probability of appearance of each topic in the answer sentence data (i) and the probability of appearance of each word in each topic (K rows and p columns matrix β: K is the number of topics, p is the number of words), as shown in FIG. 3. The topic distribution and matrix β indicated by the column vector π(i) obtained by this topic estimation process are stored as a topic model for each theme in the topic model storage means 54 (step S4 in FIG. 15).
[0210] Furthermore, the sub-theme name acquisition means 35 extracts a predetermined number of top (e.g., top 10) words with high occurrence probability for each theme and for each topic from the occurrence probability (matrix β) of each word in each topic for each theme stored in the topic model storage means 54, that is, extracts, for example, the top 10 words with high occurrence probability in each row of the matrix β, and creates sub-theme name creation request data for each theme and for each topic using the extracted multiple words (e.g., 10 words) (see FIG. 13). The created sub-theme name creation request data is transmitted via the network 1 to a service providing system 70 using a large scale language model (LLM) such as ChatGPT, and the sub-theme name (output of the large scale language model) returned from the service providing system 70 is associated with the sub-theme identification information and stored in the sub-theme information storage means 55 (step S5 in FIG. 15).
[0211] Next, as shown in FIG. 4, for each theme, the topic vector calculation means 36 selects and extracts the top (e.g., top 10) answer sentence data of a predetermined number (U items) with the highest topic value from the multiple answer sentence data for each of the multiple topics in the theme, and obtains topic vectors (in this embodiment, 1,536-dimensional answer sentence vectors created by the answer sentence vector creation means 32 and stored in the answer sentence vector storage means 53 (see FIG. 4)) obtained by vectorizing each of the selected and extracted top multiple (e.g., U=10 items) answer sentence data to obtain topic vectors (in this embodiment, 1,536-dimensional vectors) expressing each topic, and stores the obtained topic vectors in the sub-theme information storage means 55 in association with the sub-theme identification information (step S6 in FIG. 15).
[0212] Then, as shown in FIG. 5, the corporate document data creation means 37 divides each of the corporate document data, such as the integrated report data of each company stored in the corporate document data storage means 56, into sentences for each period, and stores the multiple corporate document data obtained by the division in the corporate document data storage means 57 (see FIG. 5) in association with company identification information (stock code, stock name, etc.) and company document identification information that identifies the corporate document data (step S7 in FIG. 15).
[0213] Next, as shown in Fig. 5, the company text vector creation means 38 vectorizes each of the multiple company text data stored in the company text data storage means 57 to create a company text vector, and stores the multiple created company text vectors in the company text vector storage means 58 in association with the company identification information (stock code, stock name, etc.) and the company text identification information (step S8 in Fig. 15). At this time, the company text vector creation means 38 executes vectorization processing using the same method (for example, Azure OpenAI Embeddings API text-embedding-ada-002 version2 in this embodiment) as the method used in the vectorization processing of the response text data by the response text vector creation means 32.
[0214] Then, as shown in FIG. 5, the similarity calculation means 39 calculates the similarity (in this embodiment, cosine similarity) between the multiple company sentence vectors stored in the company sentence vector storage means 58 and the topic vector of each subtheme stored in the subtheme information storage means 55 for each subtheme, thereby obtaining an association score for each of the multiple company sentence data with respect to each subtheme, and stores the obtained association score in the association score storage means 59 in association with the subtheme identification information, company identification information (stock code, stock name, etc.), and company sentence identification information (step S9 in FIG. 15).
[0215] Furthermore, as shown in FIG. 5, the representative related score calculation means 40 extracts, for each sub-theme, the maximum related score from among the multiple related scores associated with the same company identification information using the related scores for each sub-theme stored in the related score storage means 59, and sets the extracted related score as the representative related score for the company of the company identification information, or extracts the top multiple (V, for example, V=5) related scores with the highest values, and sets the average or weighted average or other calculated value calculated using the extracted top multiple (V) related scores as the representative related score for the company of the company identification information, and stores it in the representative related score storage means 60 in association with the sub-theme identification information and company identification information (stock code, stock name, etc.) (step S10 in FIG. 15).
[0216] Then, the output means 41 and the company commitment text acquisition means 42 execute various output processes (see Figures 6 to 12) including a process of displaying the company text data stored in the company text data storage means 57 (see Figure 5) on a screen using the related scores for each sub-theme for each company's company text data stored in the related score storage means 59 (see Figure 5) and the representative related scores for each sub-theme for each company stored in the representative related score storage means 60 (see Figure 5) (step S11 in Figure 15).
[0217] <Effects of this embodiment>
[0218] According to this embodiment, the following effects are obtained: That is, the dashboard system 10 performs a topic estimation process for each theme using the health management level questionnaire data or other questionnaire data to determine sub-themes that indicate each topic, obtains topic vectors that express the sub-themes, calculates the similarity between the topic vector and the corporate sentence vector of the corporate sentence data created from the integrated report data or other corporate document data, sets the calculated similarity as an association score for each sub-theme for each corporate sentence data, and performs output processing to display each corporate sentence data on the screen using this association score.
[0219] Therefore, since the themes of issues in questionnaire data such as the Health and Productivity Management Survey Data are automatically subdivided by the topic estimation process, issues can be understood in terms of sub-themes (each topic of each theme) with finer granularity than the theme, and company document data can be referenced or compared. As a result, users can reference the details of each company's efforts regarding human capital management (including non-financial situations), including health and productivity management, without spending time or effort.
[0220] Moreover, the dashboard system 10 includes a representative related score calculation means 40, and the output means 41 is configured to execute a bird's-eye view display process, so that the representative related score for each sub-theme for each company can be calculated using the related scores, and an overview (see FIG. 7) can be displayed on the screen in which one axis of the matrix-like display unit 101 formed by arranging cells vertically and horizontally represents each sub-theme and the other axis represents each company, and the intensity of the display color of each cell corresponds to the magnitude of the representative related score. Therefore, the status of efforts for each sub-theme for each company can be seen at a glance by the intensity of the color of the cell, and, for example, the status of efforts in the entire industry and differences in the status of efforts between companies can be easily grasped.
[0221] Furthermore, the dashboard system 10 includes a sub-theme name acquisition means 35 that uses a large-scale language model (LLM), which reduces the workload of the system developer and allows for the adoption of objective, more appropriate names. Conventionally, when a topic name (a name equivalent to the sub-theme name of the present invention) needs to be determined in the topic estimation process, the topic name was determined manually, but by using a large-scale language model to determine the topic name, the system developer can save time and effort, and can avoid naming that is largely dependent on human experience and thinking.
[0222] Furthermore, when corporate document data such as integrated report data is PDF data, the corporate document data creation means 37 applies OCR technology (for example, Azure Form Recognizer's OCR technology) to convert the data into text, which significantly eliminates the problem of garbled characters when converting corporate document data created in various formats into text, and makes it possible to extract character strings by block.
[0223] Furthermore, the output means 41 executes an inter-company effort comparison list display process (see FIG. 8), so that it is possible to easily compare the effort status of a plurality of companies (two companies in this embodiment) regarding a certain sub-theme.
[0224] Furthermore, the dashboard system 10 is equipped with a corporate commitment statement acquisition means 42 that utilizes a large-scale language model (LLM), so that on the inter-company commitment comparison list display screen 130 (see FIG. 8), the individual company commitment statement display screen 160 (see FIG. 9), the similar sentence search screen 200 (see FIG. 10), and the recommendation screen 230 (see FIG. 12), the large-scale language model can be used to output multiple commitment statement data showing corporate initiatives in list format using multiple corporate sentence data with high relevance scores (values indicating the relevance to each sub-theme) displayed in the corporate sentence data display sections 140, 142, 170, 204, 240 or the recommended commitment statement display section 241, or in the case of FIG. 10, the input string relevance score (values indicating the relevance to the search string data), and these can be displayed on the initiative list display sections 141, 143, 171, 206, 242, 243. This allows the user to more easily understand the details of the company's initiatives, and also allows the user to use these initiatives as search strings on the similar sentence search screen 200 (see FIG. 10).
[0225] In addition, the output means 41 executes a similar sentence search process (see FIG. 10), and therefore, for each of a plurality of company sentence data, an input string related score for any search string data input by the user is obtained, and the company sentence data can be displayed in descending order of the input string related score. This allows the user to more easily view the information that he or she wishes to refer to.
[0226] Furthermore, the output means 41 executes a similar subtheme display process (see FIG. 11) during the similar sentence search process (see FIG. 10), so that it is possible to present to the user subthemes similar to any search character string data input by the user. This allows the user to change the display method using the subtheme (such as changing the order in which the company sentence data is displayed).
[0227] Furthermore, the output means 41 can execute the integrated related score order display process (see FIG. 11) after executing the similar subtheme display process during the similar sentence search process (see FIG. 10), so that the user can refer to the company sentence data sorted in descending order of the input string related score in a state sorted in descending order of the integrated related score. In other words, the data can be referred to in a state sorted taking into account the related scores for the subthemes presented in the similar subtheme display process. This allows the user to sort and refer to the company sentence data from various perspectives.
[0228] Furthermore, the output means 41 can execute the similar subtheme display process during the similar sentence search process (see FIG. 10) and then execute the adjusted related score order display process (see FIG. 11), so that the user can refer to the company sentence data sorted in descending order of the input string related score in a state sorted in descending order of the adjusted related score. In other words, the data can be referred to in a state sorted using the topic vectors of the subthemes presented in the similar subtheme display process. This allows the user to sort and refer to the company sentence data from various perspectives.
[0229] The dashboard system 10 includes a representative related score calculation means 40, and an output means 41 is configured to execute an effort recommendation sub-theme extraction process, a focus company effort statement display process, and a recommended effort statement display process, so that it is possible to display recommended effort statements for a focus company (a company that needs to improve its efforts) that needs to improve its efforts regarding human capital management, including health management. Therefore, when a company that is not fully committed to human capital management, including health management, or a company that wants to further enhance its efforts, is a user, it can easily refer to the efforts of other companies that it should refer to, and can efficiently collect information. In addition, when a management consulting firm is a user, it can efficiently introduce the efforts of excellent companies that are making sufficient efforts to companies with insufficient efforts.
[0230] <Transformation Form>
[0231] The present invention is not limited to the above-described embodiment, and modifications within the scope of the present invention are included in the present invention.
[0232] For example, in the above embodiment, the dashboard system 10 is described as a stand-alone system, but the dashboard system of the present invention may be a server-client system connected via a network, in which case the main body 20 may be configured as a server, and the display means 80 and input means 81 may be provided in the client terminal.
[0233] In addition, in the above embodiment, the topic inference process was performed using the response data (see FIG. 2) included in the health management questionnaire data, but this is not limited to this, and for example, questionnaire data following the guidelines of ISO30414 (guidelines for disclosure of information related to human capital) may also be used. In short, it is sufficient that the response data includes data in which each company responds according to a predetermined theme of an issue, and that the response data functions as training data for building this system.
[0234] Furthermore, in the above embodiment, integrated report data (see Figure 5) was used to create corporate document data to be presented to the user, but the corporate document data used to create the corporate document data is not limited to integrated report data, and may be, for example, the section of a securities report that describes human capital management, including health management, or a sustainability report. In short, unlike questionnaire data, any data that contains information about human capital management, including health management, related to the company in a free format for each company may be used.
[0235] In the above embodiment, the topic estimation process is performed using the health management survey data or other survey data to subdivide the theme of the problem. More specifically, in the case of the health management survey data, there are 10 themes defined by the Ministry of Economy, Trade and Industry, and by performing the topic estimation process for each of these themes, each theme is further divided into topics, and each topic within each theme is called a sub-theme. Therefore, the 10 themes are subdivided into, for example, 52 sub-themes in total.
[0236] For such a sub-theme obtained by subdividing the theme of the assignment of the questionnaire by topic, another sub-theme (hereinafter referred to as "extended sub-theme") may be prepared by another process (processing without topic estimation) and added to the sub-theme obtained by subdividing the theme of the assignment of the questionnaire. Therefore, the number of sub-themes increases. That is, the topic vector obtained by the topic estimation process is stored in the sub-theme information storage means 55 of the above embodiment, but a vector indicating an extended sub-theme corresponding to this topic vector (hereinafter referred to as "extended sub-theme vector") may be prepared by the above-mentioned other process and stored in the sub-theme information storage means 55. In this case, the extended sub-theme vector is stored in the sub-theme information storage means 55 in association with the sub-theme identification information, as with the topic vector, so that the extended sub-theme is naturally given additional sub-theme identification information. In addition, the name of the extended sub-theme (hereinafter referred to as "extended sub-theme name") is also stored in the sub-theme information storage means 55 in association with the additionally given sub-theme identification information. In addition, since there is no particular need to change the reference numerals in the above embodiment, the reference numerals in the above embodiment will be used as they are in the description.
[0237] Furthermore, when such extended sub-themes are prepared, there is no equivalent to the theme of the questionnaire's assignment. This is because the extended sub-theme is not obtained by subdividing the theme. Therefore, a tentative theme name may be displayed in the theme selection section of each screen of this system (e.g., theme selection section 131 of inter-company initiative comparison list display screen 130 in FIG. 8). The tentative theme name to be displayed may be any name, such as "extended theme," "additional theme," or "theme outside questionnaire," as long as it can be distinguished from each theme of the questionnaire's assignment, and at least one added extended sub-theme may be attributed to one of these tentative themes, such as "extended theme."
[0238] The above-mentioned other process for preparing an extended sub-theme (processing without topic estimation) is executed by an extended sub-theme generating means (not shown) provided in the present system, and is the following process.
[0239] First, prepare terms related to the extended sub-theme you want to create (hereafter referred to as "keywords for extended sub-theme"). This preparation may be done in advance by the designer / developer of this system (before the user performs search / reference processing), by the administrator / operator, or by the user (a company representative who wants to know the activities of other companies, a consultant at a management consulting firm, etc.), or a user who feels the need to set up an extended sub-theme during search / reference processing may interrupt the search / reference processing and do so.
[0240] For example, extended sub-theme keywords may include labor productivity, women's participation in the workforce, financial wellness, etc. When these extended sub-theme keywords are prepared by users of this system, they are terms related to the efforts of each company that they wish to search, refer to, and check, and when they are prepared by the designers / developers or administrators / operators of this system, they are terms related to the efforts of each company that have been determined by considering and anticipating the search / reference needs of users in order to improve the usability of this system.
[0241] Next, the designer / developer, administrator / operator, or user of this system inputs the above-mentioned extended sub-theme keywords prepared by them using input means 81, and the extended sub-theme keyword input acceptance process is executed in which the input is accepted by the extended sub-theme creation means.
[0242] Then, the extended subtheme creating means uses the input multiple extended subtheme keywords to create extended subtheme challenge sentence creation request data such as "Please create 10 challenge sentences for the following keywords. {message}" (embedding multiple extended subtheme keywords in place of {message}), transmits the created extended subtheme challenge sentence creation request data to a service providing system 70 using a large-scale language model (ChatGPT, etc.) via the network 1, and executes an extended subtheme challenge sentence acquisition process that receives a predetermined number of challenge sentences (chat sentence data) (here, 10 as an example) that are generated by the service providing system 70 and sent (returned) via the network 1. This extended subtheme challenge sentence acquisition process is substantially similar to the process of the company challenge sentence acquisition means 42, and only the content of the request sentence and the content of {message} are different.
[0243] Next, the extended sub-theme creation means executes vectorization processing for each of a predetermined number (e.g., 10) of action sentences (attempt sentence data) received from the service provision system 70 using a large-scale language model (ChatGPT, etc.) using the same method (in the above embodiment, for example, Azure OpenAI Embeddings API text-embedding-ada-002 version2) used in the vectorization processing of the answer sentence data by the answer sentence vector creation means 32 in the above embodiment, and executes an action sentence vector creation process to create a predetermined number (e.g., 10) of action sentence vectors. This action sentence vector is, for example, a 1,536-dimensional vector, as in the above embodiment.
[0244] Then, the extended subtheme creation means executes an extended subtheme vector creation process in which it averages a predetermined number (e.g., 10) of the created initiative text vectors to create an extended subtheme vector, and associates the obtained extended subtheme vector with the additionally added subtheme identification information and stores it in the subtheme information storage means 55.
[0245] Furthermore, when acquiring the extended sub-theme name, the extended sub-theme creation means can use the sub-theme name acquisition means 35 (see FIG. 13) of the above embodiment. That is, the extended sub-theme creation means can acquire the extended sub-theme name by embedding the extended sub-theme keyword (e.g., labor productivity, women's participation, financial wellness, etc.) embedded in {message} of the above-mentioned extended sub-theme initiative sentence creation request data into the sub-theme name creation request data (see FIG. 13) and passing it to the service provision system 70 using a large-scale language model (ChatGPT, etc.).
[0246] In addition, since the extended subtheme vectors stored in the subtheme information storage means 55 correspond to topic vectors, they are treated in the same way as topic vectors in the processing by the similarity calculation means 39. That is, the similarity calculation means 39 calculates the similarity between the multiple company sentence vectors and the extended subtheme vectors (corresponding to the topic vectors of each subtheme in the above embodiment) stored in the subtheme information storage means 55, thereby obtaining the relevance scores for the extended subthemes (corresponding to each subtheme in the above embodiment) for each of the multiple company sentence data. Then, the relevance scores for the obtained extended subthemes are stored in the relevance score storage means 60 in association with the additionally assigned subtheme identification information, company identification information, and company sentence identification information. In addition, in the processing by the representative relevance score calculation means 40, they are treated in the same way as in the above embodiment, and the representative relevance scores are obtained using the relevance scores for the extended subthemes (corresponding to each subtheme in the above embodiment) stored in the relevance score storage means 60, and stored in the representative relevance score storage means 60. Therefore, the extended sub-theme name is additionally displayed in the sub-theme name display section 104 of the overhead view display screen 100 in FIG.
[0247] As described above, any sub-theme can be constructed and used as an extended sub-theme. For example, if the total number of sub-themes (topics within each theme) obtained by subdividing the theme of the assignment as in the above embodiment is 52, adding one extended sub-theme makes it possible to prepare and provide 53 sub-themes to the user. In addition, the number of extended sub-themes created is not limited to one, and may be two or more, so that 54th, 55th, ... sub-themes can be prepared and provided. As a result, a system can be constructed that meets the needs of the user, and the convenience of the system can be improved. [Industrial Applicability]
[0248] As described above, the dashboard system and program of the present invention are suitable for use in, for example, improving the efficiency and sophistication of the work of the human resources department or business planning department of an employer company of a health insurance association, or of consultants at management consulting firms. [Explanation of symbols]
[0249] 1 Network 10. Dashboard System 31 Method of creating response data 32. Method for creating response text vectors 34 Topic Estimation Methods 35 Method for obtaining sub-theme name 36 Topic Vector Calculation Method 37. Methods for creating corporate document data 38 Corporate document vector creation method 39 Similarity calculation means 40 Representative related score calculation method 41 Output Method 42 How to obtain a corporate statement 51 Survey data storage means 52 Answer sentence data storage means 53 Answer text vector storage means 54 Topic model storage means 55 Sub-theme information storage means 56 Corporate document data storage means 57 Corporate document data storage means 58 Corporate text vector storage means 59 Related score storage means 60 Representative related score storage means 70 Service provision system using large-scale language models (LLM)
Claims
1. A dashboard system configured by a computer that presents information on human capital management, including health management, a sub-theme information storage means for performing a topic vector calculation process in which, for health and productivity management level questionnaire data including response data of each company to a theme raised as an issue in a survey on human capital management including health and productivity management or other questionnaire data, each of the response data of each company is divided into sentences, and using the multiple answer sentence data obtained by dividing the response data, a topic estimation process is performed for each theme, which estimates multiple topics for the theme by soft clustering or a neural language model, thereby obtaining a topic value indicating the occurrence probability of each topic in each of the answer sentence data, selecting and extracting, for each of the multiple topics, a predetermined number of top answer sentence data having the largest topic value from the multiple answer sentence data, and using the multiple answer sentence vectors obtained by vectorizing each of the multiple selected and extracted top answer sentence data, a topic vector calculation process is performed to obtain a topic vector expressing each topic, and storing the obtained topic vector in association with sub-theme identification information that identifies a sub-theme indicating the topic within the theme; A corporate document data storage means for dividing each of the corporate document data of each company into sentences, and storing the multiple corporate document data obtained by dividing the data into sentences, in association with corporate identification information that identifies the company and corporate document identification information that identifies the corporate document data, for each of the integrated report data of each company, which is different from the questionnaire data and in which each of the companies describes information on human capital management, including health management, related to the company in a free description format for each company, and other corporate document data; a similarity calculation means for calculating a similarity between a plurality of company sentence vectors obtained by vectorizing each of the plurality of company sentence data by a vectorization process using the same method as that used in the vectorization process of the answer sentence data for each of the sub-themes and the topic vectors of each of the sub-themes stored in the sub-theme information storage means, thereby executing a process of obtaining a relevance score for each of the plurality of company sentence data with respect to each of the sub-themes; an association score storage means for storing the association score for each sub-theme calculated by the similarity calculation means in association with the sub-theme identification information, the company identification information, and the company sentence identification information; an output means for executing a process of displaying the company sentence data stored in the company sentence data storage means on a screen by using the relevance score for each sub-theme for the company sentence data of each company stored in the relevance score storage means; A dashboard system comprising:
2. a representative related score calculation means for executing a process of extracting, for each of the sub-themes, the maximum related score from among the multiple related scores associated with the same company identification information, using the related scores for each of the sub-themes stored in the related score storage means, and setting the extracted related score as a representative related score for the company of the company identification information, or extracting a plurality of the related scores with the highest values, and setting an average value or other calculated value calculated using the extracted multiple related scores as a representative related score for the company of the company identification information, The output means includes: The display unit is configured to execute an overhead view display process for displaying an overhead view on a screen in which one axis of a matrix-like display unit formed by arranging cells vertically and horizontally represents each of the sub-themes and the other axis represents each of the companies, and the intensity of the display color of each cell corresponds to the magnitude of the representative related score calculated by the representative related score calculation means. The dashboard system according to claim 1 .
3. A topic estimation means for executing the topic estimation process; a topic model storage means for storing, for each theme, the occurrence probability of each word in each topic within the theme obtained by executing the topic estimation process for each theme by the topic estimation means; a sub-theme name acquisition means for extracting a predetermined number of top words having a high occurrence probability for each theme and for each topic from among the occurrence probabilities of each word in each topic for each theme stored in the topic model storage means, inputting the extracted words and sub-theme name creation request data for each theme and for each topic, which contains a request for a sub-theme name to be created using the extracted words, into a chat generative pre-trained transformer (ChatGPT) or other large-scale language model, and storing the sub-theme name output from the large-scale language model in the sub-theme information storage means in association with the sub-theme identification information, The output means includes: When a user's request to display information for each sub-theme is received, or when information for each sub-theme is displayed, a process using the sub-theme name stored in the sub-theme information storage means is executed. The dashboard system according to claim 1 .
4. The output means includes: The system is configured to execute an inter-company initiative comparison list display process that accepts a user's selection of the sub-theme and a selection of multiple companies, selects and extracts, for each of the multiple selected companies, the company sentence data with the highest related score for the selected sub-theme from the multiple company sentence data, and displays the selected and extracted company sentence data on a screen together with the company identification information or together with the company identification information and the related score in a state where the selected and extracted company sentence data are compiled into a list for comparison by company. The dashboard system according to claim 1 .
5. When the output means executes the process of displaying the inter-company initiative comparison list, the output means inputs the top multiple company sentence data for each selected and extracted company and the company initiative creation request data, which includes a request to create a list of multiple initiative sentences showing the company initiatives using the multiple company sentence data, into a chat generative pretrained transformer (ChatGPT) or other large-scale language model, and outputs the multiple initiative sentences in list format from the large-scale language model. The output means includes: As the inter-company initiative comparison list display process, the process of displaying the plurality of initiative documents acquired by the company initiative document acquisition means in a list format in correspondence with the plurality of top company document data for each company is also executed.
5. The dashboard system according to claim 4.
6. The output means includes: The method is configured to receive an input of an arbitrary search string by a user, vectorize the input search string data by a vectorization process using the same method as that used in the vectorization process of the answer sentence data, calculate a similarity between the obtained search string vector and each of the plurality of company sentence vectors obtained by vectorizing the plurality of company sentence data stored in the company sentence data storage means, obtain an input string related score for the search string data for each of the plurality of company sentence data, and execute a similar sentence search process in which the company sentence data are displayed in order of the highest input string related score obtained. The dashboard system according to claim 1 .
7. The output means includes: The method is configured to execute a similar subtheme display process in which a similarity between the search character string vector and the topic vectors for the plurality of subthemes stored in the subtheme information storage means is calculated, and subthemes of the topic vectors whose similarity to the search character string vector is greater than or equal to a predetermined threshold value are extracted, or subthemes of the topic vectors whose similarity is highest among a predetermined number of cases are extracted, and the extracted subthemes are displayed on a screen to present to a user.
7. The dashboard system according to claim 6.
8. The output means includes: the input string association score; and using the related score stored in the related score storage means in association with the sub-theme identification information of at least one sub-theme selected by a user from at least one sub-theme presented by the similar sub-theme display process, For each of the plurality of corporate text data, the input string relevance score and at least one of the relevance scores are used to calculate an average value, a weighted average value, or other calculated value, which is set as an integrated relevance score; The system is configured to execute an integration related score order display process in which the company statement data is sorted and displayed in order of the highest integration related score.
8. The dashboard system according to claim 7.
9. The output means includes: Using the search character string vector and at least one of the topic vectors stored in the subtheme information storage means in association with the subtheme identification information of at least one subtheme selected by a user from at least one subtheme presented by the similar subtheme display process, an average value or a weighted average value or other calculated value of the elements of each vector is calculated to create a composite vector; A similarity between the composite vector and each of the plurality of company sentence vectors obtained by vectorizing the plurality of company sentence data stored in the company sentence data storage means is calculated to obtain an adjusted association score for each of the plurality of company sentence data; The system is configured to execute an adjusted related score order display process for sorting and displaying the company statement data in order of the adjusted related score.
8. The dashboard system according to claim 7.
10. a representative related score calculation means for executing a process of extracting, for each of the sub-themes, the maximum related score from among the multiple related scores associated with the same company identification information, using the related scores for each of the sub-themes stored in the related score storage means, and setting the extracted related score as a representative related score for the company of the company identification information, or extracting a plurality of the related scores with the highest values, and setting an average value or other calculated value calculated using the extracted multiple related scores as a representative related score for the company of the company identification information, The output means includes: a process of accepting a user's selection of one target company, calculating, for each sub-theme, a difference between the representative related score of the selected target company and the largest representative related score among the representative related scores of companies other than the target company, or calculating a difference between the representative related score of the selected target company and an average or other calculated value of a plurality of the top representative related scores among the representative related scores of companies other than the target company, thereby calculating a deviation degree of the representative related score of the target company from the representative related scores of companies other than the target company, and extracting a plurality of the top sub-themes with the largest calculated deviation degrees as recommended sub-themes and displaying them on the screen; accepting a selection by a user of at least one of the plurality of activity recommendation sub-themes extracted in the activity recommendation sub-theme extraction process; A target company initiative display process that displays the target company's initiative recommendation sub-theme stored in the target company initiative data storage means in descending order of the relevance score for the target company initiative recommendation sub-theme selected by the user for the target company initiative recommendation sub-theme stored in the relevance score storage means; and a recommended initiative display process for displaying the company statement data of the companies other than the target company stored in the company statement data storage means together with the company identification information or together with the company identification information and the related score in descending order of the related score for the initiative recommendation sub-theme related to the user selection for the company statement data of the companies other than the target company stored in the related score storage means on a screen. The dashboard system according to claim 1 .
11. an extended sub-theme creating means for creating an extended sub-theme to be added to the sub-theme obtained by the topic inference process; This method of creating an extended sub-theme is as follows: an extended sub-theme keyword input reception process for receiving an input of an extended sub-theme keyword; A process of inputting the request data for creating an extended sub-theme initiative text, which includes a request for creating multiple initiative texts showing initiatives related to the extended sub-theme keyword using the received extended sub-theme keyword, into a chat generative pre-trained transformer (ChatGPT) or other large-scale language model, and outputting multiple initiative text data from the large-scale language model; A process of creating a plurality of approach sentence vectors by performing a vectorization process for each of the plurality of output approach sentence data by the same method as that used in the vectorization process for the answer sentence data; The method is configured to execute an extended subtheme vector creation process in which an average of the created multiple initiative sentence vectors is taken to create an extended subtheme vector, and the obtained extended subtheme vector is associated with the additionally added subtheme identification information and stored in the subtheme information storage means. The dashboard system according to claim 1 .
12. A program for causing a computer to function as the dashboard system according to any one of claims 1 to 11.
Citation Information
Patent Citations
Matching system and program
JP2021026413A
Matching system and program
JP2022190557A
Evaluation method, evaluation program, and evaluation apparatus for integrated report
JP2023043426A
Natural language processing based on textual polarity
US20170293680A1