Recommendation device and recommendation method
A machine learning-based recommendation device addresses the challenge of dynamically configuring data analysis environments by analyzing user attributes and usage patterns to optimize resource allocation and cost management.
Patent Information
- Application Number
- JP2021171021
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-10-19
AI Technical Summary
Existing systems fail to recommend optimal configuration content for computer environments based on user attributes such as analysis targets, methods, and cost budgets, necessitating a solution that can dynamically adjust resources according to user needs.
A recommendation device using machine learning technology to analyze user attributes and generate configuration recommendations for data analysis environments, incorporating tools, CPU cores, memory, and disk resources based on user profiles and usage patterns.
Enables personalized and efficient construction or modification of data analysis environments tailored to user requirements, optimizing resource allocation and cost management.
Smart Images

Figure 0007713364000001 
Figure 0007713364000002 
Figure 0007713364000003
Abstract
Description
Technical Field
[0001] The present invention relates to a recommendation device and a recommendation method for recommending the configuration content of a computer environment according to a user.
Background Art
[0002] With the progress of computer processing of large amounts of data and machine learning technology, data analysis such as failure prediction, cost analysis, and fraud visualization has become widespread. In order to perform advanced and high-speed processing on large amounts of data, a high-performance and large-scale data analysis environment is required. By constructing a data analysis environment on the cloud, high-performance and large-scale computing resources can be easily utilized in a short time, and a desired data analysis environment can be easily realized.
[0003] As a technology for constructing a computer environment in the cloud, there is a system construction device described in Patent Document 1. This system construction device is a system construction device that constructs a system that provides network services by a plurality of servers, and based on a predetermined determination condition, determines to operate a certain component of one server among the plurality of servers on another server, and based on the configuration management information of the one server, selects the other server and operates the component on the other server.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] According to the above-described system construction device, in a server that provides network services, when the service requests increase and the server load increases, it is possible to scale out in units of the server's components (increasing the number of servers). However, in a data analysis environment, it is necessary to determine the tools and resources to be used according to the user attributes such as the analysis target, analysis method, cost budget (upper limit of usage fees), and data volume. Simply increasing the number of servers cannot handle the situation. This is the same not only in data analysis but also in other computer applications. It is required to recommend the configuration content such as tools and resources to the user according to the user attributes, obtain the user's consent, and then construct the computer environment.
[0006] The present invention has been made in view of such a background, and an object thereof is to provide a recommendation device and a recommendation method that enable recommendation of the configuration content of a computer environment according to a user.
Means for Solving the Problems
[0007] To solve the above-described problems, a recommendation device according to the present invention uses learning data having attribute information of a user of a computer environment as an explanatory variable and identification information of data indicating the configuration content of the computer environment as an objective variable. With reference to the generated environment learning model, the identification information is calculated from the user's attribute information, and the configuration content of the computer environment corresponding to the identification information is output. , and the usage status of the user's computer environment as an explanatory variable, and generated using learning data having identification information of data indicating the configuration content of the computer environment as an objective variable. Change Referring to the environment learning model, the identification information is calculated from the user's attribute information , and the usage status of the user's computer environment and the configuration content of the computer environment corresponding to the identification information is output. Control unit and includes , the usage status of the computer environment corresponding to the identification information that is the explanatory variable of the learning data includes at least any one of the maximum usage number of the CPU used by the tool included in the computer environment during a predetermined period, the maximum usage amount of the memory used by the tool during a predetermined period, and the maximum usage capacity of the disk used by the tool during a predetermined period .
Effects of the Invention
[0008] According to the present invention, it is possible to provide a recommendation device and a recommendation method that enable recommendation of the configuration content of a computer environment according to a user. Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Embodiment for Carrying Out the Invention
[0010] ≪Outline of Recommendation Device≫ The following describes a recommendation device in an embodiment (embodiment) for carrying out the present invention. The recommendation device uses machine learning technology to output the configuration content of a data analysis environment according to user attributes and recommend it to the user. The explanatory variables of the learning data of the machine learning model include, as user attributes, industry type, analysis purpose, analysis proficiency, and the like. The target variable of the learning data indicates the configuration content of the data analysis environment, and is information indicating tools used for data analysis, the number of CPUs (Central Processing Unit) cores assigned to the tools, the memory size, and the like.
[0011] The learning data is an environment with usage records (used for a certain period of time / used more than a certain number of times). By referring to the configuration of the recommended data analysis environment, the user can construct a data analysis environment suitable for the user regardless of the amount of expertise and experience.
[0012] ≪Overall Configuration of Data Analysis Environment≫ FIG. 1 is an overall configuration diagram of a data analysis system 500 including a recommendation device 100 according to this embodiment. The data analysis system 500 includes a Web server 510, an analysis environment system 520, an environment construction device 530, an analysis environment information collection device 540, a recommendation device 100, and an administrator terminal 570. The Web server 510 provides an interface with the user related to the construction and change of the data analysis environment. The interaction between the Web server 510 and the user will be described later with reference to FIGS. 2 to 4.
[0013] The analysis environment system 520 is a computer environment that includes the user's data analysis environment and includes computing resources and storage. The web server 510 and the analysis environment system 520 are connected to the user terminal 590 via the network 580. The environment construction device 530 deploys (arranges / constructs) the configuration of the data analysis environment instructed by the user to the analysis environment system 520. The analysis environment information collection device 540 collects the configuration and usage status of each user's data analysis environment. The administrator terminal 570 is the terminal used by the administrator of the data analysis system 500.
[0014] The recommendation device 100 obtains the user's profile (attribute information, see FIG. 2 described later) from the web server 510 and the usage status of the user's data analysis environment from the analysis environment information collection device 540. Further, the recommendation device 100 transmits the identification information of the template indicating the configuration content of the data analysis environment to be recommended to the user and the configuration content of the data analysis environment corresponding to the identification information to the web server 510. These pieces of information are transmitted from the web server 510 to the user terminal 590 and displayed (see FIGS. 3 and 4 described later).
[0015] FIG. 2 is a screen configuration diagram of a profile registration screen 610 for obtaining the user's profile (attribute information) according to the present embodiment. The profile registration screen 610 is displayed on the display of the user terminal 590. When the user inputs his / her profile and presses the "Register" button 618, the registration content is transmitted to the web server 510. The web server 510 transmits the registration content to the recommendation device 100.
[0016] Figure 3 is a screen configuration diagram of a recommendation screen 620 showing the configuration details of the data analysis environment recommended to the user according to this embodiment. As for the timing of recommendation, there are cases of newly constructing a data analysis environment and cases of changing the configuration of an existing data analysis environment. In the former case, the identification information of the template showing the configuration details recommended as a new construction is displayed on button 621. In the latter case, the identification information of the template showing the configuration details recommended as a configuration change is displayed on button 622. When buttons 621 and 622 are pressed, a recommendation content details screen 630 (see FIG. 4 described later) is displayed, which shows the details of the configuration of the data analysis environment that is the content of the template. When the "Construct" button 628 is pressed, a data analysis environment corresponding to the template is constructed (deployed) in the analysis environment system 520 (see FIG. 1).
[0017] Figure 4 is a screen configuration diagram of a recommendation content details screen 630 showing the details of the configuration of the recommended data analysis environment according to this embodiment. On the recommendation content details screen 630, the configuration details of the data analysis environment corresponding to the template identification information displayed on buttons 621 and 622 and the monthly usage fee are displayed. As the configuration details, in addition to the algorithms used for data analysis, there are the names of various tools included in the environment and the resources used by the tools. When the user wants to change a tool or a resource, after operating the column to be changed (for example, selecting another tool, changing the memory amount), the user presses the "Register" button 638. Then, the screen returns to the recommendation screen 620 (see FIG. 3), and the identification information of the template different from the initially displayed template and corresponding to the changed content is displayed on buttons 621 and 622.
[0018] <<Configuration of Recommendation Device>> Figure 5 is a functional block diagram of a recommendation device 100 according to this embodiment. The recommendation device 100 is a computer and includes a control unit 110, a storage unit 120, and a communication unit 130. The communication unit 130 includes a communication device and performs data transmission and reception with other devices such as a Web server 510, an analysis environment information collection device 540, and an administrator terminal 570. The control unit 110 includes a CPU and is provided with a user information acquisition unit 111, an environment information acquisition unit 112, a recommendation processing unit 113, a learning unit 114, and a recommendation information generation unit 115.
[0019] The storage unit 120 includes storage devices such as a ROM (Read Only Memory), a RAM (Random Access Memory), and an SSD (Solid State Drive). Stored in the storage unit 120 are a user information database 210, an environment information database 220, a newly constructed recommendation information database 230, a configuration change recommendation information database 240, a newly constructed template database 250, a configuration change template database 260, newly constructed recommendation learning data 270, configuration change recommendation learning data 280, a resource addition database 290, a newly constructed learning model 121, a configuration change learning model 122, and a program 128. In FIG. 5, the databases are denoted as DB. For example, the user information database 210 is denoted as the user information DB.
[0020] The program 128 includes descriptions of the generation processes of the newly constructed learning model 121, the configuration change learning model 122, the newly constructed recommendation information database 230, and the configuration change recommendation information database 240 (see FIGS. 15, 16, and 17 to be described later). Hereinafter, the configurations of the control unit 110 and the storage unit 120 will be described in order.
[0021] ≪User Information Acquisition≫ The user information acquisition unit 111 acquires the user attribute information (refer to the profile registration screen 610 shown in FIG. 2) transmitted by the Web server 510 and stores it in the user information database 210. FIG. 6 is a data configuration diagram of the user information database 210 according to the present embodiment. The user information database 210 is, for example, tabular data, and one row (record) indicates a user. Each record includes columns (attributes) of record identification information (described as # in FIG. 6), user identification information (described as user ID in FIG. 6), company name, department name, industry type, analysis purpose, budget, planned number of users, monthly analysis execution times, and analysis proficiency. The budget is the upper limit of the monthly usage fee, and the unit is in millions of yen. The monthly analysis execution times is the planned number of analysis executions per month.
[0022] ≪Environment Information Acquisition≫ The environment information acquisition unit 112 acquires the usage status of the data analysis environment of each user collected by the analysis environment information collection device 540 and stores it in the environment information database 220. FIG. 7 is a data configuration diagram of the environment information database 220 according to the present embodiment. The environment information database 220 is, for example, tabular data, and one row (record) indicates the configuration content of the data analysis environment. Each record includes columns (attributes) of record identification information (described as # in FIG. 7), user identification information (described as user ID in FIG. 7), environment identification information (described as environment ID in FIG. 7), construction date and time, last update date and time, and usage status of each tool.
[0023] The environment identification information is identification information assigned to each user, and is newly assigned when a user newly constructs a data analysis environment or changes the configuration of an existing data analysis environment. The construction date and time is the date and time when the data analysis environment was newly constructed or constructed after being changed. The last update date and time is the date and time when the record was last updated. Note that the record is updated when the maximum monthly CPU usage or the maximum monthly memory usage of the tool described later is updated.
[0024] The column of tool usage status further includes columns of availability, number of CPUs, maximum monthly CPU usage, memory size, and maximum monthly memory usage. The column of tool usage status may also include columns (not shown) such as disk size / capacity, maximum monthly disk usage, number of licenses, and definition files (configuration files). Availability indicates whether the data analysis environment includes the tool. The number of CPUs indicates the upper limit of the number of CPU cores used by the tool. The maximum monthly CPU usage indicates the maximum number of CPU cores used in the most recent one month. The memory size indicates the upper limit of the memory size used by the tool. The maximum monthly memory usage indicates the maximum amount of memory (used size) used in the most recent one month.
[0025] ≪Recommendation Processing≫ In response to a request from the Web server 510, the recommendation processing unit 113 transmits template identification information indicating the configuration details of the recommended data analysis environment, and the configuration details corresponding to the template identification information.
[0026] FIG. 8 is a data configuration diagram of the newly constructed recommendation information database 230 according to the present embodiment. The newly constructed recommendation information database 230 is, for example, tabular data, and one row (record) indicates information related to the configuration details of the data analysis environment recommended for a user at the time of new construction. Each record includes columns (attributes) of record identification information (denoted as # in FIG. 8), user identification information (denoted as user ID in FIG. 8), recommendation date and time, and template identification information (denoted as template ID in FIG. 8).
[0027] The template identification information is the identification information of the template recommended for a user indicated by the user identification information at the time of new construction. This template identification information corresponds to the template identification information (denoted as template ID in FIG. 10) of the newly constructed template database 250 (see FIG. 10) described later.
[0028] When the recommendation processing unit 113 receives a request for the configuration details of the data analysis environment to be recommended for a new construction, including user identification information, from the Web server 510, it refers to the new construction recommendation information database 230 to search for the template identification information corresponding to the user identification information, and then replies to the Web server 510. This process is executed when the Web server 510 transmits the data of the recommendation screen 620 (see FIG. 3) to the user terminal 590 during new construction. Note that the recommendation date and time in the new construction recommendation information database 230 is the execution date and time of this process.
[0029] FIG. 9 is a data configuration diagram of the configuration change recommendation information database 240 according to the present embodiment. The configuration change recommendation information database 240 is, for example, tabular data, and one row (record) indicates information related to the configuration details of the data analysis environment to be recommended when changing the configuration for the user. Each record includes columns (attributes) of record identification information (denoted as # in FIG. 9), user identification information (denoted as user ID in FIG. 9), environment identification information (denoted as environment ID in FIG. 9), recommendation date and time, and template identification information (denoted as template ID in FIG. 9).
[0030] The template identification information is the identification information of the template to be recommended when changing the configuration for the user indicated by the user identification information. This template identification information corresponds to the template identification information (denoted as template ID in FIG. 10) in the configuration change template database 260 (see FIG. 10) described later.
[0031] When the recommendation processing unit 113 receives a request for the configuration content of the data analysis environment to be recommended for configuration changes, including user identification information and current environment identification information, from the Web server 510, it refers to the configuration change recommendation information database 240 to search for template identification information corresponding to the user identification information and the environment identification information, and returns it to the Web server 510. This process is executed when the Web server 510 transmits the data of the recommendation screen 620 (see FIG. 3) to the user terminal 590 at the time of configuration change. Note that the recommendation date and time in the configuration change recommendation information database 240 is the execution date and time of this process.
[0032] FIG. 10 is a data configuration diagram of the newly constructed template database 250 and the configuration change template database 260 according to the present embodiment. The newly constructed template database 250 and the configuration change template database 260 have the same data configuration. The newly constructed template database 250 and the configuration change template database 260 are, for example, tabular data, and one row (record) shows the configuration content of the data analysis environment corresponding to the template identification information. Each record includes record identification information (described as # in FIG. 10), template identification information (described as template ID in FIG. 10), record registration date and time, and columns (attributes) of the configuration content of each tool.
[0033] The columns of the configuration content of the tool include columns (attributes) such as availability, CPU, memory, disk, number of licenses, definition file (configuration file), etc. Availability indicates whether the data analysis environment includes the tool. CPU indicates the number of CPU cores used by the tool. Memory indicates the memory size used by the tool. Disk indicates the area size of the disk or SSD used by the tool. The number of licenses indicates the number of licenses of the tool. The definition file indicates the definition file (configuration file) used by the tool.
[0034] When the recommendation processing unit 113 receives a request for template content for a new construction including template identification information from the Web server 510, it refers to the new construction template database 250 to search for the configuration content of each tool corresponding to the template identification information, and replies to the Web server 510. Also, when the recommendation processing unit 113 receives a request for template content for a configuration change including template identification information from the Web server 510, it refers to the configuration change template database 260 to search for the configuration content of each tool corresponding to the template identification information, and replies to the Web server 510. These processes are executed when the Web server 510 transmits the data of the recommendation content details screen 630 (see FIG. 4) to the user terminal 590.
[0035] As described so far, the new construction recommendation information database 230, the configuration change recommendation information database 240, the new construction template database 250, and the configuration change template database 260 are referred to by the recommendation processing unit 113. These databases are generated by the learning unit 114 and the recommendation information generation unit 115 described later.
[0036] Also, the new construction template database 250 and the configuration change template database 260 are data indicating the configuration content of the data analysis environment (computer environment) in the analysis environment system 520. The template identification information is the identification information of this data. The recommendation processing unit 113 (recommendation unit) transmits (outputs) the configuration content corresponding to this identification information to the Web server 510.
[0037] ≪Learning model generation≫ The learning unit 114 generates new construction recommendation learning data 270 (see FIG. 12 described later) according to the instructions of the administrator of the data analysis system 500, and trains the newly constructed learning model 121 (generating the newly constructed learning model 121 by having it learn the newly constructed recommendation learning data 270). Further, the learning unit 114 generates configuration change recommendation learning data 280 (see FIG. 13 described later) according to the instructions of the administrator of the data analysis system 500, and trains the configuration change learning model 122.
[0038] The newly constructed learning model 121 and the configuration change learning model 122 are learning models of machine learning technology, for example, learning models of neural networks. The newly constructed learning model 121 and the configuration change learning model 122 may be learning models of other machine learning technologies such as support vector machines and decision trees.
[0039] FIG. 11 is a screen configuration diagram of the recommendation learning model generation instruction screen 310 according to the present embodiment. The recommendation learning model generation instruction screen 310 is a screen displayed on the administrator terminal 570 (see FIG. 1), and is a screen operated by the administrator of the data analysis system 500. In the learning model designation area 311, it is designated whether to generate the newly constructed learning model 121 used for recommendation at the time of new construction or the configuration change learning model 122 used for recommendation at the time of configuration change.
[0040] In the learning method designation area 312, it is designated whether the learning method is automatic or manual. In the case of automatic, based on the data in the environment information database 220 (see FIG. 7), the newly constructed recommendation learning data 270 and the configuration change recommendation learning data 280 are generated. In the case of manual, the learning data is acquired from the file designated in the learning data designation area 314.
[0041] In the learning target date specification area 313, when the learning method is automatic, it is specified at what point in time the data in the environmental information database 220 before that is used as learning data. As learning data, data indicating the configuration content that has been continuously used for a certain period is desirable, and for example, three months ago or one year ago is specified. A newly constructed or recently configuration-changed data analysis environment is expected to have changes in the tools and amount of resources used and is likely to be unstable. As the recommended data analysis environment, a configuration that has been continuously used for a predetermined period and is stable is desirable. Also, a configuration with a certain number of usage records or a configuration that has not been changed for a predetermined period may be used. When the "Learning Execution" button 318 is pressed, learning data is generated, and a newly constructed learning model 121 or a configuration change learning model 122 trained using this learning data is generated.
[0042] FIG. 12 is a data configuration diagram of the newly constructed recommendation learning data 270 according to the present embodiment. The newly constructed recommendation learning data 270 is the learning data of the newly constructed learning model 121. The target variable (output of the newly constructed learning model 121) in the newly constructed recommendation learning data 270 is template identification information (described as template ID in FIG. 12). This template identification information corresponds to the template identification information in the newly constructed template database 250 (see FIG. 10).
[0043] The explanatory variables (inputs to the newly constructed learning model 121) of the newly constructed recommendation learning data 270 include the industry type, analysis purpose, budget, planned number of users, monthly analysis execution frequency, and analysis proficiency, which are the attribute information of the user. The content of these attributes is the same as the attributes in the user information database 210 (see FIG. 6).
[0044] FIG. 13 is a data configuration diagram of the configuration change recommendation learning data 280 according to the present embodiment. The configuration change recommendation learning data 280 is learning data for the configuration change learning model 122. The target variable (the output of the configuration change learning model 122) in the configuration change recommendation learning data 280 is template identification information (described as template ID in FIG. 13). This template identification information corresponds to the template identification information in the configuration change template database 260 (see FIG. 10).
[0045] The explanatory variables (the inputs of the configuration change learning model 122) of the configuration change recommendation learning data 280 include user attribute information and the configuration content of each tool (and thus the configuration content of the data analysis environment). The user attribute information includes industry type, analysis purpose, budget, planned number of users, monthly analysis execution frequency, and analysis proficiency. The contents of these attributes are the same as the attributes in the user information database 210 (see FIG. 6). The configuration content of the tool includes availability, CPU, maximum monthly CPU usage, memory, maximum monthly memory usage, disk, maximum monthly disk usage, number of licenses, and attributes of the definition file (configuration file). These attributes are the same as the attributes of the configuration content of the tool in the environment information database 220 (see FIG. 7).
[0046] FIG. 14 is a data configuration diagram of the resource increase database 290 according to the present embodiment. The resource increase database 290 is, for example, tabular data, and one row (record) shows information related to the increase of resources such as CPU, memory, and disk. Each record includes columns (attributes) of record identification information 291 (described as # in FIG. 14), resource 292, additional unit 293, and threshold 294.
[0047] When generating the newly constructed template database 250 (see FIG. 10) and the configuration change template database 260 that indicate the configuration details of the recommended data analysis environment, the learning unit 114 refers to the resource increment database 290. Specifically, if the utilization rate (monthly maximum usage amount / amount of resource) of the resource 292 shown in the environment information database 220 (see FIG. 7) exceeds the threshold 294, the learning unit 114 adds the amount of resource shown in the addition unit 293 to generate the newly constructed template database 250 and the configuration change template database 260 (see FIGS. 15 and 16 described later). Also, if the utilization rate of the resource 292 is equal to or less than the threshold 294, the learning unit 114 generates the newly constructed template database 250 and the configuration change template database 260 using the monthly maximum usage amount of the resources in the environment information database 220.
[0048] For example, if the monthly maximum usage amount / number of the CPU of tool A shown in the environment information database 220 exceeds 0.9, record 299 indicates that the number of CPUs is increased by 2. If the number of CPUs of tool A shown in the environment information database 220 is 2, then the CPU of tool A in the newly constructed template database 250 and the configuration change template database 260 is increased by 2 to become 4. If the monthly maximum usage amount / number is equal to or less than 0.9, the CPU of tool A in the newly constructed template database 250 and the configuration change template database 260 is the monthly maximum usage amount of the CPU of tool A shown in the environment information database 220.
[0049] ≪Recommendation Information Generation≫ The recommendation information generation unit 115 generates the newly constructed recommendation information database 230 (see FIG. 8) using the newly constructed learning model 121. Specifically, for each user, the recommendation information generation unit 115 inputs the attributes of the user (see the user information database 210 described in FIG. 6), and calculates the template identification information using the newly constructed learning model 121 to generate the newly constructed recommendation information database 230.
[0050] Also, the recommendation information generation unit 115 generates a configuration change recommendation information database 240 (see FIG. 9) using the configuration change learning model 122. Specifically, for each user and the user's environment, the recommendation information generation unit 115 inputs the attributes of the user and the configuration details of the current data analysis environment (see the environment information database 220 described in FIG. 7), and calculates template identification information using the configuration change learning model 122 to generate the configuration change recommendation information database 240.
[0051] ≪New construction learning model generation process≫ FIG. 15 is a flowchart of the generation process of the new construction learning model 121 according to the present embodiment. When new construction is specified in the learning model designation area 311 of the recommendation learning model generation instruction screen 310 (see FIG. 11) and the "learning execution" button 318 is pressed, this generation process starts.
[0052] In step S11, the learning unit 114 initializes (deletes all records) the new construction template database 250 (see FIG. 10) and the new construction recommendation learning data 270 (see FIG. 12). In step S12, if "automatic" is selected in the learning method designation area 312 (see FIG. 11) (step S12 → YES), the learning unit 114 proceeds to step S13, and if "manual" is selected (step S12 → NO), the learning unit 114 proceeds to step S21.
[0053] In step S13, the learning unit 114 acquires records in the environment information database 220 (see FIG. 7) whose last update date and time are before the specified date specified in the learning target date designation area 313. If no specified date is designated, the learning unit 114 acquires records whose last update date and time are before the default date, for example, three months before the present. In step S14, the learning unit 114 repeatedly executes steps S15 to S19 for each record acquired in step S13.
[0054] In step S15, the learning unit 114 adds the usage status of each tool, which is the content of the record, to the newly constructed template database 250. At the same time, the learning unit 114 assigns template identification information to the record. In step S16, the learning unit 114 adds the template identification information assigned in step S15 and the attribute information of the user to the newly constructed recommendation learning data 270. Here, the attribute information of the user is the attribute information (such as industry type) of the user corresponding to the user identification information of the record in the user information database 210 (see FIG. 6).
[0055] In step S17, for each resource (CPU, memory, disk) included in the configuration content of each tool in the record, if the utilization rate (monthly maximum usage amount / resource amount) exceeds the threshold value 294 of the resource (see the resource 292 in FIG. 14) (step S17 → YES), it proceeds to step S19, and if it does not exceed (step S17 → NO), it proceeds to step S18.
[0056] In step S18, the learning unit 114 stores the value of the monthly maximum usage amount of the corresponding resource of the record for the resources (CPU, memory, disk) of the tool in the newly constructed template database 250 added in step S15. In step S19, the learning unit 114 increases the resource of the additional unit 293 to the resources of the tool included in the record of the newly constructed template database 250 added in step S15. For example, for the CPU of tool A, if the utilization rate exceeds 0.9, it increases by 2 (see record 299). Note that the value of the resource in the newly constructed template database 250 before the increase is the value of the resource of the record (see step S15).
[0057] In step S20, the learning unit 114 trains and generates a newly constructed learning model 121 with the attribute information of the user in the newly constructed recommendation learning data 270 as the explanatory variable and the template identification information as the target variable. In step S21, the learning unit 114 acquires learning data from the file specified in the learning data specification area 314, generates new construction recommendation learning data 270, and proceeds to step S20.
[0058] ≪Configuration change learning model generation process≫ FIG. 16 is a flowchart of the generation process of the configuration change learning model 122 according to the present embodiment. When a configuration change is specified in the learning model specification area 311 of the recommendation learning model generation instruction screen 310 (see FIG. 11) and the "Learning execution" button 318 is pressed, this generation process starts.
[0059] In step S31, the learning unit 114 initializes the configuration change template database 260 (see FIG. 10) and the configuration change recommendation learning data 280 (see FIG. 13). In step S32, if "Automatic" is selected in the learning method specification area 312 (see FIG. 11) (step S32 → YES), the learning unit 114 proceeds to step S33, and if "Manual" is selected (step S32 → NO), the learning unit 114 proceeds to step S41.
[0060] In step S33, the learning unit 114 acquires records in the environment information database 220 (see FIG. 7) whose last update date and time are before the specified date specified in the learning target date specification area 313. If no specified date is specified, the learning unit 114 acquires records whose last update date and time are before the default date, for example, three months before the current time. In step S34, the learning unit 114 repeatedly executes steps S35 to S39 for each record acquired in step S33.
[0061] In step S35, the learning unit 114 adds the usage status of each tool, which is the content of the record, to the configuration change template database 260. At the same time, the learning unit 114 assigns template identification information to the record.
[0062] In step S36, the learning unit 114 adds the template identification information assigned in step S35, the user attribute information, and the usage status of each tool which is the content of the record to the configuration change recommendation learning data 280. Here, the user attribute information is the user attribute information (such as industry type) corresponding to the user identification information of the record in the user information database 210 (see FIG. 6). Also, for the CPU, memory, and disk, the values of the corresponding resource amount of the record and its monthly maximum usage amount are stored in the configuration change recommendation learning data 280.
[0063] In step S37, for each resource (CPU, memory, disk) included in the configuration content of each tool of the record, if the utilization rate (monthly maximum usage amount / resource amount) exceeds the threshold value 294 of the resource (see resource 292 in FIG. 14) (step S37 → YES), it proceeds to step S39, and if it does not exceed (step S37 → NO), it proceeds to step S38.
[0064] In step S38, the learning unit 114 stores the value of the monthly maximum usage amount of the corresponding resource of the record for the resources (CPU, memory, disk) of the tool in the configuration change template database 260 added in step S35. In step S39, the learning unit 114 increases the resource of the additional unit 293 to the resources of the tools included in the record in the configuration change template database 260 added in step S35. For example, for the CPU of tool A, if the utilization rate exceeds 0.9, it increases by 2 (see record 299). Note that the value of the resource in the configuration change template database 260 before the increase is the value of the resource of the record (see step S35).
[0065] In step S40, the learning unit 114 trains and generates the configuration change learning model 122 using the user attribute information and the usage status of each tool in the configuration change recommendation learning data 280 as explanatory variables and the template identification information as the target variable. In step S41, the learning unit 114 acquires learning data from the file specified in the learning data specification area 314, generates configuration change recommendation learning data 280, and proceeds to step S39.
[0066] ≪Recommendation Information Generation Process≫ FIG. 17 is a flowchart of the recommendation information generation process according to the present embodiment. With reference to FIG. 17, a process in which the recommendation information generation unit 115 generates a newly constructed recommendation information database 230 (see FIG. 8) and a configuration change recommendation information database 240 (see FIG. 9) will be described. The recommendation information generation process may be executed subsequent to the generation of the newly constructed learning model 121 and the configuration change learning model 122, may be executed at the instruction of the administrator of the data analysis system 500, or may be executed at other timings.
[0067] In step S51, the recommendation information generation unit 115 initializes the newly constructed recommendation information database 230 and the configuration change recommendation information database 240. In step S52, the recommendation information generation unit 115 repeatedly executes steps S52 to S59 for each user identification information in the user information database 210 (see FIG. 6). In step S53, the recommendation information generation unit 115 acquires the attribute information of the user corresponding to the user identification information from the user information database 210.
[0068] In step S54, the recommendation information generation unit 115 calculates template identification information (objective variable) using the attribute information of the user acquired in step S53 as an input (explanatory variable) using the newly constructed learning model 121. In step S55, the recommendation information generation unit 115 registers the template identification information calculated in step S54 in the newly constructed recommendation information database 230 together with the user identification information.
[0069] In step S56, the recommendation information generation unit 115 acquires records in the environment information database 220 (see FIG. 7) in which the user identification information matches. In step S57, the recommendation information generation unit 115 repeatedly executes steps S58 and S59 for each record acquired in step S56. In step S58, the recommendation information generation unit 115 calculates template identification information (objective variable) using the user attribute information acquired in step S53 and the usage status of each tool in the record of the environment information database 220 as inputs (explanatory variables). Note that the CPU, memory, and disk in the usage status of the tool as the input are the resource amount and the monthly maximum usage amount in the record of the environment information database 220.
[0070] In step S59, the recommendation information generation unit 115 registers the template identification information calculated in step S58 together with the user identification information and the environment identification information in the configuration change recommendation information database 240. Note that the environment identification information is the environment identification information included in the record of the environment information database 220.
[0071] ≪Features of the Recommendation Device≫ The recommendation device 100 recommends, as a new data analysis environment, the configuration content of the data analysis environment according to the user's attribute information using a machine learning model. In addition, the recommendation device 100 recommends, as a changed data analysis environment, the configuration content of the data analysis environment according to the user's attribute information and the configuration content of the data analysis environment used by the user.
[0072] The learning data of the learning model is an environment with usage records (used for a certain period of time or used more than a certain number of times). Regardless of the amount of specialized knowledge and experience of the user, the user can build a data analysis environment suitable for the user. For example, when the utilization rate of resources is higher than a predetermined value, resources are increased (see step S19 in FIG. 15 and step S39 in FIG. 16). Also, when the utilization rate of resources is below the predetermined value, resources are reduced (see step S18 in FIG. 15 and step S38 in FIG. 16).
[0073] <<Modification Example: Threshold of Resource Utilization Rate>> In the above-described embodiment, the recommendation device 100 generates learning data such that if the utilization rate of resources exceeds the threshold value 294 (see FIG. 14), additional unit 293 of resources is increased, and if it is below the threshold value 294, the resources are reduced to the maximum monthly usage amount (see steps S18 and S19 in FIG. 15 and steps S38 and S39 in FIG. 16). Apart from the threshold value 294, a new threshold value for reducing to the maximum monthly usage amount may be provided. In this case, when the utilization rate is between the threshold value 294 and the new threshold value, neither the increase nor the reduction of resources is performed, and the amount of resources in the environmental information database 220 remains as it is.
[0074] <<Modification Example: User Attributes>> In the above-described embodiment, the attribute information of the user is the industry type, analysis purpose, etc. (see FIGS. 12 and 13), but may further include other attribute information. For example, the attribute information of the user may include the estimated amount of data to be analyzed and the estimated amount of data that increases over a predetermined period (such as months or years). By recommending including such attribute information, the recommendation device 100 can recommend resources with higher accuracy. Also, the attribute information of the user may include the type of data (production information, equipment information, management data, etc.) and the data update cycle (such as months, half a year, year, etc.). By recommending including such attribute information, the recommendation device 100 can recommend tools according to the type of data.
[0075] <<Modification Example: Learning Data>> In the above-described embodiments, the objective variables of the novel construction learning model 121 and the configuration change learning model 122 are template identification information. Instead of the template identification information, the configuration details of the data analysis environment may be used as the objective variable. To configure in this way, the objective variables of the novel construction recommendation learning data 270 and the configuration change recommendation learning data 280 may be set to the configuration details of the data analysis environment.
[0076] ≪Other Modification Examples≫ As described above, some embodiments of the present invention have been explained. However, these embodiments are merely examples and do not limit the technical scope of the present invention. For example, although the recommendation device 100 stores each database in the storage unit 120, it may access each database stored in an external device.
[0077] In the above-described embodiment, the recommendation processing unit 113 refers to the novel construction recommendation information database 230 (see FIG. 8) and the configuration change recommendation information database 240 (see FIG. 9) generated by the recommendation information generation unit 115, and transmits the template identification information to be recommended to the Web server 510. When receiving a request from the Web server 510, the recommendation processing unit 113 (recommendation unit) may calculate the template identification information to be recommended using the novel construction learning model 121 and the configuration change learning model 122. Also, although the recommendation device 100 in the above-described embodiment recommends the configuration details of the data analysis environment, it may recommend the configuration details related to other-purpose computer environments such as software development and simulation execution.
[0078] The present invention can take various other embodiments, and furthermore, various changes such as omission and substitution can be made without departing from the gist of the present invention. These embodiments and their modifications are included in the scope and gist of the invention described in this specification and the like, and are also included in the invention described in the claims and its equivalent scope.
Explanation of Reference Numerals
[0079] 100 Recommendation Device 111 User Information Acquisition Unit 112 Environment Information Acquisition Unit 113 Recommendation Processing Unit (Recommendation Unit) 114 Learning Unit 115 Recommendation Information Generation Unit 121 Newly Constructed Learning Model (New Environment Learning Model) 122 Configuration Change Learning Model (Changed Environment Learning Model) 250 Newly Constructed Template Database (Data Indicating the Configuration Content of the Computer Environment) 260 Configuration Change Template Database (Data Indicating the Configuration Content of the Computer Environment) 270 Newly Constructed Recommendation Learning Data (Learning Data) 280 Configuration Change Recommendation Learning Data (Learning Data) 520 Analysis Environment System (Computer Environment)
Claims
1. With reference to a modified environment learning model generated using learning data that uses the attribute information of a user of a computer environment and the usage status of the user's computer environment as explanatory variables, and the identification information of data indicating the configuration content of the computer environment as the objective variable, the identification information is calculated from the attribute information of the user and the usage status of the user's computer environment, A control unit is provided that outputs the configuration content of the computer environment corresponding to the identification information. The usage status of the computer environment corresponding to the identification information that is an explanatory variable of the learning data includes at least any one of the maximum usage number of the CPU used by the tool included in the computer environment during a predetermined period, the maximum usage amount of the memory used by the tool during a predetermined period, and the maximum usage capacity of the disk used by the tool during a predetermined period. A recommendation device characterized by the above.
2. The configuration content of the computer environment corresponding to the identification information that is the objective variable of the learning data corresponds to any one of the configuration content that has not been changed for a predetermined period, the configuration content with a usage record for a predetermined number of times, and the configuration content with a usage record for a predetermined time. The recommendation device according to claim 1, characterized by the above.
3. The attribute information includes at least any one of industry type, analysis purpose, budget, planned number of users, monthly analysis execution frequency, and analysis proficiency. The recommendation device according to claim 1, characterized by the above.
4. The computer environment is an analysis environment for data, The attribute information includes at least any one of the estimated amount of the data, the estimated amount of the data that increases in a predetermined period, the type of the data, and the update period of the data. The recommendation device according to claim 1, characterized by the above.
5. The configuration content of the computer environment includes at least any one of the presence or absence of tools included in the computer environment, the number of CPUs used by the tool, the memory size used by the tool, the disk capacity used by the tool, the number of licenses of the tool, and the definition file of the tool. The recommendation device according to claim 1, characterized by the above.
6. The configuration content of the computer environment includes the presence or absence of tools included in the computer environment and the amount of resources used by the tool. In an existing configuration content that is the configuration content of the computer environment with utilization history, if the ratio of the maximum utilization amount of the resource in a predetermined period to the amount of the resource in the existing configuration content is greater than a predetermined value, the amount of the resource included in the configuration content of the computer environment corresponding to the identification information that is the target variable of the learning data is greater than the amount of the resource in the existing configuration content The recommendation device according to claim 1, characterized in that.
7. The configuration content of the computer environment includes the presence or absence of tools included in the computer environment and the amount of resources used by the tools. In an existing configuration content that is the configuration content of the computer environment with utilization history, if the ratio of the maximum utilization amount of the resource in a predetermined period to the amount of the resource in the existing configuration content is less than a predetermined value, the amount of the resource included in the configuration content of the computer environment corresponding to the identification information that is the target variable of the learning data is less than the amount of the resource in the existing configuration content The recommendation device according to claim 1, characterized in that.
8. The resources used by the tool include at least any one of a CPU, a memory, and a disk. The recommendation device according to claim 6 or 7, characterized in that.
9. The recommendation device is Using a change environment learning model generated using learning data with the attribute information of the user of the computer environment and the usage status of the user's computer environment as explanatory variables and the identification information of the data indicating the configuration content of the computer environment as the target variable, calculating the identification information from the attribute information of the user and the usage status of the user's computer environment; Outputting the configuration content of the computer environment corresponding to the identification information. The usage status of the computer environment corresponding to the identification information that is the explanatory variable of the learning data includes at least any one of the maximum usage count of the CPU used by the tool included in the computer environment in a predetermined period, the maximum usage amount of the memory used by the tool in a predetermined period, and the maximum usage capacity of the disk used by the tool in a predetermined period. The recommendation method, characterized in that.
Citation Information
Patent Citations
Server construction support system, server construction support device, server construction support method and program therefor
JP2006072772A
Management computer, computer system, and instance management method
JP2015524581A
System construction apparatus, system construction method, and program
JP2018142277A
Network requirement generation system, and network requirement generation method
JP2020140276A