Information processing system, information processing method, and program

The system identifies and explains outliers in financial data using AI, improving user comprehension by highlighting significant deviations.

JP2025161295APending Publication Date: 2025-10-24株式会社ナレッジラボ
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024064366
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-12
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Financial data lists are complex and difficult for users with low accounting literacy to interpret, often leading to important information being overlooked.

Method used

An information processing system that identifies outliers in numerical data, generates explanatory sentences for these outliers, and displays them alongside the data using AI technology.

Benefits of technology

Enhances user understanding of financial data by highlighting significant deviations through clear, concise explanations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025161295000001_ABST
    Figure 2025161295000001_ABST
Patent Text Reader

Abstract

To provide an information processing system which can promote understanding about outliers.SOLUTION: According to an aspect of the present invention, an information processing system is provided. The information processing system includes at least one processor. The processor acquires numerical value data groups including a plurality of numerical values related to each other in an acquisition step, specifies outliers in the acquired numerical data groups in a specification step, generates sentences explaining the specified outliers in a generation step, and causes the system to display a list of the acquired numerical data groups and to display the sentences explaining the specified outliers included in the list in a mode representing correspondence between the sentences and the outliers in a display control step.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system, an information processing method, and a program. [Background technology]

[0002] Patent Document 1 discloses a technique for identifying abnormal values ​​by comparing financial data of multiple companies. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-252259 Summary of the Invention [Problem to be solved by the invention]

[0004] Financial data lists contain a large amount of information, making it difficult to interpret the figures and leading to overlooking important information. In particular, users with low accounting literacy have difficulty quickly understanding the key points that should be reported from the list.

[0005] In view of the above circumstances, the present invention provides an information processing system and the like that can promote understanding of outliers. [Means for solving the problem]

[0006] According to one aspect of the present invention, there is provided an information processing system including at least one processor, wherein the processor acquires a set of numerical data including a plurality of numerical values ​​that are related to each other in an acquisition step, identifies an outlier in the acquired set of numerical data in an identification step, generates a sentence explaining the identified outlier, and displays a list of the acquired set of numerical data, and displays a sentence generated for an outlier included in the list in a manner indicating that the sentence corresponds to the outlier.

[0007] According to this embodiment, it is possible to promote understanding of outliers. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram illustrating an example of the overall configuration of an outlier presentation system 1. FIG. [Figure 2] 2 is a diagram illustrating an example of a hardware configuration of a server device 10. FIG. [Figure 3] FIG. 2 is a diagram illustrating an example of a hardware configuration of a user terminal 30. [Figure 4] FIG. 10 is an activity diagram illustrating an example of a list display process. [Figure 5] FIG. 10 is a diagram illustrating an example of a system screen. [Figure 6] FIG. 10 is a diagram showing an example of a displayed list. [Figure 7] FIG. 10 is a flowchart showing a first outlier identification process. [Figure 8] FIG. 10 is a flowchart showing a second outlier identification process. [Figure 9] FIG. 10 is a diagram showing an example of a displayed speech bubble image. [Figure 10] FIG. 10 is a diagram showing another example of a displayed speech bubble image. [Figure 11] FIG. 2 is a diagram showing another example of the configuration of the AI ​​device 20. [Figure 12] FIG. 10 is a diagram illustrating another example of the system screen. [Figure 13]FIG. 10 is a diagram illustrating another example of the system screen. [Figure 14] FIG. 10 is a diagram illustrating an example of an input data table. [Figure 15] FIG. 10 is a diagram illustrating another example of an input data table. [Figure 16] FIG. 10 is a diagram illustrating another example of an input data table. [Figure 17] FIG. 10 is a diagram illustrating an example of a character number table. [Figure 18] FIG. 10 is a diagram illustrating another example of a character number table. [Figure 19] FIG. 10 is a diagram illustrating another example of a character number table. [Figure 20] FIG. 10 is a diagram showing another example of a displayed speech bubble image. [Figure 21] FIG. 10 is a diagram showing an example of an explanatory text that is redisplayed. DETAILED DESCRIPTION OF THE INVENTION

[0009] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described below with reference to the accompanying drawings. Various features shown in the following embodiments can be combined with each other.

[0010] Incidentally, the program for realizing the software appearing in one embodiment may be provided as a non-transitory computer-readable medium, or may be provided so that it can be downloaded from an external server, or may be provided so that the program is started on an external computer and its functions are realized on a client terminal (so-called cloud computing).

[0011] Furthermore, various information processing according to an embodiment may realize input and output corresponding to the input. Here, the form of information referenced in such information processing (hereinafter referred to as reference information) is not limited as long as an output is obtained as a result of the input. The reference information may be, for example, rule-based information such as a database, a lookup table, or a predetermined function (including a decision formula such as a regression formula constructed using a statistical method), a trained model that has previously trained the correlation between input and output, or a large-scale language model that can output a desired result by inputting a prompt.

[0012] In one embodiment, a "unit" may include, for example, a combination of hardware resources implemented by a circuit in the broad sense and software information processing that can be specifically realized by these hardware resources. In one embodiment, various information is handled, and this information is represented, for example, by physical values ​​of signal values ​​representing voltage and current, high and low signal values ​​as a binary bit set consisting of 0 or 1, or quantum superposition (so-called quantum bits), and communication and calculations can be performed on a circuit in the broad sense.

[0013] Furthermore, a circuit in the broad sense is a circuit realized by at least an appropriate combination of a circuit, circuitry, processor, memory, etc. The processor may be a general-purpose processor or a dedicated circuit. That is, it includes an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)), etc.

[0014] <Embodiment> 1. System Configuration The system configuration according to the embodiment will be described below. Fig. 1 is a diagram showing an example of the overall configuration of an outlier presentation system 1. Fig. 1 shows an overview of each device included in the outlier presentation system 1 and users who use those devices. Each overview will be explained as needed with reference to other figures.

[0015] The outlier presentation system 1 is an information processing system that identifies outliers from a group of numerical data and performs information processing to present the identified outliers to a user. The group of numerical data is a group of data indicating various numerical values, such as financial data, performance data, sales data, equipment data, or experimental data. Each group of numerical data includes numerical values ​​such as monetary amounts, quantities of objects, or values ​​output by devices. An outlier is a numerical value included in a group of numerical data that deviates from the trends or rules indicated by the other numerical values. The outlier presentation system 1 includes a communication line 2, a server device 10, an AI device 20, and a user terminal 30.

[0016] The communication line 2 is not particularly limited, but may be configured, for example, by the Internet network. The communication line 2 may also include a local area network, a mobile communication network, a VPN (Virtual Private Network), etc. The communication line 2 mediates the exchange of data between devices connected to the line itself. In the example of Figure 1, the server device 10 and the AI ​​device 20 are connected to the communication line 2 by wire, and the user terminal 30 is connected wirelessly. Note that the connection of each device to the communication line 2 may be wired or wireless.

[0017] The server device 10 is an information processing device that identifies the outliers described above and provides a service of presenting the identified outliers to a user together with a set of numerical data. The server device 10 stores a numerical data set database DB1. The numerical data set database DB1 stores sets of numerical data.

[0018] The AI ​​device 20 is an information processing device that executes information processing using AI (Artificial Intelligence) technology. The AI ​​device 20 includes an artificial intelligence module 200. The artificial intelligence module 200 is a module adjusted to realize a predetermined function using AI technology, and hereinafter will also be referred to simply as "AI." The artificial intelligence module 200 has a natural language processing model whose accuracy has been improved by machine learning using a large data set called LLM (Large Language Models), for example, and realizes a sentence generation function that can generate natural sentences.

[0019] The user terminal 30 is a terminal for a user of a service provided by the server device 10, and is, for example, a personal computer, a smartphone, or a tablet terminal. The user terminal 30 is an information processing device that displays a service usage screen and accepts operations by service users. The server device 10 executes a display process for displaying an image on the user terminal 30 and an authentication process for authenticating a user who uses the user terminal 30.

[0020] The server device 10 performs processes such as generating and transmitting an HTML (Hyper Text Markup Language) file as display processing, and causes the user terminal 30 to display a web page showing a system screen using a browser function. Note that the user terminal 30 may install an application program for using the outlier presentation system 1, and the server device 10 may perform processes such as generating and transmitting display data in the application as display processing. The server device 10 controls the display of the user terminal 30 by performing these display processes.

[0021] The server device 10 stores authentication information (such as a user ID and a password) for authenticating a user who uses the outlier presentation system 1, and authenticates a user who inputs the authentication information. By authenticating a user, the server device 10 can restrict access to data or assign identification information to data input by the user to make the data identifiable.

[0022] 2. Hardware Configuration The hardware configuration according to the first embodiment will be described below. 2 is a diagram showing an example of the hardware configuration of server device 10. Server device 10 includes a control unit 11, a storage unit 12, a communication unit 13, and a bus 14. Bus 14 electrically connects the various units included in server device 10.

[0023] (Control unit 11) The control unit 11 has at least one processor. The at least one processor may be configured by, for example, a central processing unit (CPU), a micro processing unit (MPU), a graphics processing unit (GPU), one or more integrated circuits, one or more discrete circuits, or a combination thereof (not shown).

[0024] The control unit 11 is a computer that realizes various functions related to the outlier presentation system 1 by reading out predetermined programs stored in the storage unit 12. That is, information processing by software stored in the storage unit 12 is specifically realized by the control unit 11, which is an example of hardware, and can be executed as each functional unit included in the control unit 11. Note that the control unit 11 is not limited to being single, and may be implemented with multiple control units 11 for each function. Also, a combination of these may be used.

[0025] (Storage unit 12) The storage unit 12 stores various pieces of information defined above. This can be implemented, for example, as a storage device such as a solid state drive (SSD) or a hard disk drive (HDD) that stores various programs and the like related to the outlier presentation system 1 executed by the control unit 11, or as a memory such as a random access memory (RAM) that stores temporarily required information (arguments, arrays, etc.) related to program calculations. The storage unit 12 stores various programs, variables, etc. related to the outlier presentation system 1 executed by the control unit 11.

[0026] (Communications Department 13) The communication unit 13 is configured by a communication module. The communication module may be a wireless communication module conforming to standards such as IEEE802.11a / b / g / n / ac / ax, LTE, 5G, or 6G, or may be a wired communication module conforming to standards such as IEEE802.3. The communication unit 13 is configured to be able to transmit various electrical signals from the server device 10 to external components. The communication unit 13 is also configured to be able to receive various electrical signals from the external components to the server device 10. More preferably, the communication unit 13 has a network communication function, which allows various information to be communicated between the server device 10 and external devices via the communication line 2.

[0027] 2 has the same hardware configuration as the server device 10. In the following description of the AI ​​device 20, only the control unit 21 is assigned a different reference numeral from the control unit 11 of the server device 10.

[0028] Fig. 3 is a diagram showing an example of the hardware configuration of user terminal 30. User terminal 30 includes a control unit 31, a storage unit 32, a communication unit 33, an input unit 34, an output unit 35, and a bus 36. Bus 36 electrically connects the various units included in user terminal 30. Control unit 31, storage unit 32, and communication unit 33 are similar hardware to control unit 11, storage unit 12, and communication unit 13 shown in Fig. 3, although their specifications, models, etc. may differ.

[0029] (Input unit 34) The input unit 34 has keys, buttons, a touch screen, a mouse, etc., and receives input from the user. The input unit 34 may also have a microphone and have the function of receiving voice input from the user.

[0030] (Output section 35) The output unit 35 has a display, a speaker, etc., and displays visual information generated in a manner that is visible to the user, such as a screen, an image, an icon, or text, on the display surface of the display, and outputs sound including voice.

[0031] 3. Information Processing Information processing according to the embodiment will be described below. In the following description, the server device 10, the AI ​​device 20, and the user terminal 30 are described as the main actors in each information process, but the information processing is executed by at least one processor included in the outlier presentation system 1, i.e., a processor included in the control unit of each device. The outlier presentation system 1 executes a list display process that displays a list of numerical data groups and outliers included in the numerical data groups.

[0032] 4 is an activity diagram showing an example of the list display process. The list display process is executed in a state where a system screen for using the outlier presentation system 1 is displayed on the user terminal 30. Fig. 5 is a diagram showing an example of a system screen. System screen C1 shown in Fig. 5 displays a character string saying "Please specify the range of the numeric data group," an input field D11 for inputting target data, an input field D12 for inputting numeric data items, an input field D13 for inputting the period of the numeric data, and a list display button B11.

[0033] Input field D11 is a field for inputting the name of the group of numerical data for which a list is to be displayed. In the example of Figure 5, the name of the group of numerical data "Company A's Financial Data" is input as the target. Input field D12 is a field for inputting the items of numerical data to be included in the list. In the example of Figure 5, items such as "BBBB," "CCCC," "DDDD," and "EEEE" are input. Input field D13 is a field for inputting the period of the numerical data to be included in the list. Since financial data is a group of numerical data that changes over time, the range of the group of numerical data is determined by specifying the period. In the example of Figure 5, the period from "2022 / 04" to "2024 / 03" is input.

[0034] The target data entered into the input field D11 indicates a numeric data group stored in the numeric data group database DB1. Note that, in each input field, input may be performed by displaying the name, item, and period of the numeric data group stored in the numeric data group database DB1 in a pull-down list. The user terminal 30 accepts input operations into each input field and operations on the list display button B11 as operations for selecting a numeric data group (activity A11).

[0035] The user terminal 30 transmits request data indicating the selected numeric data group and a request to display a list of the numeric data groups to the server device 10. Upon receiving the request data, the server device 10 reads out the numeric data group indicated by the received request data from the numeric data group database DB1 (activity A12). Next, the server device 10 executes an outlier identification process to identify outliers from the read numeric data group (activity A13). The outlier identification process will be described later with reference to an example of a numeric data group.

[0036] Next, the server device 10 executes a sentence generation process to generate an explanatory sentence for the identified outlier (activity A14). In the example of FIG. 4, the server device 10 generates the explanatory sentence by instructing the artificial intelligence module 200 to generate the explanatory sentence. In detail, the server device 10 executes the sentence generation process by creating a sentence indicating an instruction to the artificial intelligence module 200, a so-called prompt, and transmitting the created prompt to the AI ​​device 20. The server device 10 instructs the AI ​​device 200 to generate the explanatory sentence within a predetermined number of characters, for example. The predetermined number of characters is a fixed number of characters, for example, about 100 to 200 characters.

[0037] The server device 10 also specifies words to be included in the explanatory text as input data using a prompt, and inputs the input data to the artificial intelligence module 200. The input data is, for example, numerical data that served as the basis for identifying the outlier. Specific examples of input data will be explained later. The AI ​​device 20 generates an explanatory text for the outlier using the artificial intelligence module 200 in accordance with the transmitted prompt (activity A21).

[0038] More specifically, the artificial intelligence module 200 generates explanatory text including the input data included in the prompt. Furthermore, if multiple outliers are identified, the artificial intelligence module 200 generates explanatory text for each outlier. The explanatory text may include only one sentence, or may include two or more sentences. The AI ​​device 20 transmits the explanatory text thus generated by the artificial intelligence module 200 to the server device 10.

[0039] The server device 10 stores the transmitted explanatory text in association with the outlier that the explanatory text explains (activity A22). In this way, the server device 10 generates the explanatory text using the artificial intelligence module 200. Next, the server device 10 generates a list of the numerical data groups read out in A12 (activity A23). The server device 10 transmits screen data showing the generated list to the user terminal 30. The user terminal 30 displays the list shown by the transmitted screen data (activity A24).

[0040] Fig. 6 is a diagram showing an example of a displayed list. On the system screen C2 shown in Fig. 6, a character string "List of Financial Data of Company A," a list display field D21 of numeric data groups, a vertical scroll bar B21, a horizontal scroll bar B22, and a back button B23 are displayed. A numeric data group E21 is displayed in the list display field D21. The numeric data group E21 shown in Fig. 6 is a part of the numeric data group E21, and other parts can be displayed by operating the vertical scroll bar B21 and the horizontal scroll bar B22.

[0041] In the numerical data group E21, outlier images F21 and F22 indicating the identified outliers are displayed. The outlier images F21 and F22 are thick-line frames surrounding the outliers H21 and H22, respectively. The outlier H21 is the numerical data "350000" of "2023 / 03" in the item "CCCC," and the outlier H22 is the numerical data "6400000" of "2023 / 02" in the item "DDDD." Here, the outlier identification process performed in A13 shown in FIG. 4 will be described using the outliers H21 and H22 as examples. In the example of FIG. 6, two types of outlier identification processes (first outlier identification process and second outlier identification process) are used.

[0042] 7 is a flow diagram showing the first outlier identification process. The server device 10 first extracts a numeric data group for one item from the numeric data group E21 (step Sa131). The term "numeric data group" refers to multiple pieces of numeric data, and is therefore used to refer to both the entire numeric data group E21 and multiple pieces of numeric data included in each item. In the example of FIG. 6, the server device 10 extracts, for example, 24 pieces of numeric data for each month from "2022 / 04" to "2024 / 03" as the numeric data group for the item "CCCC."

[0043] Next, the server device 10 determines whether the extracted numeric data group satisfies a first target condition (step Sa132). The first target condition is a condition that must be satisfied by a numeric data group that is to be identified as an outlier using the first outlier identification process. For example, the first target condition is a condition that is satisfied when the numeric data group indicates that it is distributed according to a certain law. This is because, if the numeric data group is distributed according to a certain law, values ​​that deviate from that distribution can be identified as outliers.

[0044] The distribution used as the first target condition is, for example, a continuous probability distribution such as a normal distribution, a continuous uniform distribution, or a gamma distribution. The following describes a case where a condition that is satisfied when a group of numerical data exhibits a normal distribution is used as the first target condition. In this case, the server device 10 determines whether the first target condition is satisfied using, for example, a well-known technique for determining whether a distribution is normal (such as the "Shapiro-Wilk test" or the "Kolmogorov-Smirnov test"). The determination of whether the distribution is normal may also be performed by calculating the mean value, standard deviation, etc. from the group of numerical data.

[0045] If the first target condition is not satisfied (if NO), the server device 10 does not identify outliers. If the first target condition is satisfied (if YES), the server device 10 selects one candidate value for the outlier from the extracted numeric data group (step Sa133). Next, the server device 10 calculates a statistical value of the numeric data group excluding the selected candidate value (step Sa134). The server device 10 calculates, for example, the average value of the numeric data group excluding the candidate value. Note that the average value may be an arithmetic average or a weighted average. Furthermore, the statistical value is not limited to the average value, and may be, for example, a median or a quartile.

[0046] Next, the server device 10 calculates the difference between the selected outlier candidate value and the calculated statistical value, and determines whether the calculated difference is equal to or greater than a threshold (step Sa135). The threshold used here may be a predetermined fixed value, or a value corresponding to the statistical value (e.g., 50% or 100% of the statistical value). The following describes a case where 50% of the statistical value is used as the threshold. If the calculated difference is equal to or greater than the threshold (YES), the server device 10 identifies the candidate value as an outlier (step Sa136). If the calculated difference is less than the threshold (NO), the server device 10 determines that the candidate value is not an outlier and skips Sa136.

[0047] In the example of FIG. 6, the group of numeric data items for the "CCCC" item contains many values ​​between 190,000 and 210,000, but the value for "2023 / 03" is significantly off at "350,000." If the server device 10 selects "350,000" as a candidate value, it calculates the average value of "200,000" as the statistical value of the group of numeric data items excluding the candidate value. In this case, since "100,000," which is 50% of "200,000," is the threshold, the difference between "350,000" and "200,000," or "150,000," is greater than the threshold. Therefore, the server device 10 identifies the numeric value "350,000" in Sa136 as an outlier H21. Furthermore, the server device 10 does not identify the other numeric values ​​for the "CCCC" item shown in FIG. 6 as outliers because their differences from the statistical values ​​are less than the threshold.

[0048] Then, server device 10 determines whether there is another candidate value for which outlier determination has not yet been performed in the group of numerical data extracted in Sa131 (step Sa137), and if it determines that there is another candidate value (YES), it returns to step Sa134 and continues operation. If it determines that there is no other candidate value (NO), or if it determines that the first target condition is not satisfied (NO) in Sa132, server device 10 determines whether there is another item for which outlier determination has not yet been performed (step Sa138). If it determines that there is another item (YES), server device 10 returns to step Sa131 and continues operation, and if it determines that there is no other item (NO), it ends the first outlier identification process.

[0049] 8 is a flow diagram showing the second outlier identification process. The server device 10 first extracts a numeric data group for one item from the numeric data group E21 (step Sb131). In the example of FIG. 6, the server device 10 extracts, for example, 24 pieces of numeric data for each month from "2022 / 04" to "2024 / 03" as the numeric data group for the item "DDDD."

[0050] Next, the server device 10 determines whether the extracted numerical data group satisfies a second target condition (step Sb132). The second target condition is a condition that must be satisfied by a numerical data group for which an outlier is to be identified using the second outlier identification process. The second target condition is, for example, not satisfying the first target condition. In this case, if the extracted numerical data group exhibits a normal distribution, i.e., satisfies the first target condition, the server device 10 does not identify an outlier because the second target condition is not satisfied (in this case, the first outlier identification process is executed). However, if the extracted numerical data group does not exhibit a normal distribution, i.e., does not satisfy the first target condition, the server device 10 satisfies the second target condition and therefore identifies an outlier.

[0051] If the group of numerical data satisfies the second target condition, the server device 10 selects one candidate value for the outlier from the extracted group of numerical data (step Sb133). Next, the server device 10 calculates a statistical value of the group of numerical data from a period prior to the selected candidate value (step Sb134). Here, the calculated statistical value may be, for example, an average value (arithmetic average value or weighted average value), a median value, or a quartile value, and the following describes the case where the average value is calculated.

[0052] If the candidate value is the outlier H22 shown in Figure 6, that is, the numerical data "6,400,000" for "2023 / 02" in the item "DDDD," the "previous period" would be the period before "2023 / 01." The period before "2023 / 01" has a numerical value of around 3.1 million, and in Sb134, for example, an average value of "3.1 million" is calculated.

[0053] Next, the server device 10 determines whether the difference between the selected outlier candidate value and the calculated statistical value is equal to or greater than a threshold value, and whether the groups of numerical data for the period before and after the candidate value both exhibit normal distributions (step Sb135). If the determination in Sb135 is YES, the server device 10 identifies the candidate value as an outlier (step Sb136), and if the determination in Sb135 is NO, the server device 10 determines that the candidate value is not an outlier and skips Sb136.

[0054] The threshold used in Sb135 is, for example, 50% of the calculated statistical value. For periods before "2023 / 01", the average value is the aforementioned "3.1 million", so 50% of that, or "1.55 million", is the threshold. Since the candidate value for "2023 / 02" is "6,400,000", the difference from the average value of "3.1 million" for periods before "2023 / 01" is greater than or equal to the threshold of "1.55 million". Furthermore, the numerical data for periods before "2023 / 01" is around 3.1 million, and the numerical data for periods after "2023 / 02" is around 6.45 million, both of which are assumed to show a normal distribution.

[0055] Therefore, when the server device 10 selects "6,400,000" for "2023 / 02" in the "DDDD" item as a candidate outlier, it determines YES in Sb135 and identifies "6,400,000" as outlier H22 in Sb136. The numerical data group for the "DDDD" item had an average of approximately 3.1 million up until "2023 / 01," but changed to an average of approximately 6.45 million from "2023 / 02" onwards. When the server device 10 executes the second outlier identification process, it identifies the numerical value for "2023 / 02" at this transition as an outlier.

[0056] For example, if "3,120,000" for "2023 / 01" is selected as the candidate value, the difference from the average value (approximately 3.1 million) for the period prior to "2022 / 12" is less than the threshold (approximately 1.55 million), so Sb135 determines "NO" and the value is not identified as an outlier. Similarly, if "6,510,000" for "2023 / 03" is selected as the candidate value, the average value will be higher because "6,400,000" for "2023 / 2" is added to the set of numerical data for the period prior to "2023 / 2." In this case, it is difficult to determine whether the difference from the statistical value exceeds the threshold. However, even if it does exceed the threshold, if the parameter of the numerical data set is not large, the set of numerical data for the period prior to "2023 / 2" will no longer show a normal distribution and will not be identified as an outlier.

[0057] Then, in step Sb137, the server device 10 performs the same operation as step Sa137 shown in Fig. 7 (determines whether there is another candidate value). In addition, in step Sb138, the server device 10 performs the same operation as step Sa138 shown in Fig. 7 (determines whether there is another item). Sb138 is also executed when it is determined in Sa132 that the second target condition is not satisfied (NO). When it is determined in Sb138 that there is no other item (NO), the server device 10 ends the second outlier identification process.

[0058] The user can know the outliers H21 and H22 included in the numerical data group from the outlier images F21 and F22 in the list displayed as shown in Fig. 6. Here, the user can perform an operation (such as tapping or hovering the mouse) to designate the displayed outlier. For example, when the user terminal 30 receives a designation operation for the outlier H21 (activity A31), the user terminal 30 transmits request data to the server device 10 requesting the display of an explanatory text that explains the designated outlier H21.

[0059] When the server device 10 receives the request data, it reads out the explanatory text stored in A22 in association with the outlier H21 indicated by the received request data (activity A32). Next, the server device 10 generates a speech bubble image including the read-out explanatory text (activity A33) and transmits it to the user terminal 30. The user terminal 30 displays the transmitted speech bubble image so as to point out the outlier H21 indicated by the user (activity A34).

[0060] Fig. 9 is a diagram showing an example of a displayed balloon image. On the system screen C2 shown in Fig. 9, a balloon image G21 is displayed to indicate the outlier H21. The balloon image G21 displays explanatory text J21. When the user next performs an operation to designate the outlier H22, operations A31 to A34 are performed, and explanatory text explaining the outlier H22 is displayed.

[0061] Fig. 10 is a diagram showing another example of a displayed speech bubble image. On the system screen C2 shown in Fig. 10, a speech bubble image G22 is displayed to indicate an outlier H22. The speech bubble image G22 displays explanatory text J22. As described above, the explanatory texts J21 and J22 are texts with a fixed number of characters, approximately 100 to 200 characters. Furthermore, the explanatory texts J21 and J22 are texts including numerical data that served as the basis for identifying the outlier.

[0062] In the case of explanatory text J21, the server device 10 uses, as input data, for example, "200,000," which is the average value of a group of numerical data in the same category as the outlier H21 and excluding the outlier H21, and the phrase "50%," which indicates a threshold percentage. As a result, the artificial intelligence module 200 generates an explanatory text including, for example, a sentence such as, "This outlier is a value that deviates by 50% or more from the average value of 200,000 for another period, ...." In the case of explanatory text J22, the server device 10 uses, as input data, for example, "3.1 million," which is the average value of a group of numerical data in the same category as the outlier H22, from a period prior to the outlier H22, and "6.45 million," which is the average value of a group of numerical data in the period after the outlier H22. As a result, the artificial intelligence module 200 generates an explanatory text including, for example, a sentence such as, "This outlier is a value when the average value changed from 3.1 million to 6.45 million, ...."

[0063] When the user operates the back button B23 displayed on the system screen C2, the user can return to the system screen C1 shown in Fig. 5 and reselect the range of the numerical data group. In this way, the user can select a range of the numerical data group to display a list, check the outliers contained therein using the outlier images, and by displaying the explanatory text for each outlier, understand the content of each outlier.

[0064] As described above, the server device 10 functions as an example of an acquisition unit that acquires a set of numerical data including multiple numerical values ​​that are related to each other. The "Financial data of Company A" specified in the example of Fig. 5 is an example of a set of numerical data. The numerical values ​​that change over time and are included in the "Financial data of Company A" are numerical values ​​for each month that are tallied for the same item, and are an example of multiple numerical values ​​that are related to each other (that is, are related as numerical values ​​tallied for the same item).

[0065] Next, the server device 10 functions as an example of an identification unit that identifies outliers in the acquired numerical data group. The server device 10 identifies outliers through the first outlier identification process and the second outlier identification process described above. The outliers H21 and H22 shown in FIG. 6 are examples of outliers. Note that in the case of a numerical data group that is not correlated with one another, it is not possible to identify an outlier because it does not show a certain trend or rule (if it does show a certain trend or rule, it would mean that there is some kind of correlation). For such a numerical data group, the server device 10 does not identify an outlier.

[0066] Next, the server device 10 functions as an example of a generation unit that generates a sentence explaining the identified outlier. In the example of FIG. 4 etc., the server device 10 generates the explanatory sentence by instructing the artificial intelligence module 200 to generate a sentence explaining the identified outlier. The artificial intelligence module 200 is an example of an artificial intelligence module having a sentence generation function. For example, as described above, the server device 10 instructs the artificial intelligence module 200 to generate the sentence by specifying the number of characters of the sentence and input data in a prompt.

[0067] The server device 10 functions as an example of a display control unit that displays a list of the acquired numerical data groups. The server device 10 (an example of a display control unit) displays explanatory text generated for outliers included in the displayed list in a manner indicating that the explanatory text corresponds to the outlier. The explanatory text J21 displayed in the speech bubble image G21 shown in FIG. 9 and the explanatory text J22 displayed in the speech bubble image G22 shown in FIG. 10 are examples of explanatory text displayed in a manner indicating that the explanatory text corresponds to the outliers H21 and H22, respectively. This aspect can promote understanding of outliers compared to when explanatory text is not displayed.

[0068] Furthermore, the numerical data group may include numerical data groups of multiple items, as shown in Fig. 6 etc. In this case, the server device 10 functions as an example of a determination unit that determines whether the distribution of the numerical data group of each item included in the multiple items shows a certain law. For example, in step Sa132 shown in Fig. 7, the server device 10 determines whether the numerical data group shows a normal distribution (an example of a distribution that shows a certain law).

[0069] The server device 10 (an example of an identification unit) then identifies outliers in the numerical data group of the item whose distribution (of the numerical data group of the item) is determined to follow a certain rule. If the numerical data group follows a certain rule, it is possible to identify numerical values ​​that deviate from that rule as outliers. However, if the numerical data group does not follow that rule, it is difficult to identify outliers. According to this embodiment, it is possible to more reliably identify outliers than when the rule of distribution of the numerical data group is not taken into consideration.

[0070] Furthermore, in the first outlier identification process described above, the server device 10 (an example of an identification unit) identifies outliers by comparing candidate values ​​that are candidates for outliers included in the acquired numerical data group with statistical values ​​(such as arithmetic mean values) of the numerical data group excluding the candidate values. If candidate values ​​are included when calculating statistical values, the statistical values ​​may become inappropriate, and values ​​that should be outliers may not be identified as outliers. Therefore, by excluding candidate values ​​as described above, outliers can be identified more accurately than when statistical values ​​are calculated including candidate values.

[0071] Furthermore, the server device 10 (an example of an identification unit) identifies a candidate value as an outlier if the difference between the statistical value of the group of numerical data for the period before the candidate value is equal to or greater than a threshold, and the distribution of the group of numerical data for the period before the candidate value and the period after the candidate value both show a certain pattern. In the example of Figure 6, the server device 10 identifies the numerical value for "2023 / 02," which is the first month in which the average value changed from approximately 3.1 million to approximately 6.45 million, as an outlier. According to this aspect, if a numerical value has changed since a certain period, it is possible to determine the time of the change.

[0072] Furthermore, the server device 10 (an example of a generation unit) instructs the artificial intelligence module 200 to generate a sentence within a predetermined number of characters. In the examples of FIGS. 9 and 10, the predetermined number of characters is a fixed number of about 100 to 200 characters. According to this embodiment, it is less likely that the sentence will be too long and other numerical values ​​will be hidden behind the explanatory sentence, and it is possible to make it less likely that the displayed sentence will interfere with the numerical values ​​compared to when the number of characters is not limited.

[0073] <Variation: Item extraction by AI> As shown in Figure 6, the acquired numerical data group may include numerical data groups for multiple items. In this case, although the example in Figure 5 shows that the numerical data group is for an item selected by the user, the numerical data group may be for an item extracted by AI.

[0074] FIG. 11 is a diagram showing another example of the configuration of the AI ​​device 20. In FIG. 11, the AI ​​device 20 includes a second artificial intelligence module 210 in addition to the artificial intelligence module 200. The second artificial intelligence module 210 realizes an item extraction function that, when multiple items and a predetermined group are input, extracts items that belong to the input group. For example, if the group of numerical data is financial data, the predetermined group is a group such as "labor costs," "legal welfare expenses," "selling and general administrative expenses," or "manufacturing costs."

[0075] For example, when "selling and general administrative expenses" is input as a group, the second artificial intelligence module 210 performs machine learning to extract items such as "salary allowances," "bonus reserve appropriations," "executive compensation," and "statutory welfare expenses" as items belonging to "selling and general administrative expenses." The user terminal 30 displays a screen for inputting a group.

[0076] Figure 12 is a diagram showing another example of a system screen. System screen C3 shown in Figure 12 displays a character string saying "Please specify the range of the numerical data group," an input field D31 for target data, an input field D32 for group, an input field D33 for item, an input field D34 for period, and a list display button B31. In the example of Figure 12, a group called "selling and general administrative expenses" has been entered in input field D32, and no item has been entered in input field D33.

[0077] When the list display button B31 is operated in this state, the user terminal 30 transmits request data indicating a request to display a list of the numerical data groups of the items belonging to the input group to the server device 10. Upon receiving the request data, the server device 10 functions as an example of an instruction unit that instructs the second artificial intelligence module 210 to extract items belonging to a predetermined group from the multiple items. Then, the server device 10 (an example of an identification unit) identifies outliers in the numerical data groups of the items extracted by the second artificial intelligence module 210.

[0078] According to this embodiment, the user can grasp the outliers of each item belonging to a group of interest by simply inputting the name of the group, without having to input each item belonging to the group one by one. Also, even if the user does not know the items belonging to a group, the AI ​​can determine the items belonging to that group and grasp the outliers of those items.

[0079] In addition, the user may be able to select further items from the items extracted by the AI. Fig. 13 is a diagram showing another example of a system screen. The system screen C4 shown in Fig. 13 displays a character string "Please select items to include in the list," an input field D41 for target data, an input field D42 for group, an input field D43 for item, an input field D44 for period, an extract item button B41, and a list display button B42.

[0080] When the target data name is entered in input field D41, the period is entered in input field D44, and the group is entered in input field D42, and then the extract button B41 is operated, the second artificial intelligence module 210 extracts items belonging to the entered group, and the extracted items are displayed in input field D43. In the example of Fig. 13, the items "salary allowance," "bonus reserve provision amount," "executive compensation," and "statutory welfare expenses" that have been extracted as items belonging to the group "selling and general administrative expenses" are displayed.

[0081] A check box B43 is displayed in association with each item. The user can select the corresponding item by checking the check box B43 (white circle in the example of FIG. 13), and can deselect the corresponding item by not checking it (black circle in the example of FIG. 13). In the example of FIG. 13, only "executive compensation" has been deselected. When the list display button B42 is operated in this state, the server device 10 identifies outliers in the group of numerical data for the item selected by the user. According to this aspect, it is possible to identify outliers in items that are in a group of interest and that have been selected by a person.

[0082] The second artificial intelligence module 210 may perform machine learning using the results of the user's selection of items as training data. For example, when the user selects a group as described above and then selects an item belonging to that group, the server device 10 instructs the second artificial intelligence module 210 to perform machine learning using training data that takes the selected group as an input and the selected item as a correct answer. According to this embodiment, the items extracted by the AI ​​as items belonging to the group become more accurate.

[0083] <Variation: Method for identifying outliers> The server device 10 may identify outliers using a method different from the method using the first outlier identification process (first method) and the method using the second outlier identification process (second method) described above. For example, the server device 10 may perform regression analysis on a group of numerical data for the same item to calculate an approximate formula, and identify values ​​that deviate from the calculated approximate formula by a predetermined value or more as outliers (regression analysis method). Furthermore, when the numerical value of item A is correlated with the numerical value of item B, the server device 10 may identify the numerical value of item A that does not show the correlation with the numerical value of item B as outliers (multi-item correlation method).

[0084] Furthermore, the AI ​​device 20 may be provided with an artificial intelligence module (preferably other than the artificial intelligence module 200 and the second artificial intelligence module 210) that realizes a function for identifying outliers from a group of numerical data, and the server device 10 may input the group of numerical data to the artificial intelligence module to identify the outliers (AI utilization method). This artificial intelligence module can improve the accuracy of identifying outliers by performing machine learning using input of the group of numerical data and the correct answers for the outliers contained in the group of numerical data as training data.

[0085] The server device 10 may also identify outliers using other well-known methods. The server device 10 may execute these multiple identification methods in parallel, or may execute only an identification method selected from the multiple identification methods. The identification method may be selected by setting conditions for using the identification method, as described with reference to FIGS. 7 and 8, or may be selected based on a user operation.

[0086] <Variation: Change in input to AI> Although the server device 10 uses the average value and threshold value of the group of numerical data of the same item as the outlier as the input data to be input to the artificial intelligence module 200 when instructing the generation of an explanatory text, the input data is not limited to this. The input data may be, for example, the outlier itself, the numerical data itself included in the group of numerical data of the same item as the outlier (e.g., the numerical value of the previous month or the numerical value of the same month of the previous year), other statistical values ​​of the group of numerical data (variance, standard deviation, minimum value, maximum value, median, etc.), or numerical data of other items that are correlated with the outlier.

[0087] These input data are preferably the numerical data that are the basis of the outliers, but are not limited to this and may be any numerical data related to the outliers. Numerical data related to the outliers is, for example, numerical data (such as variance or standard deviation) that indicates that the distribution of the numerical data group of items including the outliers follows a certain pattern. The input data may be one, or two or more. The number and type of input data may be variable.

[0088] For example, the importance of an outlier may be used as a parameter that makes input data variable. In this case, the server device 10 functions as an example of a determination unit that determines the importance of a specified outlier. For example, the server device 10 determines that the larger the numerical value of the outlier, the more important it is. The server device 10 also determines the importance depending on the time of the outlier. For example, the server device 10 determines that the importance of an outlier during a busy period is "high," the importance of an outlier during an off-season is "low," and the importance of an outlier during other periods is "medium."

[0089] Furthermore, the server device 10 determines the importance level depending on the outlier item. For example, the server device 10 determines the importance level of an outlier related to cost as "high," the importance level of an outlier related to profit as "medium," and the importance level of an outlier related to sales as "low." Note that these determination methods are merely examples, and the importance level may be determined in two or four or more stages, or other well-known determination methods may be used.

[0090] The server device 10 (an example of a generator) instructs the generation of a sentence based on input data. This input data is the numerical data that is the basis of the outlier or numerical data related to the outlier, and the higher the importance determined for the outlier, the more data there is. The server device 10 uses an input data table that associates the importance of the outlier with the type of input data to be used.

[0091] FIG. 14 is a diagram showing an example of an input data table. In the input data table TB1 shown in FIG. 14, the importance levels of "low," "medium," and "high" are associated with the types of input data of "average value of other months," "average value of other months, threshold," and "average value, minimum value, maximum value, and threshold value of other months." After determining the importance of the identified outlier, the server device 10 inputs, as input data, the type of numerical data associated with the determined importance level in the input data table TB1 into the artificial intelligence module 200. According to this aspect, the more important the outlier, the more input data there is, allowing for more detailed explanation.

[0092] The parameter that makes the input data variable is not limited to the importance. For example, the input data may be a number of pieces of data corresponding to the degree to which the outlier is different from the set of numerical data. In this case, the server device 10 uses an input data table that associates the degree of difference of the outlier with the type of input data to be used.

[0093] Fig. 15 is a diagram showing another example of an input data table. In the input data table TB2 shown in Fig. 15, the degrees of outliers, "small," "medium," and "large," are associated with the types of input data, "average value of other months," "average value and threshold of other months," and "average value, minimum value, maximum value, and threshold of other months." The server device 10 first determines the degree of outlier of the identified outlier.

[0094] For example, when the server device 10 identifies a candidate value whose difference from the statistical value is 50% or more of the statistical value as an outlier through the first outlier identification process, the server device 10 determines the degree of outlier of the outlier as "small" if the difference between the outlier and the statistical value is 50% or more and less than 100% of the statistical value, "medium" if the difference is 100% or more and less than 150% of the statistical value, and "large" if the difference is 150% or more of the statistical value. Note that the method of determining the degree of outlier is not limited to this, and other well-known methods may be used.

[0095] The server device 10 inputs, as input data, the type of numerical data associated in the input data table TB2 with the degree of deviation of the determined outlier to the artificial intelligence module 200. According to this embodiment, the greater the degree of deviation of the outlier, the more input data there is, allowing for a more detailed explanation. The greater the degree of deviation, the more likely it is that multiple factors are influencing the value. Therefore, by increasing the input data, it becomes possible to appropriately explain even values ​​with a large degree of deviation.

[0096] Furthermore, the server device 10 (an example of an identification unit) may identify an outlier using any one of a plurality of identification methods for an outlier. The plurality of identification methods include, for example, the above-described first method (a method using a first outlier identification process), second method (a method using a second outlier identification process), regression analysis method (a method using regression analysis), multi-item correlation method (a method based on a correlation with the numerical values ​​of other items), and AI utilization method (a method using AI).

[0097] When multiple identification methods are used in this way, the input data may be the numerical data that is the basis of the outlier and may be data of a type associated with the identification method used to identify the outlier. In this case, the server device 10 uses an input data table that associates the outlier identification method with the type of input data to be used.

[0098] Fig. 16 is a diagram showing another example of an input data table. In the input data table TB3 shown in Fig. 16, the outlier identification methods, namely, "first method," "second method," "regression analysis method," "multiple item correlation method," and "AI utilization method," are associated with the input data types, namely, "average value of other months," "average value up to the previous month, average value from the current month onward," "approximation formula," "item to be compared, compared numerical value," and "outlier factor output by AI."

[0099] The AI ​​can output outlier factors by, for example, inputting a set of numerical data, an outlier contained in the set of numerical data, and a correct answer for the numerical data that causes the outlier as training data into an artificial intelligence module and causing it to perform machine learning. The server device 10 identifies outliers using, for example, one of a plurality of identification methods designated by a user. The server device 10 also defines other conditions for determining whether to use each identification method, such as the first target condition shown in FIG. 7 and the second target condition shown in FIG. 8, and when these conditions are met, identifies outliers using the identification method corresponding to the condition.

[0100] When the server device 10 identifies an outlier using any of the identification methods described above, it inputs, as input data, the type of numerical data associated in the input data table TB3 with the identification method used to identify the outlier to the artificial intelligence module 200. According to this aspect, each outlier can be explained in accordance with the characteristics of the outlier, more specifically, based on appropriate input data that matches the characteristics of the outlier.

[0101] <Variation: Number of characters in the explanatory text> The server device 10 instructs the artificial intelligence module 200 to generate an explanatory text within a predetermined number of characters. While the predetermined number of characters is a fixed number of characters in the above example, it is not limited to this. For example, assume that the server device 10 (an example of a determination unit) determines the importance of the identified outlier as described above. In this case, the server device 10 (an example of a generation unit) may instruct the artificial intelligence module 200 to increase the predetermined number of characters as the determined importance of the outlier increases.

[0102] In this case, the server device 10 uses a character number table that associates the degree of outlier of an outlier with the number of characters in the explanatory text. Fig. 17 is a diagram showing an example of a character count table. In the character count table TB4 shown in Fig. 17, the importance levels of "low," "medium," and "high" are associated with the character counts of explanatory texts of "less than N11," "N11 or more but less than N12," and "N12 or more" (N11 and N12 are natural numbers).

[0103] When the server device 10 determines the importance of the identified outlier, it instructs the artificial intelligence module 200 to generate an explanatory text with the number of characters associated with the determined importance in the character number table TB4. According to this aspect, the more important the outlier, the more characters there are, allowing for a more detailed explanation.

[0104] The parameter for varying the number of characters in the explanatory text is not limited to the importance. For example, the server device 10 (an example of a generating unit) may instruct the predetermined number of characters to be a number corresponding to the degree to which the outlier deviates from the set of numerical data. In this case, the server device 10 uses a character number table that associates the degree of deviation of the outlier with the number of characters in the explanatory text.

[0105] Fig. 18 is a diagram showing another example of a character count table. In the character count table TB5 shown in Fig. 18, the degrees of outliers, "small," "medium," and "large," are associated with the numbers of characters in the explanatory text, "less than N21," "N21 or greater but less than N22," and "N22 or greater" (N21 and N22 are natural numbers). The server device 10 first determines the degree of outlier of the identified outlier as described above.

[0106] The server device 10 instructs the artificial intelligence module 200 to generate explanatory text with the number of characters associated with the degree of deviation of the determined outlier in the character number table TB5. The greater the degree of deviation, the more likely it is that multiple factors are influencing the text, and the more text is likely to require explanation. Therefore, according to the above embodiment, the amount of explanatory text can be made appropriate more easily than when the number of characters is uniform.

[0107] As described above, it is assumed that the server device 10 (an example of an identification unit) identifies an outlier using one of a plurality of identification methods for an outlier. In this case, the server device 10 (an example of a generation unit) may instruct the predetermined number of characters to be a number according to the identification method used to identify the outlier. The server device 10 uses a character number table that associates the outlier identification method with the number of characters in the explanatory text.

[0108] Fig. 19 is a diagram showing another example of a character count table. In the character count table TB6 shown in Fig. 19, the outlier identification methods "first method," "second method," "regression analysis method," "multi-item correlation method," and "AI utilization method" are associated with the character counts of the explanatory text "N31 or more and less than N32," "N41 or more and less than N42," "N51 or more and less than N52," "N61 or more and less than N62," and "N71 or more and less than N72" (N31, N32, etc. are all natural numbers).

[0109] When the server device 10 identifies an outlier using one of the identification methods as described above, it instructs the artificial intelligence module 200 to generate explanatory text using the number of characters associated with the identification method used to identify the outlier in the character count table TB6. Among the outlier identification methods, for example, the second method is more complex than the first method, and the explanatory text tends to be longer. Furthermore, compared to the first method, the second method, and the regression analysis method, which compare numerical values ​​of the same item, the multi-item correlation method, which uses correlations with numerical values ​​of other items, tends to produce longer explanatory text. Furthermore, the AI-based method does not identify outliers based on fixed rules like other methods, so the length of the explanatory text tends to vary.

[0110] In this way, the tendency of the amount of explanatory text is determined for each identification method. Note that the tendency described here is just an example, and the tendency will vary depending on the type of statistical value and the type of approximation formula used in each identification method. However, regardless of which tendency appears, by instructing the generation of explanatory text so that the number of characters matches that tendency, it is possible to explain the outlier with a number of characters that matches the tendency of the amount of explanatory text for the outlier that appears for each identification method.

[0111] <Variation: simultaneous display of text> In the above example, the server device 10 displays an explanatory text for an outlier that is specified by a user's instruction operation among multiple outliers, but multiple explanatory texts may be displayed simultaneously. For example, the server device 10 (an example of a determination unit) determines the importance of the outlier identified as described above. Then, the server device 10 (an example of a display control unit) simultaneously displays the generated text for the identified outliers whose determined importance falls within a predetermined range.

[0112] 20 is a diagram showing another example of a displayed speech bubble image. In the system screen C2 shown in FIG. 20, the server device 10 simultaneously displays a speech bubble image G21 showing explanatory text J21 explaining the outlier H21 and a speech bubble image G22 showing explanatory text J22 explaining the outlier H22. The outliers H21 and H22 are both outliers whose importance falls within a predetermined range. If the importance is expressed as a range from 0 to 100, the predetermined range here may be any range, such as a high importance range of 70 or more, a medium importance range of 30 to 70, or a low importance range of less than 30 (the numerical values ​​of the importance are an example). According to this aspect, outliers whose importance falls within the same range can be easily compared.

[0113] The server device 10 (an example of a display control unit) may simultaneously display sentences generated for outliers whose degree of deviation from the group of numerical data falls within a predetermined range. In this case, outliers with similar degrees of deviation can be easily compared. The server device 10 (an example of a display control unit) may also simultaneously display sentences generated for outliers that have been identified using a common identification method. In this case, outliers that have been identified using a common identification method can be easily compared.

[0114] <Variation: Text switching display> The server device 10 may display a new explanatory sentence in place of an explanatory sentence that has been displayed once. Specifically, when an operation to redisplay an explanatory sentence that has been displayed is performed, the server device 10 (an example of a display control unit) displays, in place of the explanatory sentence, an explanatory sentence that is separately generated for the outlier that the explanatory sentence explains.

[0115] Fig. 21 is a diagram showing an example of explanatory text that is redisplayed. Fig. 21 shows a state in which the outlier H21 shown in Fig. 9, explanatory text J21 for the outlier H21, and balloon image G21 are initially displayed. Assume that a redisplay operation is performed in this state. The redisplay operation is, for example, an operation of clicking or tapping on balloon image G21. However, the redisplay operation is not limited to this, and may also be an operation of double-clicking, double-tapping, or right-clicking to select redisplay from the displayed menu, etc.

[0116] When the redisplay operation is performed, the server device 10 instructs the artificial intelligence module 200 to generate explanatory text with a larger number of characters than the explanatory text J21. For example, if the server device 10 has generated explanatory text J21 as a text of 200 characters or less, the server device 10 generates explanatory text J21a as a text of 200 to 400 characters. The server device 10 displays the explanatory text J21a generated in response to the instruction and the speech bubble image G21a containing it in a manner indicating that it corresponds to the outlier H21.

[0117] Furthermore, when a redisplay operation is performed while explanatory text J21a is displayed, server device 10 instructs artificial intelligence module 200 to generate explanatory text with a larger number of characters than explanatory text J21a. Server device 10 generates a text of, for example, 400 to 600 characters and displays the generated explanatory text J21b and a speech bubble image G21b containing it in a manner indicating that it corresponds to outlier value H21. In this manner, a user can easily learn new information by performing a redisplay operation.

[0118] In the above example, the server device 10 generates a new explanatory text when a redisplay operation is performed, but the explanatory text may be generated in advance and switched each time a redisplay operation is performed. Also, in the above example, the server device 10 generates explanatory text with an increased number of characters for redisplay, but this is not limiting. For example, the server device 10 may generate explanatory text for redisplay by increasing the amount of input data or changing the type of input data. In either case, by performing a redisplay operation, a text explaining the new content is displayed, allowing the user to easily learn new information.

[0119] <Modification: Other Configurations> In the above example, the server device 10 displays a list of numeric data groups for multiple items as shown in Fig. 6, etc. However, it may also display a list of numeric data groups for only one item. Even in this case, it is possible to identify outliers from the numeric data group for that item.

[0120] Furthermore, although the server device 10 indicates that the explanatory text corresponds to the outlier by displaying a speech bubble image, the manner in which the correspondence between the explanatory text and the outlier is not limited to this. For example, the server device 10 may display the explanatory text in a manner in which the color of the frame in which the explanatory text is displayed is the same as the color of the frame in which the outlier is displayed, or may display the explanatory text in a manner in which the thickness or type of the frame is the same. Furthermore, the server device 10 may display the explanatory text and the outlier in a manner in which the same symbol is assigned to both the explanatory text and the outlier, or may display the explanatory text in a manner in which the explanatory text is in a fixed positional relationship with the outlier.

[0121] 4, the server device 10 generates the explanatory text by instructing the artificial intelligence module 200 to generate the explanatory text, but the invention is not limited to this and the explanatory text may be generated using a specific logic based on input information. As the specific logic, for example, when "numerical data that is the basis of the outlier" (the average value of numerical data other than the outlier) and an "outlier" are input, a logic is used in which the "difference" between the numerical data and the outlier is calculated, and a text is generated explaining that the "outlier" is due to the difference between the "numerical data that is the basis of the outlier" and the "difference."

[0122] Another specific logic is, for example, a logic that, when "numerical data related to outliers" (variance, standard deviation, etc.) is input, generates a sentence that explains the trend of the numerical data of an item that includes the outlier using the "numerical data related to the outlier." According to this aspect, it is possible to generate a more stable explanatory sentence than when using AI. These logics are merely examples, and any logic may be used as long as it generates a sentence that explains the outlier.

[0123] <Example of variation: Variation of composition> The configurations (overall configuration, hardware configuration, functional configuration, etc.) shown in FIG. 1 and elsewhere are merely examples, and other configurations may be used as long as they are not inconvenient for implementation. For example, the server device 10 and the AI ​​device 20 may each be distributed across two or more devices, or may be provided in the form of SaaS (Software as a Service) or a cloud computing system. Furthermore, the information processing performed by the server device 10, the AI ​​device 20, and the user terminal 30 may be collectively performed by a device that integrates these devices (a device that integrates two or three of them). In short, as long as the necessary information processing is performed by the outlier presentation system 1 as a whole, the devices that perform this information processing may have any configuration.

[0124] The output destination of information or data (hereinafter referred to as "information, etc.") may be another device, a display, a memory unit (including an internal memory unit and an external memory unit), an email address, an account of another system, etc. Acquisition of information, etc. includes acquiring information, etc. generated by the device itself, as well as acquiring information, etc. transmitted from another device. The table, etc. (table, database, etc.) in which parameters are associated is not limited to the illustrated table, etc., and the number of parameters may be reduced or increased. Furthermore, information, etc. corresponding to parameters may be obtained using a mathematical formula, a conditional formula, etc., without using a table, etc.

[0125] The above-described embodiments are information processing devices such as the server device 10 and the user terminal 30, and information processing systems such as the outlier presentation system 1 including the server device 10, the AI ​​device 20, and the user terminal 30. However, the embodiments may also be information processing methods. The information processing methods include the same steps as those executed by the information processing system. The above-described embodiments may also be programs. The programs cause a computer to execute the same steps as those executed by the information processing system.

[0126] <Additional Notes> Furthermore, it may be provided in the following aspects.

[0127] (1) An information processing system having at least one processor, wherein the processor acquires a group of numerical data including a plurality of mutually related numerical values ​​in an acquisition step, identifies an outlier in the acquired group of numerical data in an identification step, generates a sentence explaining the identified outlier in a generation step, and displays a list of the acquired group of numerical data and displays the sentence generated for the outlier included in the list in a manner indicating that it corresponds to the outlier.

[0128] According to this embodiment, it is possible to promote understanding of outliers.

[0129] (2) In the information processing system described in (1) above, the numerical data group includes numerical data groups of multiple items, and in the determination step, the processor determines whether the distribution of the numerical data group of each item included in the multiple items shows a certain pattern, and in the identification step, identifies outliers in the numerical data group of an item whose distribution is determined to show a certain pattern.

[0130] According to this aspect, it is possible to more reliably identify outliers.

[0131] (3) In the information processing system described in (1) or (2) above, in the identification step, the processor identifies the outlier by comparing a candidate value that is a candidate for an outlier included in the acquired group of numerical data with a statistical value of the group of numerical data excluding the candidate value.

[0132] According to this aspect, it is possible to more accurately identify outliers.

[0133] (4) In the information processing system described in (3) above, in the identification step, the processor identifies the candidate value as an outlier if the difference between the statistical value of the group of numerical data for the period before the candidate value and the candidate value is equal to or greater than a threshold, and the distributions of the group of numerical data for the period before the candidate value and the period after the candidate value both show a certain pattern.

[0134] According to this aspect, if the numerical value has been changed since a certain time, it is possible to know when the change occurred.

[0135] (5) In the information processing system described in any one of (1) to (4) above, the acquired numerical data group includes numerical data groups of multiple items, and in the instruction step, the processor instructs a second artificial intelligence module to extract items belonging to a predetermined group from the multiple items, and in the identification step, identifies outliers in the numerical data group of the items extracted by the second artificial intelligence module.

[0136] According to this aspect, it is possible to grasp outliers in a group of interest.

[0137] (6) In the information processing system described in any one of (1) to (5) above, in the generation step, the processor generates the sentence by instructing an artificial intelligence module having a sentence generation function to generate a sentence explaining the identified outlier.

[0138] According to this aspect, it is possible to generate sentences.

[0139] (7) In the information processing system described in (6) above, the processor determines the importance of the identified outlier in the determination step, and instructs the processor to generate the sentence based on input data in the generation step, the input data being numerical data that is the basis of the outlier or numerical data related to the outlier, and the higher the determined importance of the outlier, the more data there is.

[0140] According to this aspect, the more important the outlier, the more detailed the explanation can be.

[0141] (8) In the information processing system described in (6) above, in the generation step, the processor instructs the processor to generate the sentence based on input data, and the input data is numerical data that is the basis of the outlier or numerical data related to the outlier, and the number of data corresponds to the degree to which the outlier deviates from the group of numerical data.

[0142] According to this aspect, the greater the degree of deviation of an outlier, the more detailed the explanation can be.

[0143] (9) In the information processing system described in (6) above, in the identification step, the processor identifies the outlier using one of a plurality of identification methods for the outlier, and in the generation step, instructs the processor to generate the sentence based on input data, the input data being numerical data that is the basis of the outlier or numerical data related to the outlier, and being a type of data associated with the identification method used to identify the outlier.

[0144] According to this aspect, it is possible to explain the outliers in accordance with their characteristics.

[0145] (10) In the information processing system described in any one of (6) to (9) above, in the generation step, the processor instructs the artificial intelligence module to generate the sentence within a predetermined number of characters.

[0146] According to this aspect, it is possible to prevent the text from interfering with the numerical values.

[0147] (11) In the information processing system described in (10) above, in the determination step, the processor determines the importance of the identified outlier, and in the generation step, the higher the importance determined for the outlier, the more the specified number of characters is increased when issuing the instruction.

[0148] According to this aspect, the more important the outlier, the more detailed the explanation can be.

[0149] (12) In the information processing system described in (10) or (11) above, in the generation step, the processor issues the instruction with the predetermined number of characters being a number corresponding to the degree to which the outlier deviates from the group of numerical data.

[0150] According to this aspect, it is possible to make it easier to adjust the amount of explanatory text to an appropriate amount.

[0151] (13) In the information processing system described in any one of (10) to (12) above, in the identification step, the processor identifies the outlier using one of a plurality of identification methods for the outlier, and in the generation step, the processor issues the instruction with the predetermined number of characters corresponding to the identification method used to identify the outlier.

[0152] According to this aspect, the outlier can be explained using a number of characters that matches the tendency of the amount of explanatory text for the outlier that appears for each identification method.

[0153] (14) In the information processing system described in any one of (1) to (13) above, the processor determines the importance of the identified outlier in the determination step, and in the display control step, simultaneously displays the sentences generated for the identified outliers whose determined importance falls within a predetermined range.

[0154] According to this embodiment, outliers in the same range of importance can be easily compared.

[0155] (15) In the information processing system described in any one of (1) to (14) above, in the display control step, when an operation to redisplay the displayed sentence is performed on the displayed sentence, the processor displays a sentence that has been separately generated about the outlier that the sentence explains instead of the displayed sentence.

[0156] According to this embodiment, new information can be easily obtained.

[0157] (16) An information processing method, in which a processor included in an information processing system executes each step of the information processing system described in any one of (1) to (15) above.

[0158] According to this embodiment, it is possible to promote understanding of outliers.

[0159] (17) A program that causes a computer to execute each step of the information processing system according to any one of (1) to (15) above.

[0160] According to this embodiment, it is possible to promote understanding of outliers. Of course, this is not the case. Furthermore, the above-described embodiments and modifications may be combined in any desired manner.

[0161] Finally, while various embodiments of the present invention have been described, these are presented by way of example only and are not intended to limit the scope of the invention. The novel embodiments may be embodied in various other forms, and various omissions, substitutions, and modifications may be made without departing from the spirit of the invention. The embodiments and their modifications are intended to be included within the scope and spirit of the invention, as well as within the scope of the inventions and their equivalents as defined in the appended claims. [Explanation of symbols]

[0162] 1: Value detection system 2: Communication line 10: Server device 11: Control section 20:AI device 21: Control unit 30: User terminal 31: Control unit 200: Artificial Intelligence Module 210: Second AI module

Claims

1. An information processing system including at least one processor, the processor: In the acquisition step, a group of numerical data including a plurality of numerical values ​​related to each other is acquired; In the identifying step, an outlier is identified in the acquired group of numerical data; In the generating step, a sentence describing the identified outlier is generated; In the display control step, a list of the acquired numerical data groups is displayed, and the sentences generated for the outliers included in the list are displayed in a manner indicating that they correspond to the outliers. Information processing system.

2. 2. The information processing system according to claim 1, The numerical data group includes numerical data groups of a plurality of items, the processor: In the determination step, it is determined whether or not the distribution of the numerical data group of each item included in the plurality of items shows a certain rule; In the identifying step, an outlier is identified in the group of numerical data of the item whose distribution is determined to show a certain law. Information processing system.

3. 2. The information processing system according to claim 1, the processor: In the identifying step, a candidate value that is a candidate for an outlier included in the acquired group of numerical data is compared with a statistical value of the group of numerical data excluding the candidate value to identify the outlier. Information processing system.

4. 4. The information processing system according to claim 3, the processor: In the identifying step, if a difference between the candidate value and a statistical value of a group of numerical data for a period preceding the candidate value is equal to or greater than a threshold, and the distributions of the groups of numerical data for the period preceding the candidate value and the period following the candidate value both show a certain pattern, the candidate value is identified as an outlier. Information processing system.

5. 2. The information processing system according to claim 1, The acquired set of numerical data includes sets of numerical data for a plurality of items, the processor: In the instruction step, the second artificial intelligence module is instructed to extract items belonging to a predetermined group from the plurality of items; In the identifying step, the second artificial intelligence module identifies outliers in the group of numerical data of the extracted items. Information processing system.

6. 2. The information processing system according to claim 1, the processor: In the generating step, an artificial intelligence module having a sentence generation function is instructed to generate a sentence explaining the identified outlier, thereby generating the sentence. Information processing system.

7. 7. The information processing system according to claim 6, the processor: In the determination step, the importance of the identified outlier is determined; In the generating step, an instruction is given to generate the sentence based on input data; The input data is numerical data that is the basis of the outlier or numerical data related to the outlier, and the higher the importance determined for the outlier, the more data there is. Information processing system.

8. 7. The information processing system according to claim 6, the processor: In the generating step, an instruction is given to generate the sentence based on input data; The input data is numerical data that is the basis of the outlier or numerical data related to the outlier, and the number of pieces of data corresponds to the degree to which the outlier deviates from the group of numerical data. Information processing system.

9. 7. The information processing system according to claim 6, the processor: In the identifying step, the outlier is identified using one of a plurality of identification methods for the outlier; In the generating step, an instruction is given to generate the sentence based on input data; The input data is numerical data that is the basis of the outlier or numerical data related to the outlier, and is data of a type associated with the identification method used to identify the outlier. Information processing system.

10. 7. The information processing system according to claim 6, the processor: In the generating step, the artificial intelligence module is instructed to generate the sentence within a predetermined number of characters. Information processing system.

11. 11. The information processing system according to claim 10, the processor: In the determination step, the importance of the identified outlier is determined; In the generating step, the instruction is given by increasing the predetermined number of characters as the importance determined for the outlier increases. Information processing system.

12. 11. The information processing system according to claim 10, the processor: In the generating step, the instruction is given with a number corresponding to the degree to which the outlier deviates from the group of numerical data as the predetermined number of characters. Information processing system.

13. 11. The information processing system according to claim 10, the processor: In the identifying step, the outlier is identified using one of a plurality of identification methods for the outlier; In the generating step, the instruction is given with a number according to an identification method used to identify the outlier as the predetermined number of characters. Information processing system.

14. 2. The information processing system according to claim 1, the processor: In the determination step, the importance of the identified outlier is determined; In the display control step, the sentences generated for the identified outliers whose determined importance falls within a predetermined range are simultaneously displayed. Information processing system.

15. 2. The information processing system according to claim 1, the processor: In the display control step, when an operation for redisplaying the displayed sentence is performed, a sentence separately generated about the outlier explained by the sentence is displayed instead of the sentence. Information processing system.

16. An information processing method, comprising: The processor of the information processing system Executing each step of the information processing system according to any one of claims 1 to 15. Information processing methods.

17. A program, A computer is caused to execute each step of the information processing system according to any one of claims 1 to 15. program.

Citation Information

Patent Citations

  • Data analysis apparatus and method

    JP2006252259A