Device for ensuring reliability of service by data science, method for ensuring reliability of service by data science, and computer program for causing computer to execute said method
The apparatus and method ensure reliability in data science by generating self-description information units and calculating reliability features, addressing the lack of comprehensive assurance in existing technologies and maintaining stakeholder agreements.
Patent Information
- Application Number
- PCT/JP2025/003654
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-06
- Filing Date
- 2025-02-04
- Publication Date
- 2025-08-14
AI Technical Summary
Existing data science technologies lack reliability assurance, particularly in ensuring the ethical and transparent use of AI systems, leading to misunderstandings among stakeholders and challenges in maintaining agreement states due to changes in informed consent or regulations, with existing tools only providing partial reliability from a single entity's perspective.
An apparatus and method for ensuring reliability using data science, which includes generating self-description information units that organize accountability structures among stakeholders, calculating reliability features, and visualizing predictions to maintain agreement states and respond to deviations.
The solution provides a comprehensive reliability assurance mechanism that predicts and visualizes reliability changes, enabling stakeholders to confirm the basis for their confidence in data science solutions, thus stabilizing agreements and addressing the limitations of existing tools.
Smart Images

Figure JP2025003654_14082025_PF_FP_ABST
Abstract
Description
Apparatus for ensuring reliability of services using data science, method for ensuring reliability of services using data science, and computer program for causing a computer to execute the method
[0001] The present invention relates to an apparatus for ensuring the reliability of a service using data science, a method for ensuring the reliability of a service using data science, and a computer program for causing a computer to execute the method.
[0002] The IEEE Standards Association (SA) (NPL 1) has developed the 7000 series of standards, and IEEE 7000-2021 in particular defines a series of lifecycle processes for organizations to appropriately implement ethical values related to target systems. Non-Patent Document 2 describes explainable AI technology, which defines AI fairness, accountability, and transparency as social principles required of AI. Non-Patent Document 3, the IEC 62853 standard, redefines overall reliability as open systems dependability, which is reliability based on open system characteristics. Non-Patent Document 4 describes a method for supporting the automation of consensus-building process groups in the IEC 62853 standard. Non-Patent Document 5 describes a method for improving the quality of consensus building by analyzing the accountability structure when an organization implements the accountability execution process group in the IEC 62853 standard. Non-Patent Document 6 actually elucidates the accountability structure for a mobile app designed to improve the quality of life for patients with atopic dermatitis. Understanding this structure affects the quality of consensus building. Non-Patent Document 7 states that Amazon Web Services, Inc. (hereinafter referred to as AWS) provides a machine learning governance tool for its SageMaker service. This tool can be used to fulfill accountability. Patent Document 1 states that "the fault response cycle execution device and the change response cycle execution device work together to detect the occurrence of a fault or a sign of a fault in the target system in the fault response cycle, and if the shutdown of the target system is unavoidable, the dependability description data is changed in the change response cycle" (paragraph 0037). [Prior Art Documents] [Non-Patent Documents] [Non-Patent Document 1] IEEE SA - Systems and Software Engineering Standards Committee, IEEE Std 7000 TM-2021- IEEE Standard Model Process for Addressing Ethical Concerns during System Design, approved June 16, 2021. (ISBN 978-1-5044-7687-4) [Non-patent document 2] Naoki Otsubo et al., "XAI (Explainable AI) - What was Artificial Intelligence Thinking at the Time?" Rick Telecom Co., Ltd. 2021. (ISBN 978-4-86594-292-7) [Non-patent document 3] International Electrotechnical Commission, IEC 62853:2018 - Open systems dependability, 2018. (ISBN 978-2-8322-5789-0) [Non-patent document 4] Yanagisawa Y, Yokote Y. A New Approach to Better Consensus Building and Agreement Implementation for Trustworthy AI Systems. International Conference on Computer Safety, Reliability, and Security, 2021. https: / / doi.org / 10.1007 / 978-3-030-83906-2_26 [Non-patent Document 5] Yanagisawa Y, Yokote Y. An Accountability Approach to Resolve Multi-stakeholder Conflicts. International Conference on Computer Safety, Reliability, and Security, 2021. https: / / doi .org / 10.1007 / 978-3-030-83906-2_13 [Non-patent Document 6] Yanagisawa Y, Ashizaki K, Fujie Y, Toriumi Y, Ayano N, Tanese K, Kawasaki H, Amagai M, Yokote Y.Implementation of effective self-care services for atopic dermatitis by structural analysis of accountability. Medical Informatics 42 (Supplement), 881, 2022. (in Japanese) [Non-patent document 7] AWS "Machine learning governance with Amazon SageMaker" https: / / aws.amazon.com / jp / sagemaker / ml-governance / [Patent documents] [Patent document 1] Japanese Patent No. 5280587. General disclosure
[0003] A first aspect of the present invention provides an apparatus for ensuring the reliability of a service using data science, comprising: an acquisition unit that acquires, for each of a plurality of deviations that are considered to occur in the service, basis information indicating a plurality of basis for resolving each deviation that has been agreed upon among a plurality of stakeholders that provide and receive the service; a generation unit that generates a self-description information unit describing each deviation by determining a plurality of first relationships among the plurality of basis and associating the plurality of basis with each other based on the plurality of first relationships; and a calculation unit that determines a plurality of second relationships among the self-description information units generated for the plurality of deviations, calculates feature quantities for each of the plurality of self-description information units, and calculates reliability feature quantities that represent the reliability of the service by associating the feature quantities with each other based on the plurality of second relationships.
[0004] In the above-mentioned device, the generation unit may generate the self-descriptive information unit including: an accountability structure, which is a data structure that organizes, in accordance with the plurality of first relationships, a plurality of grounds, among the plurality of grounds, regarding the accountability that is imposed on each of the plurality of stakeholders to be fulfilled in order to resolve each of the deviations or the accountability that is to be fulfilled in order to confirm the validity of the accountability that is fulfilled; and an extended structure, which is a data structure that organizes, in accordance with the plurality of first relationships, a plurality of grounds, among the plurality of grounds, grounds that become a top goal that is a measure to be taken in case the result of fulfilling the accountability is inappropriate, grounds that become a strategy for taking the measure, grounds that become a plurality of sub-goals obtained by decomposing the top goal in accordance with the strategy, and grounds that become evidence supporting each of the plurality of sub-goals.
[0005] In any of the above devices, the generation unit may generate the self-descriptive information unit further including an argument structure, which is a data structure that organizes, in accordance with the plurality of first relationships, a plurality of grounds that, when there is a ground among the plurality of grounds for each deviation that is not agreed upon among the plurality of stakeholders, indicates the process of discussion until all of the plurality of stakeholders reach an agreement, or the process of argument until some of the plurality of stakeholders reach an agreement and the rest entrust the judgment of the few of them.
[0006] In any of the above devices, the generation unit may generate the self-describing information unit further including a trail structure, which is a data structure that organizes, for each deviation, a plurality of grounds that serve as a plurality of trails showing the legitimacy of the top goal, by subdividing the evidence that supports each of the plurality of subgoals among the plurality of grounds, according to the plurality of first relationships.
[0007] In any of the above devices, the calculation unit may calculate, for each deviation, an embedded feature from text data describing each of the plurality of grounds constituting the data structure for each of the plurality of data structures that are included in the self-describing information unit and have relationships with each other, and perform a convolution calculation to convolve the data structure from the calculated embedded features while maintaining the plurality of first relationships between the plurality of grounds in the data structure, thereby calculating a convolutional feature as the feature of the self-describing information unit.
[0008] Any of the above devices may further include a detection unit that detects a change in a distance from the reliability feature at a first time to the reliability feature at a second time after the first time by analyzing time-series data of the reliability feature, and thereby detects a deviation from the agreement state of the plurality of grounds among the plurality of stakeholders at the second time. Any of the above devices may further include an extraction unit that, when the detection unit detects a deviation from the agreement state, extracts a self-description information unit related to the deviation from the agreement state from the plurality of self-description information units at the second time.
[0009] In any of the above devices, the generation unit may generate the self-description information unit including grounds for indicating an expected deviation that assumes a case in which the result of the accountability imposed on each of the multiple stakeholders to resolve the deviation or to confirm the appropriateness of the accountability that has been exercised is inappropriate, and grounds for indicating measures to address the expected deviation. In any of the above devices, the extraction unit may, when the detection unit detects that a deviation has occurred in the agreement state, determine whether the deviation content of the agreement state corresponds to any of the multiple expected deviations, and, when an expected deviation corresponding to the deviation content exists, determine whether the measure to address the expected deviation is executable, and, when the measure to address the expected deviation is not executable, determine whether the deviation of the agreement state is identifiable, and, if identifiable, extract a self-description information unit related to the deviation of the agreement state from the multiple self-description information units at the second time point.
[0010] Any of the above devices may further include a prediction unit that predicts, by analyzing time-series data of the reliability features, that a mutation will occur in at least one of the plurality of grounds for resolving each of the deviations agreed upon among the plurality of stakeholders.
[0011] Any of the above devices may further include a display unit that visualizes and displays the self-descriptive information unit related to the basis for the prediction by the prediction unit that the mutation will occur.
[0012] In any of the above devices, the prediction unit may notify, via the display unit, the plurality of stakeholders of response actions required to resolve the deviation related to the basis for which the mutation is predicted to occur.
[0013] In any of the above devices, the acquisition unit may acquire text data of the plurality of grounds described by the plurality of interested parties via a user interface.
[0014] Any of the above devices may further include a shared control unit that controls the acquisition unit to cross boundaries of accountability of the multiple stakeholders to acquire reliability features of services other than the service, and controls the calculation unit to apply the reliability features of the other services to the calculation of the reliability feature of the service.
[0015] A second aspect of the present invention provides a method for ensuring the reliability of a service using data science, comprising: acquiring, for each of a plurality of deviations that are considered to occur in the service, basis information indicating a plurality of basis for resolving each deviation that has been agreed upon among a plurality of stakeholders providing and receiving the service; determining a plurality of first relationships among the plurality of basis for resolving each deviation, and generating a self-description information unit describing each deviation by relating the plurality of basis based on the plurality of first relationships; determining a plurality of second relationships among the self-description information units generated for the plurality of deviations, calculating feature quantities for each of the plurality of self-description information units, and relating the feature quantities to each other based on the plurality of second relationships, thereby calculating reliability feature quantities representing the reliability of the service.
[0016] In a second aspect of the present invention, there is provided a computer program for causing a computer to execute a method for ensuring the reliability of a service using data science. When executed by the computer, the computer program causes the computer to perform the following steps: acquire, for each of a plurality of deviations that are considered to occur in the service, basis information indicating a plurality of basis points for resolving each deviation that have been agreed upon among a plurality of stakeholders who provide and receive the service; generate a self-description information unit describing each of the deviations by determining a plurality of first relationships among the plurality of basis points and relating the plurality of basis points to one another based on the plurality of first relationships; and calculate a reliability feature that represents the reliability of the service by determining a plurality of second relationships among the self-description information units generated for the plurality of deviations, calculating a feature value for each of the plurality of self-description information units, and relating the feature values to one another based on the plurality of second relationships.
[0017] The above summary of the invention does not list all of the features of the present invention, and subcombinations of these features may also be inventions.
[0018] 1 is a functional block diagram illustrating an example configuration of an apparatus 100 for ensuring the reliability of services using data science, according to one embodiment;
[0034] FIG. 1 illustrates an example hardware configuration of the apparatus 100, according to one embodiment;
[0035] FIG. 1 illustrates an example configuration of structured data, according to one embodiment;
[0036] FIG. 1 illustrates an example of details of a self-information recording unit 300, according to one embodiment;
[0037] FIG. 1 illustrates an example of details of an accountability structure 310-01, according to one embodiment;
[0038] FIG. 1 illustrates an example of details of a GSN extension structure 310-02, according to one embodiment;
[0039] FIG. 1 illustrates an example of details of an argument structure 310-03, according to one embodiment;
[0039] FIG. 1 illustrates an example of details of a trail structure, according to one embodiment;
[0039] FIG. 1 illustrates an example of a procedure for collecting a self-description information unit 300 by the apparatus 100, according to one embodiment;
[0039] FIG. 1 illustrates an example of a procedure for analyzing an accountability structure 400-01 by the apparatus 100, according to one embodiment;
[0039] FIG. 1 illustrates an example of a procedure for describing a GSN extension structure 400-04 by the apparatus 100, according to one embodiment;
[0039] FIG. 1 illustrates an example of a procedure for verifying by argument 400-03 by the apparatus 100, according to one embodiment;
[0039] FIG. 1 illustrates an example of a procedure for reaching a consensus 400-02 by the apparatus 100, according to one embodiment. 4A illustrates an example procedure for extracting 400-05 evidence by the device 100, according to one embodiment. FIG. 4B illustrates an example procedure for determining 400-06 whether a deviation has occurred by the device 100, according to one embodiment. FIG. 4C illustrates an example procedure for extracting 400-07 self-describing information units related to a deviation by the device 100, according to one embodiment. FIG. 4D illustrates an example configuration 500 of the UI unit 100-06, according to one embodiment. FIG. 4E illustrates a diagram of an example conversation scenario 510 presented to a user by the UI unit 100-06, according to one embodiment. FIG. 4F illustrates a diagram of an example conversation scenario 510 presented to a user by the UI unit 100-06, according to one embodiment. FIG. 4F illustrates a diagram of an example conversation scenario 510 presented to a user by the UI unit 100-06, according to one embodiment. FIG. 4B illustrates a diagram of an example conversation scenario 520 related to the verification by argument 400-03 described in FIG. 4A. 1 illustrates an example structure 600 of a structured data record 100-02, according to one embodiment; 2 illustrates an example data SLA 700 in the form of a GSN extension structure 310-02, according to one embodiment; 3 illustrates an example data SLA 710 in the form of an accountability structure 310-01, according to one embodiment.1 shows an example configuration 800 of the evidence management unit 100-03 according to one embodiment.
[0044] FIG. 1 shows an example configuration 900 of the feature calculation unit 100-04 according to one embodiment.
[0045] FIG. 1 shows a conceptual diagram of an example calculation 910 of features related to a self-description information unit 300 by the device 100 according to one embodiment.
[0046] FIG. 1 shows an example procedure 920 for calculating features of a minutes starting point by the device 100 according to one embodiment.
[0047] FIG. 1 shows an example configuration 1000 of the agreement reliability prediction unit 100-05 according to one embodiment.
[0048] FIG. 1 shows an example processing procedure 1010 of the self-description information unit inter-relationship prediction unit 1000-03 according to one embodiment.
[0049] FIG. 1 shows an example processing procedure 1020 of the consensus building depth prediction unit 1000-04 according to one embodiment.
[0049] FIG. 1 shows an example processing procedure 1030 of the evidence structure variation prediction unit 1000-05 according to one embodiment.
[0049] FIG. 1 shows an example processing procedure 1040 of the argument structure variation prediction unit 1000-06 according to one embodiment. 1 illustrates an example processing procedure 1050 for a contribution calculation unit 1000-07 according to one embodiment; 2 illustrates an example processing procedure 1060 for a feature-based structured data search unit 1000-09 according to one embodiment; 3 illustrates an example configuration 1100 for a feature sharing control unit 100-07 according to one embodiment; 4 illustrates a screen snapshot 1200 as an example of a dashboard according to one embodiment; and 5 illustrates an example computer 2200 in which aspects of the present invention may be embodied in whole or in part.
[0019] The present invention will be described below through embodiments of the invention, but the following embodiments do not limit the scope of the invention as claimed. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.
[0020] 1 is a functional block diagram showing an outline of an example configuration of an apparatus 100 for ensuring the reliability of a service using data science according to one embodiment. The apparatus 100 of this embodiment makes it possible to guarantee the reliability of a problem solution derived using data science, particularly the reliability of predictions, among stakeholders involved in the problem solution. The apparatus 100 of this embodiment further makes it possible to visualize prediction results related to the reliability of the problem solution, allowing stakeholders to confirm them.
[0021] The device 100 of this embodiment is composed of seven functional blocks: a structured data collection unit 100-01, a structured data recording unit 100-02, a trail management unit 100-03, a feature calculation unit 100-04, an agreement reliability prediction unit 100-05, a UI unit 100-06, and a feature sharing control unit 100-07. Note that the functional blocks are not necessarily limited to these, but a brief description of these representative functional blocks will be given below.
[0022] The structured data collection unit 100-01 collects structured data generated when solving problems using data science, i.e., when resolving multiple deviations that are thought to occur in a data science service. The deviations referred to here may refer to, for example, deviations from a state of agreement between multiple stakeholders regarding problem solving using data science, i.e., a state of agreement between an entity that provides a data science service and an entity that uses the service.
[0023] In this embodiment, the structured data collection unit 100-01 identifies multiple stakeholders, i.e., entities that provide data science-based services and entities that use the services. The structured data collection unit 100-01 provides the identified stakeholders with a means for describing their trustworthiness in problem solving using data science. That is, the structured data collection unit 100-01 provides a means for entities that use data science-based services to describe their trustworthiness requirements for the services from the perspective of accountability, and a means for entities that provide the services to describe their trustworthiness characteristics in response to the trustworthiness requirements from the perspective of accountability.
[0024] The structured data is an example of a self-describing information unit that describes each of multiple deviations that are thought to occur in a data science-based service. The self-describing information unit is composed of multiple rationales for resolving each deviation, agreed upon among multiple stakeholders who provide and receive the service. The structured data collection unit 100-01 is an example of an acquisition unit that acquires rationale information indicating the multiple rationales.
[0025] The structured data collection unit 100-01 also provides a means for extracting a self-description information unit describing a specific deviation from the self-description information units described by the stakeholders for each deviation. The structured data collection unit 100-01 is also an example of an extraction unit that, when a deviation is detected in the agreement state of multiple grounds between multiple stakeholders, extracts a self-description information unit related to the deviation from the agreement state from the multiple self-description information units at the time of detection. The extraction of a self-description information unit related to the deviation from the agreement state will be described in detail from Figure 4H onwards.
[0026] The structured data collection unit 100-01 is also an example of a generation unit that determines multiple first relationships between the multiple grounds described above and generates self-descriptive information units that describe each deviation by associating the multiple grounds with each other based on the multiple first relationships. The generation of self-descriptive information units will be described in detail from FIG. 3C onwards. In the following description, terms such as "relationship," "correlation," and "link" may be used instead of the term "relationship." In this specification, the term "relationship" may mean "a mutual relationship in which a change in one causes a change in the other."
[0027] The structured data recording unit 100-02 records and persists the self-describing information units. The structured data recording unit 100-02 may use an SQL database, a NoSQL database, or a file system, and any persistence means may be used as long as the self-describing information units can be referenced on any time axis after they are extracted.
[0028] The trail management unit 100-03 manages the trails recorded in the structured data recording unit 100-02, for example, the trails of the GSN extended structure 310-02 shown in Figure 3D (element evidence 330-05 and the group of elements shown in Figure 3F). In this embodiment, the trail management unit 100-03 may be configured to manage information that needs to be managed as evidence, regardless of these.
[0029] The feature calculation unit 100-04 calculates reliability features representing the reliability of the above-mentioned service from the relationships between the self-description information units recorded in the structured data recording unit 100-02 and records the results. In this embodiment, the feature calculation unit 100-04 may set the relationships at the time of extracting the self-description information units. The feature calculation unit 100-04 may calculate the embedding features of the self-description information units and set the distances between them, or alternatively, may calculate the reliability features using other means. The feature calculation unit 100-04 is an example of a calculation unit that determines multiple second relationships between multiple self-description information units generated for the above-mentioned multiple deviations, calculates each feature of the multiple self-description information units, and associates the multiple feature values with each other based on the multiple second relationships to calculate the reliability features. The calculation of the reliability features by associating the features of multiple self-description information units with each other will be described in detail from FIG. 9B onwards.
[0030] The agreement reliability prediction unit 100-05 acquires a machine learning model that predicts variations in reliability from the reliability features calculated by the feature calculation unit 100-04. The agreement reliability prediction unit 100-05 predicts variations in the agreement state between the stakeholders by analyzing time-series data of the reliability features using the machine learning model. In this embodiment, the agreement reliability prediction unit 100-05 records the agreement formed between the stakeholders for one self-description information unit. The agreement reliability prediction unit 100-05 also adds a description that anticipates the possibility of deviation from the agreement and a description of preventive measures against such deviation. The agreement reliability prediction unit 100-05 further adds the record of the agreement to the self-description information unit using an information representation based on the argumentation structure between the stakeholders. The agreement reliability prediction unit 100-05 is an example of the generation unit described above.
[0031] The agreement reliability prediction unit 100-05 may predict mutations in the agreement state between stakeholders by analyzing time-series data of the argument structure to predict mutations in the argument structure, or may predict mutations in the agreement state between stakeholders using other methods. The agreement reliability prediction unit 100-05 is also an example of a prediction unit that predicts mutations in at least one of multiple rationales for resolving deviations agreed upon among multiple stakeholders by analyzing time-series data of reliability features. Predicting mutations in rationales will be described in detail from FIG. 9C onwards. The prediction unit may notify, via a display unit (described later), of response actions by multiple stakeholders required to resolve deviations in the rationales for which the mutations are predicted to occur.
[0032] The UI unit 100-06 serves as a user interface for the device 100. In this embodiment, the UI unit 100-06 may have functions such as accepting input from identified stakeholders, notifying them of necessary response actions, displaying and notifying them of the prediction results of the agreement reliability prediction unit 100-05, and visualizing the prediction process in the agreement reliability prediction unit 100-05, and may also have other functions. The UI unit 100-06 is an example of a display unit that visualizes and displays self-descriptive information units related to the basis for which the prediction unit above predicts that a mutation will occur.
[0033] The feature sharing control unit 100-07 provides a function for sharing reliability features calculated by the feature calculation unit 100-04 across accountability boundaries. Crossing accountability boundaries here refers to, for example, when a system developed and operated among stakeholders belonging to one organization involves another system built by another organization, acquiring reliability features of the other system to clarify accountability for the other system as well. The feature sharing control unit 100-07 improves prediction accuracy by using reliability features from other cases as samples for machine learning input. The feature sharing control unit 100-07 is an example of a sharing control unit that controls the acquisition unit to acquire reliability features of services other than the above-mentioned service across accountability boundaries between multiple stakeholders and controls the calculation unit to apply the reliability features of the other services to the calculation of the reliability features of the above-mentioned service.
[0034] The above has described an outline of an example configuration of the device 100 with reference to the functional block diagram in Fig. 1. The multiple functional blocks shown in Fig. 1 may be integrated into the device 100 configured with a single computer, or alternatively, may be distributed among multiple computers that configure the device 100, for example, with a computer assigned to each functional block. In this case, each computer may have some kind of communication means to enable mutual communication.
[0035] FIG. 2 shows an example hardware configuration of device 100 according to one embodiment. In its most basic configuration, device 100 is an electronic computer including an arithmetic unit 200-01, a control unit 200-02, a memory unit 200-03, and an input / output unit 200-04, all connected by an instruction bus and a data bus. In device 100 shown in FIG. 2, arithmetic unit 200-01 executes arithmetic operations, logical operations, comparison operations, shift operations, and the like, based on bit data information input from input / output unit 200-04. The executed data is stored in memory unit 200-03 and output from input / output unit 200-04 as necessary. This series of processes is controlled by control unit 200-02 in accordance with a software program stored in memory unit 200-03. Device 100 according to this embodiment is hardware equipped with the basic computer functions described above, and is controlled by a group of programs, such as an operating system, device drivers, middleware, and application software.
[0036] Recently, with the rapid advancement of AI (Artificial Intelligence) technology used in data science, due to technical and time constraints, there has been insufficient consideration given to the reliability of the insights and findings derived from data science for problem-solving. Typical examples of this are the spread of fake news and the widening gap caused by binary oppositions.
[0037] The IEEE Standards Association (IEEE SA) is developing the 7000 series of standards, and IEEE 7000-2021 in particular defines a series of lifecycle processes for organizations to properly implement ethical values for target systems. This standard defines the interaction between the concept exploration stage and the development stage using four processes based on the transparency management process: the concept of operations and context exploration process, the ethical values elicitation and prioritization process, and the ethical requirements definition process.
[0038] Since this standard indicates its compatibility with the ISO / IEC / IEEE 15288 standard in Annex A, it is also possible to align it with the IEC 62853 standard. As a result, implementing data science processes in accordance with IEC 62853 enables the appropriate implementation of ethical values based on IEEE 7000, which is particularly effective in implementing data science in the field of medical science.
[0039] Data science is considered a comprehensive academic field that encompasses multiple fields, including statistics, computer science, and data visualization. Statistics can explain "why" an output occurs in response to input based on mathematics. However, deep learning technology, which has emerged as a result of advances in machine learning technology and computers being able to process big data in a realistic amount of time, is a black box, and therefore cannot explain "why" an output occurs in response to input.
[0040] Therefore, explainable AI technology is being developed as a social principle requiring AI to meet the three items of fairness, accountability, and transparency approved by the G20 in 2019. This involves the development of technology that can identify the dominant tendencies of a target AI and understand the reasons behind its judgments regarding its prediction results.
[0041] For example, LIME is a local explanation technology that calculates the features that contributed to AI predictions based on input data. SHAP calculates the Shapley value in game theory to determine the player's contribution, and uses this to calculate the contribution of features to the input data. Other technologies that have been developed include Permutation Importance, Partial Dependence Plot, and surrogate models. Because these technologies require a certain amount of computation, there is a conflict between performance requirements and explainability requirements, and these conflicts must be resolved. Furthermore, XAI technology does not go as far as to assess the reliability of the explanation content.
[0042] The implementation of an operating system for data science, especially AI, called AI-OS, is desirable in that it would standardize and supervise all controls related to the target system, thereby bridging the gap with user expectations and providing users with reliability.
[0043] In the description of this embodiment, dependability is synonymous with dependability. The IEC 62853 standard redefines dependability as open systems dependability, which is dependability based on open system characteristics. Today's data science has open system characteristics in terms of the rapid change in technology used, the need for ethical considerations, and the involvement of complex stakeholders. Therefore, as stipulated in the IEC 62853 standard, data science execution should be managed as a lifecycle process, which is a process of processes of consensus building and accountability fulfillment.
[0044] The IEC 62853 standard relies on the ISO / IEC / IEEE 15288 standard, and therefore requires modifications for implementation in data science. By combining the above-mentioned non-patent documents 4 to 6, it is possible to solve problems in data science using a lifecycle process that conforms to the IEC 62853 standard. However, in terms of reliability in problem solving, reliability is still treated as a non-functional requirement, leading to many misunderstandings among stakeholders, and stakeholders are unaware of changes in their own or other stakeholders' reliability requirements. Furthermore, while machine learning models are often reused between organizations in data science, the issue of reliability, particularly reproducibility, remains unresolved.
[0045] AWS provides machine learning governance tools in its SageMaker service. These tools include user role management when executing machine learning, access rights policy control for required resources, visualization of evaluation results and error analysis functions, and information management for model sharing. The machine learning workflow on SageMaker is recorded as a model lineage graph, allowing model training steps to be reproduced. The Model Cards function provides a repository of model information necessary for reusing models within and between organizations, and allows the training environment, parameters, results, and other information used to train the model to be referenced when reusing the model. Hugging Face (https: / / huggingface.co / ), a platform for developing, sharing, and publishing machine learning models, also offers a similar Model Cards function.
[0046] By modifying and implementing the process defined in the IEC 62853 standard in SageMaker, it is possible to leave evidence that guarantees a certain degree of reliability for the data science performed. However, this evidence is from the perspective of a single entity, the model developer, and does not guarantee the reliability of the solution method for all stakeholders attempting to solve some problem using data science. Furthermore, when the partiality of a dataset, such as when personal information is included in the data, changes frequently, or when information not used as training data in machine learning changes frequently, it is not possible to respond to changes in the reliability of the evidence.
[0047] Furthermore, agreements on reliability formed among stakeholders involved in problem-solving, particularly in the medical science field, are difficult to maintain stably due to factors such as the revocation of informed consent (IC) or changes in regulations, making it difficult to respond to changes in the state of the agreement. The dependability maintenance system described in Patent Document 1 above makes it possible to respond to such changes to some extent, but the request for changing from a fault response cycle to a change response cycle in this patent cannot accommodate crossing of responsibility boundaries. Furthermore, in all of the above cases, reliability of data science cannot be guaranteed based on predictions.
[0048] In consideration of the above-mentioned problems, an apparatus 100 for assuring the reliability of a data science-based service according to this embodiment assures the reliability of data science based on predictions. In this specification, "assurance" may refer to stakeholders being able to present the basis for their confidence in the reliability of a problem solution. To this end, the apparatus 100 according to this embodiment may include a means for recording structured data that organizes the basis for confidence in the relationship of accountability between stakeholders, such as a structured data recording unit 100-02. The apparatus 100 according to this embodiment may also include a means for calculating, from the structured data, reliability features that enable predicting changes in the state of agreement reached between the stakeholders regarding the legitimacy of the basis, such as a feature calculation unit 100-04. The apparatus 100 according to this embodiment may further include a means for visualizing the prediction results calculated from the features so that stakeholders can confirm them, such as a UI unit 100-06.
[0049] [Embodiment of Structured Data] The JIS Q 27000 standard, "Information Technology - Security Techniques - Information Security Management Systems - Terminology," defines stakeholders in Section 3.37 as "individuals or organizations that can influence, be affected by, or perceive themselves to be affected by a decision or activity." Based on this definition, in this embodiment, stakeholders are defined as "all interested parties, including organizations and project members, that have an interest in establishing the reliability of problem-solving using data science." Identifying stakeholders means recognizing which people and organizations fit this definition in a project using data science to solve problems. Here, stakeholders include society at large. While this term is synonymous with the public or the general public, it is difficult to extract the information described below from such stakeholders, so personas and role models will be created as necessary.
[0050] FIG. 3A illustrates an example of the structure of structured data according to one embodiment. The structured data illustrated in FIG. 3A is composed of one or more self-description information units. One self-description information unit 300 is connected to multiple other self-description information units via labeled links. The self-connection of a link labeled "label" to a self-description information unit 300 in the figure illustrates a one-to-many connection between one instance of the self-description information unit 300 and one or more other instances of the self-description information unit 300. In FIG. 3A, an * symbol indicates a one-to-many connection, and this also applies to subsequent figures, so redundant explanations will be omitted. Note that in the following explanation, connecting multiple reliability features, multiple self-description information units, multiple data structures, multiple components, etc. to each other via "links" with specific labels may mean associating multiple features, multiple evidence, etc. with each other based on the relationship represented by a specific label.
[0051] In this embodiment, each component of the self-describing information unit 300 is assigned a unique identifier and is also assigned attributes "created" and "updated." The attribute "created" represents the timestamp when the element was created, and the attribute "updated" represents the timestamp when the element was updated. In this embodiment, the label is given the name "accountability_chain," which will be described later, but other names and meanings may also be given.
[0052] 3B shows an example of details of a self-information description unit 300 according to one embodiment. The self-information description unit 300 of this embodiment includes, for example, four components: an accountability structure 310-01, a GSN extension structure 310-02, an argument structure 310-03, and a trail structure 310-04. These components are examples of rationales for resolving each of a number of deviations that may occur in a data science service, agreed upon among multiple stakeholders who provide and receive the service. The rationale information described above may include one or more of these components.
[0053] Accountability structure 310-01 has a one-to-many connection with GSN extension structure 310-02 via a link with label dcase, and a one-to-many connection with argumentation structure 310-03 via a link with label discuss. GSN extension structure 310-02 has a one-to-many connection with accountability structure 310-01 via a link with label amap, a one-to-many connection with argumentation structure 310-03 via a link with label discuss, and a one-to-many connection with trail structure 310-04 via a link with label trail.
[0054] FIG. 3C illustrates an example of details of the accountability structure 310-01 according to one embodiment. The accountability structure 310-01 illustrated in FIG. 3C is, for example, composed of 12 elements, as described below, but other configurations are also possible. These elements are examples of rationales for resolving each of the multiple deviations that are considered to occur in a data science service, agreed upon between multiple stakeholders providing and receiving the service. The rationale information described above may include one or more of these elements. Similarly, in the following description, the multiple elements constituting each data structure, such as the GSN extension structure 310-02, the argument structure 310-03, and the evidence structure 310-04, are examples of rationales for resolving each deviation agreed upon between multiple stakeholders, and the rationale information described above may include one or more of these elements.
[0055] The device 100, for example, using the structured data collection unit 100-01, the structured data recording unit 100-02, the agreement reliability prediction unit 100-05, etc., defines the two or more identified stakeholders as having a one-to-one relationship of "who" and "to whom," and records information in each element from the perspective of "for what" purpose each stakeholder is to be held accountable. The device 100 records the purpose in the element "purpose" 320-01, records "who" in the element "explainer" 320-02, and records "to whom" in the element "explainee" 320-03.
[0056] In this embodiment, the device 100 allows aliases for records relating to "who," and records them in elements "alias" 320-09 and "alias" 320-11, and records representative names in elements "user" 320-10 and "user" 320-12.
[0057] In the accountability structure 310-01 shown in Figure 3C, as an example, the element explainer 320-02 is connected one-to-many to the element alias 320-09 by a link with the label alias, and the element explainer 320-02 is connected one-to-many to the element user 320-10 by a link with the label explainer. The element alias 320-09 is connected one-to-one to the element user 320-10 by a link with the label user. The element explainee 320-03 is connected one-to-many to the element alias 320-11 by a link with the label alias, and the element explainee 320-03 is connected one-to-many to the element user 320-12 by a link with the label explainee. The element alias 320-11 is connected one-to-many to the element user 320-12 by a link with the label user.
[0058] In this embodiment, the stakeholder corresponding to the above "who" may be referred to as an accountability provider, and the stakeholder corresponding to the above "to whom" may be referred to as an accountability beneficiary. The device 100 records the content of accountability that the accountability beneficiary requests from the accountability provider in the element "requirement" 320-04. The accountability beneficiary may record a request for what kind of accountability it requests the accountability provider to fulfill in the element "requirement" 320-04.
[0059] The device 100 records a means by which the accountability provider fulfills accountability for the request in the element explainer_action 320-05. The device 100 may record a means by which the accountability provider fulfills accountability upon receiving the request in the element explainer_action 320-05.
[0060] The device 100 also records a means for verifying the validity of the means asserted by the accountability beneficiary in the element explainee_action 320-06. The means for verifying the validity of the assertion of the accountability provider may be recorded in the element explainee_action 320-06.
[0061] 3C, for example, the accountability beneficiary element "explainee" 320-03 is connected one-to-many with a link labeled "why" to the element "requirement" 320-04, and the element "requirement" 320-04 is connected many-to-one with a link labeled "understand" to the element "explainer" 320-02, the accountability provider. The element "explainer" 320-02 is also connected one-to-many with a link labeled "what," the element "explainer_action" 320-05 is also connected one-to-many with a link labeled "what," the element "explainer_action" 320-05 is also connected one-to-one with a link labeled "acknowledgement," and the element "explainee_action" 320-06 is also connected many-to-one with the element "explainee" 320-03, the accountability beneficiary, with a link labeled "understand."
[0062] The accountability structure 310-01 shown in Figure 3C is an example of an accountability structure, which is a data structure that organizes, according to the plurality of first relationships, the plurality of grounds, among the plurality of grounds, relating to the accountability imposed on each of the plurality of stakeholders to be fulfilled to resolve each deviation or to verify the validity of the fulfilled accountability. The accountability structure 310-01 is included in the self-describing information unit generated by the generation unit for each deviation. The first relationships may refer, for example, to the relationships represented by the labels of the links connecting the elements of the accountability structure 310-01 shown in Figure 3C. Recording information in each element of the accountability structure 310-01 as described above is an example of the generation unit generating a self-describing information unit including an accountability structure that organizes the plurality of grounds according to the plurality of first relationships.
[0063] In this embodiment, in addition to the above information, the device 100 anticipates and records situations in which the accountability beneficiary cannot confirm the appropriateness of the accountability provider's fulfillment of accountability, or in which the fulfillment of accountability is deemed inappropriate. Such situations are undesirable, as they indicate that the accountability beneficiary's demands on the accountability provider are not satisfied. IEC 62853 prescribes that such situations should be anticipated and preventive and response measures should be planned and recorded in advance.
[0064] Therefore, in this embodiment, the device 100 records the assumption of the situation in the assumed_deviation element 320-07 and records corresponding preventive and countermeasure measures in the preventive_action element 320-08. In the accountability structure 310-01 shown in FIG. 3C , for example, the assumed_deviation element 320-07 is connected one-to-many to the purpose element 320-01 via a link labeled response, and is connected one-to-one to the preventive_action element 320-08 via a link labeled correspondence. This is an example of the generation unit generating a self-descriptive information unit including a basis for indicating an assumed deviation and a basis for indicating measures to address the assumed deviation, which assumes a case in which the results of the accountability imposed on each of the multiple stakeholders for resolving each deviation or the accountability to be performed to confirm the appropriateness of the performed accountability are inappropriate.
[0065] FIG. 3D illustrates an example of the details of the GSN extension structure 310-02 according to one embodiment. The device 100 of this embodiment uses an extension of the Goal Structuring Notation standard (https: / / scsc.uk / gsn) as an example, but other notations may also be used. Here, as an example, one argument structure is visualized. An argument structure is a structure that expresses the validity of an argument and enables the legitimacy of a claim to be inferred from evidence. The accountability structure 310-01 breaks down the accountability structure into a one-to-one relationship between stakeholders, namely, accountability providers and accountability beneficiaries. The necessary consensus building between stakeholders is recorded, for example, in the GSN extension structure 310-02 shown in FIG. 3D.
[0066] The apparatus 100 records the claim that is the top goal in the claim element 330-01, and in order to record the evidence for the claim, breaks down the argument according to the strategy and records it in the subgoals. The apparatus 100 records the strategy in the strategy element 330-04. In this case, as an example, the apparatus 100 assigns a weight to the supportedBy link connecting the claim element 330-01 and the strategy element 330-04. The apparatus 100 may use the assigned weight as an argument during the convolution calculation described below. For example, in the case of a D-Case, the link connected to the strategy element is important, so the apparatus 100 may assign a stronger weight to the link.
[0067] The apparatus 100 records the stakeholder in the claim to clarify who made the claim. More specifically, the apparatus 100 records the stakeholder in a stakeholder element 330-02. The apparatus 100 can further record a premise in the claim and records the premise in a context element 330-03. In this case, the apparatus 100 assigns a weight to the link inContextOf that connects the claim element 330-01 and the context element 330-03, for example.
[0068] When the legitimacy of the top goal is decomposed into multiple subgoals according to the argument strategy and each of the subgoals is supported by sufficiently clear evidence, the device 100 records the evidence in the element evidence 330-05. Note that the element evidence 330-05 is connected to each of the three elements (element document 350-01, element log 350-02, and element minutes 350-03) shown in FIG. 3F (described later) by links labeled trail.
[0069] In the example of the GSN extension structure 310-02 shown in Figure 3D, these individual elements connected from the top goal are interconnected by links labeled inContextOf and supportedBy. The element claim 330-01 and the element stakeholder 330-02 are connected one-to-many by a link labeled inContextOf, and the element claim 330-01 and the element context 330-03 are connected one-to-many by a link labeled inContextOf. The element claim 330-01 and the element strategy 330-04 are connected one-to-many by a link labeled supportedBy, and the element claim 330-01 and the element evidence 330-05 are connected one-to-many by a link labeled supportedBy. Furthermore, the element strategy 330-04 and the element claim 330-01 are connected one-to-many by a link labeled supportedBy. As a record of the consensus building, the element claim 330-01 is connected one-to-many to the element user 330-08 of the participant in the consensus building via a link 330-09 labeled agree / disagree / abstain. The labels agree / disagree / abstain represent agreement / disagree / abstain in the discussion regarding the element claim 330-01.
[0070] The GSN extension structure 310-02 shown in Figure 3D is an example of an extension structure, which is a data structure that organizes, according to the above-mentioned multiple first relationships, the above-mentioned multiple grounds, including the grounds that constitute a top goal, which is a countermeasure to be taken if the results of the accountability are inappropriate, the grounds that constitute a strategy for implementing the countermeasure, the grounds that constitute multiple subgoals obtained by breaking down the top goal according to the strategy, and the grounds that constitute evidence supporting each of the multiple subgoals. This first relationship may refer to, for example, the relationship represented by the label of the link connecting each element of the GSN extension structure 310-02 shown in Figure 3D. Recording information in each element of the GSN extension structure 310-02 as described above is an example of the above-mentioned generation unit generating a self-describing information unit including an extension structure that organizes the above-mentioned multiple grounds according to multiple first relationships.
[0071] 3E illustrates an example of details of argument structure 310-03 according to one embodiment. The device 100 of this embodiment uses, as an example, the DCW (if Data, then C, since W) structure described in The Uses of Argument by Stephen E. Toulmin (ISBN: 978-0521534833), although other models may be used.
[0072] The device 100 records assertions in element claim 340-01, evidence in element data 340-02, and evidence in element warrant 340-03. In the example of the argument structure 310-03 shown in Figure 3E, the element claim 340-01 is connected one-to-many to the element data 340-02 with a link labeled support. The element data 340-02 is connected one-to-one to the element warrant 340-03 with a link labeled context.
[0073] The device 100 can modify each element, and when a modification is made, a new element is added. More specifically, when the device 100 modifies element claim 340-01, it records the modified version in element claim 340-04, and establishes a one-to-many connection between the two with a link labeled "update." Similarly, when the device 100 modifies element data 340-02, it records the modified version in element data 340-05, and establishes a one-to-many connection between the two with a link labeled "update." Similarly, when the device 100 modifies element warrant 340-03, it records the modified version in element warrant 340-06, and establishes a one-to-many connection between the two with a link labeled "update."
[0074] The argument structure 310-03 shown in Figure 3E is an example of an argument structure, which is a data structure that organizes, according to the above-mentioned multiple first relationships, multiple evidences that represent the process of discussion until all stakeholders reach agreement when there are evidences that are not agreed upon among the multiple stakeholders, or the process of argumentation until some stakeholders reach agreement and the rest entrust the judgment of those few. The argument structure 310-03 is included in the self-describing information unit generated by the above-mentioned generator for each deviation. The first relationships may refer, for example, to the relationships represented by the labels of the links connecting the elements of the argument structure 310-03 shown in Figure 3E. Recording information in each element of the argument structure 310-03 as described above is an example of the above-mentioned generator generating a self-describing information unit including an argument structure that organizes the above-mentioned multiple evidences according to the above-mentioned multiple first relationships.
[0075] 3F shows details of the evidence structure 310-04 according to one embodiment. In the example of the evidence structure 310-04 shown in FIG. 3F, the element evidence 330-05 is connected one-to-many to the element document 350-01 via a link labeled tail.
[0076] The device 100 identifies, in the element document 350-01, a document appropriate as evidence for the higher-level claim / goal (element claim 330-01) of the element evidence 330-05, and records a link (URL) to the document. The device 100 also records, in the element log 350-02, a file in which appropriate log records or audit logs identified as evidence for the higher-level claim / goal are recorded. The device 100 may record, in the element log 350-02, the file itself, or link information that allows access to the file. Note that, although log records or audit logs appropriate as evidence for the higher-level claim / goal are identified outside the device 100, they may alternatively be identified by the device 100.
[0077] For the evidence, interested parties (users / stakeholders) reach a consensus on whether to accept the content as evidence, and the device 100 records this by connecting the element evidence 330-05 to the element user / stakeholder 350-11. More specifically, if the interested parties agree, the device 100 connects them with a link 350-12 labeled agree. If the interested parties disagree, the device 100 connects them with a link 350-12 labeled disagree. If the interested parties abstain from the consensus, the device 100 connects them with a link 350-12 labeled abstain.
[0078] As an example, the device 100 of this embodiment considers minutes of meetings in project execution to be one of the important evidences, and records the minutes in a group of elements that are each connected one-to-many with links labeled "content" from the element "minutes" 350-03. More specifically, the device 100 records the date, start time (begin), and end time (end) of the meeting in the element "date_time" 350-04. The device 100 records the meeting topic / agenda in the element "agenda" 350-05.
[0079] The device 100 records the contents of the discussion in the element discussion 350-06. The device 100 records the decisions made at the meeting in the element decision 350-07. The device 100 records action items (tasks) in the element todo 350-08, and records the status of the task, such as not done, in progress, or completed, in the status column. The device 100 records future issues in the element issue 350-09.
[0080] The device 100 records the date of the next meeting in the element next 350-10. If the meeting has a previous or next relationship, the device 100 connects the elements minutes 350-03 with each other using a link 350-12 labeled previous. The participants of the meeting can confirm the contents of the minutes, and when the participants of the meeting confirm them, the device 100 connects the element minutes 350-03 to the participant element participant 350-13 in a one-to-many manner using a link labeled confirmed. The device 100 may record the minutes in a JSON format, or alternatively, may record them in any format such as Markdown format or Plain format.
[0081] The evidence structure 310-04 shown in Figure 3F is an example of a data structure that organizes, according to the first relationships, the evidence supporting each of the subgoals, which constitutes the evidence supporting the top goal. The evidence structure 310-04 is included in the self-describing information unit generated by the generator for each deviation. The first relationships may refer to, for example, the relationships represented by the labels of the links connecting the elements of the evidence structure 310-04 shown in Figure 3F. Recording information in each element of the evidence structure 310-04 as described above is an example of the generator generating a self-describing information unit that includes an evidence structure that organizes the evidence supporting each of the subgoals according to the first relationships.
[0082] 4A illustrates an example of a procedure for collecting self-describing information units 300 by device 100 according to one embodiment. For a new project, the procedure begins with accountability structure analysis 400-01. As a result of the analysis, accountability structure 310-01 information that constitutes the self-describing information unit 300 is extracted. In GSN extension structure description 400-04, the GSN extension structure 310-02 is used to promote the project while reaching consensus 400-02 on reliability requirements among stakeholders, and necessary evidence is collected in evidence extraction 400-05.
[0083] The evidence is used to reach a consensus 400-02 among stakeholders that the evidence is sufficient for accountability. In the consensus building 400-02, if there is doubt about the common understanding among stakeholders, a record is kept of the fact that the common understanding has been reached using the argument structure 310-03 in the confirmation by argument 400-03, as necessary.
[0084] At least one of the steps up to this point is an example of a step included in a method for ensuring the reliability of a service using data science, which step includes obtaining, for each of a plurality of deviations that are considered to occur in the service, evidence information indicating a plurality of evidences for resolving each deviation that are agreed upon among a plurality of stakeholders who provide and receive the service. At least one of the steps up to this point is also an example of a step included in the method, which step includes determining a plurality of first relationships among the plurality of evidences and generating a self-describing information unit that describes each deviation by relating the plurality of evidences to each other based on the plurality of first relationships.
[0085] If a deviation from the consensus state is found during the process of evidence extraction 400-05 or consensus building 400-02 (deviation occurred 400-6: YES), self-description information units related to the deviation are extracted 400-07, and accountability structure analysis 400-01 begins. Note that, although this embodiment assumes the project is promoted in the above-mentioned loop format, the project may be promoted in other formats.
[0086] FIG. 4B shows an example of the procedure for accountability structure analysis 400-01 by the device 100 according to one embodiment. Because accountability is primarily required when a failure occurs, the objective of fulfilling accountability for failures is determined. To achieve this, in failure extraction 410-01, failures that have occurred in the past and anticipated failures that may occur in the future are extracted, and the scope of their impact is identified. Note that failures may be anticipated and extracted, for example, when the environment changes or when laws and regulations (such as the Personal Information Protection Act) are revised.
[0087] Next, the following steps are repeated for each problem (410-02, 410-05). First, stakeholders are identified based on the impact scope (410-03). If the problem is known, information such as the initial response, impact scope, causes, and response status is collected. If not, information such as the expected impact scope, response type, and countermeasures is collected (410-04). The information compiled through this repetition (problem list 410-06) is shared among stakeholders while achieving a common understanding. In this embodiment, the problem list 410-06 is used as the input source, but this is not a limitation. The most appropriate input source for the incident or project may be selected.
[0088] Next, in this embodiment, accountability is considered not only from the perspective of the stakeholders who should be held accountable, but also from the perspective of the stakeholders who demand accountability, and by verifying the accountability demanded by the latter, it is not a one-way fulfillment of accountability from the former to the latter, but a two-way one, so that trust in each other's accountability can be assured. In this embodiment, the former is called the accountability provider and is recorded in explainer 320-02 in Figure 3C, and the latter is called the accountability beneficiary and is recorded in explainee 320-03 in Figure 3C.
[0089] Once the obstacle list 410-06 is collected, the next step is to reach a common understanding among stakeholders regarding the accountability structure. This allows the accountability provider to achieve accountability with confidence in terms of reliability, and allows the accountability beneficiary to understand and accept it. In this embodiment, this procedure breaks down accountability into one-to-one relationships between stakeholders (who and to whom) to understand its structure. In this process, the "purpose" of accountability—why it is being fulfilled—is clarified, and the "demand" for accountability is defined. This is a request from the accountability beneficiary to the accountability provider. In response to this request, the accountability provider and the accountability beneficiary agree on the content of the implementation. Furthermore, to ensure reliability as defined by IEC 62853, the possibility of deviation from this agreement is anticipated, and preventative measures are devised. The accountability structure is a collection of these elements combined into a single information unit. The specific steps are described below.
[0090] Select any one obstacle from the obstacle list 410-06 (410-11). The following steps are repeated for all obstacles in the obstacle list 410-06 (410-10 to 410-21). Identify "who" and "to whom" the accountability for the selected obstacle will be fulfilled, with one stakeholder to one, and set the stakeholders (410-12). "Who" is the stakeholder who should be held responsible, and is set as the accountability provider (explainer) 320-02 shown in Figure 3C. "Who" is the stakeholder who seeks accountability, and is set as the accountability beneficiary (explainee) 320-03 shown in Figure 3C. If multiple combinations of accountability providers and accountability beneficiaries are identified, separate accountability structures 310-01 are used for each.
[0091] In the purpose description 410-13, the "purpose" of the accountability exercise is defined. The scope should be the range that involves the identified combination of accountability provider and accountability beneficiary. This is because the purpose will differ depending on the combination of accountability provider and accountability beneficiary, even when carrying out accountability for the same disability. This is then recorded in the purpose 320-01 in Figure 3C.
[0092] Next, in the requirement description 410-14, the accountability beneficiary extracts their hopes and expectations for the accountability provider to fulfill their accountability obligations. The accountability provider then analyzes what the accountability beneficiary specifically requires, and determines the requirements. These are then recorded in requirement 320-04 in Figure 3C.
[0093] In the description of the accountability provider's actions 410-15, the content that the accountability provider should carry out in accordance with the purpose based on the above requirements is described as the accountability provider's actions. This is then recorded in explainer_action 320-05 shown in Figure 3C. In addition, in the description of the accountability beneficiary's verification means 410-16, how to verify that the accountability provider's actions satisfy the above requirements is described as the accountability beneficiary's verification means. This is then recorded in explainee_action 320-06 shown in Figure 3C.
[0094] In the description of expected deviations 410-17, cases in which the agreement among stakeholders regarding the accountability structure for the selected fault handling is deviated are predicted, and the type of problem that will occur is described as an "expected deviation." This is then recorded in assumed_deviation 320-07 shown in FIG. 3C. In the description of preventive measures 410-18, how to prevent expected deviations is described, and the proposed countermeasures are described as "preventive measures." This is then recorded in preventive_action 320-08 shown in FIG. 3C.
[0095] After these steps, a consensus is formed (410-19), which is performed by describing the discussion in the GSN extension structure description (400-04) in FIG. 4A and performing a consensus (400-02). Next, the accountability structure (310-01) obtained in the above steps is connected (410-20) to other related accountability structures (310-01). This is recorded as a link with the label accountability_chain (320-13) shown in FIG. 3C. In this embodiment, seven types are recorded in the link attribute type: "followed-by," "depends-on," "ruled-by," "related-to," "reify," "conflict-of," and "independent-of." However, other types may also be recorded. The device 100 repeats the above steps (410-10 to 410-21) for each failure. The link attribute type of the link with the label accountability_chain 320-13 is an example of a plurality of second relationships between a plurality of self-description information units generated for a plurality of deviations.
[0096] In this stage, the accountability execution priority determination 410-40 is performed for the multiple accountability structures 310-01 recorded in the structured data recording unit 100-02. Regarding actual or anticipated failures, stakeholders must ensure prevention and early response, plan recurrence prevention measures, and fulfill their accountability. However, the number of preventive measures and recurrence prevention measures that can be implemented simultaneously is limited. High-priority measures must be identified and implemented from the accumulated accountability structures 310-01 in accordance with the organization's resources and operational plans. Furthermore, regardless of priority, creating an accountability structure and agreeing among stakeholders on when it is scheduled to be implemented and why it is prioritized can avoid a situation where nothing is prepared when the failure occurs. Furthermore, because the implementation method has already been considered in the accountability structures 310-01, it is possible to explain when and what measures were planned to be implemented.
[0097] 4C illustrates an example of a procedure for describing 400-04 a GSN extension structure by the device 100 according to one embodiment. The device 100 executes the description 400-04 of the GSN extension structure when reaching a consensus among stakeholders regarding the recorded contents of the purpose 320-01, the requirement 320-04, and the preventive action 320-08, among the elements constituting the accountability structure 310-01 in the accountability structure analysis 400-01.
[0098] First, a content asserting the legitimacy of the recorded content is set as the top goal of the GSN extension structure in top goal setting 420-01. The device 100 records the top goal in claim 330-01 shown in FIG. 3D. At that time, the device 100 may also set a context 420-02, including the stakeholders related to the legitimacy assertion and the background and premise of the assertion. The device 100 records this in stakeholder 330-02 or context 330-03 shown in FIG. 3D. The top goal claim 330-01 and the stakeholder 330-02 or context 330-03 are connected by a link labeled inContextOf.
[0099] Next, it checks whether evidence exists for all goals (420-03). The validity of all goals is based on the premise that it is proven by supporting evidence. Therefore, since a goal without evidence is insufficiently discussed (No in 420-03), one goal is selected and a discussion decomposition strategy is set (420-04). This strategy serves as an approach, perspective, or viewpoint for discussing the validity of the goal. The device 100 records the strategy in strategy 330-04 shown in Figure 3D. The goal claim 330-01 is connected to the strategy 330-04 by a link with the label supportedBy.
[0100] Subgoals are extracted 420-05 according to the above strategy. One or more subgoals are extracted. For each subgoal, a context including the background and premise related to the claim of legitimacy may be set 420-06. This is recorded in claim 330-01 shown in Figure 3D, and if the context exists, the claim and the context 330-03 are connected by a link labeled inContextOf.
[0101] For one or more of the above subgoals, it is checked in 420-07 whether evidence supporting each subgoal exists. If it does not exist (No in 420-07), a subgoal that does not exist is selected, a discussion decomposition strategy is set in 420-04, and the above procedure is repeated. If evidence exists (Yes in 420-07), evidence extraction is performed in 400-05. The extraction method differs depending on the type of evidence. For example, if the evidence is an audit log, a certificate of non-tampering is attached along with the log file. In the case of analysis results using machine learning, the accuracy, error, and contribution of the acquired features to the results are attached.
[0102] After extracting the evidence, the system checks again whether evidence exists for all goals (420-03). If it does (Yes in 420-03), the validity of the claim of the top goal is proven through discussion among stakeholders and evidence that serves as the basis for one or more subgoals. This is then reached as a consensus among stakeholders (400-02). Specifically, agreement on the proof is obtained from all stakeholders. For the top goal, a link 330-09 is connected from claim 330-01 in FIG. 3D to the stakeholder (element user 330-08). At this time, the label of the link records the agreement state. In this embodiment, one of three states is recorded: agree, disagree, or abstain. Depending on the embodiment, agreement from all stakeholders may be required, abstention may be included, or other states may be acceptable. Discussions on any claim may be recorded as a record of agreement on the claim among stakeholders using the procedure described in GSN extension structure description 400-04.
[0103] 4D shows an example of a procedure for argumentative confirmation 400-03 performed by device 100 according to one embodiment. If there is any misunderstanding or difference in understanding among stakeholders during consensus building 400-02 or during description of the GSN extension structure 400-04, or if any stakeholders express opposition or abstention to the above agreement state, argumentative confirmation 400-03 is performed to aim for unanimous agreement. In this embodiment, the basic form of the argumentative structure described in The Uses of Argument by Stephen Toulmin (ISBN: 978-0521534833) is used, but other argumentative structures may also be used.
[0104] First, a stakeholder in an issue that requires discussion identifies the participants in the discussion (430-01). The identified participants are recorded in user / participant in discussion (element user 340-07) shown in Figure 3E. Next, an opinion is expressed (430-02). In this embodiment, instead of "assertion" shown in The Uses of Argument, the expression "opinion" is used so that users can easily understand. The expressed opinion is recorded in claim / assertion (element claim 340-01) shown in Figure 3E.
[0105] Next, the trigger for expressing the opinion is expressed 430-03. This is an expression of the "basis" described in The Uses of Argument. In this embodiment, the expression "basis" is used to facilitate user understanding. The expressed trigger is recorded in data / basis (element data 340-02) described in FIG. 3E. The element data 340-02 is many-to-one connected to the element claim 340-01 via a link with the label support.
[0106] If the stakeholders participating in the discussion can agree on the basis for the claim, discussion is considered unnecessary. Discussion becomes necessary when there is a discrepancy between the claim and the basis, and there are stakeholders who cannot agree. In this case, the stakeholder who expressed their opinion makes a statement of reason 430-04 to bridge the discrepancy. This is a statement of the "reason" described in The Uses of Argument. In this embodiment, the expression "reason" is used. The stated reason is recorded in the warrant / reason (warrant element 340-03) shown in Figure 3E. The warrant element 340-03 is connected one-to-one to the data element 340-02 via a link with the label context.
[0107] Once all the statements have been collected, stakeholders confirm whether they agree with the statements 430-05. If they agree, they confirm their agreement 430-06. The claims shown in Figure 3E are connected by a link 340-08 labeled agree if they agree, a link 340-08 labeled disagree if they disagree, or a link 340-08 labeled abstain if they abstain.
[0108] If the answer is No, return to the opinion statement 430-02 and make a restatement 430-07. This is recorded in the claim / assertion (element claim 340-04) shown in Figure 3E, and is connected one-to-many by the link with the label update shown in Figure 3E. If the restatement 430-07 results in a correction 430-08 or 430-09 to the trigger statement 430-03 or the reason statement 430-09, the correction is recorded.
[0109] These are recorded in the data / basis (element data 340-05) or warrant / rationale (element warrant 340-06) shown in Figure 3E. The trigger (rationale) is connected one-to-many from element data 340-02 to element data 340-05 with a link labeled "update." The reason (rationale) is connected one-to-many from element warrant 340-03 to element warrant 340-06 with a link labeled "update." In this embodiment, once agreement is confirmed in this manner, the discussion ends and the discussion process is persisted in the structure shown in Figure 3E.
[0110] 4E illustrates an example procedure for consensus building 400-02 by device 100 according to one embodiment. First, a target for consensus building is identified 440-01. In this embodiment, the targets are the goal set in setting the top goal 420-01 in the GSN extension structure description 400-04 shown in FIG. 4B (the claim element 330-01 shown in FIG. 3D in the record) and the evidence extracted by evidence extraction 400-05 (described below) (the evidence element 330-05 shown in FIG. 3D in the record), but any target may be selected.
[0111] Next, the stakeholders who require consensus are identified 440-02. In this embodiment, the stakeholders (420-02) set in the GSN extension structure description 400-04 are identified. This corresponds to the record of the element user 330-08 shown in Figure 3D.
[0112] The following steps (440-03 to 440-09) are applied to all identified stakeholders: First, select one stakeholder and confirm whether that stakeholder agrees to reach a consensus on the subject (440-04).
[0113] If the stakeholder is in favor (440-04: Yes), the approval is recorded 440-05. A link 330-09 labeled agree is connected from the element claim 330-01 shown in Figure 3D to the element user 330-08 of the stakeholder who has expressed his or her intention to approve. If the stakeholder is not in favor (440-04: No), the stakeholder is asked 440-06 whether or not they oppose the consensus formation.
[0114] If the stakeholder is opposed (440-06: Yes), the opposition is recorded 440-07. A link 330-09 labeled disagree is connected from the element claim 330-01 shown in Figure 3D to the element user 330-08 of the stakeholder who has expressed their opposition. If the stakeholder is not opposed (440-06: No), the stakeholder is determined to have abstained, and an abstention record 440-08 is left. A link 330-09 labeled abstain is connected from the element claim 330-01 shown in Figure 3D to the element user 330-08 of the stakeholder who has expressed their abstention.
[0115] After repeating the above procedure for all identified stakeholders, the consensus building rule is to check whether a unanimous agreement has been reached (440-10). If 440-10 is Yes, check 440-11 to see if everyone is in agreement, and if 440-11 is Yes, it is considered that a consensus has been reached. On the other hand, if 440-11 is No, it is necessary to carry out argumentative confirmation (400-03) to change the minds of stakeholders who have expressed opposition or abstention to agree. The above procedure is then repeated until a unanimous agreement is reached.
[0116] 4F shows an example of a procedure for extracting 400-05 a trail by the apparatus 100 according to one embodiment. Since the procedure for extracting a trail differs depending on the corresponding goal, FIG. 4F shows an abstract procedure.
[0117] First, the type of trail is identified 450-01. In this embodiment, there are three types of trail: document, audit log, and processing result, but other types may also be used.
[0118] If the type is a document (Yes in 450-02), the corresponding document is identified (450-03). In this step, an existing document or a new document that can confirm the validity of the assertion recorded in the goal is identified. It is checked whether the document is minutes (450-04), and if it is minutes (Yes in 450-04), the structure of the minutes is extracted (450-05). The structure is recorded (450-11) in an element connected to the minutes element 350-03 shown in Figure 3F. If it is not minutes (No in 450-04), the URL where the document can be referenced is recorded (450-11) in the document element 350-01 shown in Figure 3F.
[0119] If the type is an audit log (Yes in 450-06), a trail is extracted from the log (450-07). This step differs depending on the log recording system implemented, but in this embodiment, a database that aggregates API access to the system is accessed, and log records that can prove the validity of the assertion recorded in the goal are extracted using a search expression. The actual state of the file containing the extracted log, or accessible access information, is recorded in element log 350-02 shown in Figure 3F (450-11).
[0120] If the type is a processing result (Yes in 450-08), the processing result is analyzed (450-09) to determine whether the assertion recorded in the goal is valid. For example, if the result is a result of a SHAP process, one of the XAI techniques, the contribution of the feature to the prediction can be analyzed by visualizing the calculated SHAP value. The analysis result or a URL that can reference the analysis result is recorded (450-11) in the log element 350-02 shown in Figure 3F. A URL to the document containing the analysis result may also be recorded in the document element 350-01.
[0121] If the type does not fall into any of the above three categories, an individual response 450-10 will be performed. In the case of an individual response, if the incident is caused by an unexpected incident, a deviation occurrence 400-06 will be determined. The response result or a URL at which the response result can be referenced is recorded 450-11 in the element log 350-02 shown in FIG. 3F. A URL to a document of the response result may be recorded in the element document 350-01.
[0122] Finally, a consensus 400-02 is reached among the stakeholders on the recorded evidence, thereby confirming and approving a common understanding among the stakeholders regarding the legitimacy of the goal based on the evidence.
[0123] FIG. 4G shows an example of a procedure for determining whether a deviation has occurred 400-06 by the device 100 according to one embodiment. First, an incident is detected 460-01. In this embodiment, an incident includes, but is not limited to, when a user notices an abnormality and notifies the device, when a pre-configured notification alert is triggered in a monitoring program, when a new vulnerability assigned a Common Vulnerability Exposure (CVE) is found, when the type of evidence is a processing result and the results of analyzing the processing result 450-07 are unsatisfactory, and so on. This is a state in which some adverse effect is occurring to the user and can be identified.
[0124] As will be described in detail later, the device 100 may analyze time-series data of the reliability feature to detect a change in the distance from the reliability feature at a first time to a reliability feature at a second time after the first time, thereby detecting a deviation in the consensus state of the above-described multiple evidences among the multiple stakeholders at the second time. The device 100 may also include a detection unit that detects the occurrence of the deviation. Note that detecting a change in the distance from the reliability feature at the first time to the reliability feature at the second time may refer to detecting that the distance between the reliability feature at the time immediately after the deviation occurs and the reliability feature at the time immediately before the deviation occurs becomes a second distance different from the first distance, assuming that the distance between two reliability feature values that are temporally adjacent to each other is maintained at a first distance during a period in which the deviation does not occur.
[0125] Next, it is determined whether the incident is an expected deviation (460-02). The records of the assumed_deviation element 320-07 shown in FIG. 3C are first checked. If a record matching the incident exists (Yes in 460-02), that is, if the incident is an expected deviation, it is determined whether a preventive measure can be implemented (preventive_action element 320-08) shown in FIG. 3C (460-03). If a preventive measure can be implemented (Yes in 460-03), the preventive measure is implemented (460-04).
[0126] If the determination in 460-02 whether the deviation is within the expected range is No, or if the determination in 460-03 whether preventive measures can be implemented is No, then a determination is made in 460-05 whether a fault can be identified. If a fault cannot be identified (No in 460-05), the procedure ends. In this case, stakeholders may be informed that an unexpected deviation has occurred, or that an expected deviation has occurred but preventive measures cannot be implemented, and countermeasures may be discussed. If a fault can be identified (Yes in 460-05), the procedure proceeds to extraction of self-descriptive information units related to the deviation 400-07.
[0127] The extraction unit may be configured, when the detection unit detects that a deviation from the consensus state has occurred, to extract a self-description information unit related to the deviation from the consensus state from the plurality of self-description information units at the second time. Furthermore, as described above, the extraction unit may be configured, when the detection unit detects that a deviation from the consensus state has occurred, to determine whether the deviation from the consensus state corresponds to any of a plurality of expected deviations. The extraction unit may be further configured, when an expected deviation corresponding to the deviation exists, to determine whether a countermeasure for the expected deviation is executable. The extraction unit may be further configured, when a countermeasure for the expected deviation is not executable, to determine whether the deviation from the consensus state is identifiable, and, if identifiable, to extract a self-description information unit related to the deviation from the consensus state from the plurality of self-description information units at the second time.
[0128] 4H shows an example of a processing procedure for extracting 400-07 self-descriptive information units related to a deviation by the device 100 according to one embodiment. First, a feature vector is calculated 470-01 from text describing the incident. The method for this calculation is not specified in this embodiment, but may be, for example, the TD (word frequency)-IDF (inverse document frequency) method, mutual information, or an API provided by an existing cloud service (e.g., the embeddings API provided by ChatGPT, Inc.).
[0129] Next, the structured data recording unit 100-02 is searched (470-02) using the feature vector. Then, the relevance between the search results and the deviation is calculated (470-03). The search results are elements of self-describing information units that make up the structured data, and the relevance is calculated from their position within the structure. All elements in the search results are examined (470-04) to see if they have a strong relationship with the deviation, and elements that have a weak relationship between the search results and the deviation (No in 470-04) are removed from the search results (470-05). The remaining search results are used to continue with the structural analysis of accountability (400-01).
[0130] As described above, the collection 400 of the self-description information units 300 in this embodiment is characterized by being repeatedly executed during the progress of the project.
[0131] [Input of Structured Data] Figure 5A shows an example configuration 500 of the UI unit 100-06 according to one embodiment. Operations in the structured data collection unit 100-01 described above are performed via a user interface provided by the UI unit 100-06. The UI unit 100-06 is broadly composed of three function groups (FGs): a conversation control FG 501, a recording control FG 502, and a machine learning control FG 503. The conversation control FG 501 is primarily responsible for interactions with users. The recording control FG 502 is primarily responsible for interfacing between the structured data recording unit 100-02 and the trail management unit 100-03. The machine learning control FG 503 is primarily responsible for interfacing between the trail management unit 100-03, the feature calculation unit 100-04, and the agreement reliability prediction unit 100-05.
[0132] The conversation control FG 501 consists of two elements: a Chat UI unit 501-01 and a conversation unit 501-02. The recording control FG 502 consists of three elements: a display unit 502-01, a query unit 502-02, and a search unit 502-03. The machine learning control FG 503 consists of two elements: a dashboard unit 503-01 and a history access unit 503-02. Note that an example configuration of the machine learning control FG 503 in this embodiment will be described after the details of the feature calculation unit 100-04 and the agreement reliability prediction unit 100-05 in this embodiment.
[0133] The Chat UI unit 501-01 in this embodiment cooperates with an external service to provide a user with a conversational user interface. This embodiment is characterized by a conversational user interface, but a GUI (Graphical User Interface) type user interface may also be used. In this embodiment, Slack is used as the external service, but other similar services may also be used.
[0134] The conversation unit 501-02 constructs a search query in the query unit 502-02 in accordance with instructions from the Chat UI unit 501-01, and instructs the search unit 502-03 to acquire elements necessary for the conversation from the structured data recording unit 100-02. The elements are presented to the user using the display unit 502-01, or are included in the conversation.
[0135] Figures 5B-1, 5B-2, 5B-3, and 5B-4 illustrate a segmented example of a conversation scenario 510 presented to a user by the UI unit 100-06 according to one embodiment. Figure 5C illustrates a conversation scenario 520 related to the argumentative verification 400-03 described in Figure 4A. In Figures 5B-1, 5B-2, 5B-3, and 5B-4, pairs (A), (B), (C), (D), and (E) are interconnected. As shown in Figures 5B-1 and elsewhere, a portion of the conversation scenario 520 is connected to Figure 5C. In Figures 5B-1 and elsewhere, the portion labeled A-Map corresponds to the scenario related to the accountability structural analysis 400-01 described in Figure 4A. In Figures 5B-1 and elsewhere, the portion labeled D-Case corresponds to the scenario related to the GSN extension structural description 400-04.
[0136] In Figure 5C, a discussion unfolds in a thread on a Slack channel (520-01). A ChatBot (shown as dADD by Slack in Figure 5C) posts a message to the channel saying, "Please write your opinion in a message," so the user posts their opinion (argument) (520-02). The ChatBot then posts, "What prompted your opinion?" so the user posts their reason (evidence) (520-03). The ChatBot then posts, "Please write your reason," so the user posts their reason (evidence), and the ChatBot then posts, "Are there any other reasons that you've noticed?" and presents the user with the options "Yes" or "No" (520-04).
[0137] By pressing the "Yes" button, the user can write further evidence and arguments in accordance with the ChatBot's post (520-05 and 520-06). After the user has written their argument, the ChatBot presents the option of "Yes" or "No" to write further evidence. Selecting "Yes" allows the user to write further evidence and arguments, but selecting "No" causes the ChatBot to write a summary of the user's opinions up to that point and give channel participants the option to "start a discussion to deepen understanding" (520-07). Here, users can freely discuss the opinions presented by the ChatBot in the channel and, if necessary, revise previous posts presented by the ChatBot (520-08). The ChatBot presents a "Get everyone's agreement" button to channel users for post 520-08. Once the discussion is exhausted, pressing this button causes the ChatBot to write the status of the channel users' expressed opinions (520-09). Once everyone agrees, the ChatBot posts the content of 520-10 to the channel, and the discussion ends.
[0138] 6 shows a configuration example 600 of the structured data recording unit 100-02 according to one embodiment. The configuration example 600 serves to persist the self-describing information units 300 as structured data. The configuration example 600 is composed of three elements: a search engine unit 600-01, a graph structure calculation unit 600-02, and a column-based database 600-03.
[0139] The search engine unit 600-01 searches for components of the self-descriptive information unit 300 based on an arbitrary query from the search unit 502-03. The graph structure calculation unit 600-02 reconstructs the structured data shown in FIGS. 3A to 3F as a graph structure in memory, persists it in the column-type database 600-03, and executes calculations related to the graph structure. The graph structure calculation unit 600-02 performs calculations required for graph-based network analysis, such as extracting subgraphs, calculating the connection distance between elements, and extracting matrices. The column-type database 600-03 may be replaced with an SQL database or another NoSQL database. The graph structure calculation unit 600-02 and the column-type database 600-03 may be integrated into a graph database.
[0140] [Reliability based on data SLA] The GSN extension structure 310-02 shown in Figure 3D represents the argument structure in a discussion and is used as a means of expression when reaching a consensus in this embodiment. This embodiment is characterized by using an SLA (Service Level Agreement) for data as an expression format when reaching an agreement between stakeholders. An example of this procedure is described below.
[0141] FIG. 7A shows an example of a data SLA 700 in the form of a GSN extension structure 310-02 according to one embodiment. Here, an SLA for data that presupposes the calculation of a machine learning model (simply referred to as a model in the figure) is assumed. The top goal is set to "data quality meets model requirements" 700-01. Therefore, the existence of an information source, such as a model card, that describes the parameter information and performance of the machine learning model is assumed / context 700-02.
[0142] The top goal in question breaks down and deepens the discussion using the following three strategies. The first is the discussion "Regarding model requirements" 700-03. This discussion is divided into two assertions: "Model requirements are defined" 700-04 and "Model requirements are agreed upon" 700-05. These two assertions are supported by two pieces of evidence: "Model requirements definition document" 700-06 and "Agreement record" 700-07, which demonstrate the validity of the top goal regarding model requirements.
[0143] The second is the discussion "Regarding data quality rules" 700-09. This discussion is divided into two assertions: "The data quality rules conform to the model requirements" 700-10 and "The data quality rules are applied correctly" 700-11. In this embodiment, the use of some kind of data quality support tool (AWS Glue Data Quality, etc.) is assumed 700-08. These two assertions are supported by three pieces of evidence: a "rule description file" 700-12, an "execution log" 700-13, and a "review record" 700-14, and the validity of the top goal regarding the data quality rules is demonstrated.
[0144] The third is the discussion of "Response to Errors" in 700-15. This strategy mainly assumes the case where the processing in 700-11, "Data Quality Rules Are Applied Properly," does not end normally. This discussion is divided into two claims: "The cause of the quality problem can be identified" in 700-16 and "Errors can be corrected" in 700-17. These two claims are supported by two pieces of evidence: "Rule Application History" in 700-18 and "Established Defect Handling Procedures" in 700-19, which justify the top goal of response to errors.
[0145] Ultimately, once the above evidence (700-06 / 07, 700-12 / 13 / 14, 700-18 / 19) is collected and agreed upon among stakeholders, a consensus has been reached on the top goal: "Data quality meets model requirements." As in this example, expressing the data SLA using the GSN extended structure format and forming an agreement between the data provider and data user when using data for secondary purposes ensures the reliability of the data analysis attempted by the data user. While this example assumes the use of a method that primarily controls data quality based on rules, other aspects of data quality (e.g., the impact of bias) may also be required. Therefore, the discussion may be broken down using strategies other than those described in this example.
[0146] Figure 7B shows an example of a data SLA 710 in the form of an accountability structure 310-01, according to one embodiment. It shows three simplified accountability structures. Each of the three summarizes, in card format, the accountability provider, accountability beneficiary, requests from the accountability beneficiary to the accountability provider, the accountability provider's actions in response to the requests, the accountability beneficiary's means of verifying the actions, expected deviations, and preventative measures (710-01 / 02 / 03).
[0147] Card 710-01 expresses in card format the accountability that a data provider has to a data user. The data user's request is to "confirm that data quality conforms to the model requirements." To that end, the data provider provides the data user with the quality characteristics of the data it provides, and the data user confirms that the provided quality characteristics conform to the established model requirements. At that time, in anticipation of cases where data quality is insufficient, measures to improve data quality are planned as a preventative measure. Regarding such preventative measures, specific plans will need to be discussed in the GSN extension structure, but this will be omitted here. The data provider and data user have already reached an agreement regarding this card.
[0148] Card 710-02 expresses in card format the accountability that data quality support tool providers fulfill to data users. The data user's requirement is "to be able to verify that data quality conforms to model requirements." To achieve this, the data quality support tool provider provides a description specification for rules that can define data quality, and the data user confirms that the rules can be written according to the provided specification and that the written rules can achieve the required data quality. In this case, assuming that the rule description in the specification does not result in data quality that conforms to model requirements, the provider identifies elements that can be mitigated from the given data quality as a preventative measure. Note that as a preventative measure in this case, the provider may be asked to modify the support tool. Finally, the data quality support tool provider and data user reach a consensus on the card.
[0149] Card 710-03 expresses the accountability of the data preparer to the data analyst in card format. The data analyst's requirement is to be able to explain the quality of the data used in the data analysis to a third party. To achieve this, the data preparer presents an execution history of applying quality definition rules to the provided data, and the data analyst confirms that the execution history can be used to examine the explainability of the data quality. In this case, assuming that the execution history contains errors, the cause of the error can be identified as a target for preventive measures, and a procedure for dealing with the problem can be established. Regarding preventive measures, specific procedures must be discussed in the GSN extension structure, but this is omitted here. Here, the data analyst is one of the data user roles in the two example cards above, and the other role is the data preparer. Therefore, a dependency relationship exists between card 710-03 and card 710-02, and the type of dependency is recorded using the label accountability_chain 320-13. Finally, the data quality support tool provider and the data user reach a consensus on the card.
[0150] The data SLA 710 in the accountability structure 310-01 format differs from the data SLA 700 in the GSN extension structure 310-02 format in that it organizes the accountability relationship as a one-to-one stakeholder relationship, thereby enabling stakeholders related to data quality to be assured down to the role level.
[0151] 8 shows an example configuration 800 of the trail management unit 100-03 according to one embodiment. The example configuration 800 is made up of four elements: a reference information management unit 800-01, a reference history management unit 800-02, a hash tree calculation unit 800-03, and an SQL database 800-04.
[0152] The reference information management unit 800-01 ensures reliable access to information recorded on a computer. For example, in the case of a URL for a web page, the page may be deleted, and even in that case, it must be possible to access the page as evidence by saving a copy of the page. In the case of an audit log, it is essential to guarantee the authenticity of the log. In the case of a processing result, it is desirable to be able to refer to the program code related to the processing.
[0153] The access history management unit 800-02 records who accessed the information and when. The access should be read-only; write and update operations are prohibited, and attempts to perform such operations are recorded as a log.
[0154] The hash tree calculation unit 800-03 calculates a hash value for the information, manages a hash tree (Merkle tree) for each self-description information unit 300 that includes the information, and persists the management information and reference information to the information in the SQL database 800-04. This makes it possible to guarantee the authenticity of the information, which is part of the trail. Note that in this embodiment, a Merkle tree is constructed for each self-description information unit 300, but other forms may also be used. Furthermore, the hash tree calculation unit 800-03 and the SQL database 800-04 may be integrated into a ledger database.
[0155] [Calculation of Feature Values Related to Trustworthiness] As described above, the device 100 according to this embodiment aims to ensure the trustworthiness of data science. The assurance in this regard means that stakeholders can present the basis for their confidence in the trustworthiness of problem solving. To this end, the device 100 according to this embodiment includes a means for recording structured data that organizes the basis for the confidence in terms of the accountability relationships between the stakeholders, and a means for calculating, from the data accumulated by the recording means, a reliability feature value that uses the consensus state formed among the stakeholders as the basis for the confidence.
[0156] 9A shows an example configuration 900 of the feature calculation unit 100-04 according to one embodiment. The example configuration 900 implements this means. The example configuration 900 is composed of seven elements: a data collection unit 900-01, a data shaping unit 900-02, a data quality confirmation rule unit 900-03, an algorithm execution unit 900-04, an algorithm repository unit 900-05, a feature verification unit 900-06, and a feature quality confirmation rule unit 900-07.
[0157] The data collection unit 900-01 acquires partially structured data of the self-description information unit 300 using the graph structure calculation unit 600-02 shown in Fig. 6. The data collection unit 900-01 also acquires trail data required for calculating feature quantities from the trail management unit 100-03 shown in Fig. 1. In this embodiment, to calculate the features related to the minutes, the following elements are obtained: a group of elements connected to the element minutes 350-03 shown in Figure 3F with a link labeled content; an element group of document 350-01 and log 350-02 connected to element evidence 330-05 which is connected from the element minutes 350-03 with a link labeled trail, an element participant 350-13 in a state of participant agreement connected from the element minutes 350-03 with a link labeled confirmed, and an agreement state connected from the element evidence 330-05 to element user / stakeholder 350-11 with a link labeled agree / disagree / abstain 350-12; however, any group of elements in a range that can be traversed from the element minutes 350-03 may be obtained.
[0158] The data reforming unit 900-02 checks the quality of the data from the data collection unit 900-01 by applying rules recorded in the data quality check rule unit 900-03. The data reforming unit 900-02 performs, for example, normalization check, schema check, missing value check, and noise removal check on the data to maintain consistent data quality. Note that an external service may be used to check the data quality.
[0159] The algorithm execution unit 900-04 retrieves implementation code for executing the algorithm required for feature calculation from the algorithm repository unit 900-05 and applies it to the data output by the data reforming unit 900-02. Note that an external service may be used to execute the algorithm. For example, the gensim implementation (https: / / github.com / RaRe-Technologies / gensim) may be used to embed text into a vector space, or a service provided by a cloud service provider may be used.
[0160] The feature verification unit 900-06 applies the rules recorded in the feature quality confirmation rule unit 900-07 to the feature calculated as a result of the algorithm execution unit 900-04, thereby verifying that the feature is applicable to the agreement reliability prediction unit 100-05. The feature verification unit 900-06 performs, for example, checking for abnormal values or missing values in the feature, checking traceability to the data from which the feature was calculated, checking the distance between features and semantic distance, checking the agreement state, etc.
[0161] Two example calculation procedures for reliability-related features primarily used in this embodiment are shown in Figures 9B and 9C. At least one of these calculation procedures is an example of a method for ensuring the reliability of a service using data science, which includes determining multiple second relationships between multiple self-description information units generated for the multiple deviations, calculating feature values for each of the multiple self-description information units, and correlating the feature values with each other based on the multiple second relationships to calculate a reliability feature value representing the reliability of the service. Furthermore, as described below, the acquisition unit may acquire text data of multiple evidences described by multiple stakeholders via a user interface. For each deviation, the calculation unit may calculate an embedded feature value from text data describing each of the multiple evidences constituting a data structure for each of multiple data structures that are mutually related and included in the self-description information unit. The calculation unit may further calculate a convolutional feature value as a feature value of the self-description information unit by performing a convolutional calculation to convolve the data structure from the calculated embedded feature values while maintaining the multiple first relationships for the multiple evidences in the data structure.
[0162] FIG. 9B illustrates a conceptual diagram of an example 910 of feature calculations for a self-describing information unit 300 by device 100 according to one embodiment. In FIG. 9B, a conceptual diagram of the self-describing information unit 300 as structured data is shown in a solid-line box 910-01. This is structured data at a certain time T1. Labels for links connecting components are omitted in FIG. 9B, but these are the label names shown in FIGS. 3B through 3F. Here, the four structures shown in FIG. 3B are shown as groups of components enclosed in dotted-line boxes.
[0163] Each of these four groups is connected by cross-group links between the components, resulting in a single graph structure consisting of all four groups. This graph structure is also connected to one or more other accountability structures 910-03 by links labeled accountability_chain 320-13 from the element purpose 320-01 of the accountability structure 310-01.
[0164] First, using the graph structure calculation unit 600-02 shown in Fig. 6, the accountability structure 310-01 is extracted as a subgraph 910-02 from the self-description information unit 300. To calculate embedded features (vectors) from the text recorded in each element (node) constituting the subgraph 910-02, the corresponding algorithm implementation is retrieved from the algorithm repository unit 900-05, processed by the algorithm execution unit 900-04 (using an external service in this embodiment), and the results (embedded features in the text of the element) are recorded in the elements constituting the subgraph.
[0165] Next, in order to calculate convolutional features while maintaining the structure of the subgraph 910-02, the corresponding algorithm implementation is retrieved from the algorithm repository unit 900-05 and processed by the algorithm execution unit 900-04 (in this embodiment, RelGraphConv from DGL.ai, which implements Relational Graph Convolutional Network, R-GCN, is used), and the subgraph is replaced with the element CONV(a) 910-04 that records the resulting features (910-05 in the figure).
[0166] 6 is used to extract the argument structure 310-03 as a subgraph 910-06 from the self-description information unit 300. To calculate embedded features (vectors) from the text recorded in each element (node) constituting the subgraph 910-06, the corresponding algorithm implementation is retrieved from the algorithm repository 900-05, processed by the algorithm execution unit 900-04 (using an external service in this embodiment), and the results (embedded features in the text of the element) are recorded in the elements constituting the subgraph 910-06.
[0167] Next, in order to calculate the convolutional features while maintaining the structure of the subgraph 910-06, the corresponding algorithm implementation is retrieved from the algorithm repository unit 900-05 and processed (RelGraphConv) by the algorithm execution unit 900-04, and the subgraph is replaced with the element CONV(b) 910-07 that records the resulting features (910-08 in the figure).
[0168] The above procedure is repeated for the components of the self-descriptive information unit 300, and all elements of the self-descriptive information unit 300 of 910-01 in FIG. 9B are folded 910-09 to calculate an element CONV(c) 910-10 that reflects the characteristics of the elements of that information unit. The self-descriptive information unit 300 is connected to other (one or more) self-descriptive information units 300 via the element purpose 320-01 of the accountability structure 310-01 and a link labeled accountability_chain 320-13. As a result, structured data 910-11 is calculated that reflects the characteristics of the components of the structured data at a certain time T1 and is connected by a link labeled accountability_chain 320-13. Note that this process is an example of associating multiple features with each other based on multiple second relationships, as described above.
[0169] The same procedure is repeated for time Tn (n=2, 3, ..., N) to calculate time series data of the structured data from time T1 to T. Note that this process is an example of associating multiple feature amounts with each other based on multiple second relationships described above.
[0170] The database records all structured data from the time the system was developed and designed, making it possible to identify the time and content of past rewrites.In addition, by comparing the structured data from the current point in time with the structured data from the previous point in time, it is possible to predict deviations from the consensus state.
[0171] 9C illustrates an example procedure 920 for calculating features of a minutes-starting point by the device 100 according to one embodiment. Minutes are documents created, recorded, and shared among stakeholders in most projects. In this embodiment, minutes are recorded in an element having a structure connected by a link labeled "content" from the element "minutes" 350-03 in the evidence structure illustrated in FIG. 3F.
[0172] First, minutes are identified 920-01. The minutes for which features are to be calculated are selected in chronological order. One of these minutes is selected, and subgraph identification 920-02 is performed. One element, minutes 350-03, is selected from the subgraph, and the range that can be traversed by following links from that element and the elements included within it are determined. Multiple elements are selected. One element from that range is selected, and in order to calculate the embedded features of the text string 920-03, the corresponding algorithm implementation is retrieved from the algorithm repository unit 900-05 and processed (using an external service in this embodiment) by the algorithm execution unit 900-04, and the result (embedded features of the element's text) is recorded in the element. The same process is performed on all elements in the range (920-04), while maintaining the structure including the elements in the range.
[0173] Next, to calculate convolutional features for structured data including an element minutes 350-03 selected from the subgraph, the corresponding algorithm implementation is retrieved from the algorithm repository unit 900-05, processed (RelGraphConv) by the algorithm execution unit 900-04, and the subgraph is replaced with an element that records the resulting features (920-05). If the element participant 350-13 is connected to the selected element minutes 350-03 by a link labeled confirmed, the minutes have been agreed upon among the participants (920-06).
[0174] In this step, the connection to user / stakeholder via a link labeled agree / disagree / abstain from element evidence 330-05, which is connected to the element minutes 350-03 via a link labeled tail, may be extracted as the agreement state (920-06). Similarly, the status of element todo 350-08 is extracted. At this point, a set of three features, consisting of the convolution feature of the subgraph, the agreement state, and the todo state, has been calculated for the time of the selected minutes.
[0175] By repeating the above procedure for all minutes (920-08), we finally calculate a set of three features for the time series of minutes: the subgraph convolution feature, the consensus state, and the todo state.
[0176] The features calculated by the above procedure are a time-series data set, reflecting the consensus state formed among the stakeholders at a given time. Taking this into consideration, the device 100 according to this embodiment uses the features as features related to reliability for predicting changes in the consensus state, as described below. Note that the method for calculating the reliability features is recalculated as necessary, since the target substructure changes depending on the procedure of the agreement reliability prediction unit 100-05, as described below.
[0177] In addition, the reliability features calculated by the above procedure are assigned labels as additional attributes. Since the calculation is based on the source structure (accountability structure, GSN extension structure, argument structure, or evidence structure) or a substructure, the state of the corresponding structure, including the agreement status, is assigned to the label as an attribute.
[0178] [Prediction of Changes in Agreement State] As described above, in this embodiment, guaranteeing reliability means that stakeholders can present grounds for confidence in the reliability of problem resolution. For this purpose, the above-mentioned means calculates reliability features based on the agreement state formed between the stakeholders. When a failure occurs, the agreement state can be considered to have been broken, the self-description information unit at that time is updated, and the feature is calculated. When the failure is recovered, the agreement state can be considered to have been restored, the self-description information unit at that time is updated, and the feature is calculated. The difference between the two feature values is the subject of analysis and prediction.
[0179] That is, this embodiment is characterized by analyzing the time-series data of the feature quantities to predict how the consensus state regarding reliability at a certain time will change at the next time. As a result, it is possible to take actions based on the prediction, as will be described later, and it is expected that the occurrence of a fault will be prevented. Note that the term "prediction" in this embodiment has two aspects: one is the prediction of missing relationships, and the other is the prediction of changes or trends over time, but in this specification, these terms may be used without any particular distinction.
[0180] 10A shows a configuration example 1000 of the agreement reliability prediction unit 100-05 according to one embodiment. The configuration example 1000 includes eight elements: a reliability feature collection unit 1000-01, a time-series database unit 1000-02, a self-description information unit relationship prediction unit 1000-03, a consensus formation depth prediction unit 1000-04, a trail structure variation prediction unit 1000-05, an argument structure variation prediction unit 1000-06, a contribution calculation unit 1000-07, and a feature-origin structured data search unit 1000-09.
[0181] The reliability feature collection unit 1000-01 collects features corresponding to substructures related to the self-description information units 300 as a time-series data set from the feature calculation unit 100-04 and the feature sharing control unit 100-07. The time-series database unit 1000-02 records the time-series data set in a time-series database and uses it for analysis in the self-description information unit relationship prediction unit 1000-03, the consensus building depth prediction unit 1000-04, the evidence structure variation prediction unit 1000-05, and the argument structure variation prediction unit 1000-06.
[0182] The self-description information unit relationship prediction unit 1000-03 predicts the relationship (label attribute) between the self-description information units 300 connected by labels. The consensus formation depth prediction unit 1000-04 predicts changes in the consensus state, taking into account the depth of common understanding and the time required for consensus formation, from a time series data set of reliability features corresponding to the partial structure of the self-description information units 300.
[0183] The evidence structure variation prediction unit 1000-05 predicts changes in the influence of the evidence on the agreement state from a time series data set of reliability features corresponding to the substructure of the self-description information unit 300 of the evidence structure 310-04. The argument structure variation prediction unit 1000-06 predicts changes in the agreement state from a time series data set of reliability features corresponding to the substructure of the self-description information unit 300 of the argument structure 310-03.
[0184] The contribution calculation unit 1000-07 calculates the degree of reliability features that contribute to the results of the self-description information unit relationship prediction unit 1000-03, the consensus formation depth prediction unit 1000-04, the evidence structure variation prediction unit 1000-05, and / or the argument structure variation prediction unit 1000-06. The feature-origin structured data search unit 1000-09 extracts structured data recorded in the structured data recording unit that served as the basis for calculating the features.
[0185] FIG. 10B shows an example processing procedure 1010 of the relationship prediction unit 1000-03 between self-description information units according to one embodiment. First, a model for calculating the relationship prediction is acquired. Samples of reliability features related to the self-description information units 300 are extracted 1010-01 from the time-series database unit 1000-02. A learning algorithm is selected, and parameters are selected 1010-02. Model calculation 1010-03 is performed using the samples. Since the prediction model has been acquired at this point, next, self-description information units whose relationships are of interest are identified, reliability features are extracted, and the models are applied 1010-05.
[0186] Next, the prediction result is subjected to result verification 1010-06 of the relationship between self-descriptive information units. The feature-based structured data search unit 1000-09 attempts to extract the label attribute of the self-descriptive information unit 300 with the closest information distance from the prediction result. If extraction is not possible, the result verification 1010-06 is determined to be invalid (No in 1010-07), and the learning algorithm selection is redone, along with the relevant parameter selection 1010-02. The procedure of redoing the model calculation and verifying the validity of the prediction result is repeated (1010-03 to 1010-06). If it is determined to be valid (Yes in 1010-07), the process ends. Note that a model determined to be valid may be reused. In this case, the procedure starts from 1010-08.
[0187] 10C-1 and 10C-2 show an example processing procedure 1020 of the consensus building depth prediction unit 1000-04 according to one embodiment. Two patterns are illustrated in 10C-1 and 10C-2. The example in 10C-1 starts from consensus building, while the example in 10C-2 starts from a change in a component of the self-describing information unit 300.
[0188] In the example of Figure 10C-1, first, the target consensus is identified 1020-01. While the example shown in Figure 3D is configured so that consensus can be reached for any goal claim 330-01, the example in Figure 10C-1 targets consensus on the top goal.
[0189] Next, two related structures are identified. One is the identification of a related accountability structure 1020-02, and the other is the identification of a related trail structure 1020-03. The former identifies the accountability structure 310-01 connected to the top goal by a link with the label dcase (or a link with the label amap). The latter identifies the trail structure 310-04 connected to the element evidence 330-05 of the GSN extension structure 310-02 by a link with the label trail.
[0190] The top goal (assertion) is an assertion established through the structural analysis of accountability, and consensus has been formed among stakeholders based on the evidence element evidence 330-05. By comparing these accountability structures with the evidence structure 1020-04, the depth of consensus formation can be predicted.
[0191] In this embodiment, depth prediction is performed through model acquisition and application 1020-10. First, samples of accountability structures and trail structures are extracted (steps 1020-11 and 1020-12), respectively. A learning algorithm is selected, and parameters are selected 1020-13. The learning algorithm is applied to each of the samples to calculate models for the accountability structures and the trail structures (steps 1020-14 and 1020-15).
[0192] If each model is determined to be valid (Yes in step 1020-16), the model is applied to the identified accountability structure (1020-17) to obtain feature quantities. Similarly, the model is applied to the identified trail structure (1020-18) to obtain feature quantities. The distance between the two feature quantities is calculated (1020-19). In this embodiment, if the distance is determined to be close, it is determined in step 1020-01 that the accountability structure and trail structure for the target agreement closely correspond.
[0193] In the example of Fig. 10C-2, first, the feature quantity for the self-description information unit 300 at time T1 is calculated 1020-21. This calculation can be performed using the above procedure 910 (1020-22). Next, changes to the components of the self-description information unit 300 are identified 1020-23.
[0194] As the project progresses, changes occur to the components of the self-description information unit 300, for example, when new minutes are created, evidence is updated due to an incident, an argument structure is developed to deepen the discussion, etc. This point in time is designated as time T2, and the feature quantities related to the self-description information unit 300 are calculated.
[0195] By identifying the difference in features between time T2 and time T1 (1020-26), the function of the feature-based structured data search unit 1000-09 (described later) is used to identify the substructure corresponding to the identified feature (1020-26), and the extent of the impact of the change is predicted as a change in the consensus state. Note that the depth of consensus building may be determined by a method other than that described in this embodiment. For example, the determination may be made by taking into account cases where consensus building was easily achieved and cases where consensus building was difficult.
[0196] Figure 10D shows an example processing procedure 1030 of the trail structure mutation prediction unit 1000-05 according to one embodiment. In this embodiment, a "mutation" can refer to a change in the trend of a feature value over a past time series. For example, if the trend is expressed as positive and negative, it can refer to a change from positive to negative (or vice versa). This procedure targets minutes (a group of elements related to the element minutes 350-03 shown in Figure 3F).
[0197] First, the minutes of the meeting to be processed are identified as a time-series data set (1030-01). Next, step 920 shown in FIG. 9C is executed (1030-02). As a result, a time-series data set is obtained. A distance function is identified (1030-03), and the time-series data set is clustered (1030-04).
[0198] The validity of the results is examined 1030-05, and if there is any doubt about the validity, the distance function identification 1030-03 is redone and clustering 1030-04 is redone. At this time, the possibility of redoing step 920 1030-07 is examined. If it is determined to be valid (Yes in 1030-05), the trail structure is explained 1030-06 using the feature-based structured data search unit 1000-09 from the features identified in step 1030-01 in the same cluster. In this embodiment, the trail structure 310-04 is connected to the accountability structure 310-01 and the GSN extension structure 310-02, and the relationship with the agreement state can be explained.
[0199] The evidence includes the document element document 350-01 and / or the log element log 350-02 shown in FIG. 3F. In the case of documents, mutations in the evidence structure are predicted by comparing keyword features related to the content of the consensus building with keyword features related to the document content. In the case of logs, audit logs are included, but standardizing the recording format increases the effectiveness of the method and enables the use of various predictive detection methods (parameter changes and latent structural changes). Since mutation prediction for documents and logs covers a wide range of targets, any procedure may be used as long as it can predict mutations in the evidence and confirm the consensus state from the prediction results.
[0200] FIG. 10E shows an example processing procedure 1040 of the argument structure variation prediction unit 1000-06 according to one embodiment. The record of the argument structure is the DCW structure shown in FIG. 3E, and includes the argument history. To predict variations in the structure, first, embedded features are calculated for each component of the argument structure (1040-01). Next, convolutional features are calculated for each argument structure (1040-02), and arguments that deepen understanding are identified (1040-03). That is, in this embodiment, the argument structure is connected to the purpose element 320-01 of the accountability structure 310-01 or the strategy element 310-04 of the GSN extension structure, and plays a role in deepening common understanding among stakeholders related to the element.
[0201] In step 1040-03, after identifying the elements related to the connection, argument structure connection prediction 1040-04 is performed. At this time, a model predictor obtained by sampling argument structure time series data set 1040-05 is used. By predicting the connection attributes between the argument structure and the target element, the agreement state based on the relationship between them is predicted.
[0202] FIG. 10F shows an example processing procedure 1050 of the contribution calculation unit 1000-07 according to one embodiment. First, by performing 1050-01 identification of reliability features related to prediction, reliability features that influenced the derivation of the prediction result are extracted. Next, 1050-02 identification of intermediate calculation results of the feature is performed. For example, when calculating convolutional features, embedded features are calculated in advance. Then, 1050-03 identification of structured data is performed. A substructure of the structured data that was used as input for calculating the feature is identified. At this stage, the series of feature calculation processes leading to the substructure can be determined as data lineage.
[0203] If this is confirmed (Yes in 1050-04), important features are calculated in 1050-05, and contribution explanation is performed in 1050-06. If this is not confirmed, steps 1050-02 and 1050-03 are repeated. If the lineage still cannot be confirmed, the processing of the contribution calculation unit 1000-07 is interrupted (not shown). By confirming the data lineage in this way, it becomes possible to trace the structured data related to the features. Furthermore, by examining the contribution of the features to the change prediction results, it is possible to trace the structured data from the features with the highest contribution and confirm the content of the agreement between the stakeholders. Note that the data lineage may be configured to be recorded in the structured data recording unit 100-02 shown in FIG. 1.
[0204] FIG. 10G shows an example processing procedure 1060 of the feature-based structured data search unit 1000-09 according to one embodiment. A structured data index 1060-10 is created in advance. Structured data elements are extracted 1060-01 from the structured data persisted in the structured data recording unit 100-02 shown in FIG. 1, and element IDs and reliability features are indexed. The latter indexing takes into account the distance between features. Once the index is created, the reliability features are received as input 1060-01, and the following process is executed. First, an exact match search 1060-02 is performed. If a reliability feature that exactly matches the feature is found (Yes in 1060-03), the element ID is searched for in the structured data recording unit 100-02, and the corresponding structured data is obtained. On the other hand, if no perfectly matching feature is found (No in 1060-03), an element ID that is close in distance to the input reliability feature is searched for, and the element ID is searched for in the structured data recording unit 100-02 to obtain the corresponding structured data.
[0205] In this embodiment, the agreement reliability prediction unit 100-05 is configured as shown in FIG. 10A, but this configuration is not limited to this. Other configurations are possible as long as the agreement state can be predicted using reliability features. In this embodiment, the self-descriptive information unit is configured from the four types of structures shown in FIG. 3B, but these structures may be simplified to handle more free text, and the generation AI may be configured to acquire reliability features corresponding to these structures.
[0206] [Actions Associated with Prediction] Predetermining the necessary actions for a project in advance if any change in the consensus status occurs based on the results of the consensus reliability prediction unit 100-05 contributes to improving the reliability of data science. For example, if a process is implemented using a lifecycle model described in Annex A of the international standard IEC 62853, and the substructure derived as a result of the prediction unit is related to the assumed deviation element assumed_deviation 320-07 shown in Figure 3C, the preventive action element preventive_action 320-08 may be implemented as part of the Failure Response Cycle of the standard. This may be automated to the greatest extent possible. Furthermore, if the prediction does not result in a substructure related to the preventive action, or if no substructure is derived at all, the Change Accommodation Cycle of the standard may be implemented.
[0207] 11 shows a configuration example 1100 of the feature sharing control unit 100-07 according to one embodiment. The configuration example 1100 includes four elements: an authentication unit 1100-01, an API control unit 1100-02, a responsibility boundary control unit 1100-03, and a reliability feature repository unit 1100-04.
[0208] The authentication unit 1100-01 authenticates the requester of a reliability feature sharing request. Based on the authentication unit's confirmation of the requester, the API control unit 1100-02 controls access permissions for reliability features requested by the accountability boundary control unit 1100-03 across the requester's accountability boundary. The provider of the reliability features specifies a policy for the access permissions, and the control unit controls the access permissions in accordance with the policy to share the features recorded in the reliability feature repository unit 1100-04 via API with the sharing requester. If there are concerns that project-sensitive information may be included in the sharing of reliability features across accountability boundaries, the risk of the concern can be addressed by checking the data lineage described in the processing procedure 1050 of the contribution calculation unit.
[0209] Reusing reliability features between projects is expected to contribute to improving the reliability of data science worldwide by spreading this to other organizations, other groups, and other projects. If there are no concerns about crossing responsibility boundaries, the feature sharing control unit 100-07 may be configured using AWS Feature Store (https: / / aws.amazon.com / jp / sagemaker / feature-store / ).
[0210] [Dashboard] Figure 12 shows a screen snapshot 1200 as an example of a dashboard according to one embodiment. In Figure 12, 1200-01 visualizes a graphical representation of an accountability structure to enable agreement analysis. 1200-02 is a snapshot of a GSN-extended consensus among multiple stakeholders. 1200-03 visualizes one accountability structure. 1200-04 is a snapshot of a UI related to a discussion using an argument structure. 1200-05 visualizes the embedding features of components of multiple accountability structures. 1200-06 visualizes keyword features related to meeting minutes.
[0211] Various embodiments of the present invention may be described with reference to flowcharts and block diagrams, where the blocks may represent (1) stages of a process in which operations are performed or (2) sections of apparatus responsible for performing the operations. Particular stages and sections may be implemented by dedicated circuitry, programmable circuitry provided with computer-readable instructions stored on a computer-readable medium, and / or a processor provided with computer-readable instructions stored on a computer-readable medium. Dedicated circuitry may include digital and / or analog hardware circuitry, and may include integrated circuits (ICs) and / or discrete circuits. Programmable circuitry may include reconfigurable hardware circuitry including logical AND, OR, XOR, NAND, NOR, and other logic operations, flip-flops, registers, memory elements such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), and the like.
[0212] A computer-readable medium may include any tangible device capable of storing instructions that are executed by an appropriate device, such that the computer-readable medium having instructions stored thereon comprises an article of manufacture containing instructions that can be executed to create means for performing the operations specified in the flowcharts or block diagrams. Examples of computer-readable media may include electronic, magnetic, optical, electromagnetic, and semiconductor storage media. More specific examples of computer-readable media may include floppy disks, diskettes, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), electrically erasable programmable read-only memories (EEPROMs), static random access memories (SRAMs), compact disc read-only memories (CD-ROMs), digital versatile discs (DVDs), Blu-ray discs, memory sticks, integrated circuit cards, and the like.
[0213] The computer readable instructions may include either assembler instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk®, JAVA®, C++, etc., and conventional procedural programming languages such as the “C” programming language or similar programming languages.
[0214] The computer-readable instructions may be provided to a processor or programmable circuitry of a programmable data processing apparatus, such as a general-purpose computer, special-purpose computer, or other computer, either locally or over a local area network (LAN), a wide area network (WAN) such as the Internet, etc., which executes the computer-readable instructions to create means for performing the operations specified in the flowcharts or block diagrams. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.
[0215] 13 illustrates an example of a computer 2200 in which aspects of the present invention may be embodied, in whole or in part. Programs installed on the computer 2200 may cause the computer 2200 to function as or perform operations associated with an apparatus or one or more sections of the apparatus according to embodiments of the present invention, and / or to perform a process or steps of a process according to embodiments of the present invention. Such programs may be executed by the CPU 2212 to cause the computer 2200 to perform specific operations associated with some or all of the blocks in the flowcharts and block diagrams described herein.
[0216] A computer 2200 according to this embodiment includes a CPU 2212, a RAM 2214, a graphics controller 2216, and a display device 2218, which are interconnected by a host controller 2210. The computer 2200 also includes input / output units such as a communication interface 2222, a hard disk drive 2224, a DVD-ROM drive 2226, and an IC card drive, which are connected to the host controller 2210 via an input / output controller 2220. The computer also includes legacy input / output units such as a ROM 2230 and a keyboard 2242, which are connected to the input / output controller 2220 via an input / output chip 2240.
[0217] The CPU 2212 operates according to programs stored in the ROM 2230 and RAM 2214, thereby controlling each unit. The graphics controller 2216 acquires image data generated by the CPU 2212 into a frame buffer or the like provided in the RAM 2214 or into the graphics controller 2216 itself, and causes the image data to be displayed on the display device 2218.
[0218] The communication interface 2222 communicates with other electronic devices via a network. The hard disk drive 2224 stores programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads programs or data from the DVD-ROM 2201 and provides the programs or data to the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from an IC card and / or writes programs and data to an IC card.
[0219] ROM 2230 stores therein a boot program or the like that is executed by computer 2200 upon activation, and / or programs that depend on the hardware of computer 2200. I / O chip 2240 may also connect various I / O units to I / O controller 2220 via parallel ports, serial ports, keyboard ports, mouse ports, etc.
[0220] The programs are provided by a computer-readable medium such as a DVD-ROM 2201 or an IC card. The programs are read from the computer-readable medium, installed in the hard disk drive 2224, RAM 2214, or ROM 2230, which are also examples of computer-readable media, and executed by the CPU 2212. Information processing described in these programs is read by the computer 2200, and brings about cooperation between the programs and the various types of hardware resources described above. An apparatus or method may be configured by implementing information manipulation or processing in accordance with the use of the computer 2200.
[0221] For example, when communication is performed between computer 2200 and an external device, CPU 2212 may execute a communication program loaded in RAM 2214 and instruct communication interface 2222 to perform communication processing based on the processing described in the communication program. Under the control of CPU 2212, communication interface 2222 reads transmission data stored in a transmission buffer processing area provided in RAM 2214, hard disk drive 2224, DVD-ROM 2201, or a recording medium such as an IC card, and transmits the read transmission data to the network, or writes received data received from the network to a reception buffer processing area or the like provided on the recording medium.
[0222] Furthermore, the CPU 2212 may cause all or a necessary portion of a file or database stored on an external recording medium such as the hard disk drive 2224, the DVD-ROM drive 2226 (DVD-ROM 2201), an IC card, etc. to be read into the RAM 2214, and may perform various types of processing on the data on the RAM 2214. The CPU 2212 then writes back the processed data to the external recording medium.
[0223] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and may undergo information processing. The CPU 2212 may perform various types of processing on data read from the RAM 2214, including various types of operations, information processing, conditional judgment, conditional branching, unconditional branching, information search / replacement, etc., as described throughout this disclosure and specified by the instruction sequences of the programs, and write the results back to the RAM 2214. The CPU 2212 may also search for information in a file, database, etc. on the recording medium. For example, if multiple entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored on the recording medium, the CPU 2212 may search for an entry that matches a condition specified by the attribute value of the first attribute from among the multiple entries, read the attribute value of the second attribute stored in the entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.
[0224] The above-described programs or software modules may be stored in a computer-readable medium on or near the computer 2200. A recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can also be used as a computer-readable medium, thereby providing the programs to the computer 2200 via the network.
[0225] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications and improvements can be made to the above embodiments. It is clear from the claims that such modifications and improvements can also be included within the technical scope of the present invention.
[0226] For example, the control system may be a computer housed in a single housing. That is, the controller may be realized by executing a program on the computer's processor, and each input / output device may be implemented as an I / O device of the computer. The controller may also be implemented as a virtual machine executed by one or more processors. In such a configuration, the control system does not include a network, whether a general-purpose or dedicated network, and the controller and the input / output devices may be connected by a chipset, such as a memory controller hub and an I / O controller hub, that connect the processor and the I / O devices.
[0227] It should be noted that the order of execution of each process, such as operations, procedures, steps, and stages, in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not specifically stated as "before," "prior to," etc., and that the processes can be performed in any order unless the output of a previous process is used in a subsequent process. Even if the operational flow in the claims, specifications, and drawings is described using "first," "next," etc. for convenience, this does not mean that the processes must be performed in this order.
[0228] 100 Device 100-01 Structured data collection unit 100-02 Structured data recording unit 100-03 Trail management unit 100-04 Feature calculation unit 100-05 Agreement reliability prediction unit 100-06 UI unit 100-07 Feature sharing control unit 200-01 Arithmetic unit 200-02 Control unit 200-03 Memory unit 200-04 Input / output unit 300 Self-description information unit 310-01 Accountability structure 310-02 GSN extension structure 310-03 Argumentation structure 310-04 Trail structure 320-01 Element purpose 320-02 Element explainer 320-03 Element explainee 320-04 Element requirement 320-05 Element explainer_action 320-06 Element explainee_action 320-07 Element assumed_deviation 320-08 Element preventive_action 320-09 Element alias 320-10 Element user 320-11 Element alias 320-12 Element user 320-13 Label accountability_chain 330-01 Element claim 330-02 Element stakeholder 330-03 Element context 330-04 Element strategy 330-05 Element evidence 330-08 Element user 330-09 Link 340-01 Element claim 340-02 Element data 340-03 Element warrant 340-04 Element claim 340-05 Element data 340-06 Element warrant 340-07 Element user 340-08 Link 350-01 Element document 350-02 Element log 350-03 Element minutes 350-04 Element date_time 350-05 Element agenda 350-06 Element discussion 350-07 Element decision 350-08 Element todo 350-09 Element issue 350-10 Element next 350-11 Element user / stakeholder 350-12 Link 350-13 Element participant 501 Conversation control FG502 Recording control FG 503 Machine learning control FG 501-01 Chat UI section 501-02 Conversation section 502-01 Display section 502-02 Query section 502-03 Search section 503-01 Dashboard section 503-02 History access section 600 Configuration example 600-01 Search engine section 600-02 Graph structure calculation section 600-03 Column-type database 700 Data SLA 710 Data SLA 800 Configuration example 800-01 Reference information management section 800-02 Reference history management section 800-03 Hash tree calculation section 800-04 SQL-type database 900 Configuration example 900-01 Data collection section 900-02 Data formatting section 900-03 Data quality confirmation rule section 900-04 Algorithm execution section 900-05 Algorithm repository unit 900-06 Feature verification unit 900-07 Feature quality confirmation rule unit 910-01 Frame 910-02 Subgraph 910-03 Accountability structure 910-04 Element CONV(a) 910-06 Subgraph 910-07 Element CONV(b) 910-10 Element CONV(c) 910-11 Structured data 1000 Configuration example 1000-01 Reliability feature collection unit 1000-02 Time series database unit 1000-03 Relationship prediction unit between self-descriptive information units 1000-04 Consensus formation depth prediction unit 1000-05 Evidence structure variation prediction unit 1000-06 Argument structure variation prediction unit 1000-07 Contribution calculation unit 1000-09 Feature-based structured data search unit 1100 Configuration example 1100-01 Authentication unit 1100-02 API control unit 1100-03 Responsibility boundary control unit 1100-04 Reliability feature repository unit 2200 Computer 2201 DVD-ROM 2210 Host controller 2212 CPU 2214 RAM 2216 Graphics controller 2218 Display device 2220 Input / output controller 2222 Communication interface 2224 Hard disk drive 2226 DVD-ROM drive 2230 ROM 2240 Input / output chip 2242 Keyboard
Claims
1. An apparatus for ensuring the reliability of a service using data science, comprising: an acquisition unit that acquires, for each of a plurality of deviations that are thought to be likely to occur in the service, basis information indicating a plurality of basis for resolving each deviation, agreed upon among a plurality of stakeholders who provide and receive the service; a generation unit that generates a self-description information unit that describes each of the deviations by determining a plurality of first relationships between the plurality of basis and relating the plurality of basis to each other based on the plurality of first relationships; and a calculation unit that determines a plurality of second relationships between the plurality of self-description information units generated for the plurality of deviations, calculates feature quantities for each of the plurality of self-description information units, and calculates reliability feature quantities that represent the reliability of the service by relating the plurality of feature quantities to each other based on the plurality of second relationships.
2. The device described in claim 1, wherein the generation unit generates the self-describing information unit including, for each deviation, an accountability structure which is a data structure that organizes, in accordance with the plurality of first relationships, a plurality of grounds, among the plurality of grounds, relating to the accountability imposed on each of the plurality of stakeholders to be fulfilled in order to resolve each deviation or to be fulfilled in order to confirm the validity of the accountability that is fulfilled; and an extended structure which is a data structure that organizes, in accordance with the plurality of first relationships, a plurality of grounds, among the plurality of grounds, grounds that serve as a top goal that is a measure to be taken in case the result of fulfilling the accountability is inappropriate, grounds that serve as a strategy for taking the measure, grounds that serve as a plurality of subgoals obtained by decomposing the top goal in accordance with the strategy, and grounds that serve as evidence supporting each of the plurality of subgoals.
3. The device described in claim 2, wherein the generation unit generates the self-describing information unit, for each deviation, if there is a reason among the plurality of reasons that is not agreed upon among the plurality of stakeholders, further including an argument structure that is a data structure that organizes, in accordance with the plurality of first relationships, a plurality of reasons that indicate the process of discussion until all of the plurality of stakeholders reach agreement, or the process of argument until some of the plurality of stakeholders reach agreement and the rest entrust the judgment of the some of the stakeholders to the judgment of the some of the stakeholders.
4. The device described in claim 2, wherein the generation unit generates the self-describing information unit for each deviation, further including a trail structure, which is a data structure that organizes, according to the plurality of first relationships, a plurality of grounds that serve as a trail showing the legitimacy of the top goal, by subdividing the evidence that supports each of the plurality of subgoals among the plurality of grounds.
5. The device described in any one of claims 2 to 4, wherein the calculation unit, for each of the deviations, calculates embedded features from text data describing each of the plurality of grounds constituting the data structure for each of the plurality of data structures included in the self-describing information unit that have relationships with each other, and performs a convolution calculation to convolve the data structure from the calculated embedded features while maintaining the plurality of first relationships between the plurality of grounds in the data structure, thereby calculating convolutional features as the features of the self-describing information unit.
6. The device according to claim 1, further comprising: a detection unit that detects a change in the distance from the reliability feature at a first time to the reliability feature at a second time after the first time by analyzing time series data of the reliability feature, thereby detecting a deviation in the consensus state of the multiple evidences between the multiple stakeholders at the second time; and an extraction unit that, when the detection unit detects a deviation in the consensus state, extracts a self-description information unit related to the deviation from the consensus state from the multiple self-description information units at the second time.
7. The device described in claim 6, wherein the generation unit generates the self-descriptive information unit including grounds for indicating an expected deviation that assumes a case in which the results of the accountability imposed on each of the multiple stakeholders to be fulfilled to resolve each deviation or to confirm the appropriateness of the accountability that is fulfilled are inappropriate, and grounds for indicating measures to address the expected deviation; and the extraction unit, when the detection unit detects that a deviation has occurred in the agreement state, determines whether the content of the deviation in the agreement state corresponds to any of the multiple expected deviations; when an expected deviation corresponding to the content of the deviation exists, determines whether the measures to address the expected deviation are executable; when the measures to address the expected deviation are not executable, determines whether the deviation in the agreement state is identifiable, and if identifiable, extracts a self-descriptive information unit related to the deviation in the agreement state from the multiple self-descriptive information units at the second time.
8. The device according to claim 1, further comprising a prediction unit that predicts a change in at least one of the plurality of grounds for resolving each of the deviations agreed upon among the plurality of stakeholders by analyzing time-series data of the reliability features.
9. The device according to claim 8, further comprising a display unit that visualizes and displays the self-descriptive information unit related to the basis for the prediction by the prediction unit that the mutation will occur.
10. The device according to claim 9, wherein the prediction unit notifies, via the display unit, the plurality of stakeholders of corresponding actions required to resolve the deviation related to the basis on which the mutation is predicted to occur.
11. The device according to claim 1, wherein the acquisition unit acquires text data of the plurality of grounds described by the plurality of stakeholders via a user interface.
12. The device according to claim 1, further comprising a shared control unit that controls the acquisition unit to cross boundaries of accountability of the multiple stakeholders to acquire reliability features of services other than the service, and that controls the calculation unit to apply the reliability features of the other services to the calculation of the reliability features of the service.
13. A method for ensuring the reliability of a service using data science, comprising: obtaining, for each of a plurality of deviations that are considered to be likely to occur in the service, basis information indicating a plurality of basis for resolving each deviation, which basis has been agreed upon among a plurality of stakeholders who provide and receive the service; determining a plurality of first relationships among the plurality of basis, and generating a self-describing information unit that describes each of the deviations by relating the plurality of basis to each other based on the plurality of first relationships; determining a plurality of second relationships among the plurality of self-describing information units generated for the plurality of deviations, calculating feature quantities for each of the plurality of self-describing information units, and relating the feature quantities to each other based on the plurality of second relationships, thereby calculating reliability feature quantities that represent the reliability of the service.
14. A computer program for causing a computer to execute a method for ensuring the reliability of a service using data science, which, when executed by the computer, causes the computer to execute the following steps: for each of a plurality of deviations that are thought to occur in the service, obtain basis information indicating a plurality of basis for resolving each deviation that has been agreed upon among a plurality of stakeholders who provide and receive the service; generate a self-describing information unit that describes each of the deviations by determining a plurality of first relationships among the plurality of basis and relating the plurality of basis to one another based on the plurality of first relationships; and calculate a reliability feature that represents the reliability of the service by determining a plurality of second relationships among the plurality of self-describing information units generated for the plurality of deviations, calculating a feature of each of the plurality of self-describing information units, and relating the plurality of feature to one another based on the plurality of second relationships.
Citation Information
Patent Citations
Dependability maintenance system, change response cycle execution device, fault response cycle execution device, control method for dependability maintenance system, control program, and computer-readable recording medium containing the same.
JP5280587B2